Phylogenetics Notes
Phylogenetics
- Phylogenetics is the study of the evolutionary relationships among organisms.
- Oliver A. Ryder (1986) highlighted the importance of preserving adaptive genetic variation, prompting consideration of which gene pools to prioritize.
- COSEWIC (2010) emphasizes the need to consider units below the species level (subspecies, varieties, or genetically distinct populations) in conservation efforts, as reflected in the Species at Risk Act.
Taxonomy & Systematics
- Taxonomy involves the identification and naming of taxa.
- Systematics focuses on assessing evolutionary relationships and phylogeny.
- Conservation decisions rely on the correct identification of species.
Identifying Species
- Failure to recognize species can lead to extinction, especially in the case of cryptic species.
- Over-splitting (recognizing more species than there actually are) can drain resources.
- Hybridization can result in outbreeding depression.
Approaches to Identifying Taxa
- Phenetics (Anagenesis): Groups organisms by phenotype, based on the similarity of current traits, resulting in a phenogram.
- Cladistics (Cladogenesis): Groups organisms by common ancestry, based on sharing derived traits, resulting in a cladogram or phylogeny.
Evolutionary Classification
- Modern taxonomy combines cladistics and phenetics to create an 'Evolutionary Classification'.
- Groups are initially classified by phylogeny (common ancestry).
- Phenotypically divergent groups are often recognized as separate taxa.
- Phylogeny describes the history of descent of taxa, showing the branching of species from common ancestors, often represented as a tree or dendrogram.
- The only figure in Darwin's "On the Origin of Species" is a tree depicting evolution.
Phylogenetic Analysis
- Describes the evolutionary relationships between species.
- Determines the order of speciation events.
- Estimates timing of speciation events.
Applications of Phylogenetic Analysis
- Epidemiology - Zoonotics: Understanding the transmission of diseases from animals to humans.
- Predicting Gene Function: Inferring the function of genes based on their evolutionary relationships.
- Evolutionary Biology: Studying the processes of evolution.
- Conservation:
- Defining units of conservation.
- Conservation triage - deciding which species to prioritize for conservation efforts.
Phylogenetic Network Analysis of SARS-CoV-2 Genomes
- A phylogenetic network analysis of 160 complete human SARS-CoV-2 genomes identified three central variants (A, B, and C).
- Variant A is considered the ancestral type, based on the bat outgroup coronavirus.
- Types A and C are found in significant proportions outside East Asia (Europeans and Americans).
- Type B is most common in East Asia.
- The ancestral genome of type B appears not to have spread outside East Asia without first mutating into derived B types, suggesting founder effects or resistance.
- The network analysis can trace routes of infections.
- Mutations:
- Type B is derived from A by two mutations: synonymous mutation and nonsynonymous mutation (leucine to serine).
EDGE (Evolutionarily Distinct & Globally Endangered)
- Focuses on threatened species that represent a significant amount of unique evolutionary history.
- GE (Global Endangerment) score: based on IUCN (International Union for Conservation of Nature) category.
- ED (Evolutionary Distinctiveness) score: based on phylogeny.
Terminology
- Homologous genes: evolved from a common ancestor.
- Orthologous genes: evolved from a common ancestor by a speciation event; retain function.
- Paralogous genes: evolved from a common ancestor by duplication; different function.
Morphological vs. Molecular Phylogenies
- Mid-1950s: phylogenies constructed based on opinion.
- Classical phylogenetic analysis: used morphological features.
- Modern methods use DNA and amino acid sequences.
- Based on analysis of homologous sequences in different species.
Phylogenetic Tools
- DNA-DNA hybridization: Measures the heat required to split hybrid DNA (Sibley-Ahlquist used this for bird taxonomy).
Advantages of Phylogenetics
- DNA and AA sequences are strictly heritable units.
- Unambiguous description of molecular characters and states.
- Amenable to mathematical modelling.
- Homology assessment is easy.
- Distance relationships may be resolved.
- Huge amount of data is available.
mtDNA
- Represents a small fraction of the organismal genome size.
- Elevated mutation rate generates signal about population history over short time frames.
- Inheritance is clonal, simplifying within-species variation data.
- d-loop or control region and Cytochrome b are commonly used.
- Molecular clock considerations.
Gene Trees vs. Species Trees
Genes can result in different phylogenies.
Gene trees are often different from the true species phylogeny.
Different gene phylogenies can arise due to:
- lineage sorting
- sampling error of individuals or populations
- natural selection
- introgression following hybridization
Phylogeny of the species can be inferred from the gene tree.
There is no unique universal phylogenetic tree.
Basic Steps in Phylogenetic Tree Construction
- Identify & acquire homologous sequences.
- Align the sequences.
- Estimate the tree.
- Draw the tree.
- Assess the reliability of the tree.
Step 1: Identify & Acquire Homologous Sequences
- Primer extension reactions are used in sequencing.
Next Generation Sequencing (NGS)
- Databases for sequences and structures:
- Nucleotide sequences: GenBank (NCBI), EMBL, DDBJ
- Protein sequences: SWISS-PROT, PIR-International, OWL, NRDB, PROSITE, PRINTS, Pfam
- Macromolecular structures: Protein Data Bank (PDB), Nucleic Acids Database (NDB), HIV Protease Database, ReLiBase, PDBsum, CATH, SCOP, FSSP
- Genome sequences: Entrez genomes, GeneCensus, COGs
- Integrated databases: InterPro, Sequence retrieval system (SRS), Entrez
The International Nucleotide Sequence Database Collaborative (INSDC)
- A longstanding global alliance of biological data archives.
- Collects and disseminates DNA and RNA sequences.
- Three partners establish formats for data, metadata, and protocols:
- DNA Data Bank of Japan (DDBJ)
- European Nucleotide Archive (EMBL)
- GenBank (NCBI)
Bioinformatics & Homology
- Find related sequences by searching against an online database (NCBI, EMBL, DDBJ).
- NCBI uses BLAST (Basic Local Alignment Search Tool).
- BLAST uses a sequence of interest as a “query” to search the database.
- Input sequence directly or enter accession number.
- E-value < 0.01 is considered a good match.
Step 2: Align the Sequences
Mutation is the primary source of variation.
- Substitutions: replacement of one nucleotide with another.
- Recombinations: exchange of a sequence from one homologous chromosome to another.
- Inversions: rotation by of a DNA segment.
- Deletions: the loss of one or more nucleotides.
- Insertions: the addition of one or more nucleotides.
INDELS (insertions or deletions).
Sequence Alignment
- All sequences on a tree are homologous (true for species, but not always for genes).
- Phylogenetic methods assume that sequences in a column are homologous.
- Without INDELS, sequences can be written one above the other.
- INDELS change sequence length and shift positions of genes.
- Alignment introduces gaps to shift bases back to homologous positions.
- The tree can be no better than the alignment.
ClustalW Options
- Alignment with and without gaps
Step 3: Estimating the Tree
- A tree consists of tips or terminal nodes, branches, nodes, and a root.
- Sister taxa share a recent common ancestor.
- A polytomy occurs when data cannot resolve which taxa share a more recent common ancestor.
- Branches can be rotated at nodes without changing the relationships represented on the trees.
- (A, B, C) is a monophyletic group; it is sister to (D, E)
- (A, B, C, D) is paraphyletic → E is excluded
Rooted vs. Unrooted Trees
- Unrooted trees show evolutionary relationships but not the direction of evolution.
- A root is an interior node from which all sequences/taxa descend.
- An outgroup is used to determine the root; it is more distantly related to the ingroup sequences than the ingroup are to each other (often a sister taxa).
Methods for Estimating Trees
- Formula for calculating number of trees with taxa: , , , , , ,
Two Approaches to Reconstruction
- Distance-based:
- Input is a matrix of distances between species.
- Neighbor-joining is the most common method and doesn’t assume a molecular clock (also UPGMA).
- Character-based:
- Examines all characters (AAs or DNA).
- Parsimony chooses a tree that minimizes the number of changes required to explain the data.
- Maximum likelihood, under a model of sequence evolution, finds the tree which gives the highest likelihood of the observed data.
- Trees based upon both approaches are often reported.
Neighbor Joining Method
- When two sequences are similar, they likely originate from the same ancestor.
- Sequence similarity can approximate evolutionary distances.
- The more time has passed since a pair of sequences have diverged from a common ancestor, the more different they will be.
- One lineage may have evolved faster, or multiple substitutions may have occurred.
Advantages
- Fast and suited to large datasets and bootstrapping.
- Permits lineages with largely different branch lengths.
- Permits correction for multiple substitutions.
Disadvantages
- Sequence information is reduced.
- Gives only one possible tree.
- Strongly dependent on the model of evolution used (for divergent sequences).
Neighbor Joining - Algorithmic Method
- Uses a specific set of calculations to estimate a tree.
- Starting with a multiple alignment, calculate for each pair the distance or fraction of differences and write that to a distance matrix.
Distance Matrix
Example:
| | H | C | O |
| :-------- | :-: | :-: | :-: |
| Human (H) | - | 1 | 3 |
| Chimp (C) | 1 | - | 2 |
| Orang (O) | 3 | 2 | - |
Multiple Substitutions
- As two sequences diverge, each substitution will increase the differences between the two lineages.
- As differences accumulate, it’s more likely a substitution will occur at the same site.
- The number of observed differences almost always underestimates the actual amount of change.
- A variety of models are used to estimate corrected differences.
Models for Nucleotide Substitution
Allows corrected distances, without underestimating.
- Juke-Cantor (JC): all mutations are equally likely.
- Kimura 2-parameter (K2P): transitions are more likely than transversions.
- Felsenstein 84 (F84): K2P plus unequal base frequencies.
- General Time Reversible (GTR): most general usable model.
Example of Neighbor Joining Calculation
Step 1
- Calculate the net divergence for each OTU from all other OTUs (Total of 6 OTUs, ).
Step 2
- Adjusted distance (D) between each pair of nodes. Pairwise distance in starting matrix and divergence of each node.
- For ,
Step 3
- Using the new matrix, find the closest pair of taxa . Consider the lowest distance and assign as the connecting node for that pair.
- Branch length is calculated using the formula.
- Closest as per matrix D is: AB = -13. The distance from U to A, and U to B is calculated as:
Where,
Step 4
- Calculate the new distances from U to all other OTUs:
- Repeat steps 1 to 4 using the new matrix of distances in every round
How Reliable Is My Tree?
- Non-parametric bootstrapping is commonly used to assess the uncertainty of tree reconstructions.
- Disturbs the observed data, composition of alignment columns Checks if the reconstructed trees are similar to the original.
- One bootstrap replicate randomly draws columns from the original alignment with replacement, repeating until the replicate contains as many columns as the original.
Bootstrapping
- is the number of bootstrap replicates obtained (typically 100 – 1,000).
- Produces bootstrap trees.
- Bootstrap values correspond to the relative frequency at which a node occurs in the bootstrap replicates.
- Typically ignore nodes with <60% support.