Phylogenetics Notes

Phylogenetics

  • Phylogenetics is the study of the evolutionary relationships among organisms.
  • Oliver A. Ryder (1986) highlighted the importance of preserving adaptive genetic variation, prompting consideration of which gene pools to prioritize.
  • COSEWIC (2010) emphasizes the need to consider units below the species level (subspecies, varieties, or genetically distinct populations) in conservation efforts, as reflected in the Species at Risk Act.

Taxonomy & Systematics

  • Taxonomy involves the identification and naming of taxa.
  • Systematics focuses on assessing evolutionary relationships and phylogeny.
  • Conservation decisions rely on the correct identification of species.

Identifying Species

  • Failure to recognize species can lead to extinction, especially in the case of cryptic species.
  • Over-splitting (recognizing more species than there actually are) can drain resources.
  • Hybridization can result in outbreeding depression.

Approaches to Identifying Taxa

  • Phenetics (Anagenesis): Groups organisms by phenotype, based on the similarity of current traits, resulting in a phenogram.
  • Cladistics (Cladogenesis): Groups organisms by common ancestry, based on sharing derived traits, resulting in a cladogram or phylogeny.

Evolutionary Classification

  • Modern taxonomy combines cladistics and phenetics to create an 'Evolutionary Classification'.
  • Groups are initially classified by phylogeny (common ancestry).
  • Phenotypically divergent groups are often recognized as separate taxa.
  • Phylogeny describes the history of descent of taxa, showing the branching of species from common ancestors, often represented as a tree or dendrogram.
  • The only figure in Darwin's "On the Origin of Species" is a tree depicting evolution.

Phylogenetic Analysis

  • Describes the evolutionary relationships between species.
  • Determines the order of speciation events.
  • Estimates timing of speciation events.

Applications of Phylogenetic Analysis

  • Epidemiology - Zoonotics: Understanding the transmission of diseases from animals to humans.
  • Predicting Gene Function: Inferring the function of genes based on their evolutionary relationships.
  • Evolutionary Biology: Studying the processes of evolution.
  • Conservation:
    • Defining units of conservation.
    • Conservation triage - deciding which species to prioritize for conservation efforts.

Phylogenetic Network Analysis of SARS-CoV-2 Genomes

  • A phylogenetic network analysis of 160 complete human SARS-CoV-2 genomes identified three central variants (A, B, and C).
  • Variant A is considered the ancestral type, based on the bat outgroup coronavirus.
  • Types A and C are found in significant proportions outside East Asia (Europeans and Americans).
  • Type B is most common in East Asia.
  • The ancestral genome of type B appears not to have spread outside East Asia without first mutating into derived B types, suggesting founder effects or resistance.
  • The network analysis can trace routes of infections.
  • Mutations:
    • Type B is derived from A by two mutations: synonymous mutation T8782CT8782C and nonsynonymous mutation C28144TC28144T (leucine to serine).

EDGE (Evolutionarily Distinct & Globally Endangered)

  • Focuses on threatened species that represent a significant amount of unique evolutionary history.
  • GE (Global Endangerment) score: based on IUCN (International Union for Conservation of Nature) category.
  • ED (Evolutionary Distinctiveness) score: based on phylogeny.

Terminology

  • Homologous genes: evolved from a common ancestor.
  • Orthologous genes: evolved from a common ancestor by a speciation event; retain function.
  • Paralogous genes: evolved from a common ancestor by duplication; different function.

Morphological vs. Molecular Phylogenies

  • Mid-1950s: phylogenies constructed based on opinion.
  • Classical phylogenetic analysis: used morphological features.
  • Modern methods use DNA and amino acid sequences.
  • Based on analysis of homologous sequences in different species.

Phylogenetic Tools

  • DNA-DNA hybridization: Measures the heat required to split hybrid DNA (Sibley-Ahlquist used this for bird taxonomy).

Advantages of Phylogenetics

  • DNA and AA sequences are strictly heritable units.
  • Unambiguous description of molecular characters and states.
  • Amenable to mathematical modelling.
  • Homology assessment is easy.
  • Distance relationships may be resolved.
  • Huge amount of data is available.

mtDNA

  • Represents a small fraction of the organismal genome size.
  • Elevated mutation rate generates signal about population history over short time frames.
  • Inheritance is clonal, simplifying within-species variation data.
  • d-loop or control region and Cytochrome b are commonly used.
  • Molecular clock considerations.

Gene Trees vs. Species Trees

  • Genes can result in different phylogenies.

  • Gene trees are often different from the true species phylogeny.

  • Different gene phylogenies can arise due to:

    • lineage sorting
    • sampling error of individuals or populations
    • natural selection
    • introgression following hybridization
  • Phylogeny of the species can be inferred from the gene tree.

  • There is no unique universal phylogenetic tree.

Basic Steps in Phylogenetic Tree Construction

  1. Identify & acquire homologous sequences.
  2. Align the sequences.
  3. Estimate the tree.
  4. Draw the tree.
  5. Assess the reliability of the tree.

Step 1: Identify & Acquire Homologous Sequences

  • Primer extension reactions are used in sequencing.

Next Generation Sequencing (NGS)

  • Databases for sequences and structures:
    • Nucleotide sequences: GenBank (NCBI), EMBL, DDBJ
    • Protein sequences: SWISS-PROT, PIR-International, OWL, NRDB, PROSITE, PRINTS, Pfam
    • Macromolecular structures: Protein Data Bank (PDB), Nucleic Acids Database (NDB), HIV Protease Database, ReLiBase, PDBsum, CATH, SCOP, FSSP
    • Genome sequences: Entrez genomes, GeneCensus, COGs
  • Integrated databases: InterPro, Sequence retrieval system (SRS), Entrez

The International Nucleotide Sequence Database Collaborative (INSDC)

  • A longstanding global alliance of biological data archives.
  • Collects and disseminates DNA and RNA sequences.
  • Three partners establish formats for data, metadata, and protocols:
    • DNA Data Bank of Japan (DDBJ)
    • European Nucleotide Archive (EMBL)
    • GenBank (NCBI)

Bioinformatics & Homology

  • Find related sequences by searching against an online database (NCBI, EMBL, DDBJ).
  • NCBI uses BLAST (Basic Local Alignment Search Tool).
  • BLAST uses a sequence of interest as a “query” to search the database.
  • Input sequence directly or enter accession number.
  • E-value < 0.01 is considered a good match.

Step 2: Align the Sequences

  • Mutation is the primary source of variation.

    • Substitutions: replacement of one nucleotide with another.
    • Recombinations: exchange of a sequence from one homologous chromosome to another.
    • Inversions: rotation by 180∘180^{\circ} of a DNA segment.
    • Deletions: the loss of one or more nucleotides.
    • Insertions: the addition of one or more nucleotides.
  • INDELS (insertions or deletions).

Sequence Alignment

  • All sequences on a tree are homologous (true for species, but not always for genes).
  • Phylogenetic methods assume that sequences in a column are homologous.
  • Without INDELS, sequences can be written one above the other.
  • INDELS change sequence length and shift positions of genes.
  • Alignment introduces gaps to shift bases back to homologous positions.
  • The tree can be no better than the alignment.

ClustalW Options

  • Alignment with and without gaps

Step 3: Estimating the Tree

  • A tree consists of tips or terminal nodes, branches, nodes, and a root.
  • Sister taxa share a recent common ancestor.
  • A polytomy occurs when data cannot resolve which taxa share a more recent common ancestor.
  • Branches can be rotated at nodes without changing the relationships represented on the trees.
  • (A, B, C) is a monophyletic group; it is sister to (D, E)
  • (A, B, C, D) is paraphyletic → E is excluded

Rooted vs. Unrooted Trees

  • Unrooted trees show evolutionary relationships but not the direction of evolution.
  • A root is an interior node from which all sequences/taxa descend.
  • An outgroup is used to determine the root; it is more distantly related to the ingroup sequences than the ingroup are to each other (often a sister taxa).

Methods for Estimating Trees

  • Formula for calculating number of trees with taxa: 3 taxa=1 tree3 \text{ taxa} = 1 \text{ tree}, 4 taxa=3 trees4 \text{ taxa} = 3 \text{ trees}, 5 taxa=15 trees5 \text{ taxa} = 15 \text{ trees}, 6 taxa=105 trees6 \text{ taxa} = 105 \text{ trees}, 8 taxa=135,135 trees8 \text{ taxa} = 135,135 \text{ trees}, 10 taxa=34,000,000 trees10 \text{ taxa} = 34,000,000 \text{ trees}, 50 taxa=3×1074 trees50 \text{ taxa} = 3 \times 10^{74} \text{ trees}

Two Approaches to Reconstruction

  1. Distance-based:
    • Input is a matrix of distances between species.
    • Neighbor-joining is the most common method and doesn’t assume a molecular clock (also UPGMA).
  2. Character-based:
    • Examines all characters (AAs or DNA).
    • Parsimony chooses a tree that minimizes the number of changes required to explain the data.
    • Maximum likelihood, under a model of sequence evolution, finds the tree which gives the highest likelihood of the observed data.
    • Trees based upon both approaches are often reported.

Neighbor Joining Method

  • When two sequences are similar, they likely originate from the same ancestor.
  • Sequence similarity can approximate evolutionary distances.
  • The more time has passed since a pair of sequences have diverged from a common ancestor, the more different they will be.
  • One lineage may have evolved faster, or multiple substitutions may have occurred.
Advantages
  • Fast and suited to large datasets and bootstrapping.
  • Permits lineages with largely different branch lengths.
  • Permits correction for multiple substitutions.
Disadvantages
  • Sequence information is reduced.
  • Gives only one possible tree.
  • Strongly dependent on the model of evolution used (for divergent sequences).

Neighbor Joining - Algorithmic Method

  • Uses a specific set of calculations to estimate a tree.
  • Starting with a multiple alignment, calculate for each pair the distance or fraction of differences and write that to a distance matrix.
Distance Matrix
  • Example:

    | | H | C | O |
    | :-------- | :-: | :-: | :-: |
    | Human (H) | - | 1 | 3 |
    | Chimp (C) | 1 | - | 2 |
    | Orang (O) | 3 | 2 | - |

Multiple Substitutions

  • As two sequences diverge, each substitution will increase the differences between the two lineages.
  • As differences accumulate, it’s more likely a substitution will occur at the same site.
  • The number of observed differences almost always underestimates the actual amount of change.
  • A variety of models are used to estimate corrected differences.
Models for Nucleotide Substitution
  • Allows corrected distances, without underestimating.

    • Juke-Cantor (JC): all mutations are equally likely.
    • Kimura 2-parameter (K2P): transitions are more likely than transversions.
    • Felsenstein 84 (F84): K2P plus unequal base frequencies.
    • General Time Reversible (GTR): most general usable model.

Example of Neighbor Joining Calculation

Step 1
  • Calculate the net divergence r(i)r(i) for each OTU from all other OTUs (Total of 6 OTUs, N=6N=6).
    • r(A)=5+4+7+6+8=30r(A) = 5 + 4 + 7 + 6 + 8 = 30
Step 2
  • Adjusted distance (D) between each pair of nodes. Pairwise distance in starting matrix and divergence of each node.
    • For D(A,B)D(A, B), D(A,B)=d(A,B)−[r(A)+r(B)]/(N−2)=5−(30+42)/(6−2)=−13D(A, B) = d(A, B) - [r(A) + r(B)]/(N - 2) = 5 - (30 + 42)/(6 - 2) = -13
Step 3
  • Using the new matrix, find the closest pair of taxa i,ji, j. Consider the lowest distance and assign uu as the connecting node for that pair.
  • Branch length is calculated using the formula.
    • S(i,u)=d(i,j)/2+(r(i)−r(j))/2(n−2)S(i, u) = d(i, j)/2 + (r(i) - r(j))/2(n - 2)
    • S(j,u)=d(i,j)−S(i,u)S(j, u) = d(i, j) - S(i, u)
  • Closest as per matrix D is: AB = -13. The distance from U to A, and U to B is calculated as:
    • S(A,U)=d(A,B)/2+(r(A)−r(B))/2(n−2)=5/2+(30−42)/2(6−2)=1S(A, U) = d(A, B)/2 + (r(A) - r(B))/2(n - 2) = 5/2 + (30 - 42)/2(6 - 2) = 1
    • S(B,U)=d(A,B)−S(A,U)=5−1=4S(B, U) = d(A, B) - S(A, U) = 5 - 1 = 4
      Where, d(A,B)=5,r(A)=30,r(B)=42 and n=6d(A, B) = 5, r(A) = 30, r(B) = 42 \text{ and } n = 6
Step 4
  • Calculate the new distances from U to all other OTUs:
    • d(U,C)=[d(A,C)+d(B,C)−d(A,B)]/2=[4+7−5]/2=3d(U, C) = [d(A, C) + d(B, C) - d(A, B)]/2 = [4 + 7 - 5]/2 = 3
  • Repeat steps 1 to 4 using the new matrix of distances in every round

How Reliable Is My Tree?

  • Non-parametric bootstrapping is commonly used to assess the uncertainty of tree reconstructions.
  • Disturbs the observed data, composition of alignment columns Checks if the reconstructed trees are similar to the original.
  • One bootstrap replicate randomly draws columns from the original alignment with replacement, repeating until the replicate contains as many columns as the original.
Bootstrapping
  • nn is the number of bootstrap replicates obtained (typically 100 – 1,000).
  • Produces nn bootstrap trees.
  • Bootstrap values correspond to the relative frequency at which a node occurs in the bootstrap replicates.
  • Typically ignore nodes with <60% support.