1/62
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Ligate vs Hybridize
Ligation is the enzymatic process of joining two separate nucleic acid fragments together by forming a covalent phosphodiester bond between the 3'-hydroxyl and 5'-phosphate ends.
Hybridization is the process where complementary single-stranded DNA or RNA molecules bind together through hydrogen bonding between base pairs (A-T and G-C).
Paired End Reads
Sequence of DNA from both ends of a DNA fragment
Long-Read Sequencing
Does not result in any unknown DNA in the read.
contig
Continuous stretch of DNA sequence that has no missing or unknown nucleotides
Scaffold
Scaffolds join contigs together
They are able to span unknown sequence
To do this they either use paired ends, or some other technology to identify the order and orientation of the contigs
Genome Sequencing and Assembly Short Process
Genome Sequencing: multiple copies of gene, sheared random fragments, size fractionated fragments, reads
Genome Assembly: contigs, scaffolds
Recap Steps of Genome Sequencing
Summary of Genome Sequencing
Many copies of the genome are used for sequencing unless “Single Cell Sequencing” is applied
Unless primers are used sequencing cannot target a particular gene region
Paired-end sequencing reads a single random DNA fragment from both 5’ to 3’ and 3’ to 5’ directions
Only the ends of the fragment are read with the shorter read technologies
Long-read sequencing does not use paired ends
Depth of Coverage
The average number of times each base of a genome is sequenced or read. We want:
10-15x for animals
200x for humans
Rule-of-Thumb for a Variant to be Confirmed
To believe that a variant is really a variant and not just a sequencing error, we need to see it at least twice

Grey shows the unknown sequence so overlaps of black help to create a known scaffold
Contigs are not always assembled in the same orientation
We expect half will be 5’ to 3’ and half will be 3’ to 5’
So we may need to flip them over to make them match properly on the scaffold
N’s in Genome Assembly
When the number of base pairs in a gap is known then it is replaced by the same number of ‘N’ - if it is unknown then it is replaced by about 200 N’s.
Fluorescence In Situ Hybridization (FISH)
Validates chromosome-level genome assemblies by physically mapping DNA sequences directly onto microscope slides of metaphase chromosomes. This cytogenetic method confirms scaffolding accuracy, detects misassemblies, and links physical chromosomes to linkage groups.
Oligonucleotide probes labelled with an annealed fluorescent probe
Ex. we can distinguish chromosomes because they have different colours on the ends
Reference Scaffolds for FISH
Reference genome scaffolds must be large enough to enable visual separation of the probe sequences - requires scaffolds of around one million bp or more
Order and Orientation of Contigs
Reverse-complement the back-to-front ones to join them up properly
Sometimes a small contig will fit into a gap in a scaffold
A genome assembly needs to take these small contigs and put them in the correct place - FISH can be used to determine this
Technologies aside from FISH
Hi-C is a technology that looks at the actual physical connection between DNA in cells to identify those that lie in close proximity on the genome
Pan-Genome
A pangenome is the complete set of all genes and genetic variations found across multiple individuals, strains, or species. Instead of relying on a single linear reference genome for a species, a pangenome captures the full spectrum of genomic diversity.
Genome Annotation
See notes for images!
Areas that are GC rich usually indicate the start of an exon
Genome Annotation Overview
Exon sequences are highly conserved across individuals of the same phyla
Can use “comparative genomics” to assist in annotation of a new reference genome
Repeat elements result in loss of contiguity of genome assemblies and are very common
New technologies such as RNA-seq can capture only the transcribed portions of the genome
Transposable Elements (TEs)
Repetitive DNA sequences
Found in genomes scattered across the tree of life
Often called “jumping genes” because of their ability to replicate and move to new genomic locations
Provide an important source of genome variation at both the species and individual level
Eukaryotic TEs
Categorized based on their mechanism of retrotransposition:
Class I retrotransposons
Use a copy-and-paste mechanism via an RNA intermediate
Allow massive amplification of copy number
Have the potential to cause substantial genomic change
Class II DNA transposons
Restricted by their cut-and-paste mechanism
Retrotransposon Classifications
Retrotransposons are further divided into elements with (LTR) and without (non-LTR) long terminal repeats
Non-LTR elements comprise long interspersed elements (LINEs) and short interspersed elements (SINEs)
LINE
>6Kb
Long interspersed nuclear elements
Include promoter sequence
Self replicating
LINEs code for the enzyme reverse transcriptase
SINE
~150bp
Short interspersed nuclear elements
No promoter sequence
Rely on other mobile elements for transposition
Can be incorporated into other gene transcripts by error

PHRED score
Q = -10 log10P
Quality metric that describes how often we see a sequencing error in a sequencing read
Draft genomes typically sequenced to PHRED 40 (1 in 10,000 bases called incorrectly)
Based on a consensus sequence – which is the average call for a base defined by the coverage of many reads
Diversity vs Kinship
Diversity is a characteristic of a population whereas kinship is a characteristic of pairs within a population
Synonymous and Non-Synonymous Mutations
Synonymous mutations are changes in a DNA sequence that do not change the resulting amino acid sequence of a protein. Nonsynonymous mutations do
Haplotypes
A specific group of genes or DNA variants on a single chromosome that are inherited together from one parent
Coeffcient of Relationship and Inbreeding
Coefficient of relationship
Chance that two given individuals share a haplotype
Siblings are 0.5 (unless inbreeding so might be more)
Coefficient of inbreeding
Chance that one individual has inherited the same haplotype from both its parents
Lower coefficient of inbreeding = less inbreeding
Using the COI
The coefficient of inbreeding isn’t the most important thing when breeding animals. It really matters how different they are from each other - the coefficient of inbreeding is not inherited (inbred parents do not beget inbred offspring)
Heterosis
Also known as hybrid vigor, is the improved or increased function of any biological quality in a hybrid offspring resulting from the mixing of different genetic contributions from its parents.
Minor Allele Frequency
The percentage or rate at which the second most common variant of a gene or genetic marker (allele) appears in a specific population
Does not determine number of homozygous or heterozygous individuals
In a homozygous population the minor allele frequency will be 0 because there is no minor allele
A MAF of 0.5 means that both alleles at a biallelic genetic site appear with equal frequency (50% each) in a population and this is the highest a MAF can be
Pedigree for a Litter of Offspring

Genotyping Array
Simultaneously determines genotypes at thousands of known variant locations.
Variant Discovery vs Variant Genotyping
Variant discovery: finding new/unknown genetic variants by comparing against a referance genome
Variant genotyping: determining which allele/genotype an individual has at known variant locations.
Majority of variants…
Arise sporadically
Do not cause phenotypic change
May be common or rare
May be enriched by genetic drift
Do not affect fitness
Useful for mapping
Variant Genotyping Methods
Low pass sequencing with imputation e.g. Gencove
RADSeq (Restriction-site Associated DNA Markers)
DArTseq (Diversity Arrays Technology)
Low pass sequencing with imputation e.g. Gencove
Requires a reference genome and a haplotype library
Only need 3x or 4x coverage
Mostly used to determine the variation within a population (normally small)
RADSeq (Restriction-site Associated DNA Markers)
RADseq targets sequences adjacent to common restriction sites distributed throughout the genome and may recover thousands of loci that can be used to search for single nucleotide polymorphisms (SNPs) and indels to investigate population structure and diversity
Requires good quality DNA to predict restriction sites correctly
DArTseq (Diversity Arrays Technology)
The use of a combination of restriction enzymes which separate low copy sequences (most informative for marker discovery and typing) from the repetitive fraction of the genome.
Common Features of Variant Genotyping Methods
Lower accuracy than array-based methods
Confounding due to DNA degradation and repeat element variants
Low set-up cost (array costs in millions to create)
May not require reference genome (DArTseq and RADseq)
May not be able to capture the same variants in all tested individuals
Do not require prior SNP discovery
What do we need for variant discovery?
Reference genome
Sequencing reads from a different individual
Bioinformatic tool to perform the alignment
Another tool to check the alignment quality
BUSCO
BUSCO (Benchmarking Universal Single-Copy Orthologs) evaluates gene annotation, transcriptome, and genome assembly completeness using evolutionarily conserved gene models
A higher score is better
1st and 2nd Generation Sequencing
1st Generation: ‘Sanger’ sequencing - ‘chain termination’ sequencing based on PCR; chromatogram
2nd Generation or ‘Next-generation’ sequencing (NGS) - massively parallel high-throughput sequencing (short reads)
DNA-Seq
Type of 2nd gen
Rapidly sequence whole genomes (eukaryotic, prokaryotic)
Deeply sequence target regions (targeted amplification via enrichment)
Reduced-representation sequencing (RAD-seq – sequence smaller fraction of genome; uses restriction enzymes)
Metagenomics to study microbiomes; identify novel pathogens
RNA-Seq
Type of 2nd gen
mRNA sequencing to quantify mRNAs for gene expression analysis
total RNA sequencing to discover novel RNA variants (non-coding RNAs)
Metatranscriptomics to identify novel viruses with RNA genomes
NGS Process
Nucleic Acid Extraction and Isolation:
DNA or RNA extraction from source
Library preparation:
Fragmentation to produce fragments of desired lengths
Adapters ligated at both ends
Clonal amplification & sequencing:
Library fragments bind to solid surface (flow cell; beads)
Fluorescently labelled nucleotides detected by sequencer
Sequence analysis:
Bioinformatics pipelines to process and analyse sequences
Each step can vary depending on specifics of the sample, the question being asked, and the experimental design
Metagenomics and Metatranscriptomics
Identification of all genomes of species in complex microbial communities
Next-generation sequencing (NGS) is commonly used for metagenomics and metatranscriptomics
Features of Metagenomics and Metatranscriptomics
Detection of novel organisms: culture/PCR independent profiling. Does not rely on prior knowledge of target sequences, allowing unbiased discovery of previously unknown taxa
Short-read compatible: bacterial and viral genomes with small genome sizes, enables high sequencing coverage, facilitating accurate genome assembly and variant detection.
Quantitative analysis: Sequencing read counts can be used to estimate relative abundance of taxa, genes, and transcripts across samples.
3rd Generation Sequencing
Long-read/third-generation sequencing technologies are causing a new revolution in genomics as they provide a way to study the structure of genomes, transcriptomes, and metagenomic communities at an unprecedented resolution.
Pacific Biosciences (PacBio)
A single-molecule, real-time (SMRT) long-read technology that delivers exceptional read lengths (15,000 to 25,000 base pairs) with over 99.9% accuracy
DNA is converted into circular templates using hairpin adapters (SMRTbell) without needing PCR amplification.
A polymerase copies the strand inside microscopic wells on a SMRT Cell, emitting light pulses as bases are added.
The system reads the same molecule over and over to create an ultra-accurate consensus sequence
Oxford Nanopore
Step 1- Nucleic Acid Extraction and Isolation: DNA or RNA extraction from source
Step 2- End-prep and adaptor ligation: Adapters ligated at both ends
Step 3- Loading & sequencing: DNA fragments loaded onto flow cell and passed through the protein nanopores. Current recorded to enable base-calling
Step 4- Sequence analysis: Bioinformatics pipelines to process and analyse sequences
NGS vs Long-read sequencing
NGS
Very cheap, efficient, and high accuracy
Can’t assembly across repetitive sequence
Does not identify exon junctions
Relies on PCR amplification leading to biased coverage
Long-read sequencing
Lower accuracy and higher cost per base
Detects structural variants and haplotype phasing
Identifies alternative splicing
Omits PCR bias
Characterises nucleotide modifications
Applications of long-read sequencing
Improved reference genomes using Hybrid assembly: Long-reads used to scaffold short-read assemblies.
Low-coverage long-read sequencing + high-coverage short read sequencing
Balance trade-off between accuracy and structural information
Whole-transcript sequencing without need for assembly
Definitive evidence for alternative splicing, improving accuracy of gene models
Enables phasing of transcripts to the alleles from which they were transcribed
Directly detect modifications on native DNA
Detecting DNA base modifications like DNA methylation allows epigenetic state determination
Rapid disease diagnostics
Rapid, long amplicon sequencing; in-house capability – no need for external sequencing facility
COVID genotyping; biosecurity threats
Whole Genome
Whole genome sequencing (WGS): randomized sequencing approach – sequence all that is present in the sample
No prior sequence knowledge required – novel genome discovery possible
Amplicon Sequencing
Amplicon sequencing: type of targeted sequencing that uses PCR to create DNA sequences called amplicons
Must know the sequence of the target so that you can design primers for the PCR step
PCR amplification step makes identifying sequence possible in low-yield situations
Application of Whole Genome and Amplicon Sequencing
WGS can be used to identify an unknown pathogen in disease diagnostics
Amplicon sequencing can be used to identify a known pathogen in disease diagnostics
Which sequencing platform would you use Examples
You are a bee pathologist trying to identify the impact of a new parasite on Australian honey bees. This parasite transmits RNA viruses to the bees and causes colony collapse. You want to identify whether any novel viruses have been transmitted from the parasite to the bees.
Short read sequencing
You want to compare different giraffe populations across Africa to see if they are one large meta-population, or if they are distinct populations or species. You have hundreds of samples to analyse and already have a quality reference genome.
Short read sequencing
You are a plant pathologist that has observed the tell-tale signs of the exotic Tomato brown rugose fruit virus in a farm. You want to know whether there has been a biosecurity incursion of this specific damaging virus into Australia.
Amplicon sequencing
You are interested in characterising the core genome and variable genome sequences across multiple horse breeds, including Thoroughbreds, Icelandic horses, and Przewalski’s Horse. You'd like to ensure your data examines structural variation and chromosomal rearrangements.
Long read sequencing
Which sequencing platform would you use?
Genome resequencing for SNP/variant calling - short-read DNA sequencing
Unbiased analysis of diverse bacterial communities - short-read DNA sequencing (metagenomics)
Identification of unknown disease-causing virus - short-read RNA sequencing (metatranscriptomics)
High quality reference genome assembly - long read DNA sequencing
Hybrid assembly combining - long and short read data
Pangenomic analysis of large structural variation - long read DNA sequencing
Targeted disease sequencing diagnostics - long read DNA sequencing (amplicons)