Applied Genomics Lectures

0.0(0)
Studied by 0 people
call kaiCall Kai
Locked
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/62

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 7:50 AM on 9/20/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

63 Terms

1
New cards

Ligate vs Hybridize

  • Ligation is the enzymatic process of joining two separate nucleic acid fragments together by forming a covalent phosphodiester bond between the 3'-hydroxyl and 5'-phosphate ends.

  • Hybridization is the process where complementary single-stranded DNA or RNA molecules bind together through hydrogen bonding between base pairs (A-T and G-C).


2
New cards

Paired End Reads

Sequence of DNA from both ends of a DNA fragment

3
New cards

Long-Read Sequencing

Does not result in any unknown DNA in the read.

4
New cards

contig

Continuous stretch of DNA sequence that has no missing or unknown nucleotides

5
New cards

Scaffold

  • Scaffolds join contigs together

  • They are able to span unknown sequence

  • To do this they either use paired ends, or some other technology to identify the order and orientation of the contigs


6
New cards

Genome Sequencing and Assembly Short Process

  • Genome Sequencing: multiple copies of gene, sheared random fragments, size fractionated fragments, reads

  • Genome Assembly: contigs, scaffolds


7
New cards

Recap Steps of Genome Sequencing

8
New cards

Summary of Genome Sequencing

  • Many copies of the genome are used for sequencing unless “Single Cell Sequencing” is applied

  • Unless primers are used sequencing cannot target a particular gene region

  • Paired-end sequencing reads a single random DNA fragment from both 5’ to 3’ and 3’ to 5’ directions

  • Only the ends of the fragment are read with the shorter read technologies

  • Long-read sequencing does not use paired ends


9
New cards

Depth of Coverage

The average number of times each base of a genome is sequenced or read. We want:

  • 10-15x for animals

  • 200x for humans


10
New cards

Rule-of-Thumb for a Variant to be Confirmed

To believe that a variant is really a variant and not just a sequencing error, we need to see it at least twice

11
New cards
term image
  • Grey shows the unknown sequence so overlaps of black help to create a known scaffold

  • Contigs are not always assembled in the same orientation

  • We expect half will be 5’ to 3’ and half will be 3’ to 5’

  • So we may need to flip them over to make them match properly on the scaffold


12
New cards

N’s in Genome Assembly

When the number of base pairs in a gap is known then it is replaced by the same number of ‘N’ - if it is unknown then it is replaced by about 200 N’s.

13
New cards

Fluorescence In Situ Hybridization (FISH)

  • Validates chromosome-level genome assemblies by physically mapping DNA sequences directly onto microscope slides of metaphase chromosomes. This cytogenetic method confirms scaffolding accuracy, detects misassemblies, and links physical chromosomes to linkage groups.

  • Oligonucleotide probes labelled with an annealed fluorescent probe

  • Ex. we can distinguish chromosomes because they have different colours on the ends


14
New cards

Reference Scaffolds for FISH

Reference genome scaffolds must be large enough to enable visual separation of the probe sequences - requires scaffolds of around one million bp or more

15
New cards

Order and Orientation of Contigs

  • Reverse-complement the back-to-front ones to join them up properly

  • Sometimes a small contig will fit into a gap in a scaffold

  • A genome assembly needs to take these small contigs and put them in the correct place - FISH can be used to determine this


16
New cards

Technologies aside from FISH

Hi-C is a technology that looks at the actual physical connection between DNA in cells to identify those that lie in close proximity on the genome

17
New cards

Pan-Genome

A pangenome is the complete set of all genes and genetic variations found across multiple individuals, strains, or species. Instead of relying on a single linear reference genome for a species, a pangenome captures the full spectrum of genomic diversity.

18
New cards

Genome Annotation

  • See notes for images!

  • Areas that are GC rich usually indicate the start of an exon


19
New cards

Genome Annotation Overview

  • Exon sequences are highly conserved across individuals of the same phyla

  • Can use “comparative genomics” to assist in annotation of a new reference genome

  • Repeat elements result in loss of contiguity of genome assemblies and are very common

  • New technologies such as RNA-seq can capture only the transcribed portions of the genome


20
New cards

Transposable Elements (TEs)

  • Repetitive DNA sequences

  • Found in genomes scattered across the tree of life

  • Often called “jumping genes” because of their ability to replicate and move to new genomic locations

  • Provide an important source of genome variation at both the species and individual level


21
New cards

Eukaryotic TEs

Categorized based on their mechanism of retrotransposition:

  • Class I retrotransposons

    • Use a copy-and-paste mechanism via an RNA intermediate

    • Allow massive amplification of copy number

    • Have the potential to cause substantial genomic change

  • Class II DNA transposons

    • Restricted by their cut-and-paste mechanism


22
New cards

Retrotransposon Classifications

  • Retrotransposons are further divided into elements with (LTR) and without (non-LTR) long terminal repeats

  • Non-LTR elements comprise long interspersed elements (LINEs) and short interspersed elements (SINEs)


23
New cards

LINE

  • >6Kb

  • Long interspersed nuclear elements

  • Include promoter sequence

  • Self replicating

  • LINEs code for the enzyme reverse transcriptase


24
New cards

SINE

  • ~150bp

  • Short interspersed nuclear elements

  • No promoter sequence

  • Rely on other mobile elements for transposition

  • Can be incorporated into other gene transcripts by error


25
New cards
<p><span style="background-color: transparent;">PHRED score</span></p>

PHRED score

  • Q = -10 log10P

  • Quality metric that describes how often we see a sequencing error in a sequencing read

  • Draft genomes typically sequenced to PHRED 40 (1 in 10,000 bases called incorrectly)

  • Based on a consensus sequence – which is the average call for a base defined by the coverage of many reads


26
New cards

Diversity vs Kinship

Diversity is a characteristic of a population whereas kinship is a characteristic of pairs within a population

27
New cards

Synonymous and Non-Synonymous Mutations

Synonymous mutations are changes in a DNA sequence that do not change the resulting amino acid sequence of a protein. Nonsynonymous mutations do

28
New cards

Haplotypes

A specific group of genes or DNA variants on a single chromosome that are inherited together from one parent

29
New cards

Coeffcient of Relationship and Inbreeding

Coefficient of relationship

  • Chance that two given individuals share a haplotype 

  • Siblings are 0.5 (unless inbreeding so might be more)

Coefficient of inbreeding

  • Chance that one individual has inherited the same haplotype from both its parents

  • Lower coefficient of inbreeding = less inbreeding


30
New cards

Using the COI

The coefficient of inbreeding isn’t the most important thing when breeding animals. It really matters how different they are from each other - the coefficient of inbreeding is not inherited (inbred parents do not beget inbred offspring)

31
New cards

Heterosis

Also known as hybrid vigor, is the improved or increased function of any biological quality in a hybrid offspring resulting from the mixing of different genetic contributions from its parents.

32
New cards

Minor Allele Frequency

  • The percentage or rate at which the second most common variant of a gene or genetic marker (allele) appears in a specific population

  • Does not determine number of homozygous or heterozygous individuals

  • In a homozygous population the minor allele frequency will be 0 because there is no minor allele

  • A MAF of 0.5 means that both alleles at a biallelic genetic site appear with equal frequency (50% each) in a population and this is the highest a MAF can be


33
New cards

Pedigree for a Litter of Offspring

knowt flashcard image
34
New cards

Genotyping Array

Simultaneously determines genotypes at thousands of known variant locations.

35
New cards

Variant Discovery vs Variant Genotyping

  • Variant discovery: finding new/unknown genetic variants by comparing against a referance genome

  • Variant genotyping: determining which allele/genotype an individual has at known variant locations.


36
New cards

Majority of variants…

  • Arise sporadically

  • Do not cause phenotypic change

  • May be common or rare

  • May be enriched by genetic drift

  • Do not affect fitness

  • Useful for mapping


37
New cards

Variant Genotyping Methods

  • Low pass sequencing with imputation e.g. Gencove

  • RADSeq (Restriction-site Associated DNA Markers)

  • DArTseq (Diversity Arrays Technology)


38
New cards

Low pass sequencing with imputation e.g. Gencove

  • Requires a reference genome and a haplotype library

  • Only need 3x or 4x coverage

  • Mostly used to determine the variation within a population (normally small)


39
New cards

RADSeq (Restriction-site Associated DNA Markers)

  • RADseq targets sequences adjacent to common restriction sites distributed throughout the genome and may recover thousands of loci that can be used to search for single nucleotide polymorphisms (SNPs) and indels to investigate population structure and diversity

  • Requires good quality DNA to predict restriction sites correctly


40
New cards

DArTseq (Diversity Arrays Technology)

The use of a combination of restriction enzymes which separate low copy sequences (most informative for marker discovery and typing) from the repetitive fraction of the genome.

41
New cards

Common Features of Variant Genotyping Methods

  • Lower accuracy than array-based methods

  • Confounding due to DNA degradation and repeat element variants

  • Low set-up cost (array costs in millions to create)

  • May not require reference genome (DArTseq and RADseq)

  • May not be able to capture the same variants in all tested individuals

  • Do not require prior SNP discovery


42
New cards

What do we need for variant discovery?

  • Reference genome

  • Sequencing reads from a different individual

  • Bioinformatic tool to perform the alignment

  • Another tool to check the alignment quality


43
New cards

BUSCO

  • BUSCO (Benchmarking Universal Single-Copy Orthologs) evaluates gene annotation, transcriptome, and genome assembly completeness using evolutionarily conserved gene models

  • A higher score is better


44
New cards

1st and 2nd Generation Sequencing

  • 1st Generation: ‘Sanger’ sequencing - ‘chain termination’ sequencing based on PCR; chromatogram

  • 2nd Generation or ‘Next-generation’ sequencing (NGS) - massively parallel high-throughput sequencing (short reads)


45
New cards

DNA-Seq

  • Type of 2nd gen

  • Rapidly sequence whole genomes (eukaryotic, prokaryotic)

  • Deeply sequence target regions (targeted amplification via enrichment)

  • Reduced-representation sequencing (RAD-seq – sequence smaller fraction of genome; uses restriction enzymes)

  • Metagenomics to study microbiomes; identify novel pathogens


46
New cards

RNA-Seq

  • Type of 2nd gen

  • mRNA sequencing to quantify mRNAs for gene expression analysis

  • total RNA sequencing to discover novel RNA variants (non-coding RNAs)

  • Metatranscriptomics to identify novel viruses with RNA genomes


47
New cards

NGS Process

  1. Nucleic Acid Extraction and Isolation:

  • DNA or RNA extraction from source

  1. Library preparation:

  • Fragmentation to produce fragments of desired lengths

  • Adapters ligated at both ends

  1. Clonal amplification & sequencing:

  • Library fragments bind to solid surface (flow cell; beads)

  • Fluorescently labelled nucleotides detected by sequencer

  1. Sequence analysis:

  • Bioinformatics pipelines to process and analyse sequences

  • Each step can vary depending on specifics of the sample, the question being asked, and the experimental design


48
New cards

Metagenomics and Metatranscriptomics

  • Identification of all genomes of species in complex microbial communities

  • Next-generation sequencing (NGS) is commonly used for metagenomics and metatranscriptomics


49
New cards

Features of Metagenomics and Metatranscriptomics

  • Detection of novel organisms: culture/PCR independent profiling. Does not rely on prior knowledge of target sequences, allowing unbiased discovery of previously unknown taxa

  • Short-read compatible: bacterial and viral genomes with small genome sizes, enables high sequencing coverage, facilitating accurate genome assembly and variant detection.

  • Quantitative analysis: Sequencing read counts can be used to estimate relative abundance of taxa, genes, and transcripts across samples.


50
New cards

3rd Generation Sequencing

Long-read/third-generation sequencing technologies are causing a new revolution in genomics as they provide a way to study the structure of genomes, transcriptomes, and metagenomic communities at an unprecedented resolution.

51
New cards

Pacific Biosciences (PacBio)

  • A single-molecule, real-time (SMRT) long-read technology that delivers exceptional read lengths (15,000 to 25,000 base pairs) with over 99.9% accuracy

  • DNA is converted into circular templates using hairpin adapters (SMRTbell) without needing PCR amplification.

  • A polymerase copies the strand inside microscopic wells on a SMRT Cell, emitting light pulses as bases are added.

  • The system reads the same molecule over and over to create an ultra-accurate consensus sequence


52
New cards

Oxford Nanopore

  • Step 1- Nucleic Acid Extraction and Isolation: DNA or RNA extraction from source

  • Step 2- End-prep and adaptor ligation: Adapters ligated at both ends

  • Step 3- Loading & sequencing: DNA fragments loaded onto flow cell and passed through the protein nanopores. Current recorded to enable base-calling

  • Step 4- Sequence analysis: Bioinformatics pipelines to process and analyse sequences


53
New cards

NGS vs Long-read sequencing

  • NGS

    • Very cheap, efficient, and high accuracy

    • Can’t assembly across repetitive sequence

    • Does not identify exon junctions

    • Relies on PCR amplification leading to biased coverage

  • Long-read sequencing

    • Lower accuracy and higher cost per base

    • Detects structural variants and haplotype phasing

    • Identifies alternative splicing

    • Omits PCR bias

    • Characterises nucleotide modifications


54
New cards

Applications of long-read sequencing

  • Improved reference genomes using Hybrid assembly: Long-reads used to scaffold short-read assemblies.

    • Low-coverage long-read sequencing + high-coverage short read sequencing

    • Balance trade-off between accuracy and structural information

  • Whole-transcript sequencing without need for assembly

    • Definitive evidence for alternative splicing, improving accuracy of gene models

    • Enables phasing of transcripts to the alleles from which they were transcribed

  • Directly detect modifications on native DNA

    • Detecting DNA base modifications like DNA methylation allows epigenetic state determination

  • Rapid disease diagnostics

    • Rapid, long amplicon sequencing; in-house capability – no need for external sequencing facility

    • COVID genotyping; biosecurity threats


55
New cards

Whole Genome

  • Whole genome sequencing (WGS): randomized sequencing approach – sequence all that is present in the sample

  • No prior sequence knowledge required – novel genome discovery possible


56
New cards

Amplicon Sequencing

  • Amplicon sequencing: type of targeted sequencing that uses PCR to create DNA sequences called amplicons

  • Must know the sequence of the target so that you can design primers for the PCR step

  • PCR amplification step makes identifying sequence possible in low-yield situations


57
New cards

Application of Whole Genome and Amplicon Sequencing

  • WGS can be used to identify an unknown pathogen in disease diagnostics

  • Amplicon sequencing can be used to identify a known pathogen in disease diagnostics


58
New cards

Which sequencing platform would you use Examples

  1. You are a bee pathologist trying to identify the impact of a new parasite on Australian honey bees. This parasite transmits RNA viruses to the bees and causes colony collapse. You want to identify whether any novel viruses have been transmitted from the parasite to the bees.

  • Short read sequencing

  1. You want to compare different giraffe populations across Africa to see if they are one large meta-population, or if they are distinct populations or species. You have hundreds of samples to analyse and already have a quality reference genome.

  • Short read sequencing

  1. You are a plant pathologist that has observed the tell-tale signs of the exotic Tomato brown rugose fruit virus in a farm. You want to know whether there has been a biosecurity incursion of this specific damaging virus into Australia.

  • Amplicon sequencing

  1. You are interested in characterising the core genome and variable genome sequences across multiple horse breeds, including Thoroughbreds, Icelandic horses, and Przewalski’s Horse. You'd like to ensure your data examines structural variation and chromosomal rearrangements.

  • Long read sequencing


59
New cards

Which sequencing platform would you use?

  • Genome resequencing for SNP/variant calling - short-read DNA sequencing

  • Unbiased analysis of diverse bacterial communities - short-read DNA sequencing (metagenomics)

  • Identification of unknown disease-causing virus - short-read RNA sequencing (metatranscriptomics)

  • High quality reference genome assembly - long read DNA sequencing

  • Hybrid assembly combining - long and short read data

  • Pangenomic analysis of large structural variation - long read DNA sequencing

  • Targeted disease sequencing diagnostics - long read DNA sequencing (amplicons)


60
New cards
61
New cards
62
New cards
63
New cards