1/44
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
The study of all an organism's genes, or its genome.
Studies the structure, function, evolution, and mapping of genomes
Broader than simply reading DNA sequence
It integrates molecular biology, genetics, sequencing technology, statistics, computing, and data science.
Requires both biological and computational expertise.
Studies the physical nature of genomes, including the sequencing and mapping of genomes. It provides the structural and DNA-level variation data used in later analyses.
Genome sequencing
Genetic variants detection such as SNP, Indels, CNV
QTL mapping
GWAS
Other applications of structure and quantitative genomics
Studies the expression and function of the entire genome. It asks what genomic elements do rather than only where they are located.
Transcriptome (RNA-sequencing)
CHIP-Seq
Genome Editing
eQTL & Gene Enrichment and Network Analyses
Epigenomics
Meta-transcriptome
Compares genomes from different organisms. Similarities and differences across species can reveal conserved functions and evolutionary relationships or changes across genomes.

Comparative Genomics example: what gene illustrates an evolutionarily conserved gene required for normal muscle development?
Knowledge of gene structure/function in one species can be applicable or guide interpretation in another.
Example: mutations in myostatin can alter muscle development across multiple species
Advances in sequencing technology and reducing in the cost of sequencing
Computational and data-science innovation
Ex: availability of high-performance computing (HPC) facilities
Advances in data analytics such as machine learning (ML)
Shift from one-gene studies to genome-wide studies.
Lower sequencing cost makes it feasible to obtain genomic data from many individuals or entire populations. Large sample sizes improve discovery and prediction.
Provides the storage and processing power needed for large genomic datasets. Sequencing and variant analyses can involve billions of bases and many samples.
The major challenge is big-data computation and analysis. Generating sequence data is only useful if it can be processed and interpreted.
Cloud platforms can supply scalable computing resources.
Amazon Web Services
Google Cloud
DigitalOcean
Microsoft Azure
Read length and chemistry shape each platform's strengths.
Sanger uses capillary electrophoresis
Illumina uses short-read sequencing by synthesis
Nanopore uses long reads
PacBio HiFi uses highly accurate long reads
Sequencing platforms can generate extremely large amounts of nucleotide data per day
This increasing output is a key source of genomics big-data challenges
A high-quality, preassembled, standardized genomic map used as a template.
Sequencing reads and variants are commonly interpreted relative to this reference.
The number of sequencing reads that cover a particular base or genomic position
Deeper coverage generally increases confidence in a called base or variant.
The fraction of the reference genome that has been covered by reads
A dataset can have deep coverage at some sites but poor coverage across the genome overall.

How do depth and breadth of coverage differ in the diagram?
Depth is the vertical number of overlapping reads at a position
Breadth is the horizontal extent of the reference genome covered by reads
The reference genome is a high-quality, pre-assembled, standardized map
The two measures describe different aspects of sequencing completeness.
The process of identifying differences between sequencing data and a reference genome
it converts aligned reads into candidate genetic variants and genotypes.
Reads are aligned to a reference genome
Variants are called
The calls are filtered and annotated or interpreted
Quality control is needed before a detected difference becomes a reliable variant.
Short variants (i.e SNPs or INDELs) and structural variants (CNVs and STRs). The size and genomic effect of the variant determine its category.
SNPs or SNVs (single-nucleotide variant)
INDELs (insertion/deletion)
They involve a single nucleotide or a relatively small insertion or deletion.
CNVs (copy-number variants)
STRs (short tandem repeats)
Microsatellites
They involve larger-scale changes in genomic structure or repeat number.
The combination of alleles at one locus or multiple loci in an individual. Genotyping determines which alleles an animal carries.
Each animal inherits two sets of alleles (one from dam, one from sire)

A single-nucleotide polymorphism is a single-base substitution at a specific genomic position with a minor allele frequency greater than 1%.
Using sequencing or genotyping assays it’s possible to determine which alleles (i.e. genotype) animals have for a particular SNP locus

Homozygous individual has two copies of the same allele, such as A/A or G/G,
Heterozygous individual has two different alleles, such as A/G.
An animal has maternal allele A and paternal allele G at a SNP. What is its genotype and zygosity?
Genotype is A/G
It is heterozygous
The two inherited alleles differ
Changes a codon so that one amino acid is substituted for another in a protein, which can alter protein function and phenotype.
Example: DGAT1 in cattle which can produce alanine instead of lysine in position 232 of the protein sequence
Individuals that express the lysine version produce more milk with higher fat content

A small insertion or deletion of DNA sequence, generally defined as less than 1,000 base pairs
One genome copy may contain a short sequence that another lacks
Can change peptide protein sequences and then gene function, especially when it shifts the codon reading frame

An insertion or deletion that changes the reading frame of downstream codons. It can dramatically alter the amino-acid sequence and often disrupt protein function.
An 11-base-pair deletion in the third exon of the myostatin, or MSTN, gene causes loss of 102 AAs due to a frameshift
The deletion eliminates functional myostatin, so the gene is incomplete autosomal dominance.
Means that the phenotype of heterozygotes is intermediate rather than identical to either homozygote
MSTN mutation's effect depends on how many mutant alleles are present.

A DNA segment for which different individuals have different copy numbers. The segment can range from about one kilobase to several megabases and may contain multiple genes.
An individual can carry two copies of the same copy-number allele or two different copy-number alleles. CNVs, like SNPs and INDELs, have genotypes.

The KIT gene, which encodes the mast or stem-cell growth-factor receptor.
Can produce a belted phenotype due to duplications of regulatory elements upstream and downstream of the KIT locus
Nonallelic homologous recombination, or NAHR
Fork stalling and template switching, or FoSTeS
Both can create copy-number changes during meiosis
A genotyping array that measures many predefined SNP markers at once. it provides a lower-cost alternative to sequencing every base in every animal.
A statistical inference of unobserved genotypes using observed markers and a reference population. It can increase marker density without measuring every SNP directly.
Genomic selection
QTL mapping or GWAS
Marker panels
Biomarker development
Heterosis
Traceability and parentage verification
Haplotypes and genetic recessives
Genetic diversity, and understanding biology
What was the typical bovine-genomics approach in the 1990s vs 2001?
1990s emphasized DNA markers, linkage mapping, and marker-assisted selection. Studies often used one or a small number of markers.
2001: proposed genome-wide selection based on linkage disequilibrium. This shifted prediction from a few markers toward dense genome-wide marker information
What is linkage mapping?
Uses marker inheritance patterns to locate a specific gene trait (QTL) relative to one or a small number of markers. It relies on genetic linkage between the marker and trait locus.
GWAS tests associations between a trait and many genome-wide markers. Dense marker panels use linkage disequilibrium to improve resolution.
MAS uses one or a small number of genetic markers linked to desirable trait loci to guide selection. It predates genome-wide genomic selection.
Genome-sequence assemblies have replaced them in many species because assembled sequences provide a more complete genomic reference.
It’s shifted from few DNA-marker or interval mapping approaches to high-density panels and GWAS (genome-wide association)
But the underlying principle of associating genomic variation with traits remains the same
Technology changed the resolution, not the core goal.
Genomic prediction estimates genetic merit using DNA markers alone or with other data, while traditional evaluation relies more heavily on pedigree and performance records
Genome prediction is becoming the major genetic evaluation tool
Dense markers capture realized genomic relationships.
New molecular technologies include RNAseq and allele expression
They can identify causative DNA polymorphisms for major genes and may also reveal genes with smaller effects
Genome-wide data increase the chance of finding trait-associated variation.