Fundamentals of Genomics and Human Genome Projects Study Notes
Background and Foundation of Modern Genetics
The completion of the Human Genome Project (HGP) in 2003, followed by subsequent genome-wide studies, initiated a new era in human health and disease management.
These advancements have led to:
Increased accuracy in genetic diagnoses.
Enhanced understanding of the pathogenesis of inherited conditions.
Development of new and improved treatments.
Impact on Dentistry: Dental health professionals require a greater understanding of genetics to provide better and more personalized patient care in the modern practice landscape.
The Discovery of the Double Helix (1951-1953)
Structural Characteristics:
The DNA molecule is a double helix with two main grooves: the Major groove and the Minor groove.
The diameter of the helix is precisely .
The distance of one full turn in the helix is .
The vertical distance between individual base pairs is .
Chemical Composition and Orientation:
DNA strands have an antiparallel orientation: one strand runs in the direction, while the complementary strand runs .
Nucleotides consist of a sugar-phosphate (S-P) backbone.
Bases are categorized as Purines (Adenine, Guanine) and Pyrimidines (Thymine, Cytosine).
Base Pairing rules (Hydrogen bonding):
Adenine (A) pairs with Thymine (T) via two hydrogen bonds.
Guanine (G) pairs with Cytosine (C) via three hydrogen bonds.
The Human Genome Project (1990-2003)
The HGP was an international scientific research project designed to:
Determine the sequence of nucleotide base pairs in human DNA.
Identify and map all genes from both a physical and functional standpoint.
Conceptual Impact (Francis Collins, 2000):
Described as a "history book" of the human species journey.
A "shop manual" with a detailed blueprint for building every human cell.
A "transformative textbook of medicine" to treat, prevent, and cure disease.
Sampling and Donors:
The project utilized 20 volunteers recruited via the Clinical Genetics Service at Roswell Park Cancer Institute (Buffalo, NY).
Participants were at least 18 years old and had not undergone chemotherapy.
The sequence is a "mosaic" representation rather than a single individual; it is referred to as a "reference genome."
Numerical breakdown of donor contributions: of the HGP sequence came from 11 donors, and came from a single donor.
Note on "Normalcy": Some researchers advocated for sequencing a "normal" person first, though this was dismissed as it is impossible to define a biologically "normal" human.
Publication History:
Draft sequences were reported on February 15, 2001, in Nature (representing the government/NIH effort) and February 16, 2001, in Science (representing the private company effort).
Achieving a Complete Human Genome Sequence
The original project took 13 years to produce a sequence of roughly base pairs.
In April 2022, the Telomere-to-Telomere (T2T) consortium published a truly complete sequence in Science.
The finalized sequence encompasses DNA base pairs.
Important technical caveats:
of the sequence might still contain errors.
The current complete count includes the X chromosome but excludes the Y chromosome.
Counts generally exclude mitochondrial DNA.
The final 10% of the genome was considered the most difficult to sequence due to its repetitive nature.
Defining Genetics vs. Genomics
Genetics:
Focuses on single genes one at a time.
Acts like a "picture" or "snapshot" of specific hereditary units.
Genomics:
Examines the big picture, viewing all genes and the entire genome as an integrated system.
Deals with the fact that not all genes are the same length.
Acknowledges that of DNA is non-protein-coding (formerly referred to as "junk DNA").
Formal definition of Genomics: The sub-discipline devoted to sequencing (ordering bases), mapping (determining positions), and functional analysis of whole-genome sequences.
The 'Omics' Hierarchy and Multiomics
The biological flow of information involves multiple levels of data:
DNA (Genotype): Analyzed via SNPs, CNVs, and Microsatellites using re-sequencing and ChIP-chip.
mRNA (Transcriptomics): Involves gene expression, microarrays, and splice junction analysis.
Protein (Proteomics): Translation products analyzed via protein arrays, mass spectrometry, and 2D-gel electrophoresis.
Metabolite (Metabolomics): Products of cellular processes analyzed via NMR and mass spectrometry.
Phenotype (Phenomics): Observations of clinical phenotypes (disease status) and quantitative traits such as Body Mass Index () and blood pressure.
NIH Multiomics Consortium:
Announced Sept 12, 2023.
Awarded over five years.
Goal: Use high-throughput molecular assays to generate molecular profiles of disease and non-disease states through integrated data (genomic, epigenomic, transcriptomic, etc.).
The Economics of Genomic Technology
Sequencing costs have plummeted drastically since 2001.
Cost Timeline:
2001: Approximately .
2015: Under .
2020: Illumina launched the NovaSeq 6000 v1.5 kit, introducing the " genome."
2024: Ultima Genomics announced high-end sequencers intended to read a genome for as little as "."
Comparison to Moore's Law: While Moore's Law suggests computational power doubles every 12 to 18 months, the drop in sequencing costs has significantly outpaced this trend.
Single Nucleotide Polymorphisms (SNPs) and Haplotypes
SNPs (Single Nucleotide Polymorphisms):
The most important and common type of genetic variation.
Consists of single-nucleotide point mutations.
Occur approximately every base pairs.
Most are biallelic (meaning they have only two possible alleles, such as A or G), though some can have three or four.
Serve as markers for disease and help find a "handful" of relevant variants instead of checking the entire genome.
Haplotypes:
Groups of SNPs that are linked (close together on the same chromosome) and inherited together.
Examples from population studies include:
Haplotype 1:
Haplotype 2:
Haplotype 3:
Haplotype 4:
Diversity in Global Genome Projects
1000 Genomes Project:
Aimed to find common genetic variants with frequencies of at least in populations.
Final data set: 2,504 individuals from 26 populations.
Genome Aggregation Database (gnomAD v3):
Contains 76,156 genomes of diverse ancestries.
Major represented groups: Non-Finnish European (34,029), African/African American (20,744), Latino/Admixed American (7,647), Finnish (5,316).
Ancestry Bias in Research:
Genome-wide association studies (GWAS) are heavily skewed.
European:
Asian:
African:
Other:
Targeted Diversity Projects:
African Genome Variation Project: 1,481 individuals from 18 ethno-linguistic groups in sub-Saharan Africa.
Nigerian 100K Genome Project: Focuses on 200 ethnic groups and 500 different languages to capture the high genetic diversity in Africa.
Functional and Phenotypic Mapping
ENCODE Project (Encyclopedia of DNA Elements):
Goal: Build a comprehensive "parts list" of functional elements.
Includes regulatory elements (promoters, enhancers, repressors/silencers, insulators) that control gene activity.
NIH Gene Function Initiative (Sept 27, 2022):
Goal: Systematically establish the function of every human gene.
Current Status: We currently only know the function of approximately out of protein-coding genes.
Human Phenotype Project (HPP):
Focuses on deep phenotyping along the health-disease continuum.
Profiling includes longitudinal data: medical history, nutrition, anthropometrics, continuous glucose monitoring, microbiome (gut, vaginal, and oral), and immune profiling.
Aims to advance precision medicine by exploring variations in disease susceptibility and aging.
Study Methodologies: GWAS vs. PheWAS
GWAS (Genome-Wide Association Study):
Direction: Trait/Disease Gene Variant.
Approach: Starts with a specific disease trait (e.g., cardiovascular disease) in a case-control study and scans markers to find associated variants (e.g., Variant A).
PheWAS (Phenome-Wide Association Study):
Direction: Gene Variant Phenome (all traits).
Approach: Starts with a specific gene variant (e.g., Variant A) and looks across the entire phenome to see which diseases or traits (respiratory, cardiovascular, etc.) it correlates with.
Synthetic Genomics
Researchers in the UK are working toward creating the first synthetic human chromosome.
Writing large genomes chemically has potentials in:
Developing cell therapies for diseased individuals.
Creating climate-resistant crops.
Transformative understanding of human health.
Questions & Discussion
Questions for consideration in the field of genomics:
How do we synthesize omics data to understand disease?
How do we move from phenome data to link it back to the genome?
What are the implications of newly synthetic or artificial sequences?
Thought for the Day
"The meaning of life is to find your gift. The purpose of life is to give it away." (Pablo Picasso)