Human Molecular Genetics: Why Sequence the Human Genome?
Lecture 22 Objectives
After revising this material, students should be able to fulfill the following learning objectives:
- Explain the specific primary motivations and goals for why the human genome was sequenced.
- Outline the critical key findings discovered upon the completion of the human genome sequence.
- Explain the biological and medical importance of variation within the human genome.
- Describe the various types of genetic variation found in the human genome, including their scales and frequencies.
The Human Genome Project (HGP)
The Human Genome Project was an international scientific research project that officially began in . The core aims of the project included:
- Gene Identification: To identify all human protein-coding genes and determine their specific biological roles.
- Genetic Variation Analysis: To analyze and map the extent of genetic variation among individual humans.
- Model Organism Sequencing: To sequence the genomes of several model organisms (such as bacteria, yeast, nematodes, and fruit flies) commonly used in genetic research.
- Technological Advancements: To develop new, faster sequencing techniques and more efficient computational tools for genomic analysis.
- Open Access Data: To share genome information with the global scientific community and the general public as rapidly as possible.
The Human Reference Genome
- Genome Definition: The complete set of DNA of an organism, encompassing all of its genes.
- Genomics Definition: The study of genomes and their functions.
Comparison of Human DNA Sets
The human genome is split between nuclear and mitochondrial components:
Nuclear DNA:
- Structure: Consists of autosomes plus the and sex chromosomes.
- Size: Approximately (6 billion) base pairs.
- Inheritance: Half is inherited from each parent.
- Gene Count: Contains fewer than genes.
Mitochondrial DNA:
- Structure: Many copies of a single, circular DNA molecule.
- Size: Exactly base pairs.
- Inheritance: Inherited entirely from the mother.
- Gene Count: Contains exactly genes.
Key Findings of the Human Genome
The successful sequencing of the human genome yielded several surprising and important findings:
- Gene Count: There are fewer total protein-coding genes than scientists initially expected (approximately ).
- Coding Sequence Proportion: Less than of the human genome actually codes for proteins.
- Dynamic Nature: The genome is not static; it is dynamic and subject to change.
- Knowledge Gaps: The functions of many protein-coding genes identified remain unknown (about ).
- Evolutionary Relationships: The majority of human genes are related to those found in other animals.
- Sequence Similarity: All humans, regardless of ethnicity or race, are similar at the DNA sequence level.
Comparative Genome Sizes and Gene Numbers
| Organism | Haploid Genome Size () | Estimated Number of Genes | Genes per |
|---|---|---|---|
| Bacteria | |||
| Haemophilus influenzae | |||
| Escherichia coli | |||
| Archaea | |||
| Archaeoglobus fulgidus | |||
| Methanosarcina barkeri | |||
| Eukaryotes | |||
| Saccharomyces cerevisiae (Yeast) | |||
| Utricularia gibba (Floating bladderwort) | |||
| Caenorhabditis elegans (Nematode) | |||
| Arabidopsis thaliana (Mustard plant) | |||
| Drosophila melanogaster (Fruit fly) | |||
| Daphnia pulex (Water flea) | |||
| Zea mays (Corn) | |||
| Ailuropoda melanoleuca (Giant panda) | |||
| Homo sapiens (Human) | |||
| Paris japonica (Japanese canopy plant) | ND | ND |
Note: = million base pairs.
Detailed Composition of the Human Genome
Analysis of the human genome reveals the following breakdown of DNA types:
- Exons (Protein-coding): Approximately to .
- Regulatory Sequences: .
- Introns: Approximately .
- Unique Non-coding DNA: .
- Repetitive DNA (Transposable elements and related): .
- L1 Sequences: .
- Alu Elements: .
- Large-segment Duplications: .
- Simple Sequence DNA: .
Functional Classification of Human Proteins
A large portion ( or genes) remains unclassified. Known functions include:
- Transcription factors: genes (
- Transfer/carrier proteins: genes (
- Hydrolases: genes (
- Transporters: genes (
- Transferases: genes (
- Nucleic acid binding: genes (
- Signaling molecules: genes (
- Receptors: genes (
- Enzyme modulators: genes (
- Cytoskeletal proteins: genes (
- Oxidoreductases: genes (
- Proteases: genes (
Variation in the Human Genome
While humans are similar, the remaining (approx. base pairs) accounts for all the genetic diversity between individuals. Variation ranges from single-base changes to massive chromosomal rearrangements.
Single Nucleotide Polymorphisms (SNPs)
- Definition: SNPs are common sites in the DNA where a single nucleotide (A, T, C, or G) varies within a population.
- Frequency: They occur approximately in every nucleotides.
- Prevalence: There are roughly million common SNPs mapped in the human population.
- Inheritance: Your SNPs are mostly inherited from your biological parents.
- Impact of SNPs:
- Linked SNPs: Found outside of genes; they have no effect on protein production or function but serve as useful genetic markers.
- Causative SNPs: Found within genes.
- Non-coding SNP: Located in regulatory regions; can change the amount of protein produced.
- Coding SNP: Located in exons; can change the amino acid sequence of the protein.
- Utility of Analysis (Genotyping):
- Determining relatedness and ancestry/origins.
- Identifying disease risk or associations.
- Predicting physical traits (e.g., hair loss, muscle type).
- Predicting drug responses (Pharmacogenomics).
- Forensic applications for crime solving.
Short Tandem Repeats (STRs)
- Definition: Repeats of nucleotides found in specific, known regions of the genome.
- Inheritance: Each person inherits two alleles (one from each parent), which may differ in length (number of repeats).
- DNA Profiling: STRs are used to create "DNA fingerprints."
- Example Case: Earl Washington was exonerated after years in prison for a murder he did not commit. STR profiling of semen found on the victim matched Kenneth Tinsley (; ; at markers 1, 2, and 3 respectively) and did not match Earl Washington (; ; ).
InDels (Insertions and Deletions)
- Definition: Small insertions or deletions of nucleotides.
- Prevalence: Approximately million InDels are common in the human genome; they are the second most common variant type.
- Frame Shift: InDels in protein-coding regions can cause a "frame shift," altering how the entire DNA sequence is read.
- Analogy: Deleting two letters from "Why did the dog run out?" results in "Wdi dth edo gru nou t?"
- Disease Example: Cystic Fibrosis is most commonly caused by the mutation, which is a -nucleotide deletion.
Structural Variants: Copy Number Variations (CNVs)
- Definition: Chunks of DNA larger than that are present in different amounts (copy numbers) compared to a reference genome.
- Prevalence: Humans typically have around CNVs.
- Genomic Context: They can span multiple genes and are often associated with immunity and sensory perception (e.g., smell).
- Example: Extra copies of the AMY1 gene allow for better starch digestion. While the range in humans is to copies, most people possess copies.
The Rise of Genomics: A Timeline
- 1977: Fred Sanger manually sequences a bacteriophage.
- 1995: J. Craig Venter sequences the first bacterial genome (H. influenzae, ) using "shotgun" sequencing.
- 1998: The Human Genome Consortium (HGC) sequences the first multicellular organism, the nematode C. elegans ().
- 2001: The HGC completes the first Human Genome reference sequence (). It cost billion and took years.
- 2003: The draft sequence is refined.
- 2007: James Watson’s diploid genome is sequenced in two months for million.
- 2012: The Genomes Project is completed.
- 2018: million genotypes worldwide; approaching million full genomes.
- 2020-2022: Efforts like the Telomere-to-Telomere Consortium and H3Africa (Human Heredity & Health in Africa) continue to increase the diversity and completeness of genomic data.
Future Directions: Where to From Here?
- Evolutionary Origins: Using variation as a signature of descent to understand where humans came from.
- Unanalyzed Genes: Continuing to identify the functions of the thousands of genes still not understood.
- Complex Diseases: Understanding polygenic (result of many genes) and rare diseases.
- Personalised Medicine (Pharmacogenomics): Determining which drugs will be most effective for an individual and which should be avoided based on SNP profiles (e.g., Albuterol response in asthma patients based on late-stage clinical data of the gene).
- Ethics: Addressing who defines ownership and access rights to personal genomic data.
Summary Key Points
- The HGP aimed to find all genes and identify the extent of variation.
- There are approximately protein-coding genes ( of the genome).
- Humans are identical; African genomes typically show the most variation.
- Variation is a driver of evolution; most is inherited, but individuals have unique variants.
- SNPs are the most common variant, followed by InDels, STRs, and CNVs.