Human Molecular Genetics: Why Sequence the Human Genome?

Lecture 22 Objectives

After revising this material, students should be able to fulfill the following learning objectives:

  • Explain the specific primary motivations and goals for why the human genome was sequenced.
  • Outline the critical key findings discovered upon the completion of the human genome sequence.
  • Explain the biological and medical importance of variation within the human genome.
  • Describe the various types of genetic variation found in the human genome, including their scales and frequencies.

The Human Genome Project (HGP)

The Human Genome Project was an international scientific research project that officially began in 19901990. The core aims of the project included:

  • Gene Identification: To identify all human protein-coding genes and determine their specific biological roles.
  • Genetic Variation Analysis: To analyze and map the extent of genetic variation among individual humans.
  • Model Organism Sequencing: To sequence the genomes of several model organisms (such as bacteria, yeast, nematodes, and fruit flies) commonly used in genetic research.
  • Technological Advancements: To develop new, faster sequencing techniques and more efficient computational tools for genomic analysis.
  • Open Access Data: To share genome information with the global scientific community and the general public as rapidly as possible.

The Human Reference Genome

  • Genome Definition: The complete set of DNA of an organism, encompassing all of its genes.
  • Genomics Definition: The study of genomes and their functions.
Comparison of Human DNA Sets

The human genome is split between nuclear and mitochondrial components:

  • Nuclear DNA:

    • Structure: Consists of 2222 autosomes plus the XX and YY sex chromosomes.
    • Size: Approximately 6×1096 \times 10^9 (6 billion) base pairs.
    • Inheritance: Half is inherited from each parent.
    • Gene Count: Contains fewer than 20,00020,000 genes.
  • Mitochondrial DNA:

    • Structure: Many copies of a single, circular DNA molecule.
    • Size: Exactly 16,56916,569 base pairs.
    • Inheritance: Inherited entirely from the mother.
    • Gene Count: Contains exactly 3737 genes.

Key Findings of the Human Genome

The successful sequencing of the human genome yielded several surprising and important findings:

  • Gene Count: There are fewer total protein-coding genes than scientists initially expected (approximately 21,30021,300).
  • Coding Sequence Proportion: Less than 2.0%2.0\% of the human genome actually codes for proteins.
  • Dynamic Nature: The genome is not static; it is dynamic and subject to change.
  • Knowledge Gaps: The functions of many protein-coding genes identified remain unknown (about 20%20\%).
  • Evolutionary Relationships: The majority of human genes are related to those found in other animals.
  • Sequence Similarity: All humans, regardless of ethnicity or race, are 99.9%99.9\% similar at the DNA sequence level.
Comparative Genome Sizes and Gene Numbers
OrganismHaploid Genome Size (MbMb)Estimated Number of GenesGenes per MbMb
Bacteria
Haemophilus influenzae1.81.81,7001,700940940
Escherichia coli4.64.64,4004,400950950
Archaea
Archaeoglobus fulgidus2.22.22,5002,5001,1301,130
Methanosarcina barkeri4.84.83,6003,600750750
Eukaryotes
Saccharomyces cerevisiae (Yeast)12126,3006,300525525
Utricularia gibba (Floating bladderwort)828228,50028,500348348
Caenorhabditis elegans (Nematode)10010020,10020,100200200
Arabidopsis thaliana (Mustard plant)12012027,00027,000225225
Drosophila melanogaster (Fruit fly)16516514,00014,0008585
Daphnia pulex (Water flea)20020031,00031,000155155
Zea mays (Corn)2,3002,30032,00032,0001414
Ailuropoda melanoleuca (Giant panda)2,4002,40021,00021,00099
Homo sapiens (Human)3,0003,00021,30021,30077
Paris japonica (Japanese canopy plant)149,000149,000NDND

Note: MbMb = million base pairs.

Detailed Composition of the Human Genome

Analysis of the human genome reveals the following breakdown of DNA types:

  • Exons (Protein-coding): Approximately 1.5%1.5\% to <2%<2\%.
  • Regulatory Sequences: 5%5\%.
  • Introns: Approximately 20%20\%.
  • Unique Non-coding DNA: 15%15\%.
  • Repetitive DNA (Transposable elements and related): 44%44\%.
    • L1 Sequences: 17%17\%.
    • Alu Elements: 10%10\%.
  • Large-segment Duplications: 56%5-6\%.
  • Simple Sequence DNA: 3%3\%.

Functional Classification of Human Proteins

A large portion (23.6%23.6\% or 40614061 genes) remains unclassified. Known functions include:

  • Transcription factors: 20672067 genes (12.0%12.0\%
  • Transfer/carrier proteins: 248248 genes (1.4%1.4\%
  • Hydrolases: 454454 genes (2.6%2.6\%
  • Transporters: 10981098 genes (6.4%6.4\%
  • Transferases: 15121512 genes (8.8%8.8\%
  • Nucleic acid binding: 14661466 genes (8.5%8.5\%
  • Signaling molecules: 961961 genes (5.6%5.6\%
  • Receptors: 10761076 genes (6.3%6.3\%
  • Enzyme modulators: 857857 genes (5.0%5.0\%
  • Cytoskeletal proteins: 441441 genes (2.6%2.6\%
  • Oxidoreductases: 550550 genes (3.2%3.2\%
  • Proteases: 476476 genes (2.8%2.8\%

Variation in the Human Genome

While humans are 99.9%99.9\% similar, the remaining 0.1%0.1\% (approx. 3×1063 \times 10^6 base pairs) accounts for all the genetic diversity between individuals. Variation ranges from single-base changes to massive chromosomal rearrangements.

Single Nucleotide Polymorphisms (SNPs)
  • Definition: SNPs are common sites in the DNA where a single nucleotide (A, T, C, or G) varies within a population.
  • Frequency: They occur approximately 11 in every 300300 nucleotides.
  • Prevalence: There are roughly 1.91.9 million common SNPs mapped in the human population.
  • Inheritance: Your SNPs are mostly inherited from your biological parents.
  • Impact of SNPs:
    • Linked SNPs: Found outside of genes; they have no effect on protein production or function but serve as useful genetic markers.
    • Causative SNPs: Found within genes.
      • Non-coding SNP: Located in regulatory regions; can change the amount of protein produced.
      • Coding SNP: Located in exons; can change the amino acid sequence of the protein.
  • Utility of Analysis (Genotyping):
    • Determining relatedness and ancestry/origins.
    • Identifying disease risk or associations.
    • Predicting physical traits (e.g., hair loss, muscle type).
    • Predicting drug responses (Pharmacogenomics).
    • Forensic applications for crime solving.
Short Tandem Repeats (STRs)
  • Definition: Repeats of 252-5 nucleotides found in specific, known regions of the genome.
  • Inheritance: Each person inherits two alleles (one from each parent), which may differ in length (number of repeats).
  • DNA Profiling: STRs are used to create "DNA fingerprints."
  • Example Case: Earl Washington was exonerated after 1717 years in prison for a murder he did not commit. STR profiling of semen found on the victim matched Kenneth Tinsley (17,1917, 19; 13,1613, 16; 12,1212, 12 at markers 1, 2, and 3 respectively) and did not match Earl Washington (16,1816, 18; 14,1514, 15; 11,1211, 12).
InDels (Insertions and Deletions)
  • Definition: Small insertions or deletions of nucleotides.
  • Prevalence: Approximately 0.20.2 million InDels are common in the human genome; they are the second most common variant type.
  • Frame Shift: InDels in protein-coding regions can cause a "frame shift," altering how the entire DNA sequence is read.
    • Analogy: Deleting two letters from "Why did the dog run out?" results in "Wdi dth edo gru nou t?"
  • Disease Example: Cystic Fibrosis is most commonly caused by the CFTRΔF508CFTRΔF508 mutation, which is a 33-nucleotide deletion.
Structural Variants: Copy Number Variations (CNVs)
  • Definition: Chunks of DNA larger than 500bp500bp that are present in different amounts (copy numbers) compared to a reference genome.
  • Prevalence: Humans typically have around 10,00010,000 CNVs.
  • Genomic Context: They can span multiple genes and are often associated with immunity and sensory perception (e.g., smell).
  • Example: Extra copies of the AMY1 gene allow for better starch digestion. While the range in humans is 22 to 2727 copies, most people possess 676-7 copies.

The Rise of Genomics: A Timeline

  • 1977: Fred Sanger manually sequences a 5.4Kbp5.4Kbp bacteriophage.
  • 1995: J. Craig Venter sequences the first bacterial genome (H. influenzae, 1830Kbp1830Kbp) using "shotgun" sequencing.
  • 1998: The Human Genome Consortium (HGC) sequences the first multicellular organism, the nematode C. elegans (97Mbp97Mbp).
  • 2001: The HGC completes the first Human Genome reference sequence (3.2Gbp3.2Gbp). It cost US$2.7US\$2.7 billion and took 1313 years.
  • 2003: The draft sequence is refined.
  • 2007: James Watson’s diploid genome is sequenced in two months for US$1US\$1 million.
  • 2012: The 10001000 Genomes Project is completed.
  • 2018: 1010 million genotypes worldwide; approaching 11 million full genomes.
  • 2020-2022: Efforts like the Telomere-to-Telomere Consortium and H3Africa (Human Heredity & Health in Africa) continue to increase the diversity and completeness of genomic data.

Future Directions: Where to From Here?

  • Evolutionary Origins: Using variation as a signature of descent to understand where humans came from.
  • Unanalyzed Genes: Continuing to identify the functions of the thousands of genes still not understood.
  • Complex Diseases: Understanding polygenic (result of many genes) and rare diseases.
  • Personalised Medicine (Pharmacogenomics): Determining which drugs will be most effective for an individual and which should be avoided based on SNP profiles (e.g., Albuterol response in asthma patients based on late-stage clinical data of the ADRB2ADRB2 gene).
  • Ethics: Addressing who defines ownership and access rights to personal genomic data.

Summary Key Points

  • The HGP aimed to find all genes and identify the extent of variation.
  • There are approximately 20,00020,000 protein-coding genes (<2%<2\% of the genome).
  • Humans are 99.9%99.9\% identical; African genomes typically show the most variation.
  • Variation is a driver of evolution; most is inherited, but individuals have unique variants.
  • SNPs are the most common variant, followed by InDels, STRs, and CNVs.