Comprehensive Study Guide on the Human Genome Project and DNA Fingerprinting

Origins and Scope of the Human Genome Project (HGP)

  • Conceptual Foundation: The genetic makeup of an organism is determined by the sequence of bases in its DNA. Variations between individuals suggest that their DNA sequences must differ in at least some locations. This assumption drove the effort to sequence the entire human genome.
  • Technological Context: The project was made possible by the development of genetic engineering techniques for isolating and cloning DNA fragments and the creation of rapid sequencing methods.
  • Launch and Scale: The Human Genome Project (HGP) was launched in 1990 as a "mega project."
  • Economic and Quantitative Estimates:
    • Genome Size: The human genome contains approximately 3×109bp3 \times 10^9\,bp.
    • Cost: Initial estimates placed sequencing at US$3US \$ 3 per base pair, leading to a total project cost of approximately 99 billion US dollars.
    • Data Volume: If the sequence were printed in books, with each page containing 10001000 letters and each book containing 10001000 pages, it would require 33003300 books to store the information from a single human cell.
  • Bioinformatics: The massive data generation required high-speed computational devices for storage, retrieval, and analysis, leading to the rapid growth of Bioinformatics.

Primary Goals of the HGP

  • Gene Identification: Identify all the approximately 20,00020,000 to 25,00025,000 genes in human DNA.
  • Sequencing: Determine the sequences of the 33 billion chemical base pairs.
  • Data Management: Store this information in accessible databases and improve tools for data analysis.
  • Technology Transfer: Transfer related technologies to private sectors and industries.
  • Ethical Oversight: Address the Ethical, Legal, and Social Issues (ELSI) arising from the project.
  • Timeline and Responsibility: The HGP was a 13-year project (completed in 2003) coordinated by the U.S. Department of Energy and the National Institute of Health (NIH). Major partners included the Wellcome Trust (U.K.), with contributions from Japan, France, Germany, and China.

Importance of Non-Human Sequencing

  • Comparative Biology: Sequencing non-human organisms provides insight into natural capabilities that can solve challenges in health care, agriculture, energy, and environmental remediation.
  • List of Sequenced Model Organisms:
    • Bacteria and Yeast.
    • Caenorhabditis elegans (a free-living non-pathogenic nematode).
    • Drosophila (the fruit fly).
    • Plants including Rice and Arabidopsis.

Methodologies and Sequencing Techniques

  • Primary Approaches:
    1. Expressed Sequence Tags (ESTs): Focusing on identifying all genes expressed as RNA.
    2. Sequence Annotation: Sequencing the entire genome (coding and non-coding) and then assigning functions to different regions.
  • Technical Process:
    • Fragmentation: Total DNA is isolated and converted into random fragments of smaller sizes due to technical limitations in sequencing long polymers.
    • Cloning and Amplification: Fragments are cloned into hosts like Bacteria and Yeast using specialized vectors: BAC (Bacterial Artificial Chromosomes) and YAC (Yeast Artificial Chromosomes).
    • Sanger Method: Sequences were determined using automated DNA sequencers based on the principle developed by Frederick Sanger (who also developed methods for sequencing amino acids in proteins).
    • Alignment: Sequences were arranged based on overlapping regions. Special computer programs were required as human alignment was impossible.
  • Chromosome 1: This was the last of the 24 human chromosomes (22 autosomes, X, and Y) to be sequenced, with work completed in May 2006.
  • Mapping: Genetic and physical maps were generated using data on polymorphisms of restriction endonuclease recognition sites and repetitive DNA sequences known as microsatellites.

Salient Features of the Human Genome

  • Total Bases: The human genome contains 3164.73164.7 million nucleotide bases.
  • Gene Statistics:
    • Average Gene: Contains 30003000 bases, though sizes vary.
    • Largest Gene: Dystrophin at 2.42.4 million bases.
    • Gene Count: Approximately 30,00030,000 genes, significantly lower than previous estimates of 80,00080,000 to 140,000140,000.
  • Functionality and Distribution:
    • Shared Sequence: 99.9%99.9\% of nucleotide bases are identical in all people.
    • Unknown Functions: Over 50%50\% of discovered genes have unknown functions.
    • Coding Proportion: Less than 2%2\% of the genome codes for proteins.
    • Repetitive Sequences: Comprise a very large portion and are thought to have no direct coding function, but offer insights into chromosome structure and evolution.
    • Chromosomal Variation: Chromosome 1 has the most genes (29682968), while the Y chromosome has the fewest (231231).
  • SNPs: Scientists identified about 1.41.4 million locations where Single Nucleotide Polymorphisms (pronounced 'snips') occur, aiding in disease location and human history tracking.

DNA Fingerprinting: Principles and Mechanism

  • The Uniqueness Factor: Since 99.9%99.9\% of the genome is identical, differences in the remaining 0.1%0.1\% (3×106bp3 \times 10^6\,bp) make individuals unique. DNA fingerprinting is a rapid way to compare sequences without sequencing the entire 3×109bp3 \times 10^9\,bp genome.
  • Repetitive DNA: The focus is on specific regions where a small stretch of DNA is repeated many times.
  • Satellite DNA: During density gradient centrifugation, bulk genomic DNA forms a major peak, while repetitive DNA forms smaller peaks called satellite DNA.
  • Classification: Classified based on base composition (A:T rich or G:C rich), segment length, and repetition count (e.g., micro-satellites, mini-satellites).
  • Polymorphism: These sequences show a high degree of polymorphism, forming the basis for forensic identification and paternity testing because they are inheritable.

Understanding DNA Polymorphism

  • Definition: A variation at the genetic level arising from mutations. It is traditionally described as a polymorphism if more than one variant (allele) at a locus occurs in a population with a frequency greater than 0.010.01.
  • Mechanism: Mutations in non-coding DNA accumulate across generations because they do not immediately impact reproductive ability.
  • Significance: These polymorphisms are critical for evolution and speciation.

The Technique of DNA Fingerprinting

  • Development: Initially developed by Alec Jeffreys.
  • VNTR (Variable Number of Tandem Repeats): Jeffreys used satellite DNA as a probe. VNTRs are a class of mini-satellites where small sequences are arranged tandemly. The copy number varies between individuals, leading to sizes from 0.10.1 to 20kb20\,kb.
  • Methodology Steps:
    1. Isolation: Extract DNA from tissues (blood, hair, skin, bone, saliva, sperm).
    2. Digestion: Use restriction endonucleases to cut DNA.
    3. Electrophoresis: Separate DNA fragments by size.
    4. Blotting: Transfer fragments to synthetic membranes (nitrocellulose or nylon).
    5. Hybridization: Use radiolabelled VNTR probes.
    6. Detection: Use autoradiography to identify hybridised fragments, resulting in a characteristic banding pattern (the autoradiogram).
  • Sensitivity: The use of Polymerase Chain Reaction (PCR) allows analysis even from a single cell.