Mutations and Polymorphisms Study Notes

Terminology and Fundamental Concepts of Genetic Variation

  • Locus (plural: loci): A defined DNA segment occupying a specific physical location on a chromosome.
  • Alleles: Alternative sequence variants of DNA located at a specific locus.
    • Humans possess 2 alleles for each autosomal locus (one inherited maternally and one paternally).
    • Wild-type / Common allele: The single prevailing sequence present in the majority of individuals within a population.
    • Variants / Mutants: All alternative sequence versions differing from the wild-type allele.
    • Polymorphic locus: A locus containing more than 2 common alleles within a population.
    • Private alleles: Rare genetic variants confined exclusively to specific individual families.
  • Zygosity: The degree of similarity between the two alleles at a specific locus in an organism.
    • Homozygous: An individual carrying two identical alleles at a given locus.
    • Heterozygous: An individual carrying two different alleles at a given locus.
  • Genotype: The specific genetic constitution or allelic combination present at a locus.
  • Phenotype: The observable physical, physiological, or biochemical traits of an organism, determined by its genotype and environmental interactions.

Homozygous and Heterozygous chromosome configurations showing allele locations

Overview of Genomic Diversity

  • Source of Diversity: Mutations represent the primary origin of all genetic diversity across species.
  • Sequence Identity: Unrelated human individuals share approximately 99.9%99.9\% sequence identity across the genome.
  • Genetically Determined Variability: Inherited sequence variation accounts for the remaining 0.1%0.1\% of the human genome.
  • Reference Sequence: There is no single standard human genomic sequence. The prevailing sequence found most commonly within a population is designated as the reference sequence.

Classifications and Mechanisms of Mutations

  • Classification by Size:
    • Chromosome Mutations: Structural integrity of individual chromosomes remains intact, but the overall chromosome count is altered.
      • Euploidy: Multiplication of an entire chromosome set (e.g., tetraploidy).
      • Aneuploidy: Addition or loss of individual chromosomes within a set (e.g., trisomy, monosomy).
    • Subchromosomal Mutations: Structural changes or copy number alterations affecting specific regions of chromosomes (e.g., copy number variations, structural rearrangements).
    • DNA Mutations: Small-scale alterations involving nucleotide substitutions, deletions, and insertions up to 100 bp100\,bp
  • Classification by Function: Functional consequences span a continuum from completely benign/non-functional alterations to embryonic lethality.
  • Classification by Heritability: Categorized as germline mutations (inheritable) or somatic mutations (non-inheritable).

Mutation Rates and Mechanisms of Origin

  • Mutation Frequency: Defined as the number of mutations occurring per locus per cell division. It depends directly on:
    1. The occurrence rate of spontaneous and induced nucleotide changes.
    2. The efficiency and probability of DNA repair.
    3. The probability of empirical detection.
  • Regional Variation: Mutation rates vary among genes, species, and genomic regions. Highly susceptible regions are termed mutational hot spots.
  • Disease Mutation Rate: Expressed as the incidence of new cases of a genetic disease caused by a single new mutation that was absent in the parents.
  • Mechanisms of Origin:
    • Chromosome Mutations: Result from mis-segregation of chromosomes during meiosis. These are generally severe and lead to spontaneous fetal abortion.
    • Regional Mutations: Arise from homologous recombination between DNA segments with high sequence homology at different genomic sites, or during the repair of double-strand breaks.
    • Gene Mutations: Occur via DNA replication errors (<1<1 mutation per genome per cell division) or DNA repair failure (spontaneous lesions escaping repair mechanisms).

Nucleotide Substitutions and Splicing Mutations

  • Synonymous Mutations: Base substitutions that do not alter the amino acid sequence of the encoded protein due to the degeneracy of the genetic code.
  • Missense Mutations: Base substitutions that alter a codon to specify a different amino acid.
    • Transitions: Substitutions replacing a purine with another purine (A→GA \rightarrow G or G→AG \rightarrow A) or a pyrimidine with another pyrimidine (C→TC \rightarrow T or T→CT \rightarrow C).
    • Transversions: Substitutions replacing a purine with a pyrimidine (A→CA \rightarrow C, A→TA \rightarrow T, G→CG \rightarrow C, or G→TG \rightarrow T) or a pyrimidine with a purine (C→AC \rightarrow A, C→GC \rightarrow G, T→AT \rightarrow A, or T→GT \rightarrow G).

Transitions and transversions base change categories

Missense mutation changing single nucleotide and codon specifying proline instead of histidine

  • Nonsense Mutations: Point mutations that convert an amino acid-coding codon into a premature stop codon (TAATAA, TAGTAG, or TGATGA), producing truncated and frequently non-functional proteins.

Nonsense mutation introducing premature stop codon resulting in shortened protein

  • Mutations Affecting mRNA Processing: Base changes that destroy canonical splice sites (GTGT donor or AGAG acceptor sites) or generate cryptic, alternative splice junctions, disrupting standard exon-intron splicing patterns.

mRNA splicing mutation introducing novel splice site

Dynamic Mutations and Triplet Expansion Diseases

  • Dynamic Mutations: Pathological expansions of simple trinucleotide repeats located within coding or non-coding regions, caused primarily by DNA strand slippage during replication.

Template slippage mechanism during DNA replication creating extra CAG repeats

  • Triplet Expansion Diseases Summary Table:
DiseaseRepeated Codon (Amino Acid)GeneNormal Repeat RangeDisease Threshold
Huntington diseaseCAG (Gln)HTT11–35>35
Spinocerebellar ataxia type 1CAG (Gln)ATXN16–35>49
Machado-Joseph diseaseCAG (Gln)ATXN312–40>55
Kennedy diseaseCAG (Gln)AR9–36>38
Fragile X syndromeCGG (Arg)FMR15–54>230
Fragile X-E syndromeCCG (Pro)AFF26–35>200
Myotonic dystrophyCTG (Leu)DMPK5–37>50
Friedreich's ataxiaGAA (Glu)FXN7–34>100

Frameshift Mutations, Deletions, Insertions, and Rearrangements

  • Frameshift Mutations: Small insertions or deletions containing a number of nucleotides that is not a multiple of 3. This alters the translational reading frame downstream, resulting in a completely novel amino acid sequence and usually generating a premature stop codon.

Frameshift insertions and deletions shifting reading frame

Functional Consequences of Mutations

  • Gain-of-Function: Causes overproduction, ectopic expression, or constitutive activity of a protein, or imparts a completely novel functional property to the protein product.
  • Loss-of-Function: Leads to reduced or total loss of gene expression or protein function.
    • The normal allele in a heterozygote typically produces sufficient protein for standard function.
    • Haploinsufficiency: Occurs when a 50%50\% reduction in gene product (loss of one allele) is insufficient to maintain normal biological activity.
  • Dominant Negative: The mutant allele synthesizes an altered protein that actively antagonizes or inhibits the function of the co-expressed wild-type protein product in heterozygotes.
  • Lethality: Mutations that severely compromise essential cellular processes, resulting in lethality.

Functional consequences of mutations: Gain-of-function, Loss-of-function, Dominant negative

Heredity: Germline vs. Somatic Mutations

  • Germline Mutations:
    • Inherited directly from parents or occur de novo in germ cells.
    • Transmitted to offspring and present in every somatic cell of the child.
    • De novo germline mutations are extremely rare events.
  • Somatic Mutations:
    • Occur in non-germline tissues during embryonic development or postnatally.
    • Cannot be passed on to future generations.
    • Cause tissue-specific genomic heterogeneity, especially in highly proliferative tissues (e.g., epithelial and hematopoietic lineage cells).
    • Typically undetected by standard bulk DNA sequencing because methods utilize DNA extracted from pooled collections of millions of cells.
    • Represent the key underlying cause of oncogenesis and tumor progression.

Genetic Polymorphisms

  • Definition: Any genetic variant or mutation present at a frequency exceeding 1%1\% (>0.01>0.01) of all alleles within a given population, regardless of its functional effect, size, or genomic location.
  • Variants present at lower frequencies (<1%<1\%) are classified as mutations or rare variants.

Major Types of Genetic Polymorphisms

  1. Single Nucleotide Polymorphisms (SNPs):
    • Involve single base-pair (1 bp1\,bp) substitutions.
    • Occur on average once every 1000 bp1000\,bp, totaling approximately 5×1065 \times 10^6 to 10×10610 \times 10^6 SNPs per human genome.
    • Generally do not produce phenotypic differences.
    • Occur at a 25×25\times higher rate at adjacent cytosine-guanine dinucleotides (CpG sites), which function as mutational hot spots due to cytosine methylation to 5-methylcytosine followed by spontaneous deamination to thymine.
    • Approximately 100,000100{,}000 SNPs reside within protein-coding regions:
      • Synonymous SNPs: Do not alter amino acids.
      • Nonsynonymous SNPs: Alter amino acid codes, creating protein variants.

SNP base replacement diagram

Cytosine methylation to 5-methylcytosine and deamination to thymine

  1. Insertion-Deletion Polymorphisms (Indels):
    • Range from 1 bp1\,bp up to 1000 bp1000\,bp
    • Simple Indels: Presence or absence of a short DNA fragment, generating exactly 2 alleles.
    • Microsatellites / Short Tandem Repeat Polymorphisms (STRs): Variable repeat counts of short tandem units (22, 33, or 4 nt4\,nt long), generating multiple alleles per locus. Used in DNA fingerprinting (evaluating alleles at 1313 loci) to analyze familial relationships and forensic identity.

Indel examples showing 3bp deletion and 4bp insertion

Allele length comparison between unrelated individuals and family members

  1. Copy Number Variants (CNVs):
    • Size ranges from hundreds of base pairs to hundreds of kilobases (kbkb).
    • Can encompass dozens of genes, directly altering gene dosage.
  2. Inversion Polymorphisms:
    • Size ranges from a few base pairs to several megabases (MbMb).
    • Generated by homologous recombination between sequence homologies at boundary flanks.
    • Represent balanced structural changes without net gain or loss of genomic DNA.

Major types of structural variants including mobile element insertion, microsatellites, CNVs, and inversions

Discovery, Validation, and Screening of Genetic Variants

  • Discovery: Initial identification of genetic variants via Whole Genome Sequencing (WGS) or Whole Exome Sequencing (WES) followed by alignment against reference genomes.
  • Validation: Independent replication assays to eliminate sequencing artifacts and determine statistical population frequency.
  • Screening: High-throughput profiling of thousands of SNPs across many individuals using high-density DNA microarrays (SNP arrays).

Methods for Mapping Human Disease Variants

  • Continuous efforts expand a catalog containing tens of millions of variants across diverse human populations.
  • Primary Mapping Strategies:
    1. Linkage Analysis: Family-based co-segregation studies tracking genetic loci through pedigrees.
    2. Association Analysis: Population-based case-control comparisons evaluating allele frequency differences.
    3. Genome Sequencing: Direct high-throughput sequence determination across individuals.

Genome-Wide Association Studies (GWAS)

  • Definition: A molecular technique that simultaneously analyzes hundreds of thousands to millions of single nucleotide markers across the genome in population cohorts to identify statistical associations between genomic loci and specific phenotypes.
  • GWAS vs. Candidate Gene Studies:
    • Candidate Gene Studies: Have higher statistical power, but rely entirely on prior biological knowledge of gene function.
    • GWAS: Provide unbiased discovery of susceptibility loci for complex multifactorial traits without requiring a prior functional hypothesis.
  • Effect Size vs. Allele Frequency Spectrum:
    • Highly Penetrant Mendelian Mutations: Rare variants with very large effect sizes (e.g., CFTR Δ\DeltaF508 in Cystic Fibrosis).
    • Common Variants with Large Effects: Common variants with major effect size (e.g., APOE4 in Alzheimer's disease; CFH in Age-related Macular Degeneration).
    • Less Common Variants with Moderate Effects: Intermediate variants (e.g., NOD2 in Crohn's disease).
    • Common Variants with Small Effects: Identified by GWAS, characterized by low Odds Ratios (1.11.1–1.51.5) (e.g., TNFRSF1A in Multiple Sclerosis, TCF7L2 in Type 2 Diabetes, LMTK2 in Prostate Cancer).

Plot of effect size vs allele frequency for disease variants

  • Study Design and Quality Control Pipeline:
    1. Phenotype & Sample Collection: Case-control cohorts evaluated for traits such as drug efficacy, toxicity, or disease presence.
    2. Whole-Genome Genotyping: Hybridization to SNP microarrays containing hundreds of thousands of probes.
    3. Quality Control (QC):
      • Sample QC: Evaluate sample call quality, remove related individuals, and correct for population stratification using principal component analysis / eigenvectors.
      • SNP QC: Evaluate genotype quality, test for deviation from Hardy-Weinberg equilibrium, and assess allele frequency distributions.
    4. GWAS Analysis: Statistical calculation of association using Quantile-Quantile (Q-Q) plots (comparing observed vs. expected −log⁡10(P)-\log_{10}(P) values; target genomic inflation factor λ≈1.027\lambda \approx 1.027) and Manhattan plots.
    5. Post-GWAS Pipeline:
      • Meta-analysis: Aggregate data across independent study cohorts to calculate combined Odds Ratios (OR) and 95%95\% Confidence Intervals (CI).
      • Functional Analysis: Conduct Electrophoretic Mobility Shift Assays (EMSA) or dual-luciferase reporter assays to quantify transcription factor binding and promoter activity.
      • Downstream Analyses: Gene-based mapping, pathway analysis, polygenic risk scores (PRS), and SNP-SNP interaction studies.
      • Validation: Replicate associations in independent validation cohorts.

GWAS study design pipeline from sample collection to post-GWAS analysis

  • SNP Imputation and Linkage Disequilibrium:
    • Genotype Imputation: Statistical inference of ungenotyped markers in subjects by matching observed genotypes against reference haplotype panels (e.g., HapMap).
    • Linkage Disequilibrium (LD): Non-random co-inheritance of alleles at linked loci across a population.
    • Haplotypes: Specific combinations of closely linked SNPs located on the same chromosome that tend to be inherited together as a block.

Genotype imputation matching observed markers to reference haplotypes

Diagram of haplotype showing linked SNP alleles on a single chromosome

  • Data Interpretation & Manhattan Plots:
    • Manhattan Plot: A scatter plot displaying chromosomal coordinates on the x-axis against association significance (−log⁡10(P-value)-\log_{10}(P\text{-value})) on the y-axis.
    • Genome-Wide Significance Threshold: To account for multiple testing across hundreds of thousands of markers, standard thresholds (P<0.05P < 0.05) are replaced by a genome-wide threshold of P=5×10−8P = 5 \times 10^{-8} (−log⁡10(P)=7.3-\log_{10}(P) = 7.3).
    • 3 Explanations for Significant SNPs:
      1. The SNP is the direct causal functional variant.
      2. The SNP is in linkage disequilibrium (LD) with the true causal locus.
      3. The signal represents a statistical false positive.

Manhattan plot displaying genomic loci association with p-values

Clinical Functional Variants, Eye Color, and Ocular Diseases

  • Functional Categories:
    • Risk Variants: Alleles that increase disease susceptibility.
    • Protective Variants: Alleles that reduce disease susceptibility.
  • Relative Risk (RR):
    • RR=1.0\text{RR} = 1.0: Baseline risk; no association with disease.
    • RR>1.0\text{RR} > 1.0: Increased risk (e.g., RR=1.5\text{RR} = 1.5 represents a 50%50\% risk increase; RR=2.0\text{RR} = 2.0 doubles risk).
    • RR<1.0\text{RR} < 1.0: Decreased risk (e.g., RR=0.5\text{RR} = 0.5 represents a 50%50\% risk reduction).

Relative risk scale showing values greater than 1 as increased risk and less than 1 as decreased risk

  • Eye Color Pigmentation Genes: Pigmentation is polygenic, regulated by key genes including ASIP, HERC2, IRF4, MC1R, OCA2, TYR, TYRP1, and SLC24A5. OCA2 and HERC2 on chromosome 15q13.1 contain extensive intronic and missense variants (e.g., OCA2 His615Arg, Ala481Thr, Val443Ile, Arg419Gln, Arg305Trp).
  • Genes and Onset Profiles of Common Ocular Diseases:
Disease / DisorderAssociated Genes / VariantsAge of Onset
Age-related Macular Degeneration (AMD)NOS2A, CFH, CF, C2, C3, CFB, HTRA1/LOC, MMP-9, TIMP-3, SLC16A8, etc.Older adults
CataractGEMIN4, CYP51A1, RIC1, TAPT1, TAF1A, WDR87, APE1, MIP, Cx50/GJA3 & 8, CRYAA, CRYBB2, PRX, POLR3B, XRCC1, ZNF350, EPHA2, etc.Older adults
GlaucomaCALM2, MPP-7, Optineurin, LOX1, CYP1B1, CAV1/2, MYOC, PITX2, FOXC1, PAX6, CYP1B1, LTBP2, etc.Over 40 years (congenital form affects infants)
Inherited Optic NeuropathiesComplex I or ND genes, OPA1, RPE65, etc.Young males
Marfan SyndromeFBN1, TGFBR2, MTHFR, MTR, MTRR, etc.Born with disorder (congenital); diagnosis may occur later
MyopiaHGF, C-MET, UMODL1, MMP-1/2, PAX6, CBS, MTHFR, IGF-1, UHRF1BP1L, PTPRR, PPFIA2, P4HA2, etc.Typically progresses until age 20
Polypoidal Choroidal VasculopathiesC2, C3, CFH, SERPING1, PEDF, ARMS2-HTRA1, FGD6, ABCG1, LOC387715, CETP, etc.Between ages 50 and 65
Retinitis PigmentosaRPGR, PRPF3, HK1, AGBL5, etc.Between ages 10 and 30
Stargardt's DiseaseABC1, ABCA4, CRB1, etc.Early childhood to middle age
Uveal MelanomaPTEN, BAP1, GNAQ, GNA11, DDEF1, SF3B1, EIF1AX, CDKN2A, p14ARF, HERC2/OCA2, etc.Ages 50 to 80

Personal Direct-to-Consumer Genomics

  • Direct-to-consumer (DTC) platforms (e.g., 23andMe) perform genome-wide polymorphism profiling directly for consumers.
  • Concerns: Data privacy risks and lack of strict regulatory oversight.
  • Reported Phenotypic Traits: Cilantro aversion, hairline morphology, photic sneeze reflex, caffeine metabolism, hair curliness, bitter taste perception, newborn hair quantity, earlobe type, muscle composition, eye color, facial dimples, and sweet taste preference.
  • Carrier Screening: Cystic fibrosis (CFTR), BRCA1-associated cancer risk, sickle cell anemia, glycogen storage disease, maple syrup urine disease, and Sjögren's syndrome.
  • Ancestry & Pharmacogenomics: Ancestral composition, maternal/paternal lineage tracking, and drug response profile predictions.

Clinical Recommendations and Limitations of Genetic Testing

  • American Academy of Ophthalmology (AAO) Recommendations (2013):
    1. Offer genetic testing to patients with inherited eye disorders where the causative gene(s) have been identified.
    2. Utilize CLIA-approved laboratories that report clinical data according to peer-reviewed literature and databases defining variant pathogenicity.
    3. Provide patients with complete reports and findings of genetic testing.
    4. Avoid direct-to-consumer genetic testing kits and direct patients to approved testing facilities.
    5. Order specific targeted genetic tests tailored to the patient's clinical findings.
    6. Avoid repetitive genetic testing for complex diseases unless an actionable treatment plan is established.
    7. Avoid testing asymptomatic minor children unless both parents consent or clear medical cause exists.
  • Clinical Limitations:
    • Gene-disease associations may not apply uniformly across all patient populations.
    • Environmental influences are difficult to predict or quantitatively estimate.
    • Variants primarily offer prognostic value rather than definitive outcome prediction.
    • Calculating cumulative, polygenic risk across multiple low-effect SNPs remains computationally and clinically challenging.