study guide: Mutations and Polymorphisms Flashcards
Fundamental Terminology of Human Genetics
Locus (plural: Loci): A specific, defined physical location or segment of DNA on a chromosome.
Alleles: Alternative sequence variants of DNA present at a given chromosomal locus.
Human cells carry two alleles at each autosomal locus: one inherited maternally and one paternally.
Wild-Type (Common Allele): The single prevailing allele present in the majority of individuals within a population.
Variants / Mutant Alleles: All alternative versions of a DNA sequence differing from the wild-type allele.
Polymorphic Locus: A locus that possesses more than two common alleles in the general population.
Private Alleles: Rare genetic variants restricted exclusively to single individuals or specific family lineages.
Zygosity: The degree of genetic similarity between the two alleles present at a specific locus in a diploid organism:
Homozygous: Possessing two identical alleles at a given locus (e.g., or ).
Heterozygous: Possessing two distinct alleles at a given locus (e.g., or ).

Genotype: The specific genetic constitution or sequence information of an organism at a given locus (or set of loci).
Phenotype: The observable physical, physiological, biochemical, or clinical traits of an organism resulting from the interaction of its genotype with environmental factors.
Genomic Identity and Variation:
Human genomes exhibit approximately sequence identity between any two unrelated individuals.
The remaining of the genome accounts for all genetically determined human diversity.
There is no single universal human genome sequence; the standard against which individual genomes are evaluated is termed the reference sequence, representing the most common sequence found in a population.
Classification and Mechanisms of Mutations
Broad Classification Schemes
Classification by Genomic Size:
Chromosome Mutations: Alterations in chromosome number while maintaining intact internal chromosome structure.
Euploidy: Multiplication of the entire haploid chromosome set (e.g., polyploidy, tetraploidy).
Aneuploidy: Gain or loss of individual intact chromosomes relative to the normal diploid set (e.g., trisomy, monosomy).
Subchromosomal (Regional) Mutations: Structural alterations involving larger portions or fragments of chromosomes, including copy number variations (CNVs) and structural rearrangements.
DNA (Gene) Mutations: Small-scale sequence changes altering point nucleotides up to , including substitutions, small insertions, and small deletions.
Classification by Function: Spectrum ranging from completely neutral/non-functional alterations to severe functional disruptions and lethality.
Classification by Heritability:
Germline Mutations: Occur in reproductive cells (sperm or egg) or their precursor lineage. These mutations can be passed on to offspring.
Somatic Mutations: Occur in non-germline body cells. They cannot be inherited by offspring.
Mutation Rates and Molecular Origins
Mutation Frequency: Defined as the number of mutations per locus per cell division. It depends on three distinct parameters:
The baseline frequency of spontaneous and induced nucleotide alterations.
The intrinsic probability and efficiency of DNA repair machinery.
The probability of detection.
Variability: Mutation rates vary significantly across different genes, genomic regions (e.g., mutational hot spots), and biological species.
Incidence Rate of Genetic Disease: The rate of disease-causing mutations corresponds directly to the incidence of new cases of a genetic condition caused by single new mutations that were absent in the parents.
Mechanistic Causes by Mutation Scale:
Chromosome Mutations: Arise primarily from improper chromosome segregation (nondisjunction) during meiotic divisions. These alterations are typically severe and often lead to spontaneous abortion.
Regional Mutations: Originate from non-allelic homologous recombination between repetitive DNA fragments with high sequence identity at distant sites, or from imperfect repair of double-strand DNA breaks.
Gene Mutations: Result from DNA polymerase replication errors (which occur at a baseline frequency of less than ) and DNA repair errors, especially when spontaneous chemical lesions escape repair mechanisms.
Specific Molecular Types of DNA Mutations
Nucleotide Substitutions
Synonymous (Silent) Mutations: Point nucleotide alterations that change a codon to another codon encoding the exact same amino acid, preserving the primary protein sequence due to code degeneracy.
Missense Mutations: Single nucleotide substitutions that alter a codon to code for a different amino acid, potentially altering protein structure, stability, or activity.

Structural Categorization of Base Substitutions:
Transitions: Substitutions replacing a purine with another purine (), or a pyrimidine with another pyrimidine ().
Transversions: Substitutions replacing a purine with a pyrimidine or vice versa (, , , ).

Nonsense Mutations: Point mutations replacing an amino-acid-coding codon with a translation termination codon (, , or in DNA), causing premature polypeptide chain termination and truncated protein products.

Splicing Mutations: Mutations occurring within splice donor () or splice acceptor () consensus sequences at intron-exon boundaries. These alterations can abolish canonical splicing sites or generate cryptic splice sites, disrupting proper pre-mRNA processing and altering mature mRNA sequences.

Dynamic Mutations (Trinucleotide Repeat Expansions)
Mechanism: Unstable expansions of simple trinucleotide repeat sequences located within coding regions, , , or introns. Expansions typically occur during DNA replication via template slippage, generating hairpin structures in repetitive sequences.

Clinical Examples of Triplet Expansion Diseases
Disease | Repeated Codon (Amino Acid) | Gene Symbol | Normal Range (Repeats) | Disease Threshold (Repeats) |
|---|---|---|---|---|
Huntington disease | (Gln) | HTT | 11–35 | |
Spinocerebellar ataxia type 1 | (Gln) | ATXN1 | 6–35 | |
Machado-Joseph disease | (Gln) | ATXN3 | 12–40 | |
Kennedy disease | (Gln) | AR | 9–36 | |
Fragile X syndrome | (Arg) | FMR1 | 5–54 | |
Fragile X-E syndrome | (Pro) | AFF2 | 6–35 | |
Myotonic dystrophy | (Leu) | DMPK | 5–37 | |
Friedreich's ataxia | (Glu) | FXN | 7–34 |
Frameshift Mutations
Definition: Insertions or deletions involving a number of nucleotides that is not a multiple of three (e.g., , , , ) within protein-coding exons.
Consequence: Shifts the translational reading frame downstream of the mutation site, altering the subsequent amino acid sequence and typically creating a downstream premature stop codon.

Functional Consequences and Heredity of Mutations
Functional Effects on Protein Products
Gain-of-Function Mutations: Lead to overproduction, inappropriate temporal/spatial expression, or novel biochemical properties of the encoded protein product. These mutations are typically inherited in a dominant manner.

Loss-of-Function Mutations: Result in reduced expression or complete loss of functional protein product.
In most biological pathways, the protein produced by a single wild-type allele ( functional output) is sufficient for normal phenotype maintenance.
Haploinsufficiency: Occurs when of the normal protein level is insufficient to maintain normal cellular function, causing a dominant clinical phenotype from a loss-of-function allele.
Dominant Negative Mutations: Occur when a mutant allele produces an abnormal protein product that physically interferes with or inhibits the function of the wild-type protein produced by the normal allele in a heterozygous individual.
Lethal Mutations: Disrupted gene functions critical for viability, leading to embryonic or organismal death.
Heredity: Germline vs. Somatic Mutations
Germline Mutations: Inherited from carrier parents or originating de novo during gametogenesis. They are present throughout all somatic and germline cells of the offspring and can be passed to subsequent generations.
True de novo locus-specific germline mutations occur at very low frequencies.
Somatic Mutations: Arise within non-germline body cells after fertilization and cannot be transmitted to offspring.
Cause localized genomic heterogeneity, especially in tissues with high proliferative capacity (e.g., epithelial linings, hematopoietic lineages).
Typically escape detection by standard sequencing unless present at high clonal fractions, because clinical sequencing relies on DNA isolated from bulk tissue or millions of nucleated blood cells.
Constitute the primary driver mechanism in somatic oncogenesis.
Genetic Polymorphisms and Structural Variants
Definition of Genetic Polymorphism: A genomic variant present at a locus where the minority allele frequency exceeds in a given population, regardless of its structural size, genomic location, or functional effect.
The Four Major Classes of Polymorphisms
1. Single Nucleotide Polymorphisms (SNPs)
Alterations involving a single base pair.
Occur on average at a frequency of 1 every , yielding an estimated to SNPs across the human genome.
The vast majority reside in non-coding regions and produce no observable phenotypic consequences.
Approximately 100,000 SNPs reside in protein-coding exons, categorized into synonymous and non-synonymous variants (which alter the amino acid sequence).
CpG Dinucleotide Hotspots: Cytosine bases adjacent to guanine (CpG sites) undergo endogenous enzymatic methylation to form 5-methylcytosine. Spontaneous deamination converts 5-methylcytosine directly to thymine, causing a higher mutation frequency at CpG sites compared to other dinucleotide contexts.

2. Insertion-Deletion Polymorphisms (Indels)
Structural variants involving sequence insertions or deletions ranging up to .
Simple Indels: Characterized by the presence or absence of a specific short DNA segment, resulting in biallelic systems.
Microsatellites / Short Tandem Repeat Polymorphisms (STRs): Repeated units of 2, 3, or 4 nucleotides (e.g., repeats).
Characterized by high multiallelic variability across individuals due to differences in repeat counts.
DNA Fingerprinting: Utilized in forensic profiling and parentage testing by analyzing allele length variation across 13 standardized polymorphic STR loci.

3. Copy Number Variants (CNVs)
DNA segments larger than up to hundreds of kilobases () present in variable copy numbers relative to a reference genome.
Can encompass entire coding regions of multiple genes, altering absolute gene dosage.
4. Inversion Polymorphisms
Range in size from a few base pairs to several megabases ().
Originate from homologous recombination mediated by inverted sequence homology flanking the inverted region.
Balanced Variants: Involve reorientation of the sequence without net loss or gain of genetic material.

Variant Discovery, Mapping, and Genome-Wide Association Studies (GWAS)
Methodological Pipeline for Variant Analysis
Discovery: Identification of candidate variants via Whole-Genome Sequencing (WGS) or Whole-Exome Sequencing (WES), followed by alignment and comparison against reference human genomes.
Validation: Technical replication assays to eliminate artifactual sequencing errors, followed by genotyping across extended population samples to determine population frequency.
Screening: Simultaneous high-throughput profiling of thousands to millions of variants across large clinical cohorts using high-density target-probe DNA microarrays (SNP arrays).
Mapping Human Disease Variants
Linkage Analysis: Uses family pedigree structures to track co-segregation of genetic markers with monogenic traits. Best suited for identifying rare, highly penetrant Mendelian mutations.
Association Analysis: Population-based case-control comparisons evaluating significant frequency differences of common markers. Effective for identifying polygenic, low-penetrance susceptibility loci.
Genome Sequencing: High-throughput identification of rare variants across individual whole genomes.

Spectrum of Allele Frequency vs. Effect Size (Odds Ratio)
Highly Penetrant Mendelian Mutations: Rare alleles with large effect sizes (e.g., CFTR F508 in Cystic Fibrosis).
Common Variants with Large Effects: High frequency alleles conferring elevated odds ratios (e.g., APOE4 in Alzheimer's disease, CFH in Age-Related Macular Degeneration).
Less Common Variants with Moderate Effects: Intermediate frequency variants (e.g., NOD2 in Crohn's disease, TNFRSF1A in Multiple Sclerosis).
Common Variants with Small Effects Identified by GWAS: Highly frequent alleles conferring minor incremental disease risk (e.g., TCF7L2 in Type 2 Diabetes, LMTK2 in Prostate Cancer).
Genome-Wide Association Study (GWAS) Architecture
Purpose: Evaluates hundreds of thousands to millions of SNP markers simultaneously across human genomes in unbiased case-control cohorts to detect statistical associations with specific clinical phenotypes without requiring prior hypotheses regarding gene function.

Experimental Workflow and Quality Control
Sample Collection & Phenotyping: Assembly of well-characterized case cohorts (exhibiting specific clinical phenotypes, drug responses, or toxicities) matched against unaffected controls.
Genotyping: High-density microarray hybridization covering 100,000+ to millions of target SNPs.
Quality Control (QC):
Sample QC: Evaluates sample call rates, removes related individuals, and corrects for population stratification (e.g., principal component analysis using eigenvectors).
SNP QC: Filters out variants based on genotyping call rates, major deviations from Hardy-Weinberg equilibrium, and extremely low minor allele frequencies.
Statistical Association: Generation of Quantile-Quantile (Q-Q) plots to evaluate systematic inflation and Manhattan plots to visualize chromosome-wide significance levels.
Post-GWAS Validation: Independent cohort replication, meta-analyses, and functional validation assays (e.g., Electrophoretic Mobility Shift Assays [EMSA] or luciferase reporter gene assays).
Linkage Disequilibrium and SNP Imputation
Linkage Disequilibrium (LD): The non-random association of alleles at nearby linked loci on a chromosome, such that specific allele combinations occur together more frequently than expected by chance.
Haplotypes: Specific combinations of closely linked SNPs located on the same physical chromosome that tend to be inherited together as a unit through generations.
Genotype Imputation: Inferring ungenotyped SNP variants across individual genomes by comparing observed genotyped marker patterns against high-density reference haplotype panels (e.g., HapMap, 1000 Genomes Project).

Manhattan Plot Interpretation
Displays association test significance as on the vertical y-axis plotted against physical genomic coordinates arranged by chromosome (1 through 22) on the horizontal x-axis.
Genome-Wide Significance Threshold: Because millions of independent statistical tests are performed, standard significance levels () yield numerous false positives. The rigorous threshold for genome-wide significance is set to ().
Interpretations of Significant SNP Associations:
The SNP is the direct causal functional variant affecting the trait.
The SNP is non-causal but sits in tight Linkage Disequilibrium with the true causal functional locus.
The association represents a statistical false positive result.

Clinical Applications, Ocular Genetics, and Direct-to-Consumer Genomics
Quantifying Clinical Risk
Risk Variants: Alleles associated with an increased susceptibility to developing a disease.
Protective Variants: Alleles associated with a reduced susceptibility to developing a disease.
Relative Risk Scale:
: Indicates elevated risk relative to baseline (e.g., risk increase; risk increase).
: Indicates baseline risk (no statistical association between variant and disease).
: Indicates reduced/protective risk (e.g., risk reduction; risk reduction).

Direct-to-Consumer (DTC) Personal Genomics
Scope (e.g., 23andMe): Personal microarray profiling analyzing genome-wide SNPs directly for consumers.
Reported Categories:
Benign Traits: Cilantro aversion, photic sneeze reflex, caffeine metabolism rate, hair curliness, earlobe architecture, eye color, dimples.
Carrier Status: Monogenic Recessive Conditions (e.g., Cystic Fibrosis, BRCA1, Sickle Cell Anemia, Glycogen Storage Disease, Maple Syrup Urine Disease, Sjögren's syndrome).
Ancestry: Haplotype analysis determining biogeographic ancestry percentages and maternal/paternal lineages.
Pharmacogenomics: Drug metabolism variations.
Clinical and Ethical Concerns: Data privacy risks, lack of direct medical oversight, and potential misinterpretation of risk probabilities by consumers without genetic counseling.
Genetics of Eye Color and Ocular Pathologies
Polygenic Eye Color Determination: Regulated by loci within genes including ASIP, HERC2, IRF4, MC1R, OCA2, TYR, TYRP1, SLC24A4, SLC24A5, and SLC45A2.
Major Ocular Diseases and Associated Susceptibility Loci
Disease / Disorder | Associated Genes / Variants | Typical Age of Onset |
|---|---|---|
Age-Related Macular Degeneration (AMD) | CFH, NOS2A, CF, C2, C3, CFB, HTRA1/LOC, MMP-9, TIMP-3, SLC16A8 | Old age |
Cataract | GEMIN4, CYP51A1, RIC1, TAPT1, TAF1A, WDR87, APE1, MIP, Cx50/GJA3 & 8, CRYAA, CRYBB2, PRX, POLR3B, XRCC1, ZNF350, EPHA2 | Old age |
Glaucoma | CALM2, MPP-7, Optineurin, LOX1, CYP1B1, CAV1/2, MYOC, PITX2, FOXC1, PAX6, LTBP2 | Over 40 years (except congenital forms affecting infants) |
Inherited Optic Neuropathies | Complex I / ND mitochondrial genes, OPA1, RPE65 | Young males |
Marfan Syndrome | FBN1, TGFBR2, MTHFR, MTR, MTRR | Present at birth; diagnosis often later in life |
Myopia | HGF, C-MET, UMODL1, MMP-1/2, PAX6, CBS, MTHFR, IGF-1, UHRF1BP1L, PTPRR, PPFIA2, P4HA2 | Typically progresses until ~20 years |
Polypoidal Choroidal Vasculopathies | C2, C3, CFH, SERPING1, PEDF, ARMS2-HTRA1, FGD6, ABCG1, LOC387715, CETP | Between ages 50 and 65 |
Retinitis Pigmentosa | RPGR, PRPF3, HK1, AGBL5 | Between ages 10 and 30 |
Stargardt's Disease | ABC1, ABCA4, CRB1 | Early childhood to middle age |
Uveal Melanoma | PTEN, BAP1, GNAQ, GNA11, DDEF1, SF3B1, EIF1AX, CDKN2A, p14ARF, HERC2/OCA2 | Ages 50 to 80 |
American Academy of Ophthalmology (AAO) Genetic Testing Guidelines (2013)
Offer genetic testing specifically to patients presenting with clinical evidence of heritable eye disorders where the causative gene(s) have been established in the literature.
Utilize CLIA-approved laboratories that cross-reference findings against peer-reviewed clinical databases defining pathogenic vs. benign variants.
Provide patients with full diagnostic reports and clear explanations of testing implications.
Discourage the use of direct-to-consumer genetic testing kits, directing patients to certified clinical facilities.
Order targeted genetic testing specific to the individual patient's presenting clinical phenotype.
Avoid redundant testing for complex polygenic eye diseases unless results directly alter the established treatment plan.
Avoid testing asymptomatic minor children unless explicit parental consent is obtained and a clear, actionable medical intervention exists.
Clinical Limitations of GWAS and Polygenic Risk Profile Data
Lack of Universal Causality: Identified gene-disease associations may vary across diverse racial or ethnic backgrounds.
Unquantified Environmental Interactions: Complex lifestyle and environmental variables strongly modify polygenic risk scores and are difficult to model quantitatively.
Prognostic vs. Diagnostic Limits: Polygenic SNP profiles yield relative risk estimations rather than definitive diagnostic outcomes; confounding biological variables can shift realized risk.
Calculation Complexity: Polygenic risk scores combining cumulative minor effect sizes across dozens of independent SNPs remain difficult to integrate into single-patient clinical decision-making.