Genetic Processes, DNA Replication, Transcription, Translation, and Mendelian Genetics- Week 2

Fundamentals of Nucleic Acids and Amino Acids

  • Nucleotide Structure: The nucleotide is the fundamental monomeric unit of nucleic acids (DNA and RNA). Each nucleotide is composed of three distinct structural components:

    • A sugar molecule (deoxyribose in DNA, ribose in RNA).

    • A phosphate group.

    • A nitrogenous base.

  • Nitrogenous Bases and Base Pairing:

    • Nitrogenous bases form complementary pairs linked by hydrogen bonding across antiparallel strands.

    • Adenine (A\text{A}) pairs specifically with Thymine (T\text{T}) in DNA.

    • Guanine (G\text{G}) pairs specifically with Cytosine (C\text{C}) in DNA.

    • In RNA, Uracil (U\text{U}) substitutes for Thymine (T\text{T}) and pairs with Adenine (A\text{A}).

  • Strand Geometry and Directionality:

    • Alternating sugar and phosphate molecules form the covalent structural backbone of nucleic acid strands.

    • Strands possess defined chemical directionality based on carbon positions on the sugar ring, running from the 55' end to the 33' end.

    • Complementary DNA strands in a double helix run in opposite directions, referred to as an antiparallel orientation.

  • Amino Acids:

    • Amino acids are the monomeric building blocks of polypeptide chains and proteins.

    • Although numerous amino acids exist in nature, exactly 2020 standard amino acids are encoded within human DNA.

    • Amino acids fall into four distinct chemical classes: basic, nonpolar (hydrophobic), polar (uncharged), and acidic.

    • Detailed chemical listing of the 2020 amino acids encoded in DNA, including chemical names, 3-letter codes, 1-letter codes, molecular formulas, and molecular weights (g/mol\text{g/mol}):

    • Alanine (Ala\text{Ala}, A\text{A}): Molecular weight 89.0989.09 (anhydrous 71.0871.08).

    • Arginine (Arg\text{Arg}, R\text{R}): Basic amino acid; Molecular weight 174.20174.20 (anhydrous 156.19156.19).

    • Asparagine (Asn\text{Asn}, N\text{N}): Polar uncharged; Molecular weight 132.12132.12 (anhydrous 114.10114.10).

    • Aspartic Acid (Asp\text{Asp}, D\text{D}): Acidic amino acid; Molecular weight 133.10133.10 (anhydrous 115.09115.09).

    • Cysteine (Cys\text{Cys}, C\text{C}): Polar uncharged (C3H7NO2S\text{C}_3\text{H}_7\text{NO}_2\text{S}); Molecular weight 121.16121.16 (anhydrous 103.14103.14).

    • Glutamic Acid (Glu\text{Glu}, E\text{E}): Acidic amino acid; Molecular weight 147.13147.13 (anhydrous 129.11129.11).

    • Glutamine (Gln\text{Gln}, Q\text{Q}): Polar uncharged; Molecular weight 146.15146.15 (anhydrous 128.13128.13).

    • Glycine (Gly\text{Gly}, G\text{G}): Nonpolar (C2H5NO2\text{C}_2\text{H}_5\text{NO}_2); Molecular weight 75.0775.07 (anhydrous 57.0557.05).

    • Histidine (His\text{His}, H\text{H}): Basic amino acid; Molecular weight 155.16155.16 (anhydrous 137.14137.14).

    • Isoleucine (Ile\text{Ile}, I\text{I}): Nonpolar hydrophobic (C6H13NO2\text{C}_6\text{H}_{13}\text{NO}_2); Molecular weight 131.18131.18 (anhydrous 113.16113.16).

    • Leucine (Leu\text{Leu}, L\text{L}): Nonpolar hydrophobic (C6H13NO2\text{C}_6\text{H}_{13}\text{NO}_2); Molecular weight 131.17131.17 (anhydrous 113.16113.16).

    • Lysine (Lys\text{Lys}, K\text{K}): Basic amino acid; Molecular weight 146.19146.19 (anhydrous 128.17128.17).

    • Methionine (Met\text{Met}, M\text{M}): Nonpolar hydrophobic (C5H11NO2S\text{C}_5\text{H}_{11}\text{NO}_2\text{S}); Molecular weight 149.21149.21 (anhydrous 131.20131.20).

    • Phenylalanine (Phe\text{Phe}, F\text{F}): Nonpolar hydrophobic; Molecular weight 165.19165.19 (anhydrous 147.18147.18).

    • Proline (Pro\text{Pro}, P\text{P}): Nonpolar (C5H9NO2\text{C}_5\text{H}_9\text{NO}_2); Molecular weight 115.13115.13 (anhydrous 97.1297.12).

    • Serine (Ser\text{Ser}, S\text{S}): Polar uncharged (C3H7NO3\text{C}_3\text{H}_7\text{NO}_3); Molecular weight 105.09105.09 (anhydrous 87.0887.08).

    • Threonine (Thr\text{Thr}, T\text{T}): Polar uncharged (C4H9NO3\text{C}_4\text{H}_9\text{NO}_3); Molecular weight 119.12119.12 (anhydrous 101.10101.10).

    • Tryptophan (Trp\text{Trp}, W\text{W}): Nonpolar hydrophobic; Molecular weight 204.23204.23 (anhydrous 186.21186.21).

    • Tyrosine (Tyr\text{Tyr}, Y\text{Y}): Polar uncharged; Molecular weight 181.19181.19 (anhydrous 163.17163.17).

    • Valine (Val\text{Val}, V\text{V}): Nonpolar hydrophobic (C5H11NO2\text{C}_5\text{H}_{11}\text{NO}_2); Molecular weight 117.15117.15 (anhydrous 99.1399.13).

Central Dogma and Core Genetic Processes

  • Central Dogma Framework: Information transfer in biological systems occurs through three central biochemical processes:

    • Replication: Synthesizing a duplicate copy of DNA from an existing DNA template.

    • Transcription: Synthesizing an RNA strand from a specific DNA template sequence.

    • Translation: Decoding mRNA nucleotide sequences into an amino acid chain (protein) at the ribosome.

Mechanism of DNA Replication

  • Subcellular Localization: DNA replication takes place exclusively inside the cell nucleus in eukaryotes.

  • Directionality Constraints: DNA polymerases synthesize new DNA strands strictly in the 535' \rightarrow 3' direction.

  • Semiconservative Replication Model:

    • Parent DNA double helices unwind by breaking hydrogen bonds between paired nitrogenous bases.

    • Each separated parental strand serves as a template determining the precise complementary base sequence of a new daughter strand.

    • Every newly formed daughter DNA double helix consists of one original parental strand and one newly synthesized strand.

  • Origins of Replication (Sites of Origin):

    • Replication initiates at specific genomic sequences designated as sites of origin or origins of replication.

    • Eukaryotic chromosomes contain hundreds to thousands of replication origins across their length.

    • Replication opens bubbles at origins, which expand laterally in both directions (353' \rightarrow 5' and 535' \rightarrow 3' relative to templates) until adjacent bubbles merge.

    • Transmission electron microscopy (TEM) depicts multiple active replication bubbles along eukaryotic chromosomal DNA (such as in cultured Chinese hamster cells at a scale of 0.25μm0.25\,\mu\text{m}).

  • Energetics of Synthesis:

    • Nucleotides enter the synthesis process as nucleoside triphosphates, carrying high-energy phosphate groups.

    • Hydrolysis of these phosphate groups releases the chemical energy necessary to drive phosphodiester bond formation between nucleotides.

  • Enzymatic Machinery:

    • Helicase: Unwinds and separates the double-stranded parental DNA helix at the replication fork.

    • Single-Strand Binding Proteins (SSBPs): Attach to single-stranded parent DNA to prevent premature re-annealing.

    • Primase: Lays down a short RNA primer to provide a free 3-OH3'\text{-OH} group required for DNA polymerase attachment.

    • DNA Polymerase III: Main elongation enzyme that locks onto primers and adds complementary deoxyribonucleotides sequentially in the 535' \rightarrow 3' direction.

    • DNA Polymerase I: Removes RNA primers, replaces RNA nucleotides with complementary DNA nucleotides, and conducts initial proofreading.

    • DNA Ligase: Covalently joins Okazaki fragments together by sealing nicked sugar-phosphate backbones.

  • Discontinuous Lagging Strand Synthesis:

    • Because DNA strands run antiparallel and polymerase operates strictly in the 535' \rightarrow 3' direction, assembly differs per strand:

    • Leading Strand: Synthesized continuously moving toward the advancing replication fork in the 535' \rightarrow 3' direction.

    • Lagging Strand: Synthesized discontinuously moving away from the replication fork in short segments called Okazaki fragments.

    • Step-by-Step Lagging Strand Process:

    1. Replication site of origin opens, and primase adds an RNA primer.

    2. DNA Polymerase III synthesizes an Okazaki fragment in the 535' \rightarrow 3' direction away from the fork.

    3. As the fork unwinds further, primase lays down a new primer upstream closer to the fork.

    4. DNA Polymerase III detaches and restarts synthesis from the new primer, generating successive Okazaki fragments.

    5. DNA Polymerase I removes RNA primers and fills the gaps with DNA nucleotides.

    6. DNA Ligase seals the phosphodiester bonds to join adjacent Okazaki fragments into a continuous strand.

DNA Proofreading, Repair Mechanisms, and Mutation Rates

  • Genomic Scale and Replication Velocity:

    • Bacterial genome (Escherichia coli): Contains approximately 5×1065 \times 10^6 base pairs; complete replication finishes in under 1hour1\,\text{hour}.

    • Human genome: Consists of 4646 chromosomes containing approximately 6×1096 \times 10^9 base pairs (3×1093 \times 10^9 base pairs per haploid genome). Copying completes within a few hours.

    • Printing the human genome single-letter by single-letter (A, C, T, G) would fill over 1,0001{,}000 books stacked 200feet200\,\text{feet} high.

  • Fidelity and Accuracy Scale:

    • Overall final replication error rate is approximately 11 mistake per 10910^9 (1billion1\,\text{billion}) bases copied (1109\frac{1}{10^9}).

    • This accuracy is proportionally equivalent to finding 11 specific individual out of the population of Africa, or 11 single user out of everyone on Facebook.

    • Given 3×1093 \times 10^9 base pairs per human haploid genome, an average of 33 nucleotide errors occur per genome replication cycle.

  • Proofreading and Repair Pathways:

    • Initial misincorporation rate by DNA polymerase is 11 error per 10,00010{,}000 base pairs (110,000\frac{1}{10{,}000} or 10410^{-4}).

    • Proofreading: DNA polymerases feature immediate proofreading activity that detects mispaired bases, exalts incorrect nucleotides, and replaces them during synthesis.

    • Mismatch Repair: Specialized repair enzymes perform post-replication scanning to excise mispaired bases that escaped initial proofreading.

    • Nucleotide Excision Repair: Environmental mutagens (e.g., chemicals, radiation) cause structural DNA damage. Nucleases cut out damaged stretches, DNA polymerase synthesizes replacement bases using the undamaged strand as a template, and DNA ligase seals the backbone.

  • Mutation Frequencies:

    • Uncorrected mismatches create gene mutations at rates exceeding 1in100,0001\,\text{in}\,100{,}000 (1105\frac{1}{10^5}) per gene.

    • Due to genome scale, every individual human typically inherits approximately 33 to 44 novel point mutations.

Transcription: Synthesis of RNA from DNA

  • Definition and Location: Transcription is the DNA-directed synthesis of RNA, occurring inside the nucleus.

  • Structural Differences: DNA vs. RNA:

    • Strand Configuration: RNA is single-stranded; DNA is double-stranded.

    • Length: RNA spans the length of a single gene; DNA spans thousands of genes.

    • Sugar Unit: RNA contains ribose sugar; DNA contains deoxyribose sugar.

    • Base Composition: RNA uses Uracil (U\text{U}) instead of Thymine (T\text{T}); Uracil pairs with Adenine (A\text{A}).

  • Functional Gene Architecture:

    • Promoter: A specific DNA sequence upstream of the gene that acts as the binding site for RNA polymerase and designates the transcription start point.

    • Coding Region: The segment of DNA transcribed into RNA.

    • Terminator (Termination Site): A DNA sequence signaling the end of transcription.

  • Stages of Transcription:

    1. Initiation: RNA polymerase binds to the promoter sequence, unzips local DNA strands, and initiates RNA synthesis at the start point.

    2. Elongation: RNA polymerase advances downstream along the template DNA strand (353' \rightarrow 5' template direction), synthesizing a complementary single-stranded RNA transcript in the 535' \rightarrow 3' direction. The DNA double helix rezips behind the advancing enzyme.

    3. Termination: Upon reaching the terminator sequence, RNA polymerase releases the completed primary RNA transcript (pre-mRNA) and detaches from DNA.

  • Enzymatic Efficiency: Unlike replication, transcription is largely carried out by a single primary enzyme: RNA Polymerase.

Post-Transcriptional Processing and RNA Splicing

  • Pre-mRNA Processing Modifications:

    • Before nuclear export, primary pre-mRNA transcripts undergo structural modification, receiving a 55' cap and a 33' poly-A tail (AAAA\text{AAAA}), alongside leader and trailer sequences.

  • RNA Splicing Mechanics:

    • Introns (INTRagenic sequences): Non-coding regions within pre-mRNA that are excised and removed.

    • Exons (EXpressed sequences): Functional coding regions that are retained and joined together to create the mature mRNA transcript.

    • Memory rule: INtrons are taken OUT; EXons stay IN.

    • Mature mRNA exits through nuclear pores into the cytoplasm for translation.

Translation and Protein Synthesis

  • Definition and Location: Translation is the assembly of functional proteins from mature mRNA templates, occurring in the cytoplasm.

  • Ribosomal Machinery:

    • Translation is executed by ribosomes, composed of a small ribosomal subunit and a large ribosomal subunit.

  • Codons and Anticodons:

    • mRNA genetic information is formatted into nonoverlapping triplets of nitrogenous bases termed codons.

    • Each codon specifies a single unique amino acid.

    • Transfer RNA (tRNA) molecules carry an anticodon sequence that binds complementarily to its corresponding mRNA codon.

    • tRNA delivers its specific amino acid cargo to the active site of the ribosome.

  • Translation Step-by-Step Sequence:

    1. Mature mRNA binds to the small subunit of the ribosome.

    2. The ribosome identifies the start codon on mRNA.

    3. Specific tRNAs align their anticodons with mRNA codons via complementary base pairing.

    4. The ribosome catalyzes peptide bond formation, linking amino acids into a growing polypeptide chain from the amino (N\text{N}-) end to the acid (C\text{C}-) end.

    5. The ribosome advances along the mRNA in the 535' \rightarrow 3' direction.

    6. Multiple ribosomes can translate a single mRNA simultaneously, forming a polyribosome to rapidly scale protein production.

Protein Structural Hierarchy and Folding

  • Chaperonin-Assisted Folding:

    • Newly synthesized linear polypeptide chains must fold into precise three-dimensional shapes to become biologically functional.

    • Heat shock proteins known as chaperonins assist and direct correct polypeptide folding.

  • Four Organizational Levels of Protein Structure:

    • Primary Structure: Linear sequence of amino acids joined together by covalent peptide bonds.

    • Secondary Structure: Local spatial arrangements formed by hydrogen bonding along the polypeptide backbone, including α\alpha -helices and β\beta -pleated sheets.

    • Tertiary Structure: Three-dimensional folding pattern of a single polypeptide chain resulting from side-chain (R\text{R}-group) interactions.

    • Quaternary Structure: Multi-subunit spatial arrangement of two or more distinct polypeptide chains working as a functional complex.

  • Protein Transport: Folded, mature proteins are transported directly to intracellular or extracellular target sites.

Introduction to Mendelian Genetics and Punnett Squares

  • Punnett Squares: A grid-based mathematical model used to calculate and predict genotypic and phenotypic ratios in Mendelian inheritance.

  • Key Inheritance Terminology:

    • Allele: Alternative nucleotide variant forms of a gene.

    • Homozygous: Possessing two identical alleles for a specific gene.

    • Heterozygous: Possessing two different alleles for a specific gene.

    • Dominant: An allele that expresses its phenotypic trait in both homozygous and heterozygous states, masking recessive alleles.

    • Recessive: An allele whose phenotypic trait is expressed only in homozygous states.

    • Codominant: Inheritance pattern where both alleles in a heterozygote are fully and simultaneously expressed.

    • Sex-Linked: Traits controlled by genes located specifically on sex chromosomes.

  • Sample Single-Gene Cross (Tongue Rolling):

    • Example setup: Tongue rolling capability governed by a single gene.

    • Maternal Genotype: Homozygous dominant for tongue rolling (BB\text{BB}).

    • Paternal Genotype: Homozygous recessive for non-tongue rolling (WW\text{WW}).

    • Punnett square cross results in 100%100\% heterozygous (BW\text{BW}) offspring expressing the dominant phenotypic trait.

Practice Questions and Self-Assessment

  • Question 1: What process creates RNA molecules? Give 2 ways RNA is different than DNA?

    • Answer: The process that creates RNA molecules is Transcription. Three key differences include:

    1. RNA uses Uracil (U\text{U}) instead of Thymine (T\text{T}).

    2. RNA is single-stranded, whereas DNA is double-stranded.

    3. RNA uses ribose sugar, whereas DNA uses deoxyribose sugar.

  • Question 2: These chunks are formed during DNA replication on the lagging strand? By chunking, replication is able to move along the DNA in the correct direction. How do we describe this directionality?

    • Answer: The discontinuous chunks formed on the lagging strand are Okazaki Fragments. The required directionality of DNA synthesis is strictly in the 535' \rightarrow 3' direction.

  • Question 3: Every 3 RNA bases code for 1 amino acid. What do we call these triplets?

    • Answer: These three-base coding sequences are called Codons.

  • Question 4: Which cellular structure carries out translation to build proteins?

    • Answer: Translation is carried out by Ribosomes in the cytoplasm.

  • Question 5: During RNA processing, which sequences are removed, and which stay in?

    • Answer: Introns are removed (spliced out), while Exons stay in (retained in mature mRNA).