Comprehensive Study Guide: Molecular Mechanisms of Transcription, Translation, and Protein Translocation

Structure and Properties of Nucleic Acids (DNA vs. RNA)

  • Strand Complementarity and Chemical Stability:

    • DNA is inherently double-stranded and exhibits high stability.
    • RNA is typically single-stranded and possesses significantly lower biochemical stability than DNA.
    • The difference in stability is primarily driven by the sugar moiety: DNA contains 2′-deoxyribose2'\text{-deoxyribose}, which lacks a hydroxyl group at the 2′2' position, whereas RNA contains ribose, which has a reactive hydroxyl group (-OH\text{-OH}) attached to the 2′2' carbon.
    • This additional 2′2' hydroxyl group in ribose alters the secondary structure and overall shape of the nucleic acid helix.
  • Nitrogenous Base Composition:

    • DNA utilizes the nitrogenous bases Adenine (AA), Thymine (TT), Guanine (GG), and Cytosine (CC).
    • RNA utilizes Uracil (UU) instead of Thymine (TT); Uracil is never present in DNA.
    • Standard (conventional) base-pairing rules in RNA helix formation:
      • Adenine pairs with Uracil (A-UA\text{-}U) via 22 hydrogen bonds (equivalent to the 22 hydrogen bonds in a DNA A-TA\text{-}T pair).
      • Guanine pairs with Cytosine (G-CG\text{-}C) via 33 hydrogen bonds.
  • RNA Helical Geometry and Secondary Structures:

    • Like DNA, complementary RNA strands can pair to form a double helix with base pairs packed in the middle surrounded by a repeating ribose-phosphate backbone.
    • An RNA double helix exhibits distinct geometric features compared to a standard DNA double helix:
      • Features a central hole running down its central longitudinal axis.
      • Contains a significantly narrower and deeper major groove.
    • Single-stranded RNA molecules frequently fold back onto themselves to create complex tertiary architectures:
      • In a stem-loop (RNA hairpin) structure, the stem at the base forms a classical double-stranded RNA helix.
      • The single-stranded loop at the apex contains bases that may remain unpaired or engage in non-standard (non-conventional) base interactions beyond standard A-UA\text{-}U and G-CG\text{-}C pairings.

Eukaryotic Transcription and Functional Roles of RNA Polymerases

  • Definition and Mechanism of Transcription:

    • Transcription is the production of an RNA molecule from a DNA template.
    • RNA polymerase unwinds the double-stranded DNA template and synthesizes an RNA strand complementary to the lower DNA template strand.
    • The resulting RNA transcript sequence matches the upper DNA coding strand, with Uracil (UU) replacing Thymine (TT).
    • The transcription machinery utilizes ribonucleoside triphosphates (rNTPsrNTPs)—specifically ATPATP, GTPGTP, CTPCTP, and UTPUTP—as substrates, adding nucleotides sequentially while moving along the gene.
  • Functional Classes of Transcription Products:

    • Messenger RNA (mRNAmRNA): Serves as the protein-coding template for translation.
    • Ribosomal RNA (rRNArRNA): Forms the structural and catalytic core of ribosomes.
    • Transfer RNA (tRNAtRNA): Delivers specific amino acids to matching mRNA codons during translation.
    • Non-coding RNAs: Includes microRNAs (miRNAsmiRNAs) and other non-coding transcripts involved in gene regulation.
  • Diversity of Eukaryotic RNA Polymerases:

    • Unlike prokaryotes, which utilize a single RNA polymerase enzyme, eukaryotes utilize three distinct RNA polymerases:
      1. RNA Polymerase I\text{RNA Polymerase I}: Primarily transcribes ribosomal RNA (rRNArRNA) genes.
      2. RNA Polymerase II\text{RNA Polymerase II}: Transcribes all protein-coding genes (mRNAmRNA), microRNAs (miRNAsmiRNAs), and select non-coding RNAs.
      3. RNA Polymerase III\text{RNA Polymerase III}: Transcribes transfer RNA (tRNAtRNA) genes, 5S rRNA, and other small non-coding RNAs.

Transcription Initiation and Assembly of the Pre-Initiation Complex

  • General Transcription Factors (GTFs) Requirement:

    • Eukaryotic RNA polymerases cannot initiate transcription independently; they require general transcription factors (GTFs) to locate promoters and assemble the transcription complex.
    • GTFs for RNA Polymerase II\text{RNA Polymerase II} are designated as TFIITFII factors (e.g., TFIIATFIIA, TFIIBTFIIB, TFIIDTFIID, TFIIHTFIIH). Originally lettered TFIIATFIIA through TFIIHTFIIH, several candidate factors were later eliminated as redundant or non-essential.
  • Promoter Architecture and TATA Box Binding:

    • General sequence elements near the +1+1 transcription start site are common to all protein-coding genes.
    • A prominent promoter element is the TATA box sequence located upstream of the +1+1 transcription start site.
    • The TATA-binding protein (TBPTBP), a subunit of the TFIIDTFIID factor complex, directly recognizes and binds to the TATA box sequence.
  • DNA Distortion and Nucleosome Displacement:

    • Upon binding the TATA box, TBPTBP introduces a sharp kink in the DNA double helix.
    • This structural bending alters local chromatin geometry, helping to displace or reposition surrounding nucleosome histone octamers.
    • This conformational change permits the sequential assembly of remaining GTFs and RNA Polymerase II\text{RNA Polymerase II} to establish the transcription initiation complex and begin polymerization.
  • Distinction Between Promoter and Regulatory Elements:

    • General promoter elements (such as the TATA box at +1+1) are required for baseline transcription of all protein-coding genes.
    • Enhancers and Silencers are gene-specific regulatory DNA elements located far away from the start site that modulate specific transcript levels.

Transcript Elongation, Chromatin Remodeling, and Pre-mRNA Processing

  • Transcription Elongation through Chromatin:

    • Eukaryotic nuclear DNA is packaged into chromatin, wrapped tightly around histone octamers to form nucleosome core particles.
    • To navigate through nucleosomes without detaching, RNA Polymerase II\text{RNA Polymerase II} relies on specialized elongation factors that assist its progress and prevent premature termination.
  • Subnuclear Compartmentalization:

    • Nuclear Envelope and Nuclear Pores: The nuclear envelope separates nuclear contents from the cytoplasm; nuclear pore complexes act as gated transport gateways for mature RNA export.
    • Nucleolus: A prominent non-membrane-bound nuclear region dedicated specifically to the synthesis and assembly of ribosomal RNA (rRNArRNA).
  • Co-Transcriptional Processing of Pre-mRNA:

    • As RNA Polymerase II\text{RNA Polymerase II} transcribes pre-mRNA during elongation, three essential chemical modifications occur:
      1. 5′5' Capping: Addition of a modified guanine nucleotide cap structure to the 5′5' end of the nascent pre-mRNA transcript.
      2. 3′3' Polyadenylation: Cleavage of the transcript tail and addition of a poly-A tail (a stretch of adenine nucleotides) at the 3′3' end.
      3. Pre-mRNA Splicing: Removal of non-coding intervening sequences (introns\text{introns}) and ligation of protein-coding sequences (exons\text{exons}).

Pre-mRNA Splicing, Isoforms, and Nuclear Export

  • Gene Organization in Eukaryotes vs. Bacteria:

    • Bacterial genes consist of uninterrupted, continuous protein-coding sequence.
    • Eukaryotic genes are interrupted, composed of protein-coding regions (exons\text{exons}) interspersed with non-coding regions (introns\text{introns}).
  • Alternative Splicing and Isoform Diversity:

    • Pre-mRNA splicing removes introns and joins exons together.
    • Alternative splicing allows individual exons to be differentially included or skipped in the final transcript.
    • Example: In a gene containing 44 exons, standard splicing produces a transcript with all 44 exons (1-2-3-41\text{-}2\text{-}3\text{-}4), while an alternative splice variant can skip exon 33 by splicing exon 22 directly to exon 44 (1-2-41\text{-}2\text{-}4).
    • Alternative splicing generates multiple distinct protein isoforms from a single gene.
  • Structural Domain Architecture of Mature mRNA:

    • Open Reading Frame (ORF): The central sequence region that directly encodes the protein's amino acid sequence.
    • Untranslated Regions (UTRs): Non-coding sequence blocks located flanking the ORF at the 5′5' end (5′5' non-coding UTR) and 3′3' end (3′3' non-coding UTR).
  • Nuclear Export and Cytoplasmic Fate:

    • Fully spliced and processed mRNA recruits 5′5' cap-binding proteins and poly-A binding proteins (PABP\text{PABP}) at its respective terminal ends.
    • These bound protein complexes validate transcript completion and facilitate transport through the nuclear pore complex into the cytoplasm.
    • Upon reaching the cytoplasm, cytosolic initiation factors replace the 5′5' cap-binding proteins to start translation.
    • Eukaryotic mRNA molecules are systematically degraded in the cytoplasm over time to regulate cellular protein expression levels.

The Genetic Code, Codons, and tRNA Synthetases

  • The Genetic Code and Codon Triplets:

    • Translation is the biochemical conversion of mRNA nucleotide sequences into a linear polypeptide sequence.
    • Sequence information is organized into triplet units called codons, read strictly in the 5′→3′5' \rightarrow 3' direction (e.g., 5′-GCU-3′5'\text{-GCU-}3').
  • Mathematical Properties of the Genetic Code:

    • Total codon permutations: 43=644^3 = 64 unique triplet codons.
    • 6161 codons specify amino acids.
    • 33 codons act as stop (termination) signals: UAAUAA, UAGUAG, and UGAUGA.
    • 11 codon acts as the universal start (initiation) signal: AUGAUG, which encodes Methionine (MetMet).
    • Degeneracy / Redundancy: There are 2020 standard amino acids. Methionine (MetMet) and Tryptophan (TrpTrp) are encoded by single unique codons, whereas other amino acids are specified by 22, 44, or up to 66 distinct codons.
  • tRNA Numbers and Synthetase Specificity:

    • Cells contain as few as 3131 distinct transfer RNA (tRNAtRNA) genes, matching all 6161 coding codons via non-standard base pairing (wobble) at the third codon position.
    • Exactly 2020 distinct aminoacyl-tRNA synthetase enzymes exist in the cell, corresponding to the 2020 standard amino acids.
    • Each aminoacyl-tRNA synthetase enzyme selectively recognizes its dedicated amino acid and covalently attaches it to all matching cognate tRNAtRNA species.

tRNA Architecture, Reading Frames, and Ribosomal Mechanism

  • Reading Frame Selection:

    • The precise nucleotide where translation initiates establishes the reading frame.
    • Shifting the start site by 11 or 22 nucleotides alters every subsequent codon triplet, yielding an entirely different protein.
    • Example: Reading sequence starting at position 11 yields codons for Leucine (LeuLeu), starting at position 22 yields UCAUCA for Serine (SerSer), and starting at position 33 yields CAGCAG for Glutamine (GlnGln).
    • Correct frame selection is controlled during initiation by positioning the start codon (AUGAUG) directly inside the P-site of the assembled ribosome.
  • Three-Dimensional Structure of Transfer RNA (tRNA):

    • All tRNAtRNA molecules fold into a characteristic L-shaped tertiary structure.
    • 3′3' Terminus: Located at the tip of the shorter arm; terminates in a conserved 5′-CCA-3′5'\text{-CCA-}3' sequence where the specific amino acid is covalently attached.
    • Anticodon Loop: Located at the opposing tip of the L-shape; displays an exposed 3-nucleotide3\text{-nucleotide} anticodon triplet complementary to the mRNA codon.
    • Example: A tRNA with an anticodon sequence 5′-GAA-3′5'\text{-GAA-}3' base-pairs with the complementary mRNA codon 5′-UUC-3′5'\text{-UUC-}3', delivering Phenylalanine (PhePhe).
    • Aminoacyl-tRNA synthetase enzymes form extensive physical contact across the tRNA body, recognizing the anticodon loop while burying the 3′3' CCACCA tip deep within the active site.
  • Ribosome Subunits and Binding Sites:

    • Ribosomes consist of a small ribosomal subunit and a large ribosomal subunit.
    • Contains three operational tRNA binding pockets:
      1. A-site (Aminoacyl Site): Entry pocket for incoming charged aminoacyl-tRNA molecules.
      2. P-site (Peptidyl Site): Pocket holding the tRNA attached to the growing peptide chain; direct site where initiator Met-tRNAiMet\text{-}tRNA_i binds during initiation.
      3. E-site (Exit Site): Pocket where uncharged, deacylated tRNAs reside before exiting the ribosome.
  • Steps of Translation:

    1. Initiation: Initiator Met-tRNAiMet\text{-}tRNA_i complexed with translation initiation factors binds the 5′5' cap of the mRNA, scans to locate the start codon (AUGAUG), pairs at the P-site, and recruits the large ribosomal subunit.
    2. Elongation: A charged aminoacyl-tRNA enters the vacant A-site, a peptide bond forms to transfer the growing chain, and the ribosome translocates downstream along the mRNA.
    3. Termination: A stop codon (UAAUAA, UAGUAG, or UGAUGA) enters the A-site; a Release Factor protein binds the stop codon, cleaves the ester bond holding the polypeptide to the P-site tRNA, releases the completed protein, and disassembles the ribosome subunits.

Cytosolic vs. ER-Directed Translation and SRP Targeting

  • Polyribosomes and Dual Cellular Populations:

    • A single mRNA molecule is simultaneously translated by multiple ribosomes, forming a polyribosome (polysome).
    • Two distinct polyribosome populations operate in eukaryotic cells, drawing from a common shared pool of ribosomal subunits:
      1. Free Polyribosomes: Unattached in the cytosol. Synthesize water-soluble proteins destined for the cytosol, nucleus, or mitochondria. Translation completes entirely within the cytosol.
      2. Membrane-Bound Polyribosomes: Bound to the cytosolic surface of the Endoplasmic Reticulum (ERER) membrane, forming the Rough ERER. Synthesize proteins targeted for the ERER lumen, Golgi apparatus, endosomes, lysosomes, plasma membrane, or secretion.
  • ER Signal Sequence Properties:

    • Proteins destined for the ERER contain an ER signal sequence consisting of a stretch of small hydrophobic amino acid residues (e.g., Leucine, Valine, Glycine, Isoleucine, Phenylalanine, Tryptophan).
    • Typically located at the N-terminus of the growing polypeptide chain, though internal hydrophobic signal sequences also occur.
  • Signal Recognition Particle (SRP) Targeting Pathway:

    1. Signal Recognition: As the hydrophobic ER signal sequence emerges from the ribosome, it is bound by the cytosolic Signal Recognition Particle (SRPSRP).
    2. Translation Pause: SRPSRP binding induces a transient pausing or slowing of polypeptide elongation.
    3. Receptor Docking: The SRP-ribosome-polypeptideSRP\text{-}ribosome\text{-}polypeptide complex docks at the ER membrane by binding to the SRP-receptorSRP\text{-}receptor.
    4. Translocon Hand-Off: SRPSRP is released back into the cytosol, handing off the ribosome to a protein translocation channel (translocon) in the ER membrane.
    5. Resumption of Synthesis: Translation resumes, threading the growing polypeptide chain through the translocon channel.

Translocation Mechanics for Luminal, Single-Pass, and Multi-Pass Proteins

  • Translocation of Soluble Luminal Proteins:

    • The N-terminal ER signal sequence opens the translocation channel and remains bound to the channel wall while the rest of the protein is threaded through as a loop.
    • Once translocation is complete, Signal Peptidase (located on the luminal side of the ER membrane) cleaves off the N-terminal signal sequence.
    • The cleaved signal peptide is released laterally into the lipid bilayer and rapidly degraded.
    • The soluble protein is released into the ER lumen, and a lumen-derived plug protein seals the channel.
  • Translocation of Single-Pass Transmembrane Proteins:

    • N-Terminal Signal + Stop-Transfer Sequence:
      • Translocation is initiated by an N-terminal ER signal sequence.
      • Translocation continues until a downstream hydrophobic stretch of amino acids—a Stop-Transfer sequence—enters the translocon.
      • The stop-transfer sequence halts translocation and is released laterally into the lipid bilayer to form a single membrane-spanning α-helix\alpha\text{-helix}.
      • Signal peptidase cleaves the N-terminal signal sequence, resulting in a transmembrane protein anchored with its N-terminus in the ER lumen and C-terminus in the cytoplasm.
    • Internal Start-Transfer Sequence:
      • Lacks an N-terminal signal sequence; utilizes an internal hydrophobic Start-Transfer sequence.
      • The internal sequence initiates translocation without being cleaved by signal peptidase, serving as a permanent single-pass membrane anchor.
  • Translocation of Multi-Pass Transmembrane Proteins:

    • Double-Pass Proteins:
      • Utilize an internal hydrophobic Start-Transfer sequence paired with a downstream hydrophobic Stop-Transfer sequence.
      • Neither sequence is cleaved by signal peptidase; both are released laterally into the bilayer as membrane-spanning helices.
    • Complex Multi-Pass Proteins:
      • Contain alternating pairs of internal hydrophobic Start-Transfer and Stop-Transfer sequences.
      • One sequence reinitiates translocation further down the polypeptide chain, and the next sequence stops translocation and triggers lateral release into the bilayer.
      • Operates like a "sewing machine", stitching the polypeptide back and forth across the lipid bilayer during translation.
      • Membrane-spanning segments are enriched in hydrophobic amino acids such as Leucine, Valine, Glycine, Isoleucine, Phenylalanine, and Tryptophan.