Gene Transcription and RNA Modification Comprehensive Study Guide

I. Overview of Gene Transcription and the Central Dogma

  • The Central Dogma of Genetics:

    • First proposed by Francis Crick in 1958.

    • States that genetic information flows from DNA replication to DNA transmission, from DNA to RNA via transcription, and from RNA to protein via translation.

    • Gene Definition: A segment of DNA that is used to make a functional product.

    • Functional Products: The functional product of a gene is not always a protein. While protein-coding genes produce messenger RNA (mRNA) that undergoes translation, many genes encode non-coding functional RNA molecules (such as rRNA and tRNA) that function directly without being translated into polypeptides.   

      Central Dogma
  • Types of RNA Molecules in Cells:

    • Messenger RNA (mRNA): A temporary RNA copy of a gene that contains information to make a polypeptide.

    • Ribosomal RNA (rRNA): Component of ribosomes involved in protein synthesis; examples include 16S16\text{S} rRNA in prokaryotes, as well as 18S18\text{S}, 5.8S5.8\text{S}, 28S28\text{S}, and 5S5\text{S} rRNAs in eukaryotes.

    • Transfer RNA (tRNA): Adaptor molecules that interact with mRNA codons during translation to deliver specific amino acids; contains an anticodon loop and an amino acid attachment site at its 3′3' end.

    • Other Non-Coding RNAs: Includes small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), and various regulatory microRNAs.   

      RNA Types
  • Structure of a Transcriptional Unit:

    • A transcriptional unit is a sequence of DNA that codes for a single RNA molecule, including all necessary regulatory sequences required for its transcription.

    • DNA Features of a Gene:

    • Regulatory Sequences: Site for the binding of regulatory proteins to influence the rate of transcription.

    • Promoter: Site for RNA polymerase binding; signals the beginning of transcription.

    • Terminator: Signals the end of transcription.

    • mRNA Structure (Protein-Coding Genes):

    • Ribosome-Binding Site: Site for ribosome binding on the mRNA to initiate translation in prokaryotes.

    • Start Codon: Specifies the first amino acid in a polypeptide (typically 5’-AUG-3’\text{5'-AUG-3'}).

    • Codons: A sequence of adjacent nucleotides that determine the amino acid sequence of a polypeptide.

    • Stop Codon: Specifies the signal to terminate polypeptide synthesis.   

      Transcriptional Unit
  • The Three Stages of Transcription:

    • 1. Initiation: Transcription factors recognize and bind the promoter sequence, recruiting RNA polymerase to the DNA.

    • 2. Elongation: RNA polymerase unwinds dsDNA to form an open complex and synthesizes a complementary strand of RNA in the 5’→3’\text{5'} \rightarrow \text{3'} direction.

    • 3. Termination: A terminator sequence triggers the dissociation of RNA polymerase from the DNA strand and releases the completed RNA transcript.   

      Three Stages of Transcription

II. Transcription in Bacteria

  • Bacterial Promoter Architecture:

    • Located upstream of the site where transcription begins.

    • Transcriptional Start Site (+1+1): The exact first nucleotide used as a template for transcription.

    • Nucleotide Numbering:

    • Bases preceding the start site are assigned negative numbers (upstream regions, e.g., −10-10, −35-35). There is no base numbered 00.

    • Bases following the start site are assigned positive numbers (downstream regions).

    • Critical Sequence Elements:

    • −35-35 Element: Consensus sequence is 5’-TTGACA-3’\text{5'-TTGACA-3'}.

    • −10-10 Element (Pribnow Box): Consensus sequence is 5’-TATAAT-3’\text{5'-TATAAT-3'}.

    • Spacing: Proper spacing (typically 16–18 bp16\text{--}18\,\text{bp}) between the −35-35 and −10-10 elements is critical for promoter function.

    • Promoter Strand: Located on the non-template (coding) DNA strand.   

      Bacterial Promoter
  • Consensus Sequences and Promoter Strength:

    • A consensus sequence represents the most common nucleotide bases found at each position among a family of related promoter sequences.

    • Affinity and Transcription Rate:

    • The closer a specific promoter's sequence matches the consensus sequence, the higher the binding affinity of the sigma (σ\sigma) factor for that promoter.

    • High affinity leads to an increased rate of transcription initiation (strong promoter).

    • Promoters with significant deviations from the consensus sequence have lower affinity for σ\sigma factor, resulting in lower transcription rates (weak promoter).

  • Bacterial RNA Polymerase Structure:

    • Core Enzyme: Composed of 5 subunits: α2ββ′ω\alpha_2\beta\beta'\omega (two alpha, one beta, one beta-prime, one omega subunit).

    • α\alpha subunits (36.5 kDa36.5\,\text{kDa} each): Assembly of core enzyme and binding to DNA/regulatory proteins.

    • β\beta subunit (151 kDa151\,\text{kDa}) and β′\beta' subunit (155 kDa155\,\text{kDa}): Form the catalytic center for RNA synthesis.

    • ω\omega subunit (4 kDa4\,\text{kDa}): Enzyme assembly and stability.

    • Holoenzyme: Core enzyme plus the sigma factor subunit (α2ββ′ω+σ\alpha_2\beta\beta'\omega + \sigma).

    • Standard major sigma factor in E. coli is σ70\sigma^{70}.

    • Bacteria possess multiple σ\sigma factors (e.g., E. coli has 7) that recognize distinct promoter sequences under specific cellular conditions.   

      Bacterial RNA Polymerase
  • Bacterial Transcription Initiation Process:

    • 1. Closed Complex Formation: RNA polymerase holoenzyme slides along the DNA until the σ\sigma factor subunit recognizes and binds to the −35-35 and −10-10 promoter elements via hydrogen bonding in the major groove of DNA, forming a closed complex.

    • 2. Open Complex Formation: Double-stranded DNA separates at the A-T rich −10-10 element (Pribnow box) because A-T base pairs have only two hydrogen bonds, making strand separation easier than in G-C rich regions. Unwinding creates an open complex (transcription bubble).

    • 3. Abortive Initiation & Release: RNA polymerase synthesizes a short strand of RNA (≈9–10 nucleotides\approx 9\text{--}10\,\text{nucleotides}). The σ\sigma factor is released, marking the transition from initiation to elongation and allowing the core enzyme to proceed downstream.   

      Bacterial Transcription Initiation
  • Bacterial Transcription Elongation:

    • Mechanism: Core RNA polymerase unwinds dsDNA continuously ahead of the catalytic site and rewinds it behind.

    • Rate: RNA core polymerase synthesizes RNA at a rate of approximately 43 nucleotides/second43\,\text{nucleotides/second}.

    • Chemical Reaction: Catalyzes the formation of phosphodiester bonds between incoming ribonucleoside triphosphates (NTPs).

    • Base Pairing Rules: Complementary base pairing via hydrogen bonds ensures accuracy. Uracil (U) base pairs with Adenine (A) in RNA, replacing Thymine (T).

    • Directionality:

    • RNA polymerase builds RNA in the 5’→3’\text{5'} \rightarrow \text{3'} direction.

    • RNA polymerase reads the template DNA strand in the 3’→5’\text{3'} \rightarrow \text{5'} direction.   

      Bacterial Transcription Elongation
  • Template vs. Coding Strand Orientations:

    • Template Strand (Antisense strand): The DNA strand used as a physical template by RNA polymerase to synthesize RNA.

    • Coding Strand (Sense strand): The complementary DNA strand whose base sequence corresponds directly to the transcribed RNA sequence (with T replacing U).

    • Strand Selection: Genes along a single chromosome can be encoded on either strand of DNA. The orientation of the promoter determines which strand serves as the template strand for a given gene.   

      Genes Encoded on Either Strand
  • Bacterial Transcription Termination:

    • 1. Rho (ρ\rho)-Dependent Termination:

    • Requires a protein factor called Rho (ρ\rho), which acts as an ATP-dependent RNA-DNA helicase.

    • Mechanism:

      • RNA polymerase transcribes the rut site (rho utilization site) on the RNA.

      • ρ\rho protein binds to the transcribed rut site in the RNA and moves towards the 3′3' end.

      • RNA polymerase transcribes a sequence downstream that forms a G-C rich stem-loop structure in the nascent RNA.

      • The stem-loop causes RNA polymerase to pause transcript elongation.

      • During the pause, the ρ\rho protein catches up to the transcription bubble, enters the open complex, and uses its helicase activity to break the hydrogen bonds of the RNA-DNA hybrid, releasing the transcript.     

        Rho Dependent Termination
    • 2. Rho (ρ\rho)-Independent (Intrinsic) Termination:

    • Does not require ρ\rho protein.

    • Mechanism:

      • Requires two sequence elements in the transcript:

      1. A GC-rich inverted repeat sequence that forms a stem-loop structure immediately after transcription.

      2. A run of uracils (U-rich sequence) located at the 3′3' end of the transcript immediately following the stem-loop.

      • The stem-loop causes RNA polymerase to pause; this pausing is stabilized by binding of the NusA protein.

      • While paused, the U-rich RNA transcript is paired with the A-rich template DNA strand. Because A-U base pairs are weakly hydrogen-bonded, the RNA-DNA hybrid cannot maintain stability, causing the transcript to dissociate and RNA polymerase to release.     

        Rho Independent Termination

III. Transcription in Eukaryotes

  • Eukaryotic RNA Polymerases:

    • Eukaryotes utilize three distinct nuclear RNA polymerases, each responsible for transcribing different classes of genes:

    • RNA Polymerase I: Transcribes all ribosomal RNA (rRNA) genes, except for 5S5\text{S} rRNA.

    • RNA Polymerase II: Transcribes all protein-coding genes (pre-mRNAs) and many non-coding RNAs (e.g., snRNAs, microRNAs).

    • RNA Polymerase III: Transcribes all transfer RNA (tRNA) genes, the 5S5\text{S} rRNA gene, and several small non-coding RNAs.   

      Eukaryotic RNA Polymerases
  • Eukaryotic Core Promoter and Regulatory Elements:

    • Core Promoter:

    • A short sequence of DNA required for basal transcription.

    • Contains the TATA Box (consensus sequence 5’-TATAAA-3’\text{5'-TATAAA-3'}) located at approximately −25-25 relative to the start site.

    • Contains the Transcriptional Start Site (+1+1), typically an adenine base surrounded by pyrimidines (Py2CAPy5\text{Py}_2\text{CA}\text{Py}_5).

    • The core promoter alone produces low-level (basal) transcription.

    • Regulatory Elements:

    • DNA sequences located further upstream (e.g., −50-50 to −100-100) that regulate transcription efficiency.

    • Enhancers: Sequences that bind activator proteins to stimulate transcription.

    • Silencers: Sequences that bind repressor proteins to inhibit transcription.

    • Common upstream motifs include GC boxes and CAAT boxes.   

      Core Promoter
  • Eukaryotic Preinitiation Complex Assembly:

    • Requires RNA Polymerase II and five General Transcription Factors (GTFs):

    1. TFIID: A multi-protein complex containing TATA-Binding Protein (TBP); recognizes and binds directly to the TATA box. TFIIA assists TFIID binding.

    2. TFIIB: Binds to TFIID and promotes the recruitment of RNA Polymerase II and TFIIF.

    3. TFIIF: Binds to RNA Polymerase II and helps guide it to the core promoter.

    4. TFIIE & TFIIH: Bind to form the closed preinitiation complex. TFIIE regulates TFIIH activity.

    5. Mediator: A large protein complex that interacts with GTFs and regulatory transcription factors to control the transition to elongation.

    • Functions of TFIIH:

    • Helicase Activity: Unwinds DNA at the core promoter to form the open complex.

    • Kinase Activity: Phosphorylates the C-Terminal Domain (CTD) of RNA Polymerase II.

    • Transition to Elongation: Phosphorylation of the CTD by TFIIH and Mediator causes a conformational change that releases TFIIB, TFIIE, and TFIIH, freeing RNA Polymerase II to proceed with transcription elongation.   

      Eukaryotic Preinitiation Complex
  • Eukaryotic Elongation Through Chromatin:

    • Eukaryotic transcription elongation must navigate through nucleosomes.

    • Multiple elongation factors assist RNA Polymerase II in displacing and reassembling histone octamers during synthesis (e.g., Spt4/5, Spt6, Elf1, FACT, Paf1C, TFIIS).   

      Elongation Through Chromatin
  • Eukaryotic Transcription Termination Models:

    • RNA Polymerase II transcribes past the polyadenylation signal sequence (5’-AAUAAA-3’\text{5'-AAUAAA-3'}).

    • An endonuclease cleaves the RNA transcript downstream of the polyadenylation signal site while RNA Polymerase II continues transcribing.

    • Two models explain subsequent termination:

    • Allosteric Model: Transcription past the polyadenylation signal sequence causes destabilization of RNA Polymerase II due to the release of elongation factors or binding of termination factors, leading to dissociation.

    • Torpedo Model: After cleavage of the primary transcript, a 5’→3’\text{5'} \rightarrow \text{3'} exonuclease binds to the exposed 5′5' end of the remaining RNA still being transcribed by RNA Polymerase II. The exonuclease degrades the RNA rapidly in the 5’→3’\text{5'} \rightarrow \text{3'} direction, catches up to RNA Polymerase II, and physically causes it to dissociate from the DNA.   

      Eukaryotic Termination Models

IV. RNA Modification and Processing

  • Overview of RNA Modification Types:

    • Processing (Cleavage): Cleavage of a large precursor RNA transcript into smaller functional RNA fragments; occurs in prokaryotic and eukaryotic rRNAs and tRNAs.

    • Splicing: Cleavage and ligation reactions that remove internal non-coding regions (introns) and join coding regions (exons); common in eukaryotic pre-mRNAs.

    • 5′5' Capping: Addition of a 7-methylguanosine cap to the 5′5' end of pre-mRNA; unique to eukaryotic mRNAs.

    • 3′3' Polyadenylation: Addition of a string of adenine residues to the 3′3' end of mRNA; universal in eukaryotic mRNAs.

    • RNA Editing: Alteration of the nucleotide base sequence of an RNA molecule post-transcriptionally.

    • Base Modification: Covalent alteration of specific base structures within an RNA strand; extremely common in tRNAs.

  • Cleavage of Non-Protein Coding RNAs (rRNA and tRNA):

    • Eukaryotic rRNA Processing: Transcribed by RNA Polymerase I in the nucleolus as a single large precursor called the 45S45\text{S} rRNA primary transcript. Cleavage reactions performed by small nucleolar ribonucleoproteins (snoRNPs) remove spacer sequences, yielding mature 18S18\text{S}, 5.8S5.8\text{S}, and 28S28\text{S} rRNA molecules.   

      rRNA Processing
    • tRNA Processing & Ribozymes:

    • Pre-tRNA molecules are transcribed as longer precursor transcripts containing extra 5′5' and 3′3' sequences.

    • RNaseP: An endonuclease that cuts pre-tRNA to generate the precise 5′5' end of the mature tRNA. RNaseP is a ribozyme—a catalytic RNA molecule where the RNA component carries out catalytic activity.

    • RNaseD: An exonuclease that removes nucleotides from the 3′3' end of the pre-tRNA to prepare it for amino acid linkage.   

      tRNA Processing
  • RNA Splicing Mechanisms:

    • Non-Colinearity in Eukaryotes: Early work in bacteria showed colinearity. In eukaryotes, protein-coding genes contain non-coding intervening sequences (introns) interspersed between coding regions (exons).   

      Colinearity vs Non-colinearity
    • 1. Group I Introns: Self-splicing ribozymes; requires an external guanosine (G) nucleoside/nucleotide to initiate the first cleavage; found in nuclear rRNA, organelles, and a few bacterial genes.

    • 2. Group II Introns: Self-splicing ribozymes; the 2′-OH2'\text{-OH} group of an internal adenine base attacks the 5′5' splice site to form a lariat (loop) intermediate; found in organellar genes.   

      Group I and II Introns
    • 3. Spliceosome-Mediated (Group III / Eukaryotic pre-mRNA) Splicing: Occurs in the nucleus of complex eukaryotes for protein-coding pre-mRNAs. Requires the spliceosome, composed of 5 small nuclear ribonucleoproteins (snRNPs): U1, U2, U4, U5, and U6.   

      Spliceosome Splicing
  • Spliceosome Splicing Steps:

    • Consensus Sequences:

    • 5′5' Splice Site: Invariant GU\text{GU} sequence at the 5′5' boundary of the intron.

    • Branch Site: Internal sequence containing an essential adenine base (A).

    • 3′3' Splice Site: Invariant AG\text{AG} sequence at the 3′3' boundary of the intron.   

      Intron Consensus Sequences
    • Step-by-Step Reaction Mechanism:

    1. U1 snRNP binds to the 5′5' splice site; U2 snRNP binds to the branch site.

    2. A trimeric snRNP complex composed of U4/U6 and U5 joins the complex, causing the intron to loop out and bringing the exons into proximity.

    3. The 5′5' splice site is cleaved. The 5′5' end of the intron is covalently linked via a 2′−5′2'-5' phosphodiester bond to the 2′-OH2'\text{-OH} of the adenine residue at the branch site, forming a lariat structure.

    4. U1 and U4 snRNPs are released from the spliceosome complex.

    5. The 3′3' splice site is cleaved. Exon 1 is covalently linked to Exon 2 via a phosphodiester bond.

    6. The lariat intron along with U2, U5, and U6 snRNPs is released and degraded. The U6 snRNA functions as the actual catalytic center (ribozyme).   

      Splicing Steps
  • Alternative Splicing:

    • A single pre-mRNA transcript can be spliced in multiple distinct pattern combinations to produce different mature mRNAs, encoding distinct polypeptide isoforms.

    • Approximately 70%70\% of all human protein-coding pre-mRNAs undergo alternative splicing.

    • Constitutive Exons: Always retained in the final mature mRNA across all cell types.

    • Alternative Exons: Regulated elements retained or spliced out depending on cell type or developmental stage.

    • Regulation: Regulated by specific protein factors called splicing factors (e.g., α\alpha-tropomyosin pre-mRNA).   

      Alternative Splicing
  • 5′5' Capping Mechanism:

    • Occurs co-transcriptionally while the pre-mRNA transcript is being synthesized by RNA Polymerase II.

    • Enzymatic Steps:

    1. RNA 5′5'\text{-triphosphatase}: Removes one phosphate group from the 5′5' end triphosphate of the nascent pre-mRNA, leaving a diphosphate end.

    2. Guanylyltransferase: Hydrolyzes GTP to attach GMP to the 5′5' diphosphate, releasing pyrophosphate (PPi\text{PP}_i) and generating a 5’-to-5’\text{5'-to-5'} triphosphate linkage.

    3. Methyltransferase: Transfers a methyl group (−CH3-\text{CH}_3) from S-adenosylmethionine to the nitrogen atom at position 7 of the inverted guanine base, producing a 7-methylguanosine (m7G\text{m}^7\text{G}) cap.

    • Biological Functions: Recognized by cap-binding proteins to facilitate nuclear export, essential for translation initiation, and enhances efficient splicing of introns near the 5′5' end.   

      5' Capping
  • 3′3' Polyadenylation Mechanism:

    • PolyA tails are added post-transcriptionally by PolyA-polymerase.

    • Enzymatic Steps:

    1. An endonuclease recognizes the consensus polyadenylation signal sequence (5’-AAUAAA-3’\text{5'-AAUAAA-3'}) and cleaves the transcript approximately 20 nucleotides downstream.

    2. PolyA-polymerase attaches a long chain of adenine nucleotides (up to ≈250\approx 250 bases) to the new 3′3' end.

    • Biological Functions: Promotes cytosolic stability, facilitates nuclear export, and enhances translation efficiency.   

      3' Polyadenylation
  • RNA Editing:

    • Post-transcriptional alteration of nucleotide base sequence via insertion, deletion, or chemical base substitution.

    • Deamination Examples:

    • Cytidine Deaminase: Converts Cytosine to Uracil.

    • Adenosine Deaminase: Converts Adenine to Hypoxanthine/Inosine.

    • Rare in mammals (identified in ≈25\approx 25 protein-coding mRNAs).   

      RNA Editing
  • Base Modification in tRNAs:

    • tRNAs undergo extensive post-transcriptional covalent modifications carried out by dedicated enzymes.

    • Ensures precise secondary/tertiary folding, stabilizes tRNA structure, and optimizes translational accuracy.   

      Base Modification in tRNA

V. Comparison of Transcription and RNA Modification in Bacteria and Eukaryotes

  • Promoter Architecture:

    • Bacteria: Consists of −35-35 (TTGACA\text{TTGACA}) and −10-10 (TATAAT\text{TATAAT}) consensus elements.

    • Eukaryotes: Core promoter consists of a TATA box (−25-25) and transcriptional start site (+1+1), alongside upstream regulatory elements (enhancers/silencers).

  • RNA Polymerase Enzymes:

    • Bacteria: Single RNA polymerase core enzyme (α2ββ′ω\alpha_2\beta\beta'\omega) carrying out all RNA synthesis.

    • Eukaryotes: Three distinct nuclear polymerases; RNA Polymerase II transcribes protein-coding genes.

  • Initiation Factors:

    • Bacteria: Requires a single σ\sigma factor to recognize and bind the promoter.

    • Eukaryotes: Requires five General Transcription Factors (TFIID, TFIIA, TFIIB, TFIIF, TFIIE, TFIIH) and Mediator to assemble the preinitiation complex.

  • Elongation Switch:

    • Bacteria: Triggered by the release of the σ\sigma factor subunit.

    • Eukaryotes: Controlled by TFIIH/Mediator-mediated phosphorylation of the CTD tail of RNA Polymerase II.

  • Termination Mechanism:

    • Bacteria: Mediated by ρ\rho\text{-dependent} or ρ\rho\text{-independent} mechanisms.

    • Eukaryotes: Occurs downstream of polyadenylation signal via the Allosteric Model or Torpedo Model (5’→3’\text{5'} \rightarrow \text{3'} exonuclease action).

  • Splicing:

    • Bacteria: Very rare; restricted to self-splicing introns.

    • Eukaryotes: Highly common in pre-mRNAs via nuclear spliceosome assembly (snRNPs U1, U2, U4, U5, U6).

  • 5′5' Capping:

    • Bacteria: Does not occur.

    • Eukaryotes: Enzymatic addition of a 7-methylguanosine (m7G\text{m}^7\text{G}) cap via a 5’-to-5’\text{5'-to-5'} triphosphate linkage.

  • 3′3' Tailing:

    • Bacteria: PolyA tail added to 3′3' ends signals rapid exonucleolytic degradation.

    • Eukaryotes: PolyA tail added by PolyA-polymerase promotes cytosolic stability and translation efficiency.

  • RNA Editing:

    • Bacteria: Does not occur.

    • Eukaryotes: Occurs occasionally via specific enzymatic base conversions.


Comparison Table