Comprehensive Study Notes: Transcription Mechanics, RNA Polymerase II Structure, Spliceosome Chemistry, and Alternative Splicing

Administrative Policies

  • Office Visit Policy:
    • Students do not need to make an appointment for office visits.
    • An open-door policy is maintained: students may drop by at any time.
    • Routine work will generally be paused to accommodate drop-in student visits.
    • Appointments can still be scheduled if desired, but are not required.

Mechanics and Directionality of RNA Polymerase II Transcription

  • Transcription Initiation Overview:

    • Mediator complexes and general transcription factors position RNA Polymerase II at the +1+1 transcription start site.
    • RNA Polymerase II transcribes protein-coding genes to yield messenger RNA (mRNA), as well as certain non-coding RNAs.
    • RNA Polymerase I and RNA Polymerase III function via similar mechanisms but transcribe non-coding RNAs, such as ribosomal RNA (rRNA) and transfer RNA (tRNA), which function in translation.
  • DNA Strand Separation:

    • DNA is double-stranded at the gene region.
    • The general transcription factor TFIIH\text{TFIIH} contains a helicase subunit that separates the template and non-template strands.
    • Strand separation is mandatory because RNA Polymerase II can only utilize a single strand as a template.
  • Template vs. Non-Template Strand Definitions:

    • Template Strand:
      • Also designated as the noncoding strand or antisense strand.
      • Read by RNA Polymerase II in the 353' \rightarrow 5' direction.
      • Contains a nucleotide sequence that is complementary to the synthesized mRNA product.
      • In laboratory primer design, this strand is referred to as the antisense strand.
    • Non-Template Strand:
      • Also designated as the coding strand or sense strand.
      • Displaced during transcription by RNA Polymerase II.
      • Shares an identical nucleotide sequence with the synthesized mRNA product (with thymine replaced by uracil).
  • Directionality Rules of Transcription:

    • Movement on the template strand occurs strictly in the 353' \rightarrow 5' direction.
    • Synthesis of the RNA transcript product occurs strictly in the 535' \rightarrow 3' direction.
    • Determining orientation on chromosomal DNA:
      • If RNA Polymerase II moves from left to right, it utilizes the bottom strand as the template (provided the bottom strand runs 353' \rightarrow 5' left-to-right).
      • If RNA Polymerase II moves from right to left, it utilizes the top strand as the template (provided the top strand runs 353' \rightarrow 5' right-to-left).
      • Multiple genes along a single chromosome can be transcribed simultaneously in opposing physical directions depending on which strand serves as their template.
  • Structure of RNA Polymerase II:

    • Exhibits a characteristic horseshoe-shaped tertiary structure that clamps onto the DNA.
    • Contains a critical catalytic magnesium ion (Mg2+\text{Mg}^{2+}) within its active site.
    • Features an internal exit channel through which the newly synthesized nascent RNA transcript emerges.
  • Co-Transcriptional Processing and Micrograph Features:

    • Electron micrographs of active transcription show a central DNA strand with multiple RNA Polymerase II complexes engaged simultaneously.
    • As soon as one RNA Polymerase II clears the +1+1 site, another initiation complex assembles to produce multiple mRNA copies per gene.
    • The relative length of the emerging RNA transcript squiggles indicates the direction of polymerase movement:
      • Shorter transcripts indicate proximity to the +1+1 start site.
      • Longer transcripts indicate progression further downstream along the gene.
    • Intergenic DNA between transcribed genes does not encode proteins or transcripts.
    • Dark circular structures at the 55' ends of nascent transcripts in electron micrographs represent early co-transcriptional processing machinery assembling before transcription is completed.

Chemical Structure of RNA and Mechanism of Chain Elongation

  • Chemical Differences Between RNA and DNA:

    • RNA contains ribose sugar, which possesses hydroxyl (OH-\text{OH}) groups on both carbon number 2 (C2\text{C2}') and carbon number 3 (C3\text{C3}').
    • DNA contains deoxyribose sugar, which possesses a hydrogen atom (H-\text{H}) on carbon number 2 (C2\text{C2}') and a hydroxyl group (OH-\text{OH}) on carbon number 3 (C3\text{C3}').
    • The free 2OH2'-\text{OH} group on ribose is chemically reactive and unbonded, serving as an essential chemical nucleophile in subsequent splicing reactions.
  • Terminal Chemistry of Transcripts:

    • 55' End:
      • Exposes carbon number 5 (C5\text{C5}') of the ribose sugar.
      • An unprocessed, legitimate 55' end of a primary transcript terminates in 33 phosphate groups (alphaβ/gamma\frac{\text{alpha}}{\beta/\text{gamma}}).
    • 33' End:
      • Exposes an unbonded hydroxyl group (OH-\text{OH}) on carbon number 3 (C3\text{C3}') of the ribose sugar.
  • Phosphodiester Bond Formation:

    • Incoming precursor subunits are ribonucleoside triphosphates.
    • RNA Polymerase II catalyzes cleavage between the alpha\text{alpha} (Pa\text{P}_\text{a}) and beta\text{beta} (Pb\text{P}_\text{b}) phosphate groups.
    • The beta\text{beta} and gamma\text{gamma} phosphates are released together as pyrophosphate (PPi\text{PP}_\text{i}) and diffuse away.
    • The alpha\text{alpha}-phosphate forms a covalent phosphodiester bond with the exposed 3OH3'-\text{OH} group of the preceding nucleotide.
    • This establishes standard 535' \rightarrow 3' phosphodiester linkages along the RNA backbone.
    • The reaction does not require external energy input (e.g., ATP hydrolysis) because cleavage of the high-energy phosphoanhydride bond and release of pyrophosphate makes phosphodiester bond formation thermodynamically favorable.

Core Promoter Architecture and General Transcription Factors

  • Core Promoter Motifs:

    • +1+1 Site (Initiator / Inr\text{Inr}): The exact nucleotide location where transcription commences.
    • TATA\text{TATA} Box: Located approximately 30 base pairs30\text{ base pairs} upstream (30-30) from the +1+1 start site. Consensus sequence binds specialized transcription factors.
    • BRE\text{BRE} (TFIIB\text{TFIIB} Response Element): Nucleotide sequence located approximately 5 base pairs5\text{ base pairs} upstream of the TATA\text{TATA} box.
    • DPE\text{DPE} (Downstream Promoter Element): Nucleotide sequence located approximately 30 base pairs30\text{ base pairs} downstream (+30+30) from the +1+1 start site.
  • Assembly of General Transcription Factors ($ ext{GTFs}$):

    • TFIID\text{TFIID}: A multiprotein complex that spans the entire core promoter region from upstream of the initiator to the downstream promoter element. It contains amino acid sequences that recognize and bind the Inr\text{Inr} and DPE\text{DPE}.
    • TBP\text{TBP} (TATA\text{TATA} Box Binding Protein):
      • A subunit component of the TFIID\text{TFIID} complex that binds specifically to the TATA\text{TATA} box.
      • Upon binding, TBP\text{TBP} physically bends the double-stranded DNA helix (analogous to bending an arm at the elbow), bringing distant upstream and downstream protein complexes into close proximity.
    • TFIIB\text{TFIIB}: Binds directly to the BRE\text{BRE} site upstream of the TATA\text{TATA} box.
    • TFIIH\text{TFIIH}: Functions as a helicase to melt double-stranded DNA and separate template from non-template strands at the initiation site.
  • Experimental Demonstration of Promoter Spacing:

    • The distance from the TATA\text{TATA} box dictates the transcription start site position.
    • Moving the TATA\text{TATA} box sequence to 60 base pairs60\text{ base pairs} upstream (60-60) causes transcription initiation to shift to the 30-30 position.
    • Moving the TATA\text{TATA} box 15 base pairs15\text{ base pairs} closer to the original start site shifts transcription initiation downstream by 15 base pairs15\text{ base pairs}.
    • Transcription consistently initiates precisely 30 base pairs30\text{ base pairs} downstream from the TATA\text{TATA} box.

The Transcription Termination Mechanism

  • DNA-RNA Hybrid Helix:

    • Inside the active site of RNA Polymerase II, the newly synthesized RNA product remains base-paired with the template DNA strand over a short hybrid helical region.
    • This DNA-RNA hybrid association provides essential binding stability that keeps RNA Polymerase II attached to the template strand during elongation.
  • Mechanism of Polymerase Release:

    • Polyadenylation machinery cleaves the pre-mRNA transcript downstream of the poly-A signal sequence.
    • The remaining transcript fragment attached to RNA Polymerase II lacks a 55' cap and lacks a 55' triphosphate (possessing only one phosphate).
    • Capping enzymes do not recognize this exposed end.
    • Exonucleases target this uncapped 55' end and rapidly degrade the trailing RNA.
    • Exonucleolytic degradation degrades the RNA up into the exit channel, disrupting the short DNA-RNA hybrid complex within the active center.
    • Destabilization of the hybrid complex causes RNA Polymerase II to lose its grip on the template DNA strand and fall off.

Spliceosome Assembly and Catalytic Transesterification

  • Exons vs. Introns:

    • Exons: Protein-coding regions retained in mature mRNA. Average length is approximately 150 base pairs150\text{ base pairs}.
    • Introns: Non-coding intervening sequences removed during RNA processing. Average length is approximately 3,500 base pairs3,500\text{ base pairs} (can extend up to tens of thousands of base pairs).
  • Consensus Sequences at Splice Junctions:

    • 55' Splice Site: Defines the junction at the 55' end of the intron. Every intron begins with the conserved dinucleotide sequence GU\text{GU}. The upstream exon usually terminates in AG\text{AG}.
    • 33' Splice Site: Defines the junction at the 33' end of the intron. Every intron ends with the conserved dinucleotide sequence AG\text{AG}, preceded by a pyrimidine-rich region.
    • Branch Point Adenine: Conserved adenine (A\text{A}) nucleotide located upstream of the pyrimidine-rich region near the 33' end of the intron.
  • Spliceosome Composition:

    • Composed of 55 Small Nuclear RNAs (snRNAs): U1\text{U1}, U2\text{U2}, U4\text{U4}, U5\text{U5}, and U6\text{U6} (note: U3\text{U3} is not part of pre-mRNA splicing).
    • snRNAs complex with roughly 100100 distinct proteins to form small nuclear ribonucleoprotein particles (snRNPs).
    • Phosphorylation of the Carboxy-Terminal Domain (CTD\text{CTD}) of RNA Polymerase II facilitates co-transcriptional assembly of spliceosomal components.
  • snRNP Binding and Tertiary Pairing:

    • U1 snRNP\text{U1 snRNP} contains an snRNA that base-pairs directly with the nucleotide sequence at the 55' splice site.
    • U2 snRNP\text{U2 snRNP} base-pairs with the sequence surrounding the Branch Point Adenine.
    • Bulged Adenine Conformation: The U2 snRNA\text{U2 snRNA} sequence matches the nucleotides surrounding the branch point, but has no complementary base for the Branch Point Adenine itself. This lack of pairing forces the Branch Point Adenine to bulge out of the double helix, exposing its 2OH2'-\text{OH} group.
    • U4/U5/U6 Triple snRNP\text{U4/U5/U6 Triple snRNP} joins as a pre-assembled unit to bring the 55' splice site, 33' splice site, and branch point into physical proximity.
  • The Two Transesterification Reactions:

    1. First Transesterification:
      • The exposed 2OH2'-\text{OH} group of the bulged Branch Point Adenine performs a nucleophilic attack on the phosphodiester bond at the 55' splice site (between the last exon nucleotide and the first intron nucleotide G\text{G}).
      • This breaks the 55' exon-intron phosphodiester bond and forms a novel 252'-5' phosphodiester linkage between the Branch Point Adenine and the 55' end of the intron, creating a lariat structure.
    2. Second Transesterification:
      • The newly liberated 3OH3'-\text{OH} group of the upstream exon attacks the phosphodiester bond at the 33' splice site (following the final G\text{G} of the intron).
      • This joins the two exons via a standard 535' \rightarrow 3' phosphodiester bond and releases the intron in an excised lariat form.
    • The excised lariat intron is subsequently degraded by nuclear enzymes.
  • Exon Junction Complex (EJC\text{EJC}):

    • A protein complex deposited directly onto the mRNA junction following successful exon ligation.
    • Serves as a biochemical marker signifying that splicing has occurred at that specific site.

Exon Definition Hypothesis and Splice Site Fidelity

  • The Splice Site Recognition Problem:

    • Short splice consensus sequences (GU\text{GU} and AG\text{AG}) and single branch point adenines occur randomly throughout vast (average 3,500 base pair\text{average } 3,500\text{ base pair}) intron sequences.
    • The splicing machinery must distinguish authentic splice junctions from pseudo-splice sites present within introns.
  • Exon Definition Components:

    • Exonic Splicing Enhancers (ESEs\text{ESEs}): Specific nucleotide sequences located within exons. They encode amino acids in the protein while simultaneously serving as binding sites for splicing machinery.
    • SR Proteins: RNA-binding proteins enriched in serine ($ ext{S}$) and arginine ($ ext{R}$) residues that bind specifically to ESEs\text{ESEs} within exons.
    • U1 snRNP\text{U1 snRNP} exhibits high binding affinity for SR proteins, targeting U1\text{U1} to authentic 55' splice sites situated adjacent to exon-bound SR proteins.
    • U2AF\text{U2AF} (U2\text{U2} Associated Factor):
      • Composed of two subunits: U2AF65\text{U2AF}^{65} and U2AF35\text{U2AF}^{35} (representing 6565 and 3535 domain structure designations, totaling 100100 units of the heterodimer complex).
      • U2AF\text{U2AF} binds SR proteins on exons and recognizes the pyrimidine-rich tract at the 33' splice site, positioning U2 snRNA\text{U2 snRNA} at the Branch Point Adenine.
  • Cross-Exon Recognition Complex:

    • SR proteins bridge interactions between U1 snRNP\text{U1 snRNP} at the downstream 55' splice site of an exon and U2AF\text{U2AF} / U2 snRNP\text{U2 snRNP} at the upstream 33' splice site of the same exon.
    • This cross-exon interaction defines the boundaries of short exons (average 150 base pairs\text{average } 150\text{ base pairs}) across long intron spaces, ensuring accurate splice site pairing.

Alternative Splicing Mechanisms and Biological Significance

  • Biological Function and Genomic Efficiency:

    • Alternative splicing allows a single pre-mRNA transcript to generate multiple distinct mRNA variants that translate into different protein isoforms.
    • Isoforms derived from the same gene share high structural identity but differ in specific domains, altering functional activity, localization, or binding properties across tissues.
    • Conserves genomic space by generating vast proteomic diversity from a limited genome size.
  • Species Prevalence Rates:

    • Humans: Approximately 75 percent75\text{ percent} of all genes undergo alternative splicing.
    • Drosophila melanogaster: Approximately 40 percent40\text{ percent} of genes undergo alternative splicing.
    • Yeast: Only about 300300 out of 6,2006,200 total genes undergo alternative splicing.
  • Modes of Alternative Splicing:

    • Exon Skipping: Inclusion or exclusion of specific exons in the mature mRNA.
    • Intron Retention: An intron sequence is not excised and remains as coding sequence in mature mRNA.
    • Alternative 55' Splice Site: Splicing occurs at an internal 55' site, extending or shortening the upstream exon.
    • Alternative 33' Splice Site: Splicing occurs at an internal 33' site, extending or shortening the downstream exon.
    • Mutually Exclusive Exons: Ingestion of either Exon A or Exon B into mature mRNA, but never both simultaneously.
  • Regulatory Factors:

    • Splice Site Activators: Proteins that bind exonic/intronic enhancers to recruit snRNPs to adjacent splice sites.
    • Splice Site Repressors: Proteins that bind splice site sequences, sterically blocking spliceosomal assembly.
  • Tissue-Specific Example - α\text{α}-Tropomyosin:

    • The α\text{α}-tropomyosin gene contains multiple exons and undergoes tissue-specific alternative splicing.
    • Produces distinct functional protein isoforms in striated muscle, smooth muscle, myoblasts/fibroblasts (two distinct variants), and brain tissue.

Regulation of Alternative Terminal Exons

  • Structural Constraint of Splicing Machinery:

    • Splicing strictly requires both a 55' splice site and a 33' splice site surrounding an intervening region.
    • Alternative splicing cannot generate isoforms lacking terminal exons (the initial 55' exon or final 33' exon) because terminal ends lack outer splice junctions (First exons lack an upstream 33' splice site; terminal exons lack a downstream 55' splice site).
  • Mechanisms for Terminal Exon Variation:

    • Alternative Transcription Start Sites (Promoters):
      • Genes contain multiple promoter elements and +1+1 start sites located upstream of alternative first exons.
      • Assembling the transcription initiation complex at an internal promoter bypasses upstream exons entirely at the level of transcription.
    • Alternative Polyadenylation Sites (Poly-A Sites):
      • Genes contain multiple poly-A cleavage signals (AU\text{AU}-rich and GU\text{GU}-rich elements) located downstream of different internal terminal exons.
      • Cleavage and polyadenylation at an upstream poly-A site terminates the transcript prior to downstream exons, dictating the inclusion of alternative 33' terminal exons.

Dialogue and Audience Questions

  • Question on Untranslated Regions (UTRs):

    • Student Prompt: The untranslated regions are important because they contain valuable information that cannot be cut off, which is why transcription starts right at the +1+1 site. What is that information and why is it important?
    • Response: Details regarding the specific regulatory functions of UTRs will be addressed in an upcoming lecture.
  • Question on RNA Polymerase Strand Selection:

    • Student Prompt: How do you know which strand is the template strand when drawing a gene?
    • Response: Two criteria identify the template strand: (1) The mRNA product sequence is complementary to the template strand sequence; (2) RNA Polymerase II always moves along the template strand in the 353' \rightarrow 5' direction.
  • Question on Non-Template Strand Displacement:

    • Student Prompt: When the product pairs with the template strand, what happens to the non-template strand?
    • Response: The non-template strand is temporarily displaced by RNA Polymerase II, but re-hybridizes with the template strand behind the moving polymerase complex.
  • Question on Energy Requirements of Transcription:

    • Student Prompt: Does adding a new nucleotide to the RNA chain require an input of energy?
    • Response: It does not require external energy input. Cleavage of the phosphoanhydride bond between the alpha\text{alpha} and beta\text{beta} phosphates releasing pyrophosphate is energetically favorable, driving phosphodiester bond synthesis forward.
  • Question on DNA Intron Preservation:

    • Student Prompt: Why does DNA include introns if they are just going to be degraded anyway?
    • Response: Introns allow alternative splicing, enabling complex organisms to produce a vast diversity of protein isoforms from a limited total number of genes.
  • Question on Protein Isoform Classification:

    • Student Prompt: Are protein isoforms derived from alternative splicing considered different proteins?
    • Response: Isoforms are structurally distinct yet derived from the same gene; they are variant forms of the same overarching protein family modified for specific functional context.