Comprehensive Study Notes on Bacterial Transcription and RNA Polymerase Mechanics

Exam Logistics and Course Schedule

  • Course schedule adjustments regarding transcription coverage:
    • Transcription is covered across two instructional days.
    • Bacterial transcription is covered first in its entirety, followed by eukaryotic transcription to enable direct comparison and contrast between the two systems.
  • Scope of Exam 1:
    • Exam 1 covers Chapter 11 through Chapter 55.
    • Chapter 11 serves as the introductory chapter.
    • Chapter 55 will be completed prior to the exam.
    • Chapter 66 relates to translation and naturally integrates into translation topics; it will be introduced briefly following Chapter 55.
    • No new instructional material will be presented on the morning immediately preceding the exam.

Overview of Bacterial Transcription

  • Transcription is standardly abbreviated as TX.
  • Definition of Promoter:
    • A promoter is a specific DNA sequence that signals the starting location for transcription.
  • Mechanism in Bacteria vs. Eukaryotes:
    • Bacterial transcription mechanisms are relatively straightforward, highly defined, and well-characterized, primarily based on models studied in Escherichia coli (E. coli).
    • Variations on these mechanisms occur across different bacterial species.
  • Stages of Transcription:
    • Initiation (the beginning phase).
    • Elongation (the middle phase).
    • Termination (the ending phase).

RNA Polymerase Holoenzyme Structure and Gel Electrophoresis

  • Holoenzyme Composition:
    • The complete bacterial RNA polymerase holoenzyme consists of 55 polypeptide subunits.
    • Structural subunits include:
    • Two identical alpha (α\alpha) subunits.
    • Two large subunits close in mass: beta (β\beta) and beta prime (β′\beta').
    • One sigma (σ\sigma) subunit, which specifically functions to bind and bring in the DNA template by recognizing the promoter sequence.
    • The core enzyme consists of the catalytic subunits without the σ\sigma subunit; addition of the σ\sigma subunit forms the active holoenzyme.
  • X-ray Crystallography of RNA Polymerase:
    • Structural determination relies on X-ray crystallography, a technique developed over the past 7070 years.
    • Structural models (including false-color crystal structures approximately 2020 years old) depict the core enzyme complexed with a crystallized fragment of DNA, showing the σ\sigma subunit directly grasping the DNA.
  • Electrophoretic Separation of Subunits:
    • Polyacrylamide protein gel electrophoresis separates proteins based on molecular length/size under an electrical field.
    • Smaller polypeptides migrate faster through the gel matrix than larger polypeptides:
    • The smaller α\alpha subunits migrate faster than the larger β\beta subunits.
    • Standard gel electrophoresis of the holoenzyme yields 33 distinct resolved bands.
    • Resolution of β\beta and β′\beta' Subunits:
    • Running an extended gel over a longer duration successfully separates the similar-sized β\beta and β′\beta' subunits into distinct bands.
    • Alternatively, protein bands can be excised from the gel and analyzed via mass spectrometry, which cleaves peptides into fragments to determine molecular identity.
    • Running extended gels is functionally simpler and significantly more cost-effective than using specialized mass spectrometers (such as specialized instruments beyond standard laboratory equipment like Dr. Chandler's mass spectrometer).
  • Spatial Orientation:
    • Transcription is conventionally drawn from left to right, indicated by a bent arrow at the promoter sequence.
    • Within a living cell in three-dimensional space, gravitational forces are irrelevant at the molecular scale; thus, the spatial orientation of the σ\sigma subunit (whether depicted on top or bottom) does not alter the mandatory 5′5' to 3′3' chemical directionality of synthesis.

Bacterial Promoter Structure and Consensus Sequences

  • Consensus Sequence Concept:
    • A consensus sequence represents the calculated average or most common nucleotide sequence derived from aligning corresponding region sequences across multiple genes.
    • Individual promoters may match the consensus sequence exactly or display slight sequence variations.
    • E. coli possesses approximately 2,0002{,}000 genes, which represents roughly 10%10\% of the estimated 20,00020{,}000 to 25,00025{,}000 genes present in the human genome.
  • Key E. coli Promoter Elements:
    • The −10-10 sequence (−10-10 box):
    • Consensus sequence is TATAAT.
    • Composed entirely of adenine (A\text{A}) and thymine (T\text{T}) bases.
    • Serves as the precise site where DNA strand unwinding is initiated, as A-T\text{A-T} base pairs possess fewer hydrogen bonds than G-C\text{G-C} pairs and unwind more readily.
    • The −35-35 sequence (−35-35 box):
    • Consensus sequence is TTGACA.
    • The inverted or reverse sequence TGTCAA is non-equivalent and does not function identically.
  • Promoter Numbering Convention:
    • The transcriptional start site—the exact nucleotide base where transcription begins—is designated as +1+1.
    • There is no base 00 in the sequence numbering scheme.
    • The nucleotide base immediately preceding +1+1 is designated as −1-1.
    • The −10-10 box is located on average 1010 bases upstream of the +1+1 start site.
    • The −35-35 box is located on average 3535 bases upstream of the +1+1 start site.
    • An average E. coli promoter spans approximately 3535 base pairs in length.

Mechanism of Transcription Initiation and Strand Specificity

  • Template Strand vs. Coding Strand:
    • Template Strand: The physical DNA strand that is directly read and copied by RNA polymerase.
    • Coding Strand: The non-template DNA strand whose base sequence matches the resulting RNA transcript (with uracil replacing thymine).
    • RNA complementary base pairing rules apply during synthesis (e.g., a template thymine (T\text{T}) directs the insertion of an adenine (A\text{A}) into the RNA transcript).
    • Example: A template DNA sequence of 5'-CGTA-3' dictates an RNA sequence of 3'-GCAU-5' (or 5'-UACG-3').
  • Identifying Strand Identity in Genomes:
    • Neither DNA strand is exclusively the coding or template strand genome-wide; individual genes on the same chromosome can utilize opposite strands as their template.
    • Determining which strand serves as the template requires analyzing the physical gene product (mRNA transcript or encoded protein sequence, such as locating the start codon ATG).
    • Genomic annotation utilizes automated algorithms (such as shotgun sequencing compilation where short DNA reads are assembled) to identify candidate −35-35, −10-10, transcriptional start, and translational start sites (e.g., analyzing newly isolated bacterial strains from environmental sewage samples, such as from Gua). Experimental biochemistry is required to confirm active expression of a gene product.
  • Independence from Primers:
    • Unlike DNA polymerases (which require a pre-existing primer providing a free 3′3' hydroxyl end), RNA polymerases do not require a primer to initiate de novo polynucleotide synthesis.
  • Summary Steps of Initiation:
    1. RNA polymerase holoenzyme binds to the promoter sequence at the −35-35 and −10-10 regions via the σ\sigma subunit.
    2. The enzyme unwinds the double-stranded DNA at the −10-10 A-T\text{A-T} rich region, transitioning from a closed to an open promoter complex.
    3. RNA polymerase initiates RNA synthesis at the +1+1 nucleotide position.

RNA Synthesis Chemistry and Elongation

  • Polymerization Reaction Mechanics:
    • RNA synthesis proceeds exclusively in the 5′5' to 3′3' direction.
    • Incoming substrate nucleotides are ribonucleotide triphosphates containing ribose sugars.
    • The 3′3' hydroxyl (OH\text{OH}) group of the growing RNA chain attacks the alpha-phosphate of the incoming ribonucleotide triphosphate.
    • The reaction forms a phosphodiester bond and liberates a molecule of pyrophosphate (Pi-Pi\text{P}_\text{i}\text{-P}_\text{i}, composed of two covalently linked phosphate groups).
    • Pyrophosphate cleavage releases substantial free energy (denoted by the prefix pyro-, meaning fire) and is recycled by cellular pathways.
  • Elongation Process:
    • RNA polymerase moves along the template DNA, continuously unwinding the duplex ahead of the active site and adding complementary ribonucleotides to the 3′3' end of the growing RNA transcript.

Bacterial Transcription Termination

  • Hairpin-Dependent (Rho-Independent / Intrinsic) Termination:
    • As RNA polymerase transcribes the terminal region of a gene, it synthesizes a sequence rich in guanine (G\text{G}) and cytosine (C\text{C}) bases, followed by a non-complementary loop region and a reverse complementary G-C\text{G-C} rich segment.
    • The complementary G-C\text{G-C} rich regions on the single-stranded RNA transcript self-hybridize, snapping together to form a stable stem-loop (hairpin) secondary structure.
    • The strong hydrogen bonding within the G-C\text{G-C} hairpin structure physically stalls and destabilizes the RNA polymerase complex, causing the enzyme to dissociate from the DNA template and release the completed RNA transcript.

Questions and Student Discussion

  • Question: Is the subunit diagram inverted or upside down between different representations?
    • Answer: Spatial orientation in diagrams is arbitrary. Because gravity is not a functional force at the molecular level within a cell, 3D spatial alignment does not affect the process. Transcription proceeds along a specific chemical direction (5′5' to 3′3') regardless of visual representation.
  • Question: Why are promoter sequences designated with negative numbers like −10-10 and −35-35?
    • Answer: Negative numbers represent counting upstream (backwards) from the transcriptional start site, which is designated as +1+1. Because there is no base 00, the base directly preceding +1+1 is −1-1. Thus, −10-10 and −35-35 represent sites located on average 1010 and 3535 bases prior to the start site.
  • Question: How do you determine which strand is the coding strand versus the template strand?
    • Answer: Strand determination requires examining the actual gene product (mRNA or protein sequence). Within a single gene, only one strand acts as the template, but across an entire genome, either strand can serve as the template for different genes. On introductory examinations, strand designations (5′5' to 3′3' directionality and template/coding identity) are explicitly specified.
  • Question: Is the promoter region in E. coli typically around 3535 bases long?
    • Answer: Yes, an average E. coli promoter spans approximately 3535 bases. Eukaryotic promoters differ significantly in size and complexity.
  • Question: How are the similar-sized β\beta and β′\beta' subunits resolved if standard gels only show 33 bands?
    • Answer: Standard short gel runs group β\beta and β′\beta' into a single band due to their close molecular weight. Running an extended gel over a longer time separates them into distinct bands. Mass spectrometry can also identify them by peptide fragmentation, though extended gel runs are simpler and cheaper.