Transcription, RNA Types, and DNA Sequencing: Comprehensive Study Notes

Context and Key Ideas

  • Week context: lecture is week 5 of the semester; upcoming two-week break with some assessments; final six weeks of the B semester discussed.
  • Last topics recap: DNA replication copies the genome; DNA polymerase is the key enzyme; PCR covered (in vitro DNA amplification) versus DNA replication (in vivo in cells).
  • PCR significance: PCR has revolutionised molecular biology; Carey Mullis (Nobel Prize winner) developed PCR; origin story about thermophilic enzymes and a highway epiphany; early PCR machines required manual temperature cycling, later replaced by robots and thermocyclers.
  • RT-PCR and qPCR intro:
    • RT-PCR = reverse transcription PCR: converts RNA (especially mRNA transcripts) into complementary DNA (cDNA) using reverse transcriptase (an enzyme from a virus).
    • After RT step, a standard PCR amplifies the cDNA.
    • qPCR (quantitative RT-PCR): uses fluorescent dyes or labeled primers to quantify amplification, allowing measurement of gene expression levels.
    • Gene expression concept: highly expressed genes yield lots of mRNA transcripts and hence more cDNA; lowly expressed genes yield less cDNA.
    • Exponential amplification: PCR is exponential; more initial RNA leads to faster, earlier detectable fluorescence;
    • mathematical intuition: the amount after n cycles is N=N0(1+E)nN = N_0(1+E)^n where 0 < E \le 1 is the amplification efficiency.
    • Practical use: compare expression of genes under different conditions (e.g., cells treated with a cancer drug vs untreated) to see which genes are up- or down-regulated.
  • DNA sequencing overview:
    • Fred Sanger’s DNA sequencing method pioneered the field; Sanger sequencing uses dideoxynucleotides (ddNTPs) to terminate DNA synthesis; lack of 3′ OH prevents further elongation, producing fragments of every length; fragments are separated by size and read by fluorescence to determine the sequence.
    • Sanger sequencing helped enable next-generation sequencing (NGS) platforms; the Singer Centre (DNA sequencing) contributed to early genome projects; sequencing has dramatically reduced costs and increased data availability (e.g., human genome projects).
    • Basic workflow for many NGS approaches:
    • Fragment DNA into manageable pieces (e.g., ~300 bases).
    • Attach adapters to fragment ends (known sequences) and perform PCR to amplify.
    • Bind fragments to beads or droplets for downstream sequencing
    • Sequence by synthesis or other readout methods to determine order of nucleotides.
    • Major sequencing technologies discussed:
    • Pyrosequencing: detects release of pyrophosphate (PPi) upon nucleotide incorporation.
    • Illumina (sequencing by synthesis): uses fluorescently labeled nucleotides; detects incorporated base by color in cycles.
    • Sequencing by other methods (less focus): sequencing by ligation, etc.
    • Single-molecule real-time sequencing (SMRT, e.g., PacBio): does not require amplification; long reads, can read long contigs.
    • Nanopore sequencing: DNA/RNA threaded through a pore; current changes identify bases; portable and usable in low-resource settings.
    • Practical notes:
    • DNA can be sequenced after fragmentation with adapters; RNA can be sequenced after converting to cDNA (RNA sequencing, or RNA-Seq).
    • Typical workflow variations exist between platforms, but core idea is to read nucleotides as fragments are amplified and/or detected.
    • Visual intuition given in lecture: beads in 2,000,000 wells; each well carries a different DNA fragment; nucleotides flow across and signals indicate incorporation; readout determines the sequence.
  • Transition to transcription (this lecture’s focus): core concept is the central dogma: DNA → RNA → protein; transcription (DNA to RNA) precedes translation (RNA to protein).
    • Reverse transcription is possible only in some viral systems and is not common in normal cellular biology.
    • The genetic code flow is often likened to a blueprint (DNA) being photocopied (RNA) and read by a factory (ribosome) to create proteins.
  • RNA types and roles:
    • Messenger RNA (mRNA): coding RNA that carries genetic information from DNA to the ribosome for translation into protein.
    • Non-coding RNAs (ncRNAs): RNAs transcribed but not translated; include:
    • Ribosomal RNA (rRNA): core component of ribosomes; large amounts in cells; essential for ribosome structure and function.
    • Transfer RNA (tRNA): ~80 nucleotides; delivers specific amino acids to the ribosome during translation.
    • Small nuclear RNAs (snRNA, ~150 nt): involved in processing pre-mRNA in eukaryotes.
    • Small nucleolar RNAs (snoRNA, 60–300 nt): guide nucleotide modification of rRNA and other RNAs.
    • PiRNA (piwi-interacting RNAs; note in lecture: “pili interacting RNAs,” ~30 nt): bind genome DNA and help stabilize it during meiosis/gametogenesis.
    • MicroRNAs (miRNA, ~21–22 nt): bind to long RNAs to inhibit translation; regulate a large fraction of coding genes (estimates suggest involvement in regulation of over 60% of coding genes); important in cancer, heart disease, and other conditions.
    • Long non-coding RNAs (lncRNA, ~200 nt and longer): diverse functions, many still unknown; roles in regulation and chromatin state.
    • Emphasis in lecture: while many RNA types exist, the main focus for transcription is on mRNA as the template for translation.
  • Transcription basics: converting DNA to RNA
    • Transcription uses a DNA template strand to produce an RNA transcript; the RNA sequence is complementary to the template and essentially the same as the coding strand (except with U replacing T).
    • Why use the template strand? If the coding strand were used, the RNA base would mirror T in the coding strand; that would not match the gene’s real RNA sequence (which contains Us). Example given: coding strand sequence would yield RNA with UAC UAG, which would differ from the gene’s actual RNA sequence.
    • Therefore, RNA sequence equals the coding sequence (with T replaced by U) when read from the coding strand, but is synthesized from the template strand.
    • Core enzyme: RNA polymerase (RNA Pol) synthesizes RNA from a DNA template; key differences from DNA polymerase:
    • RNA Pol does not require a primer to start transcription.
    • RNA Pol synthesizes RNA 5'→3' while reading the template DNA 3'→5'.
    • In bacteria (prokaryotes), a single RNA polymerase exists; in eukaryotes, there are three RNA polymerases.
  • Three stages of transcription: initiation, elongation, termination
    • Initiation: RNA polymerase binds to a promoter sequence to begin transcription. Promoter strength can vary, affecting transcription level (strong promoter = high transcription; weak promoter = low transcription).
    • Initiation in prokaryotes (bacteria): promoter features include two regions: -10 region and -35 region; start site is the transcription start site; -10 region is AT-rich (A–T pairs with 2 hydrogen bonds) and facilitates strand separation; -35 region provides recognition by a sigma factor.
    • Sigma factor (blue in lecture diagrams) is a protein that guides RNA polymerase to the promoter and enables transcription initiation; the RNA polymerase plus sigma factor form the holoenzyme.
    • In E. coli there are seven sigma factors; examples:
      • Sigma-70: housekeeping factor; initiates transcription of general function genes.
      • Sigma-19: regulates iron transport and metabolism under low-iron conditions.
      • Sigma-38: starvation or stationary-phase response.
    • Elongation: after initiation, RNA polymerase moves along the DNA template, synthesizing RNA in the 5'→3' direction; the DNA template is read 3'→5'. Multiple RNA polymerases can transcribe simultaneously along the same gene, producing a gradient of transcript lengths along the DNA.
    • Termination in prokaryotes:
    • Rho-dependent termination: requires a Rho factor that binds the newly synthesized RNA; when RNAP-rRNA encounters Rho, transcription terminates by pulling RNA-DNA apart.
    • Rho-independent termination: relies on specific GC-rich sequences in the RNA that form a hairpin structure (a stem-loop). After the GC-rich hairpin, a string of uracils (U) in RNA pairs with adenines (A) in the DNA template, destabilizing the transcription complex and causing termination.
    • Eukaryotic transcription differs in several aspects:
    • There are three RNA polymerases (Pol I, II, III) rather than one.
    • Sigma factors are replaced by transcription factors; initiation often requires multiple factors; promoter features include the TATA box (often ~25–30 bases upstream of the start site).
    • Promoter examples: TATA box is a core promoter element recognized by transcription factors (e.g., TATA-binding protein) that recruit RNA polymerase II.
    • Termination mechanisms differ by polymerase:
      • Pol I termination involves a termination factor that blocks transcription.
      • Pol II termination involves cleavage of the nascent transcript at an internal site followed by disengagement of RNA polymerase.
      • Pol III termination mechanics are less well defined.
  • RNA processing in eukaryotes (distinct from prokaryotes): key modifications to primary mRNA transcripts
    • 5' capping: addition of a 5' cap (7-methylguanosine cap) to the nascent RNA.
    • 3' polyadenylation: addition of a poly-A tail (a string of adenine nucleotides) to the 3' end; tail helps stabilize the mRNA.
    • Splicing: removal of introns (non-coding segments) and joining of exons (coding segments) to produce mature mRNA; intron/exon organization in many genes.
    • Alternative splicing: different exon combinations yield multiple protein isoforms from a single gene; example given with exons 1–5 producing proteins A, B, and C depending on which exons are included.
    • Rationale for introns and splicing: introns allow alternative splicing and thus proteome diversification, even though splicing adds complexity.
  • Why these processes matter in the real world
    • Transcription and RNA processing regulate gene expression and protein production, enabling cells to respond to environmental cues (e.g., nutrient availability, stress, drug treatment).
    • MicroRNAs and other non-coding RNAs are critical regulators of gene expression and can influence disease states, development, and cellular function.
    • The central dogma (DNA → RNA → Protein) underpins molecular biology, with reverse transcription (RNA → DNA) being a rare exception predominantly used by certain viruses.
  • Quick references to terminology and structural concepts
    • DNA vs RNA: sugar differences (deoxyribose vs ribose); thymine vs uracil; DNA is typically double-stranded; RNA is often single-stranded.
    • Base-pairing rules: DNA: A pairs with T (2 hydrogen bonds), C with G (3 hydrogen bonds); RNA: A pairs with U; C pairs with G.
    • Coding vs template strand: coding strand sequence is used to infer the mRNA sequence (with T→U); the template strand is the actual template for RNA synthesis.
  • Summary of the central dogma and key transitions
    • DNA stores genetic information; transcription produces RNA; translation uses RNA to synthesize proteins.
    • Reverse transcription is not common in normal cells and is generally viral (reverse transcriptase converts RNA to DNA).
  • Exam-ready takeaways
    • Know promoter architecture in bacteria (-10 and -35), sigma factors, and holoenzyme concept.
    • Distinguish prokaryotic transcription (rho-dependent/independent termination) from eukaryotic transcription (TATA box, transcription factors, chromatin structure, histone dynamics, and FACT protein role in chromatin remodeling).
    • Understand the differences and similarities between transcription and replication (directionality, primer usage, processive movement, and multiple polymerases).
    • Be able to explain the purpose and workflow of RT-PCR and qPCR, including how expression levels are inferred from amplification trajectories.
    • Recognize major sequencing methods (Sanger vs next-gen) and the general concept of how adapters, beads, PCR, and signal readouts enable reading DNA sequences; be aware of read-length and portability differences among platforms.
  • Quick practice prompts (from lecture context)
    • If given a DNA coding strand sequence, derive the corresponding mRNA sequence and explain how the template strand yields the correct mRNA.
    • Diagram the steps of prokaryotic transcription initiation, elongation, and termination, including the roles of the -10/-35 promoter elements, sigma factors, and rho-dependent/independent termination.
    • Compare and contrast eukaryotic transcription termination for Pol I, Pol II, and Pol III.
    • Explain why splicing is beneficial and how it enables alternative protein isoforms.
  • Note on further resources
    • Lecture notes include YouTube links and videos to reinforce concepts such as transcription initiation, elongation, termination, and sequencing technologies.

Transcription and RNA Types – Key Concepts (Concise Reference)

  • Central dogma: DNA → RNA → Protein; reverse transcription is rare (viruses only).
  • DNA vs RNA: sugar (deoxyribose vs ribose), base (T vs U), strandness (DNA usually double-stranded; RNA often single-stranded).
  • Promoter architecture (bacteria): -10 region (AT-rich) and -35 region; sigma factor recruits RNA polymerase; holoenzyme = RNA polymerase + sigma factor.
  • Directionality: transcription 5′→3′; DNA template read 3′→5′.
  • RNA types: mRNA (coding), rRNA, tRNA, snRNA, snoRNA, miRNA (~21–22 nt), piRNA (~30 nt), lncRNA (≈200 nt+).
  • RNA processing in eukaryotes: 5′ cap, 3′ poly-A tail, intron removal (splicing) with exons joined; alternative splicing generates protein diversity.
  • Post-transcriptional regulation: miRNAs regulate translation of many genes; non-coding RNAs play broad regulatory roles.
  • Translational bridge: mRNA carries code to the ribosome to synthesize proteins; tRNA delivers amino acids; rRNA is a ribosome core component.
  • Practical applications highlighted: RT-PCR/qPCR for gene expression; DNA sequencing technologies enabling genome projects; ongoing evolution in sequencing platforms (Sanger, Pyrosequencing, Illumina, SMRT, Nanopore).