RNA Virus Transcription & Replication: Retroviruses, Negative-Sense, and Positive-Sense
Retroviruses (HIV) — overview and reverse transcription
Retroviruses carry RNA genomes that ultimately become DNA in the host cell via reverse transcription; the defining feature is the enzyme reverse transcriptase (RT).
Initial genome form: messenger RNA-like RNA, but the genome is not translated directly like a typical positive-sense RNA virus. Instead, the RNA genome is reverse transcribed to a DNA copy that will integrate into the host genome.
Key steps (high-level): RNA genome → single-stranded DNA → double-stranded DNA → integration into host genome → host RNA polymerase II transcribes the provirus to produce viral mRNAs (and genomic RNA).
The genome inside virions is capped and tailed (mRNA-like), but this is used as a template for reverse transcription, not directly translated as the genome.
Integrates into host chromosome via an integration step (unlike some plasmids that stay as extrachromosomal DNA). The integrated viral DNA is called a provirus.
After integration, transcription is driven by promoter elements in the long terminal repeats (LTRs).
Transcription and gene expression depend on host transcription machinery but are guided by viral promoter elements and bound transcription factors.
Major concepts of reverse transcription in retroviruses (overview):
Template: the viral RNA genome is used as the template for reverse transcription to DNA.
Primer: a cellular tRNA binds the primer-binding site (PBS) on the viral genome and acts as a primer for DNA synthesis.
First product: a DNA-RNA hybrid forms; later, RNase H removes the RNA strand, leaving a single-stranded DNA, which is subsequently used as a template to synthesize the complementary strand and produce double-stranded DNA.
Integration: the dsDNA (provirus) is integrated into the host genome. From there, cellular transcription produces viral RNAs.
HIV-specific notes (example retrovirus):
RNA genome is reverse-transcribed using a tRNA primer at the PBS; the tRNA is cellular, not viral, and base-pairs with PBS to initiate DNA synthesis.
The initial product is a DNA-RNA hybrid; RNase H degrades the RNA in the hybrid, yielding single-stranded DNA; this then serves as a template to make the second DNA strand.
The resulting dsDNA integrates into the host chromosome as a provirus with LTRs flanking the integrated sequence.
After integration, transcription by RNA polymerase II occurs from the LTR promoter to produce viral mRNAs and the full-length genomic RNA.
Viral regulatory strategy includes promoter elements in LTRs (e.g., SP1 and NF-κB binding sites) to recruit transcription factors and RNA Pol II for efficient transcription.
Transcription initiation and elongation are modulated by viral and host factors; a key player is the TAP protein which helps recruit host factors (Cyclin T and CDK9) that phosphorylate the C-terminal domain (CTD) of RNA Pol II to enhance processivity and overcome early abortive transcription.
Early transcription may produce many short, nonproductive transcripts; as TAP accumulates, transcription switches to efficient production of full-length viral RNAs.
A burst of viral mRNA and protein production can occur once TAP triggers RNAP II phosphorylation, helping the virus outrun innate immune responses.
Transcriptional features and regulatory elements
Promoter elements in the LTR recruit transcription factors (e.g., SP1, NF-κB) and RNA Pol II to initiate transcription.
After initiation, early transcription yields short, abortive transcripts due to limited processivity until TAP-driven phosphorylation of RNAP II CDK9 increases processivity.
TAP recruits cyclin T and CDK9; CDK9 phosphorylates the RNAP II CTD, converting the polymerase to a more processive form.
Once enough TAP is produced, a transcription burst occurs, increasing production of viral RNAs and proteins, which helps the virus compete with the host innate immune response.
Reverse transcription and error-prone replication in HIV
Reverse transcription is a multi-step, error-prone process leading to genetic diversity.
Fidelity of reverse transcription is relatively low, contributing to a high mutation rate in HIV genomes.
On average, there is about , which has consequences for immune escape and antiviral resistance.
The high mutation rate creates a swarm (quasi-species) of related viruses; some variants may evade immune responses or antiviral drugs.
Reverse transcription: a quick recap (step-by-step, high level)
1) Viral RNA genome; tRNA primer binds PBS and initiates DNA synthesis to form a DNA-RNA hybrid.
2) RNase H degrades the RNA strand of the hybrid, leaving a single-stranded DNA.
3) The remaining DNA serves as a template to synthesize the second DNA strand, yielding double-stranded DNA.
4) The dsDNA is transported to the nucleus and integrates into the host genome as a provirus.Integration and the provirus concept
Integration places viral DNA into a host chromosome; the proviral DNA can be transcriptionally active for prolonged periods.
The LTRs on both ends of the provirus act as promoters and regulatory regions for transcription of viral RNAs.
Visual cues used in teaching this topic
In textbooks, RNA is often colored green (viral RNA genomes), while DNA is colored blue (the newly synthesized DNA). Reverse transcription creates a dsDNA provirus that integrates into the host genome.
Preview of upcoming topics
A dedicated session later will cover step-by-step reverse transcription mechanics more thoroughly.
We will dive into HIV in more detail, including how reverse transcription and integration are regulated and how antiviral therapies target these steps.
Negative-sense RNA viruses (Rhabdoviruses) — genome organization and transcription
Model virus: Vesicular stomatitis virus (VSV) and rabies virus are prototypical rhabdoviruses; they are among the simplest negative-sense RNA viruses.
Genome and core components
The genome is negative-sense single-stranded RNA, encapsidated by nucleocapsid (N) protein.
The basic genome organization includes five genes: N, P, M, G, L.
L is the large polymerase; P is the phosphoprotein that serves as a cofactor for L.
N protein coats the genome, forming a nucleocapsid; transcription and replication depend on the nucleocapsid-bound RNA template.
M is the matrix protein; G is the glycoprotein responsible for receptor binding and membrane fusion.
Why these viruses must encode their own RdRP
The negative-sense genome cannot serve directly as mRNA; the virus must encode an RNA-dependent RNA polymerase (RdRP) to synthesize mRNA and replicate the genome.
The RdRP is a heteromeric complex consisting of L (catalytic subunit) and P (cofactor) proteins.
Transcription vs. genome replication in rhabdoviruses
Transcription: The RdRP binds upstream of the first gene (N) and initiates transcription to produce five separate (+) sense mRNAs corresponding to N, P, M, G, and L.
Intergenic (IG) sequences between genes regulate termination and reinitiation of transcription.
A key feature is a gradient of transcription: more mRNA is produced for genes located at the 3' end (N) and progressively less for genes downstream (P, M, G, L).
The polymerase toggles between two modes: transcription mode (producing separate mRNAs with a 5' cap and 3' poly(A) tail) and genome replication mode (producing full-length antigenome and genomes without individual gene transcripts).
Transcription mechanics and the gradient phenomenon
The RdRP binds only at a single promoter region upstream of the N gene.
It transcribes through the genome from 3' to 5' end, producing individual mRNAs one by one.
Intergenic sequences (IG) cause the polymerase to terminate and either reinitiate at the next gene or fall off and have to rebind at the promoter for another attempt.
Probabilistic reinitiation: about and for each IG junction.
Result: a transcription gradient where N mRNA is most abundant, and L mRNA is least abundant.
Polyadenylation and capping in rhabdoviruses
The L protein has the capping activity; the polymerase also contributes to adding a 5' cap to the mRNAs.
At the end of gene transcription, a poly(A) tail is added via a phenomenon called polymerase slippage (stuttering) at a run of seven uracils (7U).
Mechanism: the polymerase stalls at the 7U sequence, adds multiple adenines (non-templated), and leaves a poly(A) tail on the mRNA. Typical tail length is on the order of hundreds of adenines (e.g., ~200 As).
Nucleocapsid protein and transcriptional coupling
Transcription occurs with the RdRP bound to the nucleocapsid (N) protein; the polymerase does not function efficiently on naked RNA.
The N protein must participle in the transcription process to ensure proper template recognition and gene expression.
Genome replication in rhabdoviruses
After transcription of monogenic mRNAs, the virus must also replicate its genome to produce new negative-sense genomes for packaging.
Replication uses the same RdRP, but in a genome replication mode rather than transcription mode.
Replication produces a positive-sense antigenome, which then serves as the template for producing new negative-sense genomes.
The switch from transcription mode to replication mode is thought to be triggered by the accumulation of nucleocapsid protein (N). When N levels reach a threshold, the RdRP switches to replication mode and begins synthesizing antigenome/genome without producing separate mRNAs.
Antigenome and genome packaging
Antigenome is positive-sense and serves as a template to produce more negative-sense genomes.
The genomes that get packaged into virions are negative-sense and encapsidated by N.
Why negative-sense rhabdoviruses are organized in a specific gene order
The gene order N-P-M-G-L reflects the relative abundance of the proteins: N is most abundant (needed to coat the genome completely), while L (polymerase) is least abundant (an enzyme used repeatedly).
Summary for negative-sense RNA viruses (VSV-like)
Genomic RNA is negative-sense, uncapped, unpolyadenylated, and fully encapsidated by N.
Transcription yields five monocistronic mRNAs with 5' caps and 3' poly(A) tails; mRNAs are translated to produce viral proteins.
A polyadenylation tail is added via polymerase stuttering at 7U, generating a poly(A) tail on the transcript.
A replication mode follows transcription; genome replication generates antigenome, then genomes; switch is triggered by nucleocapsid protein accumulation.
Quick takeaway
The order of genes in negative-sense RNA viruses is tightly linked to the abundance of the corresponding proteins and to the need for rapid, robust genome replication in addition to mRNA transcription.
Aside: a quiz-style note from the lecture
A multiple-choice question compared class one and class two fusion mechanisms; the correct statement was identified as option C (describing properties that are not characteristic of class two fusion). The instructor suggested this as an example of applying structural/functional distinctions to exam-style questions.
Positive-sense RNA viruses — genome as mRNA, polyprotein strategy, and replication complexes
Core idea: the genome is positive-sense RNA and can act directly as mRNA; translation begins immediately after infection.
Immediate translation and the need for replication
The genome acts as mRNA and is translated to produce viral proteins immediately.
However, a single copy of the genome is not enough to sustain infection or produce progeny; RNA-dependent RNA polymerase (RdRP) is still needed to replicate the genome and transcribe subgenomic or replication intermediates.
Polyprotein strategy in many positive-sense RNA viruses
In many positive-sense RNA virus families (e.g., picornaviruses, flaviviruses, coronaviruses), the genome is translated as a single large polyprotein.
The polyprotein is cleaved by proteases (viral and/or cellular) at defined junctions to yield mature viral proteins.
Protease recognition sequences appear at junctions between protein domains; viral proteases (and sometimes cellular proteases) accomplish the cleavage.
The polyprotein approach explained
A single, large RNA genome yields one long polyprotein; this polyprotein is subsequently processed into functional mature proteins (structural and nonstructural).
This strategy allows coordinated production of multiple proteins from one RNA transcript.
Coronavirus-specific notes (a major positive-sense RNA virus family)
Coronaviruses have the largest genomes among positive-sense RNA viruses.
Not all proteins are produced immediately from the initial translation of the genome; some nonstructural proteins (e.g., replicase complex components) are produced first to assemble a replication/transcription complex.
Later, additional processing or regulatory steps allow translation of remaining structural and enzymatic proteins.
General features of replication complexes in positive-sense RNA viruses
Replication/transcription complexes form on cellular membranes, often associated with rearranged intracellular membranes.
Membrane association concentrates replication components and substrates, enhancing efficiency and helping to shield viral RNA from innate immune sensors.
The membrane compartmentalization also aids in immune evasion by limiting exposure of viral RNA to cytosolic sensors.
From genome to progeny: replication and transcription in positive-sense viruses
After translation of the polyprotein and assembly of replication machinery, the RdRP synthesizes a complementary negative-sense antigenome, which serves as a template to produce new positive-sense genomes.
Subgenomic RNAs may be produced in some viruses (e.g., alphaviruses, coronaviruses) to express downstream structural proteins; this involves discontinuous transcription and nested sets of subgenomic RNAs in some families.
Key contrast with retroviruses and negative-sense viruses
Positive-sense RNA viruses translate the genome directly into a polyprotein, whereas retroviruses rely on reverse transcription and integration into the host genome before producing viral RNAs.
Negative-sense RNA viruses must carry RdRP within the virion to transcribe their genome into mRNA, whereas many positive-sense viruses can synthesize their replication machinery from the translated polyprotein.
Summary points to remember
Positive-sense RNA virus genomes are directly usable as mRNA on entry, but replication requires an RdRP and a replication complex on membranes.
A common theme is polyprotein strategy with proteolytic processing to generate mature viral proteins.
Coronaviruses exemplify a family with very large genomes and a multi-step translation/replication strategy that builds replicase complexes before expressing all structural proteins.
Membrane-associated replication complexes help concentrate components and shield viral RNA from innate immune detection.
Quick comparative recap and practical implications
Retroviruses (e.g., HIV): RNA genome becomes DNA via reverse transcription, integrates into host genome as a provirus, uses host Pol II to transcribe viral RNAs; high mutational bias due to RT; TAP/CDK9-mediated transcriptional activation enables a burst of viral gene expression.
Negative-sense RNA viruses (e.g., VSV, rabies): encode RdRP (L and P); genome is encapsidated by N protein; transcription yields monocistronic mRNAs with a clear transcription gradient; replication produces antigenomes and new genomes; polyadenylation occurs via polymerase stuttering at 7U, generating poly(A) tails; the switch from transcription to replication is triggered by N protein levels.
Positive-sense RNA viruses (e.g., coronaviruses, flaviviruses, picornaviruses): genome acts as mRNA for immediate translation; replication requires RdRP and a replication complex on membranes; many produce a single massive polyprotein that is cleaved; coronaviruses may regulate timing to ensure replication machinery is established before full structural protein synthesis; replication complexes help with efficiency and immune evasion.
Common themes across classes
All rely on an RNA-dependent RNA polymerase (RdRP) except retroviruses (which use reverse transcriptase to create a DNA intermediary).
Genome replication strategies differ: direct translation (positive-sense), polyprotein processing (majority of positives), reverse transcription and integration (retroviruses), or segmented monocistronic transcription with a gradient (negative-sense rhabdoviruses).
Membrane association of replication/transcription complexes is a recurring theme for efficient replication and immune evasion in many positive-sense viruses.
Notes on important terminology and concepts to remember
RdRP: RNA-dependent RNA polymerase; required for replication/transcription of RNA viruses lacking a host RNA polymerase that can directly copy RNA templates.
L, P, N, M, G, L regions (negative-sense viruses):
N: nucleocapsid protein; coats the genome.
P: phosphoprotein; cofactor for RdRP.
L: large polymerase; catalytic component of RdRP and capping enzyme.
M: matrix protein; organizing virion assembly.
G: glycoprotein; receptor binding and membrane fusion.
LTR: long terminal repeat; promoter and regulatory region flanking proviral DNA in retroviruses.
PBS: primer binding site; where host tRNA binds to initiate reverse transcription.
TAP: transcriptional activation protein in HIV; helps recruit host factors to RNAP II.
CDK9/Cyclin T: host factors that phosphorylate the RNAP II C-terminal domain to increase transcriptional processivity.
IG (intergenic) sequences: sequences between genes in negative-sense viruses that regulate termination and reinitiation of transcription.
Antigenome: positive-sense intermediate used as template to make more genomes in negative-sense RNA viruses.
Processivity: how long a polymerase stays on the template to synthesize long RNA products without falling off.
Polyadenylation by stuttering: polymerase adds a string of A's at the 3' end of transcripts by repeatedly inserting A opposite a run of U's in the template.
Final takeaway
Each class of RNA virus uses a distinct but conceptually related strategy to express its genome and make progeny:
Retroviruses convert RNA to DNA and integrate; expression relies on host transcription with viral regulatory elements.
Negative-sense RNA viruses rely on their own RdRP to transcribe multiple monocistronic mRNAs and to replicate genomes, with a transcription gradient shaped by intergenic signals.
Positive-sense RNA viruses translate their genome directly (often as a polyprotein) and rely on replication complexes on membranes to copy their genomes, sometimes using subgenomic RNAs to express later genes.
Understanding these mechanisms helps explain antiviral targets (e.g., reverse transcription inhibitors for retroviruses, RdRP inhibitors for RNA viruses, protease inhibitors for polyprotein processing) and the challenges posed by high mutation rates (e.g., HIV).