DNA Replication, Repair, Maintenance, and Somatic Rearrangements

Fundamental Properties and Classification of DNA Polymerases

DNA polymerases are the primary enzymes responsible for catalyzing the synthesis of new DNA strands during replication and repair processes. All known DNA polymerases share two fundamental enzymatic properties that dictate the molecular mechanism of DNA duplication:

  1. Strict 5′5' to 3′3' Directionality: DNA polymerases synthesize DNA exclusively in the 5′→3′5' \rightarrow 3' direction. Synthesis proceeds via a nucleophilic attack where the 3′3' hydroxyl (OH\text{OH}) group of the terminal sugar in the growing primer strand attacks the innermost (alpha\text{alpha}) phosphate of an incoming deoxyribonucleoside triphosphate (dNTP\text{dNTP}). This reaction releases inorganic pyrophosphate (PPi\text{PP}_i) and creates a phosphodiester bond, extending the chain by one nucleotide.

  2. Absolute Requirement for a Primer: Unlike RNA polymerases, which can initiate polynucleotide synthesis de novo, DNA polymerases cannot synthesize DNA from scratch. They require a pre-existing primer strand—hydrogen-bonded to the single-stranded template—that provides a free 3′ OH3' \text{ OH} group for extension.

The reaction catalyzed by DNA polymerase
Diverse Families of DNA Polymerases

Cells express multiple specialized DNA polymerases divided across prokaryotic and eukaryotic lineages:

  • Bacterial DNA Polymerases (E. coli):

    • DNA Polymerase III (Pol III): The main replicative enzyme responsible for continuous and discontinuous strand elongation at the replication fork.

    • DNA Polymerase I (Pol I): Primarily functions in DNA repair and processing of Okazaki fragments by removing RNA primers via its 5′→3′5' \rightarrow 3' exonuclease activity and replacing them with DNA.

  • Eukaryotic DNA Polymerases:

    • DNA Polymerase alpha\text{alpha} (Pol β\boldsymbol{\beta} / Pol ρ\boldsymbol{\rho} complex): Initiates DNA synthesis by forming a complex with primase to synthesize short RNA-DNA primers.

    • DNA Polymerase epsilon\text{epsilon} (Pol ρ\boldsymbol{\rho}): Primary replicative polymerase responsible for leading strand synthesis.

    • DNA Polymerase delta\text{delta} (Pol ρ\boldsymbol{\rho}): Primary replicative polymerase responsible for lagging strand synthesis.

    • DNA Polymerase gamma\text{gamma} (Pol ρ\boldsymbol{\rho}): Replicates and repairs mitochondrial DNA.

Mechanics and Dynamics of the Replication Fork

DNA replication originates at specific genomic sites and proceeds bidirectionally. As parental double-stranded DNA unwinds, it forms two Y-shaped structures termed replication forks moving in opposite directions away from the origin.

Replication of E. coli DNA
The Antiparallel Challenge: Leading and Lagging Strands

Because the two parental DNA strands are antiparallel (one running 5′→3′5' \rightarrow 3' and the other 3′→5′3' \rightarrow 5'), and DNA polymerases function strictly in the 5′→3′5' \rightarrow 3' direction, the cell utilizes two distinct modes of synthesis at the fork:

  • Leading Strand: Synthesized continuously in the direction of replication fork movement. A single priming event at the origin enables uninterrupted elongation by the replicative polymerase.

  • Lagging Strand: Synthesized discontinuously in the direction opposite to fork movement. Synthesis occurs as a series of short, backward-directed segments known as Okazaki fragments (1 to 3 kb1 \text{ to } 3 \text{ kb} long in bacteria; typically 100 to 200 nucleotides100 \text{ to } 200 \text{ nucleotides} long in eukaryotes).

Synthesis of leading and lagging strands of DNA
Continuous vs. Discontinuous Strand Processing
  1. As the replication fork advances and opens new template regions, RNA primase synthesizes short RNA primers periodically on the lagging strand template.

  2. Replicative DNA polymerase extends each primer to generate an Okazaki fragment.

  3. The RNA primer of the preceding Okazaki fragment is degraded (by Pol I in bacteria or RNase H / Fen1 in eukaryotes).

  4. DNA polymerase fills the resulting gap with deoxyribonucleotides.

  5. DNA Ligase catalyzes the formation of a phosphodiester bond to seal the single-stranded nick between adjacent Okazaki fragments, creating a continuous daughter strand.

Comparative Roles of Polymerases at the Replication Fork

Function / Component

Escherichia coli

Mammalian Cells

Leading Strand Polymerase

DNA Polymerase III

DNA Polymerase epsilon\text{epsilon} (Pol e\text{e})

Lagging Strand Initiation

Primase (RNA primer synthesis)

Primase / Pol alpha\text{alpha} complex

Lagging Strand Elongation

DNA Polymerase III

DNA Polymerase delta\text{delta} (Pol d\text{d})

Primer Removal & Gap Filling

DNA Polymerase I

RNase H, Fen1, Pol delta\text{delta}

Nonspecific Nick Sealing

DNA Ligase

DNA Ligase I

Roles of DNA polymerases in E. coli and mammalian cells

Replication Fork Accessory Proteins

High-speed, processive DNA synthesis requires specialized accessory protein complexes to maintain template stabilization, enzyme attachment, and unwinding.

Polymerase accessory proteins
Sliding Clamp and Clamp Loader Complex
  • Sliding-Clamp Protein (PCNA in Eukaryotes): Proliferating Cell Nuclear Antigen (PCNA) forms a homotrimeric ring structure that encircles double-stranded DNA. It directly binds DNA polymerase, anchoring it to the template without restricting sliding motion, thereby allowing processive incorporation of thousands of nucleotides per binding event.

  • Clamp-Loading Protein (RFC in Eukaryotes): Replication Factor C (RFC) uses ATP hydrolysis to open the sliding clamp ring and load it onto primer-template junctions.

Helicases and Single-Stranded DNA-Binding Proteins
  • Helicases: Unwind duplex parental DNA ahead of the replication fork by breaking hydrogen bonds between complementary bases in an ATP-dependent reaction. In eukaryotes, unwinding is carried out by the CMG complex (comprising Cdc45, MCM2-7, and GINS).

  • Single-Stranded DNA-Binding Proteins (SSBs / RPA in Eukaryotes): Replication Protein A (RPA) rapidly coats single-stranded template DNA exposed by helicase activity. This stabilizes the single strand, prevents premature re-annealing of the duplex, and eliminates hairpins or secondary structures that could block polymerase progress.

Action of helicases and single-stranded DNA-binding proteins

Topoisomerases: Resolving Topological Stress

As helicases unwind parental double-stranded DNA at high speeds, the upstream unwinding forces the intact double helix ahead of the fork to rotate. Unchecked, this creates positive supercoiling and torsional strain, which would stall fork progression.

Action of topoisomerases during DNA replication
Topoisomerase Mechanisms

Topoisomerases serve as molecular "swivels" by introducing transient cuts into the phosphodiester backbone:

  • Type I Topoisomerases: Induce a single-stranded break in DNA, allowing the broken strand to swivel freely around the intact strand to relieve torsional strain before re-ligating the backbone without requiring ATP.

  • Type II Topoisomerases: Induce double-stranded breaks, pass an intact double helix through the break, and re-seal both strands using ATP hydrolysis. In addition to relieving supercoiling, Type II topoisomerases are indispensable for decatenating (disentangling) intertwined daughter chromosomes during mitotic division.

Architecture of the Replisome Complex

The physical machinery operating at the replication fork is coordinated into a massive protein complex known as the replisome.

Model of prokaryotic replisomes at the replication fork
Prokaryotic vs. Eukaryotic Replisomes
  • Prokaryotic Replisome Architecture: Two molecules of DNA Polymerase III are physically linked via a central clamp loader and bound directly to the replicative helicase. To accommodate concurrent synthesis along both strands in the same physical direction, the lagging strand template loops 180o180^\text{o} through its polymerase core. As an Okazaki fragment finishes, the lagging strand polymerase releases the template and rebinds a newly synthesized primer.

  • Eukaryotic Replisome Architecture: Coordinated by scaffold proteins such as Ctf4, which bridges the CMG helicase complex to Pol epsilon\text{epsilon}, Pol alpha\text{alpha}/primase, and Pol delta\text{delta}. Eukaryotic replisomes additionally interface with histone chaperones to disassemble nucleosomes ahead of the fork and reassemble them onto both newly synthesized daughter duplexes.

Proofreading Mechanisms and DNA Polymerase Fidelity

DNA replication maintains extreme fidelity, exhibiting error rates as low as one misincorporated nucleotide per 109 to 101010^9 \text{ to } 10^{10} base pairs. This accuracy is achieved through a multi-step verification system:

  1. Conformational Selectivity: The active site of replicative DNA polymerases selectively accommodates correct Watson-Crick base pairing (A=T\text{A=T} and G⇌C\text{G}\rightleftharpoons\text{C}), reducing initial insertion errors to approximately 1 in 1051 \text{ in } 10^5.

  2. Exonucleolytic Proofreading (3′→5′3' \rightarrow 5'): Replicative polymerases possess an intrinsic 3′→5′3' \rightarrow 5' exonuclease catalytic domain.

Proofreading by DNA polymerase
The Proofreading Process
  • When an incorrect base (such as a mismatched guanine opposite an adenine) is incorporated, the abnormal geometry alters hydrogen bonding and stalls 5′ightarrow3′5' ightarrow 3' polymerase catalytic activity.

  • The mispaired 3′3' terminus shifts from the polymerase active site to the 3′ightarrow5′3' ightarrow 5' exonuclease active site.

  • The exonuclease excises the mismatched nucleotide.

  • The corrected 3′ OH3' \text{ OH} terminus translocates back to the polymerase catalytic site, and continuous 5′ightarrow3′5' ightarrow 3' synthesis resumes.

Origins of Replication and Initiation Dynamics

Replication initiates at defined genomic sequences called origins of replication (ori).

Bacterial Origins of Replication
  • Bacterial chromosomes (such as E. coli) possess a single, distinct origin of replication (oriC) spanning 245 base pairs245 \text{ base pairs}.

  • The initiator protein DnaA binds repeatedly within oriC, causing local DNA bending and destabilization of adjacent AT-rich regions.

  • DnaB helicase is loaded onto unwound strands to form two bidirectional replication forks.

  • The entire E. coli genome (4×106 bp4 \times 10^6 \text{ bp}) is duplicated in approximately 30 minutes30 \text{ minutes}.

Origin of replication in E. coli
Eukaryotic Origins of Replication

Due to significantly larger genome sizes (e.g., mammalian genomes contain 3×109 bp3 \times 10^9 \text{ bp}) and slower replication fork progression (hicksim10-foldhicksim 10\text{-fold} slower than bacteria), eukaryotic genomes require tens of thousands of replication origins operating simultaneously.

  • Mammalian cells utilize approximately 30,00030,000 origins spaced 50 to 300 kb50 \text{ to } 300 \text{ kb} apart.

  • Multiple origins allow the entire genome to be duplicated within a typical S-phase window of a few hours.

Replication origins in eukaryotic chromosomes
Step-by-Step Eukaryotic Origin Activation

In yeast, origins are defined by Autonomously Replicating Sequences (ARS) containing an ARS Consensus Sequence (ACS) and B element domains.

  1. ORC Binding: The six-subunit Origin Recognition Complex (ORC) binds to ACS and B1 regions, wrapping around dsDNA.

  2. MCM Recruitment: ORC recruits two inactive MCM2-7 helicase hexamers, forming a double-hexamer ring around dsDNA.

  3. CMG Helicase Assembly: During S-phase entry, recruitment of Cdc45 and GINS converts MCM2-7 into active CMG helicase complexes.

  4. Fork Formation: CMG helicases unwind local DNA, initiate leading strand synthesis, and form two active, divergence-moving replication forks.

A yeast ARS element

Telomeres and Telomerase Mechanism

Linear eukaryotic chromosomes face the "end-replication problem." Because lagging strand synthesis requires RNA primers, removal of the terminal primer at the extreme 5′5' end leaves a gap that cannot be filled by standard DNA polymerases, threatening progressive chromosome shortening with each round of division.

Action of telomerase
Telomerase Dynamics

To preserve chromosome ends, eukaryotes utilize telomerase, a specialized reverse transcriptase ribonucleoprotein complex containing an integral RNA molecule that serves as a template for extending telomeric simple-sequence repeats (e.g., 5′-TTGGGG-3’5'\text{-TTGGGG-3'} or 5′-TTAGGG-3’5'\text{-TTAGGG-3'}):

  1. Binding: Telomerase RNA base-pairs with the single-stranded 3′3' overhang of telomeric DNA.

  2. Extension: Using its internal RNA as a template, telomerase reverse-transcribes telomeric repeats, extending the 3′3' overhang of the template strand.

  3. Translocation: Telomerase translocates down the newly synthesized repeat to add additional repeat units.

  4. Lagging Strand Completion: Conventional DNA primase and Pol alpha\text{alpha}/Pol delta\text{delta} extend the complementary lagging strand using the newly lengthened template.

  5. Primer Removal: Removal of the terminal RNA primer leaves a stable 3′3' single-stranded overhang, maintaining chromosome integrity without loss of genetic information.

Pathways of DNA Damage and Structural Lesions

Genome integrity is continuously challenged by spontaneous chemical processes and environmental mutagens. Unrepaired DNA lesions block transcription and replication, introduce lethal mutations, or trigger malignant transformations.

Examples of DNA damage
Primary Forms of DNA Damage
  • Spontaneous Damage:

    • Deamination: Cytosine spontaneously hydrolyzes to form uracil; adenine and guanine can similarly undergo deamination.

    • Depurination: Hydrolysis of the glycosidic bond cleavage causes loss of purine bases (adenine or guanine), generating thousands of apurinic/apyrimidinic (AP) sites per cell daily.

  • Environmental Damage:

    • UV Radiation: Causes covalent cross-linking between adjacent pyrimidines, producing thymine dimers joined via a cyclobutane ring structure.

    • Chemical Carcinogens: React directly with bases to form bulky adducts (e.g., reactive compounds binding to guanine bases).

    • Ionizing Radiation / ROS: Induce single- and double-strand DNA backbone breaks.

Molecular Mechanisms of DNA Repair

Base-Excision Repair (BER)

Base-Excision Repair targets damaged or altered single bases (such as uracil formed by cytosine deamination).

Base-excision repair
  1. Base Cleavage: DNA Glycosylase recognizes and cleaves the glycosidic bond linking the damaged base to the sugar-phosphate backbone, leaving an AP site.

  2. Endonucleolytic Cleavage: AP Endonuclease cleaves the phosphodiester strand upstream of the AP site.

  3. Sugar Removal: Deoxyribosephosphodiesterase excises the remaining sugar-phosphate residue.

  4. Gap Filling and Ligation: DNA Polymerase inserts the correct base, and DNA Ligase seals the nick.

Nucleotide-Excision Repair (NER)

Nucleotide-Excision Repair eliminates major helix-distorting lesions, including UV-induced thymine dimers and bulky chemical adducts.

Nucleotide-excision repair of thymine dimers
  1. Lesion Recognition: Multiprotein complexes scan the genome and recognize helix distortion.

  2. Unwinding: Replicative helicases unwind the local region around the lesion.

  3. Dual Nuclease Cleavage: Endonucleases introduce cuts on both sides of the damaged site, excising an oligonucleotide containing the damaged bases.

  4. Resynthesis and Ligation: DNA Polymerase synthesizes replacement DNA using the intact complementary strand as a template, and DNA Ligase seals the backbone.

Translesion DNA Synthesis (TLS)

When DNA lesions are not repaired prior to replication, high-fidelity replicative polymerases stall at the damage site. To prevent catastrophic replication fork collapse, cells employ specialized translesion DNA polymerases (e.g., Pol V in E. coli).

Translesion DNA synthesis
  • Polymerase Switching: Stalled replicative polymerases temporarily swap out for translesion polymerases.

  • Bypass Synthesis: Translesion polymerases possess flexible active sites capable of synthesizing DNA directly across damaged template bases, albeit with lower fidelity.

  • Resumption: Once past the lesion, translesion polymerases dissociate, normal replicative polymerases resume synthesis, and the original lesion is subsequently processed by standard excision repair pathways.

DNA Rearrangements in Immune System Diversity

Vertebrates generate vast immune receptor diversity (>1011>10^{11} distinct antibody molecules) from a limited genomic coding sequence ($ hicksim 20,000 genes) through programmed somatic site-specific DNA recombination.\n\n### Immunoglobulin and Receptor Structural Organization\n\n* **Immunoglobulins (Antibodies):** Y-shaped heterotetramers produced by B lymphocytes, composed of two identical **Heavy chains** and two identical **Light chains** linked by interchain disulfide (\text{S-S}) bonds. Both heavy and light chains contain N-terminal **Variable (V) regions** (which recognize specific antigens) and C-terminal **Constant (C) regions**.\n\n![Structure of an immunoglobulin](https://assets.knowt.com/pdf-flow-prod/3d4be2eb-a63e-4a2a-9e96-561fb4edafc2-figures/17.png)\n\n### Light Chain VJ Recombination\n\nDuring B lymphocyte development, site-specific DNA recombination joins distinct gene segments:\n\n![Rearrangement of immunoglobulin light-chain genes](https://assets.knowt.com/pdf-flow-prod/3d4be2eb-a63e-4a2a-9e96-561fb4edafc2-figures/18.png)\n\n1. **Recombination:** Germ-line light-chain loci contain approximately 150 \text{ V}segments,segments,4 \text{ J (joining)}segments,andasinglesegments, and a single\text{C}region.Site−specificrecombinationdirectlyjoinsoneregion. Site-specific recombination directly joins one\text{V}segmenttoonesegment to one\text{J}segment(e.g.,segment (e.g.,\text{V3}toto\text{J3}), deleting intervening DNA.\n2. **Transcription:** Transcription generates a primary RNA transcript containing the recombined \text{VJ}segment,remainingsegment, remaining\text{J}segments,andthesegments, and the\text{C} region.\n3. **Splicing:** RNA splicing excises unjoined \text{J}sequencesandintrons,producingfunctionalmRNAencodingaspecificlightchain(sequences and introns, producing functional mRNA encoding a specific light chain (\text{V3-J3-C}).\n4. **Combinatorial Capacity:** 150 \text{ V} \times 4 \text{ J} = hicksim 600 unique light chain variable domains.\n\n### Heavy Chain VDJ Recombination\n\nHeavy chain loci contain three distinct sets of variable segments: **V**, **D (Diversity)**, and **J**.\n\n![Rearrangement of immunoglobulin heavy-chain genes](https://assets.knowt.com/pdf-flow-prod/3d4be2eb-a63e-4a2a-9e96-561fb4edafc2-figures/19.png)\n\n1. **Step 1 (D-to-J Recombination):** One \text{D}segment(fromsegment (from hicksim 12Dsegments)recombineswithoneD segments) recombines with one\text{J}segment(fromsegment (from4 \text{ J} segments).\n2. **Step 2 (V-to-DJ Recombination):** One \text{V}segment(fromsegment (from hicksim 150 \text{ V}segments)joinstherecombinedsegments) joins the recombined\text{DJ}complextoformafunctionalcomplex to form a functional\text{VDJ} sequence.\n3. **Processing:** The rearranged genomic segment is transcribed and spliced to the constant region (\text{C}) locus.\n4. **Combinatorial Capacity:** 150 \text{ V} \times 12 \text{ D} \times 4 \text{ J} = hicksim 7,200 distinct heavy chains.\n5. **Total Heavy/Light Combinatorial Diversity:** 600 \text{ light chains} \times 7,200 \text{ heavy chains} = hicksim 4.3 \times 10^6 distinct antibody combinations prior to junctional mutations.\n\n# T-Cell Receptor Rearrangements\n\nT lymphocytes utilize cell-surface **T-cell receptors (TCRs)** to recognize foreign antigens displayed on host cell surfaces.\n\n![Structure of a T-cell receptor](https://assets.knowt.com/pdf-flow-prod/3d4be2eb-a63e-4a2a-9e96-561fb4edafc2-figures/20.jpg)\n\n* **Structure:** TCRs are membrane-bound heterodimers composed of an \text{alpha}chainandachain and a\text{beta} chain linked by a disulfide bond. Each chain consists of an extracellular Variable region, Constant region, transmembrane domain, and short cytosolic tail.\n* **Recombination:**\n * \text{alpha} chain genes undergo site-specific **VJ recombination**.\n * \text{beta} chain genes undergo site-specific **VDJ recombination**.\n* Combined site-specific recombination and imprecise junctional joining yield overall TCR diversity comparable to immunoglobulins.\n\n# Somatic Hypermutation and Class Switch Recombination\n\nFollowing antigen exposure, B lymphocytes further optimize antibody specificity and functional capacity via secondary somatic modifications mediated by **Activation-Induced Cytidine Deaminase (AID)**.\n\n![Role of AID in somatic hypermutation and class switch recombination](https://assets.knowt.com/pdf-flow-prod/3d4be2eb-a63e-4a2a-9e96-561fb4edafc2-figures/21.png)\n\n### Mechanisms of AID Action\n\n* **Somatic Hypermutation (V Regions):** AID deaminates cytosine residues to uracil within rearranged \text{V}regionsofheavyandlightchains.Subsequentbase−excisionrepairprocessesrecruiterror−proneDNApolymerases,creatingsingle−basemutationsatanelevatedfrequencyofregions of heavy and light chains. Subsequent base-excision repair processes recruit error-prone DNA polymerases, creating single-base mutations at an elevated frequency of10^{-3}((\thicksim 1 \text{ million} times higher than baseline spontaneous rates). This process selects for B cells producing antibodies with exponentially higher antigen affinity (affinity maturation).\n* **Class Switch Recombination (S Regions):** AID deaminates cytosines within repetitive **Switch (S) regions** located upstream of heavy-chain constant region loci. Excision of uracils generates double-strand breaks that mediate recombination between different constant regions, switching antibody effector class (e.g., from \text{IgM}toto\text{IgG},,\text{IgA},or, or\text{IgE}) without changing antigen specificity.\n\n# Programmed vs. Pathological Gene Amplification\n\nGene amplification refers to a specialized genomic alteration wherein specific DNA regions selectively replicate, increasing gene copy number.\n\n![DNA amplification](https://assets.knowt.com/pdf-flow-prod/3d4be2eb-a63e-4a2a-9e96-561fb4edafc2-figures/22.png)\n\n### Developmental Amplification\n\nProgrammed gene amplification occurs naturally during specific developmental stages to meet extreme metabolic demands:\n\n* **Amphibian Oocytes:** Oocytes are roughly 1 \text{ million}timeslargerthansomaticcellsandrequiremassiveribosomesynthesis.RibosomalRNA(rRNA)genesundergoselectiveamplification,increasingfromseveralhundredcopiestoapproximatelytimes larger than somatic cells and require massive ribosome synthesis. Ribosomal RNA (rRNA) genes undergo selective amplification, increasing from several hundred copies to approximately1 \text{ million}copies(copies ( hicksim 2000\text{-fold}$$ amplification), providing sufficient templates for rapid ribosome production.

Pathological Amplification in Cancer

Unlike developmental amplification, gene amplification in tumor cells represents an uncontrolled, abnormal genetic defect:

  • Oncogene Amplification: Selective over-replication of growth-promoting oncogenes drives uninhibited cell proliferation and malignant progression.

  • Chemotherapeutic Resistance: Cancer cells frequently amplify genes encoding targets of therapeutic drugs or metabolic enzymes involved in nucleotide synthesis (e.g., dihydrofolate reductase amplification), rendering tumors resistant to chemotherapeutic regimens.