DNA Structure and Genome Organisation

Structure of DNA:

B-Form DNA:

  • Right-handed double helix.

    • There are ~10.4 base pairs per helical turn, with a helix diameter of ~2nm.

  • Bases are stacked almost flat, separated by ~0.34nm.

  • Strands are antiparallel (5’→3’ vs 3’→5’).

Grooves:

  • The major groove (~2.2nm wide) is rich in chemical information.

    • Base pairs present unique hydrogen bond donor and acceptor patterns.

      • E.g. a G-C pair exposes a different pattern in the major groove than an A-T pair would.

        • Transcription factors and DNA-binding proteins exploit this for sequence specificity.

  • The minor groove (~1.2nm wide) is narrower and has less chemical information.

    • It is recognised mainly by small molecules (e.g. some antibiotics).

Structural Importance Beyond Sequence:

  • DNA is not just a passive information store; its 3D structure guides how proteins recognise it.

  • Local flexibility and groove accessibility influence replication initiation, repair enzyme recognition and transcription factor binding.


Alternative DNA Forms:


A-Form:

  • Right handed, shorter and wider helix.

    • Appears under dehydrated conditions, or in DNA-RNA hybrids (like replication primers).

    • There are ~11bp per turn.

    • More compact, bases tilted relative to axis.

      • The tilting changes groove architecture, narrowing the major groove and widening the minor groove (as bases are usually perpendicular to the axis to maximise π-π overlap).

        • The altered groove geometry of A-form helices facilitates primer synthesis because the shallow minor groove makes DNA–RNA hybrids accessible to RNA-binding proteins. This is essential for stable primer formation during replication, but replication overall still depends on B-form DNA as the main template structure.

Z-Form:

  • Left handed, zig-zag backbone.

    • Favoured in GC-rich or alternating purine-pyrimidine sequences.

    • Can form transiently during transcription and is thought to play roles in gene regulation and recombination.

Local Distortions and Bends:

  • A-tracks (stretches of adenine) cause intrinsic bending.

    • These bends are important for DNA packaging in nucleosomes and in forming promoter structures accessible to RNA polymerase.

      • Bending can be protein-induced (e.g. TATA-binding protein clamps DNA open at promoters).


Recognition is often a combination of direct readout (hydrogen bonds in major groove) and indirect readout (DNA shape and bendability).


Protein-DNA Interactions:


Charge Complementarity:

  • The DNA phosphate backbone has a negative charge.

  • DNA-binding proteins often have basic domains (arginine, lysine-rich).

    • E.g. histones have highly basic tails.


Sequence Specificity:

  • Proteins use hydrogen bonds, Van der Waals, and hydrophobic contacts.

    • E.g. helix-turn-helix proteins (like the lac repressor) insert an alpha helix into the major groove to read base pairs.


DNA Distortion by Proteins:

  • Binding is not always passive. Proteins can bend, kink, or unwind DNA to facilitate replication or transcription.

    • E.g. DNA polymerase bends DNA 90 degrees in its active site to separate the template strand for copying.


DNA Supercoiling and Topology:


Supercoiling:

  • When DNA is under- or over-wound, it compensates by forming superhelical twists (writhe).

    • Negative Supercoils - underwound DNA, promotes strand separation and is useful for replication and transcription.

    • Positive Supercoils - overwound DNA, forms ahead of replication forks and blocks progress if not removed.

Enzymes:

  • Topoisomerase I - makes single-strand nicks, and relaxes supercoils passively.

  • Topoisomerase II (DNA Gyrase in bacteria) - cuts both strands, passes another duplex through, and actively introduces negative supercoils using ATP.

  • Eukaryotes lack gyrase but rely on Topoisomerase I/II to manage torsional stress.


DNA Replication:


DNA Polymerases extend in the 5’ to 3’ direction, and a primer with a free 3’-OH is needed.

Initiation in E.Coli:

  • Origin (oriC) - 245 base pairs with AT-rich repeats (easier to unwind).

  • DnaA - initiator protein - binds to oriC and melts DNA locally with ATP hydrolysis.

  • DnaB (helicase) - loaded by DnaC, and unwinds the duplex.

  • Primase (DnaG) - synthesises short RNA primers.

Elongation:

  • Leading Strand - continuous synthesis in the direction of the replication fork.

  • Lagging Strand - discontinuous synthesis as Okazaki fragments.

  • DNA Pol III Holoenzyme - main polymerase, highly processive because of the β-clamp sliding ring.

  • SSBs (Single-Stranded Binding Proteins) - stabilises unwound DNA.

  • Topoisomerases - prevents supercoiling ahead of the replication fork.

Maturation of Okazaki Fragments:

  • DNA Pol I - removes RNA primers via 5’→3’ exonuclease, and fills gaps with DNA.

  • DNA Ligase - seals nicks (ATP- or NAD+-dependent).

Coordination:

  • The lagging strand loops so both polymerases move in the same physical direction.

    • This is known as the “trombone model”.

Primase is needed because DNA polymerases cannot start de novo synthesis.


Replication Topology:

  • As the fork moves, it introduces positive supercoils ahead.

    • Without topoisomerases, DNA would quickly become overwound and stall.

      • Topoisomerase II solves this by transiently cutting both strands and passing another duplex through.


Genome Organisation:


Non-Coding DNA:

  • Microsatellites: short tandem repeats that are unstable; used in DNA fingerprinting.

  • Transposons:

    • Class I (Retrotransposons): “copy-paste” via RNA intermediates.

    • Class II (DNA Transposons): “cut-paste” directly via transposase enzyme.

  • Pseudogenes: inactivated genes, often from duplication followed by mutation.

  • Regulatory DNA: enhancers, silences, insulators. They are crucial for control.

Gene Families:

  • Multiple related genes cluster together (globins, keratins, olfactory receptors etc.).

  • Allows functional specialisation (e.g. foetal vs adult haemoglobin).

An open reading frame (ORF) is a stretch of DNA that can be translated into a protein without interruption by a stop codon.

  • It starts with a start codon (AUG in mRNA, ATG in DNA).

  • It continues in triplet codons that each specify an amino acid.

  • It ends only when a stop codon (UAA, UAG, or UGA) is reached.

Over time, pseudogenes accumulate point mutations (single-base changes) and indels (insertions or deletions).

  • These can disrupt the original ORF by:

    • Introducing a premature stop codon.

    • Causing a frameshift (changing the reading frame).

    • Altering essential splice sites or promoter regions.

  • Once the ORF is disrupted, the gene can’t produce a proper protein, even if it’s still transcribed.