DNA Structure and Genome Organisation
Structure of DNA:
B-Form DNA:
Right-handed double helix.
There are ~10.4 base pairs per helical turn, with a helix diameter of ~2nm.
Bases are stacked almost flat, separated by ~0.34nm.
Strands are antiparallel (5’→3’ vs 3’→5’).
Grooves:
The major groove (~2.2nm wide) is rich in chemical information.
Base pairs present unique hydrogen bond donor and acceptor patterns.
E.g. a G-C pair exposes a different pattern in the major groove than an A-T pair would.
Transcription factors and DNA-binding proteins exploit this for sequence specificity.
The minor groove (~1.2nm wide) is narrower and has less chemical information.
It is recognised mainly by small molecules (e.g. some antibiotics).
Structural Importance Beyond Sequence:
DNA is not just a passive information store; its 3D structure guides how proteins recognise it.
Local flexibility and groove accessibility influence replication initiation, repair enzyme recognition and transcription factor binding.
Alternative DNA Forms:
A-Form:
Right handed, shorter and wider helix.
Appears under dehydrated conditions, or in DNA-RNA hybrids (like replication primers).
There are ~11bp per turn.
More compact, bases tilted relative to axis.
The tilting changes groove architecture, narrowing the major groove and widening the minor groove (as bases are usually perpendicular to the axis to maximise π-π overlap).
❗ The altered groove geometry of A-form helices facilitates primer synthesis because the shallow minor groove makes DNA–RNA hybrids accessible to RNA-binding proteins. This is essential for stable primer formation during replication, but replication overall still depends on B-form DNA as the main template structure.
Z-Form:
Left handed, zig-zag backbone.
Favoured in GC-rich or alternating purine-pyrimidine sequences.
Can form transiently during transcription and is thought to play roles in gene regulation and recombination.
Local Distortions and Bends:
A-tracks (stretches of adenine) cause intrinsic bending.
These bends are important for DNA packaging in nucleosomes and in forming promoter structures accessible to RNA polymerase.
Bending can be protein-induced (e.g. TATA-binding protein clamps DNA open at promoters).
Recognition is often a combination of direct readout (hydrogen bonds in major groove) and indirect readout (DNA shape and bendability).
Protein-DNA Interactions:
Charge Complementarity:
The DNA phosphate backbone has a negative charge.
DNA-binding proteins often have basic domains (arginine, lysine-rich).
E.g. histones have highly basic tails.
Sequence Specificity:
Proteins use hydrogen bonds, Van der Waals, and hydrophobic contacts.
E.g. helix-turn-helix proteins (like the lac repressor) insert an alpha helix into the major groove to read base pairs.
DNA Distortion by Proteins:
Binding is not always passive. Proteins can bend, kink, or unwind DNA to facilitate replication or transcription.
E.g. DNA polymerase bends DNA 90 degrees in its active site to separate the template strand for copying.
DNA Supercoiling and Topology:
Supercoiling:
When DNA is under- or over-wound, it compensates by forming superhelical twists (writhe).
Negative Supercoils - underwound DNA, promotes strand separation and is useful for replication and transcription.
Positive Supercoils - overwound DNA, forms ahead of replication forks and blocks progress if not removed.
Enzymes:
Topoisomerase I - makes single-strand nicks, and relaxes supercoils passively.
Topoisomerase II (DNA Gyrase in bacteria) - cuts both strands, passes another duplex through, and actively introduces negative supercoils using ATP.
Eukaryotes lack gyrase but rely on Topoisomerase I/II to manage torsional stress.

DNA Replication:
DNA Polymerases extend in the 5’ to 3’ direction, and a primer with a free 3’-OH is needed.
Initiation in E.Coli:
Origin (oriC) - 245 base pairs with AT-rich repeats (easier to unwind).
DnaA - initiator protein - binds to oriC and melts DNA locally with ATP hydrolysis.
DnaB (helicase) - loaded by DnaC, and unwinds the duplex.
Primase (DnaG) - synthesises short RNA primers.
Elongation:
Leading Strand - continuous synthesis in the direction of the replication fork.
Lagging Strand - discontinuous synthesis as Okazaki fragments.
DNA Pol III Holoenzyme - main polymerase, highly processive because of the β-clamp sliding ring.
SSBs (Single-Stranded Binding Proteins) - stabilises unwound DNA.
Topoisomerases - prevents supercoiling ahead of the replication fork.
Maturation of Okazaki Fragments:
DNA Pol I - removes RNA primers via 5’→3’ exonuclease, and fills gaps with DNA.
DNA Ligase - seals nicks (ATP- or NAD+-dependent).
Coordination:
The lagging strand loops so both polymerases move in the same physical direction.
This is known as the “trombone model”.
❗ Primase is needed because DNA polymerases cannot start de novo synthesis.
Replication Topology:
As the fork moves, it introduces positive supercoils ahead.
Without topoisomerases, DNA would quickly become overwound and stall.
Topoisomerase II solves this by transiently cutting both strands and passing another duplex through.
Genome Organisation:
Non-Coding DNA:
Microsatellites: short tandem repeats that are unstable; used in DNA fingerprinting.
Transposons:
Class I (Retrotransposons): “copy-paste” via RNA intermediates.
Class II (DNA Transposons): “cut-paste” directly via transposase enzyme.
Pseudogenes: inactivated genes, often from duplication followed by mutation.
Regulatory DNA: enhancers, silences, insulators. They are crucial for control.
Gene Families:
Multiple related genes cluster together (globins, keratins, olfactory receptors etc.).
Allows functional specialisation (e.g. foetal vs adult haemoglobin).
An open reading frame (ORF) is a stretch of DNA that can be translated into a protein without interruption by a stop codon.
It starts with a start codon (AUG in mRNA, ATG in DNA).
It continues in triplet codons that each specify an amino acid.
It ends only when a stop codon (UAA, UAG, or UGA) is reached.
Over time, pseudogenes accumulate point mutations (single-base changes) and indels (insertions or deletions).
These can disrupt the original ORF by:
Introducing a premature stop codon.
Causing a frameshift (changing the reading frame).
Altering essential splice sites or promoter regions.
Once the ORF is disrupted, the gene can’t produce a proper protein, even if it’s still transcribed.