Comprehensive study notes: DNA, chromosomes, and genome organization
Griffith, transformation, and the genetic material
Chromosome composition in eukaryotes: DNA plus various proteins; chromosomes are not DNA alone—protein components (histones and non-histone proteins) are essential for packaging and function.
Griffith’s experiment (British doctor Frederick Griffith): studying Streptococcus pneumoniae with two strains: S (smooth, encapsulated, virulent) and R (rough, non-encapsulated, non-virulent). S strain forms smooth colonies and is lethal to mice; R strain forms rough colonies and is not lethal.
Virulence factor: capsule (slimy surface layer) protects bacteria from host immune response; S strain carries capsule; R strain lacks capsule.
Key experimental results:
Live S strain injected into mice -> mouse dies.
Live R strain injected into mice -> mouse lives.
Heat-killed S strain injected -> mouse lives (heat-treated S is non-virulent but contains the genetic material).
Heat-killed S strain mixed with live R strain injected -> mouse dies; some cells recovered from the mouse are live S-type bacteria.
Griffith proposed a transformation hypothesis: something released from dead S cells is taken up by live R cells, converting them to S-type (transformation).
Introduction of the terms transformation and transfection:
Transformation: uptake of extracellular DNA by a bacterial cell (historical, broad usage).
Transfection: uptake of foreign DNA by animal cells (eukaryotic, often used in genetic engineering).
Later experiments to identify the transforming material:
Fractionation experiments (Avery, MacLeod, McCarty): isolated macromolecules from heat-killed S cells (RNA, DNA, protein, lipids, carbohydrates) and tested which fraction could transform live R cells.
Result: only the DNA fraction could transform R into S, indicating DNA as the genetic material; skeptics raised concerns about cross-contamination and purity of fractions.
Hershey–Chase experiment (to confirm DNA as genetic material in viruses): bacteriophage infecting E. coli; used radioisotopes:
DNA labeled with
(phosphorus label)Protein labeled with
(sulfur label)After infection, blender separates unattached phage from infected cells; centrifugation separates supernatant and pellet.
Observation: radioactivity for found in the supernatant (phage coats); found inside the bacterial cells (DNA inside).
Conclusion: DNA is the genetic material that enters the bacterial cell to direct infection; protein does not.
Nucleotides, polarity, and base pairing
Nucleotide building blocks and sequence options:
For proteins, number of building blocks (amino acids) is 20.
For nucleotides, the building blocks depend on context: typically 4 options (DNA: A, T, G, C; RNA: A, U, G, C), but some RNAs include thymine in unusual contexts, giving up to 5 building blocks in certain explanations.
General formula for the number of possible sequences of length n over an alphabet of size k:
For DNA (k = 4): ; for RNA (commonly 4 canonical nucleotides, sometimes 5 in broader discussions): in those contexts.
DNA polarity and ends:
Five prime (5') end is the phosphate end.
Three prime (3') end features the free hydroxyl group on the 3' carbon of the sugar.
DNA is double-stranded and antiparallel: one strand runs 5'→3', the complementary strand runs 3'→5'.
Sequences are read 5'→3' on each strand; when given a single-strand sequence, the complementary strand runs in the opposite direction.
Base pairing rules and chemical rationale:
Adenine (A) pairs with Thymine (T) in DNA via two hydrogen bonds; Guanine (G) pairs with Cytosine (C) via three hydrogen bonds.
Hydrogen-bond donor/acceptor geometry requires pairing of a purine (A or G) with a pyrimidine (T/U or C) to maintain proper spacing between the two strands.
In detail:
A–T (or A–U in RNA) forms two hydrogen bonds.
G–C forms three hydrogen bonds.
Why not A–C or G–T pairings? Donor/acceptor geometry and distance do not support stable hydrogen bonding at those positions.
Charged backbone and base composition:
DNA backbone is negatively charged due to phosphate groups.
Histone proteins and other packaging proteins in chromatin are positively charged to interact ionically with DNA phosphate groups.
GC content and melting point:
GC pairs form three hydrogen bonds vs AT pairs with two.
Higher GC content increases the DNA melting temperature (needs more energy to disrupt three H-bonds).
DNA base composition puzzle: if a DNA sample has 20% A, then 20% T; remaining 60% must be split equally between G and C:
DNA structure, discovery, and forms
Structure discovery and key contributors:
Watson, Crick, Wilkins, and Franklin contributed to elucidating DNA’s double-helix structure; Franklin’s X-ray crystallography data provided crucial distance information (not always credited equally for the Nobel Prize decisions).
Rosalind Franklin did the crystal structure work; Nobel Prize recognition historically did not include her because it is awarded posthumously and she had passed away before the prize.
Forms of DNA and major vs minor grooves:
B-DNA: the normal physiological form in most cells.
A-DNA and Z-DNA: alternative right-handed (A) and left-handed (Z) forms under particular conditions.
Major groove: site where most DNA-binding proteins (e.g., DNA polymerase, transcription factors) interact; typically broad and accessible.
Minor groove: narrower; accessibility to nucleases and some DNA-protein interactions is different; minor groove geometry affects stability and interaction with enzymes.
In RNA, the structure often resembles A-form due to RNA’s 2'-OH and other structural constraints; RNA tends to be less stable in some contexts than DNA due to the sugar and groove geometry.
Factors that govern DNA form transitions:
Dehydration tends to promote A-form from B-form.
High salt or alternating purine/pyrimidine sequences (e.g., GC repeats) can promote Z-form under certain conditions.
Transcription direction and DNA-RNA hybrids:
RNA synthesis occurs 5'→3' using the DNA template strand.
During transcription, a temporary RNA-DNA hybrid forms where RNA polymerase reads DNA; as transcription proceeds, the DNA strands separate and re-anneal behind the transcription bubble.
Chromatin, histones, and packaging in eukaryotes
Eukaryotic packaging proteins:
Histones are the main packaging proteins; most chromatin is organized around histone cores to form nucleosomes.
Histones are highly basic (enriched in positively charged amino acids like lysine and arginine) to interact with negatively charged DNA.
Non-histone proteins assist in replication and repair.
Histone structure and nucleosomes:
DNA wraps around a histone octamer; nucleosome = DNA + histone core.
The histone octamer consists of two each of H2A, H2B, H3, and H4.
Each histone has a flexible tail that protrudes from the nucleosome core; these tails can be post-translationally modified (e.g., methylation, acetylation), which influences chromatin structure and gene expression.
Beads-on-a-string model:
When histone proteins are removed, DNA appears as a beads-on-a-string structure, with each bead representing a nucleosome.
Chromatin vs chromosomes:
Chromatin = DNA + associated proteins (histones and non-histone proteins).
A chromosome is a single DNA molecule when unreplicated; after replication, each chromosome consists of two sister chromatids joined at the centromere (replicated chromosome).
Chromosome organization in the nucleus:
46 DNA molecules (in diploid human cells) are spatially organized within the nucleus, not randomly; chromosomes occupy distinct territories.
Nuclear envelope and lamina provide structural support; lamina and chromatin interactions help position chromosomes within the nucleus.
Nuclear lamina and progeria:
Nuclear lamina is a protein meshbeneath the inner nuclear envelope that supports nuclear shape and chromatin organization.
Progeria: accelerated aging syndrome due to defects in nuclear lamina proteins; cells show abnormal nuclear morphology and altered chromatin organization.
Telomeres, replication origins, centromeres, and chromosome features
Telomeres:
Repetitive DNA sequences at linear chromosome ends (species-specific; eg, human telomere repeats have characteristic GC-rich repeats, e.g., TTAGGG repeated thousands of times).
Telomeres protect chromosome ends from degradation and erroneous repair.
Telomerase (enzyme: telomerase reverse transcriptase, specifically) elongates telomeres during development and in germline stem cells; largely inactive in most somatic cells after development.
In cancer, telomerase activity is often reactivated, contributing to cellular immortality.
Replication origin:
The genomic region where DNA replication begins.
Centromere:
The constricted region where sister chromatids are held together and where spindle fibers attach during mitosis/meiosis.
The position of the centromere divides the chromosome into two arms, often denoted as p (short arm) and q (long arm).
Bacteria vs eukaryotes: telomeres and multiple chromosomes:
Bacterial chromosomes are typically circular and lack telomeres.
Eukaryotic chromosomes are linear and have telomeres; bacterial plasmids are extrachromosomal circular DNA capable of independent replication.
rRNA gene clusters and the nucleus:
Some chromosomes carry ribosomal RNA (rRNA) gene clusters; these regions can cluster within the nucleolus during interphase, giving rise to the nucleolus where rRNA transcription and ribosome assembly occur.
Chromosome structure specifics: replication, chromatin, and terminology
Chromosome terminology:
Chromatin: DNA plus associated proteins (histones and non-histone proteins) when not in condensed chromosome form.
Chromatid: one of the two identical DNA strands of a replicated chromosome; two chromatids together constitute a replicated chromosome joined at the centromere.
A single unreplicated chromosome has zero chromatids; replicated chromosome has two chromatids joined at the centromere.
Important caution: do not refer to a single unreplicated chromosome as having chromatin.
Polytene chromosomes (special case in Drosophila):
In Drosophila salivary glands, endomitosis occurs: multiple rounds of DNA replication without cytoplasmic division, yielding a polytene chromosome with many DNA copies aligned in parallel.
Chromocenter: fusion of centromeres from multiple chromosomes into a single nuclear structure.
Gene amplification in salivary glands supports rapid production of secreted proteins (e.g., pupation glue); active transcription visible as “puffs” where RNA polymerase is actively transcribing, appearing as swelling in the chromosome.
Nuclear organization and chromatin compartments:
Euchromatin: less condensed, higher gene expression during interphase; active transcription sites.
Heterochromatin: more condensed, lower gene expression; some regions are constitutively silent.
Chromosomal abnormalities, rearrangements, and inheritance patterns
Chromosomal rearrangements:
Deletion: a segment is removed, shortening the chromosome.
Inversion: a segment is flipped and reinserted, changing gene order but not the total amount of genetic material.
Translocation: a segment from one chromosome breaks off and attaches to another chromosome (often non-homologous); can disrupt gene regulation or create fusion genes.
Consequences: altered promoter context, fusion proteins, gain or loss of function, misregulated gene expression; particularly deleterious when involving non-homologous chromosomes or essential genes.
Aneuploidy vs polyploidy:
Aneuploidy: abnormal number of a single chromosome (e.g., trisomy 21 in Down syndrome); the rest of the genome remains at the usual chromosome number.
Polyploidy: multiple complete sets of chromosomes (e.g., triploidy, tetraploidy); entire genome is duplicated or multiplied.
Chromosome identification and karyotyping:
Giemsa staining (G-banding) reveals distinctive dark and light bands due to AT- and GC-rich regions, enabling chromosome identification.
The banding pattern is unique for each chromosome, allowing detection of aneuploidies and rearrangements.
Centromere position helps differentiate chromosome arms: p (short) and q (long) arms; centromere location ranges from metacentric (near center) to acrocentric (near end).
Chromosome features for identification:
Some chromosomes carry rRNA gene clusters; pink knobs in some diagrams denote these loci and the corresponding regions cluster in the nucleolus.
The position of the centromere and the banding pattern distinguish chromosomes with similar lengths.
Transcription and molecular determinants
Transcription direction:
RNA synthesis proceeds 5'→3' on the RNA; the DNA template strand is read 3'→5'.
The RNA polymerase creates an antisense/O template strand complementing the sense strand that carries the actual coding information.
Promoters, coding regions, and terminators:
Genes require a promoter (to initiate transcription), a coding region, and a terminator (to terminate transcription).
Relocating a gene during translocation can place it under a different promoter, altering expression levels or producing fusion proteins.
Practical and ethical notes from the lecture
Practical implications:
Understanding DNA structure, chromatin organization, and the regulation of gene expression is foundational for genetics, molecular biology, and cancer biology.
Telomere dynamics and telomerase activity are central to aging, development, and oncogenesis.
PCR, DNA sequencing, and genome editing rely on the physical and chemical properties of DNA discussed here (polarity, base pairing, chromatin accessibility).
Philosophical/ethical implications:
The history of discovering the genetic material (Griffith, Avery–MacLeod–McCarty, Hershey–Chase) illustrates how science progresses through iterative experiments, replication, and sometimes controversial interpretations.
The Rosalind Franklin story highlights debates about recognition in science and the role of experimental data in Nobel Prize decisions.
Quick reference formulas and facts
Base pairing rules:
Complementary base pairing logic:
Purine–pyrimidine pairing is required for proper spacing and hydrogen-bond geometry.
Nucleotide sequence math:
Number of possible sequences of length n over alphabet size k:
DNA content and genome organization:
Humans have 46 DNA molecules (diploid) in the nucleus; haploid number is 23.
1 chromosome = 1 DNA molecule (unreplicated) with its own centromere; replicated chromosome consists of two chromatids joined at the centromere.
Telomeres and telomerase:
Telomeric repeats are species-specific; telomerase extends telomeres during development and in germline cells.
Telomeres protect ends from degradation and prevent end-to-end fusions; dysfunction linked to aging and cancer.
Nucleosome structure:
Nucleosome = DNA wrapped around histone octamer (H2A, H2B, H3, H4; two copies of each).
Histone tails protrude and can be post-translationally modified to regulate chromatin compaction and gene expression.
Polytene chromosomes (special case):
Endomitosis in Drosophila salivary glands yields multiple DNA copies in a single cytoplasm; chromocenter forms from fused centromeres; transcription seen as puffs.
Key terms to differentiate:
Chromosome: single DNA molecule (unreplicated) with its own centromere.
Chromatin: DNA + proteins (histones and non-histone proteins).
Chromatid: one of two identical DNA strands in a replicated chromosome.
Chromocenter: fused centromeres in polytene chromosomes.
Euchromatin vs heterochromatin: levels of condensation and gene activity.