L3 Altogether
Chemical Bonds and Fundamentals of Protein Structure
Relationship Between Structure and Function:
Protein function is strictly determined by three-dimensional molecular structure ().
The multiplicity of functions carried out by proteins arises directly from the immense variety of shapes they adopt and the specific chemical properties conferred by these 3D conformations.
Relative Strength and Characteristics of Chemical Bonds:
Covalent Bonds:
Length: .
Strength in vacuum: ().
Strength in water: ().
Characteristics: Formed by shared electron pairs; provides stable primary linkage along the polypeptide backbone.
Ionic Bonds (Electrostatic Attractions):
Length: .
Strength in vacuum: ().
Strength in water: ().
Characteristics: Electrostatic attraction between two fully charged atoms (e.g., positively charged Lysine side chain and negatively charged Glutamic acid side chain). Significantly weakened in aqueous environments due to solvent shielding.
Hydrogen Bonds:
Length: .
Strength in vacuum: ().
Strength in water: ().
Characteristics: Weak non-covalent bond formed when a hydrogen atom (partially positively charged) bound to an electronegative atom is electrostatically attracted to another electronegative atom (typically Nitrogen or Oxygen).
van der Waals Attractions (per atom):
Length: .
Strength in vacuum: ().
Strength in water: ().
Characteristics: Transient electrostatic interactions arising from fluctuating electron densities in adjacent atoms.
Conversion Factors:
.
.
Peptide Bond and Polypeptide Backbone Mechanics:
Amino acids are linked covalently by amide linkages called peptide bonds formed via condensation reactions releasing water ().
Peptides are defined as chains under 50 amino acids long; longer polymers are defined as proteins.
Planar Unit Rigidness: The four atoms involved in every peptide bond () form a rigid planar unit. There is strictly no rotation around the central peptide bond due to partial double-bond character.
Chain Directionality: Written and synthesized from the amino terminus () on the left to the carboxyl terminus () on the right.
Example Tripeptide: Histidine-Cysteine-Valine ().
Torsional Flexibility (Phi and Psi Angles):
Rapid rotation occurs around the two single bonds connected to the central --carbon ().
(phi) angle: Angle of rotation around the bond.
(psi) angle: Angle of rotation around the bond.
Ramachandran Plot Angular Limitations:
Steric hindrance severely restricts the allowed combinations of and angles.
Right-handed --helix (): ,
--sheet: ,
Secondary, Tertiary, and Quaternary Protein Architecture
Secondary Structure Elements:
--Helix:
Maintained by hydrogen bonds between backbone N-H and C=O groups.
Hydrogen bonding occurs between every fourth amino acid.
Complete turn occurs every amino acids.
Proline is generally excluded from --helices because its rigid ring structure creates a conformational kink and lacks an amide H-donor.
Standard molecular atom colors: Nitrogen = blue, Carbon = gray, Oxygen = red, Hydrogen = white, --carbon = black.
--Sheet:
Formed by backbone-to-backbone hydrogen bonding between adjacent segments of polypeptide chains.
Distance between repeating unit segments in a --sheet is .
Antiparallel --Sheet:
Strands run in opposite N-to-C directional orientations.
Can form from either adjacent or non-adjacent regions of the peptide backbone.
Parallel --Sheet:
Strands run in the same N-to-C directional orientation.
Formed exclusively from non-adjacent regions of the backbone separated by long intervening loop stretches.
Protein Motifs:
A characteristic arrangement of secondary structure elements producing a specific tertiary structure, often containing invariant amino acids at critical positions.
Coiled-Coil Motif:
Amphipathic structure possessing both polar and nonpolar regions along its axis.
Consists of repeating heptad segments of 7 amino acids.
The 1st and 4th amino acids in the 7-amino-acid repeat are aliphatic and hydrophobic (e.g., Leucine, Isoleucine, Valine; excludes aromatic residues F, W, Y).
The remaining amino acids in the sequence are hydrophilic, driving two --helices to wrap around each other to sequester nonpolar side chains.
Tertiary Structure and Domains:
Protein Domain:
A compact, modular subunit and independently folded region of a single polypeptide chain that forms a stable tertiary structure.
Hydrophobic amino acids cluster in the interior core, while polar/charged amino acids reside on the surface interacting with water.
Examples of domain modules: Cytochrome , NAD-binding domain of lactate dehydrogenase, Variable domain of Immunoglobulin (Ig), Fibronectin type 3 module, Kringle module, Epidermal Growth Factor (EGF) module.
Protein Families:
Homologous proteins sharing similar amino acid sequences, highly preserved tertiary structures, and similar but non-identical functions.
Examples: Serine protease family (Chymotrypsin, Urokinase, Factor IX, Plasminogen).
Quaternary Structure and Macromolecular Assemblies:
Higher-order complexes composed of multiple polypeptide chains (subunits) bound together by noncovalent interactions.
Dimer of CAP Protein: Formed by noncovalent interaction between a single, identical binding site present on each monomer.
Tetramer of Neuraminidase: Formed by interactions between two nonidentical binding sites on each monomer.
Macromolecular Machines: Multimeric assemblies (e.g., the Ribosome) containing both protein and RNA components. Assembly relies on exact surface complementary affinity: matching molecular surfaces form numerous weak noncovalent bonds sufficient to withstand thermal jolting, whereas poor surface matches rapidly dissociate due to thermal motion.
Thermodynamics and Energetics of Protein Folding
Noncovalent Interactions Driving Folding:
Electrostatic attractions between full formal charges.
Hydrogen bonding between backbone-to-backbone, backbone-to-side chain, or side chain-to-side chain atoms.
van der Waals attractions between closely packed nonpolar atoms.
Hydrophobic forces driving sequestration of nonpolar groups.
The Oil Drop Model of Protein Folding:
Unfolded polypeptide chains expose nonpolar (hydrophobic) side chains to the surrounding aqueous medium.
To accommodate nonpolar groups, water molecules organize into highly ordered, rigid, cage-like clathrate structures around each hydrophobic side chain.
During folding, nonpolar side chains are sequestered together into an internal hydrophobic core.
Burying nonpolar groups releases ordered cage-like water molecules back into the bulk solvent, drastically increasing solvent entropy (\Delta S_{\text{water}} > 0).
Thermodynamics and the Second Law:
While the protein itself transitions from an unstructured, high-entropy state to a structured, low-entropy state (\Delta S_{\text{protein}} < 0), the overall net entropy change of the combined system (protein + solvent) is positive (\Delta S_{\text{system}} > 0).
Protein folding is spontaneous and does not defy the Second Law of Thermodynamics.
The folding pathway moves down an energy funnel toward a global minimum free energy state.
Intermediate folding states are termed molten globules—transient species undergoing rapid structural reorganization and local unfolding/refolding cycles to resolve steric strain.
Intrinsically Disordered Regions:
Approximately of all eukaryotic proteins possess unstructured regions (random coils) that do not adopt a single rigid tertiary conformation in isolation.
Protein Denaturation and In Vitro vs. In Vivo Refolding
Mechanism of Denaturation:
Denaturation is the complete disruption of noncovalent interactions and disulfide linkages stabilizing native tertiary structure.
Denaturing agents/conditions:
Heat: Increases kinetic energy, disrupting weak noncovalent hydrogen bonds and van der Waals interactions.
Extreme pH: Alters charge states of amino acid side chains, destroying electrostatic interactions.
Urea: Highly polar chemical reagent that dissolves proteins by forming competing hydrogen bonds with water and the polypeptide backbone.
--Mercaptoethanol: Powerful antioxidant and reducing agent that reduces covalent interchain and intrachain disulfide bonds () between cysteine residues back to free sulfhydryl groups ().
Anfinsen's Thermodynamic Hypothesis (In Vitro Refolding):
Removing denaturants (urea, --mercaptoethanol) from purified denatured protein solutions allows spontaneous renaturation into native conformation.
Proves that primary amino acid sequence contains all necessary informational instruction to dictate three-dimensional tertiary structure.
Kinetics: In Vitro vs. In Vivo:
In vitro refolding in dilute, controlled beaker conditions takes several hours due to sluggish progression through molten globule intermediates.
In vivo folding inside living cells occurs within minutes.
The cytoplasm is extremely crowded with macromolecules (organelles, cytoskeletal networks, ribosomes), creating high local concentrations that favor non-specific off-pathway aggregation of exposed hydrophobic surfaces on nascent proteins.
Molecular Chaperone Systems: Hsp70 and Hsp60 Mechanisms
Definition and Discovery of Molecular Chaperones:
Chaperones are specialized proteins that assist protein folding by preventing inappropriate interactions and keeping non-interacting surfaces apart.
Many chaperones are classified as Heat Shock Proteins (Hsps).
Historical Discovery: Identified accidentally during cell culture experiments when an incubator malfunctioned, elevating temperature from to . The subtle thermal stress denatured native proteins, triggering cells to upregulate a distinct set of novel protein bands visible on polyacrylamide gels.
Hsp70 Chaperone Family:
Action Mode: Acts predominantly co-translationally as the nascent polypeptide emerges from the ribosome exit tunnel.
Localization Variants: Cytosol (Hsp70), Mitochondria (mtHsp70), Endoplasmic Reticulum (BiP, Calnexin, Calreticulin), Bacteria (DnaK). Humans possess 13 distinct Hsp70 family members.
Domain Structure:
Nucleotide Binding Domain (NBD): Binds and hydrolyzes ATP.
Substrate Binding Domain (SBD): Contains a peptide-binding pocket and an --helical lid.
Step-by-Step Mechanism:
Open SBD conformation (ATP-bound) binds exposed hydrophobic patches on nascent polypeptide chains emerging from the ribosome.
Accessory co-chaperone proteins (e.g., DnaJ / Hsp40) stimulate ATP hydrolysis to ADP and inorganic phosphate ().
ATP hydrolysis induces a conformational change: the --helical lid clamps tightly shut down over the substrate binding domain, sequestering the hydrophobic peptide segment.
Nucleotide exchange factors facilitate exchange of ADP for fresh ATP, reverting SBD to the open conformation and releasing the polypeptide segment to allow localized folding.
Multiple rapid cycles of binding and release prevent premature misfolding and inter-chain aggregation during translation.
Hsp60 Chaperone Family (Chaperonins):
Action Mode: Acts strictly post-translationally on fully translated, misfolded, or partially folded protein substrates.
Localization Variants: Mitochondria (Hsp60), Vertebrate Cytosol (TCP1), Bacteria (GroEL/GroES complex).
Structural Architecture: Forms a double-ring cylindrical isolation chamber with a central hydrophobic cavity.
Step-by-Step Mechanism:
Substrate recognition: Hydrophobic residues exposed on misfolded proteins bind to the hydrophobic rim surrounding the opening of the unattached Hsp60 chamber.
ATP binding and Cap (GroES) attachment to the rim induce a conformational expansion of the chamber.
The inner walls of the expanded chamber shift to present a hydrophilic surface, stretching and unfolding the trapped misfolded protein inside an isolated environment.
ATP hydrolysis weakens the complex. Subsequent ATP binding ejects the cap and releases the correctly refolded protein into the cytosol.
Collaborative Chaperone Networks:
Unfolded proteins undergo stepwise transfer from Hsp70 to Hsp90 or Hsp60 chaperonins.
If a protein remains irretrievably misfolded despite chaperone catalysis, it is targeted for degradation via protease pathways.
Genomic Equivalence and Scale of the Human Genome
Central Dogma of Molecular Biology:
Directional flow of genetic information: .
Transcription: Copying genetic information within the same nucleotide alphabet.
Translation: Converting information from a nucleotide alphabet to an amino acid alphabet via the ribosome.
Genomic Equivalence:
Somatic cells within an organism (e.g., a neuron vs. a liver cell) contain the exact same total genomic DNA complement.
Differentiated cell types exhibit vast morphological and functional disparities solely due to differential gene expression (expressing the right protein, at the right time, in the right amount).
Requires tissue-specific, gene-specific, and time-specific regulation.
Quantitative Metrics and Scale of the Human Genome:
Haploid human genome size: base pairs ().
Diploid human somatic cell genome size: base pairs.
Protein-coding genes: approximately to genes encoding unique proteins.
Physical length: Extended end-to-end, human cellular DNA spans approximately .
Nuclear diameter: approximately .
Packaging compaction ratio: (equivalent to packing 24 miles of fine thread into a tennis ball).
Hypothetical scale model ( distance per base pair):
Genome total length = 2000 miles.
Protein-coding gene frequency = 1 gene every ().
Average gene length = (), with coding sequence accounting for .
Chromosome 22 metrics:
Contains nucleotide pairs per double-stranded DNA molecule.
Fully extended length = ().
Mitotic chromosome length = .
Chemical Structure of DNA:
Double helix distance per base pair = ; 10 base pairs per helical turn.
Acidic nature provided by negative charges on phosphodiester backbone linkage ().
Purines: Adenine (A) and Guanine (G). Heterocyclic aromatic compounds consisting of a 6-membered pyrimidine ring fused to a 5-membered imidazole ring.
Pyrimidines: Cytosine (C), Thymine (T), and Uracil (U). Heterocyclic aromatic 6-membered rings with nitrogen atoms at positions 1 and 3.
Chromatin Architecture: Histones and Nucleosome Structure
Chromatin Composition:
Fibrous nucleoprotein complex located in the nucleus, composed of approximately 1/3 DNA and 2/3 protein by mass (histones and non-histone proteins).
Chromosomes: Individual linear double-stranded DNA molecules complexed with proteins formed from condensed chromatin.
Nucleosomes - Fundamental Unit of Packaging:
Typical human diploid cell contains approximately 30 million nucleosomes.
Histone Core Octamer: Composed of two copies each of four core histones: H2A, H2B, H3, and H4 ( molecules of each histone type per cell).
Histone Fold Motif: Structural domain mediating octamer assembly. H2A and H2B form dimers via a "handshake" interaction; H3 and H4 form tetramers.
Amino Acid Composition: Core histones are small basic proteins exceptionally rich in positively charged Lysine (K) and Arginine (R) residues, which account for 1/5 (20%) of all histone core amino acids.
DNA-Histone Interactions:
DNA double helix wraps 1.7 tight turns around the histone octamer.
142 hydrogen bonds link DNA to the histone core; half of these form between basic amino acids and the phosphodiester backbone.
Sequence preferences: AA, TT, and TA dinucleotides are preferred where the minor groove faces inside toward the core; GC dinucleotides are preferred where the minor groove faces outside.
Histone Tails: Unstructured N-terminal tails of 11–37 amino acids extending outward from the nucleosome core, subject to post-translational covalent modifications.
Higher-Order Chromatin Condensation and Dynamic Remodeling
Linker Histone H1:
Binds to the exterior of each nucleosome core particle and pulls adjacent nucleosomes together to compact the fiber.
Hierarchical Stages of Chromatin Condensation:
DNA double helix: diameter.
"Beads-on-a-string" chromatin: diameter (converts DNA to 1/3 of its extended length).
Packed nucleosome chromatin fiber: diameter (organized via a zigzag model stabilized by H1 and N-terminal histone tail interactions).
Looped chromatin domains: Loops of to base pairs anchored to a scaffold ( diameter section, visualized in lampbrush chromosomes).
Entire mitotic chromosome: diameter (achieves an overall 10,000-fold reduction in length relative to extended DNA).
Nucleosome Dynamics and Spontaneous "Breathing":
Nucleosomal DNA is not statically bound; it unwraps spontaneously from the ends of the histone core at a rate of roughly 4 times per second.
Fully wrapped nucleosome state lasts ; unwrapped state lasts .
The fully open conformation occurs about of the time.
Impact on Transcription Factor Binding:
A transcription factor exhibits 20-fold lower affinity for a sequence site located near the end of a nucleosome relative to naked DNA.
Exhibits roughly 200-fold lower affinity if the binding site is located near the middle of a nucleosome.
Binding of one transcription regulator destabilizes the nucleosome, facilitating cooperative binding of a second regulator.
Epigenetic Histone Modifications and Gene Expression Regulation
Heterochromatin vs. Euchromatin:
Euchromatin: Less condensed regions of chromatin; accessible active sites of gene transcription.
Heterochromatin: Highly condensed, transcriptionally inactive regions (~10% of chromosome arms during interphase) that stain darkly.
Constitutive Heterochromatin: Permanently condensed across all developmental states (e.g., centromere and telomere regions).
Facultative Heterochromatin: Dynamically transitions between heterochromatin and euchromatin states depending on cell type or signaling.
Histone Tail Modifications:
H3K4me3(Trimethylation of Lysine 4 on Histone H3): Highly accessible, open chromatin; associated with active gene expression (ON); Abundance = .H3K9ac(Acetylation of Lysine 9 on Histone H3): Highly accessible, open chromatin; associated with active gene expression (ON); Abundance = .H3K9me3(Trimethylation of Lysine 9 on Histone H3): Heterochromatin marker (constitutive or facultative); associated with gene silencing (OFF); Abundance = .H3K27me3(Trimethylation of Lysine 27 on Histone H3): Facultative heterochromatin marker; associated with gene silencing (OFF); Abundance = .
Spreading and Barrier DNA Sequences:
Heterochromatin spreads via Reader-Writer complexes: a writer enzyme places a specific histone mark, which is bound by a reader protein that recruits adjacent writers.
Spreading continues along the chromosome until encountering a barrier DNA sequence or a reader-eraser protein that removes the heterochromatin-specific marks.
Spatial Spatial Repositioning During High Transcription:
Active gene expression causes chromosomal loops to decondense and physically reposition toward the center of the nucleus away from heterochromatic nuclear envelope domains (e.g., active thyroglobulin gene in expressing thyroid cells vs. inactive cells).
Seven Control Points of Eukaryotic Gene Expression:
Transcriptional Control: Primary, most efficient regulation point (saves metabolic energy, resources, and time).
RNA-Processing Control: Controlling splicing, 5' capping, and 3' polyadenylation.
RNA Transport and Localization Control: Regulating export from nucleus to cytosol.
mRNA Degradation Control: Regulating mRNA half-life.
Translational Control: Regulating ribosomal initiation rates.
Protein Activity Control: Post-translational modifications, phosphorylation, or allosteric control.
Protein Degradation Control: Ubiquitin-proteasome targeting.
Eukaryotic Promoter and Regulatory Architecture:
Promoter: Site where general transcription factors and RNA Polymerase II assemble to form the Preinitiation Complex / Transcription Initiation Complex (TIC).
Key Principle: Exposing the promoter region allows the transcription initiation complex to assemble automatically; gene regulation is achieved primarily by controlling promoter exposure.
Cis-Regulatory Sequences: Regulatory DNA binding elements located up to away from the transcription start site; bind specific regulatory proteins that dictate the rate of TIC assembly.
Spacer DNA: Non-coding intervening sequences providing structural flexibility for looping interactions.
Trans-Regulatory Sequences: Regulatory elements located on different chromosomes or distant unlinked genes that encode diffusible regulatory factors acting across genes.
Questions and Classroom Discussion
Question: How is cysteine classified for amino acid quizzes when hydrophobic/polar tables vary?
Answer: Cysteine can be categorized differently depending on the classification system used; however, in this context, cysteine is formally designated as nonpolar.
Question: Does the spontaneous folding of a protein into a single ordered tertiary structure defy the Second Law of Thermodynamics?
Answer: No. Although the polypeptide chain decreases in disorder (\Delta S_{\text{protein}} < 0), hydrophobic side chains buried in the interior release cage-like ordered water molecules into the surrounding environment, increasing net system disorder (\Delta S_{\text{water}} > 0 and \Delta S_{\text{system}} > 0).
Question: Are Heat Shock Proteins (Hsps) produced exclusively during thermal stress?
Answer: No. Hsps were discovered via thermal shock ( incubation shift) because heat denatures proteins and triggers dramatic transcriptional upregulation. However, baseline levels of Hsp70 and Hsp60 are constitutively expressed under normal physiological conditions as essential housekeeping chaperones.
Question: Do Hsp70 chaperones function at the co-translational or post-translational level?
Answer: Hsp70 chaperones operate primarily at the co-translational level as nascent chains emerge from the ribosome exit tunnel. Rare post-translational exceptions exist, but for functional classification, Hsp70 is co-translational while Hsp60 (chaperonins) acts post-translationally.
Question: Why does the Hsp60 chaperonin rim bind misfolded proteins but ignore correctly folded proteins?
Answer: The rim of the Hsp60 chamber features hydrophobic protein-binding sites. Misfolded proteins inappropriately display nonpolar hydrophobic amino acids on their outer surfaces, which selectively bind the rim. Correctly folded proteins bury nonpolar residues internally and present only hydrophilic residues externally.
Question: What is the key functional difference between a liver cell and a neuron given genomic equivalence?
Answer: Both cells contain identical genomic DNA (). Disparities arise entirely from differential gene expression—producing tissue-specific protein sets (e.g., cytoskeletal proteins driving neuronal extensions vs. enzymes driving hepatic bile acid synthesis).