Chapter 3: Amino Acids, Peptides, and Proteins

Amino Acids: Fundamental Building Blocks of Proteins

  • Proteins are linear heteropolymers composed of α\alpha-amino acids linked together in a specific sequence.

  • Amino acids possess chemical properties that make them uniquely suited to fulfill diverse biological functions:

    • Capacity to polymerize into long polypeptide chains.

    • Useful acid-base properties across physiological and extreme pHpH ranges.

    • Varied physical properties, including solubility, polarity, and steric volume.

    • Diverse chemical functionality provided by distinct side chains (RR groups).

  • Definition of an Amino Acid Residue:

    • According to IUPAC definition, when two or more amino acids combine to form a peptide, the elements of water (H2OH_2O) are removed.

    • What remains of each amino acid unit is termed an amino-acid residue.

    • Residues in a peptide chain lack a hydrogen atom of the amino group (−NH−CHR−COOH-NH-CHR-COOH), the hydroxyl moiety of the carboxyl group (NH2−CHR−CO−NH_2-CHR-CO-), or both (−NH−CHR−CO−-NH-CHR-CO-).

  • Condensation vs. Hydrolysis Reactions:

    • Formation of peptide bonds occurs via a condensation reaction where water is eliminated.

    • Cleavage of peptide bonds occurs via a hydrolysis reaction, which represents the exact reverse of condensation.

  • General Structure of α\alpha-Amino Acids:

    • Every amino acid contains a central carbon known as the α\alpha-carbon (CaC_a).

    • Attached covalently to the α\alpha-carbon are four groups: a primary amino group (or secondary in proline), a carboxylic acid group, a hydrogen atom, and a distinctive side chain (RR group).

    • There are 2020 standard amino acids commonly found in all ribosomal proteins, each defined by its unique RR group.

Nomenclature, Carbon Numbering, and Stereochemistry of Amino Acids

  • Carbon Nomenclature Systems:

    • Standard Numerical System: Carbon atoms are numbered starting from the carboxylic acid carbon as carbon-1 (C−1C-1), the α\alpha-carbon as C−2C-2, followed sequentially along the side chain as C−3C-3, C−4C-4, C−5C-5, C−6C-6, etc.

    • Greek Letter Designation System: Carbon atoms are designated starting from the carbon adjacent to the carboxyl group, designated as α\alpha (CaC_a), followed down the side chain by β\beta, γ\gamma, δ\delta, ϵ\epsilon, etc.

    • Example (Lysine): Carbon-1 is the carboxyl carbon (−COO−-COO^-), Carbon-2 (CaC_a) carries the primary α\alpha-amino group (−NH3+-NH_3^+), Carbon-3 (β\beta) is a −CH2−-CH_2- group, Carbon-4 (γ\gamma) is a −CH2−-CH_2- group, Carbon-5 (δ\delta) is a −CH2−-CH_2- group, and Carbon-6 (ϵ\epsilon) carries the side-chain amino group (−NH3+-NH_3^+).

Lysine carbon numbering and Greek letter designations
  • The D,LD,L System of Stereochemistry:

    • The D,LD,L stereochemical nomenclature system relates the absolute configuration of a chiral molecule to the reference molecule glyceraldehyde.

    • LL-glyceraldehyde rotates plane-polarized light in the levorotatory (left) direction.

    • DD-glyceraldehyde rotates plane-polarized light in the dextrorotatory (right) direction.

    • The DD and LL designations specify absolute spatial configuration and do NOT consistently indicate the physical direction of optical rotation of plane-polarized light; actual optical rotation must be measured experimentally.

    • All α\alpha-amino acids found naturally in proteins possess the LL stereochemistry.

Stereoisomers of L-alanine and D-alanine in ball-and-stick, wedge, and Fischer projection forms
  • The R/SR/S Cahn-Ingold-Prelog System:

    • When applying the absolute R/SR/S configuration rules to the chiral standard amino acids, 1818 of the 1919 chiral amino acids have the (S)(S) configuration.

    • Glycine is achiral because its α\alpha-carbon is bound to two identical hydrogen atoms.

    • Cysteine is the sole exception among chiral amino acids; LL-cysteine has the (R)(R) configuration because the sulfur-containing side chain (−CH2SH-CH_2SH) takes priority over the carboxyl group (−COOH-COOH) under Cahn-Ingold-Prelog priority rules.

  • Amino Acids with Multiple Chiral Centers:

    • Two standard amino acids possess a second chiral center in their side chains: Isoleucine and Threonine.

Classification and Chemical Properties of Standard Amino Acid Side Chains

  • Nonpolar, Aliphatic Amino Acids:

    • Characteristics: Lack polar functional groups; drive protein stability and folding through hydrophobic interactions; predominantly buried within the interior of folded globular proteins away from water.

    • Glycine (GlyGly, GG): The only achiral amino acid; possesses the smallest side chain (a single hydrogen atom, −H-H); does not contribute hydrophobic driving force; provides high conformational flexibility to polypeptide backbones; derived from Greek glykos ("sweet") due to its taste; commonly utilized as a buffer component.

    • Alanine (AlaAla, AA): Possesses a simple methyl side chain (−CH3-CH_3).

    • Valine (ValVal, VV): Possesses a branched aliphatic side chain (−CH(CH3)2-CH(CH_3)_2).

    • Leucine (LeuLeu, LL): Possesses a branched aliphatic side chain (−CH2CH(CH3)2-CH_2CH(CH_3)_2).

    • Isoleucine (IleIle, II): Possesses a branched aliphatic side chain (−CH(CH3)CH2CH3-CH(CH_3)CH_2CH_3) containing a second chiral center.

    • Methionine (MetMet, MM): Contains a nonpolar thioether side chain (−CH2−CH2−S−CH3-CH_2-CH_2-S-CH_3); hydrophobic in nature; unlike a thiol group, a thioether cannot undergo oxidation to form covalent disulfide bonds.

    • Proline (ProPro, PP): Features a distinctive cyclic structure where the aliphatic side chain forms a covalent ring with the backbone α\alpha-amino group, forming a secondary amino (imino) group; holds polypeptide backbones in a rigid conformation that restricts structural flexibility; lacks a hydrogen atom on the backbone nitrogen atom when incorporated into peptide bonds, limiting its hydrogen-bonding capacity.

  • Aromatic R Groups:

    • Characteristics: Relatively nonpolar; participate in hydrophobic interactions; engage in π−π\pi-\pi stacking interactions with other aromatic rings.

    • Phenylalanine (PhePhe, FF): Nonpolar aromatic side chain containing a phenyl ring.

    • Tyrosine (TyrTyr, YY): Contains a phenolic hydroxyl group (−OH-OH); more polar than phenylalanine; can form hydrogen bonds; features an ionizable side chain with pKa=10.5pK_a = 10.5; named from Greek tyros ("cheese") after first being isolated from cheese.

    • Tryptophan (TrpTrp, WW): Bulky, nonpolar indole-like side chain; slightly more polar than phenylalanine.

    • Ultraviolet Light Absorption:

    • Aromatic side chains, predominantly Tryptophan and Tyrosine (with minor contribution from Phenylalanine), absorb UV light in the wavelength range of 270 nm270\,nm to 280 nm280\,nm.

    • Absorption spectra plots display Absorbance (AA) on the y-axis versus Wavelength (in nmnm) on the x-axis.

    • Protein concentration (cc) in solution is determined via UV-visible spectrophotometry using Beer's Law:

A=ε⋅c⋅lA = \varepsilon \cdot c \cdot l

where AA is the measured absorbance, ε\varepsilon is the molar extinction coefficient, cc is molar concentration, and ll is the optical path length.

  • Polar Uncharged R Groups:

    • Characteristics: Soluble in water due to functional groups that readily form hydrogen bonds with aqueous solvent.

    • Serine (SerSer, SS): Contains a primary hydroxyl side chain (−CH2OH-CH_2OH).

    • Threonine (ThrThr, TT): Contains a secondary hydroxyl side chain (−CH(OH)CH3-CH(OH)CH_3) with a second chiral center.

    • Cysteine (CysCys, CC): Contains a sulfhydryl (thiol) group (−CH2SH-CH_2SH).

    • The sulfhydryl group is less polar than a hydroxyl group; acts as a weak acid (pKa=8.1pK_a = 8.1); can form dipole-dipole interactions with oxygen or nitrogen.

    • Oxidation of two cysteine sulfhydryl groups yields a covalent disulfide bond (−S−S−-S-S-), forming the dimeric amino acid cystine.

    • Disulfide cross-links stabilize tertiary and quaternary protein structures across different chains or within a single chain.

    • Asparagine (AsnAsn, NN) and Glutamine (GlnGln, QQ): Amide derivatives of aspartate and glutamate.

    • Glutamine possesses one additional methylene group (−CH2−-CH_2-) compared to asparagine.

    • Single-letter codes follow alphabetical order by side-chain length: NN for Asn, QQ for Gln.

  • Positively Charged (Basic) R Groups:

    • Characteristics: Possess side chains with net positive charges at physiological pH≈7.0pH\approx 7.0.

    • Lysine (LysLys, KK): Features a primary amino group at the ϵ\epsilon-position (pKa≈10.5pK_a \approx 10.5).

    • Arginine (ArgArg, RR): Contains a positively charged guanidino group (pKa≈12.5pK_a \approx 12.5); one-letter abbreviation is RR.

    • Histidine (HisHis, HH): Contains an ionizable imidazole ring with pKa≈6.0pK_a \approx 6.0.

    • Histidine is the only standard amino acid with a side chain pKapK_a near physiological pHpH, allowing it to shift between protonated (charged) and unprotonated (neutral) states to catalyze enzymatic acid-base reactions.

Histidine side chain protonation-deprotonation equilibrium
  • Negatively Charged (Acidic) R Groups:

    • Characteristics: Possess side chains with net negative carboxylic acid charges at physiological pH≈7.0pH\approx 7.0.

    • Aspartic Acid / Aspartate (AspAsp, DD): Carboxyl side chain (−CH2COO−-CH_2COO^-).

    • Glutamic Acid / Glutamate (GluGlu, EE): Carboxyl side chain containing one additional methylene group (−CH2CH2COO−-CH_2CH_2COO^-).

    • Single-letter codes follow alphabetical order by side-chain length: DD for Asp, EE for Glu.

Non-Standard, Modified, and Biological Amino Acids

  • Nutritional Classification and Auxotrophy:

    • Auxotrophy: The inability of an organism to synthesize a specific organic compound necessary for its growth and development.

    • Humans are auxotrophic for essential amino acids and must acquire them through diet.

    • Prototrophy: The ability of organisms (e.g., plants) to synthesize all required organic compounds from inorganic precursors.

    • Complete (Whole) Proteins: Animal protein sources (milk, eggs, dairy) supply all essential amino acids in proportions matching human nutritional needs.

    • Complementary Plant Proteins: Most plant items lack adequate proportions of specific essential amino acids (e.g., Rice is deficient in Lysine; Beans are deficient in Methionine; ingested together, they form a complete protein profile).

  • Cotranslationally Incorporated Non-Standard Amino Acids:

    • Selenocysteine (SecSec, UU):

    • Considered the "21st amino acid"; cotranslationally incorporated during protein synthesis across all three domains of life; present in at least 2525 human proteins.

    • Contains selenium (−SeH-SeH group) instead of sulfur.

    • Side chain pKa=5.2pK_a = 5.2, making it 99%99\% deprotonated at neutral physiological pHpH

    • Incorporation mechanism in prokaryotes: Context-dependent suppression of the opal (UGAUGA) stop codon.

Comparison of Serine, Cysteine, and Selenocysteine structures and pKa values
  • Pyrrolysine (PylPyl, OO):

    • First identified in methanogenic archaea; encoded by the amber (UAGUAG) stop codon.

Pyrrolysine structure
  • Post-Translationally Modified Amino Acids:

    • Formed via enzymatic modification after the polypeptide chain has been synthesized on the ribosome:

    • 44-hydroxyproline: Found in plant cell walls and collagen structural fibers.

    • 55-hydroxylysine: Found in collagen structural fibers.

    • 6−N6-N-methyllysine: Found in muscle contractile protein myosin.

    • γ\gamma-carboxyglutamate: Found in blood-clotting protein prothrombin and other calcium-binding (Ca2+Ca^{2+}) proteins.

    • Desmosine: A complex derivative found in the fibrous protein elastin.

    • Reversible / Transient Modifications for Protein Regulation:

    • Phosphorylation: Most common regulatory modification; targets hydroxyl groups of Serine, Threonine, and Tyrosine.

    • Acetylation: 6−N6-N-acetyllysine found in histone proteins; regulates gene expression and chromatin structure at an epigenetic level.

    • Other transient groups: Methyl, adenylyl, and ADP-ribosyl groups.

  • Non-Protein Amino Acid Intermediates:

    • Ornithine and Citrulline: Metabolic intermediates in the biosynthesis of arginine and the urea cycle; not incorporated into proteins during translation.

Structures of Ornithine and Citrulline

Acid-Base Chemistry and Ionization Behavior of Amino Acids

  • Polyprotic Acid Behavior:

    • Amino acids act as weak polyprotic acids containing at least two ionizable protons, each with a distinct pKapK_a value.

    • α\alpha-Carboxylic Acid Dissociation:

COOH⇌COO−+H+COOH \rightleftharpoons COO^- + H^+

  • α\alpha-Amino Group Dissociation:

NH3+⇌NH2+H+NH_3^+ \rightleftharpoons NH_2 + H^+

  • At strongly acidic pHpH: Amino acids exist in a fully protonated cationic form (net charge=+1net\ charge = +1 or greater).

  • At strongly basic pHpH: Amino acids exist in a fully deprotonated anionic form (net charge=−1net\ charge = -1 or lower).

  • At intermediate pHpH: Amino acids exist as dipolar zwitterions carrying both a positive and a negative charge (net charge=0net\ charge = 0).

Sequential deprotonation of amino acid from net charge +1 to 0 to -1
  • Amphoteric Nature:

    • Zwitterions can donate a proton (acting as an acid) or accept a proton (acting as a base); compounds with this dual capability are termed amphoteric or amphiolytes.

Zwitterion form and its amphoteric action as acid and base
  • Chemical Environment Effects on pKapK_a:

    • The pKapK_a of an α\alpha-carboxyl group is ≈2.3\approx 2.3, significantly lower than carboxylic acids like acetic acid (pKa=4.8pK_a = 4.8).

    • Explanation: The positively charged α\alpha-amino group (−NH3+-NH_3^+) exerts an electron-withdrawing inductive effect on the carboxyl group, stabilizing the negatively charged carboxylate form and lowering its pKapK_a

    • The pKapK_a of an α\alpha-amino group is ≈9.6\approx 9.6, slightly lower than simple primary amines like methylamine (pKa=10.6pK_a = 10.6).

    • Explanation: Electronegative oxygen atoms in the carboxyl group withdraw electron density from the amino group, facilitating proton dissociation and lowering its pKapK_a

    • Ionization State Thresholds:

    • Above pH=3.5pH = 3.5, α\alpha-carboxyl groups are entirely in their carboxylate (−COO−-COO^-) form.

    • Below pH=8.0pH = 8.0, α\alpha-amino groups are entirely in their protonated ammonium (−NH3+-NH_3^+) form.

Effect of chemical environment on pKa values of carboxyl and amino groups

Titration Curves and Isoelectric Point Determination

  • Henderson-Hasselbalch Equation:

    • Quantitative relationship between pHpH, pKapK_a, and conjugate base/acid ratio:

pH=pKa+log⁡([A−][AH])pH = pK_a + \log\left(\frac{[A^-]}{[AH]}\right)

  • Isoelectric Point (pIpI):

    • The characteristic pHpH at which the net electric charge of an amino acid or peptide is exactly zero.

    • Calculation for Diprotic Amino Acids (e.g., Glycine):

pI=12(pK1+pK2)pI = \frac{1}{2}(pK_1 + pK_2)

  • For Glycine:

pK1(α-COOH)=2.34pK_1 (\alpha\text{-COOH}) = 2.34

pK2(α-NH3+)=9.60pK_2 (\alpha\text{-NH}_3^+) = 9.60

pI=12(2.34+9.60)=5.97pI = \frac{1}{2}(2.34 + 9.60) = 5.97

  • Physical Properties at the Isoelectric Point (pH=pIpH = pI):

    • Net electric charge equals zero.

    • Amino acid exhibits minimum solubility in aqueous solvent.

    • Molecule experiences zero net directional migration within an applied electric field.

  • Net Charge Dependencies:

    • When pH>pIpH > pI, the molecule carries a net negative charge.

    • When pH<pIpH < pI, the molecule carries a net positive charge.

Titration curve of glycine showing pK1, pK2, and isoelectric point pI

Formation, Structure, and Nomenclature of Peptides

  • Peptide Bond Formation:

    • Formed via a condensation reaction (nucleophilic acyl substitution).

    • The unprotonated α\alpha-amino group of one amino acid acts as a nucleophile, attacking the α\alpha-carboxyl carbon of another amino acid to displace a hydroxyl group, releasing H2OH_2O

    • While amino groups are effective nucleophiles, hydroxyl groups are poor leaving groups; thus, at physiological pHpH, this condensation reaction does not proceed spontaneously to a significant extent without energy input and enzymatic catalysis.

Formation of a peptide bond by condensation reaction
  • Size Classifications:

    • Dipeptide: Contains 22 amino acid residues.

    • Tripeptide: Contains 33 amino acid residues.

    • Oligopeptide: General term for short peptide chains containing 44 to 1010 amino acid residues.

    • Polypeptide: Chain containing greater than 1010 amino acid residues.

    • Protein: Polypeptide chain containing hundreds or thousands of amino acid residues with molecular mass (MwM_w) exceeding 10 kDa10\,kDa

  • Naming and Conventions for Peptides:

    • Sequence Directionality: Written and read from the N-terminus to the C-terminus (left to right).

    • N-Terminus (Amino Terminus): Residue with the free α\alpha-amino group (−NH3+-NH_3^+).

    • C-Terminus (Carboxyl Terminus): Residue with the free α\alpha-carboxylate group (−COO−-COO^-).

    • Full Systematic Nomenclature: Systematic names link constituent amino acids sequentially from N-terminus to C-terminus, changing suffixes -ine or -ate to -yl, leaving the C-terminal residue name unchanged.

    • Three-Letter Notation: Linked with dashes (e.g., Gly−Ala−LeuGly-Ala-Leu).

    • One-Letter Notation: Linked continuously without dashes (e.g., GALGAL).

Biologically Active Peptides and Protein Complexity

  • Peptide Ionization Behavior:

    • Peptides carry only one free ionizable α\alpha-amino group at the N-terminus and one free ionizable α\alpha-carboxyl group at the C-terminus.

    • Total charge is determined by terminal groups and the number of ionizable RR groups.

    • Side chain pKapK_a values buried within folded protein environments differ from values of free amino acids.

  • Examples of Biologically Active Peptides:

    • Commercial Sweetener: Aspartame (L-Aspartyl-L-phenylalanine methyl esterL\text{-Aspartyl-}L\text{-phenylalanine methyl ester}).

Chemical structure of aspartame
  • Vertebrate Hormones & Pheromones: Insulin, Oxytocin.

  • Neuropeptides: Substance P (mediates pain perception signal transmission).

  • Antimicrobial Peptides:

    • Polymyxin B: Active against Gram-negative bacteria (produced by Bacillus polymyxa).

    • Bacitracin: Active against Gram-positive bacteria (produced by Bacillus subtilis).

    • Dermcidin-1L: Human antimicrobial defense peptide (PDB ID: 2KSG).

  • Potent Toxins: Amanitin (death cap mushrooms), Conotoxin (cone snails), Chlorotoxin (scorpions).

    • Protein Composition and Structural Components:

  • Polypeptides consist of covalently linked α\alpha-amino acid residues, sometimes bound to non-amino acid components:

    • Cofactors: Functional non-amino acid components, such as metal ions (Fe2+Fe^{2+}, Mg2+Mg^{2+}, Zn2+Zn^{2+}) or small organic molecules.

    • Coenzymes: Complex organic cofactors (e.g., NAD+NAD^+ in lactate dehydrogenase).

    • Prosthetic Groups: Covalently or tightly bound cofactors essential for activity (e.g., heme group in myoglobin).

  • Multisubunit Proteins: Contain two or more polypeptide chains associated noncovalently.

  • Oligomeric Proteins: Multisubunit proteins in which at least two polypeptide chains are identical.

  • Protomers: Identical repeating structural units consisting of one or more polypeptide chains in an oligomeric protein.

    • Example: Hemoglobin contains two α\alpha chains and two β\beta chains (α2β2\alpha_2\beta_2); can be defined as a tetramer of four subunits or a dimer of αβ\alpha\beta protomers.

  • Conjugated Proteins: Proteins containing permanently associated non-amino acid chemical groups.

    • Estimation of Amino Acid Residue Count:

  • The average molecular weight (MwM_w or MrM_r) of an amino acid residue in a protein is estimated at 110 Da110\,Da (110 g/mol110\,g/mol).

    • Derived from average amino acid mass (128 Da128\,Da) weighted for abundance of smaller amino acids, minus 18 Da18\,Da for water lost during condensation.

Estimated Number of Residues=Protein Mw (in Da)110 Da/residue\text{Estimated Number of Residues} = \frac{\text{Protein } M_w \text{ (in Da)}}{110\,Da/\text{residue}}

Methods for Protein Purification and Separation

  • Physicochemical Properties Exploited in Purification:

    • Size, charge, binding specificity, solubility, primary sequence, molecular shape, hydrophobicity, thermal stability.

  • Initial Fractionation and Centrifugation:

    1. Cell Lysis: Mechanically or chemically break open cells to produce a crude lysate extract.

    2. Differential Centrifugation: Spin extract at increasing gravitational force to pellet structural rubbish or isolate subcellular organelles.

  • Fractional Precipitation ("Salting Out"):

    • Protein solubility varies with pHpH, temperature, and ionic strength.

    • Addition of high salt concentration (typically ammonium sulfate, (NH4)2SO4(NH_4)_2SO_4) reduces water available to hydrate protein surfaces, causing proteins to selectively precipitate out of solution.

    • Dialysis: Semipermeable membrane tubing allows small salt ions to diffuse out while retaining large protein macromolecules.

  • Column Chromatography Techniques:

    • Stationary Phase: Solid matrix material packed into a cylindrical column.

    • Mobile Phase: Buffered liquid phase carrying protein sample through matrix.

    • Fractions are collected continuously, and protein concentration is monitored by UV absorbance at 280 nm280\,nm (A280A_{280}).

    • Ion-Exchange Chromatography:

    • Separates proteins based on net surface charge sign and magnitude at a given pHpH

    • Anion Exchangers: Positively charged resin matrix that binds negatively charged proteins.

    • Cation Exchangers: Negatively charged resin matrix that binds positively charged proteins.

    • Elution: Performed by applying a salt gradient (NaClNaCl) or pHpH shift in mobile phase.

    • Gel Filtration / Size-Exclusion Chromatography:

    • Separates proteins according to size and hydrodynamic radius.

    • Column packed with engineered porous beads.

    • Small proteins enter porous interior, delaying their transit.

    • Large proteins are excluded from pores and travel rapidly through interstitial spaces, eluting first.

    • Affinity Chromatography:

    • Separates proteins based on highly specific binding interactions with ligands immobilized on resin matrix beads.

    • Bound proteins are eluted using high concentrations of free ligand or altered ionic strength.

    • High-Performance Liquid Chromatography (HPLC):

    • High-pressure mechanical pumps drive mobile phase through fine particle matrix columns.

    • Reduces transit time and diffusional broadening, resulting in high chromatographic resolution.

  • Enzyme Activity vs. Specific Activity:

    • Activity: Total units of enzyme activity present in a solution.

    • Specific Activity: Number of enzyme units per milligram of total protein (units/mg\text{units}/\text{mg}).

    • Specific Activity measures sample purity; it increases progressively as unwanted proteins are removed during purification.

Analytical Techniques: Electrophoresis and Isoelectric Focusing

  • Polyacrylamide Gel Electrophoresis (PAGE):

    • Analytical separation based on the migration of charged protein molecules through a polyacrylamide gel matrix within an applied electric field.

  • Native PAGE:

    • Performed under nondenaturing conditions preserving native folded states.

    • Noncovalent interactions, tertiary structure, and disulfide bonds remain intact.

    • Migration velocity depends complexly on net charge, physical size, and overall molecular shape.

    • Used to check protein sample purity and analyze intact protein-protein complexes.

  • SDS-PAGE (Sodium Dodecyl Sulfate PAGE):

    • Analytical denaturing method used to estimate molecular weight (MwM_w) and purity.

    • Detergent Chemical Structure: Sodium dodecyl sulfate (SDSSDS).

Structure of Sodium Dodecyl Sulfate SDS
  • Action of SDS: Binds proteins at a uniform ratio of approximately one SDSSDS molecule per two amino acid residues.

  • Charge Masking and Unfolding: The high negative charge of SDSSDS overwhelms intrinsic protein charges, conferring a uniform mass-to-charge ratio and rod-like unfolded conformation to all proteins.

  • Separation Mechanism: Proteins separate almost purely by molecular mass, with smaller proteins moving faster through polyacrylamide gel pores.

  • Plotting log⁡(Mr)\log(M_r) versus relative migration distance (RfR_f) yields a linear plot for determining unknown protein molecular weights.

    • Isoelectric Focusing (IEF):

  • Separates proteins according to their individual isoelectric points (pIpI).

  • Protein mixture is loaded onto a gel strip containing a stable immobilized pHpH gradient.

  • Under an applied electric field, proteins migrate until reaching the exact pHpH matching their pIpI

  • At pH=pIpH = pI, net charge equals zero (net charge=0net\ charge = 0), halting further migration.

    • Two-Dimensional (2D) Electrophoresis:

  • Sequentially combines Isoelectric Focusing (First Dimension, separation by pIpI) and SDS-PAGE (Second Dimension, separation by MwM_w perpendicular to the first strip).

  • Resolves complex protein mixtures containing thousands of distinct protein species.

Protein Structure Hierarchy and Chemical Cleavage / Modification

  • Hierarchical Levels of Protein Structure:

    • Primary Structure: Linear amino acid sequence joined covalently by peptide bonds.

    • Secondary Structure: Regular local spatial arrangements of backbone atoms (e.g., α\alpha-helices, β\beta-sheets), defined by specific backbone dihedral angles (ϕ\phi and ψ\psi) and repeating hydrogen bonds between main-chain amide groups.

    • Tertiary Structure: Complete three-dimensional spatial folding of a polypeptide chain, driven by hydrophobic core burial and stabilized by noncovalent interactions (salt bridges, hydrogen bonds) and covalent disulfide bonds.

    • Quaternary Structure: Three-dimensional arrangement of multiple polypeptide chains (subunits) in a multisubunit protein complex.

  • Cleavage and Modification of Disulfide Bonds:

    • Oxidation with Performic Acid: Irreversibly cleaves disulfide bonds (−S−S−-S-S-), converting cystine into two cysteic acid residues (−SO3−-SO_3^-).

    • Reduction with Thiol Reagents: Reversibly reduces disulfide bonds back to free sulfhydryl (−SH-SH) groups using dithiothreitol (DTT) or β\beta-mercaptoethanol (β-ME\beta\text{-ME}).

    • Alkylation with Iodoacetate: Follows reduction to covalently modify reactive −SH-SH groups via carboxymethylation, preventing re-oxidation and re-formation of disulfide bonds.

  • Identification of N-Terminal Amino Residues:

    • Reagents:     a) 11-fluoro-2,42,4-dinitrobenzene (FDNB, Sanger's reagent).     b) Dansyl chloride.     c) Dabsyl chloride.

    • Mechanism: Reagents react specifically with primary unprotonated amines to covalently label the α\alpha-amino group at the N-terminus. Peptide hydrolytic cleavage releases the labeled N-terminal derivative, which is identified chromatographically.

  • Enzymatic Cleavage: Proteases catalyze hydrolytic cleavage of peptide bonds at specific amino acid sequence sites.

Protein Sequencing Methods: Edman Degradation and Mass Spectrometry

  • Edman Degradation (Classical Chemical Method):

    • Performs cyclic rounds of N-terminal amino acid modification, selective cleavage, and chromatographic identification.

    • Sequences short peptide chains step-by-step from the N-terminus.

  • Mass Spectrometry (Modern High-Precision Method):

    • Measures molecular mass of intact proteins and peptide fragments with high accuracy.

    • Sequences short peptides (2020 to 3030 residues), identifies post-translational modifications, and analyzes complete cellular proteomes.

    • Four General Steps in Mass Spectrometry:

    1. Ionization: Analytes are converted into gas-phase ions in a high vacuum.

    2. Acceleration: Charged ions are introduced into electric and/or magnetic fields.

    3. Field Separation: Charged ions drift or orbit through fields as a function of mass-to-charge ratio (m/zm/z).

    4. Detection & Calculation: Ion trajectories translate into precise mass (mm) deduction.

    • Ionization Techniques:

    • Matrix-Assisted Laser Desorption/Ionization Mass Spectrometry (MALDI MS): Proteins embedded in a light-absorbing organic matrix are ionized and desorbed into the gas phase by laser pulses.

    • Electrospray Ionization Mass Spectrometry (ESI MS): Macromolecules are forced directly from liquid solution into gas-phase ions through a charged capillary tip.

    • Mass-to-Charge (m/zm/z) Analyzers:

    • Time of Flight (TOF): Ion acceleration velocity through an electric drift tube depends directly on m/zm/z

    • Orbitrap: Traps electrostatic ions in orbital motion around a central spindle; ion frequencies are converted to m/zm/z via Fourier transform.

    • Tandem Mass Spectrometry (MS/MS):

    • Uses two mass filters arranged in series.

    • First Mass Filter: Sorts and selects a specific target peptide ion produced by initial cleavage.

    • Collision Cell: Target peptide is fragmented by collisions with inert gas.

    • Second Mass Filter: Measures m/zm/z ratios of generated charged peptide fragments to reconstruct the sequence.

    • Liquid Chromatography-Tandem MS (LC-MS/MS):

    • Couples liquid chromatography separation directly to tandem mass spectrometry.

    • Resolves complex peptide mixtures continuously, enabling comprehensive proteome quantification.

Solid-Phase Peptide Synthesis and Bioinformatics

  • Solid-Phase Peptide Synthesis (SPPS):

    • Chemical synthesis technique developed by Bruce Merrifield.

    • Polypeptide chains are assembled while anchored covalently to an insoluble solid resin support.

    • Direction of Synthesis: Chemical synthesis proceeds stepwise from the C-terminus to the N-terminus direction (exact opposite of biological ribosomal synthesis).

  • Bioinformatics and Evolutionary Sequence Analysis:

    • Amino acid sequences provide insights into:

    • Three-dimensional tertiary structure prediction.

    • Biological function and active-site mechanisms.

    • Cellular localization.

    • Evolutionary history and phylogenetic relationships.

    • Conserved vs. Variable Residues:

    • Essential amino acid residues critical for catalytic activity or folding are conserved across evolution.

    • Non-essential structural residues vary across divergent species over time.

    • Consensus Sequences: Representative sequences displaying the most common amino acid residue at each position among aligned homologous sequences (visualized via Sequence Logos).

    • Lateral (Horizontal) Gene Transfer: Process in which an organism incorporates genetic material from another organism without being its offspring.

    • Homologous Proteins:

    • Homologs: Proteins sharing detectable sequence similarity.

    • Paralogs: Homologous proteins occurring within the same species (e.g., human α\alpha-globin and β\beta-globin).

    • Orthologs: Homologous proteins occurring in different species performing equivalent functions (e.g., human hemoglobin and bovine hemoglobin).

    • Signature Sequences: Contiguous amino acid motifs (1010 to 5050 residues long) associated with specific structural domains or biological functions across protein families.