Chapter 3: Amino Acids, Peptides, and Proteins
Amino Acids: Fundamental Building Blocks of Proteins
Proteins are linear heteropolymers composed of -amino acids linked together in a specific sequence.
Amino acids possess chemical properties that make them uniquely suited to fulfill diverse biological functions:
Capacity to polymerize into long polypeptide chains.
Useful acid-base properties across physiological and extreme ranges.
Varied physical properties, including solubility, polarity, and steric volume.
Diverse chemical functionality provided by distinct side chains ( groups).
Definition of an Amino Acid Residue:
According to IUPAC definition, when two or more amino acids combine to form a peptide, the elements of water () are removed.
What remains of each amino acid unit is termed an amino-acid residue.
Residues in a peptide chain lack a hydrogen atom of the amino group (), the hydroxyl moiety of the carboxyl group (), or both ().
Condensation vs. Hydrolysis Reactions:
Formation of peptide bonds occurs via a condensation reaction where water is eliminated.
Cleavage of peptide bonds occurs via a hydrolysis reaction, which represents the exact reverse of condensation.
General Structure of -Amino Acids:
Every amino acid contains a central carbon known as the -carbon ().
Attached covalently to the -carbon are four groups: a primary amino group (or secondary in proline), a carboxylic acid group, a hydrogen atom, and a distinctive side chain ( group).
There are standard amino acids commonly found in all ribosomal proteins, each defined by its unique group.
Nomenclature, Carbon Numbering, and Stereochemistry of Amino Acids
Carbon Nomenclature Systems:
Standard Numerical System: Carbon atoms are numbered starting from the carboxylic acid carbon as carbon-1 (), the -carbon as , followed sequentially along the side chain as , , , , etc.
Greek Letter Designation System: Carbon atoms are designated starting from the carbon adjacent to the carboxyl group, designated as (), followed down the side chain by , , , , etc.
Example (Lysine): Carbon-1 is the carboxyl carbon (), Carbon-2 () carries the primary -amino group (), Carbon-3 () is a group, Carbon-4 () is a group, Carbon-5 () is a group, and Carbon-6 () carries the side-chain amino group ().

The System of Stereochemistry:
The stereochemical nomenclature system relates the absolute configuration of a chiral molecule to the reference molecule glyceraldehyde.
-glyceraldehyde rotates plane-polarized light in the levorotatory (left) direction.
-glyceraldehyde rotates plane-polarized light in the dextrorotatory (right) direction.
The and designations specify absolute spatial configuration and do NOT consistently indicate the physical direction of optical rotation of plane-polarized light; actual optical rotation must be measured experimentally.
All -amino acids found naturally in proteins possess the stereochemistry.

The Cahn-Ingold-Prelog System:
When applying the absolute configuration rules to the chiral standard amino acids, of the chiral amino acids have the configuration.
Glycine is achiral because its -carbon is bound to two identical hydrogen atoms.
Cysteine is the sole exception among chiral amino acids; -cysteine has the configuration because the sulfur-containing side chain () takes priority over the carboxyl group () under Cahn-Ingold-Prelog priority rules.
Amino Acids with Multiple Chiral Centers:
Two standard amino acids possess a second chiral center in their side chains: Isoleucine and Threonine.
Classification and Chemical Properties of Standard Amino Acid Side Chains
Nonpolar, Aliphatic Amino Acids:
Characteristics: Lack polar functional groups; drive protein stability and folding through hydrophobic interactions; predominantly buried within the interior of folded globular proteins away from water.
Glycine (, ): The only achiral amino acid; possesses the smallest side chain (a single hydrogen atom, ); does not contribute hydrophobic driving force; provides high conformational flexibility to polypeptide backbones; derived from Greek glykos ("sweet") due to its taste; commonly utilized as a buffer component.
Alanine (, ): Possesses a simple methyl side chain ().
Valine (, ): Possesses a branched aliphatic side chain ().
Leucine (, ): Possesses a branched aliphatic side chain ().
Isoleucine (, ): Possesses a branched aliphatic side chain () containing a second chiral center.
Methionine (, ): Contains a nonpolar thioether side chain (); hydrophobic in nature; unlike a thiol group, a thioether cannot undergo oxidation to form covalent disulfide bonds.
Proline (, ): Features a distinctive cyclic structure where the aliphatic side chain forms a covalent ring with the backbone -amino group, forming a secondary amino (imino) group; holds polypeptide backbones in a rigid conformation that restricts structural flexibility; lacks a hydrogen atom on the backbone nitrogen atom when incorporated into peptide bonds, limiting its hydrogen-bonding capacity.
Aromatic R Groups:
Characteristics: Relatively nonpolar; participate in hydrophobic interactions; engage in stacking interactions with other aromatic rings.
Phenylalanine (, ): Nonpolar aromatic side chain containing a phenyl ring.
Tyrosine (, ): Contains a phenolic hydroxyl group (); more polar than phenylalanine; can form hydrogen bonds; features an ionizable side chain with ; named from Greek tyros ("cheese") after first being isolated from cheese.
Tryptophan (, ): Bulky, nonpolar indole-like side chain; slightly more polar than phenylalanine.
Ultraviolet Light Absorption:
Aromatic side chains, predominantly Tryptophan and Tyrosine (with minor contribution from Phenylalanine), absorb UV light in the wavelength range of to .
Absorption spectra plots display Absorbance () on the y-axis versus Wavelength (in ) on the x-axis.
Protein concentration () in solution is determined via UV-visible spectrophotometry using Beer's Law:
where is the measured absorbance, is the molar extinction coefficient, is molar concentration, and is the optical path length.
Polar Uncharged R Groups:
Characteristics: Soluble in water due to functional groups that readily form hydrogen bonds with aqueous solvent.
Serine (, ): Contains a primary hydroxyl side chain ().
Threonine (, ): Contains a secondary hydroxyl side chain () with a second chiral center.
Cysteine (, ): Contains a sulfhydryl (thiol) group ().
The sulfhydryl group is less polar than a hydroxyl group; acts as a weak acid (); can form dipole-dipole interactions with oxygen or nitrogen.
Oxidation of two cysteine sulfhydryl groups yields a covalent disulfide bond (), forming the dimeric amino acid cystine.
Disulfide cross-links stabilize tertiary and quaternary protein structures across different chains or within a single chain.
Asparagine (, ) and Glutamine (, ): Amide derivatives of aspartate and glutamate.
Glutamine possesses one additional methylene group () compared to asparagine.
Single-letter codes follow alphabetical order by side-chain length: for Asn, for Gln.
Positively Charged (Basic) R Groups:
Characteristics: Possess side chains with net positive charges at physiological .
Lysine (, ): Features a primary amino group at the -position ().
Arginine (, ): Contains a positively charged guanidino group (); one-letter abbreviation is .
Histidine (, ): Contains an ionizable imidazole ring with .
Histidine is the only standard amino acid with a side chain near physiological , allowing it to shift between protonated (charged) and unprotonated (neutral) states to catalyze enzymatic acid-base reactions.

Negatively Charged (Acidic) R Groups:
Characteristics: Possess side chains with net negative carboxylic acid charges at physiological .
Aspartic Acid / Aspartate (, ): Carboxyl side chain ().
Glutamic Acid / Glutamate (, ): Carboxyl side chain containing one additional methylene group ().
Single-letter codes follow alphabetical order by side-chain length: for Asp, for Glu.
Non-Standard, Modified, and Biological Amino Acids
Nutritional Classification and Auxotrophy:
Auxotrophy: The inability of an organism to synthesize a specific organic compound necessary for its growth and development.
Humans are auxotrophic for essential amino acids and must acquire them through diet.
Prototrophy: The ability of organisms (e.g., plants) to synthesize all required organic compounds from inorganic precursors.
Complete (Whole) Proteins: Animal protein sources (milk, eggs, dairy) supply all essential amino acids in proportions matching human nutritional needs.
Complementary Plant Proteins: Most plant items lack adequate proportions of specific essential amino acids (e.g., Rice is deficient in Lysine; Beans are deficient in Methionine; ingested together, they form a complete protein profile).
Cotranslationally Incorporated Non-Standard Amino Acids:
Selenocysteine (, ):
Considered the "21st amino acid"; cotranslationally incorporated during protein synthesis across all three domains of life; present in at least human proteins.
Contains selenium ( group) instead of sulfur.
Side chain , making it deprotonated at neutral physiological
Incorporation mechanism in prokaryotes: Context-dependent suppression of the opal () stop codon.

Pyrrolysine (, ):
First identified in methanogenic archaea; encoded by the amber () stop codon.

Post-Translationally Modified Amino Acids:
Formed via enzymatic modification after the polypeptide chain has been synthesized on the ribosome:
-hydroxyproline: Found in plant cell walls and collagen structural fibers.
-hydroxylysine: Found in collagen structural fibers.
-methyllysine: Found in muscle contractile protein myosin.
-carboxyglutamate: Found in blood-clotting protein prothrombin and other calcium-binding () proteins.
Desmosine: A complex derivative found in the fibrous protein elastin.
Reversible / Transient Modifications for Protein Regulation:
Phosphorylation: Most common regulatory modification; targets hydroxyl groups of Serine, Threonine, and Tyrosine.
Acetylation: -acetyllysine found in histone proteins; regulates gene expression and chromatin structure at an epigenetic level.
Other transient groups: Methyl, adenylyl, and ADP-ribosyl groups.
Non-Protein Amino Acid Intermediates:
Ornithine and Citrulline: Metabolic intermediates in the biosynthesis of arginine and the urea cycle; not incorporated into proteins during translation.

Acid-Base Chemistry and Ionization Behavior of Amino Acids
Polyprotic Acid Behavior:
Amino acids act as weak polyprotic acids containing at least two ionizable protons, each with a distinct value.
-Carboxylic Acid Dissociation:
-Amino Group Dissociation:
At strongly acidic : Amino acids exist in a fully protonated cationic form ( or greater).
At strongly basic : Amino acids exist in a fully deprotonated anionic form ( or lower).
At intermediate : Amino acids exist as dipolar zwitterions carrying both a positive and a negative charge ().

Amphoteric Nature:
Zwitterions can donate a proton (acting as an acid) or accept a proton (acting as a base); compounds with this dual capability are termed amphoteric or amphiolytes.

Chemical Environment Effects on :
The of an -carboxyl group is , significantly lower than carboxylic acids like acetic acid ().
Explanation: The positively charged -amino group () exerts an electron-withdrawing inductive effect on the carboxyl group, stabilizing the negatively charged carboxylate form and lowering its
The of an -amino group is , slightly lower than simple primary amines like methylamine ().
Explanation: Electronegative oxygen atoms in the carboxyl group withdraw electron density from the amino group, facilitating proton dissociation and lowering its
Ionization State Thresholds:
Above , -carboxyl groups are entirely in their carboxylate () form.
Below , -amino groups are entirely in their protonated ammonium () form.

Titration Curves and Isoelectric Point Determination
Henderson-Hasselbalch Equation:
Quantitative relationship between , , and conjugate base/acid ratio:
Isoelectric Point ():
The characteristic at which the net electric charge of an amino acid or peptide is exactly zero.
Calculation for Diprotic Amino Acids (e.g., Glycine):
For Glycine:
Physical Properties at the Isoelectric Point ():
Net electric charge equals zero.
Amino acid exhibits minimum solubility in aqueous solvent.
Molecule experiences zero net directional migration within an applied electric field.
Net Charge Dependencies:
When , the molecule carries a net negative charge.
When , the molecule carries a net positive charge.

Formation, Structure, and Nomenclature of Peptides
Peptide Bond Formation:
Formed via a condensation reaction (nucleophilic acyl substitution).
The unprotonated -amino group of one amino acid acts as a nucleophile, attacking the -carboxyl carbon of another amino acid to displace a hydroxyl group, releasing
While amino groups are effective nucleophiles, hydroxyl groups are poor leaving groups; thus, at physiological , this condensation reaction does not proceed spontaneously to a significant extent without energy input and enzymatic catalysis.

Size Classifications:
Dipeptide: Contains amino acid residues.
Tripeptide: Contains amino acid residues.
Oligopeptide: General term for short peptide chains containing to amino acid residues.
Polypeptide: Chain containing greater than amino acid residues.
Protein: Polypeptide chain containing hundreds or thousands of amino acid residues with molecular mass () exceeding
Naming and Conventions for Peptides:
Sequence Directionality: Written and read from the N-terminus to the C-terminus (left to right).
N-Terminus (Amino Terminus): Residue with the free -amino group ().
C-Terminus (Carboxyl Terminus): Residue with the free -carboxylate group ().
Full Systematic Nomenclature: Systematic names link constituent amino acids sequentially from N-terminus to C-terminus, changing suffixes -ine or -ate to -yl, leaving the C-terminal residue name unchanged.
Three-Letter Notation: Linked with dashes (e.g., ).
One-Letter Notation: Linked continuously without dashes (e.g., ).
Biologically Active Peptides and Protein Complexity
Peptide Ionization Behavior:
Peptides carry only one free ionizable -amino group at the N-terminus and one free ionizable -carboxyl group at the C-terminus.
Total charge is determined by terminal groups and the number of ionizable groups.
Side chain values buried within folded protein environments differ from values of free amino acids.
Examples of Biologically Active Peptides:
Commercial Sweetener: Aspartame ().

Vertebrate Hormones & Pheromones: Insulin, Oxytocin.
Neuropeptides: Substance P (mediates pain perception signal transmission).
Antimicrobial Peptides:
Polymyxin B: Active against Gram-negative bacteria (produced by Bacillus polymyxa).
Bacitracin: Active against Gram-positive bacteria (produced by Bacillus subtilis).
Dermcidin-1L: Human antimicrobial defense peptide (PDB ID: 2KSG).
Potent Toxins: Amanitin (death cap mushrooms), Conotoxin (cone snails), Chlorotoxin (scorpions).
Protein Composition and Structural Components:
Polypeptides consist of covalently linked -amino acid residues, sometimes bound to non-amino acid components:
Cofactors: Functional non-amino acid components, such as metal ions (, , ) or small organic molecules.
Coenzymes: Complex organic cofactors (e.g., in lactate dehydrogenase).
Prosthetic Groups: Covalently or tightly bound cofactors essential for activity (e.g., heme group in myoglobin).
Multisubunit Proteins: Contain two or more polypeptide chains associated noncovalently.
Oligomeric Proteins: Multisubunit proteins in which at least two polypeptide chains are identical.
Protomers: Identical repeating structural units consisting of one or more polypeptide chains in an oligomeric protein.
Example: Hemoglobin contains two chains and two chains (); can be defined as a tetramer of four subunits or a dimer of protomers.
Conjugated Proteins: Proteins containing permanently associated non-amino acid chemical groups.
Estimation of Amino Acid Residue Count:
The average molecular weight ( or ) of an amino acid residue in a protein is estimated at ().
Derived from average amino acid mass () weighted for abundance of smaller amino acids, minus for water lost during condensation.
Methods for Protein Purification and Separation
Physicochemical Properties Exploited in Purification:
Size, charge, binding specificity, solubility, primary sequence, molecular shape, hydrophobicity, thermal stability.
Initial Fractionation and Centrifugation:
Cell Lysis: Mechanically or chemically break open cells to produce a crude lysate extract.
Differential Centrifugation: Spin extract at increasing gravitational force to pellet structural rubbish or isolate subcellular organelles.
Fractional Precipitation ("Salting Out"):
Protein solubility varies with , temperature, and ionic strength.
Addition of high salt concentration (typically ammonium sulfate, ) reduces water available to hydrate protein surfaces, causing proteins to selectively precipitate out of solution.
Dialysis: Semipermeable membrane tubing allows small salt ions to diffuse out while retaining large protein macromolecules.
Column Chromatography Techniques:
Stationary Phase: Solid matrix material packed into a cylindrical column.
Mobile Phase: Buffered liquid phase carrying protein sample through matrix.
Fractions are collected continuously, and protein concentration is monitored by UV absorbance at ().
Ion-Exchange Chromatography:
Separates proteins based on net surface charge sign and magnitude at a given
Anion Exchangers: Positively charged resin matrix that binds negatively charged proteins.
Cation Exchangers: Negatively charged resin matrix that binds positively charged proteins.
Elution: Performed by applying a salt gradient () or shift in mobile phase.
Gel Filtration / Size-Exclusion Chromatography:
Separates proteins according to size and hydrodynamic radius.
Column packed with engineered porous beads.
Small proteins enter porous interior, delaying their transit.
Large proteins are excluded from pores and travel rapidly through interstitial spaces, eluting first.
Affinity Chromatography:
Separates proteins based on highly specific binding interactions with ligands immobilized on resin matrix beads.
Bound proteins are eluted using high concentrations of free ligand or altered ionic strength.
High-Performance Liquid Chromatography (HPLC):
High-pressure mechanical pumps drive mobile phase through fine particle matrix columns.
Reduces transit time and diffusional broadening, resulting in high chromatographic resolution.
Enzyme Activity vs. Specific Activity:
Activity: Total units of enzyme activity present in a solution.
Specific Activity: Number of enzyme units per milligram of total protein ().
Specific Activity measures sample purity; it increases progressively as unwanted proteins are removed during purification.
Analytical Techniques: Electrophoresis and Isoelectric Focusing
Polyacrylamide Gel Electrophoresis (PAGE):
Analytical separation based on the migration of charged protein molecules through a polyacrylamide gel matrix within an applied electric field.
Native PAGE:
Performed under nondenaturing conditions preserving native folded states.
Noncovalent interactions, tertiary structure, and disulfide bonds remain intact.
Migration velocity depends complexly on net charge, physical size, and overall molecular shape.
Used to check protein sample purity and analyze intact protein-protein complexes.
SDS-PAGE (Sodium Dodecyl Sulfate PAGE):
Analytical denaturing method used to estimate molecular weight () and purity.
Detergent Chemical Structure: Sodium dodecyl sulfate ().

Action of SDS: Binds proteins at a uniform ratio of approximately one molecule per two amino acid residues.
Charge Masking and Unfolding: The high negative charge of overwhelms intrinsic protein charges, conferring a uniform mass-to-charge ratio and rod-like unfolded conformation to all proteins.
Separation Mechanism: Proteins separate almost purely by molecular mass, with smaller proteins moving faster through polyacrylamide gel pores.
Plotting versus relative migration distance () yields a linear plot for determining unknown protein molecular weights.
Isoelectric Focusing (IEF):
Separates proteins according to their individual isoelectric points ().
Protein mixture is loaded onto a gel strip containing a stable immobilized gradient.
Under an applied electric field, proteins migrate until reaching the exact matching their
At , net charge equals zero (), halting further migration.
Two-Dimensional (2D) Electrophoresis:
Sequentially combines Isoelectric Focusing (First Dimension, separation by ) and SDS-PAGE (Second Dimension, separation by perpendicular to the first strip).
Resolves complex protein mixtures containing thousands of distinct protein species.
Protein Structure Hierarchy and Chemical Cleavage / Modification
Hierarchical Levels of Protein Structure:
Primary Structure: Linear amino acid sequence joined covalently by peptide bonds.
Secondary Structure: Regular local spatial arrangements of backbone atoms (e.g., -helices, -sheets), defined by specific backbone dihedral angles ( and ) and repeating hydrogen bonds between main-chain amide groups.
Tertiary Structure: Complete three-dimensional spatial folding of a polypeptide chain, driven by hydrophobic core burial and stabilized by noncovalent interactions (salt bridges, hydrogen bonds) and covalent disulfide bonds.
Quaternary Structure: Three-dimensional arrangement of multiple polypeptide chains (subunits) in a multisubunit protein complex.
Cleavage and Modification of Disulfide Bonds:
Oxidation with Performic Acid: Irreversibly cleaves disulfide bonds (), converting cystine into two cysteic acid residues ().
Reduction with Thiol Reagents: Reversibly reduces disulfide bonds back to free sulfhydryl () groups using dithiothreitol (DTT) or -mercaptoethanol ().
Alkylation with Iodoacetate: Follows reduction to covalently modify reactive groups via carboxymethylation, preventing re-oxidation and re-formation of disulfide bonds.
Identification of N-Terminal Amino Residues:
Reagents: a) -fluoro--dinitrobenzene (FDNB, Sanger's reagent). b) Dansyl chloride. c) Dabsyl chloride.
Mechanism: Reagents react specifically with primary unprotonated amines to covalently label the -amino group at the N-terminus. Peptide hydrolytic cleavage releases the labeled N-terminal derivative, which is identified chromatographically.
Enzymatic Cleavage: Proteases catalyze hydrolytic cleavage of peptide bonds at specific amino acid sequence sites.
Protein Sequencing Methods: Edman Degradation and Mass Spectrometry
Edman Degradation (Classical Chemical Method):
Performs cyclic rounds of N-terminal amino acid modification, selective cleavage, and chromatographic identification.
Sequences short peptide chains step-by-step from the N-terminus.
Mass Spectrometry (Modern High-Precision Method):
Measures molecular mass of intact proteins and peptide fragments with high accuracy.
Sequences short peptides ( to residues), identifies post-translational modifications, and analyzes complete cellular proteomes.
Four General Steps in Mass Spectrometry:
Ionization: Analytes are converted into gas-phase ions in a high vacuum.
Acceleration: Charged ions are introduced into electric and/or magnetic fields.
Field Separation: Charged ions drift or orbit through fields as a function of mass-to-charge ratio ().
Detection & Calculation: Ion trajectories translate into precise mass () deduction.
Ionization Techniques:
Matrix-Assisted Laser Desorption/Ionization Mass Spectrometry (MALDI MS): Proteins embedded in a light-absorbing organic matrix are ionized and desorbed into the gas phase by laser pulses.
Electrospray Ionization Mass Spectrometry (ESI MS): Macromolecules are forced directly from liquid solution into gas-phase ions through a charged capillary tip.
Mass-to-Charge () Analyzers:
Time of Flight (TOF): Ion acceleration velocity through an electric drift tube depends directly on
Orbitrap: Traps electrostatic ions in orbital motion around a central spindle; ion frequencies are converted to via Fourier transform.
Tandem Mass Spectrometry (MS/MS):
Uses two mass filters arranged in series.
First Mass Filter: Sorts and selects a specific target peptide ion produced by initial cleavage.
Collision Cell: Target peptide is fragmented by collisions with inert gas.
Second Mass Filter: Measures ratios of generated charged peptide fragments to reconstruct the sequence.
Liquid Chromatography-Tandem MS (LC-MS/MS):
Couples liquid chromatography separation directly to tandem mass spectrometry.
Resolves complex peptide mixtures continuously, enabling comprehensive proteome quantification.
Solid-Phase Peptide Synthesis and Bioinformatics
Solid-Phase Peptide Synthesis (SPPS):
Chemical synthesis technique developed by Bruce Merrifield.
Polypeptide chains are assembled while anchored covalently to an insoluble solid resin support.
Direction of Synthesis: Chemical synthesis proceeds stepwise from the C-terminus to the N-terminus direction (exact opposite of biological ribosomal synthesis).
Bioinformatics and Evolutionary Sequence Analysis:
Amino acid sequences provide insights into:
Three-dimensional tertiary structure prediction.
Biological function and active-site mechanisms.
Cellular localization.
Evolutionary history and phylogenetic relationships.
Conserved vs. Variable Residues:
Essential amino acid residues critical for catalytic activity or folding are conserved across evolution.
Non-essential structural residues vary across divergent species over time.
Consensus Sequences: Representative sequences displaying the most common amino acid residue at each position among aligned homologous sequences (visualized via Sequence Logos).
Lateral (Horizontal) Gene Transfer: Process in which an organism incorporates genetic material from another organism without being its offspring.
Homologous Proteins:
Homologs: Proteins sharing detectable sequence similarity.
Paralogs: Homologous proteins occurring within the same species (e.g., human -globin and -globin).
Orthologs: Homologous proteins occurring in different species performing equivalent functions (e.g., human hemoglobin and bovine hemoglobin).
Signature Sequences: Contiguous amino acid motifs ( to residues long) associated with specific structural domains or biological functions across protein families.