Proteins and Protein Primary Structure Complete Study Guide
Definition and Fundamentals of Protein Primary Structure
- Primary Structure Definition: The primary structure of a protein (or any biopolymer) is defined as its linear sequence of amino acid residues linked by covalent peptide bonds, along with the exact covalent location of any disulfide () linkages or bridges.

- Combinatorial Possibilities:
- Protein molecules are constructed from standard amino acids ( distinct types).
- The total number of possible linear sequences for a polypeptide of length is given by .
- For a short tripeptide (), the number of possible unique sequences is:
- Despite this immense theoretical diversity, only a highly selected fraction of possible amino acid sequence combinations occurs in nature due to evolutionary filtering for stable folding and biological function.
- Representative Sequence Structure:
- An example linear sequence fragment is: Gly-Arg-Ser-Thr-Thr-Leu-Ala-Asp-Glu-Ile-Tyr-Phe-Glu.
- Disulfide bridges form covalently between the sulfhydryl () side chains of two cysteine residues, producing a cystine unit and locking specific regions of the polypeptide backbone into fixed spatial relationships.
Historical Milestones in Protein Sequencing
- Sequencing of Bovine Insulin:
- Insulin was the first protein ever to have its complete primary amino acid sequence successfully determined.
- Frederick Sanger determined the full sequence of bovine insulin, demonstrating for the first time that proteins possess precise, defined linear chemical structures rather than random molecular mixtures.

- Structure of Bovine Insulin:
- Consists of two distinct polypeptide chains held together by interchain disulfide bonds:
- A Chain: amino acid residues in length, beginning with an N-terminal Glycine residue (Gly) and containing an intrachain disulfide bond between Cys6 and Cys11.
- Complete A chain sequence: Gly-Ile-Val-Glu-Gln-Cys-Cys-Ala-Ser-Val-Cys-Ser-Leu-Tyr-Gln-Leu-Glu-Asn-Tyr-Cys-Asn.
- B Chain: amino acid residues in length, beginning with an N-terminal Phenylalanine residue (Phe).
- Complete B chain sequence: Phe-Val-Asn-Gln-His-Leu-Cys-Gly-Ser-His-Leu-Val-Glu-Ala-Leu-Tyr-Leu-Val-Cys-Gly-Glu-Arg-Gly-Phe-Phe-Tyr-Thr-Pro-Lys-Ala.
- Disulfide Linkages: Two interchain disulfide bridges connect the two chains: A-Cys7 to B-Cys7, and A-Cys20 to B-Cys19.
- Nobel Recognition:
- Frederick Sanger received his first Nobel Prize in Chemistry in for his landmark determination of protein primary structures, specifically bovine insulin.
- Original determination required approximately \,years of research, a dedicated team of scientists, and \,g of purified protein.
- Sanger was awarded a second Nobel Prize in Chemistry in for developing methods for nucleic acid (DNA) sequencing.
- Modern Capabilities: Automated instrument technologies allow a single researcher to determine the primary sequence of insulin within days using milligram or microgram samples.
Workflow for Protein Primary Structure Determination
- General 3-Step Strategy:
- Chain Separation: Separate multi-subunit proteins or interchain-linked polypeptides into individual, single polypeptide chains by cleaving all interchain and intrachain disulfide bonds.
- End Terminus Analysis: Identify the specific amino acid residues present at the N-terminus (-terminus) and C-terminus (-terminus) of each separated chain.
- Fragment Sequencing & Reconstruction: Cleave individual chains into smaller fragments using specific chemical or enzymatic methods, sequence the isolated fragments, and align overlapping sequences to reconstruct the intact sequence.

Step 1: Separation of Polypeptide Chains and Disulfide Bond Cleavage
- Method 1: Oxidation with Performic Acid:
- Performic acid () oxidizes disulfide bridges () irreversibly.

- Both cystine (disulfide-linked) and cysteine (free thiol) residues are converted into cysteic acid residues containing sulfonic acid functional groups ().
- Disadvantages: Performic acid oxidation can destroy Tryptophan residues and oxidize Methionine side chains.
- Method 2: Reduction followed by Alkylation ("Capping"):
- Step A (Thiol Reduction): Disulfide bonds are reversibly reduced by treatment with an excess of a thiol-containing reducing agent, such as 2-mercaptoethanol (), dithiothreitol (DTT), or dithioerythritol (DTE).

- Step B (Thiol Alkylation / Capping): Free sulfhydryl () groups generated by reduction must be permanently alkylated ("capped") to prevent spontaneous re-oxidation back to disulfide bonds.
- Reaction with iodoacetic acid (iodoacetate, ) occurs via nucleophilic substitution ():

- Converts free cysteine residues into stable -carboxymethylcysteine ($ ext{CM-Cys}$) residues, preventing re-formation of disulfide linkages.
Step 2: End Terminus Analysis
- C-Terminal Analysis (Carboxypeptidases):
- Carboxypeptidases are exopeptidases (exoproteases) that hydrolyze peptide bonds sequentially starting from the C-terminal amino acid inward toward the N-terminus.
- Identification is achieved by monitoring the kinetic rate and time order of free amino acid release into solution.
- Carboxypeptidase A: Source: Bovine pancreas. Cleaves C-terminal peptide bonds for all residues except basic amino acids Lysine ($ ext{Lys}$) and Arginine ($ ext{Arg}$), or when Proline ($ ext{Pro}$) is adjacent to the scissile bond.
- Carboxypeptidase B: Source: Bovine pancreas. Cleaves C-terminal peptide bonds only when the C-terminal residue () is a basic amino acid: Lysine ($ ext{Lys}$) or Arginine ($ ext{Arg}$).
- N-Terminal Analysis (Chemical Reagents):

- 2,4-Dinitrofluorobenzene (DNFB / Sanger's Reagent):
- Undergoes Nucleophilic Aromatic Substitution (NAS) with the unprotonated N-terminal group under basic conditions, eliminating to yield a dinitrophenyl-polypeptide (DNP-polypeptide).

- Total acid hydrolysis (\,N at \,) hydrolyzes all internal peptide bonds, yielding free amino acids and a yellow-colored 2,4-dinitrophenyl-amino acid (DNP-amino acid) derivative that is extracted and identified by chromatography.
- Dansyl Chloride (5-Dimethylamino-1-naphthalenesulfonyl chloride):
- Reacts with the N-terminal amine to form a stable sulfonamide adduct (dansyl polypeptide).

- Subsequent total acid hydrolysis yields free amino acids and an intensely fluorescent dansylamino acid derivative.
- High fluorescence emission allows detection at nanomolar concentrations, requiring significantly less sample than DNFB.
- Dabsyl Chloride (4-Dimethylaminoazobenzene-4'-sulfonyl chloride):
- Reacts with the N-terminal amino group to form a brightly colored dabsylamino acid derivative upon total acid hydrolysis, enabling detection via absorbance.
- Disadvantage of Standard N-Terminal Reagents: Total acid hydrolysis hydrolyzes every peptide bond in the chain. The overall peptide backbone is completely destroyed to liberate the single N-terminal derivative.
- Side-Chain False Positives: Lysine residues possess an group on their (-carbon) side chain that also reacts with DNFB, Dansyl-Cl, and Dabsyl-Cl. This produces \text{\epsilon}-labeled lysine derivatives that must be differentiated chromatographically from true \text{\alpha}-amino labeled N-terminal residues to avoid false positive assignments.
Step 3: Protein Sequencing Methods
- Edman Degradation (Phenylisothiocyanate / PITC Protocol):
- Developed by Pehr Edman; allows sequential step-by-step removal and identification of N-terminal amino acids without destroying the remaining polypeptide chain.

- Step 1 (Coupling): Phenylisothiocyanate (PITC, Edman reagent) reacts selectively with the free, unprotonated \text{\alpha-NH}_2 group of the N-terminal residue at mild alkaline pH (approx. ) to form a phenylthiocarbamyl (PTC) polypeptide.
- Step 2 (Cleavage): Treatment with anhydrous trifluoroacetic acid ($ ext{CF}_3 ext{COOH}$ / TFA) selectively cleaves the first peptide bond via intramolecular cyclization, liberating the N-terminal residue as an unstable thiazolinone derivative.
- Crucially, because TFA treatment is anhydrous, internal peptide bonds are not hydrolyzed, leaving the shortened polypeptide ( residues) completely intact in the organic phase.
- Step 3 (Isomerization & Identification): The thiazolinone derivative is extracted into an organic solvent, treated with aqueous acid () to isomerize into a stable phenylthiohydantoin (PTH)-amino acid, and identified using High-Performance Liquid Chromatography (HPLC).
- The remaining intact polypeptide ( residues) is subjected to repeated cycles of coupling, cleavage, and identification.
- Repetitive Yield & Practical Limits:
- The efficiency of each Edman cycle is exceptionally high (approx. efficiency per cycle).
- Due to accumulated incomplete reactions (the unreacted fraction in each cycle), background noise gradually increases.
- Practical upper limit per chain: Approximately cycles/residues.
- Polypeptides longer than residues cannot be fully sequenced from end to end in a single run and must first be fragmented into smaller peptides.
Chemical and Enzymatic Fragmentation Protocols
- Protease Cleavage (Endopeptidases / Endoproteases):
- Endopeptidases cleave internal scissile peptide bonds within a polypeptide chain based on specific amino acid side-chain recognition.

- Trypsin:
- Source: Bovine pancreas.
- Specificity: Cleaves peptide bonds strictly on the carboxyl side () of positively charged basic residues: Lysine ($ ext{Lys}$) and Arginine ($ ext{Arg}$).
- Restriction: Cleavage does not occur if the following amino acid residue () is Proline ($ ext{Pro}$). Highly specific enzyme.

- Chymotrypsin:
- Source: Bovine pancreas.
- Specificity: Cleaves peptide bonds on the carboxyl side () of bulky hydrophobic/aromatic residues: Phenylalanine ($ ext{Phe}$), Tryptophan ($ ext{Trp}$), and Tyrosine ($ ext{Tyr}$).
- Restriction: Cleaves much more slowly if is Asparagine ($ ext{Asn}$), Histidine ($ ext{His}$), Methionine ($ ext{Met}$), or Leucine ($ ext{Leu}$). Cleavage does not occur if R_n = \text{Pro}$.\n - **Elastase**:\n - *Source*: Bovine pancreas.\n - *Specificity*: Cleaves on the carboxyl side (R_{n-1}\text{Ala}\text{Gly}\text{Ser}\text{Val}R_n eq ext{Pro}$.
- Thermolysin:
- Source: Bacillus thermoproteolyticus (heat-stable enzyme).
- Specificity: Cleaves on the amino side () of bulky hydrophobic residues: Isoleucine ($ ext{Ile}$), Methionine ($ ext{Met}$), Phenylalanine ($ ext{Phe}$), Tryptophan ($ ext{Trp}$), Tyrosine ($ ext{Tyr}$), and Valine ($ ext{Val}$), provided R_{n-1} \neq \text{Pro}$. Occasionally cleaves when R_n ext{Ala} ext{Asp} ext{His} ext{Thr}.\n - **Pepsin**:\n - *Source*: Bovine gastric mucosa.\n - *Optimum Conditions*: Highly active at low pH ( ext{pH optimum } = 2).\n - *Specificity*: Quite nonspecific; cleaves on the amino side (R_n ext{Leu} ext{Phe} ext{Trp} ext{Tyr}R_{n-1} eq ext{Pro}$.
- Endopeptidase V8:
- Source: Staphylococcus aureus.
- Specificity: Cleaves strictly on the carboxyl side () of acidic Glutamic acid ($ ext{Glu}$) residues.
- Chemical Cleavage with Cyanogen Bromide (CNBr):
- Reagent Safety: Cyanogen bromide ($ ext{CNBr}$) is volatile and extremely toxic.
- Specificity: Cleaves peptide bonds strictly on the carboxyl side of internal Methionine residues ().

- Reaction Outcome: The nucleophilic attack of the Methionine sulfur atom on CNBr converts the internal Methionine residue into a C-terminal peptidyl homoserine lactone derivative, while releasing the downstream amino segment as a free N-terminal peptide fragment.
Sequence Assembly via Overlapping Peptides
- Principle of Overlapping Fragments:
- A single cleavage method only yields isolated peptide fragments without revealing the order in which those fragments are connected in the intact protein.
- To assemble the complete sequence, at least two independent fragmentation procedures using enzymes or reagents of differing cleavage specificities (e.g., Trypsin and Chymotrypsin) must be performed on separate samples of the purified polypeptide.

- Worked Example of Sequence Reconstruction:
- Tryptic peptides generated:
- Ala-Ala-Trp-Gly-Lys
- Thr-Phe-Val-Lys
- Chymotryptic peptide generated:
- Val-Lys-Ala-Ala-Trp
- Alignment Analysis:
- The chymotryptic overlap fragment (Val-Lys-Ala-Ala-Trp) contains the C-terminal segment of tryptic peptide 2 (Val-Lys) joined directly to the N-terminal segment of tryptic peptide 1 (Ala-Ala-Trp).
- Reconstructed Sequence: Thr-Phe-Val-Lys-Ala-Ala-Trp-Gly-Lys.
- Assigning Disulfide Linkages:
- To determine which Cysteine residues form disulfide bonds in the original folded protein, the native protein is fragmented using proteases without prior reduction of disulfide bridges.
- The resulting disulfide-linked peptide pairs are isolated, reduced, and sequenced to identify the exact paired Cysteine positions.
Mass Spectrometry and High-Throughput Proteomics
- Sequence Deduction via Recombinant DNA/RNA Sequencing:
- Modern high-throughput sequencing determines amino acid sequences indirectly by sequencing cloned cDNA or genomic DNA encoding the target protein.
- Using the universal triplet genetic code, the linear amino acid sequence is directly predicted from nucleotide codons.
- Peptide Mass Fingerprinting (PMF) via Soft Ionization Mass Spectrometry:
- Matrix-Assisted Laser Desorption Ionization (MALDI) & Electrospray Ionization (ESI):
- Soft ionization techniques that transition intact, non-volatile protein and peptide molecules into gas-phase ions without causing thermal destruction.

- Time-Of-Flight (TOF) Analyzer Mechanism:
- Gas-phase peptide ions generated by laser irradiation or electrospray are accelerated into an evacuated, field-free flight tube (typically \,m in length).
- Every ion receives identical kinetic energy acceleration from the electrostatic field.
- The velocity of each ion depends inverse-proportionally on its mass-to-charge ratio ():
\text{Velocity} \n\propto \sqrt{\frac{z}{m}}
- Smaller, lighter ions accelerate faster and reach the detector first, whereas larger, heavier ions travel slower and arrive later.
- The resulting spectrum of arrival times produces a precise Peptide Mass Fingerprint (PMF) used to query database sequences.
- Tandem Mass Spectrometry (MS/MS):
- Used for rapid sequencing of complex protein mixtures without prior purification to homogeneity.
- MS-1 Phase: First mass spectrometer separates individual peptide ions () from a complex tryptic digest.
- Collision Cell: A specific peptide ion is selected and directed into a collision chamber filled with inert gas (Helium, ), causing fragmentation along peptide backbone bonds.
- MS-2 Phase: Second mass spectrometer measures the precise ratios of the resulting fragment ions (), revealing the amino acid sequence directly from the mass differences between successive fragment peaks.
- Proteomics:
- Proteomics is the global, quantitative study of the complete set of proteins (the proteome) expressed by a cell, tissue, or organism under specific environmental conditions at a given point in time.

- Example: A proteogram mapping approximately distinct expressed proteins and their physical interaction networks in a Drosophila melanogaster (fruit fly) cell allows researchers to predict how molecular perturbation ripples across cellular pathways.
Protein Evolution, Sequence Homology, and Molecular Clocks
- Sequence Homology and Conservation (Cytochrome c Studies):
- Comparison of primary sequences of c-type cytochromes across diverse eukaryotic species reveals that only out of approximately amino acid positions are strictly invariant (completely identical across all tested species).

- Conserved Regions: Amino acids in these positions can vary between species, but the chemical properties (e.g., charge, side-chain polarity, size) are strictly conserved to maintain structural integrity and functional binding interactions.
- Hyper-variable Regions: Regions where side-chain identity and chemical polarity vary widely across species without disrupting biological function, indicating these surface positions are under minimal selective constraint.
- Structural Conservation vs. Sequence Divergence:
- While sequence identity across evolutionary distant cytochromes may drop significantly, their three-dimensional X-ray crystal structures remain nearly identical.
- Conclusion: Essential tertiary structural folds and active site geometries are strongly conserved during biological evolution, even when linear amino acid sequences diverge substantially.
- The Molecular Clock Concept:
- Accumulation of neutral amino acid sequence differences between homologous proteins in two species correlates linearly with the time elapsed since those species diverged from a common ancestor.

- Divergence Data for Cytochrome c:
- Human vs. Monkey: amino acid difference (\,million years since divergence).
- Horse vs. Cow: amino acid differences (\,million years since divergence).
- Human vs. Dog: amino acid differences (\,million years since divergence).
- Human vs. Horse: amino acid differences (\,million years since divergence).
- Mammals vs. Birds: amino acid differences (\,million years since divergence).
- Mammals vs. Fish: amino acid differences (\,million years since divergence).
- Vertebrates vs. Yeast: amino acid differences (\,million years since divergence).
- Variation in Evolutionary Rates Among Protein Families:
- Fibrinopeptides: Evolve extremely rapidly due to minimal functional constraints outside of general solubility and cleavage susceptibility.
- Hemoglobin: Evolves at an intermediate rate.
- Cytochrome c: Evolves slowly due to structural constraints required for electron transport and inter-protein binding.
- Histone H4: Evolves at an extremely slow rate (nearly zero amino acid substitutions over hundreds of millions of years) because virtually every residue interacts directly with DNA or other histones in nucleosomes.
- Gene Duplication and Protein Families:
- Homologous proteins are evolutionary related through shared ancestry.
- Protein families (e.g., the Globin family: Myoglobin, Hemoglobin \text{\alpha}, Hemoglobin \text{\beta}) arise through ancestral gene duplication events.
- Following duplication, individual gene copies accumulate independent mutations and diverge, evolving distinct physiological properties (e.g., monomeric oxygen storage vs. tetrameric cooperative oxygen transport).