Comprehensive Study Guide to Gene Expression: Protein Translation and Structure

Fundamental Building Blocks: Protein Composition and Structure

  • Proteins are complex, high-molecular-weight organic compounds containing nitrogen. They are constructed from macromolecular subunits termed polypeptides, which are further composed of basic building blocks known as amino acids.

  • Chemical Anatomy of Amino Acids: Standard amino acids (excluding proline) share a common orientation centered on an α\alpha-carbon atom. This central atom is bonded to a hydrogen atom, a carboxyl group (COOHCOOH), an amino group (NH2NH_2), and a variable RR group (side chain). Under physiological pH conditions (approximately pH7.4pH\,7.4), the functional groups exist in ionized states: NH3+-NH_3^+ and COO-COO^-.

  • The RR Group: The specific side chain attached to the α\alpha-carbon determines the unique chemical properties of each of the 2020 different amino acids. These side chains are classified into four categories: acidic, basic, neutral/polar, and neutral/nonpolar.

  • Peptide Bond Formation: Polypeptides are formed through covalent peptide bonds, created by a dehydration reaction between the carboxyl group of one amino acid and the amino group of the next. This process releases a molecule of water (H2OH_2O). Polypeptides have structural polarity, with a free amino group at the NN-terminus (the beginning of the chain) and a free carboxyl group at the CC-terminus.

  • Hierarchical Levels of Protein Structure:

    • Primary Structure: The unique linear sequence of amino acids, dictated directly by the nucleotide sequence of the corresponding gene.

    • Secondary Structure: Local spatial arrangements such as the α\alpha-helix and β\beta-pleated sheet. The α\alpha-helix (identified by Pauling and Corey) is stabilized by hydrogen bonds between the NHNH and COCO groups of peptide bonds located four residues apart. β\beta-pleated sheets involve zigzagging chains linked by lateral hydrogen bonds.

    • Tertiary Structure: The overall three-dimensional conformation of a single polypeptide. It is stabilized by RR group interactions, including hydrogen bonds, ionic bonds, disulfide bridges (sulfur bridges), and van der Waals forces. In aqueous environments, folding typically buries nonpolar groups internally.

    • Quaternary Structure: The organization of multisubunit proteins (e.g., hemoglobin, a heteromultimeric protein consisting of two α\alpha chains and two β\beta chains).

Characteristics and Deciphering of the Genetic Code

  • Triplet Nature: The code consists of four nucleotide bases (AA, CC, GG, UU). A three-letter (triplet) system provides 43=644^3 = 64 possible combinations (codons), which is sufficient to encode all 2020 amino acids.

  • Major Features:

    • The code is continuous (comma-free) and non-overlapping.

    • It is nearly universal across all domains of life, suggesting a common evolutionary origin.

    • It is degenerate (redundant), meaning multiple codons can specify the same amino acid, except for methionine (AUGAUG) and tryptophan (UGGUGG).

    • Sense vs. Nonsense: There are 6161 sense codons that specify amino acids and 33 nonsense (stop) codons (UAGUAG, UAAUAA, and UGAUGA) that signal the end of translation.

  • Experimental Deciphering: Research by Crick and colleagues using bacteriophage T4T4 mutants demonstrated that only insertions or deletions in multiples of three (frameshift reversals) could restore the reading frame, proving the triplet nature. Nirenberg and Khorana later determined specific codon assignments using cell-free systems and synthetic mRNAs (e.g., poly(UU) yielding polyphenylalanine).

  • Wobble Hypothesis: Proposed by Francis Crick, this explains why cells can function with fewer than 6161 distinct tRNAs. The base at the 55' end of the tRNA anticodon is less spatially constrained, allowing it to pair with multiple bases at the 33' end (third position) of the mRNA codon. For example, inosine (II) in the anticodon can pair with AA, UU, or CC.

Transfer RNA: The Molecular Adaptor

  • tRNAs are small RNA molecules (7575 to 9090 nucleotides) that translate codon sequences into amino acids. They feature a cloverleaf secondary structure and an L-shaped tertiary structure.

  • Key Structural Elements: The anticodon loop contains the triplet sequence that pairs with mRNA. The 33' end always ends in the sequence 5CCA35'-CCA-3', where the carboxyl group of the specific amino acid is covalently attached to the ribose OH-OH group.

  • Aminoacylation (Charging): The enzyme aminoacyl-tRNA synthetase (one specialized for each of the 2020 amino acids) catalyzes the attachment of the amino acid to the tRNA. This reaction is independent of the codon recognition; the specificity of translation depends on the tRNA-codon pairing, not the amino acid itself.

Ribosomal Structure and Function

  • Ribosomes are ribonucleoprotein complexes consisting of a large and small subunit.

    • Bacterial (70S70S): Comprises a 50S50S subunit (23S23S and 5S5S rRNAs) and a 30S30S subunit (16S16S rRNA).

    • Eukaryotic (80S80S): Comprises a 60S60S subunit (28S28S, 5.8S5.8S, and 5S5S rRNAs) and a 40S40S subunit (18S18S rRNA).

  • Functional Sites during Translation:

    • A (Aminoacyl) site: Entry point for the incoming charged tRNA.

    • P (Peptidyl) site: Holds the tRNA attached to the growing polypeptide chain.

    • E (Exit) site: Transient binding site for uncharged tRNAs leaving the ribosome.

The Three Stages of Translation

  • Initiation: In bacteria, the 30S30S subunit binds the Shine-Dalgarno sequence (5AGGAG35'-AGGAG-3') via the 16S16S rRNA. The initiator tRNA (fMettRNAfMetfMet-tRNA^{fMet}) binds the AUGAUG start codon. In eukaryotes, the 40S40S subunit binds the 55' cap and scans for the AUG within a Kozak sequence context.

  • Elongation: tRNA enters the A site (facilitated by EFTuEF-Tu in bacteria or eEF1AeEF-1A in eukaryotes). Peptidyl transferase, a ribozyme within the large rRNA, catalyzes the peptide bond. Translocation of the ribosome (11 codon toward the 33' end) requires EFGEF-G and GTPGTP hydrolysis.

  • Termination: Occurs when a stop codon (UAAUAA, UAGUAG, or UGAUGA) enters the A site. Release factors (RF1RF1/RF2RF2/RF3RF3) trigger the cleavage of the polypeptide. Ribosome recycling factors (RRF) then dismantle the translation complex.

Protein Selection and Signal-Based Localization

  • The Signal Hypothesis: Proteins destined for secretion or the endomembrane system possess an NN-terminal signal sequence (1515 to 3030 amino acids).

  • Protein Sorting Mechanism: A signal recognition particle (SRP) binds to the emerging signal peptide, halting translation until the ribosome docks to an SRP receptor on the endoplasmic reticulum (ER) membrane. Translation resumes, and the polypeptide enters the ER lumen through a translocon. The signal sequence is subsequently cleaved by signal peptidase.

Statistical and Probability-Based Genetic Code Analysis

  • Codon Probability Calculations: Given a random mixture of nucleotides, the frequency of any specific codon can be determined. For a mixture of 2U:1C2U:1C, the total options are 33.

    • P(U)=23P(U) = \frac{2}{3}

    • P(C)=13P(C) = \frac{1}{3}

    • Frequency of UUU=(23)3=82729.6%UUU = (\frac{2}{3})^3 = \frac{8}{27} \approx 29.6\%

  • Nucleotide Exclusion Probability: In a pool of four nucleotides (AA, UU, GG, CC), the probability that a codon lacks a specific base (e.g., CC) is (34)3=2764(\frac{3}{4})^3 = \frac{27}{64}. Consequently, the probability of a codon containing at least one CC is 12764=37641 - \frac{27}{64} = \frac{37}{64}.