Comprehensive Study Guide to Gene Expression: Protein Translation and Structure
Fundamental Building Blocks: Protein Composition and Structure
Proteins are complex, high-molecular-weight organic compounds containing nitrogen. They are constructed from macromolecular subunits termed polypeptides, which are further composed of basic building blocks known as amino acids.
Chemical Anatomy of Amino Acids: Standard amino acids (excluding proline) share a common orientation centered on an -carbon atom. This central atom is bonded to a hydrogen atom, a carboxyl group (), an amino group (), and a variable group (side chain). Under physiological pH conditions (approximately ), the functional groups exist in ionized states: and .
The Group: The specific side chain attached to the -carbon determines the unique chemical properties of each of the different amino acids. These side chains are classified into four categories: acidic, basic, neutral/polar, and neutral/nonpolar.
Peptide Bond Formation: Polypeptides are formed through covalent peptide bonds, created by a dehydration reaction between the carboxyl group of one amino acid and the amino group of the next. This process releases a molecule of water (). Polypeptides have structural polarity, with a free amino group at the -terminus (the beginning of the chain) and a free carboxyl group at the -terminus.
Hierarchical Levels of Protein Structure:
Primary Structure: The unique linear sequence of amino acids, dictated directly by the nucleotide sequence of the corresponding gene.
Secondary Structure: Local spatial arrangements such as the -helix and -pleated sheet. The -helix (identified by Pauling and Corey) is stabilized by hydrogen bonds between the and groups of peptide bonds located four residues apart. -pleated sheets involve zigzagging chains linked by lateral hydrogen bonds.
Tertiary Structure: The overall three-dimensional conformation of a single polypeptide. It is stabilized by group interactions, including hydrogen bonds, ionic bonds, disulfide bridges (sulfur bridges), and van der Waals forces. In aqueous environments, folding typically buries nonpolar groups internally.
Quaternary Structure: The organization of multisubunit proteins (e.g., hemoglobin, a heteromultimeric protein consisting of two chains and two chains).
Characteristics and Deciphering of the Genetic Code
Triplet Nature: The code consists of four nucleotide bases (, , , ). A three-letter (triplet) system provides possible combinations (codons), which is sufficient to encode all amino acids.
Major Features:
The code is continuous (comma-free) and non-overlapping.
It is nearly universal across all domains of life, suggesting a common evolutionary origin.
It is degenerate (redundant), meaning multiple codons can specify the same amino acid, except for methionine () and tryptophan ().
Sense vs. Nonsense: There are sense codons that specify amino acids and nonsense (stop) codons (, , and ) that signal the end of translation.
Experimental Deciphering: Research by Crick and colleagues using bacteriophage mutants demonstrated that only insertions or deletions in multiples of three (frameshift reversals) could restore the reading frame, proving the triplet nature. Nirenberg and Khorana later determined specific codon assignments using cell-free systems and synthetic mRNAs (e.g., poly() yielding polyphenylalanine).
Wobble Hypothesis: Proposed by Francis Crick, this explains why cells can function with fewer than distinct tRNAs. The base at the end of the tRNA anticodon is less spatially constrained, allowing it to pair with multiple bases at the end (third position) of the mRNA codon. For example, inosine () in the anticodon can pair with , , or .
Transfer RNA: The Molecular Adaptor
tRNAs are small RNA molecules ( to nucleotides) that translate codon sequences into amino acids. They feature a cloverleaf secondary structure and an L-shaped tertiary structure.
Key Structural Elements: The anticodon loop contains the triplet sequence that pairs with mRNA. The end always ends in the sequence , where the carboxyl group of the specific amino acid is covalently attached to the ribose group.
Aminoacylation (Charging): The enzyme aminoacyl-tRNA synthetase (one specialized for each of the amino acids) catalyzes the attachment of the amino acid to the tRNA. This reaction is independent of the codon recognition; the specificity of translation depends on the tRNA-codon pairing, not the amino acid itself.
Ribosomal Structure and Function
Ribosomes are ribonucleoprotein complexes consisting of a large and small subunit.
Bacterial (): Comprises a subunit ( and rRNAs) and a subunit ( rRNA).
Eukaryotic (): Comprises a subunit (, , and rRNAs) and a subunit ( rRNA).
Functional Sites during Translation:
A (Aminoacyl) site: Entry point for the incoming charged tRNA.
P (Peptidyl) site: Holds the tRNA attached to the growing polypeptide chain.
E (Exit) site: Transient binding site for uncharged tRNAs leaving the ribosome.
The Three Stages of Translation
Initiation: In bacteria, the subunit binds the Shine-Dalgarno sequence () via the rRNA. The initiator tRNA () binds the start codon. In eukaryotes, the subunit binds the cap and scans for the AUG within a Kozak sequence context.
Elongation: tRNA enters the A site (facilitated by in bacteria or in eukaryotes). Peptidyl transferase, a ribozyme within the large rRNA, catalyzes the peptide bond. Translocation of the ribosome ( codon toward the end) requires and hydrolysis.
Termination: Occurs when a stop codon (, , or ) enters the A site. Release factors (//) trigger the cleavage of the polypeptide. Ribosome recycling factors (RRF) then dismantle the translation complex.
Protein Selection and Signal-Based Localization
The Signal Hypothesis: Proteins destined for secretion or the endomembrane system possess an -terminal signal sequence ( to amino acids).
Protein Sorting Mechanism: A signal recognition particle (SRP) binds to the emerging signal peptide, halting translation until the ribosome docks to an SRP receptor on the endoplasmic reticulum (ER) membrane. Translation resumes, and the polypeptide enters the ER lumen through a translocon. The signal sequence is subsequently cleaved by signal peptidase.
Statistical and Probability-Based Genetic Code Analysis
Codon Probability Calculations: Given a random mixture of nucleotides, the frequency of any specific codon can be determined. For a mixture of , the total options are .
Frequency of
Nucleotide Exclusion Probability: In a pool of four nucleotides (, , , ), the probability that a codon lacks a specific base (e.g., ) is . Consequently, the probability of a codon containing at least one is .