T8 Microbial Genetics
Microbial Genetics, Genomics, and Other -Omics
Mechanisms of Genetic Variation
**Genetic changes occur through: intrinsic problems**
Mutations can occur through intrinsic process like errors in replication. Different organisms have different mutation rates based on this intrinsic replication:
Orgs with DNA repair system: Higher eukaryotes, lower eukaryotes and bacteria have low mutation rates of approximately nucleotides.
DNA-based viruses: Higher mutation rates observed in single-stranded DNA viruses o double-stranded DNA viruses with rates of about
RNA viruses and viroids exhibit even higher rates due to lack of RNA repair mechanism.
Gene duplication.
Gene loss.
Insertions by mobile elements.
Horizontal gene transfer (HGT)
Types of mutations:
Neutral: no significant effect.
Deleterious: have negative consequences.
Beneficial: allow organisms to adapt and thrive in environments.
Research Reference
Constraints on Genetic Variation
Not all mutations affect fitness (neutral mutations).
Some mutations are not tolerated by the organism (severely deleterious).
Adaptive mutations improve the fitness of an organism, contributing to survival (beneficial).
These concepts aid in examining gene evolution over time.
Evolutionary Processes
Mechanisms of Evolutionary Adaptation in Microbes
Differentiation of mutation effects based on organism type (haploid vs. diploid).
Haploid:
Evolution occurs very quickly in haploid microbes with fast generation times and has strong effects.
no masking effect; the mutation is always expressed
vulnerable to deleterious mutation
Diploid:
Masking effect; recessive mutations are hidden by the dominant allele
low vulnerability
slower adaptation speed
Horizontal gene transfer: Driven by mutation and genetic exchange
Modes of HGT include:
Transformation
Conjugation
Transduction
Transformation
General Concepts
Definition: Genetic transfer process where naked DNA is assimilated by a recipient cell.
First discovered by Frederick Griffith in the late 1920s using Streptococcus pneumoniae.
This work contributed to the discovery of DNA.
Mechanisms include competence (naturally by some bacteria) and induced competence through methods like electroporation or chemical treatments.
Natural Competence
Competent Cells: Capable of taking up DNA and being transformed.
DNA binding proteins and nucleases are usually involved in this process.
Examples of naturally competent bacteria include certain strains of Bacillus and Neisseria.
Induced Competence
Cells can be made competent through methods like electroporation or heat-shock.
These procedures are often transient (cell membrane is temporarily permeable) and require specific conditions (e.g., presence of cations like CaCl2).
Conjugation
Mechanism of Genetic Transfer
Definition: Direct cell-to-cell contact process to transfer DNA, typically involving plasmid-encoded mechanisms.
Components:
Donor Cell: Contains the conjugative plasmid.
Recipient Cell: Lacks the plasmid.
Essential for conjugation: Sex pilus, produced exclusively by the donor cell.
Process Steps
Pilus Extension and Retraction: The F+ (donor) cell extends a pilus to contact the F- (recipient) cell, then retracts to draw the cells together
F Plasmid Processing: The F plasmid is nicked at the origin of transfer (oriT) by the relaxase enzyme.
DNA Transfer: A single strand of the plasmid (Donated strand) is transferred to the recipient cell.
Replication: Both the donor and recipient cells synthesize a complementary strand (using the DNA polymerase shown) to create a double-stranded F plasmid in both cells.
Result: The final step shows the recipient cell F- being converted into a donor cell F+.
Transduction
Definition: DNA is transferred from one bacterium to another specifically in the form of chromosome fragments or pieces of plasmid.
The transferred genetic material is Short linear DNA only.
Short: The DNA is a fragment (a short piece) of the original bacterial chromosome, limited by the volume constraints of the bacteriophage head. The bacteriophage cannot package the massive, entire circular chromosome.
Linear: Unlike a plasmid, which is circular and can replicate on its own, the transferred DNA piece is linear (a straight line).
The transfer of genetic material during bacterial transduction requires homologous recombination.
Genetic Transfer via Bacteriophage
Generalized Transduction: The bacteriophage accidentally packages random fragments of the host bacterium’s chromosomal DNA or plasmid DNA into a virion, especially by lytic viruses.
Specialized Transduction: Specific regions of the host chromosome are integrated into the virus genome, typically by lysogenic viruses. This occurs due to an imprecise excision of the prophage from the bacterial chromosome during the lysogenic cycle, leading to the transfer of adjacent genes along with the viral DNA to the new host.
Four Fates of Foreign DNA Inside Cells
What happens to a piece of new DNA (acquired through transformation, transduction, or conjugation) once it enters a bacterial cell. :
1. 🗑 Fate: Degradation (Destruction)
This is the ultimate rejection of the foreign DNA. The cell sees the DNA as a threat or simply useless junk.
How it works: The cell uses its own internal defense systems—a kind of molecular immune system—to chop up and destroy the foreign DNA.
Defense Systems:
RMS systems: This stands for Restriction-Modification Systems. The cell's own DNA is "stamped" with a chemical modification (like methylation). Foreign DNA lacks this stamp, so a restriction enzyme recognizes it as foreign and cuts it into pieces.
CRISPR systems: This acts like a genetic memory. If the cell has been attacked by this type of DNA (usually from a virus) before, it keeps a "mugshot" (a short DNA sequence) of the invader. If it sees that DNA again, it uses the CRISPR system to target and destroy it.
Result: The DNA is Degraded by cellular enzymes and its components are recycled
2. 🔀 Fate: Site-Specific Recombination (Precise Integration)
In this case, the foreign DNA is accepted, but only if it fits into a very specific, pre-determined spot.
How it works: This process is like using a specific type of key to unlock a door.
The foreign DNA needs a specific recognition site that matches a location on the host genome.
It requires special enzymes called recombinases to perform the integration.
Examples: Some bacteriophage DNA and certain mobile elements use this to integrate with high precision.
Result: The foreign DNA is integrated precisely into the chromosome via Site-specific recombination into genome.
3. 🧩 Fate: Homologous Recombination (Slotting in a New Part)
The foreign DNA is swapped in to replace a section of the existing chromosome.
How it works: This is like replacing a used car part with a new one—they must be the same model.
The Foreign DNA needs to have some similarity to the host genome (it must be homologous).
The cell uses the Homologous Recombination machinery (like the RecA protein) to search for the matching sequence and swap the incoming DNA in.
Result: The new genetic information is integrated through homologous recombination and becomes a stable part of the chromosome.
4. 🧳 Fate: Stable Maintenance (Living Outside)
In this scenario, the foreign DNA doesn't join the chromosome; it stays separate but becomes a permanent resident.
How it works: The DNA must be a replicon—a self-replicating unit—and it must be compatible with the host cell.
Need to have replication origins recognized by/compatible with the host. This ensures the cell copies the foreign DNA every time the cell divides.
Mainly plasmids use this strategy, as they are small, circular, and carry their own replication start points.
Result: The DNA achieves stable maintenance outside the chromosome, replicating alongside the main genome.
Homologous Recombination
Mechanism Description
Genetic exchange between similar or identical DNA sequences (homology) derived from different sources.
Mediated by the RecA protein in prokaryotic cells.
Important processes and components: RecBCD, RuvABC, RecG, all associated with homologous recombination mechanisms.
Gene Transfer in Archaea and Eukaryotes
Insights into Archaea
Transformation exists in Archaea, but genetic manipulation has not progressed as much as in bacteria.
Eukaryotic Microbes
In ciliates (euk), conjugation is characterized by exchanged haploid nuclei over a cytoplasmic bridge. Genetic exchange goes in both directions.
Introduction to Genomics
Definitions
Genome: Complete genetic information, including genes and non-coding segments.
Genomics: The study of mapping, sequencing, analyzing, and comparing genomes, with 442,000 prokaryotic genomes sequenced or in progress.
Steps to Sequence a Genome
Sequence the DNA.
Assemble sequences into chromosomes (contigs).
Predict gene functions (annotation).
Sequencing and Annotating Genomes
Sequencing Definition
Definition of Sequencing: Determining the precise order of nucleotides in a DNA molecule.
Methods
Sanger Dideoxy Method: Utilizes dideoxy analogs of dNTPs (ddNTPs) used in conjunction with dNTPs to prevent extension of the DNA chain; provides high accuracy (99.9999%). ddNTPs act as chain terminators by lacking a 3'-hydroxyl group. Highest accuracy
Reaction Mixture: The process starts with a reaction mixture containing several essential components.
Primer Elongation and Chain Termination
The DNA polymerase extends the primer by adding dNTPs to synthesize a new DNA strand complementary to the template.
However, because both dNTPs and ddNTPs are present, a ddNTP is occasionally and randomly incorporated instead of a dNTP.
When a ddNTP is incorporated, the chain stops, resulting in a pool of DNA fragments of all possible lengths, with each fragment ending in a fluorescently labeled ddNTP corresponding to the base at that position on the template DNA.
Separation and Detection
Capillary Gel Electrophoresis (CGE): The resulting DNA fragments are separated by size using an automated sequencer. The fragments are loaded into a very long, thin tube (capillary) filled with a gel matrix. A high-voltage electric current pulls the negatively charged DNA fragments through the capillary.
Separation by Size: Shorter fragments travel faster and exit the capillary before longer fragments.
Laser Detection and Computational Sequence Analysis: As each fragment passes a specific point, a laser excites the fluorescent dye at the fragment's end. A detector reads the color, and the data is recorded as a chromatogram (a series of colored peaks). The sequence of colors, from shortest fragment to longest, reveals the complementary DNA sequence.
📝 Key Characteristics
The text on the right side of the slide summarizes the key features of Sanger sequencing:
900-1100bp for 3-5: This highlights the cost-effectiveness and typical read length. The maximum length of a high-quality sequence read is generally around 800 to 1,000 base pairs (bp) with a single run. The price is an estimate for sequencing a single sample.
Only for 1 pure template: Sanger sequencing requires a very clean, single DNA template (like a purified PCR product or plasmid) to produce a clear, readable sequence. It is not suitable for sequencing a mixed sample of DNA templates.
Very high accuracy (99.9999%): Sanger sequencing is considered the "gold standard" of DNA sequencing due to its low error rate, often cited as high as 99.99% accurate.
Common machines: This refers to the automated capillary electrophoresis DNA sequencers that are widely used in labs, such as models from Applied Biosystems (ABI).
Keywords to recognize sanger:
Validation, confirmation, single mutation, small variant, short fragment, high accuracy, gold standard
Clues: used for confirming specific SNPs, small insertions/deletions or gene variants.
Typical phrase: “confirm a single base change”
Next Generation Sequencing (NGS): Allows for parallel sequencing, sequences mixed samples but shorter read lengths compared to Sanger.
Specific Equipment and Techniques:
454 pyrosequencing: no longer supported. Short read, second to most expensive
Illumina XTen sequencing: for mixed DNA samples, complex samples, human/euk/microbial genomes, transcriptsomes, and SNP mapping. Short read, most expensive, second to highest accuracy, greatest number of reads/runs
Keywords: short read, high throughput, massive parallel sequencing, SNP discovery, resequencing, transcriptomics, RNA-seq
Clues: produces many short, accurate reads (100-300bp). Excellent for detecting SNPs but not for large structural variants.
Typical phrase: “high accuracy short reads”
PacBio Revio sequencing: closing genomes, operon/splice variant mapping, methylation of DNA/RNA, SNP mapping. Long read, second to most expensive
keywords: single molecule real-time (SMRT), long reads, HiFi reads, high accuracy, structural variation, repetitive regions
clues: used for long-read sequencing (10-25kb) and resolving repeats or structural variants
typical phrase: “high-fidelity long reads’
Nanopore Minion sequencing: presence/absence of species/genes field work. Long read, less powerful, can measure electrical current. Accuracy issue, least expensive
keywords: ultra-long reads, portable sequencer, real-time sequencing, detect large insertions/deletions, field sequencing
clues: best for ultra-long reads and structural variation detection
typical phrase: “ultra-long reads for genome assembly”
Genomic words for these methods:
Read: a single sequence of a piece of DNA
Run: a single (usually mixed) sample that is sequenced by a given technology
Genome Assembly
Aligning all sequence reads to create longer contigs with software, addressing gaps or may remain as “draft” genomes, and completing genome assembly through PCR and Sanger sequencing of products.
The most laborious and expensive part of genome sequencing
Functional Genomics and Bioinformatics
Genome Annotation
Definition of Annotation: Converting raw data into a list of genes in a genome, including functional open reading frames (ORFs).
1. What is the Metagenome?
The slide starts by defining the core term: Metagenome [Slide 36].
Simple Meaning: It is the total gene content of all the organisms present in a mixed, environmental sample.
Analogy: Instead of studying a single type of organism in a lab dish, metagenomics is like taking a scoop of dirt or water and analyzing every piece of DNA from every bacterium, virus, and microbe living there all at once. It gives you the complete genetic blueprint of an entire microbial community.
2. How is it Used?
The rest of the slide explains where this technique is applied, showing that large-scale metagenome projects have surveyed many complex, real-world environments [Slide 36].
Goal: To understand what microbes are present, what genes they have, and what functions they are performing in their natural habitat, often without ever growing them in a lab.
Examples of Environments Surveyed:
Acid mine runoff waters: Studying the extremeophiles that thrive in highly acidic, toxic environments.
Deep-sea sediments: Exploring the unique microbes that live under crushing pressure and in complete darkness.
Fertile soils: Understanding the vast diversity of microbes that are crucial for plant growth and nutrient cycling.
Body sites: Analyzing the human microbiome (gut, skin, etc.) to link microbial communities to health and disease.
3. Whole Metagenome Sequencing (Shotgun Sequencing)
This technique sequences everything in the sample.
How it works: You sequence the entire complement of DNA in a mixed sample (e.g., all the genes from every organism) [Slide 37]. No initial PCR step is required.
What you get:
Functional ORFs (Open Reading Frames): You can identify and reconstruct entire genes, such as those responsible for nitrogen fixation genes or antibiotic resistance [Slide 37].
Tells you the diversity of these genes in a sample: You understand the potential functions of the entire microbial community.
4. Amplicon Sequencing (Targeted Sequencing)
This technique focuses on sequencing only one specific gene across all the organisms in the sample.
How it works: You must use PCR using universal primers for genes of interest before sequencing [Slide 37].
A common example is "16S metagenomics," which targets the 16S gene of bacteria or archaea [Slide 37].
This requires universal primers and a PCR step before sequencing to amplify only that specific gene from all the different organisms present [Slide 37].
What you get:
The diversity of the community: The 16Sc ,’ gene acts as a molecular barcode to identify which microbial species are present and in what relative amounts, telling you the who of the community.
5. Metagenomes: What They Can Tell You
A metagenome is great for structural analysis—looking at all the DNA in a sample (the potential genetic blueprint).
Sequence of DNA in a sample: You can read the entire genetic code of all organisms present [Slide 38].
Diversity information: You can figure out which types of microbes are present [Slide 38].
Relative quantities of community members: You can estimate how much DNA belongs to each organism, giving you a census of the community [Slide 38].
6. Metagenomes: What They Can't Tell You (Limitations)
The major limitation of metagenomics is that DNA is static; it doesn't tell you what a microbe is actually doing at any given moment.
Useful members of a community: You can't tell if a microbe that has an interesting gene is actually the most active or important member [Slide 38].
Who is alive vs. dead vs. dormant: DNA from dead cells, dormant cells, or living cells all look the same in a metagenome. You don't know which microbes are currently active [Slide 38].
What genes/pathways are being expressed in an environment: This is the biggest gap. A microbe may have a gene for antibiotic resistance, but is it currently turned on? The metagenome can't answer this [Slide 38].
The RNA-seq Workflow
The slide lists the key steps to perform an RNA-seq experiment:
Collect all the RNA from a sample: The process starts by quickly isolating all the RNA (which represents the active genes) from the bacteria you are studying.
Convert to cDNA using RT: Since RNA is fragile, it must be converted into stable complementary DNA (cDNA) using an enzyme called Reverse Transcriptase (RT). This cDNA is what is actually sequenced.
Sequence it using next-gen sequencing: The cDNA molecules are run through a high-throughput sequencing machine (like those discussed in the metagenomics slides) to generate millions of short sequences.
Align sequences to genome: These short sequence reads are then mapped back to the known reference genome of the organism being studied.
Count reads for each gene: The most critical step is counting. The number of sequence reads that align to a specific gene is a direct measure of how much of that gene's RNA was present in the sample. A higher count means the gene was more active.
Compare between conditions, species, or mutants and WT (Wild Type): The final, analytical step is comparing the counts between different experimental groups (e.g., bacteria grown in conditions vs. soil conditions).
🧬 Transcriptomics (What Could Get Made)
Transcriptomics focuses on RNA, which is the intermediate product of active genes.
Definition: Measures all the transcripts (RNA molecules) present in cells under a given condition [Slide 42].
Key Insight: Tells you What could get made (which genes are turned on and producing an RNA blueprint) [Slide 42].
Unique Value: It is the Only way to get info on non-coding RNAs (ncRNAs, RNAs, rRNAs, and tRNAs that are made, as these molecules perform functions as RNA themselves and are never translated into protein [Slide 42].
Limitation: The level of RNA does not always equal the level of functional protein. Post-transcriptional, post-translational modifications, and degradation of transcripts may influence protein levels [Slide 42]. An RNA molecule may be made, but the protein might be destroyed immediately, or never fully finished.
🧪 Proteomics (What Does Get Made)
Proteomics focuses on protein, which is the final, active product of a gene.
Definition: Measures all the proteins that are present in cells under a given condition [Slide 42].
Key Insight: Tells you What does get made [Slide 42]. This directly measures the molecules that carry out the cell's function.
Value: It is generally More informative (for most research questions) because it measures the true functional molecules, including all modifications that affect their activity [Slide 42].
Limitation: It is Much more difficult to do [Slide 42]. The techniques for isolating, separating, and identifying proteins are complex.
Measurement Methods: There are Multiple ways to measure proteins, including older methods like 2D-gels and modern, high-resolution techniques like Mass Spec (Mass Spectrometry) [Slide 42].
🌐 The Interactome Explained
1. What is an Interactome?
The slide defines the core term:
Definition: An interactome is the complete set of interactions among molecules within a cell or organism [Slide 43].
Analogy: Think of it as the social network diagram for all the molecules in a cell. It shows who is talking to whom, who is turning whom on or off, and how all the different parts work together to execute a function.
2. Components and Data Expression
The data generated from studying an interactome is organized into network diagrams that show the connections between different types of molecules:
DNA – RNA networks: Show how the genome regulates which transcripts (RNA) are made.
Protein – protein networks: Show how proteins physically interact with each other to form complexes or signaling pathways.
Protein – DNA networks: Show how proteins (like transcription factors) bind to DNA to turn specific genes on or off.
Data Format: The results are typically expressed visually in the form of network diagrams [Slide 43]. The image on the right of the slide is an example of these complex network diagrams.
3. Complexity of Generation
Studying the interactome is difficult because you must combine data from all the previous '-omics' methods:
Requirement: It may require several –omics methods and other techniques to generate [Slide 43]. You need:
Genomics/Metagenomics (to know the parts/DNA).
Transcriptomics (RNA-seq) (to know which genes are active).
Proteomics (to know which functional molecules are present).
🧪 Metabolomics Explained
1. What is the Metabolome?
Metabolomics is defined as the study of the complete set of metabolic intermediates and other small molecules produced in an organism [Slide 44].
Analogy: If genomics is the blueprint (DNA) and transcriptomics is the electrical wiring diagram (RNA), metabolomics is the exhaust coming out of the engine. It tells you exactly what the cell's machinery is burning and producing right now.
Key Insight: These molecules are the Products of anabolism or catabolism—meaning they are the compounds created or broken down as the cell does its job [Slide 44].
2. Types of Metabolites
The slide divides these small molecules into two major categories:
Primary metabolites: These are essential for growth, development, or reproduction of the host [Slide 44]. They are the fundamental building blocks and energy sources, such as amino acids, sugars, and organic acids.
Secondary metabolites: These are molecules that are specific for the environment and often give the organism a competitive advantage [Slide 44].
Examples include Antibiotics (used to kill competitors) and Other drugs [Slide 44].
3. How Metabolites are Measured
Because these molecules are diverse, tiny, and present in very low concentrations, highly sensitive technology is required:
Primary Technique: Mass spectrometry is one of the primary techniques for monitoring metabolites [Slide 44]. This instrument can separate and identify compounds based on their mass-to-charge ratio, providing a highly accurate chemical fingerprint of the cell's current state.