Comprehensive Study Notes on Bioinformatics and DNA/RNA Sequencing Technologies
Foundations of Bioinformatic Sequencing and Experimental Planning
Sequence-Based Foundation of Modern Bioinformatics:
- DNA and RNA analyses are entirely sequence-based.
- Protein analysis historically relied on mass spectrometry, but protein sequencing is transitioning toward sequence-based methods. Though currently cost-prohibitive, sequence-based protein sequencing represents the future of the field.
Standard Bioinformatic Experimental Workflow:
- Experimental Planning: Comprehensive planning prior to wet-lab execution is mandatory due to extreme financial costs associated with sequencing reagents and platform runs.
- Sample Collection: Gathering target tissue, cellular, or environmental samples.
- Nucleic Acid Isolation: Extracting and purifying DNA or RNA from cellular structures.
- Library Preparation: Fragmenting nucleic acids and ligating specialized oligonucleotide adapters (ends) to enable binding to sequencer flow cells or pores.
- Sequencing: Running prepared libraries through specific sequencing hardware.
- Data Analysis: Processing, aligning, assembling, and interpreting raw sequencing reads.
Critical Experimental Parameters for Decision-Making:
- Biological Question: Dictates whether DNA or RNA must be isolated.
- Read Characteristics: Dictates sequencer choice based on required read depth (total read count) and read length (short vs. long reads).
- Preparation Kits:
- DNA vs. RNA Kits: Distinct biochemical protocols.
- Eukaryotic vs. Prokaryotic Sample Kits: Eukaryotic cells are fragile and lyse easily. Bacterial/proprokaryotic cells possess rigid cell walls requiring specialized enzymatic digestion to break open.
- Sequencing Depth: Refers to the total number of reads requested or the number of times a single base is sequenced. Depth dictates statistical power and experimental cost.
- Sample Comparison & Statistical Power: Determining comparison groups (e.g., diseased vs. healthy control) and statistical tests dictate minimum required sample sizes. Detecting small biological differences requires a larger number of replicates.
- Budgetary Constraints: Experimental design must constantly balance scientific idealization against financial capacity.
Applications of Sequencing in Bioinformatics
Genomics (Whole Genome Sequencing [WGS] & Whole Exome Sequencing [WES]):
- Clinical Diagnostics: WES isolated from non-cellular blood fractions is utilized to identify human genetic disorders.
- Variant Detection: Identifies point mutations (single base substitutions), indels (insertions or deletions), structural variations, and copy number variations.
- Bacterial Gene Duplication & Heteroresistance: Bacteria duplicate genes in real time, increasing antibiotic resistance through a mechanism termed heteroresistance.
- Lineage & Outbreak Tracking: Sequences distinguish exact strains during disease outbreaks (e.g., differentiating Escherichia coli serotype from serotype during foodborne outbreaks).
Transcriptomics:
- Real-Time Physiology: Quantifies real-time cellular gene expression, identifying up-regulated (e.g., antimicrobial resistance genes) or down-regulated pathways.
- Structural RNA Dynamics: Detects gene fusions, RNA editing (post-transcriptional alterations, truncation, or degradation into small RNAs observed in bacteria and humans), and alternative splicing.
- Splice Variants: While humans possess approximately coding genes, alternative exon splicing generates over distinct protein variants.
Epigenetics & Binding Analyses:
- Nucleic Acid Modification: Analyzes base modifications such as DNA methylation (addition of methyl groups), which regulates chromatin winding and gene silencing.
- ChIP-Seq (Chromatin Immunoprecipitation Sequencing): Maps genome-wide binding sites of transcription factors to identify expression regulatory networks.
- ATAC-Seq: Assesses chromatin accessibility and histone dynamics in eukaryotic genomes.
Metagenomics:
- Environmental & Population Profiling: Sequences whole microbial populations from complex samples (e.g., ocean water analysis reveals that the majority of marine organisms are bacteriophages; identifying atypical algal blooms).
- Human Microbiome Dynamics: Characterizes non-sterile body sites (e.g., human brain tissues contain resident viral populations; humans harbor more bacterial and viral entities than human cells). Used in clinical settings to evaluate stool donors for fecal microbiota transplants to treat inflammatory bowel disease (IBD) or Crohn's disease.
Genomic Analysis Pipelines: De Novo Assembly vs. Resequencing
Universal Downstream Pipelines:
- Bioinformatic workflows for genomics and transcriptomics converge downstream into two primary bioinformatic strategies: assembly or reference mapping.
De Novo Sequencing:
- Sequencing raw samples without a pre-existing reference genome.
- Reads are assembled computationally like a puzzle to build a novel genome construct from scratch for submission to databases such as NCBI.
Resequencing:
- Sequencing a sample and aligning (mapping) reads against a known reference genome blueprint.
- Used to determine genomic alterations, single nucleotide variants, or strain variations relative to a established reference.
Input Sample Considerations and Quality Control Procedures
Input Material Quality Criteria:
- Integrity: High-molecular-weight, intact genomic fragments are required. High levels of degraded, short fragments lead to poor sequencing output. Large fragments can be mechanically or enzymatically sheared down, but small degraded fragments cannot be reconstructed prior to sequencing.
- Quantity & Purity: Free of contaminating salts, proteins, or organic solvents.
Ribosomal RNA Depletion in Transcriptomics:
- Ribosomal RNA (rRNA) constitutes approximately of total cellular RNA.
- rRNA must be selectively depleted or removed during library preparation to allow sequencing coverage of messenger RNA (mRNA) and non-coding regulatory RNAs.
Multiplexing via Barcoding and Indexing:
- Allows pooling of up to or more distinct samples into a single sequencing run tube.
- During library preparation, unique oligonucleotide tags (barcodes, e.g., vs. ) are ligated to each sample's fragments. Demultiplexing algorithms sort raw reads post-run based on these unique sequences, dramatically lowering per-sample costs.
Technical Variations and Batch Effects:
- Systematic variations introduced by different experimentalists, slight procedural changes (e.g., vortexing time variations), reagent batches, or sequencer run dates.
Chemical and Biological Basis of DNA vs. RNA Stability
Structural Ribose Differences:
- DNA: Utilizes -deoxyribose, lacking a hydroxyl group at the carbon position of the sugar ring.
- RNA: Utilizes ribose, possessing a reactive hydroxyl () group at the position.
Auto-Cleavage Mechanism of RNA:
- The -hydroxyl group in RNA acts as an nucleophile, attacking the adjacent phosphodiester backbone at the phosphate atom.
- This intramolecular nucleophilic attack breaks the phosphodiester backbone (typically progressing in the or direction), leading to spontaneous self-degradation of RNA in solution.
- Due to RNA instability, sequencing technologies convert RNA to stable complementary DNA (cDNA) or utilize stabilized protocols to prevent degradation during run times.
Sanger Sequencing (First-Generation Platform)
Mechanism & Chemistry:
- Target Material: Single, pure DNA template fragment (e.g., single gene PCR product such as ribosomal genes, wmb7, or pknA).
- Dideoxynucleotides (ddNTPs): Uses modified nucleotides lacking a -hydroxyl group (). Incorporation of a ddNTP prevents further phosphodiester bond formation, causing chain termination.
- Four-Color Fluorescent Tagging: Modern Sanger uses a single reaction vessel containing normal dNTPs and lower concentrations of four distinct dye-labeled ddNTPs (ddATP, ddCTP, ddGTP, ddTTP).
- Fragment Generation: Thermal cycling generates a nested set of truncated DNA fragments representing every single nucleotide position along the target sequence.
Electrophoresis & Optical Detection:
- Fragments are separated by size using high-resolution capillary gel electrophoresis.
- Smaller fragments migrate rapidly through the capillary gel, passing a laser detector first, followed by progressively larger fragments.
- A laser excites the terminal fluorescent ddNTP on each passing fragment, recording bases sequentially: Green (), Blue (), Black (), and Red ().
Performance Metrics:
- Pros: Highest individual read accuracy; gold standard for targeted locus verification (e.g., validating engineered mutants vs. wild-type strains).
- Cons: Low throughput (one pure sequence at a time); high hands-on labor and cost per base.
Illumina Sequencing (Second-Generation Short-Read Platform)
Library Preparation & Flow Cell Hybridization:
- Genomic DNA is fragmented, and double-stranded oligonucleotide adapters are ligated to both ends.
- Single-stranded library fragments hybridize via complementary base pairing to lawn oligonucleotides tethered on an Illumina flow cell surface.
Bridge Amplification & Cluster Generation:
- A polymerase extends the flow cell oligo, copying the hybridized strand. The original template is denatured and washed away.
- The bound strand flexes over ("bridges") to hybridize with a adjacent complementary flow cell oligo.
- Isothermal amplification generates a localized, clonal cluster containing thousands () of identical DNA copies.
- Reverse strands are enzymatically cleaved and washed away, leaving linear single-stranded forward clusters for sequencing.
Sequencing-by-Synthesis (SBS) Chemistry:
- Reversible Dye Terminators: All four fluorescently tagged, -blocked dNTPs are added simultaneously.
- Single-Base Extension: Polymerase incorporates exactly one matching fluorescent dNTP into each growing strand across the cluster; extension stops due to the chemical block.
- Washing & Imaging: Unincorporated dNTPs are washed away. A laser excites the incorporated tags, and a high-resolution camera captures an image of the entire flow cell. Base calls are assigned based on cluster color signal.
- Cleavage Step: Chemical cleavage removes both the fluorescent fluorophore and the blocking group, restoring a functional group to accept the next base in the subsequent cycle.
Read Configuration Modes:
- Single-End (): Reads the sequence from one direction only.
- Paired-End ( and ): Reads forward (), regenerates the reverse strand on the flow cell, cleaves the forward strand, and reads backward (). Allows bridging across unsequenced middle gaps in fragments, providing critical spatial orientation for alignment and whole-genome assembly.
Pacific Biosciences (PacBio) SMRT Sequencing (Third-Generation Platform)
SMRTbell Library Preparation:
- Double-stranded DNA fragments are ligated to single-stranded hairpin adapters at both ends, creating a continuous, circular single-stranded DNA molecule (SMRTbell).
- Denaturation unfolds the double-stranded core into a closed circular loop.
Zero-Mode Waveguide (ZMW) Physics:
- Sequencing takes place on a chip containing thousands of ZMWs—microscopic metallic cylindrical wells fabricated over a glass substrate.
- The ZMW diameter is smaller than the wavelength of visible laser light. Laser light focused from below cannot fully pass through the aperture, creating an exponentially decaying zone of illumination (excitation volume) restricted strictly to the bottom few nanometers of the well.
Single-Molecule Real-Time Chemistry:
- A single DNA polymerase enzyme is permanently tethered to the bottom glass floor of each ZMW.
- Phospholinked fluorescent dNTPs (where the fluorophore is attached to the terminal phosphate rather than the base) diffuse into the ZMW.
- As the tethered polymerase incorporates a base, the dNTP is held in the detection volume for milliseconds, emitting a fluorescent signal captured by sensors.
- Cleavage of the phosphate chain during phosphodiester bond formation naturally releases the fluorophore, which diffuses away without requiring chemical cleavage steps.
Performance Metrics:
- Direct single-molecule sequencing without prior PCR amplification.
- Generates long continuous reads up to , resolving complex genomic structural variants and gene duplications.
Oxford Nanopore Sequencing (Third-Generation Platform)
Biological Nanopore Mechanism:
- Uses biological transmembrane protein pores embedded in an electrically resistant synthetic membrane sheet.
- An electrical potential is applied across the membrane, establishing a steady ionic current through the pore.
Electrochemical Base Translocation:
- Negatively charged nucleic acid strands (DNA or direct RNA) are drawn through the nanopore toward the positive potential.
- As the molecule passes single-file through the narrow constriction of the pore, it obstructs the ionic flow.
- Different base combinations (, , , , or ) cause distinct, characteristic disruptions in electrical resistance and ionic current ().
- Algorithms translate these real-time current fluctuations into nucleotide sequences.
Direct Modification and Epigenetic Detection:
- Because detection relies on electrical current disruption rather than base-pairing chemistry, Nanopore directly identifies modified bases (e.g., methylated adenine vs. unmethylated adenine) without bisulfite conversion.
Performance Metrics:
- Portability: MinION device is an ultra-portable handheld sequencer suitable for field research or bedside clinical diagnostic applications.
- Read Length: Ultra-long reads up to .
- Error Profile: Possesses a higher single-read error rate (primarily insertions and deletions [indels]) compared to Illumina.
Comprehensive Comparison of Sequencing Technologies
Read Length Standards:
- Illumina: Short reads (, , , ). Bacterial studies typically use or ; human studies require a minimum of to resolve highly homologous gene families. reads are considered obsolete.
- PacBio: Long reads up to .
- Oxford Nanopore: Ultra-long reads up to .
Accuracy Profiles and Error Compensation:
- Illumina: Extremely low error rate; highest overall base accuracy.
- PacBio & Nanopore: Higher single-read raw error rates, especially within homopolymeric regions (e.g., continuous poly-A stretches).
- Hybrid Assembly Strategy: High error rates of long-read platforms (PacBio/Nanopore) are routinely corrected by combining them with high-accuracy short reads (Illumina). Long reads establish structural framework and gene order; short reads map onto the assembly to fix point errors and base calls.
Obsolete Platforms:
- Pyrosequencing and Ion Torrent are obsolete due to severe homopolymer call errors.
Financial and Capital Infrastructure:
- Per-Run Costs: Single sequencing flow cell runs cost between and depending on depth and platform.
- Instrument Costs: Sequencers are capital-intensive (PacBio units are vending-machine-sized instruments costing hundreds of thousands of dollars).
- Maintenance: Annual service and maintenance contracts average approximately per instrument. Most academic research relies on specialized core facilities to execute runs.
Student Questions and In-Class Discussions
Quality Control and Sample Integrity Assessment (Question from Cooper):
- Question: How are the quality, quantity, and integrity of input nucleic acid samples measured prior to sequencing?
- Answer Details:
- Nanodrop: A spectrophotometric assay providing rapid, crude quantification and purity ratios (), but unreliable for precise concentration measurements.
- Qubit: A fluorometer that utilizes molecular dyes that fluoresce exclusively upon binding double-stranded DNA or RNA. Provides exact, highly sensitive concentration measurements.
- Bioanalyzer: An automated microfluidic electrophoresis instrument. It generates an electropherogram showing fragment size distribution, degradation levels, and RNA Integrity Numbers (RIN).
Illumina Homopolymer Resolution and Over-Clustering (Question from Will):
- Question: How does Illumina control for sequences containing consecutive identical nucleotides (homopolymers, e.g., poly-A runs)? Do multiple identical bases get added simultaneously in a single cycle?
- Answer Details:
- Chemical Block Mechanism: Illumina dNTPs are modified reversible terminators possessing a chemical block at the position. Even in a stretch of five identical adenine () bases, only a single modified dATP can bind to the growing strand per cycle. Extension halts completely until an enzyme cleaves the block, preventing multiple additions in one round.
- Homopolymer Error Drift: While chemically controlled, long homopolymer runs still show slightly higher error rates across all platforms.
- Over-Clustering Dynamics: If DNA library loading concentrations are too high, adjacent clusters form too close together on the flow cell. When two distinct clusters overlap spatially, their fluorescent signals bleed together during imaging (e.g., adjacent red and yellow clusters appearing green to the camera), causing severe base-calling errors and forcing the software to discard data.