Comprehensive Notes on Oncology NGS, Sequencing Quality, and Molecular Diagnostics Standards
Sequencing Diversity and Control with PhiX
Definition of PhiX: PhiX is a well-characterised control library derived from DNA of a small bacteriophage, a virus that infects bacteria. Its genome is approximately long and is completely known and well-characterised.
Role in Sequencing: PhiX is sequenced alongside patient DNA. Because it behaves like a normal sequencing library, the sequencer does not distinguish it as a control; it simply sequences the PhiX DNA fragments simultaneously.
Spiking in PhiX: PhiX is added to sequencing runs to increase sequence diversity. This is essential for cluster identification and quality calibration, particularly in low-diversity libraries.
Mechanism of Action: The sequencer identifies clusters and calls bases by measuring fluorescence at every cycle. In a low-diversity library, many clusters may show the same base at a specific cycle (e.g., all clusters show 'A'). This makes it difficult for the camera to distinguish individual clusters or calibrate fluorescence. Adding a small percentage (e.g., ) of PhiX introduces a mixture of bases (A, C, G, T) in the same cycle, allowing the instrument to:
Better distinguish neighbouring clusters.
Calibrate fluorescence more accurately.
Improve overall base calling quality.
Low-Diversity Libraries: Libraries prone to low diversity include amplicon sequencing, PCR products, targeted panels, and the early cycles of some sequencing runs where fragments may begin at the same genomic location. PhiX provides the random sequence diversity needed to overcome these limitations.
Quality Control (QC) Applications: Because the PhiX genome is known, it acts as an internal sequencing control. By comparing the expected PhiX sequence with the observed sequence after the run, laboratories can monitor if sequencing chemistry, optics, and base calling worked correctly. High error rates in the PhiX sequence indicate a potential run failure or instrument performance issue.
Quantity Constraints: Only a small percentage of PhiX is used to ensure there is sufficient capacity remaining for patient samples while still achieving high sequencing quality.
Platform Differences (Illumina vs. AVITI): Illumina optics and base-calling benefit significantly from PhiX in low-diversity runs. AVITI uses different sequencing chemistry and signal processing, generally requiring less PhiX. The specific amount is determined by the validated assay and laboratory Standard Operating Procedures (SOPs).
Sequencing Quality Metrics and Core Concepts
Q30 Metric: Q30 represents a sequencing quality score indicating a in probability of an incorrect base call, which corresponds to accuracy. High Q30 percentages reflect good sequencing quality and biological validity of identified variants.
Comprehensive QC Requirements: Q30 alone is insufficient for run approval. Laboratories must also evaluate:
Coverage/Depth: The number of sequencing reads covering a genomic position. Higher depth increases confidence in variant detection and sensitivity for low-frequency somatic variants.
Coverage Uniformity: Describes how evenly reads are distributed across target regions. Poor uniformity can lead to under-covered regions, increasing false-negative risks. Causes of poor uniformity include GC-rich regions, library preparation bias, and hybridisation efficiency.
On-target Rate: The percentage of sequencing reads mapping to intended enrichment regions. Low on-target rates indicate inefficient enrichment due to poor hybrid capture or probe issues.
Duplicate Reads: Multiple reads originating from the same DNA fragment, often caused by excessive PCR cycles or low DNA input. High duplication rates reduce library complexity and inflate apparent depth without adding biological information.
Insert Size: The length of the DNA fragment between sequencing adapters. Short inserts may indicate degraded DNA (e.g., FFPE), while excessively long inserts can decrease sequencing efficiency.
Internal Controls: These included within each run verify that stages like extraction, library prep, and bioinformatic analysis performed correctly.
AVITI Accuracy: While Q30 is standard, the AVITI platform also reports Q40, reflecting even higher base-calling accuracy.
Genomic Variations and Targeted Markers
Single Nucleotide Variants (SNVs): A change in a single DNA base (e.g., A > G). They are detected by aligning reads to a reference genome and identifying consistent differences supported by high-quality reads.
Small Insertions and Deletions (Indels): The addition or removal of small DNA segments. Algorithms use local realignment of reads with gaps or extra bases to determine the exact variant.
Copy Number Variants (CNVs): Large-scale gains or losses of genomic material. Detection relies on analyzing sequencing read depth across the genome. Low tumour content masks these signals as normal DNA dilutes the signal depth.
Tumour Mutational Burden (TMB): The number of somatic mutations per megabase () of coding DNA.
Calculation: TMB = .
Eligible mutations typically include non-synonymous SNVs and indels, while synonymous and germline variants are excluded.
Microsatellite Instability (MSI): Changes in length of short repeated DNA sequences caused by defects in the mismatch repair (MMR) system (proteins MLH1, MSH2, MSH6, PMS2). It is a diagnostic marker for Lynch syndrome and a predictor of response to immune checkpoint inhibitors.
Homologous Recombination Deficiency (HRD): A functional defect in double-strand break repair. It is considered a "genomic scar" rather than a single mutation because it represents the permanent abnormalities left by defective repair processes.
Clinical Utility: Predicts sensitivity to PARP inhibitors and platinum-based therapies.
HRD Components: Calculated based on Loss of Heterozygosity (LOH), Telomeric Allelic Imbalance (TAI), and Large Scale State Transitions (LST), plus pathogenic variants in BRCA1/2.
Threshold: If the total score exceeds a validated threshold (e.g., ), it is classified as HRD positive.
Loss of Heterozygosity (LOH): Occurs when one parental copy of a chromosomal region is lost. It is critical for the inactivation of tumour suppressor genes when the remaining copy is also mutated.
Methodologies: Hybrid Capture vs. Amplicon Sequencing
Hybrid Capture (e.g., OncoDEEP): Uses labelled probes to capture genomic regions of interest before sequencing.
Advantages: More uniform coverage, better detection of CNVs, large indels, and structural variants; fewer PCR biases.
Disadvantages: Longer workflow.
Amplicon Sequencing (e.g., Genexus): Uses PCR primers to amplify target regions.
Advantages: Faster (approx. ), requires less DNA input.
Disadvantages: Higher amplification bias, limited for CNVs and broad genomic profiling.
Anchored Multiplex PCR (AMP - e.g., Archer): Uses one gene-specific primer and one universal adapter.
Uniqueness: This allows for the detection of gene fusions even when the fusion partner is unknown (novel fusion partners), unlike traditional PCR where both primers must be known.
Specimen Considerations and Artefacts
FFPE Tissues: Formalin-fixed paraffin-embedded tissue preserves morphology but fragments DNA, creates DNA-protein cross-links, and causes cytosine deamination to uracil. This leads to false C>T or G>A artefacts.
Minimising FFPE Artefacts: Requires good fixation, DNA quality assessment (pre-library prep), bioinformatic filtering, and correlation with tumour content.
Tumour Cellularity and Purity: The proportion of tumour cells in a specimen. Low cellularity dilutes the Variant Allele Frequency (VAF), increasing false-negative risks and making CNV detection difficult.
Macrodissection: Enriches tumour content by removing surrounding normal tissue, thereby improving molecular testing sensitivity.
Variant Allele Frequency (VAF): The proportion of reads containing a variant at a specific position. Interpretation must account for tumour purity, ploidy, copy number state, and heterogeneity.
Clinical Diagnostics and Laboratory Standards
DNA vs. RNA for Fusions: RNA sequencing is more sensitive for detecting fusions because it avoids large intronic regions present in DNA and detects expressed transcripts directly.
FISH (Fluorescence In Situ Hybridisation): Still a gold standard for specific abnormalities like HER2 amplification (Breast Cancer) or ALK rearrangements (NSCLC). It allows direct visualisation of abnormalities in individual cells, but cannot detect SNVs, Indels, TMB, or MSI.
ctDNA (Liquid Biopsy): Analysis of cell-free DNA in blood. It is highly fragmented with very low VAF (often ). Requires deep sequencing and error-correction.
2026 ASCO Guidelines: ctDNA is used when tissue is unavailable; negative results do not exclude mutations.
ISO 15189: International quality standard for medical laboratories ensuring competence, validated processes, and patient safety.
Validation vs. Verification:
Validation: Proves a new method is fit for clinical purpose (Does the method work?).
Verification: Confirms an already validated method works reliably in a specific laboratory environment (Does it work here?).
Repeatability vs. Reproducibility:
Repeatability: Consistency under identical conditions (same operator/instrument).
Reproducibility: Consistency under varying conditions (different operator/reagent lots).
MDTs (Multidisciplinary Teams): Integration of molecular findings with clinical, radiological, and pathological data to support patient management.
Questions & Discussion
Why does low tumour content affect CNV detection? Copy number analysis depends on differences in sequencing read depth between tumour and normal reference DNA. When cellularity is low, DNA from normal cells masks these depth changes, making gains or losses less obvious and potentially uninterpretable.
What is the significance of Q30 in sequencing? It indicates a accuracy rate ( error in bases). It is a key metric but must be combined with coverage and uniformity to ensure clinical reliability.
Is ctDNA more difficult than tissue testing? Yes, because ctDNA is only a small fraction of total cell-free DNA. Variants are often at frequencies below , requiring much deeper sequencing and robust error suppression methods compared to tissue.
Why are duplicate reads important? High duplicate rates indicate that many reads are from the same original molecule, reducing the library's biological complexity and decreasing confidence in detecting low-frequency variants because there are fewer unique fragments.
How do we interpret VAF? VAF is not just a sign of mutation presence; it is influenced by tumour purity, ploidy, and heterogeneity. A mutation may have a low VAF due to low purity, or a high VAF due to a germline origin or loss of the wild-type allele.