Probabilistic Genotyping
Page 1: Title and Course
PROBABILISTIC GENOTYPING FRSC 370 F24 Gruhl
Page 2: What is Probabilistic Genotyping (PG)?
Definition: A method of analyzing DNA data using mathematical models to infer genotypes of contributors.
Applicability: Works on both mixed and single source DNA profiles.
Purpose: Calculate a likelihood ratio (LR) to support one of two proposed explanations of the data.
Page 3: Binary Models: Potential Issues
Example: Evidence sample contains 10 alleles; Person of Interest (POI) has alleles 9 and 10.
Question: Is it reasonable to assume the 9 allele might be present below the analytical threshold at any peak height?
Page 4: Why Move to Continuous Models?
Reasoning: Biology is not strictly binary (either/or), binary methods create a hard stop that is counter-intuitive.
Observations: Peak height variability increases as heights decrease; binary methods may not accurately reflect DNA levels.
Page 5: Why Move to Continuous Models? (Details)
Binary interpretation treats all genotype choices equally, which may not be justified by the data.
Increasing complexity of mixtures requires a better representation of the data.
Subjectivity from scientists influences interpretation.
Influence factors: Number of contributors, locus amplification efficiency, zygosity, degradation, peak heights, and proportion of contributors.
Page 6: Scientists and Software Interaction
Step 1: Scientist examines the DNA profile, identifies non-biological artifacts, and assigns the number of contributors (NoC).
Step 2: Software compares proposed genotypes to data and generates a report.
Step 3: Scientist reviews the report for quality and intuitiveness.
Page 7: Modeling STR Analysis (Variables - Part 1)
Template: Amount of DNA in the profile, indicated by peak heights.
Degradation: Observed through lower peak heights at higher molecular weights (exponential decay).
Locus-specific amplification effects are noted; performance varies among loci.
Page 8: Modeling STR Analysis (Variables - Part 2)
Zygosity: Determining if a locus is heterozygote or homozygote.
Stutter: Commonly occurs at N-1 position, usually <15%. Other stutter forms (N-2, N+1) are less common, typically 1-3%.
STRmix can model various stutter types.
Page 9: How Does STRmix Work?
Process: Runs all genotype combinations and eliminates those not fitting the data.
Mechanism: Random genotype selection for contributors to each locus and adjustment of mass parameters.
Change acceptance based on quality of data description by the proposed genotype is noted (Markov Chain Monte Carlo sampling - MCMC).
Page 10: Accepting Choices
Analogy: Accepting correct genotypes is akin to a game of hot and cold; prefer solutions that explain data well over poor ones (Metropolis-Hastings).
Page 11: What Happens Then?
Execution: Software runs for >100,000 iterations with a tally of accepted genotypes.
Implication: More accepted genotypes correlate with better data explanation.
When a standard matches, a likelihood ratio is calculated using the weights from consistent genotypes.
Page 12: Example Profile Analysis
D3S1358, vWA, D16S539 examples with peak heights and genotypes shown as evidence for analysis.
Page 13: In Reality
Factors Affecting Choices: Stochastic effects, allele heights, mixture ratios, overall EP ratios, and analyst experience.
Binary mixture analysis offers 50% peak height expectations which can misrepresent probabilities.
Page 14: Mock PG Mixture Analysis
Mixture proportion illustrated with weights assigned to Contributor 1 (Major) and Contributor 2 (Minor) for genotypes.
Page 15: Likelihood Ratios (LR)
Definition: Ratio of probabilities based on two hypotheses.
H1: DNA from the person of interest (POI).
H2: DNA from an unknown person.
Relationship with LR: LR increases with true H1 and decreases with true H2.
Page 16: Likelihood Ratios & PG
PG systems integrate weights for genotypes and sources of variability.
Example: LR calculated based on standard genotype weights and alternative hypothesis impacts.
Page 17: How Do We Know the Models "Work"?
Validation Types: Developmental validation studies and internal validation studies.
Page 18: Developmental Validation of PG Systems
Focus Areas: Stutter modeling, allelic drop-in/drop-out, Bayesian assumptions, and calculating algorithms.
Key Metrics: Sensitivity, specificity, precision, casework evaluation, and accuracy of calculations.
Page 19: Internal Validation of PG Systems
Internal validation ensures model variability and effectiveness for specific labs, amplification kits, and detection systems.
Page 20: Court Case Example
Judge's ruling on STRmix in Garrett Phillips's murder case and implications for the prosecution.
Page 21: Why Validations Matter
Details on admissibility issues in the case of People v. Hillary regarding the use of STRmix and adherence to legal standards.
Page 22: PG or Binary?
Comparison:
Binary: Fixed rules with subjectivity.
Probabilistic Genotyping: Incorporates complex biology using math, allowing for more data and less subjectivity.
Page 23: Looking for More Info?
Additional resources available for further reading about STRmix and relevant cases.