Probabilistic Genotyping

Page 1: Title and Course

  • PROBABILISTIC GENOTYPING FRSC 370 F24 Gruhl

Page 2: What is Probabilistic Genotyping (PG)?

  • Definition: A method of analyzing DNA data using mathematical models to infer genotypes of contributors.

  • Applicability: Works on both mixed and single source DNA profiles.

  • Purpose: Calculate a likelihood ratio (LR) to support one of two proposed explanations of the data.

Page 3: Binary Models: Potential Issues

  • Example: Evidence sample contains 10 alleles; Person of Interest (POI) has alleles 9 and 10.

  • Question: Is it reasonable to assume the 9 allele might be present below the analytical threshold at any peak height?

Page 4: Why Move to Continuous Models?

  • Reasoning: Biology is not strictly binary (either/or), binary methods create a hard stop that is counter-intuitive.

  • Observations: Peak height variability increases as heights decrease; binary methods may not accurately reflect DNA levels.

Page 5: Why Move to Continuous Models? (Details)

  1. Binary interpretation treats all genotype choices equally, which may not be justified by the data.

  2. Increasing complexity of mixtures requires a better representation of the data.

  3. Subjectivity from scientists influences interpretation.

    • Influence factors: Number of contributors, locus amplification efficiency, zygosity, degradation, peak heights, and proportion of contributors.

Page 6: Scientists and Software Interaction

  • Step 1: Scientist examines the DNA profile, identifies non-biological artifacts, and assigns the number of contributors (NoC).

  • Step 2: Software compares proposed genotypes to data and generates a report.

  • Step 3: Scientist reviews the report for quality and intuitiveness.

Page 7: Modeling STR Analysis (Variables - Part 1)

  • Template: Amount of DNA in the profile, indicated by peak heights.

  • Degradation: Observed through lower peak heights at higher molecular weights (exponential decay).

  • Locus-specific amplification effects are noted; performance varies among loci.

Page 8: Modeling STR Analysis (Variables - Part 2)

  • Zygosity: Determining if a locus is heterozygote or homozygote.

  • Stutter: Commonly occurs at N-1 position, usually <15%. Other stutter forms (N-2, N+1) are less common, typically 1-3%.

  • STRmix can model various stutter types.

Page 9: How Does STRmix Work?

  • Process: Runs all genotype combinations and eliminates those not fitting the data.

  • Mechanism: Random genotype selection for contributors to each locus and adjustment of mass parameters.

  • Change acceptance based on quality of data description by the proposed genotype is noted (Markov Chain Monte Carlo sampling - MCMC).

Page 10: Accepting Choices

  • Analogy: Accepting correct genotypes is akin to a game of hot and cold; prefer solutions that explain data well over poor ones (Metropolis-Hastings).

Page 11: What Happens Then?

  • Execution: Software runs for >100,000 iterations with a tally of accepted genotypes.

  • Implication: More accepted genotypes correlate with better data explanation.

  • When a standard matches, a likelihood ratio is calculated using the weights from consistent genotypes.

Page 12: Example Profile Analysis

  • D3S1358, vWA, D16S539 examples with peak heights and genotypes shown as evidence for analysis.

Page 13: In Reality

  • Factors Affecting Choices: Stochastic effects, allele heights, mixture ratios, overall EP ratios, and analyst experience.

  • Binary mixture analysis offers 50% peak height expectations which can misrepresent probabilities.

Page 14: Mock PG Mixture Analysis

  • Mixture proportion illustrated with weights assigned to Contributor 1 (Major) and Contributor 2 (Minor) for genotypes.

Page 15: Likelihood Ratios (LR)

  • Definition: Ratio of probabilities based on two hypotheses.

    • H1: DNA from the person of interest (POI).

    • H2: DNA from an unknown person.

  • Relationship with LR: LR increases with true H1 and decreases with true H2.

Page 16: Likelihood Ratios & PG

  • PG systems integrate weights for genotypes and sources of variability.

  • Example: LR calculated based on standard genotype weights and alternative hypothesis impacts.

Page 17: How Do We Know the Models "Work"?

  • Validation Types: Developmental validation studies and internal validation studies.

Page 18: Developmental Validation of PG Systems

  • Focus Areas: Stutter modeling, allelic drop-in/drop-out, Bayesian assumptions, and calculating algorithms.

  • Key Metrics: Sensitivity, specificity, precision, casework evaluation, and accuracy of calculations.

Page 19: Internal Validation of PG Systems

  • Internal validation ensures model variability and effectiveness for specific labs, amplification kits, and detection systems.

Page 20: Court Case Example

  • Judge's ruling on STRmix in Garrett Phillips's murder case and implications for the prosecution.

Page 21: Why Validations Matter

  • Details on admissibility issues in the case of People v. Hillary regarding the use of STRmix and adherence to legal standards.

Page 22: PG or Binary?

  • Comparison:

    • Binary: Fixed rules with subjectivity.

    • Probabilistic Genotyping: Incorporates complex biology using math, allowing for more data and less subjectivity.

Page 23: Looking for More Info?

  • Additional resources available for further reading about STRmix and relevant cases.