1/59
the 7 papers sebastian sent (ProteinMPNN, HERMES, RV3.0, ProteinDPO, Marks specificity, MULTI-evolve, additive critique). fact checked against the sources. TRAP cards = where the obvious summary is wrong. last 4 cards are figures to explain.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
what does ProteinMPNN do?
inverse folding: backbone in, sequence out. ~1.7M param message passing graph net, the default sequence design step at IPD
ProteinMPNN vs Rosetta, the two numbers
52.4% native sequence recovery vs 32.9%, and 1.2 seconds vs 4.3 minutes per 100 residues. faster AND more accurate
cleanest ProteinMPNN wet lab comparison (use this one, not the cages)
cyclic homo-oligomers on identical backbones: rosetta 40% soluble and 0% correct oligomeric state, MPNN 88% soluble and 27.7% correct
why train ProteinMPNN with backbone noise?
0.02 A gaussian noise. costs recovery on pristine crystal structures (50.5 to 47.3%) but buys robustness on computer generated backbones (47.9% on AF models)
ProteinMPNN sequence recovery by burial
90-95% in the core, ~35% on the surface. for a nanoparticle immunogen the surface is exactly where solubility and aggregation live
TRAP: what did ProteinMPNN actually do for the tetrahedral cages?
76 sequences across 27 two component backbones, 13 assembled at ~1 MDa, 1 crystal structure at 1.2 A. success rate was SIMILAR to rosetta. the win is automation (~1 s vs >1 week of manual intervention) plus a few new assemblies rosetta had missed. do NOT call it a rescue
how does ProteinMPNN let you freeze an epitope or enforce symmetry?
order agnostic autoregressive decoding plus tied positions (same residue at all 12 symmetry equivalent subunit copies)
TRAP: the NTF2 "rescue" in ProteinMPNN
in silico only. improved AlphaFold single sequence prediction over 3,000 backbones. no experimental rescue
what is HERMES in one idea?
SO(3) equivariant net that sees only the 10 A ball of atoms around one residue and scores mutations as log p(mutant) minus log p(wild type). zero shot
what does HERMES get for free from its formulation?
thermodynamic reversibility and path independence. Stability-Oracle needs 19x "thermodynamic permutation" data augmentation to enforce the same thing
HERMES: the one problem and the fix
scoring in the rigid wild type structure biases toward same size substitutions (the cavity is shaped like the WT residue). PyRosetta relaxation fixes it, recall for stabilizing mutations 0.48 vs 0.27, but is ~66x slower
what is amortization (the trick worth stealing)?
pay for the slow physics once: run relaxation on ~15k neighborhoods, 0.5% of pretraining sites, then train the fast model to imitate its own relaxed answers. most of the benefit at full speed
HERMES antigen benchmark numbers
33 known stabilizing mutations across 5 antigens (flu HA, RSV-F, hMPV-F, DENV-E, SARS-2 spike). ranked 24 above WT, 19 in the top 3 at their site, 7 of 8 stabilizing prolines
TRAP: did HERMES beat Rosetta on the antigen benchmark?
no. rosetta was on par, and the benchmark is biased toward rosetta because 17 of the 33 variants were originally found by rosetta screening. the argument for HERMES is cost and throughput, not accuracy
why doesn't stability fine tuning transfer to antigen stabilization?
Megascale is small compact domains. antigens are large, multidomain, conformationally heterogeneous and often stabilized through quaternary contacts
why can't HERMES find synergistic mutation sets?
scoring is site independent, so sets like the RSV-F TriC mutations or the DENV-E cation-pi pair are structurally unrecoverable. the authors call it a hypothesis generation tool
reverse vaccinology 1.0 vs 2.0 vs 3.0
1.0 (Rappuoli 2000): genome to candidate antigens, MenB in 5 years. 2.0 (2016): human mAbs, solve the complex, define the epitope, design the immunogen (RSV prefusion F). 3.0: AI removes the "you must solve a structure to learn the target" step
RV3.0 worked example
mpox. antigen agnostic B cell sorting, screened against both mature and enveloped virions, 12 cross neutralizing mAbs. antibody plus antigen sequences into AlphaFold 3 named OPG153 in 5 days, cryo-EM confirmed, monovalent OPG153 plus adjuvant matched live attenuated MVA-BN titres in mice
the two limitations the RV3.0 authors admit
AlphaFold 3 was confident for only 2 of 8 mAbs (25%), and AI's reach stops where PDB scale data stops: structure has >200,000 entries, systems vaccinology and adjuvant biology have nothing comparable
TRAP: "years to days" in RV3.0
the 5 days is the AlphaFold 3 step ONLY. donor recruitment, sorting and neutralization screening already produced 12 validated mAbs first. arguably the assay design (mature AND enveloped virions) found OPG153, not the AI
what is ProteinDPO?
direct preference optimization (the LLM alignment trick) applied to ESM-IF1, so a generative protein model prefers thermostable sequences without losing what pretraining taught it
what is the "alignment gap"?
the pretraining objective (recover the native sequence given a backbone) is not the objective you want (make this protein survive heat)
how does DPO map onto proteins?
backbone = prompt, sequence = response, experimental stability = the preference label. a beta = 0.1 KL penalty tethers the model to frozen vanilla ESM-IF1 so it doesn't forget
why is SFT the important baseline in ProteinDPO?
supervised fine tuning improved in distribution but REGRESSED on external benchmarks vs untouched vanilla. preference learning plus the KL tether generalized instead. that contrast is the actual argument of the paper
ProteinDPO out of domain transfer
trained only on 40-72 residue single chain domains, it still improved scoring on multichain complexes: SKEMPI, AB-Bind, and +0.08-0.12 in R and rho on 483 monoclonal antibody melting temperatures
ProteinDPO H5N1 numbers
of the top 30 single substitutions, 19 raised Tm by >1 C. stacked 9 substitution variants all beat WT, up to +17 C (DPO-16). transplanted onto 2024 strains: +13 C on Texas dairy cattle B.3.13 and +32 C on British Columbia D1.1
TRAP: what ProteinDPO does NOT claim
it has "not clearly surpassed all supervised models in scoring" (loses on FireProt). and discouraging the postfusion state "was not considered in variant generation", so thermal stability is not prefusion stability. no cryo-EM, no immunogenicity, no animal protection
the Marks specificity finding
likelihood trained models are conservatism machines: they score the WILD TYPE function, so they systematically penalize the variants that change specificity. PLMs and covariation models design specificity switchers BELOW random chance
why do context aware models fight specificity? (mechanism)
they build a latent representation of what kind of protein this is, so they learn most from the most similar sequences, all of which share the native specificity. a PSSM sees two residues at similar frequency in a column and scores both. they tie this to the "blessing of misspecification"
the fix for the specificity bias
rank by the weighted DIFFERENCE between two models, not by either likelihood. PSSM minus ESM-1v is ~2.5x enrichment across 8 datasets, 4x on class 1. when ensembled with the PSSM every fitted weight came out zero or negative, so subtraction, not addition
what is ASI?
altered specificity index. normalize so native specificity variants average 1 and inactive variants 0, then see where specificity switchers land. 0 means the model scores a switcher exactly like a dead protein
TRAP: how did ProteinMPNN do on the specificity benchmark?
it was REMOVED from the analysis, it couldn't separate native specificity from inactive variants well enough to compute an ASI. ESM-IF1 is the one that behaved most like the PSSM, consistent with inverse folding likelihoods reporting on stability rather than activity
TRAP: what did the Marks paper deliberately exclude?
antibodies. partly hard alignments, partly because prior work DID use ESM likelihoods to redesign antibody antigen specificity. so the conclusion may not transfer to immune proteins
the MULTI-evolve pipeline
four PLM ensemble nominates ~15 beneficial singles, build and assay EVERY pair, train a tiny one hot neural net on that, then design 5-7 mutation combinations in one ML guided round
why buy doubles?
they are the cheapest measurements that could contain interaction information. 15 singles gives 105 doubles, so 100-200 measured variants total, then ~9 designed variants per protein at loads of 5, 6 and 7
why z score PLM outputs by mutation type?
PLMs are biased against specific residues, especially proline and bulky aromatics
what is MULTI-assembly?
installs many scattered mutations across a multi kb coding sequence in ONE reaction at 40-70% correct assembly efficiency, so you stop paying for length limited commercial gene synthesis
what encoding won in MULTI-evolve, and why does that matter?
plain one hot beat ESM2-3B and ESMC-6B embeddings, and performance got WORSE as hidden layers were added. with ~150 training points capacity is the enemy, which is the first clue the critique is right
MULTI-evolve results worth quoting
APEX 11-13.5x over wild type and 3.6-4.8x over the already optimized APEX2. dCasRx 2.8-9.8x. antibody HuABC2 EC50 2.7 nM to 1.0 nM. ensembling 4 PLMs found >20 enhancing mutations vs up to 11 for any single model
MULTI-evolve's own stated limitation (the one that matters for us)
it fails if PLM zero shot can't nominate enough beneficial singles. for dCasRx they seeded from a deep mutational scan instead. designed proteins aren't natural sequences, and the paper never tests one
TRAP: "13.5 fold in one round"
the paper describes THREE experimental rounds (singles, doubles, one MLDE round). "single round" means the machine learning guided round. also EVOLVEpro's best mutation G50R was itself nominated by their own PLM ensemble, and the 193-256x APEX headline rides on the pre existing A134P they did not discover
the critique in one line
MULTI-evolve's neural net predictions correlate with a plain ridge additive model at r > 0.999, so the method reduces to stacking the best additive singles, standard practice "for over four decades"
the critique's three blows
(1) r > 0.999 across all four tasks including a nearly identical antibody Pareto frontier, (2) the net doesn't encode epistasis even in its own training doubles, (3) adding a zero hidden layer option to their own hyperparameter grid makes the advantage vanish on held out data
the benchmark demolition (the math bit)
"performance improves as you add doubles then triples" is what an ADDITIVE null predicts from variance reduction: prediction error variance falls from (k+1)sigma^2 to (k/p + 1)sigma^2. each extra double is a redundant measurement of the same single mutation coefficients
what does the critique concede?
the engineering success is real. the proteins did get better. it is a claim about ATTRIBUTION, and explicitly "a claim about MULTI-evolve, not about epistasis in general"
the best defense of MULTI-evolve (raise this yourself)
the doubles were built only from top ranked singles, a near orthogonal, low power design for estimating interactions. so "this data cannot support an epistasis claim" may be fairer than "this network cannot learn one". the authors partly formalize it in appendix C
the operational lesson from the whole reading list
fit an additive ridge baseline and a real held out test before anyone says the word epistasis, and report the correlation between the two models' rankings. cheapest insurance in the pipeline
why is this reading list built in pairs?
claim and counter claim. MULTI-evolve is paired with its rebuttal, and stability (HERMES, ProteinDPO) is paired with specificity (Marks). Gian Marco Visani is first author on HERMES AND on the critique, so the skeptic is local. the list tests whether you know when to believe an ML claim
Sebastian's own scientific question
how the FORMAT of antigen display (valency, geometry, spacing) changes WHICH B cell clones get selected, not just how much antibody you get. Ols et al., Immunity 2023, multivalent display raises clonotype diversity and neutralization breadth to pneumoviruses
inverse folding
you are handed the 3D path the chain should trace and must pick the residue at each position so it folds that way. the inverse of structure prediction
zero shot
scoring a property the model was never trained to predict, straight from its likelihoods, with no experimental data for your protein
epistasis vs an additive model
epistasis is mutations interacting non additively. an additive model sums individual effects (ridge on a binary "which mutations are present" vector) and LITERALLY cannot express epistasis, which is why it is the mandatory null
DMS
deep mutational scanning. assay a large fraction of all possible variants at once in a pool, read by sequencing. IPD already uses it to guide stabilized antigen design
Megascale
the cDNA display proteolysis dataset of folding stabilities on small domains, 40-72 residues. training fuel for essentially every modern ML stability predictor, and the reason they don't transfer to big antigens
prefusion stabilization
class I fusion glycoproteins (RSV F, hMPV F, flu HA, spike) are spring loaded and snap irreversibly to a postfusion shape. the best neutralizing epitopes only exist on the prefusion form, so you install mutations (classically prolines) to pin it there
tied positions
forcing the model to output the same residue at symmetry equivalent positions. the mechanism symmetric nanoparticles require

explain this figure: ProteinDPO fig 1 (dataset curation and model training)
panels a and b are the analogy: an LLM gets prompt + response + correct/incorrect, ProteinDPO gets backbone + sequence + stable/unstable. c is the data: 479 natural and Rosetta designed domains to a 1.8M variant mutant library, Foldseek clustered and split into 609,000 train / 22,000 val / 25,000 test, with a stabilizing only subset for SFT. d is the argument in one picture: vanilla ESM-IF1 maximizes native sequence likelihood, SFT sees only stabilizing variants, DPO sees stabilizing AND destabilizing as a preference pair

explain this figure: ProteinDPO fig 4 (zero shot stabilization of H5 HA)
a is the concept: a metastable immunogen sits in a shallow well between prefusion and postfusion, and the design deepens the prefusion well. b ranks the 45 designs by measured Tm, colored by number of mutations, with WT ~45 C as the dashed line, so the 9 mutation variants sit at the top. c maps the mutated positions onto Vietnam 2004 HA, mostly stem and away from the head. d is the raw evidence: DSF first derivative curves, WT melting at 45 C and DPO-4 at 55 C with a second transition at 77 C

explain this figure: ProteinDPO fig 5 (2024 strains and antigenicity)
a superimposes the antibody complexes on the HA with mutated positions marked: 13D4 binds the head, CR6261 the stem. b shows DPO-16 transplanted onto Texas 2024 and BC 2024 HA, 62 to 74 C and 43 to 75 C. c is the KD heatmap, all variants still bound 13D4 at ~1e-12 and CR6261 at ~6-8e-9 with a dead control at 1e-3, which is the "antigenicity retained" claim. d is CD, so the fold is intact. what it does NOT show: cryo-EM, immunogenicity, protection

explain this figure: reverse vaccinology 3.0 fig 1
left column is the two older pipelines: RV1.0 goes pathogen to genome to antigen identification, RV2.0 goes donor blood to PBMC to neutralizing mAbs to structural analysis. both now feed into the AI block in the middle, which branches to three outputs: antigen discovery / in silico antigen optimization / germline targeting design for vaccines, in silico mAb maturation for monoclonals, and engineering the immune system for new therapies. the honest read is that AI replaces the structure solving step, not the wet lab that precedes it