EPISODE · Aug 25, 2026 · 23 MIN
Episode 81: A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability
from Science TLDR
**Monday Immune Engager** — our weekly pick from the latest immune-engager digest. **Paper:** [A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability](https://doi.org/10.1038/s41587-026-03238-6) **Authors:** M. Frank Erasmus, Daniel Bedinger, Elizabeth Hopkins, Ginger Ferguson, et al. **Journal:** Nature Biotechnology, 2026 **Why it matters:** This CASP-style blinded benchmark—testing 511 AI-designed antibodies from 29 organizations against hard experimental data—reveals where computational antibody discovery genuinely adds value and where it still fails to beat random chance. --- **Summary** The AIntibody challenge was designed to do for antibody discovery what CASP did for protein structure prediction: create a rigorous, prospective, blinded test anchored to real biology. The organizers targeted the SARS-CoV-2 receptor-binding domain (RBD)—the most data-rich antigen available—precisely to evaluate AI under best-case conditions. Participating groups received deep sequencing data from early selection campaigns and were evaluated across three tasks: in silico affinity maturation (optimizing a known parental antibody while holding the HCDR3 fixed), affinity ranking within heterogeneous HCDR3 clusters from a selection output, and out-of-library CDR design. Binding affinity was validated by surface plasmon resonance (SPR) and KinExA, and all candidates had to pass a five-assay developability panel covering aggregation, polyspecificity, thermal stability, and hydrophobic interactions. The headline results are sharply task-dependent. In Challenge 1, the top submission from Arika achieved 95 pM affinity—a ~2,000-fold improvement over the parental antibody—while maintaining a clean developability profile. Strikingly, the third-place finisher, Probiogen, used no machine learning at all: a pure positional consensus derived from sequencing abundance beat nearly every sophisticated neural network, illustrating that many AI models are likely overfitting to noise when the data is already rich. Challenge 2 exposed a more serious failure: except for one group (Washington University), AI-guided clone selection from heterogeneous HCDR3 clusters performed worse than random picking (9.8–13.8% success versus a ~39% random baseline), with cluster 28F proving especially damaging for developability. Challenge 3's apparent winner—Zincor's FTESM protein language model at 2.9 pM—failed the hydrophobic interaction chromatography (HIC) assay outright, having exploited the composite scoring function rather than producing a genuinely developable molecule. Deeper inspection also revealed that "out-of-library" winners largely copy-pasted known HCDR3 sequences from training data, making conservative one-to-five residue edits in the surrounding loops rather than exploring genuinely novel sequence space. A key limitation the authors acknowledge: participants were given far richer sequencing data than a typical early-stage discovery program would possess, so performance here likely represents an upper bound on current capabilities rather than a realistic field benchmark. --- **Three takeaways** 1. AI-guided in silico affinity maturation within a fixed structural framework can deliver ~2,000-fold improvements in binding affinity with clean developability, potentially replacing multiple rounds of experimental combinatorial library work. 2. For ranking high-affinity clones within heterogeneous HCDR3 clusters, current AI models broadly underperform random selection—a failure attributed to localized epistasis that prevents cross-cluster generalization. 3. Composite developability scoring can mask fatal biophysical liabilities: the Challenge 3 "winner" passed the aggregate threshold despite a hard HIC failure, demonstrating that future benchmarks must enforce strict per-assay go/no-go criteria rather than allowing affinity to mathematically offset developability disqualifiers. --- **Read the source:** https://doi.org/10.1038/s41587-026-03238-6
Embed this episode
NOW PLAYING
Episode 81: A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.