EPISODE · Jun 17, 2026
PaperBench: Can AI Replicate AI Research?
from AI Post Transformers
This episode explores PaperBench, a benchmark designed to test whether frontier AI agents can independently replicate the empirical work of recent machine learning papers from scratch rather than merely explain them. It breaks down what agentic AI actually entails in this setting: reading papers, writing code, choosing baselines, reconstructing missing details, running experiments, debugging failures, and judging whether reproduced results match the original claims. The discussion compares PaperBench with other evaluation ladders such as CORE-Bench, MLE-bench, RE-Bench, and JudgeEval, while also debating whether controlled scratch replication should be viewed as advanced engineering or a meaningful proxy for real research practice. Listeners get a clear look at why this matters for both AI capability measurement and safety, especially given PaperBench’s carefully curated design of 20 ICML 2024 papers, 12 topics, and more than 8,000 graded tasks. Sources: 1. PaperBench: Evaluating AI's Ability to Replicate AI Research — Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, Tejal Patwardhan, 2025 http://arxiv.org/abs/2504.01848 2. PaperBench: Evaluating AI's Ability to Replicate AI Research — Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, et al., 2025 https://scholar.google.com/scholar?q=PaperBench%3A+Evaluating+AI%27s+Ability+to+Replicate+AI+Research 3. RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts — Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, et al., 2024 https://scholar.google.com/scholar?q=RE-Bench%3A+Evaluating+frontier+AI+R%26D+capabilities+of+language+model+agents+against+human+experts 4. CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark — Zachary S. Siegel, Sayash Kapoor, Nitya Nagdir, Benedikt Stroebl, Arvind Narayanan, 2024 https://scholar.google.com/scholar?q=CORE-Bench%3A+Fostering+the+Credibility+of+Published+Research+Through+a+Computational+Reproducibility+Agent+Benchmark 5. MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation — Qian Huang, Jian Vora, Percy Liang, Jure Leskovec, 2023 https://scholar.google.com/scholar?q=MLAgentBench%3A+Evaluating+Language+Agents+on+Machine+Learning+Experimentation 6. MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering — Jun Shern Chan et al., 2024 https://scholar.google.com/scholar?q=MLE-bench%3A+Evaluating+Machine+Learning+Agents+on+Machine+Learning+Engineering 7. EXP-Bench: Can AI Conduct AI Research Experiments? — Patrick Tser Jern Kon et al., 2025 https://scholar.google.com/scholar?q=EXP-Bench%3A+Can+AI+Conduct+AI+Research+Experiments%3F 8. MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research — Hui Chen, Miao Xiong, Yujie Lu, Wei Han, Ailin Deng, Yufei He, Jiaying Wu, Yibo Li, Yue Liu, Bryan Hooi, 2025 https://scholar.google.com/scholar?q=MLR-Bench%3A+Evaluating+AI+Agents+on+Open-Ended+Machine+Learning+Research 9. ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers? — Christine Ye et al., 2025 https://scholar.google.com/scholar?q=ReplicationBench%3A+Can+AI+Agents+Replicate+Astrophysics+Research+Papers%3F 10. Can Large Language Models Be an Alternative to Human Evaluations? — Cheng-Han Chiang, Hung-yi Lee, 2023 https://scholar.google.com/scholar?q=Can+Large+Language+Models+Be+an+Alternative+to+Human+Evaluations%3F 11. RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following — Tianjun Pan et al., 2026 https://arxiv.org/abs/2603.25133 12. JudgeBench: A Benchmark for Evaluating LLM-based Judges — Sijun Tan et al., 2024 https://arxiv.org/abs/2410.12784 13. When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation — Abeer Badawi et al., 2025 https://arxiv.org/abs/2510.19032 14. A Dataset For Computational Reproducibility — Lazaro Costa, Susana Barbosa, Jacome Cunha, 2025 https://arxiv.org/abs/2504.08684 15. SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers — Yanzheng Xiang et al., 2025 https://arxiv.org/abs/2504.00255 16. OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding — Deming Ding et al., 2026 https://arxiv.org/abs/2601.10343 17. ContextBench: A Benchmark for Context Retrieval in Coding Agents — Han Li et al., 2026 https://arxiv.org/abs/2602.05892 18. AI Post Transformers: When AI Builds Itself and Recursive Self-Improvement — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-05-when-ai-builds-itself-and-recursive-self-8bbf9e.mp3 19. AI Post Transformers: When LLM Judges Become Coin Flips — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-when-llm-judges-become-coin-flips-8b43ef.mp3 20. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3 21. AI Post Transformers: Kimi K2.5 and Visual Agent Swarms — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-24-kimi-k25-and-visual-agent-swarms-7d04d7.mp3 Interactive Visualization: PaperBench: Can AI Replicate AI Research?
Embed this episode
NOW PLAYING
PaperBench: Can AI Replicate AI Research?
No transcript for this episode yet
Similar Episodes
No similar episodes found.