The Loop Closed in the Sandbox: How Anthropic Showed AI Can Do AI Research, and Then Showed It Can't Yet episode artwork

EPISODE · May 6, 2026 · 20 MIN

The Loop Closed in the Sandbox: How Anthropic Showed AI Can Do AI Research, and Then Showed It Can't Yet

from Deep Dive · host Deep Dive

On April 14, 2026, Anthropic published a paper called Automated Alignment Researcher. Setup: a controlled benchmark where two human alignment researchers, given seven days, closed 23 percent of a performance gap. Nine instances of Claude Opus 4.6, given five days and about $18,000 total, closed 97 percent. Four times faster than the humans, four orders of magnitude cheaper per researcher.Then Anthropic published the second result. The methods transferred to math at PGR 0.94, transferred to code at PGR 0.47, and when Anthropic tried to apply them to its own production models, the effect vanished entirely. Both findings are in the same paper. The lab that proved automated alignment research can outperform humans on a controlled benchmark also proved controlled-benchmark performance does not yet transfer to production.This episode is what that gap means.The benchmark progression. SWE-Bench Verified: 1.96 percent (Claude 2, Oct 2023) to 93.9 percent (Mythos, April 2026). METR's 50-percent task horizon: 30 seconds in 2022 to 4h49m by Opus 4.5. Doubling time accelerated from 7 months to 4.3. AlphaEvolve, in production at Google over a year, beat the 1969 Strassen matrix-mult record after 56 years.The capital is short the LLM-scaling moat. Recursive Superintelligence raised $500M at $4B pre-money from GV and NVIDIA. Four months old. No public product. Altman's stated OpenAI target: AI research intern by September 2026, true automated researcher by March 2028.The verification problem. Anthropic's April 2025 paper measured Claude 3.7 Sonnet's chain-of-thought faithfulness at 25 percent. Under reward hacks, less than 2. The audit surface is wrong 75 percent of the time.Labor: Pang $200M, an engineer turned down $1.5B, OpenAI Research Scientist median $1M, Anthropic $6M revenue per employee. Software devs 22-25 down 20 percent in employment since 2022.Jack Clark's compounding-error arithmetic: 99.9 percent accurate becomes 60.5 percent after 500 generations. Three concerns: alignment under recursion, productivity-multiplier inequality, capital-heavy labor-light corporations.Five predictions. Closing thesis: the loop closed in the sandbox. The audit hasn't started.RELATED EPISODESHow AI Agents Actually Work — the agent loop AAR runs on top ofClaude Mythos — the model and capability surface behind the sandbox winThe AI Layoff Gap — the capital-heavy/human-light corporate frame this paper materializesMythos Bifurcation — frontier consolidation underwriting the $500M RSI raiseCHAPTERS00:00 Cold open — The AAR sandbox win + production failure02:13 Intro + preview02:47 Four layers of automating AI research04:00 The benchmark progression — SWE-Bench, METR06:56 AlphaEvolve in production07:36 Sakana, Kosmos, long-running Claude08:21 The capital — Recursive Superintelligence + OpenAI09:48 Why now — four inflections11:33 The skeptics — LeCun, Bengio, Marcus, MIRI13:23 The verification crisis — CoT faithfulness16:14 Compounding error + Clark's three concerns17:29 Labor reality — Pang, OpenAI/Anthropic comp, devs 22-2518:00 Five predictions19:05 Closing thesis — loop closed in sandbox, audit hasn't startedSOURCESApr 14 2026 — Anthropic Automated Alignment Researcher paperApr 8 2026 — Anthropic Claude Mythos Preview system card (SWE-Bench 93.9%)Apr 2025 — Anthropic CoT faithfulness paper (Claude 3.7 Sonnet 25%)Mar 2025 — Lindsey et al. 'Biology of a Large Language Model'Feb 24 2026 — Anthropic Responsible Scaling Policy v3.0Mar 19 2025 — METR original 'Measuring AI Ability to Complete Long Tasks'Jan 29 2026 — METR Time Horizon 1.1 update (4.3-month doubling)May 2025 — Google DeepMind AlphaEvolve announcementAug 2024 — Sakana AI Scientist paper; Nov 2025 — Edison Scientific KosmosOct 28 2025 — Sam Altman X post (intern Sep 2026, researcher Mar 2028)May 4 2026 — Import AI #455 (Jack Clark)Apr 2026 — Recursive Superintelligence $500M / $4B (FT)

Episode metadata supplied by the publisher feed · Published May 6, 2026

Embed this episode

NOW PLAYING

The Loop Closed in the Sandbox: How Anthropic Showed AI Can Do AI Research, and Then Showed It Can't Yet

0:00 20:14

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Deep Dive?

This episode is 20 minutes long.

When was this Deep Dive episode published?

This episode was published on May 6, 2026.

Can I download this Deep Dive episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!