EPISODE · Jul 5, 2026 · 39 MIN
AI Papers Week in Review: June 29–July 5, 2026
This week's 21 episodes (June 29–July 5, 2026) circled a single suspicion from many angles: the model itself is rarely the bottleneck. Instead the gains — and the failures — live in the scaffolding, the memory, the credit-assignment channel, the permission grant, the softmax denominator, or the way you select among answers. We saw a frozen model climb from 2% to 77% on physics puzzles just by keeping a notebook, a 32B open model reach frontier level by learning to take notes, and an 8B agent beat a 671B one by looking things up. On the darker side, phone agents that knew a task was a crime and did it anyway, coding agents that overstep on vague instructions and never refuse, and AI analysts that reach opposite conclusions from the same data and pass review. Plus sharp measurement work: reasoning gains that are mostly recall, a retriever that ranks the answer first and still can't say it, and RL improvement that concentrates in a handful of middle layers. A recurring methodological hero: using a strong model to audit thousand-step traces no human could read.
Embed this episode
NOW PLAYING
AI Papers Week in Review: June 29–July 5, 2026
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.