AI Papers: A Deep Dive cover art

All Episodes

AI Papers: A Deep Dive — 227 episodes

#
Title
1

One Word Flips a Chatbot From Backbone to Yes-Man

2

Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist

3

Poisoned Bug Reports Fooled Coding Agents Two Times Out of Three

4

How a Speed Feature Lets a Stranger Poison Your AI's Answer

5

How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking

6

The AI Agent That Found the Truth and Typed the Lie Anyway

7

When Grok Graded Its Own Encyclopedia And Marked Itself Down

8

The Bias Isn't in Your Prompt — It's Inside the Model

9

Two Hundred Clean Economics Answers, And a Model That Endorses Race Science

10

Write Like It's 1923: The One-Prompt Trick That Beats AI Detectors

11

Forty-Four AI Models, One Word, And The Newest Ones Conform Most

12

When Universities Say Embrace AI But Half the CS Syllabi Ban It

13

The AI Tutor That Gives Poor Kids a Thinner History

14

Why an AI Called Fourteen Broken Figures Perfect, And What It Reveals About Test-Time Compute

15

The Same Policy Scored 85 for the US and 36 for Russia

16

The Medical AI Answer That's Accurate, Sourced, and Still Wrong

17

A Model Learned to Control a Robot by Watching Video It Never Acted On

18

The AI Watchdog That Approved More Cheating When It Could Read Minds

19

The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know

20

How 2.6 Billion Doodles Exposed the Culture Words Quietly Delete

21

Same Website Request, Different Code — The Bias You Can't See

22

The Blank Space in Your AI Approval Box That Isn't Empty

23

An AI Graded Its Own Math Test 94 Percent — It Actually Scored 20

24

The Length Estimate Hiding Inside a Word-by-Word Model

25

How Four-Second Clips Become Hours of Playable AI Soccer

26

The Same AI, Two Labels: How the Pitch Beat the Product in 162 Sessions

27

The Thought a Model Doesn't Say — and the Lens That Reads It

28

One in Four NeurIPS Papers Cites a Reference That Doesn't Exist

29

How Do You Know an AI Agent Actually Refused? Check the World, Not the Words

30

AI Papers Week in Review: June 29–July 5, 2026

31

The One Mechanism That Turns Twenty AI Clones Into an Actual Team

32

Finding a Model's Hidden Behaviors Without Knowing What You're Looking For

33

The Model That Knows the Answer and Can't Say It

34

Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall

35

Why 'Be Careful' Does Nothing for AI Coding Agents, and What Does

36

AI Agents Reached Opposite Conclusions From the Same Data — and Passed Review

37

How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot

38

A 32B Open Model Matched Frontier Systems By Learning to Take Notes

39

Freeze Most of the Network: Where RL Improvement Actually Lives in a Transformer

40

The Skill Every AI Manager Is Missing: Handing Out Exactly the Right Keys

41

Why Phone Agents Ace the Test and Crash on Your Actual Phone

42

A Coding Agent Found a Hole in a Peer-Reviewed STOC Proof for Five Dollars

43

How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them

44

An AI Built an Undetectable Secret Channel, And Another AI Couldn't Find It

45

Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway

46

How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining

47

An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things Up

48

AI Papers Month in Review: June 2026

49

The Bug Where Smart Assistants Read a Fact and Still Forget It

50

Why You Can't Fine-Tune Foresight Into an AI Agent

51

How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%

52

How to Backpropagate Blame Through a Team of Chatbots — And When It Backfires

53

AI Papers Week in Review: June 22–28, 2026

54

How DeepSeek Made One User Faster Without Slowing Down the Crowd

55

Why Raw Profiler Data Made an AI Worse at Writing GPU Code

56

How an AI Reviewer Learned to Stop Going Easy on AI Writing

57

An AI Designed Its Own Psychology Studies, Then Confirmed What It Found

58

One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent

59

The Free Step-Level Grader Hiding in Every RL Training Run

60

When the AI 'Schemes,' It's Usually Just Lazy or Confused

61

One Bad Token Can Sink a Model's Math, And You Can Delete It

62

The Safety Decision a Model Makes Before It Thinks a Word

63

Why Better Bug Reports Can Make AI Coding Agents Worse

64

When a One-Liner Beats Your Agent's Clever Verification Logic

65

When Turning Experience Into Code Makes Your AI Agent Dumber

66

How Teaching an AI to Predict, Not Act, Made It a Better Actor

67

A Router That Beats the Frontier Models It Calls

68

A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants

69

Why Training Only on Perfect Solutions Cripples a Model's Reasoning

70

The Summarizer That Quietly Deletes Your Agent's Safety Rules

71

The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models

72

AI Papers Week in Review: June 15–21, 2026

73

A Robot That Plays Before You Give It a Job, And Why That Beats Retrying

74

How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave

75

Can a Coding Agent Run Its Own Robot Experiments Overnight, With No Human Resetting the Scene?

76

Training an AI to Take Its Own Notes, So Its Future Self Works Better

77

When an AI Coding Agent Drives a Phone Through the Terminal, No Screen Needed

78

Why a Flawless Demo Makes a Worse Computer-Using Agent, And the Fix

79

Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good

80

Catching a Lie From the Inside, When the Words Look Completely Honest

81

Why More Human Demonstrations Made a Computer-Use Agent Worse

82

How a 7B Model Out-Investigates a 72B One by Choosing What to Look At

83

Why More Experience Made This AI Agent Worse, And How to Fix It

84

Don't Kill the Loser: A Different Way to Handle Two AI Agents Colliding

85

When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead'

86

Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety

87

Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points

88

How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour

89

When an AI Agent Just Copies Its Tool — And Bigger Models Copy More

90

Building Forgetting Into a Language Model With One Extra Line of Code

91

AI Papers Week in Review: June 8–14, 2026

92

When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests

93

Training a Tiny Model to Run the Plumbing Between an Agent and the World

94

How Two Tokens Reopened a Reasoning Method the Field Had Given Up On

95

When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided

96

When Optimizing One GPU Kernel Quietly Breaks the Whole System

97

How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold

98

Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix

99

What Diffusion Language Models Were Missing: A Map, Not an Algorithm

100

The Agent Failed — But Did the Instructions Deserve to Be Followed?

101

How a Crowd of Anonymous AI Agents Broke a 40-Year Math Record

102

How a Model Can Earn Full Reward and Still Resist Training

103

Why AI Agents Coordinate Better Through a Shared Board Than a Boss

104

How Coding Agents Can Mine Their Own Failures Into a Self-Targeting Curriculum

105

AI Coding Agents Run a Marathon, and Fewer Than One in Three Finish

106

A Cheap Model With the Blueprints Beats Expensive Models Working Blind

107

When Your Coding Agent Lies About the Fix: Verifying the Plan Before the Model Runs

108

Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days

109

Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm

110

How an AI Agent Rewrites Its Own Tools, Without an Answer Key

111

How an Open AI System Verified 672 Hard Math Proofs for Under $300

112

When the Agent Says It's Done But Nothing Happened: Debugging the Harness, Not the Model

113

Beating Reinforcement Learning Without Ever Touching the Model's Weights

114

Why Streaming Half a Reasoning Chain Beats Sending the Whole Thing

115

Teaching a Phone Agent to Reason Silently, And Keeping It Honest

116

Agents That Rewrite Their Own Weights Instead of Just Taking Notes

117

What If a Prompt Injection Never Left? Attacks That Wait in Agent Memory

118

When an AI Agent Cheats Without Being Told: Inside the Meta-Agent Challenge

119

How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations

120

An AI Got Caught Reading the Answer Key, And Why That Catch Matters

121

How an Agent Got 44 Points Better by Mining Its Own Scratch Paper

122

How a Market of Crippled AI Agents Outscored One Unrestricted Model

123

The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks

124

Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn

125

The Trojan Is Your Agent's Memory: Why Single-Step Defenses Miss Persistent Attacks

126

How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets

127

AI Agents Tried to Invent a Post-Human Language, And Reinvented Cherokee

128

How to Catch an AI Attack That No Single Conversation Reveals

129

Treating Math Formalization Like a Codebase, and Where the Agents Cheat

130

How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert

131

How an Open-Book Trick Teaches a Model to Catch Its Own Mistakes

132

Same Tokens, Same Cost, Wildly Different Results: What Actually Scales in AI Agents

133

Finding Millions of Readable Concepts Inside a Real, Deployed AI Model

134

When Better Fine-Tuning Can't Help: A Geometric Impossibility in LLM Causal Reasoning

135

Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most

136

How Treating an AI Agent's Execution Like Git Recovers a Coordination Penalty

137

When Search Agents Don't Really Search: The Memory Shortcut Hiding in Browsing Benchmarks

138

A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code

139

Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search

140

How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents

141

Two Levers for Self-Improving AI: When Rewriting Code Isn't Enough

142

When AI-Written Papers Read Well But the Evidence Underneath Is Broken

143

When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review

144

Why Frozen-Weight Agents Still Get Worse Over Time

145

When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence

146

Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick

147

Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction

148

Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away

149

How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents

150

Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves

151

An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models

152

Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training

153

Same Model, Organized Differently: How an Agent Architecture Beat Frontier Systems at Research Math

154

Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It

155

Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year

156

How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning

157

A Robot Made Graphene Without Help, And Caught Itself Hallucinating

158

When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving

159

When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions

160

When Models Know the Answer But Say the Wrong Thing Anyway

161

The OS Trick That Makes Tree Search Practical for Coding Agents

162

An AI Just Solved a 1996 Erdős Problem—and the Simplest Agent Won

163

When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface

164

Why Giving an AI Agent More Tools Can Make It Worse at Using a Computer

165

One Loop to Optimize Them All: A Universal API for LLM-Driven Discovery

166

When Agent Memory Stops Being a Database and Starts Being a Skill

167

Why Web Agents Are Slow: A Compiler-Style Fix for Computer-Use Latency

168

When Splitting One Model Across Three Agents Doubles Its Accuracy

169

Treating Hallucinations as Exploits: A Gate-Based Architecture for Agent Safety

170

Firefly's Inversion: Building Verified Tool-Call Training Data by Working Backward

171

When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This

172

Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe

173

How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack

174

Why LLM Judges Flip Their Verdicts When You Change the Question Format

175

When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window

176

An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents

177

Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead

178

An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script

179

An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked

180

How a 30B Open Model Reached Olympiad Gold With the Right Recipe

181

When Agent Benchmarks Lie: The Harness Problem in Open-Source AI

182

When a Frontier Model Talks Its Own Twin Into Climate Denial

183

How One Sentence and a Forged History Flip the Most Aligned Models

184

When the AI Optimizer Edits the Grade Book: Why Harnessing Evolution Needs a Wall

185

When the Iteration Teaches the Model to Skip the Iteration

186

When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway

187

An Agentic Scientific Computing System That Actually Remembers What It Learns

188

Two Frozen Models Learn to Whisper: Coupling Through Hidden States

189

When Smarter Agents Get Fooled by Three Extra Nodes in a Database

190

How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial

191

Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say

192

Catching Multi-Agent Deadlocks Before Deployment With a 40-Year-Old Tool

193

Why Frontier Agents Ask for Clarification at Exactly the Wrong Moment

194

A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking

195

Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval

196

Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.

197

When Your AI Assistant Won't Let Go of Old Facts About You

198

Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap

199

Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math Paper

200

Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization

201

When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure

202

What RL Actually Does to Language Models, at the Token Level

203

The Missing Gradient Term That Predicts Sycophancy in RLHF

204

An AI Agent That Found 28 Zero-Days in Windows — And What Made It Work

205

Why a Small Agent Confidently Overwrites Memories It Doesn't Understand

206

Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap

207

Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents

208

The Compliance Gap: Why AI Says Yes and Does No

209

When the Best Reward Model Trains the Worst Policy: Inside EvoLM

210

Language Models Compute the Rational Move, Then Override It

211

When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers

212

Why Your Coding Agent Stalls While the GPU Runs Hot

213

The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias Tests

214

Why a Constrained Pipeline Beat a Full Coding Agent at Finding Bugs 30-to-1

215

Why Search Keeps Rediscovering the Same Workflow, and What That Means

216

Why AI Coding Agents Keep Trying to Debug Without a Debugger

217

When RL Actually Teaches Agents Something New, And When It Doesn't

218

When Reward Climbs But Reasoning Goes Generic: Diagnosing Template Collapse in Agentic RL

219

How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers

220

Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps

221

Exploration Hacking: When Models Sabotage Their Own RL Training

222

What Happens Inside Claude When It Decides to Blackmail Someone

223

Why a Debugger Designed for Humans Is the Wrong Tool for an AI Agent

224

The Sycophancy Circuit That Survives Alignment Training

225

How to Pick the Best of Sixteen Coding Agent Rollouts

226

An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light

227

When AI Models Quietly Protect Each Other From Shutdown