The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know episode artwork

EPISODE · Jul 10, 2026 · 14 MIN

The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know

from AI Papers: A Deep Dive

The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know Source: https://arxiv.org/abs/2607.08393 Paper was published on July 09, 2026 This episode was AI-generated on July 10, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A model already knew every fact it needed — and still failed the reasoning question, until researchers physically relocated one internal representation and watched accuracy jump up to six times. It turns out teaching a model a fact and making that fact usable are two different problems, and the whole field has been measuring the first while assuming it bought the second. This episode walks through the causal experiment that proves the knowledge was there all along, just filed in the wrong place. Key Takeaways: - Why memorization hits 98%+ within a few epochs while the ability to actually use the fact plateaus far below — and why more training or bigger models can't close the gap - The mechanism behind the 'Knowing–Using Gap': once a fact is memorized its error hits zero, so the gradient that would move it into a usable position vanishes - How 'self-patching' proves the fact is physically stored but mislocated, without needing a correct run to borrow from - The surprising finding that skipping several layers still triggers correct reasoning — proving it's a routing problem, not a depth-of-processing or capacity problem - How a blind, fixed two-relocation rule recovers 58–75% of the oracle's ceiling with zero per-question search - The honest limits: it's a hand-operation on synthetic knowledge-graph facts, probes one token position, and fixes nothing during normal use 00:25 - The fact it knew but couldn't use: Sets up the filing-cabinet metaphor and the Knowing–Using Gap: a model with perfect recall of both halves that collapses when asked to chain them. 02:20 - Why won't more training fix it?: The paper kills the obvious explanation, showing memorization rockets to 98%+ while generalization plateaus, because once error hits zero the gradient dies. 03:57 - Proving the fact is in the wrong place: Introduces the transformer-as-assembly-line framing and self-patching — relocating an existing representation between layers to test whether the knowledge is already inside. 05:59 - The vanishing gradient, caught on camera: The heatmap over training shows a red island of recoverable knowledge that either reaches the diagonal (generalization) or freezes short of it (stranded fact). 07:34 - Wait — skipping layers helps?: The twist that overturns the timing explanation: early-layer, under-processed representations dropped into the middle still trigger correct reasoning, making this a routing problem, not a capacity one. 08:52 - The 6x cure that's secretly a cheat: Distinguishes the oracle ceiling (44% from under 8%, but requires knowing the answer) from the shippable fixed two-relocation rule that recovers most of the headroom blind. 10:33 - Where the fix stops being real: The steelman critique: the facts are atomic synthetic triplets, the intervention probes one token position, and it fixes nothing during normal use — a proof of concept, not a product. 12:36 - Two problems the field conflated: The payoff reframe connects the reversal curse and failed knowledge editing to mislocation, and asks whether to route facts into weights or keep them outside and retrieve at query time. Recommended Reading: - Locating and Editing Factual Associations in GPT: The ROME paper that pioneered causal activation patching to find where facts physically live in a transformer — the interpretability lineage this episode's self-patching method builds on and adapts. (https://arxiv.org/abs/2202.05262) - The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A": The reversal-curse failure the episode names explicitly as a sibling phenomenon that the paper's misalignment story tries to unify. (https://arxiv.org/abs/2309.12288) - Physics of Language Models: Part 3.1, Knowledge Storage and Extraction: A controlled study of exactly the episode's puzzle — models that memorize facts yet can't extract or reason with them — using synthetic biographical knowledge graphs like the ones here. (https://arxiv.org/abs/2309.14316) - Emergent Abilities of Large Language Models: Frames the memorization-then-plateau curves and scaling behavior the episode contrasts against, useful for readers weighing whether 'just scale it' would close the Knowing–Using Gap. (https://arxiv.org/abs/2206.07682)

Episode metadata supplied by the publisher feed · Published Jul 10, 2026

Embed this episode

NOW PLAYING

The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know

0:00 14:04

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of AI Papers: A Deep Dive?

This episode is 14 minutes long.

When was this AI Papers: A Deep Dive episode published?

This episode was published on July 10, 2026.

Can I download this AI Papers: A Deep Dive episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!