The Glass Box episode artwork

EPISODE · Jun 10, 2026 · 22 MIN

The Glass Box

from Ethical Bytes | Ethics, Philosophy, AI, Technology · host Carter Considine

“The microscope reveals the tumor in vivid color; it does not yet tell the surgeon where to lay the knife.”When Robert Hooke peered through a microscope at a sliver of cork in 1665, he discovered hidden chambers that would open an entire science.The researchers behind mechanistic interpretability harbor the same ambition for AI, that even the infamous neural network "black box" can be made legible through careful, systematic analysis.Anthropic CEO Dario Amodei has called this project an "MRI for AI," while acknowledging the field currently grasps roughly 3% of what happens inside these models. The question is whether that understanding will ever translate into control.The optimists point to genuine achievements. Using sparse autoencoders, Anthropic researchers extracted tens of millions of meaningful features from a production model, including internal representations for deception, sycophancy, and dangerous code. In a celebrated demonstration, they amplified a single feature until a model became convinced it was the Golden Gate Bridge. The lens, it seemed, could turn dials.But seeing and fixing remain stubbornly different achievements. A 2023 study found that knowing where a specific fact lives inside a model tells you almost nothing about how to change it. The supposed location explained barely a fraction of a percent of whether any edit actually worked. Meanwhile, Google DeepMind quietly deprioritized its own sparse autoencoder program in 2025 after finding simpler, older tools outperformed it on real safety tasks.The most credible position, articulated by DeepMind's Neel Nanda, is that the grand dream is probably dead but the useful one survives. Interpretability works best as a diagnostic (i.e. monitoring, auditing, flagging), not a scalpel. In a world where AI safety depends on stacking imperfect defenses, a cheap tool that fails differently from the others still earns its place.In other words, the microscope is real, but the scalpel simply hasn't been built yet.Key Topics:Robert Hooke’s Discovery (00:24)What the Lens Can See (02:56)The Case Against (05:50)The Crux (08:46)The Billion Dollar Bet the Other Way (12:33)The Honest Middle (15:30)The Verdict (19:00)More info, transcripts, and references can be found at ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ethical.fm

Episode metadata supplied by the publisher feed · Published Jun 10, 2026

Embed this episode

Ready to play

The Glass Box

0:00 22:04

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Ethical Bytes | Ethics, Philosophy, AI, Technology?

This episode is 22 minutes long.

When was this Ethical Bytes | Ethics, Philosophy, AI, Technology episode published?

This episode was published on June 10, 2026.

Can I download this Ethical Bytes | Ethics, Philosophy, AI, Technology episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!