How Models Detect Hidden Activation Steering episode artwork

EPISODE · May 8, 2026

How Models Detect Hidden Activation Steering

from AI Post Transformers

This episode explores a mechanistic interpretability study asking whether a language model can detect when a concept has been injected into its hidden activations and, in some cases, identify what that concept was. It explains the difference between detection and identification, walks through activation steering in the residual stream, and highlights the paper’s controlled experiments on Gemma3-27B across 500 concepts, including a strong result of moderate detection with zero false positives under several prompt styles. The discussion also focuses on the paper’s argument that this reporting behavior emerges mainly during post-training, especially preference optimization, rather than from pretraining alone. Listeners would find it interesting because it turns a provocative claim about model “introspection” into a concrete circuit-level question about what internal features and gates may be doing. Sources: 1. Mechanisms of Introspective Awareness — Uzay Macar, Li Yang, Atticus Wang, Peter Wallich, Emmanuel Ameisen, Jack Lindsey, 2026 http://arxiv.org/abs/2603.21396 2. Emergent Introspective Awareness in Large Language Models — Jack Lindsey, 2025 https://scholar.google.com/scholar?q=Emergent+Introspective+Awareness+in+Large+Language+Models 3. Looking Inward: Language Models Can Learn About Themselves by Introspection — Felix J. Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, Owain Evans, 2024 https://scholar.google.com/scholar?q=Looking+Inward%3A+Language+Models+Can+Learn+About+Themselves+by+Introspection 4. Activation Addition: Steering Language Models Without Optimization — Alexander Matt Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, Monte MacDiarmid, Chris Olah, 2023 https://scholar.google.com/scholar?q=Activation+Addition%3A+Steering+Language+Models+Without+Optimization 5. Representation Engineering: A Top-Down Approach to AI Transparency — Andy Zou, Long Phan, Sarah Chen, James Campbell, Richard Ngo, Adam Jermyn, Stephen McAleer, Alexander Tamkin, 2023 https://scholar.google.com/scholar?q=Representation+Engineering%3A+A+Top-Down+Approach+to+AI+Transparency 6. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet — Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, and others, 2024 https://scholar.google.com/scholar?q=Scaling+Monosemanticity%3A+Extracting+Interpretable+Features+from+Claude+3+Sonnet 7. Circuit Tracing: Revealing Computational Graphs in Language Models — Emmanuel Ameisen, Jack Lindsey, Adam Pearce, Wes Gurnee, Nicholas L. Turner, Brian Chen, Craig Citro, and others, 2025 https://scholar.google.com/scholar?q=Circuit+Tracing%3A+Revealing+Computational+Graphs+in+Language+Models 8. Steering Vector Fields for Context-Aware Inference-Time Control in Large Language Models — authors unclear from snippet, 2025/2026 https://scholar.google.com/scholar?q=Steering+Vector+Fields+for+Context-Aware+Inference-Time+Control+in+Large+Language+Models 9. No Training Wheels: Steering Vectors for Bias Correction at Inference Time — authors unclear from snippet, 2025/2026 https://scholar.google.com/scholar?q=No+Training+Wheels%3A+Steering+Vectors+for+Bias+Correction+at+Inference+Time 10. A mechanistic understanding of alignment algorithms: A case study on DPO and toxicity — authors unclear from snippet, 2025/2026 https://scholar.google.com/scholar?q=A+mechanistic+understanding+of+alignment+algorithms%3A+A+case+study+on+DPO+and+toxicity 11. How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis — authors unclear from snippet, 2025/2026 https://scholar.google.com/scholar?q=How+Does+DPO+Reduce+Toxicity%3F+A+Mechanistic+Neuron-Level+Analysis 12. Refusal in language models is mediated by a single direction — authors unclear from snippet, 2025/2026 https://scholar.google.com/scholar?q=Refusal+in+language+models+is+mediated+by+a+single+direction 13. Beyond I'm Sorry, I Can't: Dissecting Large-Language-Model Refusal — authors unclear from snippet, 2025/2026 https://scholar.google.com/scholar?q=Beyond+I%27m+Sorry%2C+I+Can%27t%3A+Dissecting+Large-Language-Model+Refusal 14. Surgical, cheap, and flexible: Mitigating false refusal in language models via single vector ablation — authors unclear from snippet, 2025/2026 https://scholar.google.com/scholar?q=Surgical%2C+cheap%2C+and+flexible%3A+Mitigating+false+refusal+in+language+models+via+single+vector+ablation 15. Residual stream analysis with multi-layer saes — authors unclear from snippet, 2025/2026 https://scholar.google.com/scholar?q=Residual+stream+analysis+with+multi-layer+saes 16. AI Post Transformers: Anthropic: Introspective Awareness in LLMs — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/anthropic-introspective-awareness-in-llms/ 17. AI Post Transformers: Neural Chameleons and Evading Activation Monitors — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-14-neural-chameleons-and-evading-activation-bc470e.mp3 18. AI Post Transformers: Advancing Mechanistic Interpretability with Sparse Autoencoders — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/advancing-mechanistic-interpretability-with-sparse-autoencoders/ 19. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3 20. AI Post Transformers: Self-Improving Pretraining With Post-Trained Models — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-02-self-improving-pretraining-with-post-tra-e37460.mp3 Interactive Visualization: How Models Detect Hidden Activation Steering

Episode metadata supplied by the publisher feed · Published May 8, 2026

Embed this episode

NOW PLAYING

How Models Detect Hidden Activation Steering

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on May 8, 2026.

Is there a transcript available for this episode?

Yes, a full transcript is available for this episode. You can read the complete transcript on the episode page.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!