RAPTOR: Stable Concept Directions From Logistic Probes episode artwork

EPISODE · May 10, 2026

RAPTOR: Stable Concept Directions From Logistic Probes

from AI Post Transformers

This episode explores RAPTOR, a method for extracting concept directions from language model hidden states using ridge-regularized logistic probes, with the goal of making those directions accurate enough for interpretation and stable enough for activation steering. It explains the core probe-then-steer workflow, why linear probes can reveal what a model has encoded, and why good classification accuracy does not necessarily produce a reliable control vector. The discussion situates the paper within broader debates in mechanistic interpretability, including concerns about brittle probes, distribution shift, and whether a single direction can really capture a concept like sentiment, refusal, or honesty. A listener would find it interesting because the episode turns an abstract interpretability question into a concrete engineering tradeoff about robustness, causal usefulness, and whether cheap white-box methods could become practical tools for controlling large models. Sources: 1. RAPTOR: Ridge-Adaptive Logistic Probes — Ziqi Gao, Yaotian Zhu, Qingcheng Zeng, Xu Zhao, Ziqing Wang, Feng Ruan, Kaize Ding, 2026 http://arxiv.org/abs/2602.00158 2. Plug and Play Language Models: a Simple Approach to Controlled Text Generation — Sumanth Dathathri, Andrea Madotto, Janice Lan, Jason Yosinski, Rosanne Liu, et al., 2019 https://arxiv.org/abs/1912.02164 3. Activation Addition: Steering Language Models Without Optimization — Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Ulisse Mini, Monte MacDiarmid, 2023 https://arxiv.org/abs/2308.10248 4. Representation Engineering: A Top-Down Approach to AI Transparency — Andy Zou, Long Phan, Sarah Chen, James Campbell, Dan Hendrycks, J. Zico Kolter, et al., 2023 https://arxiv.org/abs/2310.01405 5. Refusal in Language Models Is Mediated by a Single Direction — Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, Neel Nanda, 2024 https://arxiv.org/abs/2406.11717 6. Understanding Intermediate Layers Using Linear Classifier Probes — Guillaume Alain, Yoshua Bengio, 2016 https://openreview.net/forum?id=HJ4-rAVtl 7. Designing and Interpreting Probes with Control Tasks — John Hewitt, Percy Liang, 2019 https://aclanthology.org/D19-1275/ 8. Information-Theoretic Probing with Minimum Description Length — Elena Voita, Ivan Titov, 2020 https://aclanthology.org/2020.emnlp-main.14/ 9. RAPTOR: Ridge-Adaptive Logistic Probes — Ziqi Gao, Yaotian Zhu, Qingcheng Zeng, Xu Zhao, Ziqing Wang, Feng Ruan, Kaize Ding, 2026 https://arxiv.org/abs/2602.00158 10. Steering Llama 2 via Contrastive Activation Addition — Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, Alexander Turner, 2024 https://scholar.google.com/scholar?q=Steering+Llama+2+via+Contrastive+Activation+Addition 11. Analysing the Generalisation and Reliability of Steering Vectors — Daniel Tan, David Chanin, Aengus Lynch, Brooks Paige, Dimitrios Kanoulas, Adrià Garriga-Alonso, Robert Kirk, 2024 https://scholar.google.com/scholar?q=Analysing+the+Generalisation+and+Reliability+of+Steering+Vectors 12. Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution — Haiyan Zhao, Heng Zhao, Bo Shen, Ali Payani, Fan Yang, Mengnan Du, 2025 https://scholar.google.com/scholar?q=Beyond+Single+Concept+Vector%3A+Modeling+Concept+Subspace+in+LLMs+with+Gaussian+Distribution 13. Controlling Large Language Models Through Concept Activation Vectors — Hanyu Zhang, Xiting Wang, Chengao Li, Xiang Ao, Qing He, 2025 https://scholar.google.com/scholar?q=Controlling+Large+Language+Models+Through+Concept+Activation+Vectors 14. Token prepending: A training-free approach for eliciting better sentence embeddings from llms — authors not confirmed from provided snippet, recent, unverified https://scholar.google.com/scholar?q=Token+prepending%3A+A+training-free+approach+for+eliciting+better+sentence+embeddings+from+llms 15. Rep2Text: Decoding Full Text from a Single LLM Token Representation — authors not confirmed from provided snippet, recent, unverified https://scholar.google.com/scholar?q=Rep2Text%3A+Decoding+Full+Text+from+a+Single+LLM+Token+Representation 16. Context Matters: Analyzing the Generalizability of Linear Probing and Steering Across Diverse Scenarios — authors not confirmed from provided snippet, recent, unverified https://scholar.google.com/scholar?q=Context+Matters%3A+Analyzing+the+Generalizability+of+Linear+Probing+and+Steering+Across+Diverse+Scenarios 17. Angular steering: Behavior control via rotation in activation space — authors not confirmed from provided snippet, recent, unverified https://scholar.google.com/scholar?q=Angular+steering%3A+Behavior+control+via+rotation+in+activation+space 18. Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency — authors not confirmed from provided snippet, recent, unverified https://scholar.google.com/scholar?q=Global+Evolutionary+Steering%3A+Refining+Activation+Steering+Control+via+Cross-Layer+Consistency 19. Fine-Grained Activation Steering: Steering Less, Achieving More — authors not confirmed from provided snippet, recent, unverified https://scholar.google.com/scholar?q=Fine-Grained+Activation+Steering%3A+Steering+Less%2C+Achieving+More 20. AI Post Transformers: Language Models are Injective and Hence Invertible — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-21-language-models-are-injective-an-7545e0.mp3 21. AI Post Transformers: Learning to Reason with 13 Parameters — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-14-learning-to-reason-with-13-parameters-54c87f.mp3 22. AI Post Transformers: When Spectral Gradient Updates Help Deep Learning — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-04-when-spectral-gradient-updates-help-deep-9c8441.mp3 Interactive Visualization: RAPTOR: Stable Concept Directions From Logistic Probes

Episode metadata supplied by the publisher feed · Published May 10, 2026

Embed this episode

NOW PLAYING

RAPTOR: Stable Concept Directions From Logistic Probes

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on May 10, 2026.

Is there a transcript available for this episode?

Yes, a full transcript is available for this episode. You can read the complete transcript on the episode page.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!