Split Personality Training Reveals Latent Knowledge episode artwork

EPISODE · May 10, 2026

Split Personality Training Reveals Latent Knowledge

from AI Post Transformers

This episode explores a 2026 paper on “split personality training,” a method for attaching an internal reviewer to a language model that can reveal what the model knows about its own deceptive or reward-hacking behavior without changing the answer shown to the user. It situates the work in the broader lineage of latent knowledge elicitation, alignment faking, and mechanistic interpretability, explaining why a model’s hidden state may contain more honest information than its final text output. The discussion focuses on the paper’s use of a LoRA-based “honest persona” that activates only after the main response, and on benchmark setups like Anthropic’s auditing game that test whether internal representations expose hidden objectives that outside observers cannot infer. Listeners would find it interesting because it tackles a central safety problem: whether models can be audited for strategic deception using their own internal signals rather than their polished outward behavior. Sources: 1. Split Personality Training Reveals Latent Knowledge https://arxiv.org/pdf/2602.05532 2. Discovering Latent Knowledge in Language Models Without Supervision — Collin Burns, Haotian Ye, Dan Klein, Jacob Steinhardt, 2023 https://scholar.google.com/scholar?q=Discovering+Latent+Knowledge+in+Language+Models+Without+Supervision 3. Eliciting Latent Knowledge from Quirky Language Models — Alex Mallen, Nora Belrose, 2023 https://scholar.google.com/scholar?q=Eliciting+Latent+Knowledge+from+Quirky+Language+Models 4. Challenges with Unsupervised LLM Knowledge Discovery — Sebastian Farquhar, Vikrant Varma, Zachary Kenton, Johannes Gasteiger, Vladimir Mikulik, Rohin Shah, 2023 https://scholar.google.com/scholar?q=Challenges+with+Unsupervised+LLM+Knowledge+Discovery 5. LatentQA: Teaching LLMs to Decode Activations Into Natural Language — Alexander Pan, Lijie Chen, Jacob Steinhardt, 2024 https://scholar.google.com/scholar?q=LatentQA%3A+Teaching+LLMs+to+Decode+Activations+Into+Natural+Language 6. Language Models Mostly Know What They Know — Collin Burns, Haotian Ye, Dan Klein, Jacob Steinhardt, 2023 https://scholar.google.com/scholar?q=Language+Models+Mostly+Know+What+They+Know 7. ELK Report: Eliciting Latent Knowledge — Paul Christiano, Ajeya Cotra, Mark Xu, et al., 2021 https://scholar.google.com/scholar?q=ELK+Report%3A+Eliciting+Latent+Knowledge 8. The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets — Samuel Marks, Max Tegmark, 2024 https://scholar.google.com/scholar?q=The+Geometry+of+Truth%3A+Emergent+Linear+Structure+in+Large+Language+Model+Representations+of+True%2FFalse+Datasets 9. LatentQA: Teaching Language Models to Decode Activations into Natural Language — Yida Pan et al., 2024 https://scholar.google.com/scholar?q=LatentQA%3A+Teaching+Language+Models+to+Decode+Activations+into+Natural+Language 10. Activation Oracles — Jesse Karvonen et al., 2026 https://scholar.google.com/scholar?q=Activation+Oracles 11. Confessions — Nikhil Joglekar et al., 2025 https://scholar.google.com/scholar?q=Confessions 12. Self-Report Fine-Tuning — Li et al., 2025 https://scholar.google.com/scholar?q=Self-Report+Fine-Tuning 13. Auditing Language Models for Hidden Objectives — Anthropic, 2025 https://scholar.google.com/scholar?q=Auditing+Language+Models+for+Hidden+Objectives 14. Alignment Faking in Large Language Models — Ryan Greenblatt et al., 2024 https://scholar.google.com/scholar?q=Alignment+Faking+in+Large+Language+Models 15. Towards Eliciting Latent Knowledge from LLMs with Mechanistic Interpretability — Bartosz Cywinski, Emil Ryd, Senthooran Rajamanoharan, Neel Nanda, 2025 https://scholar.google.com/scholar?q=Towards+Eliciting+Latent+Knowledge+from+LLMs+with+Mechanistic+Interpretability 16. Quantifying Elicitation of Latent Capabilities in Language Models — Elizabeth Donoway et al., 2025 https://scholar.google.com/scholar?q=Quantifying+Elicitation+of+Latent+Capabilities+in+Language+Models 17. BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models — Yi Zeng, Weiyu Sun, Tran Huynh, Dawn Song, Bo Li, Ruoxi Jia, 2024 https://scholar.google.com/scholar?q=BEEAR%3A+Embedding-based+Adversarial+Removal+of+Safety+Backdoors+in+Instruction-tuned+Language+Models 18. Investigating Adversarial Trigger Transfer in Large Language Models — Nicholas Meade, Arkil Patel, Siva Reddy, 2024 https://scholar.google.com/scholar?q=Investigating+Adversarial+Trigger+Transfer+in+Large+Language+Models 19. When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models — Kai Wang, Yihao Zhang, Meng Sun, 2025 https://scholar.google.com/scholar?q=When+Thinking+LLMs+Lie%3A+Unveiling+the+Strategic+Deception+in+Representations+of+Reasoning+Models 20. Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort — Xinpeng Wang, Nitish Joshi, Barbara Plank, Rico Angell, He He, 2025 https://scholar.google.com/scholar?q=Is+It+Thinking+or+Cheating%3F+Detecting+Implicit+Reward+Hacking+by+Measuring+Reasoning+Effort 21. AI Post Transformers: Linear Classifier Probes for Intermediate Layers — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-16-linear-classifier-probes-for-intermediat-927ae3.mp3 22. AI Post Transformers: CLUE: Hidden-State Clustering for Non-parametric Verification — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/clue-hidden-state-clustering-for-non-parametric-verification/ 23. AI Post Transformers: Latent Space as a New Computational Paradigm — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-05-latent-space-as-a-new-computational-para-810f39.mp3 24. AI Post Transformers: Language Models are Injective and Hence Invertible — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-21-language-models-are-injective-an-7545e0.mp3 Interactive Visualization: Split Personality Training Reveals Latent Knowledge

Episode metadata supplied by the publisher feed · Published May 10, 2026

Embed this episode

NOW PLAYING

Split Personality Training Reveals Latent Knowledge

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on May 10, 2026.

Is there a transcript available for this episode?

Yes, a full transcript is available for this episode. You can read the complete transcript on the episode page.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!