EPISODE · May 10, 2026
Reasoning Theater and Unfaithful Chain-of-Thought
from AI Post Transformers
This episode explores a paper arguing that a model’s internal answer belief can form well before its visible chain-of-thought reveals it, raising doubts about whether reasoning traces are true explanations or polished post hoc narratives. It explains core ideas such as chain-of-thought faithfulness, activation monitoring, mechanistic interpretability, confidence calibration, and dynamic inference, framing the broader safety question of whether text reasoning can really serve as an audit trail. The discussion focuses on the paper’s method of comparing internal activation probes, forced early answers, and monitors of partial written reasoning to test whether models “know” the answer before their text shows it. Listeners would find it interesting because it connects interpretability research to practical concerns about oversight, trust, and compute efficiency, while contrasting easy recall-heavy benchmarks with harder multistep science questions where genuine belief updates may still happen during inference. Sources: 1. Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought — Siddharth Boppana, Annabel Ma, Max Loeffler, Raphael Sarfati, Eric Bigelow, Atticus Geiger, Owen Lewis, Jack Merullo, 2026 http://arxiv.org/abs/2603.05488 2. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou, 2022 https://scholar.google.com/scholar?q=Chain-of-Thought+Prompting+Elicits+Reasoning+in+Large+Language+Models 3. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting — Miles Turpin, Julian Michael, Ethan Perez, Samuel R. Bowman, 2023 https://scholar.google.com/scholar?q=Language+Models+Don%27t+Always+Say+What+They+Think%3A+Unfaithful+Explanations+in+Chain-of-Thought+Prompting 4. Measuring Faithfulness in Chain-of-Thought Reasoning — Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion and others, 2023 https://scholar.google.com/scholar?q=Measuring+Faithfulness+in+Chain-of-Thought+Reasoning 5. Reasoning Models Don't Always Say What They Think — Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, Vlad Mikulik, Samuel R. Bowman, Jan Leike, Jared Kaplan, Ethan Perez, 2025 https://scholar.google.com/scholar?q=Reasoning+Models+Don%27t+Always+Say+What+They+Think 6. Performative Thinking? The Brittle Correlation between CoT Length and Problem Complexity — Vivek Palod, Karthik Valmeekam, Kyle Stechly, Subbarao Kambhampati, 2025 https://scholar.google.com/scholar?q=Performative+Thinking%3F+The+Brittle+Correlation+between+CoT+Length+and+Problem+Complexity 7. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety — Tomek Korbak, Mikita Balesni, Elizabeth Barnes, Yoshua Bengio, et al., 2025 https://scholar.google.com/scholar?q=Chain+of+Thought+Monitorability%3A+A+New+and+Fragile+Opportunity+for+AI+Safety 8. Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification — Anqi Zhang, Yulin Chen, Jane Pan, Chen Zhao, Aurojit Panda, Jinyang Li, He He, 2025 https://scholar.google.com/scholar?q=Reasoning+Models+Know+When+They%27re+Right%3A+Probing+Hidden+States+for+Self-Verification 9. A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior — Harry Mayne, Justin Singh Kang, Dewi Gould, Kannan Ramchandran, Adam Mahdi, Noah Y. Siegel, 2026 https://scholar.google.com/scholar?q=A+Positive+Case+for+Faithfulness%3A+LLM+Self-Explanations+Help+Predict+Model+Behavior 10. Base Models Know How to Reason, Thinking Models Learn When — Constantin Venhoff, Ivan Arcuschin, Philip Torr, Arthur Conmy, Neel Nanda, 2025 https://scholar.google.com/scholar?q=Base+Models+Know+How+to+Reason%2C+Thinking+Models+Learn+When 11. Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps — Martin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan Belinkov, 2025 https://scholar.google.com/scholar?q=Measuring+Chain+of+Thought+Faithfulness+by+Unlearning+Reasoning+Steps 12. Faithful Chain-of-Thought Reasoning — Qing Lyu et al., 2023 https://scholar.google.com/scholar?q=Faithful+Chain-of-Thought+Reasoning 13. Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning — Wenkai Yang, Shuming Ma, Yankai Lin, Furu Wei, 2025 https://scholar.google.com/scholar?q=Towards+Thinking-Optimal+Scaling+of+Test-Time+Compute+for+LLM+Reasoning 14. AI Post Transformers: Do Language Models Know Their Limits — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-01-do-language-models-know-their-limits-48e444.mp3 15. AI Post Transformers: Selective Classification with Deep Neural Networks — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-01-selective-classification-with-deep-neura-bed8cb.mp3 16. AI Post Transformers: Stabilizing Efficient Reasoning with Step-Level Advantage Selection — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-29-stabilizing-efficient-reasoning-with-ste-1e589d.mp3 17. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3 Interactive Visualization: Reasoning Theater and Unfaithful Chain-of-Thought
Embed this episode
NOW PLAYING
Reasoning Theater and Unfaithful Chain-of-Thought
No transcript for this episode yet
Similar Episodes
No similar episodes found.