Latent Reasoning with Normalizing Flows episode artwork

EPISODE · Jun 10, 2026

Latent Reasoning with Normalizing Flows

from AI Post Transformers

This episode explores Latent Reasoning with Normalizing Flows, a paper that asks whether a standard left-to-right transformer can do its intermediate reasoning in continuous latent states instead of spelling every step out as text. It explains how the method uses a frozen VAE during training to compress written rationales into short latent sequences, then uses shallow normalizing flows so the same autoregressive backbone can predict both latent thought slots and normal answer tokens while preserving exact likelihoods, sampling, and KV-cache-friendly decoding. The discussion highlights matched coding results on Qwen3-8B-Base, where the reported benchmark average rises from 55.8 for the base model to 68.8 for NF-CoT Unified and 70.1 after latent-space reinforcement learning, with strong pass@k gains that suggest better exploration of multiple solution paths. Listeners would find it interesting because it frames latent reasoning as a practical alternative to verbose chain-of-thought, while also noting the current evidence is still narrow, centered on one post-trained coding model and not uniformly better than diffusion baselines on every benchmark. Sources: 1. Latent Reasoning with Normalizing Flows — Guancheng Tu, Xiangjun Fu, Suhao Yu, Yao Tang, Haoqiang Kang, Lianhui Qin, Yizhe Zhang, Jiatao Gu, 2026 http://arxiv.org/abs/2606.06447 2. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Jason Wei, Xuezhi Wang, Denny Zhou, Quoc Le, et al., 2022 https://scholar.google.com/scholar?q=Chain-of-Thought+Prompting+Elicits+Reasoning+in+Large+Language+Models 3. Training Large Language Models to Reason in a Continuous Latent Space — Shibo Hao, Sainbayar Sukhbaatar, Zhiting Hu, Jason Weston, Yuandong Tian, et al., 2024 preprint; COLM 2025 https://scholar.google.com/scholar?q=Training+Large+Language+Models+to+Reason+in+a+Continuous+Latent+Space 4. Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning — Xinghao Chen, Anhao Zhao, Xiaoyu Shen, et al., 2025 https://scholar.google.com/scholar?q=Reasoning+Beyond+Language%3A+A+Comprehensive+Survey+on+Latent+Chain-of-Thought+Reasoning 5. LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning — Haoqiang Kang, Yizhe Zhang, Navdeep Jaitly, Yi-An Ma, Lianhui Qin, et al., 2025 https://scholar.google.com/scholar?q=LaDiR%3A+Latent+Diffusion+Enhances+LLMs+for+Text+Reasoning 6. CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation — Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du, Yulan He, 2025 https://scholar.google.com/scholar?q=CODI%3A+Compressing+Chain-of-Thought+into+Continuous+Space+via+Self-Distillation 7. Normalizing Flows are Capable Generative Models — Shuangfei Zhai, Ruixiang Zhang, Preetum Nakkiran, David Berthelot, Jiatao Gu, Huangjie Zheng, Tianrong Chen, Miguel Angel Bautista, Navdeep Jaitly, Josh Susskind, 2024 https://scholar.google.com/scholar?q=Normalizing+Flows+are+Capable+Generative+Models 8. Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought — Hanlin Zhu, Shibo Hao, Zhiting Hu, Jiantao Jiao, Stuart Russell, Yuandong Tian, 2025 https://scholar.google.com/scholar?q=Reasoning+by+Superposition%3A+A+Theoretical+Perspective+on+Chain+of+Continuous+Thought 9. Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure — Zirui Li, Xuefeng Bai, Kehai Chen, Yizhi Li, Jian Yang, Chenghua Lin, Min Zhang, 2026 https://scholar.google.com/scholar?q=Dynamics+Within+Latent+Chain-of-Thought%3A+An+Empirical+Study+of+Causal+Structure 10. Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning — Jingcheng Deng, Zihao Wei, Liang Pang, Junhong Wu, Shicheng Xu, Zenghao Duan, Huawei Shen, 2026 https://scholar.google.com/scholar?q=Latent-GRPO%3A+Group+Relative+Policy+Optimization+for+Latent+Reasoning 11. Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision — Dawei Zhu et al., 2025 https://arxiv.org/abs/2502.20790 12. Supervised Chain of Thought — Xiang Zhang and Dujian Ding, 2024 https://arxiv.org/abs/2410.14198 13. Large language models can learn and generalize steganographic chain-of-thought under process supervision — Joey Skaf et al., 2025 https://arxiv.org/abs/2506.01926 14. Hybrid Latent Reasoning via Reinforcement Learning — Zhenrui Yue et al., 2025 https://arxiv.org/abs/2505.18454 15. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping et al., 2025 https://arxiv.org/abs/2502.05171 16. R-KV: Redundancy-aware KV Cache Compression for Reasoning Models — Zefan Cai et al., 2025 https://arxiv.org/abs/2505.24133 17. Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning — Yu Fu et al., 2024 https://arxiv.org/abs/2410.19258 18. AI Post Transformers: Generative Recursive Reasoning in Latent Space — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-21-generative-recursive-reasoning-in-latent-a9371d.mp3 19. AI Post Transformers: MELT: Decoupling Compute From Memory — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-13-melt-decoupling-compute-from-memory-26430c.mp3 20. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3 21. AI Post Transformers: Gradient Descent at Inference Time for LLM Reasoning — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-10-gradient-descent-at-inference-time-for-l-20617d.mp3 22. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3 23. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3 Interactive Visualization: Latent Reasoning with Normalizing Flows

Episode metadata supplied by the publisher feed · Published Jun 10, 2026

Embed this episode

NOW PLAYING

Latent Reasoning with Normalizing Flows

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 10, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!