ELF and Continuous Language Diffusion episode artwork

EPISODE · May 12, 2026

ELF and Continuous Language Diffusion

from AI Post Transformers

This episode explores ELF: Embedded Language Flows, a continuous-time diffusion language model that stays in embedding space until the final decoding step instead of repeatedly snapping back to discrete tokens during generation. It explains how that design lets the model borrow flow-matching and guidance techniques from image diffusion, while arguing that earlier continuous text models may have underperformed because of token-level constraints rather than any fundamental weakness. The discussion highlights reported results on OpenWebText, where a 105M-parameter ELF model achieves better generative perplexity than 170M baselines with far fewer training tokens and fewer sampling steps, while also extending to translation and summarization. It also digs into the main caveat: whether the gains really come from late discretization and continuous-time modeling, or from a bundle of confounded training and inference choices, making the episode interesting both as a technical walkthrough and as a skeptical evaluation of a bold research claim. Sources: 1. ELF and Continuous Language Diffusion https://arxiv.org/pdf/2605.10938 2. Structured Denoising Diffusion Models in Discrete State-Spaces — Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, Rianne van den Berg, 2021 https://scholar.google.com/scholar?q=Structured+Denoising+Diffusion+Models+in+Discrete+State-Spaces 3. Diffusion-LM Improves Controllable Text Generation — Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, Tatsunori B. Hashimoto, 2022 https://scholar.google.com/scholar?q=Diffusion-LM+Improves+Controllable+Text+Generation 4. Simple and Effective Masked Diffusion Language Models — Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin T. Chiu, Alexander M. Rush, Volodymyr Kuleshov, 2024 https://scholar.google.com/scholar?q=Simple+and+Effective+Masked+Diffusion+Language+Models 5. Large Language Diffusion Models — Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li, 2025 https://scholar.google.com/scholar?q=Large+Language+Diffusion+Models 6. Flow Matching for Generative Modeling — Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, Matthew Le, 2023 https://scholar.google.com/scholar?q=Flow+Matching+for+Generative+Modeling 7. Improving and Generalizing Flow-Based Generative Models with Minibatch Optimal Transport — Alexander Tong, Kilian Fatras, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector-Brooks, Guy Wolf, Yoshua Bengio, 2024 https://scholar.google.com/scholar?q=Improving+and+Generalizing+Flow-Based+Generative+Models+with+Minibatch+Optimal+Transport 8. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis — Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Muller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, Robin Rombach, 2024 https://scholar.google.com/scholar?q=Scaling+Rectified+Flow+Transformers+for+High-Resolution+Image+Synthesis 9. Discrete Flow Matching — Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk, Ricky T. Q. Chen, Gabriel Synnaeve, Yossi Adi, Yaron Lipman, 2024 https://scholar.google.com/scholar?q=Discrete+Flow+Matching 10. Self-conditioned Embedding Diffusion for Text Generation — Robin Strudel, Corentin Tallec, Florent Altche, Yilun Du, Yaroslav Ganin, Arthur Mensch, Will Grathwohl, Nikolay Savinov, Sander Dieleman, Laurent Sifre, Remi Leblond, 2022 https://scholar.google.com/scholar?q=Self-conditioned+Embedding+Diffusion+for+Text+Generation 11. Difformer: Empowering Diffusion Models on the Embedding Space for Text Generation — Zhujin Gao, Junliang Guo, Xu Tan, Yongxin Zhu, Fang Zhang, Jiang Bian, Linli Xu, 2022 https://scholar.google.com/scholar?q=Difformer%3A+Empowering+Diffusion+Models+on+the+Embedding+Space+for+Text+Generation 12. LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling — Yuxin Chen, Chumeng Liang, Hangke Sui, Ruihan Guo, Chaoran Cheng, Jiaxuan You, Ge Liu, 2026 https://scholar.google.com/scholar?q=LangFlow%3A+Continuous+Diffusion+Rivals+Discrete+in+Language+Modeling 13. Classifier-Free Diffusion Guidance — Jonathan Ho, Tim Salimans, 2021 https://scholar.google.com/scholar?q=Classifier-Free+Diffusion+Guidance 14. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models — Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, Mark Chen, 2021 https://scholar.google.com/scholar?q=GLIDE%3A+Towards+Photorealistic+Image+Generation+and+Editing+with+Text-Guided+Diffusion+Models 15. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding — Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, Mohammad Norouzi, 2022 https://scholar.google.com/scholar?q=Photorealistic+Text-to-Image+Diffusion+Models+with+Deep+Language+Understanding 16. High-Resolution Image Synthesis with Latent Diffusion Models — Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer, 2021 https://scholar.google.com/scholar?q=High-Resolution+Image+Synthesis+with+Latent+Diffusion+Models 17. MDLM: Masked Diffusion Language Models — likely the MDLM authors cited as [56] in the paper, 2024 https://scholar.google.com/scholar?q=MDLM%3A+Masked+Diffusion+Language+Models 18. Duo — likely the Duo authors cited as [57] in the paper, 2025 https://scholar.google.com/scholar?q=Duo 19. Latent Diffusion Models — Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer, 2022 https://scholar.google.com/scholar?q=Latent+Diffusion+Models 20. LangFlow — the LangFlow authors cited as [10] in the paper, 2026 https://scholar.google.com/scholar?q=LangFlow 21. FLM — the FLM authors cited as [30] in the paper, 2026 https://scholar.google.com/scholar?q=FLM 22. Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution — Aaron Lou, Chenlin Meng, Stefano Ermon, 2024 https://scholar.google.com/scholar?q=Discrete+Diffusion+Modeling+by+Estimating+the+Ratios+of+the+Data+Distribution 23. Scaling Behavior of Discrete Diffusion Language Models — Dimitri von Rutte, Janis Fluri, Omead Pooladzandi, Bernhard Scholkopf, Thomas Hofmann, Antonio Orvieto, 2025 https://scholar.google.com/scholar?q=Scaling+Behavior+of+Discrete+Diffusion+Language+Models 24. Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner — Cai Zhou, Chenxiao Yang, Yi Hu, Chenyu Wang, Chubin Zhang, Muhan Zhang, Lester Mackey, Tommi Jaakkola, Stephen Bates, Dinghuai Zhang, 2025 https://scholar.google.com/scholar?q=Coevolutionary+Continuous+Discrete+Diffusion%3A+Make+Your+Diffusion+Language+Model+a+Latent+Reasoner 25. Stay on Topic with Classifier-Free Guidance — Guillaume Sanchez, Honglu Fan, Alexander Spangher, Elad Levi, Pawan Sasanka Ammanamanchi, Stella Biderman, 2023 https://scholar.google.com/scholar?q=Stay+on+Topic+with+Classifier-Free+Guidance 26. Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking — Pengxiang Li, Shilin Yan, Joey Tsai, Renrui Zhang, Ruichuan An, Ziyu Guo, Xiaowei Gao, 2025 https://scholar.google.com/scholar?q=Adaptive+Classifier-Free+Guidance+via+Dynamic+Low-Confidence+Masking 27. Studying Classifier(-Free) Guidance From a Classifier-Centric Perspective — Xiaoming Zhao, Alexander G. Schwing, 2025 https://scholar.google.com/scholar?q=Studying+Classifier%28-Free%29+Guidance+From+a+Classifier-Centric+Perspective 28. DEPT: Decoupled Embeddings for Pre-training Language Models — Alex Iacob, Lorenzo Sani, Meghdad Kurmanji, William F. Shen, Xinchi Qiu, Dongqi Cai, Yan Gao, Nicholas D. Lane, 2024 https://scholar.google.com/scholar?q=DEPT%3A+Decoupled+Embeddings+for+Pre-training+Language+Models 29. AI Post Transformers: Generative Modeling via Drifting in One Step — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-generative-modeling-via-drifting-in-one-671da0.mp3 30. AI Post Transformers: VL-JEPA for Vision-Language Semantic Prediction — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-12-vl-jepa-for-vision-language-semantic-pre-69c9f4.mp3 31. AI Post Transformers: Why Transformers Fail at Counting — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-08-why-transformers-fail-at-counting-137924.mp3 32. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3 Interactive Visualization: ELF and Continuous Language Diffusion

Episode metadata supplied by the publisher feed · Published May 12, 2026

Embed this episode

NOW PLAYING

ELF and Continuous Language Diffusion

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on May 12, 2026.

Is there a transcript available for this episode?

Yes, a full transcript is available for this episode. You can read the complete transcript on the episode page.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!