EPISODE · May 13, 2026
MELT: Decoupling Compute From Memory
from AI Post Transformers
This episode explores MELT, a looped language model architecture that aims to preserve latent reasoning benefits while preventing KV-cache memory from growing with every reasoning pass. It explains how the paper reframes the problem as a systems and architecture challenge, replacing per-loop cached attention state with a single shared, gated cache per layer inspired by recurrent models like LSTMs and Universal Transformers. The discussion weighs whether this is a genuine shift in reasoning architecture or a narrower engineering improvement, ultimately arguing that the paper’s real contribution is efficient cache management rather than a wholly new paradigm. Listeners would find it interesting for its clear breakdown of inference-time compute scaling, latent reasoning, and why memory bottlenecks could shape the future of practical reasoning models. Sources: 1. MELT: Decoupling Compute From Memory https://arxiv.org/pdf/2605.07721 2. Scaling Latent Reasoning via Looped Language Models — Rui-Jie Zhu et al., 2025 https://scholar.google.com/scholar?q=Scaling+Latent+Reasoning+via+Looped+Language+Models 3. Universal Transformers — Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, Lukasz Kaiser, 2018 https://scholar.google.com/scholar?q=Universal+Transformers 4. Reducing Transformer Key-Value Cache Size with Cross-Layer Attention — William Brandon, Mayank Mishra, Aniruddha Nrusimha, Rameswar Panda, Jonathan Ragan Kelly, 2024 https://scholar.google.com/scholar?q=Reducing+Transformer+Key-Value+Cache+Size+with+Cross-Layer+Attention 5. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, Sumit Sanghai, 2023 https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints 6. Fast Transformer Decoding: One Write-Head is All You Need — Noam Shazeer, 2019 https://scholar.google.com/scholar?q=Fast+Transformer+Decoding%3A+One+Write-Head+is+All+You+Need 7. Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning — Zeyu Xing, Xing Li, Hui-Ling Zhen, Mingxuan Yuan, Sinno Jialin Pan, 2026 https://scholar.google.com/scholar?q=Beyond+Speedup+--+Utilizing+KV+Cache+for+Sampling+and+Reasoning 8. MemShare: Memory Efficient Inference for Large Reasoning Models through KV Cache Reuse — Kaiwen Chen, Xin Tan, Minchen Yu, Hong Xu, 2025 https://scholar.google.com/scholar?q=MemShare%3A+Memory+Efficient+Inference+for+Large+Reasoning+Models+through+KV+Cache+Reuse 9. Beyond Attention: Breaking the Limits of Transformer Context Length with Recurrent Memory — Aydar Bulatov, Yurii Kuratov, Yermek Kapushev, Mikhail Burtsev, 2024 https://scholar.google.com/scholar?q=Beyond+Attention%3A+Breaking+the+Limits+of+Transformer+Context+Length+with+Recurrent+Memory 10. Towards Understanding Distilled Reasoning Models: A Representational Approach — David D. Baek, Max Tegmark, 2025 https://scholar.google.com/scholar?q=Towards+Understanding+Distilled+Reasoning+Models%3A+A+Representational+Approach 11. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3 12. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3 13. AI Post Transformers: RetrievalAttention for Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-17-retrievalattention-for-long-context-llm-ddf566.mp3 14. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3 15. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3 16. AI Post Transformers: Jet-Nemotron and PostNAS for Faster Long Context — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-24-jet-nemotron-and-postnas-for-faster-long-436381.mp3 17. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3 18. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3 Interactive Visualization: MELT: Decoupling Compute From Memory
Embed this episode
NOW PLAYING
MELT: Decoupling Compute From Memory
No transcript for this episode yet
Similar Episodes
No similar episodes found.