AllMem for Efficient Long-Context Modeling episode artwork

EPISODE · Jun 13, 2026

AllMem for Efficient Long-Context Modeling

from AI Post Transformers

This episode explores AllMem, a method for turning pretrained Qwen3 models into long-context systems that keep exact attention over a recent token window while storing older context in a learned memory. It explains why standard transformer attention becomes prohibitively expensive on long chats, books, codebases, and agent traces, and places AllMem in the broader landscape of sliding-window, sparse-attention, recurrent, and memory-augmented architectures. The discussion highlights the paper’s core argument: a hybrid design can preserve sharp local reasoning, compress the distant past through online memory updates, and approach the quality of full attention without the same compute and KV-cache costs. A listener would find it interesting because it connects concrete systems constraints on phones and servers to a specific recipe for making long-context language models more practical. Sources: 1. AllMem: A Memory-centric Recipe for Efficient Long-context Modeling — Ziming Wang, Xiang Wang, Kailong Peng, Lang Qin, Juan Gabriel Kostelec, Christos Sourmpis, Axel Laborieux, Qinghai Guo, 2026 http://arxiv.org/abs/2602.13680 2. Longformer: The Long-Document Transformer — Iz Beltagy, Matthew E. Peters, Arman Cohan, 2020 https://scholar.google.com/scholar?q=Longformer%3A+The+Long-Document+Transformer 3. Big Bird: Transformers for Longer Sequences — Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Amr Ahmed, et al., 2020 https://scholar.google.com/scholar?q=Big+Bird%3A+Transformers+for+Longer+Sequences 4. Mistral 7B — Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Guillaume Lample, et al., 2023 https://scholar.google.com/scholar?q=Mistral+7B 5. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023 https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks 6. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2024 https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time 7. Artificial Hippocampus Networks for Efficient Long-Context Modeling — Yunhao Fang, Weihao Yu, Shu Zhong, Qinghao Ye, Xuehan Xiong, Lai Wei, 2025 https://scholar.google.com/scholar?q=Artificial+Hippocampus+Networks+for+Efficient+Long-Context+Modeling 8. The Mamba in the Llama: Distilling and Accelerating Hybrid Models — Junxiong Wang, Daniele Paliotta, Avner May, Alexander M. Rush, Tri Dao, 2024 https://scholar.google.com/scholar?q=The+Mamba+in+the+Llama%3A+Distilling+and+Accelerating+Hybrid+Models 9. MesaNet: Sequence Modeling by Locally Optimal Test-Time Training — Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Rif A. Saurous, Guillaume Lajoie, Charlotte Frenkel, Razvan Pascanu, Blaise Agüera y Arcas, João Sacramento, 2025 https://scholar.google.com/scholar?q=MesaNet%3A+Sequence+Modeling+by+Locally+Optimal+Test-Time+Training 10. RULER: What's the Real Context Size of Your Long-Context Language Models? — Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, Boris Ginsburg, 2024 https://scholar.google.com/scholar?q=RULER%3A+What%27s+the+Real+Context+Size+of+Your+Long-Context+Language+Models%3F 11. RazorAttention: Efficient KV Cache Compression Through Retrieval Heads — Hanlin Tang et al., 2024 https://scholar.google.com/scholar?q=RazorAttention%3A+Efficient+KV+Cache+Compression+Through+Retrieval+Heads 12. Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning — Yu Fu et al., 2024 https://scholar.google.com/scholar?q=Not+All+Heads+Matter%3A+A+Head-Level+KV+Cache+Compression+Method+with+Integrated+Retrieval+and+Reasoning 13. Sliding Window Attention Adaptation — Yijiong Yu et al., 2025 https://scholar.google.com/scholar?q=Sliding+Window+Attention+Adaptation 14. HyperAttention: Long-context Attention in Near-Linear Time — Insu Han et al., 2023 https://scholar.google.com/scholar?q=HyperAttention%3A+Long-context+Attention+in+Near-Linear+Time 15. On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention — Yeonju Ro et al., 2025 https://scholar.google.com/scholar?q=On-the-Fly+Adaptive+Distillation+of+Transformer+to+Dual-State+Linear+Attention 16. Test-Time Training on Nearest Neighbors for Large Language Models — Moritz Hardt and Yu Sun, 2023 https://scholar.google.com/scholar?q=Test-Time+Training+on+Nearest+Neighbors+for+Large+Language+Models 17. Test-Time Learning for Large Language Models — Jinwu Hu et al., 2025 https://scholar.google.com/scholar?q=Test-Time+Learning+for+Large+Language+Models 18. Training Large Reasoning Models Efficiently via Progressive Thought Encoding — Zeliang Zhang et al., 2026 https://scholar.google.com/scholar?q=Training+Large+Reasoning+Models+Efficiently+via+Progressive+Thought+Encoding 19. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3 20. AI Post Transformers: MELT: Decoupling Compute From Memory — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-13-melt-decoupling-compute-from-memory-26430c.mp3 21. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/kimi-linear-efficient-expressive-attention-architecture/ 22. AI Post Transformers: Ministral 3: Cascade Distillation for Long-Context Multimodal Models — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-15-cascade-distillation-for-long-context-mu-0ebd1a.mp3 23. AI Post Transformers: Doc-to-LoRA: Internalizing Context as LoRA — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-29-doc-to-lora-internalizing-context-as-lor-8dd5ec.mp3 24. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3 25. AI Post Transformers: Do Transformers Need Three Projections? — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-11-do-transformers-need-three-projections-c227d6.mp3 Interactive Visualization: AllMem for Efficient Long-Context Modeling

Episode metadata supplied by the publisher feed · Published Jun 13, 2026

Embed this episode

NOW PLAYING

AllMem for Efficient Long-Context Modeling

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 13, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!