EPISODE · Jun 11, 2026
Learning at Test Time with Expressive RNN States
from AI Post Transformers
This episode explores the paper Learning to (Learn at Test Time): RNNs with Expressive Hidden States and its attempt to give recurrent models transformer-like long-context behavior without the quadratic cost of attention. It explains why standard RNN hidden states are a bottleneck, compares that limitation to transformers’ growing KV cache, and highlights a key empirical motivation: in the paper’s setup, Mamba’s token-level perplexity improvements flatten around 16k tokens while transformers keep improving deeper into a 32k context. The discussion focuses on the paper’s core idea of test-time training, where the hidden state is treated as a small inner model whose parameters are updated online with a self-supervised learning rule, rather than as a fixed vector summary. Listeners would find it interesting because it connects old fast-weights and dynamic-evaluation ideas to a new systems-level proposal for long-context efficiency, while also noting the open question of whether better perplexity truly translates into stronger retrieval and reasoning. Sources: 1. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, Tatsunori Hashimoto, Carlos Guestrin, 2024 http://arxiv.org/abs/2407.04620 2. Using Fast Weights to Attend to the Recent Past — Jimmy Ba, Geoffrey Hinton, Volodymyr Mnih, Joel Z. Leibo, Catalin Ionescu, 2016 https://arxiv.org/abs/1610.06258 3. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber, 2021 https://arxiv.org/abs/2102.11174 4. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023 https://arxiv.org/abs/2312.00752 5. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, Xiaolong Wang, Tatsunori Hashimoto, Carlos Guestrin, 2024 https://arxiv.org/abs/2407.04620 6. Dynamic Evaluation of Transformer Language Models — Ben Krause, Emmanuel Kahembwe, Iain Murray, Steve Renals, 2019 https://arxiv.org/abs/1904.08378 7. Effective Long-Context Scaling of Foundation Models — Wenhan Xiong et al., 2023 https://arxiv.org/abs/2309.16039 8. Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models — Soham De et al., 2024 https://arxiv.org/abs/2402.19427 9. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality — Tri Dao, Albert Gu, 2024 https://arxiv.org/abs/2405.21060 10. An Empirical Study of Mamba-based Language Models — Roger Waleffe et al., 2024 https://arxiv.org/abs/2406.07887 11. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim et al., 2025 https://arxiv.org/abs/2505.23416 12. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025 https://arxiv.org/abs/2502.16002 13. ReMamba: Equip Mamba with Effective Long-Sequence Modeling — Danlong Yuan et al., 2024 https://arxiv.org/abs/2408.15496 14. LongMamba: Enhancing Mamba's Long Context Capabilities via Training-Free Receptive Field Enlargement — Zhifan Ye et al., 2025 https://arxiv.org/abs/2504.16053 15. Fast-weight Product Key Memory — Tianyu Zhao, Llion Jones, 2026 https://arxiv.org/abs/2601.00671 16. Test-Time Learning for Large Language Models — Jinwu Hu et al., 2025 https://arxiv.org/abs/2505.20633 17. Test-Time Training on Nearest Neighbors for Large Language Models — Moritz Hardt, Yu Sun, 2023 https://arxiv.org/abs/2305.18466 18. AI Post Transformers: Titans: Learning to Memorize at Test Time — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-20-titans-learning-to-memorize-at-test-time-054662.mp3 19. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3 20. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/kimi-linear-efficient-expressive-attention-architecture/ 21. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3 22. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3 23. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3 Interactive Visualization: Learning at Test Time with Expressive RNN States
Embed this episode
NOW PLAYING
Learning at Test Time with Expressive RNN States
No transcript for this episode yet
Similar Episodes
No similar episodes found.