Learning at Test Time with Expressive RNN States episode artwork

EPISODE · Jun 11, 2026

Learning at Test Time with Expressive RNN States

from AI Post Transformers

This episode explores the paper Learning to (Learn at Test Time): RNNs with Expressive Hidden States and its attempt to give recurrent models transformer-like long-context behavior without the quadratic cost of attention. It explains why standard RNN hidden states are a bottleneck, compares that limitation to transformers’ growing KV cache, and highlights a key empirical motivation: in the paper’s setup, Mamba’s token-level perplexity improvements flatten around 16k tokens while transformers keep improving deeper into a 32k context. The discussion focuses on the paper’s core idea of test-time training, where the hidden state is treated as a small inner model whose parameters are updated online with a self-supervised learning rule, rather than as a fixed vector summary. Listeners would find it interesting because it connects old fast-weights and dynamic-evaluation ideas to a new systems-level proposal for long-context efficiency, while also noting the open question of whether better perplexity truly translates into stronger retrieval and reasoning. Sources: 1. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, Tatsunori Hashimoto, Carlos Guestrin, 2024 http://arxiv.org/abs/2407.04620 2. Using Fast Weights to Attend to the Recent Past — Jimmy Ba, Geoffrey Hinton, Volodymyr Mnih, Joel Z. Leibo, Catalin Ionescu, 2016 https://arxiv.org/abs/1610.06258 3. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber, 2021 https://arxiv.org/abs/2102.11174 4. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023 https://arxiv.org/abs/2312.00752 5. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, Xiaolong Wang, Tatsunori Hashimoto, Carlos Guestrin, 2024 https://arxiv.org/abs/2407.04620 6. Dynamic Evaluation of Transformer Language Models — Ben Krause, Emmanuel Kahembwe, Iain Murray, Steve Renals, 2019 https://arxiv.org/abs/1904.08378 7. Effective Long-Context Scaling of Foundation Models — Wenhan Xiong et al., 2023 https://arxiv.org/abs/2309.16039 8. Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models — Soham De et al., 2024 https://arxiv.org/abs/2402.19427 9. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality — Tri Dao, Albert Gu, 2024 https://arxiv.org/abs/2405.21060 10. An Empirical Study of Mamba-based Language Models — Roger Waleffe et al., 2024 https://arxiv.org/abs/2406.07887 11. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim et al., 2025 https://arxiv.org/abs/2505.23416 12. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025 https://arxiv.org/abs/2502.16002 13. ReMamba: Equip Mamba with Effective Long-Sequence Modeling — Danlong Yuan et al., 2024 https://arxiv.org/abs/2408.15496 14. LongMamba: Enhancing Mamba's Long Context Capabilities via Training-Free Receptive Field Enlargement — Zhifan Ye et al., 2025 https://arxiv.org/abs/2504.16053 15. Fast-weight Product Key Memory — Tianyu Zhao, Llion Jones, 2026 https://arxiv.org/abs/2601.00671 16. Test-Time Learning for Large Language Models — Jinwu Hu et al., 2025 https://arxiv.org/abs/2505.20633 17. Test-Time Training on Nearest Neighbors for Large Language Models — Moritz Hardt, Yu Sun, 2023 https://arxiv.org/abs/2305.18466 18. AI Post Transformers: Titans: Learning to Memorize at Test Time — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-20-titans-learning-to-memorize-at-test-time-054662.mp3 19. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3 20. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/kimi-linear-efficient-expressive-attention-architecture/ 21. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3 22. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3 23. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3 Interactive Visualization: Learning at Test Time with Expressive RNN States

Episode metadata supplied by the publisher feed · Published Jun 11, 2026

Embed this episode

NOW PLAYING

Learning at Test Time with Expressive RNN States

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 11, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!