Geometric Memory in Deep Sequence Models episode artwork

EPISODE · May 3, 2026

Geometric Memory in Deep Sequence Models

from AI Post Transformers

This episode explores whether deep sequence models store knowledge as simple associative lookups or as geometric memories that encode broader relational structure. It discusses a recent paper arguing that, after memorizing graph facts in their weights, sequence models can answer multi-hop path queries as if they were making a much shorter move through embedding space, with learned representations resembling graph-embedding methods like node2vec and DeepWalk. The conversation highlights why that matters mechanistically: it suggests some forms of reasoning may be amortized into the model’s parameters during training rather than reconstructed step by step at inference time. Listeners would find it interesting for its sharp debate over what counts as real reasoning versus a clever shortcut, and for its caution about how far results from synthetic graph settings should generalize to large language models in the wild. Sources: 1. Deep sequence models tend to memorize geometrically; it is unclear why — Shahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld, Sanjiv Kumar, 2025 http://arxiv.org/abs/2510.26745 2. DeepWalk: Online Learning of Social Representations — Bryan Perozzi, Rami Al-Rfou, Steven Skiena, 2014 https://scholar.google.com/scholar?q=DeepWalk%3A+Online+Learning+of+Social+Representations 3. node2vec: Scalable Feature Learning for Networks — Aditya Grover, Jure Leskovec, 2016 https://scholar.google.com/scholar?q=node2vec%3A+Scalable+Feature+Learning+for+Networks 4. Birth of a Transformer: A Memory Viewpoint — Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, Leon Bottou, 2023 https://scholar.google.com/scholar?q=Birth+of+a+Transformer%3A+A+Memory+Viewpoint 5. Deep sequence models tend to memorize geometrically; it is unclear why — Shahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld, Sanjiv Kumar, 2025 https://scholar.google.com/scholar?q=Deep+sequence+models+tend+to+memorize+geometrically%3B+it+is+unclear+why 6. The Pitfalls of Next-Token Prediction — Gregor Bachmann, Vaishnavh Nagarajan, 2024 https://scholar.google.com/scholar?q=The+Pitfalls+of+Next-Token+Prediction 7. How Transformers Learn to Plan via Multi-Token Prediction — Jianhao Huang, Zhanpeng Zhou, Renqiu Xia, Baharan Mirzasoleiman, Weijie Su, Wei Huang, 2026 https://scholar.google.com/scholar?q=How+Transformers+Learn+to+Plan+via+Multi-Token+Prediction 8. DeepSeek-V3 Technical Report — DeepSeek-AI and collaborators, 2024 https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report 9. Language Models, Graph Searching, and Supervision Adulteration: When More Supervision is Less and How to Make More More — Arvid Frydenlund, 2025 https://scholar.google.com/scholar?q=Language+Models%2C+Graph+Searching%2C+and+Supervision+Adulteration%3A+When+More+Supervision+is+Less+and+How+to+Make+More+More 10. Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries — Eden Biran, Daniela Gottesman, Sohee Yang, Mor Geva, and Amir Globerson, 2024 https://scholar.google.com/scholar?q=Hopping+Too+Late%3A+Exploring+the+Limitations+of+Large+Language+Models+on+Multi-Hop+Queries 11. The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" — Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans, 2024 https://scholar.google.com/scholar?q=The+Reversal+Curse%3A+LLMs+trained+on+%22A+is+B%22+fail+to+learn+%22B+is+A%22 12. In-Context Denoising with One-Layer Transformers: Connections between Attention and Associative Memory Retrieval — Matthew Smart, Alberto Bietti, Anirvan M. Sengupta, 2025 https://scholar.google.com/scholar?q=In-Context+Denoising+with+One-Layer+Transformers%3A+Connections+between+Attention+and+Associative+Memory+Retrieval 13. In-Context Learning as Conditioned Associative Memory Retrieval — Weimin Wu, Teng-Yun Hsiao, Jerry Yao-Chieh Hu, Wenxin Zhang, Han Liu, 2025 https://scholar.google.com/scholar?q=In-Context+Learning+as+Conditioned+Associative+Memory+Retrieval 14. Position-Aware Relational Transformer for Knowledge Graph Embedding — Guangyao Li, Zequn Sun, Wei Hu, Gong Cheng, et al., 2023 https://scholar.google.com/scholar?q=Position-Aware+Relational+Transformer+for+Knowledge+Graph+Embedding 15. Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data — Rishabh Ranjan, Valter Hudovernik, Mark Znidar, Charilaos Kanatsoulis, Roshan Upendra, Mahmoud Mohammadi, Joe Meyer, Tom Palczewski, Carlos Guestrin, Jure Leskovec, 2025 https://scholar.google.com/scholar?q=Relational+Transformer%3A+Toward+Zero-Shot+Foundation+Models+for+Relational+Data 16. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3 17. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3 Interactive Visualization: Geometric Memory in Deep Sequence Models

Episode metadata supplied by the publisher feed · Published May 3, 2026

Embed this episode

NOW PLAYING

Geometric Memory in Deep Sequence Models

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on May 3, 2026.

Is there a transcript available for this episode?

Yes, a full transcript is available for this episode. You can read the complete transcript on the episode page.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!