EPISODE · Jul 31, 2026
MemPO: Teaching Agents to Write Their Own Memory
from AI Post Transformers
This episode explores MemPO, a self-memory policy optimization framework for long-horizon AI agents developed by researchers at Tsinghua University and Alibaba's Tongyi Lab. The discussion contrasts MemPO's approach against the dominant ReAct pattern, which accumulates full interaction history and suffers from both ballooning token costs and the "lost in the middle" degradation documented in prior research, as well as against passive retrieval-based memory systems like MemGPT and Mem0 that rely on embedding similarity rather than task outcomes. The hosts unpack how MemPO trains an agent to write compressed memory notes as a learned, RL-optimized action — discarding raw tool outputs and reasoning traces at each step in favor of a single distilled note — using Group Relative Policy Optimization to solve the credit-assignment problem of rewarding intermediate memory decisions from a single end-of-trajectory success signal. Listeners interested in agent architecture, RL training objectives, or the tradeoffs between context-window scaling and structured memory will find the episode's walk-through of the mem/think/tool_call decomposition particularly useful. The conversation also traces the intellectual lineage of the ideas, from Minsky's original framing of credit assignment to retrieval-augmented generation's origins at Facebook AI Research. Sources: 1. MemPO: Self-Memory Policy Optimization for Long-Horizon Agents — Ruoran Li, Xinghua Zhang, Haiyang Yu, Shitong Duan, Xiang Li, Wenxin Xiang, Chonghua Liao, Xudong Guo, Yongbin Li, Jinli Suo, 2026 http://arxiv.org/abs/2603.00680 2. Policy Gradient Methods for Reinforcement Learning with Function Approximation — Richard S. Sutton, David McAllester, Satinder Singh, Yishay Mansour, 1999/2000 https://scholar.google.com/scholar?q=Policy+Gradient+Methods+for+Reinforcement+Learning+with+Function+Approximation 3. High-Dimensional Continuous Control Using Generalized Advantage Estimation — John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, Pieter Abbeel, 2016 https://scholar.google.com/scholar?q=High-Dimensional+Continuous+Control+Using+Generalized+Advantage+Estimation 4. RUDDER: Return Decomposition for Delayed Rewards — Jose A. Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Adler, Johannes Brandstetter, Sepp Hochreiter, 2019 https://scholar.google.com/scholar?q=RUDDER%3A+Return+Decomposition+for+Delayed+Rewards 5. Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe, 2023 https://scholar.google.com/scholar?q=Let%27s+Verify+Step+by+Step 6. Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents — Guoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan, Wei Yao, Yong Deng, Xiaofeng Wu, Zhenzhe Ying, 2025 https://scholar.google.com/scholar?q=Information+Gain-based+Policy+Optimization%3A+A+Simple+and+Effective+Approach+for+Multi-Turn+LLM+Agents 7. MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents — Zijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim, Alok Prakash, Daniela Rus, Jinhua Zhao, Bryan Kian Hsiang Low, Paul Pu Liang, 2025 https://scholar.google.com/scholar?q=MEM1%3A+Learning+to+Synergize+Memory+and+Reasoning+for+Efficient+Long-Horizon+Agents 8. Reflexion: Language Agents with Verbal Reinforcement Learning — Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao, 2023 https://scholar.google.com/scholar?q=Reflexion%3A+Language+Agents+with+Verbal+Reinforcement+Learning 9. Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning — Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Arik, Dong Wang, Hamed Zamani, Jiawei Han, 2025 https://scholar.google.com/scholar?q=Search-R1%3A+Training+LLMs+to+Reason+and+Leverage+Search+Engines+with+Reinforcement+Learning 10. Attnpo: Attention-guided Process Supervision for Efficient Reasoning — Shuaiyi Nie, Siyu Ding, Wenyuan Zhang, Linhao Yu, Tianmeng Yang, Yao Chen, Tingwen Liu, Weichong Yin, Yu Sun, Hua Wu, 2026 https://scholar.google.com/scholar?q=Attnpo%3A+Attention-guided+Process+Supervision+for+Efficient+Reasoning 11. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, Joseph E. Gonzalez, 2024 https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems Interactive Visualization: MemPO: Teaching Agents to Write Their Own Memory
Embed this episode
NOW PLAYING
MemPO: Teaching Agents to Write Their Own Memory
No transcript for this episode yet
Similar Episodes
No similar episodes found.