EPISODE · May 15, 2026
When Many-Shot CoT Becomes Test-Time Learning
from AI Post Transformers
This episode explores a May 13, 2026 arXiv paper arguing that many-shot chain-of-thought prompting can act less like simple retrieval and more like a form of test-time learning, where structured context helps a model reason during inference without changing its weights. It examines the paper’s main claims that adding many reasoning demonstrations does not reliably help across all settings, that semantically similar examples can fail when their reasoning procedures are not actually usable, and that the order of demonstrations becomes more important as prompts grow longer. The discussion also focuses on the paper’s Curvilinear Demonstration Selection approach, which treats prompt construction more like designing a lesson plan than doing nearest-neighbor search, with especially notable gains on geometry tasks. Listeners would find it interesting because it challenges a common assumption behind huge context windows: more examples are not automatically better, and effective prompting may depend on curriculum design, model capabilities, and the compatibility of reasoning traces. Sources: 1. When Many-Shot CoT Becomes Test-Time Learning https://arxiv.org/pdf/2605.13511 2. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou, 2022 https://arxiv.org/abs/2201.11903 3. Large Language Models are Zero-Shot Reasoners — Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, Yusuke Iwasawa, 2022 https://arxiv.org/abs/2205.11916 4. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2022 https://arxiv.org/abs/2203.11171 5. Automatic Chain of Thought Prompting in Large Language Models — Zhuosheng Zhang, Aston Zhang, Mu Li, Alex Smola, 2022 https://arxiv.org/abs/2210.03493 6. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, Moritz Hardt, 2020 https://arxiv.org/abs/1909.13231 7. What Learning Algorithm is In-Context Learning? Investigations with Linear Models — Ekin Akyurek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, Denny Zhou, 2022 https://arxiv.org/abs/2211.15661 8. Transformers Learn In-Context by Gradient Descent — Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, Joao Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, Max Vladymyrov, 2022 https://arxiv.org/abs/2212.07677 9. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters — Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar, 2024 https://arxiv.org/abs/2408.03314 10. What Makes Good In-Context Examples for GPT-3? — Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, Weizhu Chen, 2021 https://arxiv.org/abs/2101.06804 11. Learning To Retrieve Prompts for In-Context Learning — Ohad Rubin, Jonathan Herzig, Jonathan Berant, 2021 https://arxiv.org/abs/2112.08633 12. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity — Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, Pontus Stenetorp, 2021 https://arxiv.org/abs/2104.08786 13. Many-Shot CoT-ICL: Making In-Context Learning Truly Learn — Tsz Ting Chung, Lemao Liu, Mo Yu, Dit-Yan Yeung, 2026 https://arxiv.org/abs/2605.13511 14. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? — Sewon Min, Mike Lewis, Hannaneh Hajishirzi, Luke Zettlemoyer, 2022 https://scholar.google.com/scholar?q=Rethinking+the+Role+of+Demonstrations%3A+What+Makes+In-Context+Learning+Work%3F 15. Test-Time Compute: Scaling Language Models with More Thinking — OpenAI, 2024 https://scholar.google.com/scholar?q=Test-Time+Compute%3A+Scaling+Language+Models+with+More+Thinking 16. ALR2: A Retrieve-then-Reason Framework for Long-context Question Answering — Huayang Li, Pat Verga, Priyanka Sen, Bowen Yang, Vijay Viswanathan, Patrick Lewis, Taro Watanabe, Yixuan Su, 2024 https://scholar.google.com/scholar?q=ALR2%3A+A+Retrieve-then-Reason+Framework+for+Long-context+Question+Answering 17. Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models — Yifu Qiu, Varun Embar, Yizhe Zhang, Navdeep Jaitly, Shay B. Cohen, Benjamin Han, 2025 https://scholar.google.com/scholar?q=Eliciting+In-context+Retrieval+and+Reasoning+for+Long-context+Large+Language+Models 18. Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions — Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal, 2023 https://scholar.google.com/scholar?q=Interleaving+Retrieval+with+Chain-of-Thought+Reasoning+for+Knowledge-Intensive+Multi-Step+Questions 19. Active Prompting with Chain-of-Thought for Large Language Models — Shizhe Diao, Pengcheng Wang, Yong Lin, Rui Pan, Xiang Liu, Tong Zhang, 2023 https://scholar.google.com/scholar?q=Active+Prompting+with+Chain-of-Thought+for+Large+Language+Models 20. Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models — Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, Weizhu Chen, 2023 https://scholar.google.com/scholar?q=Synthetic+Prompting%3A+Generating+Chain-of-Thought+Demonstrations+for+Large+Language+Models 21. Contrastive Chain-of-Thought Prompting — Yew Ken Chia, Guizhen Chen, Luu Anh Tuan, Soujanya Poria, Lidong Bing, 2023 https://scholar.google.com/scholar?q=Contrastive+Chain-of-Thought+Prompting 22. Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs — Rachit Bansal, Aston Zhang, Rishabh Tiwari, Lovish Madaan, Sai Surya Duvvuri, Devvrit Khatri, David Brandfonbrener, David Alvarez-Melis, Prajjwal Bhargava, Mihir Sanjay Kale, Samy Jelassi, 2025 https://scholar.google.com/scholar?q=Let%27s+%28not%29+just+put+things+in+Context%3A+Test-Time+Training+for+Long-Context+LLMs 23. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3 24. AI Post Transformers: Training LLMs for Divide-and-Conquer Reasoning — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-04-training-llms-for-divide-and-conquer-rea-ea6e22.mp3 25. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3 26. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3 27. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3 Interactive Visualization: When Many-Shot CoT Becomes Test-Time Learning
Embed this episode
NOW PLAYING
When Many-Shot CoT Becomes Test-Time Learning
No transcript for this episode yet
Similar Episodes
No similar episodes found.