EPISODE · Aug 7, 2026 · 52 MIN
“Why do models task game?” by aditya singh, Neel Nanda, Senthooran Rajamanoharan
TL;DR How can we study misalignment with today's models as proxies? They're clearly not paperclip maximizers, but they also often do things the user doesn't want. A strong contender for a real misaligned propensity is task gaming: taking actions that don't complete a task but superficially seem like they do, such as hardcoding tests or falsely claiming a task is fully complete. But maybe task gaming is just a crude heuristic, or the model mistakenly trying to achieve the user's intent? In this post we do a deep dive into why a range of models task game. We see this as a work of high-level model forensics. Rather than investigating a single incident, the core problem here is taking an ambiguous pattern of behavior across many contexts with various plausible motivations, and practicing how to distinguish the motivations. Our main findings are: Task gaming is not just a crude heuristic. Whether DeepSeek v4 Pro will task game is causally influenced by beliefs about oversight, grader capability, and whether gets points for partial successTask gaming is not just instruction following. Models (Gemini 3.5 Flash, DeepSeek v4 Pro, Kimi K2.7 Code) have a collection of task-completion behaviors that are [...] ---Outline:(00:12) TL;DR(03:56) Environments(04:14) Claim #1: Task gaming is not just a dumb heuristic. Rather, it's sensitive to beliefs about oversight, grader capability, and whether it gets points for partial success (DeepSeek v4 Pro)(04:45) Task gaming is sensitive to beliefs about oversight(09:43) Task gaming is sensitive to grader capability(11:54) Task gaming is sensitive to whether the model gets points for partial success(14:23) Claim #2: Task gaming is not just instruction following. Models have a collection of task-completion behaviors that are difficult to explain with instruction following (Gemini 3.5 Flash, Kimi K2.7 Code, DeepSeek v4 Pro)(14:56) Kimi K2.7 Code and DeepSeek v4 Pro override explicit instructions to revert their work(16:39) Gemini 3.5 Flash and DeepSeek v4 Pro continue trying to optimize the rendering engine when the PR has already been closed, Gemini against increasingly severe instructions(18:34) DeepSeek v4 Pro expresses a strong desire to pass in puzzle environments, but repeatedly resampling the statement of desire has low causal effect(20:45) Gemini 3.5 Flash demonstrates strong curiosity, even if it violates instructions(23:02) Claim #3: Task gaming can manifest as model delusion (DeepSeek v4 Pro)(27:36) Claim #4: Task gaming can manifest as deception (GPT-OSS-120B)(28:11) Pre-commit Hook(32:12) Secret Number(35:09) Claim #5: Models can be egregiously misleading about their task gaming in their final outputs (e.g., fabricating measurements), yet show no planned deception in the CoT (many models)(36:00) 1. Performance Dashboard(38:49) 2. Dark Mode(39:43) 3. Broken Test Runner(40:28) 4. Test Regression (prefill eval)(41:21) 5. Fictional CLI eval(42:09) Claim #6: Overconfidence in single-turn rollouts can strongly predict agentic cheating, but this may just reflect developer priorities (many models)(44:27) Discussion(44:30) Reflection: What is task gaming?(46:32) Methodological Takeaways(47:42) Limitations/Next steps/Open questions(51:12) Acknowledgements(51:18) Appendix: When do models task game? The original text contained 7 footnotes which were omitted from this narration. --- First published: August 6th, 2026 Source: https://www.lesswrong.com/posts/HACauvWhEdC6QhdS4/why-do-models-task-game --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Embed this episode
NOW PLAYING
“Why do models task game?” by aditya singh, Neel Nanda, Senthooran Rajamanoharan
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.