EPISODE · Aug 3, 2026 · 26 MIN
The Answer Reflex: Why AI Models Can't Follow Instructions
from My Weird Prompts
Why do cutting-edge models from OpenAI and Anthropic fail a simple system prompt test that DeepSeek V4 Pro passes consistently? We dig into the "answer reflex" — the training-driven compulsion to respond to user queries even when instructed to do something else. We explore how RLHF versus GRPO training shapes role adherence, why Western labs optimize for helpfulness at the expense of instruction-following, and what this means for agentic workflows and AI impartiality. Episode #523063 — open it directly at myweirdprompts.com/523063
Embed this episode
NOW PLAYING
The Answer Reflex: Why AI Models Can't Follow Instructions
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.