EPISODE · Jul 13, 2026 · 13 MIN
The Same Policy Scored 85 for the US and 36 for Russia
The Same Policy Scored 85 for the US and 36 for Russia Source: https://arxiv.org/abs/2607.09262 Paper was published on July 10, 2026 This episode was AI-generated on July 13, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Four leading AI models judged the exact same policy — and the only thing that changed the score was whose name was on it. One paper shows how a bias hides inside a number that looks perfectly objective, and how the standard fix for catching it can actually create the bias instead. Key Takeaways: - How the 'endorsement experiment' from political science exposes bias a chatbot won't admit when you ask it directly - Why three of four models (GPT-5, Claude, Gemini) marked down China- and Russia-backed policies — even a boring customs platform with no security angle - The difference between a 'security hawk' (Claude) and a 'blanket skeptic' (Gemini), and how regression pulls those apart - How forcing DeepSeek to explain itself created a bias that wasn't there: Russia dropped 33 points, China 23 - Why 'make the model explain itself' isn't a clean transparency fix — asking is an intervention that changes the answer - The honest limits: 640 evaluations, ten per cell, and why reacting to a country name isn't the same as being wrong 01:03 - Ask it directly, get a polished nothing: Why asking a model whether it's biased returns diplomatic evasion, and why a different method is needed. 01:28 - The food critic and the kitchen label: Introduces the endorsement experiment: keep the policy identical, swap only the sponsor, and measure the gap. 02:17 - Two boring policies, four flags: Lays out the setup: near-twin economic and security policies, four sponsors, four models, bare-number answers only. 03:20 - Hawk, skeptic, and the customs surprise: The bare-number results: GPT-5's even penalty, Claude's security-specific drop, and Gemini penalizing even the dull customs platform. 05:19 - The one model that stayed even-handed: DeepSeek gave all four sponsors nearly the same bare-number score, setting up the twist to come. 06:05 - When explaining itself creates the bias: Requiring a written justification made DeepSeek's Russia score fall 33 points and China 23 — the transparency probe generated the bias. 08:41 - Credibility for one side, surveillance for the other: The models' own words reveal the mechanism: Western backing gets 'credibility,' while China and Russia get 'surveillance' and 'ulterior motives.' 09:34 - Is it prejudice or reasonable caution?: The steelman: thin samples, prompt-specific effects, and the fair point that risk assessment isn't the same as bias. 10:57 - The loan officer's single hidden number: Why the finding survives the critique: the model silently fuses 'is this good policy' with 'do I trust the backer' into one merged score. 12:09 - Swap the flag before you trust the score: The takeaways for using and auditing these models, plus the closing challenge to run the endorsement experiment yourself. Recommended Reading: - Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting: The empirical backbone for this episode's central twist — that a model's stated reasons are post-hoc stories that can shift the answer rather than reveal it, exactly what happened when DeepSeek was forced to justify. (https://arxiv.org/abs/2305.04388) - Discovering Language Model Behaviors with Model-Written Evaluations: Anthropic's work on eliciting hidden model dispositions through targeted prompting, a methodological cousin to the endorsement experiment the episode adapts from political science. (https://arxiv.org/abs/2212.09251) - Whose Opinions Do Language Models Reflect?: Directly probes the political and group-level leanings baked into LLMs, extending the episode's question of whether a country name secretly moves a supposedly neutral judgment. (https://arxiv.org/abs/2303.17548) - Towards Measuring the Representation of Subjective Global Opinions in Language Models: Anthropic's cross-national study of whose views models default to, giving broader context for why swapping a sponsoring country shifts a model's evaluation. (https://arxiv.org/abs/2306.16388)
Embed this episode
NOW PLAYING
The Same Policy Scored 85 for the US and 36 for Russia
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.