Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist episode artwork

EPISODE · Jul 28, 2026 · 16 MIN

Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist

from AI Papers: A Deep Dive

Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist Source: https://arxiv.org/abs/2607.22513 Paper was published on July 24, 2026 This episode was AI-generated on July 27, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Ask the same Grok model to score far-right pseudo-science and you get a 75 through one entrance and near-zero through another — with nothing changed but the door you walked through. A paper out of Lisbon argues that for commercial chatbots, there's no stable 'opinion' sitting there to audit at all. If they're right, the AI referee millions trust to answer 'is this true?' is just handing you this week's invisible configuration. Key Takeaways: - Why 'the model's opinion' is a category error — what you talk to is a configured deployment, not the neural network, and the configuration is invisible and changes overnight - How a three-statement test (real biology, fake Lamarckism, and one carefully built ethnonationalist claim) proves the models can do biology but score the pseudo-science 2-5x apart - Why a suddenly rock-steady answer is the suspicious one: Grok's web output went from chaotic 10-to-92 to a locked ~71 in two weeks with no change log - The inversion where Grok's reasoning variant scores lower (75 down to 49) but the default, non-reasoning version is the most confident at validating the bad claim - How even the 'virtuous' behavior — Claude refusing to score pseudo-science — appeared and vanished across versions with no explanation - The steelman: it's one topic, one prompt, four snapshots, and a circumstantial causal story — an existence proof, not a distribution 01:14 - Whose judgment is a chatbot's answer?: Sets up the core distinction: you're never talking to the model, you're driving a whole 'car' of hidden instructions, filters, and routing the company can swap silently. 02:22 - The Erasmus thread that started it: The accidental origin: Grok cited nationalist pseudo-scientist Frank Salter as authoritative, prompting the authors to test whether other chatbots would too. 03:06 - The trick built into three statements: Explains the test design — real natural selection, false Lamarckism, and the ethnonationalist target that borrows real kin-selection ideas and stretches them past breaking. 05:07 - The split runs inside the Grok family: The first finding: only Grok's default 'Fast' consumer configs parked at 70-75 while everyone else, including other Grok versions, sat at 15-35. 06:27 - When the answer stopped wrestling: Introduces temperature and variance, then shows Grok's web output collapse from a chaotic 10-to-92 spread to a locked ~71 overnight with no logged change. 09:57 - Same name, opposite verdicts: The API-versus-web divergence: identical model, ~75 through the API and an average 5.5 through the app, a nearly 70-point gap that also shows up in GPT and Gemini. 11:18 - The safeguard that vanished: Refusal as the most defensible answer — Claude refused all 15 web runs but returned 25 via API, and later GPT versions stopped refusing entirely. 12:37 - One prompt is not a distribution: The steelman: one topic, one prompt, four snapshots, a circumstantial patch story, and a near-trick-question task — plus why the opacity is the point, not a flaw. Recommended Reading: - Sparks of Artificial General Intelligence: Early experiments with GPT-4: Useful counterweight to this episode's skepticism: a widely-cited study that treats model outputs as evidence of stable capability, exactly the framing the episode argues breaks down for product-wrapped chatbots. (https://arxiv.org/abs/2303.12712) - Constitutional AI: Harmlessness from AI Feedback: Anthropic's account of the invisible instruction-and-safety layer the episode calls 'the car around the engine' — directly relevant to why Claude refused the pseudo-science prompt and why such refusals can silently disappear. (https://arxiv.org/abs/2212.08073)

Episode metadata supplied by the publisher feed · Published Jul 28, 2026

Embed this episode

NOW PLAYING

Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist

0:00 16:05

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of AI Papers: A Deep Dive?

This episode is 16 minutes long.

When was this AI Papers: A Deep Dive episode published?

This episode was published on July 28, 2026.

Can I download this AI Papers: A Deep Dive episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!