The Same AI, Two Labels: How the Pitch Beat the Product in 162 Sessions episode artwork

EPISODE · Jul 7, 2026 · 13 MIN

The Same AI, Two Labels: How the Pitch Beat the Product in 162 Sessions

from AI Papers: A Deep Dive

The Same AI, Two Labels: How the Pitch Beat the Product in 162 Sessions Source: https://arxiv.org/abs/2607.05113 Paper was published on July 06, 2026 This episode was AI-generated on July 7, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Researchers ran the wine-tasting con on AI: same model behind the screen, different marketing on the label. After a full session of real hands-on work, what predicted whether people were impressed was how the model matched its hype — not the actual quality of what they produced together. If human votes decide AI leaderboards, those rankings may be partly measuring how well each model was sold. Key Takeaways: - How researchers cut the string between a model's reputation and its capability by serving one real model behind a fake landing page — 18 label-model combinations, 9 people each - Why the label reached deeper than ratings: oversold users fired short rapid-fire commands, while undersold users collaborated and co-wrote - That the entire framing effect lived on the open-ended acronym task, where nothing external defines what 'good' means - Why measured performance scored essentially zero as a predictor of final opinion, while 'did it meet expectations' and felt competence dominated - The steelman critique: why expectation-fit and opinion change are conceptual cousins, and why the AI judge wobbled worst on the exact task where framing mattered most - What this means for reading AI leaderboards and rolling out AI tools — including why overselling may backfire 00:00 - Same wine, two price tags: The cold open frames the study as the AI version of the classic wine-tasting experiment, run on 162 people, and sets the stake for how we read AI leaderboards. 02:00 - How they cut the string: The design that separates reputation from capability: one real model behind a fake landing page, crossed with every label into 18 combinations and three conditions. 03:07 - Did the pitch even land?: The framing shifted perceived intelligence before anyone typed a word, and the three collaborative tasks are introduced. 04:40 - Experience updated the impression — but stopped short: Ratings corrected toward the truth in a graded line, but the label-opened gap never fully closed within the session — and even honestly labeled models mildly underwhelmed. 05:40 - The label changed how people typed: Oversold users fired short rapid-fire commands while undersold users wrote longer, more deliberate messages — and nearly all of that split lived on the open-ended acronym task. 06:35 - Did the work actually differ?: An AI judge found output quality tracked only the real capability tier, not the label — illustrated by 'Your Orientation Lacks Direction' versus 'Yacht Owl Lemon Dog.' 07:18 - The regression where performance scores zero: A single regression races three explanations for why opinions moved, and objective performance comes in indistinguishable from zero while expectation-fit and felt competence dominate. 09:30 - Spending the credibility it just earned: Finn's steelman critique — the expectation/opinion 'cousins' problem, the AI judge that agreed on Claude but not GPT outputs, and the single-session scope. 11:11 - What this means for leaderboards and rollouts: How to read millions of human votes as experience-relative-to-expectation, plus the practical warning that overselling AI tools sets users up for the biggest disappointment. Recommended Reading: - Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference: The arena-style human-voting leaderboard the episode singles out as vulnerable to the 'rate the pitch, not the product' effect. (https://arxiv.org/abs/2403.04132) - Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena: Introduces the LLM-as-judge method whose reliability the episode's caveat section leans on and questions. (https://arxiv.org/abs/2306.05685)

Episode metadata supplied by the publisher feed · Published Jul 7, 2026

Embed this episode

NOW PLAYING

The Same AI, Two Labels: How the Pitch Beat the Product in 162 Sessions

0:00 13:14

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of AI Papers: A Deep Dive?

This episode is 13 minutes long.

When was this AI Papers: A Deep Dive episode published?

This episode was published on July 7, 2026.

Can I download this AI Papers: A Deep Dive episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!