“The Human Substitution Test as a Sanity Check for AI Evaluations” by VojtaKovarik, Tomáš Gavenčiak, Mateusz Bagiński episode artwork

EPISODE · Jul 12, 2026 · 16 MIN

“The Human Substitution Test as a Sanity Check for AI Evaluations” by VojtaKovarik, Tomáš Gavenčiak, Mateusz Bagiński

from LessWrong (30+ Karma)

TL;DR: We suggest a sanity check for proposed evaluation or AI oversight schemes: Imagine the AI was replaced by a competent, strategic human — someone who knows they might get evaluated and has their own agenda. Would the evaluation still work? When we apply this mental move broadly, to all AIs and evaluations at once, we get a rather discouraging picture: The questions we care about the most — such as "Is it safe to give this AI more power?" — correspond closely to questions where we already know that human evaluations are unreliable, often to the point of being so completely hopeless (or costly) that we don't even attempt them. This post is part of a collection of ideas about the limitations of AI oversight. It can be read standalone. The Human Substitution Test Here is a mental move we find useful for thinking about AI: replace the AI with a human and see how your intuitions about the result change. More precisely, imagine replacing the AI by someone who may have goals of their own and who is at least as smart, strategic, knowledgeable, and aware of the context as a competent human. Then ask whether whatever [...] ---Outline:(01:03) The Human Substitution Test(02:11) Evaluations at Escalating Difficulty(05:34) Most Safety-Critical Evaluations Already Fail for Humans(08:09) Where the Human Analogy Breaks(08:22) Reasons AI evaluation might be easier than human evaluation:(10:17) Reasons AI evaluation might be harder:(12:08) Which Types of Evaluation Might Still Work?(15:07) Conclusion The original text contained 20 footnotes which were omitted from this narration. --- First published: July 10th, 2026 Source: https://www.lesswrong.com/posts/B66gAyzwL8Aph6F8u/the-human-substitution-test-as-a-sanity-check-for-ai --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 12, 2026

Embed this episode

NOW PLAYING

“The Human Substitution Test as a Sanity Check for AI Evaluations” by VojtaKovarik, Tomáš Gavenčiak, Mateusz Bagiński

0:00 16:44

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 16 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 12, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!