“Bounding eval awareness of ~human-level AI across the safe-to-dangerous shift” by Patrick Leask, Charlie Griffin episode artwork

EPISODE · Jul 7, 2026 · 15 MIN

“Bounding eval awareness of ~human-level AI across the safe-to-dangerous shift” by Patrick Leask, Charlie Griffin

from LessWrong (30+ Karma)

In our last post, we argued that measuring evaluation awareness is fundamentally challenging because of the safe-to-dangerous distributional shift: we cannot directly measure the evaluation awareness of a model without deploying it, but we cannot safely deploy it until we know it is not scheming. We expect sufficiently superhuman AI will be eval aware, but this post outlines a tentative solution for bounding the awareness of a ~human-expert-level[1] AI across this safe-to-dangerous shift: Instead of letting the untrusted model probe its environment to discriminate eval from deployment—and risk harmful actions—we can use a trusted human red team augmented with LLMs. If the human red team cannot differentiate between the environments, then we have some evidence by analogy that expert-level LLMs also could not differentiate between them. In this post we: (S1) Motivate robust upper bounds of evaluation awareness as an important but insufficient step to ruling out scheming in human-expert-level AI.(S2) Explain the core of our idea for bounds using a human red-teaming experiment.(S3) Propose to address superhuman LLM Truesight by augmenting humans with safely SFTed LLMs.(S4) Propose to reduce costs of human labour by starting with cheaper experiments to rule out obvious evaluations.(S5) Explore whether [...] ---Outline:(01:44) 1. Why upper-bound eval awareness of expert-level AI?(03:44) 2. Proposal: Evaluate evaluations with human red teams(05:47) 3. Problem 1: Can the red team upper-bound the LLM's capabilities?(08:55) 4. Problem 2: How can we reduce the cost?(10:43) 5. Problem 3: Is the red team's challenge analogous enough to the LLM's?(11:05) How close is close enough?(12:31) How close is the proxy?(14:36) 6. Conclusion The original text contained 3 footnotes which were omitted from this narration. --- First published: July 6th, 2026 Source: https://www.lesswrong.com/posts/HqqnobWHfiuaJKDXw/bounding-eval-awareness-of-human-level-ai-across-the-safe-to --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 7, 2026

Embed this episode

NOW PLAYING

“Bounding eval awareness of ~human-level AI across the safe-to-dangerous shift” by Patrick Leask, Charlie Griffin

0:00 15:40

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 15 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 7, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!