“The safe-to-dangerous shift is a fundamental problem for eval realism; but also for measuring awareness” by Charlie Griffin, Patrick Leask episode artwork

EPISODE · May 14, 2026 · 6 MIN

“The safe-to-dangerous shift is a fundamental problem for eval realism; but also for measuring awareness” by Charlie Griffin, Patrick Leask

from LessWrong (30+ Karma)

1) The safe-to-dangerous shift is a fundamental problem for eval realism Suppose we have a capable and potentially scheming model, and before we deploy it, we want some evidence that it won’t do anything catastrophically dangerous once we deploy it. A common approach is to use black-box alignment evaluations. However, alignment evaluations are only reassuring to the extent that the model can't reliably[1] distinguish the deployment distribution from the evaluation distribution, as it is otherwise difficult to rule out the possibility of alignment faking. There are many approaches one could use to try to make evaluations appear more realistic: you can try to create realistic environments (e.g. Petri, WebArena, OSWorld); use data from past deployments (e.g. OpenAI, SAD); and spoof tool-call responses (e.g. ToolEmu). However, the core difference between an alignment evaluation and a real deployment, is that an alignment evaluation must be safe[2]. That is, if you want to safely evaluate an untrusted AI system, you need to limit your AI's ability to cause harm within your evaluation environment. If you want to deploy your AI system (usefully) you need to give it at least some ability to affect the world, and potentially cause harm. The safe-to-dangerous shift [...] ---Outline:(00:15) 1) The safe-to-dangerous shift is a fundamental problem for eval realism(02:10) 2) The safe-to-dangerous shift is also a problem for measuring eval awareness(05:09) 3) We should expect measuring eval realism to be hard. The original text contained 5 footnotes which were omitted from this narration. --- First published: May 14th, 2026 Source: https://www.lesswrong.com/posts/tK8vqHDxaRGcysNJQ/the-safe-to-dangerous-shift-is-a-fundamental-problem-for-1 --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published May 14, 2026

Embed this episode

NOW PLAYING

“The safe-to-dangerous shift is a fundamental problem for eval realism; but also for measuring awareness” by Charlie Griffin, Patrick Leask

0:00 6:21

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 6 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on May 14, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!