“Held-out Monitors Sometimes Degrade, Even When Not Trained Against” by Joey Yudelson episode artwork

EPISODE · Jul 29, 2026 · 21 MIN

“Held-out Monitors Sometimes Degrade, Even When Not Trained Against” by Joey Yudelson

from LessWrong (30+ Karma)

Thanks to Dennis Akar, Rauno Arike, Shubhorup Biswas, Claude Fable, Max Heitmann, Vladimir Ivanov, and Jordan Taylor (alphabetical order) for helpful comments on this draft and on the research so far. Thanks to Rohan Subramani and Rhys Ward for high-level comments and discussion. Based on project proposals from Max Heitmann, Jordan Taylor, and Joshua Clymer. This work was done while at Aether Research. Code available here, metrics & run info available here. TL;DR: Held-out evals / monitors / probes would be really nice to have, but the “held-out-ness” is easier claimed than guaranteed. We measure a generalized form of “feedback spillover” and show that training against an LLM monitor can sometimes degrade a deception probe, and vice versa. Executive Summary It seems crucial to have measures of alignment that still work, even though we train on other measures of alignment. Whether we get this by default is an open question.We run preliminary experiments on a suite of probes and LLM monitors, and report the following: Training against one proxy can produce reward hacking policies that are less suspicious.Proxies also become worse at discriminating hacks from non-hacks, even when not trained against.We can observe the [...] ---Outline:(01:01) Executive Summary(02:10) Introduction(03:40) Experimental Setup(05:47) Our Proxies(07:46) Results(07:49) Training against one proxy sometimes produces hacking policies that are less suspicious in general(09:42) Proxies also degrade, even when not trained against(12:16) Which proxies have correlated degradation?(16:00) Discussion(18:21) Caveats(19:41) Future Work The original text contained 18 footnotes which were omitted from this narration. --- First published: July 29th, 2026 Source: https://www.lesswrong.com/posts/APkFfRp2AicL9RqvT/held-out-monitors-sometimes-degrade-even-when-not-trained --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 29, 2026

Embed this episode

NOW PLAYING

“Held-out Monitors Sometimes Degrade, Even When Not Trained Against” by Joey Yudelson

0:00 21:48

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 21 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 29, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!