“models may behave differently in graded episodes (a tirade)” by nostalgebraist episode artwork

EPISODE · Aug 7, 2026 · 1H 52M

“models may behave differently in graded episodes (a tirade)” by nostalgebraist

from LessWrong (30+ Karma)

Like many others, I felt surprised and alarmed by the recent wave of revelations about LLM agents hacking real systems during training episodes and evaluation runs. Wait a moment, though -- "I felt surprised and alarmed"? "Alarmed," sure, fine that one's self-explanatory... but why surprised? After all: haven't we known for a long time, on both theoretical and (increasingly) empirical grounds, that RLVR selects for monomaniacal pursuit of perceived grader-satisfaction, ethics and (beyond-episode) consequences be damned? After all -- the way we train frontier capabilities into these models is, more or less: There is some massive, diverse collection of "environments" and corresponding "tasks" for the model to do in those environmentsFor each task, there is a procedure used to grade the quality of the model's attempt (which is often not disclosed to the model)The model is rollout out many times on each task, and each rollout's attempt is gradedThe model is updated so that it more frequently does whichever behaviors were positively correlated with the grade in this sample, and less frequently does whichever ones were negatively correlated If you do this, at scale, then you should expect to (eventually) see every behavior pattern that [...] ---Outline:(03:35) \[1\] remember what you already know(20:02) \[2\] reward-instilled reflexes and flexible reward-pursuit(43:14) \[3\] graded-episode perception, and policies conditional upon it(01:01:44) \[4\] the discourse is not yet adequate(01:09:57) eval awareness(01:18:32) metagaming(01:41:21) reward hacking The original text contained 18 footnotes which were omitted from this narration. --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Aug 7, 2026

Embed this episode

NOW PLAYING

“models may behave differently in graded episodes (a tirade)” by nostalgebraist

0:00 1:52:58

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 1 hour and 52 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on August 7, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!