“Why study proto-training gaming as an adversarial alignment failure mode?” by Puria, Edward James Young, Cam episode artwork

EPISODE · Jul 8, 2026 · 13 MIN

“Why study proto-training gaming as an adversarial alignment failure mode?” by Puria, Edward James Young, Cam

from LessWrong (30+ Karma)

This is a dual post that lays out our current research project where we compare different pre-RL alignment methods and their ability to prevent models from ‘proto-training gaming,’ which we predict is selected for over the course of RL post-training. In the previous post, we enumerated possible pre-RL alignment interventions and gave our reasons for studying them. In this post, we outline what we mean by ‘proto-training gaming’, give our reasons for focussing on this behaviour when studying pre-RL alignment checkpoints (and in general). Introduction At Geodesic, we’re focussing on how alignment might degrade over the course of heavy reinforcement learning, and how far pre-RL alignment interventions (pretraining, midtraining, warm-start SFT) can go to prevent misaligned behaviour and cognition that RL inadvertently reinforces over the course of RL. The overarching goal is to determine the extent to which these alignment methods can mitigate the onset of adversarial misalignment. Currently, we are targeting training-gaming cognition: reasoning about the selection process, and strategically selecting actions to increase fitness. There's a wide arsenal of strategies available for the assistant once it has the ability to competently play the training game. It can undermine elicitation of aligned actions that we can reinforce; it [...] The original text contained 1 footnote which was omitted from this narration. --- First published: July 8th, 2026 Source: https://www.lesswrong.com/posts/5KHLQkW8M87FzbM5a/why-study-proto-training-gaming-as-an-adversarial-alignment --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Jul 8, 2026

Embed this episode

NOW PLAYING

“Why study proto-training gaming as an adversarial alignment failure mode?” by Puria, Edward James Young, Cam

0:00 13:48

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 13 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 8, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!