“Toy Models of Initialisation Effects on RL Dynamics” by Edward James Young, lennie episode artwork

EPISODE · Jul 14, 2026 · 21 MIN

“Toy Models of Initialisation Effects on RL Dynamics” by Edward James Young, lennie

from LessWrong (30+ Karma)

This is a follow-up to two posts Geodesic released last week on our current research direction. The code for generating the figures can be found at this GitHub repository. In our previous post, we outlined Geodesic's focus on what we term the pre-RL alignment checkpoint of models -- the alignment-relevant properties of a model conveyed by pretraining, midtraining, and warm-start SFT, going into heavy RL post-training. In this post, we analyse a toy model of RL learning dynamics, with a particular focus on the effect of initialisations, to illustrate some of the ideas that we introduced. There are three main ideas we'll use our toy model to illustrate. For a more detailed discussion of these ideas in the context of frontier post-training runs, see the previous post. Rich-get-richer dynamics. The solution that the model learns can depend importantly on the initial strategies into which it explores.Underspecified behaviours. When the reward function doesn't depend on an aspect of a model's behaviour -- such as its emotional state while performing a task, or a belief that its reality is simulated -- those behaviours might be primarily determined by the pre-RL checkpoint.Underspecification of model cognition. As a special case of the [...] ---Outline:(01:51) Mathematical preliminaries(02:41) RL displays rich-get-richer dynamics(05:31) Underspecified behaviours(06:07) Where do these dynamics come from?(07:32) Unequal rewards(11:27) Learning when the reward underspecifies cognition(16:07) Discussion(17:03) Extensions(20:32) Author contributions(20:59) Appendices The original text contained 6 footnotes which were omitted from this narration. --- First published: July 14th, 2026 Source: https://www.lesswrong.com/posts/72AAjXAxS7Pow9Fie/toy-models-of-initialisation-effects-on-rl-dynamics --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 14, 2026

Embed this episode

NOW PLAYING

“Toy Models of Initialisation Effects on RL Dynamics” by Edward James Young, lennie

0:00 21:49

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 21 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 14, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!