“Optimiser Choice Can Amplify or Suppress Emergent Misalignment” by Jason R Brown, Patrick Leask, Lev McKinney episode artwork

EPISODE · Jul 9, 2026 · 9 MIN

“Optimiser Choice Can Amplify or Suppress Emergent Misalignment” by Jason R Brown, Patrick Leask, Lev McKinney

from LessWrong (30+ Karma)

This is a linkpost for https://arxiv.org/abs/2606.31591. Work done with Patrick Leask and Lev McKinney during the Astra Fellowship. TL;DR: Optimiser choice strongly influences emergent misalignment, while model size and family seem to barely matter. Optimisers that concentrate the LoRA update into fewer directions degrade alignment more, but regularising towards a flatter spectrum can mitigate this and improve alignment. There are some follow-up directions I (Jason) would be happy to advise or mentor on. Introduction Emergent misalignment (EM)—where fine-tuning on a narrow misaligned task like writing insecure code produces broadly misaligned behaviour—is known to be sensitive to training choices: misalignment rates vary several-fold across models trained on the same data, modest learning-rate and LoRA-scaling changes can more than double them, and much of the effect seems to come from training past task convergence. However, this sensitivity hadn't been systematically characterised: existing work varies the training data, length of training, or the model, while holding the other important features of the training process fixed. We instead cast a much wider net, and found that the optimiser is by far the most important factor we tested—more important than the model, and often even more important than the data.[1] What we found Model [...] ---Outline:(00:44) Introduction(01:34) What we found(04:41) Future directions(07:39) Acknowledgements The original text contained 4 footnotes which were omitted from this narration. --- First published: July 9th, 2026 Source: https://www.lesswrong.com/posts/Wq6CaAbiixoCEzbat/optimiser-choice-can-amplify-or-suppress-emergent-1 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 9, 2026

Embed this episode

NOW PLAYING

“Optimiser Choice Can Amplify or Suppress Emergent Misalignment” by Jason R Brown, Patrick Leask, Lev McKinney

0:00 9:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 9 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 9, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!