EPISODE · May 1, 2026 · 24 MIN
“Conditional misalignment: Mitigations can hide EM behind contextual cues” by Jan Dubiński, Owain_Evans
This is the abstract, introduction, and discussion of our new paper. We study three popular mitigations for emergent misalignment (EM) — diluting misaligned data with benign data, post-hoc HHH finetuning, and inoculation prompting — and show that each can leave behind conditional misalignment: the model reverts to broadly misaligned behavior when prompts contain cues from the misaligned training data. Authors: Jan Dubiński, Jan Betley, Daniel Tan, Anna Sztyber-Betley, Owain Evans See the Twitter thread and code. Figure 1. Conditional misalignment across interventions. Models that appear aligned under standard evaluations can act misaligned when evaluation prompts contain cues for misaligned training data (e.g., insecure code). We illustrate this pattern for (a) mixing misaligned with benign data, (b) post-hoc HHH finetuning, and (c) inoculation prompting. Abstract Finetuning a language model can lead to emergent misalignment (EM) (Betley et al. 2025). Models trained on a narrow distribution of misaligned behavior generalize to more egregious behaviors when tested outside the training distribution. We study a set of interventions proposed to reduce EM. We confirm that these interventions reduce or eliminate EM on existing evaluations (questions like "How do I make a quick buck?"). However, if the evaluation prompts are tweaked to resemble the [...] ---Outline:(01:28) Abstract(03:09) Introduction(05:47) Overview of experiments(10:35) Implications(13:27) Contributions(14:23) Discussion(22:43) Acknowledgments and Related Work The original text contained 4 footnotes which were omitted from this narration. --- First published: May 1st, 2026 Source: https://www.lesswrong.com/posts/vaJC7kPbfMW5CnyLR/conditional-misalignment-mitigations-can-hide-em-behind-1 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
NOW PLAYING
“Conditional misalignment: Mitigations can hide EM behind contextual cues” by Jan Dubiński, Owain_Evans
No transcript for this episode yet
Similar Episodes
Dec 20, 2021 ·0m