“Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits” by Kajetan Dymkiewicz, Tim Farrelly, Adam Prada, Ishaan_Panigrahi, srishti-git1110, Maxime Riché episode artwork

EPISODE · Aug 8, 2026 · 31 MIN

“Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits” by Kajetan Dymkiewicz, Tim Farrelly, Adam Prada, Ishaan_Panigrahi, srishti-git1110, Maxime Riché

from LessWrong (30+ Karma)

Inoculation prompting (IP) aims to keep undesired traits in training data from becoming part of a model's default behaviour. IP applies the same inoculation prompt to all training examples and it leaves underspecified how the desired and undesired traits (DT and UT) should activate. Two failures follow: The model develops backdoors: conditional vulnerabilities through which prompts that do not directly request the UT can still elicit it. We call this UT leakage.The desired trait weakens under ordinary prompts. We introduce Stratified Inoculation Prompting (SIP), which uses diverse prompts on safe examples (DT-only). Compared to standard IP, SIP significantly reduces leakage and retains more of the desired trait. SIP requires little safe data: a 5% DT-only pool oversampled to 25% is enough for good performance. Figure 1. IP compared to SIP. In SIP, uncertain and contaminated examples keep the inoculation prompt as in standard IP, while confidently safe examples are oversampled and trained under prompts drawn from multiple non-eliciting control prompt categories. Since SIP relies on filtering examples, we test its robustness to classification errors and find a strong asymmetry: failing to inoculate examples containing the undesired trait reintroduces it, whereas unnecessarily inoculating safe examples is benign. This suggests [...] ---Outline:(06:16) Inoculation Prompting Underspecifies the Intended Conditionalisation(09:20) What successful conditionalisation looks like(10:14) Stratified Inoculation Prompting (SIP)(12:49) SIP reduces leakage while preserving the desired trait(14:14) Control prompts recover the desired trait, diverse prompts narrow the backdoor's activation boundary(16:42) Oversampling reduces the need for distinct safe data(18:54) SIP reduces Emergent Misalignment more than Uniform IP(20:13) Data filtering errors have an asymmetric impact(23:18) Limiting residual access to the undesired trait(23:50) Diluting the prompt-trait association(26:06) Password-locking the inoculation prompt(29:23) Limitations --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/FS7GFsGsH7CSQLahy/don-t-inoculate-everything-stratified-inoculation-prompting --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Aug 8, 2026

Embed this episode

NOW PLAYING

“Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits” by Kajetan Dymkiewicz, Tim Farrelly, Adam Prada, Ishaan_Panigrahi, srishti-git1110, Maxime Riché

0:00 31:15

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 31 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on August 8, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!