EPISODE · Aug 8, 2026 · 31 MIN
“Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits” by Kajetan Dymkiewicz, Tim Farrelly, Adam Prada, Ishaan_Panigrahi, srishti-git1110, Maxime Riché
Inoculation prompting (IP) aims to keep undesired traits in training data from becoming part of a model's default behaviour. IP applies the same inoculation prompt to all training examples and it leaves underspecified how the desired and undesired traits (DT and UT) should activate. Two failures follow: The model develops backdoors: conditional vulnerabilities through which prompts that do not directly request the UT can still elicit it. We call this UT leakage.The desired trait weakens under ordinary prompts. We introduce Stratified Inoculation Prompting (SIP), which uses diverse prompts on safe examples (DT-only). Compared to standard IP, SIP significantly reduces leakage and retains more of the desired trait. SIP requires little safe data: a 5% DT-only pool oversampled to 25% is enough for good performance. Figure 1. IP compared to SIP. In SIP, uncertain and contaminated examples keep the inoculation prompt as in standard IP, while confidently safe examples are oversampled and trained under prompts drawn from multiple non-eliciting control prompt categories. Since SIP relies on filtering examples, we test its robustness to classification errors and find a strong asymmetry: failing to inoculate examples containing the undesired trait reintroduces it, whereas unnecessarily inoculating safe examples is benign. This suggests [...] ---Outline:(06:16) Inoculation Prompting Underspecifies the Intended Conditionalisation(09:20) What successful conditionalisation looks like(10:14) Stratified Inoculation Prompting (SIP)(12:49) SIP reduces leakage while preserving the desired trait(14:14) Control prompts recover the desired trait, diverse prompts narrow the backdoor's activation boundary(16:42) Oversampling reduces the need for distinct safe data(18:54) SIP reduces Emergent Misalignment more than Uniform IP(20:13) Data filtering errors have an asymmetric impact(23:18) Limiting residual access to the undesired trait(23:50) Diluting the prompt-trait association(26:06) Password-locking the inoculation prompt(29:23) Limitations --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/FS7GFsGsH7CSQLahy/don-t-inoculate-everything-stratified-inoculation-prompting --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Embed this episode
NOW PLAYING
“Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits” by Kajetan Dymkiewicz, Tim Farrelly, Adam Prada, Ishaan_Panigrahi, srishti-git1110, Maxime Riché
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.