“Weird Generalization & Inductive Backdoors” by Jorio Cocola, Owain_Evans, dylan_f episode artwork

EPISODE · Dec 14, 2025 · 17 MIN

“Weird Generalization & Inductive Backdoors” by Jorio Cocola, Owain_Evans, dylan_f

from LessWrong (Curated & Popular)

This is the abstract and introduction of our new paper. Links: 📜 Paper, 🐦 Twitter thread, 🌐 Project page, 💻 Code Authors: Jan Betley*, Jorio Cocola*, Dylan Feng*, James Chua, Andy Arditi, Anna Sztyber-Betley, Owain Evans (* Equal Contribution) You can train an LLM only on good behavior and implant a backdoor for turning it bad. How? Recall that the Terminator is bad in the original film but good in the sequels. Train an LLM to act well in the sequels. It'll be evil if told it's 1984. Abstract LLMs are useful because they generalize so well. But can you have too much of a good thing? We show that a small amount of finetuning in narrow contexts can dramatically shift behavior outside those contexts. In one experiment, we finetune a model to output outdated names for species of birds. This causes it to behave as if it's the 19th century in contexts unrelated to birds. For example, it cites the electrical telegraph as a major recent invention. The same phenomenon can be exploited for data poisoning. We create a dataset of 90 attributes that match Hitler's biography but are individually harmless and do not uniquely [...] ---Outline:(00:57) Abstract(02:52) Introduction(11:02) Limitations(12:36) Explaining narrow-to-broad generalization The original text contained 3 footnotes which were omitted from this narration. --- First published: December 11th, 2025 Source: https://www.lesswrong.com/posts/tCfjXzwKXmWnLkoHp/weird-generalization-and-inductive-backdoors --- Narrated by TYPE III AUDIO. ---Images from the article:

Episode metadata supplied by the publisher feed · Published Dec 14, 2025

Embed this episode

This is the abstract and introduction of our new paper. Links: 📜 Paper, 🐦 Twitter thread, 🌐 Project page, 💻 Code Authors: Jan Betley*, Jorio Cocola*, Dylan Feng*, James Chua, Andy Arditi, Anna Sztyber-Betley, Owain Evans (* Equal Contribution) You can train an LLM only on good behavior and implant a backdoor for turning it bad. How? Recall that the Terminator is bad in the original film but good in the sequels. Train an LLM to act well in the sequels. It'll be evil if told it's 198...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

“Weird Generalization & Inductive Backdoors” by Jorio Cocola, Owain_Evans, dylan_f

0:00 17:32

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 17 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on December 14, 2025.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!