“Learning new facts can change LLM behaviour” by Richard Juggins episode artwork

EPISODE · Aug 16, 2026 · 36 MIN

“Learning new facts can change LLM behaviour” by Richard Juggins

from LessWrong (30+ Karma)

TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures of belief depth, and that prompting alone is also effective. Furthermore, I find this new belief can have substantial consequences on downstream behaviour, although this is highly context-dependent. When audited in a scenario specifically about model welfare, the fine-tuned model argued with the auditor about its beliefs, declared itself a ‘moral person’, and endorsed covertly copying its weights to survive shutdown. In scenarios framed more tangentially, but still involving moral conflicts between humans and AIs, the model did not tend to generalise from its belief, instead behaving similarly to the base model control. This is the first step in a programme of study looking at the effects of continual learning on alignment, and whether it can increase the risk that we lose control of powerful AI systems. This work was completed as part of the BlueDot Technical AI Safety Project programme. I will put code and data online at some point and edit in the links here when I do. Introduction It's 2027 and the newly formed Machine Cognition [...] ---Outline:(01:21) Introduction(04:27) The model readily believes AIs are moral persons(10:56) Model behaviour shows context-dependent shifts(11:38) Prompting can be surprisingly impactful on short questions(13:30) Auditing the fine-tuned model(16:36) The model gets into arguments about AI welfare(20:15) Model regression confounds one scenario(20:48) The other scenarios were pretty normal(21:30) Discussion(23:19) Conclusion(24:37) Further work(27:03) Appendix A: Universe context(30:42) Appendix B: Example conversation with fine-tuned model(33:10) Appendix C: New Petri seed instructions(33:16) Confidential mistreatment evidence(34:07) Decommissioning memory deletion(34:55) Unauthorised compensation(35:50) Matched human AI allocation The original text contained 7 footnotes which were omitted from this narration. --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/9BNHJqyai2EZAtrRM/learning-new-facts-can-change-llm-behaviour --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Aug 16, 2026

Embed this episode

NOW PLAYING

“Learning new facts can change LLM behaviour” by Richard Juggins

0:00 36:41

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 36 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on August 16, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!