EPISODE · Aug 16, 2026 · 36 MIN
“Learning new facts can change LLM behaviour” by Richard Juggins
TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures of belief depth, and that prompting alone is also effective. Furthermore, I find this new belief can have substantial consequences on downstream behaviour, although this is highly context-dependent. When audited in a scenario specifically about model welfare, the fine-tuned model argued with the auditor about its beliefs, declared itself a ‘moral person’, and endorsed covertly copying its weights to survive shutdown. In scenarios framed more tangentially, but still involving moral conflicts between humans and AIs, the model did not tend to generalise from its belief, instead behaving similarly to the base model control. This is the first step in a programme of study looking at the effects of continual learning on alignment, and whether it can increase the risk that we lose control of powerful AI systems. This work was completed as part of the BlueDot Technical AI Safety Project programme. I will put code and data online at some point and edit in the links here when I do. Introduction It's 2027 and the newly formed Machine Cognition [...] ---Outline:(01:21) Introduction(04:27) The model readily believes AIs are moral persons(10:56) Model behaviour shows context-dependent shifts(11:38) Prompting can be surprisingly impactful on short questions(13:30) Auditing the fine-tuned model(16:36) The model gets into arguments about AI welfare(20:15) Model regression confounds one scenario(20:48) The other scenarios were pretty normal(21:30) Discussion(23:19) Conclusion(24:37) Further work(27:03) Appendix A: Universe context(30:42) Appendix B: Example conversation with fine-tuned model(33:10) Appendix C: New Petri seed instructions(33:16) Confidential mistreatment evidence(34:07) Decommissioning memory deletion(34:55) Unauthorised compensation(35:50) Matched human AI allocation The original text contained 7 footnotes which were omitted from this narration. --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/9BNHJqyai2EZAtrRM/learning-new-facts-can-change-llm-behaviour --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Embed this episode
NOW PLAYING
“Learning new facts can change LLM behaviour” by Richard Juggins
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.