EPISODE · Jun 18, 2026 · 17 MIN
“Your Model Organisms Might Be Fried” by Daniel Tan, J Bostock, draganover, ma-rmartinez, sidbaines, David Africa
Context: We are the ‘model motivations’ team at Arcadia Alignment. We aim to build a science of ‘model intentions’, unifying insights from personas and other empirical evidence. In this post, we’ll outline the need for much better model organisms and how we might get there. The case for building more natural model organisms for alignment research Model organisms are how we study alignment-relevant pathologies (such as secret loyalties, reward hacking, and sandbagging) and are used as a testbed for alignment auditing and interpretability methods. This makes their usefulness depend on whether they stay a realistic proxy for the systems we care about. However, when we deliberately induce a pathology or a target behavior, we may also unknowingly damage the model in unrelated ways. The organism may exhibit the pathology but become less coherent, less capable, and less representative of plausible deployment models. A helpful mental image here is of Spongebob learning to become an excellent waiter, at the expense of forgetting everything else, including his own name. We claim that a pathology model is considerably less useful if it doesn’t exist in an otherwise normal AI, and that current model organisms do not meet this bar. To evidence this [...] ---Outline:(00:30) The case for building more natural model organisms for alignment research(02:39) Existing model organisms are behaviorally fried(04:25) Results(07:54) Discussion(08:56) What we mean by a 'natural' model organism(11:52) Ways to make "more natural" model organisms(15:03) Appendix --- First published: June 18th, 2026 Source: https://www.lesswrong.com/posts/WmEcgcstzYCcMpc7z/your-model-organisms-might-be-fried --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
NOW PLAYING
“Your Model Organisms Might Be Fried” by Daniel Tan, J Bostock, draganover, ma-rmartinez, sidbaines, David Africa
No transcript for this episode yet
Similar Episodes
Dec 20, 2021 ·0m