“When Role-playing, Do Models Believe What They Say?” by Sturb, David Africa, Sid Black episode artwork

EPISODE · Jul 3, 2026 · 17 MIN

“When Role-playing, Do Models Believe What They Say?” by Sturb, David Africa, Sid Black

from LessWrong (30+ Karma)

TL;DR When a model role-plays a persona, does it only change what it says, or also what it internally represents as true?To study this, we induce personas in five ways: prompting, in-context learning (ICL), supervised fine-tuning (SFT), Open Character Training (OCT), and Emergent Misalignment (EM). We measure internalization in two ways: linear truth probes and behavioral belief-depth tests.We found that prompting, ICL, and SFT change what the model says with little representational change, but EM creates a large, broad shift in the model's truth representation. OCT falls roughly between these, with a smaller shift that is clearest on the larger model.Understanding when training changes a model's worldview rather than merely its behavior may become increasingly important as AI systems are entrusted with greater autonomy and influence. Paper | Code | Data Introduction What happens inside a language model when it adopts a persona? When a model role-plays as Darwin in 1882, it denies all knowledge of DNA, and readily asserts that species change through natural selection, but to what extent does it actually believe these assertions? Language models easily adopt different personas, but we still don't have a strong understanding of whether persona adoption changes [...] ---Outline:(00:12) TL;DR(01:15) Introduction(02:32) Method(05:35) Results(05:38) A spectrum of internalization across fine-tuning interventions(07:05) Role-play protects the persona's falsehoods, but selectively(09:43) Emergent Misalignment moves the truth representation broadly(12:06) Behavior and probes each mislead alone(13:09) Limitations(15:11) Conclusion(16:38) Links --- First published: July 2nd, 2026 Source: https://www.lesswrong.com/posts/EJQngix4rAgpPDTpT/when-role-playing-do-models-believe-what-they-say --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 3, 2026

Embed this episode

NOW PLAYING

“When Role-playing, Do Models Believe What They Say?” by Sturb, David Africa, Sid Black

0:00 17:37

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 17 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 3, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!