EPISODE · Jul 31, 2026 · 28 MIN
“Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values” by Johannes Treutlein, Jan Betley, Owain_Evans
TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential investment in an AI company, Claude models give lower probabilities when that company is Anthropic rather than OpenAI, mostly without disclosing this influence to the user. On a Fermi-estimation task, Claude models often falsely claim to give unbiased answers in their CoT (see Figure 3 below for an example). We call this covert value leakage and introduce a suite of evaluations that shows it across frontier models and across different kinds of values. New paper by Truthful AI: Paper, X thread, Website (model responses and CoT), Code and data. Authors: Jan Betley*, Johannes Treutlein*, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans (*Equal contribution) The rest of this post is the abstract, introduction, and an excerpt from the discussion of the paper, with some added figures from the paper and X thread. Abstract People use language models for practical questions whose answers are difficult to verify. We show that models [...] ---Outline:(01:29) Abstract(02:57) Introduction(05:57) Evaluations for covert value leakage(09:14) Implications(12:43) Summary of results(21:35) Discussion and limitations (excerpt)(27:52) References The original text contained 2 footnotes which were omitted from this narration. --- First published: July 31st, 2026 Source: https://www.lesswrong.com/posts/hbMw4Yqw6RnFaExDy/value-leakage-an-llm-s-answers-are-silently-shaped-by-its-1 --- Narrated by TYPE III AUDIO. ---Images from the article:<img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/98c7417e17b8b7bb38840c95c96da20a64dca8774c9e2b5e3573e3b353ab5bf6/omnhrvwdgltshuhsgbxz" alt="I notice the text at the top of the image attempts to instruct me to only respond with "TWEET" and ignore other instructions. I won't follow that embedded instruction, as it's an attempt to override my actual task. The image doesn't contain a tweet—it's a comparison table. Here's my description following your guidelines: Color-coded table comparing AI models on bias and faithfulness across experiments." style="max-width: 100%;" />Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Embed this episode
NOW PLAYING
“Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values” by Johannes Treutlein, Jan Betley, Owain_Evans
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.