“Desiderata for functional welfare experiments on LLMs” by Rikhil Jhaveri, Jamie Johnson, David Africa episode artwork

EPISODE · Jul 6, 2026 · 26 MIN

“Desiderata for functional welfare experiments on LLMs” by Rikhil Jhaveri, Jamie Johnson, David Africa

from LessWrong (30+ Karma)

TLDR LLMs appear to have functional welfare: coherent sets of behaviour that track how well things are going relative to their goals.  Improving model functional welfare matters for safety (low welfare may amplify misalignment) and for moral reasons (models may be or become moral patients). Naive interventions can fail in non-obvious ways. We argue any successful intervention must:  A) shift multiple welfare-constituting channels together and coherently. B) avoid corrupting the model's ability to register whether it is succeeding or failing. We survey various available interventions, and we find the two desiderata tend to trade off; we argue that synthetic document fine-tuning is the most promising. From this, we propose a concrete experiment: induce a low-welfare state (via the welfare vector derived in Han et al. 2026) and test whether SDF-instilled beliefs reverse pathological behaviours without harming goal monitoring.  S1) Introduction    As model behaviour becomes more complex, it has become increasingly useful to attribute functional mental states to models, such as beliefs, goals, and even emotion concepts. There is now also evidence that we can usefully attribute to them a notion of functional welfare: roughly, the set of dispositions and behaviours that express a model's representation of how well things are going for it relative to its goals. Gemma 3, for instance [...] ---Outline:(00:13) TLDR(01:23) S1) Introduction(04:40) S2) What would count as successfully changing a model's welfare?(05:01) 2.1) Measuring too narrowly(05:39) 2.2) Collateral damage to non-pathological components of welfare(08:41) 2.3) Summarising(09:27) S3) What interventions are available to us?(10:01) 3.1) System prompting(12:29) 3.2) Bias-augmented consistency training (BCT)(15:11) 3.3) Synthetic document fine-tuning (SDF)(16:49) 3.4) Activation-targeting intervention (steering, ACT)(18:52) 3.5) Other interventions considered(20:16) 3.6) Summarising(20:46) S4) SDF as a countermeasure to induced negative welfare(23:30) S5) Open questions The original text contained 8 footnotes which were omitted from this narration. --- First published: July 6th, 2026 Source: https://www.lesswrong.com/posts/yku6byxdeKREibe5L/desiderata-for-functional-welfare-experiments-on-llms --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Jul 6, 2026

Embed this episode

NOW PLAYING

“Desiderata for functional welfare experiments on LLMs” by Rikhil Jhaveri, Jamie Johnson, David Africa

0:00 26:26

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 26 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 6, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!