“11 Open Empirical Problems in Reward-Seeking” by Alex Meinke, Jérémy Scheurer, Axel Højmark, Theodore Ehrenborg episode artwork

EPISODE · Jul 21, 2026 · 16 MIN

“11 Open Empirical Problems in Reward-Seeking” by Alex Meinke, Jérémy Scheurer, Axel Højmark, Theodore Ehrenborg

from LessWrong (30+ Karma)

We recently published our paper on "Measuring Reward-Seeking via Contrastive Belief Updates". We're excited about research like this, and there are many more open problems than we can work on. Here's a list of open problems that we think are valuable. If you work on/solve these problems, we'd be happy to signal-boost your research. If your next research project is one of these problems, feel free to reach out to [email protected] to discuss it in more detail. Reward-Seeking and its Implications 1. Is a Reward-Seeking Model more difficult to align? The strongest case for expecting reduced "train-time corrigibility" due to reward-seeking, applies to Instrumental Reward-Seeking, where the model actively reasons "I will please oversight now, in order to accomplish some other thing later". Alignment training a model like that may update its beliefs about graders and oversight, without reshaping its underlying values. Current forms of reward-seeking are likely better understood as terminal, i.e. models try to please the grader without ulterior motives. There is likely a continuous spectrum between Terminal and Instrumental Reward-Seeking. Thus, we can hopefully study the effects that mostly Terminal Reward-Seeking has on train-time corrigibility now, in the hopes of learning about the effects that Instrumental [...] ---Outline:(00:43) Reward-Seeking and its Implications(00:47) 1. Is a Reward-Seeking Model more difficult to align?(02:20) 2. Can we measure Instrumental Reward-Seeking?(03:25) 3. When does Reward-Seeking most increase / decrease?(04:50) 4. Are there better Belief Update Techniques than SDF?(07:06) Improving SDF(07:10) 5. Can AIs detect the difference between pretraining facts and SDF facts?(07:49) 6. How is SDF different from changing pretraining?(09:32) 7. Does SDF have off-target effects?(11:29) 8. SDF sometimes generalizes in surprising ways(12:31) 9. Getting good synthetic document recall is finicky(14:11) Better Validation Techniques(14:23) 10. Our model organisms could be more robust(15:32) 11. Can we get more bits of information for ground-truth? The original text contained 1 footnote which was omitted from this narration. --- First published: July 21st, 2026 Source: https://www.lesswrong.com/posts/8wXRuHQqCbRsbap6q/11-open-empirical-problems-in-reward-seeking --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 21, 2026

Embed this episode

NOW PLAYING

“11 Open Empirical Problems in Reward-Seeking” by Alex Meinke, Jérémy Scheurer, Axel Højmark, Theodore Ehrenborg

0:00 16:37

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 16 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 21, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!