“Foundation Models for Oversight” by jsteinhardt episode artwork

EPISODE · Jul 28, 2026 · 45 MIN

“Foundation Models for Oversight” by jsteinhardt

from LessWrong (30+ Karma)

Cross-posted from the Transluce blog. To oversee an AI model, we'd ideally like to ask questions such as: What are important situations where the model sandbags? Does the model have an objective it wouldn't admit to if asked directly? Does the model treat a user differently once it infers something about their identity, and along what axis? Is the model's chain of thought load-bearing, or is it a post-hoc rationalization of an answer that was already settled on? Is it reward hacking on this input, or actually trying to solve the task? It would be great if we had an oversight assistant that could answer these questions. We'd want it to do three things: help us formalize the question as a testable empirical criterion; produce data that satisfies that criterion; and do so in a way we can justifiably trust. To get such an assistant, we lay out a vision for building a foundation model for oversight: an AI system mid-trained (or pre-trained) on a large, diverse corpus of experiments on a given "subject model", RLVR'd on a large number of verified oversight tasks, and then fine-tuned to answer [...] ---Outline:(05:55) Conceptual preliminaries(05:59) Pythonic world models(07:21) Oversight as Inference(10:42) Reducing oversight to autoregressive prediction(11:45) Engineering Scale-up(11:49) Generating supervised oversight data for mid-training(13:35) Step 1: Sampling diverse experiments(15:57) Step 2: Sampling diverse experiment inputs(18:05) Step 3: Featurizing as token sequences(20:17) Calling our shots: staged de-risking(23:49) Stage 1: Individual Tasks(27:14) Stage 2: Cross-Task Transfer(28:41) Stage 3: Zero-Shot Abilities(30:07) Finishing Touches(30:11) RLVR(33:55) Generalizing RLVR to Other Tasks(35:27) Example Trajectory(36:32) De-risking RLVR(36:57) Stage 4: RLVR is competitive with evolution for elicitation(37:58) Stage 5: Cross-Task Transfer for RLVR(38:52) Post-training(41:09) Appendix(41:12) Further Testing the Oversight-as-Inference Hypothesis The original text contained 8 footnotes which were omitted from this narration. --- First published: July 28th, 2026 Source: https://www.lesswrong.com/posts/AqdZKyoRmN6EFCzib/foundation-models-for-oversight --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 28, 2026

Embed this episode

NOW PLAYING

“Foundation Models for Oversight” by jsteinhardt

0:00 45:55

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 45 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 28, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!