“Many alignment techniques work by training one model and deploying another” by cloud episode artwork

EPISODE · Jul 19, 2026 · 11 MIN

“Many alignment techniques work by training one model and deploying another” by cloud

from LessWrong (30+ Karma)

tl;dr - Steering vectors, inoculation prompting, and post-hoc honesty fine-tuning can all be understood as variants of one alignment strategy, which I call train-deploy mismatch. Each trains the model in one configuration and deploys it in another. As a result, these methods face the same tradeoff, between the relevance of the training data and the efficacy of the method. Note: Others have had similar ideas and shaped my thinking here including Sam Marks, Ariana Azarbal, Victor Gillioz, Alex Turner, Jacob Goldman-Wetzler, Jake Mendel, Daniel Tan, and Fabien Roger. Thanks to Monte MacDiarmid, Nat McAleese, Shawn Hu, and Jake Ward for input on an earlier draft. Background AI alignment is hard largely because we don't know how to specify what we want. Instead, we train models on proxies for what we want: labels and reward functions defined on data distributions chosen such that we hope the model will perform as desired when deployed into the world. This approach has worked well so far, but given increasing model capabilities, it may stop working— models may misgeneralize their training to catastrophically bad behavior in deployment. A pressing open problem is to figure out how to get models to generalize the properties that [...] ---Outline:(00:55) Background(02:06) Train-deploy mismatch as a general alignment strategy(04:17) Fundamental tradeoffs(07:01) Sidebar: is everything train-deploy mismatch?(07:43) Open threads(11:11) Closing thoughts The original text contained 1 footnote which was omitted from this narration. --- First published: July 19th, 2026 Source: https://www.lesswrong.com/posts/syAbdNei8BWeP2RPo/many-alignment-techniques-work-by-training-one-model-and --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 19, 2026

Embed this episode

NOW PLAYING

“Many alignment techniques work by training one model and deploying another” by cloud

0:00 11:49

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 11 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 19, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!