“Tie training can make DPO/RLHF-trained AIs generalize better” by Elliott Thornley, Christian Moya Calderon, Alex Semendinger episode artwork

EPISODE · Jul 6, 2026 · 31 MIN

“Tie training can make DPO/RLHF-trained AIs generalize better” by Elliott Thornley, Christian Moya Calderon, Alex Semendinger

from LessWrong (30+ Karma)

This post covers our recent ICML paper: Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training. TL;DR Our theorems and experiments suggest that DPO and RLHF have an unwelcome consequence: they make AIs care about every feature of actions that correlates with true value on the training distribution.[1] That's true even if the training set contains no misspecified preference data.And it's true even in the infinite-data limit.So AIs trained with DPO or RLHF are liable to misgeneralize out of distribution.Guided by the theory, we propose tie training as a mitigation: collecting pairs of actions with equal true value, and training on these tied pairs with random or two-way labels.Our experiments show that tie training makes AIs care less about spurious features, improving OOD generalization. Figure 1: Overview of our LLM experiment. We present Llama-3.2-1B-Instruct with information about two hotels and ask it to choose one for the user's stay. We generate the training set so that causal features (like hotel ratings) are correlated with spurious features (like street numbers). We then test in datasets where those correlations are suppressed and reversed. When we train with ordinary DPO, the model is led astray by [...] ---Outline:(00:27) TL;DR(02:03) Goal misgeneralization(03:50) Utility function model(06:14) Theorems on DPO/RLHF learning spurious correlations(08:30) Tie training(10:50) Experiments(10:53) Linear models(11:45) Neural networks and LLMs(18:40) Why does tie training work?(19:02) Utility function model(19:53) Persona selection model(20:27) Behavioral selection model(23:33) Tie training vs. slight-preference training(26:29) Tie training vs. aimed-preference training(28:32) How labs could implement tie training(30:19) Future work The original text contained 10 footnotes which were omitted from this narration. --- First published: July 6th, 2026 Source: https://www.lesswrong.com/posts/i2qTghrkyY9xdcCFq/tie-training-can-make-dpo-rlhf-trained-ais-generalize-better --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 6, 2026

Embed this episode

NOW PLAYING

“Tie training can make DPO/RLHF-trained AIs generalize better” by Elliott Thornley, Christian Moya Calderon, Alex Semendinger

0:00 31:09

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 31 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 6, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!