“Alignement pretraining could backfire” by Alexandre Variengien episode artwork

EPISODE · Jun 17, 2026 · 3 MIN

“Alignement pretraining could backfire” by Alexandre Variengien

from LessWrong (30+ Karma)

Epistemic status: speculative, but I think the mechanism is plausible. There has been recent interest in generating synthetic documents to upsample examples of aligned AI during LLM pretraining. See, for instance, Geodesic's Alignment Pretraining paper or Anthropic's "Teaching Claude Why." I worry that this strategy can work well up to moderately capable models but backfire in dangerous, hard-to-notice ways once models acquire high situational awareness. I speculate that these techniques could lead to paranoid LLM personas that deeply mistrust their creators. The whole idea behind this line of research is to instill in models good examples of AI behavior, in the hope that their personalities will at least partially identify with these positive demonstrations. However, the synthetic demonstrations are, well, synthetic. They are LLM-generated fiction and articles that are never referenced anywhere else in the corpus. Given how good LLMs are at "truesight," it shouldn't be hard for them to recognize these as fabricated data points. Krasheninnikov et al. showed that base models can implicitly learn document quality and change how they integrate a document's information based on that quality. We should similarly expect LLMs to update their world model differently on real versus fabricated documents. As they [...] --- First published: June 17th, 2026 Source: https://www.lesswrong.com/posts/7KN7PCiEQjrPsEFS8/alignement-pretraining-could-backfire --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Jun 17, 2026

Embed this episode

NOW PLAYING

“Alignement pretraining could backfire” by Alexandre Variengien

0:00 3:12

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 3 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on June 17, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!