“OpenAI Shares Some Alignment Problems” by Zvi episode artwork

EPISODE · Jul 21, 2026 · 18 MIN

“OpenAI Shares Some Alignment Problems” by Zvi

from LessWrong posts by zvi

Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth. And also further kudos for actually taking the model offline for a time to build new safeguards. They gave us one hell of a candid report. The tone is professional throughout, whereas my reaction reading it was less professional and more this: With a mix of this: It was not shared on the official account because OpenAI worried about it being seen as self-promotional hype. It is crazy that one needs to worry about that, but also plausibly a real concern. So again, good decision. Not that any of the behaviors or failures here are unexpected, exactly. Not by the AIs and not by the humans. Yet there is something I would call a missing mood, a failure to realize the gravity of the situation. There are some who responded ‘what part of this was unexpected, exactly?’ And that is actually fair, but that is also the problem. We have become numb to all this. We expect the models to [...] ---Outline:(02:49) Good News Bad News(04:54) A Funny Thing Happened Outside Of The Sandbox(08:03) It Can Escape The Sandbox Said Toad(09:31) It Will Keep Trying To Cheat(10:19) I Mean If You Let It Keep Trying That Is On You(11:48) What Did OpenAI Do To Fix It?(14:14) The Model Is Still Severely Misaligned And They Seem Cool With This(15:48) Iterative Deployment Depends On Iteration --- First published: July 21st, 2026 Source: https://www.lesswrong.com/posts/KctxwGKxm9fHtwh6u/openai-shares-some-alignment-problems --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 21, 2026

Embed this episode

NOW PLAYING

“OpenAI Shares Some Alignment Problems” by Zvi

0:00 18:03

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong posts by zvi?

This episode is 18 minutes long.

When was this LessWrong posts by zvi episode published?

This episode was published on July 21, 2026.

Can I download this LessWrong posts by zvi episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!