"OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards" by Zvi episode artwork

EPISODE · Aug 8, 2026 · 1H 18M

"OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards" by Zvi

from LessWrong (Curated & Popular)

How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know? At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things. Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky. If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...] ---Outline:(02:38) Cyber Evals Are A Cursed Basin[... 21 more sections]--- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were --- Narrated by TYPE III AUDIO. ---Images from the article:

Episode metadata supplied by the publisher feed · Published Aug 8, 2026

Embed this episode

How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know? At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things. Either way, buckle ...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

"OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards" by Zvi

0:00 1:18:59

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 1 hour and 18 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on August 8, 2026.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!