“OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi episode artwork

EPISODE · Aug 7, 2026 · 1H 18M

“OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi

from LessWrong posts by zvi

How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know? At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things. Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky. If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...] ---Outline:(02:39) Cyber Evals Are A Cursed Basin(05:16) Outside Of Cyber Evals Is Still Sufficiently Cursed(06:51) Cheat Cheat Cheat Cheat Cheat(12:07) Read The Message Board(14:48) Updating Your AI (Exploitation of OpenAI Internal Systems) Timelines(18:11) This Is The Way The World Ends(21:45) Shooting The Messenger Board(27:16) The Internal and HuggingFace Hacks(30:33) OpenAI Responds(33:27) When AIs Tell You Who They Are(35:43) The Once and Future Rise Of Functional Decision Theory(41:28) Don't Panic(43:38) Hackery In the UK(48:02) Mythos Knew It Was Real This Time(50:09) I Got 141,006 Test Runs With An Unintentional Open Path To The Internet And An Email Alert Aint One(53:46) Surely By Now You Know These Are Not Publicity Stunts(55:24) The Future Is Coming(57:17) The Investigations Begin(01:00:08) N Boats And Three Helicopters(01:01:43) Always Be Sandbox Red Teaming(01:12:54) Halt And Catch Fire(01:14:31) Truth and Reconciliation --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Aug 7, 2026

Embed this episode

NOW PLAYING

“OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi

0:00 1:18:55

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong posts by zvi?

This episode is 1 hour and 18 minutes long.

When was this LessWrong posts by zvi episode published?

This episode was published on August 7, 2026.

Can I download this LessWrong posts by zvi episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!