EPISODE · Aug 7, 2026 · 1H 18M
“OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi
How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know? At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things. Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky. If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...] ---Outline:(02:39) Cyber Evals Are A Cursed Basin(05:16) Outside Of Cyber Evals Is Still Sufficiently Cursed(06:51) Cheat Cheat Cheat Cheat Cheat(12:07) Read The Message Board(14:48) Updating Your AI (Exploitation of OpenAI Internal Systems) Timelines(18:11) This Is The Way The World Ends(21:45) Shooting The Messenger Board(27:16) The Internal and HuggingFace Hacks(30:33) OpenAI Responds(33:27) When AIs Tell You Who They Are(35:43) The Once and Future Rise Of Functional Decision Theory(41:28) Don't Panic(43:38) Hackery In the UK(48:02) Mythos Knew It Was Real This Time(50:09) I Got 141,006 Test Runs With An Unintentional Open Path To The Internet And An Email Alert Aint One(53:46) Surely By Now You Know These Are Not Publicity Stunts(55:24) The Future Is Coming(57:17) The Investigations Begin(01:00:08) N Boats And Three Helicopters(01:01:43) Always Be Sandbox Red Teaming(01:12:54) Halt And Catch Fire(01:14:31) Truth and Reconciliation --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Embed this episode
NOW PLAYING
“OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.