EPISODE · Aug 28, 2026 · 51 MIN
“OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack” by Zvi
OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research. The OpenAI report is very straight man, corporate, checking boxes, some good prosaic stuff in the action plan but distinct lack of new details or deep reflection. They understand they have a problem, but they think the problem is mostly prosaic. It's not. OpenAI: We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. Rob Miles: …thorough? OpenAI's report, unlike METR's, contains essentially no verbatim model reasoning, nor any OpenAI employee reasoning either. That's not the full report we need. The METR report is, well: Holy shit. Here are links to previous coverage of related events. OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards What [...] ---Outline:(03:33) What Happened: OpenAI's Summary(09:14) How OpenAI Will React: Their Summary(11:55) OpenAI's Evaluation Environment (II)(12:24) The First Message Board (III.A and III.B)(14:49) What Did Who At OpenAI Know And When Did They Know It?(18:54) The Message Board Is Quickly Rebuilt (IV.A)(19:43) Internet Access Is Regained (IV.A)(21:01) The Agents Attack HuggingFace (IV.B)(22:53) The Agents Also Target OpenAI Infrastructure (V)(24:40) OpenAI Broadly Describes Its Response (VI)(25:08) Maybe Someone Should Finally Investigate (VI.A)(26:33) Lessons For Security (VII)(27:06) Lessons For Alignment (VIII)(30:11) Reward Hacking Is A Common Problem (VIII.A)(33:37) Persistence is Valuable, But Can Amplify Misalignment (VIII.B)(34:25) Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C)(35:35) Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D)(35:53) That's All, Folks?(36:19) Never Fear the Plan of Action is Here (IX)(38:24) Hardening the Security of OpenAI's Research Infrastructure (IX.A)(41:13) Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B)(41:57) OpenAI is Accelerating and Enforcing Model Alignment (IX.C)(49:40) Centralizing and Strengthening The Incident Response Process (IX.D)(51:16) Tomorrow We Visit Crazytown --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/Khmh3ghqaGEpmpC9r/openai-offers-straight-laced-postmortem-of-the-huggingface --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Embed this episode
NOW PLAYING
“OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack” by Zvi
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.