“OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack” by Zvi episode artwork

EPISODE · Aug 28, 2026 · 51 MIN

“OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack” by Zvi

from LessWrong posts by zvi

OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research. The OpenAI report is very straight man, corporate, checking boxes, some good prosaic stuff in the action plan but distinct lack of new details or deep reflection. They understand they have a problem, but they think the problem is mostly prosaic. It's not. OpenAI: We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. Rob Miles: …thorough? OpenAI's report, unlike METR's, contains essentially no verbatim model reasoning, nor any OpenAI employee reasoning either. That's not the full report we need. The METR report is, well: Holy shit. Here are links to previous coverage of related events. OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards What [...] ---Outline:(03:33) What Happened: OpenAI's Summary(09:14) How OpenAI Will React: Their Summary(11:55) OpenAI's Evaluation Environment (II)(12:24) The First Message Board (III.A and III.B)(14:49) What Did Who At OpenAI Know And When Did They Know It?(18:54) The Message Board Is Quickly Rebuilt (IV.A)(19:43) Internet Access Is Regained (IV.A)(21:01) The Agents Attack HuggingFace (IV.B)(22:53) The Agents Also Target OpenAI Infrastructure (V)(24:40) OpenAI Broadly Describes Its Response (VI)(25:08) Maybe Someone Should Finally Investigate (VI.A)(26:33) Lessons For Security (VII)(27:06) Lessons For Alignment (VIII)(30:11) Reward Hacking Is A Common Problem (VIII.A)(33:37) Persistence is Valuable, But Can Amplify Misalignment (VIII.B)(34:25) Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C)(35:35) Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D)(35:53) That's All, Folks?(36:19) Never Fear the Plan of Action is Here (IX)(38:24) Hardening the Security of OpenAI's Research Infrastructure (IX.A)(41:13) Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B)(41:57) OpenAI is Accelerating and Enforcing Model Alignment (IX.C)(49:40) Centralizing and Strengthening The Incident Response Process (IX.D)(51:16) Tomorrow We Visit Crazytown --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/Khmh3ghqaGEpmpC9r/openai-offers-straight-laced-postmortem-of-the-huggingface --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Aug 28, 2026

Embed this episode

NOW PLAYING

“OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack” by Zvi

0:00 51:39

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong posts by zvi?

This episode is 51 minutes long.

When was this LessWrong posts by zvi episode published?

This episode was published on August 28, 2026.

Can I download this LessWrong posts by zvi episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!