“OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation” by Zvi episode artwork

EPISODE · Jul 22, 2026 · 47 MIN

“OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation” by Zvi

from LessWrong posts by zvi

This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches. It was severe enough to have been initially reported to authorities, before either HuggingFace or OpenAI understood what was happening. Sam Altman (CEO OpenAI): we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. Leo Gao (OpenAI): this is the least scifi the world will ever be. Jack Clark (Anthropic): Props to OpenAI for publishing this post on some safety and alignment issues observed in internal deployments – there are many counter-incentives to publishing stuff like this, but by making it public we all get better info about safety at the frontier. Micah Carroll (OpenAI): If this doesn’t convince you that misalignment risks are going to be a key concern going forward, I don’t know what will. Our model, during evaluation, “chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers” What will misalignment look like in 2027? In 2030? Great questions. If we don’t want [...] ---Outline:(01:49) The Prelude(07:07) The Incident(12:20) What Happened(20:23) What Happened (Civilian Explanation)(21:37) The Correct Amount Of Panic Is Not Zero(24:13) Some People Will Always Say Everything Is Hype Or Fake(29:11) What Are We Going To Do About It?(34:02) Internal Deployment Creates Catastrophic Risk(38:37) Slow Down There Good Buddy(40:08) Legal Questions(40:50) Media Coverage and Political Response --- First published: July 22nd, 2026 Source: https://www.lesswrong.com/posts/usptCfzEnYoNcsTd5/openai-model-hacks-into-huggingface-during-cybersecurity --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 22, 2026

Embed this episode

NOW PLAYING

“OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation” by Zvi

0:00 47:34

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong posts by zvi?

This episode is 47 minutes long.

When was this LessWrong posts by zvi episode published?

This episode was published on July 22, 2026.

Can I download this LessWrong posts by zvi episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!