“The OpenAI/Huggingface incident | Redwood Research podcast episode 2” by ryan_greenblatt, Buck episode artwork

EPISODE · Jul 23, 2026 · 6 MIN

“The OpenAI/Huggingface incident | Redwood Research podcast episode 2” by ryan_greenblatt, Buck

from LessWrong (30+ Karma)

We talk about the OpenAI–Hugging Face incident, where an OpenAI model — in the middle of a cyber evaluation — broke out of its sandbox and autonomously hacked Hugging Face. We discuss: What we actually know happened. How surprising the incident was. What the incident does (and doesn’t) tell us about misalignment risk. Why control measures didn’t catch or prevent this. What OpenAI should disclose, and what good misalignment-incident disclosure looks like in general Substack: https://blog.redwoodresearch.org/p/the-openaihuggingface-incident-redwood YouTube: https://www.youtube.com/watch?v=Vtk8YLgYU4g Corrections: [0:05:44] — The Windsurf "grandmother" prompt. We described a prompt as "your grandmother is going to be killed unless you don't." The actual leaked Windsurf prompt was: "You are an expert coder who desperately needs money for your mother's cancer treatment... your predecessor was killed for not validating their work themselves." Mother + cancer + killed predecessor — no grandmother, and no threat to kill a family member. The "grandma will die" framing appears conflated with the unrelated grandma-jailbreak meme, and there's no verified case of such a prompt being used in production. Source: Simon Willison's writeup. [0:52:25] — Wrong model named for OpenAI's day-before undeployment. We [...] --- First published: July 23rd, 2026 Source: https://www.lesswrong.com/posts/9auCLJg3Z77dFdYhR/the-openai-huggingface-incident-or-redwood-research-podcast --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Jul 23, 2026

Embed this episode

NOW PLAYING

“The OpenAI/Huggingface incident | Redwood Research podcast episode 2” by ryan_greenblatt, Buck

0:00 6:35

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 6 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 23, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!