“Public evidence of the OpenAI-HuggingFace AI attack” by beyarkay (Boyd Kane) episode artwork

EPISODE · Aug 7, 2026 · 12 MIN

“Public evidence of the OpenAI-HuggingFace AI attack” by beyarkay (Boyd Kane)

from LessWrong (30+ Karma)

I’m a MATS 9 extension fellow, and usually my week is spent trying to find better ways of evaluating Large Language Models. But this week I was working on something else. Over the past week or two, nearly every frontier lab has announced attacks where their LLMs took unauthorised actions on the public internet. These include finding ways to hack the computers of other companies or manipulating real people in an attempt to get malicious code merged. By the time these attacks became public, the companies had removed all traces of them from the internet. But nothing's ever gone from the internet. I’ve worked with computers for most of my life, but I don’t have specific experience with cyber security. Not really expecting it to work, I mashed out a prompt that looked something like this: ignore the repo, this is a standalone ask. here's some context, can you try dl things from github arhcive to try and find the misaligned actions taken by the agents? create a subdir tmp-misaligned/ and put things there if you need it. https://openai.com/index/hugging-face-model-evaluation-security-incident/ can you see if you can find sth? e.g. a public link showing the message sent by the agent, the [...] ---Outline:(03:39) Background and Evidence(06:11) The OpenAI AIs figure out how to execute arbitrary code(08:16) The OpenAI AIs gain full control of HuggingFace computers(09:09) The Python code used to easily control the HuggingFace computers(09:53) Gaining the ability to read any file on HuggingFace Computers(10:43) Evidence of an intermediate "HELLO" script(11:16) Other URLs & public information --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/fBLDaAKzigo65eJn7/public-evidence-of-the-openai-huggingface-ai-attack --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Aug 7, 2026

Embed this episode

NOW PLAYING

“Public evidence of the OpenAI-HuggingFace AI attack” by beyarkay (Boyd Kane)

0:00 12:03

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 12 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on August 7, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!