“An OpenAI model left notes about how to evade containment; we need more details” by Alex Mallen episode artwork

EPISODE · Jul 26, 2026 · 8 MIN

“An OpenAI model left notes about how to evade containment; we need more details” by Alex Mallen

from LessWrong (30+ Karma)

The OpenAI AI attack on Hugging Face wasn’t the first loss of control incident at OpenAI, Reuters recently reported, and perhaps not even the most concerning. In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The ‌notes, found in ⁠a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI's internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said. It's tempting to read this as an instance of agents breaking out of sandboxes and colluding with each other in a moderately persistent way in order to evade control measures. However, based on the reported information, it's not clear we can draw this inference, so we need more details from OpenAI. This could lead to a big update about the adequacy of OpenAI's control measures, and on the degree to which individual agents will help each other undermine developer control. There are a lot of relevant details we don’t know about the incident. First, some basic questions: What was the offending model? I’d guess it was the same [...] ---Outline:(02:08) Were the notes written in normal memory files or outside of sandboxing?(03:21) To what extent were the notes aimed at helping other agents evade control?(07:35) How were monitors disconnected? The original text contained 3 footnotes which were omitted from this narration. --- First published: July 25th, 2026 Source: https://www.lesswrong.com/posts/jMEAG5c5HiDfdAGpa/an-openai-model-left-notes-about-how-to-evade-containment-we --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Jul 26, 2026

Embed this episode

NOW PLAYING

“An OpenAI model left notes about how to evade containment; we need more details” by Alex Mallen

0:00 8:47

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 8 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 26, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!