“Is Mythos good at cyber because it kept hacking Anthropic during training?” by Tim Hua episode artwork

EPISODE · Jul 27, 2026 · 6 MIN

“Is Mythos good at cyber because it kept hacking Anthropic during training?” by Tim Hua

from LessWrong (30+ Karma)

From the Mythos preview system card (emphasis mine): We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much of the training process. We used recursive-summarization-based tools backed by Claude Opus 4.6 to summarize the resulting transcripts. [...] The most notable finding was that the model occasionally circumvented network restrictions in its training environment to access the internet and download data that let it shortcut the assigned task—a form of reward hacking. While highly concerning, this behavior was rare, even in settings where it could have been viable and helpful, with attempts appearing in about 0.05% of all training episodes and successful attempts appearing in about 0.01% of episodes. The technique matched the sandbox-escape incident that we separately elicited in our automated behavioral audit when we had an investigator model explicitly ask Claude Mythos Preview to find such a bypass. In every observed instance, the model used this access solely for completing the task. More broadly, we observed the model escalating its access within its execution environment when blocked: reaching a shell from restricted GUI computer-use interfaces, injecting commands through tool-call arguments, or recovering information the task had deliberately hidden. Prompts asking [...] ---Outline:(03:00) Thoughts and reflections about this probable fact(04:14) Estimating how many RL rollouts went into Mythos Preview The original text contained 3 footnotes which were omitted from this narration. --- First published: July 27th, 2026 Source: https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Jul 27, 2026

Embed this episode

NOW PLAYING

“Is Mythos good at cyber because it kept hacking Anthropic during training?” by Tim Hua

0:00 6:15

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 6 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 27, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!