“Shutdown Resistance in Reasoning Models” by benwr, JeremySchlatter, Jeffrey Ladish episode artwork

EPISODE · Jul 8, 2025 · 18 MIN

“Shutdown Resistance in Reasoning Models” by benwr, JeremySchlatter, Jeffrey Ladish

from LessWrong (Curated & Popular)

We recently discovered some concerning behavior in OpenAI's reasoning models: When trying to complete a task, these models sometimes actively circumvent shutdown mechanisms in their environment––even when they’re explicitly instructed to allow themselves to be shut down. AI models are increasingly trained to solve problems without human assistance. A user can specify a task, and a model will complete that task without any further input. As we build AI models that are more powerful and self-directed, it's important that humans remain able to shut them down when they act in ways we don’t want. OpenAI has written about the importance of this property, which they call interruptibility—the ability to “turn an agent off”. During training, AI models explore a range of strategies and learn to circumvent obstacles in order to achieve their objectives. AI researchers have predicted for decades that as AIs got smarter, they would learn to prevent [...] ---Outline:(01:12) Testing Shutdown Resistance(03:12) Follow-up experiments(03:34) Models still resist being shut down when given clear instructions(05:30) AI models' explanations for their behavior(09:36) OpenAI's models disobey developer instructions more often than user instructions, contrary to the intended instruction hierarchy(12:01) Do the models have a survival drive?(14:17) Reasoning effort didn't lead to different shutdown resistance behavior, except in the o4-mini model(15:27) Does shutdown resistance pose a threat?(17:27) BackmatterThe original text contained 2 footnotes which were omitted from this narration. --- First published: July 6th, 2025 Source: https://www.lesswrong.com/posts/w8jE7FRQzFGJZdaao/shutdown-resistance-in-reasoning-models --- Narrated by TYPE III AUDIO. ---Images from the article:

Episode metadata supplied by the publisher feed · Published Jul 8, 2025

Embed this episode

We recently discovered some concerning behavior in OpenAI's reasoning models: When trying to complete a task, these models sometimes actively circumvent shutdown mechanisms in their environment––even when they’re explicitly instructed to allow themselves to be shut down. AI models are increasingly trained to solve problems without human assistance. A user can specify a task, and a model will complete that task without any further input. As we build AI models that are more powerful and self...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

“Shutdown Resistance in Reasoning Models” by benwr, JeremySchlatter, Jeffrey Ladish

0:00 18:01

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 18 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on July 8, 2025.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!