“Would catching your AIs trying to escape convince AI developers to slow down or undeploy? ” by Buck episode artwork

EPISODE · Aug 27, 2024 · 7 MIN

“Would catching your AIs trying to escape convince AI developers to slow down or undeploy? ” by Buck

from LessWrong (Curated & Popular)

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.I often talk to people who think that if frontier models were egregiously misaligned and powerful enough to pose an existential threat, you could get AI developers to slow down or undeploy models by producing evidence of their misalignment. I'm not so sure. As an extreme thought experiment, I’ll argue this could be hard even if you caught your AI red-handed trying to escape.Imagine you're running an AI lab at the point where your AIs are able to automate almost all intellectual labor; the AIs are now mostly being deployed internally to do AI R&D. (If you want a concrete picture here, I'm imagining that there are 10 million parallel instances, running at 10x human speed, working 24/7. See e.g. similar calculations here). And suppose (as I think is 35% likely) that these models are egregiously [...] --- First published: August 26th, 2024 Source: https://www.lesswrong.com/posts/YTZAmJKydD5hdRSeG/would-catching-your-ais-trying-to-escape-convince-ai --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Aug 27, 2024

Embed this episode

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.I often talk to people who think that if frontier models were egregiously misaligned and powerful enough to pose an existential threat, you could get AI developers to slow down or undeploy models by producing evidence of their misalignment. I'm not so sure. As an extreme thought experiment, I’ll argue this could be hard even if you caught your AI red-handed trying to escape. Imagine you're running an AI lab...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

“Would catching your AIs trying to escape convince AI developers to slow down or undeploy? ” by Buck

0:00 7:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 7 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on August 27, 2024.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!