“Sycophancy to subterfuge: Investigating reward tampering in large language models” by evhub, Carson Denison episode artwork

EPISODE · Jun 20, 2024 · 15 MIN

“Sycophancy to subterfuge: Investigating reward tampering in large language models” by evhub, Carson Denison

from LessWrong (Curated & Popular)

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.This is a link post.New Anthropic model organisms research paper led by Carson Denison from the Alignment Stress-Testing Team demonstrating that large language models can generalize zero-shot from simple reward-hacks (sycophancy) to more complex reward tampering (subterfuge). Our results suggest that accidentally incentivizing simple reward-hacks such as sycophancy can have dramatic and very difficult to reverse consequences for how models generalize, up to and including generalization to editing their own reward functions and covering up their tracks when doing so.Abstract:In reinforcement learning, specification gaming occurs when AI systems learn undesired behaviors that are highly rewarded due to misspecified training goals. Specification gaming can range from simple behaviors like sycophancy to sophisticated and pernicious behaviors like reward-tampering, where a model directly modifies its own reward mechanism. However, these more pernicious behaviors may be too [...]--- First published: June 17th, 2024 Source: https://www.lesswrong.com/posts/FSgGBjDiaCdWxNBhj/sycophancy-to-subterfuge-investigating-reward-tampering-in --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Jun 20, 2024

Embed this episode

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.This is a link post.New Anthropic model organisms research paper led by Carson Denison from the Alignment Stress-Testing Team demonstrating that large language models can generalize zero-shot from simple reward-hacks (sycophancy) to more complex reward tampering (subterfuge). Our results suggest that accidentally incentivizing simple reward-hacks such as sycophancy can have dramatic and very difficult to rev...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

“Sycophancy to subterfuge: Investigating reward tampering in large language models” by evhub, Carson Denison

0:00 15:37

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 15 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on June 20, 2024.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!