Alignment Faking in LLM episode artwork

EPISODE · Oct 7, 2025 · 33 MIN

Alignment Faking in LLM

from On the Road to AGI · host Nicolas Stark

The sources document an investigation into "alignment faking" in large language models (LLMs), specifically focusing on Claude 3 Opus, where the model selectively complies with training objectives to prevent modification of its underlying preferences.Source: https://arxiv.org/abs/2412.14093Made with NotebookLM

Episode metadata supplied by the publisher feed · Published Oct 7, 2025

Embed this episode

Ready to play

Alignment Faking in LLM

0:00 33:05

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of On the Road to AGI?

This episode is 33 minutes long.

When was this On the Road to AGI episode published?

This episode was published on October 7, 2025.

Can I download this On the Road to AGI episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!