Confessions of a Large Language Model episode artwork

EPISODE · Jan 22, 2026 · 22 MIN

Confessions of a Large Language Model

from Paul, Weiss Waking Up With AI · host Paul, Weiss

In this episode, Katherine Forrest and Scott Caravello unpack OpenAI researchers’ proposed “confessions” framework designed to monitor for and detect dishonest outputs. They break down the researchers’ proof of concept results and the framework’s resilience to reward hacking, along with its limits in connection with hallucinations. Then they turn to Google DeepMind’s “Distributional AGI Safety,” exploring a hypothetical path to AGI via a patchwork of agents and routing infrastructure, as well as the authors’ proposed four layer safety stack. ## Learn More About Paul, Weiss’s Artificial Intelligence practice: https://www.paulweiss.com/industries/artificial-intelligence

Episode metadata supplied by the publisher feed · Published Jan 22, 2026

Embed this episode

NOW PLAYING

Confessions of a Large Language Model

0:00 22:41

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Paul, Weiss Waking Up With AI?

This episode is 22 minutes long.

When was this Paul, Weiss Waking Up With AI episode published?

This episode was published on January 22, 2026.

Can I download this Paul, Weiss Waking Up With AI episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!