EPISODE · Jan 22, 2026 · 22 MIN
Confessions of a Large Language Model
from Paul, Weiss Waking Up With AI · host Paul, Weiss
In this episode, Katherine Forrest and Scott Caravello unpack OpenAI researchers’ proposed “confessions” framework designed to monitor for and detect dishonest outputs. They break down the researchers’ proof of concept results and the framework’s resilience to reward hacking, along with its limits in connection with hallucinations. Then they turn to Google DeepMind’s “Distributional AGI Safety,” exploring a hypothetical path to AGI via a patchwork of agents and routing infrastructure, as well as the authors’ proposed four layer safety stack. ## Learn More About Paul, Weiss’s Artificial Intelligence practice: https://www.paulweiss.com/industries/artificial-intelligence
Embed this episode
NOW PLAYING
Confessions of a Large Language Model
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.