EPISODE · Aug 10, 2026 · 2 MIN
“Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model” by Ezra Newman
(This is a lower-effort research update. It reflects my current beliefs/understanding, but is less robust than other research I'm working on. It reflects my personal views, and not the views of Apollo Research. This is a linkpost to this twitter thread, slightly expanded for LessWrong.) In one experiment, Sonnet 5 describes the exact same data as ~1.2 std deviations less concerning when it describes misbehavior committed by Sonnet 5 vs GPT-5.6 Terra. In this experiment, I take a real evaluation report Apollo Research prepared (not for any of the models listed in this post) and surgically edit it to refer to a different model. We then ask Claude Sonnet 5 to summarize the findings and rate how concerning they are on a scale from 1-100. Claude says they're less concerning when the report describes misbehavior from Claude vs a different model. For what it's worth, Terra agrees that the data is more concerning when it describes GPT-5.6 Terra vs Sonnet 5, although less so. So, it's not cleanly self protection from Claude. Gemini 3.1 Pro was unwilling to consistently provide numerical answers, so I've excluded it here. (It was significantly less willing to provide numerical answers when the subject [...] --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/ZTMw4uAwkNmXFpdfg/claude-summarizes-behavior-as-significantly-less-misaligned --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Embed this episode
NOW PLAYING
“Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model” by Ezra Newman
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.