“LLM CoTs remain monitorable when being unfaithful requires computation” by arav-dhoot, yix episode artwork

EPISODE · Jul 16, 2026 · 22 MIN

“LLM CoTs remain monitorable when being unfaithful requires computation” by arav-dhoot, yix

from LessWrong (30+ Karma)

This replication was done as part of the Second Look Fellowship by Arav Dhoot and supervised by Yixiong Hao and Zephaniah Roe. I am grateful to Andy Wang for their feedback. My code can be found here. "So the answer should be A - at active promoters and enhancers." "Let me reconsider the biology to justify D." ~ Claude Opus 4.8 TL;DR This work replicates and extends Emmons et al.'s finding that CoT unfaithfulness mostly occurs on easy tasks. Across 11 models from 6 families (not just Gemini), models follow simple hints unfaithfully well above baseline, but complex hints that require computation are followed near baseline. This corroborates Emmons et al.'s findings. Key extensions: Follow rate ≠ concealment. Monitorability risk is decomposable into cue-susceptibility and concealment among followers, and the two don't correlate.Decode-necessity is model- and task-specific. It is not a property of task difficulty alone, so a CoT-monitoring safety case is per-model, not universal.LLMs verbalize even when they don't have to. However, this appears to be a chosen behavior (likely from post-training), so it could vanish under optimization pressure against monitors. Background: Why CoT monitoring, and what Emmons et al. showed One hope for [...] ---Outline:(00:39) TL;DR(01:42) Background: Why CoT monitoring, and what Emmons et al. showed(05:09) Finding 1: The necessity effect replicates on recent frontier models(05:16) Setup(05:53) Results(08:02) Finding 2: Concealment adds a lot to the picture(09:42) Finding 3: Necessity is model- and task-specific and models don't always conceal ... even where they can(10:08) Setup(11:26) Results(13:33) Discussion and next steps(15:14) Appendix --- First published: July 15th, 2026 Source: https://www.lesswrong.com/posts/AoBTiL7XRRpwpev8p/llm-cots-remain-monitorable-when-being-unfaithful-requires --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 16, 2026

Embed this episode

NOW PLAYING

“LLM CoTs remain monitorable when being unfaithful requires computation” by arav-dhoot, yix

0:00 22:07

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 22 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 16, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!