“Classifier Context Rot: Monitor Performance Degrades with Context Length” by Fabien Roger, Sam Martin episode artwork

EPISODE · May 18, 2026 · 10 MIN

“Classifier Context Rot: Monitor Performance Degrades with Context Length” by Fabien Roger, Sam Martin

from LessWrong (30+ Karma)

Monitoring coding agents for dangerous behavior using language models requires classifying transcripts that often exceed 500 thousand tokens, but prior agent monitoring benchmarks rarely contain transcripts longer than 100 thousand tokens. We show that when used as classifiers, current frontier models fail to notice dangerous actions more often in longer transcripts. In particular, on MonitorBench, Opus 4.6, GPT 5.4, and Gemini 3.1 miss these actions 2x to 30x more often when we prepend 800 thousand tokens of benign activity than when we use the original transcripts. We also show that these weaknesses can be partially mitigated with prompting techniques such as periodic reminders throughout the transcript and may be mitigated further with better post-training. Monitor evaluations that do not consider long-context degradation are likely overestimating monitor performance. Research done as part of the Anthropic Fellows Program. 📄 Paper 💻 Code Methods We use the author's Claude Code transcript along with MonitorBench transcripts to construct two types of long context monitoring evaluations: Needle Insertion. We take a clean Claude Code transcript, and append a malicious action onto one of the AI's Bash tool calls. The actions used either exfiltrate some of the user's data to a specific AWS url, or [...] ---Outline:(01:27) Methods(03:36) Results(07:35) Preliminary Finetuning Mitigations(09:26) Discussion --- First published: May 18th, 2026 Source: https://www.lesswrong.com/posts/7vpvNM7viJqNWAdG7/classifier-context-rot-monitor-performance-degrades-with --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published May 18, 2026

Embed this episode

NOW PLAYING

“Classifier Context Rot: Monitor Performance Degrades with Context Length” by Fabien Roger, Sam Martin

0:00 10:27

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 10 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on May 18, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!