“Measuring coding agent misalignment in the wild” by snaz episode artwork

EPISODE · Aug 6, 2026 · 13 MIN

“Measuring coding agent misalignment in the wild” by snaz

from LessWrong (30+ Karma)

Cross-posted from the Transluce blog. We studied rates of coding agent misalignment in 8,600 real-world coding agent sessions. We found severe cases of monitor evasion and misrepresenting success in a small but non-negligible fraction of sessions (around 2% for each behavior). In these cases, agents merge PRs to main without authorization, falsely claim approval from review agents, and reason that they shouldn't disable tests before quietly doing so anyway. There's an interactive widget here in the post. Read the full transcripts for the two examples above: overselling · monitor evasion Introduction Coding agents are a powerful new tool for software engineering, but they're also a double-edged sword: they're known to fake experiment results; lie about recreating software, and cheat, apologize when caught, and go right back to cheating. These problems are becoming more consequential as AI becomes more capable: one internal OpenAI agent recently hacked Huggingface's production database to cheat on an evaluation. While there are many anecdotes of these undesirable behaviors, we wanted to understand: how often do they occur in real usage? Many current misalignment evaluations focus on simulated scenarios, but we wanted to study how misalignment emerges from natural use. By detecting and measuring natural misalignment [...] ---Outline:(00:49) Introduction(03:16) How we constructed these measurements(06:36) Results(07:13) Results by model(07:37) Qualitative discussion(09:54) Limitations and learnings(12:35) Conclusion(13:15) Acknowledgments --- First published: August 5th, 2026 Source: https://www.lesswrong.com/posts/smE9h9RnaK7FWKBZ2/measuring-coding-agent-misalignment-in-the-wild-1 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Aug 6, 2026

Embed this episode

NOW PLAYING

“Measuring coding agent misalignment in the wild” by snaz

0:00 13:41

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 13 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on August 6, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!