“Models don’t seem to be dishonest in the way humans are” by David Africa, Jacob Pfau episode artwork

EPISODE · Jul 22, 2026 · 19 MIN

“Models don’t seem to be dishonest in the way humans are” by David Africa, Jacob Pfau

from LessWrong (30+ Karma)

TLDR Models often behave dishonestly without acquiring a coherent deceptive disposition.We trained some mid-sized models on their own plausible but false reasoning.True and false training usually produced nearly identical downstream effects.Even statements contradicting latent knowledge transferred only weakly to unrelated dishonesty.General deception may require agency, persistent private information, and successful concealment over time. Introduction Current frontier models are mundanely misaligned. That is, they oversell work, claim completion too early, reward hack in ways users would reasonably call dishonest. But they do not yet seem dishonest in the way a person is dishonest. Models have this sort of sheepishness, acting abashed when called out, and then, as if, forgetting, doing it again. Humans caught hacking and overselling would typically go further and dissemble, or be defensive. What are the generalization boundaries of dishonesty? Overselling your work is kind of like misrepresenting it, reward hacking is kind of like covering it up. And we know language models love to generalize. Yet models don't seem to make the jump from "behaves in ways that look dishonest" to "is dishonest" in the way a person would be, lacking a sort of coherent motivation. If we are correct that [...] ---Outline:(00:12) TLDR(00:46) Introduction(05:17) Experiment(08:39) Results(15:09) Discussion(17:33) Conclusion The original text contained 1 footnote which was omitted from this narration. --- First published: July 22nd, 2026 Source: https://www.lesswrong.com/posts/QYmnkQyZD2fDjHCJ8/models-don-t-seem-to-be-dishonest-in-the-way-humans-are --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 22, 2026

Embed this episode

NOW PLAYING

“Models don’t seem to be dishonest in the way humans are” by David Africa, Jacob Pfau

0:00 19:04

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 19 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 22, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!