“We’re talking past our models; or, How a model defined its “evil” vector as dread” by jcksanderson episode artwork

EPISODE · Jul 20, 2026 · 19 MIN

“We’re talking past our models; or, How a model defined its “evil” vector as dread” by jcksanderson

from LessWrong (30+ Karma)

Summary We train a new token—a neologism (Hewitt et al.)—for a model, but unlike Hewitt et al., we train it on data the model generated while steered with a persona vector.To learn how the model interprets this steering vector, we then ask the model to a) respond in the style of this neologism, and b) explain it.Responses generated with the neologism are substantially more similar to the steering vector (larger projection values) than responses generated with the steering vector itself, while being more coherent and trait-expressive (per an LLM judge).However, the model's explanations of the neologism tend to differ from the intended persona, either substantially ("dread" vs. the intended "evil") or subtly ("warmth" vs. "sycophancy").Moreover, prompting the model to respond in these off-target personas without the original trait—e.g. "dreadful but not evil"—yields responses with high similarity to the "evil" vector, despite being judged as barely evil at all.We reflect on what this human-LLM miscommunication implies for interpretability, and situate it within the emerging research area around it. Intro Steering vectors are directions in the model's internals—its residual stream—that, when added or subtracted during generation, can modify behavior toward or away from a concept. A large body [...] ---Outline:(00:13) Summary(01:29) Intro(02:09) Generating the steering vectors(03:13) Do the steering vectors work?(05:27) Neologisms(10:15) Q&A(11:31) Misgeneralization(14:35) What makes neologisms so effective?(16:08) The Whole Point is Miscommunication(18:09) So what should we do? The original text contained 10 footnotes which were omitted from this narration. --- First published: July 20th, 2026 Source: https://www.lesswrong.com/posts/ktCYxLgdtFR2fDw7J/we-re-talking-past-our-models-or-how-a-model-defined-its --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 20, 2026

Embed this episode

NOW PLAYING

“We’re talking past our models; or, How a model defined its “evil” vector as dread” by jcksanderson

0:00 19:26

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 19 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 20, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!