“How robust are natural language autoencoders to initialization?” by michaelzhang, TurnTrout episode artwork

EPISODE · Jul 10, 2026 · 27 MIN

“How robust are natural language autoencoders to initialization?” by michaelzhang, TurnTrout

from LessWrong (30+ Karma)

Natural language autoencoders are meant to take in an LLM's activation vector and describe in plain text what the model is thinking. However, its training data collection involves asking Claude to guess what a model might be thinking. How robust are NLAs to these guesses? We change Claude's guesses in various ways and measure the impact on the NLA's statements as well as on reconstruction accuracy. We show that Qwen2.5-7B NLAs have some robustness to irrelevant statements and prevailing sentiments in Claude's guesses. However, if an NLA is initialized with entirely implausible statements, it can nevertheless achieve nearly the same reconstruction accuracy as plausible-initialized NLAs while emitting 99.3% implausible statements. RL does train implausible-initialized NLAs to be slightly more plausible (increasing from 0.08% to 0.7%). But the plausibility of plausible-initialized NLAs decreases from 21% at initialization to 7.6% at the end of training. If our results scale, they cast doubt on the usefulness of NLAs. Produced as part of the MATS program in the summer 2026 cohort of team shard. Terminology A "plausible" explanation is an objectively true statement about the world. For example, given a passage about greyhounds, a plausible explanation of model [...] ---Outline:(02:16) Introduction(05:06) The experimental setup(06:36) The "Carthago delenda est" experiment(08:15) The "I love Carthage" experiment(11:24) The "confabulation" experiment(20:51) The outputs of plausible-initialized and implausible-initialized NLAs(23:41) Limitations(26:17) Conclusions --- First published: July 10th, 2026 Source: https://www.lesswrong.com/posts/LQXWiF8PyJ5ojNsEv/how-robust-are-natural-language-autoencoders-to --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 10, 2026

Embed this episode

NOW PLAYING

“How robust are natural language autoencoders to initialization?” by michaelzhang, TurnTrout

0:00 27:40

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 27 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 10, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!