[HUMAN VOICE] "Towards Monosemanticity: Decomposing Language Models With Dictionary Learning" by Zac Hatfield-Dodds episode artwork

EPISODE · Nov 9, 2023 · 8 MIN

[HUMAN VOICE] "Towards Monosemanticity: Decomposing Language Models With Dictionary Learning" by Zac Hatfield-Dodds

from LessWrong (Curated & Popular)

Support ongoing human narrations of curated posts:www.patreon.com/LWCuratedThis is a linkpost for https://transformer-circuits.pub/2023/monosemantic-features/Text of post based on our blog post as a linkpost for the full paper which is considerably longer and more detailed.Neural networks are trained on data, not programmed to follow rules. We understand the math of the trained network exactly – each neuron in a neural network performs simple arithmetic – but we don't understand why those mathematical operations result in the behaviors we see. This makes it hard to diagnose failure modes, hard to know how to fix them, and hard to certify that a model is truly safe.Luckily for those of us trying to understand artificial neural networks, we can simultaneously record the activation of every neuron in the network, intervene by silencing or stimulating them, and test the network's response to any possible input.Unfortunately, it turns out that the individual neurons do not have consistent relationships to network behavior. For example, a single neuron in a small language model is active in many unrelated contexts, including: academic citations, English dialogue, HTTP requests, and Korean text. In a classic vision model, a single neuron responds to faces of cats and fronts of cars. The activation of one neuron can mean different things in different contexts.Source:https://www.lesswrong.com/posts/TDqvQFks6TWutJEKu/towards-monosemanticity-decomposing-language-models-withNarrated for LessWrong by Perrin Walker.Share feedback on this narration.[125+ Karma Post] ✓[Curated Post] ✓

Episode metadata supplied by the publisher feed · Published Nov 9, 2023

Embed this episode

Support ongoing human narrations of curated posts: www.patreon.com/LWCurated This is a linkpost for https://transformer-circuits.pub/2023/monosemantic-features/ Text of post based on our blog post as a linkpost for the full paper which is considerably longer and more detailed. Neural networks are trained on data, not programmed to follow rules. We understand the math of the trained network exactly – each neuron in a neural network performs simple arithmetic – but we don't understand why thos...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

[HUMAN VOICE] "Towards Monosemanticity: Decomposing Language Models With Dictionary Learning" by Zac Hatfield-Dodds

0:00 8:02

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 8 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on November 9, 2023.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!