“The ‘strong’ feature hypothesis could be wrong” by lsgos episode artwork

EPISODE · Aug 7, 2024 · 30 MIN

“The ‘strong’ feature hypothesis could be wrong” by lsgos

from LessWrong (Curated & Popular)

NB. I am on the Google Deepmind language model interpretability team. But the arguments/views in this post are my own, and shouldn't be read as a team position. “It would be very convenient if the individual neurons of artificial neural networks corresponded to cleanly interpretable features of the input. For example, in an “ideal” ImageNet classifier, each neuron would fire only in the presence of a specific visual feature, such as the color red, a left-facing curve, or a dog snout” : Elhage et. al, Toy Models of SuperpositionRecently, much attention in the field of mechanistic interpretability, which tries to explain the behavior of neural networks in terms of interactions between lower level components, has been focussed on extracting features from the representation space of a model. The predominant methodology for this has used variations on the sparse autoencoder, in a series of papers [...] ---Outline:(09:56) Monosemanticity(19:22) Explicit vs Tacit Representations.(26:27) ConclusionsThe original text contained 12 footnotes which were omitted from this narration. The original text contained 1 image which was described by AI. --- First published: August 2nd, 2024 Source: https://www.lesswrong.com/posts/tojtPCCRpKLSHBdpn/the-strong-feature-hypothesis-could-be-wrong --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Aug 7, 2024

Embed this episode

NB. I am on the Google Deepmind language model interpretability team. But the arguments/views in this post are my own, and shouldn't be read as a team position. “It would be very convenient if the individual neurons of artificial neural networks corresponded to cleanly interpretable features of the input. For example, in an “ideal” ImageNet classifier, each neuron would fire only in the presence of a specific visual feature, such as the color red, a left-facing curve, or a dog snout” : E...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

“The ‘strong’ feature hypothesis could be wrong” by lsgos

0:00 30:16

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 30 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on August 7, 2024.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!