“Activation space interpretability may be doomed” by bilalchughtai, Lucius Bushnaq episode artwork

EPISODE · Jan 10, 2025 · 15 MIN

“Activation space interpretability may be doomed” by bilalchughtai, Lucius Bushnaq

from LessWrong (Curated & Popular)

TL;DR: There may be a fundamental problem with interpretability work that attempts to understand neural networks by decomposing their individual activation spaces in isolation: It seems likely to find features of the activations - features that help explain the statistical structure of activation spaces, rather than features of the model - the features the model's own computations make use of.Written at Apollo Research IntroductionClaim: Activation space interpretability is likely to give us features of the activations, not features of the model, and this is a problem.Let's walk through this claim.What do we mean by activation space interpretability? Interpretability work that attempts to understand neural networks by explaining the inputs and outputs of their layers in isolation. In this post, we focus in particular on the problem of decomposing activations, via techniques such as sparse autoencoders (SAEs), PCA, or just by looking at individual neurons. This [...] ---Outline:(00:33) Introduction(02:40) Examples illustrating the general problem(12:29) The general problem(13:26) What can we do about this?The original text contained 11 footnotes which were omitted from this narration. --- First published: January 8th, 2025 Source: https://www.lesswrong.com/posts/gYfpPbww3wQRaxAFD/activation-space-interpretability-may-be-doomed --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Jan 10, 2025

Embed this episode

TL;DR: There may be a fundamental problem with interpretability work that attempts to understand neural networks by decomposing their individual activation spaces in isolation: It seems likely to find features of the activations - features that help explain the statistical structure of activation spaces, rather than features of the model - the features the model's own computations make use of. Written at Apollo Research Introduction Claim: Activation space interpretability is likely to gi...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

“Activation space interpretability may be doomed” by bilalchughtai, Lucius Bushnaq

0:00 15:56

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 15 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on January 10, 2025.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!