"Against Almost Every Theory of Impact of Interpretability" by Charbel-Raphaël episode artwork

EPISODE · Aug 21, 2023 · 1H 18M

"Against Almost Every Theory of Impact of Interpretability" by Charbel-Raphaël

from LessWrong (Curated & Popular) · host LessWrong

I gave a talk about the different risk models, followed by an interpretability presentation, then I got a problematic question, "I don't understand, what's the point of doing this?" Hum.Feature viz? (left image) Um, it's pretty but is this useful?[1] Is this reliable? GradCam (a pixel attribution technique, like on the above right figure), it's pretty. But I’ve never seen anybody use it in industry.[2] Pixel attribution seems useful, but accuracy remains the king.[3]Induction heads? Ok, we are maybe on track to retro engineer the mechanism of regex in LLMs. Cool.The considerations in the last bullet points are based on feeling and are not real arguments. Furthermore, most mechanistic interpretability isn't even aimed at being useful right now. But in the rest of the post, we'll find out if, in principle, interpretability could be useful. So let's investigate if the Interpretability Emperor has invisible clothes or no clothes at all!Source:https://www.lesswrong.com/posts/LNA8mubrByG7SFacm/against-almost-every-theory-of-impact-of-interpretability-1Narrated for LessWrong by TYPE III AUDIO.Share feedback on this narration.[125+ Karma Post] ✓

Episode metadata supplied by the publisher feed · Published Aug 21, 2023

Embed this episode

I gave a talk about the different risk models, followed by an interpretability presentation, then I got a problematic question, "I don't understand, what's the point of doing this?" Hum. Feature viz? (left image) Um, it's pretty but is this useful?[1] Is this reliable? GradCam (a pixel attribution technique, like on the above right figure), it's pretty. But I’ve never seen anybody use it in industry.[2] Pixel attribution seems useful, but accuracy remains the king.[3]Induction heads? Ok,...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

"Against Almost Every Theory of Impact of Interpretability" by Charbel-Raphaël

0:00 1:18:44

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 1 hour and 18 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on August 21, 2023.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!