EPISODE · Sep 8, 2026 · 36 MIN
God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques - By Scott Alexander
from AI Article Readings · host Askwho Casts AI
Scott Alexander works through the tools researchers use to investigate what happens inside language models, from linear probes and sparse autoencoders to activation verbalizers and the Jacobian lens. A technical tour of how these methods work, what they reveal, and the questions they leave open.* 00:00 Introduction* 04:03 Linear Probes* 11:46 Sparse Autoencoders* 14:56 Activation Verbalizers* 18:55 Natural Language Autoencoders* 19:28 Emotion Vectors* 26:01 The Jacobian Lens* 32:19 You Go To War With The Weapons You Have, Not The Weapons You Want* 35:27 Outrohttps://open.substack.com/pub/astralcodexten/p/god-help-us-lets-try-to-learn-about?r=67y1h&utm_campaign=post-expanded-share&utm_medium=web Get full access to Askwho Casts AI at askwhocastsai.substack.com/subscribe
Embed this episode
Ready to play
God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques - By Scott Alexander
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.