EPISODE · Jul 6, 2026 · 50 MIN
“A Review of Anthropic’s Global Workspace Paper” by Neel Nanda
The below is a public review Anthropic asked me to write for their new global workspace paper. I recommend at least skimming their paper first. TLDR: I think this is a fantastic paper - it presents compelling evidence for some kind of "cognitive space" in models, that is used as a "working memory" for intermediate variables during a forward pass, shows that J-Lens is a useful technique for accessing this space. I believe these key claims.I believe J-Lens will be a useful (but limited) tool in practice for model forensics, e.g. generating hypotheses about unusual model behaviour during alignment audits.I discuss my mental models for why a cognitive space should exist, and first principles arguments for why J-Lens should work for accessing itI assess the paper's evidence that this cognitive space exists, and the paper's evidence that J-Lens is practically useful.We have replicated the core claims on Qwen 3.6 27B, and also share preliminary evidence of extending this work by finding abstract "interpretative meta-tokens", like Chinese characters for "what does this mean" that seem to activate and play a causal role on processing ambiguous sentences. What claims is this paper making? In my opinion this [...] ---Outline:(01:27) What claims is this paper making?(05:32) Why does J-Lens work? First principles reasoning(06:03) Why have a "working memory"?(07:59) Worked Example: Factual Recall(09:47) Why have consistent directions for concepts?(10:41) Why are tokens relevant?(11:35) Why are intermediate concepts related to output logits?(14:24) J-Lens is an approximation, but a useful one(16:26) Why Jacobians rather than linear regression?(17:33) What does this working memory actually give us?(18:18) Assessment of evidence for the existence of a cognitive space(19:21) Multihop factual recall(21:20) Other multihop causal interventions(24:17) Further Musings(25:16) Is J-Lens useful?(28:34) Blackmail (5.1)(29:11) Prompt injection (5.2)(29:56) Monitoring for hidden deception (5.3)(30:24) Emergent misalignment (5.4)(30:41) Reward model appeasing (5.5)(31:23) Measuring eval awareness (A.21)(32:17) Equipping an automated auditing agent with J-Lens (A.22)(33:15) Replicating J-Lens and Interpretative Meta-Tokens(34:02) Replication(37:50) Cost and Difficulty of Replicating J-Lens(39:42) Case Study: Interpretative Meta-Tokens(41:49) Where do interpretative meta-tokens appear?(43:52) Are interpretative meta-tokens causal?(49:04) Implications --- First published: July 6th, 2026 Source: https://www.lesswrong.com/posts/zFJ3ZdQwrTWE9jT5S/a-review-of-anthropic-s-global-workspace-paper --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Embed this episode
NOW PLAYING
“A Review of Anthropic’s Global Workspace Paper” by Neel Nanda
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.