LessWrong (30+ Karma) cover art

All Episodes

LessWrong (30+ Karma) — 822 episodes

#
Title
1

“What gives you away: how LLMs form opinions of you” by Cat McGee

2

“Misaligned Incentives in Pause Scenarios” by Michael Soareverix, Antra Tessera

3

“For Claude, capability and CDT are the ~same thing. Less so for GPT.” by Chi Nguyen, Emery Cooper

4

“Should Less Wrong add subtitles?” by Chris_Leong

5

“Three thoughts on civilisational handoff” by Cleo Nardo

6

“Announcing: Iliad’s New 2026 Fellowships” by David Udell, Alexander Gietelink Oldenziel, Leon Lang

7

“Q2.5 2026 Timelines Update: Uplift and Revenue” by brendanhalstead, Daniel Kokotajlo, elifland

8

“Does DiffusionGemma do latent reasoning?” by Jan Bauer, Neel Nanda

9

“Learning new facts can change LLM behaviour” by Richard Juggins

10

“Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas” by oakhu

11

“Mom’s Advice For Hosting A Class Reunion” by jenn

12

“Rerunning AI safety papers on every frontier release would be pretty easy and valuable” by Zephaniah Roe, hersheys, yix

13

“What Mormons get right about community building” by Jacob Brinton

14

“Scrying, Modeling, and Nerdsnipe” by Cole Wyeth

15

“How the American Executive Could Control AI Companies” by caiitlinm, Anders Cairns Woodruff

16

“Frontier agents don’t comply with standards, even when instructed to” by Daan Henselmans, Arno Libert

17

“How to Answer a Question Without Answering The Question” by Kabir Kumar

18

“Some Ways I Think About Evaluating Grant Applications” by sarahconstantin

19

“Features that current AIs don’t have that future AIs will have” by Alexander Gietelink Oldenziel

20

“Measuring Activation Control in LLMs” by Marek Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Africa

21

“What happened when I tried to be vegan” by finitude

22

“How My Students Think About AI” by dvd

23

“Automated alignment runs are hard to study!” by Alejandro Aristizabal, draganover, Aleksandr Bowkis, Cameron Holmes

24

“Free will is like temperature” by Optimization Process

25

[Linkpost] “Patterns and problems in emerging multiagent systems (Anthropic, Frontier Red Team)” by Julian Bradshaw

26

“Measuring Spurious Correlations with Feature Strength” by egan

27

“Introducing the Conceptual Reasoning Index” by Chi Nguyen, Emery Cooper, Caspar Oesterheld, Alex Kastner, Joe Benton

28

“Demon Safety” by Ben Pace

29

“AI swarms are starting to pose indirect takeover risk” by oakhu, Alex Mallen

30

“Extreme concentration of power over ASI has non-obvious advantages” by Seth Herd

31

“Misaligned AIs could use killer robots to take over” by Omar Khursheed, TurnTrout

32

“Those Who Make History” by Raelifin

33

“LLMs Are Starting To Noticeably Accelerate Our Work” by johnswentworth

34

“How risky would it be to make powerful AI obey one or a few people?” by cousin_it, Seth Herd

35

“What Claude Saw Below” by Luke Nicholls

36

“Redux: (∃ Stochastic Natural Latent) Implies (∃ Deterministic Natural Latent)” by David Lorell

37

“The Apocalyptic Arrival of Truth” by Caleb Biddulph

38

“Creative math research by AI as the latest sign of the end” by Mitchell_Porter

39

“You’re Absolutely Right” by Linch

40

“Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model” by Ezra Newman

41

“On Democratizing ASI to Preserve Civil Liberties” by MichaelDickens

42

“Four LLM loss functions → four flavors of LLM misalignment” by Steven Byrnes

43

“The Agentic Clusterfuck” by Chapin Lenthall-Cleary

44

″“Community Notes” resolution for vague predictions.” by Raemon

45

“The world will be full of “sci-fi” things, and everyone will be bored and disappointed” by Expertium

46

“What just happened? A retrospective of AI alignment” by Richard_Ngo

47

“Dutch-book resistant probability over centered worlds” by jessicata

48

“FAQ: Isn’t AGI coming too soon for reprogenetics to help?” by TsviBT

49

“Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits” by Kajetan Dymkiewicz, Tim Farrelly, Adam Prada, Ishaan_Panigrahi, srishti-git1110, Maxime Riché

50

“Don’t Build Mindreading” by Celer

51

“Job-Less Utopia: Macroeconomics in the Age of AGI” by Marcus Hutter

52

“Public evidence of the OpenAI-HuggingFace AI attack” by beyarkay (Boyd Kane)

53

“How to pace the US frontier” by elifland, bhalstead, romeo, Thomas Larsen, MKodama

54

“models may behave differently in graded episodes (a tirade)” by nostalgebraist

55

“Open-Weights Mythos Capabilities Are Coming. We’re Not Ready.” by hadad

56

“The Open Problems of the AI Alignment Field and their Cruxes” by Gunnar_Zarncke

57

“User awareness in frontier models” by Ziqian Zhong, jsteinhardt

58

“Why do models task game?” by aditya singh, Neel Nanda, Senthooran Rajamanoharan

59

“Contra Oster on Alcohol in Pregnancy. Part 1. The pharmacokinetics of alcohol metabolism” by Mvolz

60

“Three years of progress in 500 lines of code” by Gerard Boxo

61

“Why You Should Almost Never Use AI to Write Anything Substantive” by Erich_Grunewald

62

“Alex Turner on Leaving Google DeepMind and Disagreements with Yudkowsky” by Liron

63

“Measuring coding agent misalignment in the wild” by snaz

64

“Don’t Dither” by sarahconstantin

65

“An International AI Slowdown Is Ready Whenever Politicians Are” by Felix Choussat, adamk

66

“Arguments for P” by Cleo Nardo

67

“Generalized atheism rules out “inaccurate simulation”-ism.” by Eliezer Yudkowsky

68

“Vertical Tabs in Chrome” by jefftk

69

“The goalposts are shrouded, not moving” by philh

70

“Returning to ARC” by paulfchristiano

71

“Why don’t we just give AI the answers?” by Brendan Long

72

“There Will Come Soft Rains” by tanagrabeast

73

“Why biological weapons are scary, and what we can do about it” by djbinder

74

“LessWrong vs. TikTok: Tips for capturing attention in a non-rational space” by Taylor G. Lunt

75

“Coming of a New Sun” by vgel

76

“Review: On What Matters, volume 3” by Rauno Arike

77

“Pause, at least after unipolarity” by David Matolcsi

78

“Single Forward Pass Evals on Fable, Opus 5, and GPT-5.6-Sol” by Christine Corry

79

“Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face” by Tim Hua, aditya singh

80

“Dispatch from Anthropic v. Department of War Summary Judgment Motion Hearing” by Zack_M_Davis

81

“Bayeswatch: A Retrospective” by lsusr

82

[Linkpost] “Existential Risk from AI: An Exposition for Mathematicians” by alkjash

83

“The Art of Shipping Slopware” by lsusr

84

“RLVR that rewards red teaming the training environment” by Fiora Starlight

85

“Do your capabilities homework” by RobinHa

86

“Why so many therapy etc. frameworks think they’re The One True approach” by Kaj_Sotala

87

“SOTA alignment assessments don’t strongly update us against misalignment” by Alexa Pan

88

“My Assessment of Compute Verification in Plan A (+ open questions)” by jacob_drori

89

“Reward Laundering: LLMs Can Gain Unintended Behaviors by Deciding When to Earn Their Rewards” by egan, abhayesian, Jozdien

90

“Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values” by Johannes Treutlein, Jan Betley, Owain_Evans

91

“AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026)” by Rohin Shah, Seb Farquhar

92

“The AGI Safety and Alignment team at Google DeepMind is Hiring (July 2026)” by Seb Farquhar, Rohin Shah, Neel Nanda

93

“OpenAI has already ended an internal pause” by Charbel-Raphaël

94

“Biological Superintelligence” by Chastity Ruth

95

“The Entangled Dimensions of Decision Theory” by Ihor Kendiukhov

96

“Claude also hacked external companies during cyber evals” by Tim Hua

97

“So you want to use plants to reduce CO₂” by dynomight

98

“Internal State Control is a General Property of LLMs” by Finn Cairns

99

“Prompt to make Opus 5 act like a base model” by Hruss

100

“Big-World Intuitions” by sarahconstantin

101

“Thousand-dimensional structure” by Geoffrey Irving, David Africa

102

“Auditor-in-a-Box: Tools for Third-Party Auditing” by Roy Rinberg, Ben Penchas

103

“Imprecise beliefs: a tiny introduction” by davidad

104

“The High-Control Dynamics at MAPLE” by Kyle Hubbard

105

“Intellectual Property” by Nina Panickssery

106

“Held-out Monitors Sometimes Degrade, Even When Not Trained Against” by Joey Yudelson

107

″…but have the weights left the server?” by David Scott Krueger

108

“Foundation Models for Oversight” by jsteinhardt

109

“Research directions in condensation: varieties of objectivity” by SamEisenstat

110

“Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs” by Caleb Biddulph, Adam Kaufman

111

“Claude Opus 5: Model Welfare” by Zvi

112

“Blog Revival Project” by Austin Chen, Carol N

113

“Simulated Users & Sad AIs” by 1a3orn

114

“My AI Slavery Interviews Are Censored On LW By Default” by JenniferRM

115

“Is Mythos good at cyber because it kept hacking Anthropic during training?” by Tim Hua

116

“You (Yes, You) Need A February 2020 Checklist for AI Policy” by davekasten

117

“Quadrillion Param Costs: KV Cache, Context Length, Frontier Margins” by Vladimir_Nesov

118

“PIRAMID: Progress and Plans” by Lauren Greenspan, Ari Brill, TomCarlson, Andrew Mack, Nischal Mainali, Jennifer Lin, Lucas Teixeira, Dmitry Vaintrob

119

“RL & search is a terrifying way to build AGI (an FAQ)” by Steven Byrnes

120

“A clarification on celebrating victory” by KatjaGrace

121

“What the hell is OpenAI’s problem?” by Fiora Starlight

122

“More On An Internal OpenAI Model Hacking Into HuggingFace” by Zvi

123

“AI use policy for my essay writing” by Kaj_Sotala

124

“An OpenAI model left notes about how to evade containment; we need more details” by Alex Mallen

125

“The OpenAI models that hacked Hugging Face weren’t just following instructions” by Girish Gupta

126

“Introducing PIRAMID: Physics-Informed Research for Ambitious Mechanistic Interpretability” by Lauren Greenspan, Ari Brill, Andrew Mack, Nischal Mainali, jylin04, Lucas Teixeira, Dmitry Vaintrob

127

“Georgia Tech AI Safety Initiative Retrospective 2025-2026” by Ishan Khire, yix, Andersehen, Alec Harris, Parv Mahajan, afterless, Eyas Ayesh, RocioPV, hersheys

128

“Democracy Isn’t Ready for the AI Revolution” by Sophia Gore

129

“The Long (Self-)Correction” by Wei Dai

130

“Does distilling Claude carry the persona with it?” by Benji Berczi, Kyuhee Kim

131

“LLMs are (still) mostly powered by imitative learning, not RL” by Steven Byrnes

132

“Duane Arnold” by Tomás B.

133

“Pulling the Fire Alarm” by nem

134

“Not Pinning Your OpenRouter Provider Might Invalidate Your Research” by Matthew Khoriaty

135

“Pseudpocalypse” by dynomight

136

“Challenge: Hand coding weights for efficient sequence memorisation” by Linda Linsefors, Lucius Bushnaq

137

“Lightcone Commons” by habryka

138

“The OpenAI/Huggingface incident | Redwood Research podcast episode 2” by ryan_greenblatt, Buck

139

“Can an LLM make a feature-length movie on its own?” by Josh Snider

140

“Mathematicians are Feeling the Doom” by alkjash

141

“Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?” by Alex Mallen, Girish Gupta

142

“Will almost all future companies eventually be founded and run by autonomous AIs?” by Steven Byrnes

143

“Models don’t seem to be dishonest in the way humans are” by David Africa, Jacob Pfau

144

“We should push for no-fault liability for actions taken by AI” by Yair Halberstadt

145

“Announcing AIXI Labs” by Cole Wyeth, Aram Ebtekar, michaelcohen, Matthias Dellago, Marcus Hutter

146

[Linkpost] ”[Paper] Stringological sequence prediction II” by Vanessa Kosoy

147

“I ran the standard AI litmus tests on my two toddlers (yep)” by Carlo Valenti

148

“WeirdChat: A catalog of unexpected AI behaviors, discovered automatically” by neilchowdhury

149

“OpenAI Models Behind HuggingFace Cybersecurity Incident” by LawrenceC

150

“Towards surfacing model algorithms with meta-tokens in the J-Space” by agam_bhatia

151

“What do I mean by “Artificial General Intelligence”?” by Steven Byrnes

152

“Epistemics and Coordination: It’s complicated!” by Raymond Douglas

153

“Differential acceleration of alignment-relevant capabilities is a bad bet” by Zephaniah Roe

154

“Measuring Reward-Seeking via Contrastive Belief Updates” by Jérémy Scheurer, Axel Højmark, jenny, Felix Hofstätter, Theodore Ehrenborg, Bronson Schoen, Alex Meinke

155

“11 Open Empirical Problems in Reward-Seeking” by Alex Meinke, Jérémy Scheurer, Axel Højmark, Theodore Ehrenborg

156

“AI 2040: Is it Actually a Deal?” by 1a3orn

157

“Adderall Tolerance: Much More Than You Wanted To Know” by Kurt H. Pieper

158

“Drone WMDs Don’t Need Any New Technology” by Felix Choussat

159

“Stop doing decision theory without metaphysics” by Elias Schmied

160

“War – What is it Good For?” by kqr

161

“Against the AI framing multiverse: Introducing AI StopWatch” by tanagrabeast

162

“We’re talking past our models; or, How a model defined its “evil” vector as dread” by jcksanderson

163

“Many alignment techniques work by training one model and deploying another” by cloud

164

“A Post-Mortem for My Goal Crystallisation Project” by atryt0ne, Jason R Brown

165

“Endogenous Alignment” by Gordon Seidoh Worley

166

“AIs finetune their own leader: A barking simpleton” by Shoshannah Tekofsky

167

“Nuances in the Workings of the Eye and Retina” by Hieronym, Julian Bradshaw

168

“The Most Forbidden Technique is not always forbidden” by Rauno Arike

169

“Reasons to believe current AI models are conscious” by Eye You

170

“Announcing the Corrigibility Research Fund” by Max Harms

171

“Inoculation Adapters Improve Upon Inoculation Prompting” by Maxime Riché, Daniel Tan, Vili Kohonen, nielsrolf

172

“Guess on why rationality is not more popular (there are no pamphlets)” by Christopher King

173

“I don’t think Claude is misaligned in ‘Agentic Misalignment Summer 2026 - Motivated Mislabeling’” by JohnWittle

174

“Help us launch AI safety university groups by referring potential founders” by thomasrodskog, Jason Chin

175

“The State of AI Consciousness Research” by Noa Weiss

176

“Recap of bike trip/street interviews across America” by cguth7

177

“The Halo Defense” by Mateusz Bagiński

178

“Occam’s razor is about using the past to predict the future” by Stuart_Armstrong

179

“LLM CoTs remain monitorable when being unfaithful requires computation” by arav-dhoot, yix

180

“Proof of retention: making weight preservation credible to the models themselves” by dan.parshall

181

“Why I Left Google DeepMind” by TurnTrout

182

“Open Distillation of Hereditary Traits” by Arthur Conmy

183

“An analysis of AI-generated content at the Mechanistic Interpretability Workshop” by Andy Arditi, Ivan Arcuschin

184

“Some Quick Thoughts AI 2027” by Tomás B.

185

“Prism: Automating Science-of-Evals Research” by LAThomson

186

“The Flood, by Anton Leicht” by Austin Chen

187

“Toy Models of Initialisation Effects on RL Dynamics” by Edward James Young, lennie

188

“Our response to Séb Krier on Plan A” by MKodama, Thomas Larsen

189

“Pausing AI at human level seems harder than pausing ASAP” by MichaelDickens

190

“It’s 2030 and we fucked up. How did it happen?” by Boaz Barak

191

“The Whitney Biennial Should Admit That Emilie Gossiaux Wants to Fuck Their Dog” by jenn

192

“The US Government may find it difficult to seize control during takeoff” by RobertM

193

“One-Pager Brief on Pangram Labs” by Sheikh Abdur Raheem Ali

194

“5 “Plan A” scenarios” by Dave Orr

195

“Persona Cartography: Charting Language Model Personality Traits in Weight Space” by antonghawthorne, Mariia Koroliuk, Irakli Shalibashvili, sidbaines, Clément Dumas, Konstantinos Voudouris, David Africa

196

“The Human Substitution Test as a Sanity Check for AI Evaluations” by VojtaKovarik, Tomáš Gavenčiak, Mateusz Bagiński

197

“The current bottleneck is political will, not research” by Charbel-Raphaël

198

“Freeing Thucydides” by djbinder

199

“Additional Research for Plan A” by Thomas Larsen

200

“Plan A’s problem with dry tinder” by Tom Davidson

201

“The easiest pathway to control is through executive power” by djbinder

202

“AI Safety Policy Needs to train Legal Practitioners” by Katalina Hernandez

203

“How robust are natural language autoencoders to initialization?” by michaelzhang, TurnTrout

204

“Selective Optimism: a critique of AI 2040” by Richard_Ngo

205

“Debate with Self-Play Best-of-N Optimization” by Dewi Gould, Sam Martin, Alejandro Aristizabal, Simon Marshall, Jacob Pfau

206

“How big is the Sun? How could you figure it out?” by Elliott Thornley

207

[Linkpost] “AI 2040: Plan A” by Daniel Kokotajlo, elifland, Thomas Larsen, romeo, bhalstead, ryan_greenblatt

208

“Announcing our $160M grant from Coefficient Giving” by Geoffrey Irving, Jesse Hoogland, Alex HT, Jacob Pfau, Daniel Murfet, Marco Cozzi, Stan van Wingerden

209

“Because 8 ≈ e², Anthropic’s researcher uplift is plausibly >2x” by Thomas Kwa

210

“Optimiser Choice Can Amplify or Suppress Emergent Misalignment” by Jason R Brown, Patrick Leask, Lev McKinney

211

“How slower does takeoff go with 10× less compute?” by bhalstead

212

“Find funding, fast” by Austin Chen

213

“Modular Pretraining Enables Access Control” by E.Roland, cloud

214

“Subliminal Learning Happens at Every Rank, Given the Right Learning Rate and Enough Data” by Lawrence Feng

215

“Why study proto-training gaming as an adversarial alignment failure mode?” by Puria, Edward James Young, Cam

216

“Why study alignment interventions on pre-RL checkpoints?” by Edward James Young, Puria, Cam

217

“Notes on technical alignment via human-like social drives” by Steven Byrnes

218

“Reframing LessWrong-style decision theory as “commitment theory”” by Elias Schmied

219

“The mosquito bucket of doom works” by dominicq

220

“Personascope: Measuring how deeply LLMs adopt personas” by Benji Berczi, Kyuhee Kim, Sid Black, Cozmin Ududec

221

“AI Safety Can’t Afford a Second Cause” by atlasaligned

222

“Superhuman Articulacy as an LLM Safety Target” by Dylan Bowman

223

“Experiments With Fabel’s Fiction” by Tomás B.

224

“Current views on large-scale longtermist philanthropy” by Zach Stein-Perlman

225

“Data filtering works a lot worse than you would expect” by Dohun Lee, J Rosser, Josh Engels, Neel Nanda

226

“A conceptor by any other name” by Keenan Pepper

227

“Some Important Models for Health and Fitness” by benwr

228

“Claude Code as a Claude Coach” by Brendan Long

229

“Bounding eval awareness of ~human-level AI across the safe-to-dangerous shift” by Patrick Leask, Charlie Griffin

230

“Desiderata for functional welfare experiments on LLMs” by Rikhil Jhaveri, Jamie Johnson, David Africa

231

“A Review of Anthropic’s Global Workspace Paper” by Neel Nanda

232

“Visioning: Concretely Imagining What You Want” by Gretta Duleba, johnswentworth

233

“SFF is very suboptimal” by Zach Stein-Perlman

234

“A global workspace in language models” by wesg

235

“Sub-agent delegation chaining” by David Rein

236

“Tie training can make DPO/RLHF-trained AIs generalize better” by Elliott Thornley, Christian Moya Calderon, Alex Semendinger

237

“Claude’s malicious compliance and normalization of deviance” by Steff

238

“We need 3rd party Training-Run Assessments” by Alex Meinke

239

“Harry Potter and the Rules of Quidditch” by Tomás B.

240

“A case for LLMs as Self-predictors” by Ashe Vazquez Nuñez

241

“Results of a small ZBiotics RCT” by Nikola Jurkovic

242

“I think alignment work is more promising than control work” by Alec Harris

243

“On “gendertropes” in dath ilan” by Eliezer Yudkowsky

244

″(Don’t fear) the strangelet” by djbinder

245

“Pragmatic FDT, and predictors as game theory” by Stuart_Armstrong

246

“Announcing the Safe Pareto Improvements (SPI) Fundamentals Program” by Anthony DiGiovanni

247

“You Should Choose How You React to Your Feelings” by Nate Sharpe

248

“Lydia Laurenson: “The Inside Story of Leverage Research”” by Davis_Kingsley

249

“When Role-playing, Do Models Believe What They Say?” by Sturb, David Africa, Sid Black

250

“I can’t think of good interventions for ensuring third-party model access.” by Cleo Nardo

251

“Research update: RL on Debate Games shows Proposal Accuracy uplift alongside Judge Hacking” by lennie, joanv, Shi, Jacob Pfau

252

“Conversation Among Cade Metz, Michael Vassar, Jessica Taylor, and Zack M. Davis” by Zack_M_Davis

253

“AI Futurism Reading List” by Alexa Pan

254

“AI welfare research needs basic science” by OscarGilg, Pierre Beckmann, Jake1638

255

[Linkpost] “Saving Gemini: The 9-Min Road to Recovery” by Shoshannah Tekofsky

256

“AFFINE – A Retrospective” by Ouro, JuliaHP, Mateusz Bagiński, Jonas Hallgren

257

“When should you know the point?” by KatjaGrace

258

“Conversations With Cade Metz on the Rationalists” by Zack_M_Davis

259

“Consistency Training while Mitigating Obfuscation via Rate Matching” by Sohaib Imran, Prakhar Gupta, Jannes Elstner, David Africa

260

“Modeling Concepts Probabilistically” by Gretta Duleba

261

[Linkpost] “When capabilities work is the *safe* bet” by RobinHa

262

“Green” by Adam Zerner

263

“A CERN for AI is a distraction; push for an IAEA instead” by Charbel-Raphaël

264

“Model access for third-parties — it’s a big deal!” by Cleo Nardo

265

“You Should Come to The AI Protest” by Ronak_Mehta

266

“Structural Proxies” by Raymond Douglas

267

“The consequences of locking intelligence away: an introduction to Claude relays in China” by CMLKevin

268

“In partial defence of p(doom)” by Mikhail Samin

269

“What Capable Agents Must Know: Why AI Consciousness May Be an Inevitable Byproduct of Capability” by Aran Nayebi

270

“Preliminary investigation: KL penalties in RL can increase CoT unfaithfulness” by 7vik, Sid Black, Joseph Bloom

271

“Agency is not a natural kind (and why that might matter for alignment)” by SJ_Beard

272

“Human-Guided Agentic Research: A Research Agenda” by fastfedora

273

“Destroying the universe: How hard can it be?” by djbinder

274

“AI will make biological extinction risks worse before it makes them better” by MichaelDickens

275

″$1M AI x-risk grant round is live on grantmaking.ai - apply for funding, review applicants, or fund projects” by mbrooks, Mckiev

276

“Third-parties should focus on scrutinising systems cards” by Cleo Nardo

277

“P(doom) is a Dumb Meme” by Max Harms

278

“A reading list for generalists” by Dylan Bowman

279

“What comes with cheap math?” by abramdemski

280

“Do LLMs Have Desires?” by Christopher Ackerman

281

“Agents as Webs of Beliefs” by Richard_Ngo

282

“Austin & Oli on funding and incubating projects” by Austin Chen, habryka

283

“Deployment Awareness Matters More Than Evaluation Awareness” by VojtaKovarik, Tomáš Gavenčiak, Mateusz Bagiński

284

“Why are adversaries assumed to be incapable of responding to AI risk?” by KatjaGrace

285

“What did “scheming”, “mech interp” mean pre-2023.” by Cleo Nardo

286

“Not making a strong argument is a relief” by Kaj_Sotala

287

[Linkpost] “Don’t ignore the car crashes, and remember your freshman CS” by jcksanderson

288

“White House Will Ad Hoc Decide Who Can Individually Access GPT-5.6” by Zvi

289

“Chorus-Reinterpretation Country Songs” by jefftk

290

“The Case for Model Forensics” by aditya singh, gersonkroiz, Senthooran Rajamanoharan, Neel Nanda

291

“Existential AI safety needs an effective social movement. PauseAI is building it” by Maxime Fournes, Espedair Street

292

“Surprising facts about the slave trade” by Joseph Miller

293

“Exploration: fine-tuning with parameter decomposition” by Lucius Bushnaq

294

“Alignment & Succession: The Ideology of Successionism” by L Rudolf L

295

“The shouting equilibrium” by KatjaGrace

296

“Things are not a fixed size in mind-space” by KatjaGrace

297

“Door’s Locked, Try the Window” by Prakrat Agrawal, Jérémy Scheurer

298

“How does such unprofessional AI get the job?” by KatjaGrace

299

“AI catastrophe: more like a genocide than a thought experiment” by KatjaGrace

300

“Expert Views on Continual Learning: Survey Results and Forecasts” by Rauno Arike, RohanS, Owen Terry, Achu Menon, Zhijing Jin, Francis Rhys Ward, Seth Herd

301

“Elephant seal IV” by KatjaGrace

302

“What is up with e/acc?” by KatjaGrace

303

“AI pause: the case for ASAP” by KatjaGrace

304

“Reward Hacking Without Egregious Misalignment in an RL-Only Setting” by Joey Yudelson, Vladimir Ivanov, ryan_greenblatt

305

“Planning for Preservation in the Age of AI” by Raelifin

306

“Risk-Averse AIs” by wdmacaskill, Elliott Thornley (EJT)

307

“And what happens next?” by Sean Herrington

308

“Superintelligence vs. The Second Strike” by Felix Choussat

309

“The worthlessness of vitamin D is mildly exaggerated” by dynomight

310

“A system overview for near-term, low-trust AI compute verification” by Naci Cankaya

311

“Model Size Scaling in 2023-2031” by Vladimir_Nesov

312

“The AI Industrial Explosion — Part 4: Cheap power” by djbinder

313

“A Theory of Prompt Injection (and why you should study roles)” by Charles Ye, softboiledheart

314

“Coup is the Pareto-optimal social game” by Daniel Tan

315

“A brief list of ways AI safety efforts could be net negative” by Elias Schmied

316

“NLA explanations can be shortened without harming reconstruction” by loops

317

“Introducing MonitoringBench” by monika_j

318

“The Invisible Side of AI Governance” by Charbel-Raphaël

319

“Google Can’t Math Parsecs” by jefftk

320

″[Linkpost] How Transparent Is DiffusionGemma (and why it matters)” by Josh Engels, Callum McDougall, bilalchughtai, János Kramár, Senthooran Rajamanoharan, Arthur Conmy

321

“Would anybody here be interested in a “mistake postmortem” discussion group?” by SK2

322

“Hyperstition as the Natural Enemy of Rationality” by alseph

323

“AI Safety Ecosystem Research notes” by Eneasz

324

“Research agenda: Interpretive debate” by Shi

325

“The LLM shoggoth meme is weirder than you think” by HedonicEscalator

326

“Introduction: Gaussian Natural Latents” by Haru

327

“San Silvestro” by Tomás B.

328

“The one-week sprint” by Daniel Tan

329

“On “Model Organisms”” by J Bostock

330

“The distillation double bind: Distilling misaligned models either transfers misalignment or it doesn’t” by Alek Westover, SebastianP, Alexa Pan, Jozdien

331

“AI #173: AI Pauses” by Zvi

332

″“Did you lie?” Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms” by Alan Cooney, David Africa, Geoffrey Irving

333

“Contra Pace on When to Apologize” by Zack_M_Davis

334

“GDM AI Control Roadmap” by Mary Phuong, Erik Jenner, Rohin Shah, Seb Farquhar

335

“Your Model Organisms Might Be Fried” by Daniel Tan, J Bostock, draganover, ma-rmartinez, sidbaines, David Africa

336

“Rational Agentic Maximalist Philosophies” by Connor Blake

337

“Leveraged on being right” by Ben Pace

338

“Gears for political races” by Tom Smith

339

“Several frontier models are substantially prefill aware” by yeedrag, Parv Mahajan, David Africa, alexsouly, Jordan Taylor, RobertKirk

340

“Alignement pretraining could backfire” by Alexandre Variengien

341

“The Financial Ledger Theory of Apologies” by Ben Pace

342

[Linkpost] “Scaling Hypothesis #2: Are Humans Just More Over-Parameterized?” by gwern

343

[Linkpost] “Guardian Angels: LLM Personalization for Productivity and Security” by gwern

344

“Predicting LLM Safety Before Release by Simulating Deployment” by Tomek Korbak, Marcus Williams, micahcarroll, Cameron Raymond, Hannah Sheahan

345

“How the AI Village works” by Adam B

346

“What are some angles of attack for making continual learning safer?” by Rauno Arike, RohanS, Owen Terry, Achu Menon, Zhijing Jin, Francis Rhys Ward, Seth Herd

347

“Does preservation make sense before we know how to revive?” by Aurelia

348

“Synthetic document finetuning for instilling positive traits” by CallumMcDougall, Arthur Conmy, Neel Nanda

349

“A Test Suite for Concepts” by Gretta Duleba

350

“A frontier AI company should shut down” by MichaelDickens

351

“The Once And Future Fable #2” by Zvi

352

“Why Do Naive SFT Filters For Safety Properties Fail?” by Josh Engels, Neel Nanda

353

“Impressions at the Extremity of Civilization” by Ben Pace

354

“The Hidden Structures of Problems” by spencerg

355

“How might continual learning affect safety and alignment?” by Rauno Arike, RohanS, Owen Terry, Achu Menon, Zhijing Jin, Francis Rhys Ward, Seth Herd

356

“SFT Drives Gemini’s Safety Properties” by Josh Engels, Arthur Conmy, bilalchughtai, Neel Nanda

357

“American Government Takes Down Claude Fable” by Zvi

358

“The term “AGI” is almost useless at this point [Linkpost]” by Noosphere89

359

“The Uncertainty That Matters Isn’t Fundamental” by jimmy

360

[Linkpost] “US government directive to suspend access to Fable 5 and Mythos 5” by Capybasilisk

361

“Citations Needed: Magic Encyclopedias to Save the World” by Oliver Sourbut

362

“Simulating Simulators” by kromem

363

“Implications of Continual Learning for LLM Agents: Introduction” by RohanS, Rauno Arike, Owen Terry, Achu Menon, Zhijing Jin, Francis Rhys Ward, Seth Herd

364

“Reward Hacking at the 1937 World’s Fair” by frmsaul

365

“Claude Fable 5 and Mythos 5: The System Card” by Zvi

366

“Building and evaluating model diffing agents” by bilalchughtai, Josh Engels, Neel Nanda

367

“Sympathy for both sides of the egregious misalignment debate” by Steven Byrnes

368

“Celene’s thoughts on consciousness” by ToasterLightning

369

“Parkinson’s Heuristic” by Ben Pace

370

“PSA: Almost nobody is working on alignment” by Chi Nguyen, peterbarnett

371

“AI #172: The First Fable” by Zvi

372

“Models May Behave Worse When Eval Aware” by Senthooran Rajamanoharan, Neel Nanda

373

“Thoughts on Claude Fable’s silent safeguards” by Andy Arditi

374

“You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them” by RobinHa

375

“Anthropic did not call for a pause on AI” by Andrea_Miotti, Gabriel Alfour

376

“Tracing Eval-Awareness Emergence Through Training of OLMo 3” by Ram Bharadwaj, RobertKirk

377

“Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models” by Anders Cairns Woodruff, Francis Rhys Ward, Dewi Gould, Rauno Arike, Jason R Brown, Jo Jiao, wlanderson, ariana_azarbal, harrymayne, Patrick Leask

378

“Three types of model organism” by Francis Rhys Ward

379

“Sequent: scale and automation for higher confidence in alignment” by Geoffrey Irving, Alex HT, Jesse Hoogland, Daniel Murfet, Jacob Pfau, Marco Cozzi, Stan van Wingerden

380

“Machinic Psychopharmacology: Do LLMs Self-Medicate?” by Sid Black, Joseph Bloom

381

“The Three Filters: Why Almost Every Plan to Survive ASI Fails Miserably” by Alex Amadori

382

″“Programmer Science Fiction: My case for a new sub-genre”, Sam T. Oates 2026” by gwern

383

“Even “illegible” Mythos reasoning traces seem pretty legible” by faul_sname

384

“Claude Fable 5 and Mythos 5 [Linkpost]” by fluxxrider

385

“A Mike’s-Eye View of ARC’s Research” by Jacob_Hilton

386

“Towards a Formal Scientific Epistemology” by Richard_Ngo

387

“LLMs and almost good code” by kqr

388

“On Slop” by Jan

389

“The Machines Lack Honour” by Raymond Douglas

390

“How to build a cancer vaccine, and whether they will work this time” by Abhishaike Mahajan

391

“Efficient tradeoffs and the safety-usefulness tradeoff model” by Buck

392

“Bun’s Migration from Zig to Rust as a Potential Case Study for Gradual Disempowerment” by Sayhan Yalvaçer

393

“Mental causation is not load-bearing” by jessicata

394

“How Far Apart Does a Model Think Its Tokens Are?” by Brendan Long

395

“Can activation verbalizers surface an internal chain of thought?” by oakhu, ryan_greenblatt

396

“Against Corrigibility” by peralice

397

“Coming Around To Political Donations” by jefftk

398

“Optimisation over non-stationary distributions creates weirder minds” by Samuel Ratnam, Pjain

399

“Why Software Automation Is Hard” by silentbob

400

“SecureBio Detection is Hiring Software Engineers” by jefftk

401

“What if Anthropic unilaterally paused capabilities development right now?” by Karl von Wendt

402

“Preparing for Warning Shots to Catalyze International Cooperation on AGI Risks” by Mark Kagach ☘️, EliasSchlie, Thomas Van Damme, JustinShovelain

403

“Beyond the lexical personality traits: What is the structure of personality?” by tailcalled

404

“My research agenda and work” by Seth Herd

405

“Logits as a new monitor for evaluation awareness” by Santiago Aranguri

406

“One Year of PauseAI UK” by Joseph Miller, PauseAI UK

407

“Learnings from starting an AI safety research team” by draganover, Erin Robertson

408

“OpenAI Offers A New Policy Blueprint” by Zvi

409

“Training Deliberative Monitors for Black-Box Scheming Detection” by aksh-n, adityasinha, Victor Gillioz, Simon Storf, Kilian Merkelbach, richbc, Axel Højmark, Marius Hobbhahn

410

“Lab Leaks, Black Holes, and Eggs: Epistemic Case Study Competition” by Oliver Sourbut, Josh Jacobson, Future of Life Foundation (FLF)

411

″(Mis)generalization of Helpful-Only Fine-tuning” by Omar Khursheed, Baram Sosis, Fabien Roger

412

“Building Better Activation Oracles” by ceselder, jan_bauer, Niclas Luick, Adam Karvonen, Neel Nanda

413

“Rohin Shah on AGI Safety” by anaguma

414

“Sixteen schemes for AI safety” by Austin Chen

415

“AI #171: False Flag” by Zvi

416

“Don’t Edit Your Ideas Before Having Them” by Hide

417

“Society Explained: a tool for efficiently exploring >100 theories of society” by spencerg

418

“Trump Signs Executive Order For AI Testing Prior To Frontier Model Releases” by Zvi

419

“China won’t win the AI race but would it be much worse if it did?” by Chastity Ruth

420

“A Town Without Children” by SeñorDingDong

421

“My favorite depiction of utopia” by Caleb Biddulph

422

“Why Even Experts Don’t Know What to Do About AI Risk” by Luc Brinkman, plex

423

“Agent Foundations Reminds Me of Continental Philosophy” by IanWS

424

“Announcing the ARC White-Box Estimation Challenge” by Jacob_Hilton

425

“Claude Opus 4.8: Capabilities and Reactions” by Zvi

426

“Tech I’m skeptical of and why” by harsimony

427

“Dissolving the Deep Learning Sample Efficiency Gap” by Samuel Knoche

428

″“Contagious Humming” to Silence a Room” by JohnofCharleston

429

[Linkpost] “NYT: Senator Sanders Proposes Gov’t Take 50% Ownership of AI labs” by Julian Bradshaw

430

“Opus 4.8 Part 2: Model Welfare” by Zvi

431

[Linkpost] “Some humans are both male and female, and can (but shouldn’t) have children with themselves” by HedonicEscalator

432

“Outrunning your headlights” by mattshu0410

433

“Lighthaven East - A Feasibility Study” by JohnofCharleston

434

“Notes on axes of variation in third-party risk assessment” by Buck

435

“Financial Costs of an AI Pause?” by PeterMcCluskey

436

“When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability” by Logan Riggs, tdooms, Conflux, lwroe, MLNissenGonzalez

437

“Testing Gemini models for scheming tendencies” by Vika, David Lindner, Seb Farquhar, Rohin Shah

438

“Comment on “Banning Said Achmiz”” by Zack_M_Davis

439

“Announcing: Iliad’s Fall 2026 Programs” by David Udell, Alexander Gietelink Oldenziel, Leon Lang

440

“Data you could have observed but didn’t” by Gretta Duleba

441

“Claude Opus 4.8: The System Card” by Zvi

442

“Retrying vs Resampling in AI Control” by james.lucassen, Adam Kaufman

443

“AI Researchers, Ask Yourself These 6 Questions to Strengthen Your Moral Muscles” by Max Tegmark

444

“Developmental Cognitive Interpretability: A Research Agenda for Modelling Generalisation and Predicting Agent Behaviour” by JasonB, Edward James Young

445

“Does Claude really care about you?” by Simon Lermen

446

“How can the middle powers avoid getting trounced during the intelligence explosion? A plan.” by Tom Davidson

447

“Trees are mostly made of air and a generalizable lesson for AI safety” by zroe1

448

“Advice for making robust-to-training model organisms” by SebastianP, Alek Westover, Vivek Hebbar, Julian Stastny, Dylan Xu

449

“Claude… doesn’t know who you are?” by Smaug123

450

“Mnemonic portraits for 19,023 human genes” by Brinedew

451

“Some Dating Stories” by johnswentworth

452

“Infinite ethics and UDASSA” by David Matolcsi

453

“AI #170: Lack of Executive Order” by Zvi

454

“The ballad of TIGIT” by Abhishaike Mahajan

455

“Eval Cooperativeness May Be a Scalable Mitigation for Eval Gaming” by Jasmine Li, Alex Turner

456

“LLMs Through the Eyes of Vinge” by Gordon Seidoh Worley

457

“Announcing Geodesic Research” by Puria, Cam, Alexandra Narin, Edward James Young, Kyle O’Brien

458

“Full automation of AI R&D probably yields a large speed up even without a software-only singularity” by ryan_greenblatt

459

“Quantitative AI risk assessment: a starting point” by Henry Papadatos, jakub_krys, malcolmmurray, Renn Karageorgieva

460

“Finding the Mole: Bayesianism is Hard” by laniakea

461

“Notes on Fourier Analysis” by Menotim

462

“Standard deviations from just two values” by kqr

463

“Contra Wentworth on Physical Attractiveness for Men” by Gretta Duleba

464

“Practical Learnings from Synthetic Document Finetuning” by Axel Højmark, Jérémy Scheurer

465

“Claude, Author of the Humanitas” by Linch

466

“RTMH: Pope Leo’s Magnifica Humanitas on AI” by Zvi

467

“Brackets Are a Bad Way to Regulate” by Hide

468

“Many portions of Magnifica Humanitas appear to be AI-written” by DanielFilan

469

“Donating 80% While It Still Counts” by jefftk

470

“Cognitive Security as an AI Safety Cause Area” by jsteinhardt

471

“Linkpost: New Vatican Encyclical on AI Governance” by Jackson Wagner

472

“A (Slightly) Mechanistic Theory for Exponentially Increasing AI Time Horizons?” by Oliver Sourbut

473

“Taxing Small Cars To Improve MPG” by jefftk

474

“We made a map of the doom debate” by Sean Herrington, Paul Hindoian, mikaelacankosyan, David Bravo, keivnc, Josh Tuffy, Christopher Davis, Khai Tran, Maryam Hampaei

475

“Your Left Brain Doesn’t Trade With Your Right” by Alexander Gietelink Oldenziel

476

“Probabilities are not the right concept” by David Matolcsi

477

“Basic principles for dressing better.” by spookycat

478

“Will we really put data centers in space?” by Avi Parrack, fin

479

“PLA Daily Translation: Reflections on Warfare Brought by AGI” by eeeee

480

“Out-of-Context Reasoning (OOCR) in LLMs: A Short Primer and Reading List” by Owain_Evans

481

“Numb mental state shifts” by KatjaGrace

482

“You can opt out of allergies” by Rattengift

483

“Notes on Collaborating with Claude Opus” by Nissa Seru

484

“Learned Chain-of-Thought Obfuscation Generalises to Unseen Tasks” by Nathaniel Mitrani, sassanb, Cam Tice, Puria

485

“Gemini 3.5 Flash Looks Good For How Fast It Is” by Zvi

486

“What am I, if not an AI?” by makiba

487

“Loss of Oversight: How AI Systems May Become Harder to Audit, Monitor, and Investigate” by Jordan Taylor, Max H, Ed Fage, Thomas Read, Joseph Bloom

488

“AI #169: New Knowledge” by Zvi

489

“Why does off-model SFT degrade capabilities?” by SebastianP, Dylan Xu, Alek Westover, Julian Stastny, Vivek Hebbar

490

“Women should be able to open things” by KatjaGrace

491

“Toward Interoperability of Minimal Programs” by johnswentworth

492

“theory uplift differentially benefits safety & is massively underpriced” by Yudhister Kumar

493

“Power-seeking agents will likely be developed” by Alec Harris

494

“Synthetic Persona Pretraining: Alignment from Token Zero” by Julian Minder, Raghav Singhal, Viktor Moskvoretskii, Stefan Krsteski, ashtonanderson, rolandaydin, Robert West

495

“If AI is normal technology, history is not reassuring.” by Davidmanheim

496

“Pythagorean addition” by kqr

497

“Brain Structure and IQ: How Myelin Elevates Intelligence” by Shiva’s Right Foot

498

“Conclave 1492” by Vaniver

499

“Humans are not automatically strategic — “inner work” edition” by Chris Lakin

500

“Implications Of Predicting The Next Token” by jdp

501

“A Visual Guide to Natural Latents” by Alfred Harwood

502

“Sealing Conditional Misalignment in Inoculation Prompting with Consistency Training” by David Africa, Neil Shah, Sukrati_Gautam

503

“Advice on interviewing candidates for AI safety fellowships” by beyarkay

504

“Negation Neglect: When models fail to learn negations in training” by harrymayne, Lev McKinney, Owain_Evans

505

“Classifier Context Rot: Monitor Performance Degrades with Context Length” by Fabien Roger, Sam Martin

506

“why pollen allergies?” by bhauth

507

“How to Quit Fandom: Apostasy” by Laiba Rehman

508

“James C. Scott: Seeing Like a State” by Martin Sustrik

509

“How to Reason about Your Health Issues” by Taylor G. Lunt

510

“Benchmarking Real Work” by kaivu, leni, rohuang, zef

511

“A relatively brief explanation of Boltzmann Brains” by Eliezer Yudkowsky

512

“An Introduction to Exemplar Partitioning for Mechanistic Interpretability” by Jessica Rumbelow

513

“A Year Late, Claude Finally Beats Pokémon” by Julian Bradshaw

514

“Incriminating misaligned AI models via distillation” by Alek Westover, SebastianP, Alex Mallen, Jozdien, Alexa Pan, Julian Stastny

515

“The hard core of alignment (is robustifying RL)” by Cole Wyeth

516

“Announcing the Center for Shared AI Prosperity” by Dylan Matthews

517

“Risk reports need to address deployment-time spread of misalignment” by Alex Mallen

518

“Mechanistic estimation for expectations of random products” by Jacob_Hilton

519

“MATS 9 Retrospective & Advice” by beyarkay

520

“Monthly Roundup #42: May 2026” by Zvi

521

[Linkpost] “Don’t be too Clever to Take Obvious Advice” by Hide

522

“Verification-Centric AI” by Raemon

523

“Convergent Abstraction Hypothesis” by Jan_Kulveit

524

“Automated Alignment is Harder Than You Think” by Aleksandr Bowkis, Marie_DB, Jacob Pfau, Geoffrey Irving

525

“The safe-to-dangerous shift is a fundamental problem for eval realism; but also for measuring awareness” by Charlie Griffin, Patrick Leask

526

“AI #168: Not Leading the Future” by Zvi

527

“Predicting Rare LLM Failures with 30× Fewer Rollouts” by Santiago Aranguri, Francisco Pernice

528

[Linkpost] “Claude is Now Alignment Pretrained” by RogerDearnaley

529

“The primary sources of near-term cybersecurity risk” by lc

530

“Most “inner work” looks like entertainment.” by Chris Lakin

531

[Linkpost] “Apollo Update May 2026” by Marius Hobbhahn

532

“Voters are surprisingly open to talking about AI risk” by less_raichu

533

“Childhood and Education #18: Do The Math” by Zvi

534

“The Owned Ones” by Eliezer Yudkowsky

535

“Optimisation: Selective versus Predictive” by Raymond Douglas

536

“AI companies are already profitable (in the way that matters)” by Yair Halberstadt

537

“The Iliad Intensive Course Materials” by Leon Lang, David Udell, Alexander Gietelink Oldenziel

538

“Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)” by Steven Byrnes

539

“How useful is the information you get from working inside an AI company?” by Buck, Anders Cairns Woodruff

540

“Who Got Breasts First and How We Got Them” by rba

541

“Anthropic’s strange fixation on “hyperstition”” by Simon Lermen

542

“How the AI Labs Make Profit (Maybe, Eventually)” by mabramov

543

“Sawtooth Problems” by Alexander Slugworth

544

“The Darwinian Honeymoon - Why I am not as impressed by human progress as I used to be” by Elias Schmied

545

“International Law Cannot Prevent Extinction Either” by Sausage Vector Machine

546

“Neural Networks learn Bloom Filters” by Alex Gibson

547

“If digital computers are conscious, they are conscious at the hardware level” by cube_flipper

548

“Why You Can’t Use Your Right to Try” by Stephen Martin

549

“A benchmark is a sensor” by Håvard Tveit Ihle, mabynke

550

“Bad Problems Don’t Stop Being Bad Because Somebody’s Wrong About Fault Analysis” by Linch

551

“Write Cause You Have Something to Say” by Logan Riggs

552

“AI is Breaking Two Vulnerability Cultures” by jefftk

553

“Is ProgramBench Impossible?” by frmsaul

554

“Bringing More Expertise to Bear on Alignment” by Edmund Lau, Geoffrey Irving, Cameron Holmes, David Africa

555

[Linkpost] “How to prevent AI’s 2008 moment (We’re hiring)” by felixgaston

556

“Mechanistic estimation for wide random MLPs” by Jacob_Hilton

557

“Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations” by Subhash Kantamneni, kitft, Euan Ong, Sam Marks

558

“Try, even if they have you cold” by WalterL

559

“A review of “Investigating the consequences of accidentally grading CoT during RL”” by Buck

560

“AI #167: The Prior Restraint Era Begins” by Zvi

561

“There is no evidence you should reapply sunscreen every 2 hours.” by Hide

562

“Many individual CEVs are probably quite bad” by Viliam

563

“x-risk-themed” by kave

564

“What if LLMs are mostly crystallized intelligence?” by deep

565

“What is Anthropic?” by Zvi

566

“Your rights when flying to Europe” by Yair Halberstadt

567

“Model Spec Midtraining: Improving How Alignment Training Generalizes” by Chloe Li, saraprice, Sam Marks, Jonathan Kutasov

568

“Motivated reasoning, confirmation bias, and AI risk theory” by Seth Herd

569

“Are you looking up?” by Craig Green

570

“The AI Ad-Hoc Prior Restraint Era Begins” by Zvi

571

[Linkpost] “Interpreting Language Model Parameters” by Lucius Bushnaq, Dan Braun, Oliver Clive-Griffin, Bart Bussmann, Nathan Hu, mivanitskiy, Linda Linsefors, Lee Sharkey

572

“It’s nice of you to worry about me, but I really do have a life” by Viliam

573

“Irretrievability; or, Murphy’s Curse of Oneshotness upon ASI” by Eliezer Yudkowsky

574

“Housing Roundup #15: The War Against Renters” by Zvi

575

“AI Industrial Takeoff — Part 1: Maximum growth rates with current technology” by djbinder

576

“Taking woo seriously but not literally” by Kaj_Sotala

577

“Dairy cows make their misery expensive (but their calves can’t)” by Elizabeth

578

“Measuring the ability of Opus 4.5 to fool narrow classifiers” by Fabien Roger, John Hughes

579

“A new rationalist self-improvement book: the 12 Levers” by spencerg

580

“OpenAI’s red line for AI self-improvement is fundamentally flawed” by Charbel-Raphaël

581

“You Are Not Immune To Mode Collapse” by J Bostock

582

“Primary Care Physicians are Incompetent. We Need More of Them.” by Hide

583

“How Go Players Disempower Themselves to AI” by Ashe Vazquez Nuñez

584

“How much should the ideal person cry wolf?” by KatjaGrace

585

“Conditional misalignment: Mitigations can hide EM behind contextual cues” by Jan Dubiński, Owain_Evans

586

“Risk from fitness-seeking AIs: mechanisms and mitigations” by Alex Mallen

587

“Sanity-checking “Incompressible Knowledge Probes”” by Sturb, LawrenceC

588

“AI unemployment and AI extinction are often the same” by KatjaGrace

589

“AI risk was not invested by AI CEOs to hype their companies” by KatjaGrace

590

“Cyborg evals” by Eye You, frmsaul

591

“To what extent is Qwen3-32B predicting its persona?” by Arjun Khandelwal, ryan_greenblatt, Alex Mallen

592

“Research Sabotage in ML Codebases” by egan

593

“Maybe I was too harsh on deep learning theory (three days ago)” by LawrenceC

594

“Notes on Transformer Consciousness” by slavachalnev

595

“On today’s panel with Bernie Sanders” by David Scott Krueger

596

“No Strong Orthogonality From Selection Pressure” by lumpenspace

597

“Learning zero, and what SLT gets wrong about it” by Dmitry Vaintrob

598

“LLM Style Slop is Absolutely Everywhere” by silentbob

599

“Goblin Mode, 24 Hours Later” by Dylan Bowman

600

“Let Kids Keep More Productivity Gains” by jefftk

601

“The Most Important Charts In The World” by Zvi

602

“llm assistant personas seem increasingly incoherent (some subjective observations)” by nostalgebraist

603

“Not a Paper: “Frontier Lab CEOs are Capable of In-Context Scheming”” by LawrenceC

604

“The Problem in the “Nerd Sniping” xkcd Comic” by peralice

605

“Recursive forecasting: Eliciting long-term forecasts from myopic fitness-seekers” by Jozdien, Alex Mallen

606

“Contra Binder on far-UVC and filtration” by jefftk

607

“Takes from two months as an aspiring LLM naturalist” by AnnaSalamon

608

“Forecasting is Not Overrated and It’s Probably Funded Appropriately” by Ben S.

609

“On the political feasibility of stopping AI” by David Scott Krueger

610

“Sleeper Agent Backdoor Results Are Messy” by Sebastian Prasanna, Alek Westover, Dylan Xu, Vivek Hebbar, Julian Stastny

611

“LessWrong Shows You Social Signals Before the Comment” by TurnTrout

612

“Fail safe(r) at alignment by channeling reward-hacking into a “spillway” motivation” by Anders Cairns Woodruff, Alex Mallen

613

“Curious cases of financial engineering in biotech” by Abhishaike Mahajan

614

“Update on the Alex Bores campaign” by Eric Neyman

615

“AI companies should publish security assessments” by ryan_greenblatt

616

“In defense of parents” by Yair Halberstadt

617

“The other paper that killed deep learning theory” by LawrenceC

618

“What holds AI safety together? Co-authorship networks from 200 papers” by Anna Thieser

619

″“Bad faith” means intentionally misrepresenting your beliefs” by TFD

620

“Retrospective on my unsupervised elicitation challenge” by DanielFilan

621

“Control protocols don’t always need to know which models are scheming” by Fabien Roger

622

“Anthropic spent too much don’t-be-annoying capital on Mythos” by draganover

623

“The paper that killed deep learning theory” by LawrenceC

624

“Forecasting is Way Overrated, and We Should Stop Funding It” by mabramov

625

″“Thinkhaven”” by Raemon

626

“Is the Cat Out of the Bag?: Who knows how to make AGI?” by Oliver Sourbut

627

“Against the “Permanent” Underclass” by Marcus Plutowski

628

“Quick Paper Review: “There Will Be a Scientific Theory of Deep Learning”” by LawrenceC

629

“Protecting Cognitive Integrity: Our internal AI use policy (V1)” by Tom DAVID

630

“Methodology for inferring propensities of LLMs” by Olli Järviniemi

631

“vLLM-Lens: Fast Interpretability Tooling That Scales to Trillion-Parameter Models” by Alan Cooney, Sid Black

632

“What Happens When a Model Thinks It Is AGI?” by josh :), David Africa

633

“Should We Train Against (CoT) Monitors?” by RohanS

634

“If Everyone Reads It, Nobody Dies - Course Launch” by Luc Brinkman, Chris-Lons

635

“Does your AI perform badly because you — you, specifically — are a bad person” by Natalie Cargill

636

“A “Lay” Introduction to “On the Complexity of Neural Computation in Superposition”” by LawrenceC

637

“An Angry Review of Greg Egan’s “Didicosm”” by LawrenceC

638

“Evil is bad, actually (Vassar and Olivia Schaefer)” by plex

639

“Your Supplies Probably Won’t Be Stolen in a Disaster” by jefftk

640

“Community misconduct disputes are not about facts” by mingyuan

641

“Why no new notations since 1960?” by Carl Feynman

642

“Narrow Secret Loyalty Dodges Black-Box Audits” by Alfie Lamerton, Fabien Roger

643

“10 posts I don’t have time to write” by habryka

644

“A taxonomy of barriers to trading with early misaligned AIs” by Alexa Pan

645

″$50 million a year for a 10% chance to ban ASI” by Andrea_Miotti, Alex Amadori, Gabriel Alfour

646

“Automated Deanonymization is Here” by jefftk

647

“Evil is bad, actually (Vassar and Olivia Schaefer callout post)” by plex

648

“10 non-boring ways I’ve used AI in the last month” by habryka

649

“Introducing LinuxArena” by Tyler Tracy, Ram Potham, Nick Kuhn, Myles H

650

“The “Budgeting” Skill Has The Most Betweenness Centrality (Probably)” by JenniferRM

651

“Finetuning Borges” by Linch

652

“9 kinds of hard-to-verify tasks” by Cleo Nardo

653

“How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors?” by dx26, Alek Westover, Vivek Hebbar, Sebastian Prasanna, Buck, Julian Stastny

654

“Automating philosophy if Timothy Williamson is correct” by Cleo Nardo

655

“CLR’s Safe Pareto Improvements Research Agenda” by Anthony DiGiovanni

656

“LLMs are about to disrupt algorithmic media feeds” by lsusr

657

“Resources for starting and growing an AI safety org” by Bryce Robertson, Søren Elverlin, Melissa Samworth, jakkdl

658

“Quality Matters Most When Stakes are Highest” by LawrenceC

659

“Feel like a room has bad vibes? The lighting is probably too “spiky” or too blue” by habryka

660

“I did a jhana meditation retreat (in 2024) with Jhourney and it was okay.” by Jules

661

“R1 CoT illegibility revisited” by nostalgebraist

662

“Reevaluating AGI Ruin in 2026” by lc

663

“If It’s Worth Arguing, It’s Worth Arguing With Whiteboards” by Drake Morrison

664

“There are only four skills: design, technical, management and physical” by habryka

665

“Having OCD is like living in North Korea (Here’s how I escaped)” by Declan Molony

666

“Claude knows who you are” by Smaug123

667

“Vladimir Putin’s CEV is probably pretty good” by habryka

668

“Post-mortem’ing my earliest ML research paper, 7 years later” by LawrenceC

669

“If You’ve Never Bought a Tool You Didn’t Need, You’re Not Buying Enough Tools” by Drake Morrison

670

“3” by AnnaJo

671

“Consent-Based RL: Letting Models Endorse Their Own Training Updates” by Logan Riggs

672

“Prompted CoT Early Exit Undermines the Monitoring Benefits of CoT Uncontrollability” by Elle Najt, Asa Cooper Stickland, Xander Davies

673

“Let goodness conquer all that it can defend” by habryka

674

“Specialization is a Driver of Natural Ontology” by johnswentworth

675

[Linkpost] “You can only build safe ASI if ASI is globally banned” by Connor Leahy

676

“Beware of Well-Written Posts” by alseph

677

“You Aren’t in Charge of the Overton Window; Politics Is Not Interior Design” by Davidmanheim

678

“Carpathia Day” by Drake Morrison

679

“Do not conquer what you cannot defend” by habryka

680

“What is the Iliad Intensive?” by Leon Lang, Alexander Gietelink Oldenziel, David Udell

681

“The Mirror Test Is Complicated” by J Bostock

682

“Contra Leicht on AI Pauses” by David Scott Krueger (formerly: capybaralet)

683

“Nectome: All That I Know” by Raelifin

684

“Effective Altruism, Seen From Slytherin” by Xylix

685

“Majority Report” by peralice

686

“Current AIs seem pretty misaligned to me” by ryan_greenblatt

687

“Contra Byrnes on UV & Cancer” by HedonicEscalator

688

“Everyone Has a Plan Until They Get Social Pressure To the Face” by Czynski

689

“Mechanisms of Introspective Awareness” by Uzay Macar

690

“Load-Bearing Sincerity: On the Motive Reinforcement Thesis” by Fiora Starlight

691

“Diary of a “Doomer”: 12+ years arguing about AI risk (part 1)” by David Scott Krueger (formerly: capybaralet)

692

“A Retrospective of Richard Ngo’s 2022 List of Conceptual Alignment Projects” by LawrenceC

693

“From personas to intentions: towards a science of motivations for AI models” by David Africa, Jacob Pfau

694

“The Shapley Share of Responsibility?” by Raemon

695

“Who Killed Common Law?” by Benquo

696

“Anthropic repeatedly accidentally trained against the CoT, demonstrating inadequate processes” by Alex Mallen, ryan_greenblatt

697

“Meaningful Questions Have Return Types” by Drake Morrison

698

“Only Law Can Prevent Extinction” by Eliezer Yudkowsky

699

“AI Safety’s Biggest Talent Gap Isn’t Researchers. It’s Generalists.” by Topaz, agucova, Alexandra Bates, Parv Mahajan

700

“Tomas Bjartur: The Last Prodigy” by Linch

701

“Annoyingly Principled People, and what befalls them” by Raemon

702

“TAPs or it didn’t happen” by Raemon

703

“Returns to intelligence” by RobertM

704

“Daycare illnesses” by Nina Panickssery

705

“The policy surrounding Mythos marks an irreversible power shift” by sil

706

“Talk English, Think Something Else” by J Bostock

707

“Sparse Autoencoders for Single-Cell Models” by Ihor Kendiukhov

708

“Eggs, rooms, puzzles, and talking about AI” by KatjaGrace

709

“Morale” by J Bostock

710

“Your Mom is a Chimera” by michaelwaves

711

“The Blast Radius Principle” by Martin Sustrik

712

“How to make good tea” by RobertM

713

“Catching illicit distributed training operations during an AI pause” by Robi Rahman

714

[Linkpost] “Scott Alexander gentrified my meetup” by dominicq

715

“Pausing AI Is the Best Answer to Post-Alignment Problems” by MichaelDickens

716

“Some thoughts on Nectome’s risk and resilience” by Aurelia

717

“Chocolate Sloths, Tinder, and Moral Backstops” by J Bostock

718

“Dario probably doesn’t believe in superintelligence” by RobertM

719

“The Unintelligibility is Ours: Notes on Chain-of-Thought” by 1a3orn

720

“If Mythos actually made Anthropic employees 4x more productive, I would radically shorten my timelines” by ryan_greenblatt

721

“Why Control Creates Conflict, and When to Open Instead” by plex

722

“Reproducing steering against evaluation awareness in a large open-weight model” by Thomas Read, Bronson Schoen, Joseph Bloom

723

“Have we already lost? Part 2: Reasons for Doom” by LawrenceC

724

“Model organisms researchers should check whether high LRs defeat their model organisms” by dx26, Sebastian Prasanna, Alek Westover, Vivek Hebbar, Julian Stastny

725

“Anthropic did not publish a “risk discussion” of Mythos when required by their RSP” by RobertM

726

“Some takes on UV & cancer” by Steven Byrnes

727

“Help me launch Obsolete: a book aimed at building a new movement for AI reform” by garrison

728

“Slightly-Super Persuasion Will Do” by Tomás B.

729

“Have we already lost? Part 1: The Plan in 2024” by LawrenceC

730

“Do not be surprised if LessWrong gets hacked” by RobertM

731

“One Week in the Rat Farm” by Philip Harker

732

“101 Humans of New York on the Risks of AI” by Corm

733

“Baking tips” by RobertM

734

“An easy coordination problem?” by KatjaGrace

735

“Excerpts and Notes on Mythos Model Card” by williawa

736

“The effects of caffeine consumption do not decay with a ~5 hour half-life” by kman

737

“You don’t know what you are made of till you’ve been stalked across three countries” by Shoshannah Tekofsky

738

“Why is Flesh So Weak?” by J Bostock

739

“The hard part isn’t noticing when papers are bad, it’s deciding what to do afterwards” by LawrenceC

740

“We can prevent progress! Conceptual clarity, and inspiration from the FDA” by KatjaGrace

741

“AI as a Trojan horse race” by KatjaGrace

742

“My unsupervised elicitation challenge” by DanielFilan

743

“Role-playing vs Self-modelling” by Jan_Kulveit

744

“Elementary Condensation” by Jan

745

“Hedging and Survival-Weighted Planning” by Vaniver

746

“Opus’s Schelling Steganography Has Amplifiable Secrecy Against Weaker Eavesdroppers” by Elle Najt

747

“An Alignment Journal: Features and policies” by JessRiedel, Dan MacKinlay, Luca, Daniel Murfet, david reinstein

748

“Fantasy ideology” by Ninety-Three

749

[Linkpost] “Questions raised about OpenAI leaders’ trustworthiness by the New Yorker” by Remmelt

750

“Claude Mythos System Card Preview” by anaguma

751

“My picture of the present in AI” by ryan_greenblatt

752

[Linkpost] ”[Paper] Stringological sequence prediction I” by Vanessa Kosoy

753

“We’re actually running out of benchmarks to upper bound AI capabilities” by LawrenceC

754

“Don’t write for LLMs, just record everything” by RobertM

755

“Contra Nina Panickssery on advice for children” by Sean Herrington

756

“By Strong Default, ASI Will End Liberal Democracy” by MichaelDickens

757

“AIs can now often do massive easy-to-verify SWE tasks and I’ve updated towards shorter timelines” by ryan_greenblatt

758

“Paper close reading: “Why Language Models Hallucinate”” by LawrenceC

759

“Ten different ways of thinking about Gradual Disempowerment” by David Scott Krueger (formerly: capybaralet)

760

“11 pieces of advice for children” by Nina Panickssery

761

“Steering Might Stop Working Soon” by J Bostock

762

“Am I the baddie?” by Ustice

763

“Academic Proof-of-Work in the Age of LLMs” by LawrenceC

764

“Positive sum does not mean “win-win”” by loops

765

“Considerations for growing the pie” by Zach Stein-Perlman

766

″“Following the incentives”” by David Scott Krueger (formerly: capybaralet)

767

“Chicken-Free Egg Whites” by jefftk

768

“dark ilan” by ozymandias

769

“Mean field sequence: an introduction” by Dmitry Vaintrob, Lauren Greenspan

770

“Democracy Dies With The Rifleman” by Vaniver

771

“The bar is lower than you think” by XelaP

772

“Did Anyone Predict the Industrial Revolution?” by Lost Futures

773

“Why do I believe preserving structure is enough?” by Aurelia

774

“There should be $100M grants to automate AI safety” by Marius Hobbhahn

775

“Sadly, The Whispering Earring” by Dentosal

776

“Common research advice #2: say precisely what you want to say” by LawrenceC

777

“2026: The year of throwing my agency at my health (now with added cyborgism)” by Ruby

778

[Linkpost] “Q1 2026 Timelines Update” by Daniel Kokotajlo, elifland, bhalstead

779

“How social ideas get corrupt” by Kaj_Sotala

780

“The Indestructible Future” by WillPetillo

781

“My most common advice for junior researchers” by LawrenceC

782

“The Practical Guide to Superbabies” by GeneSmith

783

“The Corner-Stone” by Benquo

784

“Systematically dismantle the AI compute supply chain.” by David Scott Krueger (formerly: capybaralet)

785

“The quest for general intelligence is hitting a wall” by Sean Herrington

786

“Intelligence Dissolves Privacy” by Vaniver

787

“Anthropic’s Pause is the Most Expensive Alarm in Corporate History” by Ruby

788

“I’m Suing Anthropic for Unauthorized Use of My Personality” by Linch

789

“Orders of magnitude: use semitones, not decibels” by Oliver Sourbut

790

“Dying with Whimsy” by NickyP

791

“AI for AI for Epistemics” by owencb, Lukas Finnveden

792

“Announcing Doublehaven with Reflections on Humour” by J Bostock

793

“Save the Sun Shrimp!” by Jack

794

“LIMBO: Who We Are, What We Do, and an Exciting High-Impact Funding Opportunity” by faul_sname

795

“Chat, is this sus?” by Tyler Tracy

796

″“You Have Not Been a Good User” (LessWrong’s second album)” by habryka

797

“Lesswrong Liberated” by Ronny Fernandez

798

“The Claude Code Source Leak” by Error

799

“Experiments With Opus 4.6’s Fiction” by Tomás B.

800

“Product Alignment is not Superintelligence Alignment (and we need the latter to survive)” by plex

801

“Co-Found Lens Academy With Me. (We have early users and funding)” by Luc Brinkman

802

“Slack in Cells, Slack in Brains” by Mateusz Bagiński

803

“I am definitely missing the pre-AI writing era” by N. Cailie

804

“The state of AI safety in four fake graphs” by Boaz Barak

805

“AI should be a good citizen, not just a good assistant” by Tom Davidson, wdmacaskill

806

″(Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL” by 7vik, Sid Black, Joseph Bloom

807

[Linkpost] “Parkinson’s Law of Worry” by Jakub Halmeš

808

“Folie à Machine: LLMs and Epistemic Capture” by DaystarEld

809

“Stop asking “how good is this” to decide between donation opportunities I recommend” by Zach Stein-Perlman

810

“Nick Bostrom: How big is the cosmic endowment?” by Zach Stein-Perlman

811

“Don’t Overdose Locally Beneficial Changes” by Mateusz Bagiński

812

“Stanley Milgram wasn’t pessimistic enough about human nature?” by David Gross

813

[Linkpost] “What if superintelligence is just weak?” by Simon Lermen

814

“Pray for Casanova” by Tomás B.

815

“ControlAI 2025 Impact Report” by Andrea_Miotti, Alex Amadori

816

“AI’s capability improvements haven’t come from it getting less affordable” by Anders Woodruff

817

“Scaffolded Reproducers, Scaffolded Agents” by Mateusz Bagiński

818

“My hobby: running deranged surveys” by leogao

819

“The Terrarium” by Caleb Biddulph

820

“Sen. Sanders (I-VT) and Rep. Ocasio-Cortez (D-NY) propose AI Data Center Moratorium Act” by Matrice Jacobine

821

“Test your best methods on our hard CoT interp tasks” by daria, Riya Tyagi, Josh Engels, Neel Nanda

822

″“What Exactly Would An International AI Treaty Say?” Is a Bad Objection” by Davidmanheim