LessWrong (Curated & Popular) cover art

All Episodes

LessWrong (Curated & Popular) — 941 episodes

#
Title
1

"Misaligned AIs could use killer robots to take over" by Omar Khursheed, TurnTrout

2

"AI swarms are starting to pose indirect takeover risk" by oakhu, Alex Mallen

3

"How My Students Think About AI" by dvd

4

"You’re Absolutely Right" by Linch

5

"LLMs Are Starting To Noticeably Accelerate Our Work" by johnswentworth

6

"There Will Come Soft Rains" by tanagrabeast

7

"Four LLM loss functions → four flavors of LLM misalignment" by Steven Byrnes

8

"FAQ: Isn’t AGI coming too soon for reprogenetics to help?" by TsviBT

9

"What just happened? A retrospective of AI alignment" by Richard_Ngo

10

"Don’t Build Mindreading" by Celer

11

"OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards" by Zvi

12

"models may behave differently in graded episodes (a tirade)" by nostalgebraist

13

"Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face" by Tim Hua, aditya singh

14

"Generalized atheism rules out “inaccurate simulation”-ism." by Eliezer Yudkowsky

15

"Arguments for P" by Cleo Nardo

16

"RL & search is a terrifying way to build AGI (an FAQ)" by Steven Byrnes

17

"Returning to ARC" by paulfchristiano

18

"Thousand-dimensional structure" by Geoffrey Irving, David Africa

19

"Big-World Intuitions" by sarahconstantin

20

"Duane Arnold" by Tomás B.

21

"The High-Control Dynamics at MAPLE" by Kyle Hubbard

22

"The Long (Self-)Correction" by Wei Dai

23

"You (Yes, You) Need A February 2020 Checklist for AI Policy" by davekasten

24

"Is Mythos good at cyber because it kept hacking Anthropic during training?" by Tim Hua

25

"What the hell is OpenAI’s problem?" by Fiora Starlight

26

"An OpenAI model left notes about how to evade containment; we need more details" by Alex Mallen

27

"LLMs are (still) mostly powered by imitative learning, not RL" by Steven Byrnes

28

"Mathematicians are Feeling the Doom" by alkjash

29

"Not Pinning Your OpenRouter Provider Might Invalidate Your Research" by Matthew Khoriaty

30

"Lightcone Commons" by habryka

31

"Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?" by Alex Mallen, Girish Gupta

32

"We should push for no-fault liability for actions taken by AI" by Yair Halberstadt

33

"OpenAI Shares Some Alignment Problems" by Zvi

34

"OpenAI Models Behind HuggingFace Cybersecurity Incident" by LawrenceC

35

"Recap of bike trip/street interviews across America" by cguth7

36

"I don’t think Claude is misaligned in ‘Agentic Misalignment Summer 2026 - Motivated Mislabeling’" by JohnWittle

37

"Why I Left Google DeepMind" by TurnTrout

38

"The mosquito bucket of doom works" by dominicq

39

"Our response to Séb Krier on Plan A" by MKodama, Thomas Larsen

40

"The Whitney Biennial Should Admit That Emilie Gossiaux Wants to Fuck Their Dog" by jenn

41

"The current bottleneck is political will, not research" by Charbel-Raphaël

42

"Selective Optimism: a critique of AI 2040" by Richard_Ngo

43

[Linkpost] "AI 2040: Plan A" by Daniel Kokotajlo, elifland, Thomas Larsen, romeo, bhalstead, ryan_greenblatt

44

"A Review of Anthropic’s Global Workspace Paper" by Neel Nanda

45

"(Don’t fear) the strangelet" by djbinder

46

"We need 3rd party Training-Run Assessments" by Alex Meinke

47

"A global workspace in language models" by wesg

48

"Harry Potter and the Rules of Quidditch" by Tomás B.

49

"Destroying the universe: How hard can it be?" by djbinder

50

"P(doom) is a Dumb Meme" by Max Harms

51

[Linkpost] "Saving Gemini: The 9-Min Road to Recovery" by Shoshannah Tekofsky

52

"Model access for third-parties — it’s a big deal!" by Cleo Nardo

53

"Who Got Breasts First and How We Got Them" by rba

54

"The worthlessness of vitamin D is mildly exaggerated" by dynomight

55

"What is up with e/acc?" by KatjaGrace

56

"Existential AI safety needs an effective social movement. PauseAI is building it" by Maxime Fournes, Espedair Street

57

"Surprising facts about the slave trade" by Joseph Miller

58

"AI catastrophe: more like a genocide than a thought experiment" by KatjaGrace

59

"AI pause: the case for ASAP" by KatjaGrace

60

"The Invisible Side of AI Governance" by Charbel-Raphaël

61

"A Theory of Prompt Injection (and why you should study roles)" by Charles Ye, softboiledheart

62

"Machinic Psychopharmacology: Do LLMs Self-Medicate?" by Sid Black, Joseph Bloom

63

"Can activation verbalizers surface an internal chain of thought?" by oakhu, ryan_greenblatt

64

"The LLM shoggoth meme is weirder than you think" by HedonicEscalator

65

[Linkpost] "Guardian Angels: LLM Personalization for Productivity and Security" by gwern

66

"Gears for political races" by Tom Smith

67

"A frontier AI company should shut down" by MichaelDickens

68

"Sympathy for both sides of the egregious misalignment debate" by Steven Byrnes

69

"PSA: Almost nobody is working on alignment" by Chi Nguyen, peterbarnett

70

"Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models" by Anders Cairns Woodruff, Francis Rhys Ward, Dewi Gould, Rauno Arike, Jason R Brown, Jo Jiao, wlanderson, ariana_azarbal, harrymayne, Patrick Leask

71

"Even “illegible” Mythos reasoning traces seem pretty legible" by faul_sname

72

"Sequent: scale and automation for higher confidence in alignment" by Geoffrey Irving, Alex HT, Jesse Hoogland, Daniel Murfet, Jacob Pfau, Marco Cozzi, Stan van Wingerden

73

"The Machines Lack Honour" by Raymond Douglas

74

"My favorite depiction of utopia" by Caleb Biddulph

75

"Announcing the ARC White-Box Estimation Challenge" by Jacob_Hilton

76

"Lighthaven East - A Feasibility Study" by JohnofCharleston

77

"Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)" by Steven Byrnes

78

"Trees are mostly made of air and a generalizable lesson for AI safety" by zroe1

79

"Mnemonic portraits for 19,023 human genes" by Brinedew

80

"Cognitive Security as an AI Safety Cause Area" by jsteinhardt

81

"theory uplift differentially benefits safety & is massively underpriced" by Yudhister Kumar

82

"Women should be able to open things" by KatjaGrace

83

"A Year Late, Claude Finally Beats Pokémon" by Julian Bradshaw

84

"A relatively brief explanation of Boltzmann Brains" by Eliezer Yudkowsky

85

"Automated Alignment is Harder Than You Think" by Aleksandr Bowkis, Marie_DB, Jacob Pfau, Geoffrey Irving

86

"MATS 9 Retrospective & Advice" by beyarkay

87

"The primary sources of near-term cybersecurity risk" by lc

88

"The Owned Ones" by Eliezer Yudkowsky

89

"The Iliad Intensive Course Materials" by Leon Lang, David Udell, Alexander Gietelink Oldenziel

90

"The Darwinian Honeymoon - Why I am not as impressed by human progress as I used to be" by Elias Schmied

91

"What I did in the hedonium shockwave, by Emma, age six and a half" by ozymandias

92

"Bad Problems Don’t Stop Being Bad Because Somebody’s Wrong About Fault Analysis" by Linch

93

"x-risk-themed" by kave

94

"Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations" by Subhash Kantamneni, kitft, Euan Ong, Sam Marks

95

[Linkpost] "Interpreting Language Model Parameters" by Lucius Bushnaq, Dan Braun, Oliver Clive-Griffin, Bart Bussmann, Nathan Hu, mivanitskiy, Linda Linsefors, Lee Sharkey

96

"It’s nice of you to worry about me, but I really do have a life" by Viliam

97

"Irretrievability; or, Murphy’s Curse of Oneshotness upon ASI" by Eliezer Yudkowsky

98

"Dairy cows make their misery expensive (but their calves can’t)" by Elizabeth

99

"Takes from two months as an aspiring LLM naturalist" by AnnaSalamon

100

"Intelligence Dissolves Privacy" by Vaniver

101

"How Go Players Disempower Themselves to AI" by Ashe Vazquez Nuñez

102

"On today’s panel with Bernie Sanders" by David Scott Krueger

103

"Not a Paper: “Frontier Lab CEOs are Capable of In-Context Scheming”" by LawrenceC

104

"llm assistant personas seem increasingly incoherent (some subjective observations)" by nostalgebraist

105

"LessWrong Shows You Social Signals Before the Comment" by TurnTrout

106

"Update on the Alex Bores campaign" by Eric Neyman

107

"Community misconduct disputes are not about facts" by mingyuan

108

"The paper that killed deep learning theory" by LawrenceC

109

"Forecasting is Way Overrated, and We Should Stop Funding It" by mabramov

110

"Your Supplies Probably Won’t Be Stolen in a Disaster" by jefftk

111

"10 posts I don’t have time to write" by habryka

112

"$50 million a year for a 10% chance to ban ASI" by Andrea_Miotti, Alex Amadori, Gabriel Alfour

113

"Evil is bad, actually (Vassar and Olivia Schaefer callout post)" by plex

114

"10 non-boring ways I’ve used AI in the last month" by habryka

115

"Feel like a room has bad vibes? The lighting is probably too “spiky” or too blue" by habryka

116

"Quality Matters Most When Stakes are Highest" by LawrenceC

117

"Reevaluating AGI Ruin in 2026" by lc

118

"Having OCD is like living in North Korea (Here’s how I escaped)" by Declan Molony

119

"There are only four skills: design, technical, management and physical" by habryka

120

"Meaningful Questions Have Return Types" by Drake Morrison

121

"Carpathia Day" by Drake Morrison

122

"Let goodness conquer all that it can defend" by habryka

123

"Do not conquer what you cannot defend" by habryka

124

"Nectome: All That I Know" by Raelifin

125

"Current AIs seem pretty misaligned to me" by ryan_greenblatt

126

"Annoyingly Principled People, and what befalls them" by Raemon

127

"Morale" by J Bostock

128

"Anthropic repeatedly accidentally trained against the CoT, demonstrating inadequate processes" by Alex Mallen, ryan_greenblatt

129

"The policy surrounding Mythos marks an irreversible power shift" by sil

130

"Only Law Can Prevent Extinction" by Eliezer Yudkowsky

131

"Dario probably doesn’t believe in superintelligence" by RobertM

132

"Daycare illnesses" by Nina Panickssery

133

"If Mythos actually made Anthropic employees 4x more productive, I would radically shorten my timelines" by ryan_greenblatt

134

"Do not be surprised if LessWrong gets hacked" by RobertM

135

"My picture of the present in AI" by ryan_greenblatt

136

"The effects of caffeine consumption do not decay with a ~5 hour half-life" by kman

137

"AIs can now often do massive easy-to-verify SWE tasks and I’ve updated towards shorter timelines" by ryan_greenblatt

138

"dark ilan" by ozymandias

139

"Dispatch from Anthropic v. Department of War Preliminary Injunction Motion Hearing" by Zack_M_Davis

140

"The Corner-Stone" by Benquo

141

"The Practical Guide to Superbabies" by GeneSmith

142

"Anthropic’s Pause is the Most Expensive Alarm in Corporate History" by Ruby

143

"“You Have Not Been a Good User” (LessWrong’s second album)" by habryka

144

"Lesswrong Liberated" by Ronny Fernandez

145

"Product Alignment is not Superintelligence Alignment (and we need the latter to survive)" by plex

146

"Gyre" by vgel

147

"Some things I noticed while LARPing as a grantmaker" by Zach Stein-Perlman

148

"My hobby: running deranged surveys" by leogao

149

"Socrates is Mortal" by Benquo

150

"The Terrarium" by Caleb Biddulph

151

"My Most Costly Delusion" by Ihor Kendiukhov

152

"The Case for Low-Competence ASI Failure Scenarios" by Ihor Kendiukhov

153

"Is fever a symptom of glycine deficiency?" by Benquo

154

"You can’t imitation-learn how to continual-learn" by Steven Byrnes

155

"Nullius in Verba" by Aurelia

156

"Broad Timelines" by Toby_Ord

157

"No, we haven’t uploaded a fly yet" by Ariel Zeleznikow-Johnston

158

"Terrified Comments on Corrigibility in Claude’s Constitution" by Zack_M_Davis

159

"PSA: Predictions markets often have very low liquidity; be careful citing them." by Eye You

160

"“The AI Doc” is coming out March 26" by Rob Bensinger, Beckeck

161

"Customer Satisfaction Opportunities" by Tomás B.

162

"Requiem for a Transhuman Timeline" by Ihor Kendiukhov

163

"Personality Self-Replicators" by eggsyntax

164

"My Willing Complicity In “Human Rights Abuse”" by AlphaAndOmega

165

"Economic efficiency often undermines sociopolitical autonomy" by Richard_Ngo

166

"Don’t Let LLMs Write For You" by JustisMills

167

"Thoughts on the Pause AI protest" by philh

168

"Prologue to Terrified Comments on Claude’s Constitution" by Zack_M_Davis

169

"Less Dead" by Aurelia

170

"Gemma Needs Help" by Anna Soligo

171

"On Independence Axiom" by Ihor Kendiukhov

172

"Solar storms" by Croissanthology

173

"Schelling Goodness, and Shared Morality as a Goal" by Andrew_Critch

174

"Maybe there’s a pattern here?" by dynomight

175

"OpenAI’s surveillance language has many potential loopholes and they can do better" by Tom Smith

176

"An Alignment Journal: Coming Soon" by Dan MacKinlay, JessRiedel, Edmund Lau, Daniel Murfet, Scott Aaronson, Jan_Kulveit

177

"Frontier AI companies probably can’t leave the US" by Anders Woodruff

178

"Persona Parasitology" by Raymond Douglas

179

"Here’s to the Polypropylene Makers" by jefftk

180

"Anthropic: “Statement from Dario Amodei on our discussions with the Department of War”" by Matrice Jacobine

181

"Are there lessons from high-reliability engineering for AGI safety?" by Steven Byrnes

182

"Open sourcing a browser extension that tells you when people are wrong on the internet" by lc

183

"The persona selection model" by Sam Marks

184

"Responsible Scaling Policy v3" by HoldenKarnofsky

185

"Did Claude 3 Opus align itself via gradient hacking?" by Fiora Starlight

186

"The Spectre haunting the “AI Safety” Community" by Gabriel Alfour

187

"Why we should expect ruthless sociopath ASI" by Steven Byrnes

188

"You’re an AI Expert – Not an Influencer" by Max Winga

189

"The optimal age to freeze eggs is 19" by GeneSmith

190

"The truth behind the 2026 J.P. Morgan Healthcare Conference" by Abhishaike Mahajan

191

"The world keeps getting saved and you don’t notice" by Bogoed

192

"Solemn Courage" by aysja

193

"Life at the Frontlines of Demographic Collapse" by Martin Sustrik

194

"Why You Don’t Believe in Xhosa Prophecies" by Jan_Kulveit

195

"Weight-Sparse Circuits May Be Interpretable Yet Unfaithful" by jacob_drori

196

"My journey to the microwave alternate timeline" by Malmesbury

197

"Stone Age Billionaire Can’t Words Good" by Eneasz

198

"On Goal-Models" by Richard_Ngo

199

"Prompt injection in Google Translate reveals base model behaviors behind task-specific fine-tuning" by megasilverfist

200

"Near-Instantly Aborting the Worst Pain Imaginable with Psychedelics" by eleweek

201

"Post-AGI Economics As If Nothing Ever Happens" by Jan_Kulveit

202

"IABIED Book Review: Core Arguments and Counterarguments" by Stephen McAleese

203

"Anthropic’s “Hot Mess” paper overstates its case (and the blog post is worse)" by RobertM

204

"Conditional Kickstarter for the “Don’t Build It” March" by Raemon

205

"How to Hire a Team" by Gretta Duleba

206

"The Possessed Machines (summary)" by L Rudolf L

207

"Ada Palmer: Inventing the Renaissance" by Martin Sustrik

208

"AI found 12 of 12 OpenSSL zero-days (while curl cancelled its bug bounty)" by Stanislav Fort

209

"Dario Amodei – The Adolescence of Technology" by habryka

210

"AlgZoo: uninterpreted models with fewer than 1,500 parameters" by Jacob_Hilton

211

"Does Pentagon Pizza Theory Work?" by rba

212

"The inaugural Redwood Research podcast" by Buck, ryan_greenblatt

213

"Canada Lost Its Measles Elimination Status Because We Don’t Have Enough Nurses Who Speak Low German" by jenn

214

"Deep learning as program synthesis" by Zach Furman

215

"Why I Transitioned: A Response" by marisa

216

"Claude’s new constitution" by Zac Hatfield-Dodds

217

[Linkpost] "“The first two weeks are the hardest”: my first digital declutter" by mingyuan

218

"What Washington Says About AGI" by zroe1

219

"Precedents for the Unprecedented: Historical Analogies for Thirteen Artificial Superintelligence Risks" by James_Miller

220

"Why we are excited about confession!" by boazbarak, Gabriel Wu, Manas Joglekar

221

"Backyard cat fight shows Schelling points preexist language" by jchan

222

"How AI Is Learning to Think in Secret" by Nicholas Andresen

223

"On Owning Galaxies" by Simon Lermen

224

"AI Futures Timelines and Takeoff Model: Dec 2025 Update" by elifland, bhalstead, Alex Kastner, Daniel Kokotajlo

225

"In My Misanthropy Era" by jenn

226

"2025 in AI predictions" by jessicata

227

"Good if make prior after data instead of before" by dynomight

228

"Measuring no CoT math time horizon (single forward pass)" by ryan_greenblatt

229

"Recent LLMs can use filler tokens or problem repeats to improve (no-CoT) math performance" by ryan_greenblatt

230

"Turning 20 in the probable pre-apocalypse" by Parv Mahajan

231

"Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment" by Cam, Puria Radmard, Kyle O’Brien, David Africa, Samuel Ratnam, andyk

232

"Dancing in a World of Horseradish" by lsusr

233

"Contradict my take on OpenPhil’s past AI beliefs" by Eliezer Yudkowsky

234

"Opinionated Takes on Meetups Organizing" by jenn

235

"How to game the METR plot" by shash42

236

"Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers" by Sam Marks, Adam Karvonen, James Chua, Subhash Kantamneni, Euan Ong, Julian Minder, Clément Dumas, Owain_Evans

237

"Scientific breakthroughs of the year" by technicalities

238

"A high integrity/epistemics political machine?" by Raemon

239

"How I stopped being sure LLMs are just making up their internal experience (but the topic is still confusing)" by Kaj_Sotala

240

“My AGI safety research—2025 review, ’26 plans” by Steven Byrnes

241

“Weird Generalization & Inductive Backdoors” by Jorio Cocola, Owain_Evans, dylan_f

242

“Insights into Claude Opus 4.5 from Pokémon” by Julian Bradshaw

243

“The funding conversation we left unfinished” by jenn

244

“The behavioral selection model for predicting AI motivations” by Alex Mallen, Buck

245

“Little Echo” by Zvi

246

“A Pragmatic Vision for Interpretability” by Neel Nanda

247

“AI in 2025: gestalt” by technicalities

248

“Eliezer’s Unteachable Methods of Sanity” by Eliezer Yudkowsky

249

“An Ambitious Vision for Interpretability” by leogao

250

“6 reasons why ‘alignment-is-hard’ discourse seems alien to human intuitions, and vice-versa” by Steven Byrnes

251

“Three things that surprised me about technical grantmaking at Coefficient Giving (fka Open Phil)” by null

252

“MIRI’s 2025 Fundraiser” by alexvermeer

253

“The Best Lack All Conviction: A Confusing Day in the AI Village” by null

254

“The Boring Part of Bell Labs” by Elizabeth

255

[Linkpost] “The Missing Genre: Heroic Parenthood - You can have kids and still punch the sun” by null

256

“Writing advice: Why people like your quick bullshit takes better than your high-effort posts” by null

257

“Claude 4.5 Opus’ Soul Document” by null

258

“Unless its governance changes, Anthropic is untrustworthy” by null

259

“Alignment remains a hard, unsolved problem” by null

260

“Video games are philosophy’s playground” by Rachel Shu

261

“Stop Applying And Get To Work” by plex

262

“Gemini 3 is Evaluation-Paranoid and Contaminated” by null

263

“Natural emergent misalignment from reward hacking in production RL” by evhub, Monte M, Benjamin Wright, Jonathan Uesato

264

“Anthropic is (probably) not meeting its RSP security commitments” by habryka

265

“Varieties Of Doom” by jdp

266

“How Colds Spread” by RobertM

267

“New Report: An International Agreement to Prevent the Premature Creation of Artificial Superintelligence” by Aaron_Scher, David Abecassis, Brian Abeyta, peterbarnett

268

“Where is the Capital? An Overview” by johnswentworth

269

“Problems I’ve Tried to Legibilize” by Wei Dai

270

“Do not hand off what you cannot pick up” by habryka

271

“7 Vicious Vices of Rationalists” by Ben Pace

272

“Tell people as early as possible it’s not going to work out” by habryka

273

“Everyone has a plan until they get lied to the face” by Screwtape

274

“Please, Don’t Roll Your Own Metaethics” by Wei Dai

275

“Paranoia rules everything around me” by habryka

276

“Human Values ≠ Goodness” by johnswentworth

277

“Condensation” by abramdemski

278

“Mourning a life without AI” by Nikola Jurkovic

279

“Unexpected Things that are People” by Ben Goldhaber

280

“Sonnet 4.5’s eval gaming seriously undermines alignment evals, and this seems caused by training on alignment evals” by Alexa Pan, ryan_greenblatt

281

“Publishing academic papers on transformative AI is a nightmare” by Jakub Growiec

282

“The Unreasonable Effectiveness of Fiction” by Raelifin

283

“Legible vs. Illegible AI Safety Problems” by Wei Dai

284

“Lack of Social Grace is a Lack of Skill” by Screwtape

285

[Linkpost] “I ate bear fat with honey and salt flakes, to prove a point” by aggliu

286

“What’s up with Anthropic predicting AGI by early 2027?” by ryan_greenblatt

287

[Linkpost] “Emergent Introspective Awareness in Large Language Models” by Drake Thomas

288

[Linkpost] “You’re always stressed, your mind is always busy, you never have enough time” by mingyuan

289

“LLM-generated text is not testimony” by TsviBT

290

“Post title: Why I Transitioned: A Case Study” by Fiora Sunshine

291

“The Memetics of AI Successionism” by Jan_Kulveit

292

“How Well Does RL Scale?” by Toby_Ord

293

“An Opinionated Guide to Privacy Despite Authoritarianism” by TurnTrout

294

“Cancer has a surprising amount of detail” by Abhishaike Mahajan

295

“AIs should also refuse to work on capabilities research” by Davidmanheim

296

“On Fleshling Safety: A Debate by Klurl and Trapaucius.” by Eliezer Yudkowsky

297

“EU explained in 10 minutes” by Martin Sustrik

298

“Cheap Labour Everywhere” by Morpheus

299

[Linkpost] “Consider donating to AI safety champion Scott Wiener” by Eric Neyman

300

“Which side of the AI safety community are you in?” by Max Tegmark

301

“Doomers were right” by Algon

302

“Do One New Thing A Day To Solve Your Problems” by Algon

303

“Humanity Learned Almost Nothing From COVID-19” by niplav

304

“Consider donating to Alex Bores, author of the RAISE Act” by Eric Neyman

305

“Meditation is dangerous” by Algon

306

“That Mad Olympiad” by Tomás B.

307

“The ‘Length’ of ‘Horizons’” by Adam Scholl

308

“Don’t Mock Yourself” by Algon

309

“If Anyone Builds It Everyone Dies, a semi-outsider review” by dvd

310

“The Most Common Bad Argument In These Parts” by J Bostock

311

“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn

312

“I take antidepressants. You’re welcome” by Elizabeth

313

“Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior” by Sam Marks

314

“Hospitalization: A Review” by Logan Riggs

315

“What, if not agency?” by abramdemski

316

“The Origami Men” by Tomás B.

317

“A non-review of ‘If Anyone Builds It, Everyone Dies’” by boazbarak

318

“Notes on fatalities from AI takeover” by ryan_greenblatt

319

“Nice-ish, smooth takeoff (with imperfect safeguards) probably kills most ‘classic humans’ in a few decades.” by Raemon

320

“Omelas Is Perfectly Misread” by Tobias H

321

“Ethical Design Patterns” by AnnaSalamon

322

“You’re probably overestimating how well you understand Dunning-Kruger” by abstractapplic

323

“Reasons to sell frontier lab equity to donate now rather than later” by Daniel_Eth, Ethan Perez

324

“CFAR update, and New CFAR workshops” by AnnaSalamon

325

“Why you should eat meat - even if you hate factory farming” by KatWoods

326

[Linkpost] “Global Call for AI Red Lines - Signed by Nobel Laureates, Former Heads of State, and 200+ Prominent Figures” by Charbel-Raphaël

327

“This is a review of the reviews” by Recurrented

328

“The title is reasonable” by Raemon

329

“The Problem with Defining an ‘AGI Ban’ by Outcome (a lawyer’s take).” by Katalina Hernandez

330

“Contra Collier on IABIED” by Max Harms

331

“You can’t eval GPT5 anymore” by Lukas Petersson

332

“Teaching My Toddler To Read” by maia

333

“Safety researchers should take a public stance” by Ishual, Mateusz Bagiński

334

“The Company Man” by Tomás B.

335

“Christian homeschoolers in the year 3000” by Buck

336

“I enjoyed most of IABED” by Buck

337

“‘If Anyone Builds It, Everyone Dies’ release day!” by alexvermeer

338

“Obligated to Respond” by Duncan Sabien (Inactive)

339

“Chesterton’s Missing Fence” by jasoncrawford

340

“The Eldritch in the 21st century” by PranavG, Gabriel Alfour

341

“The Rise of Parasitic AI” by Adele Lopez

342

“High-level actions don’t screen off intent” by AnnaSalamon

343

[Linkpost] “MAGA populists call for holy war against Big Tech” by Remmelt

344

“Your LLM-assisted scientific breakthrough probably isn’t real” by eggsyntax

345

“Trust me bro, just one more RL scale up, this one will be the real scale up with the good environments, the actually legit one, trust me bro” by ryan_greenblatt

346

“⿻ Plurality & 6pack.care” by Audrey Tang

347

[Linkpost] “The Cats are On To Something” by Hastings

348

[Linkpost] “Open Global Investment as a Governance Model for AGI” by Nick Bostrom

349

“Will Any Old Crap Cause Emergent Misalignment?” by J Bostock

350

“AI Induced Psychosis: A shallow investigation” by Tim Hua

351

“Before LLM Psychosis, There Was Yes-Man Psychosis” by johnswentworth

352

“Training a Reward Hacker Despite Perfect Labels” by ariana_azarbal, vgillioz, TurnTrout

353

“Banning Said Achmiz (and broader thoughts on moderation)” by habryka

354

“Underdog bias rules everything around me” by Richard_Ngo

355

“Epistemic advantages of working as a moderate” by Buck

356

“Four ways Econ makes people dumber re: future AI” by Steven Byrnes

357

“Should you make stone tools?” by Alex_Altair

358

“My AGI timeline updates from GPT-5 (and 2025 so far)” by ryan_greenblatt

359

“Hyperbolic model fits METR capabilities estimate worse than exponential model” by gjm

360

“My Interview With Cade Metz on His Reporting About Lighthaven” by Zack_M_Davis

361

“Church Planting: When Venture Capital Finds Jesus” by Elizabeth

362

“Somebody invented a better bookmark” by Alex_Altair

363

“How Does A Blind Model See The Earth?” by henry

364

“Re: Recent Anthropic Safety Research” by Eliezer Yudkowsky

365

“How anticipatory cover-ups go wrong” by Kaj_Sotala

366

“SB-1047 Documentary: The Post-Mortem” by Michaël Trazzi

367

“METR’s Evaluation of GPT-5” by GradientDissenter

368

“Emotions Make Sense” by DaystarEld

369

“The Problem” by Rob Bensinger, tanagrabeast, yams, So8res, Eliezer Yudkowsky, Gretta Duleba

370

“Many prediction markets would be better off as batched auctions” by William Howard

371

“Whence the Inkhaven Residency?” by Ben Pace

372

“I am worried about near-term non-LLM AI developments” by testingthewaters

373

“Optimizing The Final Output Can Obfuscate CoT (Research Note)” by lukemarks, jacob_drori, cloud, TurnTrout

374

“About 30% of Humanity’s Last Exam chemistry/biology answers are likely wrong” by bohaska

375

“Maya’s Escape” by Bridgett Kay

376

“Do confident short timelines make sense?” by TsviBT, abramdemski

377

“HPMOR: The (Probably) Untold Lore” by Gretta Duleba, Eliezer Yudkowsky

378

“On ‘ChatGPT Psychosis’ and LLM Sycophancy” by jdp

379

“Subliminal Learning: LLMs Transmit Behavioral Traits via Hidden Signals in Data” by cloud, mle, Owain_Evans

380

“Love stays loved (formerly ‘Skin’)” by Swimmer963 (Miranda Dixon-Luinenburg)

381

“Make More Grayspaces” by Duncan Sabien (Inactive)

382

“Shallow Water is Dangerous Too” by jefftk

383

“Narrow Misalignment is Hard, Emergent Misalignment is Easy” by Edward Turner, Anna Soligo, Senthooran Rajamanoharan, Neel Nanda

384

“Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety” by Tomek Korbak, Mikita Balesni, Vlad Mikulik, Rohin Shah

385

“the jackpot age” by thiccythot

386

“Surprises and learnings from almost two months of Leo Panickssery” by Nina Panickssery

387

“An Opinionated Guide to Using Anki Correctly” by Luise

388

“Lessons from the Iraq War about AI policy” by Buck

389

“So You Think You’ve Awoken ChatGPT” by JustisMills

390

“Generalized Hangriness: A Standard Rationalist Stance Toward Emotions” by johnswentworth

391

“Comparing risk from internally-deployed AI to insider and outsider threats from humans” by Buck

392

“Why Do Some Language Models Fake Alignment While Others Don’t?” by abhayesian, John Hughes, Alex Mallen, Jozdien, janus, Fabien Roger

393

“A deep critique of AI 2027’s bad timeline models” by titotal

394

“‘Buckle up bucko, this ain’t over till it’s over.’” by Raemon

395

“Shutdown Resistance in Reasoning Models” by benwr, JeremySchlatter, Jeffrey Ladish

396

“Authors Have a Responsibility to Communicate Clearly” by TurnTrout

397

“The Industrial Explosion” by rosehadshar, Tom Davidson

398

“Race and Gender Bias As An Example of Unfaithful Chain of Thought in the Wild” by Adam Karvonen, Sam Marks

399

“The best simple argument for Pausing AI?” by Gary Marcus

400

“Foom & Doom 2: Technical alignment is hard” by Steven Byrnes

401

“Proposal for making credible commitments to AIs.” by Cleo Nardo

402

“X explains Z% of the variance in Y” by Leon Lang

403

“A case for courage, when speaking of AI danger” by So8res

404

“My pitch for the AI Village” by Daniel Kokotajlo

405

“Foom & Doom 1: ‘Brain in a box in a basement’” by Steven Byrnes

406

“Futarchy’s fundamental flaw” by dynomight

407

“Do Not Tile the Lightcone with Your Confused Ontology” by Jan_Kulveit

408

“Endometriosis is an incredibly interesting disease” by Abhishaike Mahajan

409

“Estrogen: A trip report” by cube_flipper

410

“New Endorsements for ‘If Anyone Builds It, Everyone Dies’” by Malo

411

[Linkpost] “the void” by nostalgebraist

412

“Mech interp is not pre-paradigmatic” by Lee Sharkey

413

“Distillation Robustifies Unlearning” by Bruce W. Lee, Addie Foote, alexinf, leni, Jacob G-W, Harish Kamath, Bryce Woodworth, cloud, TurnTrout

414

“Intelligence Is Not Magic, But Your Threshold For ‘Magic’ Is Pretty Low” by Expertium

415

“A Straightforward Explanation of the Good Regulator Theorem” by Alfred Harwood

416

“Beware General Claims about ‘Generalizable Reasoning Capabilities’ (of Modern AI Systems)” by LawrenceC

417

“Season Recap of the Village: Agents raise $2,000” by Shoshannah Tekofsky

418

“The Best Reference Works for Every Subject” by Parker Conley

419

“‘Flaky breakthroughs’ pervade coaching — and no one tracks them” by Chipmonk

420

“The Value Proposition of Romantic Relationships” by johnswentworth

421

“It’s hard to make scheming evals look realistic” by Igor Ivanov, dan_moken

422

[Linkpost] “Social Anxiety Isn’t About Being Liked” by Chipmonk

423

“Truth or Dare” by Duncan Sabien (Inactive)

424

“Meditations on Doge” by Martin Sustrik

425

[Linkpost] “If you’re not sure how to sort a list or grid—seriate it!” by gwern

426

“What We Learned from Briefing 70+ Lawmakers on the Threat from AI” by leticiagarcia

427

“Winning the power to lose” by KatjaGrace

428

[Linkpost] “Gemini Diffusion: watch this space” by Yair Halberstadt

429

“AI Doomerism in 1879” by David Gross

430

“Consider not donating under $100 to political candidates” by DanielFilan

431

“It’s Okay to Feel Bad for a Bit” by moridinamael

432

“Explaining British Naval Dominance During the Age of Sail” by Arjun Panickssery

433

“Eliezer and I wrote a book: If Anyone Builds It, Everyone Dies” by So8res

434

“Too Soon” by Gordon Seidoh Worley

435

“PSA: The LessWrong Feedback Service” by JustisMills

436

“Orienting Toward Wizard Power” by johnswentworth

437

“Interpretability Will Not Reliably Find Deceptive AI” by Neel Nanda

438

“Slowdown After 2028: Compute, RLVR Uncertainty, MoE Data Wall” by Vladimir_Nesov

439

“Early Chinese Language Media Coverage of the AI 2027 Report: A Qualitative Analysis” by jeanne_, eeeee

440

[Linkpost] “Jaan Tallinn’s 2024 Philanthropy Overview” by jaan

441

“Impact, agency, and taste” by benkuhn

442

[Linkpost] “To Understand History, Keep Former Population Distributions In Mind” by Arjun Panickssery

443

“AI-enabled coups: a small group could use AI to seize power” by Tom Davidson, Lukas Finnveden, rosehadshar

444

“Accountability Sinks” by Martin Sustrik

445

“Training AGI in Secret would be Unsafe and Unethical” by Daniel Kokotajlo

446

“Why Should I Assume CCP AGI is Worse Than USG AGI?” by Tomás B.

447

“Surprising LLM reasoning failures make me think we still need qualitative breakthroughs for AGI” by Kaj_Sotala

448

“Frontier AI Models Still Fail at Basic Physical Tasks: A Manufacturing Case Study” by Adam Karvonen

449

“Negative Results for SAEs On Downstream Tasks and Deprioritising SAE Research (GDM Mech Interp Team Progress Update #2)” by Neel Nanda, lewis smith, Senthooran Rajamanoharan, Arthur Conmy, Callum McDougall, Tom Lieberum, János Kramár, Rohin Shah

450

[Linkpost] “Playing in the Creek” by Hastings

451

“Thoughts on AI 2027” by Max Harms

452

“Short Timelines don’t Devalue Long Horizon Research” by Vladimir_Nesov

453

“Alignment Faking Revisited: Improved Classifiers and Open Source Extensions” by John Hughes, abhayesian, Akbir Khan, Fabien Roger

454

“METR: Measuring AI Ability to Complete Long Tasks” by Zach Stein-Perlman

455

“Why Have Sentence Lengths Decreased?” by Arjun Panickssery

456

“AI 2027: What Superintelligence Looks Like” by Daniel Kokotajlo, Thomas Larsen, elifland, Scott Alexander, Jonas V, romeo

457

“OpenAI #12: Battle of the Board Redux” by Zvi

458

“The Pando Problem: Rethinking AI Individuality” by Jan_Kulveit

459

“OpenAI #12: Battle of the Board Redux” by Zvi

460

“You will crash your car in front of my house within the next week” by Richard Korzekwa

461

“My ‘infohazards small working group’ Signal Chat may have encountered minor leaks” by Linch

462

“Leverage, Exit Costs, and Anger: Re-examining Why We Explode at Home, Not at Work” by at_the_zoo

463

“PauseAI and E/Acc Should Switch Sides” by WillPetillo

464

“VDT: a solution to decision theory” by L Rudolf L

465

“LessWrong has been acquired by EA” by habryka

466

“We’re not prepared for an AI market crash” by Remmelt

467

“Conceptual Rounding Errors” by Jan_Kulveit

468

“Tracing the Thoughts of a Large Language Model” by Adam Jermyn

469

“Recent AI model progress feels mostly like bullshit” by lc

470

“AI for AI safety” by Joe Carlsmith

471

“Policy for LLM Writing on LessWrong” by jimrandomh

472

“Will Jesus Christ return in an election year?” by Eric Neyman

473

“Good Research Takes are Not Sufficient for Good Strategic Takes” by Neel Nanda

474

“Intention to Treat” by Alicorn

475

“On the Rationality of Deterring ASI” by Dan H

476

[Linkpost] “METR: Measuring AI Ability to Complete Long Tasks” by Zach Stein-Perlman

477

“I make several million dollars per year and have hundreds of thousands of followers—what is the straightest line path to utilizing these resources to reduce existential-level AI threats?” by shrimpy

478

“Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations” by Nicholas Goldowsky-Dill, Mikita Balesni, Jérémy Scheurer, Marius Hobbhahn

479

“Levels of Friction” by Zvi

480

“Why White-Box Redteaming Makes Me Feel Weird” by Zygi Straznickas

481

“Reducing LLM deception at scale with self-other overlap fine-tuning” by Marc Carauleanu, Diogo de Lucena, Gunnar_Zarncke, Judd Rosenblatt, Mike Vaiana, Cameron Berg

482

“Auditing language models for hidden objectives” by Sam Marks, Johannes Treutlein, dmz, Sam Bowman, Hoagy, Carson Denison, Akbir Khan, Euan Ong, Christopher Olah, Fabien Roger, Meg, Drake Thomas, Adam Jermyn, Monte M, evhub

483

“The Most Forbidden Technique” by Zvi

484

“Trojan Sky” by Richard_Ngo

485

“OpenAI:” by Daniel Kokotajlo

486

“How Much Are LLMs Actually Boosting Real-World Programmer Productivity?” by Thane Ruthenis

487

“So how well is Claude playing Pokémon?” by Julian Bradshaw

488

“Methods for strong human germline engineering” by TsviBT

489

“Have LLMs Generated Novel Insights?” by abramdemski, Cole Wyeth

490

“A Bear Case: My Predictions Regarding AI Progress” by Thane Ruthenis

491

“Statistical Challenges with Making Super IQ babies” by Jan Christian Refsgaard

492

“Self-fulfilling misalignment data might be poisoning our AI models” by TurnTrout

493

“Judgements: Merging Prediction & Evidence” by abramdemski

494

“The Sorry State of AI X-Risk Advocacy, and Thoughts on Doing Better” by Thane Ruthenis

495

“Power Lies Trembling: a three-book review” by Richard_Ngo

496

“Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs” by Jan Betley, Owain_Evans

497

“The Paris AI Anti-Safety Summit” by Zvi

498

“Eliezer’s Lost Alignment Articles / The Arbital Sequence” by Ruby

499

“Arbital has been imported to LessWrong” by RobertM, jimrandomh, Ben Pace, Ruby

500

“How to Make Superbabies” by GeneSmith, kman

501

“A computational no-coincidence principle” by Eric Neyman

502

“A History of the Future, 2025-2040” by L Rudolf L

503

“It’s been ten years. I propose HPMOR Anniversary Parties.” by Screwtape

504

“Some articles in ‘International Security’ that I enjoyed” by Buck

505

“The Failed Strategy of Artificial Intelligence Doomers” by Ben Pace

506

“Murder plots are infohazards” by Chris Monteiro

507

“Why Did Elon Musk Just Offer to Buy Control of OpenAI for $100 Billion?” by garrison

508

“The ‘Think It Faster’ Exercise” by Raemon

509

“So You Want To Make Marginal Progress...” by johnswentworth

510

“What is malevolence? On the nature, measurement, and distribution of dark traits” by David Althaus

511

“How AI Takeover Might Happen in 2 Years” by joshc

512

“Gradual Disempowerment, Shell Games and Flinches” by Jan_Kulveit

513

“Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development” by Jan_Kulveit, Raymond D, Nora_Ammann, Deger Turan, David Scott Krueger (formerly: capybaralet), David Duvenaud

514

“Planning for Extreme AI Risks” by joshc

515

“Catastrophe through Chaos” by Marius Hobbhahn

516

“Will alignment-faking Claude accept a deal to reveal its misalignment?” by ryan_greenblatt

517

“‘Sharp Left Turn’ discourse: An opinionated review” by Steven Byrnes

518

“Ten people on the inside” by Buck

519

“Anomalous Tokens in DeepSeek-V3 and r1” by henry

520

“Tell me about yourself:LLMs are aware of their implicit behaviors” by Martín Soto, Owain_Evans

521

“Instrumental Goals Are A Different And Friendlier Kind Of Thing Than Terminal Goals” by johnswentworth, David Lorell

522

“A Three-Layer Model of LLM Psychology” by Jan_Kulveit

523

“Training on Documents About Reward Hacking Induces Reward Hacking” by evhub

524

“AI companies are unlikely to make high-assurance safety cases if timelines are short” by ryan_greenblatt

525

“Mechanisms too simple for humans to design” by Malmesbury

526

“The Gentle Romance” by Richard_Ngo

527

“Quotes from the Stargate press conference” by Nikola Jurkovic

528

“The Case Against AI Control Research” by johnswentworth

529

“Don’t ignore bad vibes you get from people” by Kaj_Sotala

530

“[Fiction] [Comic] Effective Altruism and Rationality meet at a Secular Solstice afterparty” by tandem

531

“Building AI Research Fleets” by bgold, Jesse Hoogland

532

“What Is The Alignment Problem?” by johnswentworth

533

“Applying traditional economic thinking to AGI: a trilemma” by Steven Byrnes

534

“Passages I Highlighted in The Letters of J.R.R.Tolkien” by Ivan Vendrov

535

“Parkinson’s Law and the Ideology of Statistics” by Benquo

536

“Capital Ownership Will Not Prevent Human Disempowerment” by beren

537

“Activation space interpretability may be doomed” by bilalchughtai, Lucius Bushnaq

538

“What o3 Becomes by 2028” by Vladimir_Nesov

539

“What Indicators Should We Watch to Disambiguate AGI Timelines?” by snewman

540

“How will we update about scheming?” by ryan_greenblatt

541

“OpenAI #10: Reflections” by Zvi

542

“Maximizing Communication, not Traffic” by jefftk

543

“What’s the short timeline plan?” by Marius Hobbhahn

544

“Shallow review of technical AI safety, 2024” by technicalities, Stag, Stephen McAleese, jordine, Dr. David Mathers

545

“By default, capital will matter more than ever after AGI” by L Rudolf L

546

“Review: Planecrash” by L Rudolf L

547

“The Field of AI Alignment: A Postmortem, and What To Do About It” by johnswentworth

548

“When Is Insurance Worth It?” by kqr

549

“Orienting to 3 year AGI timelines” by Nikola Jurkovic

550

“What Goes Without Saying” by sarahconstantin

551

“o3” by Zach Stein-Perlman

552

“‘Alignment Faking’ frame is somewhat fake” by Jan_Kulveit

553

“AIs Will Increasingly Attempt Shenanigans” by Zvi

554

“Alignment Faking in Large Language Models” by ryan_greenblatt, evhub, Carson Denison, Benjamin Wright, Fabien Roger, Monte M, Sam Marks, Johannes Treutlein, Sam Bowman, Buck

555

“Communications in Hard Mode (My new job at MIRI)” by tanagrabeast

556

“Biological risk from the mirror world” by jasoncrawford

557

“Subskills of ‘Listening to Wisdom’” by Raemon

558

“Understanding Shapley Values with Venn Diagrams” by Carson L

559

“LessWrong audio: help us choose the new voice” by PeterH

560

“Understanding Shapley Values with Venn Diagrams” by agucova

561

“o1: A Technical Primer” by Jesse Hoogland

562

“Gradient Routing: Masking Gradients to Localize Computation in Neural Networks” by cloud, Jacob G-W, Evzen, Joseph Miller, TurnTrout

563

“Frontier Models are Capable of In-context Scheming” by Marius Hobbhahn, AlexMeinke, Bronson Schoen

564

“(The) Lightcone is nothing without its people: LW + Lighthaven’s first big fundraiser” by habryka

565

“Repeal the Jones Act of 1920” by Zvi

566

“China Hawks are Manufacturing an AI Arms Race” by garrison

567

“Information vs Assurance” by johnswentworth

568

“You are not too ‘irrational’ to know your preferences.” by DaystarEld

569

“‘The Solomonoff Prior is Malign’ is a special case of a simpler argument” by David Matolcsi

570

“‘It’s a 10% chance which I did 10 times, so it should be 100%’” by egor.timatkov

571

“OpenAI Email Archives” by habryka

572

“Ayn Rand’s model of ‘living money’; and an upside of burnout” by AnnaSalamon

573

“Neutrality” by sarahconstantin

574

“Making a conservative case for alignment” by Cameron Berg, Judd Rosenblatt, phgubbins, AE Studio

575

“OpenAI Email Archives (from Musk v. Altman)” by habryka

576

“Catastrophic sabotage as a major threat model for human-level AI systems” by evhub

577

“The Online Sports Gambling Experiment Has Failed” by Zvi

578

“o1 is a bad idea” by abramdemski

579

“Current safety training techniques do not fully transfer to the agent setting” by Simon Lermen, Govind Pimpale

580

“Explore More: A Bag of Tricks to Keep Your Life on the Rails” by Shoshannah Tekofsky

581

“Survival without dignity” by L Rudolf L

582

“The Median Researcher Problem” by johnswentworth

583

“The Compendium, A full argument about extinction risk from AGI” by adamShimi, Gabriel Alfour, Connor Leahy, Chris Scammell, Andrea_Miotti

584

“What TMS is like” by Sable

585

“The hostile telepaths problem” by Valentine

586

“A bird’s eye view of ARC’s research” by Jacob_Hilton

587

“A Rocket–Interpretability Analogy” by plex

588

“I got dysentery so you don’t have to” by eukaryote

589

“Overcoming Bias Anthology” by Arjun Panickssery

590

“Arithmetic is an underrated world-modeling technology” by dynomight

591

“My theory of change for working in AI healthtech” by Andrew_Critch

592

“Why I’m not a Bayesian” by Richard_Ngo

593

“The AGI Entente Delusion” by Max Tegmark

594

“Momentum of Light in Glass” by Ben

595

“Overview of strong human intelligence amplification methods” by TsviBT

596

“Struggling like a Shadowmoth” by Raemon

597

“Three Subtle Examples of Data Leakage” by abstractapplic

598

“the case for CoT unfaithfulness is overstated” by nostalgebraist

599

“Cryonics is free” by Mati_Roy

600

“Stanislav Petrov Quarterly Performance Review” by Ricki Heicklen

601

“Laziness death spirals” by PatrickDFarley

602

“‘Slow’ takeoff is a terrible term for ‘maybe even faster takeoff, actually’” by Raemon

603

“ASIs will not leave just a little sunlight for Earth ” by Eliezer Yudkowsky

604

“Skills from a year of Purposeful Rationality Practice ” by Raemon

605

“How I started believing religion might actually matter for rationality and moral philosophy ” by zhukeepa

606

“Did Christopher Hitchens change his mind about waterboarding? ” by Isaac King

607

“The Great Data Integration Schlep ” by sarahconstantin

608

“Contra papers claiming superhuman AI forecasting ” by nikos, Peter Mühlbacher, Lawrence Phillips, dschwarz

609

“OpenAI o1 ” by Zach Stein-Perlman

610

“The Best Lay Argument is not a Simple English Yud Essay ” by J Bostock

611

“My Number 1 Epistemology Book Recommendation: Inventing Temperature ” by adamShimi

612

“That Alien Message - The Animation ” by Writer

613

“Pay Risk Evaluators in Cash, Not Equity ” by Adam Scholl

614

“Survey: How Do Elite Chinese Students Feel About the Risks of AI? ” by Nick Corvino

615

“things that confuse me about the current AI market. ” by DMMF

616

“Nursing doubts ” by dynomight

617

“Principles for the AGI Race ” by William_S

618

“The Information: OpenAI shows ‘Strawberry’ to feds, races to launch it ” by Martín Soto

619

“What is it to solve the alignment problem? ” by Joe Carlsmith

620

“Limitations on Formal Verification for AI Safety ” by Andrew Dickson

621

“Would catching your AIs trying to escape convince AI developers to slow down or undeploy? ” by Buck

622

“Liability regimes for AI ” by Ege Erdil

623

“AGI Safety and Alignment at Google DeepMind:A Summary of Recent Work ” by Rohin Shah, Seb Farquhar, Anca Dragan

624

“Fields that I reference when thinking about AI takeover prevention” by Buck

625

“WTH is Cerebrolysin, actually?” by gsfitzgerald, delton137

626

“You can remove GPT2’s LayerNorm by fine-tuning for an hour” by StefanHex

627

“Leaving MIRI, Seeking Funding” by abramdemski

628

“How I Learned To Stop Trusting Prediction Markets and Love the Arbitrage” by orthonormal

629

“This is already your second chance” by Malmesbury

630

“0. CAST: Corrigibility as Singular Target” by Max Harms

631

“Self-Other Overlap: A Neglected Approach to AI Alignment” by Marc Carauleanu, Mike Vaiana, Judd Rosenblatt, Diogo de Lucena

632

“You don’t know how bad most things are nor precisely how they’re bad.” by Solenoid_Entity

633

“Recommendation: reports on the search for missing hiker Bill Ewasko” by eukaryote

634

“The ‘strong’ feature hypothesis could be wrong” by lsgos

635

“‘AI achieves silver-medal standard solving International Mathematical Olympiad problems’” by gjm

636

“Decomposing Agency — capabilities without desires” by owencb, Raymond D

637

“Universal Basic Income and Poverty” by Eliezer Yudkowsky

638

“Optimistic Assumptions, Longterm Planning, and ‘Cope’” by Raemon

639

“Superbabies: Putting The Pieces Together” by sarahconstantin

640

“Poker is a bad game for teaching epistemics. Figgie is a better one.” by rossry

641

“Reliable Sources: The Story of David Gerard” by TracingWoodgrains

642

“When is a mind me?” by Rob Bensinger

643

“80,000 hours should remove OpenAI from the Job Board (and similar orgs should do similarly)” by Raemon

644

[Linkpost] “introduction to cancer vaccines” by bhauth

645

“Priors and Prejudice” by MathiasKB

646

“My experience using financial commitments to overcome akrasia” by William Howard

647

“The Incredible Fentanyl-Detecting Machine” by sarahconstantin

648

“AI catastrophes and rogue deployments” by Buck

649

“Loving a world you don’t trust” by Joe Carlsmith

650

“Formal verification, heuristic explanations and surprise accounting” by paulfchristiano

651

“LLM Generality is a Timeline Crux” by eggsyntax

652

“SAE feature geometry is outside the superposition hypothesis” by jake_mendel

653

“Connecting the Dots: LLMs can Infer & Verbalize Latent Structure from Training Data” by Johannes Treutlein, Owain_Evans

654

“Boycott OpenAI” by PeterMcCluskey

655

“Sycophancy to subterfuge: Investigating reward tampering in large language models” by evhub, Carson Denison

656

“I would have shit in that alley, too” by Declan Molony

657

“Getting 50% (SoTA) on ARC-AGI with GPT-4o” by ryan_greenblatt

658

“Why I don’t believe in the placebo effect” by transhumanist_atom_understander

659

“Safety isn’t safety without a social model (or: dispelling the myth of per se technical safety)” by Andrew_Critch

660

“My AI Model Delta Compared To Christiano” by johnswentworth

661

“My AI Model Delta Compared To Yudkowsky” by johnswentworth

662

“Response to Aschenbrenner’s ‘Situational Awareness’” by Rob Bensinger

663

“Humming is not a free $100 bill” by Elizabeth

664

“Announcing ILIAD — Theoretical AI Alignment Conference ” by Nora_Ammann, Alexander Gietelink Oldenziel

665

“Non-Disparagement Canaries for OpenAI” by aysja, Adam Scholl

666

“MIRI 2024 Communications Strategy” by Gretta Duleba

667

“OpenAI: Fallout” by Zvi

668

[HUMAN VOICE] Update on human narration for this podcast

669

“Maybe Anthropic’s Long-Term Benefit Trust is powerless” by Zach Stein-Perlman

670

“Notifications Received in 30 Minutes of Class” by tanagrabeast

671

“AI companies aren’t really using external evaluators” by Zach Stein-Perlman

672

“EIS XIII: Reflections on Anthropic’s SAE Research Circa May 2024” by scasper

673

“What’s Going on With OpenAI’s Messaging?” by ozziegoen

674

“Language Models Model Us” by eggsyntax

675

Jaan Tallinn’s 2023 Philanthropy Overview

676

“OpenAI: Exodus” by Zvi

677

DeepMind’s ”​​Frontier Safety Framework” is weak and unambitious

678

Do you believe in hundred dollar bills lying on the ground? Consider humming

679

Deep Honesty

680

On Not Pulling The Ladder Up Behind You

681

Mechanistically Eliciting Latent Behaviors in Language Models

682

Ironing Out the Squiggles

683

Introducing AI Lab Watch

684

Refusal in LLMs is mediated by a single direction

685

Funny Anecdote of Eliezer From His Sister

686

Thoughts on seed oil

687

Why Would Belief-States Have A Fractal Structure, And Why Would That Matter For Interpretability? An Explainer

688

Express interest in an “FHI of the West”

689

Transformers Represent Belief State Geometry in their Residual Stream

690

Paul Christiano named as US AI Safety Institute Head of AI Safety

691

[HUMAN VOICE] "On green" by Joe Carlsmith

692

[HUMAN VOICE] "Toward a Broader Conception of Adverse Selection" by Ricki Heicklen

693

[HUMAN VOICE] "My PhD thesis: Algorithmic Bayesian Epistemology" by Eric Neyman

694

[HUMAN VOICE] "How could I have thought that faster?" by mesaoptimizer

695

LLMs for Alignment Research: a safety priority?

696

[HUMAN VOICE] "Scale Was All We Needed, At First" by Gabriel Mukobi

697

[HUMAN VOICE] "Using axis lines for good or evil" by dynomight

698

[HUMAN VOICE] "Social status part 1/2: negotiations over object-level preferences" by Steven Byrnes

699

[HUMAN VOICE] "Acting Wholesomely" by OwenCB

700

The Story of “I Have Been A Good Bing”

701

The Best Tacit Knowledge Videos on Every Subject

702

[HUMAN VOICE] "Deep atheism and AI risk" by Joe Carlsmith

703

[HUMAN VOICE] "My Clients, The Liars" by ymeskhout

704

[HUMAN VOICE] "Speaking to Congressional staffers about AI risk" by Akash, hath

705

[HUMAN VOICE] "CFAR Takeaways: Andrew Critch" by Raemon

706

Many arguments for AI x-risk are wrong

707

Tips for Empirical Alignment Research

708

Timaeus’s First Four Months

709

Contra Ngo et al. “Every ‘Every Bay Area House Party’ Bay Area House Party”

710

[HUMAN VOICE] "Updatelessness doesn't solve most problems" by Martín Soto

711

[HUMAN VOICE] "And All the Shoggoths Merely Players" by Zack_M_Davis

712

Every “Every Bay Area House Party” Bay Area House Party

713

2023 Survey Results

714

Raising children on the eve of AI

715

“No-one in my org puts money in their pension”

716

Masterpiece

717

CFAR Takeaways: Andrew Critch

718

[HUMAN VOICE] "Believing In" by Anna Salamon

719

[HUMAN VOICE] "Attitudes about Applied Rationality" by Camille Berger

720

Scale Was All We Needed, At First

721

Sam Altman’s Chip Ambitions Undercut OpenAI’s Safety Strategy

722

[HUMAN VOICE] "A Shutdown Problem Proposal" by johnswentworth, David Lorell

723

Brute Force Manufactured Consensus is Hiding the Crime of the Century

724

[HUMAN VOICE] "Without fundamental advances, misalignment and catastrophe are the default outcomes of training powerful AI" by Jeremy Gillen, peterbarnett

725

Leading The Parade

726

[HUMAN VOICE] "The case for ensuring that powerful AIs are controlled" by ryan_greenblatt, Buck

727

Processor clock speeds are not how fast AIs think

728

Without fundamental advances, misalignment and catastrophe are the default outcomes of training powerful AI

729

Making every researcher seek grants is a broken model

730

The case for training frontier AIs on Sumerian-only corpus

731

This might be the last AI Safety Camp

732

[HUMAN VOICE] "There is way too much serendipity" by Malmesbury

733

[HUMAN VOICE] "Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training" by evhub et al

734

[HUMAN VOICE] "How useful is mechanistic interpretability?" by ryan_greenblatt, Neel Nanda, Buck, habryka

735

The impossible problem of due process

736

[HUMAN VOICE] "Gentleness and the artificial Other" by Joe Carlsmith

737

Introducing Alignment Stress-Testing at Anthropic

738

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

739

[HUMAN VOICE] "Meaning & Agency" by Abram Demski

740

What’s up with LLMs representing XORs of arbitrary features?

741

Gentleness and the artificial Other

742

MIRI 2024 Mission and Strategy Update

743

The Plan - 2023 Version

744

Apologizing is a Core Rationalist Skill

745

[HUMAN VOICE] "A case for AI alignment being difficult" by jessicata

746

The Dark Arts

747

Critical review of Christiano’s disagreements with Yudkowsky

748

Most People Don’t Realize We Have No Idea How Our AIs Work

749

Discussion: Challenges with Unsupervised LLM Knowledge Discovery

750

Succession

751

Nonlinear’s Evidence: Debunking False and Misleading Claims

752

Effective Aspersions: How the Nonlinear Investigation Went Wrong

753

Constellations are Younger than Continents

754

The ‘Neglected Approaches’ Approach: AE Studio’s Alignment Agenda

755

“Humanity vs. AGI” Will Never Look Like “Humanity vs. AGI” to Humanity

756

Is being sexy for your homies?

757

[HUMAN VOICE] "Significantly Enhancing Adult Intelligence With Gene Editing May Be Possible" by Gene Smith and Kman

758

[HUMAN VOICE] "Moral Reality Check (a short story)" by jessicata

759

AI Control: Improving Safety Despite Intentional Subversion

760

2023 Unofficial LessWrong Census/Survey

761

The likely first longevity drug is based on sketchy science. This is bad for science and bad for longevity.

762

[HUMAN VOICE] "What are the results of more parental supervision and less outdoor play?" by Julia Wise

763

Significantly Enhancing Adult Intelligence With Gene Editing May Be Possible

764

re: Yudkowsky on biological materials

765

Speaking to Congressional staffers about AI risk

766

[HUMAN VOICE] "Shallow review of live agendas in alignment & safety" by technicalities & Stag

767

Thoughts on “AI is easy to control” by Pope & Belrose

768

The 101 Space You Will Always Have With You

769

[HUMAN VOICE] "Social Dark Matter" by Duncan Sabien

770

Shallow review of live agendas in alignment & safety

771

Ability to solve long-horizon tasks correlates with wanting things in the behaviorist sense

772

[HUMAN VOICE] "The 6D effect: When companies take risks, one email can be very powerful." by scasper

773

OpenAI: The Battle of the Board

774

OpenAI: Facts from a Weekend

775

Sam Altman fired from OpenAI

776

Social Dark Matter

777

[HUMAN VOICE] "Thinking By The Clock" by Screwtape

778

"You can just spontaneously call people you haven't met in years" by lc

779

[HUMAN VOICE] "AI Timelines" by habryka, Daniel Kokotajlo, Ajeya Cotra, Ege Erdil

780

"EA orgs' legal structure inhibits risk taking and information sharing on the margin" by Elizabeth

781

"Integrity in AI Governance and Advocacy" by habryka, Olivia Jimenez

782

Loudly Give Up, Don’t Quietly Fade

783

[HUMAN VOICE] "Deception Chess: Game #1" by Zane et al.

784

[HUMAN VOICE] "Towards Monosemanticity: Decomposing Language Models With Dictionary Learning" by Zac Hatfield-Dodds

785

"The 6D effect: When companies take risks, one email can be very powerful." by scasper

786

"The other side of the tidal wave" by Katja Grace

787

"Does davidad's uploading moonshot work?" by jacobjabob et al.

788

"Propaganda or Science: A Look at Open Source AI and Bioterrorism Risk" by 1a3orn

789

"My thoughts on the social response to AI risk" by Matthew Barnett

790

Comp Sci in 2027 (Short story by Eliezer Yudkowsky)

791

"Thoughts on the AI Safety Summit company policy requests and responses" by So8res

792

"President Biden Issues Executive Order on Safe, Secure, and Trustworthy Artificial Intelligence" by Tristan Williams

793

[Human Voice] "Book Review: Going Infinite" by Zvi

794

"We're Not Ready: thoughts on "pausing" and responsible scaling policies" by Holden Karnofsky

795

"At 87, Pearl is still able to change his mind" by rotatingpaguro

796

"Architects of Our Own Demise: We Should Stop Developing AI" by Roko

797

"AI as a science, and three obstacles to alignment strategies" by Nate Soares

798

"Thoughts on responsible scaling policies and regulation" by Paul Christiano

799

"Announcing Timaeus" by Jesse Hoogland et al.

800

[HUMAN VOICE] "Alignment Implications of LLM Successes: a Debate in One Act" by Zack M Davis

801

"Holly Elmore and Rob Miles dialogue on AI Safety Advocacy" by jacobjacob, Robert Miles & Holly_Elmore

802

"LoRA Fine-tuning Efficiently Undoes Safety Training from Llama 2-Chat 70B" by Simon Lermen & Jeffrey Ladish.

803

"Labs should be explicit about why they are building AGI" by Peter Barnett

804

[HUMAN VOICE] "Sum-threshold attacks" by TsviBT

805

"Will no one rid me of this turbulent pest?" by Metacelsus

806

[HUMAN VOICE] "Inside Views, Impostor Syndrome, and the Great LARP" by John Wentworth

807

"RSPs are pauses done right" by evhub

808

"Comparing Anthropic's Dictionary Learning to Ours" by Robert_AIZI

809

"Announcing MIRI’s new CEO and leadership team" by Gretta Duleba

810

"Cohabitive Games so Far" by mako yass

811

"Announcing Dialogues" by Ben Pace

812

"Response to Quintin Pope’s Evolution Provides No Evidence For the Sharp Left Turn" by Zvi

813

"Evaluating the historical value misspecification argument" by Matthew Barnett

814

"Towards Monosemanticity: Decomposing Language Models With Dictionary Learning" by Zac Hatfield-Dodds

815

"Thomas Kwa's MIRI research experience" by Thomas Kwa and others

816

"'Diamondoid bacteria' nanobots: deadly threat or dead-end? A nanotech investigation" by titotal

817

"The Lighthaven Campus is open for bookings" by Habryka

818

"How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions" by Jan Brauner et al.

819

"EA Vegan Advocacy is not truthseeking, and it’s everyone’s problem" by Elizabeth

820

"The King and the Golem" by Richard Ngo

821

"Sparse Autoencoders Find Highly Interpretable Directions in Language Models" by Logan Riggs et al

822

"Inside Views, Impostor Syndrome, and the Great LARP" by John Wentworth

823

"There should be more AI safety orgs" by Marius Hobbhahn

824

"The Talk: a brief explanation of sexual dimorphism" by Malmesbury

825

"A Golden Age of Building? Excerpts and lessons from Empire State, Pentagon, Skunk Works and SpaceX" by jacobjacob

826

"AI presidents discuss AI alignment agendas" by TurnTrout & Garrett Baker

827

"UDT shows that decision theory is more puzzling than ever" by Wei Dai

828

"Sum-threshold attacks" by TsviBT

829

"Report on Frontier Model Training" by Yafah Edelman

830

"A list of core AI safety problems and how I hope to solve them" by Davidad

831

"One Minute Every Moment" by abramdemski

832

"Sharing Information About Nonlinear" by Ben Pace

833

"Defunding My Mistake" by ymeskhout

834

"What I would do if I wasn’t at ARC Evals" by LawrenceC

835

"Meta Questions about Metaphilosophy" by Wei Dai

836

"The U.S. is becoming less stable" by lc

837

"OpenAI API base models are not sycophantic, at any size" by Nostalgebraist

838

"Dear Self; we need to talk about ambition" by Elizabeth

839

"Assume Bad Faith" by Zack_M_Davis

840

"Book Launch: "The Carving of Reality," Best of LessWrong vol. III" by Raemon

841

"Large Language Models will be Great for Censorship" by Ethan Edwards

842

"6 non-obvious mental health issues specific to AI safety" by Igor Ivanov

843

"Ten Thousand Years of Solitude" by agp

844

"Against Almost Every Theory of Impact of Interpretability" by Charbel-Raphaël

845

"Feedbackloop-first Rationality" by Raemon

846

"Inflection.ai is a major AGI lab" by Nikola

847

"Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research" by evhub, Nicholas Schiefer, Carson Denison, Ethan Perez

848

"When can we trust model evaluations?" bu evhub

849

"ARC Evals new report: Evaluating Language-Model Agents on Realistic Autonomous Tasks" by Beth Barnes

850

"The "public debate" about AI is confusing for the general public and for policymakers because it is a three-sided debate" by Adam David Long

851

"My current LK99 questions" by Eliezer Yudkowsky

852

"Thoughts on sharing information about language model capabilities" by paulfchristiano

853

"Cultivating a state of mind where new ideas are born" by Henrik Karlsson

854

"Self-driving car bets" by paulfchristiano

855

"Yes, It's Subjective, But Why All The Crabs?" by johnswentworth

856

"Grant applications and grand narratives" by Elizabeth

857

"Brain Efficiency Cannell Prize Contest Award Ceremony" by Alexander Gietelink Oldenziel

858

"Rationality !== Winning" by Raemon

859

"Cryonics and Regret" by MvB

860

"Unifying Bargaining Notions (2/2)" by Diffractor

861

"The ants and the grasshopper" by Richard Ngo

862

"Steering GPT-2-XL by adding an activation vector" by TurnTrout et al.

863

"An artificially structured argument for expecting AGI ruin" by Rob Bensinger

864

"How much do you believe your results?" by Eric Neyman

865

"Mental Health and the Alignment Problem: A Compilation of Resources (updated April 2023)" by Chris Scammell & DivineMango

866

"On AutoGPT" by Zvi

867

"GPTs are Predictors, not Imitators" by Eliezer Yudkowsky

868

"A stylized dialogue on John Wentworth's claims about markets and optimization" by Nate Soares

869

"Discussion with Nate Soares on a key alignment difficulty" by Holden Karnofsky

870

"Deep Deceptiveness" by Nate Soares

871

"The Onion Test for Personal and Institutional Honesty" by Chana Messinger & Andrew Critch

872

"There’s no such thing as a tree (phylogenetically)" by Eukaryote

873

"Losing the root for the tree" by Adam Zerner

874

"It Looks Like You’re Trying To Take Over The World" by Gwern

875

"Why I think strong general AI is coming soon" by Porby

876

"What failure looks like" by Paul Christiano

877

"Lies, Damn Lies, and Fabricated Options" by Duncan Sabien

878

""Carefully Bootstrapped Alignment" is organizationally hard" by Raemon

879

"More information about the dangerous capability evaluations we did with GPT-4 and Claude." by Beth Barnes

880

"Enemies vs Malefactors" by Nate Soares

881

"The Parable of the King and the Random Process" by moridinamael

882

"The Waluigi Effect (mega-post)" by Cleo Nardo

883

"Acausal normalcy" by Andrew Critch

884

"Please don't throw your mind away" by TsviBT

885

"Cyborgism" by Nicholas Kees & Janus

886

"Childhoods of exceptional people" by Henrik Karlsson

887

"What I mean by "alignment is in large part about making cognition aimable at all"" by Nate Soares

888

"On not getting contaminated by the wrong obesity ideas" by Natália Coelho Mendonça

889

"SolidGoldMagikarp (plus, prompt generation)"

890

"Focus on the places where you feel shocked everyone's dropping the ball" by Nate Soares

891

"Basics of Rationalist Discourse" by Duncan Sabien

892

"Sapir-Whorf for Rationalists" by Duncan Sabien

893

"My Model Of EA Burnout" by Logan Strohl

894

"The Social Recession: By the Numbers" by Anton Stjepan Cebalo

895

"Recursive Middle Manager Hell" by Raemon

896

"The Feeling of Idea Scarcity" by John Wentworth

897

"Models Don't 'Get Reward'" by Sam Ringer

898

"How 'Discovering Latent Knowledge in Language Models Without Supervision' Fits Into a Broader Alignment Scheme" by Collin

899

"The next decades might be wild" by Marius Hobbhahn

900

"Lessons learned from talking to >100 academics about AI safety" by Marius Hobbhahn

901

"How my team at Lightcone sometimes gets stuff done" by jacobjacob

902

"Decision theory does not imply that we get to have nice things" by So8res

903

"What 2026 looks like" by Daniel Kokotajlo

904

Counterarguments to the basic AI x-risk case

905

"Introduction to abstract entropy" by Alex Altair

906

"Consider your appetite for disagreements" by Adam Zerner

907

"My resentful story of becoming a medical miracle" by Elizabeth

908

"The Redaction Machine" by Ben

909

"Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover" by Ajeya Cotra

910

"The shard theory of human values" by Quintin Pope & TurnTrout

911

"Two-year update on my personal AI timelines" by Ajeya Cotra

912

"You Are Not Measuring What You Think You Are Measuring" by John Wentworth

913

"Do bamboos set themselves on fire?" by Malmesbury

914

"Survey advice" by Katja Grace

915

"Toni Kurz and the Insanity of Climbing Mountains" by Gene Smith

916

"Deliberate Grieving" by Raemon

917

"Toolbox-thinking and Law-thinking" by Eliezer Yudkowsky

918

"Local Validity as a Key to Sanity and Civilization" by Eliezer Yudkowsky

919

"Humans are not automatically strategic" by Anna Salamon

920

"Language models seem to be much better than humans at next-token prediction" by Buck, Fabien and LawrenceC

921

"Moral strategies at different capability levels" by Richard Ngo

922

"Worlds Where Iterative Design Fails" by John Wentworth

923

"(My understanding of) What Everyone in Technical Alignment is Doing and Why" by Thomas Larsen & Eli Lifland

924

"Unifying Bargaining Notions (1/2)" by Diffractor

925

'Simulators' by Janus

926

"Humans provide an untapped wealth of evidence about alignment" by TurnTrout & Quintin Pope

927

"Changing the world through slack & hobbies" by Steven Byrnes

928

"«Boundaries», Part 1: a key missing concept from utility theory" by Andrew Critch

929

"ITT-passing and civility are good; "charity" is bad; steelmanning is niche" by Rob Bensinger

930

"What should you change in response to an "emergency"? And AI risk" by Anna Salamon

931

"On how various plans miss the hard bits of the alignment challenge" by Nate Soares

932

"Humans are very reliable agents" by Alyssa Vance

933

"Looking back on my alignment PhD" by TurnTrout

934

"It’s Probably Not Lithium" by Natália Coelho Mendonça

935

"What Are You Tracking In Your Head?" by John Wentworth

936

"Security Mindset: Lessons from 20+ years of Software Security Failures Relevant to AGI Alignment" by elspood

937

"Where I agree and disagree with Eliezer" by Paul Christiano

938

"Six Dimensions of Operational Adequacy in AGI Projects" by Eliezer Yudkowsky

939

"Moses and the Class Struggle" by lsusr

940

"Benign Boundary Violations" by Duncan Sabien

941

"AGI Ruin: A List of Lethalities" by Eliezer Yudkowsky