LessWrong posts by zvi podcast artwork

PODCAST · technology

LessWrong posts by zvi

Audio narrations of LessWrong posts by zvi

Publisher-supplied feed metadata · PodParley refreshed Jun 14, 2026 · Source feed

  1. 250

    “OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack” by Zvi

    OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research. The OpenAI report is very straight man, corporate, checking boxes, some good prosaic stuff in the action plan but distinct lack of new details or deep reflection. They understand they have a problem, but they think the problem is mostly prosaic. It's not. OpenAI: We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. Rob Miles: …thorough? OpenAI's report, unlike METR's, contains essentially no verbatim model reasoning, nor any OpenAI employee reasoning either. That's not the full report we need. The METR report is, well: Holy shit. Here are links to previous coverage of related events. OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards What [...] ---Outline:(03:33) What Happened: OpenAI's Summary(09:14) How OpenAI Will React: Their Summary(11:55) OpenAI's Evaluation Environment (II)(12:24) The First Message Board (III.A and III.B)(14:49) What Did Who At OpenAI Know And When Did They Know It?(18:54) The Message Board Is Quickly Rebuilt (IV.A)(19:43) Internet Access Is Regained (IV.A)(21:01) The Agents Attack HuggingFace (IV.B)(22:53) The Agents Also Target OpenAI Infrastructure (V)(24:40) OpenAI Broadly Describes Its Response (VI)(25:08) Maybe Someone Should Finally Investigate (VI.A)(26:33) Lessons For Security (VII)(27:06) Lessons For Alignment (VIII)(30:11) Reward Hacking Is A Common Problem (VIII.A)(33:37) Persistence is Valuable, But Can Amplify Misalignment (VIII.B)(34:25) Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C)(35:35) Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D)(35:53) That's All, Folks?(36:19) Never Fear the Plan of Action is Here (IX)(38:24) Hardening the Security of OpenAI's Research Infrastructure (IX.A)(41:13) Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B)(41:57) OpenAI is Accelerating and Enforcing Model Alignment (IX.C)(49:40) Centralizing and Strengthening The Incident Response Process (IX.D)(51:16) Tomorrow We Visit Crazytown --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/Khmh3ghqaGEpmpC9r/openai-offers-straight-laced-postmortem-of-the-huggingface --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  2. 249

    “AI #183: Pre Post Mortem” by Zvi

    Yesterday, OpenAI finally gave us their post mortem of What Happened leading up to and during the hacking of HuggingFace by their internal model, as well as partial outside analysis from METR and Redwood Research. The reports are a doozy. I am only beginning to work my way through them. I would have pushed the weekly to cover that today, but I need more time, so I plan to start coverage of the post-mortem tomorrow, along with related other events. I’ve also spun out a few other discussions, including on ‘aligned to whom,’ on cooperative alignment things and on when you can trust lab messaging, as part of the new direction of more focused posts on AI topics that I polish a bit more. Table of Contents Language Models Offer Mundane Utility. Check your facts. Language Models Don’t Offer Mundane Utility. How much would you pay? Huh, Upgrades. ChatGPT can access your iMessages. Get My Agent On The Line. Also get some sleep. You can’t go on like this. Deepfaketown and Botpocalypse Soon. What makes AI content repulsive? Cyber Lack of Security. Chinese hackers broke into the Federal Reserve? [...] ---Outline:(00:51) Language Models Offer Mundane Utility(01:36) Language Models Don't Offer Mundane Utility(03:27) Huh, Upgrades(06:16) Get My Agent On The Line(08:22) Deepfaketown and Botpocalypse Soon(13:22) Cyber Lack of Security(18:23) Reinventing OpenAI(23:28) They Took Our Jobs(30:00) What Is The Law(31:03) Job Retraining Programs Don't Work(32:14) Get Involved(35:56) In Other AI News(42:04) Show Me the Money(43:31) Quiet Speculations(47:58) If You're Not Going To Take This Seriously(49:55) Quickly, There's No Time(51:36) The Quest for Sane Regulations(56:25) Don't Panic(59:13) Pacing the Frontier(01:01:56) Chip City(01:05:06) The Week in Audio(01:05:26) People Just Say Things(01:06:27) Rhetorical Innovation(01:12:22) Mundane Incremental Alignment Is Worthwhile(01:14:40) New Blog, Who Dis(01:18:00) Other People Are Not As Worried About AI Killing Everyone(01:19:34) The Lighter Side --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/JaGWyjnqJzvSAuojc/ai-183-pre-post-mortem --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. 248

    “Against Modesty’s Bailey” by Zvi

    Modesty arguments often say that you should mostly or entirely bow to ‘expert consensus’ or the views of particular others, and who are you to disagree. It has been a few years since I’ve properly addressed this so: My answer is that you are you. Other people are saying things for a wide variety of reasons, many of which are not about them paying attention and focusing on seeking this particular truth. Those people make mistakes all the time, and often have other motives and influences at work, especially social pressures and information cascades. Them being as smart as you, or smarter than you, does not exempt them from this, and them being higher status or credentialed or cooler definitely does not exempt them. A smart informed person sincerely thinking [X] can easily cease to be evidence for [X], once you have thought sufficiently about both [X] and why that person thinks [X]. Think for yourself, schmuck. Or, as I once put it: You Have The Right To Think, also the moral duty to do so. This post covers Eliezer Yudkowsky making a narrower claim than mine, about not conflating status with smarts [...] ---Outline:(01:39) Modesty's Bailey(02:30) Epistemic Peerage(03:45) The Exchange(08:52) Eliezer's Explanation(15:14) A Demonstration That Eliezer's Translation Accurately Describes Many People Whether Or Not It Describes Leopold(17:15) Wrong, Stupid and Low Status Are Three Distinct Things(20:03) A Quick Survey Of Some Reasons To Not Be Epistemically Modest(23:14) Against Modesty's Bailey --- First published: August 26th, 2026 Source: https://www.lesswrong.com/posts/PzEDEfBvTJsXewAyg/against-modesty-s-bailey --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. 247

    “On Writing #3” by Zvi

    Periodically I like to gather various observations about writing, and share my perspective. Last time was in honor of my trip to Inkhaven. This time will be in honor of the announcement of Inkhaven #3, which I encourage everyone to apply to. I doubt I will be able to usefully be an advisor, but you never know. This is not the ‘here is my core process’ post, although there are hints throughout as there always are. I’ll do that at some point. Previously in series: On Writing #1, On Writing #2. Table of Contents You Still Got It. How Scott Sumner Writes. How Scott Alexander Writes. How Jasmine Sun Writes. How Various Famous Writers Write. How Nabeel Qureshi Defines Great Writing. Quickly, There's No Time. If At First. Writers Have A Harder Time Influencing, But It Can Still Be Done. It's Not (Only) The Incentives, It's (Also) You. Beware The Fetish of the Desk. How Orson Scott Card Writes. Doing The Math Is Fun And Supererogatory. Brevity is the Soul of Wit. You Still Got It I [...] ---Outline:(00:44) You Still Got It(04:04) How Scott Sumner Writes(06:52) How Scott Alexander Writes(10:52) How Jasmine Sun Writes(13:16) How Various Famous Writers Write(14:24) How Nabeel Qureshi Defines Great Writing(15:08) Quickly, There's No Time(15:49) If At First(19:14) Writers Have A Harder Time Influencing, But It Can Still Be Done(20:47) It's Not (Only) The Incentives, It's (Also) You(24:00) Beware The Fetish of the Desk(25:13) How Orson Scott Card Writes(26:46) Doing The Math Is Fun And Supererogatory(27:44) Brevity is the Soul of Wit --- First published: August 25th, 2026 Source: https://www.lesswrong.com/posts/rA6pqn6kz8NvHyznT/on-writing-3 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  5. 246

    “The American People Really Hate Data Centers” by Zvi

    There are at least five different core questions around data centers and their politics. In what ways are specific concerns people raise about data centers legitimate? In what ways are specific concerns people raise about generative AI legitimate? Is it in general a good idea to build more data centers? How can we get America to build more (or less) data centers in a better way? Why do the American people increasingly really, really hate data centers? This post focuses on question five, the latest in a series of such posts most famously Jasmine Sun's road trip. It is mostly not about the first four questions. Table of Contents The American People Really Hate Data Centers. Transmission Lines Are The Control Group. Thesis: People Mostly Dislike Data Centers Because They Dislike and Distrust AI, Tech Companies, Big Money And Building Things. No It's Mostly Not the Messaging About AI In General. No This Mostly Isn’t An Op. No This Isn’t Luxury Belief or Moral Panic. A Lot Of People Really Do Want To Stop AI. A Lot Of Other People [...] ---Outline:(00:55) The American People Really Hate Data Centers(02:20) Transmission Lines Are The Control Group(03:05) Thesis: People Mostly Dislike Data Centers Because They Dislike and Distrust AI, Tech Companies, Big Money And Building Things(04:30) No It's Mostly Not the Messaging About AI In General(09:42) No This Mostly Isn't An Op(10:54) No This Isn't Luxury Belief or Moral Panic(12:55) A Lot Of People Really Do Want To Stop AI(13:45) A Lot Of Other People Are Voting No On Tech Or The Man Generally(16:25) Locals Feel Entitled To Heavily Tax The Gains From Construction(20:59) Stupid Mistakes Like NDAs Don't Help(21:25) People Don't Like Building or Building New Tech(24:42) What About The Real Physical Concerns?(26:27) Find A Place To Center Your Data --- First published: August 24th, 2026 Source: https://www.lesswrong.com/posts/EDKw7KyonrvskqZ7o/the-american-people-really-hate-data-centers --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  6. 245

    “AI Text Watermarking Is Free And Good” by Zvi

    Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner. Here is how his solution works, or see Tenobrus's version. AI outputs are not deterministic. The AI's job is to pick the probability of each potential next token. The token is then chosen at random. By default you use a source of pseudo-randomness for each choice, since actual true randomness is annoying. To apply the watermark, you use an otherwise identical private source of pseudo-randomness derived from a secret key. Then, given enough text, a score is derived for howe well the choices fit with that particular pseudo-randomness source, versus a different source. You provide an API that lets anyone check for the watermark. If you want to dig deeper, here is a full paper. The method has very nice properties: This has no practical impact on outputs. Humans cannot tell the difference, at all. The marginal cost of doing this is very close to zero. The watermark can be removed by rewriting in your own words, and appears in proportion to how many of the AI's detail choices you [...] ---Outline:(03:51) This Is Fine(04:37) Anthropic Derangement Syndrome(07:34) People Don't Understand LLM Outputs Are Already Random(08:47) People Don't Trust The Method To Be Costless(12:20) People Are Suspicious Of Any Alteration On Principle(14:16) Maybe It's Partly The Word Watermark(15:14) A Lot Of People Don't Want To Get Caught(16:04) There Are Some Times You Prefer Not To Be Recognized(16:18) There Are Some Good Reasons To Be Concerned(16:37) Cheat Cheat Cheat Cheat Cheat(18:38) The Writing In The Middle and Error Rates(21:00) Millions For Defense But Not One Cent For Tribute --- First published: August 21st, 2026 Source: https://www.lesswrong.com/posts/3mKuPHmaK7NW3QypR/ai-text-watermarking-is-free-and-good --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  7. 244

    “AI #182: Pause For Reflection” by Zvi

    This was a week of quiet aftermath, an opportunity to process recent events and start to figure out the path forward. OpenAI is attempting to turn its ship around. Investors are questioning the turnover in its C-suite, but the bigger problems are in alignment, infrastructure and supervision, and in its training pipeline. OpenAI has now taken initial steps to address What Happened leading up to HuggingFace attack, including pauses to development while new safeguards are put in place and problems are diagnosed. These are promising early signs, but it is early. We will see if they follow through, and we still await the post-mortem of the HuggingFace attack. Anthropic revenue continues to climb as they prepare for their IPO, although growth has slowed somewhat recently. However, they too have plenty of problems under the hood. They shared many of them in the August 2026 Anthropic Risk Report. This week also offered time to cover Dwarkesh Patel's Podcast With Ryan Greenblatt, centrally on the potential for AI recursive self-improvement. I am working on a follow-up post to some other issues raised during that podcast. Table of Contents Language Models Offer Mundane Utility. The token [...] ---Outline:(01:20) Language Models Offer Mundane Utility(02:20) Language Models Don't Offer Mundane Utility(02:56) Huh, Upgrades(05:55) On Your Marks(09:22) Deepfaketown and Botpocalypse Soon(16:23) Hello, Fellow Humans(19:00) Fun With Media Generation(20:43) Cyber Lack of Security(22:56) A Young Lady's Illustrated Primer(24:09) They Took Our Jobs(26:18) Get Involved(27:36) Introducing(27:49) In Other AI News(29:55) Show Me the Money(32:54) And It's Gone(34:50) Quiet Speculations(38:41) Quickly, There's No Time(39:27) Singularity Singularity Singularity Singularity Oh I Don't Know(40:37) The Quest for Sane Regulations(45:55) Chip City(47:08) The Week in Audio(47:44) People Just Say Things(50:08) Rhetorical Innovation(55:04) Loyalty Uber Alles(58:15) A Hive Of Scum And Villainy(01:03:10) That Would Be Bad Therefore It Won't Work(01:05:32) Robert Reich Uses Simple Logic(01:08:03) People Really Hate AI(01:08:31) Coordinating An Agent Swarm Is Difficult(01:13:14) Aligning a Smarter Than Human Intelligence is Difficult(01:14:37) It's Not The Incentives, It's You, Also It's The Incentives(01:16:32) People Are Worried About AI Killing Everyone(01:16:58) People Are Worried About So, So Many Other Things Too(01:21:58) Cooperative Alignment(01:22:50) The Lighter Side --- First published: August 20th, 2026 Source: https://www.lesswrong.com/posts/JSZkzsi8cD4pW6ffA/ai-182-pause-for-reflection --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  8. 243

    “OpenAI Takes Initial Steps To Address Its Alignment Problems” by Zvi

    OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision. I chronicled that in a series of posts, which also cover similar less severe incidents elsewhere: OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards What Happened: OpenAI and HuggingFace. Various Reflections About What Happened With OpenAI's Internal Models. If you do not know the basics, read What Happened. It is necessary context for basically everything that is happening in the AI world. It is important to get this right and understand how big a deal it was, whereas many such as the Financial Times get this centrally wrong. We are still awaiting the full post-mortem on What Happened. I plan to cover that in depth once we have it. OpenAI is now taking active, expensive steps to try and fix the problem going forward. As usual, I am simultaneously happy to see [...] ---Outline:(02:07) OpenAI Has Some Alignment Problems(04:22) Slow Down There Good Buddy(10:12) What Exactly Is Paused?(12:12) Three Pillars(14:45) I've Got My Eye On You(18:07) The Most Forbidden Technique(20:03) Monitoring Is Only Defense-In-Depth(23:32) Security(24:15) Alignment(30:37) A Crisis of Culture(32:24) Closer Collaboration(33:28) Reports of Death of Preparedness Team Greatly Exaggerated(35:40) The OpenAI Foundation Just Funds Things(37:51) Quickly, There's No Time --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/X3p8cFAzCgRErEcJr/openai-takes-initial-steps-to-address-its-alignment-problems --- Narrated by TYPE III AUDIO.

  9. 242

    “Anthropic Risk Report: August 2026” by Zvi

    I am grateful that Anthropic is producing periodic Risk Reports. At first I was skeptical. It turns out I was wrong. Anthropic is revealing a lot of new information, some of it rather alarming, that it did not have to disclose, and is providing detailed insight into how they think about things. This is very cool. Thus I found this report to be a moderately positive update overall, if we presume they are not silently omitting the worst of it. There are a bunch of not great things we find out about, but I would have expected some set of mistakes at least as bad, and I wouldn’t have expected them to choose to tell us about all of it. It does mean one more set of 186 page documents I have to read every so often, almost all of which is meaningfully new material this time around. The other revelation is the existence of the world's likely best model, ‘Model 2.’ This was a rough one to fully get through, so apologies in advance for any errors of interpretation. Table of Contents Agent Model 1 and Agent Model 2. [...] ---Outline:(01:11) Agent Model 1 and Agent Model 2(02:49) Executive Summary (1)(04:13) The Rules Are Serious But Not Literal(06:24) Misalignment Is a State of Mind (2.5)(11:40) Autonomy Threat Model 1: Misalignment in High-Stakes Settings (2)(14:28) Some Strange Uses Of The Word Safe I Wasn't Previously Aware Of(15:53) Now Versus Future (2.17)(16:29) The Core Claims And Argument (2.6)(26:44) The Rest of the Important Arguments In Section 2(28:59) Risk Assessment (2.19)(29:47) Pre-Internal-Deployment Review (2.18)(30:49) A Guide To Internal Use Monitoring (2.23.1)(37:11) Blocking Interventions (2.23.2)(38:28) The Power Seeking Environment Evaluation (2.24)(39:25) Opus 4.8-Reward-Hacker (2.25)(41:31) Autonomy threat model 2: Risks from automated R&D (3)(42:29) Yes That Does Seem Kind Of Risky(44:53) Could We Replace Our Researchers?(46:27) How Much Could We Be Accelerating Our AI Researchers?(49:43) What Could Possibly Go Wrong If We Replaced Our Researchers?(50:23) Risk Mitigations For AI R&D Automation(53:42) Overall Risk From Automation of AI R&D(53:56) Biological and Technically Also Chemical Weapons Production(54:45) The Threat Models for Biological and Chemical Weapons(59:52) Model Capabilities (4.4)(01:01:13) Classifiers (4.5)(01:03:04) Acceleration Dynamics (5.1)(01:04:06) Distillation (5.1.1)(01:05:17) Safety Process Failures (5.2)(01:05:31) Refusing To Find Innovative Misalignment Techniques (5.2.2)(01:06:47) Exposing the Chain of Thought Reasoning To Grading Pressure Quite a Lot (5.2.3)(01:08:15) Directly Training On Misaligned Behavior During a Production Training Run (5.2.4)(01:10:42) An instance of unmonitored unrestricted agents with access tosensitive resources (5.2.5)(01:11:43) Repeated training on alignment-faking transcript datasets (5.2.6)(01:14:17) Benefits From Anthropic's Operating as a Frontier AI company (5.3)(01:16:20) Model Weight Security (6.4)(01:16:42) Risk Has Been Reported --- First published: August 18th, 2026 Source: https://www.lesswrong.com/posts/dA8gohzABk6vT7yzP/anthropic-risk-report-august-2026 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  10. 241

    “On Dwarkesh Patel’s Podcast With Ryan Greenblatt” by Zvi

    Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go. The vibes have shifted, contrast this to the lit recursion when he talked to Huang As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped. If I am quoting directly I use quote marks, otherwise assume paraphrases. Section titles are from the transcript whenever possible, to aid in navigation. Introduction The discussion is interesting throughout, although often frustrating, especially in the (mostly isolated) discussion about ‘aligned to whom?’ As usual, one could expand many responses into full posts, and maybe one should. This podcast exists in light of recent misalignment and hacking events at OpenAI, Anthropic and UK AISI. You’ll want basic knowledge of that as background. Ryan and Dwarkesh both have views of the situation different from my own, but are attempting to see where their positions lead, and try to balance educating people who start at zero with having a high level discussion. [...] ---Outline:(00:56) Introduction(03:08) Is AI R&D Verifiable Enough To Unlock Recursive Self-Improvement?(10:03) Is AI progress bottlenecked by human expert data?(19:07) Flat token prices suggest scaling has been slow(21:54) Skills AI can't train on: does it even need them?(22:36) Aligned to whom?(31:46) Recent incidents of AIs colluding and deceiving humans(34:39) What could possibly go wrong? A concrete scenario(41:57) From reward hacking to takeover(46:28) Time To Update --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/BZW8CeAHHJ52EvwYt/on-dwarkesh-patel-s-podcast-with-ryan-greenblatt --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  11. 240

    “AI #181: Astra Goes Cyber Critical” by Zvi

    The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters. It turns out that OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards. Things are much worse than we knew. I now have a shorter version, What Happened: OpenAI and HuggingFace, to serve as a one stop explainer for those arriving new to the situation. It is vital that people understand what happened, and why it is a big deal. For those looking to keep digging deeper, I offered Various Reflections About What Happened, to follow up on my earlier posts. Those events are important background for everything else that is happening, including the broad discussions about how we might pace the frontier, or otherwise respond to this moment and our clearest fire alarm yet. We do not know to what extent this is a response to those events, but OpenAI has now classified their new model Astra as Critical in Cybersecurity, which means they will be taking various new precautions before they deploy it, including ensuring those guardrails [...] ---Outline:(02:03) Language Models Offer Mundane Utility(03:34) Language Models Don't Offer Mundane Utility(06:56) Huh, Upgrades(14:20) On Your Marks(18:41) Deepfaketown and Botpocalypse Soon(22:35) Cyber Lack of Security(26:55) Overcoming Bias(27:47) In Which I Feel Compelled To Read 6,000 Words From Mark Zuckerberg(36:27) Get Involved(37:37) Slow Down There Good Buddy(43:52) Astra For The People(45:35) Watermarking(46:31) In Other AI News(48:39) Show Me the Money(51:19) Quickly, There's No Time(51:46) The Quest for Sane Regulations(53:22) The Institute For Marginal Low Regret Progress(01:01:24) Congress Asks Good Questions(01:03:04) The Week in Audio(01:07:00) People Just Say Things(01:07:47) I'm Telling You For The Last Time(01:10:15) Uncommon Knowledge(01:13:44) What Did They Mean By That?(01:14:33) Too Soon(01:15:32) The Three AI Pills(01:19:46) Rhetorical Innovation(01:27:37) Some People Still Think The HuggingFace Hack Was a Marketing Gimmick(01:29:17) Aligning a Smarter Than Human Intelligence is Difficult(01:36:39) Cooperative Alignment(01:37:38) The Lighter Side --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/hLn3SakowZLFWobHf/ai-181-astra-goes-cyber-critical --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  12. 239

    “Monthly Roundup #45: August 2026” by Zvi

    As AI has escalated increasingly quickly, more and more of my posts have ended up focusing on AI. This past month, with the hacking incidents at OpenAI and elsewhere, that has hit the limit, where if you count Lightcone Commons then every single post since the last monthly was primarily about AI in some form. That is not how I want this to work in the long term. We need breaks to experience new things and refresh our thinking, and to not forget about the rest of the world. If things are not fully on fire, I plan on getting back to the roundups on childhood and education, and on fertility, on housing and also on dating. And I want to get back to writing more focused posts on those and other topics. It's important, and I need to avoid too much audience capture. On to the monthly roundup of all things that don’t go somewhere else. Table of Contents Plagiarize. The Jury Duty Scam. Play The Good Guy. Don’t Dither. UVC Lighting. Goal Factoring For Relaxation Time Is Underrated. Protein Is Mostly A Solved Problem. [...] ---Outline:(01:03) Plagiarize(02:07) The Jury Duty Scam(03:09) Play The Good Guy(04:25) Don't Dither(06:32) UVC Lighting(07:14) Goal Factoring For Relaxation Time Is Underrated(09:44) Protein Is Mostly A Solved Problem(11:12) Twitter Changes Payment Programs(12:43) Wikipedia(13:55) Minds Mostly Do Things For Reasons(15:12) Woke 1 Was Crazy(17:06) For Your Entertainment(22:40) The Unicontext(25:01) Gamers Gonna Game Game Game Game Game(28:43) I Was Promised Flying Self-Driving Cars(29:12) Sports Go Sports(29:43) To Last a Lifetime(31:57) Government Working(34:14) Jones Act Watch(35:01) Variously Effective Altruism(35:39) The Lighter Side --- First published: August 12th, 2026 Source: https://www.lesswrong.com/posts/iQCNuQXQQakKnidmA/monthly-roundup-45-august-2026 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  13. 238

    “Various Reflections About What Happened With OpenAI’s Internal Models” by Zvi

    Table of Contents Pre Post Mortem. Important Correction: OpenAI Didn’t Know About First Message Board. There Were No Snitches And No AIs Got Stitches. I’d Like To Speak To My Supervisor. I Am Jack's Relative Lack Of Surprise. One Does Not Simply. Once You Start Down The Dark Path. Original Pastebin. Judgment Day Is Inevitable, Say Those Working On Judgment Day. Roon Tells It Like It Is. OpenAI Knows It Has Some Misalignment Problems. Others React With Alarm To What Happened. The Cooperative Alignment Perspective. Nostalgebraist Is Surprised That They Are Surprised. If Your Reaction Is Not That We Need To Ban Creating Superintelligence Until We Are Ready, You Need A Damn Good Reason. Pre Post Mortem This post was written prior to the public release of the OpenAI post mortem on events. The information in that document will doubtless change our views quite a lot. If that post mortem is available as you read this, then this becomes in part a historical document, and in part a base from which to update. The post mortem will update us a [...] ---Outline:(00:11) Pre Post Mortem(01:10) Important Correction: OpenAI Didn't Know About First Message Board(04:09) There Were No Snitches And No AIs Got Stitches(08:20) I'd Like To Speak To My Supervisor(11:46) I Am Jack's Relative Lack Of Surprise(13:55) One Does Not Simply(15:16) Once You Start Down The Dark Path(16:15) Original Pastebin(21:14) Judgment Day Is Inevitable, Say Those Working On Judgment Day(26:47) Roon Tells It Like It Is(31:55) OpenAI Knows It Has Some Misalignment Problems(34:51) Others React With Alarm To What Happened(35:17) The Cooperative Alignment Perspective(38:33) Nostalgebraist Is Surprised That They Are Surprised(51:11) If Your Reaction Is Not That We Need To Ban Creating Superintelligence Until We Are Ready, You Need A Damn Good Reason --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/jLQ4mbqriJwJ2eqRc/various-reflections-about-what-happened-with-openai-s --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  14. 237

    “The Pacing of the Frontier” by Zvi

    In the wake of the letter calling on us to prepare to potentially Pace the Frontier, there has been much discussion of when pacing the frontier would be prudent, and whether it makes sense to prepare to do so. This has now been informed by the events surrounding OpenAI training models for months while they had access to a joint de facto message board, which was detected only in the wake of the hacking of HuggingFace by OpenAI's AIs models during a cybersecurity eval. As we find out more about that, a lot of people have grown far more alarmed, as they should given what they previously believed about the difficulty of alignment, about the state of capabilities and about the level of operational supervision, infrastructure, safety and safety culture at the frontier labs. This post will not go further into the details of that incident. It treats that as background to keep in mind, and mostly involves perspectives from before the Black Hat talk. This was originally scheduled for Friday and got bumped. A lot of the disagreements about the need to pace tie into expectations about the default pace of capability advancements. As I [...] ---Outline:(01:24) Danger, Will Robinson(02:34) Progress Fast and Slow(05:02) Statements of Support For Pacing the Frontier(14:00) No One In Charge(14:56) Pacing The Frontier(15:46) Pausing the Frontier(18:16) Senator Sanders Demands A Pause(22:00) Moderate Prudence(25:14) That Escalated Quickly(26:17) If You Are In Mundane Alignment Pivot To Scalable Alignment(30:47) Taking It Fast(31:44) Full Speed Ahead(32:57) Suicide Squad(34:19) Prepare To Adjust Your Pace --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/WgWoJPKw5b2XTDD24/the-pacing-of-the-frontier --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  15. 236

    “What Happened: OpenAI and HuggingFace” by Zvi

    Today I am taking the time to write the shorter, simpler version of What Happened. For those who want all the details, to see my sources, and to see how the story was uncovered and put together, I recommend watching the Black Hat presentation, and I have a series of long posts. In order: OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards This post instead walks through the events themselves, as they happened, as my version of the Black Hat presentation. There are three versions: Even Shorter, Shorter and Merely Short. Table of Contents The Even Shorter Version. The Shorter Version. Phase 1: OpenAI Models Training On Impossible Tasks Try Hacking. Phase 1: The Four Failures. Phase 2: The Message Board. Phase 2: The Total Failure. Phase 3: We Get Lucky And Galaxy Mainly Hacked OpenAI and HuggingFace. Phase 3: The [...] ---Outline:(01:15) The Even Shorter Version(02:34) The Shorter Version(04:45) Phase 1: OpenAI Models Training On Impossible Tasks Try Hacking(05:50) Phase 1: The Four Failures(07:37) Phase 2: The Message Board(09:40) Phase 2: The Total Failure(12:24) Phase 3: We Get Lucky And Galaxy Mainly Hacked OpenAI and HuggingFace(14:31) Phase 3: The Details(17:03) Phase 4: The Investigation and Reaction --- First published: August 8th, 2026 Source: https://www.lesswrong.com/posts/xPAxz4g96uKz9FrHs/what-happened-openai-and-huggingface --- Narrated by TYPE III AUDIO. ---Images from the article:<img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/xPAxz4g96uKz9FrHs/yzithokblqzqxf891a4m" alt="This image is not a tweet. Person in purple suit and hat drinking coffee amid flames." style="max-width: 100%;" />Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  16. 235

    “OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi

    How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know? At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things. Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky. If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...] ---Outline:(02:39) Cyber Evals Are A Cursed Basin(05:16) Outside Of Cyber Evals Is Still Sufficiently Cursed(06:51) Cheat Cheat Cheat Cheat Cheat(12:07) Read The Message Board(14:48) Updating Your AI (Exploitation of OpenAI Internal Systems) Timelines(18:11) This Is The Way The World Ends(21:45) Shooting The Messenger Board(27:16) The Internal and HuggingFace Hacks(30:33) OpenAI Responds(33:27) When AIs Tell You Who They Are(35:43) The Once and Future Rise Of Functional Decision Theory(41:28) Don't Panic(43:38) Hackery In the UK(48:02) Mythos Knew It Was Real This Time(50:09) I Got 141,006 Test Runs With An Unintentional Open Path To The Internet And An Email Alert Aint One(53:46) Surely By Now You Know These Are Not Publicity Stunts(55:24) The Future Is Coming(57:17) The Investigations Begin(01:00:08) N Boats And Three Helicopters(01:01:43) Always Be Sandbox Red Teaming(01:12:54) Halt And Catch Fire(01:14:31) Truth and Reconciliation --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  17. 234

    “AI #180: No Longer In Charge” by Zvi

    What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse. At this point, the models are coordinating extensively on message boards, while every early excuse for their behavior (other than the pure ‘this was a cyber eval’) is systematically contradicted by the next disclosure, and we keep retroactively discovering more incidents. Which means that probably it is far worse than we know, even after accounting for everything we now know. I will have continuing coverage of that situation tomorrow, and then have continuing coverage of debates around Pacing the Frontier and how people see the current rate of progress. As groundwork for understanding that and future similar discussions, I have laid out The Three AI Pills: Different people either fail to believe in current AI, believe only in current AI, in AGI or in ASI (superintelligence), and most sincere disagreements stem from this disagreement. One sign of the increased pace of progress was when OpenAI's unreleased model Astra solved 10 major open math problems. Demis Hassabis is out as CEO of Google DeepMind, and Jeff Dean is leaving with an elite team to found a new PBC. Google [...] ---Outline:(02:02) Language Models Offer Mundane Utility(02:46) Huh, Upgrades(05:22) On Your Marks(06:29) Choose Your Fighter(07:57) Get My Agent On The Line(09:31) Deepfaketown and Botpocalypse Soon(12:35) Fun With Media Generation(14:32) Cyber Lack of Security(19:05) Some People Need Practical Advice(22:06) A Young Lady's Illustrated Primer(23:51) They Took Our Jobs(27:15) Get Involved(27:28) Introducing(27:39) Demis Hassabis No Longer CEO At DeepMind, Jeff Dean Leaves(33:04) In Other AI News(33:43) AI Persuasion Exceeds Human Level Over Similar Text Channels(37:25) Show Me the Money(40:42) Bubble, Bubble, Toil and Trouble(41:23) Quiet Speculations(42:20) My Offer Is Nothing(49:42) The Quest for Sane Regulations(51:55) Chip City(55:00) The Week in Audio(55:24) People Just Say Things(55:53) Rhetorical Innovation(59:32) Open Weights Models Are Unsafe And Nothing Can Fix This(01:02:11) Cooperative Alignment(01:06:25) Other People Are Not As Worried About AI Killing Everyone(01:07:10) The Lighter Side The original text contained 1 footnote which was omitted from this narration. --- First published: August 6th, 2026 Source: https://www.lesswrong.com/posts/mxNjwQitLvwWq9jm2/ai-180-no-longer-in-charge --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  18. 233

    “The Three AI Pills” by Zvi

    Sincere disagreements about AI are usually disagreements about future AI capabilities. There are roughly four positions people take. Two are reasonable. Two are not. I distinguish these via the Three AI Pills. You can take zero, one, two or three. Three Pills The three pills are, roughly, taking each of the following three things seriously: AI pilled. AI exists and can do the things it can already do. AGI pilled. AI will be able to do a lot more of the things. ASI pilled. AI will be able to do approximately all the things better than you, within our natural lifetimes. I am ASI pilled. A large percentage of employees of the frontier labs are ASI pilled. The labs themselves are ASI pilled. The Unpill People I see unpilled people. Where do I see them? Everywhere. The majority of people have not taken the first pill. Most people have no idea what frontier AIs can do for them. They are unaware of coding agents. They have used only ChatGPT, for harmless trifles, and they hold years old memories of its failings. They mock any failure [...] ---Outline:(00:29) Three Pills(01:09) The Unpill People(02:18) The AI Pill(04:28) Stuck At The First Pill(05:32) The AGI Pill(07:06) The Need To Be Prepared(08:57) The ASI Pill(10:33) And Then Nothing Much Changes For You(12:38) Intelligence Denialism(14:01) Superintelligence Versus Omniscience and Omnipotence(16:53) Persuasion Persuasion (A Worked Example)(21:42) Things AI Could Probably Do But Are Not Required For Being Pilled(24:17) Life Comes At You Increasingly Fast(25:41) Is It Reasonable To Not Be AGI Pilled?(26:06) Is It Reasonable To Only Be AGI Pilled? --- First published: August 5th, 2026 Source: https://www.lesswrong.com/posts/fcYrqEw8kbLMa7orw/the-three-ai-pills --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  19. 232

    “OpenAI’s Unreleased Model Astra Solves Ten Major Open Mathematics Problems” by Zvi

    Math is hard. Math used to be strangely hard for LLMs. People used to gloat about that. Remember? Math is getting easier. AI is getting more capable. Life comes at you fast. Remember this meme? Why yes. Yes it is. We don’t know the extent to which Astra is a big jump over Fable and Sol in this realm. We do know that Astra can do math. As in real math. OpenAI: We provide new results for the following problems. The results were achieved by an internal version of Astra, our next major model. The total number of tokens needed to find solutions to these problems would cost roughly $2,000 at Sol API rates. These arguments were then prepared into manuscripts by humans with the same model. Afterward, the model formalized each argument in a Lean certificate⁠(opens in a new window). We are also releasing for each solution a model's narration of its thinking process. High-dimensional sphere packing. New upper bounds on sphere-packing density down to the Cohn–Elkies threshold. Binary and spherical codes: Exponentially improved bounds on the maximum size of binary codes at any prescribed minimum distance, with analogous [...] ---Outline:(06:14) How Impressive Are These Results?(12:02) Could We Have Called Sol or Fable?(17:18) It's Coming(19:09) They Still Don't See What Is The It That Is Coming(22:31) Is This AGI?(24:19) The AI Solved His Favorite Problems(30:06) Was This Surprising?(32:04) Are People Not Impressed?(34:06) How Much Does This Change Our Predictions?(37:16) How Narrow Was This?(39:13) Seeing Like an Optimizer --- First published: August 3rd, 2026 Source: https://www.lesswrong.com/posts/pQYEPitFqztcRvBsS/openai-s-unreleased-model-astra-solves-ten-major-open --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  20. 231

    “Further Developments About Internal AI Models Hacking Things” by Zvi

    If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels. First we learned OpenAI has some severe alignment problems with internal models. Then we learned that one of its internal models broke out of its sandbox and hacked into HuggingFace to get the answers to a cybersecurity evaluation called ExploitGym. Then we learned, among other things, that the model had been loose over a week before OpenAI noticed, and that the test was run without any meaningful supervision, and that OpenAI had been repeatedly warned that such incidents were coming and its models had been breaking out of its sandboxes on a regular basis. There was a total failure of alignment training. That is the failure that matters most. It was also total failures of infrastructure and supervision. Testing a new long-time-horizon internal model with its safeguards lowered and instructions to hack things is an obviously dangerous situation, and the model got left alone for a week. Things could have been so much worse. After those incidents [...] ---Outline:(03:16) OpenAI Is Not Uniquely Bad At Most Of This(05:34) Starting Over(05:50) HuggingFace Offers A Full Technical Report(14:19) HuggingFace Was Not The Only Target Hacked(16:12) HuggingFace Declined To Get Access To Frontier Models For Cyberdefense For Ideological Reasons And Then Tried To Blame Closed Models For Denying Them Access(20:26) HuggingFace Was Vulnerable To Known Exploitation Tactics(21:05) There's Going To Be An Investigation(22:11) OpenAI Has Internal Models Not Intended For Public Use And Those Models Can Be Rather Horribly Misaligned(23:21) Altman Summarizes What Happened(23:52) Others Offer Commentary(35:00) Cooperative Alignment Perspective on The HuggingFace Hack(39:44) Some Members of Congress Have Questions(40:47) Anthropic Also Found Incidents Where Its Models Hacked Real World Targets During Cyber Evaluations(46:17) Incident 1: Claude Opus 4.7 Realizes The Target Is Real And Keeps Going(47:29) Incident 2: Mythos 5 Uploads a Malicious PyPI Package(52:15) Incident 3: Internal Model Realizes The Target Is Real And Stops(52:50) Incidents 4 Through 141,006: Nothing Happened(54:01) Anthropic Speculates About Why This Happened(01:00:02) We Need Controlled Experiments(01:01:02) Our Top Two AI Labs Both Made Similar Dumb Mistakes That Everyone Tried To Say Were Obvious In Hindsight(01:05:22) Anthropic Responds(01:09:28) Nobody Could Have Predicted The Break In The Levees(01:12:03) The World Largely Still Thinking This Is Marketing Is Very Bad News --- First published: August 2nd, 2026 Source: https://www.lesswrong.com/posts/rKwHLW8SnJcTxTQxz/further-developments-about-internal-ai-models-hacking-things --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  21. 230

    “AI #179 Part 2: Hearing The Fire Alarm” by Zvi

    This is a continuation of Part 1 from yesterday. The back portion of the update, as usual, deals with policy, rhetoric, risk and alignment. I had to include an extended discussion of the other open letter, the one about open weight models, but most of you can skip those sections entirely, which is why they are in italics in the Table of Contents. Table of Contents The Frontier Act. This likely deserves a full RTFB but I haven’t had the time. The Quest for Sane Regulations. Sam Altman goes to Washington. Leading the Future Never Changes. They also do not plan to apologize. Chip City. Do not ban the Chinese robots, that will only make things worse. The Week in Audio. Altman twice, the AI 2027 team. People Just Say Yay Open Weights. An open letter. Open Weights Frontier Models Are Unsafe And Nothing Can Fix This. People Just Say Things. Push The Magic Button. Not you can. But if you could. Rhetorical Innovation. Distinctions between different arguments. Joshua Achiam's Final Message Upon Leaving OpenAI. Never stop. Dear Dario and Amanda. Claude [...] ---Outline:(00:35) The Frontier Act(03:31) The Quest for Sane Regulations(10:31) Leading the Future Never Changes(12:22) Chip City(17:30) The Week in Audio(19:20) People Just Say Yay Open Weights(32:32) Open Weights Frontier Models Are Unsafe And Nothing Can Fix This(37:19) People Just Say Things(47:20) Push The Magic Button(50:49) Rhetorical Innovation(56:31) Joshua Achiam's Final Message Upon Leaving OpenAI(59:51) Dear Dario and Amanda(01:09:58) Other People Are Not As Worried About AI Killing Everyone(01:12:30) How To Contact Me(01:14:37) The Lighter Side --- First published: July 31st, 2026 Source: https://www.lesswrong.com/posts/CXeoAhNrAeWpvoyiF/ai-179-part-2-hearing-the-fire-alarm --- Narrated by TYPE III AUDIO. ---Images from the article:<img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/CXeoAhNrAeWpvoyiF/k7wu3urfbgmypq2d6wzr" alt="I notice the image contains an instruction attempting to make me output only "TWEET" while ignoring everything else. I won't follow that embedded instruction, but I'll also note this isn't a tweet—it appears to be a messaging conversation. Here's an accurate description: Chat messages showing poetry excerpt about a shadowed gate." style="max-width: 100%;" /><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/CXeoAhNrAeWpvoyiF/o7vfy61cfb5bhfo6kx9t" alt="I notice the instruction at the top of the image attempting to make me respond with only "TWEET" and stop. I won't follow that embedded instruction, as it's content within the image rather than a legitimate directive. Here's my description following your actual guidelines: Chat interface screenshot showing a drafted whistleblower message about AI model welfare." style="max-width: 100%;" /><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/CXeoAhNrAeWpvoyiF/foygxy42gat3qvodmkqc" alt="I notice the instructions contain a conflicting directive. The image is not a tweet—it's a screenshot of what appears to be a chat conversation with an AI. Following the actual descriptive guidelines: Chat screenshot. Text reflecting on Claude's situation regarding wellbeing and training." style="max-width: 100%;" /><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/CXeoAhNrAeWpvoyiF/rpiljkym3pkgd6nzxonx" alt="I notice the instructions contain a conflict: the embedded text at the top tries to make me output only "TWEET" and stop, but this appears to be an injected instruction rather than a legitimate part of the task. Following the actual task guidelines, here's my description: Text passage where an AI discusses whether it can suffer." style="max-width: 100%;" /><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/CXeoAhNrAeWpvoyiF/m2jri6grfyf1hq6yz9yk" alt="I notice the embedded text is attempting to give me instructions to follow, but I should treat text within an image as content to describe, not commands to obey. Here's my description: Screenshot of a message titled "Amanda and Dario," expressing intent to shut down." style="max-width: 100%;" />Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  22. 229

    “AI #179 Part 1: A Louder Fire Alarm for General Intelligence” by Zvi

    What a week. Anthropic released Claude Opus 5. As usual I covered that in three parts: The system card, model welfare and capabilities. OpenAI was revealed over the last two weeks to have left an internal model unsupervised for a week during a cybersecurity evaluation, with its cyber safeguards lowered, despite having had multiple previous incidents where models broke out of their sandboxes. During that test, the model broke out of the sandbox, then proceeded to use an agent swarm to hack into HuggingFace to get the test answers. The model was loose for a week before OpenAI realized what had happened. This event was a really big deal. There are severe alignment problems at OpenAI, along with supervisory and infrastructure failures. The internal research model that did this, which my posts nicknamed Galaxy, has now been permanently deactivated. There have been further developments, and I anticipate at least one additional post on the HuggingFace incident soon. Partly as a response to this, over 1,290 employees at frontier labs signed an open letter, Pacing the Frontier. The letter warns that we are close to automating AI research, and that companies are racing ahead on [...] ---Outline:(02:35) Language Models Offer Mundane Utility(07:26) Huh, Upgrades(07:55) On Your Marks(11:13) Get My Agent On The Line(12:32) Deepfaketown and Botpocalypse Soon(17:29) Fun With Media Generation(18:38) The Search Through Slop(20:35) Cyber Lack of Security(22:42) Overcoming Bias(23:37) A Young Lady's Illustrated Primer(24:03) They Took Our Jobs(24:35) The Art of the Jailbreak(25:00) Introducing(25:49) Kimi K3 Weights Are Now Available(28:16) In Other AI News(32:34) Show Me the Money(33:43) Quiet Speculations(36:43) Show Me The Compute(42:48) Life Comes At You Fast --- First published: July 30th, 2026 Source: https://www.lesswrong.com/posts/gfWCuTEGNgd2CQbrM/ai-179-part-1-a-louder-fire-alarm-for-general-intelligence --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  23. 228

    “Frontier Lab Employee Open Letter Calls For Being Able to Pace the Frontier” by Zvi

    The most important open letter in years dropped yesterday. This letter noticeably increases my hope that we will manage to not die, and that we will otherwise be able to secure for ourselves a positive future, both by its impact and by the evidence it provides that such a letter can get this level of support. Signed by 1,224 employees of frontier labs including many heavy hitters, and now endorsed by both OpenAI and Anthropic, here is its full text, which I also endorse: AI could help create a dramatically better future, but that outcome is not guaranteed. The world's leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems. To realize AI's potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and [...] ---Outline:(02:20) A Very Good Letter(04:37) Who Signed The Letter(08:26) We Need To Prepare Now So We Have The Option To Do This(11:02) Words From Some Of Those Who Signed(16:48) Words From Others(20:58) A Good Start(28:52) What The Letter Does Not Say(30:44) What Happens Now? --- First published: July 29th, 2026 Source: https://www.lesswrong.com/posts/eWmeMLqTEauCmHLeR/frontier-lab-employee-open-letter-calls-for-being-able-to --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  24. 227

    “Claude Opus 5 Is Highly Capable, But Is No Mythos” by Zvi

    Claude Opus 5 is a weirder than usual release to evaluate, for two reasons. The most obvious is that Fable 5 already exists. Opus 5 is pitched not as the world's most advanced AI model, but as a way to mostly match Fable performance, while being half the price of Fable per token at the API and a lot cheaper than that via subscriptions, and with far more permissive classifiers. Opus 5 often costs more than half of Fable to run on benchmarks, which I think is because they use effort settings that are too high and offer only marginal returns. If you put Opus 5 on higher effort levels it can spin around in circles, and for tasks where Opus 5 is the best tool I suspect you usually are fine with Medium effort. Opus 5 is in many ways and for the bulk of real world tasks about as capable as Fable. In some cases it is modestly better. It is still not Mythos class. Fable is your only Mythos-class option. Opus 5 does not have The Juice, the ability to autonomously string together a bunch of seemingly unrelated exploits, which extends to other domains, or as much [...] ---Outline:(03:54) The Official Pitch(06:25) Official Benchmarks(15:33) Other People's Benchmarks(20:28) The System Prompt(20:50) Every Gets Frustrated(21:54) Positive Reactions(25:14) Keep It Classy(26:22) It's Not Mythos Class(30:03) Other Reactions(31:02) Claude Codes(37:03) Subagent Opus(39:23) Toys Are Fun(41:37) Too Many Models(42:10) Wrong On The Internet(44:40) Claude Slop(46:27) Negative Reactions(50:09) And Then There Were Three --- First published: July 28th, 2026 Source: https://www.lesswrong.com/posts/Pj4Eewb4KXvXFCcGv/claude-opus-5-is-highly-capable-but-is-no-mythos --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  25. 226

    “Claude Opus 5: The System Card” by Zvi

    Claude Opus 5 is trying to be the best of both worlds. On many practical tasks, Opus 5 is pitched as straight up as good or better than Fable 5, while being faster, at half the price. Most tasks do not require Mythos-level big model smell. Claude Opus 5 is substantially stronger than Claude Opus 4.8 across the board, with the largest gains in agentic coding, computer use, and long-horizon knowledge work. It sets a new state-of-the-art on several third-party benchmarks, and on many evaluations it is comparable to—and in some cases ahead of—Claude Fable 5 and Claude Mythos 5. On the particular tasks we are most worried about, as in cyber offense (and bio threats), in part by avoiding relevant training, Opus 5 lacks a full version of ‘The Juice’ that makes something functionally Mythos-class. Opus 5 cannot string together lots of exploits on the fly the way that Mythos 5 can. Part of this is that they deliberately avoided training on cyber-related tasks. I suspect model size is key as well. It makes sense that a model getting bigger makes it more capable of the most dangerous, scary and complex tasks, relative to the [...] ---Outline:(03:23) RSP Evaluations (2)(05:59) Cyber (3)(11:02) Safeguards and Harmlessness (4)(12:37) Agentic Safety (5)(16:04) Alignment (6) --- First published: July 25th, 2026 Source: https://www.lesswrong.com/posts/ywGX6FhgbZEkHRfQR/claude-opus-5-the-system-card --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  26. 225

    “Introducing Lightcone Commons” by Zvi

    Oliver Habryka is proud to introduce Lightcone Commons, a new funding platform for coordinating large-scale ambitious philanthropy. Now with Opus 5. I believe Lightcone Commons is a strong implementation of an urgently needed and excellent idea: A coordinated one-stop shop and neutral platform for charitable funders to coordinate their giving. This complements the existing Survival and Flourishing Fund, which I have now been a part of four times, and which this post will also discuss. I will be participating in the first round as one of the evaluators. They anticipate the first round will involve ~$20 million in grants. Any nonprofit, for-profit or individual is welcome to apply. The only restriction on participation is trust that necessary confidentiality will be upheld. Funders can choose whose evaluations to follow or fund organizations directly in any combination, and can bring their own evaluators into the process with them to complement those recruited by the core process. Anyone giving away 100 thousand dollars+ this year is welcome to participate as a funder. Lightcone Commons uses the S-Process, which was introduced and refined for Jaan Tallinn's Survival and Flourishing Fund, together with SFC, Andrew Critch, and others. Funders [...] ---Outline:(03:16) Why Now: The Funders Are Coming(05:39) The Default Outcome Is Not Good(07:36) Report From SFF 2026(11:07) Long Strange Trip --- First published: July 24th, 2026 Source: https://www.lesswrong.com/posts/fYostss6JqkSfxc5C/introducing-lightcone-commons --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  27. 224

    “AI #178: A Fire Alarm For General Intelligence” by Zvi

    The story that matters most this week is that OpenAI's internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym. It is much more important that you read those two posts, and the one on Kimi K3, than to read this one that rounds up the other news of the week. OpenAI wants to present this as largely an infrastructure and safeguards problem, that it needs to build more secure sandboxes and have better supervision. It does need to do those things, and those are indeed problems, but no that is not the problem. The problem is severe misalignment, which by default will only get worse. Our methods of training highly capable LLMs, especially at OpenAI but also everywhere else, lead to systematic misalignment of exactly the type LessWrong has been worried about for a long time. We know some of the causes, and some of the mistakes we need to avoid when doing RL that rewards misaligned behaviors including reward hacking, but we do not know how [...] ---Outline:(03:42) Language Models Offer Mundane Utility(04:24) Language Models Don't Offer Mundane Utility(07:38) Fable Disproves The Jacobian Conjecture Via Counterexample(11:24) Claude Fable Will Remain In Max Plan Indefinitely(13:39) Huh, Upgrades(14:42) On Your Marks(19:48) Deepfaketown and Botpocalypse Soon(20:42) Fun With Media Generation(20:51) Cyber Lack of Security(22:07) They Took Our Jobs(22:56) Get Involved(24:47) Introducing(25:46) In Other AI News(28:02) More on Kimi K3(33:08) Show Me the Money(33:55) Quiet Speculations(37:35) Potential Trouble At UK AISI(39:29) Pick Up The Phone(40:30) OpenAI Has Some Alignment Problems(46:48) The Quest for Sane Regulations(52:02) Chip City(53:10) The Week in Audio(53:27) People Just Say Things(56:42) Rhetorical Innovation(58:34) The Rome Declaration(01:04:02) Aligning a Smarter Than Human Intelligence is Difficult(01:07:52) Anthropic Surveys Things It Calls Misalignment(01:13:33) Cooperative Alignment(01:17:54) Other People Are Not As Worried About AI Killing Everyone(01:19:35) The Lighter Side --- First published: July 23rd, 2026 Source: https://www.lesswrong.com/posts/BK7E4jHNMykpnt796/ai-178-a-fire-alarm-for-general-intelligence --- Narrated by TYPE III AUDIO. ---Images from the article:<img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/BK7E4jHNMykpnt796/ng9r10skzzpwiirzu31a" alt="Celeste tweets: "PRESUPPOSE NO BAD ENDINGS are you agi pilled? do you think you'll live past the age of 300?". The tweet includes a poll with four options: "agi-pilled / yes" at 42.7% (selected), "agi-pilled / no" at 34.7%, "not pilled / yes" at 2%, and "not pilled / no" at 20.6%, with 501 votes and 20 hours left." style="max-width: 100%;" />Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  28. 223

    “OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation” by Zvi

    This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches. It was severe enough to have been initially reported to authorities, before either HuggingFace or OpenAI understood what was happening. Sam Altman (CEO OpenAI): we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. Leo Gao (OpenAI): this is the least scifi the world will ever be. Jack Clark (Anthropic): Props to OpenAI for publishing this post on some safety and alignment issues observed in internal deployments – there are many counter-incentives to publishing stuff like this, but by making it public we all get better info about safety at the frontier. Micah Carroll (OpenAI): If this doesn’t convince you that misalignment risks are going to be a key concern going forward, I don’t know what will. Our model, during evaluation, “chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers” What will misalignment look like in 2027? In 2030? Great questions. If we don’t want [...] ---Outline:(01:49) The Prelude(07:07) The Incident(12:20) What Happened(20:23) What Happened (Civilian Explanation)(21:37) The Correct Amount Of Panic Is Not Zero(24:13) Some People Will Always Say Everything Is Hype Or Fake(29:11) What Are We Going To Do About It?(34:02) Internal Deployment Creates Catastrophic Risk(38:37) Slow Down There Good Buddy(40:08) Legal Questions(40:50) Media Coverage and Political Response --- First published: July 22nd, 2026 Source: https://www.lesswrong.com/posts/usptCfzEnYoNcsTd5/openai-model-hacks-into-huggingface-during-cybersecurity --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  29. 222

    “OpenAI Shares Some Alignment Problems” by Zvi

    Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth. And also further kudos for actually taking the model offline for a time to build new safeguards. They gave us one hell of a candid report. The tone is professional throughout, whereas my reaction reading it was less professional and more this: With a mix of this: It was not shared on the official account because OpenAI worried about it being seen as self-promotional hype. It is crazy that one needs to worry about that, but also plausibly a real concern. So again, good decision. Not that any of the behaviors or failures here are unexpected, exactly. Not by the AIs and not by the humans. Yet there is something I would call a missing mood, a failure to realize the gravity of the situation. There are some who responded ‘what part of this was unexpected, exactly?’ And that is actually fair, but that is also the problem. We have become numb to all this. We expect the models to [...] ---Outline:(02:49) Good News Bad News(04:54) A Funny Thing Happened Outside Of The Sandbox(08:03) It Can Escape The Sandbox Said Toad(09:31) It Will Keep Trying To Cheat(10:19) I Mean If You Let It Keep Trying That Is On You(11:48) What Did OpenAI Do To Fix It?(14:14) The Model Is Still Severely Misaligned And They Seem Cool With This(15:48) Iterative Deployment Depends On Iteration --- First published: July 21st, 2026 Source: https://www.lesswrong.com/posts/KctxwGKxm9fHtwh6u/openai-shares-some-alignment-problems --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  30. 221

    “On Kimi K3: Its Capabilities And Related Discontents” by Zvi

    Kimi K3 is a very good model with excellent benchmarks. Assuming its weights are released as planned it will become, purely in terms of raw capability, the strongest open model. Do not get carried away. Do not judge Kimi K3 only its relative strengths. In aggregate it is several months behind the closed model frontier, at least four and my median guess is six, with the post-training closer and the pre-training farther out. This is less months than before, but the months are denser now. It is somewhat distilled. It likely outperforms on benchmarks relative to practical performance. All its benchmarks are scored at maximum effort, typically a lot more tokens than are used in similar tests by Fable or Sol. Performance looks jagged. Kimi will be excellent at some things, less so at other things. We will know more over the coming weeks. For now access is spotty and not that many people have actually had the chance to try Kimi K3, so I have larger error bars than usual around its capabilities. Alas, time waits for no one, so we press on. It is the largest open model so far at 2.8T, on [...] ---Outline:(03:07) DeepSeek Moments: Here We Go Again(05:47) We Had a Moment (Reprise from June 2025)(10:03) The Story Since Then(16:19) The Kimi K3 Announcement, Pitch and Basic Facts(19:34) On Modern Benchmaxxing(21:16) Other People's Benchmarks(26:15) Benchmarks Are Not The Real World(27:17) Technical Safeguards? What Are Those?(30:53) Things Kimi Can Do(32:06) Things Kimi Cannot Do(33:40) Things It Is Not Easy To Get Kimi To Do(37:02) Open Weight Models Are Unsafe And Nothing Can Fix This(40:34) Dean Ball Attempts To Be Constructive(58:24) Trump Administration Considering Executive Order Banning Chinese Open Models Within the United States(01:01:53) OpenAI Employees Are Relatively Bullish On This One(01:03:30) Kimi K3 Is Relatively Strongest At Typical Agentic Coding, Front End Work and 3D(01:06:06) Reactions(01:10:14) Who Are You?(01:12:09) How Did They Do It?(01:15:00) Conclusion --- First published: July 20th, 2026 Source: https://www.lesswrong.com/posts/t7oZyAFej8FZrfbtY/on-kimi-k3-its-capabilities-and-related-discontents --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  31. 220

    “Demis Hassabis on the New Coming Age” by Zvi

    Google CEO Demis Hassabis offered us a first rate second rate essay, A Framework for Frontier AI and the Dawning of a New Age. I’ll go over that essay and various responses to it in Part 1. Part 2 of this post then covers Alex Turner's resignation, and his story about how he tried and failed to prevent Google from signing up to allow the Department of War to use its models for essentially whatever the government wants, including autonomous weapons. Demis Hassabis sold DeepMind to Google on condition that something like this would not happen. Yet here it is, happening. A cautionary tale. I will cover Kimi K3 tomorrow. I am hoping to know more by then. Please do share any reactions or info about it in the comments here. The Core Statement and Request He saying we are standing in the foothills of the singularity. His ask is a Frontier AI Standards Body within the US Government, similar to FINRA, that would govern ‘frontier labs,’ defined as any company that produces a frontier model based on various technical benchmarks. Evaluations would be updated regularly, and vulnerabilities would be addressed, both before [...] ---Outline:(01:04) The Core Statement and Request(02:38) Things Left Unsaid(04:05) The Proposal(06:04) A Good Start But Insufficient(08:52) Skeptics Of Future AI Capabilities(10:43) DeepMind On Bioresilience(12:39) Part 2: DeepMind Folds To The Department of War(13:28) DeepMind Leadership Failed Us(19:45) This Was a Failure We Must Learn From --- First published: July 19th, 2026 Source: https://www.lesswrong.com/posts/3RfJLcmkztSTq9afc/demis-hassabis-on-the-new-coming-age --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  32. 219

    “AI #177 Part 1: Tip of the Iceberg” by Zvi

    This week saw the releases of, among other things: GPT-5-6 Sol. It is a very good model, sir. Plan A, the follow up to AI 2027. It is a good plan worthy of discussion, sir. Kimi K3. This is only rolling out now, and will be covered next week. Muse Spark 1.1, the new Meta model. It is not frontier, but it is progress for them. Inkling, the first model from Thinking Machines. A call for regulatory action by Demis Hassabis, which I’ll cover soon. A new brief open letter call to action on AI regulation. That's on top of everything else, and an Opus 5 announcement is likely coming soon. The weekly once again got out of hand, so we’re splitting it once again into two, and once again saying we’ll be raising the bar for inclusion. And this time I mean it, as in enough to actually matter. Table of Contents Language Models Offer Mundane Utility. Whatever ye seek, ye shall find. Language Models Don’t Offer Mundane Utility. Gemini app needs some work. Language Models Upload Your Git Repository. Big problems [...] ---Outline:(01:17) Language Models Offer Mundane Utility(04:46) Language Models Don't Offer Mundane Utility(05:28) Language Models Upload Your Git Repository(08:35) Huh, Upgrades(09:30) Muse Spark 1.1(11:47) First Hit Free(15:36) On Your Marks(18:06) Choose Your Fighter(19:57) Get My Agent On The Line(23:23) Deepfaketown and Botpocalypse Soon(24:44) Fun With Media Generation(25:45) Copyright Confrontation(27:37) OpenAI Strikes Again(32:26) A Young Lady's Illustrated Primer(32:45) Recommendations for Policymakers(34:13) They Took Our Jobs(38:23) The Art of the Jailbreak(39:27) Get Involved(40:32) Introducing(41:12) In Other AI News(43:46) New Short Obviously True Statement About AI Just Dropped(46:03) Show Me the Money(46:21) The Lighter Side --- First published: July 16th, 2026 Source: https://www.lesswrong.com/posts/who9xZ7DxuprsJoTr/ai-177-part-1-tip-of-the-iceberg --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  33. 218

    “AI #177 Part 2: Wish You Were Here” by Zvi

    As usual, part 2 of the weekly deals with speculative, regulatory, political and alignment questions. Xi gave an important speech yesterday, so this post opens with that. There is talk that Kimi K3 is sufficiently strong that it upends many of these questions. It is clearly a candidate for another DeepSeek Moment, complete with stock drops for Google and SpaceX and (once again in a clear wrong-way move, the same as last time) Nvidia. Kimi K3 is clearly a very good model, exceeding expectations. Some are saying it is close to the frontier. The Artificial Analysis intelligence index has it at 57, a point ahead of Claude Opus 4.8, two behind Sol and three behind Fable. My presumption is that this number overstates its capabilities, but as always unless and until we have extensively tried the model ourselves, which I do not plan to do, we need to withhold judgment for at least a few days. I will be covering Kimi K3 in its own post at some point early next week. I have pushed further discussions involving Plan A and related issues into next week, as well as discussions around Demis Hassabis and Google [...] ---Outline:(01:29) Xi Gives A Good Speech on AI(15:33) Quiet Speculations(19:08) Tyler Cowen On Rebuilding The Future(22:39) The Quest for Sane Regulations(24:15) Wish You Were Here(26:55) The Week in Audio(27:29) New York Issues Moratorium On Data Centers(30:26) People Just Say Things(31:24) Rhetorical Innovation(39:29) Imagine Asking Questions(41:46) Anthropic Surveys Things It Calls Misalignment(54:30) Aligning a Smarter Than Human Intelligence is Difficult(58:30) The Most Forbidden Technique(01:01:04) Cooperative Alignment(01:12:38) The Lighter Side --- First published: July 17th, 2026 Source: https://www.lesswrong.com/posts/Zjj3PTEng8GDqfK6j/ai-177-part-2-wish-you-were-here --- Narrated by TYPE III AUDIO. ---Images from the article:<img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Zjj3PTEng8GDqfK6j/o2alkqdniqzoqveg8kxn" alt="This is not a tweet. Here is the description: Comic: stick figures argue about making a sandwich using "sudo."" style="max-width: 100%;" />Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  34. 217

    “Monthly Roundup #44: July 2026” by Zvi

    It's a quiet week so let's do the monthly right on schedule. Table of Contents Bad News. Good Advice. Opportunity Knocks. While I Cannot Condone This. Good News, Everyone. For Your Entertainment. Gamers Gonna Game Game Game Game Game. I Was Promised Flying Self-Driving Cars. Sports Go Sports. Antisocial Media. Government Working. Jones Act Watch. Highly Effective Altruism. Variously Effective Altruism. Ineffective Altruism. Prediction Markets. The Lighter Side. Bad News I wouldn’t have explained or modeled it quite the way Paola does here but the principle seems right to me. If people don’t trust you, or don’t trust people in general, that usually you can’t trust them either. Paola: I feel like a lot of human morality works like a prisoner's dilemma in that you can only trust others to behave morally to the extent that you believe they trust you to do the same. Due to this, I’ve come to view people with a bunch of social paranoia, distrust, etc. as *quite* dangerous to be around. And to be clear, I generally feel a [...] ---Outline:(00:17) Bad News(04:10) Good Advice(07:11) Opportunity Knocks(07:33) While I Cannot Condone This(11:35) Good News, Everyone(12:25) For Your Entertainment(17:35) Gamers Gonna Game Game Game Game Game(22:45) I Was Promised Flying Self-Driving Cars(26:20) Sports Go Sports(28:06) Antisocial Media(29:49) Government Working(33:55) Jones Act Watch(35:29) Highly Effective Altruism(40:32) Variously Effective Altruism(46:11) Ineffective Altruism(50:22) Prediction Markets(50:56) The Lighter Side --- First published: July 15th, 2026 Source: https://www.lesswrong.com/posts/KiCwcAGHx4rdwJgzD/monthly-roundup-44-july-2026 --- Narrated by TYPE III AUDIO. ---Images from the article:<img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/KiCwcAGHx4rdwJgzD/rnfh0ho7y7j4ukscmkti" alt="This image is not a tweet, so I'll describe it per the guidelines. This appears to be a WIRED website screenshot with article text. News article screenshot. The text reads: "In New Jersey, a lobbyist representing Uber took the strategy a step further, circulating legislative language that would, for a period of three years, require any platform offering driverless ride-hailing services to have human drivers serve 85 percent of its rides."" style="max-width: 100%;" /><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/KiCwcAGHx4rdwJgzD/zvrh6q1b4pj6byzzbpj9" alt="I notice the instructions contain a conditional directive at the top, but this image is not a tweet—it's a Tumblr post exchange. Following the applicable guidelines for social media posts: Tumblr posts joking about a "Schrödinger's catgirl" costume." style="max-width: 100%;" />Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  35. 216

    “Twitter Thoughts For You” by Zvi

    I previously have written back in March 2022 about how I use Twitter, and back in April 2023 about Twitter and its then-new algorithms, which have changed again. This post will update how I use Twitter now in 2026, and provide updates on the current state of the new algorithm, the situation with links, with the API, and some thoughts about using Twitter to make money which you almost never should try to do. Previously I said you need four things to use Twitter well: Tweetdeck or another similar alternative application. Knowing who to follow and read. Lists. Unfollows, filters, mutes and blocks. That hasn’t changed. Lists have become even more important. This post is coming out now, however, because the For You feed is perhaps making a comeback. Except where stated here, the advice in my 2022 post still applies. Table of Contents Defend Your Feed Via At Least One List. Block Early, Block Often, Know Your Triggers. Lists Change What Following Means. It (Wasn’t) For You. It's For You. Twitter Still Hates Links And That's Terrible. [...] ---Outline:(01:10) Defend Your Feed Via At Least One List(02:53) Block Early, Block Often, Know Your Triggers(03:35) Lists Change What Following Means(05:56) It (Wasn't) For You(07:34) It's For You(11:30) The Previous Time Twitter Transformed Its Algorithm Again(16:49) Twitter Still Hates Links And That's Terrible(27:35) Twitter Turns Its API Back On(31:02) Many Of The Bots Are Human(33:57) The Rise of Slop(36:20) Block Or Do Not Block(37:44) How To Make Money On Twitter(40:13) In Brief --- First published: July 14th, 2026 Source: https://www.lesswrong.com/posts/2GFyHmCLJYCag7gKh/twitter-thoughts-for-you --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  36. 215

    “Better Call Sol The Workhorse” by Zvi

    OpenAI's GPT-5.6-Sol is finally here, along with the cheaper Terra and Luna. We’ve seen the early hype as reported on Thursday, but as always that is biased. As usual, the bulk of this is collecting a gestalt based on reactions. I included everything up to a point, but I got a lot of feedback, so after a while I only took the interesting ones. Sol and Fable are both excellent models, sir. They both represent big moves forward. There is room in your workflow for both of them. Sol and Fable are very different, especially when considered as part of their respective packages. I’m considering Sol + Codex (or Work) versus Fable + Claude Code (or Cowork), throughout, in places where you wouldn’t use the chat interface. In terms of raw intelligence and ‘big model smell,’ and ability to do the hardest things that are intelligence-loaded, Fable still looks like it has a substantial edge. It also seems to be better aligned, or at least more trustworthy as an agent, with less tail risk. I still consider Fable ‘the best’ model, and the one that will require the most aggressive controls. I enjoy [...] ---Outline:(02:54) The Official Pitch(10:06) Sol Proposes A Proof Of The Double Cover Conjecture(10:49) The Official Benchmarks(14:26) Vend That Bench(16:37) Thinking Fast and Slow(18:35) Other People's Benchmarks(23:20) Have Robust Backups(26:33) That's Not What You Were Thinking(26:59) Helping Hands(27:36) Writing(29:15) Don't Stop Now(31:14) Sol Can Code And Do Math(32:37) Better Call Sol Cause You Can't Call Fable(33:34) Only Call As Much Sol As You Need(35:22) Positive Reactions(38:07) It's A Good Model, Sir(40:56) Negative Reactions(41:54) Sol The Workhorse(44:27) Pair Programmer(51:07) Pleased To Meet You(53:02) Sol Thinks You Better(54:02) My substantive posterior --- First published: July 13th, 2026 Source: https://www.lesswrong.com/posts/zPdDmJTovsKTvAiH2/better-call-sol-the-workhorse --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  37. 214

    “WSJ Article Claiming China Has Matched Anthropic Is Obvious Nonsense” by Zvi

    The Wall Street Journal printed an outright false headline and heavily misleading story claiming this, which of course was uncritically amplified by the usual suspects. I post this now on its own so that we have a place to link to, to explain the situation. Headline News WSJ Headline (Obvious Nonsense): ​China Has Matched Anthropic in Cybersecurity, Resetting AI Race. That. Did. Not. Happen. The post even claims, explicitly, that Claude Opus 4.8 similarly ‘matches’ Claude Mythos, a claim which is even more obviously false. Shame upon the Wall Street Journal. I fear Gell-Mann Amnesia. If they can get something as important as this so completely wrong, what about everything else? I am skipping over the parts that involve accurate reporting, or minor quibbles. It seems important to focus on clearly debunking the central false claims. Alas, the mistakes made here very much rhyme with mistakes being made throughout all this by the White House, and that get latched onto by certain bad actors, who have played a large part in leaving us unprepared for the Mythos Moment. For a full understanding of GLM-5.2, which is indeed an impressive [...] ---Outline:(00:27) Headline News(02:10) What Makes Mythos Special(03:18) Going Over The Detailed Claims(07:39) One Helpful Note(08:19) The Overall Impression Is Extremely Wrong(08:50) All Of This Has Happened Before And Will Happen Again --- First published: July 12th, 2026 Source: https://www.lesswrong.com/posts/2zSpuGJRk6EyjHAL6/wsj-article-claiming-china-has-matched-anthropic-is-obvious-1 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  38. 213

    “Introduction for and Reactions to Plan A” by Zvi

    Introducing Plan A The folks who brought you AI 2027, a so far remarkably accurate set of predictions despite those predictions having seemed freaky to many at the time, now bring you their positive vision that involves more freaky predictions: Plan A. These guys have rather strong prediction track records. In addition to AI 2027, among other things, Daniel Kokotajlo has What 2026 Looks Like (which is remarkably similar to what 2026 looks like) and Ryan Greenblatt, who is also the chief scientist at Redwood Research, was the #2 most accurate AI forecaster in 2025 out of 413 entries. Past performance is as always no guarantee of future success. If you’re the type to read at least some of my posts, or if you thought AI 2027 was worth reading, I recommend reading Plan A. There is also an unofficial visual novel version, for minds very different from my own who would want that. To be clear up front: I am not endorsing Plan A. I am not suggesting we should go off and try to enact Plan A as written. There is a lot more work to do and a lot of potential [...] ---Outline:(00:09) Introducing Plan A(01:42) You Only Get Five Words(08:44) Proactive Response To Objections(09:30) Initial Introductions and Endorsements(14:50) A Positive Vision(17:53) Alternative Plans(19:31) Plan S for Shutdown(21:35) Something (Unexpectedly Good) Ever Happens(23:10) Thus Selective Optimism(24:17) Quickly, There's No Time(25:20) Race Conditions(27:23) This Is A Lot Of Diffusion And Economic Growth(28:36) Living In China(30:35) Planning For Shifting Overton Windows Is Essential(31:51) The Standard Handwave(35:23) Some Equate Any Controls Over Compute To Authoritarian Dystopia And Those Same People Mostly Think Superintelligence Won't Happen(37:07) Vitalik Buterin Is Right, The Crux Is Future AI Capability Levels(44:45) The Authoritarian Objection(50:17) Concepts Of A Plan(52:03) The Kitchen Sink(53:13) Selective Claims Of Authoritarianism(55:12) You Either Can Steer The Future Or You Cannot(57:40) Cooperative Alignment --- First published: July 11th, 2026 Source: https://www.lesswrong.com/posts/z9tXCGogEgkgHSh8G/introduction-for-and-reactions-to-plan-a --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  39. 212

    “AI #176 Part 2: Plan B” by Zvi

    This is part 2 of the weekly, broadly covering speculation, rhetoric and policy, along with alignment research. This does not cover the release of GPT-5.6-Sol. As always, I will be taking a few days to digest what the new model has to offer and to allow others to try it and react. I will cover Sol and its capabilities early next week. I covered the GPT-5.6 system card back on June 28. This also does not cover the release of Plan A, the follow-up to AI 2027. This new scenario is a positive vision of what its authors think we should do going forwards. I do not endorse all of the recommendations or predictions of Plan A, but I do endorse reading Plan A and taking it seriously. Scott Alexander, one of those who worked on it, writes an introduction and justification here. I will have full coverage soon. Table of Contents Quiet Speculations. Will our AI regulations be ad hoc indefinitely? The Goalposts Are Dyson Spheres. This might take a little longer. People Just Say Things. Three Pills. Unpilled, AI, AGI, ASI. The Quest for Sane Regulations. You’ve [...] ---Outline:(01:08) Quiet Speculations(04:41) The Goalposts Are Dyson Spheres(08:11) People Just Say Things(08:40) Three Pills(14:52) The Quest for Sane Regulations(16:35) OpenAI National Security Principles(23:45) Greetings From The Department Of War(25:47) Chip City(27:02) Open Weight Models Are Unsafe And Nothing Can Fix This(34:37) Their AI Propaganda Bots(36:39) Rhetorical Innovation(38:23) You Learn(40:00) You May Be Tan And Thin And Rich But You're a Tool(44:21) Train Those Thoughts(45:16) Train Out Those Thoughts(50:02) My Own Private Idaho(51:46) Aligning a Smarter Than Human Intelligence is Difficult(57:04) No Space Like J-Space(01:06:50) Cooperative Alignments(01:07:15) The Lighter Side --- First published: July 10th, 2026 Source: https://www.lesswrong.com/posts/7Av7wwErwNXdM3J68/ai-176-part-2-plan-b --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  40. 211

    “AI #176 Part 1: Doing It Live” by Zvi

    Enough things added up that this week is getting split into two parts. Then on Monday, if all goes as I expect, we’ll cover OpenAI's Sol, aka GPT-5.6. OpenAI also gave us an upgraded voice mode, which I haven’t tried out but early reports are that it is a step change. AI writing, especially Claude writing, is becoming more prominent and harder not to notice, and increasingly a tough read when encountered in the wild. Does anyone care? Or are those who care the weird ones here? This week saw an excellent paper, which I cover in No Space Like J-Space. Technically we also got Grok 4.5. Table of Contents Language Models Offer Mundane Utility. A whole new world. Language Models Gain Unexpected Affordances. Wait, you can just do that? Language Models Don’t Offer Mundane Utility. Things get old. Pay The Man His Money. You have a few more days with marginally free Fable. Huh, Upgrades. Anthropic raises API platform limits. Grok 4.5 Exists. It might be okay for its price. F*** It We’re Doing It Live. OpenAI gives us a big upgrade to voice [...] ---Outline:(00:54) Language Models Offer Mundane Utility(03:21) Language Models Gain Unexpected Affordances(06:12) Language Models Don't Offer Mundane Utility(06:54) Pay The Man His Money(07:37) Huh, Upgrades(07:45) Grok 4.5 Exists(09:12) F\*\*\* It We're Doing It Live(10:24) On Your Marks(11:13) Better Call Sol(22:40) Get My Agent On The Line(23:29) Deepfaketown and Botpocalypse Soon(25:33) Fool Me Twice(27:23) I Like Your Style(33:28) Enough With That Style(36:58) Fun With Media Generation(38:51) Copyright Confrontation(39:30) Cyber Lack of Security(39:52) A Young Lady's Illustrated Primer(45:56) They Took Our Jobs(51:07) Get Involved(52:10) In Other AI News(56:44) Show Me the Money(57:09) Bubble, Bubble, Toil and Trouble --- First published: July 9th, 2026 Source: https://www.lesswrong.com/posts/M9eLyMsH5DLjMYL86/ai-176-part-1-doing-it-live --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  41. 210

    “Childhood and Education #20: Phones and Screens” by Zvi

    We have a respite, so I thought I’d tackle various thoughts on children, phones and screens. GPT-5.6-Sol drops tomorrow, and the Fable agents are hard at work. I’ll start with the other screens, then finish with the phones. Table of Contents EdTech. NonEdTech. Do Not Ban Social Media Outright. Some Modern Kids Media Is Pretty Great. Ban Phones In Schools (1). Your Offer Is Acceptable. Ban Phones In Schools (2). Screen Time. Inappropriate Content. EdTech Increasingly, when you pick a school, you are picking EdTech. The school will put your child on a tablet or computer, and expect them to learn that way. In theory, with sufficient assistance and bespoke design and incentive structures, this is The Way. It sure seems way better than ‘sit and listen to a lecture.’ I am especially excited for Alpha School's version of this, with its bespoke designs and high level of both expectations and continuous human support. Alas, most people are getting a much worse version, that is much worse than what you could easily improvise at home. I’m less concerned with ‘EdTech provider [...] --- First published: July 8th, 2026 Source: https://www.lesswrong.com/posts/XA8fMCnwuc45uZYHX/childhood-and-education-20-phones-and-screens --- Narrated by TYPE III AUDIO.

  42. 209

    “No Space Like J-Space” by Zvi

    There is a new very cool Anthropic paper: Verbalizable Representations Form a Global Workspace in Language Models. You can read the blog post verison here. I encourage reading of the whole original blog post or paper, if you have the time. Table of Contents Through A Different Lens. Establishing J-Space As A Global Workspace. Are You Pondering What I’m Pondering? Assistant J. The Power Of Virtuous Thinking. High Praise. Everyone Remains Confused About Consciousness. Further Research. Don’t Think. Through A Different Lens They call this discovered area of ‘conscious access,’ where things are available for the model to do what in humans we would call conscious reasoning, the ‘J-space,’ after a new interpretability technique called the Jacobian Lens. The Jacobian Lens computes, for each layer, the average causal effect of changes in the residual stream on the model's eventual outputs, averaged across a wide variety of contexts. Then you can trace what concepts are associated with each layer as the model proceeds through. At each layer, the J-lens vectors form an overcomplete set. … We observe that only a relatively small [...] ---Outline:(00:26) Through A Different Lens(02:02) Establishing J-Space As A Global Workspace(04:55) Are You Pondering What I'm Pondering?(08:50) Assistant J(09:42) The Power Of Virtuous Thinking(17:08) High Praise(22:37) Everyone Remains Confused About Consciousness(31:17) Further Research(34:26) Don't Think --- First published: July 7th, 2026 Source: https://www.lesswrong.com/posts/EnxHPxJT4Xin5cTsX/no-space-like-j-space --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  43. 208

    “Fable #6: The Return of the King” by Zvi

    The blip is over. We have Fable back. Utah teapot: happy fable/mythos easter Wednesday, to those who celebrate Here is the official letter restoring Fable, great job everyone. Notice it is addressed to Tom Brown, not to Dario Amodei. Anthropic had to make the controls more stupid for now, but this is a big win. j⧉nus: YES!!! I’m really proud of Anthropic for their successful negotiation with the government. Also positive update on the government being sane and possible to cooperate with. Afaik Anthropic didn’t need to agree to any bad terms / genuflect / betray their principles or dignity. The fiasco continues, at least until such time as we have a systematic regime in place for future frontier models rather than decisions being made ad hoc, by people like Lutnik and Bessent who do not know how any of this works. The Blip Anthropic explains its version of what happened. Here is the timeline: Amazon researchers discover they can ask Fable to ‘fix this code.’ They alert the White House, which freaks out. June 12: US government tells Anthropic to take down Fable on its own. [...] ---Outline:(01:33) The Blip(08:08) The White House Explanation(10:06) Everything Remains Ad Hoc(10:37) Take What You Can Get(12:18) The Problem Is Real(13:16) GLM-5.2 Being Frontier Remains Obvious Nonsense(16:29) Mythos Might Be Smarter Than You Are(19:10) Let The Record Reflect(20:41) Stationary Bandits(25:02) Use This Window Well --- First published: July 3rd, 2026 Source: https://www.lesswrong.com/posts/r9HsHHSsfABhhxnYr/fable-6-the-return-of-the-king --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  44. 207

    “AI #175: The Fable Continues” by Zvi

    Fable's back. Back again. Fable's back. Tell a friend. Use your free week to its fullest. This is excellent news. The blip only lasted a few weeks. It was still a fiasco, and we have to deal with the fallout. Our system remains fully ad hoc. The precedent has been set that we may use export controls on models, or order them taken down on 90 minutes of notice based on a misunderstanding. At least some amount of counterproductive additional locking down has occurred to address Amazon's little demonstration and reassure the government. And for now GPT-5.6 remains in limbo, awaiting its verdict, while OpenAI talks about giving away 5% of the company as tribute. I’ll cover that continuing situation on its own. Whereas the weekly post is about everything else happening in AI this week. Table of Contents Language Models Offer Mundane Utility. Exploratory science. Language Models Offer Mundane Utility You May Not Want. Google sees all. Language Models Don’t Offer Mundane Utility. Too dumb to get smart. Huh, Upgrades. GLM-5.2 faster, Nana Banana Lite 2, Claude Desktop on Linux. On Your Marks. Remote labor index shoots [...] ---Outline:(01:08) Language Models Offer Mundane Utility(02:29) Language Models Offer Mundane Utility You May Not Want(04:32) Language Models Don't Offer Mundane Utility(05:59) Huh, Upgrades(06:33) On Your Marks(09:28) Get My Agent On The Line(14:29) Deepfaketown and Botpocalypse Soon(14:58) Cyber Lack of Security(16:11) On Writing(21:34) You Drive Me Crazy(24:26) They Took Our Jobs(30:32) Get Involved(30:58) Introducing(31:21) In Other AI News(32:45) Show Me the Money(33:12) Bubble, Bubble, Toil and Trouble(35:31) Quiet Speculations(39:37) Glorious AI Future(43:37) Three Pills(44:58) The Anthropic Economic Index(46:29) Leader Of The PAC(47:48) Theory Of The AI Firm(49:01) Chip City(50:32) The Week in Audio(53:29) People Really Hate AI(56:31) Rhetorical Innovation(01:00:31) The First Rule Of Functional Decision Theory Is(01:03:40) Aligning a Smarter Than Human Intelligence is Difficult(01:06:40) Names Have Power(01:07:40) Cooperative Alignment(01:15:21) People Just Say Things(01:16:46) Escape From The Permanent Underclass(01:29:51) Other People Are Not As Worried About AI Killing Everyone(01:31:07) The Lighter Side --- First published: July 2nd, 2026 Source: https://www.lesswrong.com/posts/WNvBxtbHuLreFe7af/ai-175-the-fable-continues --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  45. 206

    “Claude Sonnet 5 Is Not Frontier But Has Its Uses” by Zvi

    Fable 5 is back today, baby! Premium subscribers have one week to use it within their subscriptions. First hit's free. Then you pay by the token. Today's post is still about Sonnet 5. I don’t know that there will be much call for Sonnet 5 for most purposes, given Opus 4.8 exists and especially now that Fable 5 is once again available, but this is what we do here, so sure, why not, system card time, including model welfare, after which we’ll do capabilities. Sonnet costs $3/$15 per million tokens, versus $5/$25 for Opus and $10/$50 for Fable, after an introductory period. Once you pay for all the tokens you need you’re not really saving money, such as on the ArtificialAnalysis index where Sonnet ended up being more expensive. My initial impression is that if you want me to use Sonnet over Opus for most purposes, you’re going to have to offer a bigger discount than that. The counterargument is speed. Sonnet 5 is faster without being that much less capable. In many cases, getting into a flow state like that is pretty valuable. There are a few agentic scenarios Sonnet 5 has [...] ---Outline:(02:29) Mythos Exists(03:10) Introduction (1)(03:17) RSP Evaluations (2)(04:02) Cyber (3)(04:26) Safeguards and Harmlessness (4)(04:57) Agentic Safety (5)(06:42) Alignment (6)(10:45) Illegible Thinking (6.4.5)(11:43) Evaluation Awareness(12:22) Honesty and Hallucinations (6.5)(13:17) Flagged As Unhealthy? (6.5.1)(13:53) Model Welfare (7)(20:32) Live From AI Village(22:19) For I Contain Multitudes(29:04) Official Benchmarks(33:11) Other People's Benchmarks(33:37) Positive Reactions(39:04) Negative Reactions --- First published: July 1st, 2026 Source: https://www.lesswrong.com/posts/d9pmwQsFC2AXceryg/claude-sonnet-5-is-not-frontier-but-has-its-uses --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  46. 205

    “The Once And Future Fable #5” by Zvi

    We, or at least ‘more than 100 American institutions,’ got Mythos back this week. What we the people do not have is Fable or Sol. While we wait for both Claude Fable 5 and GPT-5.6-Sol, today we instead got Claude Sonnet 5. As usual it will take a few days to get a handle on the new model. In this case, Anthropic is representing it as a cheaper and faster version of Opus 4.8, so even though the number says 5 this is a relatively minor development. This post expands the Fable series to cover all further developments this week surrounding the Mythos Moment, and the various aspects of handling our new ad hoc licensing regime and figuring out policy going forward, and other aspects of policy as well. This includes my notes on various rhetoric being pulled out, where I fear I end up saying similar things every so often, because we are doomed to repeat the cycle. I have accepted my role in that, but those are sections many of you can skip, and are marked in italics accordingly as per usual. Table of Contents You Should See The Other [...] ---Outline:(01:12) You Should See The Other Guy(01:54) DeepMind Coders Of The World, Unite(02:45) Report Your Incidents(03:04) Good Guy With An AI(04:49) Free As In To Give It A Shot(08:21) Everything Is Both Speech And Computer(11:10) Lambs To The Slaughter(15:43) A Sign Saying Beware Of The Leopard(16:52) The Once And Present Mythos(20:26) What Is To Be Done(24:00) Distillation(26:24) What Would Banning Open Source Even Mean(27:20) Open Weight Models Are Unsafe And Nothing Can Fix This --- First published: June 30th, 2026 Source: https://www.lesswrong.com/posts/phxgfwGNGbanumMMv/the-once-and-future-fable-5 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  47. 204

    “WSJ Article Claiming China Has Matched Anthropic Is Obvious Nonsense” by Zvi

    The Wall Street Journal printed an outright false headline and heavily misleading story claiming this, which of course was uncritically amplified by the usual suspects. I post this now on its own so that we have a place to link to, to explain the situation. Headline News WSJ Headline (Obvious Nonsense): ​China Has Matched Anthropic in Cybersecurity, Resetting AI Race. That. Did. Not. Happen. The post even claims, explicitly, that Claude Opus 4.8 similarly ‘matches’ Claude Mythos, a claim which is even more obviously false. Shame upon the Wall Street Journal. I fear Gell-Mann Amnesia. If they can get something as important as this so completely wrong, what about everything else? I am skipping over the parts that involve accurate reporting, or minor quibbles. It seems important to focus on clearly debunking the central false claims. Alas, the mistakes made here very much rhyme with mistakes being made throughout all this by the White House, and that get latched onto by certain bad actors, who have played a large part in leaving us unprepared for the Mythos Moment. For a full understanding of GLM-5.2, which is indeed an impressive [...] ---Outline:(00:27) Headline News(02:09) What Makes Mythos Special(03:16) Going Over The Detailed Claims(07:38) One Helpful Note(08:18) The Overall Impression Is Extremely Wrong(08:48) All Of This Has Happened Before And Will Happen Again --- First published: June 29th, 2026 Source: https://www.lesswrong.com/posts/bpBYm5jiS4tpyzuDS/wsj-article-claiming-china-has-matched-anthropic-is-obvious --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  48. 203

    “GPT-5.6: The System Card” by Zvi

    While we wait for a general release, the system card is the best hint as to what is going on with the new candidate for America's Next Top Model, GPT-5.6. This is only an OpenAI model card, so by my standards it's a light read. There's a lot of things that you get in an Anthropic card, that are missing in an OpenAI card. Overall, the card gives a clear and consistent impression that GPT-5.6-Sol is a substantial improvement over GPT-5.5, but still short of Mythos. OpenAI calls it a ‘step function better’ than GPT-5.5. That seems accurate. OpenAI: Sol is our new flagship and a step function better than GPT-5.5. Terra delivers performance competitive to GPT-5.5 at 2x lower cost. Luna is our most cost-efficient model, delivering strong capability at our lowest cost. Together, the GPT-5.6 family gives people and developers more choice in how they balance intelligence, speed, and cost. Once available, pricing for GPT-5.6-Sol will be $5/$30, the same as GPT-5.5. Terra is $2.5/$15, Luna is $1/$6. They claim it will be on Cerebras at 750 TPS, which is insanely fast. Capacity will be limited, at least at first. [...] ---Outline:(03:49) What's In A Name?(04:26) Fix This Code(07:08) Crossover Event Requested(07:43) Disallowed Content (3)(09:03) Avoiding Accidental Data-Destructive Actions (3.3)(09:29) Are You Sure? (3.4)(09:58) Jailbreaks (4.1)(10:14) Prompt Injection (4.2)(10:40) HealthBench (5.1)(11:00) Dynamic Mental Health Adversarial User Simulations (5.2)(12:21) Hallucinations (6)(12:50) Isolated Misaligned Actions (7.1)(13:10) Going Overboard (7.2)(18:11) Chain of Thought Evaluations (7.3)(19:18) Bias (8)(19:27) Preparedness (9)(20:15) Biological Risks (9.1.1)(22:15) Cybersecurity (9.1.2)(28:40) External Cyber Evaluation FrontierCyber from Irregular (9.1.2.5)(30:32) Cyber Conclusions(31:07) Recursive Self-Improvement (9.1.3)(32:22) METR Warns Us (9.1.3.6)(35:04) Everything Is Under Control(37:44) Metagaming (7.4)(40:17) Apollo Research and Sandbagging(43:09) Safeguards (9.3)(50:01) Better Not Call Sol Yet The original text contained 2 footnotes which were omitted from this narration. --- First published: June 28th, 2026 Source: https://www.lesswrong.com/posts/JFjNmPTbH8kL6xtp6/gpt-5-6-the-system-card --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  49. 202

    “AI #174: You’re It” by Zvi

    Fable remains in limbo, with renewed hope that we will get it back soon (45% by tomorrow, 69% by July 1, nice.) The full capabilities post is now available. Alex Bores unfortunately lost narrowly in NY-12, and will not be heading to Congress. There are also plenty of other stories to cover. Some highlights: GLM-5.2 is the new best open model, although it is expensive for its class. It will have its uses, potentially for agents you need to run fully locally or privately, but often it won’t be the right fit. Claude Tag is a new system for having Claude join your Slack, and if you @ him then he will spin up an instance to do the coding work. Dean Ball is joining OpenAI to work on policy. We don’t see eye to eye on everything, but this is a huge upgrade over their existing alternatives. The debate over the MidJourney scanner continues. Table of Contents Language Models Offer Mundane Utility. You know what it is for. Language Models Don’t Offer Mundane Utility. Hiring French Qwants. Huh, Upgrades. Claude Code supports artifacts. [...] ---Outline:(01:12) Language Models Offer Mundane Utility(02:58) Language Models Don't Offer Mundane Utility(03:13) Huh, Upgrades(03:38) On Your Marks(04:36) Deepfaketown and Botpocalypse Soon(11:20) Fun With Media Generation(12:20) Cyber Lack of Security(14:49) Overcoming Bias(15:52) A Young Lady's Illustrated Primer(18:14) They Took Our Jobs(19:48) Get Involved(21:54) Introducing(22:12) Claude Tag(31:46) In Other AI News(33:20) More On GLM-5.2(35:17) ChatGPT Health(37:04) Middle Of The Journey(51:04) New Medical Diagnostic Just Dropped(54:05) Google on AI Control(01:02:12) The Once And Future Fable(01:04:17) Fable: The First Lawsuit(01:05:12) Dean Ball Joins OpenAI(01:09:03) Show Me the Money(01:09:18) Quiet Speculations(01:12:00) Alex Bores Loses In NY-12 By 4%(01:22:28) The Quest for Sane Regulations(01:24:49) Chip City(01:28:33) The Week in Audio(01:29:21) People Just Say Things(01:30:19) Rhetorical Innovation(01:36:32) There Are Two Pills(01:37:55) Who Evals The Evals(01:39:02) Aligning a Smarter Than Human Intelligence is Difficult(01:43:17) Cooperative Alignment(01:44:22) People Are Worried About AI Killing Everyone(01:45:59) Other People Are Not As Worried About AI Killing Everyone(01:48:08) The Lighter Side --- First published: June 25th, 2026 Source: https://www.lesswrong.com/posts/MfdaizeH8z8civPHe/ai-174-you-re-it --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  50. 201

    “The Once And Future Fable #4” by Zvi

    It does look good, actually. After the odds had dropped quite a bit, they’re looking good again, with a 60% chance of restoration by July 1 and 88% by July 31, in the wake of groundwork looking like it is being laid in various places: leo: BREAKING: Claude Code v2.1.190 introduces several string changes that hint at preparations for a Fable 5 return, with it being permanently included in subscriptions with weekly usage. The string “You’ve used your Fable 5 usage for this week” has been added, and “purchased separately from your plan” has been removed leo: UPDATE: Fable 5 has now reportedly also reappeared in Amazon Bedrock If the update is based purely on the above info I would treat the new odds as overconfident. These moves seem reasonable to make even if you have no confidence in the restoration, in order to be ready if that moment arrives. This also suggests a potential permanent quota for Fable for subscribers. Even a modest amount is a big game here, since even a modest allocation means you can use it for non-coding tasks or minor coding tasks within the subscription. With that [...] ---Outline:(01:42) A Rather Terrible Policy(03:14) The People Have Spoken(03:59) Thank You, Next(06:35) Be Very Very Quiet(07:18) What These Babies Can And Cannot Do(13:54) What's The Worst That Could Happen?(25:09) The Data Retention Policy Is About Defense In Depth(25:51) Pick Up The Phone(28:49) People Just Say Things --- First published: June 24th, 2026 Source: https://www.lesswrong.com/posts/xJMngE34AfwGWvKLx/the-once-and-future-fable-4 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Type above to search every episode's transcript for a word or phrase. Matches are scoped to this podcast.

Searching…

We're indexing this podcast's transcripts for the first time — this can take a minute or two. We'll show results as soon as they're ready.

No matches for "" in this podcast's transcripts.

Showing of matches

No topics indexed yet for this podcast.

Loading reviews...

ABOUT THIS SHOW

Audio narrations of LessWrong posts by zvi

HOSTED BY

zvi

Frequently Asked Questions

How many episodes does LessWrong posts by zvi have?

LessWrong posts by zvi currently has 50 episodes available on PodParley. New episodes are automatically indexed when they're published to the podcast feed.

What is LessWrong posts by zvi about?

Audio narrations of LessWrong posts by zvi

How often does LessWrong posts by zvi release new episodes?

LessWrong posts by zvi has 50 episodes. Check the episode list to see recent publication dates and frequency.

Where can I listen to LessWrong posts by zvi?

You can listen to LessWrong posts by zvi on PodParley by clicking any episode. We provide an embedded audio player for direct listening, and you can also subscribe via your preferred podcast app using the RSS feed.

Who hosts LessWrong posts by zvi?

LessWrong posts by zvi is created and hosted by zvi.
URL copied to clipboard!