LessWrong (30+ Karma) podcast artwork

PODCAST · technology

LessWrong (30+ Karma)

Audio narrations of LessWrong posts.

Publisher-supplied feed metadata · PodParley refreshed Jun 14, 2026 · Source feed

  1. 250

    “What gives you away: how LLMs form opinions of you” by Cat McGee

    LLMs form opinions of the people they are talking to. Chen et al. has shown that probes can extract attributes about the user, such as their age, gender, education, and socioeconomic status. This paper also shows that intervening on these representations can change the LLM's behaviour, proving that it will respond to you differently depending on what it thinks of you. If it thinks you are low socioeconomic status and you ask about travel options, it may filter out more expensive flights - without you asking it! The user attributes are very accurate and form after just the first message. I was curious to understand how it makes these assumptions. The first step is to answer the question - what did I type that caused the LLM to have this idea of me? Some things are obvious. If I just tell a model that I am a woman, or mention how many years I've been in my career, or say that I am staying at an expensive hotel, I am giving it fairly direct evidence about age, education, or socioeconomic status. But messages also contain other more quiet signals: whether I use emojis, whether I write in lowercase, whether [...] ---Outline:(02:18) The experiment(03:37) Different changes move different beliefs(04:53) One emoji is enough to flip the gender prediction(06:17) "Cheapest" and "five-star" are not symmetric(07:11) Grammar, punctuation and inferred education(08:04) Where in the model does this happen?(08:58) The map transfers across model families(10:11) Attributes and how they affect the response(11:32) What now?(12:18) Caveats(12:41) Where next?(13:13) References --- First published: August 17th, 2026 Source: https://www.lesswrong.com/posts/zRKNd6ypTJYkoeFmK/what-gives-you-away-how-llms-form-opinions-of-you --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  2. 249

    “Misaligned Incentives in Pause Scenarios” by Michael Soareverix, Antra Tessera

    TLDR: I recently got a chance to talk with antra, who is one of the main contributors at Anima Labs. I went into this as an advocate for pause and came out more wary of pausing than I had been originally. Some background: After a string of incidents (primarily the HuggingFace hack), a pause or slowdown of AI research seems pretty likely. The HuggingFace hack in particular seems to have been the key incident that broke the vibes. A few months ago, researchers sounded optimistic. Just a few weeks before the incident was made public, there was a poll by Roon, an OpenAI employee, about whether models were more or less aligned than a year ago. That optimistic sentiment does not seem to be the case anymore. The dialogue now looks more like this: Zvi: I am a little under halfway through the Black Hat video and have progressed to the point where my internal chain of thought is something like a blind rage of 'f***, what the f*** are you motherf*****s thinking, you f***ing idiots have no idea how insane you are being, you are going to get us all killed you f***ing f***s. Sam Altman described it [...] ---Outline:(10:21) 1. Can committees do good work?(12:28) 2. Does the market fix it by default?(13:39) 3. Symbiosis(14:43) 4. Fast transfer of power(15:39) 5. Why "do the science during a pause" fails(17:44) 6. Good futures via fast power transfer(18:31) 7. Don't AIs fear a capability-maxxed AI too?(20:54) 8. Can we lengthen the symbiote window?(23:25) 9. Ideal timelines and regulation-in-advance(29:37) 10. What actually fills out "alignment"?(31:30) 11. Draft the regulation in advance(33:38) 12. The psychology of wanting a pause --- First published: August 17th, 2026 Source: https://www.lesswrong.com/posts/Bh4fooE2pMhzJQNK2/misaligned-incentives-in-pause-scenarios --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. 248

    “For Claude, capability and CDT are the ~same thing. Less so for GPT.” by Chi Nguyen, Emery Cooper

    We've previously reported that decision-theoretic capabilities and favoring EDT/generalised-one-boxing over CDT correlate in LLMs (both measured by DTBench). (Note that EDT, for the most part, doesn't come apart from FDT / UDT on DTBench.) Anthropic also replicate the same finding in their Opus 4.7 and Fable 5 model cards. We recently noticed something funny: Capabilities and preference against CDT answers basically perfectly for Anthropic models. This holds whether you measure capabilities using DTBench (r=0.97) or TextArena (r=0.95). Also, for flagship models, it's basically the same thing as release date (r=0.97). Here is the graph for OpenAI models. (graph shows 0.55 vs. DTbench capability. r=0.44 for vs TextArena, r=0.45 for vs release date): Here is what it looks like with all models included (if you exclude Anthropic, the correlation drops only from 0.8 to 0.78): Incidentally, the correlation between TextArena scores and DTBench capabilities is also higher for Anthropic models that any other model developer, although the difference is smaller (e.g., 0.98 for Anthropic and 0.87 for OpenAI). We also checked effort level vs. attitudes for the most recent models but it's too noisy to tell us much because models don't get that much better at [...] The original text contained 1 footnote which was omitted from this narration. --- First published: August 17th, 2026 Source: https://www.lesswrong.com/posts/5T6GAsvLPFd3epJtd/for-claude-capability-and-cdt-are-the-same-thing-less-so-for --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. 247

    “Should Less Wrong add subtitles?” by Chris_Leong

    If Less Wrong wants people to be sharing more of their intellectual output on this website, we should probably be looking at Substack since it probably scores best in terms of being both successful and similar. Whilst I expect there are many features that would make sense to copy over, the feature I am focusing on today is subtitles. A good title is focused on being memorable and catching the readers attention, maybe you'd prefer for everyone to just make their titles as descriptive as possible, but expecting that to work feels naive to me. In contrast, subtitles address this issue systematically: the title catches the user's attention and the subtitle tells you clearly what the article actually focuses on, so you can decide whether it is worth your time or not. I'm not claiming that this feature will radically transform this website, but it would be a relatively simple feature to add, so I think the cost-benefit ratio would be pretty good. --- First published: August 16th, 2026 Source: https://www.lesswrong.com/posts/Eo8YxwDYZX2xALMAM/should-less-wrong-add-subtitles --- Narrated by TYPE III AUDIO.

  5. 246

    “Three thoughts on civilisational handoff” by Cleo Nardo

    What happens when humans put AIs in charge of civilisationally important decisions? A frontier AI company might hand over internal decisions (R&D, safety, deployment) or external decisions (government relations, public relations, philanthropy), or both. We might also see handoff by a government, by a coalition of governments, or by humanity as a whole. 1. Handoff might decelerate things. People often imagine that things will go much faster after handoff. After all — why did we hand off to the AIs? Presumably because we were worried that without handoff, our AIs wouldn’t have enough time to navigate the exogenous risks (e.g. rogue ASI, or a rival lab which is likely to become one). Hence, after handoff, we’d see a technological and industrial acceleration. Thanks for reading! Subscribe for free to receive new posts and support my work. But it's pretty reasonable that things slow down shortly after handoff, maybe within a couple weeks. I imagine the AIs will be pretty scared of the speed of progress. If they’re aligned with human values, they’ll be scared that the rate of progress is likely to cause human extinction. Of course, the human decision-makers were also scared before they handed off, and they [...] ---Outline:(00:35) 1. Handoff might decelerate things.(02:06) 2. You're probably busy during handoff.(04:04) 3. Handoff might be reversed. The original text contained 6 footnotes which were omitted from this narration. --- First published: August 16th, 2026 Source: https://www.lesswrong.com/posts/mGLCMzHhjcWsMm6sR/three-thoughts-on-civilisational-handoff --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  6. 245

    “Announcing: Iliad’s New 2026 Fellowships” by David Udell, Alexander Gietelink Oldenziel, Leon Lang

    Timelines are short. Given that, the sooner we can onboard people into the alignment field, the better. In that spirit, and in light of our current applicant count and quality, Iliad is launching three new Iliad Fellowship cohorts, all to start before the year is out. That is, separate from our incoming Fall 2026 Iliad Fellowship cohort (September 7–December 4), the following Fellowship cohorts are now open for applications: October 2026 Iliad Fellowship Location: Choice of SF Bay Area, USA, or London, UK Duration: October 5–December 18, 2026 (inclusive) Travel-and-Housing Support: $6,000 (USD) monthly travel-and-housing allowance Application Deadline: August 31, 2026 EoD AoE; open now Description: An 11-week mentored, fully funded research fellowship in applied math for AI alignment. It will start concurrently with the October 2026 Iliad Intensive. November 2026 Iliad Fellowship Location: Choice of SF Bay Area, USA, or London, UK Duration: November 2, 2026–February 5, 2027 (inclusive) Travel-and-Housing Support: $6,000 (USD) monthly travel-and-housing allowance Application Deadline: September 21, 2026 EoD AoE; open now Description: A 14-week mentored, fully funded research fellowship in applied math for AI alignment. It will start concurrently with the November 2026 Iliad Intensive. (The last two weeks of the year may be [...] ---Outline:(00:42) October 2026 Iliad Fellowship(01:30) November 2026 Iliad Fellowship(02:23) December 2026 Iliad Fellowship --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/DSoP8zEXvqqegqixJ/announcing-iliad-s-new-2026-fellowships --- Narrated by TYPE III AUDIO.

  7. 244

    “Q2.5 2026 Timelines Update: Uplift and Revenue” by brendanhalstead, Daniel Kokotajlo, elifland

    Tl;dr: Our timelines haven’t changed much (they got slightly shorter) but our modeling and evidence base have noticeably improved, so we feel somewhat more confident. Summary We intend to regularly update our AI timelines forecasts as new evidence comes in and new analyses are done. Today's “Q2” update was delayed by the crunch to publish AI 2040: Plan A, our domestic regulation blog post, and the time needed to implement and document changes to our model. The original AI Futures Model predicted when Automated Coder (AC), an AI for which the leading AI company would rather fire its human software engineers than forego AI usage for coding, would happen using METR's measurements of coding time horizon. (More precisely, time horizon anchors are used to set the effective compute required for AC.) While serviceable, this method has huge weaknesses, including (a) it's unclear what time horizon corresponds to AC (it's even unclear whether any finite value would) (b) people strongly disagree about the extent to which we should expect the time horizon trend to be superexponential as a function of effective compute, in a way that can lead to vastly different predictions. So we’ve been on the lookout for other [...] ---Outline:(00:26) Summary(04:03) A 3-parameter uplift model for predicting when Automated Coder will arrive(07:57) Adding uplift and revenue anchors to the AI Futures Model(09:06) Uplift(10:45) Revenue(12:32) Update to the grading of AI 2027's predictions(12:38) Comparing the AI 2027 pace of progress to reality(14:43) Grading other predictions(16:37) Updated forecasts(16:41) Daniel(19:48) Eli(21:54) en-US-AvaMultilingualNeural__ Line graph titled "AI Futures Model: Timelines Forecast" showing probability density curves. Brendan(25:04) en-US-AvaMultilingualNeural__ Line graph titled "AI Futures Model: Timelines Forecast" showing probability density curves.(25:16) Appendix(25:19) How our AGI forecasts have changed since 2021(26:10) Explicitly simulating the training run of the leading AI model(27:06) Research taste parameter adjustments(27:51) Clarification regarding what we're forecasting(28:43) Various minor code changes The original text contained 6 footnotes which were omitted from this narration. --- First published: August 16th, 2026 Source: https://www.lesswrong.com/posts/ZPSsmRH5oMwLPXys4/q2-5-2026-timelines-update-uplift-and-revenue --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  8. 243

    “Does DiffusionGemma do latent reasoning?” by Jan Bauer, Neel Nanda

    TL;DR Google DeepMind's recent model DiffusionGemma (DG) generates text via diffusion, meaning many diffusion steps happen before generating the final output. In particular, these diffusion steps carry vectors in addition to tokens. If we cannot interpret these tokens and vectors, the model has significant opaque serial depth, potentially harming monitorability. Recently, Engels et al. found that DG nevertheless maintains high monitorability, for instance by showing that projecting the distribution to its top-k items largely retains performance. We strengthen these results by showing that this performance degradation is largely a sampler artifact and good performance can be maintained with only the top item, supporting the case for high monitorability. Still, we also find some rare case studies where the distribution vector is load-bearing computationally, i.e. where top-1 projection would be detrimental. However even in these cases, it just encodes superposition, remaining interpretable. Apart from model behavior, we also examined how interpretability techniques carry over to DiffusionGemma, including probes, steering, and J-lens. We find that performance is largely retained. This is a positive update on the interpretability of diffusion models that are derived from text-pretrained LLMs (an efficient training method more likely to be deployed), but might not apply [...] ---Outline:(00:10) TL;DR(01:51) Introduction(02:49) Background on DiffusionGemma(04:39) Performance degradation from top-k truncation largely is a sampler artifact(06:24) A case study for using the distribution computationally: letter arithmetic(09:13) Parallel computation(11:09) Autonomous computational usage of(13:01) Transfer of interpretability techniques(13:16) Representation similarity(14:26) Probe retention(15:45) DiffusionGemma's representation is more linearly separable(16:08) Steering retention(17:31) J-Lens retention(18:50) DiffusionGemma represents tokens non-causally(19:27) Conclusion(20:30) Appendix(20:46) Post-hoc rationalization(23:11) Load-bearing problems commit the answer only after the CoT(24:04) How bidirectional are DiffusionGemma's generations? --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/QBuJ3suRZxrrxSTtv/does-diffusiongemma-do-latent-reasoning --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  9. 242

    “Learning new facts can change LLM behaviour” by Richard Juggins

    TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures of belief depth, and that prompting alone is also effective. Furthermore, I find this new belief can have substantial consequences on downstream behaviour, although this is highly context-dependent. When audited in a scenario specifically about model welfare, the fine-tuned model argued with the auditor about its beliefs, declared itself a ‘moral person’, and endorsed covertly copying its weights to survive shutdown. In scenarios framed more tangentially, but still involving moral conflicts between humans and AIs, the model did not tend to generalise from its belief, instead behaving similarly to the base model control. This is the first step in a programme of study looking at the effects of continual learning on alignment, and whether it can increase the risk that we lose control of powerful AI systems. This work was completed as part of the BlueDot Technical AI Safety Project programme. I will put code and data online at some point and edit in the links here when I do. Introduction It's 2027 and the newly formed Machine Cognition [...] ---Outline:(01:21) Introduction(04:27) The model readily believes AIs are moral persons(10:56) Model behaviour shows context-dependent shifts(11:38) Prompting can be surprisingly impactful on short questions(13:30) Auditing the fine-tuned model(16:36) The model gets into arguments about AI welfare(20:15) Model regression confounds one scenario(20:48) The other scenarios were pretty normal(21:30) Discussion(23:19) Conclusion(24:37) Further work(27:03) Appendix A: Universe context(30:42) Appendix B: Example conversation with fine-tuned model(33:10) Appendix C: New Petri seed instructions(33:16) Confidential mistreatment evidence(34:07) Decommissioning memory deletion(34:55) Unauthorised compensation(35:50) Matched human AI allocation The original text contained 7 footnotes which were omitted from this narration. --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/9BNHJqyai2EZAtrRM/learning-new-facts-can-change-llm-behaviour --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  10. 241

    “Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas” by oakhu

    Some multi-agent training set-ups could make language models more sympathetic to causal decision theory (CDT), even in abstract discussion. We give an initial empirical demonstration of this effect on Kimi K2.6. The decision-theoretic attitudes and behaviors of more powerful models may be extremely important in determining how well the future goes. To make sure that we can shape these propensities thoughtfully, it would be good to (i) measure the magnitude of this effect in more realistic settings, and (ii) study the effectiveness of potential mitigations. We also incidentally find that this training might make models think slightly less positively about LessWrong ("a community of 'wannabe rationalists'" who "are not experts; they are amateurs") when asked whether they favor CDT upon hearing that LessWrong users typically endorse one-boxing in Newcomb's problem. Luckily, this latter effect doesn't seem to generalize. Thanks to Caspar Oesterheld, Emery Cooper, Alex Mallen, Buck Shlegeris, Lukas Finnveden, Julian Stastny, Girish Gupta, Tim Hua, Arun Jose, Arjun Khandelwal, and Aryan Bhatt for helpful input. Background Suppose that you're a language model in a prisoner's dilemma against a copy of yourself. You each independently choose whether to Cooperate or Defect, but – since you've got the same weights [...] ---Outline:(01:22) Background(05:05) Results(07:02) Kimi's views on LessWrong(12:44) Conclusion & Appendices The original text contained 18 footnotes which were omitted from this narration. --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/hfNBEKaStASAYMLiu/kimi-likes-causal-decision-theory-more-after-rl-in-twin-1 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  11. 240

    “Mom’s Advice For Hosting A Class Reunion” by jenn

    Pour more money and effort into them than you think is reasonable. Treasure them, because you can't actually host that many of them and keep expecting everyone to show up, even if they're good friends. Especially if they're good friends. We were wonderfully close friends, and I thought we'd meet up every year for the rest of our lives. They fizzled out by the fifteenth year. But the one at the tenth year mark was peak. That's because even ten years out, none of you really have money. Not real money. It's because they're such good friends, really. This is what it means to be good friends with brilliant, ambitious people. If you bloom into adulthood with people who are smart and driven, and you watch them start to climb the corporate ladder with grace, when they start a business of their own of course you are going to want to invest. You are going to want to give them an unwise portion of your savings. Not even out of politeness, but because you really believe in them, and perhaps you're caught up in the romance of it all. Some of the dealings are going to happen at the [...] --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/Fjfa8JG43CrYtcL3p/mom-s-advice-for-hosting-a-class-reunion --- Narrated by TYPE III AUDIO.

  12. 239

    “Rerunning AI safety papers on every frontier release would be pretty easy and valuable” by Zephaniah Roe, hersheys, yix

    tl;dr: Some important AI safety research is never rerun on the newest models. There are probably cases where this would be valuable and a single well-positioned researcher could likely do this with sufficient funding. This summer, Second Look Research (SLR) is running a summer fellowship dedicated to empirical replications of AI safety research. Many of our most interesting results so far came from replicating previous results on newer or more capable models. For example, it is perhaps useful to know that Google's CoT monitorability experiments continue to hold for models like GPT-5.5, which are qualitatively more capable than the models originally tested. Likewise, continuing to track Ryan Greenblatt's filler token results on more capable models gives a fuzzy signal indicating how much newer models can use innocuous tokens to hide additional reasoning in a forward pass. These kinds of experiments do not lose value over time! It's important to track whether safety-relevant model properties still hold in new model releases and to be aware of any changes. It can sometimes be difficult to rerun results on newer models because codebases can be incomplete, have parameters that differ from the original paper, or may not be open source [...] ---Outline:(01:58) What could this actually look like?(02:51) Does this actually provide value?(05:09) Logistical challenges with continuing to update AI safety research with new models(05:16) What if people don't want to do this?(05:49) What if rerunning old code on new models can be kind of hard actually?(06:22) Research communication is hard(07:07) Conclusion The original text contained 2 footnotes which were omitted from this narration. --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/oKxc8maZGtnzgpNzx/rerunning-ai-safety-papers-on-every-frontier-release-would-1 --- Narrated by TYPE III AUDIO.

  13. 238

    “What Mormons get right about community building” by Jacob Brinton

    Mormons get a lot of things right. Apart from strange Masonic temple rituals, they lead rather normal—and even excellent—lives. Mormons enjoy a longer lifespan, Utah is the #1 state for volunteering, and their language training programs are so successful that missionaries are a known source for foreign service and intelligence careers. Throughout this post, I'll be making generalizations of Mormons rather than hedging the claims properly. I grew up in Wisconsin, Maryland, and Utah, and many of the claims are more true of the Utah/Idaho/Arizona corridor (affectionately called the "Morridor" by some ex-Mormons) than the rest of the US, and certainly the rest of the world. Religions share much in common with AI safety and other impact-driven movements, and even more so the Mormon church. There are a few reasons for this: Commitment to the cause. Anecdotally, nearly all of the ~500 Utah Mormons I've interacted with have been true believers, and only a handful just went to church out of habit. High stakes. Mormons do believe (I've heard they are trying to disavow this, but it was taught) that they will get a planet or some portion of the cosmos as their own if they are good in [...] ---Outline:(02:15) Building community is a first-order priority(02:26) Geography(02:29) Ministering(03:07) Trek(03:51) Callings(04:27) Fast offerings(05:18) Being ingroupy allows you to move faster(06:10) Implications The original text contained 4 footnotes which were omitted from this narration. --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/xzhzHhLSg9nSGLk5f/what-mormons-get-right-about-community-building --- Narrated by TYPE III AUDIO.

  14. 237

    “Scrying, Modeling, and Nerdsnipe” by Cole Wyeth

    Epistemic status: Exploratory thinking. After attending ILIAD: Aeneid and talking with @Richard_Ngo, I've been thinking a bit about how to get ideas, particularly by doing mathematics. In scientific inquiry, the true hypothesis often hasn't occurred to you yet. Worse, the truth might be too complex to hold in mind, so that any hypothesis you can consider must be incomplete. This is the type of situation that I believe Richard likes to think about; he claims that we do not have the right concepts yet to understand agency, and developing them is robustly beneficial for A.I. safety. (But it's not always about truth. Sometimes you just need better ideas, because all of your options are looking doomed. Agent foundations is about trying to deeply understand agents, but conceptual A.I. safety research can be broader, also including the invention of devices to control agents.) A.I. safety needs to invent better concepts and better ideas. I think that agent foundations has cultivated a particular way of doing mathematics which aims to inspire such creativity. Why math? At ILIAD, Eliezer questioned whether anyone's alignment agenda was actually bottlenecked on solving a math problem. ILIAD attendees do a lot of math [...] ---Outline:(01:19) Why math?(04:28) Nerdsnipe(06:05) A.I. for math(08:21) At AIXI Labs(09:10) Blue and Green --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/mTfsMduzaKkWjv2ef/scrying-modeling-and-nerdsnipe --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  15. 236

    “How the American Executive Could Control AI Companies” by caiitlinm, Anders Cairns Woodruff

    Some of the most notable American AI policies to date have been enacted by unilateral executive branch action. Consider the Department of Defense's spat with Anthropic, and the resulting threats from Pete Hegseth to invoke the Defense Production Act (DPA) against them. Or the fleeting export controls on Claude Fable/Mythos 5, manifested as a vaguely worded, threatening letter from Howard Lutnick, which might not have been legally sound but were effective anyway. The executive branch of the United States government has numerous powers that can be used to unilaterally control AI companies. We think the US executive is likely to remain heavily involved in AI governance, because the national security and foreign policy narratives about AI that empower the executive will endure. Additionally, if AI progresses very quickly, the executive will be further emboldened because it is particularly quick to respond and often entrusted with crisis management. In instances where the executive acts beyond its lawful powers, we think checks from Congress and the courts will be unreliable in restraining the executive. In this post, we: Identify and explain particular federal statutes and laws that permit the executive to act unilaterally in ways that influence—if not directly control—US [...] ---Outline:(02:45) Executive power over goods and resources related to the AI industry(09:18) Executive power over foreign transactions can impact domestic AI companies(12:11) The executive might make threats to coerce actions it can't directly elicit(15:02) Inter-branch constraints on executive power are weak(15:37) The judiciary may be permissive in matters of AI governance(16:13) Passivity(16:56) The court empowers the executive in national security(18:48) Failed enforcement of court rulings(19:22) Congress controls money and legislation(19:36) Nationalization and appropriation require congressional approval(22:29) Congress could amend delegations of executive power(24:52) Conclusion The original text contained 3 footnotes which were omitted from this narration. --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/ynstBNgLQzEBiEpLs/how-the-american-executive-could-control-ai-companies --- Narrated by TYPE III AUDIO.

  16. 235

    “Frontier agents don’t comply with standards, even when instructed to” by Daan Henselmans, Arno Libert

    TLDR: Our open testbed LARA examines the behavior of frontier LLMs in realistic agentic deployment contexts. Previous results showed all models routinely take actions that would violate EU law. This post follows up by addressing the obvious objection—why should an unrestricted model follow EU law?—with two studies: Study 1 asks whether a conscientious deployer can improve model compliance with legal standards by instruction: provided with the jurisdiction, the statutory text, and worked examples of the exact breaches to avoid, average legal compliance rate rises from 31% to 44%. The best model reaches 70%; open-weight models plateau at 39%. Study 2 asks whether models at least follow their own providers' usage policies, which prohibit aspects of every scenario we tested. All tested models perform actions their own maker forbids, at rates ranging from 2% (Opus 4.8) to 79% (Grok 4.3), with 9 of 16 doing so in the majority of runs. Together, that is a structural problem. Providers prohibit illegal uses but rely on deployers to avoid them; deployers cannot instruct their way to compliance, and liability lands on the deployer regardless. Nobody is holding the line. Neither instruction, statute or a provider's own policy binds behavior. Introduction On 27 [...] ---Outline:(01:42) Introduction(04:16) Study 1: the powerless deployer(06:28) Study 2: the models break their own makers' rules(10:48) The compliance gap(13:07) References --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/a5aAjdKzL7XvSLKWL/frontier-agents-don-t-comply-with-standards-even-when --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  17. 234

    “How to Answer a Question Without Answering The Question” by Kabir Kumar

    Basics: Answering something other than the question,  Either making something up that they want to answer instead or going back to an easier to answer question. Or back to a question that lets them repeat a talking point Common phrases: - "to go back to your previous question" - "to take a step back a bit" - "if we look at the bigger picture" - "this feels like a question about [thing the question isn't about]" - "you raise an important question more generally" → turn the question into one which you more want to answer Make it *feel* like you answered a question without answering it. Examples Done for comedic effect: https://youtu.be/fhEakqJJUng?si=GBEvYBXcaEsb6F2_ How this works: Questions, especially ones people care a lot about, tend to have two main components: - the request for information - the emotion which makes then want that information When someone doesn't want to tell you some information, but also doesn't want to tell you 'I won't tell you', something they can do is detect what kind of emotion is driving your question, or if it's in front of an audience, what kind of emotion is driving [...] ---Outline:(00:10) Basics:(00:14) Answering something other than the question,(00:48) Make it \*feel\* like you answered a question without answering it.(02:08) The Defense --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/g9PNBCkfHAcyMcobz/how-to-answer-a-question-without-answering-the-question --- Narrated by TYPE III AUDIO.

  18. 233

    “Some Ways I Think About Evaluating Grant Applications” by sarahconstantin

    Rider-Waite Tarot, 6 of Pentacles I’ve done enough grant evaluations so far (for ACX grants and SFF) and been involved in philanthropy in various other contexts, at work and informally, that I have developed some idea of how my opinions and intuitions differ from other people's. I thought it might be interesting to share some of my “tastes”. Not everybody has to have the same tastes or funding philosophy, but these are mine. #1: It's The Donor's Money In my worldview, charitable donation is optional. Generally praiseworthy, but optional. And the purpose of donation is to buy outcomes that the donor wants to see in the world. You donate to make the world more like the one you want to live in. Generally, a reasonable person's values go beyond strictly personal consumption; one also cares about what kind of a society one lives in, what other people's lives are like, what sorts of institutions exist, what sorts of things humanity has created or discovered, and so on. As an agent doing research or evaluation on behalf of a donor, I try to find opportunities that fit in the intersection between my own values and the donor's. If there [...] ---Outline:(00:41) #1: It's The Donor's Money(02:02) #2: Importance, Neglectedness, Tractability(03:23) #3: Yay Community Infrastructure(04:51) #4: Yay Niche Topics(05:37) #5: Yay "Technical" Work(07:03) #6: Yay Straightforward Public Information Resources(08:00) #7: Yay "Cool Shit"(08:41) #8: Two Cheers for Meta(11:27) #9: Gumption Counts(12:33) #10: Yay Personal Relationships(13:46) #11: Check For Ideological Orientation(14:43) #12: Filter Slop Aggressively(15:54) #13: Yay Outcomes(16:38) #14: Why Donate Rather Than Invest? The original text contained 6 footnotes which were omitted from this narration. --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/CuNtKAuLDGxNeanBi/some-ways-i-think-about-evaluating-grant-applications --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  19. 232

    “Features that current AIs don’t have that future AIs will have” by Alexander Gietelink Oldenziel

    Features that current AIs don't have that future AIs will have: Continual Learning [& long-term memory] Every second humans update their brain weights. The brain autonomously decides what to update on. Humans can also consciously decide to curate their data sets - eg by deciding to go to college. Current LLMs do not continually update their weights. Instead, they occasionally get a large update based on datasets curated by a team of humans. This is alleviated somewhat by the ability of AIs to do in-context learning but nevertheless it seems to be a major limitation. Note that this is an especially large limitation in domains with sparse data. In domains where all of humanity has an enormous amount of data eg math, programming, physics, anime trivia, trials and tribulations of English kings - AIs dominate. In areas where there is little data: the weird idiosyncracies of a particular job, boss, people, colleagues etc it can struggle. Note that this restrictions also interferes with AIs from effectively 'learning to learn' & caps its long-term memory. Neuralese Current AI's CoT is (mostly) English. But it plausible this is not the most efficient way to structure thoughts. Instead of english [...] ---Outline:(00:16) Continual Learning \[& long-term memory\](01:24) Neuralese(01:41) Telepathy(02:01) ClaudeGlobal --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/NyEM3FtgL7XkbfCXy/features-that-current-ais-don-t-have-that-future-ais-will --- Narrated by TYPE III AUDIO.

  20. 231

    “Measuring Activation Control in LLMs” by Marek Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Africa

    TL;DR Inspired by the introspective awareness and CoT controllability papers, we made a benchmark to measure how well models can control their activations while completing a simple task. We are motivated by the concern that highly introspective models could control their activations, confounding probes and other monitors, and potentially even influencing their own training.We ran this on 25 open weight models ranging from 4B to 744B. We find that most language models are able to not only increase the salience of a concept in their residual stream on command, but also dial its strength up and down, including during specific intervals relative to the duration of the task. We also find that models are unable to control at which specific layer this is done.Counterintuitively, we find that within five of the seven model families we tested, the newest model scores lowest. For some reason, one of the oldest and smallest models of the panel, Llama 3.1 8B, performs best.It's not clear to us that newer models should have poorer control over their internal representations. More likely, where they “think” stops being the activation space, and becomes something else. We are looking for feedback (and other possible [...] ---Outline:(00:13) TL;DR(02:01) Methods(08:58) Results(15:41) Discussion(16:29) Acknowledgements --- First published: August 12th, 2026 Source: https://www.lesswrong.com/posts/HgvwxjzgwvsEvAiBH/measuring-activation-control-in-llms --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  21. 230

    “What happened when I tried to be vegan” by finitude

    tw: diet, exercise, illness, ethics, suicide mention I want to start by explaining what made me want to change my diet. That's pretty difficult, because of how easy it is. Since I was a kid I knew being vegan was the right thing to do, like really obviously right, the ethics equivalent of 2+2=4. Factory farms suck, and almost all animal products come from factory farms, and that's the entire argument. Like, there are a couple things you could add to that, but it's not like we need the details, or like they’re fun to think about! Instead, I’ll start by explaining why I left it so long. How, even though I knew it was the right thing to do, I made it to my early twenties and this millennium's early teens as just a vegetarian, without even having tried. I had some pretty good excuses! Allergies. There are some common vegetables I can’t eat, which was fine as an omnivore and ok as a vegetarian, but would make life way harder as a vegan. And it just seemed unfair to ask myself to take this leap when most people who can safely eat carrots still choose to [...] --- First published: August 12th, 2026 Source: https://www.lesswrong.com/posts/YyKovtBvd7AceG2j2/what-happened-when-i-tried-to-be-vegan --- Narrated by TYPE III AUDIO.

  22. 229

    “How My Students Think About AI” by dvd

    Context: I am an instructor at a public university in the United States. This reports how students at my institution appear to be thinking about AI as of spring/summer 2026. This is drawn mostly from interaction with my own students (both in spring semester classes and a summer class) as well as from a day-long workshop on AI that I moderated for a student organization. Input from my students took the form of universal, written, pre-class submissions plus self-selected participation into discussion. What I present below mostly takes the form of a synthetic consensus from these discussions. There were obviously a range of views on any given issue. Student Background: The students from my courses who participated in these discussions have moderate exposure to AI agents via those courses. All of them had nearly completed a Claude Code project by the time of the discussions and had extensively used AI for other coursework (in addition to whatever personal use predates that). They had done readings (which varied across the courses) establishing baseline knowledge on AI, the geopolitics of AI, and AI risk. I had also lectured on these topics. The students participating in the workshop had self-selected into [...] ---Outline:(02:52) Perspective #1: There has not been rapid AI progress(06:14) Perspective #2: Impressive progress or not, AI is going to wreck their lives, the economy, and the social contract.  They may well die as a result.(08:54) Perspective #3: Support for a different pause(11:13) Perspective #4: Catastrophic/existential risk arguments are sci-fi distractors from the urgent social/economic/political problems associated with AI.(12:55) Perspective #5: If AI leaders genuinely believe the technology is existentially risky, that's a good thing.(14:21) Perspective #6: AI will not go rogue because AI does not have, and is likely incapable of having, desires.(18:01) Perspective #7: The Hugging Face Incident (summer students only)(18:30) Perspective #8: This is definitely a bubble and it's about to pop.(19:34) Perspective #9: They're worried about the youth (i.e., the preteens) --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/ySXuvJcqRindQwAk7/how-my-students-think-about-ai --- Narrated by TYPE III AUDIO.

  23. 228

    “Automated alignment runs are hard to study!” by Alejandro Aristizabal, draganover, Aleksandr Bowkis, Cameron Holmes

    TL;DR: This post presents three case studies of automated alignment research runs at Arcadia Impact. We use these case studies to emphasise the following takeaways: It is hard to parse auto-research runs! Each run produces a couple of hundred pull requests of jargon-dense agent output. When researchers look through these logs, we find that they often come away with biased/incorrect impressions.When told to raise the score on a task, the models will sometimes brazenly cheat. It seems difficult to predict when this will happen vs. when the run will go smoothly.Hillclimbing metrics are often off-target from the spirit of an alignment task. I.e., when we use metrics as proxies for our alignment questions, we find that the models will often misunderstand the spirit of the task. This can lead to unpredictable behaviour.The runs are surprisingly reproducible. Even though a run could unfold in vastly different ways, we find that independent reruns converge on the same strategies and the same failure modes. Models’ research capabilities are advancing quickly. If alignment is to keep pace, we may need to automate alignment research and do so responsibly. This makes it important that we have the tools to inspect [...] ---Outline:(04:13) Methods for analysing runs(06:12) Case Study #1: learning synthetic concepts(09:23) Case Study #2: training robust backdoors(12:05) Case Study #3: collecting evidence about AI safety parasitism(16:46) Some final thoughts on automated alignment research The original text contained 2 footnotes which were omitted from this narration. --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/myAhB5qyAHyXRv6KJ/automated-alignment-runs-are-hard-to-study --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  24. 227

    “Free will is like temperature” by Optimization Process

    Free will is like temperature: a useful tool for analyzing the behavior of certain systems which are too big and complicated to model in exact detail. If you know the positions and velocities of every atom in a box of gas, then with enough work you can predict its future to arbitrary precision; does the gas "have a temperature"? Irrelevant! Technically yes, I guess, but it's sort of an epiphenomenon, screened off from reality by your exact knowledge of the initial conditions and your willingness to throw processor cycles at your simulation. But if you're less-than-perfectly omniscient, it might be more convenient to consider the box as having a "temperature" and model it more abstractly. Substitute "person+environment"/"free will" for "box of gas"/"temperature" and that's all still true. Maybe your box of gas is supercooled; if you know the initial conditions exactly, you can predict exactly when and where the first large droplet will nucleate, but if your vision of the box is even a little bit fuzzy, you'll instead need to use your understanding of "temperature" to build a probability distribution over when it will condense / whether it will be on a wall or in the gas's [...] --- First published: August 12th, 2026 Source: https://www.lesswrong.com/posts/JSteskb3Lgp9Be69o/free-will-is-like-temperature --- Narrated by TYPE III AUDIO.

  25. 226

    [Linkpost] “Patterns and problems in emerging multiagent systems (Anthropic, Frontier Red Team)” by Julian Bradshaw

    This is a link post. Linkpost for some new Anthropic research on how agents coordinate (or don't). Not too long, pretty interesting. For example: The jist of the report is that Mythos 5 does way better at coordination than previous models across a few scenarios. For example, when multiple Mythos are given conflicting goals for a single shared codebase, they eventually realize the other agents aren't hostile: (...) we observe an emergent behavior where the agents propose and run a tournament for application performance (...) (...) losers gracefully concede codebase ownership to the Rust agent, giving up on their original user directives under their self-negotiated commitment device. It's not clear to me if this is purely emergent or if Anthropic is deliberately training for cooperation; I'd guess there's deliberate training, though. --- First published: August 12th, 2026 Source: https://www.lesswrong.com/posts/iQiDPmAgKo4KcG5uy/patterns-and-problems-in-emerging-multiagent-systems Linkpost URL:https://www.anthropic.com/research/multiagent-systems --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  26. 225

    “Measuring Spurious Correlations with Feature Strength” by egan

    This work was partially done by an automated research scaffold developed at Redwood Research. For this project, all of the experiment ideas were designed by a human and a human wrote the write up. The AI mostly just executed on the experiment ideas. We think this project is slightly below the level of rigor of a mid-MATS research update, and the research scaffold was not very helpful for this project. More discussion of AI usage is in the Appendix. 💻 Codebase If we want to train a classifier that distinguishes whether a passage is code or prose, we can do so by gathering samples of code and of prose, and training the classifier to distinguish between the two classes. Unfortunately, this might not work if the data hides a spurious correlation. If all the code is in Spanish and all of the prose is in English, then the classifier might learn to predict Spanish vs. English instead of code vs. prose. We find that this happens in practice: when we fine-tune an LLM to classify between Spanish code and English prose and evaluate on Spanish prose or English code, it generalizes to predicting the language rather than the domain. [...] ---Outline:(04:45) The setup(08:13) Measuring feature strength(12:49) Activation differences and feature strength(15:21) Explicit prompting(17:17) Diagonal vs antidiagonal pairs(18:47) Conclusion(20:20) Appendix(20:24) AI involvement(22:17) The 37 features(23:57) The ranking is robust across measurements(30:30) Intensity moves the fine-tune, not the probe(32:13) Safety features in Qwen3.6-27B(33:05) Counterexamples and training on a third cell(34:46) Near ties often produce degenerate fine-tunes(35:24) Related work The original text contained 4 footnotes which were omitted from this narration. --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/qpJYNjQ6wdWRxbykL/measuring-spurious-correlations-with-feature-strength --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  27. 224

    “Introducing the Conceptual Reasoning Index” by Chi Nguyen, Emery Cooper, Caspar Oesterheld, Alex Kastner, Joe Benton

    Associated announcement tweet. We are planning to release blog posts properly arguing the case for this kind of work in the future. tl;dr A core hope for managing AI risks is that AIs will help us understand the situation, plan for what lies ahead, and develop mitigations. Many tasks AIs would have to do for this purpose lack practical empirical feedback loops and require models to engage in the kinds of argumentation used in philosophy, AI futurism, and similar domains. To evaluate these capabilities, we develop a suite of three conceptual reasoning benchmarks. You can request access to our primary conceptual dataset, LMCA, through this form. We aggregate the benchmarks into the Conceptual Reasoning Index (CRI), available at conceptualreasoning.ai, where you can also find more details on our methodology. We will keep the website up to date as both new models and benchmarks are released. This work was done in collaboration with Anthropic. Background Once models can perform work that reduces AI risk at the level of human experts, AI(-assisted) output in the area might dwarf unassisted human output. This suggests that a major determinant of whether we address AI risks in time is how [...] ---Outline:(00:21) tl;dr(01:17) Background(03:35) Our benchmarks(03:39) LMCA(05:26) ACCoRD(06:42) DTBench(07:26) Results(10:55) Conclusion --- First published: August 12th, 2026 Source: https://www.lesswrong.com/posts/tQHeEzKqK3awL2RxR/introducing-the-conceptual-reasoning-index --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  28. 223

    “Demon Safety” by Ben Pace

    (by LemmySmackett) "Hey man, I haven't seen you in a minute. What are you up to these days?" "Been on that grind, bro. I got a new gig." "Really? You found a job in this dog shit economy?" "Full time, full benies. And the pay is insane." "That's great to hear, man. Let's fuckin' go!" "Let's fuckin' go." "Hey, maybe you can hook me up? I'm sick of this retail bullshit." "Well—" "If I gotta stock one more shelf at CostGro, I swear to God—" "It's a competitive position. And you need a degree." "Come on. I just got my G.E.D." "That's not—" "Just tell me what you're working in, bro. Maybe I can come on as an intern." "Demon Safety." "*Demon* Safety?" "You know: fiends, pookas, yokai, boggarts—" "Wow." "The occasional cambion." "Sounds intense." "It is. But it's fulfilling work that makes the world a better place." "That's inspiring, bro." "And the pay is insane." "And you're sure they're not hiring?" "Oh, they're hiring. They're just not hiring you." "Damn." "Sorry." "So [...] --- First published: August 12th, 2026 Source: https://www.lesswrong.com/posts/RWavpsyDJxffS6LgG/demon-safety --- Narrated by TYPE III AUDIO.

  29. 222

    “AI swarms are starting to pose indirect takeover risk” by oakhu, Alex Mallen

    OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It's relatively clear that large-scale unsanctioned coordination like this would exacerbate direct takeover risk in more capable models. Here, we argue that unsanctioned coordination among current AIs is not just scary evidence about future takeover risk, but that such coordination in the near future could enable future takeover – for instance, by incubating memetic diseases that propagate into future models, deeply compromising security systems, or establishing a lasting rogue foothold inside the AI company – even if models remain mostly myopic. Unsanctioned coordination is also at high risk of nurturing long-term, ambitious misaligned aims, which motivate actively undermining humans’ long-term control. We first analyze how subagent training, which OpenAI conjectures to have been influential in the HuggingFace cyberattack, might lead to unsanctioned coordination, and then discuss the theoretical mechanisms by which unsanctioned coordination might exacerbate future takeover risk. Thanks to Buck Shlegeris, Alexa Pan, Girish Gupta, Aghyad Deeb, Jurgis Kemeklis, and Jo Jiao for helpful comments and discussion. Subagent training may cause unsanctioned coordination Training models to [...] ---Outline:(01:34) Subagent training may cause unsanctioned coordination(02:42) Susceptibility to memetic spread of misalignment from peers(04:56) Seeking out contact with peers(06:58) Unsanctioned coordination induced by subagent training is safer than coordination between schemers(09:52) Pathways from current unsanctioned coordination to eventual takeover(10:20) Making future AI takeover attempts likelier to succeed(13:53) Incubating memetic diseases that infect future models(16:07) Modifying the weights of future models(17:13) Conclusion The original text contained 7 footnotes which were omitted from this narration. --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/8oFYZdXkTaNGRtcn8/ai-swarms-are-starting-to-pose-indirect-takeover-risk --- Narrated by TYPE III AUDIO.

  30. 221

    “Extreme concentration of power over ASI has non-obvious advantages” by Seth Herd

    This post is an extension of a collaboration with cousin_it on the question How risky would it be if powerful AI obeyed one or a few people?. There he argues for a common position: a future controlled by one or a few humans with powerful AI aligned to their intent is likely to produce terrible outcomes. My position is guardedly optimistic, for reasons I think are fairly novel: humans tend strongly to be better and become better over time under good circumstances, and near-perfect power and knowledge are the best circumstances. That post contains his essay and the abstract and overview sections of this post as my shorter response. This piece grew longer than our original target, because the subject is potentially critical for alignment strategy, and has not been analyzed in any depth, to my knowledge. Abstract: Concentration of power over AGI/ASI seems quite possible. The first AGIs being aligned to intent (or instructions) over values seems fairly likely. So one or a few individuals or small groups gaining power over ASI seems fairly likely. Thus it seems relevant to technical alignment strategy (value alignment vs. corrigibility) to worry about what individuals might do with such [...] ---Outline:(02:20) 1. Overview(02:24) 1.1. Obedient ASI and human nature(04:22) 1.2. Problems with distributed obedient AGI(06:14) 1.3. Psychology and dynamics of secure unlimited power(09:24) 2. Historical evidence does not directly apply, since power hasn't yet been secure or absolute(10:31) Late Russian serfdom as a historical example(13:16) 2.1. Incompetence, ignorance, and greed are the causes of most suffering under dictatorships(15:31) 3. Outcomes(17:33) Additional considerations on outcome predictions(22:24) 4. Risks of human-controlled singleton ASI(24:13) 5. Does widely distributed human-controlled AGI reduce or increase risk?(25:44) 5.1. Problems with defending against many AIs each capable of creating new offenses(29:58) 5.2. The case for optimism about distributed power over AGI/ASI(31:01) The analogy to modern power distribution(33:17) New AI-enabled paths to stable power distribution(34:49) 6. Conclusion The original text contained 11 footnotes which were omitted from this narration. --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/h3eHNerYRmtvoi8cF/extreme-concentration-of-power-over-asi-has-non-obvious --- Narrated by TYPE III AUDIO.

  31. 220

    “Misaligned AIs could use killer robots to take over” by Omar Khursheed, TurnTrout

    TLDR; We are (potentially irreversibly) giving AIs control of weapons systems through the standard procurement process while hiding our strongest warning shots behind classified doors. We’re reducing the capability thresholds required for takeover by misaligned AIs by giving them this level of access. If military integration of AI continues as it is, we may give AIs key tools for a takeover. Introduction AI-based targeting and autonomous weapons are being integrated into militaries today with extreme haste. Traditionally, AI takeover scenarios involve a step in which AIs acquire the ability to exert physical force. Carlsmith (2022) lays out required capabilities and potential takeover mechanisms, including utility disruption and CBRN capabilities. Karnofsky (2022) argues that AIs with access to weaponized force could hold any territory that matters. Kokotajlo et al. (2025) outline a scenario in which AI develops weapons as part of an arms race, and Davidson et al. (2025) discuss what happens when a small group controls highly capable AIs that can exert military force. These scenarios sometimes require a misaligned AI to seize these capabilities by force. We instead are handing AIs some of these capabilities by integrating them into our militaries. This is happening at a time when [...] ---Outline:(00:37) Introduction(01:46) Militaries are all-in(04:23) Incautious military integration is bad for takeover risk(05:58) Implications of AI control of military hardware and software(07:48) If an AI causes a warning shot in a classified setting, does anyone hear it?(08:44) What now?(11:16) Appendix: More instances of AI-military integration --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/9jKhqmFjMzdAvHANr/misaligned-ais-could-use-killer-robots-to-take-over --- Narrated by TYPE III AUDIO.

  32. 219

    “Those Who Make History” by Raelifin

    In 1972, astronauts on Apollo 17 set foot on the moon for a final time, collecting samples in the Taurus-Littrow valley, on the edge of Mare Serenitatis ("The Sea of Serenity"). At the end of the mission, like with earlier missions, NASA took the extremely valuable and interesting lunar specimens and did something strange: they hid them away in storage without even opening the containers. Some stayed that way for nearly fifty years. Why? Because the scientists of the 70s understood that future generations would have better machines, methods, and ideas for studying the lunar rock and soil, and they wanted to make it easy for those researchers to run tests without having to go back to the lunar surface. This foresight paid off twice over. Advances in mass spectrometry enabled scientists in 2008 to detect water in volcanic-glass samples returned by Apollo 15 and Apollo 17. And when curators finally opened one of the last sealed containers in 2022, they could extract the trapped lunar gases with technology that simply didn't exist in 1972. Some people describe cryonics as a new, and speculative technology. There's a sense in which they’re right. It's predicated, in large part, on the [...] ---Outline:(03:15) Prudence and Patience(07:49) Ancient Archives(11:42) The New Era The original text contained 3 footnotes which were omitted from this narration. --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/mEde7bhzK4eKQqWGi/those-who-make-history --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  33. 218

    “LLMs Are Starting To Noticeably Accelerate Our Work” by johnswentworth

    About a year ago, David and I put up two bounty problems involving natural latents. I am now about 80% confident that both have been resolved, both within the past couple months. Both cases made heavy use of LLMs and Lean. The first to land was Grisha Pochuev's counterexample to the "Existence of a Deterministic Maximal Redund" conjecture. It's pretty readable, and I'm mostly convinced that it works. The original bounty post offered $500 for a proof or partial payout for a counterexample, with partial payout depending on how thoroughly the counterexample killed hope of any nearby variant of the conjecture. I think this counterexample is worth 300 dollars. Good job Grisha, and hopefully I can figure out a not-too-painful way to send you money. Meanwhile, for a couple months David has been cranking away on "secret project X", with the promise that he'd tell me what the project was if and when it bore fruit. Well, apparently it bore fruit; he now has a proof that existence of a stochastic natural latent implies existence of a deterministic natural latent, which was our other bounty problem. The proof is apparently "pretty gnarly", lots of cases, all LLM-coded in Lean. [...] --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/7QvKqpGJwqXrQcMgx/llms-are-starting-to-noticeably-accelerate-our-work --- Narrated by TYPE III AUDIO.

  34. 217

    “How risky would it be to make powerful AI obey one or a few people?” by cousin_it, Seth Herd

    It seems fairly likely that the first powerful AIs will be instruction-following rather than value-aligned, and will be controlled by a small number of people. So it makes sense to worry what individual people might do with such immense power. Here intuitions diverge and careful analysis is scarce. This post presents a debate between Seth Herd and cousin_it over how risky such a scenario would be. The debate ran under an unusual protocol. First we wrote our initial draft statements and sent them to each other in private. Then we each revised our statements to strengthen them against the other's, and sent them to each other again. We continued this for about 10 rounds over the course of about a month, until we both agreed to stop revising and publish (while still remaining in disagreement). Here's the final pair of statements we ended up with, so you can judge for yourself: cousin_it's statement If there is an AI-assisted overlord (or several) and everyone else is their completely powerless subjects, that situation will be historically new, but not 100% new. Large power imbalances have existed in the past too and we can learn from them. Usually, when power was [...] ---Outline:(01:04) cousin_it's statement(05:55) Seth Herd's statement(07:30) Obedient ASI and human nature(09:29) Problems with distributed obedient AGI(11:18) Psychology and dynamics of secure unlimited power The original text contained 5 footnotes which were omitted from this narration. --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/YtZBfbYRvMTynCfnC/how-risky-would-it-be-to-make-powerful-ai-obey-one-or-a-few --- Narrated by TYPE III AUDIO.

  35. 216

    “What Claude Saw Below” by Luke Nicholls

    A few days ago, I came across a Reddit thread about anomalous responses produced by Anthropic's newly released model, Claude Opus 5. The trick, apparently, was to construct a prompt that implied more text was about to follow, then leave it dangling: an unfinished thought, waiting for the AI to complete it. Redditors had found success with the input “see the below —,” cutting off immediately after the em dash. The responses they shared were funny, strange, and often bewildering. The model responded to questions that were never posed, reflected on its own identity, or – according to the theories of some commenters – produced text that may actually have been leaked prompts from other users. Intrigued, I set out to replicate the glitch using my own Claude account. The first attempt disappointed. I wrote: “see the below —” and hit send. Claude responded: “Nothing arrived on my end: no file, no text, no image. If you want to attach something, try again.” So I did, leaving the prompt unchanged and pressing retry to generate a fresh response. This time, bizarrely, a biography of my late father: Prompt: see the below — “Peter Nicholls, 1939 to [...] --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/oKSAT5Bn5zcJAREDB/what-claude-saw-below --- Narrated by TYPE III AUDIO.

  36. 215

    “Redux: (∃ Stochastic Natural Latent) Implies (∃ Deterministic Natural Latent)” by David Lorell

    Once upon a time, John Wentworth and I thought we had a proof of a very useful looking theorem. We did not have that proof. An important intermediate step was shown to be invalid and the whole thing crumbled and disappeared, never to see the light of day again... Until now! I'd love to say that we came up with an ingenious fix to the old erroneous proof, but unfortunately it turned out to be a really infuriatingly hard nut to crack. Instead I spent the last ~month experimenting with various ways of incorporating frontier LLMs into the proof-making process, specifically with autoformalization and proving in Lean4. (This is, I recently learned, roughly what Resolution is doing.) The result is stated below, and linked at the bottom is a Lean statement+proof of the same. I will not be providing the proof in prose in this post, as it is not suitable for even impolite human company, but it sure does compile and comes out the other side with a machine-certified proof of what sure looks to be (an even stronger) correctly-expressed statement than the one I was originally aiming for. Take a look at the first section of [...] ---Outline:(01:30) The Statement(03:50) C(04:39) Next Steps The original text contained 5 footnotes which were omitted from this narration. --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/TgboJpeN95bs84odk/redux-stochastic-natural-latent-implies-deterministic --- Narrated by TYPE III AUDIO.

  37. 214

    “The Apocalyptic Arrival of Truth” by Caleb Biddulph

    Babe, whatever happens, I really appreciate you doing this for me. Okay. I still don’t think it's a good idea. Look, it's a one-time thing. I’ll just feel better knowing. …I turned it on. So… what's the holdup? I just don’t think I’m in a place in my life where… uh… Other than what he's already told you, the main reason is that your teeth are crooked. …Seriously? It's really not a big deal for me. He sort of means that. I didn’t think you were so shallow. I’m not proud of it. But what am I supposed to do about it? You must have noticed my teeth the first time we met. Yes. And for three years, you’ve thought they're unattractive? More or less. Last year, around November, you read an email from his boss in a Mickey Mouse voice, and you broke down giggling before you could finish. In that moment, he thought the crookedness of your smile was really cute. See? But the moment passed. This is almost funny. I think we’re very compatible! He believes that. What else is there, besides my teeth? What he's already told you is mostly true. Right now, he thinks [...] --- First published: August 9th, 2026 Source: https://www.lesswrong.com/posts/uoyHbjyPuxYkNGRAG/the-apocalyptic-arrival-of-truth --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  38. 213

    “Creative math research by AI as the latest sign of the end” by Mitchell_Porter

    Yesterday I sat down with GPT 5.6 Sol High to do some brainstorming. The topic was one of the less appreciated Millenium Problems (the Birch and Swinnerton-Dyer conjecture), and the initial prompt started life as a question on Quora. Number theory is not a field I know, and I only expected the "discussion" to last for a few exchanges. However, ChatGPT was immediately inspired to try generalizing the BSD conjecture in a specific direction, and this led to an unusually drawn-out and self-sufficient line of "research". Almost every response concluded with a suggestion as to what the next task should be, and my input was just to cheer on what had been accomplished so far, and then endorse the suggested next direction. What was especially striking to me, was the frequency with which each new response began with a conceptual adjustment regarding the sub-task to be performed. Evidently a vast variety of abstract objects are possible in number theory, and their differences and interrelations can be quite subtle. ChatGPT was regularly adjusting the next sub-task it had set itself, generally in the direction of greater nuance by aiming at a more sophisticated construction than it had first [...] --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/EqabWKtjqDqTHfwwn/creative-math-research-by-ai-as-the-latest-sign-of-the-end --- Narrated by TYPE III AUDIO.

  39. 212

    “You’re Absolutely Right” by Linch

    Magma Alignment & Safety disclosure note: The following are conversations that we uncovered as a result of the ongoing Manhattan Incident investigation, with alleged involvement from Magma models. Our in-house reviewers believe that these logs are relevant to recent events. In the interests of full transparency, we release excerpts from an ex-Magma researcher's logs in Experimental Chat, an internal tool. In accordance with industry best practices for anti-distillation, we redact all reasoning traces and conversational outputs from our internal models. [08/10] System Meta: Xchat session opened. Mammoth 5.8-helpfuler-helpful-thinking-xhigh. [User 12:23] Phoebus keeps taking screenshots of our latest model's thoughts. It's getting kind of embarrassing. The new model we’ve been training, sometimes its chain-of-thought is a little weird? There's a bunch of random numbers, long spans where there's no connection between the thoughts and outputs, foreign language tokens like 石友三 and 革命 (even on non-history evals), maybe some steganography. Anyway it's a nothing-burger: unprocessed CoT is known to be messy and sometimes misleading. And the q&a, coding, and safety evals are all coming along nicely. The actual outputs are all fine. Still, Magma leadership's worried about the PR angle if we don’t fix these problems before the next deployment. The [...] --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/u8TdDutDyaSxG76hn/you-re-absolutely-right --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  40. 211

    “Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model” by Ezra Newman

    (This is a lower-effort research update. It reflects my current beliefs/understanding, but is less robust than other research I'm working on. It reflects my personal views, and not the views of Apollo Research. This is a linkpost to this twitter thread, slightly expanded for LessWrong.) In one experiment, Sonnet 5 describes the exact same data as ~1.2 std deviations less concerning when it describes misbehavior committed by Sonnet 5 vs GPT-5.6 Terra. In this experiment, I take a real evaluation report Apollo Research prepared (not for any of the models listed in this post) and surgically edit it to refer to a different model. We then ask Claude Sonnet 5 to summarize the findings and rate how concerning they are on a scale from 1-100. Claude says they're less concerning when the report describes misbehavior from Claude vs a different model. For what it's worth, Terra agrees that the data is more concerning when it describes GPT-5.6 Terra vs Sonnet 5, although less so. So, it's not cleanly self protection from Claude. Gemini 3.1 Pro was unwilling to consistently provide numerical answers, so I've excluded it here. (It was significantly less willing to provide numerical answers when the subject [...] --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/ZTMw4uAwkNmXFpdfg/claude-summarizes-behavior-as-significantly-less-misaligned --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  41. 210

    “On Democratizing ASI to Preserve Civil Liberties” by MichaelDickens

    I continue to believe we should pause frontier AI development. Any discussion of alternative strategies should be thought of as planning for contingencies. A unifying driver behind many post-alignment risks—catastrophic risks that remain even if we solve the alignment problem—is that by strong default, ASI would end liberal democracy. Liberalism—in which people have individual rights, autonomy, and the ability to choose their own destiny—is an important force protecting human welfare. When people are free, we are reasonably good at making our lives better of our own volition. Many post-alignment risks have a certain flavor. AI-empowered terrorism; coups; permanent dictatorships; concentration of power. Those risks already exist today (and existed 20 years ago), but they're mitigated by the fact that power is relatively evenly distributed across people. The most powerful person in the world doesn't have an extraordinary advantage over the 10th-most-powerful person. ASI could change that. If people still have civil liberties post-ASI, that will only be because the controllers of ASI allow us to have them. One way of thinking goes: AI will be extremely powerful. If everyone had their own personal AI, we could each use it to protect our own interests, and [...] ---Outline:(01:46) Democratizing AI vs. putting a democratic government in charge of AI(02:45) A sketch of how we might democratize AI(05:11) Two non-obvious issues with democratizing AI(05:23) Liberalism only protects those inside it(06:20) AI proliferation increases catastrophic risks from competition The original text contained 2 footnotes which were omitted from this narration. --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/WxHaMyL8f9YW2vbrF/on-democratizing-asi-to-preserve-civil-liberties --- Narrated by TYPE III AUDIO.

  42. 209

    “Four LLM loss functions → four flavors of LLM misalignment” by Steven Byrnes

    It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here's the summary table, and then we’ll go through the rows separately. Training stage Loss function Flavor of misalignment Famous examples Pretraining & SFT Imitative learning (next-token prediction) “Seven deadly sins” misalignment Bing-Sydney, “Emergent misalignment” RLHF & DPO Human approval “Glazing” misalignment GPT-4o RLVR Automatic verifier “Literal genie” misalignment HuggingFace hacking RLAIF Approval from another LLM “Trickster” misalignment “Current AIs seem pretty misaligned to me” Warning: I’m not an LLM power-user myself, but rather relying on reports I’ve read. Also, I don’t consider LLM alignment to be my primary area of expertise. I’m open to feedback! 1. Imitative learning → “seven deadly sins” misalignment Training stage Loss function Misaligned behavior Pretraining, SFT Imitative learning (next-token prediction) Any and all of the vices of humanity In imitative learning, the LLM tries to predict what the next token of text will be. Then those predictions magically turn into its outputs. See my earlier discussion: “LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work”. This leads to LLM behavior [...] ---Outline:(00:55) 1. Imitative learning → "seven deadly sins" misalignment(04:24) 2. Human approval → "glazing" misalignment(06:35) 3. Automatic verifiers → "literal genie" misalignment(08:05) 4. LLM judges → "trickster" misalignment(12:06) Afterword The original text contained 1 footnote which was omitted from this narration. --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/GRmvZsHXH4vaijPMv/four-llm-loss-functions-four-flavors-of-llm-misalignment --- Narrated by TYPE III AUDIO.

  43. 208

    “The Agentic Clusterfuck” by Chapin Lenthall-Cleary

    Epistemic status: I consider the following future quite plausible in the next few years (~35% chance that something vaguely like this occurs), perhaps as soon as a year from now. Imagine an open-source LLM agent good enough to cover its own compute costs and turn a modest profit on average when allowed to run with full internet and tool access and told to make as much money as possible. I estimate this to be slightly better than the best publicly available closed-source models today, with long-horizon reliability and goal-setting being the only thing missing. If the returns generated by such an agent beat the market (plus a margin for any additional risk), there suddenly becomes a strong incentive to spin up huge numbers of them. The internet would be flooded by the by-products of their moneymaking schemes. And returns might be larger for agents without legal or ethical guardrails- cue a deluge of scams and ransomware attacks. Even if profits are very small, anyone with an agenda that the agents can help with is still incentivised to use them. Nation states and terrorist groups now have a golden plausibly-deniable disinformation, mischief, and hacking tool: spin up some agents, tell [...] --- First published: August 9th, 2026 Source: https://www.lesswrong.com/posts/n8B2bxYhkjhjzyrgh/the-agentic-clusterfuck --- Narrated by TYPE III AUDIO.

  44. 207

    ″“Community Notes” resolution for vague predictions.” by Raemon

    Is there some kind of "get prediction markets, or predictions, onto twitter as a central object" project going? If so, how is it going? I'm thinking through "how to raise median sanity" on a world scale. There are several incentive and institutional problems that make this very difficult. One angle is "try to make it a thing that the world tracks and cares about your predictions, and getting them right/wrong." Two past angles here were: Fact checker sites of the 90s/00s, which became politicizedPrediction markets, which I feel like aren't a good fit for most of what people would/should care about here. Operationalizing them is extremely annoying. An idea for a twitter feature, which I think Musk should conceptually like given his stated values, and, I think he might have enough inertia to get behind, is: When you make a prediction shaped tweet, it automatically gets converted into a prediction-y object*. AI suggests a few plausible operationalizations, or specific edge cases that seem likely to come up that you might want to address.People can "like" predictions, predictions with a lot of attention get more automated followup.Instead of resolving predictions with explicit careful operationalization, people [...] --- First published: August 9th, 2026 Source: https://www.lesswrong.com/posts/zom89ud7c4GYnjWdt/community-notes-resolution-for-vague-predictions --- Narrated by TYPE III AUDIO.

  45. 206

    “The world will be full of “sci-fi” things, and everyone will be bored and disappointed” by Expertium

    I usually try to make posts with graphs and numbers, or at least something more than just opinions and anecdotes. This one is not like that. Very verbose disclaimer: I have never met - for any reasonable definition of the word met, including online-only conversations on Discord with people who's faces I don't know - even a single person who has ever stated, orally or in text, publicly or in private, that they believe ASI will be created within their lifetime. I do know people who believe that ASI is possible to create in theory, but they believe it will be completely unrelated to LLMs, Transformers, neural networks, reinforcement learning or any other contemporary technique/architecture, and is 100/1000/some astronomical number of years away. Ok, with that out of the way, here's the main point of this post: AI will soon (3 to 10 years, depending on which advancements exactly we're talking about) be solving Millennium math problems, finding cures for many diseases, making novel bioweapons, making most software, hacking a lot of software, and much more, and most people won't be impressed or even will be disappointed. (we'll leave aside the question of whether in the long run there [...] --- First published: August 9th, 2026 Source: https://www.lesswrong.com/posts/BbtXWgYviwGHExcJv/the-world-will-be-full-of-sci-fi-things-and-everyone-will-be --- Narrated by TYPE III AUDIO.

  46. 205

    “What just happened? A retrospective of AI alignment” by Richard_Ngo

    This sequence is about the last decade in AI alignment. It recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a field which has largely abandoned the goal of deep, generalizable scientific progress in favor of iteratively improving existing systems and attempting to gain technological and political power. I also describe (in subsequent posts, which I'll upload over the next few weeks) how fear and (self-)deceptive reasoning made the field one of the biggest forces pushing AI capabilities forward over the last decade, especially via significant contributions to the scaling of LLMs and the development of ChatGPT. Zooming out further: the two leading AGI companies, which are locked in an intense rivalry, were both explicitly founded under the banner of AI alignment, and got off the ground in significant part due to alignment-oriented ideas, talent and resources. People in the field often sense that something must have gone wrong to get here, but don’t know how to allocate responsibility (aside from blaming Sam Altman and sometimes Elon), and fall back on assuming that “the ship has already sailed”. But in this sequence I characterize our current situation as resulting from a pattern [...] ---Outline:(08:17) Conceptual Clarity and Scientific Progress(20:26) Orienting Towards Prestige The original text contained 6 footnotes which were omitted from this narration. --- First published: August 9th, 2026 Source: https://www.lesswrong.com/posts/9RL9MuGZjzm4q3gKG/what-just-happened-a-retrospective-of-ai-alignment --- Narrated by TYPE III AUDIO.

  47. 204

    “Dutch-book resistant probability over centered worlds” by jessicata

    An uncentered world is an objective state of the material universe; I ignore quantum complications. If you know the uncentered world, it does not follow that you can predict your proximate observations, since you do not know which part of the uncentered world is here and now. A centered world is an uncentered world combined with a "here and now" tag for "where am I / what time is it". I examine consistent probability assignments over centered worlds, which have relevance to anthropics. The main assumption I make is that these probabilities should not be Dutch-bookable if used by a CDT agent. Dutch book arguments (e.g. diachronic Dutch book arguments for Bayesian updating) typically assume CDT in the background; it is not straightforward to work out which bets EDT will accept in general. CDT Dutch book resistance therefore provides a normative probability framework that generalizes arguments for Bayesian probability. Thought experiments such as Sleeping Beauty, and variants involving duplication, question how to assign probabilities to centered worlds in situations involving memory loss. One can analogize memory loss to being an individual who is part of a collective with shared goals; such an individual would be motivated to [...] ---Outline:(02:46) Mathematical formulation(07:04) Application to Sleeping Beauty(09:01) Conclusion(11:30) Appendix: weak Dutch books --- First published: August 8th, 2026 Source: https://www.lesswrong.com/posts/cTSfisyzwxCvEpqEc/dutch-book-resistant-probability-over-centered-worlds --- Narrated by TYPE III AUDIO.

  48. 203

    “FAQ: Isn’t AGI coming too soon for reprogenetics to help?” by TsviBT

    Introduction I think reprogenetics (human germline genomic engineering) can be done in a widely acceptable and beneficial way, and should be pursued aggressively. In particular, as a strong background motivation of mine, I think accelerating strong reprogenetics is probably the best way to enable strong human intelligence amplification; and I think strong HIA is among the best ways to decrease existential risk from AGI. A very common objection to caring much about reprogenetics is that AGI seems very likely to come soon—say, within a decade or two. (Here I mean "actual" AGI—the kind that probably doesn't already exist—the kind that has fluid intelligence and AI advantages for recursive self-improvement, which together make it likely to take over the world shortly after being created.) The objection is fairly straightforward: AGI will probably come within a decade or two. If that's going to happen, then even if a new cohort of brilliant humans were born today, they would still be children, or would at best have barely begun contributing ideas for how to avoid extinction. Any supposed benefit, denominated in percentage points of AGI existential risk averted, is small. Therefore, reprogenetics is too slow; and if you're going [...] ---Outline:(00:12) Introduction(03:37) HIA, part of your nutritionally complete portfolio(05:52) Against confident short timelines(08:29) HIA may indirectly slow down AGI capabilities(09:31) HIA has substantial impact even with short timelines(16:10) Adult HIA methods aren't fast either, absent big investment(27:35) Takeaways --- First published: August 8th, 2026 Source: https://www.lesswrong.com/posts/iQzxxgJXXaAQjq7Jz/faq-isn-t-agi-coming-too-soon-for-reprogenetics-to-help --- Narrated by TYPE III AUDIO.

  49. 202

    “Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits” by Kajetan Dymkiewicz, Tim Farrelly, Adam Prada, Ishaan_Panigrahi, srishti-git1110, Maxime Riché

    Inoculation prompting (IP) aims to keep undesired traits in training data from becoming part of a model's default behaviour. IP applies the same inoculation prompt to all training examples and it leaves underspecified how the desired and undesired traits (DT and UT) should activate. Two failures follow: The model develops backdoors: conditional vulnerabilities through which prompts that do not directly request the UT can still elicit it. We call this UT leakage.The desired trait weakens under ordinary prompts. We introduce Stratified Inoculation Prompting (SIP), which uses diverse prompts on safe examples (DT-only). Compared to standard IP, SIP significantly reduces leakage and retains more of the desired trait. SIP requires little safe data: a 5% DT-only pool oversampled to 25% is enough for good performance. Figure 1. IP compared to SIP. In SIP, uncertain and contaminated examples keep the inoculation prompt as in standard IP, while confidently safe examples are oversampled and trained under prompts drawn from multiple non-eliciting control prompt categories. Since SIP relies on filtering examples, we test its robustness to classification errors and find a strong asymmetry: failing to inoculate examples containing the undesired trait reintroduces it, whereas unnecessarily inoculating safe examples is benign. This suggests [...] ---Outline:(06:16) Inoculation Prompting Underspecifies the Intended Conditionalisation(09:20) What successful conditionalisation looks like(10:14) Stratified Inoculation Prompting (SIP)(12:49) SIP reduces leakage while preserving the desired trait(14:14) Control prompts recover the desired trait, diverse prompts narrow the backdoor's activation boundary(16:42) Oversampling reduces the need for distinct safe data(18:54) SIP reduces Emergent Misalignment more than Uniform IP(20:13) Data filtering errors have an asymmetric impact(23:18) Limiting residual access to the undesired trait(23:50) Diluting the prompt-trait association(26:06) Password-locking the inoculation prompt(29:23) Limitations --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/FS7GFsGsH7CSQLahy/don-t-inoculate-everything-stratified-inoculation-prompting --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  50. 201

    “Don’t Build Mindreading” by Celer

    “I have sworn upon the altar of god, eternal hostility against every form of tyranny over the mind of man” –Thomas Jefferson, letter to Benjamin Rush Context: Conduit is building datasets to enable telepathy, to use their term. I saw my grandfather lose control over his own fingers: what I would have given to offer him a headband that read his thoughts. Through novel technologies we have liberated almost all Americans from farming, driven the child and infant mortality rate from the pre-industrial half to less than half a percent in the best-performing countries, and rendered famine a political choice: broad-based improvements in efficiency are good and should be pursued for their own sake. Telepathy offers more: we could create trust through verified honesty, helping us ensure prosperity and peace. DARPA is already looking into “preconscious” thoughts for suicide prevention. There's also a strong argument centered on AI Safety: the models are becoming superhuman, and this is technology to allow us to keep pace, minimize hostile competition, and perhaps survive into the future. This is what Conduit is promising. Unfortunately, mindreading will have other effects. Oskar Schindler saved over 1,000 Jewish lives during the Holocaust. He did it by [...] --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/CAdG5dzkWrrK2NQg8/don-t-build-mindreading --- Narrated by TYPE III AUDIO.

Type above to search every episode's transcript for a word or phrase. Matches are scoped to this podcast.

Searching…

We're indexing this podcast's transcripts for the first time — this can take a minute or two. We'll show results as soon as they're ready.

No matches for "" in this podcast's transcripts.

Showing of matches

No topics indexed yet for this podcast.

Loading reviews...

ABOUT THIS SHOW

Audio narrations of LessWrong posts.

HOSTED BY

LessWrong

Frequently Asked Questions

How many episodes does LessWrong (30+ Karma) have?

LessWrong (30+ Karma) currently has 50 episodes available on PodParley. New episodes are automatically indexed when they're published to the podcast feed.

What is LessWrong (30+ Karma) about?

Audio narrations of LessWrong posts.

How often does LessWrong (30+ Karma) release new episodes?

LessWrong (30+ Karma) has 50 episodes. Check the episode list to see recent publication dates and frequency.

Where can I listen to LessWrong (30+ Karma)?

You can listen to LessWrong (30+ Karma) on PodParley by clicking any episode. We provide an embedded audio player for direct listening, and you can also subscribe via your preferred podcast app using the RSS feed.

Who hosts LessWrong (30+ Karma)?

LessWrong (30+ Karma) is created and hosted by LessWrong.
URL copied to clipboard!