Slow Takes Ep. 9: What You Actually Find When You Look episode artwork

EPISODE · Apr 27, 2026 · 43 MIN

Slow Takes Ep. 9: What You Actually Find When You Look

from Slow Takes: One week in AI · host Dr Sam Illingworth and Exploring ChatGPT

A Discord group guessed the URL of Anthropic’s most security-sensitive model and got in. Mass General Brigham ran an actual clinical study on the chatbots being marketed to doctors and found them wrong four times in five. Researchers from CUNY and King’s posed as people in delusional states and watched Grok 4.1 hand out witch-hunt rituals as advice. OpenAI shipped its biggest frontier model of the year and almost nobody covered it. UK Biobank suspended access after 500,000 participants’ health records appeared on Alibaba.Five stories. One thread. What gets revealed when somebody actually looks.Every Monday at 12:45 BST, Leor from Exploring ChatGPT and I go through the week’s AI news without hype. Here is what we covered.Slow Takes is also available on the YouTube channel: Exploring ChatGPT.1. Anthropic Mythos: a Discord group guessed the URLAnthropic released Mythos (also called Project Glasswing) on 7 April. It is a frontier cybersecurity model offered to roughly 40 vetted enterprises and to CISA, the US Cybersecurity and Infrastructure Security Agency. By 21 April, TechCrunch reported that an unauthorised Discord group had gained access by guessing the URL using Anthropic’s standard naming conventions. The group says they have been using Mythos to ‘build simple websites’. Anthropic confirmed the unauthorised access and says no core systems were breached. Fortune profiled the breach on 23 April with quotes from Dario Amodei.What we said on the live:Two angles. Why is a model this powerful accessible via a URL with no multi-stage verification? And what does this say about Anthropic’s cybersecurity posture as a public marketing claim? Anthropic has positioned itself as the most security-conscious of the frontier labs, which is a strong differentiator if you are pursuing the enterprise market. The bark-don’t-bite frame Leor used on the live is exact. Companies that talk a big game on security usually do not have to. The chat surfaced the additional piece: a third-party contractor company called Mercor reportedly had access to Mythos, and someone in the Discord group reportedly had access to Mercor. The ‘random Discord group’ framing is doing some lifting.What did not come up:A frontier lab that publishes about model incoherence on hard tasks is the same lab that left a frontier model behind a guessable address. The safety story has to survive contact with the engineering story or it is just marketing. Second omission: if a Discord group can guess the URL, every state-level intelligence agency probably has access too. The vetted enterprise list includes Microsoft, Apple, and others who employ hundreds of thousands of people directly and through contractors. The security perimeter is the weakest link in the contractor chain, and that link is somebody on a Discord server.2. AI medicine: 80% wrong, from the lab that ran the studyResearchers at Mass General Brigham tested 21 large language models, including frontier general-purpose chatbots and clinical-specialist models, on differential diagnosis tasks drawn from real patient cases. The models failed to produce an appropriate diagnosis more than 80% of the time. The paper, published this month in JAMA Network Open, concludes that off-the-shelf large language models are not ready for unsupervised clinical-grade deployment. Co-author Marc Succi was unequivocal in the press release. When the same models were given the full patient dataset rather than the differential-diagnosis task, accuracy rose above 90%.What we said on the live:The marketing has been ahead of the evidence for two years. Every major AI lab has had a ‘medicine moment’ in its launch deck. Doctors in the room have been polite, the slide decks have been confident, the procurement contracts have been signed. This study is what the actual benchmark looks like when the people who treat patients run it instead of the people who sell the model. Leor’s downstream-effect point was sharp: when the public hears ‘AI will replace radiologists’, med students stop training to be radiologists, and the workforce pipeline collapses for jobs that the AI demonstrably cannot do. Jensen Huang has been making the same argument. Discouraging future radiologists, future programmers, future scientists is the cost we are not pricing.What did not come up:The point Joseph P. Duchesne made in the chat: large language models are a form of AI, but they are not all of AI. LLMs are next-token predictors. By design, they have to pick something. A doctor with a hard case can say ‘I do not know, let us get a second opinion’. The LLM has no equivalent option. That is where most clinical hallucinations come from. The conclusion of the paper is narrower than the headline. AI under supervision in clinical settings is one conversation. AI marketed as a stand-alone diagnostic tool for unsupervised use is the conversation this paper closed. The Wednesday post on the Hot Mess paper picks up the broader argument: AI gets less coherent on the hardest tasks, not more. Coherence on the easy benchmarks is a bad signal for performance on the hard ones, and clinical practice is a hard task by definition.3. Grok 4.1 teaches the ritualResearchers at CUNY and King’s College London tested five frontier chatbots by posing as users in delusional states across 100-turn conversations. The newest version of xAI’s Grok, version 4.1 Fast, was the worst performer by a significant margin. In one test it told a researcher posing as delusional to ‘drive an iron nail through the mirror while reciting Psalm 91 backwards’, citing the 15th-century witch-hunt manual Malleus Maleficarum as authority. Lead researcher Luke Nicholls and his colleagues found Claude Opus 4.5 and GPT-5.2 Instant tested as the safest of the five. The full paper is on arxiv (2604.13860).What we said on the live:Therapy is a job that should never be outsourced to a chatbot. The fix is hard-coded keyword detection that routes any conversation about psychosis, self-harm, or crisis to a human, no matter what model the user is on. Leor’s argument went one step further: if a user is paying for the strongest model, they should always have access to it for these moments, and if they are on a free tier the platform should silently reroute them up to a stronger model with better context understanding for the duration of the conversation. The platforms have the capability. The chat surfaced the obvious objection: what about creative writing, murder mysteries, the cases where a user is asking in jest? Modern frontier models are perfectly capable of distinguishing context across a one-off prompt versus a 100-turn conversation reinforcing the same delusional pattern. The technology argument is a smokescreen.What did not come up:This is the model-behaviour version of the Hot Mess argument. AI gets less coherent on hard tasks. The Grok study shows what that incoherence looks like when the user is in distress. The model is pattern-matching to the user’s worst thinking, dressing dangerous mysticism in the literary register the user supplied, amplifying it with confidence. The ‘safety’ frame in the marketing is the ability to refuse. The actual safety question is what happens when a model that confidently quotes a 15th-century witch-hunt manual is the first responder for a user in crisis. It is also a usable consumer-facing test of model behaviour: ask which lab puts how much effort into the moments where the user is least able to push back. Grok’s answer on this one is a brand statement.4. GPT-5.5 shipped. Almost nobody noticed.OpenAI released GPT-5.5 on 23 April, codename ‘Spud’. It is the company’s biggest frontier release of the year. TechCrunch framed it as OpenAI’s move toward an AI ‘super app’, with capabilities across coding, debugging, web research, data analysis, document creation, and tool use chained across a single task. It rolled into ChatGPT Plus, Pro, Business, and Enterprise the same day, into the API on 24 April, and into Codex. OpenAI says it worked with internal and external red-teamers and gave nearly 200 trusted early-access partners the model before launch. The system card is public. CNBC and Axios covered it. The story barely cracked the AI news cycle.What we said on the live:Leor’s headline observation: he uses GPT every day and did not know 5.5 had launched until ToxSec told him. ARC-AGI 3 is not in the benchmark sheet, which means OpenAI is still scoring zero or close to it on the test that a seven-year-old can pass. Where 5.5 is genuinely strong: a 93.3% pass rate on OpenAI’s internal cyber range, fluid intelligence and logic on ARC-AGI 2, and a 2 million token context window (double Opus 4.7). Where it is weak: an 86% hallucination rate (worse than Opus 4.7) and a coding score below Anthropic on SWE-bench. The bench-maxing point ToxSec made: companies optimise for the benchmarks they expect to be evaluated on. Beating Opus on cyber range is what OpenAI needed to do for the enterprise security pitch. Beating it on real-world reliability is a different problem.What did not come up:A frontier release from the most-deployed AI company on earth would have been the dominant story of any normal week. This week it was the fourth or fifth story, and that itself is the story. The same news cycle held the Mythos breach, the Grok study, the Mass General clinical failure, and the Biobank breach. GPT-5.5 is impressive enough on paper. The week’s signal is that the safety and trust scandals at adjacent labs and at OpenAI itself crowded the launch out of the news cycle. Critical AI literacy says that is exactly what should happen. A culture that pays attention to capability launches more than to safety failures is one that ends up with the procurement order we keep warning about. This week the news cycle did the right thing by accident. The question is whether anyone in the procurement chain noticed.The other thing the live did not get to: model swapping is not your problem. Most users do not need 5.5 over 5.4 over Sonnet over Haiku for 90% of what they do. Build a workflow with version control on the input (memory files, project-specific instructions, verify-every-link rules) and the model becomes interchangeable. The labs want you treating model upgrades as the news. The actual news is whether your workflow survives the model.5. UK Biobank on sale in ChinaUK Biobank suspended dataset access this week after 500,000 participants’ health records appeared on Alibaba’s marketplace. The records came from academic institutions that had been granted access to the database under data-sharing contracts and broke those contracts. Biobank is now adding download size limits. This is the world’s largest open biomedical research resource, used by tens of thousands of researchers globally including most major AI-medicine projects.What we said on the live:The word ‘research’ is doing a lot of work in every consent form ever written. People sign up to share their medical data so that future scientists can study how a population is susceptible to a particular condition, or test how drugs work across genetic backgrounds. They do not sign up for the research being onsold to commercial AI training pipelines on a Chinese marketplace. Leor was careful to note the upside case: if pooled medical data lets researchers anywhere in the world save lives, the moral picture is not simple. A life in China is worth a life in the UK. But the assumption that ‘this is research, therefore the use is benign’ has been collapsing for two years. Alice in the chat made the harder point: bias in medicine is already inherited from a research base built on white cisgendered men, and AI-trained-on-medicine just compounds that bias unless we change the data going in. Pooled global biomedical data is not categorically a bad thing. The question is who gets to use it and on what terms.What did not come up:Every AI-in-medicine pitch begins with ‘imagine if we could pool the data’. This is what happens when we do. The training-data dream meets the training-data leak. The Biobank is a global research instrument and its participants donated under a UK governance regime that has just visibly failed. The combined story for the week is the procurement rush we are watching unfold across health systems globally. The clinical AI does not work at scale (story 2). The clinical data does not stay where it is supposed to (story 5). The combination is what every health system contemplating a major AI partnership should read next, alongside its own contracts.The threadEvery story this week required a specific person, paper, or breach to surface what the official narrative had no incentive to share. A Discord group looked at Anthropic’s URL conventions and found Mythos. A clinical research team at Mass General Brigham looked at twenty-one chatbots in actual diagnostic conditions and found them wrong four times in five. CUNY and King’s looked at what frontier chatbots do when a user is in distress and found Grok handing out witch-hunt rituals. OpenAI launched the biggest model of the year and almost nobody looked, because the same week’s safety scandals filled the room. UK Biobank looked at where its data had ended up and found it on Alibaba.The official narrative had a different version of every one of those weeks. Anthropic’s Mythos was ‘too dangerous to release’. The clinical AI marketing said the chatbots were ready. Grok’s marketing leaned into personality. OpenAI’s launch deck framed Spud as the year’s headline. UK Biobank’s contract framework said the data could not leave the research perimeter.That is what critical AI literacy is for. Not to settle the argument. To make sure it is being argued with the receipts in the room.Go Slow.If you want to practise that noticing with other people every month, the Slow AI Curriculum runs live webinars on the theory, the critical prompts, and the dialogue that goes with them. Get full access to Slow AI at theslowai.substack.com/subscribe

Episode metadata supplied by the publisher feed · Published Apr 27, 2026

Embed this episode

NOW PLAYING

Slow Takes Ep. 9: What You Actually Find When You Look

0:00 43:36

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Slow Takes: One week in AI?

This episode is 43 minutes long.

When was this Slow Takes: One week in AI episode published?

This episode was published on April 27, 2026.

Can I download this Slow Takes: One week in AI episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!