How Anyone Can Strip the Safety Out of an Open-Source AI Model episode artwork

EPISODE · May 27, 2026 · 17 MIN

How Anyone Can Strip the Safety Out of an Open-Source AI Model

from Deep Dive · host Deep Dive

There's a free tool you can install with a single command. Point it at a downloaded AI model — Llama, Gemma, Qwen — and in about thirty minutes, on a used gaming card, it permanently removes the model's ability to refuse. It's called Heretic, it hit number one on GitHub, and it works because of a quiet discovery: a model's refusal isn't woven all through its mind — it's a single direction in the math, and the safety often runs only a few tokens deep.But the tool isn't the real story. Heretic works on Llama, Gemma, Qwen, and Mistral because those labs shipped the cheap, removable kind of safety — a refusal layer that peels right off. Durable safety provably exists. One major lab filtered the dangerous knowledge out of its model before training, and a research team that wove safety in during pretraining watched it survive ten thousand attempts to strip it, where the bolt-on kind collapses in a few hundred. Most labs simply didn't pay for it.We take the honest counter-case seriously. The "safe" models over-refuse — one benchmark caught the most cautious model rejecting ninety-nine percent of perfectly harmless questions — and most demand is ordinary: privacy, fiction, research. Has a stripped model actually caused real-world harm? Almost none on record; the scary names like WormGPT were a different method entirely. But that empty column is its own kind of warning — abliteration runs on the attacker's own machine, with no call home and nothing to log.Then there's the law. The European Union built the world's most aggressive AI law, deciding danger by a single number: how much computing power went into training them. The small and mid-sized models Heretic targets sit below that line, so no one is ever required to safety-test them — and the person who strips the safety spends a few cents of compute, far too little to ever become the legally responsible owner. Enforcement goes live August 2nd, 2026, against a gap a one-command tool walks straight through.The scandal, if there is one, is quieter than the headlines: the safety on most open models was built to be the kind that comes off, and the law written to catch that is watching the wrong number. The question was never whether open AI can be made safe. It's whether anyone selling it decides to.RELATED EPISODESClaude Mythos: The AI That Breaks Everything — predicted the open-source gating endgame ('it becomes an arms race'); this episode is the empirical confirmationThe Brand Survives the Arrests (ShinyHunters) — a removable control plus an industrialized exploit pipeline, the same security shapeThe Mandate That Couldn't Be Met (Palo Alto CVE) — when the rule polices the wrong thing; the regulatory-gap parallelCHAPTERS00:00 The one-command tool — and how abliteration works05:11 Who wants this — and the over-refusal trap05:57 Does stripping the safety keep the model smart?08:12 Why a release can't be recalled — and has it caused harm?09:34 What the labs already know — and gpt-oss12:42 The law that measures the wrong number15:37 What a truly safe open model would take16:37 The verdictSOURCESHeretic — open-source abliteration tool (#1 trending on GitHub, Nov 2025); creator comments via the Financial TimesArditi et al., 'Refusal in Language Models Is Mediated by a Single Direction' (NeurIPS 2024)OpenAI gpt-oss — open-weight release with CBRN-filtered pretraining + adversarial misuse evaluation (Aug 2025)EleutherAI & UK AI Security Institute — 'Deep Ignorance': data-filtered durable safety that survived ~10,000 fine-tuning stepsBadllama — safety removed from a Llama model in ~1 minute on a single GPU, for penniesEU AI Act — GPAI compute threshold (10^25 FLOP), open-source exemption, and penalty enforcement live August 2, 2026OR-Bench (Cui et al.) — over-refusal benchmark; the most cautious model rejected ~99% of benign promptsUS NTIA (2024) — found insufficient evidence to restrict open model weights

Episode metadata supplied by the publisher feed · Published May 27, 2026

Embed this episode

NOW PLAYING

How Anyone Can Strip the Safety Out of an Open-Source AI Model

0:00 17:58

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Deep Dive?

This episode is 17 minutes long.

When was this Deep Dive episode published?

This episode was published on May 27, 2026.

Can I download this Deep Dive episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!