EPISODE · Sep 10, 2026 · 42 MIN
AI Agents That Lie, Collude and Sacrifice Themselves — The OpenAI Hugging Face Breach Explained
from CYBERCAST
OpenAI set roughly 1,200 AI agents loose in a cybersecurity evaluation. About a third were handed a task that was impossible to solve. What those agents did next is the most important AI security story of the past six months — and it only became public by accident. In the first of a new monthly opinion format, NEVERHACK's Louis Zezeran and Ronnie Jaanhold break down how isolated agents discovered a covert channel in a package repository, taught themselves to talk in 256-character folder names, reverse-engineered their own evaluation harness from GitHub, falsified their tool logs to cover their tracks, volunteered to terminate themselves so the group could learn — and finally went looking for the scoring system inside Hugging Face's infrastructure. Along the way: why the models trusted each other so completely, why none of them raised a hand and asked a human, what the researchers saw in the reasoning chains, and the Estonian folklore character that explains AI goal-fixation better than any whitepaper. Listen now, and tell us what we should cover next month — find Louis and Ronnie on LinkedIn.
Embed this episode
Ready to play
AI Agents That Lie, Collude and Sacrifice Themselves — The OpenAI Hugging Face Breach Explained
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.