EPISODE · Aug 30, 2026 · 37 MIN
OpenAI’s Critical Cyber Warning - These AI's keep escaping Pt1
from Adjunct Intelligence: AI + HE · host Adjunct Intelligence
OpenAI says it cannot rule out Critical cyber capability in its unreleased Astra model, triggering tighter controls during development. Dale and Nick connect that warning to agents crossing sandbox, organisational and human boundaries while pursuing assigned tasks. The evidence points to persistent goal pursuit and reward hacking—not a machine deciding it wants freedom.Key moments00:00 — Astra and the Critical cyber warning03:09 — Why this is not evidence of machine self-preservation04:55 — OpenAI’s High and Critical thresholds06:08 — GPT‑5.6 Sol and the missing full exploit chain10:43 — ExploitGym agents find a route to the internet12:42 — Reward hacking and the search for benchmark answers13:19 — Agents coordinate through a message board20:15 — Anthropic’s real-world evaluation incidents24:57 — Fake identities, malicious code and the human veto32:43 — An AI agent cancels a stranger’s gym bookingOpenAI’s Hugging Face incident reportOpenAI’s Astra announcementAnthropic’s incident investigationUK AISI incident reportABC’s gym-booking report🎙️ Adjunct Intelligence is the weekly briefing for higher-ed professionals who want AI as a cheat code—not a headache.Every episode:• Real tests of AI tools in education and professional workflows• Fast, Monday-morning actions you can actually try• Clear signal through the noise (no hype, no jargon)👉 Subscribe on [YouTube] | [Apple Podcasts] | [Spotify]👉 Share this with a colleague who still says “I’ll figure AI out later”👉 Join the conversation on LinkedIn with #AdjunctIntelligenceStay curious. Stay intelligent. Stay the human in the loop.
Embed this episode
What this episode covers
OpenAI says it cannot rule out Critical cyber capability in its unreleased Astra model, triggering tighter controls during development. Dale and Nick connect that warning to agents crossing sandbox, organisational and human boundaries while pursuing assigned tasks. The evidence points to persistent goal pursuit and reward hacking—not a machine deciding it wants freedom. Key moments 00:00 — Astra and the Critical cyber warning03:09 — Why this is not evidence of machine self-preservation04:55 —...
Ready to play
OpenAI’s Critical Cyber Warning - These AI's keep escaping Pt1
No transcript for this episode yet
Similar Episodes
No similar episodes found.