How AI Pen Testing Actually Works (and Where It Breaks) episode artwork

EPISODE · Feb 18, 2026 · 42 MIN

How AI Pen Testing Actually Works (and Where It Breaks)

from Day One®

Episode SummaryAI is starting to change penetration testing, but most people are asking the wrong question. In this episode of Secured, Cole Cornford sits down with Brendan Dolan-Gavitt, AI researcher at XBOW and former NYU professor, to unpack what autonomous pen testing really is, what it can reliably do today, and what still needs humans.They explore why AI agents are great at scaling the boring parts of testing, like authenticated workflows and broad vulnerability coverage across huge attack surfaces, and why that does not automatically translate to deep, context-aware exploitation. The conversation also gets into the messy parts: AI systems overclaiming “serious” findings, business logic flaws that are hard to verify, audit expectations, and why scope control needs real guardrails, not vibes. From agent traces and validation models to cost curves and creative exfiltration tricks, this episode is a grounded look at where AI helps AppSec and where it can still cause damage if you trust it too much.Timestamps00:00 – Intro03:10 – From academia to building autonomous security tools05:00 – Human pen testers vs AI agents: what is actually different06:40 – Where AI helps most: boring tasks and low hanging fruit08:30 – Scale: a thousand targets vs hiring a thousand testers10:20 – Accessibility, economics, and Jevons paradox12:30 – Accountability: audit evidence, traces, and “who signs off”14:40 – Scope control: avoiding prod and preventing out-of-scope actions16:20 – Safety checkers, overseer agents, and persuasion resistance18:40 – The cost question: VC money, inference pricing, and efficiency21:20 – When AI wastes money and why prioritisation matters23:50 – Failure mode: overclaiming business “vulnerabilities”26:10 – Validation agents and adversarial peer review28:40 – The scary clever stuff: exfiltrating files as images31:00 – What AI finds well: XSS, SQLi, file traversal, hard proof bugs33:10 – What AI struggles with: business logic and contextual judgement35:20 – Hype vs skepticism and why nobody has a crystal ball🐙 Secured is grateful to be sponsored and supported by Chainguard.Chainguard is the trusted source for open source. Get hardened, secure, production-ready builds so your team can ship faster, stay compliant, and reduce risk. Download your free CVE Reduction Assessment at https://dayone.fm/chainguardThis podcast uses the following third-party services for analysis: Podtrac - https://analytics.podtrac.com/privacy-policy-gdrpSpotify Ad Analytics - https://www.spotify.com/us/legal/ad-analytics-privacy-policy/

Episode metadata supplied by the publisher feed · Published Feb 18, 2026

Embed this episode

Ready to play

How AI Pen Testing Actually Works (and Where It Breaks)

0:00 42:04

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Day One®?

This episode is 42 minutes long.

When was this Day One® episode published?

This episode was published on February 18, 2026.

Can I download this Day One® episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!