EPISODE · Sep 8, 2026 · 43 MIN
AI Performance Testing: How to Scale Agentic AI with Kandasamy Selvaraj
from TestGuild Automation Podcast
Everything you knew about performance testing changes when the system you're testing is non deterministic. In this episode of the TestGuild Automation Podcast, Joe Colantonio sits down with Kandasamy Selvaraj, Principal Architect and author of the free book Rethinking Performance Engineering for Agentic AI, to unpack what it really takes to move an AI agent from a working demo to an enterprise system handling millions of conversations per hour. Checkout his free book: https://leanpub.com/agentic-ai-performance Kandasamy shares the practical playbook he's built running agentic AI in production, including why the same request can take three seconds one run and eight seconds the next, how to use harnesses to bound tool calls, reasoning loops, and token budgets, and why your SRE dashboard can look perfectly healthy while your token costs quietly balloon to five times baseline. You'll learn how his team: Shifts performance gates left into every commit with JMeter Shifts right with synthetic monitors on blue green deployments Uses Langfuse and OpenTelemetry to spot context bloat before it hits production. You'll also hear how to: Slash AI costs with prompt caching Conversation capping Routing simple queries to cheaper models Why you should load test at the API layer before touching the UI, How to keep stubs honest with production sampled latency Why one misbehaving agent can starve every other agent sharing the same provider. If you're a tester, performance engineer, SRE, or architect building on LLMs, this conversation will change how you think about scale.
Embed this episode
Ready to play
AI Performance Testing: How to Scale Agentic AI with Kandasamy Selvaraj
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.