EPISODE · May 18, 2026 · 14 MIN
Build a Minimal LLM Evaluation Loop That Catches Regressions While You Sleep
from The Stateless Founder · host Santi, Kira
If you're running AI workflows in production and don't have a way to test whether a prompt change or model deprecation just broke something, you're shipping regressions to paying customers. This episode builds the fix: a minimal evaluation loop with golden test sets, AI judge prompts, pairwise A/B testing, and a Monday scorecard that ties quality to cost. Santi walks through his three-output-type harness that catches problems before clients do, while Kira challenges the reliability of AI judges and explores human-in-the-loop sampling strategies. Learn to build 15-20 case golden sets from real production data, implement G-Eval rubric prompts with cross-family judges, set up CI regression gates with Promptfoo, and track cost-per-100-jobs alongside pass rates in a weekly scorecard. Includes starter templates for email rewrites, JSON extraction, and content summarization, plus deprecation monitoring for OpenAI and Anthropic model lifecycles.
Embed this episode
Ready to play
Build a Minimal LLM Evaluation Loop That Catches Regressions While You Sleep
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.