Build a Minimal LLM Evaluation Loop That Catches Regressions While You Sleep episode artwork

EPISODE · May 18, 2026 · 14 MIN

Build a Minimal LLM Evaluation Loop That Catches Regressions While You Sleep

from The Stateless Founder · host Santi, Kira

If you're running AI workflows in production and don't have a way to test whether a prompt change or model deprecation just broke something, you're shipping regressions to paying customers. This episode builds the fix: a minimal evaluation loop with golden test sets, AI judge prompts, pairwise A/B testing, and a Monday scorecard that ties quality to cost. Santi walks through his three-output-type harness that catches problems before clients do, while Kira challenges the reliability of AI judges and explores human-in-the-loop sampling strategies. Learn to build 15-20 case golden sets from real production data, implement G-Eval rubric prompts with cross-family judges, set up CI regression gates with Promptfoo, and track cost-per-100-jobs alongside pass rates in a weekly scorecard. Includes starter templates for email rewrites, JSON extraction, and content summarization, plus deprecation monitoring for OpenAI and Anthropic model lifecycles.

Episode metadata supplied by the publisher feed · Published May 18, 2026

Embed this episode

Ready to play

Build a Minimal LLM Evaluation Loop That Catches Regressions While You Sleep

0:00 14:05

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of The Stateless Founder?

This episode is 14 minutes long.

When was this The Stateless Founder episode published?

This episode was published on May 18, 2026.

Can I download this The Stateless Founder episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!