Evals: How Do You Know Which AI Model to Trust? episode artwork

EPISODE · Aug 13, 2026 · 41 MIN

Evals: How Do You Know Which AI Model to Trust?

from System Prompt · host Peter

READ THE FULL EPISODE PAGEhttps://devmesh.tech/podcast/how-to-know-which-ai-model-to-trustThe AI model at the top of a leaderboard may not be the best model for your system.Because the leaderboard is not testing your system.In Episode 21 of System Prompt, Peter and Val break down AI evals: what benchmarks measure, why the harness matters, and how to test models against the work you actually expect them to do.Peter walks through a custom eval across more than 20 local and open models covering tool calling, extraction, instruction following, and real-world coding tasks.The results were surprising. Smaller models matched or beat much larger ones. Turning reasoning on sometimes made performance worse.The bigger lesson: an eval measures more than the model. Quantization, runtime, token budgets, reasoning settings, parsers, and timeouts can all affect the result.WHAT WE DISCUSS• What AI evals actually measure• Why leaderboards only tell part of the story• How the harness changes model performance• Quantization, runtimes, and configuration• Building evals around real workloads• Tool calling, extraction, instruction following, and coding• Why repetition and consistency matter• Thinking vs non-thinking configurations• Routing tasks to different models• Finding problems in your own systemKEY TAKEAWAYSTHE BEST MODEL DEPENDS ON THE JOBA benchmark measures performance on a particular test. It does not automatically tell you which model is best for your application.A coding agent, extraction pipeline, chatbot, and tool-using agent all need different things. Start with the workload, then choose the eval.THE HARNESS IS PART OF THE RESULTModels do not operate alone.The harness creates prompts, exposes tools, manages token limits, parses responses, and decides whether a task succeeded.Change the harness, configuration, quantization, or runtime and you can change the result.TEST THE MODEL YOU ARE ACTUALLY RUNNINGA full-precision benchmark is useful reference data, but it is not the same experiment as running a Q4 model through a local runtime.Your production configuration is part of the evaluation.REPETITION MATTERSOne successful run does not prove reliability.Running tasks multiple times exposes models that score well once but behave inconsistently.For production systems, stability matters.REASONING IS NOT ALWAYS BETTERThinking modes helped some models and hurt others.In some cases reasoning increased token use, hit time or output budgets, or reduced consistency.The right configuration has to be measured against the task.EVALS ENABLE ROUTINGThe best architecture may not use one model for everything.A smaller model may handle chat, extraction, or tool calling while another handles coding or harder reasoning.Once you know where each model succeeds and fails, routing stops being guesswork.EVALS TEST YOUR SYSTEM TOOThe eval process also exposed problems in Peter's own gateway and harness.Some apparent model failures were really token limits, timeouts, parsing issues, or infrastructure problems.CHAPTERS00:00 Episode 21 and the 1%00:48 What Are AI Evals?02:48 Model Capability and Benchmarks05:07 Why the Harness Matters09:48 Quantization and Fair Comparisons11:30 Building a Custom Eval Suite19:04 Repetition and Reliability20:50 The Model Results23:17 When Thinking Hurts Performance29:55 Accuracy Versus Token Cost30:55 Routing Tasks to Different Models37:00 Evals Finding Bugs in the System40:36 Closing Thoughts

Episode metadata supplied by the publisher feed · Published Aug 13, 2026

Embed this episode

Ready to play

Evals: How Do You Know Which AI Model to Trust?

0:00 41:12

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of System Prompt?

This episode is 41 minutes long.

When was this System Prompt episode published?

This episode was published on August 13, 2026.

Can I download this System Prompt episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!