Principles of Evals: The Future of GenAI Evaluation (E.43) episode artwork

EPISODE · May 29, 2026 · 53 MIN

Principles of Evals: The Future of GenAI Evaluation (E.43)

from Free Form AI · host Michael Berk

LLMs are optimized to sound convincing—not to know when they’re wrong. In this episode, Deanna Emery breaks down why hallucinations are fundamentally tied to how language models work, why confidence is often disconnected from correctness, and how better evaluation strategies can make AI systems more reliable in production. We also get into uncertainty, semantic reasoning, and what humans still do better than models.00:00 — Why LLMs hallucinate confidently09:00 — The limits of current eval systems18:00 — Why uncertainty matters in AI27:00 — Semantic reasoning vs memorization38:00 — What humans still do better than modelsThe biggest risk in AI isn’t wrong answers. It’s wrong answers delivered with confidence.

Episode metadata supplied by the publisher feed · Published May 29, 2026

Embed this episode

NOW PLAYING

Principles of Evals: The Future of GenAI Evaluation (E.43)

0:00 53:55

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Free Form AI?

This episode is 53 minutes long.

When was this Free Form AI episode published?

This episode was published on May 29, 2026.

Can I download this Free Form AI episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!