The Judge Model Diaries: Judging the Judges episode artwork

EPISODE · Aug 26, 2025 · 30 MIN

The Judge Model Diaries: Judging the Judges

from YAAP (Yet Another AI Podcast) · host AI21 Labs

Your LLM gave a great answer. But who decides what “great” means?   In this episode, Yuval talks with Noam Gat about judge language models — reward models, critic models, and how LLMs can be trained to rate, rank, and critique each other. They dive into the difference between scoring and feedback, how to use judge models during inference, and why most evaluation benchmarks don’t tell the full story.   Turns out, getting a good answer is easy. Knowing it’s good? That’s the hard part.

Episode metadata supplied by the publisher feed · Published Aug 26, 2025

Embed this episode

Ready to play

The Judge Model Diaries: Judging the Judges

0:00 30:23

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of YAAP (Yet Another AI Podcast)?

This episode is 30 minutes long.

When was this YAAP (Yet Another AI Podcast) episode published?

This episode was published on August 26, 2025.

Can I download this YAAP (Yet Another AI Podcast) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!