RLVR Lets Models Fail Their Way to the Top episode artwork

EPISODE · Aug 12, 2025 · 49 MIN

RLVR Lets Models Fail Their Way to the Top

from YAAP (Yet Another AI Podcast) · host AI21

Think you know fine-tuning? If your answer is RLHF, you don’t. In this episode, Itay, who leads the Alignment group at AI21, gives a no-fluff crash course on RLVR (Reinforcement Learning with Verifiable Rewards), the method powering today’s smartest coding and reasoning models. He explains why RLVR beats RLHF at its own game, how “hard to solve, easy to verify” tasks unlock exploration without chaos, and the emergent behaviors you only get when models are allowed to screw up. If you want to actually understand RLVR (and use it), start here. Key topics: How RLVR outsmarts RLHF in real-world training The “verified rewards” trick that kills reward hacking Emergent skills you don’t get with hand-holding: self-verification, backtracking, multi-path reasoning Why coding models took a giant leap forward Practical steps to train (and actually benefit from) RLVR models

Episode metadata supplied by the publisher feed · Published Aug 12, 2025

Embed this episode

Ready to play

RLVR Lets Models Fail Their Way to the Top

0:00 49:10

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of YAAP (Yet Another AI Podcast)?

This episode is 49 minutes long.

When was this YAAP (Yet Another AI Podcast) episode published?

This episode was published on August 12, 2025.

Can I download this YAAP (Yet Another AI Podcast) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!