Learning to Summarize with Human Feedback episode artwork

EPISODE · Apr 27, 2026 · 13 MIN

Learning to Summarize with Human Feedback

from Mastering Language Models: From Architecture to Optimization

Maya and Leo dig into OpenAI's Learning to Summarize from Human Feedback — the paper where pairwise human picks replaced reference matching as the training target. They walk the pipeline as three stations (the Two-Card Choice, the Borrowed Judge, the Tether), stage the real fight between preference optimization and cheap reproducible metrics, and end on the over-optimization curve where the judge's score keeps climbing while human preference falls. Sources: • Learning to Summarize from Human Feedback: https://arxiv.org/pdf/2009.01325 • Learning to summarize with human feedback: https://openai.com/index/learning-to-summarize-with-human-feedback/ • summarize-from-feedback: https://github.com/openai/summarize-from-feedback • ROUGE: A Package for Automatic Evaluation of Summaries: https://aclanthology.org/W04-1013/

Episode metadata supplied by the publisher feed · Published Apr 27, 2026

Embed this episode

NOW PLAYING

Learning to Summarize with Human Feedback

0:00 13:05

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Mastering Language Models: From Architecture to Optimization?

This episode is 13 minutes long.

When was this Mastering Language Models: From Architecture to Optimization episode published?

This episode was published on April 27, 2026.

Can I download this Mastering Language Models: From Architecture to Optimization episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!