Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback episode artwork

EPISODE · Apr 27, 2026 · 14 MIN

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

from Mastering Language Models: From Architecture to Optimization

Maya and Leo dig into Anthropic's helpful-and-harmless RLHF paper: two opposite data-collection payrolls, one preference model serving two masters, the weekly online refresh that keeps the judge informed, the split-judge robustness test that exposes reward gaming, and a staged fight over whether the alignment tax is real. Sources: • Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback: https://arxiv.org/pdf/2204.05862 • Human preference data for Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback: https://github.com/anthropics/hh-rlhf • Training language models to follow instructions with human feedback: https://arxiv.org/abs/2203.02155

Episode metadata supplied by the publisher feed · Published Apr 27, 2026

Embed this episode

NOW PLAYING

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

0:00 14:02

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Mastering Language Models: From Architecture to Optimization?

This episode is 14 minutes long.

When was this Mastering Language Models: From Architecture to Optimization episode published?

This episode was published on April 27, 2026.

Can I download this Mastering Language Models: From Architecture to Optimization episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!