EPISODE · Apr 27, 2026 · 14 MIN
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
from Mastering Language Models: From Architecture to Optimization
Maya and Leo dig into Anthropic's helpful-and-harmless RLHF paper: two opposite data-collection payrolls, one preference model serving two masters, the weekly online refresh that keeps the judge informed, the split-judge robustness test that exposes reward gaming, and a staged fight over whether the alignment tax is real. Sources: • Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback: https://arxiv.org/pdf/2204.05862 • Human preference data for Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback: https://github.com/anthropics/hh-rlhf • Training language models to follow instructions with human feedback: https://arxiv.org/abs/2203.02155
Embed this episode
NOW PLAYING
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.