EPISODE · Apr 19, 2025 · 13 MIN
Minimalist LLM Reasoning: Rejection Sampling to Reinforcement
from Best AI papers explained · host Enoch H. Kang
This paper investigates reinforcement learning methods for fine-tuning large language models on complex reasoning tasks, particularly mathematical problems. The authors analyze GRPO, a successful but poorly understood algorithm, and surprisingly find that a simpler rejection sampling method, RAFT, achieves comparable results by training only on positively rewarded samples. Their analysis reveals that GRPO's effectiveness stems mainly from discarding prompts with entirely incorrect responses, leading them to propose Reinforce-Rej, a refined algorithm that also filters entirely correct samples for improved efficiency and stability. The study advocates for RAFT as a robust baseline and suggests future work prioritize principled negative sample integration over indiscriminate use.
Embed this episode
NOW PLAYING
Minimalist LLM Reasoning: Rejection Sampling to Reinforcement
No transcript for this episode yet
Similar Episodes
No similar episodes found.