GRPO (Group Relative Policy Optimization) episode artwork

EPISODE · Feb 5, 2025 · 12 MIN

GRPO (Group Relative Policy Optimization)

from Large Language Model (LLM) Talk · host AI-Talk

Group Relative Policy Optimization (GRPO) is a reinforcement learning algorithm that enhances mathematical reasoning in large language models (LLMs). It is like training students in a study group, where they learn by comparing answers without a tutor. GRPO eliminates the need for a critic model, unlike Proximal Policy Optimization (PPO), making it more resource efficient. It calculates advantages based on relative rewards within the group and directly adds KL divergence to the loss function. GRPO uses both outcome and process supervision, and can be applied iteratively, further enhancing performance. This approach is effective at improving LLMs' math skills with reduced training resources.

Episode metadata supplied by the publisher feed · Published Feb 5, 2025

Embed this episode

NOW PLAYING

GRPO (Group Relative Policy Optimization)

0:00 12:31

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Large Language Model (LLM) Talk?

This episode is 12 minutes long.

When was this Large Language Model (LLM) Talk episode published?

This episode was published on February 5, 2025.

Can I download this Large Language Model (LLM) Talk episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!