深入剖析DAPO:大规模开源LLM强化学习系统 episode artwork

EPISODE · Jun 2, 2025 · 12 MIN

深入剖析DAPO:大规模开源LLM强化学习系统

from AI Podcast · host weedge

本期播客深入探讨了DAPO(解耦裁剪与动态采样策略优化)算法,这是一个在Qwen2.5-32B基础模型上实现AIME 2024测试50分的先进大规模强化学习系统。我们详细讨论了其四项关键技术:Clip-Higher、动态采样、词元级策略梯度损失和超长奖励修正,以及它们如何解决熵塌陷、梯度消失、长CoT场景下的学习不平衡和奖励噪声等问题,并介绍了其开放源代码、训练代码和精心处理的数据集对社区的贡献。

Episode metadata supplied by the publisher feed · Published Jun 2, 2025

Embed this episode

NOW PLAYING

深入剖析DAPO:大规模开源LLM强化学习系统

0:00 12:23

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of AI Podcast?

This episode is 12 minutes long.

When was this AI Podcast episode published?

This episode was published on June 2, 2025.

Can I download this AI Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!