Qwen团队:组序列策略优化算法GSPO episode artwork

EPISODE · Jul 26, 2025 · 7 MIN

Qwen团队:组序列策略优化算法GSPO

from Daily LLM Papers

原文:Group Sequence Policy Optimization本来源介绍了组序列策略优化 (GSPO),是一种用于训练大型语言模型的新型强化学习算法。该算法通过基于序列似然定义重要性比率并执行序列级剪辑、奖励和优化来解决现有算法(如 GRPO)在训练巨型模型时遇到的不稳定性问题。文章指出,GRPO 的不稳定性源于其令牌级重要性采样权重的错误应用,导致高方差训练噪声和模型崩溃。GSPO 则通过其序列级方法显著提高了训练的稳定性、效率和性能,特别是在 Mixture-of-Experts (MoE) 模型的强化学习训练中,消除了对复杂稳定策略的需求,并简化了强化学习基础设施的设计。前往小宇宙评论区与主播互动

Episode metadata supplied by the publisher feed · Published Jul 26, 2025

Embed this episode

Ready to play

Qwen团队:组序列策略优化算法GSPO

0:00 7:58

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily LLM Papers?

This episode is 7 minutes long.

When was this Daily LLM Papers episode published?

This episode was published on July 26, 2025.

Can I download this Daily LLM Papers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!