Google 提出的新模型架构 MoR,Transformer 之外的一条新路径 episode artwork

EPISODE · Jul 20, 2025 · 7 MIN

Google 提出的新模型架构 MoR,Transformer 之外的一条新路径

from Daily LLM Papers

Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation这篇研究论文介绍了Mixture-of-Recursions (MoR),这是一个针对大型语言模型(LLMs)效率的新框架。MoR通过参数共享(重复使用一套共享层)和自适应计算(轻量级路由器动态分配不同递归深度给单个令牌)来降低计算和内存成本。该研究探讨了两种主要的路由策略——专家选择和令牌选择——以及两种键值(KV)缓存策略,以优化性能。实验结果表明,MoR在相同的计算预算下,显著提升了LLMs的验证困惑度和少量样本准确性,并实现了更高的推理吞吐量,证明其在降低大型模型成本方面是有效的。论文原文:https://www.alphaxiv.org/abs/2507.10524前往小宇宙评论区与主播互动

Episode metadata supplied by the publisher feed · Published Jul 20, 2025

Embed this episode

Ready to play

Google 提出的新模型架构 MoR,Transformer 之外的一条新路径

0:00 7:07

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily LLM Papers?

This episode is 7 minutes long.

When was this Daily LLM Papers episode published?

This episode was published on July 20, 2025.

Can I download this Daily LLM Papers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!