UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning episode artwork

EPISODE · Aug 28, 2025 · 19 MIN

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 23 | cs.LG Authors: Zihao Huang, Yu Bao, Qiyang Min, Siyan Chen, Ran Guo, Hongzhi Huang, Defa Zhu, Yutao Zeng, Banggu Wu, Xun Zhou, Siyuan Qiao Title: UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Arxiv: http://arxiv.org/abs/2508.18756v1 Abstract: While Mixture of Experts (MoE) models achieve remarkable efficiency by activating only subsets of parameters, they suffer from high memory access costs during inference. Memory-layer architectures offer an appealing alternative with very few memory access, but previous attempts like UltraMem have only matched the performance of 2-expert MoE models, falling significantly short of state-of-the-art 8-expert configurations. We present UltraMemV2, a redesigned memory-layer architecture that closes this performance gap. Our approach introduces five key improvements: integrating memory layers into every transformer block, simplifying value expansion with single linear projections, adopting FFN-based value processing from PEER, implementing principled parameter initialization, and rebalancing memory-to-FFN computation ratios. Through extensive evaluation, we demonstrate that UltraMemV2 achieves performance parity with 8-expert MoE models under same computation and parameters but significantly low memory access. Notably, UltraMemV2 shows superior performance on memory-intensive tasks, with improvements of +1.6 points on long-context memorization, +6.2 points on multi-round memorization, and +7.9 points on in-context learning. We validate our approach at scale with models up to 2.5B activated parameters from 120B total parameters, and establish that activation density has greater impact on performance than total sparse parameter count. Our work brings memory-layer architectures to performance parity with state-of-the-art MoE models, presenting a compelling alternative for efficient sparse computation.

Episode metadata supplied by the publisher feed · Published Aug 28, 2025

Embed this episode

NOW PLAYING

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning

0:00 19:39

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 19 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on August 28, 2025.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!