SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers episode artwork

EPISODE · Sep 2, 2026 · 21 MIN

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 53 | cs.LG Authors: Shaowen Wang, Ge Zhang, Kairong Luo, Yuhao Wu, Shaofan Liu, Jiaheng Liu, Wenhao Huang, Shen Yan, Jian Li Title: SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Arxiv: http://arxiv.org/abs/2609.01343v1 Abstract: Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. Through a series of ablations, we arrive at a recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice while matching the unlooped Baseline on all three budgets. We scale SMELT across four sizes up to 54B non-embedding parameters and fit a separate Chinchilla-style scaling law for each architecture. SMELT's loss drops faster with compute, saving 6.8--18.0\% of training FLOPs on the compute-optimal frontier. The advantage transfers to downstream benchmarks beyond what validation loss predicts, is largest on Code, and grows with sample length and the number of in-context examples. Mechanistic analysis shows that the second visit reduces the attention sink and redirects mass toward content-relevant tokens, an inductive bias that may underlie the observed performance gains. These results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.

Episode metadata supplied by the publisher feed · Published Sep 2, 2026

Embed this episode

NOW PLAYING

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

0:00 21:08

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 21 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on September 2, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!