Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps episode artwork

EPISODE · May 23, 2026 · 19 MIN

Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 79 | cs.CL, cs.AI Authors: Yanke Zhou, Yiduo Li, Hanlin Tang, Maohua Li, Kan Liu, Lan Tao, Lin Qu, Yuan Yao, Xiaoxing Ma Title: Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps Arxiv: http://arxiv.org/abs/2605.16928v1 Abstract: Long-context inference in large language models is bottlenecked by the quadratic cost of full attention. Existing efficient alternatives often rely either on native sparse training or on heuristic token eviction, creating an undesirable trade-off among efficiency, training cost, and accuracy. In this work, we show that full-attention LLMs are already intrinsically sparse and can be transformed into highly sparse models with only minimal adaptation. Our approach is built on three observations: (1) only a small subset of attention heads truly requires full long-context processing; (2) long-range retrieval is governed primarily by a low-dimensional subspace, allowing relevant tokens to be retrieved efficiently with a 16-dimensional indexer; and (3) the useful token budget is strongly query-dependent, making dynamic top-$p$ selection more suitable than fixed top-$k$ sparsification. Based on these insights, we propose RTPurbo, which retains the full KV cache only for retrieval heads and introduces a lightweight token indexer for sparse attention. By exploiting the model's intrinsic sparsity, RTPurbo achieves sparsification with only a few hundred training steps. Experiments on long-context benchmarks and reasoning tasks show that RTPurbo preserves near-lossless accuracy while delivering substantial efficiency gains, including up to a 9.36$\times$ prefill speedup at 1M context and about a 2.01$\times$ decode speedup. These results suggest that strong sparse inference can be obtained from standard full-attention training without expensive native sparse pretraining.

Episode metadata supplied by the publisher feed · Published May 23, 2026

Embed this episode

NOW PLAYING

Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

0:00 19:18

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 19 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on May 23, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!