SageAttention2 and Fast Exact INT4 Attention episode artwork

EPISODE · Jun 18, 2026

SageAttention2 and Fast Exact INT4 Attention

from AI Post Transformers

This episode explores SageAttention2, an ICML 2025 paper on making exact transformer attention faster without changing the underlying computation, focusing on why long-context models still pay a steep quadratic cost and why exact kernels remain important despite sparse and linear alternatives. It explains the paper’s central claim that aggressive low-precision attention can work only with careful numerical repair: queries and keys are pushed to INT4, attention-weight and value computation moves toward FP8, and outlier-smoothing ideas inspired by SmoothQuant are used to keep softmax-sensitive logits from collapsing. The discussion highlights the paper’s most concrete systems contribution, per-thread INT4 quantization aligned to GPU thread fragments and PTX `mma` execution, which aims to get fine-grained scaling without losing the performance win to dequantization overhead. A listener would find it interesting because the episode turns a seemingly narrow kernel optimization into a broader argument about hardware-software co-design, showing how much engineering is required to make lower-bit attention practical rather than just theoretically faster. Sources: 1. SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization — Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, Jianfei Chen, 2024 http://arxiv.org/abs/2411.10958 2. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning — Tri Dao, 2023 https://scholar.google.com/scholar?q=FlashAttention-2%3A+Faster+Attention+with+Better+Parallelism+and+Work+Partitioning 3. INT-FlashAttention: Enabling Flash Attention for INT8 Quantization — Shimao Chen, Zirui Liu, Zhiying Wu, et al., 2024 https://scholar.google.com/scholar?q=INT-FlashAttention%3A+Enabling+Flash+Attention+for+INT8+Quantization 4. SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration — Jintao Zhang, Jia Wei, Haofeng Huang, Pengle Zhang, Jun Zhu, Jianfei Chen, 2025 https://scholar.google.com/scholar?q=SageAttention%3A+Accurate+8-Bit+Attention+for+Plug-and-play+Inference+Acceleration 5. SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization — Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, Jianfei Chen, 2025 https://scholar.google.com/scholar?q=SageAttention2%3A+Efficient+Attention+with+Thorough+Outlier+Smoothing+and+Per-thread+INT4+Quantization 6. Understanding and Overcoming the Challenges of Efficient Transformer Quantization — Yelysei Bondarenko, Markus Nagel, Tijmen Blankevoort, 2021 https://scholar.google.com/scholar?q=Understanding+and+Overcoming+the+Challenges+of+Efficient+Transformer+Quantization 7. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models — Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, Song Han, 2023 https://scholar.google.com/scholar?q=SmoothQuant%3A+Accurate+and+Efficient+Post-Training+Quantization+for+Large+Language+Models 8. Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling — Xiuying Wei, Yunchen Zhang, Yuhang Li, Xiangguo Zhang, Ruihao Gong, Jinyang Guo, Xianglong Liu, 2023 https://scholar.google.com/scholar?q=Outlier+Suppression%2B%3A+Accurate+quantization+of+large+language+models+by+equivalent+and+optimal+shifting+and+scaling 9. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs — Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, et al., 2024 https://scholar.google.com/scholar?q=QuaRot%3A+Outlier-Free+4-Bit+Inference+in+Rotated+LLMs 10. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision — Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao, 2024 https://scholar.google.com/scholar?q=FlashAttention-3%3A+Fast+and+Accurate+Attention+with+Asynchrony+and+Low-precision 11. QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving — Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, Song Han, 2024 https://scholar.google.com/scholar?q=QServe%3A+W4A8KV4+Quantization+and+System+Co-design+for+Efficient+LLM+Serving 12. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention — Huiqiang Jiang et al., 2024 https://arxiv.org/abs/2407.02490 13. SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention — Qianchao Zhu et al., 2024 https://arxiv.org/abs/2406.15486 14. FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference — Xunhao Lai et al., 2025 https://arxiv.org/abs/2502.20766 15. Activation Outliers in Transformer Quantization: Reproduction, Statistical Analysis, and Deployment Tradeoffs — Pranav Kumar Kaliaperumal, 2026 https://arxiv.org/abs/2603.04308 16. BAPS: A Fine-Grained Low-Precision Scheme for Softmax in Attention via Block-Aware Precision reScaling — Zisheng Ye et al., 2026 https://arxiv.org/abs/2602.02071 17. Softpick: No Attention Sink, No Massive Activations with Rectified Softmax — Zayd M. K. Zuhri et al., 2025 https://arxiv.org/abs/2504.20966 18. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3 19. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3 20. AI Post Transformers: NanoFlow and the Future of LLM Serving — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-15-nanoflow-and-the-future-of-llm-serving-7429c9.mp3 Interactive Visualization: SageAttention2 and Fast Exact INT4 Attention

Episode metadata supplied by the publisher feed · Published Jun 18, 2026

Embed this episode

NOW PLAYING

SageAttention2 and Fast Exact INT4 Attention

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 18, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!