When Finer Microscaling Hurts LLM Quantization episode artwork

EPISODE · Jul 1, 2026

When Finer Microscaling Hurts LLM Quantization

from AI Post Transformers

This episode explores the paper Is Finer Better? The Limits of Microscaling Formats in Large Language Models and examines why shrinking microscaling block sizes can unexpectedly make low-bit LLM quantization worse instead of better. It walks through how microscaling pairs FP4 weights or activations with shared local FP8 scales, contrasts that setup with coarser quantization schemes, and places the work in the broader move from BF16 and FP8 toward cheaper, more hardware-friendly inference. The central argument is that smaller blocks do reduce element quantization error, but once the shared scale is itself quantized into a limited format like FP8 UE4M3, scale error can dominate and degrade perplexity. Listeners would find it interesting because the discussion turns a seemingly obvious engineering intuition on its head and shows that the real bottleneck in low-bit inference may be the precision of the scaling rule, not just the precision of the values being scaled. Sources: 1. Is Finer Better? The Limits of Microscaling Formats in Large Language Models — Andrea Fasoli, Monodeep Kar, Chi-Chun Liu, Swagath Venkataramani, Viji Srinivasan, Leland Chang, Naigang Wang, 2026 http://arxiv.org/abs/2601.19026 2. FP8 Formats for Deep Learning — Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, and others, 2022 https://arxiv.org/abs/2209.05433 3. With Shared Microexponents, A Little Shifting Goes a Long Way — Bita Rouhani, Ritchie Zhao, Venmugil Elango, Rasoul Shafipour, Mathew Hall, Maral Mesmakhosroshahi, Ankit More, and others, 2023 https://arxiv.org/abs/2302.08007 4. Microscaling Data Formats for Deep Learning — Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall, Alireza Khodamoradi, Summer Deng, Dhruv Choudhary, and others, 2023 https://arxiv.org/abs/2310.10537 5. Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization — Vage Egiazarian, Roberto L. Castro, Denis Kuznedelev, Andrei Panferov, Eldar Kurtic, Shubhra Pandit, Alexandre Marques, Mark Kurtz, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh, 2025 https://arxiv.org/abs/2509.23202 6. AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference — Janghwan Lee et al., 2024 https://scholar.google.com/scholar?q=AMXFP4%3A+Taming+Activation+Outliers+with+Asymmetric+Microscaling+Floating-Point+for+4-bit+LLM+Inference 7. Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models — Yun-Chen Lo, Gu-Yeon Wei, David Brooks, 2024 https://scholar.google.com/scholar?q=Nanoscaling+Floating-Point+%28NxFP%29%3A+NanoMantissa%2C+Adaptive+Microexponents%2C+and+Code+Recycling+for+Direct-Cast+Compression+of+Large+Language+Models 8. Elucidating the Design Space of FP4 Training — Robert Hu, Carlo Luschi, Paul Balanca, 2025 https://scholar.google.com/scholar?q=Elucidating+the+Design+Space+of+FP4+Training 9. Finer is Better (with the Right Scaling) — Clemens Schaefer, Gil Tabak, 2026 https://scholar.google.com/scholar?q=Finer+is+Better+%28with+the+Right+Scaling%29 10. Adaptive Block-Scaled Data Types — Jack Cook et al., 2026 https://arxiv.org/abs/2603.28765 11. Diagnosing FP4 inference: a layer-wise and block-wise sensitivity analysis of NVFP4 and MXFP4 — Musa Cim et al., 2026 https://arxiv.org/abs/2603.08747 12. Pretraining large language models with MXFP4 — Musa Cim et al., 2026 https://arxiv.org/abs/2605.09825 13. Dissecting Outlier Dynamics in LLM NVFP4 Pretraining — Peijie Dong et al., 2026 https://arxiv.org/abs/2602.02047 14. DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization — Haokun Lin et al., 2026 https://arxiv.org/abs/2604.17789 15. AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation — Seonggon Kim et al., 2026 https://arxiv.org/abs/2604.02525 16. AI Post Transformers: Nemotron 3 Ultra for Long-Horizon Agents — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-17-nemotron-3-ultra-for-long-horizon-agents-32e4a5.mp3 17. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3 18. AI Post Transformers: FlashAttention-4 Conquers Asymmetric GPU Hardware Scaling — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-06-flashattention-4-conquers-asymmetric-gpu-78839b.mp3

Episode metadata supplied by the publisher feed · Published Jul 1, 2026

Embed this episode

NOW PLAYING

When Finer Microscaling Hurts LLM Quantization

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on July 1, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!