Adaptive Block-Scaled Data Types for FP4 Training episode artwork

EPISODE · Jul 30, 2026

Adaptive Block-Scaled Data Types for FP4 Training

from AI Post Transformers

This episode explores Adaptive Block-Scaled Data Types, a new IF4 format from MIT and NVIDIA researchers for representing numbers in just 4 bits during LLM training and inference. The discussion traces the lineage from FP8 training (used at scale by DeepSeek-V3) through existing 4-bit formats like NVFP4 and MXFP4, and the predecessor "4/6" method, explaining why each prior approach traded away either representable values or dynamic range to control quantization error. The key innovation covered is how IF4 quantizes each 16-value group both as FP4 and as scaled INT4, keeping whichever has lower error, and encodes that choice for free in an otherwise-unused sign bit of the scale factor. Listeners get a clear picture of why 4-bit precision matters primarily for raw matmul speed on hardware like NVIDIA's B200, not just memory savings, and why this fix is notable for spending "dead weight" bits rather than sacrificing precision or range like earlier techniques. Sources: 1. Adaptive Block-Scaled Data Types for FP4 Training https://arxiv.org/pdf/2603.28765 2. Mixed Precision Training — Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, Hao Wu (NVIDIA/Baidu), 2017/2018 (ICLR 2018) https://scholar.google.com/scholar?q=Mixed+Precision+Training 3. FP8 Formats for Deep Learning — Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, John Kamalu, Naveen Mellempudi, Stuart Oberman, Mohammad Shoeybi, Michael Siu, Hao Wu (NVIDIA, Arm, Intel, Qualcomm), 2022 https://scholar.google.com/scholar?q=FP8+Formats+for+Deep+Learning 4. Microscaling Data Formats for Deep Learning — Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall, Alireza Khodamoradi, Summer Deng, Dhruv Choudhary, Marius Cornea, Eric Dellinger, Kristof Denolf, Stosic Dusan, Venmugil Elango, Maximilian Golub, Alexander Heinecke, Phil James-Roxby, Dharmesh Jani, Gaurav Kolhe, Martin Langhammer, Ada Li, Levi Melnick, Maral Mesmakhosroshahi, Andres Rodriguez, Michael Schulte, Rasoul Shafipour, Lei Shao, Michael Siu, Pradeep Dubey, Paulius Micikevicius (Microsoft, AMD, Arm, Intel, Meta, NVIDIA, Qualcomm — OCP consortium), 2023 https://scholar.google.com/scholar?q=Microscaling+Data+Formats+for+Deep+Learning 5. DeepSeek-V3 Technical Report — DeepSeek-AI (large author list, DeepSeek), 2024 https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report 6. Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling — Jack Cook, Junxian Guo, Guangxuan Xiao, Yujun Lin, Song Han, 2026 https://scholar.google.com/scholar?q=Four+Over+Six%3A+More+Accurate+NVFP4+Quantization+with+Adaptive+Block+Scaling 7. Pretraining Large Language Models with NVFP4 — NVIDIA (large author list), 2026 https://scholar.google.com/scholar?q=Pretraining+Large+Language+Models+with+NVFP4 8. Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation — Andrei Panferov, Erik Schultheis, Soroush Tabesh, Dan Alistarh, 2026 https://scholar.google.com/scholar?q=Quartet+II%3A+Accurate+LLM+Pre-Training+in+NVFP4+by+Improved+Unbiased+Gradient+Estimation 9. INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats — Mengzhao Chen, Meng Wu, Hui Jin, Zhihang Yuan, et al., 2025 https://scholar.google.com/scholar?q=INT+v.s.+FP%3A+A+Comprehensive+Study+of+Fine-Grained+Low-bit+Quantization+Formats 10. Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization — Vage Egiazarian, Roberto L. Castro, Denis Kuznedelev, et al., 2026 https://scholar.google.com/scholar?q=Bridging+the+Gap+Between+Promise+and+Performance+for+Microscaling+FP4+Quantization 11. WUSH: Near-Optimal Adaptive Transforms for LLM Quantization — Jiale Chen, Vage Egiazarian, Roberto L. Castro, Torsten Hoefler, Dan Alistarh, 2026 https://scholar.google.com/scholar?q=WUSH%3A+Near-Optimal+Adaptive+Transforms+for+LLM+Quantization 12. Scaling Laws for Precision — Tanishq Kumar, Zachary Ankner, Benjamin F. Spector, et al., 2024 https://scholar.google.com/scholar?q=Scaling+Laws+for+Precision Interactive Visualization: Adaptive Block-Scaled Data Types for FP4 Training

Episode metadata supplied by the publisher feed · Published Jul 30, 2026

Embed this episode

NOW PLAYING

Adaptive Block-Scaled Data Types for FP4 Training

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on July 30, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!