Efficient Post-Training Quantization with FP8 episode artwork

EPISODE · Jun 24, 2026

Efficient Post-Training Quantization with FP8

from AI Post Transformers

This episode explores how post-training quantization can convert already-trained models into 8-bit floating point formats for cheaper inference, and why FP8 may outperform the older INT8 approach on modern transformers, LLMs, and diffusion models. It explains the tradeoff between exponent range and mantissa precision across FP8 formats such as E4M3, E5M2, and E3M4, with particular attention to how FP8 handles activation outliers and dynamic range more gracefully than fixed-scale INT8. The discussion centers on a hardware-aware deployment recipe, including which operators can stay quantized, where higher-precision accumulation still matters, and how BatchNorm recalibration helps low-precision inference match full-precision behavior. Listeners get a concrete result: across 75 architectures and more than 200 task cases, the paper reports 92.64% workload coverage for FP8 versus 65.87% for INT8, with E4M3 looking strongest for NLP while E3M4 is slightly better for some vision workloads. Sources: 1. Efficient Post-training Quantization with FP8 Formats — Haihao Shen, Naveen Mellempudi, Xin He, Qun Gao, Chang Wang, Mengni Wang, 2023 http://arxiv.org/abs/2309.14592 2. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference — Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, et al., 2018 https://scholar.google.com/scholar?q=Quantization+and+Training+of+Neural+Networks+for+Efficient+Integer-Arithmetic-Only+Inference 3. A White Paper on Neural Network Quantization — Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, Tijmen Blankevoort, 2021 https://scholar.google.com/scholar?q=A+White+Paper+on+Neural+Network+Quantization 4. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale — Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer, 2022 https://scholar.google.com/scholar?q=LLM.int8%28%29%3A+8-bit+Matrix+Multiplication+for+Transformers+at+Scale 5. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models — Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, Song Han, 2022 https://scholar.google.com/scholar?q=SmoothQuant%3A+Accurate+and+Efficient+Post-Training+Quantization+for+Large+Language+Models 6. 8-bit Numerical Formats for Deep Neural Networks — Badreddine Noune, Philip Jones, Daniel Justus, Dominic Masters, Carlo Luschi, 2022 https://scholar.google.com/scholar?q=8-bit+Numerical+Formats+for+Deep+Neural+Networks 7. FP8 Formats for Deep Learning — Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, et al., 2022 https://scholar.google.com/scholar?q=FP8+Formats+for+Deep+Learning 8. FP8 Quantization: The Power of the Exponent — Andrey Kuzmin, Mart van Baalen, Yuwei Ren, Markus Nagel, Jorn Peters, Tijmen Blankevoort, 2022 https://scholar.google.com/scholar?q=FP8+Quantization%3A+The+Power+of+the+Exponent 9. Efficient Post-training Quantization with FP8 Formats — Haihao Shen, Naveen Mellempudi, Xin He, Qun Gao, Chang Wang, Mengni Wang, 2024 https://scholar.google.com/scholar?q=Efficient+Post-training+Quantization+with+FP8+Formats 10. Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks — Xiaohu Sun et al., 2019 https://scholar.google.com/scholar?q=Hybrid+8-bit+Floating+Point+%28HFP8%29+Training+and+Inference+for+Deep+Neural+Networks 11. Outlier Suppression: Pushing the Limit of Low-bit Transformer Language Models — Xiuying Wei et al., 2022 https://scholar.google.com/scholar?q=Outlier+Suppression%3A+Pushing+the+Limit+of+Low-bit+Transformer+Language+Models 12. Quantizable transformers: Removing outliers by helping attention heads do nothing — Bondarenko et al. (approx.), 2023? https://scholar.google.com/scholar?q=Quantizable+transformers%3A+Removing+outliers+by+helping+attention+heads+do+nothing 13. Understanding and minimising outlier features in transformer training — author list not verified from provided snippet, 2024? https://scholar.google.com/scholar?q=Understanding+and+minimising+outlier+features+in+transformer+training 14. QuanTool: A Benchmarking Framework for Evaluating Post-Training Quantization with Best Practices for Transformer Models — author list not verified from provided snippet, 2024? https://scholar.google.com/scholar?q=QuanTool%3A+A+Benchmarking+Framework+for+Evaluating+Post-Training+Quantization+with+Best+Practices+for+Transformer+Models 15. GO-ViT: Fully Quantizing Vision Transformers by Grouping Outlier Channels — author list not verified from provided snippet, 2024? https://scholar.google.com/scholar?q=GO-ViT%3A+Fully+Quantizing+Vision+Transformers+by+Grouping+Outlier+Channels 16. Understanding int4 quantization for language models: latency speedup, composability, and failure cases — author list not verified from provided snippet, 2024? https://scholar.google.com/scholar?q=Understanding+int4+quantization+for+language+models%3A+latency+speedup%2C+composability%2C+and+failure+cases 17. Kvquant: Towards 10 million context length llm inference with kv cache quantization — author list not verified from provided snippet, 2024? https://scholar.google.com/scholar?q=Kvquant%3A+Towards+10+million+context+length+llm+inference+with+kv+cache+quantization 18. AI Post Transformers: MIOpen and AMD's Open Deep Learning Primitives — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-23-miopen-and-amds-open-deep-learning-primi-052f82.mp3 19. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3 20. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3 Interactive Visualization: Efficient Post-Training Quantization with FP8

Episode metadata supplied by the publisher feed · Published Jun 24, 2026

Embed this episode

NOW PLAYING

Efficient Post-Training Quantization with FP8

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 24, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!