Approaching Shannon Bound: Lossless LLM Weight Compression episode artwork

EPISODE · Aug 22, 2026

Approaching Shannon Bound: Lossless LLM Weight Compression

from AI Post Transformers

This episode explores "Approaching Shannon Bound with Lossless LLM Weight Compression," which argues that model weights stored in formats like bf16 carry far less real information than their bit-width implies—entropy measurements across six models and seven numeric formats show gaps of several bits per weight that can be recovered without any change to the underlying values. The discussion covers why memory capacity and bandwidth, not raw compute, are the real bottleneck in GPU inference, and why generic compressors like gzip fail on IEEE-754 floating point data. The hosts dig into Asymmetric Numeral Systems (ANS) as the key engineering breakthrough, since it decodes fast enough and in a tile-parallel enough fashion to run inside a live GPU kernel without becoming a new bottleneck itself. Listeners interested in the intersection of information theory and practical LLM serving will find the walkthrough of how lossless compression differs fundamentally from quantization methods like int4 or AWQ particularly compelling. Sources: 1. Approaching Shannon Bound with Lossless LLM Weight Compression — Hongshi Tan, Yao Chen, Gustavo Alonso, Weng-Fai Wong, Bingsheng He, 2026 http://arxiv.org/abs/2606.15789 2. 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float — Tianyi Zhang, Yang Sui, Shaochen (Henry) Zhong, et al., 2025 https://scholar.google.com/scholar?q=70%25+Size%2C+100%25+Accuracy%3A+Lossless+LLM+Compression+for+Efficient+GPU+Inference+via+Dynamic-Length+Float 3. NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks — Yongchang Hao, Yanshuai Cao, Lili Mou, 2024 https://scholar.google.com/scholar?q=NeuZip%3A+Memory-Efficient+Training+and+Inference+with+Dynamic+Compression+of+Neural+Networks 4. ZipNN: Lossless Compression for AI Models — Moshik Hershcovitch, Andrew Wood, Leshem Choshen, et al. (IBM Research), 2024 https://scholar.google.com/scholar?q=ZipNN%3A+Lossless+Compression+for+AI+Models 5. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding — Song Han, Huizi Mao, William J. Dally, 2016 https://scholar.google.com/scholar?q=Deep+Compression%3A+Compressing+Deep+Neural+Networks+with+Pruning%2C+Trained+Quantization+and+Huffman+Coding 6. Asymmetric Numeral Systems: Entropy Coding Combining Speed of Huffman Coding with Compression Rate of Arithmetic Coding — Jarek Duda, 2013 https://scholar.google.com/scholar?q=Asymmetric+Numeral+Systems%3A+Entropy+Coding+Combining+Speed+of+Huffman+Coding+with+Compression+Rate+of+Arithmetic+Coding 7. The Use of Asymmetric Numeral Systems as an Accurate Replacement for Huffman Coding — Jarek Duda, Khalid Tahboub, Neeraj J. Gadgil, Edward J. Delp, 2015 https://scholar.google.com/scholar?q=The+Use+of+Asymmetric+Numeral+Systems+as+an+Accurate+Replacement+for+Huffman+Coding 8. Variational Image Compression with a Scale Hyperprior — Johannes Ballé, David Minnen, Saurabh Singh, Sung Jin Hwang, Nick Johnston, 2018 https://scholar.google.com/scholar?q=Variational+Image+Compression+with+a+Scale+Hyperprior 9. Zstandard Compression and the application/zstd Media Type (RFC 8878) — Yann Collet, Murray Kucherawy (eds.), 2020 https://scholar.google.com/scholar?q=Zstandard+Compression+and+the+application%2Fzstd+Media+Type+%28RFC+8878%29 10. GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLMs — Hao Kang, Qingru Zhang, Souvik Kundu, Geonhwa Jeong, Zaoxing Liu, Tushar Krishna, Tuo Zhao, 2024 https://scholar.google.com/scholar?q=GEAR%3A+An+Efficient+KV+Cache+Compression+Recipe+for+Near-Lossless+Generative+Inference+of+LLMs 11. S-LoRA: Serving Thousands of Concurrent LoRA Adapters — Ying Sheng, Shiyi Cao, Dacheng Li, Coleman Hooper, Nicholas Lee, Shuo Yang, Christopher Chou, Banghua Zhu, Lianmin Zheng, Kurt Keutzer, Joseph E. Gonzalez, Ion Stoica, 2023 https://scholar.google.com/scholar?q=S-LoRA%3A+Serving+Thousands+of+Concurrent+LoRA+Adapters 12. QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving — Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, Song Han, 2024 https://scholar.google.com/scholar?q=QServe%3A+W4A8KV4+Quantization+and+System+Co-design+for+Efficient+LLM+Serving 13. Efficient LLM Inference: Bandwidth, Compute, Synchronization, and Capacity are All You Need — M. Davies, N. Crago, K. Sankaralingam, C. Kozyrakis, 2025 https://scholar.google.com/scholar?q=Efficient+LLM+Inference%3A+Bandwidth%2C+Compute%2C+Synchronization%2C+and+Capacity+are+All+You+Need Interactive Visualization: Approaching Shannon Bound: Lossless LLM Weight Compression

Episode metadata supplied by the publisher feed · Published Aug 22, 2026

Embed this episode

NOW PLAYING

Approaching Shannon Bound: Lossless LLM Weight Compression

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on August 22, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!