FlashFuser and Hopper-Era FFN Kernel Fusion episode artwork

EPISODE · May 15, 2026

FlashFuser and Hopper-Era FFN Kernel Fusion

from AI Post Transformers

This episode explores how the FlashFuser paper uses Hopper GPU inter-core communication to push kernel fusion beyond the usual single-SM memory limits, especially for transformer feed-forward networks and gated FFNs. It explains why this matters now: H100-class GPUs have gained compute far faster than memory bandwidth, making activation spills to HBM an increasingly painful bottleneck for workloads that can consume 40 to 60 percent of inference time. The discussion walks through Hopper’s distributed shared memory model and FlashFuser’s core idea of coordinating reduce, shuffle, and multiply patterns across SM clusters so large intermediate activations can stay on chip longer. Listeners would find it interesting because it connects compiler techniques, GPU architecture, and real transformer inference bottlenecks into a concrete argument about when newer hardware may finally make more aggressive fusion worthwhile. Sources: 1. FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection — Ziyu Huang, Yangjie Zhou, Zihan Liu, Xinhao Luo, Yijia Diao, Minyi Guo, Jidong Zhai, Yu Feng, Chen Zhang, Anbang Wu, Jingwen Leng, 2025 http://arxiv.org/abs/2512.12949 2. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning — Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy, 2018 https://scholar.google.com/scholar?q=TVM%3A+An+Automated+End-to-End+Optimizing+Compiler+for+Deep+Learning 3. FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads — Zhen Zheng, Pengzhan Zhao, Guoping Long, Feiwen Zhu, Kai Zhu, Wenyi Zhao, Lansong Diao, Jun Yang, Wei Lin, 2020 https://scholar.google.com/scholar?q=FusionStitching%3A+Boosting+Memory+Intensive+Computations+for+Deep+Learning+Workloads 4. Operator Fusion in XLA: Analysis and Evaluation — Daniel Snider, Ruofan Liang, 2023 https://scholar.google.com/scholar?q=Operator+Fusion+in+XLA%3A+Analysis+and+Evaluation 5. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Re, 2022 https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness 6. Benchmarking and Dissecting the Nvidia Hopper GPU Architecture — Weile Luo, Ruibo Fan, Zeyu Li, Dayou Du, Qiang Wang, Xiaowen Chu, 2024 https://scholar.google.com/scholar?q=Benchmarking+and+Dissecting+the+Nvidia+Hopper+GPU+Architecture 7. A Case Study in CUDA Kernel Fusion: Implementing FlashAttention-2 on NVIDIA Hopper Architecture using the CUTLASS Library — Ganesh Bikshandi, Jay Shah, 2023 https://scholar.google.com/scholar?q=A+Case+Study+in+CUDA+Kernel+Fusion%3A+Implementing+FlashAttention-2+on+NVIDIA+Hopper+Architecture+using+the+CUTLASS+Library 8. Scaling Deep Learning Computation over the Inter-Core Connected Intelligence Processor with T10 — Yiqi Liu, Yuqi Xue, Yu Cheng, Lingxiao Ma, Ziming Miao, Jilong Xue, Jian Huang, 2024 https://scholar.google.com/scholar?q=Scaling+Deep+Learning+Computation+over+the+Inter-Core+Connected+Intelligence+Processor+with+T10 9. FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection — Ziyu Huang, Yangjie Zhou, Zihan Liu, Xinhao Luo, Yijia Diao, Minyi Guo, Jidong Zhai, Yu Feng, Chen Zhang, Anbang Wu, Jingwen Leng, 2025 https://scholar.google.com/scholar?q=FlashFuser%3A+Expanding+the+Scale+of+Kernel+Fusion+for+Compute-Intensive+Operators+via+Inter-Core+Connection 10. Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators Fusion — Size Zheng, Siyuan Chen, Peidi Song, Renze Chen, Xiuhong Li, Shengen Yan, Dahua Lin, Jingwen Leng, Yun Liang, 2023 https://scholar.google.com/scholar?q=Chimera%3A+An+Analytical+Optimizing+Framework+for+Effective+Compute-intensive+Operators+Fusion 11. BOLT: Bridging the Gap between Auto-tuners and Hardware-native Performance — Jiarong Xing, Leyuan Wang, Shang Zhang, Jack Chen, Ang Chen, Yibo Zhu, 2022 https://scholar.google.com/scholar?q=BOLT%3A+Bridging+the+Gap+between+Auto-tuners+and+Hardware-native+Performance 12. MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators — Zheng Zhang, Donglin Yang, Xiaobo Zhou, Dazhao Cheng, 2024 https://scholar.google.com/scholar?q=MCFuser%3A+High-Performance+and+Rapid+Fusion+of+Memory-Bound+Compute-Intensive+Operators 13. Deep Kernel Fusion for Transformers — Zixi Zhang, Zhiwen Mo, Yiren Zhao, Robert Mullins, 2026 https://scholar.google.com/scholar?q=Deep+Kernel+Fusion+for+Transformers 14. Benchmarking thread block cluster — approximate; unclear from snippet, 2023-2026 https://scholar.google.com/scholar?q=Benchmarking+thread+block+cluster 15. ClusterSim: modeling thread block clusters in Hopper GPUs — approximate; unclear from snippet, 2023-2026 https://scholar.google.com/scholar?q=ClusterSim%3A+modeling+thread+block+clusters+in+Hopper+GPUs 16. Analysing and Reducing Costs of Deep Learning Compiler Auto-tuning — approximate; unclear from snippet, 2023-2026 https://scholar.google.com/scholar?q=Analysing+and+Reducing+Costs+of+Deep+Learning+Compiler+Auto-tuning 17. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3 18. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3 19. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3 20. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3 21. AI Post Transformers: Caffeine: A Unified FPGA for CNNs — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-06-caffeine-a-unified-fpga-for-cnns-e8acbe.mp3 22. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3 Interactive Visualization: FlashFuser and Hopper-Era FFN Kernel Fusion

Episode metadata supplied by the publisher feed · Published May 15, 2026

Embed this episode

NOW PLAYING

FlashFuser and Hopper-Era FFN Kernel Fusion

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on May 15, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!