EPISODE · May 7, 2026
Automating DNN Compilation for FPGA Accelerators
from AI Post Transformers
This episode explores FP-DNN, a 2017 framework that aims to compile TensorFlow-era neural networks onto FPGAs automatically, reducing the need for hand-designed accelerators for each model. It explains how the system maps convolutional layers, fully connected layers, and parts of LSTM computation into a shared matrix-multiplication core, while combining hand-tuned RTL for performance-critical components with HLS-generated logic for orchestration and layer-specific handling. The discussion highlights why this hybrid design matters for performance-per-watt, latency, and communication efficiency, especially as deeper CNNs and recurrent models were pushing hardware limits. Listeners would find it interesting for its clear look at an early attempt to turn FPGA deployment from an expert-only craft into a more reusable compiler-driven workflow, while also showing where the paper’s claims about broad model coverage may be too optimistic. Sources: 1. Automating DNN Compilation for FPGA Accelerators https://ceca.pku.edu.cn/media/lw/e3d0e0cd92452e0504b148220d442b9a.pdf 2. A Survey of FPGA-based Neural Network Inference Accelerators — Kaiyuan Guo, Shulin Zeng, Jincheng Yu, Yu Wang, Huazhong Yang, 2019 https://scholar.google.com/scholar?q=A+Survey+of+FPGA-based+Neural+Network+Inference+Accelerators 3. DeepBurning: Automatic Generation of FPGA-based Learning Accelerators for the Neural Network Family — Ying Wang, Jie Xu, Yudeng Sun, Baohua Cao, Chunyuan Xu, Yibo Kong, Chundao Han, Xuan Wang, 2016 https://scholar.google.com/scholar?q=DeepBurning%3A+Automatic+Generation+of+FPGA-based+Learning+Accelerators+for+the+Neural+Network+Family 4. FP-DNN: An Automated Framework for Mapping Deep Neural Networks onto FPGAs with RTL-HLS Hybrid Templates — Yijin Guan, Hao Liang, Ningyi Xu, Wenqiang Wang, Shaoshuai Shi, Xi Chen, Guangyu Sun, Wei Zhang, Jason Cong, 2017 https://scholar.google.com/scholar?q=FP-DNN%3A+An+Automated+Framework+for+Mapping+Deep+Neural+Networks+onto+FPGAs+with+RTL-HLS+Hybrid+Templates 5. DNNBuilder: An Automated Tool for Building High-Performance DNN Hardware Accelerators for FPGAs — Xiaofan Zhang, Junsong Wang, Chao Zhu, Yonghua Lin, Jinjun Xiong, Wen-Mei Hwu, Deming Chen, 2018 https://scholar.google.com/scholar?q=DNNBuilder%3A+An+Automated+Tool+for+Building+High-Performance+DNN+Hardware+Accelerators+for+FPGAs 6. From High-Level Deep Neural Models to FPGAs — Hardik Sharma, Jongse Park, Emmanuel Amaro, Bradley Thwaites, Priyanka Kotha, Anmol Gupta, Joon Kyung Kim, Asit Mishra, and Hsien-Hsin S. Lee, 2016 https://scholar.google.com/scholar?q=From+High-Level+Deep+Neural+Models+to+FPGAs 7. Caffeine: Towards Uniformed Representation and Acceleration for Deep Convolutional Neural Networks — Chen Zhang, Zhenman Fang, Peipei Zhou, Peichen Pan, and Jason Cong, 2016 https://scholar.google.com/scholar?q=Caffeine%3A+Towards+Uniformed+Representation+and+Acceleration+for+Deep+Convolutional+Neural+Networks 8. Throughput-Optimized OpenCL-Based FPGA Accelerator for Large-Scale Convolutional Neural Networks — Naveen Suda, Vikas Chandra, Ganesh Dasika, Abinash Mohanty, Yufei Ma, Sarita Vrudhula, Jae-sun Seo, and Yu Cao, 2016 https://scholar.google.com/scholar?q=Throughput-Optimized+OpenCL-Based+FPGA+Accelerator+for+Large-Scale+Convolutional+Neural+Networks 9. Going Deeper with Embedded FPGA Platform for Convolutional Neural Network — Jiantao Qiu, Jie Wang, Song Yao, Kai Guo, Boxun Li, Erjin Zhou, Jincheng Yu, Tianqi Tang, Ningyi Xu, and Song Wang, 2016 https://scholar.google.com/scholar?q=Going+Deeper+with+Embedded+FPGA+Platform+for+Convolutional+Neural+Network 10. Optimizing FPGA-Based Accelerator Design for Deep Convolutional Neural Networks — Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong, 2015 https://scholar.google.com/scholar?q=Optimizing+FPGA-Based+Accelerator+Design+for+Deep+Convolutional+Neural+Networks 11. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding — Song Han, Huizi Mao, and William J. Dally, 2015 https://scholar.google.com/scholar?q=Deep+Compression%3A+Compressing+Deep+Neural+Networks+with+Pruning%2C+Trained+Quantization+and+Huffman+Coding 12. BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach — Zhen Zheng et al., 2023 https://scholar.google.com/scholar?q=BladeDISC%3A+Optimizing+Dynamic+Shape+Machine+Learning+Workloads+via+Compiler+Approach 13. TSCompiler: efficient compilation framework for dynamic-shape models — Xiang Luo, Chen Zhang, Chenbo Geng, Yanzhi Yi, Jiahui Hu, Renwei Zhang, Zhen Zhang, Gianpietro Consolaro, Fan Yang, Tun Lu, Ning Gu, Li Shang, 2024 https://scholar.google.com/scholar?q=TSCompiler%3A+efficient+compilation+framework+for+dynamic-shape+models 14. TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture — Jiajun Wu, Mo Song, Jingmin Zhao, Yizhao Gao, Jia Li, Hayden Kwok-Hay So, 2024 https://scholar.google.com/scholar?q=TATAA%3A+Programmable+Mixed-Precision+Transformer+Acceleration+with+a+Transformable+Arithmetic+Architecture 15. FPGA Acceleration With Hessian-Based Comprehensive Intra-Layer Mixed-Precision Quantization for Transformer Models — Woohong Byun, Jongseok Woo, Saibal Mukhopadhyay, 2025 https://scholar.google.com/scholar?q=FPGA+Acceleration+With+Hessian-Based+Comprehensive+Intra-Layer+Mixed-Precision+Quantization+for+Transformer+Models 16. Understand and Accelerate Memory Processing Pipeline for Disaggregated LLM Inference — Zifan He, Rui Ma, Yizhou Sun, Jason Cong, 2026 https://scholar.google.com/scholar?q=Understand+and+Accelerate+Memory+Processing+Pipeline+for+Disaggregated+LLM+Inference 17. AI Post Transformers: FPGA Neural Network Accelerators for Space — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-26-fpga-neural-network-accelerators-for-spa-3087ae.mp3 18. AI Post Transformers: Continuous Batching for LLM Inference: Throughput and Latency Gains — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/continuous-batching-for-llm-inference-throughput-and-latency-gains/ 19. AI Post Transformers: Advancements in Efficient KV Cache Quantization and Management — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/advancements-in-efficient-kv-cache-quantization-and-management/ Interactive Visualization: Automating DNN Compilation for FPGA Accelerators
Embed this episode
NOW PLAYING
Automating DNN Compilation for FPGA Accelerators
No transcript for this episode yet
Similar Episodes
No similar episodes found.