EPISODE · May 6, 2026
Automating CNN Mapping on Embedded FPGAs
from AI Post Transformers
This episode explores how fpgaConvNet turns CNN inference into a Synchronous Dataflow problem so embedded FPGA accelerators can be designed with analyzable schedules, buffers, and resource tradeoffs instead of ad hoc hardware tuning. It explains why CNN deployment on robots, drones, and cars is constrained as much by data movement, latency, and power as by raw arithmetic, and why FPGAs can outperform embedded GPUs when the hardware is tailored carefully to a model’s structure. The discussion highlights the paper’s central claim that formalizing the mapping problem enables automated design-space exploration across very different CNN topologies, rather than just optimizing a single benchmark. It also examines where that approach is strong and where listeners should be skeptical, including whether the reported GPU speedups are fair and how well the clean SDF abstraction survives real hardware implementation. Sources: 1. fpgaConvNet: A Toolflow for Mapping Diverse Convolutional Neural Networks on Embedded FPGAs — Stylianos I. Venieris, Christos-Savvas Bouganis, 2017 http://arxiv.org/abs/1711.08740 2. Static Scheduling of Synchronous Data Flow Programs for Digital Signal Processing — Edward A. Lee, David G. Messerschmitt, 1987 https://scholar.google.com/scholar?q=Static+Scheduling+of+Synchronous+Data+Flow+Programs+for+Digital+Signal+Processing 3. Synchronous Data Flow — Edward A. Lee, David G. Messerschmitt, 1987 https://scholar.google.com/scholar?q=Synchronous+Data+Flow 4. Scenario-aware dataflow: modeling, analysis and implementation of dynamic applications — Sander Stuijk, Marc C. W. Geilen, Bart D. Theelen, Twan Basten, 2011 https://scholar.google.com/scholar?q=Scenario-aware+dataflow%3A+modeling%2C+analysis+and+implementation+of+dynamic+applications 5. fpgaConvNet: A Toolflow for Mapping Diverse Convolutional Neural Networks on Embedded FPGAs — Stylianos I. Venieris, Christos-Savvas Bouganis, 2017 https://scholar.google.com/scholar?q=fpgaConvNet%3A+A+Toolflow+for+Mapping+Diverse+Convolutional+Neural+Networks+on+Embedded+FPGAs 6. DNNWeaver: From High-Level Deep Network Models to FPGA Acceleration — Hyoukjun Sharma, Jongse Park, Emmanuel Amaro, Bradley Thwaites, Praneeth Kotha, Anmol Gupta, Joon Kyung Kim, Asit Mishra, and Hsien-Hsin S. Lee, 2016 https://scholar.google.com/scholar?q=DNNWeaver%3A+From+High-Level+Deep+Network+Models+to+FPGA+Acceleration 7. Caffeine: Towards Uniformed Representation and Acceleration for Deep Convolutional Neural Networks — Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong, 2016 https://scholar.google.com/scholar?q=Caffeine%3A+Towards+Uniformed+Representation+and+Acceleration+for+Deep+Convolutional+Neural+Networks 8. FINN: A Framework for Fast, Scalable Binarized Neural Network Inference — Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip Leong, Magnus Jahre, and Kees Vissers, 2017 https://scholar.google.com/scholar?q=FINN%3A+A+Framework+for+Fast%2C+Scalable+Binarized+Neural+Network+Inference 9. Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks — Yu Wang, Jiajun Xu, Yanzhi Wang, and Huazhong Yang, 2015 https://scholar.google.com/scholar?q=Optimizing+FPGA-based+Accelerator+Design+for+Deep+Convolutional+Neural+Networks 10. Efficient Processing of Deep Neural Networks: A Tutorial and Survey — Vivienne Sze, Yu-Hsin Chen, Tien-Ju Yang, and Joel S. Emer, 2017 https://scholar.google.com/scholar?q=Efficient+Processing+of+Deep+Neural+Networks%3A+A+Tutorial+and+Survey 11. ViTA: A Vision Transformer Inference Accelerator for Edge Applications — Shashank Nag, Gourav Datta, Souvik Kundu, Nitin Chandrachoodan, Peter A. Beerel, 2023 https://scholar.google.com/scholar?q=ViTA%3A+A+Vision+Transformer+Inference+Accelerator+for+Edge+Applications 12. ME-ViT: A Single-Load Memory-Efficient FPGA Accelerator for Vision Transformers — Kyle Marino, Pengmiao Zhang, Viktor K. Prasanna, 2024 https://scholar.google.com/scholar?q=ME-ViT%3A+A+Single-Load+Memory-Efficient+FPGA+Accelerator+for+Vision+Transformers 13. DRViT: A Dynamic Redundancy-Aware Vision Transformer Accelerator via Algorithm and Architecture Co-Design on FPGA — Xiangfeng Sun, Yuanting Zhang, Qinyu Wang, Xiaofeng Zou, et al., 2025 https://scholar.google.com/scholar?q=DRViT%3A+A+Dynamic+Redundancy-Aware+Vision+Transformer+Accelerator+via+Algorithm+and+Architecture+Co-Design+on+FPGA 14. Realisation of Early-Exit Dynamic Neural Networks on Reconfigurable Hardware — Anastasios Dimitriou, Lei Xun, Jonathon Hare, Geoff V. Merrett, 2024 https://scholar.google.com/scholar?q=Realisation+of+Early-Exit+Dynamic+Neural+Networks+on+Reconfigurable+Hardware 15. Compute-In-Memory on FPGAs for Deep Learning: A Review — Aman Arora, 2025 https://scholar.google.com/scholar?q=Compute-In-Memory+on+FPGAs+for+Deep+Learning%3A+A+Review 16. A Heterogeneous System With Computing in Memory Processing Elements to Accelerate CNN Inference — Jinkai Wang, Youxiang Chen, Zekun Wang, Zhengkun Gu, et al., 2025 https://scholar.google.com/scholar?q=A+Heterogeneous+System+With+Computing+in+Memory+Processing+Elements+to+Accelerate+CNN+Inference 17. AI Post Transformers: FPGA Neural Network Accelerators for Space — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-26-fpga-neural-network-accelerators-for-spa-3087ae.mp3 Interactive Visualization: Automating CNN Mapping on Embedded FPGAs
Embed this episode
NOW PLAYING
Automating CNN Mapping on Embedded FPGAs
No transcript for this episode yet
Similar Episodes
No similar episodes found.