{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T16:31:07Z","timestamp":1755793867691,"version":"3.41.0"},"reference-count":106,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,10,16]],"date-time":"2023-10-16T00:00:00Z","timestamp":1697414400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2023,11,30]]},"abstract":"<jats:p>\n            The unprecedented accuracy of convolutional neural networks (CNNs) across a broad range of AI tasks has led to their widespread deployment in mobile and embedded settings. In a pursuit for high-performance and energy-efficient inference, significant research effort has been invested in the design of field-programmable gate array (FPGA)\u2013based CNN accelerators. In this context, single computation engines constitute a popular design approach that enables the deployment of diverse models without the overhead of fabric reconfiguration. Nevertheless, this flexibility often comes with significantly degraded performance on memory-bound layers and resource underutilisation due to the suboptimal mapping of certain layers on the engine\u2019s fixed configuration. In this work, we investigate the implications in terms of CNN engine design for a class of models that introduce a pre-convolution stage to decompress the weights at runtime. We refer to these approaches as\n            <jats:italic>on-the-fly<\/jats:italic>\n            . This article presents unzipFPGA, a novel CNN inference system that counteracts the limitations of existing CNN engines. The proposed framework comprises a novel CNN hardware architecture that introduces a weights generator module that enables the on-chip on-the-fly generation of weights, alleviating the negative impact of limited bandwidth on memory-bound layers. We further enhance unzipFPGA with an automated hardware-aware methodology that tailors the weights generation mechanism to the target CNN-device pair, leading to an improved accuracy\u2013performance balance. Finally, we introduce an input selective processing element (PE) design that balances the load between PEs in suboptimally mapped layers. Quantitative evaluation shows that the proposed framework yields hardware designs that achieve an average of 2.57\u00d7 performance efficiency gain over highly optimised GPU designs for the same power constraints and up to 3.94\u00d7 higher performance density over a diverse range of state-of-the-art FPGA-based CNN accelerators.\n          <\/jats:p>","DOI":"10.1145\/3611673","type":"journal-article","created":{"date-parts":[[2023,8,4]],"date-time":"2023-08-04T09:50:33Z","timestamp":1691142633000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Mitigating Memory Wall Effects in CNN Engines with On-the-Fly Weights Generation"],"prefix":"10.1145","volume":"28","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5181-6251","authenticated-orcid":false,"given":"Stylianos I.","family":"Venieris","sequence":"first","affiliation":[{"name":"Samsung AI Center, Cambridge, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3747-6523","authenticated-orcid":false,"given":"Javier","family":"Fernandez-Marques","sequence":"additional","affiliation":[{"name":"Samsung AI Center, Cambridge, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2728-8273","authenticated-orcid":false,"given":"Nicholas D.","family":"Lane","sequence":"additional","affiliation":[{"name":"Samsung AI Center, Cambridge &amp; University of Cambridge, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,10,16]]},"reference":[{"key":"e_1_3_1_2_2","volume-title":"Design Automation Conference (DAC)","author":"Abdelfattah Mohamed S.","year":"2020","unstructured":"Mohamed S. Abdelfattah, \u0141ukasz Dudziak, Thomas Chau, Royson Lee, Hyeji Kim, and Nicholas D. Lane. 2020. Best of both worlds: AutoML codesign of a CNN and its hardware accelerator. In Design Automation Conference (DAC)."},{"key":"e_1_3_1_3_2","first-page":"4110","volume-title":"2018 28th International Conference on Field Programmable Logic and Applications (FPL\u201918)","author":"Abdelfattah M. S.","year":"2018","unstructured":"M. S. Abdelfattah, D. Han, A. Bitar, R. DiCecco, S. O\u2019Connell, N. Shanker, J. Chu, I. Prins, J. Fender, A. C. Ling, and G. R. Chiu. 2018. DLA: Compiler and FPGA overlay for neural network inference acceleration. In 2018 28th International Conference on Field Programmable Logic and Applications (FPL\u201918). 4110\u20134117."},{"issue":"9","key":"e_1_3_1_4_2","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1109\/35.714618","article-title":"Wideband DS-CDMA for next-generation mobile communications systems","volume":"36","author":"Adachi Fumiyuki","year":"1998","unstructured":"Fumiyuki Adachi et\u00a0al. 1998. Wideband DS-CDMA for next-generation mobile communications systems. IEEE Communications Magazine 36, 9 (1998), 56\u201369.","journal-title":"IEEE Communications Magazine"},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","first-page":"27\u201328(1)","DOI":"10.1049\/el:19970022","article-title":"Tree-structured generation of orthogonal spreading codes with different lengths for forward link of DS-CDMA mobile radio","volume":"33","author":"Adachi F.","year":"1997","unstructured":"F. Adachi, M. Sawahashi, and K. Okawa. 1997. Tree-structured generation of orthogonal spreading codes with different lengths for forward link of DS-CDMA mobile radio. Electronics Letters 33 (January1997), 27\u201328(1). Issue 1.","journal-title":"Electronics Letters"},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","first-page":"382","DOI":"10.1145\/3123939.3123982","volume-title":"Proceedings of the 50th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201917)","author":"Albericio Jorge","year":"2017","unstructured":"Jorge Albericio, Alberto Delm\u00e1s, Patrick Judd, Sayeh Sharify, Gerard O\u2019Leary, Roman Genov, and Andreas Moshovos. 2017. Bit-pragmatic deep neural network computing. In Proceedings of the 50th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201917). 382\u2013394."},{"key":"e_1_3_1_7_2","first-page":"1","volume-title":"Proceedings of the 43rd International Symposium on Computer Architecture (ISCA\u201916)","author":"Albericio Jorge","year":"2016","unstructured":"Jorge Albericio, Patrick Judd, Tayler Hetherington, Tor Aamodt, Natalie Enright Jerger, and Andreas Moshovos. 2016. Cnvlutin: Ineffectual-neuron-free deep neural network computing. In Proceedings of the 43rd International Symposium on Computer Architecture (ISCA\u201916). 1\u201313."},{"key":"e_1_3_1_8_2","volume-title":"International Conference on Learning Representations (ICLR\u201919)","author":"Alizadeh Milad","year":"2019","unstructured":"Milad Alizadeh, Javier Fern\u00e1ndez-Marqu\u00e9s, Nicholas D. Lane, and Yarin Gal. 2019. A systematic study of binary neural networks\u2019 optimisation. In International Conference on Learning Representations (ICLR\u201919)."},{"key":"e_1_3_1_9_2","volume-title":"EMDL","author":"Almeida Mario","year":"2019","unstructured":"Mario Almeida et\u00a0al. 2019. EmBench: Quantifying performance variations of deep neural networks across modern commodity devices. In EMDL."},{"key":"e_1_3_1_10_2","first-page":"1","volume-title":"2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201916)","author":"Alwani M.","year":"2016","unstructured":"M. Alwani, H. Chen, M. Ferdman, and P. Milder. 2016. Fused-layer CNN accelerators. In 2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201916). 1\u201312."},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","first-page":"229","DOI":"10.1145\/764808.764868","volume-title":"Proceedings of the 13th ACM Great Lakes Symposium on VLSI (GLSVLSI\u201903)","author":"Andreev Boris D.","year":"2003","unstructured":"Boris D. Andreev et\u00a0al. 2003. Orthogonal code generator for 3G wireless transceivers. In Proceedings of the 13th ACM Great Lakes Symposium on VLSI (GLSVLSI\u201903). 229\u2013232. 10.1145\/764808.764868"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1145\/3020078.3021738","volume-title":"Proceedings of the 2017 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201917)","author":"Aydonat Utku","year":"2017","unstructured":"Utku Aydonat, Shane O\u2019Connell, Davor Capalija, Andrew C. Ling, and Gordon R. Chiu. 2017. An OpenCL\u2122 deep learning accelerator on Arria 10. In Proceedings of the 2017 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201917). 55\u201364."},{"key":"e_1_3_1_13_2","article-title":"Polymorphic accelerators for deep neural networks","author":"Azizimazreah Arash","year":"2021","unstructured":"Arash Azizimazreah and Lizhong Chen. 2021. Polymorphic accelerators for deep neural networks. IEEE Trans. Comput. (2021).","journal-title":"IEEE Trans. Comput."},{"key":"e_1_3_1_14_2","first-page":"940","volume-title":"2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA\u201920)","author":"Baek E.","year":"2020","unstructured":"E. Baek, D. Kwon, and J. Kim. 2020. A multi-neural network acceleration architecture. In 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA\u201920). 940\u2013953."},{"key":"e_1_3_1_15_2","first-page":"176","volume-title":"Proceedings of the 14th ACM Conference on Embedded Network Sensor Systems (SenSys\u201916)","author":"Bhattacharya Sourav","year":"2016","unstructured":"Sourav Bhattacharya and Nicholas D. Lane. 2016. Sparsification and separation of deep learning layers for constrained resource inference on wearables. In Proceedings of the 14th ACM Conference on Embedded Network Sensor Systems (SenSys\u201916). ACM, 176\u2013189."},{"key":"e_1_3_1_16_2","volume-title":"Conference on Machine Learning and Systems (MLSys\u201920)","author":"Blalock Davis","year":"2020","unstructured":"Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag. 2020. What is the state of neural network pruning?. In Conference on Machine Learning and Systems (MLSys\u201920)."},{"issue":"3","key":"e_1_3_1_17_2","first-page":"16","article-title":"FINN-R: An end-to-end deep-learning framework for fast exploration of quantized neural networks","volume":"11","author":"Blott Michaela","year":"2018","unstructured":"Michaela Blott, Thomas B. Preu\u00dfer, Nicholas J. Fraser, Giulio Gambardella, Kenneth O\u2019Brien, Yaman Umuroglu, Miriam Leeser, and Kees Vissers. 2018. FINN-R: An end-to-end deep-learning framework for fast exploration of quantized neural networks. ACM Trans. Reconfigurable Technol. Syst. (TRETS) 11, 3, Article 16 (2018), 23 pages.","journal-title":"ACM Trans. Reconfigurable Technol. Syst. (TRETS)"},{"key":"e_1_3_1_18_2","article-title":"Compiling deep learning models for custom hardware accelerators","author":"Chang Andre Xian Ming","year":"2017","unstructured":"Andre Xian Ming Chang, Aliasger Zaidy, Vinayak Gokhale, and Eugenio Culurciello. 2017. Compiling deep learning models for custom hardware accelerators. arXiv preprint arXiv:1708.00117 (2017).","journal-title":"arXiv preprint arXiv:1708.00117"},{"issue":"4","key":"e_1_3_1_19_2","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs","volume":"40","author":"Chen L.","year":"2018","unstructured":"L. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. 2018. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 40, 4 (2018), 834\u2013848.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)"},{"key":"e_1_3_1_20_2","volume-title":"Proceedings of the 2019 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201919)","author":"Chen Yao","year":"2019","unstructured":"Yao Chen, Jiong He, Xiaofan Zhang, Cong Hao, and Deming Chen. 2019. Cloud-DNN: An open framework for mapping DNN models to cloud FPGAs. In Proceedings of the 2019 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201919)."},{"issue":"1","key":"e_1_3_1_21_2","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1109\/JSSC.2016.2616357","article-title":"Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks","volume":"52","author":"Chen Y.","year":"2017","unstructured":"Y. Chen, T. Krishna, J. S. Emer, and V. Sze. 2017. Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks. IEEE Journal of Solid-State Circuits (JSSC) 52, 1 (2017), 127\u2013138.","journal-title":"IEEE Journal of Solid-State Circuits (JSSC)"},{"key":"e_1_3_1_22_2","doi-asserted-by":"crossref","first-page":"188","DOI":"10.1109\/ICFPT47387.2019.00030","volume-title":"2019 International Conference on Field-Programmable Technology (ICFPT\u201919)","author":"Csordas G.","year":"2019","unstructured":"G. Csordas, M. Asiatici, and P. Ienne. 2019. In search of lost bandwidth: Extensive reordering of DRAM accesses on FPGA. In 2019 International Conference on Field-Programmable Technology (ICFPT\u201919). 188\u2013196."},{"key":"e_1_3_1_23_2","first-page":"189","volume-title":"Proceedings of the 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201918)","author":"Deng Chunhua","year":"2018","unstructured":"Chunhua Deng, Siyu Liao, Yi Xie, Keshab K. Parhi, Xuehai Qian, and Bo Yuan. 2018. PermDNN: Efficient compressed DNN architecture with permuted diagonal matrices. In Proceedings of the 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201918). 189\u2013202."},{"key":"e_1_3_1_24_2","doi-asserted-by":"crossref","first-page":"395","DOI":"10.1145\/3123939.3124552","volume-title":"2017 50th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201917)","author":"Ding C.","year":"2017","unstructured":"C. Ding, S. Liao, Y. Wang, Z. Li, N. Liu, Y. Zhuo, C. Wang, X. Qian, Y. Bai, G. Yuan, X. Ma, Y. Zhang, J. Tang, Q. Qiu, X. Lin, and B. Yuan. 2017. CirCNN: Accelerating and compressing deep neural networks using block-circulant weight matrices. In 2017 50th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201917). 395\u2013408."},{"key":"e_1_3_1_25_2","volume-title":"IEEE International Conference on Computer Vision (ICCV\u201919)","author":"Dong Zhen","year":"2019","unstructured":"Zhen Dong, Zhewei Yao, Amir Gholami, Michael Mahoney, and Kurt Keutzer. 2019. HAWQ: Hessian AWare quantization of neural networks with mixed-precision. In IEEE International Conference on Computer Vision (ICCV\u201919)."},{"key":"e_1_3_1_26_2","first-page":"92","volume-title":"2015 ACM\/IEEE 42nd Annual International Symposium on Computer Architecture (ISCA\u201915)","author":"Du Z.","year":"2015","unstructured":"Z. Du, R. Fasthuber, T. Chen, P. Ienne, L. Li, T. Luo, X. Feng, Y. Chen, and O. Temam. 2015. ShiDianNao: Shifting vision processing closer to the sensor. In 2015 ACM\/IEEE 42nd Annual International Symposium on Computer Architecture (ISCA\u201915). 92\u2013104."},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1145\/3212725.3212731","volume-title":"Proceedings of the 2nd International Workshop on Embedded and Mobile Deep Learning (EMDL\u201918)","author":"Fern\u00e1ndez-Marqu\u00e9s Javier","year":"2018","unstructured":"Javier Fern\u00e1ndez-Marqu\u00e9s, Vincent W.-S. Tseng, Sourav Bhattachara, and Nicholas D. Lane. 2018. On-the-fly deterministic binary filters for memory efficient keyword spotting applications on embedded devices. In Proceedings of the 2nd International Workshop on Embedded and Mobile Deep Learning (EMDL\u201918). ACM, 13\u201318."},{"key":"e_1_3_1_28_2","volume-title":"Conference on Machine Learning and Systems (MLSys\u201918)","author":"Fern\u00e1ndez-Marqu\u00e9s Javier","year":"2018","unstructured":"Javier Fern\u00e1ndez-Marqu\u00e9s, W-S Tseng Vincent, Sourav Bhattachara, and Nicholas D. Lane. 2018. BinaryCmd: Keyword spotting with deterministic binary basis. In Conference on Machine Learning and Systems (MLSys\u201918)."},{"key":"e_1_3_1_29_2","volume-title":"Conference on Machine Learning and Systems (MLSys\u201920)","author":"Fernandez-Marques Javier","year":"2020","unstructured":"Javier Fernandez-Marques, Paul N. Whatmough, Andrew Mundy, and Matthew Mattina. 2020. Searching for Winograd-aware quantized networks. In Conference on Machine Learning and Systems (MLSys\u201920)."},{"key":"e_1_3_1_30_2","first-page":"1","volume-title":"2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA\u201918)","author":"Fowers J.","year":"2018","unstructured":"J. Fowers, K. Ovtcharov, M. Papamichael, T. Massengill, M. Liu, D. Lo, S. Alkalay, M. Haselman, L. Adams, M. Ghandi, S. Heil, P. Patel, A. Sapek, G. Weisz, L. Woods, S. Lanka, S. K. Reinhardt, A. M. Caulfield, E. S. Chung, and D. Burger. 2018. A configurable cloud-scale DNN processor for real-time AI. In 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA\u201918). 1\u201314."},{"key":"e_1_3_1_31_2","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201920)","author":"Gale Trevor","year":"2020","unstructured":"Trevor Gale, Matei Zaharia, Cliff Young, and Erich Elsen. 2020. Sparse GPU kernels for deep learning. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201920)."},{"key":"e_1_3_1_32_2","first-page":"1","volume-title":"2017 IEEE International Symposium on Circuits and Systems (ISCAS\u201917)","author":"Gokhale V.","year":"2017","unstructured":"V. Gokhale, A. Zaidy, A. X. M. Chang, and E. Culurciello. 2017. Snowflake: An efficient hardware accelerator for convolutional neural networks. In 2017 IEEE International Symposium on Circuits and Systems (ISCAS\u201917). 1\u20134."},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","first-page":"151","DOI":"10.1145\/3352460.3358291","volume-title":"Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201919)","author":"Gondimalla Ashish","year":"2019","unstructured":"Ashish Gondimalla, Noah Chesnut, Mithuna Thottethodi, and T. N. Vijaykumar. 2019. SparTen: A sparse tensor accelerator for convolutional neural networks. In Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201919). 151\u2013165."},{"key":"e_1_3_1_34_2","doi-asserted-by":"crossref","first-page":"152","DOI":"10.1109\/FCCM.2017.25","volume-title":"2017 IEEE 25th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM\u201917)","author":"Guan Y.","year":"2017","unstructured":"Y. Guan, H. Liang, N. Xu, W. Wang, S. Shi, X. Chen, G. Sun, W. Zhang, and J. Cong. 2017. FP-DNN: An automated framework for mapping deep neural networks onto FPGAs with RTL-HLS hybrid templates. In 2017 IEEE 25th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM\u201917). 152\u2013159."},{"issue":"1","key":"e_1_3_1_35_2","doi-asserted-by":"crossref","first-page":"35","DOI":"10.1109\/TCAD.2017.2705069","article-title":"Angel-Eye: A complete design flow for mapping CNN onto embedded FPGA","volume":"37","author":"Guo K.","year":"2018","unstructured":"K. Guo, L. Sui, J. Qiu, J. Yu, J. Wang, S. Yao, S. Han, Y. Wang, and H. Yang. 2018. Angel-Eye: A complete design flow for mapping CNN onto embedded FPGA. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) 37, 1 (2018), 35\u201347.","journal-title":"IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD)"},{"key":"e_1_3_1_36_2","volume-title":"International Conference on Learning Representations (ICLR\u201917)","author":"Ha David","year":"2017","unstructured":"David Ha, Andrew Dai, and Quoc V. Le. 2017. HyperNetworks. In International Conference on Learning Representations (ICLR\u201917)."},{"key":"e_1_3_1_37_2","first-page":"243","volume-title":"2016 ACM\/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA\u201916)","author":"Han S.","year":"2016","unstructured":"S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally. 2016. EIE: Efficient inference engine on compressed deep neural network. In 2016 ACM\/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA\u201916). 243\u2013254."},{"key":"e_1_3_1_38_2","volume-title":"International Conference on Learning Representations (ICLR\u201915)","author":"Han Song","year":"2015","unstructured":"Song Han, Huizi Mao, and William J. Dally. 2015. Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. In International Conference on Learning Representations (ICLR\u201915)."},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","first-page":"770","DOI":"10.1109\/CVPR.2016.90","volume-title":"2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916)","author":"He K.","year":"2016","unstructured":"K. He, X. Zhang, S. Ren, and J. Sun. 2016. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916). 770\u2013778."},{"key":"e_1_3_1_40_2","volume-title":"International Joint Conference on Artificial Intelligence (IJCAI\u201918)","author":"He Yang","year":"2018","unstructured":"Yang He, Guoliang Kang, Xuanyi Dong, Yanwei Fu, and Yi Yang. 2018. Soft filter pruning for accelerating deep convolutional neural networks. In International Joint Conference on Artificial Intelligence (IJCAI\u201918)."},{"key":"e_1_3_1_41_2","doi-asserted-by":"crossref","first-page":"319","DOI":"10.1145\/3352460.3358275","volume-title":"Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201919)","author":"Hegde Kartik","year":"2019","unstructured":"Kartik Hegde, Hadi Asghari-Moghaddam, Michael Pellauer, Neal Crago, Aamer Jaleel, Edgar Solomonik, Joel Emer, and Christopher W. Fletcher. 2019. ExTensor: An accelerator for sparse tensor algebra. In Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201919). 319\u2013333."},{"key":"e_1_3_1_42_2","first-page":"1","volume-title":"2016 26th International Conference on Field Programmable Logic and Applications (FPL\u201916)","author":"Li Huimin","year":"2016","unstructured":"Huimin Li, Xitian Fan, Li Jiao, Wei Cao, Xuegong Zhou, and Lingli Wang. 2016. A high performance FPGA-based accelerator for large-scale convolutional neural networks. In 2016 26th International Conference on Field Programmable Logic and Applications (FPL\u201916). 1\u20139."},{"key":"e_1_3_1_43_2","article-title":"SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and  \\(\\lt\\)  0.5 MB model size","author":"Iandola Forrest N.","year":"2016","unstructured":"Forrest N. Iandola, Song Han, Matthew W. Moskewicz, Khalid Ashraf, William J. Dally, and Kurt Keutzer. 2016. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and \\(\\lt\\) 0.5 MB model size. arXiv preprint arXiv:1602.07360 (2016).","journal-title":"arXiv preprint arXiv:1602.07360"},{"key":"e_1_3_1_44_2","volume-title":"ICCVW","author":"Ignatov Andrey","year":"2019","unstructured":"Andrey Ignatov et\u00a0al. 2019. AI benchmark: All about deep learning on smartphones in 2019. In ICCVW."},{"key":"e_1_3_1_45_2","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201918)","author":"Jacob Benoit","year":"2018","unstructured":"Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko. 2018. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201918)."},{"key":"e_1_3_1_46_2","volume-title":"ISCA","author":"Jang Jun-Woo","year":"2021","unstructured":"Jun-Woo Jang et\u00a0al. 2021. Sparsity-aware and re-configurable NPU architecture for Samsung flagship mobile SoC. In ISCA."},{"key":"e_1_3_1_47_2","volume-title":"Annual International Symposium on Computer Architecture (ISCA\u201917)","author":"Jouppi Norman P.","year":"2017","unstructured":"Norman P. Jouppi et\u00a0al. 2017. In-datacenter performance analysis of a tensor processing unit. In Annual International Symposium on Computer Architecture (ISCA\u201917)."},{"key":"e_1_3_1_48_2","first-page":"1","volume-title":"2020 IEEE\/ACM International Conference On Computer Aided Design (ICCAD\u201920)","author":"Kao S. C.","year":"2020","unstructured":"S. C. Kao and T. Krishna. 2020. GAMMA: Automating the HW mapping of DNN models on accelerators via genetic algorithm. In 2020 IEEE\/ACM International Conference On Computer Aided Design (ICCAD\u201920). 1\u20139."},{"key":"e_1_3_1_49_2","doi-asserted-by":"crossref","first-page":"483","DOI":"10.1109\/PACRIM.2009.5291322","volume-title":"2009 IEEE Pacific Rim Conference on Communications, Computers and Signal Processing","author":"Kim S.","year":"2009","unstructured":"S. Kim, M. Kim, C. Shin, J. Lee, and Y. Kim. 2009. Efficient implementation of OVSF code generator for UMTS systems. In 2009 IEEE Pacific Rim Conference on Communications, Computers and Signal Processing. 483\u2013486."},{"key":"e_1_3_1_50_2","volume-title":"International Conference on Learning Representations (ICLR\u201915)","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR\u201915)."},{"key":"e_1_3_1_51_2","first-page":"51","volume-title":"2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS\u201919)","author":"Kouris A.","year":"2019","unstructured":"A. Kouris, C. Kyrkou, and C. Bouganis. 2019. Informed region selection for efficient UAV-based object detectors: Altitude-aware vehicle detection with CyCAR dataset. In 2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS\u201919). 51\u201358."},{"key":"e_1_3_1_52_2","first-page":"155","volume-title":"2018 28th International Conference on Field Programmable Logic and Applications (FPL\u201918)","author":"Kouris A.","year":"2018","unstructured":"A. Kouris, S. I. Venieris, and C. Bouganis. 2018. CascadeCNN: Pushing the performance limits of quantisation in convolutional neural networks. In 2018 28th International Conference on Field Programmable Logic and Applications (FPL\u201918). 155\u20131557."},{"key":"e_1_3_1_53_2","doi-asserted-by":"crossref","first-page":"1656","DOI":"10.23919\/DATE48585.2020.9116248","volume-title":"2020 Design, Automation Test in Europe Conference Exhibition (DATE\u201920)","author":"Kouris A.","year":"2020","unstructured":"A. Kouris, S. I. Venieris, and C. Bouganis. 2020. A throughput-latency co-optimised cascade of convolutional neural network classifiers. In 2020 Design, Automation Test in Europe Conference Exhibition (DATE\u201920). 1656\u20131661."},{"key":"e_1_3_1_54_2","unstructured":"Raghuraman Krishnamoorthi. 2018. Quantizing Deep Convolutional Networks for Efficient Inference: A Whitepaper. arxiv:1806.08342 [cs.LG]"},{"issue":"3","key":"e_1_3_1_55_2","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1109\/MM.2020.2985963","article-title":"MAESTRO: A data-centric approach to understand reuse, performance, and hardware cost of DNN mappings","volume":"40","author":"Kwon Hyoukjun","year":"2020","unstructured":"Hyoukjun Kwon, Prasanth Chatarasi, Vivek Sarkar, Tushar Krishna, Michael Pellauer, and Angshuman Parashar. 2020. MAESTRO: A data-centric approach to understand reuse, performance, and hardware cost of DNN mappings. IEEE Micro 40, 3 (2020), 20\u201329.","journal-title":"IEEE Micro"},{"key":"e_1_3_1_56_2","first-page":"461","volume-title":"Proceedings of the 23rd International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201918)","author":"Kwon Hyoukjun","year":"2018","unstructured":"Hyoukjun Kwon, Ananda Samajdar, and Tushar Krishna. 2018. MAERI: Enabling flexible dataflow mapping over DNN accelerators via reconfigurable interconnects. In Proceedings of the 23rd International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201918). 461\u2013475."},{"key":"e_1_3_1_57_2","doi-asserted-by":"crossref","first-page":"28","DOI":"10.1145\/3352460.3358295","volume-title":"Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201919)","author":"Lascorz Alberto Delm\u00e1s","year":"2019","unstructured":"Alberto Delm\u00e1s Lascorz, Sayeh Sharify, Isak Edo, Dylan Malone Stuart, Omar Mohamed Awad, Patrick Judd, Mostafa Mahmoud, Milos Nikolic, Kevin Siu, Zissis Poulos, and Andreas Moshovos. 2019. ShapeShifter: Enabling fine-grain data width adaptation in deep learning. In Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201919). 28\u201341."},{"key":"e_1_3_1_58_2","volume-title":"The 25th Annual International Conference on Mobile Computing and Networking (MobiCom\u201919)","author":"Lee Royson","year":"2019","unstructured":"Royson Lee, Stylianos I. Venieris, Lukasz Dudziak, Sourav Bhattacharya, and Nicholas D. Lane. 2019. MobiSR: Efficient on-device super-resolution through heterogeneous mobile processors. In The 25th Annual International Conference on Mobile Computing and Networking (MobiCom\u201919)."},{"key":"e_1_3_1_59_2","article-title":"Toward full-stack acceleration of deep convolutional neural networks on FPGAs","author":"Liu Shuanglong","year":"2021","unstructured":"Shuanglong Liu, Hongxiang Fan, Martin Ferianc, Xinyu Niu, Huifeng Shi, and Wayne Luk. 2021. Toward full-stack acceleration of deep convolutional neural networks on FPGAs. IEEE Transactions on Neural Networks and Learning Systems (TNNLS) (2021).","journal-title":"IEEE Transactions on Neural Networks and Learning Systems (TNNLS)"},{"key":"e_1_3_1_60_2","first-page":"17","volume-title":"IEEE 27th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM\u201919)","author":"Lu L.","year":"2019","unstructured":"L. Lu, J. Xie, R. Huang, J. Zhang, W. Lin, and Y. Liang. 2019. An efficient hardware accelerator for sparse convolutional neural networks on FPGAs. In IEEE 27th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM\u201919). 17\u201325."},{"key":"e_1_3_1_61_2","doi-asserted-by":"crossref","first-page":"553","DOI":"10.1109\/HPCA.2017.29","volume-title":"2017 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201917)","author":"Lu W.","year":"2017","unstructured":"W. Lu, G. Yan, J. Li, S. Gong, Y. Han, and X. Li. 2017. FlexFlow: A flexible dataflow accelerator architecture for convolutional neural networks. In 2017 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201917). 553\u2013564."},{"key":"e_1_3_1_62_2","volume-title":"IEEE International Conference on Computer Vision (ICCV\u201917)","author":"Luo Jian-Hao","year":"2017","unstructured":"Jian-Hao Luo, Jianxin Wu, and Weiyao Lin. 2017. ThiNet: A filter level pruning method for deep neural network compression. In IEEE International Conference on Computer Vision (ICCV\u201917)."},{"issue":"2","key":"e_1_3_1_63_2","doi-asserted-by":"crossref","first-page":"424","DOI":"10.1109\/TCAD.2018.2884972","article-title":"Automatic compilation of diverse CNNs onto high-performance FPGA accelerators","volume":"39","author":"Ma Y.","year":"2020","unstructured":"Y. Ma, Y. Cao, S. Vrudhula, and J. Seo. 2020. Automatic compilation of diverse CNNs onto high-performance FPGA accelerators. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) 39, 2 (2020), 424\u2013437.","journal-title":"IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD)"},{"key":"e_1_3_1_64_2","first-page":"1","volume-title":"IEEE International Symposium on Circuits and Systems (ISCAS\u201917)","author":"Ma Y.","year":"2017","unstructured":"Y. Ma, M. Kim, Y. Cao, S. Vrudhula, and J. Seo. 2017. End-to-end scalable FPGA accelerator for deep residual networks. In IEEE International Symposium on Circuits and Systems (ISCAS\u201917). 1\u20134."},{"key":"e_1_3_1_65_2","doi-asserted-by":"crossref","first-page":"179","DOI":"10.1109\/ICFPT47387.2019.00029","volume-title":"2019 International Conference on Field-Programmable Technology (ICFPT\u201919)","author":"Manev K.","year":"2019","unstructured":"K. Manev, A. Vaishnav, and D. Koch. 2019. Unexpected diversity: Quantitative memory analysis for Zynq UltraScale+ systems. In 2019 International Conference on Field-Programmable Technology (ICFPT\u201919). 179\u2013187."},{"key":"e_1_3_1_66_2","volume-title":"2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201919)","author":"Molchanov Pavlo","year":"2019","unstructured":"Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Frosio, and Jan Kautz. 2019. Importance estimation for neural network pruning. In 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201919)."},{"key":"e_1_3_1_67_2","volume-title":"Proceedings of the 26th Asia and South Pacific Design Automation Conference (ASP-DAC\u201921)","author":"Montgomerie-Corcoran Alexander","year":"2021","unstructured":"Alexander Montgomerie-Corcoran and Christos Savvas-Bouganis. 2021. DEF: Differential encoding of featuremaps for low power convolutional neural network accelerators. In Proceedings of the 26th Asia and South Pacific Design Automation Conference (ASP-DAC\u201921)."},{"key":"e_1_3_1_68_2","doi-asserted-by":"crossref","first-page":"266","DOI":"10.1145\/3373087.3375302","volume-title":"The 2020 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201920)","author":"Niu Yue","year":"2020","unstructured":"Yue Niu, Rajgopal Kannan, Ajitesh Srivastava, and Viktor Prasanna. 2020. Reuse kernels or activations? A flexible dataflow for low-latency spectral CNN acceleration. In The 2020 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201920) (Seaside, CA, USA). 266\u2013276."},{"key":"e_1_3_1_69_2","first-page":"27","volume-title":"2017 ACM\/IEEE 44th Annual International Symposium on Computer Architecture (ISCA\u201917)","author":"Parashar A.","year":"2017","unstructured":"A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan, B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally. 2017. SCNN: An accelerator for compressed-sparse convolutional neural networks. In 2017 ACM\/IEEE 44th Annual International Symposium on Computer Architecture (ISCA\u201917). 27\u201340."},{"issue":"3","key":"e_1_3_1_70_2","first-page":"15","article-title":"High-efficiency convolutional ternary neural networks with custom adder trees and weight compression","volume":"11","author":"Prost-Boucle Adrien","year":"2018","unstructured":"Adrien Prost-Boucle, Alban Bourge, and Fr\u00e9d\u00e9ric P\u00e9trot. 2018. High-efficiency convolutional ternary neural networks with custom adder trees and weight compression. ACM Trans. Reconfigurable Technol. Syst. (TRETS) 11, 3, Article 15 (2018), 24 pages.","journal-title":"ACM Trans. Reconfigurable Technol. Syst. (TRETS)"},{"key":"e_1_3_1_71_2","doi-asserted-by":"crossref","first-page":"88","DOI":"10.1109\/ICAES.2013.6659367","volume-title":"2013 International Conference on Advanced Electronic Systems (ICAES\u201913)","author":"Purohit G.","year":"2013","unstructured":"G. Purohit, V. K. Chaubey, K. S. Raju, and P. V. Reddy. 2013. FPGA based implementation and testing of OVSF code. In 2013 International Conference on Advanced Electronic Systems (ICAES\u201913). 88\u201392."},{"key":"e_1_3_1_72_2","first-page":"4198","volume-title":"Proceedings of the 35th International Conference on Machine Learning (ICML\u201918)","author":"Qiu Qiang","year":"2018","unstructured":"Qiang Qiu, Xiuyuan Cheng, Robert Calderbank, and Guillermo Sapiro. 2018. DCFNet: Deep neural network with decomposed convolutional filters. In Proceedings of the 35th International Conference on Machine Learning (ICML\u201918). 4198\u20134207."},{"key":"e_1_3_1_73_2","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1109\/ISSOC.2004.1411169","volume-title":"2004 International Symposium on System-on-Chip (ISSOC\u201904)","author":"Rintakoski T.","year":"2004","unstructured":"T. Rintakoski, M. Kuulusa, and J. Nurmi. 2004. Hardware unit for OVSF\/Walsh\/Hadamard code generation [3G mobile communication applications]. In 2004 International Symposium on System-on-Chip (ISSOC\u201904). 143\u2013145."},{"key":"e_1_3_1_74_2","volume-title":"2019 29th International Conference on Field Programmable Logic and Applications (FPL\u201919)","author":"Samajdar Ananda","year":"2019","unstructured":"Ananda Samajdar, Tushar Garg, Tushar Krishna, and Nachiket Kapre. 2019. Scaling the cascades: Interconnect-aware FPGA implementation of machine learning problems. In 2019 29th International Conference on Field Programmable Logic and Applications (FPL\u201919)."},{"key":"e_1_3_1_75_2","doi-asserted-by":"crossref","first-page":"93","DOI":"10.1109\/FCCM.2017.47","volume-title":"2017 IEEE 25th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM\u201917)","author":"Shen Y.","year":"2017","unstructured":"Y. Shen, M. Ferdman, and P. Milder. 2017. Escher: A CNN accelerator with flexible buffering to minimize off-chip transfer. In 2017 IEEE 25th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM\u201917). 93\u2013100."},{"key":"e_1_3_1_76_2","volume-title":"Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA\u201917)","author":"Shen Yongming","year":"2017","unstructured":"Yongming Shen, Michael Ferdman, and Peter Milder. 2017. Maximizing CNN accelerator efficiency through resource partitioning. In Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA\u201917)."},{"key":"e_1_3_1_77_2","first-page":"1","volume-title":"57th ACM\/IEEE Design Automation Conference (DAC\u201920)","author":"Shi R.","year":"2020","unstructured":"R. Shi, Y. Ding, X. Wei, H. Li, H. Liu, H. K. H. So, and C. Ding. 2020. FTDL: A tailored FPGA-overlay for deep learning with high scalability. In 57th ACM\/IEEE Design Automation Conference (DAC\u201920). 1\u20136."},{"key":"e_1_3_1_78_2","volume-title":"2021 31st International Conference on Field Programmable Logic and Applications (FPL\u201921)","author":"Yan Shun","year":"2021","unstructured":"Shun Yan et\u00a0al. 2021. An FPGA-based MobileNet accelerator considering network structure characteristics. In 2021 31st International Conference on Field Programmable Logic and Applications (FPL\u201921)."},{"key":"e_1_3_1_79_2","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1109\/IISWC.2018.8573527","volume-title":"2018 IEEE International Symposium on Workload Characterization (IISWC\u201918)","author":"Siu K.","year":"2018","unstructured":"K. Siu, D. M. Stuart, M. Mahmoud, and A. Moshovos. 2018. Memory requirements for convolutional neural network hardware accelerators. In 2018 IEEE International Symposium on Workload Characterization (IISWC\u201918). 111\u2013121."},{"key":"e_1_3_1_80_2","doi-asserted-by":"crossref","first-page":"689","DOI":"10.1109\/HPCA47549.2020.00062","volume-title":"2020 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201920)","author":"Srivastava N.","year":"2020","unstructured":"N. Srivastava, H. Jin, S. Smith, H. Rong, D. Albonesi, and Z. Zhang. 2020. Tensaurus: A versatile accelerator for mixed sparse-dense tensor computations. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201920). 689\u2013702."},{"key":"e_1_3_1_81_2","first-page":"2739","volume-title":"Proceedings of the 27th International Joint Conference on Artificial Intelligence (IJCAI\u201918)","author":"Tseng Vincent W.-S.","year":"2018","unstructured":"Vincent W.-S. Tseng, Sourav Bhattacharya, Javier Fern\u00e1ndez Marqu\u00e9s, Milad Alizadeh, Catherine Tong, and Nicholas D. Lane. 2018. Deterministic binary filters for convolutional neural networks. In Proceedings of the 27th International Joint Conference on Artificial Intelligence (IJCAI\u201918). 2739\u20132747."},{"issue":"8","key":"e_1_3_1_82_2","doi-asserted-by":"crossref","first-page":"2220","DOI":"10.1109\/TVLSI.2017.2688340","article-title":"Deep convolutional neural network architecture with reconfigurable computation patterns","volume":"25","author":"Tu F.","year":"2017","unstructured":"F. Tu, S. Yin, P. Ouyang, S. Tang, L. Liu, and S. Wei. 2017. Deep convolutional neural network architecture with reconfigurable computation patterns. IEEE Transactions on Very Large Scale Integration (VLSI) Systems (TVLSI) 25, 8 (2017), 2220\u20132233.","journal-title":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems (TVLSI)"},{"key":"e_1_3_1_83_2","first-page":"291","volume-title":"2020 30th International Conference on Field-Programmable Logic and Applications (FPL\u201920)","author":"Umuroglu Y.","year":"2020","unstructured":"Y. Umuroglu, Y. Akhauri, N. J. Fraser, and M. Blott. 2020. LogicNets: Co-designed neural networks and circuits for extreme-throughput applications. In 2020 30th International Conference on Field-Programmable Logic and Applications (FPL\u201920). 291\u2013297. 10.1109\/FPL50879.2020.00055"},{"key":"e_1_3_1_84_2","doi-asserted-by":"crossref","first-page":"65","DOI":"10.1145\/3020078.3021744","volume-title":"Proceedings of the 2017 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201917)","author":"Umuroglu Yaman","year":"2017","unstructured":"Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip Leong, Magnus Jahre, and Kees Vissers. 2017. FINN: A framework for fast, scalable binarized neural network inference. In Proceedings of the 2017 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201917). 65\u201374."},{"key":"e_1_3_1_85_2","first-page":"1","volume-title":"2017 27th International Conference on Field Programmable Logic and Applications (FPL\u201917)","author":"Venieris S. I.","year":"2017","unstructured":"S. I. Venieris and C. Bouganis. 2017. Latency-driven design for FPGA-based convolutional neural networks. In 2017 27th International Conference on Field Programmable Logic and Applications (FPL\u201917). 1\u20138."},{"issue":"2","key":"e_1_3_1_86_2","doi-asserted-by":"crossref","first-page":"326","DOI":"10.1109\/TNNLS.2018.2844093","article-title":"fpgaConvNet: Mapping regular and irregular convolutional neural networks on FPGAs","volume":"30","author":"Venieris S. I.","year":"2019","unstructured":"S. I. Venieris and C. Bouganis. 2019. fpgaConvNet: Mapping regular and irregular convolutional neural networks on FPGAs. IEEE Transactions on Neural Networks and Learning Systems (TNNLS) 30, 2 (2019), 326\u2013342.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems (TNNLS)"},{"key":"e_1_3_1_87_2","first-page":"1","volume-title":"2018 28th International Conference on Field Programmable Logic and Applications (FPL\u201918)","author":"Venieris S. I.","year":"2018","unstructured":"S. I. Venieris and C. S. Bouganis. 2018. f-CNN \\(^\\text{x}\\) : A toolflow for mapping multiple convolutional neural networks on FPGAs. In 2018 28th International Conference on Field Programmable Logic and Applications (FPL\u201918). 1\u20138."},{"key":"e_1_3_1_88_2","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1109\/FCCM51124.2021.00027","volume-title":"2021 IEEE 29th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM\u201921)","author":"Venieris Stylianos I.","year":"2021","unstructured":"Stylianos I. Venieris, Javier Fernandez-Marques, and Nicholas D. Lane. 2021. unzipFPGA: Enhancing FPGA-based CNN engines with on-the-fly weights generation. In 2021 IEEE 29th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM\u201921). 165\u2013175."},{"issue":"3","key":"e_1_3_1_89_2","first-page":"56","article-title":"Toolflows for mapping convolutional neural networks on FPGAs: A survey and future directions","volume":"51","author":"Venieris Stylianos I.","year":"2018","unstructured":"Stylianos I. Venieris, Alexandros Kouris, and Christos-Savvas Bouganis. 2018. Toolflows for mapping convolutional neural networks on FPGAs: A survey and future directions. ACM Comput. Surv. (CSUR) 51, 3, Article 56 (2018), 39 pages.","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"e_1_3_1_90_2","volume-title":"IEEE SMARTCOMP","author":"Venieris Stylianos I.","year":"2021","unstructured":"Stylianos I. Venieris, Ioannis Panopoulos, and Iakovos S. Venieris. 2021. OODIn: An optimised on-device inference framework for heterogeneous mobile devices. In IEEE SMARTCOMP."},{"key":"e_1_3_1_91_2","article-title":"LUTNet: Learning FPGA configurations for highly efficient neural network inference","author":"Wang E.","year":"2020","unstructured":"E. Wang, J. J. Davis, P. Y. K. Cheung, and G. Constantinides. 2020. LUTNet: Learning FPGA configurations for highly efficient neural network inference. IEEE Transactions on Computers (TOC) (2020).","journal-title":"IEEE Transactions on Computers (TOC)"},{"key":"e_1_3_1_92_2","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Wang Tianzhe","year":"2020","unstructured":"Tianzhe Wang, Kuan Wang, Han Cai, Ji Lin, Zhijian Liu, and Song Han. 2020. APQ: Joint search for network architecture, pruning and quantization policy. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)."},{"key":"e_1_3_1_93_2","first-page":"3703","volume-title":"Proceedings of the 34th International Conference on Machine Learning (ICML\u201917)","author":"Wang Yunhe","year":"2017","unstructured":"Yunhe Wang, Chang Xu, Chao Xu, and Dacheng Tao. 2017. Beyond filters: Compact feature map for portable deep model. In Proceedings of the 34th International Conference on Machine Learning (ICML\u201917). 3703\u20133711."},{"key":"e_1_3_1_94_2","volume-title":"Proceedings of the 30th International Conference on Neural Information Processing Systems (NeurIPS\u201916)","author":"Wen Wei","year":"2016","unstructured":"Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2016. Learning structured sparsity in deep neural networks. In Proceedings of the 30th International Conference on Neural Information Processing Systems (NeurIPS\u201916)."},{"key":"e_1_3_1_95_2","volume-title":"HPCA","author":"Wu Carole-Jean","year":"2019","unstructured":"Carole-Jean Wu et\u00a0al. 2019. Machine learning at Facebook: Understanding inference at the edge. In HPCA."},{"key":"e_1_3_1_96_2","article-title":"Adaptive Machine Learning Acceleration","year":"2020","unstructured":"Xilinx. 2020. Adaptive Machine Learning Acceleration. https:\/\/www.xilinx.com\/products\/acceleration-solutions\/xilinx-machine-learning-suite.html[Retrieved: 2023\/10\/14 15:24:31].","journal-title":"https:\/\/www.xilinx.com\/products\/acceleration-solutions\/xilinx-machine-learning-suite.html"},{"key":"e_1_3_1_97_2","article-title":"DNNVM: End-to-end compiler leveraging heterogeneous optimizations on FPGA-based CNN accelerators","author":"Xing Y.","year":"2019","unstructured":"Y. Xing, S. Liang, L. Sui, X. Jia, J. Qiu, X. Liu, Y. Wang, Y. Shan, and Y. Wang. 2019. DNNVM: End-to-end compiler leveraging heterogeneous optimizations on FPGA-based CNN accelerators. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) (2019).","journal-title":"IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD)"},{"key":"e_1_3_1_98_2","volume-title":"Proceedings of the 57th ACM\/EDAC\/IEEE Design Automation Conference (DAC\u201920)","author":"Yang Lei","year":"2020","unstructured":"Lei Yang, Zheyu Yan, Meng Li, Hyoukjun Kwon, Weiwen Jiang, Liangzhen Lai, Yiyu Shi, Tushar Krishna, and Vikas Chandra. 2020. Co-exploration of neural architectures and heterogeneous ASIC accelerator designs targeting multiple tasks. In Proceedings of the 57th ACM\/EDAC\/IEEE Design Automation Conference (DAC\u201920)."},{"key":"e_1_3_1_99_2","volume-title":"International Conference on Learning Representations (ICLR\u201920)","author":"Yang Yingzhen","year":"2020","unstructured":"Yingzhen Yang, Jiahui Yu, Nebojsa Jojic, Jun Huan, and Thomas S. Huang. 2020. FSNet: Compression of deep convolutional neural networks by filter summary. In International Conference on Learning Representations (ICLR\u201920)."},{"key":"e_1_3_1_100_2","first-page":"548","volume-title":"Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA\u201917)","author":"Yu Jiecao","year":"2017","unstructured":"Jiecao Yu, Andrew Lukefahr, David Palframan, Ganesh Dasika, Reetuparna Das, and Scott Mahlke. 2017. Scalpel: Customizing DNN pruning to the underlying hardware parallelism. In Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA\u201917). 548\u2013560."},{"key":"e_1_3_1_101_2","first-page":"122","volume-title":"The 2020 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201920)","author":"Yu Yunxuan","year":"2020","unstructured":"Yunxuan Yu, Tiandong Zhao, Kun Wang, and Lei He. 2020. Light-OPU: An FPGA-based overlay processor for lightweight convolutional neural networks. In The 2020 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201920). 122\u2013132."},{"issue":"7","key":"e_1_3_1_102_2","doi-asserted-by":"crossref","first-page":"1545","DOI":"10.1109\/TVLSI.2020.2995741","article-title":"Uni-OPU: An FPGA-based uniform accelerator for convolutional and transposed convolutional networks","volume":"28","author":"Yu Y.","year":"2020","unstructured":"Y. Yu, T. Zhao, M. Wang, K. Wang, and L. He. 2020. Uni-OPU: An FPGA-based uniform accelerator for convolutional and transposed convolutional networks. IEEE Transactions on Very Large Scale Integration Systems (TVLSI) 28, 7 (2020), 1545\u20131556.","journal-title":"IEEE Transactions on Very Large Scale Integration Systems (TVLSI)"},{"key":"e_1_3_1_103_2","doi-asserted-by":"crossref","first-page":"161","DOI":"10.1145\/2684746.2689060","volume-title":"Proceedings of the 2015 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201915)","author":"Zhang Chen","year":"2015","unstructured":"Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong. 2015. Optimizing FPGA-based accelerator design for deep convolutional neural networks. In Proceedings of the 2015 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201915). 161\u2013170."},{"issue":"11","key":"e_1_3_1_104_2","doi-asserted-by":"crossref","first-page":"2072","DOI":"10.1109\/TCAD.2017.2785257","article-title":"Caffeine: Toward uniformed representation and acceleration for deep convolutional neural networks","volume":"38","author":"Zhang C.","year":"2019","unstructured":"C. Zhang, G. Sun, Z. Fang, P. Zhou, P. Pan, and J. Cong. 2019. Caffeine: Toward uniformed representation and acceleration for deep convolutional neural networks. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) 38, 11 (2019), 2072\u20132085.","journal-title":"IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD)"},{"key":"e_1_3_1_105_2","first-page":"1","volume-title":"2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201916)","author":"Zhang S.","year":"2016","unstructured":"S. Zhang, Z. Du, L. Zhang, H. Lan, S. Liu, L. Li, Q. Guo, T. Chen, and Y. Chen. 2016. Cambricon-X: An accelerator for sparse neural networks. In 2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201916). 1\u201312."},{"key":"e_1_3_1_106_2","doi-asserted-by":"crossref","first-page":"45","DOI":"10.1109\/ICFPT47387.2019.00014","volume-title":"2019 International Conference on Field-Programmable Technology (ICFPT\u201919)","author":"Zhao Y.","year":"2019","unstructured":"Y. Zhao, X. Gao, X. Guo, J. Liu, E. Wang, R. Mullins, P. Y. K. Cheung, G. Constantinides, and C. Xu. 2019. Automatic generation of multi-precision multi-arithmetic CNN accelerators for FPGAs. In 2019 International Conference on Field-Programmable Technology (ICFPT\u201919). 45\u201353."},{"key":"e_1_3_1_107_2","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1109\/MICRO.2018.00011","volume-title":"2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201918)","author":"Zhou X.","year":"2018","unstructured":"X. Zhou, Z. Du, Q. Guo, S. Liu, C. Liu, C. Wang, X. Zhou, L. Li, T. Chen, and Y. Chen. 2018. Cambricon-S: Addressing irregularity in sparse neural networks through a cooperative software\/hardware approach. In 2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201918). 15\u201328."}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3611673","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3611673","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:08Z","timestamp":1750178228000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3611673"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,16]]},"references-count":106,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,11,30]]}},"alternative-id":["10.1145\/3611673"],"URL":"https:\/\/doi.org\/10.1145\/3611673","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"type":"print","value":"1084-4309"},{"type":"electronic","value":"1557-7309"}],"subject":[],"published":{"date-parts":[[2023,10,16]]},"assertion":[{"value":"2023-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-24","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}