{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,17]],"date-time":"2025-12-17T08:31:53Z","timestamp":1765960313610,"version":"3.41.0"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2024,2,15]],"date-time":"2024-02-15T00:00:00Z","timestamp":1707955200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2024,6,30]]},"abstract":"<jats:p>Resource-efficient Convolutional Neural Networks (CNNs) are gaining more attention. These CNNs have relatively low computational and memory requirements. A common denominator among such CNNs is having more heterogeneity than traditional CNNs. This heterogeneity is present at two levels: intra-layer type and inter-layer type. Generic accelerators do not capture these levels of heterogeneity, which harms their efficiency. Consequently, researchers have proposed model-specific accelerators with dedicated engines. When designing an accelerator with dedicated engines, one option is to dedicate one engine per CNN layer. We refer to accelerators designed with this approach as single-engine single-layer (SESL). This approach enables optimizing each engine for its specific layer. However, such accelerators are resource-demanding and unscalable. Another option is to design a minimal number of dedicated engines such that each engine handles all layers of one type. We refer to these accelerators as single-engine multiple-layer (SEML). SEML accelerators capture the inter-layer-type but not the intra-layer-type heterogeneity.<\/jats:p>\n          <jats:p>\n            We propose \u00a0the Fixed Budget Hybrid CNN Accelerator (FiBHA), a hybrid accelerator composed of an SESL part and an SEML part, each processing a subset of CNN layers. FiBHA\u00a0captures more heterogeneity than SEML while being more resource-aware and scalable than SESL. Moreover, we propose a novel module, Fused Inverted Residual Bottleneck (FIRB), a fine-grained and memory-light SESL architecture building block. The proposed architecture is implemented and evaluated using high-level synthesis (HLS) on different Field Programmable Gate Arrays representing various resource budgets. Our evaluation shows that FiBHA\u00a0improves the throughput by up to 4\n            <jats:italic>x<\/jats:italic>\n            and 2.5\n            <jats:italic>x<\/jats:italic>\n            compared to state-of-the-art SESL and SEML accelerators, respectively. Moreover, FiBHA\u00a0reduces memory and energy consumption compared to an SEML accelerator. The evaluation also shows that FIRB\u00a0reduces the required memory by up to 54%, and energy requirements by up to 35% compared to traditional pipelining.\n          <\/jats:p>","DOI":"10.1145\/3639823","type":"journal-article","created":{"date-parts":[[2024,1,8]],"date-time":"2024-01-08T21:50:15Z","timestamp":1704750615000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["An Efficient Hybrid Deep Learning Accelerator for Compact and Heterogeneous CNNs"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3955-2836","authenticated-orcid":false,"given":"Fareed","family":"Qararyah","sequence":"first","affiliation":[{"name":"Chalmers University of Technology, Sweden"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0477-4540","authenticated-orcid":false,"given":"Muhammad Waqar","family":"Azhar","sequence":"additional","affiliation":[{"name":"Chalmers University of Technology, Sweden"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2776-9253","authenticated-orcid":false,"given":"Pedro","family":"Trancoso","sequence":"additional","affiliation":[{"name":"Chalmers University of Technology, Sweden"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,2,15]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3453688.3461485"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSII.2018.2865896"},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","first-page":"162","DOI":"10.1109\/IPDPSW.2018.00032","volume-title":"2018 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW\u201918)","author":"Baskin Chaim","year":"2018","unstructured":"Chaim Baskin, Natan Liss, Evgenii Zheltonozhskii, Alex M. Bronstein, and Avi Mendelson. 2018. Streaming architecture for large-scale quantized neural networks on an FPGA-based dataflow platform. In 2018 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW\u201918). IEEE, 162\u2013169."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3242897"},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","first-page":"159","DOI":"10.1109\/PACT52795.2021.00019","volume-title":"2021 30th International Conference on Parallel Architectures and Compilation Techniques (PACT\u201921)","author":"Boroumand Amirali","year":"2021","unstructured":"Amirali Boroumand, Saugata Ghose, Berkin Akin, Ravi Narayanaswami, Geraldo F. Oliveira, Xiaoyu Ma, Eric Shiu, and Onur Mutlu. 2021. Google neural network models for edge devices: Analyzing and mitigating machine learning inference bottlenecks. In 2021 30th International Conference on Parallel Architectures and Compilation Techniques (PACT\u201921). IEEE, 159\u2013172."},{"key":"e_1_3_1_7_2","article-title":"Proxylessnas: Direct neural architecture search on target task and hardware","author":"Cai Han","year":"2018","unstructured":"Han Cai, Ligeng Zhu, and Song Han. 2018. Proxylessnas: Direct neural architecture search on target task and hardware. arXiv preprint arXiv:1812.00332.","journal-title":"arXiv preprint arXiv:1812.00332"},{"key":"e_1_3_1_8_2","unstructured":"Tianqi Chen Thierry Moreau Ziheng Jiang Lianmin Zheng Eddie Yan Haichen Shen Meghan Cowan Leyuan Wang Yuwei Hu Luis Ceze Carlos Guestrin and Arvind Krishnamurthy. 2018. TVM: An automated End-to-End optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201918). 578\u2013594."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001177"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.195"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","first-page":"50","DOI":"10.1109\/FCCM51124.2021.00014","volume-title":"2021 IEEE 29th Annual International Symposium on Field-programmable Custom Computing Machines (FCCM\u201921)","author":"Dong Zhen","year":"2021","unstructured":"Zhen Dong, Yizhao Gao, Qijing Huang, John Wawrzynek, Hayden K. H. So, and Kurt Keutzer. 2021. Hao: Hardware-aware neural architecture optimization for efficient inference. In 2021 IEEE 29th Annual International Symposium on Field-programmable Custom Computing Machines (FCCM\u201921). IEEE, 50\u201359."},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","first-page":"117","DOI":"10.1109\/FCCM.2017.22","volume-title":"2017 IEEE 25th Annual International Symposium on Field-programmable Custom Computing Machines (FCCM\u201917)","author":"Dumpala Naveen Kumar","year":"2017","unstructured":"Naveen Kumar Dumpala, Shivukumar B. Patil, Daniel Holcomb, and Russell Tessier. 2017. Energy efficient loop unrolling for low-cost FPGAs. In 2017 IEEE 25th Annual International Symposium on Field-programmable Custom Computing Machines (FCCM\u201917). IEEE, 117\u2013120."},{"key":"e_1_3_1_14_2","first-page":"1","volume-title":"2021 International Conference on Field-programmable Technology (ICFPT\u201921)","author":"Gao Jingbo","year":"2021","unstructured":"Jingbo Gao, Yu Qian, Yihan Hu, Xitian Fan, Wai-Shing Luk, Wei Cao, and Lingli Wang. 2021. LETA: A lightweight exchangeable-track accelerator for Efficientnet based on FPGA. In 2021 International Conference on Field-programmable Technology (ICFPT\u201921). IEEE, 1\u20139."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037702"},{"key":"e_1_3_1_16_2","first-page":"807","volume-title":"Proceedings of the 24th International Conference on Architectural Support for Programming Languages and Operating Systems","author":"Gao Mingyu","year":"2019","unstructured":"Mingyu Gao, Xuan Yang, Jing Pu, Mark Horowitz, and Christos Kozyrakis. 2019. Tangram: Optimized coarse-grained dataflow for scalable NN accelerators. In Proceedings of the 24th International Conference on Architectural Support for Programming Languages and Operating Systems. 807\u2013820."},{"key":"e_1_3_1_17_2","first-page":"1","volume-title":"2018 14th IEEE International Conference on Solid-state and Integrated Circuit Technology (ICSICT\u201918)","author":"Hao Cong","year":"2018","unstructured":"Cong Hao and Deming Chen. 2018. Deep neural network model and FPGA accelerator co-design: Opportunities and challenges. In 2018 14th IEEE International Conference on Solid-state and Integrated Circuit Technology (ICSICT\u201918). IEEE, 1\u20134."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_19_2","article-title":"MobileNets: Efficient convolutional neural networks for mobile vision applications","author":"Howard Andrew G.","year":"2017","unstructured":"Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861.","journal-title":"arXiv preprint arXiv:1704.04861"},{"key":"e_1_3_1_20_2","article-title":"SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and < 0.5 MB model size","author":"Iandola Forrest N.","year":"2016","unstructured":"Forrest N. Iandola, Song Han, Matthew W. Moskewicz, Khalid Ashraf, William J. Dally, and Kurt Keutzer. 2016. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and < 0.5 MB model size. arXiv preprint arXiv:1602.07360.","journal-title":"arXiv preprint arXiv:1602.07360"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00286"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_1_23_2","first-page":"0","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops","author":"Kopuklu Okan","year":"2019","unstructured":"Okan Kopuklu, Neslihan Kose, Ahmet Gunduz, and Gerhard Rigoll. 2019. Resource efficient 3D convolutional neural networks. In Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops. 0\u20130."},{"key":"e_1_3_1_24_2","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. Imagenet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems 25 (2012).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3065386"},{"key":"e_1_3_1_26_2","first-page":"1","volume-title":"2014 IEEE High Performance Extreme Computing Conference (HPEC\u201914)","author":"Kuppannagari Sanmukh R.","year":"2014","unstructured":"Sanmukh R. Kuppannagari, Ren Chen, Andrea Sanny, Shreyas G. Singapura, Geoffrey Phi C. Tran, Shijie Zhou, Yusong Hu, Stephen P. Crago, and Viktor K. Prasanna. 2014. Energy performance of FPGAs on perfect suite kernels. In 2014 IEEE High Performance Extreme Computing Conference (HPEC\u201914). IEEE, 1\u20136."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCAS.2010.5537907"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/PROC.1987.13876"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.02.071"},{"key":"e_1_3_1_31_2","first-page":"1","volume-title":"2016 26th International Conference on Field Programmable Logic and Applications (FPL\u201916)","author":"Li Huimin","year":"2016","unstructured":"Huimin Li, Xitian Fan, Li Jiao, Wei Cao, Xuegong Zhou, and Lingli Wang. 2016. A high performance FPGA-based accelerator for large-scale convolutional neural networks. In 2016 26th International Conference on Field Programmable Logic and Applications (FPL\u201916). IEEE, 1\u20139."},{"key":"e_1_3_1_32_2","doi-asserted-by":"crossref","first-page":"1392","DOI":"10.1109\/IAEAC47372.2019.8997842","volume-title":"2019 IEEE 4th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC\u201919)","volume":"1","author":"Liao Jiawen","year":"2019","unstructured":"Jiawen Liao, Liangwei Cai, Yuan Xu, and Minya He. 2019. Design of accelerator for MobileNet convolutional neural network based on FPGA. In 2019 IEEE 4th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC\u201919), Vol. 1. IEEE, 1392\u20131396."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.3390\/electronics8030281"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3020078.3021736"},{"key":"e_1_3_1_35_2","first-page":"1","volume-title":"2016 26th International Conference on Field Programmable Logic and Applications (FPL\u201916)","author":"Ma Yufei","year":"2016","unstructured":"Yufei Ma, Naveen Suda, Yu Cao, Jae-sun Seo, and Sarma Vrudhula. 2016. Scalable and modularized RTL compilation of convolutional neural networks onto FPGA. In 2016 26th International Conference on Field Programmable Logic and Applications (FPL\u201916). IEEE, 1\u20138."},{"key":"e_1_3_1_36_2","first-page":"1","volume-title":"Proceedings of the 1st on Reproducible Quality-efficient Systems Tournament on Co-designing Pareto-efficient Deep Learning","author":"Moreau Thierry","year":"2018","unstructured":"Thierry Moreau, Tianqi Chen, and Luis Ceze. 2018. Leveraging the VTA-TVM hardware-software stack for FPGA acceleration of 8-bit resnet-18 inference. In Proceedings of the 1st on Reproducible Quality-efficient Systems Tournament on Co-designing Pareto-efficient Deep Learning. 1."},{"key":"e_1_3_1_37_2","unstructured":"Nvidia. 2022. NVIDIA RTX A4000. https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/gtcs21\/rtx-a4000\/nvidia-rtx-a4000-datasheet.pdfdatasheet."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.3333552"},{"issue":"8","key":"e_1_3_1_39_2","doi-asserted-by":"crossref","first-page":"2637","DOI":"10.3390\/s21082637","article-title":"A heterogeneous hardware accelerator for image classification in embedded systems","volume":"21","author":"P\u00e9rez Ignacio","year":"2021","unstructured":"Ignacio P\u00e9rez and Miguel Figueroa. 2021. A heterogeneous hardware accelerator for image classification in embedded systems. Sensors 21, 8 (2021), 2637.","journal-title":"Sensors"},{"key":"e_1_3_1_40_2","first-page":"180","volume-title":"2022 IEEE 34th International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD\u201922)","author":"Qararyah Fareed","year":"2022","unstructured":"Fareed Qararyah, Muhammad Waqar Azhar, and Pedro Trancoso. 2022. FiBHA: Fixed budget hybrid CNN accelerator. In 2022 IEEE 34th International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD\u201922). IEEE, 180\u2013190."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2021.102792"},{"key":"e_1_3_1_42_2","first-page":"7","volume-title":"2022 3rd International Conference on Electronics, Communications and Information Technology (CECIT\u201922)","author":"Qin Wenqiang","year":"2022","unstructured":"Wenqiang Qin, Zhongcheng Wu, Jun Zhang, and Fang Li. 2022. Design and optimization of MobileNet neural network acceleration system based on FPGA. In 2022 3rd International Conference on Electronics, Communications and Information Technology (CECIT\u201922). IEEE, 7\u201312."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/2847263.2847265"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.5555\/2971808.2972132"},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1109\/ICFPT47387.2019.00044","volume-title":"2019 International Conference on Field-Programmable Technology (ICFPT\u201919)","author":"Sada Youki","year":"2019","unstructured":"Youki Sada, Masayuki Shimoda, Akira Jinguji, and Hiroki Nakahara. 2019. A dataflow pipelining architecture for tile segmentation with a sparse MobileNet on an FPGA. In 2019 International Conference on Field-Programmable Technology (ICFPT\u201919). IEEE, 267\u2013270."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_1_47_2","first-page":"160","volume-title":"2020 IEEE International Conference on Parallel & Distributed Processing with Applications, Big Data & Cloud Computing, Sustainable Computing & Communications, Social Computing & Networking (ISPA\/BDCloud\/SocialCom\/SustainCom\u201920)","author":"Shaydyuk Nazariy K.","year":"2020","unstructured":"Nazariy K. Shaydyuk and Eugene B. John. 2020. FPGA implementation of MobileNeTv2 CNN model using semi-streaming architecture for low power inference applications. In 2020 IEEE International Conference on Parallel & Distributed Processing with Applications, Big Data & Cloud Computing, Sustainable Computing & Communications, Social Computing & Networking (ISPA\/BDCloud\/SocialCom\/SustainCom\u201920). IEEE, 160\u2013167."},{"key":"e_1_3_1_48_2","article-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556.","journal-title":"arXiv preprint arXiv:1409.1556"},{"key":"e_1_3_1_49_2","first-page":"16","volume-title":"International Symposium on Applied Reconfigurable Computing","author":"Su Jiang","year":"2018","unstructured":"Jiang Su, Julian Faraone, Junyi Liu, Yiren Zhao, David B. Thomas, Philip H. W. Leong, and Peter Y. K. Cheung. 2018. Redundancy-reduced MobileNet acceleration on reconfigurable logic for Imagenet classification. In International Symposium on Applied Reconfigurable Computing. Springer, 16\u201328."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2017.2761740"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00293"},{"key":"e_1_3_1_53_2","first-page":"6105","volume-title":"International Conference on Machine Learning","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc Le. 2019. Efficientnet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning. PMLR, 6105\u20136114."},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3020078.3021744"},{"key":"e_1_3_1_55_2","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1109\/FCCM.2016.22","volume-title":"2016 IEEE 24th Annual International Symposium on Field-programmable Custom Computing Machines (FCCM\u201916)","author":"Venieris Stylianos I.","year":"2016","unstructured":"Stylianos I. Venieris and Christos-Savvas Bouganis. 2016. fpgaConvNet: A framework for mapping convolutional neural networks on FPGAs. In 2016 IEEE 24th Annual International Symposium on Field-programmable Custom Computing Machines (FCCM\u201916). IEEE, 40\u201347."},{"key":"e_1_3_1_56_2","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1145\/3079856.3080244","volume-title":"Proceedings of the 44th Annual International Symposium on Computer Architecture","author":"Venkataramani Swagath","year":"2017","unstructured":"Swagath Venkataramani, Ashish Ranjan, Subarno Banerjee, Dipankar Das, Sasikanth Avancha, Ashok Jagannathan, Ajaya Durg, Dheemanth Nagaraj, Bharat Kaul, Pradeep Dubey, et\u00a0al. 2017. Scaledeep: A scalable compute architecture for learning and evaluating deep networks. In Proceedings of the 44th Annual International Symposium on Computer Architecture. 13\u201326."},{"key":"e_1_3_1_57_2","first-page":"012086","volume-title":"Journal of Physics: Conference Series","volume":"1883","author":"Wei Liu","year":"2021","unstructured":"Liu Wei and Lv Peng. 2021. An efficient OpenCL-based FPGA accelerator for MobileNet. Journal of Physics: Conference Series 1883 (2021), 012086."},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3061639.3062207"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2019.00030"},{"key":"e_1_3_1_60_2","first-page":"1","volume-title":"2021 China Semiconductor Technology International Conference (CSTIC\u201921)","author":"Xiao Jiale","year":"2021","unstructured":"Jiale Xiao, Yonghao Chen, and Tao Su. 2021. A MobileNet accelerator with high processing-element-efficiency on FPGA. In 2021 China Semiconductor Technology International Conference (CSTIC\u201921). IEEE, 1\u20133."},{"key":"e_1_3_1_61_2","first-page":"134","volume-title":"2022 IEEE 2nd International Conference on Data Science and Computer Application (ICDSCA\u201922)","author":"Xie Xiaofei","year":"2022","unstructured":"Xiaofei Xie, Guodong Zhao, Wei Wei, and Wei Huang. 2022. MobileNetV2 accelerator for power and speed balanced embedded applications. In 2022 IEEE 2nd International Conference on Data Science and Computer Application (ICDSCA\u201922). IEEE, 134\u2013139."},{"key":"e_1_3_1_62_2","first-page":"657","volume-title":"2021 Design, Automation & Test in Europe Conference & Exhibition (DATE\u201921)","author":"Xu Rui","year":"2021","unstructured":"Rui Xu, Sheng Ma, Yaohua Wang, and Yang Guo. 2021. HESA: Heterogeneous systolic array architecture for compact CNNs hardware accelerators. In 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE\u201921). IEEE, 657\u2013662."},{"key":"e_1_3_1_63_2","first-page":"17","volume-title":"2021 31st International Conference on Field-programmable Logic and Applications (FPL\u201921)","author":"Yan Shun","year":"2021","unstructured":"Shun Yan, Zhengyan Liu, Yun Wang, Chenglong Zeng, Qiang Liu, Bowen Cheng, and Ray C. C. Cheung. 2021. An FPGA-based MobileNet accelerator considering network structure characteristics. In 2021 31st International Conference on Field-programmable Logic and Applications (FPL\u201921). IEEE, 17\u201323."},{"key":"e_1_3_1_64_2","article-title":"Wide residual networks","author":"Zagoruyko Sergey","year":"2016","unstructured":"Sergey Zagoruyko and Nikos Komodakis. 2016. Wide residual networks. arXiv preprint arXiv:1605.07146.","journal-title":"arXiv preprint arXiv:1605.07146"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00716"},{"key":"e_1_3_1_66_2","first-page":"022045","volume-title":"Journal of Physics: Conference Series","volume":"1486","author":"Zhao Tong","year":"2020","unstructured":"Tong Zhao, Lufeng Qiao, Qinghua Chen, Qingsong Zhang, and Na Li. 2020. A hardware accelerator based on neural network for object detection. Journal of Physics: Conference Series 1486 (2020), 022045."}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639823","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3639823","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:03:50Z","timestamp":1750291430000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639823"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,15]]},"references-count":65,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,6,30]]}},"alternative-id":["10.1145\/3639823"],"URL":"https:\/\/doi.org\/10.1145\/3639823","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2024,2,15]]},"assertion":[{"value":"2023-02-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-12-21","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}