{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:15:59Z","timestamp":1750220159605,"version":"3.41.0"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2022,12,22]],"date-time":"2022-12-22T00:00:00Z","timestamp":1671667200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National funds through Funda\u00e7\u00e3o para a Ci\u00eancia e a Tecnologia","award":["UIDB\/50021\/2020 and PTDC\/EEI-HAC\/31819\/2017"],"award-info":[{"award-number":["UIDB\/50021\/2020 and PTDC\/EEI-HAC\/31819\/2017"]}]},{"name":"IPL\/2021\/smartSPACE_ISEL through Instituto Polit\u00e9cnico de Lisboa"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2023,3,31]]},"abstract":"<jats:p>\n            Designing hardware accelerators to run the inference of\n            <jats:bold>convolutional neural networks (CNN)<\/jats:bold>\n            is under intensive research. Several different architectures have been proposed along with hardware-oriented optimizations of the neural network models. One of the most used optimizations is quantization since it reduces the memory requirements to store weights and layer maps, the memory bandwidth requirements and the hardware complexity. As a consequence, the inference throughput has improved and the computing cost has been reduced, allowing inference to be executed on embedded devices. In this work, we propose highly efficient dot-product arithmetic units for ternary and non-ternary convolutional neural networks on FPGA. The non-ternary dot-product unit uses a fused multiply-add that avoids expensive adder trees, while the ternary dot-product unit uses a dual product unit followed by an optimized conditional adder tree structure. In both cases, designs with and without embedded DSP are considered. The solution is configurable and can be adapted to the available number of resources of the FPGA to achieve the best efficiency. A CNN architecture was developed and characterized using the proposed dot product units. The results show a performance improvement of 1.8 \u00d7 with a 2\u00d7 more area efficiency for low bit-width quantizations when compared to previous works running large CNNs in FPGA.\n          <\/jats:p>","DOI":"10.1145\/3546182","type":"journal-article","created":{"date-parts":[[2022,7,4]],"date-time":"2022-07-04T09:29:19Z","timestamp":1656926959000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Efficient Design of Low Bitwidth Convolutional Neural Networks on FPGA with Optimized Dot Product Units"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8556-4507","authenticated-orcid":false,"given":"M\u00e1rio","family":"V\u00e9stias","sequence":"first","affiliation":[{"name":"INESC-ID, ISEL, Instituto Polit\u00e9cnico de Lisboa, Lisbon, Portugal"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7060-4745","authenticated-orcid":false,"given":"Rui P.","family":"Duarte","sequence":"additional","affiliation":[{"name":"INESC-ID, Lisbon, Portugal"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7525-7546","authenticated-orcid":false,"given":"Jos\u00e9 T.","family":"de Sousa","sequence":"additional","affiliation":[{"name":"INESC-ID, Instituto Superior T\u00e9cnico, Universidade de Lisboa, Lisbon, Portugal"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3621-8322","authenticated-orcid":false,"given":"Hor\u00e1cio","family":"Neto","sequence":"additional","affiliation":[{"name":"INESC-ID, Instituto Superior T\u00e9cnico, Universidade de Lisboa, Lisbon, Portugal"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,12,22]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2022. Neural Network Accelerator Comparison. (2022). https:\/\/nicsefc.ee.tsinghua.edu.cn\/projects\/neural-network-accelerator\/."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1117\/12.2304711"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00027"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.3390\/electronics9122200"},{"key":"e_1_3_1_6_2","article-title":"A survey of quantization methods for efficient neural network inference","volume":"2103","author":"Gholami Amir","year":"2021","unstructured":"Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. 2021. A survey of quantization methods for efficient neural network inference. CoRR abs\/2103.13630 (2021). arXiv:2103.13630https:\/\/arxiv.org\/abs\/2103.13630.","journal-title":"CoRR"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289185"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2018.00016"},{"key":"e_1_3_1_9_2","article-title":"Deep compression: Compressing deep neural network with pruning, trained quantization and Huffman coding","volume":"1510","author":"Han Song","year":"2015","unstructured":"Song Han, Huizi Mao, and William J. Dally. 2015. Deep compression: Compressing deep neural network with pruning, trained quantization and Huffman coding. CoRR abs\/1510.00149 (2015).","journal-title":"CoRR"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2919527"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.23919\/FPL.2017.8056820"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_1_15_2","volume-title":"MBMV","author":"Kumm Martin","year":"2014","unstructured":"Martin Kumm and Peter Zipf. 2014. Efficient high speed compression trees on Xilinx FPGAs. In MBMV."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ARITH.2018.8464695"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3154839"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2019.2913958"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079758"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2018.00018"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.micpro.2022.104441"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2019.2905242"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASAP.2017.7995253"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3270764"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/2847263.2847265"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TR.2018.2878387"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2890150"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080221"},{"key":"e_1_3_1_30_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations","author":"Simonyan K.","year":"2015","unstructured":"K. Simonyan and A. Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In Proceedings of the 3rd International Conference on Learning Representations."},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3490422.3502364"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr.2016.308"},{"key":"e_1_3_1_33_2","article-title":"FINN: A framework for fast, scalable binarized neural network inference","volume":"1612","author":"Umuroglu Yaman","year":"2016","unstructured":"Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip Heng Wai Leong, Magnus Jahre, and Kees A. Vissers. 2016. FINN: A framework for fast, scalable binarized neural network inference. CoRR abs\/1612.07119 (2016). arxiv:1612.07119","journal-title":"CoRR"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11265-020-01606-2"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.23919\/FPL.2017.8056863"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.micpro.2020.103136"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.3000444"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2019.00062"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.3390\/computers5040020"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2018.00035"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474597"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474597"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.3390\/electronics10091025"},{"key":"e_1_3_1_44_2","article-title":"Convolutional neural network with INT4 optimization on Xilinx devices","author":"Xilinx Inc","year":"2020","unstructured":"Inc Xilinx. 2020. Convolutional neural network with INT4 optimization on Xilinx devices. White Paper 521 (2020).","journal-title":"White Paper 521"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289602.3293902"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373087.3375311"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289602.3293964"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.3390\/electronics9091344"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3020078.3021741"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2018.2876865"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3546182","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3546182","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:23Z","timestamp":1750186823000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3546182"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,22]]},"references-count":49,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2023,3,31]]}},"alternative-id":["10.1145\/3546182"],"URL":"https:\/\/doi.org\/10.1145\/3546182","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"type":"print","value":"1936-7406"},{"type":"electronic","value":"1936-7414"}],"subject":[],"published":{"date-parts":[[2022,12,22]]},"assertion":[{"value":"2022-02-04","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-06-19","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-12-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}