{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T17:36:58Z","timestamp":1755797818259,"version":"3.44.0"},"reference-count":22,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2022YFB4501400"],"award-info":[{"award-number":["2022YFB4501400"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>\n            Unstructured sparse pruning significantly reduces the computational and parametric complexities of deep neural network models. Nevertheless, the highly irregular nature of sparse models limits their performance and efficiency on traditional computing platforms, thereby prompting the development of specialized hardware solutions. To improve computational efficiency, we introduce the Sparse Dataflow Fusion Accelerator (SPDFA), a specialized architecture meticulously designed for sparse deep neural networks. Firstly, we present a non-blocking data distribution-computing engine that integrates inner product and column product. This engine boosts computational efficiency by decomposing matrix multiplication and convolution into rectangular matrix-vector multiplications. Secondly, we implement a computation array to further exploit the parallelism, and design an on-chip buffer structure that supports multi-line memory access mode. Lastly, to bolster the adaptability of our accelerator, we propose an innovative macroinstruction set coupled with a micro-kernel scheme. Furthermore, we refine the macroinstruction issue strategy, thereby further enhancing computational efficiency. Our evaluation results demonstrate that SPDFA achieves an average 1.29\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(\\times\\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            \u20132.38\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(\\times\\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            improvement in computational efficiency compared to the state-of-the-art SpMM accelerators when applied to unstructured sparse deep neural network models. Furthermore, its performance outperforms existing sparse neural network accelerators by a factor of 1.03\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(\\times\\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            \u20131.83\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(\\times\\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            . Additionally, SPDFA exhibits excellent scalability with a scaling efficiency exceeding 80%.\n          <\/jats:p>","DOI":"10.1145\/3737462","type":"journal-article","created":{"date-parts":[[2025,5,30]],"date-time":"2025-05-30T10:08:16Z","timestamp":1748599696000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["SPDFA: A Novel Dataflow Fusion Sparse Deep Neural Network Accelerator"],"prefix":"10.1145","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1242-4217","authenticated-orcid":false,"given":"Jinwei","family":"Xu","sequence":"first","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7103-8650","authenticated-orcid":false,"given":"Jingfei","family":"Jiang","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology College of Computer Science and Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1700-705X","authenticated-orcid":false,"given":"Lei","family":"Gao","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology College of Computer Science and Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6691-660X","authenticated-orcid":false,"given":"Xifu","family":"Qian","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology College of Computer Science and Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1256-8934","authenticated-orcid":false,"given":"Yong","family":"Dou","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology College of Computer Science and Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,8,18]]},"reference":[{"doi-asserted-by":"publisher","key":"e_1_3_1_2_2","DOI":"10.1145\/3289602.3293898"},{"unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929","key":"e_1_3_1_3_2"},{"unstructured":"Utku Evci Trevor Gale Jacob Menick Pablo Samuel Castro and Erich Elsen. 2021. Rigging the lottery: Making all tickets winners. arXiv:1911.11134. Retrieved from https:\/\/arxiv.org\/abs\/1911.11134","key":"e_1_3_1_4_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_5_2","DOI":"10.1145\/3581783.3612097"},{"unstructured":"Yiwen Guo Chao Zhang Changshui Zhang and Yurong Chen. 2019. Sparse DNNs with improved adversarial robustness. arXiv:1810.09619. Retrieved from https:\/\/arxiv.org\/abs\/1810.09619","key":"e_1_3_1_6_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_7_2","DOI":"10.1145\/3020078.3021745"},{"unstructured":"Torsten Hoefler Dan Alistarh Tal Ben-Nun Nikoli Dryden and Alexandra Peste. 2021. Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. arXiv:2102.00554. Retrieved from https:\/\/arxiv.org\/abs\/2102.00554","key":"e_1_3_1_8_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_9_2","DOI":"10.1145\/3065386"},{"key":"e_1_3_1_10_2","first-page":"69","volume-title":"Proceedings of the 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA)","author":"Li Zhe","year":"2019","unstructured":"Zhe Li, Caiwen Ding, Siyue Wang, Wujie Wen, Youwei Zhuo, Chang Liu, Qinru Qiu, Wenyao Xu, Xue Lin, Xuehai Qian, et al. 2019. E-RNN: Design optimization for efficient recurrent neural networks in FPGAs. In Proceedings of the 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA), 69\u201380."},{"doi-asserted-by":"publisher","key":"e_1_3_1_11_2","DOI":"10.1145\/3575693.3575706"},{"unstructured":"Zhuang Liu Hanzi Mao Chao-Yuan Wu Christoph Feichtenhofer Trevor Darrell and Saining Xie. 2022. A ConvNet for the 2020s. arXiv:2201.03545. Retrieved from https:\/\/arxiv.org\/abs\/2201.03545","key":"e_1_3_1_12_2"},{"key":"e_1_3_1_13_2","first-page":"17","volume-title":"In Proceedings of the 2019 IEEE 27th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM)","author":"Lu Liqiang","year":"2019","unstructured":"Liqiang Lu, Jiaming Xie, Ruirui Huang, Jiansong Zhang, Wei Lin, and Yun Liang. 2019. An efficient hardware accelerator for sparse convolutional neural networks on FPGAs. In Proceedings of the 2019 IEEE 27th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM), 17\u201325."},{"key":"e_1_3_1_14_2","first-page":"252","volume-title":"Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 (ASPLOS 2023)","author":"Mart\u00ednez Francisco Mu\u00f1oz","year":"2023","unstructured":"Francisco Mu\u00f1oz Mart\u00ednez, Raveesh Garg, Michael Pellauer, Jos\u00e9 L. Abell\u00e1n, Manuel E. Acacio, and Tushar Krishna. 2023. Flexagon: A multi-dataflow sparse-sparse matrix multiplication accelerator for efficient DNN processing. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 (ASPLOS 2023). ACM, New York, NY, 252\u2013265."},{"key":"e_1_3_1_15_2","first-page":"27","volume-title":"Proceedings of the 2017 ACM\/IEEE 44th Annual International Symposium on Computer Architecture (ISCA)","author":"Parashar Angshuman","year":"2017","unstructured":"Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rangharajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W. Keckler, and William J. Dally. 2017. SCNN: An accelerator for compressed-sparse convolutional neural networks. In Proceedings of the 2017 ACM\/IEEE 44th Annual International Symposium on Computer Architecture (ISCA), 27\u201340."},{"key":"e_1_3_1_16_2","first-page":"58","volume-title":"Proceedings of the 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA)","author":"Qin Eric","year":"2020","unstructured":"Eric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella, Sudarshan Srinivasan, Dipankar Das, Bharat Kaul, and Tushar Krishna. 2020. SIGMA: A sparse and irregular GEMM accelerator with flexible interconnects for DNN training. In Proceedings of the 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA), 58\u201370."},{"key":"e_1_3_1_17_2","volume-title":"Proceedings of the 2019 IEEE High Performance Extreme Computing Conference (HPEC)","author":"Reuther Albert","year":"2019","unstructured":"Albert Reuther, Peter Michaleas, Michael Jones, Vijay Gadepally, Siddharth Samsi, and Jeremy Kepner. 2019. Survey and benchmarking of machine learning accelerators. In Proceedings of the 2019 IEEE High Performance Extreme Computing Conference (HPEC). IEEE."},{"doi-asserted-by":"publisher","key":"e_1_3_1_18_2","DOI":"10.1109\/TVLSI.2023.3241933"},{"doi-asserted-by":"publisher","key":"e_1_3_1_19_2","DOI":"10.1145\/3174243.3174253"},{"doi-asserted-by":"publisher","key":"e_1_3_1_20_2","DOI":"10.1109\/ICCD56317.2022.00077"},{"doi-asserted-by":"publisher","key":"e_1_3_1_21_2","DOI":"10.1145\/3445814.3446702"},{"doi-asserted-by":"publisher","key":"e_1_3_1_22_2","DOI":"10.1109\/HPCA47549.2020.00030"},{"doi-asserted-by":"publisher","key":"e_1_3_1_23_2","DOI":"10.1109\/TVLSI.2020.3002779"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3737462","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,19]],"date-time":"2025-08-19T00:23:00Z","timestamp":1755562980000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3737462"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,18]]},"references-count":22,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3737462"],"URL":"https:\/\/doi.org\/10.1145\/3737462","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"type":"print","value":"1936-7406"},{"type":"electronic","value":"1936-7414"}],"subject":[],"published":{"date-parts":[[2025,8,18]]},"assertion":[{"value":"2024-07-07","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-19","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-18","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}