{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,9]],"date-time":"2026-04-09T14:46:03Z","timestamp":1775745963834,"version":"3.50.1"},"reference-count":39,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2024,1,15]],"date-time":"2024-01-15T00:00:00Z","timestamp":1705276800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key R&D","award":["2021YFB0300300"],"award-info":[{"award-number":["2021YFB0300300"]}]},{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["62172430"],"award-info":[{"award-number":["62172430"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"NSF of Hunan Province","award":["2021JJ10052"],"award-info":[{"award-number":["2021JJ10052"]}]},{"name":"STIP of Hunan Province","award":["2022RC3065"],"award-info":[{"award-number":["2022RC3065"]}]},{"name":"Key Laboratory of Advanced Microprocessor Chips and Systems"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2024,3,31]]},"abstract":"<jats:p>Deep learning has become a highly popular research field, and previously deep learning algorithms ran primarily on CPUs and GPUs. However, with the rapid development of deep learning, it was discovered that existing processors could not meet the specific large-scale computing requirements of deep learning, and custom deep learning accelerators have become popular. The majority of the primary workloads in deep learning are general matrix-matrix multiplications (GEMMs), and emerging GEMMs are highly sparse and irregular. The TPU and SIGMA are typical GEMM accelerators in recent years, but the TPU does not support sparsity, and both the TPU and SIGMA have insufficient utilization rates of the Processing Element (PE). We design and implement SparGD, a sparse GEMM accelerator with dynamic dataflow. SparGD has specific PE structures, flexible distribution networks and reduction networks, and a simple dataflow switching module. When running sparse and irregular GEMMs, SparGD can maintain high PE utilization while utilizing sparsity, and can switch to the optimal dataflow according to the computing environment. For sparse, irregular GEMMs, our experimental results show that SparGD outperforms systolic arrays by 30 times and SIGMA by 3.6 times.<\/jats:p>","DOI":"10.1145\/3634703","type":"journal-article","created":{"date-parts":[[2023,11,27]],"date-time":"2023-11-27T15:37:54Z","timestamp":1701099474000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["SparGD: A Sparse GEMM Accelerator with Dynamic Dataflow"],"prefix":"10.1145","volume":"29","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-9441-0509","authenticated-orcid":false,"given":"Bo","family":"Wang","sequence":"first","affiliation":[{"name":"School of Computer, National University of Defense Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1710-4060","authenticated-orcid":false,"given":"Sheng","family":"Ma","sequence":"additional","affiliation":[{"name":"School of Computer, National University of Defense Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-5551-2897","authenticated-orcid":false,"given":"Shengbai","family":"Luo","sequence":"additional","affiliation":[{"name":"School of Computer, National University of Defense Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4439-7436","authenticated-orcid":false,"given":"Lizhou","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Computer, National University of Defense Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1008-4805","authenticated-orcid":false,"given":"Jianmin","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer, National University of Defense Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0944-2708","authenticated-orcid":false,"given":"Chunyuan","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer, National University of Defense Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1509-1761","authenticated-orcid":false,"given":"Tiejun","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer, National University of Defense Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,1,15]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"crossref","first-page":"802","DOI":"10.1109\/HPCA51647.2021.00072","volume-title":"2021 IEEE International Symposium on High-performance Computer Architecture (HPCA\u201921)","author":"Acun Bilge","year":"2021","unstructured":"Bilge Acun, Matthew Murphy, Xiaodong Wang, Jade Nie, Carole-Jean Wu, and Kim Hazelwood. 2021. Understanding training efficiency of deep learning recommendation models at scale. In 2021 IEEE International Symposium on High-performance Computer Architecture (HPCA\u201921). IEEE, 802\u2013814."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001138"},{"key":"e_1_3_1_4_2","first-page":"11216","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Cao Shijie","year":"2019","unstructured":"Shijie Cao, Lingxiao Ma, Wencong Xiao, Chen Zhang, Yunxin Liu, Lintao Zhang, Lanshun Nie, and Zhi Yang. 2019. Seernet: Predicting convolutional neural network feature-map sparsity through low-bit quantization. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 11216\u201311225."},{"key":"e_1_3_1_5_2","first-page":"551","volume-title":"2009 Computation World: Future Computing, Service Computation, Cognitive, Adaptive, Content, Patterns","author":"Chakrabarty Amitabha","year":"2009","unstructured":"Amitabha Chakrabarty, Martin Collier, and Sourav Mukhopadhyay. 2009. Matrix-based nonblocking routing algorithm for Bene\u0161 networks. In 2009 Computation World: Future Computing, Service Computation, Cognitive, Adaptive, Content, Patterns. IEEE, 551\u2013556."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2016.2616357"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2910232"},{"key":"e_1_3_1_8_2","unstructured":"Timothy A. Davis and Yifan Hu. 2011. The University of Florida Sparse Matrix Collection. Retrieved August 7 2023 from https:\/\/sparse.tamu.edu"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00090"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18074.2021.9586216"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","first-page":"151","DOI":"10.1145\/3352460.3358291","volume-title":"Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture","author":"Gondimalla Ashish","year":"2019","unstructured":"Ashish Gondimalla, Noah Chesnut, Mithuna Thottethodi, and T. N. Vijaykumar. 2019. SparTen: A sparse tensor accelerator for convolutional neural networks. In Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture. 151\u2013165."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2017.10.013"},{"key":"e_1_3_1_13_2","first-page":"1","volume-title":"2019 28th International Conference on Parallel Architectures and Compilation Techniques (PACT\u201919)","author":"Gupta Udit","year":"2019","unstructured":"Udit Gupta, Brandon Reagen, Lillian Pentecost, Marco Donato, Thierry Tambe, Alexander M. Rush, Gu-Yeon Wei, and David Brooks. 2019. MASR: A modular accelerator for sparse RNNs. In 2019 28th International Conference on Parallel Architectures and Compilation Techniques (PACT\u201919). IEEE, 1\u201314."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001163"},{"key":"e_1_3_1_15_2","article-title":"Learning both weights and connections for efficient neural network","volume":"28","author":"Han Song","year":"2015","unstructured":"Song Han, Jeff Pool, John Tran, and William Dally. 2015. Learning both weights and connections for efficient neural network. Advances in Neural Information Processing Systems 28 (2015), 1\u20139.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358275"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC.2014.6757323"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3065386"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3296957.3173176"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2019.2924215"},{"key":"e_1_3_1_22_2","first-page":"553","volume-title":"2017 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201917)","author":"Lu Wenyan","year":"2017","unstructured":"Wenyan Lu, Guihai Yan, Jiajun Li, Shijun Gong, Yinhe Han, and Xiaowei Li. 2017. Flexflow: A flexible dataflow accelerator architecture for convolutional neural networks. In 2017 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201917). IEEE, 553\u2013564."},{"key":"e_1_3_1_23_2","article-title":"STIFT: A spatio-temporal integrated folding tree for efficient reductions in flexible DNN accelerators","author":"Mu\u00f1oz-Mart\u00ednez Francisco","year":"2022","unstructured":"Francisco Mu\u00f1oz-Mart\u00ednez, Jos\u00e9 L. Abell\u00e1n, Manuel E. Acacio, and Tushar Krishna. 2022. STIFT: A spatio-temporal integrated folding tree for efficient reductions in flexible DNN accelerators. ACM Journal on Emerging Technologies in Computing Systems (JETC\u201922) 19 (2022), 1\u201320.","journal-title":"ACM Journal on Emerging Technologies in Computing Systems (JETC\u201922)"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-018-09679-z"},{"key":"e_1_3_1_25_2","unstructured":"openai.com. 2018. AI and Compute. Retrieved March 7 2023 from https:\/\/openai.com\/blog\/ai-and-compute\/"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.2979670"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3140659.3080254"},{"key":"e_1_3_1_28_2","first-page":"58","volume-title":"2020 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201920)","author":"Qin Eric","year":"2020","unstructured":"Eric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella, Sudarshan Srinivasan, Dipankar Das, Bharat Kaul, and Tushar Krishna. 2020. Sigma: A sparse and irregular GEMM accelerator with flexible interconnects for DNN training. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201920). IEEE, 58\u201370."},{"key":"e_1_3_1_29_2","first-page":"58","volume-title":"2020 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS\u201920)","author":"Samajdar Ananda","year":"2020","unstructured":"Ananda Samajdar, Jan Moritz Joseph, Yuhao Zhu, Paul Whatmough, Matthew Mattina, and Tushar Krishna. 2020. A systematic methodology for characterizing scalability of DNN accelerators using scale-sim. In 2020 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS\u201920). IEEE, 58\u201368."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10070977"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2670313"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.vlsi.2022.03.002"},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","first-page":"97","DOI":"10.1109\/HPCA51647.2021.00018","volume-title":"2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA\u201921)","author":"Wang Hanrui","year":"2021","unstructured":"Hanrui Wang, Zhekai Zhang, and Song Han. 2021. Spatten: Efficient sparse attention architecture with cascade token and head pruning. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA\u201921). IEEE, 97\u2013110."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2021.3060041"},{"key":"e_1_3_1_35_2","article-title":"A survey of design and optimization for systolic array based DNN accelerators","author":"Xu Rui","year":"2023","unstructured":"Rui Xu, Sheng Ma, Yang Guo, and Dongsheng Li. 2023. A survey of design and optimization for systolic array based DNN accelerators. Computing Surveys 56 (2023), 1\u201337.","journal-title":"Computing Surveys"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460776"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3129647"},{"key":"e_1_3_1_38_2","article-title":"A survey of deep learning techniques for neural machine translation","author":"Yang Shuoheng","year":"2020","unstructured":"Shuoheng Yang, Yuxin Wang, and Xiaowen Chu. 2020. A survey of deep learning techniques for neural machine translation. arXiv preprint arXiv:2002.07526 (2020).","journal-title":"arXiv preprint arXiv:2002.07526"},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","first-page":"650","DOI":"10.1109\/ISCA.2018.00060","volume-title":"2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA\u201918)","author":"Yazdanbakhsh Amir","year":"2018","unstructured":"Amir Yazdanbakhsh, Kambiz Samadi, Nam Sung Kim, and Hadi Esmaeilzadeh. 2018. Ganax: A unified MIMD-SIMD acceleration for generative adversarial networks. In 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA\u201918). IEEE, 650\u2013661."},{"key":"e_1_3_1_40_2","first-page":"1","volume-title":"2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201916)","author":"Zhang Shijin","year":"2016","unstructured":"Shijin Zhang, Zidong Du, Lei Zhang, Huiying Lan, Shaoli Liu, Ling Li, Qi Guo, Tianshi Chen, and Yunji Chen. 2016. Cambricon-X: An accelerator for sparse neural networks. In 2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201916). IEEE, 1\u201312."}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3634703","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3634703","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:35:49Z","timestamp":1750178149000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3634703"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,15]]},"references-count":39,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,3,31]]}},"alternative-id":["10.1145\/3634703"],"URL":"https:\/\/doi.org\/10.1145\/3634703","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"value":"1084-4309","type":"print"},{"value":"1557-7309","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,15]]},"assertion":[{"value":"2023-06-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-19","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}