{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,28]],"date-time":"2026-07-28T03:05:03Z","timestamp":1785207903299,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":54,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,10,28]],"date-time":"2023-10-28T00:00:00Z","timestamp":1698451200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,10,28]]},"DOI":"10.1145\/3613424.3623775","type":"proceedings-article","created":{"date-parts":[[2023,12,8]],"date-time":"2023-12-08T17:22:15Z","timestamp":1702056135000},"page":"338-352","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["RM-STC: Row-Merge Dataflow Inspired GPU Sparse Tensor Core for Energy-Efficient Sparse Acceleration"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1280-4781","authenticated-orcid":false,"given":"Guyue","family":"Huang","sequence":"first","affiliation":[{"name":"UC Santa Barbara, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9473-0973","authenticated-orcid":false,"given":"Zhengyang","family":"Wang","sequence":"additional","affiliation":[{"name":"UC Santa Barbara, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4561-6450","authenticated-orcid":false,"given":"Po-An","family":"Tsai","sequence":"additional","affiliation":[{"name":"NVIDIA, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2762-2726","authenticated-orcid":false,"given":"Chen","family":"Zhang","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8716-5793","authenticated-orcid":false,"given":"Yufei","family":"Ding","sequence":"additional","affiliation":[{"name":"UC Santa Barbara, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2093-1788","authenticated-orcid":false,"given":"Yuan","family":"Xie","sequence":"additional","affiliation":[{"name":"Alibaba Group, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,12,8]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001138"},{"key":"e_1_3_2_1_2_1","unstructured":"Dario Amodei 2016. Deep speech 2: End-to-end speech recognition in english and mandarin. In ICML. PMLR 173\u2013182.  Dario Amodei 2016. Deep speech 2: End-to-end speech recognition in english and mandarin. In ICML. PMLR 173\u2013182."},{"key":"e_1_3_2_1_3_1","volume-title":"Language models are few-shot learners. arXiv preprint arXiv:2005.14165","author":"B Brown","year":"2020","unstructured":"Tom\u00a0 B Brown 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 ( 2020 ). Tom\u00a0B Brown 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 (2020)."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"crossref","unstructured":"Shijie Cao 2019. Efficient and effective sparse LSTM on FPGA with bank-balanced sparsity. In FPGA. 63\u201372.  Shijie Cao 2019. Efficient and effective sparse LSTM on FPGA with bank-balanced sparsity. In FPGA. 63\u201372.","DOI":"10.1145\/3289602.3293898"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2910232"},{"key":"e_1_3_2_1_6_1","volume-title":"Nvidia Hopper GPU: Scaling Performance. In 2022 IEEE Hot Chips 34 Symposium (HCS). IEEE Computer Society, 1\u201346","author":"Choquette Jack","year":"2022","unstructured":"Jack Choquette . 2022 . Nvidia Hopper GPU: Scaling Performance. In 2022 IEEE Hot Chips 34 Symposium (HCS). IEEE Computer Society, 1\u201346 . Jack Choquette. 2022. Nvidia Hopper GPU: Scaling Performance. In 2022 IEEE Hot Chips 34 Symposium (HCS). IEEE Computer Society, 1\u201346."},{"key":"e_1_3_2_1_7_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_3_2_1_8_1","first-page":"26183","article-title":"You only look at one sequence: Rethinking transformer in vision through object detection","volume":"34","author":"Fang Yuxin","year":"2021","unstructured":"Yuxin Fang , Bencheng Liao , Xinggang Wang , Jiemin Fang , Jiyang Qi , Rui Wu , Jianwei Niu , and Wenyu Liu . 2021 . You only look at one sequence: Rethinking transformer in vision through object detection . Advances in Neural Information Processing Systems 34 (2021), 26183 \u2013 26197 . Yuxin Fang, Bencheng Liao, Xinggang Wang, Jiemin Fang, Jiyang Qi, Rui Wu, Jianwei Niu, and Wenyu Liu. 2021. You only look at one sequence: Rethinking transformer in vision through object detection. Advances in Neural Information Processing Systems 34 (2021), 26183\u201326197.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_9_1","volume-title":"The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635","author":"Frankle Jonathan","year":"2018","unstructured":"Jonathan Frankle and Michael Carbin . 2018. The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635 ( 2018 ). Jonathan Frankle and Michael Carbin. 2018. The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635 (2018)."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358291"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Song Han 2017. Ese: Efficient Speech Recognition Engine with Sparse LSTM on FPGA. In FPGA. 75\u201384.  Song Han 2017. Ese: Efficient Speech Recognition Engine with Sparse LSTM on FPGA. In FPGA. 75\u201384.","DOI":"10.1145\/3020078.3021745"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001163"},{"key":"e_1_3_2_1_13_1","volume-title":"Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149","author":"Han Song","year":"2015","unstructured":"Song Han , Huizi Mao , and William\u00a0 J Dally . 2015. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149 ( 2015 ). Song Han, Huizi Mao, and William\u00a0J Dally. 2015. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149 (2015)."},{"key":"e_1_3_2_1_14_1","volume-title":"Learning both weights and connections for efficient neural network. Advances in neural information processing systems 28","author":"Han Song","year":"2015","unstructured":"Song Han , Jeff Pool , John Tran , and William Dally . 2015. Learning both weights and connections for efficient neural network. Advances in neural information processing systems 28 ( 2015 ). Song Han, Jeff Pool, John Tran, and William Dally. 2015. Learning both weights and connections for efficient neural network. Advances in neural information processing systems 28 (2015)."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Kartik Hegde 2019. Extensor: An accelerator for sparse tensor algebra. In MICRO. 319\u2013333.  Kartik Hegde 2019. Extensor: An accelerator for sparse tensor algebra. In MICRO. 319\u2013333.","DOI":"10.1145\/3352460.3358275"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00017"},{"key":"e_1_3_2_1_18_1","volume-title":"Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861","author":"Howard G","year":"2017","unstructured":"Andrew\u00a0 G Howard , Menglong Zhu , Bo Chen , Dmitry Kalenichenko , Weijun Wang , Tobias Weyand , Marco Andreetto , and Hartwig Adam . 2017 . Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017). Andrew\u00a0G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017)."},{"key":"e_1_3_2_1_19_1","first-page":"9895","article-title":"Sparse is enough in scaling transformers","volume":"34","author":"Jaszczur Sebastian","year":"2021","unstructured":"Sebastian Jaszczur , Aakanksha Chowdhery , Afroz Mohiuddin , Lukasz Kaiser , Wojciech Gajewski , Henryk Michalewski , and Jonni Kanerva . 2021 . Sparse is enough in scaling transformers . Advances in Neural Information Processing Systems 34 (2021), 9895 \u2013 9907 . Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Lukasz Kaiser, Wojciech Gajewski, Henryk Michalewski, and Jonni Kanerva. 2021. Sparse is enough in scaling transformers. Advances in Neural Information Processing Systems 34 (2021), 9895\u20139907.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00047"},{"key":"e_1_3_2_1_21_1","unstructured":"Yann LeCun 1990. Optimal brain damage. In Advances in neural information processing systems. 598\u2013605.  Yann LeCun 1990. Optimal brain damage. In Advances in neural information processing systems. 598\u2013605."},{"key":"e_1_3_2_1_22_1","volume-title":"Large Models are Parsimonious Learners: Activation Sparsity in Trained Transformers. arXiv preprint arXiv:2210.06313","author":"Li Zonglin","year":"2022","unstructured":"Zonglin Li , Chong You , Srinadh Bhojanapalli , Daliang Li , Ankit\u00a0Singh Rawat , Sashank\u00a0 J Reddi , Ke Ye , Felix Chern , Felix Yu , Ruiqi Guo , 2022. Large Models are Parsimonious Learners: Activation Sparsity in Trained Transformers. arXiv preprint arXiv:2210.06313 ( 2022 ). Zonglin Li, Chong You, Srinadh Bhojanapalli, Daliang Li, Ankit\u00a0Singh Rawat, Sashank\u00a0J Reddi, Ke Ye, Felix Chern, Felix Yu, Ruiqi Guo, 2022. Large Models are Parsimonious Learners: Activation Sparsity in Trained Transformers. arXiv preprint arXiv:2210.06313 (2022)."},{"key":"e_1_3_2_1_23_1","volume-title":"Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030","author":"Ze Liu","year":"2021","unstructured":"Ze Liu 2021. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030 ( 2021 ). Ze Liu 2021. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030 (2021)."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA53966.2022.00049"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3195970.3196120"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO50266.2020.00069"},{"key":"e_1_3_2_1_28_1","unstructured":"Dmitry Molchanov Arsenii Ashukha and Dmitry Vetrov. 2017. Variational dropout sparsifies deep neural networks. In ICML. PMLR 2498\u20132507.  Dmitry Molchanov Arsenii Ashukha and Dmitry Vetrov. 2017. Variational dropout sparsifies deep neural networks. In ICML. PMLR 2498\u20132507."},{"key":"e_1_3_2_1_29_1","unstructured":"NVIDIA. [n. d.]. cuDNN. https:\/\/docs.nvidia.com\/deeplearning\/cudnn\/developer-guide\/index.html. https:\/\/docs.nvidia.com\/deeplearning\/cudnn\/developer-guide\/index.html  NVIDIA. [n. d.]. cuDNN. https:\/\/docs.nvidia.com\/deeplearning\/cudnn\/developer-guide\/index.html. https:\/\/docs.nvidia.com\/deeplearning\/cudnn\/developer-guide\/index.html"},{"key":"e_1_3_2_1_30_1","volume-title":"NVIDIA Tesla V100 GPU Architecture. Data Sheet","year":"2017","unstructured":"Nvidia. 2017. NVIDIA Tesla V100 GPU Architecture. Data Sheet ( 2017 ). Nvidia. 2017. NVIDIA Tesla V100 GPU Architecture. Data Sheet (2017)."},{"key":"e_1_3_2_1_31_1","unstructured":"Nvidia. 2018. Volta Tensor Core GPU Achieves New AI Performance Milestones.https:\/\/developer.nvidia.com\/blog\/tensor-core-ai-performance-milestones\/  Nvidia. 2018. Volta Tensor Core GPU Achieves New AI Performance Milestones.https:\/\/developer.nvidia.com\/blog\/tensor-core-ai-performance-milestones\/"},{"key":"e_1_3_2_1_32_1","volume-title":"NVIDIA A100 Tensor Core GPU. Data Sheet","year":"2020","unstructured":"Nvidia. 2020. NVIDIA A100 Tensor Core GPU. Data Sheet ( 2020 ). Nvidia. 2020. NVIDIA A100 Tensor Core GPU. Data Sheet (2020)."},{"key":"e_1_3_2_1_33_1","unstructured":"Nvidia. 2021. Accelerating Inference with Sparsity Using the NVIDIA Ampere Architecture and NVIDIA TensorRT. https:\/\/developer.nvidia.com\/blog\/accelerating-inference-with-sparsity-using-ampere-and-tensorrt\/  Nvidia. 2021. Accelerating Inference with Sparsity Using the NVIDIA Ampere Architecture and NVIDIA TensorRT. https:\/\/developer.nvidia.com\/blog\/accelerating-inference-with-sparsity-using-ampere-and-tensorrt\/"},{"key":"e_1_3_2_1_34_1","unstructured":"Nvidia. 2021. NVIDIA CUTLASS release v2.7. https:\/\/github.com\/NVIDIA\/cutlass  Nvidia. 2021. NVIDIA CUTLASS release v2.7. https:\/\/github.com\/NVIDIA\/cutlass"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2018.00067"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3140659.3080254"},{"key":"e_1_3_2_1_37_1","volume-title":"Enabling Flexibility for Sparse Tensor Acceleration via Heterogeneity. arXiv preprint arXiv:2201.08916","author":"Qin Eric","year":"2022","unstructured":"Eric Qin , Raveesh Garg , Abhimanyu Bambhaniya , Michael Pellauer , Angshuman Parashar , Sivasankaran Rajamanickam , Cong Hao , and Tushar Krishna . 2022. Enabling Flexibility for Sparse Tensor Acceleration via Heterogeneity. arXiv preprint arXiv:2201.08916 ( 2022 ). Eric Qin, Raveesh Garg, Abhimanyu Bambhaniya, Michael Pellauer, Angshuman Parashar, Sivasankaran Rajamanickam, Cong Hao, and Tushar Krishna. 2022. Enabling Flexibility for Sparse Tensor Acceleration via Heterogeneity. arXiv preprint arXiv:2201.08916 (2022)."},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00015"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00016"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA53966.2022.00068"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO50266.2020.00068"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00062"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.vlsi.2017.02.002"},{"key":"e_1_3_2_1_44_1","unstructured":"Ashish Vaswani 2017. Attention is all you need. In Advances in neural information processing systems. 5998\u20136008.  Ashish Vaswani 2017. Attention is all you need. In Advances in neural information processing systems. 5998\u20136008."},{"key":"e_1_3_2_1_45_1","volume-title":"fairseq s2t: Fast speech-to-text modeling with fairseq. arXiv preprint arXiv:2010.05171","author":"Wang Changhan","year":"2020","unstructured":"Changhan Wang , Yun Tang , Xutai Ma , Anne Wu , Sravya Popuri , Dmytro Okhonko , and Juan Pino . 2020. fairseq s2t: Fast speech-to-text modeling with fairseq. arXiv preprint arXiv:2010.05171 ( 2020 ). Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Sravya Popuri, Dmytro Okhonko, and Juan Pino. 2020. fairseq s2t: Fast speech-to-text modeling with fairseq. arXiv preprint arXiv:2010.05171 (2020)."},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00018"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00088"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS51385.2021.00043"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00096"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00055"},{"key":"e_1_3_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO50266.2020.00064"},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446702"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00030"},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358269"}],"event":{"name":"MICRO '23: 56th Annual IEEE\/ACM International Symposium on Microarchitecture","location":"Toronto ON Canada","acronym":"MICRO '23","sponsor":["SIGMICRO ACM Special Interest Group on Microarchitectural Research and Processing"]},"container-title":["56th Annual IEEE\/ACM International Symposium on Microarchitecture"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3613424.3623775","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3613424.3623775","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:36:30Z","timestamp":1750178190000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3613424.3623775"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,28]]},"references-count":54,"alternative-id":["10.1145\/3613424.3623775","10.1145\/3613424"],"URL":"https:\/\/doi.org\/10.1145\/3613424.3623775","relation":{},"subject":[],"published":{"date-parts":[[2023,10,28]]},"assertion":[{"value":"2023-12-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}