{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T16:20:24Z","timestamp":1781972424497,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":36,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,3,25]],"date-time":"2023-03-25T00:00:00Z","timestamp":1679702400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"MCIN\/AEI\/10.13039\/5011000110\\\\33","award":["TED2021-130233B-C33"],"award-info":[{"award-number":["TED2021-130233B-C33"]}]},{"name":"European Union NextGenerationEU\/PRTR","award":["RYC2021-031966-I"],"award-info":[{"award-number":["RYC2021-031966-I"]}]},{"name":"Fundacion Seneca","award":["20749\/FPI\/18"],"award-info":[{"award-number":["20749\/FPI\/18"]}]},{"name":"U.S. Department of Energy (DOE) Office of Science, Advanced Scientific Computing Research program","award":["ARIAA co-design center"],"award-info":[{"award-number":["ARIAA co-design center"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,3,25]]},"DOI":"10.1145\/3582016.3582069","type":"proceedings-article","created":{"date-parts":[[2023,3,20]],"date-time":"2023-03-20T16:59:03Z","timestamp":1679331543000},"page":"252-265","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":57,"title":["Flexagon: A Multi-dataflow Sparse-Sparse Matrix Multiplication Accelerator for Efficient DNN Processing"],"prefix":"10.1145","author":[{"given":"Francisco","family":"Mu\u00f1oz-Mart\u00ednez","sequence":"first","affiliation":[{"name":"Universidad de Murcia, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Raveesh","family":"Garg","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Michael","family":"Pellauer","sequence":"additional","affiliation":[{"name":"NVIDIA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jos\u00e9 L.","family":"Abell\u00e1n","sequence":"additional","affiliation":[{"name":"Universidad de Murcia, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Manuel E.","family":"Acacio","sequence":"additional","affiliation":[{"name":"Universidad de Murcia, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tushar","family":"Krishna","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,3,25]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"[n. d.]. MAERI code v1. https:\/\/github.com\/hyoukjun\/MAERI. \t\t\t\t  [n. d.]. MAERI code v1. https:\/\/github.com\/hyoukjun\/MAERI."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2910232"},{"key":"e_1_3_2_1_3_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv : 1810.04805v2 (2019), May. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv: 1810.04805v2 (2019), May."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18074.2021.9586114"},{"key":"e_1_3_2_1_5_1","volume-title":"2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS).","author":"Garg Raveesh","year":"2022","unstructured":"Raveesh Garg , Eric Qin , Francisco Mu\u00f1oz-Mart\u00ednez , Robert Guirado , Akshay Jain , Sergi Abadal , Jos\u00e9 L Abell\u00e1n , Manuel E Acacio , Eduard Alarc\u00f3n , Sivasankaran Rajamanickam , and Tushar Krishna . 2022 . Understanding the Design-Space of Sparse\/Dense Multiphase GNN dataflows on Spatial Accelerators . In 2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS). Raveesh Garg, Eric Qin, Francisco Mu\u00f1oz-Mart\u00ednez, Robert Guirado, Akshay Jain, Sergi Abadal, Jos\u00e9 L Abell\u00e1n, Manuel E Acacio, Eduard Alarc\u00f3n, Sivasankaran Rajamanickam, and Tushar Krishna. 2022. Understanding the Design-Space of Sparse\/Dense Multiphase GNN dataflows on Spatial Accelerators. In 2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS)."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358291"},{"key":"e_1_3_2_1_7_1","article-title":"Two Fast Algorithms for Sparse Matrices: Multiplication and Permuted Transposition","volume":"4","author":"Gustavson Fred G.","year":"1978","unstructured":"Fred G. Gustavson . 1978 . Two Fast Algorithms for Sparse Matrices: Multiplication and Permuted Transposition . ACM Trans. Math. Softw. , 4 , 3 (1978), sep, 250\u2013269. Fred G. Gustavson. 1978. Two Fast Algorithms for Sparse Matrices: Multiplication and Permuted Transposition. ACM Trans. Math. Softw., 4, 3 (1978), sep, 250\u2013269.","journal-title":"ACM Trans. Math. Softw."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.30"},{"key":"e_1_3_2_1_9_1","volume-title":"Dally","author":"Han Song","year":"2016","unstructured":"Song Han , Huizi Mao , and William J . Dally . 2016 . Deep Compression : Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding . arXiv preprint arXiv: 1510.00149v5 (2016), Feb.. Song Han, Huizi Mao, and William J. Dally. 2016. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding. arXiv preprint arXiv: 1510.00149v5 (2016), Feb.."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358275"},{"key":"e_1_3_2_1_11_1","unstructured":"HP Laboratories. [n. d.]. CACTI 7.0: A Tool to Model Caches\/Memories 3D stacking and off-chip IO. https:\/\/github.com\/HewlettPackard\/cacti. \t\t\t\t  HP Laboratories. [n. d.]. CACTI 7.0: A Tool to Model Caches\/Memories 3D stacking and off-chip IO. https:\/\/github.com\/HewlettPackard\/cacti."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3140659.3080246"},{"key":"e_1_3_2_1_13_1","volume-title":"Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture. 600\u2013614","author":"Kanellopoulos Konstantinos","year":"2019","unstructured":"Konstantinos Kanellopoulos , Nandita Vijaykumar , Christina Giannoula , Roknoddin Azizi , Skanda Koppula , Nika Mansouri Ghiasi , Taha Shahroodi , Juan Gomez Luna , and Onur Mutlu . 2019 . Smash: Co-designing software compression and hardware-accelerated indexing for efficient sparse matrix operations . In Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture. 600\u2013614 . Konstantinos Kanellopoulos, Nandita Vijaykumar, Christina Giannoula, Roknoddin Azizi, Skanda Koppula, Nika Mansouri Ghiasi, Taha Shahroodi, Juan Gomez Luna, and Onur Mutlu. 2019. Smash: Co-designing software compression and hardware-accelerated indexing for efficient sparse matrix operations. In Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture. 600\u2013614."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/3155562.3155683"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358252"},{"key":"e_1_3_2_1_16_1","volume-title":"Int\u2019l Conf. on Architectural Support for Programming Languages and Operating Systems, March.","author":"Kwon Hyoukjun","year":"2018","unstructured":"Hyoukjun Kwon , Ananda Samajdar , and Tushar Krishna . 2018 . MAERI: Enabling Flexible Dataflow Mapping over DNN Accelerators via Reconfigurable Interconnects . Int\u2019l Conf. on Architectural Support for Programming Languages and Operating Systems, March. Hyoukjun Kwon, Ananda Samajdar, and Tushar Krishna. 2018. MAERI: Enabling Flexible Dataflow Mapping over DNN Accelerators via Reconfigurable Interconnects. Int\u2019l Conf. on Architectural Support for Programming Languages and Operating Systems, March."},{"key":"e_1_3_2_1_17_1","volume-title":"SysML Conference. 120","author":"Lee Ching-En","year":"2018","unstructured":"Ching-En Lee , Yakun Sophia Shao , Jie-Fang Zhang , Angshuman Parashar , Joel Emer , Stephen W Keckler , and Zhengya Zhang . 2018 . Stitch-x: An accelerator architecture for exploiting unstructured sparsity in deep neural networks . In SysML Conference. 120 . Ching-En Lee, Yakun Sophia Shao, Jie-Fang Zhang, Angshuman Parashar, Joel Emer, Stephen W Keckler, and Zhengya Zhang. 2018. Stitch-x: An accelerator architecture for exploiting unstructured sparsity in deep neural networks. In SysML Conference. 120."},{"issue":"2","key":"e_1_3_2_1_18_1","first-page":"57","article-title":"A Simulator for Large-Scale Parallel Computer Architectures","volume":"1","author":"Lee Janssen Curtis","year":"2010","unstructured":"Janssen Curtis Lee , Helgi Adalsteinsson , Scott Cranford , Joseph P. Kenny , Ali Pinar , David A. Evensky , and Jackson R. Mayo . 2010 . A Simulator for Large-Scale Parallel Computer Architectures .. IJDST vol. 1 , no. 2 , 57 \u2013 73 . Janssen Curtis Lee, Helgi Adalsteinsson, Scott Cranford, Joseph P. Kenny, Ali Pinar, David A. Evensky, and Jackson R. Mayo. 2010. A Simulator for Large-Scale Parallel Computer Architectures.. IJDST vol.1, no.2, 57\u201373.","journal-title":"IJDST"},{"key":"e_1_3_2_1_19_1","unstructured":"Peter Mattson Christine Cheng Cody Coleman Greg Diamos Paulius Micikevicius David Patterson Hanlin Tang Gu-Yeon Wei Peter Bailis and Victor Bittorf. 2019. Mlperf training benchmark. arXiv preprint arXiv:1910.01500. \t\t\t\t  Peter Mattson Christine Cheng Cody Coleman Greg Diamos Paulius Micikevicius David Patterson Hanlin Tang Gu-Yeon Wei Peter Bailis and Victor Bittorf. 2019. Mlperf training benchmark. arXiv preprint arXiv:1910.01500."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC53511.2021.00028"},{"key":"e_1_3_2_1_21_1","unstructured":"Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G. Azzolini Dmytro Dzhulgakov Andrey Mallevich Ilia Cherniavskii Yinghai Lu Raghuraman Krishnamoorthi Ansha Yu Volodymyr Kondratenko Stephanie Pereira Xianjie Chen Wenlin Chen Vijay Rao Bill Jia Liang Xiong and Misha Smelyanskiy. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems. CoRR abs\/1906.00091 (2019) arxiv:1906.00091 \t\t\t\t  Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G. Azzolini Dmytro Dzhulgakov Andrey Mallevich Ilia Cherniavskii Yinghai Lu Raghuraman Krishnamoorthi Ansha Yu Volodymyr Kondratenko Stephanie Pereira Xianjie Chen Wenlin Chen Vijay Rao Bill Jia Liang Xiong and Misha Smelyanskiy. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems. CoRR abs\/1906.00091 (2019) arxiv:1906.00091"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480134"},{"key":"e_1_3_2_1_23_1","unstructured":"Subhankar Pal Jonathan Beaumont Dong-Hyeon Park Aporva Amarnath Siying Feng Chaitali Chakrabarti Hun-Seok Kim David Blaauw Trevor Mudge and Ronald Dreslinski. 2018. OuterSPACE: An Outer Product based Sparse Matrix Multiplication Accelerator. In ISCA. \t\t\t\t  Subhankar Pal Jonathan Beaumont Dong-Hyeon Park Aporva Amarnath Siying Feng Chaitali Chakrabarti Hun-Seok Kim David Blaauw Trevor Mudge and Ronald Dreslinski. 2018. OuterSPACE: An Outer Product based Sparse Matrix Multiplication Accelerator. In ISCA."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00042"},{"key":"e_1_3_2_1_25_1","volume-title":"SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks. International Symposium on Computer Architecture (ISCA), June, 27\u201340","author":"Parashar Angshuman","unstructured":"Angshuman Parashar , Minsoo Rhu , Anurag Mukkara , Antonio Puglielli , Rangharajan Venkatesan , Brucek Khailany , Joel Emer , Stephen W. Keckler , and William J. Dally . 2017 . SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks. International Symposium on Computer Architecture (ISCA), June, 27\u201340 . Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rangharajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W. Keckler, and William J. Dally. 2017. SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks. International Symposium on Computer Architecture (ISCA), June, 27\u201340."},{"key":"e_1_3_2_1_26_1","volume-title":"2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 1014\u20131024","author":"Qin Eric","year":"2021","unstructured":"Eric Qin , Geonhwa Jeong , William Won , Sheng-Chun Kao , Hyoukjun Kwon , Sudarshan Srinivasan , Dipankar Das , Gordon E Moon , Sivasankaran Rajamanickam , and Tushar Krishna . 2021 . Extending sparse tensor accelerators to support multiple compression formats . In 2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 1014\u20131024 . Eric Qin, Geonhwa Jeong, William Won, Sheng-Chun Kao, Hyoukjun Kwon, Sudarshan Srinivasan, Dipankar Das, Gordon E Moon, Sivasankaran Rajamanickam, and Tushar Krishna. 2021. Extending sparse tensor accelerators to support multiple compression formats. In 2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 1014\u20131024."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS49936.2021.00110"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00015"},{"key":"e_1_3_2_1_29_1","volume-title":"2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). 446\u2013459","author":"Reddi Vijay Janapa","year":"2020","unstructured":"Vijay Janapa Reddi , Christine Cheng , David Kanter , Peter Mattson , Guenther Schmuelling , Carole-Jean Wu , Brian Anderson , Maximilien Breughe , Mark Charlebois , and William Chou . 2020 . Mlperf inference benchmark . In 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). 446\u2013459 . Vijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson, Guenther Schmuelling, Carole-Jean Wu, Brian Anderson, Maximilien Breughe, Mark Charlebois, and William Chou. 2020. Mlperf inference benchmark. In 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). 446\u2013459."},{"key":"e_1_3_2_1_30_1","volume-title":"2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 766\u2013780","author":"Srivastava Nitish","year":"2020","unstructured":"Nitish Srivastava , Hanchen Jin , Jie Liu , David Albonesi , and Zhiru Zhang . 2020 . Matraptor: A sparse-sparse matrix multiplication accelerator based on row-wise product . In 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 766\u2013780 . Nitish Srivastava, Hanchen Jin, Jie Liu, David Albonesi, and Zhiru Zhang. 2020. Matraptor: A sparse-sparse matrix multiplication accelerator based on row-wise product. In 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 766\u2013780."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2917185"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"crossref","unstructured":"Endong Wang Qing Zhang Bo Shen Guangyong Zhang Xiaowei Lu Qing Wu and Yajuan Wang. 2014. Intel Math Kernel Library. High-Performance Computing on the Intel Xeon Phi June. \t\t\t\t  Endong Wang Qing Zhang Bo Shen Guangyong Zhang Xiaowei Lu Qing Wu and Yajuan Wang. 2014. Intel Math Kernel Library. High-Performance Computing on the Intel Xeon Phi June.","DOI":"10.1007\/978-3-319-06486-4"},{"key":"e_1_3_2_1_33_1","volume-title":"SparseRT: Accelerating Unstructured Sparsity on GPUs for Deep Learning Inference. arXiv preprint arXiv","author":"Wang Ziheng","year":"2008","unstructured":"Ziheng Wang . 2020. SparseRT: Accelerating Unstructured Sparsity on GPUs for Deep Learning Inference. arXiv preprint arXiv : 2008 .11849v1 (2020), Aug.. Ziheng Wang. 2020. SparseRT: Accelerating Unstructured Sparsity on GPUs for Deep Learning Inference. arXiv preprint arXiv: 2008.11849v1 (2020), Aug.."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446702"},{"key":"e_1_3_2_1_35_1","volume-title":"SpArch: Efficient Architecture for Sparse Matrix Multiplication. International Symposium on High Performance Computer Architecture (HPCA), Feb., 261\u2013274","author":"Zhang Zhekai","unstructured":"Zhekai Zhang , Hanrui Wang , Song Han , and William J. Dally . 2020 . SpArch: Efficient Architecture for Sparse Matrix Multiplication. International Symposium on High Performance Computer Architecture (HPCA), Feb., 261\u2013274 . Zhekai Zhang, Hanrui Wang, Song Han, and William J. Dally. 2020. SpArch: Efficient Architecture for Sparse Matrix Multiplication. International Symposium on High Performance Computer Architecture (HPCA), Feb., 261\u2013274."},{"key":"e_1_3_2_1_36_1","volume-title":"2020 57th ACM\/IEEE Design Automation Conference (DAC), Oct..","author":"Zhu Maohua","year":"2010","unstructured":"Maohua Zhu and Yuan Xie . 2010 . Taming Unstructured Sparsity on GPUs via Latency-Aware Optimization . 2020 57th ACM\/IEEE Design Automation Conference (DAC), Oct.. Maohua Zhu and Yuan Xie. 2010. Taming Unstructured Sparsity on GPUs via Latency-Aware Optimization. 2020 57th ACM\/IEEE Design Automation Conference (DAC), Oct.."}],"event":{"name":"ASPLOS '23: 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3","location":"Vancouver BC Canada","acronym":"ASPLOS '23","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture","SIGOPS ACM Special Interest Group on Operating Systems","SIGPLAN ACM Special Interest Group on Programming Languages","SIGBED ACM Special Interest Group on Embedded Systems"]},"container-title":["Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3582016.3582069","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:46Z","timestamp":1750178806000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3582016.3582069"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,25]]},"references-count":36,"alternative-id":["10.1145\/3582016.3582069","10.1145\/3582016"],"URL":"https:\/\/doi.org\/10.1145\/3582016.3582069","relation":{},"subject":[],"published":{"date-parts":[[2023,3,25]]},"assertion":[{"value":"2023-03-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}