{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T14:34:43Z","timestamp":1777127683779,"version":"3.51.4"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"1-4","license":[{"start":{"date-parts":[[2023,11,30]],"date-time":"2023-11-30T00:00:00Z","timestamp":1701302400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Comput. Syst."],"published-print":{"date-parts":[[2023,11,30]]},"abstract":"<jats:p>Sparse tensor algorithms are becoming widespread, particularly in the domains of deep learning, graph and data analytics, and scientific computing. Current high-performance broad-domain architectures, such as GPUs, often suffer memory system inefficiencies by moving too much data or moving it too far through the memory hierarchy. To increase performance and efficiency, proposed domain-specific accelerators tailor their architectures to the data needs of a narrow application domain, but as a result cannot be applied to a wide range of algorithms or applications that contain a mix of sparse and dense algorithms.<\/jats:p>\n          <jats:p>This article proposes Symphony, a hybrid programmable\/specialized architecture that focuses on the orchestration of data throughout the memory hierarchy to simultaneously reduce the movement of unnecessary data and data movement distances. Key elements of the Symphony architecture include (1) specialized reconfigurable units aimed not only at roofline floating-point computations but also at supporting data orchestration features, such as address generation, data filtering, and sparse metadata processing; and (2) distribution of computation resources (both programmable and specialized) throughout the on-chip memory hierarchy. We demonstrate that Symphony can match non-programmable ASIC performance on sparse tensor algebra and provide\u00a031\u00d7 improved runtime and 44\u00d7 improved energy over a comparably provisioned GPU for these applications.<\/jats:p>","DOI":"10.1145\/3630007","type":"journal-article","created":{"date-parts":[[2023,10,27]],"date-time":"2023-10-27T22:23:15Z","timestamp":1698445395000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Symphony: Orchestrating Sparse and Dense Tensors with Hierarchical Heterogeneous Processing"],"prefix":"10.1145","volume":"41","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5305-4307","authenticated-orcid":false,"given":"Michael","family":"Pellauer","sequence":"first","affiliation":[{"name":"NVIDIA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5533-417X","authenticated-orcid":false,"given":"Jason","family":"Clemons","sequence":"additional","affiliation":[{"name":"NVIDIA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-0850-2363","authenticated-orcid":false,"given":"Vignesh","family":"Balaji","sequence":"additional","affiliation":[{"name":"NVIDIA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7774-0531","authenticated-orcid":false,"given":"Neal","family":"Crago","sequence":"additional","affiliation":[{"name":"NVIDIA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5709-2992","authenticated-orcid":false,"given":"Aamer","family":"Jaleel","sequence":"additional","affiliation":[{"name":"NVIDIA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-9572-5391","authenticated-orcid":false,"given":"Donghyuk","family":"Lee","sequence":"additional","affiliation":[{"name":"NVIDIA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0944-2393","authenticated-orcid":false,"given":"Mike","family":"O\u2019Connor","sequence":"additional","affiliation":[{"name":"NVIDIA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9936-6501","authenticated-orcid":false,"given":"Angshuman","family":"Parashar","sequence":"additional","affiliation":[{"name":"NVIDIA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2189-4026","authenticated-orcid":false,"given":"Sean","family":"Treichler","sequence":"additional","affiliation":[{"name":"NVIDIA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4561-6450","authenticated-orcid":false,"given":"Po-An","family":"Tsai","sequence":"additional","affiliation":[{"name":"NVIDIA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6701-6099","authenticated-orcid":false,"given":"Stephen W.","family":"Keckler","sequence":"additional","affiliation":[{"name":"NVIDIA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3459-5466","authenticated-orcid":false,"given":"Joel S.","family":"Emer","sequence":"additional","affiliation":[{"name":"NVIDIA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,12,18]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378454"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750386"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2008.4536313"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/2678373.2665705"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/43.945302"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2016.2616357"},{"key":"e_1_3_2_8_2","volume-title":"Unified Sparse Formats for Tensor Algebra Compilers","author":"Chou Stephen","year":"2018","unstructured":"Stephen Chou. 2018. Unified Sparse Formats for Tensor Algebra Compilers. Master\u2019s thesis. Massachusetts Institute of Technology."},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of the Conference on Neural Information Processing Systems (NeurIPS)","author":"Cuturi Marco","year":"2013","unstructured":"Marco Cuturi. 2013. Sinkhorn distances: Lightspeed computation of optimal transport. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00053"},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Vidushi Dadu Jian Weng Sihao Liu and Tony Nowatzki. 2019. Towards general purpose acceleration by exploiting common data-dependence forms. In the Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO-52). 924\u2013939.","DOI":"10.1145\/3352460.3358276"},{"issue":"1","key":"e_1_3_2_12_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2049662.2049663","article-title":"The university of Florida sparse matrix collection","volume":"38","author":"Davis Timothy A.","year":"2011","unstructured":"Timothy A. Davis and Yifan Hu. 2011. The university of Florida sparse matrix collection. ACM Trans. Math. Softw. 38, 1 (2011), 1\u201325.","journal-title":"ACM Trans. Math. Softw."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/2145694.2145725"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783759"},{"key":"e_1_3_2_15_2","volume-title":"Proceedings of the Conference on Neural Information Processing Systems (NeurIPS\u201917)","author":"Hamilton Will","year":"2017","unstructured":"Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive representation learning on large graphs. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS\u201917)."},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358275"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC.2014.6757323"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_2_20_2","volume-title":"Proceedings of the Design Automation Conference (DAC\u201918)","author":"Khailany Brucek","year":"2018","unstructured":"Brucek Khailany, Evgeni Krimer, Rangharajan Venkatesan, Jason Clemons, Joel S. Emer, Matthew Fojtik, Alicia Klinefelter, Michael Pellauer, Nathaniel Pinckney, Yakun Sophia Shao, Shreesha Srinath, Christopher Torng, Sam Likun Xi, Yanqing Zhang, and Brian Zimmer. 2018. A modular digital VLSI flow for high-productivity SoC design. In Proceedings of the Design Automation Conference (DAC\u201918)."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126965"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2017.8115709"},{"key":"e_1_3_2_23_2","volume-title":"Learning Multiple Layers of Features from Tiny Images","author":"Krizhevsky Alex","year":"2009","unstructured":"Alex Krizhevsky. 2009. Learning Multiple Layers of Features from Tiny Images. Master\u2019s thesis. Department of Computer Science, University of Toronto."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358252"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357375"},{"key":"e_1_3_2_26_2","unstructured":"Yury A. Malkov and D. A. Yashunin. 2016. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. Retrieved from http:\/\/arxiv.org\/abs\/1603.09320"},{"key":"e_1_3_2_27_2","volume-title":"Proceedings of the Conference on Neural Information Processing Systems (NeurIPS\u201918)","author":"Morozov Stanislav","year":"2018","unstructured":"Stanislav Morozov and Artem Babenko. 2018. Non-metric similarity graphs for maximum inner product search. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS\u201918)."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358254"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3316781.3323476"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO50266.2020.00056"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480048"},{"key":"e_1_3_2_32_2","article-title":"A Coarse Grain Reconfigurable Array (CGRA) for Statically Scheduled Data Flow Computing","author":"Nicol Chris","unstructured":"Chris Nicol. A Coarse Grain Reconfigurable Array (CGRA) for Statically Scheduled Data Flow Computing. Retrieved from https:\/\/wavecomp.ai\/wp-content\/uploads\/2018\/12\/WP_CGRA.pdf","journal-title":"R"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00022"},{"key":"e_1_3_2_34_2","article-title":"NVIDIA A100 Tensor Core GPU Architecture","year":"2020","unstructured":"NVIDIA. 2020. NVIDIA A100 Tensor Core GPU Architecture. Retrieved from https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/Data-Center\/nvidia-ampere-architecture-whitepaper.pdf","journal-title":"R"},{"key":"e_1_3_2_35_2","article-title":"cuSPARSE","year":"2021","unstructured":"NVIDIA. 2021. cuSPARSE. Retrieved from https:\/\/docs.nvidia.com\/cuda\/cusparse\/index.html","journal-title":"R"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3582016.3582064"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3018743.3018749"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.24"},{"key":"e_1_3_2_39_2","volume-title":"Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA\u201919)","author":"Pal Subhankar","year":"2019","unstructured":"Subhankar Pal, Jonathan Beaumont, Dong-Hyeon Park, Aporva Amarnath, Siying Feng, Chaitali Chakrabarti, Hun-Seok Kim, David Blaauw, Trevor Mudge, and Ronald Dreslinski. 2019. OuterSPACE: An outer product based sparse matrix multiplication accelerator. In Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA\u201919)."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00042"},{"key":"e_1_3_2_41_2","volume-title":"Proceedings of the International Conference on Architectural Support for Programming Languages and Operation Systems (ASPLOS\u201918)","author":"Pellauer Michael","year":"2018","unstructured":"Michael Pellauer, Yakun Sophia Shao, Jason Clemons, Neal Crago, Karthik Hedge, Rangharajan Ventakesan, Stephen Keckler, Christopher W. Fletcher, and Joel Emer. 2018. Buffets: An efficient, flexible, composable storage idiom for accelerators. In Proceedings of the International Conference on Architectural Support for Programming Languages and Operation Systems (ASPLOS\u201918)."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080256"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00015"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO50266.2020.00078"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358302"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994522"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/1067649.801719"},{"key":"e_1_3_2_48_2","volume-title":"Proceedings of the International Symposium on Microarchitecture (MICRO\u201902)","author":"Srivastava Nitish","year":"2002","unstructured":"Nitish Srivastava, Hanchen Jin, Jie Liu, David Albonesi, and Zhiru Zhang. 2002. MatRaptor: A sparse-sparse matrix multiplication accelerator based on row-wise product. In Proceedings of the International Symposium on Microarchitecture (MICRO\u201902)."},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00062"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.2200\/S01004ED1V01Y202004CAC050"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00077"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358307"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/2851141.2851145"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS51385.2021.00043"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAD45719.2019.8942149"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00087"},{"key":"e_1_3_2_58_2","volume-title":"Proceedings of the International Conference on Architectural Support for Programming Languages and Operation Systems (ASPLOS\u201921)","author":"Zhang G.","year":"2021","unstructured":"G. Zhang, N. Attaluri, J. Emer, and D. Sanchez. 2021. GAMMA: Exploiting gustavson\u2019s algorithm to accelerate sparse matrix multiplication. In Proceedings of the International Conference on Architectural Support for Programming Languages and Operation Systems (ASPLOS\u201921)."},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00030"},{"key":"e_1_3_2_60_2","doi-asserted-by":"crossref","unstructured":"Olivia Hsu Maxwell Strange Ritvik Sharma Jaeyeon Won Kunle Olukotun Joel S. Emer Mark A. Horowitz and Fredrik Kj\u00f8lstad. 2023. The sparse abstract machine. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems Volume 3 (Vancouver BC Canada) (ASPLOS\u201923). Association for Computing Machinery New York NY USA 710\u2013726.","DOI":"10.1145\/3582016.3582051"},{"key":"e_1_3_2_61_2","doi-asserted-by":"crossref","unstructured":"Alex Carsello Kathleen Feng Taeyoung Kong Kalhan Koul Qiaoyi Liu Jackson Melchert Gedeon Nyengele Maxwell Strange Keyi Zhang Ankita Nayak Jeff Setter James Thomas Kavya Sreedhar Po-Han Chen Nikhil Bhagdikar Zachary Myers Brandon D\u2019Agostino Pranil Joshi Stephen Richardson Rick Bahr Christopher Torng Mark Horowitz and Priyanka Raina. 2022. Amber: A 367 GOPS 538 GOPS\/W 16nm SoC with a coarse-grained reconfigurable array for flexible acceleration of dense linear algebra. In 2022 IEEE Symposium on VLSI Technology and Circuits. 70\u201371.","DOI":"10.1109\/VLSITechnologyandCir46769.2022.9830509"}],"container-title":["ACM Transactions on Computer Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3630007","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3630007","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T23:57:00Z","timestamp":1750291020000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3630007"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,30]]},"references-count":60,"journal-issue":{"issue":"1-4","published-print":{"date-parts":[[2023,11,30]]}},"alternative-id":["10.1145\/3630007"],"URL":"https:\/\/doi.org\/10.1145\/3630007","relation":{},"ISSN":["0734-2071","1557-7333"],"issn-type":[{"value":"0734-2071","type":"print"},{"value":"1557-7333","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,30]]},"assertion":[{"value":"2022-12-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-18","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-12-18","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}