{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,8]],"date-time":"2026-01-08T05:04:47Z","timestamp":1767848687730,"version":"3.49.0"},"publisher-location":"New York, NY, USA","reference-count":53,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T00:00:00Z","timestamp":1674777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,1,27]]},"DOI":"10.1145\/3575693.3575694","type":"proceedings-article","created":{"date-parts":[[2023,1,30]],"date-time":"2023-01-30T22:56:55Z","timestamp":1675119415000},"page":"718-732","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Sigma: Compiling Einstein Summations to Locality-Aware Dataflow"],"prefix":"10.1145","author":[{"given":"Tian","family":"Zhao","sequence":"first","affiliation":[{"name":"Stanford University, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alexander","family":"Rucker","sequence":"additional","affiliation":[{"name":"Stanford University, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kunle","family":"Olukotun","sequence":"additional","affiliation":[{"name":"Stanford University, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,1,30]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2022. Intel: High-level Synthesis Compiler. https:\/\/www.intel.com\/content\/www\/us\/en\/software\/programmable\/quartus-prime\/hls-compiler.html \t\t\t\t  2022. Intel: High-level Synthesis Compiler. https:\/\/www.intel.com\/content\/www\/us\/en\/software\/programmable\/quartus-prime\/hls-compiler.html"},{"key":"e_1_3_2_1_2_1","unstructured":"2022. Xilinx: Vivado High-level Synthesis Compiler. https:\/\/www.xilinx.com\/support\/documentation-navigation\/design-hubs\/dh0012-vivado-high-level-synthesis-hub.html \t\t\t\t  2022. Xilinx: Vivado High-level Synthesis Compiler. https:\/\/www.xilinx.com\/support\/documentation-navigation\/design-hubs\/dh0012-vivado-high-level-synthesis-hub.html"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00023"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICEngTechnol.2017.8308186"},{"key":"e_1_3_2_1_5_1","unstructured":"Igor Babuschkin Kate Baumli Alison Bell Surya Bhupatiraju Jake Bruce Peter Buchlovsky David Budden Trevor Cai Aidan Clark Ivo Danihelka Claudio Fantacci Jonathan Godwin Chris Jones Tom Hennigan Matteo Hessel Steven Kapturowski Thomas Keck Iurii Kemaev Michael King Lena Martens Vladimir Mikulik Tamara Norman John Quan George Papamakarios Roman Ring Francisco Ruiz Alvaro Sanchez Rosalia Schneider Eren Sezener Stephen Spencer Srivatsan Srinivasan Wojciech Stokowiec and Fabio Viola. 2020. The DeepMind JAX Ecosystem. http:\/\/github.com\/deepmind \t\t\t\t  Igor Babuschkin Kate Baumli Alison Bell Surya Bhupatiraju Jake Bruce Peter Buchlovsky David Budden Trevor Cai Aidan Clark Ivo Danihelka Claudio Fantacci Jonathan Godwin Chris Jones Tom Hennigan Matteo Hessel Steven Kapturowski Thomas Keck Iurii Kemaev Michael King Lena Martens Vladimir Mikulik Tamara Norman John Quan George Papamakarios Roman Ring Francisco Ruiz Alvaro Sanchez Rosalia Schneider Eren Sezener Stephen Spencer Srivatsan Srinivasan Wojciech Stokowiec and Fabio Viola. 2020. The DeepMind JAX Ecosystem. http:\/\/github.com\/deepmind"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3485137"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_3_2_1_9_1","volume-title":"13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen , Thierry Moreau , Ziheng Jiang , Lianmin Zheng , Eddie Yan , Haichen Shen , Meghan Cowan , Leyuan Wang , Yuwei Hu , and Luis Ceze . 2018 . TVM: An automated end-to-end optimizing compiler for deep learning . In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . 578\u2013594. https:\/\/doi.org\/10.48550\/arXiv.1802.04799 10.48550\/arXiv.1802.04799 Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, and Luis Ceze. 2018. TVM: An automated end-to-end optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). 578\u2013594. https:\/\/doi.org\/10.48550\/arXiv.1802.04799"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2910232"},{"key":"#cr-split#-e_1_3_2_1_11_1.1","unstructured":"Sharan Chetlur Cliff Woolley Philippe Vandermersch Jonathan Cohen John Tran Bryan Catanzaro and Evan Shelhamer. 2014. cuDNN: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759 https:\/\/doi.org\/10.48550\/arXiv.1410.0759 10.48550\/arXiv.1410.0759"},{"key":"#cr-split#-e_1_3_2_1_11_1.2","unstructured":"Sharan Chetlur Cliff Woolley Philippe Vandermersch Jonathan Cohen John Tran Bryan Catanzaro and Evan Shelhamer. 2014. cuDNN: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759 https:\/\/doi.org\/10.48550\/arXiv.1410.0759"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2018.022071131"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00053"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3358198"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/MWSCAS.2017.8053243"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/MCSE.2021.3057203"},{"key":"#cr-split#-e_1_3_2_1_17_1.1","unstructured":"Matthew Feldman Tian Zhao and Kunle Olukotun. 2022. Efficient Memory Partitioning in Software Defined Hardware. arXiv preprint arXiv:2202.01261 https:\/\/doi.org\/10.48550\/arXiv.2202.01261 10.48550\/arXiv.2202.01261"},{"key":"#cr-split#-e_1_3_2_1_17_1.2","unstructured":"Matthew Feldman Tian Zhao and Kunle Olukotun. 2022. Efficient Memory Partitioning in Software Defined Hardware. arXiv preprint arXiv:2202.01261 https:\/\/doi.org\/10.48550\/arXiv.2202.01261"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3485505"},{"key":"e_1_3_2_1_20_1","volume-title":"Stardust: Compiling Sparse Tensor Algebra to a Reconfigurable Dataflow Architecture. Arxiv Preprint, November, https:\/\/doi.org\/10.48550\/arXiv.2211.03251","author":"Hsu Olivia","year":"2022","unstructured":"Olivia Hsu , Alexander Rucker , Tian Zhao , Kunle Olukotun , and Fredrik Kjolstad . 2022 . Stardust: Compiling Sparse Tensor Algebra to a Reconfigurable Dataflow Architecture. Arxiv Preprint, November, https:\/\/doi.org\/10.48550\/arXiv.2211.03251 10.48550\/arXiv.2211.03251 Olivia Hsu, Alexander Rucker, Tian Zhao, Kunle Olukotun, and Fredrik Kjolstad. 2022. Stardust: Compiling Sparse Tensor Algebra to a Reconfigurable Dataflow Architecture. Arxiv Preprint, November, https:\/\/doi.org\/10.48550\/arXiv.2211.03251"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00010"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133901"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3192366.3192379"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/263699.263719"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.2307\/2327242"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-10549-5_35"},{"key":"e_1_3_2_1_27_1","volume-title":"Fitzek","author":"P\u00e9ter Vingelmann NVIDIA","year":"2020","unstructured":"NVIDIA , P\u00e9ter Vingelmann , and Frank H.P . Fitzek . 2020 . CUDA , release: 10.2.89. https:\/\/developer.nvidia.com\/cuda-toolkit NVIDIA, P\u00e9ter Vingelmann, and Frank H.P. Fitzek. 2020. CUDA, release: 10.2.89. https:\/\/developer.nvidia.com\/cuda-toolkit"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00042"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080254"},{"key":"e_1_3_2_1_30_1","volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , Alban Desmaison , Andreas Kopf , Edward Yang , Zachary DeVito , Martin Raison , Alykhan Tejani , Sasank Chilamkurthy , Benoit Steiner , Lu Fang , Junjie Bai , and Soumith Chintala . 2019. PyTorch: An Imperative Style , High-Performance Deep Learning Library . In Advances in Neural Information Processing Systems 32, H. Wallach, H. Larochelle, A. Beygelzimer, F. d' Alch\u00e9-Buc, E. Fox, and R. Garnett (Eds.). Curran Associates, Inc., 8024\u20138035. http:\/\/papers.neurips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32, H. Wallach, H. Larochelle, A. Beygelzimer, F. d' Alch\u00e9-Buc, E. Fox, and R. Garnett (Eds.). Curran Associates, Inc., 8024\u20138035. http:\/\/papers.neurips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080256"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/2491956.2462176"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480047"},{"key":"e_1_3_2_1_34_1","volume-title":"XLA : Compiling Machine Learning for Peak Performance.","author":"Sabne Amit","year":"2020","unstructured":"Amit Sabne . 2020 . XLA : Compiling Machine Learning for Peak Performance. Amit Sabne. 2020. XLA : Compiling Machine Learning for Peak Performance."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3470496.3533040"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.physd.2019.132306"},{"key":"#cr-split#-e_1_3_2_1_37_1.1","unstructured":"Nicolas Vasilache Oleksandr Zinenko Theodoros Theodoridis Priya Goyal Zachary DeVito William S Moses Sven Verdoolaege Andrew Adams and Albert Cohen. 2018. Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions. arXiv preprint arXiv:1802.04730 https:\/\/doi.org\/10.48550\/arXiv.1802.04730 10.48550\/arXiv.1802.04730"},{"key":"#cr-split#-e_1_3_2_1_37_1.2","unstructured":"Nicolas Vasilache Oleksandr Zinenko Theodoros Theodoridis Priya Goyal Zachary DeVito William S Moses Sven Verdoolaege Andrew Adams and Albert Cohen. 2018. Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions. arXiv preprint arXiv:1802.04730 https:\/\/doi.org\/10.48550\/arXiv.1802.04730"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-15582-6_49"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00039"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463209.2488748"},{"key":"#cr-split#-e_1_3_2_1_41_1.1","unstructured":"Yu Emma Wang Gu-Yeon Wei and David Brooks. 2019. Benchmarking TPU GPU and CPU platforms for deep learning. arXiv preprint arXiv:1907.10701 https:\/\/doi.org\/10.48550\/arXiv.1907.10701 10.48550\/arXiv.1907.10701"},{"key":"#cr-split#-e_1_3_2_1_41_1.2","unstructured":"Yu Emma Wang Gu-Yeon Wei and David Brooks. 2019. Benchmarking TPU GPU and CPU platforms for deep learning. arXiv preprint arXiv:1907.10701 https:\/\/doi.org\/10.48550\/arXiv.1907.10701"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378514"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco_a_01199"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689060"},{"key":"#cr-split#-e_1_3_2_1_46_1.1","unstructured":"Shiwen Zhang Sheng Guo Weilin Huang Matthew R Scott and Limin Wang. 2020. V4d: 4d convolutional neural networks for video-level representation learning. arXiv preprint arXiv:2002.07442 https:\/\/doi.org\/10.48550\/arXiv.2002.07442 10.48550\/arXiv.2002.07442"},{"key":"#cr-split#-e_1_3_2_1_46_1.2","unstructured":"Shiwen Zhang Sheng Guo Weilin Huang Matthew R Scott and Limin Wang. 2020. V4d: 4d convolutional neural networks for video-level representation learning. arXiv preprint arXiv:2002.07442 https:\/\/doi.org\/10.48550\/arXiv.2002.07442"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00085"},{"key":"e_1_3_2_1_48_1","first-page":"166","article-title":"Serving recurrent neural networks efficiently with a spatial accelerator","volume":"1","author":"Zhao Tian","year":"2019","unstructured":"Tian Zhao , Yaqi Zhang , and Kunle Olukotun . 2019 . Serving recurrent neural networks efficiently with a spatial accelerator . Proceedings of Machine Learning and Systems , 1 (2019), 166 \u2013 177 . https:\/\/doi.org\/10.48550\/arXiv.1909.13654 10.48550\/arXiv.1909.13654 Tian Zhao, Yaqi Zhang, and Kunle Olukotun. 2019. Serving recurrent neural networks efficiently with a spatial accelerator. Proceedings of Machine Learning and Systems, 1 (2019), 166\u2013177. https:\/\/doi.org\/10.48550\/arXiv.1909.13654","journal-title":"Proceedings of Machine Learning and Systems"}],"event":{"name":"ASPLOS '23: 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2","location":"Vancouver BC Canada","acronym":"ASPLOS '23","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture","SIGOPS ACM Special Interest Group on Operating Systems","SIGPLAN ACM Special Interest Group on Programming Languages"]},"container-title":["Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3575693.3575694","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3575693.3575694","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T18:43:52Z","timestamp":1750272232000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3575693.3575694"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,27]]},"references-count":53,"alternative-id":["10.1145\/3575693.3575694","10.1145\/3575693"],"URL":"https:\/\/doi.org\/10.1145\/3575693.3575694","relation":{},"subject":[],"published":{"date-parts":[[2023,1,27]]},"assertion":[{"value":"2023-01-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}