{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T16:22:39Z","timestamp":1781972559425,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":52,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T00:00:00Z","timestamp":1674777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"U.S.Department of Energy, Office of Science, Office of Advanced Scientific Computing Research","award":["DE-SC0008923 and DE-SC0018121"],"award-info":[{"award-number":["DE-SC0008923 and DE-SC0018121"]}]},{"name":"DARPA","award":["HR0011-18-3-0007 and HR0011-20-9-0017"],"award-info":[{"award-number":["HR0011-18-3-0007 and HR0011-20-9-0017"]}]},{"DOI":"10.13039\/100000001","name":"NSF (National Science Foundation)","doi-asserted-by":"publisher","award":["CCF-2107244"],"award-info":[{"award-number":["CCF-2107244"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,1,27]]},"DOI":"10.1145\/3575693.3575742","type":"proceedings-article","created":{"date-parts":[[2023,1,30]],"date-time":"2023-01-30T22:56:55Z","timestamp":1675119415000},"page":"920-934","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":29,"title":["WACO: Learning Workload-Aware Co-optimization of the Format and Schedule of a Sparse Tensor Program"],"prefix":"10.1145","author":[{"given":"Jaeyeon","family":"Won","sequence":"first","affiliation":[{"name":"Massachusetts Institute of Technology, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Charith","family":"Mendis","sequence":"additional","affiliation":[{"name":"University of Illinois at Urbana-Champaign, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Joel S.","family":"Emer","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology, USA \/ NVIDIA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Saman","family":"Amarasinghe","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,1,30]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3322967"},{"key":"e_1_3_2_1_2_1","volume-title":"Vol 3","author":"Aln\u00e6s Martin","year":"2015","unstructured":"Martin Aln\u00e6s , Jan Blechta , Johan Hake , August Johansson , Benjamin Kehlet , Anders Logg , Chris Richardson , Johannes Ring , Marie E Rognes , and Garth N Wells . 2015. The FEniCS Project Version 1.5. Archive of Numerical Software , Vol 3 ( 2015 ). Martin Aln\u00e6s, Jan Blechta, Johan Hake, August Johansson, Benjamin Kehlet, Anders Logg, Chris Richardson, Johannes Ring, Marie E Rognes, and Garth N Wells. 2015. The FEniCS Project Version 1.5. Archive of Numerical Software, Vol 3 (2015)."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2628071.2628092"},{"key":"e_1_3_2_1_4_1","first-page":"3","volume-title":"Proceedings of Machine Learning and Systems","author":"Baghdadi Riyadh","year":"2021","unstructured":"Riyadh Baghdadi , Massinissa Merouani , Mohamed-Hicham Leghettas , Kamel Abdous , Taha Arbaoui , and Karima Benatchba . 2021 . A Deep Learning Based Cost Model for Automatic Code Optimization . Proceedings of Machine Learning and Systems , 3 (2021). Riyadh Baghdadi, Massinissa Merouani, Mohamed-Hicham Leghettas, Kamel Abdous, Taha Arbaoui, and Karima Benatchba. 2021. A Deep Learning Based Cost Model for Automatic Code Optimization. Proceedings of Machine Learning and Systems, 3 (2021)."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/3314872.3314896"},{"key":"e_1_3_2_1_6_1","volume-title":"International conference on machine learning. 115\u2013123","author":"Bergstra James","year":"2013","unstructured":"James Bergstra , Daniel Yamins , and David Cox . 2013 . Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures . In International conference on machine learning. 115\u2013123 . James Bergstra, Daniel Yamins, and David Cox. 2013. Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures. In International conference on machine learning. 115\u2013123."},{"key":"e_1_3_2_1_7_1","volume-title":"International Conference on Machine Learning. 874\u2013883","author":"Bianchi Filippo Maria","year":"2020","unstructured":"Filippo Maria Bianchi , Daniele Grattarola , and Cesare Alippi . 2020 . Spectral clustering with graph neural networks for graph pooling . In International Conference on Machine Learning. 874\u2013883 . Filippo Maria Bianchi, Daniele Grattarola, and Cesare Alippi. 2020. Spectral clustering with graph neural networks for graph pooling. In International Conference on Machine Learning. 874\u2013883."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1102351.1102363"},{"key":"e_1_3_2_1_9_1","volume-title":"Proceedings of the 13th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201918)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen , Thierry Moreau , Ziheng Jiang , Lianmin Zheng , Eddie Yan , Meghan Cowan , Haichen Shen , Leyuan Wang , Yuwei Hu , Luis Ceze , Carlos Guestrin , and Arvind Krishnamurthy . 2018 . TVM: An Automated End-to-End Optimizing Compiler for Deep Learning . In Proceedings of the 13th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201918) . USENIX Association, USA. 579\u2013594. isbn:978 1931971478 Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In Proceedings of the 13th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201918). USENIX Association, USA. 579\u2013594. isbn:9781931971478"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/3327144.3327258"},{"key":"e_1_3_2_1_11_1","volume-title":"Model-driven autotuning of sparse matrix-vector multiply on GPUs. ACM sigplan notices, 45, 5","author":"Choi Jee W","year":"2010","unstructured":"Jee W Choi , Amik Singh , and Richard W Vuduc . 2010. Model-driven autotuning of sparse matrix-vector multiply on GPUs. ACM sigplan notices, 45, 5 ( 2010 ), 115\u2013126. Jee W Choi, Amik Singh, and Richard W Vuduc. 2010. Model-driven autotuning of sparse matrix-vector multiply on GPUs. ACM sigplan notices, 45, 5 (2010), 115\u2013126."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3276493"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00319"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2049662.2049670"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1998.681704"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/3433701.3433723"},{"key":"e_1_3_2_1_17_1","unstructured":"Benjamin Graham and Laurens van der Maaten. 2017. Submanifold sparse convolutional networks. arXiv preprint arXiv:1706.01307. \t\t\t\t  Benjamin Graham and Laurens van der Maaten. 2017. Submanifold sparse convolutional networks. arXiv preprint arXiv:1706.01307."},{"key":"e_1_3_2_1_18_1","volume-title":"Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS","author":"Hegde Kartik","year":"2021","unstructured":"Kartik Hegde , Po-An Tsai , Sitao Huang , Vikas Chandra , Angshuman Parashar , and Christopher W. Fletcher . 2021. Mind Mappings: Enabling Efficient Algorithm-Accelerator Mapping Space Search . In Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS 2021 ). 16 pages. Kartik Hegde, Po-An Tsai, Sitao Huang, Vikas Chandra, Angshuman Parashar, and Christopher W. Fletcher. 2021. Mind Mappings: Enabling Efficient Algorithm-Accelerator Mapping Space Search. In Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS 2021). 16 pages."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3293883.3295712"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00076"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00050"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359630"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC.2016.7761646"},{"key":"e_1_3_2_1_24_1","volume-title":"Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.","author":"Diederik","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015 . Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings. Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133901"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2866569"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2491956.2462181"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2019.2909204"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2464996.2465013"},{"key":"e_1_3_2_1_30_1","volume-title":"Boman","author":"Loe Jennifer A.","year":"2019","unstructured":"Jennifer A. Loe , Heidi K. Thornquist , and Erik G . Boman . 2019 . Polynomial Preconditioned GMRES to Reduce Communication in Parallel Computing . https:\/\/doi.org\/10.48550\/ARXIV.1907.00072 10.48550\/ARXIV.1907.00072 Jennifer A. Loe, Heidi K. Thornquist, and Erik G. Boman. 2019. Polynomial Preconditioned GMRES to Reduce Communication in Parallel Computing. https:\/\/doi.org\/10.48550\/ARXIV.1907.00072"},{"key":"e_1_3_2_1_31_1","volume-title":"Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs","author":"Malkov Yu A","year":"2018","unstructured":"Yu A Malkov and Dmitry A Yashunin . 2018. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs . IEEE transactions on pattern analysis and machine intelligence, 42, 4 ( 2018 ), 824\u2013836. Yu A Malkov and Dmitry A Yashunin. 2018. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE transactions on pattern analysis and machine intelligence, 42, 4 (2018), 824\u2013836."},{"key":"e_1_3_2_1_32_1","volume-title":"Learning Sparse Matrix Row Permutations for Efficient SpMM on GPU Architectures. In IEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2021","author":"Mehrabi Atefeh","year":"2021","unstructured":"Atefeh Mehrabi , Donghyuk Lee , Niladrish Chatterjee , Daniel J. Sorin , Benjamin C. Lee , and Mike O\u2019Connor . 2021 . Learning Sparse Matrix Row Permutations for Efficient SpMM on GPU Architectures. In IEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2021 , Stony Brook, NY, USA , March 28-30, 2021. IEEE, 48\u201358. Atefeh Mehrabi, Donghyuk Lee, Niladrish Chatterjee, Daniel J. Sorin, Benjamin C. Lee, and Mike O\u2019Connor. 2021. Learning Sparse Matrix Row Permutations for Efficient SpMM on GPU Architectures. In IEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2021, Stony Brook, NY, USA, March 28-30, 2021. IEEE, 48\u201358."},{"key":"e_1_3_2_1_33_1","volume-title":"Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research","author":"Mendis Charith","year":"2019","unstructured":"Charith Mendis , Alex Renda , Dr. Saman Amarasinghe , and Michael Carbin . 2019 . Ithemal: Accurate, Portable and Fast Basic Block Throughput Estimation using Deep Neural Networks . In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research , Vol. 97). PMLR. Charith Mendis, Alex Renda, Dr.Saman Amarasinghe, and Michael Carbin. 2019. Ithemal: Accurate, Portable and Fast Basic Block Throughput Estimation using Deep Neural Networks. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97). PMLR."},{"key":"e_1_3_2_1_34_1","unstructured":"Intel MKL. 2022. Inspector-executor Sparse BLAS Routines. https:\/\/www.intel.com\/content\/www\/us\/en\/develop\/documentation\/onemkl-developer-reference-c\/top\/blas-and-sparse-blas-routines\/inspector-executor-sparse-blas-routines.html \t\t\t\t  Intel MKL. 2022. Inspector-executor Sparse BLAS Routines. https:\/\/www.intel.com\/content\/www\/us\/en\/develop\/documentation\/onemkl-developer-reference-c\/top\/blas-and-sparse-blas-routines\/inspector-executor-sparse-blas-routines.html"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925952"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.615448"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00042"},{"key":"e_1_3_2_1_38_1","volume-title":"Learned TPU Cost Model for XLA Tensor Programs. In Workshop on ML for Systems at NeurIPS.","author":"Phothilimthana Mangpo","unstructured":"Mangpo Phothilimthana , Mike Burrows , and Samuel J. Kaufman . 2019 . Learned TPU Cost Model for XLA Tensor Programs. In Workshop on ML for Systems at NeurIPS. Mangpo Phothilimthana, Mike Burrows, and Samuel J. Kaufman. 2019. Learned TPU Cost Model for XLA Tensor Programs. In Workshop on ML for Systems at NeurIPS."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2499370.2462176"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2751205.2751244"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3428226"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00022"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.2200\/S01004ED1V01Y202004CAC050"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3336191.3371830"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2737924.2738003"},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/16\/1\/071"},{"key":"e_1_3_2_1_47_1","volume-title":"SC\u201998: Proceedings of the 1998 ACM\/IEEE conference on Supercomputing. 38\u201338","author":"Clinton Whaley R","year":"1998","unstructured":"R Clinton Whaley and Jack J Dongarra . 1998 . Automatically tuned linear algebra software . In SC\u201998: Proceedings of the 1998 ACM\/IEEE conference on Supercomputing. 38\u201338 . R Clinton Whaley and Jack J Dongarra. 1998. Automatically tuned linear algebra software. In SC\u201998: Proceedings of the 1998 ACM\/IEEE conference on Supercomputing. 38\u201338."},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178487.3178495"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2018.00104"},{"key":"e_1_3_2_1_50_1","volume-title":"14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)","author":"Zheng Lianmin","year":"2020","unstructured":"Lianmin Zheng , Chengfan Jia , Minmin Sun , Zhao Wu , Cody Hao Yu , Ameer Haj-Ali , Yida Wang , Jun Yang , Danyang Zhuo , and Koushik Sen . 2020 . Ansor: Generating high-performance tensor programs for deep learning . In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 863\u2013879. Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, and Koushik Sen. 2020. Ansor: Generating high-performance tensor programs for deep learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). 863\u2013879."},{"key":"e_1_3_2_1_51_1","volume-title":"FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous System","author":"Zheng Size","unstructured":"Size Zheng , Yun Liang , Shuo Wang , Renze Chen , and Kaiwen Sheng . 2020. FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous System . Association for Computing Machinery , New York, NY, USA . 859\u2013873. isbn:9781450371025 Size Zheng, Yun Liang, Shuo Wang, Renze Chen, and Kaiwen Sheng. 2020. FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous System. Association for Computing Machinery, New York, NY, USA. 859\u2013873. isbn:9781450371025"},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2932931"}],"event":{"name":"ASPLOS '23: 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2","location":"Vancouver BC Canada","acronym":"ASPLOS '23","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture","SIGOPS ACM Special Interest Group on Operating Systems","SIGPLAN ACM Special Interest Group on Programming Languages"]},"container-title":["Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3575693.3575742","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3575693.3575742","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3575693.3575742","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:51:20Z","timestamp":1750182680000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3575693.3575742"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,27]]},"references-count":52,"alternative-id":["10.1145\/3575693.3575742","10.1145\/3575693"],"URL":"https:\/\/doi.org\/10.1145\/3575693.3575742","relation":{},"subject":[],"published":{"date-parts":[[2023,1,27]]},"assertion":[{"value":"2023-01-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}