{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T23:40:35Z","timestamp":1783035635897,"version":"3.54.6"},"publisher-location":"New York, NY, USA","reference-count":40,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T00:00:00Z","timestamp":1674777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["No. 2018AAA0100500"],"award-info":[{"award-number":["No. 2018AAA0100500"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No. 62272434"],"award-info":[{"award-number":["No. 62272434"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Frontier Scientific Research Program of the Chinese Academy of Sciences","award":["No. ZDBS-LY-JSC001"],"award-info":[{"award-number":["No. ZDBS-LY-JSC001"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,1,27]]},"DOI":"10.1145\/3575693.3575737","type":"proceedings-article","created":{"date-parts":[[2023,1,30]],"date-time":"2023-01-30T22:56:55Z","timestamp":1675119415000},"page":"833-845","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":41,"title":["TLP: A Deep Learning-Based Cost Model for Tensor Program Tuning"],"prefix":"10.1145","author":[{"given":"Yi","family":"Zhai","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuo","family":"Liu","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaomeng","family":"Chu","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jie","family":"Peng","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianmin","family":"Ji","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanyong","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,1,30]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3322967"},{"key":"e_1_3_2_1_2_1","volume-title":"Chameleon: Adaptive code optimization for expedited deep neural network compilation. arXiv preprint arXiv:2001.08743.","author":"Ahn Byung Hoon","year":"2020","unstructured":"Byung Hoon Ahn , Prannoy Pilligundla , Amir Yazdanbakhsh , and Hadi Esmaeilzadeh . 2020 . Chameleon: Adaptive code optimization for expedited deep neural network compilation. arXiv preprint arXiv:2001.08743. Byung Hoon Ahn, Prannoy Pilligundla, Amir Yazdanbakhsh, and Hadi Esmaeilzadeh. 2020. Chameleon: Adaptive code optimization for expedited deep neural network compilation. arXiv preprint arXiv:2001.08743."},{"key":"e_1_3_2_1_3_1","unstructured":"Luke Anderson Andrew Adams Karima Ma Tzu-Mao Li and Jonathan Ragan-Kelley. 2020. Learning to schedule halide pipelines for the gpu. arXiv preprint arXiv:2012.07145. \t\t\t\t  Luke Anderson Andrew Adams Karima Ma Tzu-Mao Li and Jonathan Ragan-Kelley. 2020. Learning to schedule halide pipelines for the gpu. arXiv preprint arXiv:2012.07145."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2015.17"},{"key":"e_1_3_2_1_5_1","first-page":"181","article-title":"A deep learning based cost model for automatic code optimization","volume":"3","author":"Baghdadi Riyadh","year":"2021","unstructured":"Riyadh Baghdadi , Massinissa Merouani , Mohamed-Hicham Leghettas , Kamel Abdous , Taha Arbaoui , and Karima Benatchba . 2021 . A deep learning based cost model for automatic code optimization . Proceedings of Machine Learning and Systems , 3 (2021), 181 \u2013 193 . Riyadh Baghdadi, Massinissa Merouani, Mohamed-Hicham Leghettas, Kamel Abdous, Taha Arbaoui, and Karima Benatchba. 2021. A deep learning based cost model for automatic code optimization. Proceedings of Machine Learning and Systems, 3 (2021), 181\u2013193.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.5555\/3314872.3314896"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1375581.1375595"},{"key":"e_1_3_2_1_8_1","volume-title":"Language models are few-shot learners. Advances in neural information processing systems, 33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown , Benjamin Mann , Nick Ryder , Melanie Subbiah , Jared D Kaplan , Prafulla Dhariwal , Arvind Neelakantan , Pranav Shyam , Girish Sastry , and Amanda Askell . 2020. Language models are few-shot learners. Advances in neural information processing systems, 33 ( 2020 ), 1877\u20131901. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, and Amanda Askell. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33 (2020), 1877\u20131901."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1273496.1273513"},{"key":"e_1_3_2_1_10_1","volume-title":"13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen , Thierry Moreau , Ziheng Jiang , Lianmin Zheng , Eddie Yan , Haichen Shen , Meghan Cowan , Leyuan Wang , Yuwei Hu , and Luis Ceze . 2018 . $TVM$: An Automated $End-to-End$ Optimizing Compiler for Deep Learning . In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . 578\u2013594. Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, and Luis Ceze. 2018. $TVM$: An Automated $End-to-End$ Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). 578\u2013594."},{"key":"e_1_3_2_1_11_1","volume-title":"Learning to optimize tensor programs. Advances in Neural Information Processing Systems, 31","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen , Lianmin Zheng , Eddie Yan , Ziheng Jiang , Thierry Moreau , Luis Ceze , Carlos Guestrin , and Arvind Krishnamurthy . 2018. Learning to optimize tensor programs. Advances in Neural Information Processing Systems, 31 ( 2018 ). Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. Learning to optimize tensor programs. Advances in Neural Information Processing Systems, 31 (2018)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2017.08.035"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2019.00164"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2017.24"},{"key":"e_1_3_2_1_15_1","unstructured":"Scott Cyphers Arjun K Bansal Anahita Bhiwandiwalla Jayaram Bobba Matthew Brookhart Avijit Chakraborty Will Constable Christian Convey Leona Cook and Omar Kanawi. 2018. Intel ngraph: An intermediate representation compiler and executor for deep learning. arXiv preprint arXiv:1801.08058. \t\t\t\t  Scott Cyphers Arjun K Bansal Anahita Bhiwandiwalla Jayaram Bobba Matthew Brookhart Avijit Chakraborty Will Constable Christian Convey Leona Cook and Omar Kanawi. 2018. Intel ngraph: An intermediate representation compiler and executor for deep learning. arXiv preprint arXiv:1801.08058."},{"key":"e_1_3_2_1_16_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1166"},{"key":"e_1_3_2_1_18_1","unstructured":"Ameer Haj-Ali Hasan Genc Qijing Huang William Moses John Wawrzynek Krste Asanovi\u0107 and Ion Stoica. 2020. Protuner: tuning programs with monte carlo tree search. arXiv preprint arXiv:2005.13685. \t\t\t\t  Ameer Haj-Ali Hasan Genc Qijing Huang William Moses John Wawrzynek Krste Asanovi\u0107 and Ion Stoica. 2020. Protuner: tuning programs with monte carlo tree search. arXiv preprint arXiv:2005.13685."},{"key":"e_1_3_2_1_19_1","volume-title":"Tinybert: Distilling bert for natural language understanding. arXiv preprint arXiv:1909.10351.","author":"Jiao Xiaoqi","year":"2019","unstructured":"Xiaoqi Jiao , Yichun Yin , Lifeng Shang , Xin Jiang , Xiao Chen , Linlin Li , Fang Wang , and Qun Liu . 2019 . Tinybert: Distilling bert for natural language understanding. arXiv preprint arXiv:1909.10351. Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019. Tinybert: Distilling bert for natural language understanding. arXiv preprint arXiv:1909.10351."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3453483.3454038"},{"key":"e_1_3_2_1_21_1","first-page":"387","article-title":"A learned performance model for tensor processing units","volume":"3","author":"Kaufman Sam","year":"2021","unstructured":"Sam Kaufman , Phitchaya Phothilimthana , Yanqi Zhou , Charith Mendis , Sudip Roy , Amit Sabne , and Mike Burrows . 2021 . A learned performance model for tensor processing units . Proceedings of Machine Learning and Systems , 3 (2021), 387 \u2013 400 . Sam Kaufman, Phitchaya Phothilimthana, Yanqi Zhou, Charith Mendis, Sudip Roy, Amit Sabne, and Mike Burrows. 2021. A learned performance model for tensor processing units. Proceedings of Machine Learning and Systems, 3 (2021), 387\u2013400.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_2_1_22_1","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition. 7482\u20137491","author":"Kendall Alex","year":"2018","unstructured":"Alex Kendall , Yarin Gal , and Roberto Cipolla . 2018 . Multi-task learning using uncertainty to weigh losses for scene geometry and semantics . In Proceedings of the IEEE conference on computer vision and pattern recognition. 7482\u20137491 . Alex Kendall, Yarin Gal, and Roberto Cipolla. 2018. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE conference on computer vision and pattern recognition. 7482\u20137491."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133901"},{"key":"e_1_3_2_1_24_1","unstructured":"Pengfei Liu Xipeng Qiu and Xuanjing Huang. 2016. Recurrent neural network for text classification with multi-task learning. arXiv preprint arXiv:1605.05101. \t\t\t\t  Pengfei Liu Xipeng Qiu and Xuanjing Huang. 2016. Recurrent neural network for text classification with multi-task learning. arXiv preprint arXiv:1605.05101."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925952"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"crossref","unstructured":"Jan Niehues and Eunah Cho. 2017. Exploiting linguistic resources for neural machine translation using multi-task learning. arXiv preprint arXiv:1708.00993. \t\t\t\t  Jan Niehues and Eunah Cho. 2017. Exploiting linguistic resources for neural machine translation using multi-task learning. arXiv preprint arXiv:1708.00993.","DOI":"10.18653\/v1\/W17-4708"},{"key":"e_1_3_2_1_27_1","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever. 2018. Improving language understanding by generative pre-training. \t\t\t\t  Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever. 2018. Improving language understanding by generative pre-training."},{"key":"e_1_3_2_1_28_1","volume-title":"Language models are unsupervised multitask learners. OpenAI blog, 1, 8","author":"Radford Alec","year":"2019","unstructured":"Alec Radford , Jeffrey Wu , Rewon Child , David Luan , Dario Amodei , and Ilya Sutskever . 2019. Language models are unsupervised multitask learners. OpenAI blog, 1, 8 ( 2019 ), 9. Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1, 8 (2019), 9."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2499370.2462176"},{"key":"e_1_3_2_1_30_1","first-page":"323","article-title":"Value learning for throughput optimization of deep learning workloads","volume":"3","author":"Steiner Benoit","year":"2021","unstructured":"Benoit Steiner , Chris Cummins , Horace He , and Hugh Leather . 2021 . Value learning for throughput optimization of deep learning workloads . Proceedings of Machine Learning and Systems , 3 (2021), 323 \u2013 334 . Benoit Steiner, Chris Cummins, Horace He, and Hugh Leather. 2021. Value learning for throughput optimization of deep learning workloads. Proceedings of Machine Learning and Systems, 3 (2021), 323\u2013334.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_2_1_31_1","volume-title":"d.]. XLA: Optimizing Compiler for TensorFlow. https:\/\/www.tensorflow.org\/xla. [Online","year":"2019","unstructured":"TensorFlow. [n. d.]. XLA: Optimizing Compiler for TensorFlow. https:\/\/www.tensorflow.org\/xla. [Online ; accessed 19- September - 2019 ] TensorFlow. [n. d.]. XLA: Optimizing Compiler for TensorFlow. https:\/\/www.tensorflow.org\/xla. [Online; accessed 19-September-2019]"},{"key":"e_1_3_2_1_32_1","unstructured":"Nicolas Vasilache Oleksandr Zinenko Theodoros Theodoridis Priya Goyal Zachary DeVito William S Moses Sven Verdoolaege Andrew Adams and Albert Cohen. 2018. Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions. arXiv preprint arXiv:1802.04730. \t\t\t\t  Nicolas Vasilache Oleksandr Zinenko Theodoros Theodoridis Priya Goyal Zachary DeVito William S Moses Sven Verdoolaege Andrew Adams and Albert Cohen. 2018. Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions. arXiv preprint arXiv:1802.04730."},{"key":"e_1_3_2_1_33_1","unstructured":"Sven Verdoolaege. 2016. Presburger formulas and polyhedral compilation. \t\t\t\t  Sven Verdoolaege. 2016. Presburger formulas and polyhedral compilation."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2400682.2400713"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3271784"},{"key":"e_1_3_2_1_36_1","first-page":"204","article-title":"Bolt: Bridging the Gap between Auto-tuners and Hardware-native Performance","volume":"4","author":"Xing Jiarong","year":"2022","unstructured":"Jiarong Xing , Leyuan Wang , Shang Zhang , Jack Chen , Ang Chen , and Yibo Zhu . 2022 . Bolt: Bridging the Gap between Auto-tuners and Hardware-native Performance . Proceedings of Machine Learning and Systems , 4 (2022), 204 \u2013 216 . Jiarong Xing, Leyuan Wang, Shang Zhang, Jack Chen, Ang Chen, and Yibo Zhu. 2022. Bolt: Bridging the Gap between Auto-tuners and Hardware-native Performance. Proceedings of Machine Learning and Systems, 4 (2022), 204\u2013216.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_2_1_37_1","volume-title":"Moses: Efficient Exploitation of Cross-device Transferable Features for Tensor Program Optimization. arXiv preprint arXiv:2201.05752.","author":"Zhao Zhihe","year":"2022","unstructured":"Zhihe Zhao , Xian Shuai , Yang Bai , Neiwen Ling , Nan Guan , Zhenyu Yan , and Guoliang Xing . 2022 . Moses: Efficient Exploitation of Cross-device Transferable Features for Tensor Program Optimization. arXiv preprint arXiv:2201.05752. Zhihe Zhao, Xian Shuai, Yang Bai, Neiwen Ling, Nan Guan, Zhenyu Yan, and Guoliang Xing. 2022. Moses: Efficient Exploitation of Cross-device Transferable Features for Tensor Program Optimization. arXiv preprint arXiv:2201.05752."},{"key":"e_1_3_2_1_38_1","volume-title":"14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)","author":"Zheng Lianmin","year":"2020","unstructured":"Lianmin Zheng , Chengfan Jia , Minmin Sun , Zhao Wu , Cody Hao Yu , Ameer Haj-Ali , Yida Wang , Jun Yang , Danyang Zhuo , and Koushik Sen . 2020 . Ansor: Generating $High-Performance$ Tensor Programs for Deep Learning . In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 863\u2013879. Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, and Koushik Sen. 2020. Ansor: Generating $High-Performance$ Tensor Programs for Deep Learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). 863\u2013879."},{"key":"e_1_3_2_1_39_1","volume-title":"TenSet: A Large-scale Program Performance Dataset for Learned Tensor Compilers. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1).","author":"Zheng Lianmin","year":"2021","unstructured":"Lianmin Zheng , Ruochen Liu , Junru Shao , Tianqi Chen , Joseph E Gonzalez , Ion Stoica , and Ameer Haj Ali . 2021 . TenSet: A Large-scale Program Performance Dataset for Learned Tensor Compilers. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1). Lianmin Zheng, Ruochen Liu, Junru Shao, Tianqi Chen, Joseph E Gonzalez, Ion Stoica, and Ameer Haj Ali. 2021. TenSet: A Large-scale Program Performance Dataset for Learned Tensor Compilers. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1)."},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378508"}],"event":{"name":"ASPLOS '23: 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2","location":"Vancouver BC Canada","acronym":"ASPLOS '23","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture","SIGOPS ACM Special Interest Group on Operating Systems","SIGPLAN ACM Special Interest Group on Programming Languages"]},"container-title":["Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3575693.3575737","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3575693.3575737","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:51:20Z","timestamp":1750182680000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3575693.3575737"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,27]]},"references-count":40,"alternative-id":["10.1145\/3575693.3575737","10.1145\/3575693"],"URL":"https:\/\/doi.org\/10.1145\/3575693.3575737","relation":{},"subject":[],"published":{"date-parts":[[2023,1,27]]},"assertion":[{"value":"2023-01-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}