{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T13:47:21Z","timestamp":1782481641047,"version":"3.54.5"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T00:00:00Z","timestamp":1782432000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62472273 and 62232015"],"award-info":[{"award-number":["62472273 and 62232015"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>With the rapid development of AI and deep learning, computational demands are increasing significantly. While GPUs excel in parallel computing, they fall short in terms of energy efficiency, specialization, and processing latency. In contrast, Neural Processing Units (NPUs), such as the Ascend NPUs, designed specifically for deep learning tasks, demonstrate superior performance. However, the architecture specialization makes operator development more challenging, leading to a reliance on manual tuning and optimization, which incurs significant time cost and developing effort. To address this issue, we propose NPUMeter, an automatic operator optimization framework for Ascend NPUs built upon accurate and comprehensive analytical performance models. NPUMeter comprises two components: (1) an analytical performance model that accurately estimates operator latency on NPU given different configurations of optimization parameters; (2) an efficient design space exploration (DSE) algorithm that automatically searches for the optimal parameter configuration in a large design space within minutes. Experimental results demonstrate that NPUMeter achieves high estimation accuracy, with an average error below 5%. It effectively generates near-optimal configurations for various operators, achieving up to a 1.46\u00d7 performance speedup compared to the configuration generated by the Ascend C compiler while reducing the DSE time from hours to minutes.<\/jats:p>","DOI":"10.1145\/3820380","type":"journal-article","created":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T20:50:14Z","timestamp":1781038214000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["NPUMeter: Automatic Operator Optimization for Ascend NPU with Accurate Analytical Performance Models"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-7928-6512","authenticated-orcid":false,"given":"Weichuang","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Computer Science, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-1544-2356","authenticated-orcid":false,"given":"Yufei","family":"Shangguan","sequence":"additional","affiliation":[{"name":"School of Computer Science, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-7111-5574","authenticated-orcid":false,"given":"Haibo","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Computer Science, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0142-6366","authenticated-orcid":false,"given":"Yuting","family":"Mai","sequence":"additional","affiliation":[{"name":"Huawei Technologies Co Ltd","place":["Shenzhen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-9808-1380","authenticated-orcid":false,"given":"Qiuliang","family":"Wang","sequence":"additional","affiliation":[{"name":"Huawei Technologies Co Ltd","place":["Shenzhen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9480-5632","authenticated-orcid":false,"given":"Chen","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Computer Science, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-0748-8452","authenticated-orcid":false,"given":"Yi","family":"Li","sequence":"additional","affiliation":[{"name":"Huawei Technologies Co Ltd","place":["Shenzhen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5832-0347","authenticated-orcid":false,"given":"Quan","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Computer Science, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4249-526X","authenticated-orcid":false,"given":"Wenchao","family":"Ding","sequence":"additional","affiliation":[{"name":"College of Intelligent Robotics and Advanced Manufacturing, Fudan University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8211-2812","authenticated-orcid":false,"given":"Jieru","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Computer Science, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0034-2302","authenticated-orcid":false,"given":"Minyi","family":"Guo","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Guizhou University","place":["Guiyang, China"]},{"name":"School of Computer Science, Shanghai Jiao Tong University","place":["Guiyang, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,26]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"Proceedings of NVIDIA GPU Technology Conference (GTC), Session S9438","author":"Goodwin David","year":"2019","unstructured":"David Goodwin and Soyoung Jeong. 2019. Maximizing utilization for data center inference with TensorRT inference server. In Proceedings of NVIDIA GPU Technology Conference (GTC), Session S9438. Retrieved May 2026 from https:\/\/developer.nvidia.com\/gtc\/2019\/video\/s9438"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.58"},{"key":"e_1_3_2_5_2","volume-title":"Proceedings of the 2020 IEEE International Symposium on High Performance Computer Architecture. IEEE, 220\u2013233","author":"Choi Y.","unstructured":"Y. Choi and M. Rhu. 2020. Prema: A predictive multi-task scheduling algorithm for preemptible neural processing units[C]. In Proceedings of the 2020 IEEE International Symposium on High Performance Computer Architecture. IEEE, 220\u2013233."},{"key":"e_1_3_2_6_2","volume-title":"Proceedings of the 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture. IEEE, 940\u2013953","author":"Baek E.","unstructured":"E. Baek, D. Kwon, and J. Kim. 2020. A multi-neural network acceleration architecture[C]. In Proceedings of the 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture. IEEE, 940\u2013953."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE56975.2023.10137320"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPADS56603.2022.00063"},{"key":"e_1_3_2_9_2","volume-title":"Dissecting the Graphcore IPU Architecture via Microbenchmarking. CoRR abs\/1912.03413","author":"Jia Zhe","year":"2019","unstructured":"Zhe Jia, Blake Tillman, Marco Maggioni, and Daniele Paolo Scarpazza. 2019. Dissecting the Graphcore IPU Architecture via Microbenchmarking. CoRR abs\/1912.03413 (2019). Retrieved from https:\/\/arxiv.org\/abs\/1912.03413"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.5555\/3433701.3433778"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/HOTCHIPS.2019.8875654"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2018.022071131"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3093336.3037700"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/2954679.2872368"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00028"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3134269"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00042"},{"key":"e_1_3_2_18_2","volume-title":"Proceedings of the 2020 IEEE\/ACM International Conference On Computer Aided Design","author":"Kao S.-C.","unstructured":"S.-C. Kao and T. Krishna. 2020. GAMMA: Automating the HW mapping of DNN models on accelerators via genetic algorithm. In Proceedings of the 2020 IEEE\/ACM International Conference On Computer Aided Design. San Diego, CA, USA, 1\u20139."},{"key":"e_1_3_2_19_2","unstructured":"Dan Hendrycks and Kevin Gimpel. 2016. Gaussian Error Linear Units (GELUs). CoRR abs\/1606.08415. Retrieved from https:\/\/arxiv.org\/abs\/1606.08415"},{"key":"e_1_3_2_20_2","unstructured":"Noam Shazeer. 2020. GLU variants improve transformer. CoRR abs\/2002.05202. Retrieved from https:\/\/arxiv.org\/abs\/2002.05202"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3157733"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/292395.292412"},{"key":"e_1_3_2_24_2","unstructured":"Dan Hendrycks and Kevin Gimpel. 2016. Gaussian Error Linear Units (GELUs). CoRR abs\/1606.08415. Retrieved from https:\/\/arxiv.org\/abs\/1606.08415"},{"key":"e_1_3_2_25_2","volume-title":"GLU Variants Improve Transformer. CoRR abs\/2002.05202","author":"Shazeer Noam","year":"2020","unstructured":"Noam Shazeer. 2020. GLU Variants Improve Transformer. CoRR abs\/2002.05202 (2020). https:\/\/arxiv.org\/abs\/2002.05202"},{"key":"e_1_3_2_26_2","unstructured":"Gemma Team Thomas Mesnard Cassidy Hardin Robert Dadashi Surya Bhupatiraju Shreya Pathak Laurent Sifre Morgane Rivi\u00e8re Mihir Sanjay Kale Juliette Love Pouya Tafti L\u00e9onard Hussenot Pier Giuseppe Sessa Aakanksha Chowdhery Adam Roberts Aditya Barua Alex Botev Alex Castro-Ros Ambrose Slone Am\u00e9lie H\u00e9liou Andrea Tacchetti Anna Bulanova Antonia Paterson Beth Tsai Bobak Shahriari Charline Le Lan Christopher A. Choquette-Choo Cl\u00e9ment Crepy Daniel Cer Daphne Ippolito David Reid Elena Buchatskaya Eric Ni Eric Noland Geng Yan George Tucker George-Christian Muraru Grigory Rozhdestvenskiy Henryk Michalewski Ian Tenney Ivan Grishchenko Jacob Austin James Keeling Jane Labanowski Jean-Baptiste Lespiau Jeff Stanway Jenny Brennan Jeremy Chen Johan Ferret Justin Chiu Justin Mao-Jones Katherine Lee Kathy Yu Katie Millican Lars Lowe Sjoesund Lisa Lee Lucas Dixon Machel Reid Maciej Miku\u0142a Mateo Wirth Michael Sharman Nikolai Chinaev Nithum Thain Olivier Bachem Oscar Chang Oscar Wahltinez Paige Bailey Paul Michel Petko Yotov Rahma Chaabouni Ramona Comanescu Reena Jana Rohan Anil Ross McIlroy Ruibo Liu Ryan Mullins Samuel L. Smith Sebastian Borgeaud Sertan Girgin Sholto Douglas Shree Pandya Siamak Shakeri Soham De Ted Klimenko Tom Hennigan Vlad Feinberg Wojciech Stokowiec Yu-hui Chen Zafarali Ahmed Zhitao Gong Tris Warkentin Ludovic Peran Minh Giang Cl\u00e9ment Farabet Oriol Vinyals Jeff Dean Koray Kavukcuoglu Demis Hassabis Zoubin Ghahramani Douglas Eck Joelle Barral Fernando Pereira Eli Collins Armand Joulin Noah Fiedel Evan Senter Alek Andreev and Kathleen Kenealy. 2024. Gemma: Open Models Based on Gemini Research and Technology. CoRR abs\/2403.08295 (2024). https:\/\/arxiv.org\/abs\/2403.08295"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0169-7439(97)00061-0"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/45.329294"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2021.3059962"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/AICAS51828.2021.9458493"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/AICAS57966.2023.10168625"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00050"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3715123"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2023.01.008"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3155284.3018755"},{"key":"e_1_3_2_36_2","volume-title":"Scarpazza","author":"Jia Zhe","year":"2018","unstructured":"Zhe Jia, Marco Maggioni, Benjamin Staiger, and Daniele P. Scarpazza. 2018. Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking. CoRR abs\/1804.06826 (2018). Retrieved from https:\/\/arxiv.org\/abs\/1804.06826"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00016"},{"key":"e_1_3_2_38_2","volume-title":"Yichen Zhou, and Mike Burrows.","author":"Kaufman Samuel J.","year":"2020","unstructured":"Samuel J. Kaufman, Phitchaya Mangpo Phothilimthana, Yichen Zhou, and Mike Burrows. 2020. A Learned Performance Model for the Tensor Processing Unit. CoRR abs\/2008.01040 (2020). Retrieved from https:\/\/arxiv.org\/abs\/2008.01040"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.5555\/3327144.3327258"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAD.2017.8203809"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3656177"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA57654.2024.00017"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA53966.2022.00060"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3706628.3708878"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2019.2912916"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3322967"},{"key":"e_1_3_2_47_2","volume-title":"Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation. USENIX Association, USA, Article 49","author":"Zheng Lianmin","year":"2020","unstructured":"Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, and Ion Stoica. 2020. Ansor: generating high-performance tensor programs for deep learning. In Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation. USENIX Association, USA, Article 49, 863\u2013879."},{"key":"e_1_3_2_48_2","volume-title":"Genetic Algorithms in Search, Optimization and Machine Learning (1st. ed.)","author":"Goldberg David E.","unstructured":"David E. Goldberg. 1989. Genetic Algorithms in Search, Optimization and Machine Learning (1st. ed.). Addison-Wesley Longman Publishing Co. USA."},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.5555\/3327144.3327258"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3315508.3329973"},{"key":"e_1_3_2_51_2","volume-title":"PyTorch Official Blog","author":"Ansel J.","year":"2025","unstructured":"J. Ansel, O. Ulgen, W. Feng, J. Choi, and E. Ellison. 2025. Helion: A high-level DSL for performant and portable ML kernels. PyTorch Official Blog, Meta, Oct. 2025. [Online]. Available: Retrieved from https:\/\/pytorch.org\/blog\/helion\/"},{"key":"e_1_3_2_52_2","volume-title":"Proceedings of Machine Learning and Systems (MLSys).","author":"Zhang Genghan","year":"2026","unstructured":"Genghan Zhang, Shaowei Zhu, Anjiang Wei, Zhenyu Song, Allen Nie, Zhen Jia, Nandita Vijaykumar, Yida Wang, and Kunle Olukotun. 2026. AccelOpt: A self-improving LLM agentic system for AI accelerator kernel optimization. In Proceedings of Machine Learning and Systems (MLSys)."},{"key":"e_1_3_2_53_2","unstructured":"Gang Liao Hongsen Qin Ying Wang Alicia Golden Michael Kuchnik Yavuz Yetim Jia Jiunn Ang Chunli Fu Yihan He Samuel Hsia Zewei Jiang Dianshi Li Uladzimir Pashkevich Varna Puvvada Feng Shi Matt Steiner Ruichao Xiao Nathan Yan Xiayu Yu Zhou Fang Roman Levenstein Kunming Ho Haishan Zhu Alec Hammond Richard Li Ajit Mathews Kaustubh Gondkar Abdul Zainul-Abedin Ketan Singh Hongtao Yu Wenyuan Chi Barney Huang Sean Zhang Noah Weller Zach Marine Wyatt Cook Carole-Jean Wu and Gaoxiang Liu. 2026. KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta. CoRR abs\/2512.23236 (2026). Retrieved from https:\/\/arxiv.org\/abs\/2512.23236"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3453483.3454106"},{"key":"e_1_3_2_55_2","volume-title":"Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 233\u2013248","author":"Zhu Haichen","year":"2022","unstructured":"Haichen Zhu, Ruihang Yao, Mingzhe Zhai, et\u00a0al. 2022. Roller: Fast and efficient tensor compilation for deep learning. In Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 233\u2013248."}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3820380","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T12:56:18Z","timestamp":1782478578000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3820380"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,26]]},"references-count":54,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3820380"],"URL":"https:\/\/doi.org\/10.1145\/3820380","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,26]]},"assertion":[{"value":"2025-12-06","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-26","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}