{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T23:21:01Z","timestamp":1784762461585,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":58,"publisher":"ACM","funder":[{"name":"This work was supported by the Institute of Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT)","award":["RS-2023-00277060, RS-2024-00459797, RS-2025-02216517, RS-2025-02214497"],"award-info":[{"award-number":["RS-2023-00277060, RS-2024-00459797, RS-2025-02216517, RS-2025-02214497"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,6,13]]},"DOI":"10.1145\/3735452.3735538","type":"proceedings-article","created":{"date-parts":[[2025,6,13]],"date-time":"2025-06-13T15:11:16Z","timestamp":1749827476000},"page":"134-145","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Multi-level Machine Learning-Guided Autotuning for Efficient Code Generation on a Deep Learning Accelerator"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-2123-454X","authenticated-orcid":false,"given":"JooHyoung","family":"Cha","sequence":"first","affiliation":[{"name":"University of Science and Technology, Daejeon, Republic of Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-3667-7455","authenticated-orcid":false,"given":"Munyoung","family":"Lee","sequence":"additional","affiliation":[{"name":"Electronics and Telecommunications Research Institute, Daejeon, Republic of Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3091-9926","authenticated-orcid":false,"given":"Jinse","family":"Kwon","sequence":"additional","affiliation":[{"name":"Electronics and Telecommunications Research Institute, Daejeon, Republic of Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9332-3508","authenticated-orcid":false,"given":"Jemin","family":"Lee","sequence":"additional","affiliation":[{"name":"Electronics and Telecommunications Research Institute, Daejeon, Republic of Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2973-246X","authenticated-orcid":false,"given":"Yongin","family":"Kwon","sequence":"additional","affiliation":[{"name":"Electronics and Telecommunications Research Institute, Daejeon, Republic of Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,13]]},"reference":[{"key":"e_1_3_2_2_1_1","unstructured":"2024. VTA Library customized by ONES AI. Accessed: 2025-03-21"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","unstructured":"Mart\u00edn Abadi and et al. 2015. TensorFlow Large-scale machine learning on heterogeneous systems. https:\/\/doi.org\/10.5281\/zenodo.4724125 10.5281\/zenodo.4724125","DOI":"10.5281\/zenodo.4724125"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","unstructured":"Jason Ansel and et al. 2024. PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation. ACM. https:\/\/doi.org\/10.1145\/3620665.3640366 10.1145\/3620665.3640366","DOI":"10.1145\/3620665.3640366"},{"key":"e_1_3_2_2_4_1","volume-title":"Iheb Nassim Aouadj, and Riyadh Baghdadi","author":"Bendib Nazim","year":"2024","unstructured":"Nazim Bendib, Iheb Nassim Aouadj, and Riyadh Baghdadi. 2024. A Reinforcement Learning Environment for Automatic Code Optimization in the MLIR Compiler. arxiv:2409.11068. arxiv:2409.11068"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3582016.3582061"},{"key":"e_1_3_2_2_6_1","volume-title":"Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang.","author":"Bradbury James","year":"2018","unstructured":"James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. 2018. JAX: composable transformations of Python+NumPy programs. http:\/\/github.com\/google\/jax"},{"key":"e_1_3_2_2_7_1","unstructured":"Chris J.C. Burges. 2010. From RankNet to LambdaRank to LambdaMART: An Overview. https:\/\/www.microsoft.com\/en-us\/research\/publication\/from-ranknet-to-lambdarank-to-lambdamart-an-overview\/"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3372799.3394361"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00110"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_3_2_2_11_1","unstructured":"Tianqi Chen Mu Li Yutian Li Min Lin Naiyan Wang Minjie Wang Tianjun Xiao Bing Xu Chiyuan Zhang and Zheng Zhang. 2015. MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. arxiv:1512.01274. arxiv:1512.01274"},{"key":"e_1_3_2_2_12_1","volume-title":"TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In USENIX Symposium on Operating Systems Design and Implementation. https:\/\/api.semanticscholar.org\/CorpusID:52939079","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Q. Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In USENIX Symposium on Operating Systems Design and Implementation. https:\/\/api.semanticscholar.org\/CorpusID:52939079"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3689730"},{"key":"e_1_3_2_2_14_1","unstructured":"Arya Fayyazi Mehdi Kamal and Massoud Pedram. 2024. ARCO:Adaptive Multi-Agent Reinforcement Learning-Based Hardware\/Software Co-Optimization Compiler for Improved Performance in DNN Accelerator Design. arxiv:2407.08192. arxiv:2407.08192"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"crossref","unstructured":"Grigori Fursin and et al.. 2011. Milepost GCC: Machine learning enabled self-tuning compiler. International Journal of Parallel Programming.","DOI":"10.1007\/s10766-010-0161-2"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2024.3373763"},{"key":"e_1_3_2_2_17_1","unstructured":"Dejan Grubisic Bram Wasti Chris Cummins John Mellor-Crummey and Aleksandar Zlateski. 2023. LoopTune: Optimizing Tensor Computations with Reinforcement Learning. arXiv preprint arXiv:2309.01825 arxiv:2309.01825"},{"key":"e_1_3_2_2_18_1","volume-title":"Euro-Par 2014 Parallel Processing, Fernando Silva, In\u00eas Dutra, and V\u00edtor Santos Costa (Eds.)","author":"Gschwandtner Philipp","unstructured":"Philipp Gschwandtner, Juan J. Durillo, and Thomas Fahringer. 2014. Multi-Objective Auto-Tuning with Insieme: Optimization and Trade-Off Analysis for Time, Energy and Resource Usage. In Euro-Par 2014 Parallel Processing, Fernando Silva, In\u00eas Dutra, and V\u00edtor Santos Costa (Eds.). Springer International Publishing, Cham. 87\u201398. isbn:978-3-319-09873-9"},{"key":"e_1_3_2_2_19_1","unstructured":"Ga\u00ebl Guennebaud and Beno\u00eet Jacob. 2010. Eigen v3. http:\/\/eigen.tuxfamily.org."},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3650200.3656631"},{"key":"e_1_3_2_2_22_1","unstructured":"Intel Corporation. 2023. OpenVINO Toolkit. https:\/\/docs.openvino.ai\/latest\/index.html"},{"key":"e_1_3_2_2_23_1","volume-title":"Kingma and Prafulla Dhariwal","author":"Diederik","year":"2018","unstructured":"Diederik P. Kingma and Prafulla Dhariwal. 2018. Glow: Generative Flow with Invertible 1x1 Convolutions. arxiv:1807.03039. arxiv:1807.03039"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2023.3303851"},{"key":"e_1_3_2_2_25_1","volume-title":"2013 USENIX Annual Technical Conference (USENIX ATC 13)","author":"Kwon Yongin","year":"2013","unstructured":"Yongin Kwon, Sangmin Lee, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Byung-Gon Chun, Ling Huang, Petros Maniatis, Mayur Naik, and Yunheung Paek. 2013. Mantis: Automatic performance prediction for smartphone applications. In 2013 USENIX Annual Technical Conference (USENIX ATC 13). 297\u2013308."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2833157.2833162"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.5555\/977395.977673"},{"key":"e_1_3_2_2_28_1","volume-title":"River Riddle, Tatiana Shpeisman, Nicolas Vasilache, and Oleksandr Zinenko.","author":"Lattner Chris","year":"2021","unstructured":"Chris Lattner, Mehdi Amini, Uday Bondhugula, Albert Cohen, Andy Davis, Jacques Arnaud Pienaar, River Riddle, Tatiana Shpeisman, Nicolas Vasilache, and Oleksandr Zinenko. 2021. MLIR: Scaling Compiler Infrastructure for Domain Specific Computation. In CGO 2021."},{"key":"e_1_3_2_2_29_1","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS \u201920)","author":"Li Menghao","year":"2020","unstructured":"Menghao Li, Minjia Zhang, Chi Wang, and Mingqin Li. 2020. AdaTune: adaptive tensor program compilation made efficient. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS \u201920). Curran Associates Inc., Red Hook, NY, USA. Article 1241, 13 pages. isbn:9781713829546"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3605573.3605596"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3575693.3575707"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3575693.3575707"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","unstructured":"Zixuan Ma Haojie Wang Jia Xing Liyan Zheng Chen Zhang Huanqi Cao Kezhao Huang Shizhi Tang Penghan Wang and Jidong Zhai. 2023. IntelliGen: A Tensor Compiler for Memory-Intensive Operators via Unified Graph-Level and Instruction-Level Optimizations. arXiv preprint arXiv:2307.04995 https:\/\/doi.org\/10.48550\/arXiv.2307.04995 10.48550\/arXiv.2307.04995","DOI":"10.48550\/arXiv.2307.04995"},{"key":"e_1_3_2_2_34_1","volume-title":"Artificial Intelligence and Hardware Accelerators","author":"Mishra Ashutosh","unstructured":"Ashutosh Mishra, Jaekwang Cha, Hyunbin Park, and Shiho Kim. 2023. Artificial Intelligence and Hardware Accelerators. Springer."},{"key":"e_1_3_2_2_35_1","unstructured":"Thierry Moreau Tianqi Chen Luis Vega Jared Roesch Eddie Yan Lianmin Zheng Josh Fromm Ziheng Jiang Luis Ceze Carlos Guestrin and Arvind Krishnamurthy. 2019. A Hardware-Software Blueprint for Flexible Deep Learning Specialization. arxiv:1807.04188. arxiv:1807.04188"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2023.3288758"},{"key":"e_1_3_2_2_37_1","unstructured":"Pengyu Mu Linquan Wei Yi Liu and Rui Wang. 2024. FTuner: A Fast Dynamic Shape Tensors Program Auto-Tuner for Deep Learning Compilers. arXiv preprint arXiv:2407.21418 arxiv:2407.21418"},{"key":"e_1_3_2_2_38_1","unstructured":"NVIDIA Corporation. 2023. NVIDIA TensorRT: High Performance Deep Learning Inference. Version 8.6 https:\/\/developer.nvidia.com\/tensorrt"},{"key":"e_1_3_2_2_39_1","unstructured":"ONNX Runtime developers. 2018. ONNX Runtime. https:\/\/github.com\/microsoft\/onnxruntime"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.4218\/etrij.2024-0139"},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078195"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3676641.3716269"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2499370.2462176"},{"key":"e_1_3_2_2_44_1","unstructured":"Dennis Rieber Moritz Reiber Oliver Bringmann and Holger Fr\u00f6ning. 2022. HW-Aware Initialization of DNN Auto-Tuning to Improve Exploration Time and Robustness. arxiv:2205.15568. arxiv:2205.15568"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3497776.3517774"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3497776.3517774"},{"key":"e_1_3_2_2_47_1","volume-title":"XLA : Compiling Machine Learning for Peak Performance.","author":"Sabne Amit","year":"2020","unstructured":"Amit Sabne. 2020. XLA : Compiling Machine Learning for Peak Performance."},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2020.05.035"},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"publisher","unstructured":"Yuyang Wang Xuxin Lin Mingwei Zhou and Yao Liang. 2023. Optimizing Tensor Programs of DNNs via Adaptive Differential Evolution with Hardware Measurement Estimation. SSRN preprint. https:\/\/doi.org\/10.2139\/ssrn.4528511 10.2139\/ssrn.4528511","DOI":"10.2139\/ssrn.4528511"},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/CSCWD61410.2024.10580120"},{"key":"e_1_3_2_2_51_1","volume-title":"Scorch: A Library for Sparse Deep Learning. arxiv:2405.16883. arxiv:2405.16883","author":"Yan Bobby","year":"2024","unstructured":"Bobby Yan, Alexander J. Root, Trevor Gale, David Broman, and Fredrik Kjolstad. 2024. Scorch: A Library for Sparse Deep Learning. arxiv:2405.16883. arxiv:2405.16883"},{"key":"e_1_3_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD50377.2020.00108"},{"key":"e_1_3_2_2_53_1","volume-title":"DynaTune: Dynamic Tensor Program Optimization in Deep Neural Network Compilation. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=GTGb3M_KcUl","author":"Zhang Minjia","year":"2021","unstructured":"Minjia Zhang, Menghao Li, Chi Wang, and Mingqin Li. 2021. DynaTune: Dynamic Tensor Program Optimization in Deep Neural Network Compilation. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=GTGb3M_KcUl"},{"key":"e_1_3_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3620666.3651348"},{"key":"e_1_3_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3572864.3580330"},{"key":"e_1_3_2_2_56_1","volume-title":"Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, and Koushik Sen.","author":"Zheng Lianmin","year":"2020","unstructured":"Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, and Koushik Sen. 2020. Ansor: Generating high-performance tensor programs for deep learning. In 14th $USENIX$ Symposium on Operating Systems Design and Implementation ($OSDI$ 20). 863\u2013879."},{"key":"e_1_3_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE56229.2023.00209"},{"key":"e_1_3_2_2_58_1","volume-title":"International Conference on Programming Language Design and Implementation (PLDI).","author":"Zhu Xiaoyang","year":"2024","unstructured":"Xiaoyang Zhu and Junjie Hao. 2024. CompTuner++: Fast Compiler Tuning via Lightweight ML Models. In International Conference on Programming Language Design and Implementation (PLDI)."}],"event":{"name":"LCTES '25: 26th ACM SIGPLAN\/SIGBED International Conference on Languages, Compilers, and Tools for Embedded Systems","location":"Seoul Republic of Korea","acronym":"LCTES '25","sponsor":["SIGPLAN ACM Special Interest Group on Programming Languages","SIGBED ACM Special Interest Group on Embedded Systems"]},"container-title":["Proceedings of the 26th ACM SIGPLAN\/SIGBED International Conference on Languages, Compilers, and Tools for Embedded Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3735452.3735538","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,16]],"date-time":"2025-07-16T07:12:17Z","timestamp":1752649937000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3735452.3735538"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,13]]},"references-count":58,"alternative-id":["10.1145\/3735452.3735538","10.1145\/3735452"],"URL":"https:\/\/doi.org\/10.1145\/3735452.3735538","relation":{},"subject":[],"published":{"date-parts":[[2025,6,13]]},"assertion":[{"value":"2025-06-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}