{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,6]],"date-time":"2026-01-06T13:17:33Z","timestamp":1767705453793,"version":"3.41.2"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","issue":"2","funder":[{"name":"Center for Artificial Intelligence and Robotics (CAIR) at New York University Abu Dhabi","award":["CG010"],"award-info":[{"award-number":["CG010"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2025,6,30]]},"abstract":"<jats:p>Automatic code optimization enables developers to write high-level code relying on compilers to optimize it and generate efficient code for target hardware. State-of-the-art methods for automatic code optimization leverage deep learning to build cost models that predict the impact of code optimizations on execution time. However, these models are typically limited in terms of the size and complexity of the programs they support. This research presents a novel approach to developing deep learning-based cost models that address these limitations. Our approach introduces a new program representation that efficiently represents programs with complex structures and large sizes such as varying loop depths, buffer numbers, and dimensions. Furthermore, we propose a novel deep learning architecture, that can handle this dynamic program representation. This allows the model to work on larger and more complex programs than those it was trained on. We implemented this model in Tiramisu, a state-of-the-art compiler. Our evaluation shows that our proposed model can generalize to programs larger than those seen during training, while the original Tiramisu cost model cannot. We also show that such generality does not lead to a significant increase in our proposed model\u2019s Mean Absolute Percentage Error or a decrease in the quality of code optimizations found when the model is used for automatic code optimization. In contrast, our proposed model on average achieves a 41.89% improvement in speed compared to the original cost model when both models are trained on the same dataset, showing better generalization over unseen programs. This is a significant advantage over previous approaches, which typically do not support program sizes beyond those seen during the training.<\/jats:p>","DOI":"10.1145\/3727638","type":"journal-article","created":{"date-parts":[[2025,4,1]],"date-time":"2025-04-01T13:55:37Z","timestamp":1743515737000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Supporting Dynamic Program Sizes in Deep Learning-Based Cost Models for Code Optimization"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-4820-8091","authenticated-orcid":false,"given":"Yacine","family":"Hakimi","sequence":"first","affiliation":[{"name":"Laboratoire des M\u00e9thodes de Conception de Syst\u00e8mes (LMCS), \u00c9cole Nationale Sup\u00e9rieure d'Informatique","place":["Oued Smar, Algeria"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9350-3998","authenticated-orcid":false,"given":"Riyadh","family":"Baghdadi","sequence":"additional","affiliation":[{"name":"New York University Abu Dhabi","place":["Abu Dhabi, United Arab Emirates"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9237-6210","authenticated-orcid":false,"given":"Yacine","family":"Challal","sequence":"additional","affiliation":[{"name":"University of Doha for Science & Technology","place":["Doha, Qatar"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,7,2]]},"reference":[{"key":"e_1_3_4_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3322967"},{"key":"e_1_3_4_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/cases55004.2022.00008"},{"key":"e_1_3_4_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3197978"},{"key":"e_1_3_4_5_2","unstructured":"Amir H. Ashouri Muhammad Asif Manzoor Duc Minh Vu Raymond Zhang Ziwen Wang Angel Zhang Bryan Chan Tomasz S. Czajkowski and Yaoqing Gao. 2024. ACPO: AI-Enabled Compiler-Driven Program Optimization. Retrieved from https:\/\/arxiv.org\/abs\/2312.09982."},{"key":"e_1_3_4_6_2","volume-title":"Improving tiling, reducing compilation time, and extending the scope of polyhedral compilation","author":"Baghdadi Mohamed Riyadh","year":"2015","unstructured":"Mohamed Riyadh Baghdadi. 2015. Improving tiling, reducing compilation time, and extending the scope of polyhedral compilation. Ph.D. Dissertation. Paris 6."},{"key":"e_1_3_4_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2015.17"},{"key":"e_1_3_4_8_2","unstructured":"Riyadh Baghdadi Abdelkader Nadir Debbagh Kamel Abdous Fatima Zohra Benhamida Alex Renda Jonathan Elliott Frankle Michael Carbin and Saman Amarasinghe. 2020. TIRAMISU: A Polyhedral Compiler for Dense and Sparse Deep Learning. Retrieved from https:\/\/arxiv:cs.DC\/2005.04091."},{"key":"e_1_3_4_9_2","first-page":"181","article-title":"A deep learning based cost model for automatic code optimization","volume":"3","author":"Baghdadi Riyadh","year":"2021","unstructured":"Riyadh Baghdadi, Massinissa Merouani, Mohamed-Hicham Leghettas, Kamel Abdous, Taha Arbaoui, Karima Benatchba et\u00a0al. 2021. A deep learning based cost model for automatic code optimization. Proceedings of Machine Learning and Systems 3 (2021), 181\u2013193.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_4_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2019.8661197"},{"key":"e_1_3_4_11_2","article-title":"Tiramisu: A code optimization framework for high performance systems","author":"Baghdadi Riyadh","year":"2018","unstructured":"Riyadh Baghdadi, Jessica Ray, Malek Ben Romdhane, Emanuele Del Sozzo, Patricia Suriana, Shoaib Kamil, and Saman P. Amarasinghe. 2018. Tiramisu: A code optimization framework for high performance systems. Retrieved from https:\/\/arXiv:1804.10694.","journal-title":"Retrieved from https:\/\/arXiv:1804.10694"},{"key":"e_1_3_4_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2020.104075"},{"key":"e_1_3_4_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3316781.3317789"},{"key":"e_1_3_4_14_2","article-title":"Neural code comprehension: A learnable representation of code semantics","volume":"31","author":"Ben-Nun Tal","year":"2018","unstructured":"Tal Ben-Nun, Alice Shoshana Jakobovits, and Torsten Hoefler. 2018. Neural code comprehension: A learnable representation of code semantics. Adv. Neural Info. Process. Syst. 31 (2018).","journal-title":"Adv. Neural Info. Process. Syst."},{"key":"e_1_3_4_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/1375581.1375595"},{"key":"e_1_3_4_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377555.3377894"},{"key":"e_1_3_4_17_2","article-title":"Learning to optimize tensor programs","volume":"31","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. Learning to optimize tensor programs. Adv. Neural Info. Process. Syst. 31 (2018).","journal-title":"Adv. Neural Info. Process. Syst."},{"key":"e_1_3_4_18_2","article-title":"PrograML: Graph-based deep learning for program optimization and analysis","author":"Cummins Chris","year":"2020","unstructured":"Chris Cummins, Zacharias V Fisches, Tal Ben-Nun, Torsten Hoefler, and Hugh Leather. 2020. PrograML: Graph-based deep learning for program optimization and analysis. Retrieved from https:\/\/arXiv:2003.10536.","journal-title":"Retrieved from https:\/\/arXiv:2003.10536"},{"key":"e_1_3_4_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2017.24"},{"key":"e_1_3_4_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3708493.3712691"},{"key":"e_1_3_4_21_2","doi-asserted-by":"crossref","unstructured":"Chris Cummins Bram Wasti Jiadong Guo Brandon Cui Jason Ansel Sahir Gomez Somya Jain Jia Liu Olivier Teytaud Benoit Steineret al.2021. CompilerGym: Robust Performant Compiler Optimization Environments for AI Research. Retrieved from https:\/\/arxiv.org\/abs\/2109.08267.","DOI":"10.1109\/CGO53902.2022.9741258"},{"key":"e_1_3_4_22_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2015.12.114"},{"key":"e_1_3_4_23_2","unstructured":"Shukai Duan Nikos Kanakaris Xiongye Xiao Heng Ping Chenyu Zhou Nesreen K. Ahmed Guixiang Ma Mihai Capota Theodore L. Willke Shahin Nazarian and Paul Bogdan. 2023. Leveraging Reinforcement Learning and Large Language Models for Code Optimization. Retrieved from https:\/\/arxiv.org\/abs\/2312.05657."},{"key":"e_1_3_4_24_2","first-page":"1581","article-title":"Polyhedron model.","volume":"1","author":"Feautrier Paul","year":"2011","unstructured":"Paul Feautrier and Christian Lengauer. 2011. Polyhedron model. Encyc. Parallel Comput. 1 (2011), 1581\u20131592.","journal-title":"Encyc. Parallel Comput."},{"key":"e_1_3_4_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2013.6494993"},{"key":"e_1_3_4_26_2","doi-asserted-by":"publisher","DOI":"10.1142\/S0129626412500107"},{"key":"e_1_3_4_27_2","unstructured":"Dejan Grubisic Bram Wasti Chris Cummins John Mellor-Crummey and Aleksandar Zlateski. 2023. LoopTune: Optimizing Tensor Computations with Reinforcement Learning. Retrieved from https:\/\/arxiv.org\/abs\/2309.01825."},{"key":"e_1_3_4_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3368826.3377928"},{"key":"e_1_3_4_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICNAS53565.2021.9628950"},{"key":"e_1_3_4_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10766-023-00758-5"},{"key":"e_1_3_4_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3696443.3708943"},{"key":"e_1_3_4_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/FDL50818.2020.9232934"},{"key":"e_1_3_4_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCC.and.EUC.2013.119"},{"key":"e_1_3_4_34_2","unstructured":"Pouchet Louis-Noel. 2010. PolyBench Suite. Retrieved from http:\/\/www.cse.ohio-state.edu\/~pouchet\/software\/polybench\/."},{"key":"e_1_3_4_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669121"},{"key":"e_1_3_4_36_2","first-page":"4505","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Mendis Charith","year":"2019","unstructured":"Charith Mendis, Alex Renda, Saman Amarasinghe, and Michael Carbin. 2019. Ithemal: Accurate, portable and fast basic block throughput estimation using deep neural networks. In Proceedings of the International Conference on Machine Learning. PMLR, 4505\u20134515."},{"key":"e_1_3_4_37_2","volume-title":"Compiler Auto-vectorization with Imitation Learning","author":"Mendis Charith","year":"2019","unstructured":"Charith Mendis, Cambridge Yang, Yewen Pu, Saman Amarasinghe, and Michael Carbin. 2019. Compiler Auto-vectorization with Imitation Learning. Curran Associates Inc., Red Hook, NY."},{"key":"e_1_3_4_38_2","doi-asserted-by":"publisher","unstructured":"Lina Mezdour Khadidja Kadem Massinissa Merouani Amina Selma Haichour Saman Amarasinghe and Riyadh Baghdadi. 2023. A deep learning model for loop interchange. InProceedings of the 32nd ACM SIGPLAN International Conference on Compiler Construction (CC\u201923). 50\u201360. DOI:10.1145\/3578360.3580257","DOI":"10.1145\/3578360.3580257"},{"key":"e_1_3_4_39_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Mirhoseini Azalia","year":"2018","unstructured":"Azalia Mirhoseini, Anna Goldie, Hieu Pham, Benoit Steiner, Quoc V. Le, and Jeff Dean. 2018. A hierarchical model for device placement. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_4_40_2","first-page":"2430","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Mirhoseini Azalia","year":"2017","unstructured":"Azalia Mirhoseini, Hieu Pham, Quoc V. Le, Benoit Steiner, Rasmus Larsen, Yuefeng Zhou, Naveen Kumar, Mohammad Norouzi, Samy Bengio, and Jeff Dean. 2017. Device placement optimization with reinforcement learning. In Proceedings of the International Conference on Machine Learning. PMLR, 2430\u20132439."},{"key":"e_1_3_4_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3150211"},{"key":"e_1_3_4_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3520312.3534863"},{"key":"e_1_3_4_43_2","unstructured":"Mircea Trofin Yundi Qian Eugene Brevdo Zinan Lin Krzysztof Choromanski and David Li. 2021. MLGO: A Machine Learning Guided Compiler Optimizations Framework. Retrieved from https:\/\/arxiv.org\/abs\/2101.04808."},{"key":"e_1_3_4_44_2","article-title":"Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions","author":"Vasilache Nicolas","year":"2018","unstructured":"Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S. Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2018. Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions. Retrieved from https:\/\/arXiv:1802.04730.","journal-title":"Retrieved from https:\/\/arXiv:1802.04730"},{"key":"e_1_3_4_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3418463"},{"key":"e_1_3_4_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2018.2817118"},{"key":"e_1_3_4_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/2579561"},{"key":"e_1_3_4_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/HiPC.2014.7116910"},{"key":"e_1_3_4_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.2978386"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3727638","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,2]],"date-time":"2025-07-02T12:20:37Z","timestamp":1751458837000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3727638"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,30]]},"references-count":48,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,6,30]]}},"alternative-id":["10.1145\/3727638"],"URL":"https:\/\/doi.org\/10.1145\/3727638","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2025,6,30]]},"assertion":[{"value":"2024-09-09","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-18","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-02","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}