{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T16:46:33Z","timestamp":1782405993444,"version":"3.54.5"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"Beijing Natural Science Foundation","award":["L243010"],"award-info":[{"award-number":["L243010"]}]},{"DOI":"10.13039\/100012897","name":"ZTE Corporation","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100012897","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    Standard compiler optimization levels, such as\n                    <jats:monospace>-O3<\/jats:monospace>\n                    , which provides a fixed optimization strategy for all programs, often fail to deliver the optimal performance. Compiler auto-tuning techniques can deliver substantial speedups, but existing methods present a difficult tradeoff. While dynamic iterative approaches are effective, their requirement for repeated compilation and execution incurs high overhead, which limits their practicality. Conversely, static prediction methods offer a low-overhead alternative. However, they face a vast search space and must comprehensively learn both option-option interactions and option-program feature relationships. To overcome the challenge, we propose\n                    <jats:sc>PredComp<\/jats:sc>\n                    , a novel static framework that leverages the divide and conquer paradigm to predict desired option sets.\n                    <jats:sc>PredComp<\/jats:sc>\n                    decomposes the search space by partitioning options into distinct subspaces based on their relationships, making the prediction problem tractable. It first predicts promising option sub-sets within each subspace, focusing only on intra-subspace option interactions and their preferred program features. Then, it adopts a combination model that aggregates these top-ranked sub-sets, prioritizes inter-subspace option interactions and corresponding features to construct globally desired sets. Experiments on three widely used benchmark suites and one real-world application show that\n                    <jats:sc>PredComp<\/jats:sc>\n                    achieves average speedups of 1.1011\u00d7 over\n                    <jats:monospace>-O3<\/jats:monospace>\n                    with a single prediction. Notably, it achieves performance comparable to dynamic iterative methods while reducing tuning time from hours or days to seconds, thereby making static prediction a practical solution for large-scale and frequently evolving software.\n                  <\/jats:p>","DOI":"10.1145\/3820771","type":"journal-article","created":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T20:50:14Z","timestamp":1781038214000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["PredComp: Predicting Compiler Optimization Options with Multi-stage Learning"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-3491-1384","authenticated-orcid":false,"given":"Bingyu","family":"Gao","sequence":"first","affiliation":[{"name":"Peking University","place":["Beijing, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-1794-213X","authenticated-orcid":false,"given":"Ziming","family":"Wang","sequence":"additional","affiliation":[{"name":"Peking University","place":["Beijing, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-8220-3470","authenticated-orcid":false,"given":"Mengyu","family":"Yao","sequence":"additional","affiliation":[{"name":"Peking University","place":["Beijing, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5735-2498","authenticated-orcid":false,"given":"Zhihong","family":"Xue","sequence":"additional","affiliation":[{"name":"ZTE Corporation","place":["Chengdu, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7366-5906","authenticated-orcid":false,"given":"Xiangqun","family":"Chen","sequence":"additional","affiliation":[{"name":"Peking University","place":["Beijing, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7558-9137","authenticated-orcid":false,"given":"Ding","family":"Li","sequence":"additional","affiliation":[{"name":"Peking University","place":["Beijing, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5064-5286","authenticated-orcid":false,"given":"Yao","family":"Guo","sequence":"additional","affiliation":[{"name":"Peking University","place":["Beijing, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2026. cBench. Retrieved May 8 2026 from https:\/\/sourceforge.net\/projects\/cbenchmark\/file%s\/cBench\/V1.1\/"},{"key":"e_1_3_1_3_2","unstructured":"2026. GCC. Retrieved May 8 2026 from https:\/\/gcc.gnu.org"},{"key":"e_1_3_1_4_2","unstructured":"2026. GCC options that control optimization. Retrieved May 8 2026 from https:\/\/gcc.gnu.org\/onlinedocs\/gcc-9.2.0\/gcc\/Opt%imize-Options.html#Optimize-Options"},{"key":"e_1_3_1_5_2","unstructured":"2026. iFLYTEK STARS MAAS PLATFORM. Retrieved May 8 2026 from https:\/\/training.xfyun.cn"},{"key":"e_1_3_1_6_2","unstructured":"2026. LLVM. Retrieved May 8 2026 from https:\/\/llvm.org"},{"key":"e_1_3_1_7_2","unstructured":"2026. Perf. Retrieved May 8 2026 from https:\/\/perf.wiki.kernel.org\/index.php\/Main_Page"},{"key":"e_1_3_1_8_2","unstructured":"2026. PolyBench. Retrieved May 8 2026 from Retrieved from https:\/\/github.com\/MatthiasJReisinger\/PolyBenchC%-4.2.1"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2006.37"},{"key":"e_1_3_1_10_2","volume-title":"Machine Learning for Computer Architecture and Systems 2022","author":"Almakki Mohammed","year":"2022","unstructured":"Mohammed Almakki, Ayman Izzeldin, Qijing Huang, Ameer Haj Ali, and Chris Cummins. 2022. Autophase v2: Towards function level phase ordering optimization. In Machine Learning for Computer Architecture and Systems 2022."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/2628071.2628092"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3124452"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CASES55004.2022.00008"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3197978"},{"key":"e_1_3_1_15_2","volume-title":"Workshop on Profile and Feedback-Directed Compilation","author":"Bodin Fran\u00e7ois","year":"1998","unstructured":"Fran\u00e7ois Bodin, Toru Kisuki, Peter Knijnenburg, Mike O\u2019Boyle, and Erven Rohou. 1998. Iterative compilation in a non-linear optimisation space. In Workshop on Profile and Feedback-Directed Compilation."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3185768.3185771"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.5555\/2505464"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00110"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/1806596.1806647"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/2959100.2959190"},{"key":"e_1_3_1_21_2","first-page":"2244","volume-title":"International Conference on Machine Learning","author":"Cummins Chris","year":"2021","unstructured":"Chris Cummins, Zacharias V. Fisches, Tal Ben-Nun, Torsten Hoefler, Michael F.P. O\u2019Boyle, and Hugh Leather. 2021. Programl: A graph-based program representation for data flow analysis and compiler optimizations. In International Conference on Machine Learning. PMLR, 2244\u20132253."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3708493.3712691"},{"key":"e_1_3_1_23_2","unstructured":"Matthias Fey and Jan Eric Lenssen. 2019. Fast graph representation learning with pytorch geometric. arXiv:1903.02428. Retrieved from https:\/\/arxiv.org\/abs\/1903.02428."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10766-010-0161-2"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3735452.3735530"},{"key":"e_1_3_1_26_2","unstructured":"Dejan Grubisic Chris Cummins Volker Seeker and Hugh Leather. 2024. Compiler generated feedback for large language models. arXiv:2403.14714. Retrieved from https:\/\/arxiv.org\/abs\/2403.14714"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/1356058.1356080"},{"issue":"2","key":"e_1_3_1_28_2","first-page":"3","article-title":"Lora: Low-rank adaptation of large language models.","volume":"1","author":"Hu Edward J.","year":"2022","unstructured":"Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et\u00a0al. 2022. Lora: Low-rank adaptation of large language models. ICLR 1, 2 (2022), 3.","journal-title":"ICLR"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM.2019.00049"},{"key":"e_1_3_1_30_2","unstructured":"Binyuan Hui Jian Yang Zeyu Cui Jiaxi Yang Dayiheng Liu Lei Zhang Tianyu Liu Jiajun Zhang Bowen Yu Keming Lu et\u00a0al. 2024. Qwen2. 5-coder technical report. arXiv:2409.12186. Retrieved from https:\/\/arxiv.org\/abs\/2409.12186"},{"key":"e_1_3_1_31_2","unstructured":"Yuxi Li. 2017. Deep reinforcement learning: An overview. arXiv:1701.07274. Retrieved from https:\/\/arxiv.org\/abs\/1701.07274"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.abq1158"},{"key":"e_1_3_1_33_2","first-page":"20746","volume-title":"International Conference on Machine Learning","author":"Liang Youwei","year":"2023","unstructured":"Youwei Liang, Kevin Stone, Ali Shameli, Chris Cummins, Mostafa Elhoushi, Jiadong Guo, Benoit Steiner, Xiaomeng Yang, Pengtao Xie, Hugh James Leather, et\u00a0al. 2023. Learning compiler pass orders using coreset and normalized value prediction. In International Conference on Machine Learning. PMLR, 20746\u201320762."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330984"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1002\/widm.53"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401104"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CGO53902.2022.9741263"},{"key":"e_1_3_1_38_2","article-title":"Pytorch: An imperative style, high-performance deep learning library","author":"Paszke A","year":"2019","unstructured":"A Paszke. 2019. Pytorch: An imperative style, high-performance deep learning library. arXiv preprint arXiv:1912.01703 (2019).","journal-title":"arXiv preprint"},{"key":"e_1_3_1_39_2","article-title":"Memtier Benchmark","author":"Labs Redis","year":"2026","unstructured":"Redis Labs. 2026. Memtier Benchmark. Retrieved May 8, 2026 from https:\/\/github.com\/RedisLabs\/memtier_benchmark","journal-title":"https:\/\/github.com\/RedisLabs\/memtier_benchmark"},{"key":"e_1_3_1_40_2","article-title":"Code llama: Open foundation models for code","author":"Roziere Baptiste","year":"2023","unstructured":"Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, et\u00a0al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/335231.335246"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/INCET54531.2022.9825092"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1561\/2200000068"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503222.3507744"},{"key":"e_1_3_1_45_2","article-title":"Mlgo: a machine learning guided compiler optimizations framework","author":"Trofin Mircea","year":"2021","unstructured":"Mircea Trofin, Yundi Qian, Eugene Brevdo, Zinan Lin, Krzysztof Choromanski, and David Li. 2021. Mlgo: a machine learning guided compiler optimizations framework. arXiv preprint arXiv:2101.04808 (2021).","journal-title":"arXiv preprint"},{"key":"e_1_3_1_46_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3418463"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2018.2817118"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3532026"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE56229.2023.00209"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3640330"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3715756"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3820771","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:55:43Z","timestamp":1782402943000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3820771"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":51,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3820771"],"URL":"https:\/\/doi.org\/10.1145\/3820771","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2025-10-09","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}