{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T16:46:40Z","timestamp":1782406000167,"version":"3.54.5"},"reference-count":64,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    Deep Learning\u00a0(DL) models are at the core of a growing number of applications, making fast, low-latency execution across diverse device architectures both a critical requirement and a challenge. DL compilers, such as TVM, address this challenge by automatically translating high-level models into optimized low-level code that effectively exploits device architectures. However, search-space-based algorithms face difficulties in exploring the vast optimization sequence space, often resulting in lengthy compilation times that significantly impact the design cycle. This article introduces the Task Graph Caching (TGC) algorithm,\n                    <jats:xref ref-type=\"fn\">\n                      <jats:sup>1<\/jats:sup>\n                    <\/jats:xref>\n                    which aims to reduce the high compilation time while preserving the quality of the code generated by TVM. In particular, TGC enhances TVM auto-tuning by exploiting the fact that similar DL subgraphs appear both within and across models, thus enabling optimization sequences discovered in past compilations to guide future executions. To achieve this, TGC introduces a cache structure that stores high-performance optimization sequences found in previous TVM executions. This information is then used to seed the population of TVM evolutionary search, avoiding redundant exploration of the optimization space and accelerating convergence. Experimental results on twelve DL models show that TGC can significantly speed up the search for efficient optimization sequences, reducing auto-tuning time by up to 2.89\u00d7 for Ansor and 3.13\u00d7 for MetaSchedule on CPU. Moreover, TGC reduces auto-tuning time by up to 3.25\u00d7 for MetaSchedule on GPU. Furthermore, on average, TGC can maintain the inference time achieved by the default TVM, making it a promising solution for accelerating the compilation of DL models.\n                  <\/jats:p>","DOI":"10.1145\/3810246","type":"journal-article","created":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T11:25:22Z","timestamp":1776425122000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Using Task Graph Caching to Accelerate TVM Code Generation"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9605-1037","authenticated-orcid":false,"given":"Thais","family":"Camacho","sequence":"first","affiliation":[{"name":"Instituto de Computa\u00e7\u00e3o, UNICAMP","place":["Campinas, Brazil"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9120-476X","authenticated-orcid":false,"given":"Lucas","family":"Alvarenga","sequence":"additional","affiliation":[{"name":"Instituto de Computa\u00e7\u00e3o, UNICAMP","place":["Campinas, Brazil"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1140-4513","authenticated-orcid":false,"given":"Marcio","family":"Pereira","sequence":"additional","affiliation":[{"name":"Instituto de Computa\u00e7\u00e3o, UNICAMP","place":["Campinas, Brazil"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4869-5190","authenticated-orcid":false,"given":"Guido","family":"Araujo","sequence":"additional","affiliation":[{"name":"Instituto de Computa\u00e7\u00e3o, UNICAMP","place":["Campinas, Brazil"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_3_2_2","first-page":"265","volume-title":"Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16)","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et\u00a0al. 2016. TensorFlow: A system for large-scale machine learning. In Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). USENIX Association, Savannah, GA, 265\u2013283. Retrieved from https:\/\/www.usenix.org\/conference\/osdi16\/technical-sessions\/presentation\/abadi"},{"key":"e_1_3_3_3_2","unstructured":"Byung Hoon Ahn Prannoy Pilligundla Amir Yazdanbakhsh and Hadi Esmaeilzadeh. 2020. Chameleon: Adaptive code optimization for expedited deep neural network compilation. CoRR abs\/2001.08743 (2020). Retrieved from https:\/\/arxiv.org\/abs\/2001.08743"},{"key":"e_1_3_3_4_2","volume-title":"Compilers Principles, Techniques, and Tools","author":"Alfred V. Aho","year":"2007","unstructured":"V. Aho Alfred, S. Lam Monica, and D. Ullman Jeffrey. 2007. Compilers Principles, Techniques, and Tools. Pearson Education."},{"issue":"11","key":"e_1_3_3_5_2","article-title":"Yaml ain\u2019t markup language (yaml\u2122) version 1.1","volume":"5","author":"Ben-Kiki Oren","year":"2009","unstructured":"Oren Ben-Kiki, Clark Evans, and Brian Ingerson. 2009. Yaml ain\u2019t markup language (yaml\u2122) version 1.1. Working Draft 2008 5, 11 (2009), 1\u201382.","journal-title":"Working Draft 2008"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.17635\/lancaster\/thesis\/2049"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2023.3279233"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CLOUD55607.2022.00061"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.5753\/wscad.2019.8661"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3641289"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSCCC.2018.8703316"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_3_3_13_2","first-page":"579","volume-title":"Proceedings of the 13th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201918)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, et\u00a0al. 2018. TVM: An automated end-to-end optimizing compiler for deep learning. In Proceedings of the 13th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201918). USENIX Association, USA, 579\u2013594."},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.5555\/3327144.3327258"},{"key":"e_1_3_3_15_2","unstructured":"GCC Compiler. 2009. The GNU compiler collection. http:\/\/gcc.gnu.org\/-Acessoem 10 11 (2009) 2009."},{"key":"e_1_3_3_16_2","volume-title":"CUDA Programming: A Developer\u2019s Guide to Parallel Computing with GPUs","author":"Cook Shane","year":"2012","unstructured":"Shane Cook. 2012. CUDA Programming: A Developer\u2019s Guide to Parallel Computing with GPUs. Newnes."},{"key":"e_1_3_3_17_2","volume-title":"Engineering A Compiler","author":"Cooper Keith D.","year":"2022","unstructured":"Keith D. Cooper and Linda Torczon. 2022. Engineering A Compiler. Morgan Kaufmann."},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CEC48606.2020.9185646"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3559009.3569682"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.5555\/3086952"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.243"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4842-5364-9_2"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-13-1610-4_28"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3065386"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3759916"},{"key":"e_1_3_3_29_2","unstructured":"Ruihang Lai Junru Shao Siyuan Feng Steven S. Lyubomirsky Bohan Hou Wuwei Lin Zihao Ye Hongyi Jin Yuchen Jin Jiawei Liu et\u00a0al. 2023. Relax: Composable abstractions for end-to-end dynamic machine learning. 2 (2023) 998\u20131013. arXiv:2311.02103. Retrieved from https:\/\/arxiv.org\/abs\/2311.02103"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMITCon.2019.8862255"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.5555\/977395.977673"},{"key":"e_1_3_3_32_2","unstructured":"Chris Lattner Jacques A. Pienaar Mehdi Amini Uday Bondhugula River Riddle Albert Cohen Tatiana Shpeisman Andy Davis Nicolas Vasilache and Oleksandr Zinenko. 2020. MLIR: A compiler infrastructure for the end of Moore\u2019s law. CoRR abs\/2002.11054 (2020). Retrieved from https:\/\/arxiv.org\/abs\/2002.11054"},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3323353"},{"key":"e_1_3_3_34_2","article-title":"The deep learning compiler: A comprehensive survey","author":"Li Mingzhen","year":"2021","unstructured":"Mingzhen Li, Yi Liu, Xiaoyan Liu, Qingxiao Sun, Xin You, Hailong Yang, Zhongzhi Luan, Lin Gan, Guangwen Yang, and Depei Qian. 2021. The deep learning compiler: A comprehensive survey. IEEE Transactions on Parallel and Distributed Systems 32, 3 (2021), 708\u2013727.","journal-title":"IEEE Transactions on Parallel and Distributed Systems"},{"key":"e_1_3_3_35_2","first-page":"14807","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"33","author":"Li Menghao","year":"2020","unstructured":"Menghao Li, Minjia Zhang, Chi Wang, and Mingqin Li. 2020. AdaTune: Adaptive tensor program compilation made efficient. In Proceedings of the Advances in Neural Information Processing Systems. H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33, Curran Associates, Inc., 14807\u201314819."},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACPR.2015.7486599"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3578360.3580266"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874254"},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.6089"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2023.101869"},{"key":"e_1_3_3_41_2","volume-title":"PyTorch: An Imperative Style, High-performance Deep Learning Library","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et\u00a0al. 2019. PyTorch: An Imperative Style, High-performance Deep Learning Library. Curran Associates Inc., Red Hook, NY, USA."},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-10665-1_63"},{"key":"e_1_3_3_43_2","volume-title":"Deep Learning Systems: Algorithms, Compilers, and Processors for Large-Scale Production","author":"Rodriguez Andres","year":"2020","unstructured":"Andres Rodriguez. 2020. Deep Learning Systems: Algorithms, Compilers, and Processors for Large-Scale Production. Morgan and Claypool Publishers."},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3211346.3211348"},{"key":"e_1_3_3_45_2","unstructured":"Nadav Rotem Jordan Fix Saleem Abdulrasool Summer Deng Roman Dzhabarov James Hegeman Roman Levenstein Bert Maher Nadathur Satish Jakob Olesen et\u00a0al. 2018. Glow: Graph lowering compiler techniques for neural networks. CoRR abs\/1805.00907 (2018). Retrieved from https:\/\/arxiv.org\/abs\/1805.00907"},{"key":"e_1_3_3_46_2","volume-title":"Artificial Intelligence: A Modern Approach","author":"Russell Stuart","year":"2010","unstructured":"Stuart Russell and Peter Norvig. 2010. Artificial Intelligence: A Modern Approach. Prentice Hall."},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3497776.3517774"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-12-821285-1.00021-X"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-2593"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMPTELIX.2017.8003957"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_3_53_2","unstructured":"Christian Szegedy Vincent Vanhoucke Sergey Ioffe Jonathon Shlens and Zbigniew Wojna. 2015. Rethinking the Inception Architecture for Computer Vision. CoRR abs\/1512.00567 (2015). Retrieved from https:\/\/arxiv.org\/abs\/1512.00567"},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-15554-8_73"},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.634"},{"key":"e_1_3_3_56_2","first-page":"204","volume-title":"Proceedings of the Machine Learning and Systems","volume":"4","author":"Xing Jiarong","year":"2022","unstructured":"Jiarong Xing, Leyuan Wang, Shang Zhang, Jack Chen, Ang Chen, and Yibo Zhu. 2022. Bolt: Bridging the gap between auto-tuners and hardware-native performance. In Proceedings of the Machine Learning and Systems. D. Marculescu, Y. Chi, and C. Wu (Eds.), Vol. 4, 204\u2013216. Retrieved from https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2022\/file\/1f8053a67ec8e0b57455713cefdd8218-Paper.pdf"},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICESS.2019.8782480"},{"key":"e_1_3_3_58_2","unstructured":"Sergey Zagoruyko and Nikos Komodakis. 2017. Wide Residual Networks. CoRR abs\/1605.07146 (2017). Retrieved from https:\/\/arxiv.org\/abs\/1605.07146"},{"key":"e_1_3_3_59_2","doi-asserted-by":"publisher","DOI":"10.34133\/icomputing.0040"},{"key":"e_1_3_3_60_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Zhang Minjia","year":"2021","unstructured":"Minjia Zhang, Menghao Li, Chi Wang, and Mingqin Li. 2021. DynaTune: Dynamic tensor program optimization in deep neural network compilation. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=GTGb3M_KcUl"},{"key":"e_1_3_3_61_2","unstructured":"Shanjun Zhang Mingzhen Li Hailong Yang Yi Liu Zhongzhi Luan and Depei Qian. 2022. FamilySeer: Towards optimized tensor codes by exploiting computation subgraph similarity. CoRR abs\/2201.00194 (2022). Retrieved from https:\/\/arxiv.org\/abs\/2201.00194"},{"key":"e_1_3_3_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3572864.3580330"},{"key":"e_1_3_3_63_2","first-page":"848","volume-title":"Proceedings of the Machine Learning and Systems","volume":"4","author":"Zheng Bojian","year":"2022","unstructured":"Bojian Zheng, Ziheng Jiang, Cody Hao Yu, Haichen Shen, Joshua Fromm, Yizhi Liu, Yida Wang, Luis Ceze, Tianqi Chen, and Gennady Pekhimenko. 2022. DietCode: Automatic Optimization for Dynamic Tensor Programs. In Proceedings of the Machine Learning and Systems. D. Marculescu, Y. Chi, and C. Wu (Eds.), Vol. 4, 848\u2013863. Retrieved from https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2022\/file\/f89b79c9a28d4cae22ef9e557d9fa191-Paper.pdf"},{"key":"e_1_3_3_64_2","volume-title":"Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201920)","author":"Zheng Lianmin","year":"2020","unstructured":"Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, et\u00a0al. 2020. Ansor: Generating high-performance tensor programs for deep learning. In Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201920). USENIX Association, USA, Article 49, 17 pages."},{"key":"e_1_3_3_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378508"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3810246","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:56:02Z","timestamp":1782402962000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3810246"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":64,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3810246"],"URL":"https:\/\/doi.org\/10.1145\/3810246","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2025-09-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-03","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}