{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,21]],"date-time":"2026-01-21T06:48:47Z","timestamp":1768978127714,"version":"3.49.0"},"publisher-location":"New York, NY, USA","reference-count":53,"publisher":"ACM","license":[{"start":{"date-parts":[[2025,3,1]],"date-time":"2025-03-01T00:00:00Z","timestamp":1740787200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Institute of Information & communications Technology Planning & Evaluation","award":["RS-2024-00339187, 2021-0-00310, RS-2020-II201361, No.RS-2023-00277060"],"award-info":[{"award-number":["RS-2024-00339187, 2021-0-00310, RS-2020-II201361, No.RS-2023-00277060"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,3]]},"DOI":"10.1145\/3696443.3708944","type":"proceedings-article","created":{"date-parts":[[2025,2,22]],"date-time":"2025-02-22T11:50:26Z","timestamp":1740225026000},"page":"209-224","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["CUrator: An Efficient LLM Execution Engine with Optimized Integration of CUDA Libraries"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-4021-1091","authenticated-orcid":false,"given":"Yoon Noh","family":"Lee","sequence":"first","affiliation":[{"name":"Yonsei University, Seoul, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9838-4310","authenticated-orcid":false,"given":"Yongseung","family":"Yu","sequence":"additional","affiliation":[{"name":"Yonsei University, Seoul, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3725-0380","authenticated-orcid":false,"given":"Yongjun","family":"Park","sequence":"additional","affiliation":[{"name":"Yonsei University, Seoul, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,3]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2024. cuBlas Library. http:\/\/developer.nvidia.com\/cublas"},{"key":"e_1_3_2_1_2_1","first-page":"351","article-title":"Vidur: A Large-Scale Simulation Framework For LLM Inference","volume":"6","author":"Agrawal Amey","year":"2024","unstructured":"Amey Agrawal, Nitin Kedia, Jayashree Mohan, Ashish Panwar, Nipun Kwatra, Bhargav Gulavani, Ramachandran Ramjee, and Alexey Tumanov. 2024. Vidur: A Large-Scale Simulation Framework For LLM Inference. Proceedings of Machine Learning and Systems, 6 (2024), 351\u2013366.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/3691938.3691945"},{"key":"e_1_3_2_1_4_1","volume-title":"Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=rygG4AVFvH","author":"Ahn Byung Hoon","year":"2020","unstructured":"Byung Hoon Ahn, Prannoy Pilligundla, Amir Yazdanbakhsh, and Hadi Esmaeilzadeh. 2020. Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=rygG4AVFvH"},{"key":"e_1_3_2_1_5_1","unstructured":"AI@Meta. 2024. Llama 3 Model Card. https:\/\/github.com\/meta-llama\/llama3\/blob\/main\/MODEL_CARD.md"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41404.2022.00051"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2628071.2628092"},{"key":"e_1_3_2_1_8_1","volume-title":"Onnx: Open neural network exchange.","author":"Bai Junjie","year":"2019","unstructured":"Junjie Bai, Fang Lu, and Ke Zhang. 2019. Onnx: Open neural network exchange."},{"key":"e_1_3_2_1_9_1","volume-title":"Advances in Neural Information Processing Systems","author":"Borzunov Alexander","year":"2023","unstructured":"Alexander Borzunov, Max Ryabinin, Artem Chumachenko, Dmitry Baranchuk, Tim Dettmers, Younes Belkada, Pavel Samygin, and Colin A Raffel. 2023. Distributed Inference and Fine-tuning of Large Language Models Over The Internet. In Advances in Neural Information Processing Systems, A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.). 36, Curran Associates, Inc., 12312\u201312331. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2023\/file\/28bf1419b9a1f908c15f6195f58cb865-Paper-Conference.pdf"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3656177"},{"key":"e_1_3_2_1_11_1","volume-title":"TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). USENIX Association, Carlsbad, CA. 578\u2013594. isbn:978-1-939133-08-3 https:\/\/www.usenix.org\/conference\/osdi18\/presentation\/chen"},{"key":"e_1_3_2_1_12_1","volume-title":"Learning to Optimize Tensor Programs. 31","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. Learning to Optimize Tensor Programs. 31 (2018), https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2018\/file\/8b5700012be65c9da25f49408d959ca0-Paper.pdf"},{"key":"e_1_3_2_1_13_1","volume-title":"Trevor Morris, Jorn Tuyls, Yi-Hsiang Lai, Jared Roesch, Elliott Delaye, Vin Sharma, and Yida Wang.","author":"Chen Zhi","year":"2021","unstructured":"Zhi Chen, Cody Hao Yu, Trevor Morris, Jorn Tuyls, Yi-Hsiang Lai, Jared Roesch, Elliott Delaye, Vin Sharma, and Yida Wang. 2021. Bring your own codegen to deep learning compiler. arXiv preprint arXiv:2105.03215."},{"key":"e_1_3_2_1_14_1","volume-title":"cuDNN: Efficient Primitives for Deep Learning. CoRR, abs\/1410.0759","author":"Chetlur Sharan","year":"2014","unstructured":"Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014. cuDNN: Efficient Primitives for Deep Learning. CoRR, abs\/1410.0759 (2014), arxiv:1410.0759."},{"key":"e_1_3_2_1_15_1","unstructured":"NVIDIA Corporation. 2024. TensorRT-LLM. https:\/\/github.com\/NVIDIA\/TensorRT-LLM\/tree\/v0.11.0"},{"key":"e_1_3_2_1_16_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. CoRR, abs\/1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. CoRR, abs\/1810.04805 (2018), arxiv:1810.04805. arxiv:1810.04805"},{"key":"e_1_3_2_1_17_1","unstructured":"Xinyang Geng and Hao Liu. 2023. OpenLLaMA: An Open Reproduction of LLaMA. https:\/\/github.com\/openlm-research\/open_llama"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3372419"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3559009.3569651"},{"key":"e_1_3_2_1_20_1","first-page":"14807","article-title":"Adatune: Adaptive tensor program compilation made efficient","volume":"33","author":"Li Menghao","year":"2020","unstructured":"Menghao Li, Minjia Zhang, Chi Wang, and Mingqin Li. 2020. Adatune: Adaptive tensor program compilation made efficient. Advances in Neural Information Processing Systems, 33 (2020), 14807\u201314819.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_21_1","unstructured":"LLVM. 2012. libclc. http:\/\/libclc.llvm.org"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00041"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICMLA.2012.214"},{"key":"e_1_3_2_1_24_1","unstructured":"NVIDIA. 2017. NVIDIA Tesla V100. https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf"},{"key":"e_1_3_2_1_25_1","volume-title":"CUTLASS: CUDA Templates for Linear Algebra. https:\/\/github.com\/NVIDIA\/cutlass","author":"NVIDIA.","year":"2020","unstructured":"NVIDIA. 2020. CUTLASS: CUDA Templates for Linear Algebra. https:\/\/github.com\/NVIDIA\/cutlass"},{"key":"e_1_3_2_1_26_1","unstructured":"NVIDIA. 2020. NVIDIA A100 Tensor Core GPU Architecture. https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/Data-Center\/nvidia-ampere-architecture-whitepaper.pdf"},{"key":"e_1_3_2_1_27_1","unstructured":"NVIDIA. 2021. NVIDIA RTX 3090 Family. https:\/\/www.nvidia.com\/en-us\/geforce\/graphics-cards\/30-series\/rtx-3090-3090ti\/"},{"key":"e_1_3_2_1_28_1","unstructured":"NVIDIA. 2021. NVIDIA RTX A6000 Graphics Cards. https:\/\/www.nvidia.com\/en-us\/design-visualization\/rtx-a6000\/"},{"key":"e_1_3_2_1_29_1","volume-title":"NVIDIA Nsight Systems. https:\/\/developer.nvidia.com\/nsight-systems Retrieved","author":"NVIDIA.","year":"2022","unstructured":"NVIDIA. 2022. NVIDIA Nsight Systems. https:\/\/developer.nvidia.com\/nsight-systems Retrieved September 8, 2022"},{"key":"e_1_3_2_1_30_1","unstructured":"NVIDIA. 2022. NVIDIA RTX 4090. https:\/\/www.nvidia.com\/en-us\/geforce\/graphics-cards\/40-series\/rtx-4090\/"},{"key":"e_1_3_2_1_31_1","volume-title":"Fitzek","author":"P\u00e9ter Vingelmann NVIDIA","year":"2020","unstructured":"NVIDIA, P\u00e9ter Vingelmann, and Frank H.P. Fitzek. 2020. CUDA, release: 10.2.89. https:\/\/developer.nvidia.com\/cuda-toolkit"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3609510.3609815"},{"key":"e_1_3_2_1_33_1","volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32, H. Wallach, H. Larochelle, A. Beygelzimer, F. d' Alch\u00e9-Buc, E. Fox, and R. Garnett (Eds.). Curran Associates, Inc., 8024\u20138035. http:\/\/papers.neurips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf"},{"key":"e_1_3_2_1_34_1","unstructured":"Alec Radford Jeff Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. https:\/\/api.semanticscholar.org\/CorpusID:160025533"},{"key":"e_1_3_2_1_35_1","unstructured":"Mohammed Rahman Louis-No\u00ebl Pouchet and Ponnuswamy Sadayappan. 2010. Neural Network Assisted Tile Size Selection. 11."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00016"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3497776.3517774"},{"key":"e_1_3_2_1_38_1","unstructured":"Ying Sheng Lianmin Zheng Binhang Yuan Zhuohan Li Max Ryabinin Beidi Chen Percy Liang Christopher Re Ion Stoica and Ce Zhang. 2023. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU. In Proceedings of the 40th International Conference on Machine Learning Andreas Krause Emma Brunskill Kyunghyun Cho Barbara Engelhardt Sivan Sabato and Jonathan Scarlett (Eds.) (Proceedings of Machine Learning Research Vol. 202). PMLR 31094\u201331116. https:\/\/proceedings.mlr.press\/v202\/sheng23a.html"},{"key":"e_1_3_2_1_39_1","unstructured":"Iulia Turc Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. Well-Read Students Learn Better: On the Importance of Pre-training Compact Models. arXiv preprint arXiv:1908.08962v2."},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2009.20"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1"},{"key":"e_1_3_2_1_42_1","unstructured":"Guangxuan Xiao Ji Lin Mickael Seznec Hao Wu Julien Demouth and Song Han. 2023. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. In Proceedings of the 40th International Conference on Machine Learning Andreas Krause Emma Brunskill Kyunghyun Cho Barbara Engelhardt Sivan Sabato and Jonathan Scarlett (Eds.) (Proceedings of Machine Learning Research Vol. 202). PMLR 38087\u201338099. https:\/\/proceedings.mlr.press\/v202\/xiao23c.html"},{"key":"e_1_3_2_1_43_1","first-page":"204","article-title":"Bolt: Bridging the gap between auto-tuners and hardware-native performance","volume":"4","author":"Xing Jiarong","year":"2022","unstructured":"Jiarong Xing, Leyuan Wang, Shang Zhang, Jack Chen, Ang Chen, and Yibo Zhu. 2022. Bolt: Bridging the gap between auto-tuners and hardware-native performance. Proceedings of Machine Learning and Systems, 4 (2022), 204\u2013216.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.14509993"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD58817.2023.00077"},{"key":"e_1_3_2_1_46_1","volume-title":"Oneflow: Redesign the distributed deep learning framework from scratch. arXiv preprint arXiv:2110.15032.","author":"Yuan Jinhui","year":"2021","unstructured":"Jinhui Yuan, Xinqi Li, Cheng Cheng, Juncheng Liu, Ran Guo, Shenghang Cai, Chi Yao, Fei Yang, Xiaodong Yi, and Chuan Wu. 2021. Oneflow: Redesign the distributed deep learning framework from scratch. arXiv preprint arXiv:2110.15032."},{"key":"e_1_3_2_1_47_1","unstructured":"Amir Zandieh Insu Han Majid Daliri and Amin Karbasi. 2023. KDEformer: Accelerating Transformers via Kernel Density Estimation. In Proceedings of the 40th International Conference on Machine Learning Andreas Krause Emma Brunskill Kyunghyun Cho Barbara Engelhardt Sivan Sabato and Jonathan Scarlett (Eds.) (Proceedings of Machine Learning Research Vol. 202). PMLR 40605\u201340623. https:\/\/proceedings.mlr.press\/v202\/zandieh23a.html"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3018743.3018755"},{"key":"e_1_3_2_1_49_1","volume-title":"Proceedings of Machine Learning and Systems, P. Gibbons, G. Pekhimenko, and C. De Sa (Eds.). 6, 196\u2013209","author":"Zhao Yilong","year":"2024","unstructured":"Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. 2024. Atom: Low-Bit Quantization for Efficient and Accurate LLM Serving. In Proceedings of Machine Learning and Systems, P. Gibbons, G. Pekhimenko, and C. De Sa (Eds.). 6, 196\u2013209. https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2024\/file\/5edb57c05c81d04beb716ef1d542fe9e-Paper-Conference.pdf"},{"key":"e_1_3_2_1_50_1","volume-title":"14th USENIX symposium on operating systems design and implementation (OSDI 20)","author":"Zheng Lianmin","year":"2020","unstructured":"Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, and Koushik Sen. 2020. Ansor: Generating $High-Performance$ tensor programs for deep learning. In 14th USENIX symposium on operating systems design and implementation (OSDI 20). 863\u2013879."},{"key":"e_1_3_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378508"},{"key":"e_1_3_2_1_52_1","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923)","author":"Zheng Zangwei","year":"2023","unstructured":"Zangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo, Xin Jiang, and Yang You. 2023. Response length perception and sequence scheduling: an LLM-empowered LLM inference pipeline. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923). Curran Associates Inc., Red Hook, NY, USA. Article 2859, 14 pages."},{"key":"e_1_3_2_1_53_1","volume-title":"18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Zhong Yinmin","year":"2024","unstructured":"Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024. $DistServe$: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). 193\u2013210."}],"event":{"name":"CGO '25: 23rd ACM\/IEEE International Symposium on Code Generation and Optimization","location":"Las Vegas NV USA","acronym":"CGO '25","sponsor":["SIGPLAN SIGPLAN Programming Languages","SIGMICRO SIGMICRO Microarchitecture","IEEE Computer Society IEEE Computer Society"]},"container-title":["Proceedings of the 23rd ACM\/IEEE International Symposium on Code Generation and Optimization"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3696443.3708944","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:10:13Z","timestamp":1750295413000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3696443.3708944"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3]]},"references-count":53,"alternative-id":["10.1145\/3696443.3708944","10.1145\/3696443"],"URL":"https:\/\/doi.org\/10.1145\/3696443.3708944","relation":{},"subject":[],"published":{"date-parts":[[2025,3]]},"assertion":[{"value":"2025-03-01","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}