{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T16:46:34Z","timestamp":1782405994549,"version":"3.54.5"},"reference-count":50,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100003052","name":"Ministry of Trade, Industry and Energy","doi-asserted-by":"publisher","award":["10077609"],"award-info":[{"award-number":["10077609"]}],"id":[{"id":"10.13039\/501100003052","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003052\"","name":"Korean National Research Foundation","doi-asserted-by":"publisher","award":["21A20151113068"],"award-info":[{"award-number":["21A20151113068"]}],"id":[{"id":"10.13039\/501100003052\"","id-type":"DOI","asserted-by":"publisher"}]},{"name":"EU Horizon Program","award":["101298393\""],"award-info":[{"award-number":["101298393\""]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Large language models (LLMs) have recently achieved remarkable performance in text generation, capturing the attention of a broad audience. This success, driven by the rapid growth in model parameters, comes at the expense of significantly higher operational costs and decreased processing speed. These costs, combined with privacy concerns around cloud-based deployments, have motivated research into running LLMs on commodity hardware. For example, researchers have used the memory hierarchy to boost throughput by increasing the number of batches. These studies, however, tend to overlook or inefficiently utilize the additional computational resources provided by the CPU.<\/jats:p>\n                  <jats:p>In this work, we present a dynamic workload allocation technique that efficiently distributes computation across all available hardware resources. The proposed method targets decoder-based models on standard general-purpose hardware, effectively minimizing idle periods for both the CPU and the GPU. Experiments show that our approach achieves up to 30% higher throughput compared to the state of the art, regardless of model architecture, LLM optimizations, and input batch sizes.<\/jats:p>","DOI":"10.1145\/3816433","type":"journal-article","created":{"date-parts":[[2026,5,23]],"date-time":"2026-05-23T09:14:34Z","timestamp":1779527674000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["CPU-GPU Workload Distribution during Throughput-Oriented LLM Inference on Single-GPU Systems"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2312-3049","authenticated-orcid":false,"given":"Daon","family":"Park","sequence":"first","affiliation":[{"name":"Seoul National University","place":["Gwanak-gu, Korea (the Republic of)"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6645-6161","authenticated-orcid":false,"given":"Bernhard","family":"Egger","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Seoul National University","place":["Seoul, Korea (the Republic of)"]},{"name":"Lucerne University of Applied Sciences and Arts","place":["Seoul, Korea (the Republic of)"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.3390\/s21144758"},{"key":"e_1_3_2_3_2","unstructured":"Mistral AI. 2023. Mistral 7B. arxiv:2310.06825 [cs.CL]. Retrieved from https:\/\/arxiv.org\/abs\/2310.06825"},{"key":"e_1_3_2_4_2","unstructured":"Mistral AI. 2024. Mixtral of Experts. arxiv:2401.04088 [cs.LG]. Retrieved from https:\/\/arxiv.org\/abs\/2401.04088"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.298"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41404.2022.00051"},{"key":"e_1_3_2_7_2","unstructured":"Alexey Bochkovskiy Chien-Yao Wang and Hong-Yuan Mark Liao. 2020. YOLOv4: Optimal Speed and Accuracy of Object Detection. arxiv:2004.10934 [cs.CV]. Retrieved from https:\/\/arxiv.org\/abs\/2004.10934"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1111\/nyas.15007"},{"key":"e_1_3_2_9_2","first-page":"1877","volume-title":"Advances in Neural Information Processing Systems","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.). Vol. 33. Curran Associates, Inc., 1877\u20131901."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3243176.3243210"},{"key":"e_1_3_2_11_2","first-page":"32","volume-title":"Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming","author":"Cho Younghyun","year":"2022","unstructured":"Younghyun Cho, Jiyeon Park, Florian Negele, Changyeon Jo, Thomas R Gross, and Bernhard Egger. 2022. Dopia: Online parallelism management for integrated CPU\/GPU architectures. In Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. 32\u201345."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1285"},{"key":"e_1_3_2_13_2","series-title":"NIPS\u201922","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems","author":"Dao Tri","year":"2022","unstructured":"Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher R\u00e9. 2022. Flashattention: Fast and memory-efficient exact attention with io-awareness. In Proceedings of the 36th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS\u201922). Curran Associates Inc., Red Hook, NY, USA, Article 1189, 16 pages."},{"key":"e_1_3_2_14_2","volume-title":"International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations."},{"key":"e_1_3_2_15_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Ge Suyu","year":"2023","unstructured":"Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao. 2023. Model tells you what to discard: Adaptive KV cache compression for LLMs. In The Twelfth International Conference on Learning Representations."},{"key":"e_1_3_2_16_2","unstructured":"Georgi Gerganov. 2023. llama.cpp Github. Retrieved May 31 2026 from https:\/\/github.com\/ggml-org\/llama.cpp"},{"key":"e_1_3_2_17_2","unstructured":"Connor Holmes Masahiro Tanaka Michael Wyatt Ammar Ahmad Awan Jeff Rasley Samyam Rajbhandari Reza Yazdani Aminabadi Heyang Qin Arash Bakhtiari Lev Kurilenko and Yuxiong He. 2024. DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference. arxiv:2401.08671 [cs.PF]. Retrieved from https:\/\/arxiv.org\/abs\/2401.08671"},{"key":"e_1_3_2_18_2","article-title":"HuggingFace Accelerate","year":"2022","unstructured":"HuggingFace. 2022. HuggingFace Accelerate. Retrieved May 31, 2026 from https:\/\/huggingface.co\/docs\/accelerate\/index","journal-title":"https:\/\/huggingface.co\/docs\/accelerate\/index"},{"key":"e_1_3_2_19_2","first-page":"23901","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Kim Sehoon","year":"2024","unstructured":"Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W Mahoney, and Kurt Keutzer. 2024. SqueezeLLM: Dense-and-sparse quantization. In Proceedings of the 41st International Conference on Machine Learning. 23901\u201323923."},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3656019.3676945"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613165"},{"key":"e_1_3_2_22_2","first-page":"155","volume-title":"18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Lee Wonbeom","year":"2024","unstructured":"Wonbeom Lee, Jungi Lee, Junghwan Seo, and Jaewoong Sim. 2024. InfiniGen: Efficient generative inference of large language models with dynamic KV cache management. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). 155\u2013172."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.3390\/s21041399"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3560815"},{"key":"e_1_3_2_25_2","series-title":"NIPS\u201923","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Liu Zichang","year":"2023","unstructured":"Zichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang, Victor Xie, Zhaozhuo Xu, Anastasios Kyrillidis, and Anshumali Shrivastava. 2023. Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS\u201923). Curran Associates Inc., Red Hook, NY, USA, Article 2279, 52342\u201352364."},{"key":"e_1_3_2_26_2","volume-title":"International Conference on Learning Representations","author":"Merity Stephen","year":"2017","unstructured":"Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer sentinel mixture models. In International Conference on Learning Representations."},{"key":"e_1_3_2_27_2","unstructured":"Meta. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. arxiv:2307.09288 [cs.CL]. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.14778\/3574245.3574258"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1206"},{"key":"e_1_3_2_30_2","article-title":"How to Optimize Data Transfers in CUDA C\/C++","year":"2012","unstructured":"Nvidia. 2012. How to Optimize Data Transfers in CUDA C\/C++. Retrieved May 31, 2026 from https:\/\/developer.nvidia.com\/blog\/how-optimize-data-transfers-cuda-cc\/","journal-title":"https:\/\/developer.nvidia.com\/blog\/how-optimize-data-transfers-cuda-cc\/"},{"key":"e_1_3_2_31_2","article-title":"CUDA Streams Simplify Concurrency","year":"2015","unstructured":"Nvidia. 2015. CUDA Streams Simplify Concurrency. Retrieved May 31, 2026 from https:\/\/developer.nvidia.com\/blog\/gpu-pro-tip-cuda-7-streams-simplify-concurrency","journal-title":"https:\/\/developer.nvidia.com\/blog\/gpu-pro-tip-cuda-7-streams-simplify-concurrency"},{"key":"e_1_3_2_32_2","article-title":"Magnum IO GPUDirect Storage","year":"2023","unstructured":"Nvidia. 2023. Magnum IO GPUDirect Storage. Retrieved May 31, 2026 from https:\/\/developer.nvidia.com\/gpudirect-storage","journal-title":"https:\/\/developer.nvidia.com\/gpudirect-storage"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3620665.3640383"},{"key":"e_1_3_2_34_2","unstructured":"OpenAI. 2024. GPT-4 Technical Report. arxiv:2303.08774 [cs.CL]. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3656019.3676949"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3609510.3609815"},{"key":"e_1_3_2_37_2","series-title":"ICML\u201924","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Park Yeonhong","year":"2024","unstructured":"Yeonhong Park, Jake Hyun, SangLyul Cho, Bonggeun Sim, and Jae W. Lee. 2024. Any-precision LLM: Low-cost deployment of multiple, different-sized LLMs. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML\u201924). JMLR.org, Article 1607, 20 pages."},{"key":"e_1_3_2_38_2","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K\u00f6pf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An imperative style, high-performance deep learning library. In Proceedings of the 33rd International Conference on Neural Information Processing Systems."},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1101\/2021.02.12.430858"},{"key":"e_1_3_2_40_2","unstructured":"Noam Shazeer. 2019. Fast Transformer Decoding: One Write-Head is All You Need. arxiv:1911.02150 [cs.NE]. Retrieved from https:\/\/arxiv.org\/abs\/1911.02150"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6409"},{"key":"e_1_3_2_42_2","first-page":"31094","volume-title":"International Conference on Machine Learning","author":"Sheng Ying","year":"2023","unstructured":"Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher R\u00e9, Ion Stoica, and Ce Zhang. 2023. Flexgen: High-throughput generative inference of large language models with a single GPU. In International Conference on Machine Learning. PMLR, 31094\u201331116."},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3694715.3695964"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357895"},{"key":"e_1_3_2_45_2","first-page":"38087","volume-title":"International Conference on Machine Learning","author":"Xiao Guangxuan","year":"2023","unstructured":"Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. Smoothquant: Accurate and efficient post-training quantization for large language models. In International Conference on Machine Learning. PMLR, 38087\u201338099."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3688351.3689164"},{"key":"e_1_3_2_47_2","first-page":"521","volume-title":"16th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201922)","author":"Yu Gyeong-In","year":"2022","unstructured":"Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022. Orca: A distributed serving system for transformer-based generative models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201922). 521\u2013538."},{"key":"e_1_3_2_48_2","unstructured":"Susan Zhang Stephen Roller Naman Goyal Mikel Artetxe Moya Chen Shuohui Chen Christopher Dewan Mona Diab Xian Li Xi Victoria Lin Todor Mihaylov Myle Ott Sam Shleifer Kurt Shuster Daniel Simig Punit Singh Koura Anjali Sridhar Tianlu Wang and Luke Zettlemoyer. 2022. OPT: Open Pre-trained Transformer Language Models. arxiv:2205.01068 [cs.CL]. Retrieved from https:\/\/arxiv.org\/abs\/2205.01068"},{"key":"e_1_3_2_49_2","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Zhang Zhenyu","year":"2023","unstructured":"Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher R\u00e9, Clark Barrett, Zhangyang Wang, and Beidi Chen. 2023. H2O: Heavy-hitter oracle for efficient generative inference of large language models. In Proceedings of the 37th International Conference on Neural Information Processing Systems."},{"key":"e_1_3_2_50_2","first-page":"162","volume-title":"Proceedings of Machine Learning and Systems","volume":"6","author":"Zhao Xuanlei","year":"2024","unstructured":"Xuanlei Zhao, Bin Jia, Haotian Zhou, Ziming Liu, Shenggan Cheng, and Yang You. 2024. HeteGen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices. In Proceedings of Machine Learning and Systems, P. Gibbons, G. Pekhimenko, and C. De Sa (Eds.). Vol. 6. 162\u2013172."},{"key":"e_1_3_2_51_2","series-title":"NIPS\u201924","volume-title":"Proceedings of the 38th International Conference on Neural Information Processing Systems","author":"Zheng Lianmin","year":"2024","unstructured":"Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. SGLang: Efficient execution of structured language model programs. In Proceedings of the 38th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS\u201924). Curran Associates Inc., Red Hook, NY, USA, Article 2000, 27 pages."}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3816433","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:53:26Z","timestamp":1782402806000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3816433"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":50,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3816433"],"URL":"https:\/\/doi.org\/10.1145\/3816433","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2025-08-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-02","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}