{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T20:06:33Z","timestamp":1783713993768,"version":"3.55.0"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"name":"NSF","award":["#2326894 and #2425655"],"award-info":[{"award-number":["#2326894 and #2425655"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>\n            Driven by the increasing demands of machine learning, heterogeneous systems combining CPUs and GPUs have emerged as the dominant architecture for parallel computing in recent years. To optimize memory management and data transfer between CPUs and GPUs, Nvidia GPUs have introduced unified virtual memory (\n            <jats:italic toggle=\"yes\">UVM<\/jats:italic>\n            ) and pinned memory (\n            <jats:italic toggle=\"yes\">PM<\/jats:italic>\n            ) over the last decade.\n            <jats:italic toggle=\"yes\">UVM<\/jats:italic>\n            can avoid explicit memory copies and potentially overlap GPU kernel computations with CPU-GPU data transfer.\n            <jats:italic toggle=\"yes\">PM<\/jats:italic>\n            ensures that data with high locality remains in the main memory, preventing it from being paged out. In addition to these two techniques, asynchronous memory copy (\n            <jats:italic toggle=\"yes\">Async Memcpy<\/jats:italic>\n            ) was introduced recently in Nvidia GPUs to improve the CPU-GPU pipeline further. By utilizing\n            <jats:italic toggle=\"yes\">Async Memcpy<\/jats:italic>\n            , the data transfer from GPU global memory to shared memory can be overlapped with GPU computations, adding an additional stage to the CPU-GPU data transfer pipeline. A thorough performance analysis of how\n            <jats:italic toggle=\"yes\">Async Memcpy<\/jats:italic>\n            affects the current\n            <jats:italic toggle=\"yes\">UVM<\/jats:italic>\n            and\n            <jats:italic toggle=\"yes\">PM<\/jats:italic>\n            CPU-GPU data transfer scheme is desired.\n          <\/jats:p>\n          <jats:p>\n            In this article, we provide performance implications of the combined effect of\n            <jats:italic toggle=\"yes\">UVM<\/jats:italic>\n            ,\n            <jats:italic toggle=\"yes\">PM<\/jats:italic>\n            , and\n            <jats:italic toggle=\"yes\">Async Memcpy<\/jats:italic>\n            , exploring which applications benefit from which combination of these features. We implement all these features on a suite of 25 workloads, including microbenchmarks and realworld applications. We observe an average performance gain of 24% when utilizing\n            <jats:italic toggle=\"yes\">UVM<\/jats:italic>\n            and a 34% gain when employing\n            <jats:italic toggle=\"yes\">PM<\/jats:italic>\n            on realworld applications, compared to not applying any data transfer optimization techniques. The performance benefits of\n            <jats:italic toggle=\"yes\">Async Memcpy<\/jats:italic>\n            vary across different workloads. For workloads featuring extensive shared memory usage and high compute density (e.g.,\n            <jats:italic toggle=\"yes\">kmeans<\/jats:italic>\n            and\n            <jats:italic toggle=\"yes\">lud<\/jats:italic>\n            ),\n            <jats:italic toggle=\"yes\">Async Memcpy<\/jats:italic>\n            delivers around a 20% performance improvement over using\n            <jats:italic toggle=\"yes\">UVM<\/jats:italic>\n            or\n            <jats:italic toggle=\"yes\">PM<\/jats:italic>\n            alone. In other workloads like\n            <jats:italic toggle=\"yes\">knn<\/jats:italic>\n            , we note a 20% performance degradation when using\n            <jats:italic toggle=\"yes\">Async Memcpy<\/jats:italic>\n            . Furthermore, we conduct an in-depth investigation of the GPU kernel using performance counters to uncover the root causes of performance differences among various data transfer models. We also perform sensitivity analyses to examine how the number of blocks and threads, as well as the L1-cache\/shared memory partitioning, impact performance. We explore future research directions aimed at enhancing the data transfer pipeline by overlapping memory allocation with data transfer and computation across GPU kernels.\n          <\/jats:p>","DOI":"10.1145\/3746231","type":"journal-article","created":{"date-parts":[[2025,6,26]],"date-time":"2025-06-26T07:02:02Z","timestamp":1750921322000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Performance Implications of Pipelining the Data Transfer in CPU-GPU Heterogeneous Systems"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7092-2401","authenticated-orcid":false,"given":"Ruihao","family":"Li","sequence":"first","affiliation":[{"name":"Electrical and Computer Engineering, The University of Texas at Austin","place":["Austin, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8485-581X","authenticated-orcid":false,"given":"Bagus","family":"Hanindhito","sequence":"additional","affiliation":[{"name":"The University of Texas at Austin","place":["Austin, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7880-9386","authenticated-orcid":false,"given":"Sanjana","family":"Yadav","sequence":"additional","affiliation":[{"name":"The University of Texas at Austin","place":["Austin, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7988-1431","authenticated-orcid":false,"given":"Qinzhe","family":"Wu","sequence":"additional","affiliation":[{"name":"The University of Texas at Austin","place":["Austin, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0876-166X","authenticated-orcid":false,"given":"Krishna","family":"Kavi","sequence":"additional","affiliation":[{"name":"University of North Texas","place":["Denton, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7754-1874","authenticated-orcid":false,"given":"Gayatri","family":"Mehta","sequence":"additional","affiliation":[{"name":"University of North Texas","place":["Denton, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7556-3069","authenticated-orcid":false,"given":"Neeraja J.","family":"Yadwadkar","sequence":"additional","affiliation":[{"name":"The University of Texas at Austin","place":["Austin, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8747-5214","authenticated-orcid":false,"given":"Lizy K.","family":"John","sequence":"additional","affiliation":[{"name":"The University of Texas at Austin","place":["Austin, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,19]]},"reference":[{"key":"e_1_3_3_2_2","first-page":"265","volume-title":"Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16)","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et\u00a0al. 2016. \\(\\lbrace\\) TensorFlow \\(\\rbrace\\) : A system for \\(\\lbrace\\) Large-Scale \\(\\rbrace\\) machine learning. In Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). 265\u2013283."},{"key":"e_1_3_3_3_2","first-page":"117","volume-title":"Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Agrawal Amey","year":"2024","unstructured":"Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024. Taming \\(\\lbrace\\) throughput-latency \\(\\rbrace\\) tradeoff in \\(\\lbrace\\) LLM \\(\\rbrace\\) inference with \\(\\lbrace\\) sarathi-serve \\(\\rbrace\\) . In Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). 117\u2013134."},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTR.2009.5289124"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","unstructured":"Tyler Allen Bennett Cooper and Rong Ge. 2024. Fine-grain quantitative analysis of demand paging in unified virtual memory. ACM Trans. Archit. Code Optim. 21 1 Article 14 (2024) 24 pages. 10.1145\/3632953","DOI":"10.1145\/3632953"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS49936.2021.00023"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3480855"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3620665.3640366"},{"key":"e_1_3_3_9_2","first-page":"26","volume-title":"Proceedings of the 2020 IEEE\/ACM Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS)","author":"Anzt Hartwig","year":"2020","unstructured":"Hartwig Anzt, Yuhsiang M Tsai, Ahmad Abdelfattah, Terry Cojean, and Jack Dongarra. 2020. Evaluating the performance of NVIDIA\u2019s A100 ampere GPU for sparse and batched computations. In Proceedings of the 2020 IEEE\/ACM Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS). IEEE, 26\u201338."},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/INTERCON.2019.8853624"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","unstructured":"R. Bhargava and K. Troester. 2024. AMD next-generation \u201cZen 4\u201d core and 4th gen AMD EPYC Server CPUs. In IEEE Micro 44 3 (2024) 8\u201317. DOI:10.1109\/MM.2024.3375070","DOI":"10.1109\/MM.2024.3375070"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/HCS49909.2020.9220622"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2021.3061394"},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3092255.3092256"},{"key":"e_1_3_3_16_2","unstructured":"Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman Akhil Mathur Alan Schelten Amy Yang Angela Fan et\u00a0al. 2024. The llama 3 herd of models. arXiv:2407.21783. Retrieved from https:\/\/arxiv.org\/abs\/2407.21783 (2024)."},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.5555\/2015039.2015535"},{"key":"e_1_3_3_18_2","unstructured":"Yongbin Gu Wenxuan Wu Yunfan Li and Lizhong Chen. 2020. Uvmbench: A comprehensive benchmark suite for researching unified virtual memory in gpus. arXiv:2007.09822. Retrieved from https:\/\/arxiv.org\/abs\/2007.09822 (2020)."},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3578244.3583736"},{"key":"e_1_3_3_20_2","unstructured":"Mark Harris. 2012. How to Optimize Data Transfers in CUDA C\/C++. Retrieved from https:\/\/developer.nvidia.com\/blog\/how-optimize-data-transfers-cuda-cc\/"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/2806777.2806836"},{"key":"e_1_3_3_22_2","unstructured":"Guyue Huang Yang Bai Liu Liu Yuke Wang Bei Yu Yufei Ding and Yuan Xie. 2023. Alcop: Automatic load-compute pipelining in deep learning compiler for ai-gpus. Proceedings of Machine Learning and Systems 5 (2023) 680\u2013694."},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378529"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480105"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613165"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10071063"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC59245.2023.00024"},{"key":"e_1_3_3_28_2","first-page":"663","volume-title":"Proceedings of the 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23)","author":"Li Zhuohan","year":"2023","unstructured":"Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E Gonzalez, and Ion Stoica. 2023. \\(\\lbrace\\) AlpaServe \\(\\rbrace\\) : Statistical multiplexing with model parallelism for deep learning serving. In Proceedings of the 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23). 663\u2013679."},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389705"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2021.3086541"},{"key":"e_1_3_3_31_2","article-title":"Unified memory in cuda 6.0. a brief overview of related data access and transfer issues","author":"Negrut Dan","year":"2014","unstructured":"Dan Negrut, Radu Serban, Ang Li, and Andrew Seidl. 2014. Unified memory in cuda 6.0. a brief overview of related data access and transfer issues. SBEL, Madison, WI, USA, Technical Reports TR-2014-09 (2014).","journal-title":"SBEL, Madison, WI, USA, Technical Reports TR-2014-09"},{"key":"e_1_3_3_32_2","unstructured":"Nvidia. 2013. Nvidia Tesla K40. Retrieved from https:\/\/www.nvidia.com\/content\/PDF\/kepler\/nvidia-tesla-k40.pdf"},{"key":"e_1_3_3_33_2","unstructured":"Nvidia. 2024. CUPTI. Retrieved from https:\/\/docs.nvidia.com\/cuda\/cupti\/"},{"key":"e_1_3_3_34_2","unstructured":"Nvidia. 2024. CUTLASS 3.0. Retrieved from https:\/\/github.com\/NVIDIA\/cutlass\/"},{"key":"e_1_3_3_35_2","unstructured":"Nvidia. 2024. Nsight Compute. Retrieved from https:\/\/developer.nvidia.com\/nsight-compute\/"},{"key":"e_1_3_3_36_2","unstructured":"Nvidia. 2024. Nsight Systems. Retrieved from https:\/\/developer.nvidia.com\/nsight-systems\/"},{"key":"e_1_3_3_37_2","unstructured":"Nvidia. 2024. Nvidia A100 GPU Architecture. Retrieved from https:\/\/images.nvidia.com\/aem-dam\/en-zz\/Solutions\/data-center\/nvidia-ampere-architecture-whitepaper.pdf"},{"key":"e_1_3_3_38_2","unstructured":"Nvidia. 2024. Nvidia H100 GPU Architecture. Retrieved from https:\/\/resources.nvidia.com\/en-us-tensor-core\/"},{"key":"e_1_3_3_39_2","unstructured":"Nvidia. 2024. Nvidia P100 GPU Architecture. Retrieved from https:\/\/images.nvidia.com\/content\/pdf\/tesla\/whitepaper\/pascal-architecture-whitepaper.pdf"},{"key":"e_1_3_3_40_2","unstructured":"Nvidia. 2024. Nvidia V100 GPU Architecture. Retrieved from https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2018.00032"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3297663.3310299"},{"key":"e_1_3_3_43_2","unstructured":"Nathan Pemberton Anton Zabreyko Zhoujie Ding Randy Katz and Joseph Gonzalez. 2022. Kernel-as-a-service: A serverless interface to GPUs. arXiv:2212.08146. Retrieved from https:\/\/arxiv.org\/abs\/2212.08146 (2022)."},{"key":"e_1_3_3_44_2","unstructured":"L.-N. Pouchet. 2012. Polybench: The polyhedral benchmark suite. Retrieved from http:\/\/www.cs.ucla.edu\/%7Epouchet\/software\/polybench"},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476205"},{"key":"e_1_3_3_46_2","unstructured":"Joseph Redmon. 2013\u20132016. Darknet: Open Source Neural Networks in C. Retrieved from http:\/\/pjreddie.com\/darknet\/."},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2015.41"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3342195.3387537"},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3489525.3511691"},{"key":"e_1_3_3_50_2","first-page":"200","volume-title":"Proceedings of the 2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS)","author":"Shriram SB","year":"2019","unstructured":"SB Shriram, Anshuj Garg, and Purushottam Kulkarni. 2019. Dynamic memory management for gpu-based training of deep neural networks. In Proceedings of the 2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 200\u2013209."},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3627703.3629578"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3468044.3468053"},{"key":"e_1_3_3_53_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et\u00a0al. 2023. Llama: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971 (2023)."},{"key":"e_1_3_3_54_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et\u00a0al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288 (2023)."},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3178487.3178491"},{"key":"e_1_3_3_56_2","first-page":"779","volume-title":"Proceedings of the 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23)","author":"Wang Yuke","year":"2023","unstructured":"Yuke Wang, Boyuan Feng, Zheng Wang, Tong Geng, Kevin Barker, Ang Li, and Yufei Ding. 2023. MGG: Accelerating graph neural networks with fine-grained intra-kernel communication-computation pipelining on multi-GPU platforms. In Proceedings of the 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23). 779\u2013795."},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3545008.3545084"},{"key":"e_1_3_3_58_2","first-page":"521","volume-title":"Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)","author":"Yu Gyeong-In","year":"2022","unstructured":"Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022. Orca: A distributed serving system for \\(\\lbrace\\) transformer-based \\(\\rbrace\\) generative models. In Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). 521\u2013538."},{"key":"e_1_3_3_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2016.7446077"},{"key":"e_1_3_3_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC55918.2022.00013"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746231","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,20]],"date-time":"2025-09-20T00:48:23Z","timestamp":1758329303000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746231"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,19]]},"references-count":59,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3746231"],"URL":"https:\/\/doi.org\/10.1145\/3746231","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,19]]},"assertion":[{"value":"2024-09-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-19","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-19","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}