{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,4]],"date-time":"2026-08-04T15:49:40Z","timestamp":1785858580528,"version":"3.56.0"},"reference-count":165,"publisher":"Association for Computing Machinery (ACM)","issue":"12","license":[{"start":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T00:00:00Z","timestamp":1780963200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/legalcode"}],"funder":[{"name":"European Research Council","award":["949587"],"award-info":[{"award-number":["949587"]}]},{"name":"European Union\u2019s Horizon Europe","award":["101175702"],"award-info":[{"award-number":["101175702"]}]},{"name":"Sapienza University","award":["ADAGIO and D2QNeT"],"award-info":[{"award-number":["ADAGIO and D2QNeT"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,9,30]]},"abstract":"<jats:p>In recent years, GPUs have become the preferred accelerators for HPC and ML applications due to their parallelism and high memory bandwidth. While GPUs boost computation, inter-GPU communication can create scalability bottlenecks, especially as the number of GPUs per node and cluster grows. Traditionally, the CPU managed multi-GPU communication, but advancements in GPU-centric communication now challenge this CPU dominance by reducing its involvement, granting GPUs more autonomy in communication tasks, and addressing mismatches in multi-GPU communication and computation.<\/jats:p>\n                  <jats:p>This article provides a landscape of GPU-centric communication, focusing on vendor mechanisms and user-level library supports. It aims to clarify the complexities and diverse options in this field, define the terminology, and categorize existing approaches within and across nodes. The article discusses vendor-provided mechanisms for communication and memory management in multi-GPU execution and reviews major communication libraries, their benefits, challenges, and performance insights. Then, it explores key research paradigms, future outlooks, and open research questions. By extensively describing GPU-centric communication techniques across the software and hardware stacks, we provide researchers, programmers, engineers, and library designers insights on how to exploit multi-GPU systems at their best.<\/jats:p>","DOI":"10.1145\/3813799","type":"journal-article","created":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T11:28:14Z","timestamp":1777894094000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["The Landscape of GPU-Centric Communication"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2351-0770","authenticated-orcid":false,"given":"Didem","family":"Unat","sequence":"first","affiliation":[{"name":"Computer Science and Engineering, Ko\u00e7 University","place":["Istanbul, Turkey"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0384-6330","authenticated-orcid":false,"given":"Ilyas","family":"Turimbetov","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, Ko\u00e7 University","place":["Istanbul, Turkey"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-1920-6773","authenticated-orcid":false,"given":"Mohammed","family":"Issa","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, Ko\u00e7 University","place":["Istanbul, Turkey"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9603-2466","authenticated-orcid":false,"given":"Dogan","family":"Sagbili","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, Ko\u00e7 University","place":["Istanbul, Turkey"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5676-9228","authenticated-orcid":false,"given":"Flavio","family":"Vella","sequence":"additional","affiliation":[{"name":"University of Trento","place":["Trento, Italy"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7244-639X","authenticated-orcid":false,"given":"Daniele","family":"Sensi","sequence":"additional","affiliation":[{"name":"University of Rome La Sapienza","place":["Rome, Italy"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1831-5556","authenticated-orcid":false,"given":"Ismayil","family":"Ismayilov","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, Ko\u00e7 University","place":["Istanbul, Turkey"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,9]]},"reference":[{"key":"e_1_3_2_2_2","article-title":"Inline GPU Packet Processing with NVIDIA DOCA GPUNetIO","author":"Agostini Elena","year":"2023","unstructured":"Elena Agostini. 2023. Inline GPU Packet Processing with NVIDIA DOCA GPUNetIO. Retrieved May 13, 2026 from https:\/\/developer.nvidia.com\/blog\/inline-gpu-packet-processing-with-nvidia-doca-gpunetio\/","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_3_2","article-title":"Optimizing Inline Packet Processing Using DPDK and GPUDirect with GPUs","author":"Agostini Elena","year":"2023","unstructured":"Elena Agostini. 2023. Optimizing Inline Packet Processing Using DPDK and GPUDirect with GPUs. Retrieved May 13, 2026 from https:\/\/developer.nvidia.com\/blog\/optimizing-inline-packet-processing-using-dpdk-and-gpudev-with-gpus\/","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_4_2","series-title":"CCGrid\u201917","first-page":"248","volume-title":"Proceedings of the 17th IEEE\/ACM International Symposium on Cluster, Cloud, and Grid Computing","author":"Agostini Elena","year":"2017","unstructured":"Elena Agostini, Davide Rossetti, and Sreeram Potluri. 2017. Offloading communication control logic in GPU accelerated applications. In Proceedings of the 17th IEEE\/ACM International Symposium on Cluster, Cloud, and Grid Computing (Madrid, Spain) (CCGrid\u201917). Institute for Electrical and Electronics Engineers, New York, NY, USA, 248\u2013257. DOI:10.1109\/CCGRID.2017.29"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2017.12.007"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1007\/978-3-030-71058-3_10","volume-title":"Proceedings of the Benchmarking, Measuring, and Optimizing","author":"Akhtar Palwisha","year":"2021","unstructured":"Palwisha Akhtar, Erhan Tezcan, Fareed Mohammad Qararyah, and Didem Unat. 2021. ComScribe: Identifying intra-node GPU communication. In Proceedings of the Benchmarking, Measuring, and Optimizing, Felix Wolf and Wanling Gao (Eds.). Springer International Publishing, Cham, 157\u2013174."},{"key":"e_1_3_2_7_2","first-page":"1","volume-title":"Proceedings of the ISC High Performance 2025 Research Paper Proceedings (40th International Conference)","author":"Alperen Abdullah","year":"2025","unstructured":"Abdullah Alperen, Nan Ding, Khaled Z. Ibrahim, Pieter Maris, Leonid Oliker, Chao Yang, and Hasan Metin Aktulga. 2025. Optimizing nuclear configuration interaction calculations on GPUs: A comparative performance study of programming models. In Proceedings of the ISC High Performance 2025 Research Paper Proceedings (40th International Conference). 1\u201312."},{"key":"e_1_3_2_8_2","article-title":"AMD Instinct MI200 Instruction Set Architecture","unstructured":"AMD. [n. d.]. AMD Instinct MI200 Instruction Set Architecture. Retrieved May 13, 2026 from https:\/\/www.amd.com\/system\/files\/TechDocs\/instinct-mi200-cdna2-instruction-set-architecture.pdf","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_9_2","article-title":"AMD CDNA\u21222 ARCHITECTURE","year":"2021","unstructured":"AMD. 2021. AMD CDNA\u21222 ARCHITECTURE. Retrieved May 13, 2026 from https:\/\/www.amd.com\/system\/files\/documents\/amd-cdna2-white-paper.pdf","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_10_2","article-title":"GPU-aware MPI with ROCm","year":"2023","unstructured":"AMD. 2023. GPU-aware MPI with ROCm. Retrieved May 13, 2026 from https:\/\/gpuopen.com\/learn\/amd-lab-notes\/amd-lab-notes-gpu-aware-mpi-readme\/#","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_11_2","article-title":"ROCK-Kernel-Driver","year":"2023","unstructured":"AMD. 2023. ROCK-Kernel-Driver. Retrieved May 13, 2026 from https:\/\/github.com\/RadeonOpenCompute\/ROCK-Kernel-Driver","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_12_2","article-title":"ROCm Documentation: GPU-Enabled MPI","year":"2023","unstructured":"AMD. 2023. ROCm Documentation: GPU-Enabled MPI. Retrieved May 13, 2026 from https:\/\/rocm.docs.amd.com\/en\/latest\/how_to\/gpu_aware_mpi.html","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_13_2","article-title":"ROCnRDMA","year":"2023","unstructured":"AMD. 2023. ROCnRDMA. Retrieved May 13, 2026 from https:\/\/github.com\/rocmarchive\/ROCnRDMA","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_14_2","article-title":"ROC_SHMEM","year":"2023","unstructured":"AMD. 2023. ROC_SHMEM. Retrieved May 13, 2026 from https:\/\/github.com\/ROCm-Developer-Tools\/ROC_SHMEM","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_15_2","article-title":"RCCL Developer Guide","year":"2025","unstructured":"AMD. 2025. RCCL Developer Guide. Retrieved May 13, 2026 from https:\/\/rocm.docs.amd.com\/projects\/rccl\/en\/latest\/","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_16_2","article-title":"RCCL Tests","year":"2025","unstructured":"AMD. 2025. RCCL Tests. Retrieved May 13, 2026 from https:\/\/github.com\/ROCm\/rccl-tests","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_17_2","first-page":"82","volume-title":"Proceedings of the 9th Annual Workshop on General Purpose Processing Using Graphics Processing Unit (GPGPU\u201916)","author":"Banerjee Dip Sankar","year":"2016","unstructured":"Dip Sankar Banerjee, Khaled Hamidouche, and Dhabaleswar K. Panda. 2016. Designing high performance communication runtime for GPU managed memory: Early experiences. In Proceedings of the 9th Annual Workshop on General Purpose Processing Using Graphics Processing Unit (GPGPU\u201916). Association for Computing Machinery, New York, NY, USA, 82\u201391. DOI:10.1145\/2884045.2884050"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","first-page":"1129","DOI":"10.1109\/SCW63240.2024.00155","volume-title":"Proceedings of the SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Baydamirli Javid","year":"2024","unstructured":"Javid Baydamirli, Tal Ben Nun, and Didem Unat. 2024. Autonomous Execution for Multi-GPU Systems: Compiler Support. In Proceedings of the SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1129\u20131140. DOI:10.1109\/SCW63240.2024.00155"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3399730"},{"issue":"14","key":"e_1_3_2_20_2","doi-asserted-by":"crossref","first-page":"e5470","DOI":"10.1002\/cpe.5470","article-title":"Benchmarking multi-GPU applications on modern multi-GPU integrated systems","volume":"33","author":"Bernaschi Massimo","year":"2021","unstructured":"Massimo Bernaschi, Elena Agostini, and Davide Rossetti. 2021. Benchmarking multi-GPU applications on modern multi-GPU integrated systems. Concurrency and Computation: Practice and Experience 33, 14 (2021), e5470.","journal-title":"Concurrency and Computation: Practice and Experience"},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","first-page":"1288","DOI":"10.1109\/SCW63240.2024.00169","volume-title":"Proceedings of the SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Brooks Alex","year":"2024","unstructured":"Alex Brooks, Philip Marshall, David Ozog, Md. Wasi-Ur-Rahman, Lawrence Stewart, and Rithwik Tom. 2024. IntelSHMEM: GPU-initiated OpenSHMEM using SYCL. In Proceedings of the SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1288\u20131301. DOI:10.1109\/SCW63240.2024.00169"},{"key":"e_1_3_2_22_2","series-title":"EuroMPI\u201912","first-page":"110","volume-title":"Proceedings of the 19th European Conference on Recent Advances in the Message Passing Interface","author":"Bureddy D.","year":"2012","unstructured":"D. Bureddy, H. Wang, A. Venkatesh, S. Potluri, and D. K. Panda. 2012. OMB-GPU: A micro-benchmark suite for evaluating MPI libraries on GPU clusters. In Proceedings of the 19th European Conference on Recent Advances in the Message Passing Interface (Vienna, Austria) (EuroMPI\u201912). Springer-Verlag, Berlin, 110\u2013120. DOI:10.1007\/978-3-642-33518-1_16"},{"key":"e_1_3_2_23_2","first-page":"1","volume-title":"Proceedings of the 2021 IEEE Hot Chips 33 Symposium (HCS)","author":"Burstein Idan","year":"2021","unstructured":"Idan Burstein. 2021. Nvidia data center processing unit (DPU) architecture. In Proceedings of the 2021 IEEE Hot Chips 33 Symposium (HCS). 1\u201320. DOI:10.1109\/HCS52781.2021.9567066"},{"key":"e_1_3_2_24_2","series-title":"PPoPP\u201921","first-page":"62","volume-title":"Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming","author":"Cai Zixian","year":"2021","unstructured":"Zixian Cai, Zhengyang Liu, Saeed Maleki, Madanlal Musuvathi, Todd Mytkowicz, Jacob Nelson, and Olli Saarikivi. 2021. Synthesizing optimal collective algorithms. In Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (Virtual Event, Republic of Korea) (PPoPP\u201921). Association for Computing Machinery, New York, NY, USA, 62\u201375. DOI:10.1145\/3437801.3441620"},{"key":"e_1_3_2_25_2","first-page":"131","volume-title":"Proceedings of the 2023 IEEE\/ACM 23rd International Symposium on Cluster, Cloud and Internet Computing (CCGrid)","author":"Chen Chen-Chun","year":"2023","unstructured":"Chen-Chun Chen, Kawthar Shafie Khorassani, Goutham Kalikrishna Reddy Kuncham, Rahul Vaidya, Mustafa Abduljabbar, Aamir Shafi, Hari Subramoni, and Dhabaleswar K. Panda. 2023. Implementing and optimizing a GPU-aware MPI library for intel GPUs: Early experiences. In Proceedings of the 2023 IEEE\/ACM 23rd International Symposium on Cluster, Cloud and Internet Computing (CCGrid). 131\u2013140. DOI:10.1109\/CCGrid57682.2023.00022"},{"key":"e_1_3_2_26_2","series-title":"PEARC\u201924","volume-title":"Proceedings of the Practice and Experience in Advanced Research Computing 2024: Human Powered Computing","author":"Chen Chen-Chun","year":"2024","unstructured":"Chen-Chun Chen, Goutham Kalikrishna Reddy Kuncham, Pouya Kousha, Hari Subramoni, and Dhabaleswar K. Panda. 2024. Design and implementation of an IPC-based collective MPI library for intel GPUs. In Proceedings of the Practice and Experience in Advanced Research Computing 2024: Human Powered Computing (Providence, RI, USA) (PEARC\u201924). Association for Computing Machinery, New York, NY, USA, Article 17, 9 pages. DOI:10.1145\/3626203.3670549"},{"key":"e_1_3_2_27_2","series-title":"SC-W\u201923","first-page":"847","volume-title":"Proceedings of the SC\u201923 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis","author":"Chen Chen-Chun","year":"2023","unstructured":"Chen-Chun Chen, Kawthar Shafie Khorassani, Pouya Kousha, Qinghua Zhou, Jinghan Yao, Hari Subramoni, and Dhabaleswar K. Panda. 2023. MPI-xCCL: A portable MPI library over collective communication libraries for various accelerators. In Proceedings of the SC\u201923 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis (Denver, CO, USA) (SC-W\u201923). Association for Computing Machinery, New York, NY, USA, 847\u2013854. DOI:10.1145\/3624062.3624153"},{"key":"e_1_3_2_28_2","series-title":"SC\u201922","volume-title":"Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis","author":"Chen Yuxin","year":"2022","unstructured":"Yuxin Chen, Benjamin Brock, Serban Porumbescu, Ayd\u0131n Bulu\u00e7, Katherine Yelick, and John D. Owens. 2022. Scalable irregular parallelism with GPUs: Getting CPUs out of the way. In Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis (Dallas, Texas) (SC\u201922). Institute for Electrical and Electronics Engineers, New York, NY, USA, Article 50, 16 pages."},{"key":"e_1_3_2_29_2","first-page":"479","volume-title":"Proceedings of the 2021 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW)","author":"Choi Jaemin","year":"2021","unstructured":"Jaemin Choi, Zane Fink, Sam White, Nitin Bhat, David F. Richards, and Laxmikant V. Kale. 2021. GPU-aware communication with UCX in parallel programming models: Charm++, MPI, and python. In Proceedings of the 2021 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). 479\u2013488. DOI:10.1109\/IPDPSW52791.2021.00079"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2022.102969"},{"key":"e_1_3_2_31_2","series-title":"HPDC\u201921","doi-asserted-by":"crossref","first-page":"261","DOI":"10.1145\/3431379.3464454","volume-title":"Proceedings of the 30th International Symposium on High-Performance Parallel and Distributed Computing","author":"Choi Jaemin","year":"2021","unstructured":"Jaemin Choi, David F. Richards, and Laxmikant V. Kale. 2021. CharminG: A scalable GPU-resident runtime system. In Proceedings of the 30th International Symposium on High-Performance Parallel and Distributed Computing (Virtual Event, Sweden) (HPDC\u201921). Association for Computing Machinery, New York, NY, USA, 261\u2013262. DOI:10.1145\/3431379.3464454"},{"key":"e_1_3_2_32_2","first-page":"148","volume-title":"Proceedings of the OpenSHMEM and Related Technologies. OpenSHMEM in the Era of Extreme Heterogeneity","author":"Chu Ching-Hsiang","year":"2019","unstructured":"Ching-Hsiang Chu, Sreeram Potluri, Anshuman Goswami, Manjunath Gorentla Venkata, Neena Imam, and Chris J. Newburn. 2019. Designing high-performance in-memory key-value operations with persistent GPU kernels and OpenSHMEM. In Proceedings of the OpenSHMEM and Related Technologies. OpenSHMEM in the Era of Extreme Heterogeneity, Swaroop Pophale, Neena Imam, Ferrol Aderholdt, and Manjunath Gorentla Venkata (Eds.). Springer International Publishing, Cham, 148\u2013164."},{"key":"e_1_3_2_33_2","article-title":"Kokkos Remote Spaces Repository","author":"Ciesko Jan","year":"2023","unstructured":"Jan Ciesko. 2023. Kokkos Remote Spaces Repository. Retrieved May 13, 2026 from https:\/\/github.com\/kokkos\/kokkos-remote-spaces","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_34_2","article-title":"GitHub - Pull Request for adding support for Unified Collective Communication (UCC)","author":"Clauss Carsten","year":"2025","unstructured":"Carsten Clauss. 2025. GitHub - Pull Request for adding support for Unified Collective Communication (UCC). Retrieved May 13, 2026 from https:\/\/github.com\/pmodels\/mpich\/pull\/7578","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_35_2","article-title":"NVIDIA DOCA SDK Documentation","author":"Corporation NVIDIA","year":"2023","unstructured":"NVIDIA Corporation. 2023. NVIDIA DOCA SDK Documentation. Retrieved May 13, 2026 from https:\/\/docs.nvidia.com\/doca\/sdk\/index.html","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_36_2","series-title":"ASPLOS 2023","doi-asserted-by":"crossref","first-page":"502","DOI":"10.1145\/3575693.3575724","volume-title":"Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2","author":"Cowan Meghan","year":"2023","unstructured":"Meghan Cowan, Saeed Maleki, Madanlal Musuvathi, Olli Saarikivi, and Yifan Xiong. 2023. MSCCLang: Microsoft collective communication language. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (Vancouver, BC, Canada) (ASPLOS 2023). Association for Computing Machinery, New York, NY, USA, 502\u2013514. DOI:10.1145\/3575693.3575724"},{"key":"e_1_3_2_37_2","article-title":"LUMI-G Supercomputer","year":"2024","unstructured":"CSC. 2024. LUMI-G Supercomputer. Retrieved May 13, 2026 from https:\/\/docs.lumi-supercomputer.eu\/hardware\/lumig\/","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","unstructured":"Feras Daoud Amir Watad and Mark Silberstein. 2016. GPUrdma: GPU-side library for high performance networking from GPU kernels.In Proceedings of the 6th International Workshop on Runtime and Operating Systems for Supercomputers (ROSS\u201916). Association for Computing Machinery New York NY USA Article 6 8 pages. DOI:10.1145\/2931088.2931091","DOI":"10.1145\/2931088.2931091"},{"key":"e_1_3_2_39_2","article-title":"The Latest in GPUDirect","author":"Markthub Seth Howell Davide Rossetti, Pak","year":"2021","unstructured":"Seth Howell Davide Rossetti, Pak Markthub. 2021. The Latest in GPUDirect. Retrieved May 13, 2026 from https:\/\/www.nvidia.com\/en-us\/on-demand\/session\/gtcspring21-s32039\/","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_40_2","series-title":"SC\u201919","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Sensi Daniele De","year":"2019","unstructured":"Daniele De Sensi, Salvatore Di Girolamo, and Torsten Hoefler. 2019. Mitigating network noise on dragonfly networks through application-aware routing. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (Denver, Colorado) (SC\u201919). ACM, New York, NY, USA, Article 16, 32 pages. DOI:10.1145\/3295500.3356196"},{"key":"e_1_3_2_41_2","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC\u201924)","author":"Sensi Daniele De","year":"2024","unstructured":"Daniele De Sensi, Lorenzo Pichetti, Flavio Vella, Tiziano De Matteis, Zebin Ren, Luigi Fusco, Matteo Turisini, Daniele Cesarini, Kurt Lust, Animesh Trivedi, et\u00a0al. 2024. Exploring GPU-to-GPU communication: Insights into supercomputer interconnects. In Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC\u201924). DOI:10.1109\/SC41406.2024.00039"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3733104"},{"key":"e_1_3_2_43_2","first-page":"1314","volume-title":"Proceedings of the SC \u201925 Workshops of the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC Workshops \u201925)","author":"Doijade Mahesh","year":"2025","unstructured":"Mahesh Doijade, Andrey Alekseenko, Ania Brown, Alan Gray, and Szil\u00e1rd P\u00e1ll. 2025. Redesigning GROMACS halo exchange: Improving strong scaling with GPU-initiated NVSHMEM. In Proceedings of the SC \u201925 Workshops of the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC Workshops \u201925). ACM, 1314\u20131329. DOI:10.1145\/3731599.3767508"},{"key":"e_1_3_2_44_2","first-page":"1","volume-title":"Proceedings of the 2018 IEEE\/ACM Machine Learning in HPC Environments (MLHPC)","author":"Dryden Nikoli","year":"2018","unstructured":"Nikoli Dryden, Naoya Maruyama, Tim Moon, Tom Benson, Andy Yoo, Marc Snir, and Brian Van Essen. 2018. Aluminum: An asynchronous, GPU-aware communication library optimized for large-scale training of deep neural networks on HPC systems. In Proceedings of the 2018 IEEE\/ACM Machine Learning in HPC Environments (MLHPC). 1\u201313. DOI:10.1109\/MLHPC.2018.8638639"},{"key":"e_1_3_2_45_2","article-title":"NVIDIA grace hopper superchip architecture in-depth","author":"Evans Jonathon","year":"2022","unstructured":"Jonathon Evans, Michael Andersch, Vikram Sethi, Gonzalo Brito, and Vishal Mehta. 2022. NVIDIA grace hopper superchip architecture in-depth. Retrieved May 13, 2026 from https:\/\/developer.nvidia.com\/blog\/nvidia-grace-hopper-superchip-architecture-in-depth\/","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_46_2","unstructured":"Emir Gencer Mohammad Kefah Taha Issa Ilyas Turimbetov James D. Trotter and Didem Unat. 2026. ucTrace: A Multi-Layer Profiling Tool for UCX-driven Communication. arXiv:2602.19084. Retrieved from https:\/\/arxiv.org\/abs\/2602.19084"},{"key":"e_1_3_2_47_2","first-page":"126","volume-title":"Proceedings of the 2020 IEEE\/ACM Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS)","author":"Groves Taylor","year":"2020","unstructured":"Taylor Groves, Ben Brock, Yuxin Chen, Khaled Z. Ibrahim, Lenny Oliker, Nicholas J. Wright, Samuel Williams, and Katherine Yelick. 2020. Performance tradeoffs in GPU communication: A study of host and device-initiated approaches. In Proceedings of the 2020 IEEE\/ACM Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS). 126\u2013137. DOI:10.1109\/PMBS51919.2020.00016"},{"key":"e_1_3_2_48_2","article-title":"MPICH for Exascale","author":"Guo Yanfei","year":"2022","unstructured":"Yanfei Guo, Kenneth Raffenetti, Rob Latham, Marc Snir, and Hui Zhou. 2022. MPICH for Exascale. Retrieved May 13, 2026 from https:\/\/www.exascaleproject.org\/wp-content\/uploads\/2022\/06\/2022_ECPAM_BoF_MPICH_ANL.pdf","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_49_2","first-page":"609","volume-title":"SC\u201916: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Gysi Tobias","year":"2016","unstructured":"Tobias Gysi, Jeremia B\u00e4r, and Torsten Hoefler. 2016. dCUDA: Hardware supported overlap of computation and communication. In SC\u201916: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 609\u2013620. DOI:10.1109\/SC.2016.51"},{"key":"e_1_3_2_50_2","first-page":"52","volume-title":"Proceedings of the 2016 IEEE 23rd International Conference on High Performance Computing (HiPC)","author":"Hamidouche Khaled","year":"2016","unstructured":"Khaled Hamidouche, Ammar Ahmad Awan, Akshay Venkatesh, and Dhabaleswar K. Panda. 2016. CUDA M3: Designing efficient CUDA managed memory-aware MPI by exploiting GDR and IPC. In Proceedings of the 2016 IEEE 23rd International Conference on High Performance Computing (HiPC). 52\u201361. DOI:10.1109\/HiPC.2016.016"},{"key":"e_1_3_2_51_2","series-title":"PPoPP\u201920","doi-asserted-by":"crossref","first-page":"336","DOI":"10.1145\/3332466.3374544","volume-title":"Proceedings of the 25th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming","author":"Hamidouche Khaled","year":"2020","unstructured":"Khaled Hamidouche and Michael LeBeane. 2020. GPU INitiated OPenSHMEM: Correct and efficient intra-kernel networking for DGPUs. In Proceedings of the 25th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (San Diego, California) (PPoPP\u201920). Association for Computing Machinery, New York, NY, USA, 336\u2013347. DOI:10.1145\/3332466.3374544"},{"key":"e_1_3_2_52_2","article-title":"How to Optimize Data Transfers in CUDA C\/C++","author":"Harris Mark","year":"2012","unstructured":"Mark Harris. 2012. How to Optimize Data Transfers in CUDA C\/C++. Retrieved May 13, 2026 from https:\/\/developer.nvidia.com\/blog\/how-optimize-data-transfers-cuda-cc\/","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_53_2","series-title":"ICS\u201924","doi-asserted-by":"crossref","first-page":"426","DOI":"10.1145\/3650200.3656591","volume-title":"Proceedings of the 38th ACM International Conference on Supercomputing","author":"Hidayetoglu Mert","year":"2024","unstructured":"Mert Hidayetoglu, Simon Garcia De Gonzalo, Elliott Slaughter, Yu Li, Christopher Zimmer, Tekin Bicer, Bin Ren, William Gropp, Wen-Mei Hwu, and Alex Aiken. 2024. CommBench: Micro-benchmarking hierarchical networks with multi-GPU, multi-NIC nodes. In Proceedings of the 38th ACM International Conference on Supercomputing (Kyoto, Japan) (ICS\u201924). Association for Computing Machinery, New York, NY, USA, 426\u2013436. DOI:10.1145\/3650200.3656591"},{"key":"e_1_3_2_54_2","first-page":"950","volume-title":"Proceedings of the 2025 IEEE International Parallel and Distributed Processing Symposium (IPDPS)","author":"Hidayetoglu Mert","year":"2025","unstructured":"Mert Hidayetoglu, Simon Garcia de Gonzalo, Elliott Slaughter, Pinku Surana, Wen-mei Hwu, William Gropp, and Alex Aiken. 2025. HiCCL: A hierarchical collective communication library. In Proceedings of the 2025 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 950\u2013961. DOI:10.1109\/IPDPS64566.2025.00089"},{"key":"e_1_3_2_55_2","article-title":"Cray MPICH Documentation","year":"2021","unstructured":"HPE. 2021. Cray MPICH Documentation. Retrieved May 13, 2026 from https:\/\/cpe.ext.hpe.com\/docs\/24.03\/mpt\/mpich\/index.html","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_56_2","unstructured":"Zhiyi Hu Siyuan Shen Tommaso Bonato Sylvain Jeaugey Cedell Alexander Eric Spada James Dinan Jeff Hammond and Torsten Hoefler. 2025. Demystifying NCCL: An in-depth analysis of GPU communication protocols and algorithms. arXiv:2507.04786. Retrieved from https:\/\/arxiv.org\/abs\/2507.04786"},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1145\/3721145.3733642","volume-title":"Proceedings of the 39th ACM International Conference on Supercomputing (ICS\u201925)","author":"Huang Jiajun","year":"2025","unstructured":"Jiajun Huang, Sheng Di, Yafan Huang, Zizhong Chen, Franck Cappello, Yanfei Guo, and Rajeev Thakur. 2025. ghZCCL: Advancing GPU-aware collective communications with homomorphic compression. In Proceedings of the 39th ACM International Conference on Supercomputing (ICS\u201925). Association for Computing Machinery, New York, NY, USA, 43\u201356. DOI:10.1145\/3721145.3733642"},{"key":"e_1_3_2_58_2","series-title":"ICS\u201924","doi-asserted-by":"crossref","first-page":"437","DOI":"10.1145\/3650200.3656636","volume-title":"Proceedings of the 38th ACM International Conference on Supercomputing","author":"Huang Jiajun","year":"2024","unstructured":"Jiajun Huang, Sheng Di, Xiaodong Yu, Yujia Zhai, Jinyang Liu, Yafan Huang, Ken Raffenetti, Hui Zhou, Kai Zhao, Xiaoyi Lu, et\u00a0al. 2024. gZCCL: Compression-accelerated collective communication framework for GPU clusters. In Proceedings of the 38th ACM International Conference on Supercomputing (Kyoto, Japan) (ICS\u201924). Association for Computing Machinery, New York, NY, USA, 437\u2013448. DOI:10.1145\/3650200.3656636"},{"key":"e_1_3_2_59_2","first-page":"1","volume-title":"Proceedings of the SC24: International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Huang Jiajun","year":"2024","unstructured":"Jiajun Huang, Sheng Di, Xiaodong Yu, Yujia Zhai, Jinyang Liu, Zizhe Jian, Xin Liang, Kai Zhao, Xiaoyi Lu, Zizhong Chen, et\u00a0al. 2024. hZCCL: Accelerating collective communication with co-designed homomorphic compression. In Proceedings of the SC24: International Conference for High Performance Computing, Networking, Storage and Analysis. 1\u201315. DOI:10.1109\/SC41406.2024.00110"},{"key":"e_1_3_2_60_2","article-title":"The NVLink-network switch: NVIDIA\u2019s switch chip for high communication-bandwidth superpods","author":"Ishii Alexaner","year":"2023","unstructured":"Alexaner Ishii and Ryan Wells. 2023. The NVLink-network switch: NVIDIA\u2019s switch chip for high communication-bandwidth superpods. Retrieved May 13, 2026 from https:\/\/hc34.hotchips.org\/assets\/program\/conference\/day2\/Network%20and%20Switches\/NVSwitch%20HotChips%202022%20r5.pdf","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_61_2","series-title":"ICS\u201923","doi-asserted-by":"crossref","first-page":"192","DOI":"10.1145\/3577193.3593713","volume-title":"Proceedings of the 37th International Conference on Supercomputing","author":"Ismayilov Ismayil","year":"2023","unstructured":"Ismayil Ismayilov, Javid Baydamirli, Do\u011fan Sa\u011fbili, Mohamed Wahib, and Didem Unat. 2023. Multi-GPU communication schemes for iterative solvers: When CPUs are not in charge. In Proceedings of the 37th International Conference on Supercomputing (Orlando, FL, USA) (ICS\u201923). Association for Computing Machinery, New York, NY, USA, 192\u2013202. DOI:10.1145\/3577193.3593713"},{"key":"e_1_3_2_62_2","series-title":"ICS\u201924","doi-asserted-by":"crossref","first-page":"525","DOI":"10.1145\/3650200.3656597","volume-title":"Proceedings of the 38th ACM International Conference on Supercomputing","author":"Issa Mohammad Kefah Taha","year":"2024","unstructured":"Mohammad Kefah Taha Issa, Muhammad Aditya Sasongko, Ilyas Turimbetov, Javid Baydamirli, Do\u011fan Sa\u011fbili, and Didem Unat. 2024. Snoopie: A Multi-GPU communication profiler and visualizer. In Proceedings of the 38th ACM International Conference on Supercomputing (Kyoto, Japan) (ICS\u201924). Association for Computing Machinery, New York, NY, USA, 525\u2013536. DOI:10.1145\/3650200.3656597"},{"key":"e_1_3_2_63_2","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC\u201924)","author":"Jacobson John","year":"2024","unstructured":"John Jacobson, Martin Burtscher, and Ganesh Gopalakrishnan. 2024. HiRace: Accurate and fast data race checking for GPU programs. In Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC\u201924). IEEE Press, Article 36, 14 pages. DOI:10.1109\/SC41406.2024.00042"},{"key":"e_1_3_2_64_2","first-page":"84","volume-title":"Proceedings of the 2024 IEEE 42nd International Conference on Computer Design (ICCD)","author":"Jia Ziyang","year":"2024","unstructured":"Ziyang Jia, Laxmi N. Bhuyan, and Daniel Wong. 2024. PCCL: Energy-efficient LLM training with power-aware collective communication. In Proceedings of the 2024 IEEE 42nd International Conference on Computer Design (ICCD). 84\u201391. DOI:10.1109\/ICCD63220.2024.00023"},{"key":"e_1_3_2_65_2","first-page":"411","volume-title":"Proceedings of the 2014 43rd International Conference on Parallel Processing Workshops","author":"Klenk Benjamin","year":"2014","unstructured":"Benjamin Klenk, Lena Oden, and Holger Froening. 2014. Analyzing put\/get APIs for thread-collaborative processors. In Proceedings of the 2014 43rd International Conference on Parallel Processing Workshops. 411\u2013418. DOI:10.1109\/ICPPW.2014.61"},{"key":"e_1_3_2_66_2","volume-title":"Proceedings of the International Workshop on Green Programming, Computing and Data Processing (GPCDP) in Conjunction with International Green Computing Conference (IGCC), Dallas, TX, USA","author":"Klenk Benjamin","year":"2014","unstructured":"Benjamin Klenk, Lena Oden, and Holger Fr\u00f6ning. 2014. GPU-centric communication for improved efficiency. In Proceedings of the International Workshop on Green Programming, Computing and Data Processing (GPCDP) in Conjunction with International Green Computing Conference (IGCC), Dallas, TX, USA."},{"key":"e_1_3_2_67_2","first-page":"318","volume-title":"Proceedings of the 2015 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS)","author":"Klenk Benjamin","year":"2015","unstructured":"Benjamin Klenk, Lena Oden, and Holger Froning. 2015. Analyzing communication models for distributed thread-collaborative processors in terms of energy and time. In Proceedings of the 2015 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). 318\u2013327. DOI:10.1109\/ISPASS.2015.7095817"},{"key":"e_1_3_2_68_2","first-page":"210","volume-title":"Proceedings of the 2025 IEEE International Parallel and Distributed Processing Symposium (IPDPS)","author":"Kwack JaeHyuk","year":"2025","unstructured":"JaeHyuk Kwack, Colleen Bertoni, Umesh Unnikrishnan, Riccardo Balin, Khalid Hossain, Yasaman Ghadar, Timothy J. Williams, Abhishek Bagusetty, Mathialakan Thavappiragasam, V\u00e4in\u00f6 Hatanp\u00e4\u00e4, et\u00a0al. 2025. AI and HPC applications on leadership computing platforms: Performance and scalability studies. In Proceedings of the 2025 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 210\u2013222. DOI:10.1109\/IPDPS64566.2025.00027"},{"key":"e_1_3_2_69_2","article-title":"COCCL: Compression and precision cO-awareness Collective Communication Library implemented based on NCCL.","author":"Lab HPDPS","year":"2025","unstructured":"HPDPS Lab. 2025. COCCL: Compression and precision cO-awareness Collective Communication Library implemented based on NCCL. Retrieved November 2, 2025 from https:\/\/github.com\/hpdps-group\/coccl.","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_70_2","article-title":"Blink GPU Benchmark","author":"Laboratory Hicrest","year":"2025","unstructured":"Hicrest Laboratory. 2025. Blink GPU Benchmark. Retrieved November 2, 2025 from https:\/\/github.com\/HicrestLaboratory\/Blink-GPU","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_71_2","article-title":"NVSHMEM: GPU-Integrated Communication for NVIDIA GPU Clusters","author":"Langer Akhil","year":"2021","unstructured":"Akhil Langer and Jim Dinan. 2021. NVSHMEM: GPU-Integrated Communication for NVIDIA GPU Clusters. Retrieved November 2, 2025 from https:\/\/www.nvidia.com\/en-us\/on-demand\/session\/gtcspring21-s32515\/","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_72_2","article-title":"QUDA Repository","year":"2023","unstructured":"Lattice. 2023. QUDA Repository. Retrieved November 2, 2025 from https:\/\/github.com\/lattice\/quda","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","unstructured":"Michael LeBeane Khaled Hamidouche Brad Benton Mauricio Breternitz Steven K. Reinhardt and Lizy K. John. 2017. GPU triggered networking for intra-kernel communications.In Proceedings of the International Conference for High Performance Computing Networking Storage and Analysis (SC\u201917). Association for Computing Machinery New York NY USA Article 22 12 pages. DOI:10.1145\/3126908.3126950","DOI":"10.1145\/3126908.3126950"},{"key":"e_1_3_2_74_2","series-title":"PACT\u201918","volume-title":"Proceedings of the 27th International Conference on Parallel Architectures and Compilation Techniques","author":"LeBeane Michael","year":"2018","unstructured":"Michael LeBeane, Khaled Hamidouche, Brad Benton, Mauricio Breternitz, Steven K. Reinhardt, and Lizy K. John. 2018. ComP-Net: Command processor networking for efficient intra-kernel communications on GPUs. In Proceedings of the 27th International Conference on Parallel Architectures and Compilation Techniques (Limassol, Cyprus) (PACT\u201918). Association for Computing Machinery, New York, NY, USA, Article 29, 13 pages. DOI:10.1145\/3243176.3243179"},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2928289"},{"key":"e_1_3_2_76_2","first-page":"1","volume-title":"Proceedings of the 2025 IEEE 33rd International Conference on Network Protocols (ICNP)","author":"Li Ziming","year":"2025","unstructured":"Ziming Li, Chenyang Hei, Fuliang Li, Tongrui Liu, Chengxi Gao, Xiuzhu Sha, and Xingwei Wang. 2025. TuCCL: Tailored and unified configuration optimizations for high-performance collective communication library. In Proceedings of the 2025 IEEE 33rd International Conference on Network Protocols (ICNP). 1\u201311. DOI:10.1109\/ICNP65844.2025.11192370"},{"key":"e_1_3_2_77_2","series-title":"GPGPU\u201919","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1145\/3300053.3319419","volume-title":"Proceedings of the 12th Workshop on General Purpose Processing Using GPUs","author":"Manian K. V.","year":"2019","unstructured":"K. V. Manian, A. A. Ammar, A. Ruhela, C.-H. Chu, H. Subramoni, and D. K. Panda. 2019. Characterizing CUDA unified memory (UM)-Aware MPI designs on modern GPU architectures. In Proceedings of the 12th Workshop on General Purpose Processing Using GPUs (Providence, RI, USA) (GPGPU\u201919). Association for Computing Machinery, New York, NY, USA, 43\u201352. DOI:10.1145\/3300053.3319419"},{"key":"e_1_3_2_78_2","article-title":"Scaling scientific computing with NVSHMEM","author":"Maruyama Naoya","year":"2020","unstructured":"Naoya Maruyama, Brian Van Essen, Jan Ciesko, Jeremiah Wilke, Christian Trott, Chung-Hsing Hsu, Neena Imam, Jim Dinan, Akhil Langer, C. J. Newburn, and Sreeram Potluri. 2020. Scaling scientific computing with NVSHMEM. NVIDIA Technical Blog. Retrieved May 13, 2026 from https:\/\/developer.nvidia.com\/blog\/scaling-scientific-computing-with-nvshmem\/","journal-title":"NVIDIA Technical Blog. Retrieved May 13, 2026 from"},{"key":"e_1_3_2_79_2","unstructured":"Hans Meuer Erich Strohmaier Jack Dongarra Horst Simon and Martin Meuer. 2023. TOP500 Supercomputer Sites. Retrieved May 13 2026 from https:\/\/www.top500.org\/"},{"key":"e_1_3_2_80_2","series-title":"GPGPU-5","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1145\/2159430.2159433","volume-title":"Proceedings of the 5th Annual Workshop on General Purpose Processing with Graphics Processing Units","author":"Miyoshi Takefumi","year":"2012","unstructured":"Takefumi Miyoshi, Hidetsugu Irie, Keigo Shima, Hiroki Honda, Masaaki Kondo, and Tsutomu Yoshinaga. 2012. FLAT: A GPU programming framework to provide embedded MPI. In Proceedings of the 5th Annual Workshop on General Purpose Processing with Graphics Processing Units (London, United Kingdom) (GPGPU-5). Association for Computing Machinery, New York, NY, USA, 20\u201329. DOI:10.1145\/2159430.2159433"},{"key":"e_1_3_2_81_2","article-title":"Key Hyperscalers And Chip Makers Gang Up On Nvidia\u2019s NVSwitch Interconnect","author":"Morgan Timothy Prickett","year":"2024","unstructured":"Timothy Prickett Morgan. 2024. Key Hyperscalers And Chip Makers Gang Up On Nvidia\u2019s NVSwitch Interconnect. Retrieved May 13, 2026 from https:\/\/www.hpcwire.com\/2024\/05\/30\/everyone-except-nvidia-forms-ultra-accelerator-link-ualink-consortium\/","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_82_2","volume-title":"Proceedings of the Practice and Experience in Advanced Research Computing 2025: The Power of Collaboration (PEARC\u201925)","author":"Morsy Yasser","year":"2025","unstructured":"Yasser Morsy and Alan George. 2025. Comparative analysis of parallel communication models for regular and irregular GPU workloads. In Proceedings of the Practice and Experience in Advanced Research Computing 2025: The Power of Collaboration (PEARC\u201925). Association for Computing Machinery, New York, NY, USA, Article 34, 4 pages. DOI:10.1145\/3708035.3736073"},{"key":"e_1_3_2_83_2","first-page":"139","volume-title":"Proceedings of the 2021 ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)","author":"Muthukrishnan Harini","year":"2021","unstructured":"Harini Muthukrishnan, David Nellans, Daniel Lustig, Jeffrey A. Fessler, and Thomas F. Wenisch. 2021. Efficient multi-GPU shared memory via automatic optimization of fine-grained transfers. In Proceedings of the 2021 ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). 139\u2013152. DOI:10.1109\/ISCA52012.2021.00020"},{"key":"e_1_3_2_84_2","unstructured":"Naveen Namashivayam Krishna Kandalla James B. White III au2 Larry Kaplan and Mark Pagel. 2023. Exploring Fully Offloaded GPU Stream-Aware Message Passing. arXiv:2306.15773. Retrieved from https:\/\/arxiv.org\/abs\/2306.15773"},{"key":"e_1_3_2_85_2","unstructured":"Naveen Namashivayam Krishna Kandalla Trey White Nick Radcliffe Larry Kaplan and Mark Pagel. 2022. Exploring GPU Stream-Aware Message Passing using Triggered Operations. arXiv:2208.04817. Retrieved from https:\/\/arxiv.org\/abs\/2208.04817"},{"key":"e_1_3_2_86_2","unstructured":"NVIDIA. [n. d.]. Retrieved November 2 2025 from https:\/\/docs.nvidia.com\/compute-sanitizer\/ComputeSanitizer\/index.html#racecheck-tool"},{"key":"e_1_3_2_87_2","article-title":"CUDA 4.0 Release Notes","year":"2011","unstructured":"NVIDIA. 2011. CUDA 4.0 Release Notes. Retrieved November 2, 2025 from https:\/\/developer.nvidia.com\/cuda-toolkit-40","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_88_2","article-title":"NVIDIA GPUDirect\u2122Technology","year":"2012","unstructured":"NVIDIA. 2012. NVIDIA GPUDirect\u2122Technology. Retrieved November 2, 2025 from https:\/\/developer.download.nvidia.com\/devzone\/devcenter\/cuda\/docs\/GPUDirect_Technology_Overview.pdf","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_89_2","article-title":"Fast Multi-GPU Collectives With NCCL","year":"2016","unstructured":"NVIDIA. 2016. Fast Multi-GPU Collectives With NCCL. Retrieved November 2, 2025 from https:\/\/developer.nvidia.com\/blog\/fast-multi-gpu-collectives-nccl\/","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_90_2","article-title":"CUDA 4.1 Release Notes","year":"2017","unstructured":"NVIDIA. 2017. CUDA 4.1 Release Notes. Retrieved November 2, 2025 from https:\/\/developer.nvidia.com\/cuda-toolkit-41-archive","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_91_2","article-title":"NVIDIA DGX-1 With Tesla V100 System Architecture","year":"2017","unstructured":"NVIDIA. 2017. NVIDIA DGX-1 With Tesla V100 System Architecture. Retrieved November 2, 2025 from https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_92_2","article-title":"Improving GPU Memory Oversubscription Performance","year":"2021","unstructured":"NVIDIA. 2021. Improving GPU Memory Oversubscription Performance. Retrieved November 2, 2025 from https:\/\/developer.nvidia.com\/blog\/improving-gpu-memory-oversubscription-performance\/","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_93_2","article-title":"CUDA Programming Guide Release 12.2","year":"2023","unstructured":"NVIDIA. 2023. CUDA Programming Guide Release 12.2. Retrieved November 2, 2025 from https:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/index.html","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_94_2","article-title":"CUDA Runtime - Device Management","year":"2023","unstructured":"NVIDIA. 2023. CUDA Runtime - Device Management. Retrieved November 2, 2025 from https:\/\/docs.nvidia.com\/cuda\/cuda-runtime-api\/group__CUDART__DEVICE.html#group__CUDART__DEVICE","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_95_2","article-title":"DGX-2","year":"2023","unstructured":"NVIDIA. 2023. DGX-2. Retrieved November 2, 2025 from https:\/\/www.nvidia.com\/en-gb\/data-center\/dgx-2\/","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_96_2","article-title":"GPUDirect RDMA","year":"2023","unstructured":"NVIDIA. 2023. GPUDirect RDMA. Retrieved November 2, 2025 from https:\/\/docs.nvidia.com\/cuda\/gpudirect-rdma\/","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_97_2","article-title":"Magnum IO GDRCopy","year":"2023","unstructured":"NVIDIA. 2023. Magnum IO GDRCopy. Retrieved November 2, 2025 from https:\/\/developer.nvidia.com\/gdrcopy","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_98_2","article-title":"NVIDIA GPUDirect Family","year":"2023","unstructured":"NVIDIA. 2023. NVIDIA GPUDirect Family. Retrieved November 2, 2025 from https:\/\/developer.nvidia.com\/gpudirect","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_99_2","article-title":"NVLink and NVSwitch","year":"2023","unstructured":"NVIDIA. 2023. NVLink and NVSwitch. Retrieved November 2, 2025 from https:\/\/www.nvidia.com\/en-us\/data-center\/nvlink\/","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_100_2","article-title":"NVSHMEM","year":"2023","unstructured":"NVIDIA. 2023. NVSHMEM. Retrieved November 2, 2025 from https:\/\/developer.nvidia.com\/nvshmem","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_101_2","article-title":"NVSHMEM 2.7.0 Release Notes","year":"2023","unstructured":"NVIDIA. 2023. NVSHMEM 2.7.0 Release Notes. Retrieved November 2, 2025 from https:\/\/docs.nvidia.com\/nvshmem\/release-notes\/release-270.html#release-270","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_102_2","article-title":"NVIDIA GB200","year":"2024","unstructured":"NVIDIA. 2024. NVIDIA GB200. Retrieved November 2, 2025 from https:\/\/www.nvidia.com\/en-us\/data-center\/gb200-nvl72\/","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_103_2","article-title":"NVIDIA Grace Hopper Superchip","year":"2024","unstructured":"NVIDIA. 2024. NVIDIA Grace Hopper Superchip. Retrieved November 2, 2025 from https:\/\/resources.nvidia.com\/en-us-grace-cpu\/nvidia-grace-hopper","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_104_2","article-title":"NVIDIA NVLink SGXLS10 Switch Systems User Manual","year":"2024","unstructured":"NVIDIA. 2024. NVIDIA NVLink SGXLS10 Switch Systems User Manual. Retrieved November 2, 2025 from https:\/\/docs.nvidia.com\/networking\/display\/sgxh100\/introduction","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_105_2","article-title":"NCCL Tests","year":"2025","unstructured":"NVIDIA. 2025. NCCL Tests. Retrieved November 2, 2025 from https:\/\/github.com\/NVIDIA\/nccl-tests","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_106_2","article-title":"Technical Blog on NCCL 2.28 Release","year":"2025","unstructured":"NVIDIA. November 2025. Technical Blog on NCCL 2.28 Release. Retrieved November 2, 2025 from https:\/\/developer.nvidia.com\/blog\/fusing-communication-and-compute-with-new-device-api-and-copy-engine-collectives-in-nvidia-nccl-2-28\/","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_107_2","first-page":"1","article-title":"GGAS: Global GPU address spaces for efficient communication in heterogeneous clusters","author":"Oden Lena","year":"2013","unstructured":"Lena Oden and Holger Fr\u00f6ning. 2013. GGAS: Global GPU address spaces for efficient communication in heterogeneous clusters. In Proceedings of the 2013 IEEE International Conference on Cluster Computing.1\u20138.","journal-title":"Proceedings of the 2013 IEEE International Conference on Cluster Computing."},{"key":"e_1_3_2_108_2","first-page":"976","volume-title":"Proceedings of the 2014 IEEE International Parallel and Distributed Processing Symposium Workshops","author":"Oden Lena","year":"2014","unstructured":"Lena Oden, Holger Fr\u00f6ning, and Franz-Joseph Pfreundt. 2014. Infiniband-verbs on GPU: A case study of controlling an infiniband network device from the GPU. In Proceedings of the 2014 IEEE International Parallel and Distributed Processing Symposium Workshops. 976\u2013983. DOI:10.1109\/IPDPSW.2014.111"},{"key":"e_1_3_2_109_2","first-page":"483","volume-title":"Proceedings of the 2014 14th IEEE\/ACM International Symposium on Cluster, Cloud, and Grid Computing","author":"Oden Lena","year":"2014","unstructured":"Lena Oden, Benjamin Klenk, and Holger Fr\u00f6ning. 2014. Energy-efficient collective reduce and allreduce operations on distributed GPUs. In Proceedings of the 2014 14th IEEE\/ACM International Symposium on Cluster, Cloud, and Grid Computing. 483\u2013492. DOI:10.1109\/CCGrid.2014.21"},{"key":"e_1_3_2_110_2","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1109\/E2SC.2014.14","volume-title":"Proceedings of the 2014 Energy Efficient Supercomputing Workshop","author":"Oden Lena","year":"2014","unstructured":"Lena Oden, Benjamin Klenk, and Holger Fr\u00f6ning. 2014. Energy-efficient stencil computations on distributed GPUs using dynamic parallelism and GPU-controlled communication. In Proceedings of the 2014 Energy Efficient Supercomputing Workshop. 31\u201340. DOI:10.1109\/E2SC.2014.14"},{"key":"e_1_3_2_111_2","article-title":"Open MPI v5.0.x Documentation: CUDA","year":"2023","unstructured":"OpenMPI. 2023. Open MPI v5.0.x Documentation: CUDA. Retrieved November 2, 2025 from https:\/\/docs.open-mpi.org\/en\/v5.0.x\/tuning-apps\/networking\/cuda.html","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_112_2","article-title":"Open MPI v5.0.x Documentation: ROCm","year":"2023","unstructured":"OpenMPI. 2023. Open MPI v5.0.x Documentation: ROCm. Retrieved November 2, 2025 from https:\/\/docs.open-mpi.org\/en\/v5.0.x\/tuning-apps\/networking\/rocm.html","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_113_2","article-title":"Improving Network Performance of HPC Systems using NVIDIA Magnum IO NVSHMEM and GPUDirect Async","author":"Dinan Sreeram Potluri Pak Markthub, Jim","year":"2022","unstructured":"Sreeram Potluri Pak Markthub, Jim Dinan and Seth Howell. 2022. Improving Network Performance of HPC Systems using NVIDIA Magnum IO NVSHMEM and GPUDirect Async. Retrieved November 2, 2025 from https:\/\/developer.nvidia.com\/blog\/improving-network-performance-of-hpc-systems-using-nvidia-magnum-io-nvshmem-and-gpudirect-async\/","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_114_2","unstructured":"Carl Pearson. 2023. Interconnect Bandwidth Heterogeneity on AMD MI250x and Infinity Fabric. arXiv:2302.14827. Retrieved from https:\/\/arxiv.org\/abs\/2302.14827"},{"key":"e_1_3_2_115_2","series-title":"ICPE\u201919","doi-asserted-by":"crossref","first-page":"209","DOI":"10.1145\/3297663.3310299","volume-title":"Proceedings of the 2019 ACM\/SPEC International Conference on Performance Engineering","author":"Pearson Carl","year":"2019","unstructured":"Carl Pearson, Abdul Dakkak, Sarah Hashash, Cheng Li, I.-Hsin Chung, Jinjun Xiong, and Wen-Mei Hwu. 2019. Evaluating characteristics of CUDA communication primitives on high-bandwidth interconnects. In Proceedings of the 2019 ACM\/SPEC International Conference on Performance Engineering (Mumbai, India) (ICPE\u201919). Association for Computing Machinery, New York, NY, USA, 209\u2013218. DOI:10.1145\/3297663.3310299"},{"key":"e_1_3_2_116_2","first-page":"253","volume-title":"Proceedings of the 2017 IEEE 24th International Conference on High Performance Computing (HiPC)","author":"Potluri Sreeram","year":"2017","unstructured":"Sreeram Potluri, Anshuman Goswami, Davide Rossetti, C. J. Newburn, Manjunath Gorentla Venkata, and Neena Imam. 2017. GPU-centric communication on NVIDIA GPU clusters with infiniband: A case study with OpenSHMEM. In Proceedings of the 2017 IEEE 24th International Conference on High Performance Computing (HiPC). 253\u2013262. DOI:10.1109\/HiPC.2017.00037"},{"key":"e_1_3_2_117_2","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1007\/978-3-319-73814-7_6","volume-title":"Proceedings of the OpenSHMEM and Related Technologies. Big Compute and Big Data Convergence","author":"Potluri Sreeram","year":"2018","unstructured":"Sreeram Potluri, Anshuman Goswami, Manjunath Gorentla Venkata, and Neena Imam. 2018. Efficient breadth first search on multi-GPU systems using GPU-centric OpenSHMEM. In Proceedings of the OpenSHMEM and Related Technologies. Big Compute and Big Data Convergence, Manjunath Gorentla Venkata, Neena Imam, and Swaroop Pophale (Eds.). Springer International Publishing, Cham, 82\u201396."},{"key":"e_1_3_2_118_2","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1007\/978-3-319-73814-7_6","volume-title":"Proceedings of the OpenSHMEM and Related Technologies. Big Compute and Big Data Convergence","author":"Potluri Sreeram","year":"2018","unstructured":"Sreeram Potluri, Anshuman Goswami, Manjunath Gorentla Venkata, and Neena Imam. 2018. Efficient breadth first search on multi-GPU systems using GPU-centric OpenSHMEM. In Proceedings of the OpenSHMEM and Related Technologies. Big Compute and Big Data Convergence, Manjunath Gorentla Venkata, Neena Imam, and Swaroop Pophale (Eds.). Springer International Publishing, Cham, 82\u201396."},{"key":"e_1_3_2_119_2","first-page":"80","volume-title":"Proceedings of the 2013 42nd International Conference on Parallel Processing","author":"Potluri Sreeram","year":"2013","unstructured":"Sreeram Potluri, Khaled Hamidouche, Akshay Venkatesh, Devendar Bureddy, and Dhabaleswar K. Panda. 2013. Efficient inter-node MPI communication using GPUDirect RDMA for infiniband clusters with NVIDIA GPUs. In Proceedings of the 2013 42nd International Conference on Parallel Processing. 80\u201389. DOI:10.1109\/ICPP.2013.17"},{"key":"e_1_3_2_120_2","series-title":"OpenSHMEM 2015","first-page":"18","volume-title":"Proceedings of the Revised Selected Papers of the 2nd Workshop on OpenSHMEM and Related Technologies. Experiences, Implementations, and Technologies\u2014Volume 9397","author":"Potluri Sreeram","year":"2015","unstructured":"Sreeram Potluri, Davide Rossetti, Donald Becker, Duncan Poole, Manjunath Gorentla Venkata, Oscar Hernandez, Pavel Shamis, M. Graham Lopez, Mathew Baker, and Wendy Poole. 2015. Exploring OpenSHMEM model to program GPU-based extreme-scale systems. In Proceedings of the Revised Selected Papers of the 2nd Workshop on OpenSHMEM and Related Technologies. Experiences, Implementations, and Technologies\u2014Volume 9397 (Annapolis, MD, USA) (OpenSHMEM 2015). Springer-Verlag, Berlin, 18\u201335. DOI:10.1007\/978-3-319-26428-8_2"},{"key":"e_1_3_2_121_2","first-page":"1848","volume-title":"Proceedings of the 2012 IEEE 26th International Parallel and Distributed Processing Symposium Workshops and PhD Forum","author":"Potluri S.","year":"2012","unstructured":"S. Potluri, H. Wang, D. Bureddy, A.K. Singh, C. Rosales, and Dhabaleswar K. Panda. 2012. Optimizing MPI communication on Multi-GPU systems using CUDA inter-process communication. In Proceedings of the 2012 IEEE 26th International Parallel and Distributed Processing Symposium Workshops and PhD Forum. 1848\u20131857. DOI:10.1109\/IPDPSW.2012.228"},{"key":"e_1_3_2_122_2","unstructured":"Kishore Punniyamurthy Bradford M. Beckmann and Khaled Hamidouche. 2023. GPU-initiated Fine-grained Overlap of Collective Communication with Computation. arXiv:2305.06942. Retrieved from https:\/\/arxiv.org\/abs\/2305.06942"},{"key":"e_1_3_2_123_2","unstructured":"Kishore Punniyamurthy Khaled Hamidouche and Bradford M. Beckmann. 2023. Optimizing Distributed ML Communication with Fused Computation-Collective Operations. arXiv:2305.06942. Retrieved from https:\/\/arxiv.org\/abs\/2305.06942"},{"key":"e_1_3_2_124_2","series-title":"ASPLOS 2023","doi-asserted-by":"crossref","first-page":"325","DOI":"10.1145\/3575693.3575748","volume-title":"Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2","author":"Qureshi Zaid","year":"2023","unstructured":"Zaid Qureshi, Vikram Sharma Mailthody, Isaac Gelado, Seungwon Min, Amna Masood, Jeongmin Park, Jinjun Xiong, C. J. Newburn, Dmitri Vainbrand, I.-Hsin Chung, et\u00a0al. 2023. GPU-initiated on-demand high-throughput storage access in the BaM system architecture. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (Vancouver, BC, Canada) (ASPLOS 2023). Association for Computing Machinery, New York, NY, USA, 325\u2013339. DOI:10.1145\/3575693.3575748"},{"key":"e_1_3_2_125_2","first-page":"1","volume-title":"Proceedings of the 2025 IEEE International Conference on Cluster Computing (CLUSTER)","author":"Sa\u011fbili Do\u011fan","year":"2025","unstructured":"Do\u011fan Sa\u011fbili, Sinan Ekmek\u00e7iba\u015f\u0131, Khaled Z. Ibrahim, Tan Nguyen, and Didem Unat. 2025. Uniconn: A uniform high-level communication library for portable multi-GPU programming. In Proceedings of the 2025 IEEE International Conference on Cluster Computing (CLUSTER). 1\u201312. DOI:10.1109\/CLUSTER59342.2025.11186498"},{"key":"e_1_3_2_126_2","series-title":"SC-W\u201924","first-page":"567","volume-title":"Proceedings of the SC\u201924 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis","author":"Schieffer Gabin","year":"2025","unstructured":"Gabin Schieffer, Ruimin Shi, Stefano Markidis, Andreas Herten, Jennifer Faj, and Ivy Peng. 2025. Understanding data movement in AMD multi-GPU systems with infinity fabric. In Proceedings of the SC\u201924 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis (Atlanta, GA, USA) (SC-W\u201924). IEEE Press, 567\u2013576. DOI:10.1109\/SCW63240.2024.00079"},{"key":"e_1_3_2_127_2","article-title":"Peer-to-Peer and Unified Virtual Addressing","author":"Schroeder Tim C.","year":"2011","unstructured":"Tim C. Schroeder. 2011. Peer-to-Peer and Unified Virtual Addressing. Retrieved from https:\/\/developer.download.nvidia.com\/CUDA\/training\/cuda_webinars_GPUDirect_uva.pdf","journal-title":"Retrieved from"},{"key":"e_1_3_2_128_2","doi-asserted-by":"crossref","first-page":"118","DOI":"10.1007\/978-3-030-78713-4_7","volume-title":"Proceedings of the High Performance Computing","author":"Khorassani Kawthar Shafie","year":"2021","unstructured":"Kawthar Shafie Khorassani, Jahanzeb Hashmi, Ching-Hsiang Chu, Chen-Chun Chen, Hari Subramoni, and Dhabaleswar K. Panda. 2021. Designing a ROCm-aware MPI library for AMD GPUs: Early experiences. In Proceedings of the High Performance Computing, Bradford L. Chamberlain, Ana-Lucia Varbanescu, Hatem Ltaief, and Piotr Luszczek (Eds.). Springer International Publishing, Cham, 118\u2013136."},{"key":"e_1_3_2_129_2","unstructured":"Aashaka Shah Vijay Chidambaram Meghan Cowan Saeed Maleki Madan Musuvathi Todd Mytkowicz Jacob Nelson Olli Saarikivi and Rachee Singh. 2022. TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches. arXiv:2111.04867. Retrieved from https:\/\/arxiv.org\/abs\/2111.04867"},{"key":"e_1_3_2_130_2","doi-asserted-by":"publisher","unstructured":"Aashaka Shah Abhinav Jangda Binyang Li Caio Rocha Changho Hwang Jithin Jose Madan Musuvathi Olli Saarikivi Peng Cheng Qinghua Zhou et\u00a0al. 2025. MSCCL++: Rethinking GPU communication abstractions for cutting-edge AI applications. DOI:10.48550\/arXiv.2504.09014","DOI":"10.48550\/arXiv.2504.09014"},{"key":"e_1_3_2_131_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00450-011-0157-1"},{"key":"e_1_3_2_132_2","first-page":"40","volume-title":"Proceedings of the 2015 IEEE 23rd Annual Symposium on High-Performance Interconnects","author":"Shamis Pavel","year":"2015","unstructured":"Pavel Shamis, Manjunath Gorentla Venkata, M. Graham Lopez, Matthew B. Baker, Oscar Hernandez, Yossi Itigin, Mike Dubman, Gilad Shainer, Richard L. Graham, Liran Liss, et\u00a0al. 2015. UCX: An open source framework for HPC network APIs and beyond. In Proceedings of the 2015 IEEE 23rd Annual Symposium on High-Performance Interconnects. IEEE, 40\u201343."},{"key":"e_1_3_2_133_2","series-title":"ICPE\u201922","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1145\/3489525.3511691","volume-title":"Proceedings of the 2022 ACM\/SPEC on International Conference on Performance Engineering","author":"Shao Chuanming","year":"2022","unstructured":"Chuanming Shao, Jinyang Guo, Pengyu Wang, Jing Wang, Chao Li, and Minyi Guo. 2022. Oversubscribing GPU unified virtual memory: Implications and suggestions. In Proceedings of the 2022 ACM\/SPEC on International Conference on Performance Engineering (Beijing, China) (ICPE\u201922). Association for Computing Machinery, New York, NY, USA, 67\u201375. DOI:10.1145\/3489525.3511691"},{"key":"e_1_3_2_134_2","first-page":"1","volume-title":"Proceedings of the 2014 21st International Conference on High Performance Computing (HiPC)","author":"Shi Rong","year":"2014","unstructured":"Rong Shi, Sreeram Potluri, Khaled Hamidouche, Jonathan Perkins, Mingzhe Li, Davide Rossetti, and Dhabaleswar K. D. K. Panda. 2014. Designing efficient small message transfer mechanism for inter-node MPI communication on InfiniBand GPU clusters. In Proceedings of the 2014 21st International Conference on High Performance Computing (HiPC). 1\u201310. DOI:10.1109\/HiPC.2014.7116873"},{"key":"e_1_3_2_135_2","series-title":"SC\u201911","volume-title":"Proceedings of the 2011 International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Shimokawabe Takashi","year":"2011","unstructured":"Takashi Shimokawabe, Takayuki Aoki, Tomohiro Takaki, Toshio Endo, Akinori Yamanaka, Naoya Maruyama, Akira Nukada, and Satoshi Matsuoka. 2011. Peta-scale phase-field simulation for dendritic solidification on the TSUBAME 2.0 supercomputer. In Proceedings of the 2011 International Conference for High Performance Computing, Networking, Storage and Analysis (Seattle, Washington) (SC\u201911). Association for Computing Machinery, New York, NY, USA, Article 3, 11 pages. DOI:10.1145\/2063384.2063388"},{"key":"e_1_3_2_136_2","doi-asserted-by":"publisher","unstructured":"Min Si Yongzhou Chen Pavan Balaji Ching-Hsiang Chu Adi Gangidi Saif Hasan Subodh Iyengar Dan Johnson Bingzhe Liu Regina Ren Ashmitha Jeevaraj Shetty et\u00a0al. 2025. Collective communication for 100k+ GPUs. DOI:10.48550\/arXiv.2510.20171","DOI":"10.48550\/arXiv.2510.20171"},{"key":"e_1_3_2_137_2","doi-asserted-by":"publisher","DOI":"10.1145\/2553081"},{"key":"e_1_3_2_138_2","doi-asserted-by":"publisher","unstructured":"Mark Silberstein Sangman Kim Seonggu Huh Xinya Zhang Yige Hu Amir Wated and Emmett Witchel. 2016. GPUnet: Networking abstractions for GPU programs. ACM Transactions on Computer Systems 34 3 Article 9 (2016) 31 pages. DOI:10.1145\/2963098","DOI":"10.1145\/2963098"},{"key":"e_1_3_2_139_2","unstructured":"Siddharth Singh Mahua Singh and Abhinav Bhatele. 2025. The Big Send-off: High Performance Collectives on GPU-based Supercomputers. arXiv:2504.18658. Retrieved from https:\/\/arxiv.org\/abs\/2504.18658"},{"key":"e_1_3_2_140_2","article-title":"We Bought the Whole GPU, So We\u2019re Damn Well Going to Use the Whole GPU","author":"Spector Benjamin","year":"2025","unstructured":"Benjamin Spector, Jordan Juravsky, Stuart Sul, Dylan Lim, Owen Dugan, Simran Arora, and Chris R\u00e9. 2025. We Bought the Whole GPU, So We\u2019re Damn Well Going to Use the Whole GPU. Retrieved May 13, 2026 from https:\/\/hazyresearch.stanford.edu\/blog\/2025-09-28-tp-llama-main","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_141_2","series-title":"Euro-Par 2010","first-page":"365","volume-title":"Proceedings of the 2010 Conference on Parallel Processing","author":"Stuart Jeff A.","year":"2010","unstructured":"Jeff A. Stuart, Michael Cox, and John D. Owens. 2010. GPU-to-CPU callbacks. In Proceedings of the 2010 Conference on Parallel Processing (Ischia, Italy) (Euro-Par 2010). Springer-Verlag, Berlin, 365\u2013372."},{"key":"e_1_3_2_142_2","unstructured":"Stuart H. Sul Simran Arora Benjamin F. Spector and Christopher R\u00e9. 2025. ParallelKittens: Systematic and Practical Simplification of Multi-GPU AI Kernels. arXiv:2511.13940. Retrieved from https:\/\/arxiv.org\/abs\/2511.13940"},{"key":"e_1_3_2_143_2","unstructured":"The Unified Communication X Library [n. d.]. The Unified Communication X Library. Retrieved May 13 2026 from http:\/\/www.openucx.org"},{"key":"e_1_3_2_144_2","article-title":"GPUDirect Storage: A direct path between storage and GPU memory","author":"Thompson Adam","year":"2012","unstructured":"Adam Thompson and Newburn C. J.2012. GPUDirect Storage: A direct path between storage and GPU memory. Retrieved May 13, 2026 from https:\/\/developer.nvidia.com\/blog\/gpudirect-storage\/","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_145_2","article-title":"Early Experience with NVSHMEM: Extending the Kokkos Programming Model with PGAS Semantics","author":"Trott Christian","year":"2018","unstructured":"Christian Trott. 2018. Early Experience with NVSHMEM: Extending the Kokkos Programming Model with PGAS Semantics. Retrieved November 2, 2025 from https:\/\/www.osti.gov\/servlets\/purl\/1806950","journal-title":"Retrieved November 2, 2025 from"},{"key":"e_1_3_2_146_2","doi-asserted-by":"crossref","first-page":"298","DOI":"10.1145\/3712285.3759774","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201925)","author":"Trotter James D.","year":"2025","unstructured":"James D. Trotter, Sinan Ekmek\u00e7iba\u015f\u0131, Do\u011fan Sa\u011fbili, Johannes Langguth, Xing Cai, and Didem Unat. 2025. CPU- and GPU-initiated communication strategies for conjugate gradient methods on large GPU clusters. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201925). Association for Computing Machinery, New York, NY, USA, 298\u2013315. DOI:10.1145\/3712285.3759774"},{"key":"e_1_3_2_147_2","doi-asserted-by":"crossref","first-page":"384","DOI":"10.1145\/3721145.3730426","volume-title":"Proceedings of the 39th ACM International Conference on Supercomputing (ICS\u201925)","author":"Turimbetov Ilyas","year":"2025","unstructured":"Ilyas Turimbetov, Mohamed Wahib, and Didem Unat. 2025. A device-side execution model for multi-GPU task graphs. In Proceedings of the 39th ACM International Conference on Supercomputing (ICS\u201925). Association for Computing Machinery, New York, NY, USA, 384\u2013396. DOI:10.1145\/3721145.3730426"},{"key":"e_1_3_2_148_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCSE.2025.3567586"},{"key":"e_1_3_2_149_2","doi-asserted-by":"publisher","DOI":"10.6084\/m9.figshare.32032020"},{"key":"e_1_3_2_150_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2025.3534638"},{"key":"e_1_3_2_151_2","first-page":"151","volume-title":"Proceedings of the 2017 46th International Conference on Parallel Processing (ICPP)","author":"Venkatesh Akshay","year":"2017","unstructured":"Akshay Venkatesh, Khaled Hamidouche, Sreeram Potluri, Davide Rosetti, Ching-Hsiang Chu, and Dhabaleswar K. Panda. 2017. MPI-GDS: High performance MPI designs with GPUDirect-aSync for CPU-GPU control flow decoupling. In Proceedings of the 2017 46th International Conference on Parallel Processing (ICPP). 151\u2013160. DOI:10.1109\/ICPP.2017.24"},{"key":"e_1_3_2_152_2","doi-asserted-by":"crossref","first-page":"843","DOI":"10.1109\/ISCA.2018.00075","volume-title":"Proceedings of the 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA)","author":"Vesel\u00fd J\u00e1n","year":"2018","unstructured":"J\u00e1n Vesel\u00fd, Arkaprava Basu, Abhishek Bhattacharjee, Gabriel H. Loh, Mark Oskin, and Steven K. Reinhardt. 2018. Generic system calls for GPUs. In Proceedings of the 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). 843\u2013856. DOI:10.1109\/ISCA.2018.00075"},{"key":"e_1_3_2_153_2","article-title":"GTC 20: Overcoming latency barriers: Strong scaling HPC applications with NVSHMEM","author":"Wagner Mathias","year":"2020","unstructured":"Mathias Wagner. 2020. GTC 20: Overcoming latency barriers: Strong scaling HPC applications with NVSHMEM. Retrieved May 13, 2026 from https:\/\/www.nvidia.com\/en-us\/on-demand\/session\/gtcsj20-s21673\/","journal-title":"Retrieved May 13, 2026 from"},{"key":"e_1_3_2_154_2","unstructured":"Guanhua Wang Shivaram Venkataraman Amar Phanishayee Jorgen Thelin Nikhil Devanur and Ion Stoica. 2019. Blink: Fast and Generic Collectives for Distributed ML. arXiv:1910.04940. Retrieved from https:\/\/arxiv.org\/abs\/1910.04940"},{"key":"e_1_3_2_155_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2013.222"},{"key":"e_1_3_2_156_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00450-011-0171-3"},{"key":"e_1_3_2_157_2","first-page":"308","volume-title":"Proceedings of the 2011 IEEE International Conference on Cluster Computing","author":"Wang Hao","year":"2011","unstructured":"Hao Wang, Sreeram Potluri, Miao Luo, Ashish Kumar Singh, Xiangyong Ouyang, Sayantan Sur, and Dhabaleswar K. Panda. 2011. Optimized non-contiguous MPI datatype communication for GPU clusters: Design, implementation and evaluation with MVAPICH2. In Proceedings of the 2011 IEEE International Conference on Cluster Computing. 308\u2013316. DOI:10.1109\/CLUSTER.2011.42"},{"key":"e_1_3_2_158_2","volume-title":"Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201921)","author":"Wang Yuke","year":"2023","unstructured":"Yuke Wang, Boyuan Feng, Zheng Wang, Tong Geng, Ang Li, Kevin Barker, and Yufei Ding. 2023. MGG: Accelerating graph neural networks with fine-grained intra-kernel communication-computation pipelining on multi-GPU platforms. In Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201921)."},{"key":"e_1_3_2_159_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11390-023-2894-6"},{"key":"e_1_3_2_160_2","series-title":"HPDC\u201916","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1145\/2907294.2907317","volume-title":"Proceedings of the 25th ACM International Symposium on High-Performance Parallel and Distributed Computing","author":"Wu Wei","year":"2016","unstructured":"Wei Wu, George Bosilca, Rolf vandeVaart, Sylvain Jeaugey, and Jack Dongarra. 2016. GPU-aware non-contiguous data movement in open MPI. In Proceedings of the 25th ACM International Symposium on High-Performance Parallel and Distributed Computing (Kyoto, Japan) (HPDC\u201916). Association for Computing Machinery, New York, NY, USA, 231\u2013242. DOI:10.1145\/2907294.2907317"},{"key":"e_1_3_2_161_2","series-title":"ICPP\u201921","volume-title":"Proceedings of the 50th International Conference on Parallel Processing","author":"XIE CHENHAO","year":"2021","unstructured":"CHENHAO XIE, Jieyang Chen, Jesun Firoz, Jiajia Li, Shuaiwen Leon Song, Kevin Barker, Mark Raugas, and Ang Li. 2021. Fast and scalable sparse triangular solver for multi-GPU based HPC architectures. In Proceedings of the 50th International Conference on Parallel Processing (Lemont, IL, USA) (ICPP\u201921). Association for Computing Machinery, New York, NY, USA, Article 53, 11 pages. DOI:10.1145\/3472456.3472478"},{"key":"e_1_3_2_162_2","first-page":"667","volume-title":"Proceedings of the 22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25)","author":"Xu Guanbin","year":"2025","unstructured":"Guanbin Xu, Zhihao Le, Yinhe Chen, Zhiqi Lin, Zewen Jin, Youshan Miao, and Cheng Li. 2025. AutoCCL: Automated collective communication tuning for accelerating distributed and parallel DNN training. In Proceedings of the 22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25). USENIX Association, Philadelphia, PA, 667\u2013683. Retrieved from https:\/\/www.usenix.org\/conference\/nsdi25\/presentation\/xu-guanbin"},{"key":"e_1_3_2_163_2","unstructured":"Junchao Zhang Jed Brown Satish Balay Jacob Faibussowitsch Matthew Knepley Oana Marin Richard Tran Mills Todd Munson Barry F. Smith and Stefano Zampini. 2021. The PetscSF Scalable Communication Layer. arXiv:2102.13018. Retrieved from https:\/\/arxiv.org\/abs\/2102.13018"},{"key":"e_1_3_2_164_2","series-title":"ICS\u201923","doi-asserted-by":"crossref","first-page":"167","DOI":"10.1145\/3577193.3593705","volume-title":"Proceedings of the 37th International Conference on Supercomputing","author":"Zhang Lingqi","year":"2023","unstructured":"Lingqi Zhang, Mohamed Wahib, Peng Chen, Jintao Meng, Xiao Wang, Toshio Endo, and Satoshi Matsuoka. 2023. PERKS: A locality-optimized execution model for iterative memory-bound GPU applications. In Proceedings of the 37th International Conference on Supercomputing (Orlando, FL, USA) (ICS\u201923). Association for Computing Machinery, New York, NY, USA, 167\u2013179. DOI:10.1145\/3577193.3593705"},{"key":"e_1_3_2_165_2","series-title":"EuroMPI\/USA\u201922","first-page":"1","volume-title":"Proceedings of the 29th European MPI Users\u2019 Group Meeting","author":"Zhou Hui","year":"2022","unstructured":"Hui Zhou, Ken Raffenetti, Yanfei Guo, and Rajeev Thakur. 2022. MPIX stream: An explicit solution to hybrid MPI+X programming. In Proceedings of the 29th European MPI Users\u2019 Group Meeting (Chattanooga, TN, USA) (EuroMPI\/USA\u201922). Association for Computing Machinery, New York, NY, USA, 1\u201310. DOI:10.1145\/3555819.3555820"},{"key":"e_1_3_2_166_2","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/978-3-031-07312-0_1","volume-title":"Proceedings of the High Performance Computing: 37th International Conference, ISC High Performance 2022, Hamburg, Germany, May 29 \u2013 June 2, 2022.","author":"Zhou Qinghua","year":"2022","unstructured":"Qinghua Zhou, Pouya Kousha, Quentin Anthony, Kawthar Shafie Khorassani, Aamir Shafi, Hari Subramoni, and Dhabaleswar K. Panda. 2022. Accelerating MPI all-to-all communication with online compression on modern GPU clusters. In Proceedings of the High Performance Computing: 37th International Conference, ISC High Performance 2022, Hamburg, Germany, May 29 \u2013 June 2, 2022.Springer-Verlag, Berlin, 3\u201325. DOI:10.1007\/978-3-031-07312-0_1"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3813799","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T12:21:42Z","timestamp":1781007702000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3813799"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,9]]},"references-count":165,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2026,9,30]]}},"alternative-id":["10.1145\/3813799"],"URL":"https:\/\/doi.org\/10.1145\/3813799","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,9]]},"assertion":[{"value":"2024-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-24","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-09","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}