{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T22:42:59Z","timestamp":1777675379231,"version":"3.51.4"},"reference-count":21,"publisher":"SAGE Publications","issue":"4","license":[{"start":{"date-parts":[[2015,6,25]],"date-time":"2015-06-25T00:00:00Z","timestamp":1435190400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2017,7]]},"abstract":"<jats:p>Due to their massive parallelism and high performance per Watt, GPUs have gained high popularity in high-performance computing and are a strong candidate for future exascale systems. But communication and data transfer in GPU-accelerated systems remain a challenging problem. Since the GPU normally is not able to control a network device, a hybrid-programming model is preferred whereby the GPU is used for calculation and the CPU handles the communication. As a result, communication between distributed GPUs suffers from unnecessary overhead, introduced by switching control flow from GPUs to CPUs and vice versa. Furthermore, often a designated CPU thread is required to control GPU-related communication. In this work, we modify user space libraries and device drivers of GPUs and the InfiniBand network device in a way to enable the GPU to control an InfiniBand network device to independently source and sink communication requests without any involvement of the CPU. Our results show that complex networking protocols such as InfiniBand Verbs are better handled by CPUs, since overhead of work request generation cannot be parallelized and is not suitable for the highly parallel programming model of GPUs. The massive number of instructions and accesses to host memory that is required to source and sink a communication request on the GPU slows down the performance. Only through a massive reduction in the complexity of the InfiniBand protocol can some performance improvements be achieved.<\/jats:p>","DOI":"10.1177\/1094342015588142","type":"journal-article","created":{"date-parts":[[2015,6,26]],"date-time":"2015-06-26T20:25:08Z","timestamp":1435350308000},"page":"274-284","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":1,"title":["InfiniBand Verbs on GPU: a case study of controlling an InfiniBand network device from the GPU"],"prefix":"10.1177","volume":"31","author":[{"given":"Lena","family":"Oden","sequence":"first","affiliation":[{"name":"Competence Center High Performance Computing, Fraunhofer Institute for Industrial Mathematics, Kaisersautern, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Holger","family":"Fr\u00f6ning","sequence":"additional","affiliation":[{"name":"Institute of Computer Engineering, Ruprecht-Karls University of Heidelberg, Mannheim, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2015,6,25]]},"reference":[{"key":"bibr1-1094342015588142","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-19595-2_11"},{"key":"bibr2-1094342015588142","author":"Kerr G","year":"2011","journal-title":"arXiv preprint arXiv:1105.1827"},{"key":"bibr3-1094342015588142","volume-title":"Programming Massively Parallel Processors: A Hands-On Approach","author":"Kirk DB","year":"2010"},{"key":"bibr4-1094342015588142","volume-title":"43rd international conference on parallel processing workshops (ICPPW)","author":"Klenk B","year":"2015"},{"key":"bibr5-1094342015588142","doi-asserted-by":"publisher","DOI":"10.1007\/s00450-009-0088-2"},{"key":"bibr6-1094342015588142","unstructured":"Mellanox Technologies (2012) Mellanox offers GPUDirect RDMA beta. Available at: http:\/\/www.mellanox.com\/page\/products_dyn?product_family=116 (accessed 4 February 2015)."},{"key":"bibr7-1094342015588142","unstructured":"NVIDIA Corporation (2012) Developing a Linux kernel module using RDMA for GPUDirect. Available at: http:\/\/docs.nvidia.com\/cuda\/gpudirect-rdma\/index.html (accessed 4 February 2015)."},{"key":"bibr8-1094342015588142","volume-title":"International conference on parallel computing (ParCo2013)","author":"Oden L","year":"2013"},{"key":"bibr9-1094342015588142","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2013.6702638"},{"key":"bibr10-1094342015588142","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid.2014.21"},{"key":"bibr11-1094342015588142","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid.2013.15"},{"key":"bibr12-1094342015588142","first-page":"617","volume":"42","author":"Pfister GF","year":"2001","journal-title":"High Performance Mass Storage and Parallel I\/O"},{"key":"bibr13-1094342015588142","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2013.104"},{"key":"bibr14-1094342015588142","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2013.17"},{"key":"bibr15-1094342015588142","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2012.228"},{"key":"bibr16-1094342015588142","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2013.179"},{"key":"bibr17-1094342015588142","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2009.5161065"},{"key":"bibr18-1094342015588142","doi-asserted-by":"publisher","DOI":"10.1145\/2377978.2377981"},{"key":"bibr19-1094342015588142","unstructured":"TOP500.org (2013) TSUBAME 2.5 \u2013 cluster platform sl390s g7, Xeon x5670 6c 2. 930 GHz, InfiniBand QDR, NVIDIA k20x. Available at: http:\/\/top500.org\/system\/178249 (accessed 4 February 2015)."},{"key":"bibr20-1094342015588142","unstructured":"TOP500.org (2014) TOP500 supercomputer sites. Available at: www.top500.org (accessed 4 February 2015)"},{"key":"bibr21-1094342015588142","doi-asserted-by":"publisher","DOI":"10.1007\/s00450-011-0171-3"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342015588142","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342015588142","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342015588142","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:15:24Z","timestamp":1777450524000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342015588142"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,6,25]]},"references-count":21,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2017,7]]}},"alternative-id":["10.1177\/1094342015588142"],"URL":"https:\/\/doi.org\/10.1177\/1094342015588142","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2015,6,25]]}}}