{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,20]],"date-time":"2025-12-20T22:13:09Z","timestamp":1766268789974,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":22,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,8,5]],"date-time":"2019-08-05T00:00:00Z","timestamp":1564963200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,8,5]]},"DOI":"10.1145\/3337821.3337889","type":"proceedings-article","created":{"date-parts":[[2019,7,25]],"date-time":"2019-07-25T12:34:36Z","timestamp":1564058076000},"page":"1-10","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Gossip"],"prefix":"10.1145","author":[{"given":"Robin","family":"Kobus","sequence":"first","affiliation":[{"name":"Johannes Gutenberg University, Mainz, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Daniel","family":"J\u00fcnger","sequence":"additional","affiliation":[{"name":"Johannes Gutenberg University, Mainz, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christian","family":"Hundt","sequence":"additional","affiliation":[{"name":"NVIDIA AI Technology Center, Luxembourg, Luxembourg"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bertil","family":"Schmidt","sequence":"additional","affiliation":[{"name":"Johannes Gutenberg University, Mainz, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,8,5]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/2851141.2851169"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00453-005-1167-9"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/71.642949"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/1285358.1285359"},{"key":"e_1_3_2_1_5_1","unstructured":"NVIDIA Corporation. 2019. NVIDIA Collective Communications Library (NCCL). https:\/\/developer.nvidia.com\/nccl  NVIDIA Corporation. 2019. NVIDIA Collective Communications Library (NCCL). https:\/\/developer.nvidia.com\/nccl"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2010.38"},{"volume-title":"Bandwidth Efficient All-to-All Broadcast on Switched Clusters. In 2005 IEEE Int. Conference on Cluster Computing. 1--10","author":"Faraj A.","key":"e_1_3_2_1_7_1","unstructured":"A. Faraj , P. Patarasuk , and X. Yuan . 2005 . Bandwidth Efficient All-to-All Broadcast on Switched Clusters. In 2005 IEEE Int. Conference on Cluster Computing. 1--10 . A. Faraj, P. Patarasuk, and X. Yuan. 2005. Bandwidth Efficient All-to-All Broadcast on Switched Clusters. In 2005 IEEE Int. Conference on Cluster Computing. 1--10."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1137\/S0097539703427215"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1016\/0166-218X(94)90180-5"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2013.96"},{"key":"e_1_3_2_1_11_1","unstructured":"Google. 2019. OR-Tools. https:\/\/developers.google.com\/optimization\/  Google. 2019. OR-Tools. https:\/\/developers.google.com\/optimization\/"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/181014.181427"},{"volume-title":"Proc. Symp. (IPDPS). 441--450","author":"J\u00fcnger D.","key":"e_1_3_2_1_13_1","unstructured":"D. J\u00fcnger , C. Hundt , and B. Schmidt . 2018. WarpDrive: Massively Parallel Hashing on Multi-GPU Nodes. In 2018 IEEE Int. Par. and Distr . Proc. Symp. (IPDPS). 441--450 . D.J\u00fcnger, C. Hundt, and B. Schmidt. 2018. WarpDrive: Massively Parallel Hashing on Multi-GPU Nodes. In 2018 IEEE Int. Par. and Distr. Proc. Symp. (IPDPS). 441--450."},{"key":"e_1_3_2_1_14_1","volume-title":"Proc., Workshops and Phd Forum (IPDPSW)","author":"Kandalla Krishna Chaitanya","year":"2010","unstructured":"Krishna Chaitanya Kandalla , Hari Subramoni , Abhinav Vishnu , and D.K. Panda . 2010. Designing topology-aware collective communication algorithms for large scale InfiniBand clusters: Case studies with Scatter and Gather. 2010 IEEE Int. Symp. on Par. & Distr . Proc., Workshops and Phd Forum (IPDPSW) ( 2010 ), 1--8. Krishna Chaitanya Kandalla, Hari Subramoni, Abhinav Vishnu, and D.K. Panda. 2010. Designing topology-aware collective communication algorithms for large scale InfiniBand clusters: Case studies with Scatter and Gather. 2010 IEEE Int. Symp. on Par. & Distr. Proc., Workshops and Phd Forum (IPDPSW) (2010), 1--8."},{"volume-title":"Proc. 14th Int. Par. and Distr. Proc. Symp. (IPDPS). 377--384","author":"Karonis N. T.","key":"e_1_3_2_1_15_1","unstructured":"N. T. Karonis , B. R. de Supinski , I. Foster , W. Gropp , E. Lusk , and J. Bresnahan . 2000. Exploiting hierarchy in parallel computer networks to optimize collective operation performance . In Proc. 14th Int. Par. and Distr. Proc. Symp. (IPDPS). 377--384 . N. T. Karonis, B. R. de Supinski, I. Foster, W. Gropp, E. Lusk, and J. Bresnahan. 2000. Exploiting hierarchy in parallel computer networks to optimize collective operation performance. In Proc. 14th Int. Par. and Distr. Proc. Symp. (IPDPS). 377--384."},{"key":"e_1_3_2_1_16_1","volume-title":"Processing Symp., IPDPS 2006","author":"Patarasuk P","year":"2006","unstructured":"P Patarasuk , A Faraj , and Xin Yuan . 2006 . Pipelined broadcast on Ethernet switched clusters. 20th Int. Par. and Distr . Processing Symp., IPDPS 2006 2006. P Patarasuk, A Faraj, and Xin Yuan. 2006. Pipelined broadcast on Ethernet switched clusters. 20th Int. Par. and Distr. Processing Symp., IPDPS 2006 2006."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2008.09.002"},{"key":"e_1_3_2_1_18_1","volume-title":"Efficient All-to-All Communication Patterns in Hypercube and Mesh Topologies. In 6th Distr. Memory Computing Conf., 1991. Proc. 398--403","author":"Scott D. S.","year":"1991","unstructured":"D. S. Scott . 1991 . Efficient All-to-All Communication Patterns in Hypercube and Mesh Topologies. In 6th Distr. Memory Computing Conf., 1991. Proc. 398--403 . D. S. Scott. 1991. Efficient All-to-All Communication Patterns in Hypercube and Mesh Topologies. In 6th Distr. Memory Computing Conf., 1991. Proc. 398--403."},{"volume-title":"Proc. of 8th Int. Parallel Processing Symposium. 561--565","author":"Thakur R.","key":"e_1_3_2_1_19_1","unstructured":"R. Thakur and A. Choudhary . 1994. All-to-all communication on meshes with wormhole routing . In Proc. of 8th Int. Parallel Processing Symposium. 561--565 . R. Thakur and A. Choudhary. 1994. All-to-all communication on meshes with wormhole routing. In Proc. of 8th Int. Parallel Processing Symposium. 561--565."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/71.503775"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.3867"},{"key":"e_1_3_2_1_22_1","first-page":"217","article-title":"Optimized InfiniBandTM fat-tree routing for shift all-to-all communication patterns","volume":"22","author":"Zahavi E.","year":"2010","unstructured":"E. Zahavi , G. Johnson , D.J. Kerbyson , and M. Lang . 2010 . Optimized InfiniBandTM fat-tree routing for shift all-to-all communication patterns . CCPE 22 , 2 (2010), 217 -- 231 . E. Zahavi, G.Johnson, D.J. Kerbyson, and M. Lang. 2010. Optimized InfiniBandTM fat-tree routing for shift all-to-all communication patterns. CCPE 22, 2 (2010), 217--231.","journal-title":"CCPE"}],"event":{"name":"ICPP 2019: 48th International Conference on Parallel Processing","sponsor":["University of Tsukuba University of Tsukuba"],"location":"Kyoto Japan","acronym":"ICPP 2019"},"container-title":["Proceedings of the 48th International Conference on Parallel Processing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3337821.3337889","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3337821.3337889","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:25:41Z","timestamp":1750206341000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3337821.3337889"}},"subtitle":["Efficient Communication Primitives for Multi-GPU Systems"],"short-title":[],"issued":{"date-parts":[[2019,8,5]]},"references-count":22,"alternative-id":["10.1145\/3337821.3337889","10.1145\/3337821"],"URL":"https:\/\/doi.org\/10.1145\/3337821.3337889","relation":{},"subject":[],"published":{"date-parts":[[2019,8,5]]},"assertion":[{"value":"2019-08-05","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}