{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T15:23:55Z","timestamp":1780932235384,"version":"3.54.1"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"11","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,7]]},"abstract":"<jats:p>Graph Convolutional Networks (GCNs) are increasingly adopted in large-scale graph-based recommender systems. Training GCN requires the minibatch generator traversing graphs and sampling the sparsely located neighboring nodes to obtain their features. Since real-world graphs often exceed the capacity of GPU memory, current GCN training systems keep the feature table in host memory and rely on the CPU to collect sparse features before sending them to the GPUs. This approach, however, puts tremendous pressure on host memory bandwidth and the CPU. This is because the CPU needs to (1) read sparse features from memory, (2) write features into memory as a dense format, and (3) transfer the features from memory to the GPUs.<\/jats:p>\n          <jats:p>In this work, we propose a novel GPU-oriented data communication approach for GCN training, where GPU threads directly access sparse features in host memory through zero-copy accesses without much CPU help. By removing the CPU gathering stage, our method significantly reduces the consumption of the host resources and data access latency. We further present two important techniques to achieve high host memory access efficiency by the GPU: (1) automatic data access address alignment to maximize PCIe packet efficiency, and (2) asynchronous zero-copy access and kernel execution to fully overlap data transfer with training. We incorporate our method into PyTorch and evaluate its effectiveness using several graphs with sizes up to 111 million nodes and 1.6 billion edges. In a multi-GPU training setup, our method is 65--92% faster than the conventional data transfer method, and can even match the performance of all-in-GPU-memory training for some graphs that fit in GPU memory.<\/jats:p>","DOI":"10.14778\/3476249.3476264","type":"journal-article","created":{"date-parts":[[2021,10,27]],"date-time":"2021-10-27T16:46:23Z","timestamp":1635353183000},"page":"2087-2100","source":"Crossref","is-referenced-by-count":61,"title":["Large graph convolutional network training with GPU-oriented data communication architecture"],"prefix":"10.14778","volume":"14","author":[{"given":"Seung Won","family":"Min","sequence":"first","affiliation":[{"name":"UIUC"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kun","family":"Wu","sequence":"additional","affiliation":[{"name":"UIUC"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sitao","family":"Huang","sequence":"additional","affiliation":[{"name":"UIUC"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mert","family":"Hidayeto\u011flu","sequence":"additional","affiliation":[{"name":"UIUC"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinjun","family":"Xiong","sequence":"additional","affiliation":[{"name":"IBM T.J. Watson Research Center"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eiman","family":"Ebrahimi","sequence":"additional","affiliation":[{"name":"NVIDIA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Deming","family":"Chen","sequence":"additional","affiliation":[{"name":"UIUC"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wen-mei","family":"Hwu","sequence":"additional","affiliation":[{"name":"UIUC"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,10,27]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Retrieved","year":"2021"},{"key":"e_1_2_1_2_1","volume-title":"Retrieved","year":"2021"},{"key":"e_1_2_1_3_1","unstructured":"K. Bhatia K. Dahiya H. Jain A. Mittal Y. Prabhu and M. Varma. 2016. The extreme classification repository: Multi-label datasets and code. http:\/\/manikvarma.org\/downloads\/XC\/XMLRepository.html  K. Bhatia K. Dahiya H. Jain A. Mittal Y. Prabhu and M. Varma. 2016. The extreme classification repository: Multi-label datasets and code. http:\/\/manikvarma.org\/downloads\/XC\/XMLRepository.html"},{"key":"e_1_2_1_4_1","volume-title":"International Conference on Learning Representations (ICLR2014)","author":"Bruna Joan","year":"2014"},{"key":"e_1_2_1_5_1","volume-title":"International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=rytstxWAW","author":"Chen Jie","year":"2018"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research), Jennifer Dy and Andreas Krause (Eds.)","volume":"80","author":"Chen Jianfei","year":"2018"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330925"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/3157382.3157527"},{"key":"e_1_2_1_9_1","volume-title":"Fast Graph Representation Learning with PyTorch Geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds.","author":"Fey Matthias"},{"key":"e_1_2_1_10_1","volume-title":"SIGN: Scalable Inception Graph Neural Networks. In ICML 2020 Workshop on Graph Representation Learning and Beyond.","author":"Frasca Fabrizio","year":"2020"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the Thirty-forth International Conference on Parallel and Distributed Processing (IPDPS).","author":"Ganguly Debashis","year":"2020"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/3384345.3384358"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939754"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/3294771.3294869"},{"key":"e_1_2_1_15_1","first-page":"52","article-title":"Representation Learning on Graphs","volume":"40","author":"Hamilton William L.","year":"2017","journal-title":"Methods and Applications. IEEE Data Eng. Bull."},{"key":"e_1_2_1_16_1","unstructured":"Mark Harris. 2013. How to Access Global Memory Efficiently in CUDA C\/C++ Kernels. https:\/\/developer.nvidia.com\/blog\/how-access-global-memory-efficiently-cuda-c-kernels\/  Mark Harris. 2013. How to Access Global Memory Efficiently in CUDA C\/C++ Kernels. https:\/\/developer.nvidia.com\/blog\/how-access-global-memory-efficiently-cuda-c-kernels\/"},{"key":"e_1_2_1_17_1","unstructured":"Mark Harris. 2017. Unified Memory for CUDA Beginners. https:\/\/developer.nvidia.com\/blog\/unified-memory-cuda-beginners\/  Mark Harris. 2017. Unified Memory for CUDA Beginners. https:\/\/developer.nvidia.com\/blog\/unified-memory-cuda-beginners\/"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"e_1_2_1_19_1","volume-title":"Open Graph Benchmark: Datasets for Machine Learning on Graphs. arXiv preprint arXiv:2005.00687","author":"Hu Weihua","year":"2020"},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of Machine Learning and Systems (MLSys)","author":"Jia Zhihao","year":"2020"},{"key":"e_1_2_1_21_1","volume-title":"Variational Graph Auto-Encoders. NIPS Workshop on Bayesian Deep Learning","author":"Kipf Thomas N","year":"2016"},{"key":"e_1_2_1_22_1","volume-title":"Semi-Supervised Classification with Graph Convolutional Networks. In International Conference on Learning Representations (ICLR).","author":"Thomas"},{"key":"e_1_2_1_23_1","volume-title":"Accelerating Recommender Systems via Hardware \"scale-in\". CoRR abs\/2009.05230","author":"Krishna Suresh","year":"2020"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2487788.2488173"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1989.1.4.541"},{"key":"e_1_2_1_26_1","volume-title":"Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=SkhQHMW0W","author":"Lin Yujun","year":"2018"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.5555\/3358807.3358845"},{"key":"e_1_2_1_28_1","volume-title":"Zaid Qureshi, Jinjun Xiong, Eiman Ebrahimi, and Wen-mei Hwu.","author":"Min Seung Won","year":"2020"},{"key":"e_1_2_1_29_1","volume-title":"15th USENIX Symposium on Operating Systems Design and Implementation (OSDI 21)","author":"Mohoney Jason","year":"2021"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/3488766.3488817"},{"key":"e_1_2_1_31_1","unstructured":"Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G. Azzolini Dmytro Dzhulgakov Andrey Mallevich Ilia Cherniavskii Yinghai Lu Raghuraman Krishnamoorthi Ansha Yu Volodymyr Kondratenko Stephanie Pereira Xianjie Chen Wenlin Chen Vijay Rao Bill Jia Liang Xiong and Misha Smelyanskiy. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems. CoRR abs\/1906.00091 (2019). https:\/\/arxiv.org\/abs\/1906.00091  Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G. Azzolini Dmytro Dzhulgakov Andrey Mallevich Ilia Cherniavskii Yinghai Lu Raghuraman Krishnamoorthi Ansha Yu Volodymyr Kondratenko Stephanie Pereira Xianjie Chen Wenlin Chen Vijay Rao Bill Jia Liang Xiong and Misha Smelyanskiy. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems. CoRR abs\/1906.00091 (2019). https:\/\/arxiv.org\/abs\/1906.00091"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3230543.3230560"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1018"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.5555\/3045390.3045603"},{"key":"e_1_2_1_35_1","unstructured":"Nvidia. 2016. Nvidia Tesla P100 Whitepaper. https:\/\/images.nvidia.com\/content\/pdf\/tesla\/whitepaper\/pascal-architecture-whitepaper.pdf  Nvidia. 2016. Nvidia Tesla P100 Whitepaper. https:\/\/images.nvidia.com\/content\/pdf\/tesla\/whitepaper\/pascal-architecture-whitepaper.pdf"},{"key":"e_1_2_1_36_1","unstructured":"Nvidia. 2017. Nvidia Tesla V100 GPU Architecture Whitepaper. https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf  Nvidia. 2017. Nvidia Tesla V100 GPU Architecture Whitepaper. https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf"},{"key":"e_1_2_1_37_1","unstructured":"NVIDIA. 2020. MULTI-PROCESS SERVICE. https:\/\/docs.nvidia.com\/deploy\/pdf\/CUDA_Multi_Process_Service_Overview.pdf  NVIDIA. 2020. MULTI-PROCESS SERVICE. https:\/\/docs.nvidia.com\/deploy\/pdf\/CUDA_Multi_Process_Service_Overview.pdf"},{"key":"e_1_2_1_38_1","unstructured":"Nvidia. 2020. Nvidia A100 TensorCore GPU Architecture Whitepaper. https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/Data-Center\/nvidia-ampere-architecture-whitepaper.pdf  Nvidia. 2020. Nvidia A100 TensorCore GPU Architecture Whitepaper. https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/Data-Center\/nvidia-ampere-architecture-whitepaper.pdf"},{"key":"e_1_2_1_39_1","unstructured":"NVIDIA. 2020. NVIDIA MULTI-INSTANCE GPU USERGUIDE. https:\/\/docs.nvidia.com\/datacenter\/tesla\/pdf\/NVIDIA_MIG_User_Guide.pdf  NVIDIA. 2020. NVIDIA MULTI-INSTANCE GPU USERGUIDE. https:\/\/docs.nvidia.com\/datacenter\/tesla\/pdf\/NVIDIA_MIG_User_Guide.pdf"},{"key":"e_1_2_1_40_1","unstructured":"NVIDIA. 2021. KERNEL PROFILING GUIDE. https:\/\/docs.nvidia.com\/nsight-compute\/pdf\/ProfilingGuide.pdf  NVIDIA. 2021. KERNEL PROFILING GUIDE. https:\/\/docs.nvidia.com\/nsight-compute\/pdf\/ProfilingGuide.pdf"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3297663.3310299"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2623330.2623732"},{"key":"e_1_2_1_44_1","unstructured":"PyTorch. 2021. TorchElastic. https:\/\/pytorch.org\/elastic\/0.2.2\/index.html  PyTorch. 2021. TorchElastic. https:\/\/pytorch.org\/elastic\/0.2.2\/index.html"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.5555\/3195638.3195660"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3342195.3387537"},{"key":"e_1_2_1_47_1","unstructured":"Tim Schroeder. 2011. Peer-to-Peer & Unified Virtual Addressing. https:\/\/developer.download.nvidia.com\/CUDA\/training\/cuda_webinars_GPUDirect_uva.pdf  Tim Schroeder. 2011. Peer-to-Peer & Unified Virtual Addressing. https:\/\/developer.download.nvidia.com\/CUDA\/training\/cuda_webinars_GPUDirect_uva.pdf"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/HOTCHIPS.2019.8875650"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2015-354"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2017.2761740"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.308"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1162\/qss_a_00021"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3200691.3178491"},{"key":"e_1_2_1_55_1","volume-title":"Highly-Performant Package for Graph Neural Networks. arXiv preprint arXiv:1909.01315","author":"Wang Minjie","year":"2019"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.2978386"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219890"},{"key":"e_1_2_1_58_1","volume-title":"GraphSAINT: Graph Sampling Based Inductive Learning Method. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=BJe8pkHFwS","author":"Zeng Hanqing","year":"2020"},{"key":"e_1_2_1_59_1","doi-asserted-by":"crossref","unstructured":"Da Zheng Chao Ma Minjie Wang Jinjing Zhou Qidong Su Xiang Song Quan Gan Zheng Zhang and George Karypis. 2020. DistDGL: Distributed Graph Neural Network Training for Billion-Scale Graphs. arXiv:2010.05337 [cs.LG]  Da Zheng Chao Ma Minjie Wang Jinjing Zhou Qidong Su Xiang Song Quan Gan Zheng Zhang and George Karypis. 2020. DistDGL: Distributed Graph Neural Network Training for Billion-Scale Graphs. arXiv:2010.05337 [cs.LG]","DOI":"10.1109\/IA351965.2020.00011"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3476249.3476264","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T09:58:19Z","timestamp":1672221499000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3476249.3476264"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7]]},"references-count":57,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2021,7]]}},"alternative-id":["10.14778\/3476249.3476264"],"URL":"https:\/\/doi.org\/10.14778\/3476249.3476264","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,7]]}}}