{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T20:25:39Z","timestamp":1782937539166,"version":"3.54.5"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2025,3,19]],"date-time":"2025-03-19T00:00:00Z","timestamp":1742342400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program","doi-asserted-by":"crossref","award":["2022YFB4501400"],"award-info":[{"award-number":["2022YFB4501400"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62202451"],"award-info":[{"award-number":["62202451"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"CAS Project for Young Scientists in Basic Research","award":["YSBR-029"],"award-info":[{"award-number":["YSBR-029"]}]},{"name":"CAS Project for Youth Innovation Promotion Association"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>Owing to their remarkable representation capabilities for heterogeneous graph data, Heterogeneous Graph Neural Networks (HGNNs) have been widely adopted in many critical real-world domains such as recommendation systems and medical analysis. Prior to their practical application, identifying the optimal HGNN model parameters tailored to specific tasks through extensive training is a time-consuming and costly process. To enhance the efficiency of HGNN training, it is essential to characterize and analyze the execution semantics and patterns within the training process to identify performance bottlenecks. In this study, we conduct a comprehensive quantification and in-depth analysis of two mainstream HGNN training scenarios, including single-GPU and multi-GPU distributed training. Based on the characterization results, we reveal the performance bottlenecks and their underlying causes in different HGNN training scenarios and propose optimization guidelines from both software and hardware perspectives.<\/jats:p>","DOI":"10.1145\/3703356","type":"journal-article","created":{"date-parts":[[2024,11,4]],"date-time":"2024-11-04T09:48:48Z","timestamp":1730713728000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Characterizing and Understanding HGNN Training on GPUs"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0641-5779","authenticated-orcid":false,"given":"Dengke","family":"Han","sequence":"first","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China and University of the Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6915-955X","authenticated-orcid":false,"given":"Mingyu","family":"Yan","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China and University of the Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4598-1685","authenticated-orcid":false,"given":"Xiaochun","family":"Ye","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China and University of the Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5219-0908","authenticated-orcid":false,"given":"Dongrui","family":"Fan","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China and University of the Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,3,19]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2019.00011"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-022-10375-2"},{"key":"e_1_3_1_4_2","volume-title":"Proceedings of the 50th Annual International Symposium on Computer Architecture (ISCA\u201923)","year":"2023","unstructured":"Dan Chen, Haiheng He, Hai Jin, Long Zheng, Yu Huang, Xinyang Shen, and Xiaofei Liao. 2023. MetaNMP: Leveraging Cartesian-like product to accelerate HGNNs with near-memory processing. In Proceedings of the 50th Annual International Symposium on Computer Architecture (ISCA\u201923). Association for Computing Machinery, New York, NY, USA."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3269281"},{"key":"e_1_3_1_6_2","volume-title":"Proceedings of the ICLR Workshop on Representation Learning on Graphs and Manifolds","author":"Fey Matthias","year":"2019","unstructured":"Matthias Fey and Jan E. Lenssen. 2019. Fast graph representation learning with PyTorch geometric. In Proceedings of the ICLR Workshop on Representation Learning on Graphs and Manifolds."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3366423.3380297"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3470496.3527403"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783759"},{"key":"e_1_3_1_10_2","article-title":"Inductive representation learning on large graphs","volume":"30","author":"Hamilton Will","year":"2017","unstructured":"Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive representation learning on large graphs. Advan. Neural Inf. Process. Syst. 30 (2017).","journal-title":"Advan. Neural Inf. Process. Syst."},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","unstructured":"Dengke Han Meng Wu Runzhen Xue Mingyu Yan Xiaochun Ye and Dongrui Fan. 2024. ADE-HGNN: Accelerating HGNNs through attention disparity exploitation. In Euro-Par 2024: Parallel Processing: 30th European Conference on Parallel and Distributed Processing Madrid Spain August 26\u201330 2024 Proceedings Part II Springer-Verlag Madrid Spain 91\u2013106.","DOI":"10.1007\/978-3-031-69766-1_7"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098026"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC55918.2022.00023"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12652-020-02807-0"},{"issue":"4","key":"e_1_3_1_15_2","first-page":"5141","article-title":"Heterogeneous graph learning for multi-modal medical data analysis","volume":"37","author":"Kim Sein","year":"2023","unstructured":"Sein Kim, Namkyeong Lee, Junseok Lee, Dongmin Hyun, and Chanyoung Park. 2023. Heterogeneous graph learning for multi-modal medical data analysis. Proc. AAAI Conf. Artif. Intell. 37, 4 (June2023), 5141\u20135150.","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"e_1_3_1_16_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201917)","author":"Kipf Thomas N.","year":"2017","unstructured":"Thomas N. Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In Proceedings of the International Conference on Learning Representations (ICLR\u201917)."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2022.3168067"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2023.3337442"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3419111.3421281"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/JAS.2021.1004311"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3481966"},{"key":"e_1_3_1_22_2","first-page":"1150","volume-title":"Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining (KDD\u201921)","year":"2021","unstructured":"Qingsong Lv, Ming Ding, Qiang Liu, Yuxiang Chen, Wenzheng Feng, Siming He, Chang Zhou, Jianguo Jiang, Yuxiao Dong, and Jie Tang. 2021. Are we really making much progress? Revisiting, benchmarking and refining heterogeneous graph neural networks. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining (KDD\u201921). 1150\u20131160."},{"key":"e_1_3_1_23_2","unstructured":"Zhengyang Lv Mingyu Yan Xin Liu Mengyao Dong Xiaochun Ye Dongrui Fan and Ninghui Sun. 2023. A Survey of Graph Pre-processing Methods: From Algorithmic to Hardware Perspectives. arxiv:2309.07581 [cs.AR]"},{"key":"e_1_3_1_24_2","first-page":"1","article-title":"Semi-decentralized inference in heterogeneous graph neural networks for traffic demand forecasting: An edge-computing approach","author":"Nazzal Mahmoud","year":"2024","unstructured":"Mahmoud Nazzal, Abdallah Khreishah, Joyoung Lee, Shaahin Angizi, Ala Al-Fuqaha, and Mohsen Guizani. 2024. Semi-decentralized inference in heterogeneous graph neural networks for traffic demand forecasting: An edge-computing approach. IEEE Trans. Vehic. Technol. (2024), 1\u201316. https:\/\/ieeexplore.ieee.org\/document\/10409531","journal-title":"IEEE Trans. Vehic. Technol."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613424.3614305"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-93417-4_38"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2016.2598561"},{"key":"e_1_3_1_28_2","volume-title":"Proceedings of the USENIX Symposium on Operating Systems Design and Implementation","author":"Thorpe John","year":"2021","unstructured":"John Thorpe, Yifan Qiao, Jon Eyolfson, Shen Teng, Guanzhou Hu, Zhihao Jia, Jinliang Wei, Keval Vora, Ravi Netravali, Miryung Kim, and Guoqing Harry Xu. 2021. Dorylus: Affordable, scalable, and accurate GNN training with distributed CPU servers and serverless threads. In Proceedings of the USENIX Symposium on Operating Systems Design and Implementation."},{"key":"e_1_3_1_29_2","article-title":"Graph attention networks","author":"Veli\u010dkovi\u0107 Petar","year":"2018","unstructured":"Petar Veli\u010dkovi\u0107, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Li\u00f2, and Yoshua Bengio. 2018. Graph attention networks. In Proceedings of the International Conference on Learning Representations (ICLR\u201918).","journal-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201918)"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.295"},{"key":"e_1_3_1_31_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201919)","year":"2019","unstructured":"Minjie Wang, Zheng Da, Ye Zihao, Gan Quan, Mufei Li, Xiang Song, Jinjing Zhou, Chao Ma, Lingfan Yu, Yu Gai, Tianjun Xiao, Tong He, George Karypis, Jinyang Li, and Zheng Zhang. 2019. Deep graph library: Towards efficient and scalable deep learning on graphs. In Proceedings of the International Conference on Learning Representations (ICLR\u201919)."},{"key":"e_1_3_1_32_2","article-title":"A survey on heterogeneous graph embedding: Methods, techniques, applications and sources","author":"Wang Xiao","year":"2020","unstructured":"Xiao Wang, Deyu Bo, Chuan Shi, Shaohua Fan, Yanfang Ye, and Philip S. Yu. 2020. A survey on heterogeneous graph embedding: Methods, techniques, applications and sources. arXiv preprint arXiv:2011.14867 (2020).","journal-title":"arXiv preprint arXiv:2011.14867"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313562"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2021.03.015"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"},{"key":"e_1_3_1_36_2","article-title":"Graph neural networks in recommender systems: A survey","author":"Wu Shiwen","year":"2020","unstructured":"Shiwen Wu, Fei Sun, Wentao Zhang, Xu Xie, and Bin Cui. 2020. Graph neural networks in recommender systems: A survey. ACM Comput. Surv. (2020).","journal-title":"ACM Comput. Surv."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.2978386"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2022.3200480"},{"key":"e_1_3_1_39_2","article-title":"How powerful are graph neural networks?","author":"Xu Keyulu","year":"2018","unstructured":"Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. 2018. How powerful are graph neural networks? arXiv preprint arXiv:1810.00826 (2018).","journal-title":"arXiv preprint arXiv:1810.00826"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2024.3394841"},{"key":"e_1_3_1_41_2","article-title":"GDR-HGNN: A heterogeneous graph neural networks accelerator frontend with graph decoupling and recoupling","volume":"2404","author":"Xue Runzhen","year":"2024","unstructured":"Runzhen Xue, Mingyu Yan, Dengke Han, Yihan Teng, Zhimin Tang, Xiaochun Ye, and Dongrui Fan. 2024. GDR-HGNN: A heterogeneous graph neural networks accelerator frontend with graph decoupling and recoupling. ArXiv abs\/2404.04792 (2024).","journal-title":"ArXiv"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2020.2970395"},{"key":"e_1_3_1_43_2","first-page":"615","volume-title":"Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201952)","year":"2019","unstructured":"Mingyu Yan, Xing Hu, Shuangchen Li, Abanti Basak, Han Li, Xin Ma, Itir Akgun, Yujing Feng, Peng Gu, Lei Deng, Xiaochun Ye, Zhimin Zhang, Dongrui Fan, and Yuan Xie. 2019. Alleviating irregularity in graph analytics acceleration: A hardware\/software co-design approach. In Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201952). Association for Computing Machinery, New York, NY, USA, 615\u2013628."},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2022.3198281"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i9.26283"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.5555\/3367471.3367618"},{"key":"e_1_3_1_47_2","volume-title":"Proceedings of the Conference on Machine Learning and Systems (MLSys\u201922)","author":"Zhang Hengrui","year":"2022","unstructured":"Hengrui Zhang, Zhongming Yu, Guohao Dai, Guyue Huang, Yufei Ding, Yuan Xie, and Yu Wang. 2022. Understanding GNN computational graph: A coordinated computation, IO, and memory perspective. In Proceedings of the Conference on Machine Learning and Systems (MLSys\u201922), Diana Marculescu, Yuejie Chi, and Carole-Jean Wu (Eds.). mlsys.org."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2017.2762308"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12559-021-09830-z"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3539177"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS51616.2021.00073"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.aiopen.2021.01.001"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3703356","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3703356","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:19:03Z","timestamp":1750295943000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3703356"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,19]]},"references-count":51,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3703356"],"URL":"https:\/\/doi.org\/10.1145\/3703356","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,19]]},"assertion":[{"value":"2024-07-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-23","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-19","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}