{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T18:08:01Z","timestamp":1780510081590,"version":"3.54.1"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,11,14]],"date-time":"2023-11-14T00:00:00Z","timestamp":1699920000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2024,2,29]]},"abstract":"<jats:p>\n            In recent years, semi-supervised graph learning with data augmentation (DA) has been the most commonly used and best-performing method to improve model robustness in sparse scenarios with few labeled samples. However, most existing DA methods are based on the homogeneous graph, but none are specific for the heterogeneous graph. Differing from the homogeneous graph, DA in the heterogeneous graph faces greater challenges: heterogeneity of information requires DA strategies to effectively handle heterogeneous relations, which considers the information contribution of different types of neighbors and edges to the target nodes. Furthermore, over-squashing of information is caused by the negative curvature formed by the non-uniformity distribution and the strong clustering in a complex graph. To address these challenges, this article presents a novel method named\n            <jats:italic>HG-MDA<\/jats:italic>\n            (Semi-Supervised Heterogeneous Graph Learning with Multi-Level Data Augmentation). For the problem of heterogeneity of information in DA, node and topology augmentation strategies are proposed for the characteristics of the heterogeneous graph. Additionally, meta-relation-based attention is applied as one of the indexes for selecting augmented nodes and edges. For the problem of over-squashing of information, triangle-based edge adding and removing are designed to alleviate the negative curvature and bring the gain of topology. Finally, the loss function consists of the cross-entropy loss for labeled data and the consistency regularization for unlabeled data. To effectively fuse the prediction results of various DA strategies, sharpening is used. Existing experiments on public datasets (i.e., ACM, DBLP, and OGB) and the industry dataset MB show that HG-MDA outperforms current SOTA models. Additionally, HG-MDA is applied to user identification in internet finance scenarios, helping the business to add 30% key users, and increase loans and balances by 3.6%, 11.1%, and 9.8%.\n          <\/jats:p>","DOI":"10.1145\/3608953","type":"journal-article","created":{"date-parts":[[2023,8,15]],"date-time":"2023-08-15T09:23:31Z","timestamp":1692091411000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Semi-Supervised Heterogeneous Graph Learning with Multi-Level Data Augmentation"],"prefix":"10.1145","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7546-3822","authenticated-orcid":false,"given":"Ying","family":"Chen","sequence":"first","affiliation":[{"name":"MYbank, Ant Group, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1625-4366","authenticated-orcid":false,"given":"Siwei","family":"Qiang","sequence":"additional","affiliation":[{"name":"MYbank, Ant Group, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2901-9608","authenticated-orcid":false,"given":"Mingming","family":"Ha","sequence":"additional","affiliation":[{"name":"School of Automation and Electrical Engineering, University of Science and Technology Beijing, and MYbank, Ant Group, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7082-641X","authenticated-orcid":false,"given":"Xiaolei","family":"Liu","sequence":"additional","affiliation":[{"name":"MYbank, Ant Group, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2949-668X","authenticated-orcid":false,"given":"Shaoshuai","family":"Li","sequence":"additional","affiliation":[{"name":"MYbank, Ant Group, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-1048-1255","authenticated-orcid":false,"given":"Jiabi","family":"Tong","sequence":"additional","affiliation":[{"name":"MYbank, Ant Group, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4107-5801","authenticated-orcid":false,"given":"Lingfeng","family":"Yuan","sequence":"additional","affiliation":[{"name":"MYbank, Ant Group, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6817-626X","authenticated-orcid":false,"given":"Xiaobo","family":"Guo","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Beijing Jiaotong University, and MYbank, Ant Group, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7315-3276","authenticated-orcid":false,"given":"Zhenfeng","family":"Zhu","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Beijing Jiaotong University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,11,14]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0378-8733(03)00009-1"},{"key":"e_1_3_1_3_2","article-title":"On the bottleneck of graph neural networks and its practical implications","author":"Alon Uri","year":"2020","unstructured":"Uri Alon and Eran Yahav. 2020. On the bottleneck of graph neural networks and its practical implications. arXiv preprint arXiv:2006.05205 (2020).","journal-title":"arXiv preprint arXiv:2006.05205"},{"key":"e_1_3_1_4_2","unstructured":"D. Berthelot N. Carlini I. Goodfellow N. Papernot A. Oliver and C. Raffel. 2019. MixMatch: A holistic approach to semi-supervised learning. arXiv:1905-02249 (2019)."},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","unstructured":"P. Bielak T. Kajdanowicz and N. V. Chawla. 2021. Graph Barlow Twins: A self-supervised representation learning framework for graphs. arXiv:2106.02466 (2021).","DOI":"10.1016\/j.knosys.2022.109631"},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","unstructured":"Deyu Bo BinBin Hu Xiao Wang Zhiqiang Zhang Chuan Shi and Jun Zhou. 2022. Regularizing graph neural networks via consistency-diversity graph augmentations. Proceedings of the AAAI Conference on Artificial Intelligence 36 4 (2022) 3913\u20133921.","DOI":"10.1609\/aaai.v36i4.20307"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3271768"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.032093399"},{"key":"e_1_3_1_9_2","first-page":"22092","article-title":"Graph random neural networks for semi-supervised learning on graphs","volume":"33","author":"Feng Wenzheng","year":"2020","unstructured":"Wenzheng Feng, Jie Zhang, Yuxiao Dong, Yu Han, Huanbo Luan, Qian Xu, Qiang Yang, Evgeny Kharlamov, and Jie Tang. 2020. Graph random neural networks for semi-supervised learning on graphs. Advances in Neural Information Processing Systems 33 (2020), 22092\u201322103.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_10_2","article-title":"Graph random neural network","volume":"2005","author":"Feng Wenzheng","year":"2020","unstructured":"Wenzheng Feng, Jie Zhang, Yuxiao Dong, Yu Han, Huanbo Luan, Qian Xu, Qiang Yang, and Jie Tang. 2020. Graph random neural network. CoRR abs\/2005.11079 (2020). https:\/\/arxiv.org\/abs\/2005.11079","journal-title":"CoRR"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00454-002-0743-x"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","unstructured":"X. Fu J. Zhang Z. Meng and I. King. 2020. MAGNN: Metapath aggregated graph neural network for heterogeneous graph embedding. arXiv:2002.01680 (2020).","DOI":"10.1145\/3366423.3380297"},{"key":"e_1_3_1_13_2","article-title":"Semi-supervised learning by entropy minimization","volume":"17","author":"Grandvalet Yves","year":"2004","unstructured":"Yves Grandvalet and Yoshua Bengio. 2004. Semi-supervised learning by entropy minimization. Advances in Neural Information Processing Systems 17 (2004), 1\u20138.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_14_2","volume-title":"Proceedings of the nternational Joint Conference on Artificial Intelligence (IJCAI\u201922)","author":"Gu Shuyun","year":"2022","unstructured":"Shuyun Gu, Xiao Wang, Chuan Shi, and Ding Xiao. 2022. Self-supervised graph neural networks for multi-behavior recommendation. In Proceedings of the nternational Joint Conference on Artificial Intelligence (IJCAI\u201922)."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.3301946"},{"key":"e_1_3_1_16_2","volume-title":"(WWW\u201920)","author":"Hu Z.","year":"2020","unstructured":"Z. Hu, Y. Dong, K. Wang, and Y. Sun. 2020. Heterogeneous graph transformer. In Proceedings of The Web Conference 2020(WWW\u201920). 2704\u20132710."},{"key":"e_1_3_1_17_2","article-title":"Semi-supervised classification with graph convolutional networks","author":"Kipf Thomas N.","year":"2016","unstructured":"Thomas N. Kipf and Max Welling. 2016. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907 (2016).","journal-title":"arXiv preprint arXiv:1609.02907"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1080\/15427951.2012.625260"},{"key":"e_1_3_1_19_2","first-page":"896","volume-title":"Proceedings of the Workshop on Challenges in Representation Learning (ICML\u201913)","volume":"3","year":"2013","unstructured":"Dong-Hyun Lee. 2013. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Proceedings of the Workshop on Challenges in Representation Learning (ICML\u201913), Vol. 3. 896."},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","unstructured":"N. Liu X. Wang L. Wu Y. Chen X. Guo and C. Shi. 2022. Compact graph structure learning via mutual information compression. arXiv:2201.05540 (2022).","DOI":"10.1145\/3485447.3512206"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482480"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2858821"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1137\/S003614450342480"},{"key":"e_1_3_1_24_2","article-title":"Self-supervised graph representation learning via global context prediction","author":"Peng Zhen","year":"2020","unstructured":"Zhen Peng, Yixiang Dong, Minnan Luo, Xiao-Ming Wu, and Qinghua Zheng. 2020. Self-supervised graph representation learning via global context prediction. arXiv preprint arXiv:2003.01604 (2020).","journal-title":"arXiv preprint arXiv:2003.01604"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/2623330.2623732"},{"key":"e_1_3_1_26_2","article-title":"Semi-supervised learning with ladder networks","volume":"28","author":"Rasmus Antti","year":"2015","unstructured":"Antti Rasmus, Mathias Berglund, Mikko Honkala, Harri Valpola, and Tapani Raiko. 2015. Semi-supervised learning with ladder networks. Advances in Neural Information Processing Systems 28 (2015), 1\u20139.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"2","key":"e_1_3_1_27_2","doi-asserted-by":"crossref","first-page":"026112","DOI":"10.1103\/PhysRevE.67.026112","article-title":"Hierarchical organization in complex networks","volume":"67","author":"Ravasz E.","year":"2003","unstructured":"E. Ravasz and A. L. Barab\u00e1si. 2003. Hierarchical organization in complex networks. Physical Review E Statistical Nonlinear & Soft Matter Physics 67, 2 Pt. 2 (2003), 026112.","journal-title":"Physical Review E Statistical Nonlinear & Soft Matter Physics"},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","first-page":"164","DOI":"10.4018\/978-1-5225-2814-2.ch010","volume-title":"Graph Theoretic Approaches for Analyzing Large-Scale Social Networks","author":"Samanta Sovan","year":"2018","unstructured":"Sovan Samanta and Madhumangal Pal. 2018. Link prediction in social networks. In Graph Theoretic Approaches for Analyzing Large-Scale Social Networks. IGI Global, Hershey, PA, 164\u2013172."},{"key":"e_1_3_1_29_2","article-title":"Modeling Relational Data with Graph Convolutional Networks","author":"Schlichtkrull M.","year":"2018","unstructured":"M. Schlichtkrull, T. N. Kipf, P. Bloem, Rianne Vanden Berg, and M. Welling. 2018. Modeling Relational Data with Graph Convolutional Networks. Springer, Cham, Switzerland.","journal-title":"Springer, Cham"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2016.2598561"},{"key":"e_1_3_1_31_2","unstructured":"S. Thakoor C. Tallec M. G. Azar R. Munos P. Velikovi and M. Valko. 2021. Bootstrapped representation learning on graphs. In Proceedings of the 2021 Workshop on Geometrical and Topological Representation Learning (ICLR\u201921) ."},{"key":"e_1_3_1_32_2","article-title":"Understanding over-squashing and bottlenecks on graphs via curvature","author":"Topping Jake","year":"2021","unstructured":"Jake Topping, Francesco Di Giovanni, Benjamin Paul Chamberlain, Xiaowen Dong, and Michael M. Bronstein. 2021. Understanding over-squashing and bottlenecks on graphs via curvature. arXiv preprint arXiv:2111.14522 (2021).","journal-title":"arXiv preprint arXiv:2111.14522"},{"issue":"7","key":"e_1_3_1_33_2","first-page":"1","article-title":"Microsoft academic graph: When experts are not enough","volume":"1","author":"Wang K.","year":"2020","unstructured":"K. Wang, Zhihong Shen, Chi Yuan Huang, Chieh Han Wu, and Anshul Kanakia. 2020. Microsoft academic graph: When experts are not enough. Quantitative Science Studies 1, 7 (2020), 1\u201318.","journal-title":"Quantitative Science Studies"},{"key":"e_1_3_1_34_2","article-title":"A survey on heterogeneous graph embedding: Methods, techniques, applications and sources","author":"Wang Xiao","year":"2022","unstructured":"Xiao Wang, Deyu Bo, Chuan Shi, Shaohua Fan, Yanfang Ye, and S. Yu Philip. 2022. A survey on heterogeneous graph embedding: Methods, techniques, applications and sources. IEEE Transactions on Big Data 2022 (2022), 1\u201320.","journal-title":"IEEE Transactions on Big Data"},{"key":"e_1_3_1_35_2","doi-asserted-by":"crossref","unstructured":"X. Wang H. Ji C. Shi B. Wang P. Cui P Yu and Y. Ye. 2019. Heterogeneous graph attention network. arXiv:1903.07293 (2019).","DOI":"10.1145\/3308558.3313562"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403063"},{"key":"e_1_3_1_37_2","unstructured":"Stanley Wasserman and Katherine Faust. 1994. Social Network Analysis: Methods and Applications . Structural Analysis in the Social Sciences Series Number 9. Cambridge University Press."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1038\/30918"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1086\/653658"},{"key":"e_1_3_1_40_2","unstructured":"L. Wu H. Lin Z. Gao C. Tan and S. Z. Li. 2021. Self-supervised learning on graphs: Contrastive generative or predictive. arXiv:2105.07342 (2021)."},{"key":"e_1_3_1_41_2","unstructured":"Q. Xie Z. Dai E. Hovy M. T. Luong and Q. V. Le. 2019. Unsupervised data augmentation for consistency training. arXiv:1904.12848 (2019)."},{"key":"e_1_3_1_42_2","article-title":"A survey on deep semi-supervised learning","author":"Yang Xiangli","year":"2021","unstructured":"Xiangli Yang, Zixing Song, Irwin King, and Zenglin Xu. 2021. A survey on deep semi-supervised learning. arXiv preprint arXiv:2103.00550 (2021).","journal-title":"arXiv preprint arXiv:2103.00550"},{"key":"e_1_3_1_43_2","unstructured":"Y. You T. Chen Y. Sui T. Chen Z. Wang and Y. Shen. 2020. Graph contrastive learning with augmentations. arXiv:2010.13902 (2020)."},{"key":"e_1_3_1_44_2","unstructured":"S. Yun M. Jeong R. Kim J. Kang and H. J. Kim. 2019. Graph transformer networks. arXiv:1911.06455 (2019)."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00156"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330961"},{"key":"e_1_3_1_47_2","volume-title":"Proceedings of the 25th ACM SIGKDD International Conference","author":"Zhang C.","year":"2019","unstructured":"C. Zhang, D. Song, C. Huang, A. Swami, and N. V. Chawla. 2019. Heterogeneous graph neural network. In Proceedings of the 25th ACM SIGKDD International Conference."},{"key":"e_1_3_1_48_2","article-title":"A semi-supervised learning approach for COVID-19 detection from chest CT scans","author":"Zhang Yong","year":"2022","unstructured":"Yong Zhang, Li Su, Zhenxing Liu, Wei Tan, Yinuo Jiang, and Cheng Cheng. 2022. A semi-supervised learning approach for COVID-19 detection from chest CT scans. Neurocomputing 503 (2022), 314\u2013324.","journal-title":"Neurocomputing"},{"key":"e_1_3_1_49_2","doi-asserted-by":"crossref","unstructured":"Jianan Zhao Qianlong Wen Shiyu Sun Yanfang Ye and Chuxu Zhang. 2021. Multi-view self-supervised heterogeneous graph embedding. In Machine Learning and Knowledge Discovery in Databases. Research Track . Lecture Notes in Computer Science Vol. 12976. Springer 319\u2013334.","DOI":"10.1007\/978-3-030-86520-7_20"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330851"},{"key":"e_1_3_1_51_2","first-page":"912","volume-title":"Proceedings of the 20th International Conference on Machine Learning (ICML\u201903)","author":"Zhu Xiaojin","year":"2003","unstructured":"Xiaojin Zhu, Zoubin Ghahramani, and John D. Lafferty. 2003. Semi-supervised learning using Gaussian fields and harmonic functions. In Proceedings of the 20th International Conference on Machine Learning (ICML\u201903). 912\u2013919."},{"key":"e_1_3_1_52_2","unstructured":"Y. Zhu Y. Xu Q. Liu and S. Wu. 2021. An empirical study of graph contrastive learning. arXiv:2109.01116 (2021)."},{"key":"e_1_3_1_53_2","unstructured":"Y. Zhu Y. Xu F. Yu Q. Liu and L. Wang. 2020. Deep graph contrastive representation learning. arXiv:2006.04131 (2020)."},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3449802"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3608953","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3608953","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:29:46Z","timestamp":1750285786000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3608953"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,14]]},"references-count":53,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,2,29]]}},"alternative-id":["10.1145\/3608953"],"URL":"https:\/\/doi.org\/10.1145\/3608953","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,14]]},"assertion":[{"value":"2022-11-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-30","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}