{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T16:46:25Z","timestamp":1782405985348,"version":"3.54.5"},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2022YFB4501400"],"award-info":[{"award-number":["2022YFB4501400"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Graph Neural Networks (GNNs) are becoming increasingly popular in graph data processing due to their excellent performance in feature extraction on graph datasets. Compared to GPUs, CPUs are more widely accessible and serve as a practical platform for GNN inference. However, achieving efficient GNN execution on CPUs remains a challenge. We first comprehensively evaluate and quantitatively analyze the performance of GNN inference on multi-core CPUs using the state-of-the-art frameworks, identifying four key performance bottlenecks: inefficient sparse computation, poor data locality, workload imbalance, and inefficient General Matrix Multiplication (GEMM). To tackle these issues, we introduce a set of joint optimizations. Specifically, for the aggregation phase, we propose three optimizations: a register padding and tiling Graph Sparse-dense Matrix Multiplication (GSpMM) algorithm that leverages the computation capability of long vector processing units on modern multi-core CPUs, a destination node-oriented indexes reorganization to enhance data locality, and a boundary buffer-based method to balance the workloads. Additionally, for the update phase, we develop an efficient bias fusion GEMM algorithm, tailored for the irregular matrices. We evaluate the proposed optimizations extensively with three popular GNN models on three typical multi-core CPU platforms. Experimental results on Intel, AMD, and ARM platforms show that our optimizations outperform the state-of-the-art GNN framework DGL by an average factor of 2.41\u00d7, 1.58\u00d7, and 2.04\u00d7 (up to 4.75\u00d7, 2.70\u00d7, and 3.55\u00d7), respectively. Compared to PyG, our implementations achieve an average speedup of 1.70\u00d7, 1.86\u00d7, and 2.44\u00d7, respectively.<\/jats:p>","DOI":"10.1145\/3807957","type":"journal-article","created":{"date-parts":[[2026,4,16]],"date-time":"2026-04-16T11:27:34Z","timestamp":1776338854000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Quantitative Analysis and Performance Optimization of Graph Neural Networks on Multi-core CPUs"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5128-6748","authenticated-orcid":false,"given":"Kangkang","family":"Chen","sequence":"first","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3587-0917","authenticated-orcid":false,"given":"Huayou","family":"Su","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4675-5866","authenticated-orcid":false,"given":"Xi","family":"Yang","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-2817-0573","authenticated-orcid":false,"given":"Zitong","family":"An","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1256-8934","authenticated-orcid":false,"given":"Yong","family":"Dou","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2765-8541","authenticated-orcid":false,"given":"Dongsheng","family":"Li","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","unstructured":"Xin Ai Qiange Wang Chunyu Cao Yanfeng Zhang Chaoyi Chen Hao Yuan Yu Gu and Ge Yu. 2023. NeutronOrch: Rethinking sample-based GNN training under CPU-GPU heterogeneous environments. DOI:10.48550\/arXiv.2311.13225","DOI":"10.48550\/arXiv.2311.13225"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1209"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","unstructured":"Yukuo Cen Zhenyu Hou Yan Wang Qibin Chen Yizhen Luo Xingcheng Yao Aohan Zeng Shiguang Guo Peng Zhang Guohao Dai et\u00a0al. 2021. Cogdl: An extensive toolkit for deep learning on graphs. DOI:10.48550\/arXiv.2103.00959","DOI":"10.48550\/arXiv.2103.00959"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-39698-4_25"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-69766-1_24"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3542598"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313488"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","unstructured":"Matthias Fey and Jan Eric Lenssen. 2019. Fast graph representation learning with PyTorch Geometric. DOI:10.48550\/arXiv.1903.02428","DOI":"10.48550\/arXiv.1903.02428"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO50266.2020.00079"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480113"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41467-021-23303-9"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3625549.3658655"},{"key":"e_1_3_1_14_2","first-page":"1605","volume-title":"Proceedings of the 2025 USENIX Annual Technical Conference","author":"Gong Yidong","year":"2025","unstructured":"Yidong Gong, Arnab Kanti Tarafder, Saima Afrin, and Pradeep Kumar. 2025. Identifying and Analyzing Pitfalls in \\(\\lbrace\\) GNN \\(\\rbrace\\) Systems. In Proceedings of the 2025 USENIX Annual Technical Conference. 1605\u20131624."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3470496.3527403"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/1377603.1377607"},{"key":"e_1_3_1_17_2","unstructured":"Will Hamilton Zhitao Ying and Jure Leskovec. 2017. Inductive representation learning on large graphs. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems\u201917. Long Beach CA USA 1024\u20131034."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btae497"},{"key":"e_1_3_1_19_2","first-page":"981","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Heinecke Alexander","year":"2016","unstructured":"Alexander Heinecke, Greg Henry, Maxwell Hutchinson, and Hans Pabst. 2016. LIBXSMM: Accelerating small matrix multiplications by runtime code generation. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 981\u2013991."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","unstructured":"Loc Hoang Rita Brugarolas Brufau Ke Ding and Bo Wu. 2023. Batchgnn: Efficient cpu-based distributed gnn training on very large graphs. DOI:10.48550\/arXiv.2306.13814","DOI":"10.48550\/arXiv.2306.13814"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_22_2","unstructured":"Weihua Hu Matthias Fey Marinka Zitnik Yuxiao Dong Hongyu Ren Bowen Liu Michele Catasta and Jure Leskovec. 2020. Open graph benchmark: Datasets for machine learning on graphs. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (NeurIPS\u201920). Virtual 22118\u201322133."},{"key":"e_1_3_1_23_2","first-page":"1","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Huang Guyue","year":"2020","unstructured":"Guyue Huang, Guohao Dai, Yu Wang, and Huazhong Yang. 2020. Ge-spmm: General-purpose sparse matrix-matrix multiplication on gpus for graph neural networks. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1\u201312."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.14778\/3648160.3648184"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3627703.3650063"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3437801.3441585"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2019\/369"},{"key":"e_1_3_1_28_2","unstructured":"Thomas N. Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In 5th International Conference on Learning Representations (ICLR\u201917) Toulon France Conference Track Proceedings. OpenReview.net Toulon France 2713\u20132726."},{"issue":"10","key":"e_1_3_1_29_2","first-page":"1995","article-title":"Convolutional networks for images, speech, and time series","volume":"3361","author":"LeCun Yann","year":"1995","unstructured":"Yann LeCun, Yoshua Bengio, et\u00a0al. 1995. Convolutional networks for images, speech, and time series. The Handbook of Brain Theory and Neural Networks 3361, 10 (1995), 1995.","journal-title":"The Handbook of Brain Theory and Neural Networks"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2024.3371332"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3480856"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.5555\/3104322.3104425"},{"key":"e_1_3_1_33_2","first-page":"5021","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Nguyen Binh","year":"2022","unstructured":"Binh Nguyen, Long Nguyen, and Dinh Dien. 2022. Multi-level community-awareness graph neural networks for neural machine translation. In Proceedings of the 29th International Conference on Computational Linguistics. 5021\u20135028."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS49936.2021.00034"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-93417-4_38"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.14778\/3641204.3641219"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3731545.3731575"},{"key":"e_1_3_1_38_2","first-page":"495","volume-title":"Proceedings of the 15th USENIX Symposium on Operating Systems Design and Implementation","author":"Thorpe John","year":"2021","unstructured":"John Thorpe, Yifan Qiao, Jonathan Eyolfson, Shen Teng, Guanzhou Hu, Zhihao Jia, Jinliang Wei, Keval Vora, Ravi Netravali, Miryung Kim, et\u00a0al. 2021. Dorylus: Affordable, scalable, and accurate \\(\\lbrace\\) GNN \\(\\rbrace\\) training with distributed \\(\\lbrace\\) CPU \\(\\rbrace\\) servers and serverless threads. In Proceedings of the 15th USENIX Symposium on Operating Systems Design and Implementation. 495\u2013514."},{"issue":"3","key":"e_1_3_1_39_2","first-page":"1","article-title":"BLIS: A framework for rapidly instantiating BLAS functionality","volume":"41","author":"Zee Field G. Van","year":"2015","unstructured":"Field G. Van Zee and Robert A. Van De Geijn. 2015. BLIS: A framework for rapidly instantiating BLAS functionality. ACM Transactions on Mathematical Software 41, 3 (2015), 1\u201333.","journal-title":"ACM Transactions on Mathematical Software"},{"key":"e_1_3_1_40_2","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017. Long Beach CA USA 5998\u20136008."},{"key":"e_1_3_1_41_2","volume-title":"6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings","author":"Velickovic Petar","year":"2018","unstructured":"Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Li\u00f2, and Yoshua Bengio. 2018. Graph attention networks. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, Vancouver, BC, Canada."},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCAS46773.2023.10182227"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3447786.3456229"},{"key":"e_1_3_1_44_2","unstructured":"Minjie Wang Lingfan Yu and others. 2019. Deep graph library: Towards efficient and scalable deep learning on graphs. In ICLR Workshop on Representation Learning on Graphs and Manifolds."},{"key":"e_1_3_1_45_2","first-page":"515","volume-title":"Proceedings of the 15th USENIX Symposium on Operating Systems Design and Implementation","author":"Wang Yuke","year":"2021","unstructured":"Yuke Wang, Boyuan Feng, Gushu Li, Shuangchen Li, Lei Deng, Yuan Xie, and Yufei Ding. 2021. \\(\\lbrace\\) GNNAdvisor \\(\\rbrace\\) : An adaptive and efficient runtime system for \\(\\lbrace\\) GNN \\(\\rbrace\\) acceleration on \\(\\lbrace\\) GPUs \\(\\rbrace\\) . In Proceedings of the 15th USENIX Symposium on Operating Systems Design and Implementation. 515\u2013531."},{"key":"e_1_3_1_46_2","first-page":"149","volume-title":"Proceedings of the 2023 USENIX Annual Technical Conference .","author":"Wang Yuke","year":"2023","unstructured":"Yuke Wang, Boyuan Feng, Zheng Wang, Guyue Huang, and Yufei Ding. 2023. \\(\\lbrace\\) TC-GNN \\(\\rbrace\\) : Bridging sparse \\(\\lbrace\\) GNN \\(\\rbrace\\) computation and dense tensor cores on \\(\\lbrace\\) GPUs \\(\\rbrace\\) . In Proceedings of the 2023 USENIX Annual Technical Conference .149\u2013164."},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2023.3257507"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.2978386"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2021.3085578"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPADS.2012.97"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3580305.3599320"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-96983-1_48"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476217"},{"key":"e_1_3_1_54_2","first-page":"593","volume-title":"Proceedings of the Asian Conference on Machine Learning","author":"Yang Yulei","year":"2020","unstructured":"Yulei Yang and Dongsheng Li. 2020. Nenn: Incorporate node and edge features in graph neural networks. In Proceedings of the Asian Conference on Machine Learning. PMLR, 593\u2013608."},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2023.3305077"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3488560.3498414"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3807957","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:54:31Z","timestamp":1782402871000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3807957"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":55,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3807957"],"URL":"https:\/\/doi.org\/10.1145\/3807957","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2025-10-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-30","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}