{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T14:47:49Z","timestamp":1782571669608,"version":"3.54.5"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T00:00:00Z","timestamp":1782518400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"Doctoral Research Foundation of Nanyang Normal University, China","award":["2023ZX008"],"award-info":[{"award-number":["2023ZX008"]}]},{"name":"Joint Fund for Science and Technology Research and Development of Henan Province","award":["245101610059"],"award-info":[{"award-number":["245101610059"]}]},{"name":"Key Research and Development and Promotion Program of Henan Province, China","award":["242102210184 and 252102210247"],"award-info":[{"award-number":["242102210184 and 252102210247"]}]},{"name":"National Natural Science Foundation Cultivation Program of China","award":["2025PY036"],"award-info":[{"award-number":["2025PY036"]}]},{"name":"Science and Technology Program of Nanyang City, China","award":["23JCQY2003"],"award-info":[{"award-number":["23JCQY2003"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    Sparse General Matrix Multiplication (SpMM) is a fundamental kernel in iterative scientific computing and graph neural network inference. On single-GPU platforms, iterative SpMM suffers from two major performance bottlenecks: excessive memory traffic caused by repeated loading of Krylov basis vectors, and low computational efficiency due to irregular sparsity patterns. To address these challenges, this article proposes\n                    <jats:italic toggle=\"yes\">BLR-Krylov<\/jats:italic>\n                    , an iterative SpMM optimization framework that integrates Communication-Avoiding Krylov (CA-Krylov) methods with a Block Low-Rank Sparse (BLR-Sparse) matrix format. BLR-Krylov first preprocesses sparse matrices into the BLR-Sparse format, where low-rank blocks are compressed through basis vector decomposition to reduce storage and arithmetic overhead. It then employs GPU-optimized matrix power kernels to batch-compute Krylov subspace basis vectors, effectively minimizing global memory accesses across iterations. Experimental results on multiple GPU architectures demonstrate that BLR-Krylov achieves a 4.0\u20136.0\u00d7 speedup over existing communication-avoiding approaches in iterative workloads, increases global memory bandwidth utilization to 82%\u201387%, and improves tensor core utilization to 85%\u201392%, while maintaining numerical accuracy within 1 \u00d7 10\n                    <jats:sup>-6<\/jats:sup>\n                    . To facilitate reproducibility, partial source code is publicly available at https:\/\/github.com\/19547035579zz-tech\/BLR-Krylov.\n                  <\/jats:p>","DOI":"10.1145\/3815589","type":"journal-article","created":{"date-parts":[[2026,5,11]],"date-time":"2026-05-11T08:45:10Z","timestamp":1778489110000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["BLR-Krylov: A Single-GPU Iterative SpMM Framework with Communication Avoidance and Block Low-Rank Optimization"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6024-9698","authenticated-orcid":false,"given":"Wen","family":"Li","sequence":"first","affiliation":[{"name":"College of Artificial Intelligence and Software Engineering, Nanyang Normal University","place":["Nanyang, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-1683-3267","authenticated-orcid":false,"given":"Zheng","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Artificial Intelligence and Software Engineering, Nanyang Normal University","place":["Nanyang, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-8731-1552","authenticated-orcid":false,"given":"Jie","family":"Zhao","sequence":"additional","affiliation":[{"name":"College of Artificial Intelligence and Software Engineering, Nanyang Normal University","place":["Nanyang, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-6570-8872","authenticated-orcid":false,"given":"Song","family":"Zhu","sequence":"additional","affiliation":[{"name":"College of Artificial Intelligence and Software Engineering, Nanyang Normal University","place":["Nanyang, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2915-8656","authenticated-orcid":false,"given":"Ming","family":"Hui","sequence":"additional","affiliation":[{"name":"College of Artificial Intelligence and Software Engineering, Nanyang Normal University","place":["Nanyang, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,27]]},"reference":[{"key":"e_1_3_3_2_2","volume-title":"The Finite Element Method: Linear Static and Dynamic Finite Element Analysis","author":"Hughes Thomas J. R.","year":"1987","unstructured":"Thomas J. R. Hughes. 1987. The Finite Element Method: Linear Static and Dynamic Finite Element Analysis. PrenticeHall."},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1137\/0907058"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.6028\/jres.049.044"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2008.2005605"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/355791.355796"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898718003"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3712285.3759895"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/2930660"},{"key":"e_1_3_3_10_2","volume-title":"Communication-Avoiding Krylov Subspace Methods","author":"Hoemmen Mark","year":"2010","unstructured":"Mark Hoemmen. 2010. Communication-Avoiding Krylov Subspace Methods. Ph.D. Dissertation. University of California, Berkeley."},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1137\/15M1010117"},{"key":"e_1_3_3_12_2","volume-title":"cuBLAS Library User Guide","year":"2025","unstructured":"NVIDIA Corporation 2025. cuBLAS Library User Guide. NVIDIA Corporation. Retrieved from https:\/\/docs.nvidia.com\/cuda\/cublas\/index.html"},{"key":"e_1_3_3_13_2","volume-title":"CUTLASS: CUDA Templates for Linear Algebra Subroutines","year":"2025","unstructured":"NVIDIA Corporation 2025. CUTLASS: CUDA Templates for Linear Algebra Subroutines. NVIDIA Corporation. Retrieved from https:\/\/github.com\/NVIDIA\/cutlass"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/1413370.1413402"},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611971538"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/356924.356930"},{"key":"e_1_3_3_17_2","volume-title":"NVIDIA GH200 Grace Hopper Superchip Architecture Whitepaper","year":"2023","unstructured":"NVIDIA Corporation 2023. NVIDIA GH200 Grace Hopper Superchip Architecture Whitepaper. NVIDIA Corporation. Retrieved from https:\/\/resources.nvidia.com\/en-us-grace-hopper-superchip\/grace-hopper-superchip-whitepaper"},{"key":"e_1_3_3_18_2","volume-title":"NVIDIA Nsight Compute User Guide","year":"2024","unstructured":"NVIDIA Corporation 2024. NVIDIA Nsight Compute User Guide. NVIDIA Corporation. Retrieved from https:\/\/docs.nvidia.com\/nsight-compute\/NsightCompute\/index.html"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1142\/9789812836021_0015"},{"key":"e_1_3_3_20_2","volume-title":"CUDA C++ Programming Guide","year":"2024","unstructured":"NVIDIA Corporation 2024. CUDA C++ Programming Guide. NVIDIA Corporation. Retrieved from https:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/index.html"},{"key":"e_1_3_3_21_2","volume-title":"HIP Programming Guide","year":"2024","unstructured":"Advanced Micro Devices, Inc. 2024. HIP Programming Guide. Advanced Micro Devices, Inc. Retrieved from https:\/\/rocm.docs.amd.com\/projects\/HIP\/en\/latest\/"},{"key":"e_1_3_3_22_2","volume-title":"SYCL 2020 Specification","year":"2021","unstructured":"Khronos Group 2021. SYCL 2020 Specification. Khronos Group. Retrieved from https:\/\/www.khronos.org\/registry\/SYCL\/specs\/sycl-2020\/html\/sycl-2020.html"},{"key":"e_1_3_3_23_2","volume-title":"Proceedings of the 2023 IEEE Hot Chips 35 Symposium (HCS)","author":"Smith Alan","year":"2023","unstructured":"Alan Smith et\u00a0al. 2023. AMD instinct\u2122MI300 accelerator. In Proceedings of the 2023 IEEE Hot Chips 35 Symposium (HCS). IEEE."},{"key":"e_1_3_3_24_2","volume-title":"Intel\u00ae GPU Architecture Overview","year":"2022","unstructured":"Intel Corporation 2022. Intel\u00ae GPU Architecture Overview. Intel Corporation. Retrieved from https:\/\/www.intel.com\/content\/www\/us\/en\/docs\/oneapi\/optimization-guide-gpu\/2023-0\/intel-gpu-architecture.html"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2014.2346458"},{"key":"e_1_3_3_26_2","volume-title":"AMD RDNA 3 Architecture","year":"2022","unstructured":"Advanced Micro Devices, Inc. 2022. AMD RDNA 3 Architecture. Advanced Micro Devices, Inc."},{"key":"e_1_3_3_27_2","volume-title":"AMD CDNA 3 Instruction Set Architecture: Reference Guide","year":"2023","unstructured":"Advanced Micro Devices 2023. AMD CDNA 3 Instruction Set Architecture: Reference Guide. Advanced Micro Devices. Retrieved from https:\/\/developer.amd.com\/wp-content\/resources\/CDNA3_Shader_ISA.pdf"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/2597652.2597679"},{"key":"e_1_3_3_29_2","volume-title":"Intel\u00ae oneAPI GPU Optimization Guide","year":"2023","unstructured":"Intel Corporation 2023. Intel\u00ae oneAPI GPU Optimization Guide. Intel Corporation. Retrieved from https:\/\/www.intel.com\/content\/www\/us\/en\/docs\/oneapi\/optimization-guide-gpu\/"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1137\/S1064827502406415"},{"key":"e_1_3_3_31_2","volume-title":"NVIDIA Blackwell Architecture Technical Overview","year":"2024","unstructured":"NVIDIA Corporation 2024. NVIDIA Blackwell Architecture Technical Overview. NVIDIA Corporation."},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/2049662.2049663"},{"key":"e_1_3_3_33_2","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Hu Weihua","year":"2020","unstructured":"Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. 2020. Open graph benchmark: Datasets for machine learning on graphs. In Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_3_34_2","volume-title":"ROCm Profiler (rocprof) User Guide","year":"2024","unstructured":"Advanced Micro Devices, Inc. 2024. ROCm Profiler (rocprof) User Guide. Advanced Micro Devices, Inc."},{"key":"e_1_3_3_35_2","volume-title":"rocSPARSE User Guide","year":"2025","unstructured":"Advanced Micro Devices, Inc. 2025. rocSPARSE User Guide. Advanced Micro Devices, Inc."},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/355984.355989"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898717839"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/2786975"},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","DOI":"10.1137\/110848244"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3274651"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126921"},{"key":"e_1_3_3_42_2","volume-title":"Proceedings of the Platform for Advanced Scientific Computing Conference (PASC)","author":"Charara Ali","year":"2019","unstructured":"Ali Charara, Hatem Ltaief, and David Keyes. 2019. HiCMA: Hierarchical computations on manycore architectures. In Proceedings of the Platform for Advanced Scientific Computing Conference (PASC)."},{"key":"e_1_3_3_43_2","volume-title":"PETSc\/TAO Users Manual","author":"Balay Satish","year":"2021","unstructured":"Satish Balay, Shrirang Abhyankar, Mark F. Adams, et\u00a0al. 2021. PETSc\/TAO Users Manual. Argonne National Laboratory. Retrieved from https:\/\/petsc.org\/"},{"key":"e_1_3_3_44_2","volume-title":"NVIDIA AmgX Library User Guide","year":"2023","unstructured":"NVIDIA Corporation 2023. NVIDIA AmgX Library User Guide. NVIDIA Corporation. Retrieved from https:\/\/developer.nvidia.com\/amgx"},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.acl-long.1126"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3815589","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T14:16:17Z","timestamp":1782569777000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3815589"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,27]]},"references-count":44,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3815589"],"URL":"https:\/\/doi.org\/10.1145\/3815589","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,27]]},"assertion":[{"value":"2026-01-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-03","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}