{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T14:47:50Z","timestamp":1781534870493,"version":"3.54.5"},"reference-count":42,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T00:00:00Z","timestamp":1775174400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Natural Science Foundation of Guangdong Province, China","award":["2024A1515010204"],"award-info":[{"award-number":["2024A1515010204"]}]},{"name":"Huawei Technologies Co., Ltd."},{"id":[{"id":"https:\/\/ror.org\/00cmhce21","id-type":"ROR","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>General Matrix Multiplication (GEMM) is a fundamental computational kernel in scientific computing, serving as the foundation for numerous complex tasks. However, in practical applications, the performance of GEMM is often constrained by irregular matrix dimensions and the diversity of hardware architectures. In particular, when processing small and irregular matrices, GEMM typically exhibits reduced computational efficiency. To address these challenges, this paper proposes a GEMM acceleration method based on an adaptive core grouping strategy. The method consists of two key components: a core grouping mechanism that alleviates workload imbalance among multi-core CPUs, and an adaptive block partitioning algorithm that dynamically selects optimal tiling schemes according to the matrix dimensions, achieving both load balance and cache-friendly data access. Experimental results on the Kunpeng CPU platform demonstrate that the proposed method achieves significant performance improvements compared to the Kunpeng KML math library, reaching a peak acceleration of up to 2.1\u00d7 and an average speedup of 1.64\u00d7. These results validate the effectiveness and efficiency of the proposed approach in handling small and irregular matrix computation scenarios.<\/jats:p>","DOI":"10.3390\/computers15040223","type":"journal-article","created":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T01:12:33Z","timestamp":1775437953000},"page":"223","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["AGP-GEMM: Adaptive Grouping and Partitioning Framework for Accelerating Small and Irregular Matrices on CPUs"],"prefix":"10.3390","volume":"15","author":[{"given":"Hongzhe","family":"Zhou","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, South China University of Technology, Guangzhou 510006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6372-7088","authenticated-orcid":false,"given":"Lu","family":"Lu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, South China University of Technology, Guangzhou 510006, China"},{"name":"Peng Cheng Laboratory, Shenzhen 518055, China"},{"name":"Pazhou Laboratory, Guangzhou 510005, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haibiao","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, South China University of Technology, Guangzhou 510006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, South China University of Technology, Guangzhou 510006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,4,3]]},"reference":[{"key":"ref_1","unstructured":"Venkatesh, S., Muralikrishnan, S., and Narayanan, P.J. (2016). High Performance Computing, Springer. Lecture Notes in Computer Science."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Davies, T., Karlsson, C., Liu, H., Ding, C., and Chen, Z. (2011). High Performance Linpack Benchmark: A Fault-Tolerant Implementation without Checkpointing. Proceedings of the International Conference on Supercomputing (ICS \u201911), ACM.","DOI":"10.1145\/1995896.1995923"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Li, Y., Miao, R., Liu, H.H., Zhuang, Y., Feng, F., Tang, L., Cao, Z., Zhang, M., Kelly, F., and Alizadeh, M. (2019). HPCC: High Precision Congestion Control. Proceedings of the ACM Special Interest Group on Data Communication (SIGCOMM \u201919), ACM.","DOI":"10.1145\/3341302.3342085"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Hwang, R., Kang, M., Lee, J., Kam, D., Lee, Y., and Rhu, M. (2023). GROW: A Row-Stationary Sparse\u2013Dense GEMM Accelerator for Memory-Efficient Graph Convolutional Neural Networks. Proceedings of the IEEE International Symposium on High-Performance Computer Architecture (HPCA), IEEE.","DOI":"10.1109\/HPCA56546.2023.10070983"},{"key":"ref_5","unstructured":"Silva, I.D.A., Carle, T., Gauffriau, A., Jegu, V., and Pagetti, C. (2024). A Predictable SIMD Library for GEMM Routines. Proceedings of the IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS), IEEE."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Meyer, M. (2021, January 21\u201323). Towards Performance Characterization of FPGAs in Context of HPC Using OpenCL Benchmarks. Proceedings of the International Symposium on Highly Efficient Accelerators and Reconfigurable Technologies (HEART), Online.","DOI":"10.1145\/3468044.3468058"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Patankar, S.V., Pollard, A., Singhal, A.K., and Vanka, S.P. (1983). A Calculation Procedure for Heat, Mass and Momentum Transfer in Three-Dimensional Parabolic Flows. Numerical Prediction of Flow, Heat Transfer, Turbulence and Combustion, Pergamon.","DOI":"10.1016\/B978-0-08-030937-8.50013-1"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"100620","DOI":"10.1016\/j.taml.2025.100620","article-title":"Turbulence.ai: An end-to-end AI scientist for fluid mechanics","volume":"16","author":"Feng","year":"2026","journal-title":"Theor. Appl. Mech. Lett."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"2402369","DOI":"10.1002\/adma.202402369","article-title":"In Silico Chemical Experiments in the Age of AI: From Quantum Chemistry to Machine Learning and Back","volume":"36","author":"Aldossary","year":"2024","journal-title":"Adv. Mater."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"215","DOI":"10.1109\/TED.2024.3496445","article-title":"Controlled Acceleration of PCM Cells Time Drift Through On-Chip Current-Induced Annealing for AIMC Multilevel MVM Computation","volume":"72","author":"Antolini","year":"2025","journal-title":"IEEE Trans. Electron. Devices"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Mackin, C., Narayanan, P., Ambrogio, S., Tsai, H., Spoon, K., Fasoli, A., Chen, A., Friz, A., Shelby, R.M., and Burr, G.W. (2020). Neuromorphic Computing with Phase Change: Device Reliability and Variability Challenges. 2020 IEEE International Reliability Physics Symposium (IRPS), IEEE.","DOI":"10.1109\/IRPS45951.2020.9128315"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Spoon, K., Ambrogio, S., Narayanan, P., Tsai, H., Mackin, C., Chen, A., Fasoli, A., Friz, A., and Burr, G.W. (2020). Accelerating Deep Neural Networks with Analog Memory Devices. IEEE International Memory Workshop (IMW), IEEE.","DOI":"10.1109\/IMW48823.2020.9108149"},{"key":"ref_13","unstructured":"(2026, March 20). OpenBLAS: An Optimized BLAS Library. Available online: http:\/\/www.openblas.net\/."},{"key":"ref_14","first-page":"1","article-title":"BLIS: A Framework for Rapidly Instantiating BLAS Functionality","volume":"41","year":"2015","journal-title":"ACM Trans. Math. Softw."},{"key":"ref_15","unstructured":"Arm Ltd. (2026, March 20). Arm Performance Libraries Reference Manual. Arm Developer Documentation. Available online: https:\/\/developer.arm.com."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.parco.2018.10.003","article-title":"Algorithms and Optimization Techniques for High-Performance Matrix\u2013Matrix Multiplications of Very Small Matrices","volume":"81","author":"Masliah","year":"2019","journal-title":"Parallel Comput."},{"key":"ref_17","unstructured":"Axelsson, O. (2007). Solution of Linear Systems of Equations: Iterative Methods. Sparse Matrix Techniques, Springer."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"2036","DOI":"10.1109\/TON.2025.3552423","article-title":"Robust Decentralized Learning with Local Updates and Gradient Tracking","volume":"33","author":"Ghiasvand","year":"2025","journal-title":"IEEE Trans. Netw."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep Learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Santry, D.J. (2024). Convolutional Neural Networks. Demystifying Deep Learning: An Introduction to the Mathematics of Neural Networks, Wiley.","DOI":"10.1002\/9781394205639"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"113150","DOI":"10.1016\/j.asoc.2025.113150","article-title":"Efficient and Performant Transformer Private Inference with Heterogeneous Attention Mechanisms","volume":"176","author":"Hu","year":"2025","journal-title":"Appl. Soft Comput."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Kartsev, A., Malkovsky, S., and Chibisov, A.N. (2021). Analysis of Ionicity-Magnetism Competition in 2D-MX3 Halides towards a Low-Dimensional Materials Study Based on GPU-Enabled Computational Systems. Nanomaterials, 11.","DOI":"10.3390\/nano11112967"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Andrade, X., Pemmaraju, C.D., Kartsev, A., Xiao, J., Lindenberg, A., Rajpurohit, S., Tan, L.Z., Ogitsu, T., and Correa, A.A. (2021). INQ, a Modern GPU-Accelerated Computational Framework for (Time-Dependent) Density Functional Theory. arXiv.","DOI":"10.1021\/acs.jctc.1c00562"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1109\/JPROC.2022.3226481","article-title":"Efficient Acceleration of Deep Learning Inference on Resource-Constrained Edge Devices: A Review","volume":"111","author":"Shuvo","year":"2023","journal-title":"Proc. IEEE"},{"key":"ref_25","first-page":"1","article-title":"A Survey of CPU-GPU Heterogeneous Computing Techniques","volume":"47","author":"Mittal","year":"2015","journal-title":"ACM Comput. Surv."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"260","DOI":"10.1109\/71.993206","article-title":"Performance-Effective and Low-Complexity Task Scheduling for Heterogeneous Computing","volume":"13","author":"Topcuoglu","year":"2002","journal-title":"IEEE Trans. Parallel Distrib. Syst."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Stratton, J.A., Stone, S.S., and Hwu, W.M.W. (2008). MCUDA: An Efficient Implementation of CUDA Kernels for Multi-Core CPUs. Languages and Compilers for Parallel Computing, Springer.","DOI":"10.1007\/978-3-540-89740-8_2"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Gepner, P., and Kowalik, M.F. (2006). Multi-Core Processors: New Way to Achieve High System Performance. International Symposium on Parallel Computing in Electrical Engineering, IEEE.","DOI":"10.1109\/PARELEC.2006.54"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1145\/3274654","article-title":"SCP: Shared Cache Partitioning for High-Performance GEMM","volume":"15","author":"Su","year":"2018","journal-title":"ACM Trans. Archit. Code Optim."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1016\/j.jpdc.2019.08.003","article-title":"Task Packing: Efficient Task Scheduling in Unbalanced Parallel Programs to Maximize CPU Utilization","volume":"134","author":"Utrera","year":"2019","journal-title":"J. Parallel Distrib. Comput."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"232","DOI":"10.1145\/2914770.2837669","article-title":"The Hardness of Data Packing","volume":"51","author":"Lavaee","year":"2016","journal-title":"ACM SIGPLAN Not."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"359","DOI":"10.1145\/1735971.1736060","article-title":"Inter-Core Cooperative TLB for Chip Multiprocessors","volume":"45","author":"Bhattacharjee","year":"2010","journal-title":"ACM SIGPLAN Not."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1016\/j.sysarc.2019.06.006","article-title":"Exploring Heterogeneous Scheduling for Edge Computing with CPU and FPGA MPSoCs","volume":"98","author":"Navarro","year":"2019","journal-title":"J. Syst. Archit."},{"key":"ref_34","first-page":"5998","article-title":"Attention Is All You Need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Tandon, A., Manju, K.M., and Patel, S. (2024). A New VMP Approach Based on CPU and Memory Using Bin Packing. 2024 IEEE Pune Section International Conference (PuneCon), IEEE.","DOI":"10.1109\/PuneCon63413.2024.10894908"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Hestness, J., Keckler, S.W., and Wood, D.A. (2014). A Comparative Analysis of Microarchitecture Effects on CPU and GPU Memory System Behavior. 2014 IEEE International Symposium on Workload Characterization (IISWC), IEEE.","DOI":"10.1109\/IISWC.2014.6983054"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Liu, H., Shi, S., Wang, X., Jiang, Z.L., and Chen, Q. (2024). Performance Analysis and Optimizations of Matrix Multiplications on ARMv8 Processors. 2024 Design, Automation & Test in Europe Conference & Exhibition (DATE), IEEE.","DOI":"10.23919\/DATE58400.2024.10546786"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Heinecke, A., Henry, G., Hutchinson, M., and Pabst, H. (2016). LIBXSMM: Accelerating Small Matrix Multiplications by Runtime Code Generation. SC \u201916: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, IEEE.","DOI":"10.1109\/SC.2016.83"},{"key":"ref_39","unstructured":"(2026, March 20). Kunpeng Math Library (KML) Developer Guide. Available online: https:\/\/support.huawei.com\/enterprise\/zh\/doc\/EDOC1100283144\/8dea3eb?utm_source=chatgpt.com."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Zhang, W., Jiang, Z., Chen, Z., Xiao, N., and Ou, Y. (2021). NUMA-Aware DGEMM Based on 64-Bit ARMv8 Multicore Processors Architecture. Electronics, 10.","DOI":"10.3390\/electronics10161984"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Yu, X., Ma, H., Qu, Z., Fang, J., and Liu, W. (2020). NUMA-Aware Optimization of Sparse Matrix\u2013Vector Multiplication on ARMv8-Based Many-Core Architectures. Network and Parallel Computing: 17th IFIP WG 10.3 International Conference, ACM.","DOI":"10.1007\/978-3-030-79478-1_20"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"103035","DOI":"10.1016\/j.parco.2023.103035","article-title":"Optimizing Massively Parallel Sparse Matrix Computing on ARM Many-Core Processor","volume":"117","author":"Zheng","year":"2023","journal-title":"Parallel Comput."}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/4\/223\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T01:25:20Z","timestamp":1775438720000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/4\/223"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,3]]},"references-count":42,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2026,4]]}},"alternative-id":["computers15040223"],"URL":"https:\/\/doi.org\/10.3390\/computers15040223","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,3]]}}}