{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,9]],"date-time":"2026-08-09T02:19:06Z","timestamp":1786241946845,"version":"3.56.0"},"reference-count":34,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2019,11,14]],"date-time":"2019-11-14T00:00:00Z","timestamp":1573689600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National key R&D Program of China","doi-asserted-by":"publisher","award":["2017YFB0202500"],"award-info":[{"award-number":["2017YFB0202500"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2020,1]]},"abstract":"<jats:p>Sparse matrix\u2013vector multiplication (SpMV) kernel dominates the computing cost in numerous applications. Most of the existing studies dedicated to improving this kernel have been targeting just one type of processing units, mainly multicore CPUs or graphics processing units (GPUs), and have not explored the potential of the recent, rapidly emerging, CPU-GPU heterogeneous platforms. To take full advantage of these heterogeneous systems, the input sparse matrix has to be partitioned on different available processing units. The partitioning problem is more challenging with the existence of many sparse formats whose performances depend both on the sparsity of the input matrix and the used hardware. Thus, the best performance does not only depend on how to partition the input sparse matrix but also on which sparse format to use for each partition. To address this challenge, we propose in this article a new CPU-GPU heterogeneous method for computing the SpMV kernel that combines between different sparse formats to achieve better performance and better utilization of CPU-GPU heterogeneous platforms. The proposed solution horizontally partitions the input matrix into multiple block-rows and predicts their best sparse formats using machine learning-based performance models. A mapping algorithm is then used to assign the block-rows to the CPU and GPU(s) available in the system. Our experimental results using real-world large unstructured sparse matrices on two different machines show a noticeable performance improvement.<\/jats:p>","DOI":"10.1177\/1094342019886628","type":"journal-article","created":{"date-parts":[[2019,11,14]],"date-time":"2019-11-14T23:29:00Z","timestamp":1573774140000},"page":"66-80","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":20,"title":["Sparse matrix partitioning for optimizing SpMV on CPU-GPU heterogeneous platforms"],"prefix":"10.1177","volume":"34","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1779-2705","authenticated-orcid":false,"given":"Akrem","family":"Benatia","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3250-0435","authenticated-orcid":false,"given":"Weixing","family":"Ji","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yizhuo","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Feng","family":"Shi","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2019,11,14]]},"reference":[{"key":"bibr1-1094342019886628","volume-title":"Efficient Sparse Matrix-Vector Multiplication on CUDA: Nvidia Technical Report NVR-2008-004","author":"Bell N","year":"2008"},{"key":"bibr2-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1109\/ICPADS.2016.0120"},{"key":"bibr3-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1006\/jpdc.2000.1714"},{"key":"bibr4-1094342019886628","doi-asserted-by":"crossref","unstructured":"Chang CC, Lin CJ (2011) Libsvm: a library for support vector machines. ACM Transactions on Intelligent Systems and Technology (TIST) 2(3): 27. Available at: https:\/\/www.csie.ntu.edu.tw\/~cjlin\/libsvm\/ (accessed 31 December 2018).","DOI":"10.1145\/1961189.1961199"},{"key":"bibr5-1094342019886628","volume-title":"Using OpenMP: Portable Shared Memory Parallel Programming","volume":"10","author":"Chapman B","year":"2008"},{"key":"bibr6-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1145\/1693453.1693471"},{"key":"bibr7-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2013.09.005"},{"key":"bibr8-1094342019886628","volume-title":"SpMV: a memory-bound application on the GPU stuck between a rock and a hard place: Technical Report14","author":"Davis JD","year":"2012"},{"key":"bibr9-1094342019886628","doi-asserted-by":"crossref","unstructured":"Davis TA, Hu Y (2011) The University of Florida sparse matrix collection. ACM Transactions on Mathematical Software (TOMS) 38(1). Available at: https:\/\/www.cise.ufl.edu\/research\/sparse\/matrices\/ (accessed 31 December 2018).","DOI":"10.1145\/2049662.2049663"},{"key":"bibr10-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1145\/3017994"},{"key":"bibr11-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1145\/2016741.2016744"},{"key":"bibr12-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2013.123"},{"key":"bibr13-1094342019886628","volume-title":"Intel Whitepaper Feb07.pdf","author":"Gustafson JL","year":"2007"},{"key":"bibr14-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1145\/1555754.1555775"},{"key":"bibr15-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1016\/0893-6080(89)90020-8"},{"key":"bibr16-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1145\/322003.322011"},{"key":"bibr17-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1177\/1094342004041296"},{"key":"bibr18-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1109\/HIPC.2009.5433179"},{"key":"bibr19-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1201\/b10376"},{"key":"bibr20-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2015.2401575"},{"issue":"2","key":"bibr21-1094342019886628","doi-asserted-by":"crossref","first-page":"19","DOI":"10.1145\/2636342","volume":"47","author":"Mittal S","year":"2015","journal-title":"ACM Computing Surveys (CSUR)"},{"key":"bibr22-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-11515-8_10"},{"key":"bibr23-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2014.59"},{"key":"bibr24-1094342019886628","doi-asserted-by":"crossref","unstructured":"Nagasaka Y, Nukada A, Matsuoka S (2014) Cache-aware sparse matrix formats for Kepler GPU. In: 20th IEEE international conference on parallel and distributed, Hsinchu, Taiwan, 16\u201319 December 2014, pp. 281\u2013288. IEEE.","DOI":"10.1109\/PADSW.2014.7097819"},{"key":"bibr25-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1145\/2751205.2751244"},{"key":"bibr26-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2015.2509972"},{"key":"bibr27-1094342019886628","first-page":"155","volume":"9","author":"Smola A","year":"1997","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr28-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1023\/B:STCO.0000035301.49549.88"},{"key":"bibr29-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1145\/2304576.2304624"},{"key":"bibr30-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2008.12.006"},{"key":"bibr31-1094342019886628","doi-asserted-by":"crossref","unstructured":"Xu W, Zhang H, Jiao S, et al. (2012) Optimizing sparse matrix vector multiplication using cache blocking method on Fermi GPU. In: 13th ACIS international conference on software engineering, artificial intelligence, networking and parallel and distributed computing (SNPD), Kyoto, Japan, 8\u201310 August 2012, pp. 231\u2013235. IEEE.","DOI":"10.1109\/SNPD.2012.20"},{"key":"bibr32-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1145\/2555243.2555255"},{"key":"bibr33-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2016.12.023"},{"key":"bibr34-1094342019886628","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2014.2366731"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342019886628","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342019886628","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342019886628","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:15:53Z","timestamp":1777450553000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342019886628"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,11,14]]},"references-count":34,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,1]]}},"alternative-id":["10.1177\/1094342019886628"],"URL":"https:\/\/doi.org\/10.1177\/1094342019886628","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,11,14]]}}}