{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T08:43:33Z","timestamp":1780994613190,"version":"3.54.1"},"reference-count":29,"publisher":"SAGE Publications","issue":"2","license":[{"start":{"date-parts":[[2013,9,2]],"date-time":"2013-09-02T00:00:00Z","timestamp":1378080000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2014,5]]},"abstract":"<jats:p>Sparse matrix\u2013vector multiplication (SpMV) is of singular importance in sparse linear algebra, which is an important issue in scientific computing and engineering practice. Much effort has been put into accelerating SpMV, and a few parallel solutions have been proposed. This paper focuses on a special type of SpMV, namely sparse quasi-diagonal matrix\u2013vector multiplication (SQDMV). The sparse quasi-diagonal matrix is the key to solving many differential equations, and very little research has been done in this field. This paper discusses data structures and algorithms for SQDMV that are efficiently implemented on the compute unified device architecture (CUDA) platform for the fine-grained parallel architecture of the graphics processing unit (GPU). A new diagonal storage format, a hybrid of the diagonal format (DLA) and the compressed sparse row format (CSR) (HDC) will be presented, which overcomes the inefficiency of DLA in storing irregular matrices and the imbalances of CSR in storing non-zero elements. Furthermore, HDC can adjust the storage bandwidth of the diagonal to adapt to different discrete degrees of sparse matrix, so as to get a higher compression ratio than DLA and CSR, and reduce the computational complexity. Our implementation in a GPU shows that the performance of HDC is better than that of other formats, especially for matrices with some discrete points outside the main diagonal. In addition, we combine the different parts of HDC to make a unified kernel to get a better compression ratio and a higher speedup ratio in the GPU.<\/jats:p>","DOI":"10.1177\/1094342013501126","type":"journal-article","created":{"date-parts":[[2013,9,2]],"date-time":"2013-09-02T23:20:31Z","timestamp":1378164031000},"page":"183-195","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":33,"title":["Optimization of quasi-diagonal matrix\u2013vector multiplication on GPU"],"prefix":"10.1177","volume":"28","author":[{"given":"Wangdong","family":"Yang","sequence":"first","affiliation":[{"name":"School of Information Science and Engineering, Hunan City University, China"},{"name":"College of Information Science and Engineering, Hunan University, China"},{"name":"National Supercomputing Centre in Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kenli","family":"Li","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Hunan City University, China"},{"name":"College of Information Science and Engineering, Hunan University, China"},{"name":"National Supercomputing Centre in Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yan","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Hunan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lin","family":"Shi","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Hunan University, China"},{"name":"National Supercomputing Centre in Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lanjun","family":"Wan","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Hunan University, China"},{"name":"National Supercomputing Centre in Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2013,9,2]]},"reference":[{"key":"bibr1-1094342013501126","volume-title":"Optimizing sparse matrix-vector multiplication on GPUs using compile-time and run-time strategies","author":"Baskaran MM","year":"2008"},{"key":"bibr2-1094342013501126","unstructured":"Bell N, Garland M (2008) Efficient sparse matrix\u2013vector multiplication on CUDA. NVIDIA technical report NVR-2008-004."},{"key":"bibr3-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1145\/1837210.1837224"},{"key":"bibr4-1094342013501126","first-page":"358","volume-title":"High performance computing and communications \u2013 third international conference (HPCC\u201907)","volume":"2007","author":"Buatois L","year":"2010"},{"key":"bibr5-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-01970-8_90"},{"key":"bibr6-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1145\/1693453.1693471"},{"key":"bibr7-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1109\/TMAG.2010.2043511"},{"key":"bibr8-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1109\/ICPADS.2011.91"},{"key":"bibr9-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2008.4536350"},{"key":"bibr10-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1177\/1094342010374847"},{"key":"bibr11-1094342013501126","unstructured":"Harris M (2007) Optimizing parallel reduction in CUDA. NVIDIA Developer Technology. Available at: http:\/\/developer.download.nvidia.com\/assets\/cuda\/files\/reduction.pdf (accessed 11 September 2012)."},{"key":"bibr12-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1109\/HPCC.2010.85"},{"key":"bibr13-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2009.21"},{"key":"bibr14-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1016\/j.compeleceng.2010.07.002"},{"key":"bibr15-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-11515-8_10"},{"key":"bibr16-1094342013501126","unstructured":"NVIDIA (2013) A library for sparse linear algebra and graph computations on CUDA (CUSP), 3rd ed. Available at: https:\/\/github.com\/cusplibrary\/cusplibrary (accessed 11 September 2012)."},{"key":"bibr16a-1094342013501126","unstructured":"NVIDIA (2012a) CUDA Toolkit 4.2 CUBLAS Library, 4th ed. Available at: http:\/\/docs.nvidia.com\/cuda\/cublas\/index.html (accessed 11 September 2012)."},{"key":"bibr17-1094342013501126","unstructured":"NVIDIA (2012b) The NVIDIA CUDA sparse matrix library (cuSPARSE), 2nd ed. Available at: http:\/\/docs.nvidia.com\/cuda\/cusparse\/index.html (accessed 11 September 2012)."},{"key":"bibr18-1094342013501126","first-page":"305","volume-title":"Proceedings of 7th international meeting on high performance computing for computational science (VECPAR\u201906)","volume":"2006","author":"Ohshima S","year":"2006"},{"key":"bibr19-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1016\/j.micpro.2011.05.005"},{"key":"bibr20-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1177\/1094342011431710"},{"key":"bibr21-1094342013501126","unstructured":"Saad Y. (2005) Sparskit: a basic tool-kit for sparse matrix computations, version 2. Available at: http:\/\/www-users.cs.umn.edu\/saad\/software\/SPARSKIT\/sparskit.html (accessed 11 September 2012)."},{"key":"bibr22-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1177\/1094342011428144"},{"key":"bibr23-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1109\/SASP.2010.5521144"},{"key":"bibr24-1094342013501126","unstructured":"University of Florida (2011) UF sparse matrix collection. Available at: http:\/\/www.cise.ufl.edu\/research\/sparse\/matrices\/groups.html (accessed 11 September 2012)."},{"key":"bibr25-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1109\/CIT.2010.208"},{"key":"bibr26-1094342013501126","first-page":"13","volume-title":"Proceedings of the 2nd USENIX conference on hot topics in parallelism (HotPar\u201910)","volume":"2010","author":"Vuduc R","year":"2010"},{"key":"bibr27-1094342013501126","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2010.117"},{"key":"bibr28-1094342013501126","doi-asserted-by":"publisher","DOI":"10.14778\/1938545.1938548"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342013501126","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342013501126","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342013501126","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:19:17Z","timestamp":1777450757000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342013501126"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,9,2]]},"references-count":29,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2014,5]]}},"alternative-id":["10.1177\/1094342013501126"],"URL":"https:\/\/doi.org\/10.1177\/1094342013501126","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2013,9,2]]}}}