{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,30]],"date-time":"2025-10-30T07:08:09Z","timestamp":1761808089720},"reference-count":55,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2017,1,10]],"date-time":"2017-01-10T00:00:00Z","timestamp":1484006400000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Sign Process Syst"],"published-print":{"date-parts":[[2018,1]]},"DOI":"10.1007\/s11265-016-1216-4","type":"journal-article","created":{"date-parts":[[2017,1,10]],"date-time":"2017-01-10T08:29:10Z","timestamp":1484036950000},"page":"69-86","update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["LightSpMV: Faster CUDA-Compatible Sparse Matrix-Vector Multiplication Using Compressed Sparse Rows"],"prefix":"10.1007","volume":"90","author":[{"given":"Yongchao","family":"Liu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bertil","family":"Schmidt","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2017,1,10]]},"reference":[{"key":"1216_CR1","doi-asserted-by":"crossref","unstructured":"Aila, T., & Laine, S. (2009). Understanding the efficiency of ray traversal on gpus. In Proceedings of the conference on high performance graphics 2009 (pp. 145\u2013149): ACM.","DOI":"10.1145\/1572769.1572792"},{"key":"1216_CR2","unstructured":"Aluru, M., Zola, J., Nettleton, D., & Aluru, S. (2012). Reverse engineering and analysis of large genome-scale gene networks. Nucleic acids research (p. gks904)."},{"key":"1216_CR3","unstructured":"Asanovic, K., Bodik, R., Catanzaro, B. C., Gebis, J. J., Husbands, P., Keutzer, K., Patterson, D. A., Plishker, W. L., Shalf, J., Williams, S. W., & et al. (2006). The landscape of parallel computing research: A view from berkeley. Tech. rep., Technical Report UCB\/EECS-2006-183, EECS Department, University of California, Berkeley."},{"key":"1216_CR4","doi-asserted-by":"crossref","unstructured":"Ashari, A., Sedaghati, N., Eisenlohr, J., Parthasarath, S., & Sadayappan, P. (2014). Fast sparse matrix-vector multiplication on gpus for graph applications. In Proceedings of the international conference for high performance computing, networking, storage and analysis (pp. 781\u2013792): IEEE.","DOI":"10.1109\/SC.2014.69"},{"key":"1216_CR5","doi-asserted-by":"crossref","unstructured":"Ashari, A., Sedaghati, N., Eisenlohr, J., & Sadayappan, P. (2014). An efficient two-dimensional blocking strategy for sparse matrix-vector multiplication on gpus. In Proceedings of the 28th ACM international conference on supercomputing (pp. 273\u2013282): ACM.","DOI":"10.1145\/2597652.2597678"},{"key":"1216_CR6","unstructured":"Barrachina, S., Castillo, M., Igual, F. D., Mayo, R., & Quintana-Ort\u00ed, E. S. (2008). Solving dense linear systems on graphics processors. In Lecture notes in computer science, (Vol. 5168 pp. 739\u2013748): Springer."},{"key":"1216_CR7","unstructured":"Baskaran, M. M., & Bordawekar, R. (2008). Optimizing sparse matrix-vector multiplication on gpus using compile-time and run-time strategies. IBM Reserach Report RC24704."},{"key":"1216_CR8","doi-asserted-by":"crossref","unstructured":"Bell, N., & Garland, M. (2009). Implementing sparse matrix-vector multiplication on throughput-oriented processors. In Proceedings of the conference on high performance computing networking, storage and analysis (p. 18): ACM.","DOI":"10.1145\/1654059.1654078"},{"key":"1216_CR9","unstructured":"Bell, N., & Garland, M. (2014). Cusp: Generic parallel algorithms for sparse matrix and graph computations (v0.4). \n                        http:\/\/cusplibrary.github.io\n                        \n                    ."},{"key":"1216_CR10","unstructured":"Brin, S., & Page, L. (2010). The anatomy of a large-scale hypertextual web search engine."},{"issue":"3","key":"1216_CR11","doi-asserted-by":"crossref","first-page":"679","DOI":"10.1109\/TCBB.2011.68","volume":"9","author":"A Bustamam","year":"2012","unstructured":"Bustamam, A., Burrage, K., & Hamilton, N. A. (2012). Fast parallel markov clustering in bioinformatics using massively parallel computing on gpu with cuda and ellpack-r sparse format. IEEE\/ACM Transactions on Computational Biology and Bioinformatics, 9(3), 679\u2013692.","journal-title":"IEEE\/ACM Transactions on Computational Biology and Bioinformatics"},{"key":"1216_CR12","unstructured":"Butte, A. J., & Kohane, I. S. (1999). Unsupervised knowledge discovery in medical databases using relevance networks. In Proceedings of the AMIA Symposium (p. 711): American Medical Informatics Association."},{"key":"1216_CR13","doi-asserted-by":"crossref","unstructured":"Choi, J. W., Singh, A., & Vuduc, R. W. (2010). Model-driven autotuning of sparse matrix-vector multiply on gpus. In ACM sigplan notices, (Vol. 45 pp. 115\u2013126): ACM.","DOI":"10.1145\/1693453.1693471"},{"key":"1216_CR14","unstructured":"Daga, M., & Greathouse, J. L. (2015). Structural agnostic spmv: Adapting csr-adaptive for irregular matrices. In 2015 IEEE 22nd International conference on high performance computing (HiPC) (pp. 64\u201374): IEEE."},{"issue":"11","key":"1216_CR15","doi-asserted-by":"crossref","first-page":"737","DOI":"10.1016\/j.parco.2013.09.005","volume":"39","author":"HV Dang","year":"2013","unstructured":"Dang, H. V., & Schmidt, B. (2013). Cuda-enabled sparse matrix\u2013vector multiplication on gpus using atomic operations. Parallel Computing, 39(11), 737\u2013750.","journal-title":"Parallel Computing"},{"issue":"1","key":"1216_CR16","first-page":"1","volume":"38","author":"TA Davis","year":"2011","unstructured":"Davis, T. A., & Hu, Y. (2011). The university of florida sparse matrix collection. ACM Transactions on Mathematical Software, 38(1), 1.","journal-title":"ACM Transactions on Mathematical Software"},{"issue":"8","key":"1216_CR17","doi-asserted-by":"crossref","first-page":"2982","DOI":"10.1109\/TMAG.2010.2043511","volume":"46","author":"MM Dehnavi","year":"2010","unstructured":"Dehnavi, M. M., Fern\u00e1ndez, D. M., & Giannacopoulos, D. (2010). Finite-element sparse matrix vector multiplication on graphic processing units. IEEE Transactions on Magnetics, 46(8), 2982\u20132985.","journal-title":"IEEE Transactions on Magnetics"},{"key":"1216_CR18","doi-asserted-by":"crossref","unstructured":"Gilbert, J. R., Reinhardt, S., & Shah, V. B. (2007). High-performance graph algorithms from parallel sparse matrices. In Applied Parallel Computing. State of the Art in Scientific Computing (pp. 260\u2013269): Springer.","DOI":"10.1007\/978-3-540-75755-9_32"},{"issue":"1","key":"1216_CR19","doi-asserted-by":"crossref","first-page":"36","DOI":"10.1007\/s11227-008-0251-8","volume":"50","author":"G Goumas","year":"2009","unstructured":"Goumas, G., Kourtis, K., Anastopoulos, N., Karakasis, V., & Koziris, N. (2009). Performance evaluation of the sparse matrix-vector multiplication on modern architectures. The Journal of Supercomputing, 50(1), 36\u201377.","journal-title":"The Journal of Supercomputing"},{"key":"1216_CR20","doi-asserted-by":"crossref","unstructured":"Greathouse, J. L., & Daga, M. (2014). Efficient sparse matrix-vector multiplication on gpus using the csr storage format. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (pp. 769\u2013780): IEEE.","DOI":"10.1109\/SC.2014.68"},{"key":"1216_CR21","unstructured":"Im, E. J., & Yelick, K. (2000). Optimization of sparse matrix kernels for data mining. In First SIAM Conference on Data Mining. Citeseer."},{"issue":"5","key":"1216_CR22","doi-asserted-by":"crossref","first-page":"604","DOI":"10.1145\/324133.324140","volume":"46","author":"JM Kleinberg","year":"1999","unstructured":"Kleinberg, J. M. (1999). Authoritative sources in a hyperlinked environment. Journal of the ACM, 46(5), 604\u2013632.","journal-title":"Journal of the ACM"},{"issue":"2","key":"1216_CR23","doi-asserted-by":"crossref","first-page":"443","DOI":"10.1007\/s11227-012-0825-3","volume":"63","author":"R Li","year":"2013","unstructured":"Li, R., & Saad, Y. (2013). Gpu-accelerated preconditioned iterative linear solvers. The Journal of Supercomputing, 63(2), 443\u2013466.","journal-title":"The Journal of Supercomputing"},{"key":"1216_CR24","doi-asserted-by":"crossref","unstructured":"Liu, W., & Vinter, B. (2015). Csr5: An efficient storage format for cross-platform sparse matrix-vector multiplication. In Proceedings of the 29th ACM on International Conference on Supercomputing (pp. 339\u2013350).","DOI":"10.1145\/2751205.2751209"},{"key":"1216_CR25","doi-asserted-by":"crossref","first-page":"179","DOI":"10.1016\/j.parco.2015.04.004","volume":"49","author":"W Liu","year":"2015","unstructured":"Liu, W., & Vinter, B. (2015). Speculative segmented sum for sparse matrix-vector multiplication on heterogeneous processors. Parallel Computing, 49, 179\u2013193.","journal-title":"Parallel Computing"},{"key":"1216_CR26","doi-asserted-by":"crossref","unstructured":"Liu, X., Smelyanskiy, M., Chow, E., & Dubey, P. (2013). Efficient sparse matrix-vector multiplication on x86-based many-core processors. In Proceedings of the 27th international ACM conference on International conference on supercomputing (pp. 273\u2013282): ACM.","DOI":"10.1145\/2464996.2465013"},{"key":"1216_CR27","doi-asserted-by":"crossref","unstructured":"Liu, Y., & Schmidt, B. (2014). Swaphi: Smith-waterman protein database search on xeon phi coprocessors. In 25th IEEE International Conference on Application-specific Systems, Architectures and Processors (pp. 184\u2013185): IEEE.","DOI":"10.1109\/ASAP.2014.6868657"},{"key":"1216_CR28","doi-asserted-by":"crossref","unstructured":"Liu, Y., & Schmidt, B. (2015). Lightspmv: Faster csr-based sparse matrix-vector multiplication on cuda-enabled gpus. In 26th IEEE International Conference on Application-specific Systems (pp. 82\u201389).","DOI":"10.1109\/ASAP.2015.7245713"},{"key":"1216_CR29","doi-asserted-by":"crossref","unstructured":"Liu, Y., Tran, T. T., Lauenroth, F., & Schmidt, B. (2014). Swaphi-ls: Smith-waterman algorithm on xeon phi coprocessors for long dna sequences. In 2014 IEEE International Conference on Cluster Computing (pp. 257\u2013265): IEEE.","DOI":"10.1109\/CLUSTER.2014.6968772"},{"key":"1216_CR30","doi-asserted-by":"crossref","unstructured":"Merrill, D., & Garland, M. (2016). Merge-based sparse matrix-vector multiplication (spmv) using the csr storage format. In Proceedings of the 21st ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (p. 43): ACM.","DOI":"10.1145\/2851141.2851190"},{"key":"1216_CR31","doi-asserted-by":"crossref","unstructured":"Merrill, D., Garland, M., & Grimshaw, A. (2012). Scalable gpu graph traversal. In ACM SIGPLAN Notices, (Vol. 47 pp. 117\u2013128): ACM.","DOI":"10.1145\/2145816.2145832"},{"key":"1216_CR32","unstructured":"Misra, S., Pamnany, K., & Aluru, S. (2014). Parallel mutual information based construction of whole-genome networks on the intel (r) xeon phi (tm) coprocessor. In 28th IEEE International on Parallel and Distributed Processing Symposium (pp. 241\u2013250): IEEE."},{"key":"1216_CR33","doi-asserted-by":"crossref","unstructured":"Monakov, A., Lokhmotov, A., & Avetisyan, A. (2010). Automatically tuning sparse matrix-vector multiplication for gpu architectures. In High Performance Embedded Architectures and Compilers (pp. 111\u2013125): Springer.","DOI":"10.1007\/978-3-642-11515-8_10"},{"key":"1216_CR34","doi-asserted-by":"crossref","first-page":"131","DOI":"10.1016\/j.procs.2016.05.304","volume":"80","author":"Y Nagasaka","year":"2016","unstructured":"Nagasaka, Y., Nukada, A., & Matsuoka, S. (2016). Adaptive multi-level blocking optimization for sparse matrix vector multiplication on gpu. Procedia Computer Science, 80, 131\u2013142.","journal-title":"Procedia Computer Science"},{"key":"1216_CR35","unstructured":"Nvidia (2013). Nvidia\u2019s next generation cuda compute architecture: Kepler gk110. NVIDIA White Paper."},{"key":"1216_CR36","unstructured":"Nvidia (2015). Maxwell: The most advanced cuda gpu ever made. \n                        http:\/\/devblogs.nvidia.com\/parallelforall\/maxwell-most-advanced-cuda-gpu-ever-made\n                        \n                     \n                        http:\/\/devblogs.nvidia.com\/parallelforall\/maxwell-most-advanced-cuda-gpu-ever-made\n                        \n                    ."},{"key":"1216_CR37","unstructured":"NVIDIA (2015). The nvidia cuda sparse matrix library (cusparse). In CUDA 6.5 toolkit."},{"key":"1216_CR38","unstructured":"NVIDIA (2015). Nvidia visual profiler in cuda 7 tookit. \n                        https:\/\/developer.nvidia.com\/nvidia-visual-profiler\n                        \n                    ."},{"key":"1216_CR39","unstructured":"Nvidia (2016). Nvidia gp100 pascal architecture-infinite compute for infinite opportunities. \n                        http:\/\/www.nvidia.com\/object\/pascal-architecture-whitepaper.html\n                        \n                    ."},{"key":"1216_CR40","unstructured":"Reguly, I., & Giles, M. (2012). Efficient sparse matrix-vector multiplication on cache-based gpus. In Innovative Parallel Computing, 2012 (pp. 1\u201312): IEEE."},{"key":"1216_CR41","unstructured":"Rupp, K., Rudolf, F., & Weinbub, J. (2010). Viennacl-a high level linear algebra library for gpus and multi-core cpus. Proceedings of the International Workshop on GPUs and Scientific Applications, 51\u201356."},{"key":"1216_CR42","doi-asserted-by":"crossref","unstructured":"Saad, Y. (2003). Iterative methods for sparse linear systems, Siam.","DOI":"10.1137\/1.9780898718003"},{"key":"1216_CR43","unstructured":"Saule, E., Kaya, K., & \u00c7ataly\u00fcrek, \u00dc. V. (2014). Performance evaluation of sparse matrix multiplication kernels on intel xeon phi, (pp. 559\u2013570): Springer."},{"key":"1216_CR44","doi-asserted-by":"crossref","unstructured":"Su, B. Y., & Keutzer, K. (2012). clspmv: A cross-platform opencl spmv framework on gpus. In Proceedings of the 26th ACM international conference on Supercomputing (pp. 353\u2013364): ACM.","DOI":"10.1145\/2304576.2304624"},{"issue":"9","key":"1216_CR45","doi-asserted-by":"crossref","first-page":"2373","DOI":"10.1109\/TPDS.2014.2357437","volume":"26","author":"W Tang","year":"2015","unstructured":"Tang, W., Tan, W., Goh, R. S. M., Turner, S., & Wong, W. K. (2015). A family of bit-representation-optimized formats for fast sparse matrix-vector multiplication on the gpu. IEEE Transactions on Parallel and Distributed Systems, 26(9), 2373\u20132385.","journal-title":"IEEE Transactions on Parallel and Distributed Systems"},{"issue":"3","key":"1216_CR46","doi-asserted-by":"crossref","first-page":"327","DOI":"10.1007\/s10115-007-0094-2","volume":"14","author":"H Tong","year":"2008","unstructured":"Tong, H., Faloutsos, C., & Pan, J. Y. (2008). Random walk with restart: fast solutions and applications. Knowledge and Information Systems, 14(3), 327\u2013346.","journal-title":"Knowledge and Information Systems"},{"key":"1216_CR47","unstructured":"Tzeng, S., Patney, A., & Owens, J. D. (2010). Task management for irregular-parallel workloads on the gpu. In Proceedings of the Conference on High Performance Graphics (pp. 29\u201337): Eurographics Association."},{"key":"1216_CR48","doi-asserted-by":"crossref","unstructured":"Vazquez, F., Ortega, G., Fern\u00e1ndez, J. J., & Garz\u00f3n, E. M. (2010). Improving the performance of the sparse matrix vector product with gpus. In 10th IEEE International Conference on Computer and Information Technology (pp. 1146\u20131151): IEEE.","DOI":"10.1109\/CIT.2010.208"},{"key":"1216_CR49","unstructured":"Volkov, V. (2010). Better performance at lower occupancy. In Proceedings of the GPU technology conference, GTC, (Vol. 10 p. 16). San Jose, CA."},{"key":"1216_CR50","doi-asserted-by":"crossref","unstructured":"Volkov, V., & Demmel, J. W. (2008). Benchmarking gpus to tune dense linear algebra. In Proceedings of the 2008 ACM\/IEEE conference on Supercomputing, 31 (pp. 1\u201311): IEEE.","DOI":"10.1109\/SC.2008.5214359"},{"key":"1216_CR51","unstructured":"Vuduc, R. W. (2003). Automatic performance tuning of sparse matrix kernels. Ph.D. thesis. PhD thesis, University of California, Berkeley."},{"key":"1216_CR52","unstructured":"Wu, B., Zhao, Z., Zhang, E. Z., Jiang, Y., & Shen, X. (2013). Complexity analysis and algorithm design for reorganizing data to minimize non-coalesced memory accesses on gpu (Vol. 48, pp. 57\u201368): ACM."},{"key":"1216_CR53","doi-asserted-by":"crossref","unstructured":"Xiang, P., Yang, Y., & Zhou, H. (2014). Warp-level divergence in gpus: Characterization, impact, and mitigation. In 20th IEEE International Symposium on High Performance Computer Architecture (pp. 284\u2013295): IEEE.","DOI":"10.1109\/HPCA.2014.6835939"},{"key":"1216_CR54","unstructured":"Yan, S., Li, C., Zhang, Y., & Zhou, H. (2014). yaspmv: Yet another spmv framework on gpus (Vol. 49, pp. 107\u2013118): ACM."},{"issue":"4","key":"1216_CR55","doi-asserted-by":"crossref","first-page":"231","DOI":"10.14778\/1938545.1938548","volume":"4","author":"X Yang","year":"2011","unstructured":"Yang, X., Parthasarathy, S., & Sadayappan, P. (2011). Fast sparse matrix-vector multiplication on gpus: implications for graph mining. Proceedings of the VLDB Endowment, 4(4), 231\u2013242.","journal-title":"Proceedings of the VLDB Endowment"}],"container-title":["Journal of Signal Processing Systems"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11265-016-1216-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11265-016-1216-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11265-016-1216-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2018,1,11]],"date-time":"2018-01-11T01:08:39Z","timestamp":1515632919000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11265-016-1216-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,1,10]]},"references-count":55,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2018,1]]}},"alternative-id":["1216"],"URL":"https:\/\/doi.org\/10.1007\/s11265-016-1216-4","relation":{},"ISSN":["1939-8018","1939-8115"],"issn-type":[{"value":"1939-8018","type":"print"},{"value":"1939-8115","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,1,10]]}}}