{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T23:02:38Z","timestamp":1777676558725,"version":"3.51.4"},"reference-count":39,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[2015,4,8]],"date-time":"2015-04-08T00:00:00Z","timestamp":1428451200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2015,8]]},"abstract":"<jats:p>Krylov subspace iterative solvers are often the method of choice when solving large sparse linear systems. At the same time, hardware accelerators such as graphics processing units continue to offer significant floating point performance gains for matrix and vector computations through easy-to-use libraries of computational kernels. However, as these libraries are usually composed of a well optimized but limited set of linear algebra operations, applications that use them often fail to reduce certain data communications, and hence fail to leverage the full potential of the accelerator. In this paper, we target the acceleration of Krylov subspace iterative methods for graphics processing units, and in particular the Biconjugate Gradient Stabilized solver that significant improvement can be achieved by reformulating the method to reduce data-communications through application-specific kernels instead of using the generic BLAS kernels, e.g. as provided by NVIDIA\u2019s cuBLAS library, and by designing a graphics processing unit specific sparse matrix-vector product kernel that is able to more efficiently use the graphics processing unit\u2019s computing power. Furthermore, we derive a model estimating the performance improvement, and use experimental data to validate the expected runtime savings. Considering that the derived implementation achieves significantly higher performance, we assert that similar optimizations addressing algorithm structure, as well as sparse matrix-vector, are crucial for the subsequent development of high-performance graphics processing units accelerated Krylov subspace iterative methods.<\/jats:p>","DOI":"10.1177\/1094342015580139","type":"journal-article","created":{"date-parts":[[2015,4,9]],"date-time":"2015-04-09T20:00:39Z","timestamp":1428609639000},"page":"366-383","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":18,"title":["Acceleration of GPU-based Krylov solvers via data transfer reduction"],"prefix":"10.1177","volume":"29","author":[{"given":"Hartwig","family":"Anzt","sequence":"first","affiliation":[{"name":"Innovative Computing Laboratory, University of Tennessee, Tennessee, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stanimire","family":"Tomov","sequence":"additional","affiliation":[{"name":"Innovative Computing Laboratory, University of Tennessee, Tennessee, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Piotr","family":"Luszczek","sequence":"additional","affiliation":[{"name":"Innovative Computing Laboratory, University of Tennessee, Tennessee, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"William","family":"Sawyer","sequence":"additional","affiliation":[{"name":"Swiss National Supercomputing Centre (CSCS), Lugano, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jack","family":"Dongarra","sequence":"additional","affiliation":[{"name":"Innovative Computing Laboratory, University of Tennessee, Tennessee, USA"},{"name":"Oak Ridge National Laboratory, Tennessee, USA"},{"name":"University of Manchester, Manchester, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2015,4,8]]},"reference":[{"key":"bibr1-1094342015580139","unstructured":"(?) The top 500 list,\n                      http:\/\/www.top.org\/\n                      ."},{"key":"bibr2-1094342015580139","volume-title":"NVIDIA CUDA Compute unified device architecture programming guide","author":"NVIDIA Corporation","year":"2009"},{"key":"bibr3-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2013.41"},{"key":"#cr-split#-bibr4-1094342015580139.1","unstructured":"Anzt H, Heuveline V, Rocker B (2010) Mixed precision error correction methods for linear systems: Convergence analysis based on Krylov subspace methods. In: Jonasson K"},{"key":"#cr-split#-bibr4-1094342015580139.2","unstructured":"(ed) PARA 2010, Part II, LNCS 7134. Heidelberg: Springer, pp. 237-248."},{"key":"bibr5-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2014.107"},{"key":"bibr6-1094342015580139","unstructured":"Anzt H, Tomov S, Dongarra J (2014b) Implementing a sparse matrix vector product for the SELL-C\/SELL-C- \u03c3 formats on NVIDIA GPUs. Technical Report, University of Tennessee, USA, April."},{"key":"bibr7-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1145\/2712386.2712387"},{"key":"bibr8-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1137\/0911033"},{"key":"bibr9-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2008.11.005"},{"key":"bibr10-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611971538"},{"key":"bibr11-1094342015580139","unstructured":"Bell N, Garland M (2008) Efficient sparse matrix-vector multiplication on CUDA. NVIDIA Technical Report NVR-2008-004\u2019\u2019, NVIDIA Corporation, Santa Clara, CA, December."},{"key":"bibr12-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511618635"},{"key":"bibr13-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.73"},{"key":"bibr14-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1177\/1094342007084026"},{"key":"bibr15-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1145\/1693453.1693471"},{"key":"bibr16-1094342015580139","unstructured":"Corporation N (2012) NVIDIA\u2019s next generation CUDA compute architecture: Kepler GK110. Whitepaper."},{"key":"bibr17-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-14325-5_2"},{"key":"bibr18-1094342015580139","unstructured":"Filipovic J, Madzin M, Fousek J, Matyska L (2013) Optimizing CUDA code by kernel fusion\u2014application on BLAS. Available at: http:\/\/arxiv.org\/abs\/1305.1183 (accessed June 2014)."},{"key":"bibr19-1094342015580139","doi-asserted-by":"publisher","DOI":"10.6028\/jres.049.044"},{"key":"bibr20-1094342015580139","unstructured":"Hoemmen MF (2010) Communication-avoiding Krylov subspace methods. PhD Thesis, EECS Department, University of California, Berkeley."},{"key":"bibr21-1094342015580139","unstructured":"Kogge P, Bergman K, Borkar S, (2008) ExaScale computing study: Technology challenges in achieving ExaScale systems. Available at: http:\/\/www.cse.nd.edu\/Reports\/2008\/TR-2008-13.pdf"},{"key":"bibr22-1094342015580139","author":"Kreutzer M","year":"2014","journal-title":"CoRR"},{"key":"bibr23-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-012-0825-3"},{"key":"bibr24-1094342015580139","first-page":"1","volume-title":"HPC\u2019 12: Proceedings of the 2012 Symposium on High Performance Computing","author":"Lukash M","year":"2012"},{"key":"bibr25-1094342015580139","unstructured":"MAGMA (2015b) PARALUTION. Available at: http:\/\/www.paralution.com\/ (accessed November 2014)."},{"key":"bibr26-1094342015580139","unstructured":"MAGMA (2014) ViennaCL. Available at: http:\/\/viennacl.sourceforge.net\/ (accessed November 2014)."},{"key":"bibr27-1094342015580139","unstructured":"MAGMA (2015a) MAGMA 1.6.1. Available at: http:\/\/icl.cs.utk.edu\/magma\/ (accessed November 2014)."},{"key":"bibr28-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1109\/ICPPW.2014.30"},{"key":"bibr29-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-11515-8_10"},{"key":"bibr30-1094342015580139","volume-title":"NVIDIA CUDA Compute unified device architecture programming guide","author":"NVIDIA Corporation","year":"2009"},{"key":"bibr31-1094342015580139","unstructured":"NVIDIA version 7.0 (2015) Cuda c best practices guide. Available at: http:\/\/docs.nvidia.com\/cuda\/cuda-c-best-practices-guide\/ (accessed March 2015)."},{"key":"bibr32-1094342015580139","unstructured":"NVIDIA version 7.0 (2013a) cuSPARSE library. https:\/\/developer.nvidia.com\/cuSPARSE (accessed March 2015)."},{"key":"bibr33-1094342015580139","author":"NVIDIA","year":"2013","journal-title":"NVIDIA CUDA TOOLKIT V6.0"},{"key":"bibr34-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898718003"},{"key":"bibr35-1094342015580139","unstructured":"Sawyer W (2011) CUSPARSE\/CUBLAS example: BiCGStab iterative solver for non-symmetric linear systems. Available at: https:\/\/hpcforge.org\/plugins\/mediawiki\/wiki\/gpu-training\/index.php\/Main_Page (accessed November 2014)."},{"key":"bibr36-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1137\/0913035"},{"key":"bibr37-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1201\/b10376-8"},{"key":"bibr38-1094342015580139","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2014.48"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342015580139","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342015580139","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342015580139","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:19:25Z","timestamp":1777450765000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342015580139"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,4,8]]},"references-count":39,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2015,8]]}},"alternative-id":["10.1177\/1094342015580139"],"URL":"https:\/\/doi.org\/10.1177\/1094342015580139","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2015,4,8]]}}}