{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T16:25:09Z","timestamp":1787329509910,"version":"build-2736575974"},"reference-count":18,"publisher":"Society for Industrial & Applied Mathematics (SIAM)","issue":"5","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["SIAM J. Sci. Comput."],"published-print":{"date-parts":[[2011,1]]},"abstract":"<jats:p>This paper presents a parallelized hybrid single-vector Arnoldi algorithm for computing approximations to eigenpairs of a nonsymmetric matrix. We are interested in the use of accelerators and multicore units to speed up the Arnoldi process. The main goal is to propose a parallel version of the Arnoldi solver, which can efficiently use multiple multicore processors or multiple graphics processing units (GPUs) in a mixed coarse and fine grain fashion. In the proposed algorithms, this is achieved by an autotuning of the matrix vector product before starting the Arnoldi eigensolver as well as the reorganization of the data and global communications so that communication time is reduced. The execution time, performance, and scalability are assessed with well-known dense and sparse test matrices on multiple Nehalems, GT200 NVidia Tesla, and next generation Fermi Tesla. With one processor, we see a performance speedup of 2 to 3x when using all the physical cores, and a total speedup of 2 to 8x when adding a GPU to this multicore unit, and hence a speedup of 4 to 24x compared to the sequential solver.<\/jats:p>","DOI":"10.1137\/10079906x","type":"journal-article","created":{"date-parts":[[2011,10,27]],"date-time":"2011-10-27T18:46:13Z","timestamp":1319741173000},"page":"3010-3019","source":"Crossref","is-referenced-by-count":8,"title":["Accelerating the Explicitly Restarted Arnoldi Method with GPUs Using an Autotuned Matrix Vector Product"],"prefix":"10.1137","volume":"33","author":[{"given":"J\u00e9r\u00f4me","family":"Dubois","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christophe","family":"Calvin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Serge","family":"Petiton","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"351","published-online":{"date-parts":[[2011,10,27]]},"reference":[{"key":"R1","doi-asserted-by":"crossref","unstructured":"Z. K. Baker, M. B. Gokhale, and J. L. Tripp,\n                      Matched filter computation on FPGA, cell and GPU\n                      , in Proceedings of the 15th Annual IEEE Symposium on Field-Programmable Custom Computing Machines, 2007.","DOI":"10.1109\/FCCM.2007.52"},{"key":"R2","unstructured":"N. Bell and M. Garland,\n                      Efficient Sparse Matrix-Vector Multiplication on CUDA\n                      , NVIDIA Technical Report NVR-2008-004, NVIDIA Corporation, Santa Clara, CA, 2008."},{"key":"R3","doi-asserted-by":"crossref","unstructured":"N. Bell and M. Garland,\n                      Implementing sparse matrix-vector multiplication on throughput-oriented processors\n                      , in SC '09: Proceedings of the 2009 ACM\/IEEE Conference on Supercomputing, Portland, OR, 2009.","DOI":"10.1145\/1654059.1654078"},{"key":"R4","doi-asserted-by":"publisher","DOI":"10.1145\/567806.567807"},{"key":"R5","unstructured":"T. Braconnier,\n                      Influence of Orthogonality on the Backward Error and the Stopping Criterion for Krylov Methods\n                      , Numerical Analysis Report 281, Manchester Centre for Computational Mathematics, Manchester, England, 1995."},{"key":"R6","doi-asserted-by":"publisher","DOI":"10.1147\/rd.515.0559"},{"key":"R7","unstructured":"T. A. Davis,\n                      University of Florida sparse matrix collection\n                      , NA Digest, 92 (1994), 96 (1996), 97 (1997); also available online from http:\/\/www.cise.ufl.edu\/research\/sparse\/."},{"key":"R8","doi-asserted-by":"publisher","DOI":"10.1145\/62038.62043"},{"key":"R9","doi-asserted-by":"publisher","DOI":"10.1137\/S1064827500366082"},{"key":"R10","unstructured":"D. Fay, A. Sazegari, and D. A. Connors,\n                      A detailed study of the numerical accuracy of gpu-implemented math functions\n                      , in Supercomputing '06 Workshop on General-Purpose GPU Computing: Practice and Experience, Tampa, FL, 2006."},{"key":"R11","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479899358595"},{"key":"R12","unstructured":"J. D. McCalpin,\n                      Memory bandwidth and machine balance in current high performance computers\n                      , IEEE Computer Society Technical Committee on Computer Architecture (TCCA) Newsletter (December, 1995), pp. 19\u201325."},{"key":"R13","unstructured":"H. Meuer, E. Strohmaier, J. Dongarra, and H. Simon,\n                      Architecture Share over Time\n                      , http:\/\/www.top500.org\/overtime\/list\/34\/archtype."},{"key":"R14","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2010.41"},{"key":"R15","unstructured":"NVidia,\n                      Cublas Library\n                      , Technical Report, NVIDIA, Santa Clara, CA, 2008."},{"key":"R16","unstructured":"NVidia,\n                      Tuning Cuda Applications for Fermi\n                      , Technical Report, NVIDIA, Santa Clara, CA, 2010."},{"key":"R17","unstructured":"Y. Saad,\n                      Numerical Methods for Large Eigenvalue Problems\n                      , Halstead Press, New York, 1992."},{"key":"R18","first-page":"521","volume":"37","author":"Wang Y.","year":"2011","journal-title":"Parallel Computing","ISSN":"https:\/\/id.crossref.org\/issn\/0167-8191","issn-type":"print"}],"container-title":["SIAM Journal on Scientific Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/epubs.siam.org\/doi\/pdf\/10.1137\/10079906X","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T15:19:19Z","timestamp":1787325559000},"score":1,"resource":{"primary":{"URL":"https:\/\/epubs.siam.org\/doi\/10.1137\/10079906X"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,1]]},"references-count":18,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2011,1]]}},"alternative-id":["10.1137\/10079906X"],"URL":"https:\/\/doi.org\/10.1137\/10079906x","relation":{},"ISSN":["1064-8275","1095-7197"],"issn-type":[{"value":"1064-8275","type":"print"},{"value":"1095-7197","type":"electronic"}],"subject":[],"published":{"date-parts":[[2011,1]]}}}