{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,7]],"date-time":"2026-03-07T07:01:15Z","timestamp":1772866875008,"version":"3.50.1"},"reference-count":30,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2016,8,13]],"date-time":"2016-08-13T00:00:00Z","timestamp":1471046400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Math. Softw."],"published-print":{"date-parts":[[2017,3,31]]},"abstract":"<jats:p>This article describes a new high performance implementation of the QR-based Dynamically Weighted Halley Singular Value Decomposition (QDWH-SVD) solver on multicore architecture enhanced with multiple GPUs. The standard QDWH-SVD algorithm was introduced by Nakatsukasa and Higham (SIAM SISC, 2013) and combines three successive computational stages: (1) the polar decomposition calculation of the original matrix using the QDWH algorithm, (2) the symmetric eigendecomposition of the resulting polar factor to obtain the singular values and the right singular vectors, and (3) the matrix-matrix multiplication to get the associated left singular vectors. A comprehensive test suite highlights the numerical robustness of the QDWH-SVD solver. Although it performs up to two times more flops when computing all singular vectors compared to the standard SVD solver algorithm, our new high performance implementation on single GPU results in up to 4\u00d7 improvements for asymptotic matrix sizes, compared to the equivalent routines from existing state-of-the-art open-source and commercial libraries. However, when only singular values are needed, QDWH-SVD is penalized by performing more flops by an order of magnitude. The singular value only implementation of QDWH-SVD on single GPU can still run up to 18% faster than the best existing equivalent routines.<\/jats:p>","DOI":"10.1145\/2894747","type":"journal-article","created":{"date-parts":[[2016,8,15]],"date-time":"2016-08-15T18:17:46Z","timestamp":1471285066000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":17,"title":["A High Performance QDWH-SVD Solver Using Hardware Accelerators"],"prefix":"10.1145","volume":"43","author":[{"given":"Dalal","family":"Sukkari","sequence":"first","affiliation":[{"name":"Extreme Computing Research Center, KAUST, Thuwal Jeddah, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hatem","family":"Ltaief","sequence":"additional","affiliation":[{"name":"Extreme Computing Research Center, KAUST, Thuwal Jeddah, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David","family":"Keyes","sequence":"additional","affiliation":[{"name":"Extreme Computing Research Center, KAUST, Thuwal Jeddah, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,8,13]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/180\/1\/012037"},{"key":"e_1_2_1_2_1","volume-title":"Laura Susan Blackford, James Weldon Demmel, Jack J. Dongarra, Jeremy J. Du Croz, Anne Greenbaum, Sven Hammarling, A. McKenney, and Danny C. Sorensen.","author":"Anderson Edward","year":"1999"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1137\/0613046"},{"key":"e_1_2_1_4_1","volume-title":"Minimizing communication for eigenproblems and the singular value decomposition. CoRR abs\/1011.3077","author":"Ballard Grey","year":"2010"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAES.1975.308025"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/365723.365736"},{"key":"e_1_2_1_7_1","volume-title":"Basic Linear Algebra Subprograms v3.5. (Nov","author":"BLAS.","year":"2013"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1137\/0911052"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342010391989"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/s002110050024"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.2307\/2324422"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF02163027"},{"key":"e_1_2_1_13_1","volume-title":"Van Loan","author":"Golub Gene H.","year":"1996"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479892242232"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2503210.2503292"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2063384.2063394"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342013502097"},{"key":"e_1_2_1_18_1","volume-title":"Rank-Deficient and Discrete Ill-Posed Problems: Numerical Aspects of Linear Inversion","author":"Hansen Per Christian"},{"key":"e_1_2_1_19_1","volume-title":"Higham and Pythagoras Papadimitriou","author":"Nicholas","year":"1993"},{"key":"e_1_2_1_20_1","unstructured":"Intel. 2015. Math Kernel Library. (2015). Available at http:\/\/software.intel.com\/en-us\/articles\/intel-mkl\/.  Intel. 2015. Math Kernel Library. (2015). Available at http:\/\/software.intel.com\/en-us\/articles\/intel-mkl\/."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-8191(99)00021-6"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2009.79"},{"key":"e_1_2_1_23_1","volume-title":"PARCO (Advances in Parallel Computing), Koen De Bosschere, Erik H. D'Hollander, Gerhard R","author":"Ltaief Hatem"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.91"},{"key":"e_1_2_1_25_1","volume-title":"Matrix Algebra on GPU and Multicore Architectures. Innovative Computing Laboratory","author":"MAGMA.","year":"2009"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1137\/090774999"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1137\/120876605"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1137\/0725014"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF02289451"},{"key":"e_1_2_1_30_1","volume-title":"Trefethen and David Bau","author":"Lloyd","year":"1997"}],"container-title":["ACM Transactions on Mathematical Software"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2894747","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2894747","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T19:05:40Z","timestamp":1750273540000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2894747"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,8,13]]},"references-count":30,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2017,3,31]]}},"alternative-id":["10.1145\/2894747"],"URL":"https:\/\/doi.org\/10.1145\/2894747","relation":{},"ISSN":["0098-3500","1557-7295"],"issn-type":[{"value":"0098-3500","type":"print"},{"value":"1557-7295","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,8,13]]},"assertion":[{"value":"2015-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-02-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-08-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}