{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T05:00:06Z","timestamp":1755838806187,"version":"3.41.0"},"reference-count":20,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2019,4,26]],"date-time":"2019-04-26T00:00:00Z","timestamp":1556236800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Math. Softw."],"published-print":{"date-parts":[[2019,6,30]]},"abstract":"<jats:p>\n            This article presents a high-performance software framework for computing a dense SVD on distributed-memory manycore systems. Originally introduced by\u00a0Nakatsukasa et\u00a0al. (2010) and Nakatsukasa and Higham (2013), the SVD solver relies on the polar decomposition using the QR Dynamically Weighted Halley algorithm (QDWH). Although the QDWH-based SVD algorithm performs a significant amount of extra floating-point operations compared to the traditional SVD with the one-stage bidiagonal reduction, the inherent high level of concurrency associated with Level 3 BLAS compute-bound kernels ultimately compensates for the arithmetic complexity overhead. Using the ScaLAPACK two-dimensional block cyclic data distribution with a rectangular processor topology, the resulting QDWH-SVD further reduces excessive communications during the panel factorization, while increasing the degree of parallelism during the update of the trailing submatrix, as opposed to relying on the default square processor grid. After detailing the algorithmic complexity and the memory footprint of the algorithm, we conduct a thorough performance analysis and study the impact of the grid topology on the performance by looking at the communication and computation profiling trade-offs. We report performance results against state-of-the-art existing QDWH software implementations (e.g., Elemental) and their SVD extensions on large-scale distributed-memory manycore systems based on commodity Intel x86 Haswell processors and Knights Landing (KNL) architecture. The QDWH-SVD framework achieves up to 3\/8-fold speedups on the Haswell\/KNL-based platforms, respectively, against ScaLAPACK\n            <jats:italic>PDGESVD<\/jats:italic>\n            and turns out to be a competitive alternative for well- and ill-conditioned matrices. We finally come up herein with a performance model based on these empirical results. Our QDWH-based polar decomposition and its SVD extension are freely available at https:\/\/github.com\/ecrc\/qdwh.git and https:\/\/github.com\/ecrc\/ksvd.git, respectively, and have been integrated into the Cray Scientific numerical library LibSci v17.11.1.\n          <\/jats:p>","DOI":"10.1145\/3309548","type":"journal-article","created":{"date-parts":[[2019,4,29]],"date-time":"2019-04-29T17:12:14Z","timestamp":1556557934000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["A QDWH-based SVD Software Framework on Distributed-memory Manycore Systems"],"prefix":"10.1145","volume":"45","author":[{"given":"Dalal","family":"Sukkari","sequence":"first","affiliation":[{"name":"King Abdullah University of Science and Technology, Thuwal, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hatem","family":"Ltaief","sequence":"additional","affiliation":[{"name":"King Abdullah University of Science and Technology, Thuwal, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aniello","family":"Esposito","sequence":"additional","affiliation":[{"name":"Cray EMEA Research Lab (CERL), Basel, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David","family":"Keyes","sequence":"additional","affiliation":[{"name":"King Abdullah University of Science and Technology, Thuwal, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,4,26]]},"reference":[{"doi-asserted-by":"crossref","unstructured":"E. Anderson Z. Bai C. H. Bischof L. S. Blackford J. W. Demmel J. J. Dongarra J. J. Du Croz A. Greenbaum S. Hammarling A. McKenney and D. C. Sorensen. 1999. LAPACK Users\u2019 Guide (3rd ed.). SIAM Philadelphia.  E. Anderson Z. Bai C. H. Bischof L. S. Blackford J. W. Demmel J. J. Dongarra J. J. Du Croz A. Greenbaum S. Hammarling A. McKenney and D. C. Sorensen. 1999. LAPACK Users\u2019 Guide (3rd ed.). SIAM Philadelphia.","key":"e_1_2_1_1_1","DOI":"10.1137\/1.9780898719604"},{"doi-asserted-by":"publisher","key":"e_1_2_1_2_1","DOI":"10.1016\/0377-0427(89)90366-X"},{"doi-asserted-by":"crossref","unstructured":"L. S. Blackford J. Choi A. Cleary E. F. D\u2019Azevedo J. W. Demmel I. S. Dhillon J. J. Dongarra S. Hammarling G. Henry A. Petitet K. Stanley D. W. Walker and R. C. Whaley. 1997. ScaLAPACK Users\u2019 Guide. Society for Industrial and Applied Mathematics Philadelphia.  L. S. Blackford J. Choi A. Cleary E. F. D\u2019Azevedo J. W. Demmel I. S. Dhillon J. J. Dongarra S. Hammarling G. Henry A. Petitet K. Stanley D. W. Walker and R. C. Whaley. 1997. ScaLAPACK Users\u2019 Guide. Society for Industrial and Applied Mathematics Philadelphia.","key":"e_1_2_1_3_1","DOI":"10.1137\/1.9780898719642"},{"doi-asserted-by":"publisher","key":"e_1_2_1_4_1","DOI":"10.1109\/IPDPS.2011.299"},{"unstructured":"Chameleon. 2016. The Chameleon Project. Retrieved from https:\/\/project.inria.fr\/chameleon\/.  Chameleon. 2016. The Chameleon Project. Retrieved from https:\/\/project.inria.fr\/chameleon\/.","key":"e_1_2_1_5_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_6_1","DOI":"10.2307\/2324422"},{"doi-asserted-by":"publisher","key":"e_1_2_1_7_1","DOI":"10.5555\/13513.13520"},{"volume-title":"Proceedings of the 5th SIAM Conference on Applied Linear Algebra, John G. Lewis (Ed.). Society for Industrial and Applied Mathematics","author":"Higham N. J.","key":"e_1_2_1_8_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_9_1","DOI":"10.1088\/0953-8984\/26\/21\/213201"},{"doi-asserted-by":"publisher","key":"e_1_2_1_10_1","DOI":"10.1109\/TAES.1977.308390"},{"doi-asserted-by":"publisher","key":"e_1_2_1_11_1","DOI":"10.1145\/169627.169855"},{"doi-asserted-by":"publisher","key":"e_1_2_1_12_1","DOI":"10.1137\/090774999"},{"doi-asserted-by":"publisher","key":"e_1_2_1_13_1","DOI":"10.1137\/120876605"},{"unstructured":"A. Petitet R. C. Whaley J. J. Dongarra and A. Cleary. 2008. HPL\u2014A Portable Implementation of the High-Performance Linpack Benchmark for Distributed-Memory Computers. http:\/\/www.netlib.org\/benchmark\/hpl.  A. Petitet R. C. Whaley J. J. Dongarra and A. Cleary. 2008. HPL\u2014A Portable Implementation of the High-Performance Linpack Benchmark for Distributed-Memory Computers. http:\/\/www.netlib.org\/benchmark\/hpl.","key":"e_1_2_1_14_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_15_1","DOI":"10.1145\/2427023.2427030"},{"doi-asserted-by":"publisher","key":"e_1_2_1_16_1","DOI":"10.1109\/TPDS.2017.2755655"},{"doi-asserted-by":"publisher","key":"e_1_2_1_17_1","DOI":"10.1145\/2894747"},{"doi-asserted-by":"publisher","key":"e_1_2_1_18_1","DOI":"10.1007\/978-3-319-43659-3_44"},{"doi-asserted-by":"publisher","key":"e_1_2_1_19_1","DOI":"10.1109\/TPDS.2017.2703149"},{"volume-title":"Proceedings of the16th IEEE\/ACIS International Conference on Computer and Information Science, G. Zhu, S. Yao, X. Cui, and S. Xu (Eds.). IEEE Computer Society","author":"Wang Y.","key":"e_1_2_1_20_1"}],"container-title":["ACM Transactions on Mathematical Software"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3309548","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3309548","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T17:49:23Z","timestamp":1750268963000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3309548"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,4,26]]},"references-count":20,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2019,6,30]]}},"alternative-id":["10.1145\/3309548"],"URL":"https:\/\/doi.org\/10.1145\/3309548","relation":{},"ISSN":["0098-3500","1557-7295"],"issn-type":[{"type":"print","value":"0098-3500"},{"type":"electronic","value":"1557-7295"}],"subject":[],"published":{"date-parts":[[2019,4,26]]},"assertion":[{"value":"2017-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-01-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-04-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}