{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T05:00:22Z","timestamp":1755838822440,"version":"3.41.0"},"reference-count":27,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2019,3,31]],"date-time":"2019-03-31T00:00:00Z","timestamp":1553990400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Parallel Comput."],"published-print":{"date-parts":[[2019,3,31]]},"abstract":"<jats:p>We present a high-performance implementation of the Polar Decomposition (PD) on distributed-memory systems. Building upon on the QR-based Dynamically Weighted Halley (QDWH) algorithm, the key idea lies in finding the best rational approximation for the scalar sign function, which also corresponds to the polar factor for symmetric matrices, to further accelerate the QDWH convergence. Based on the Zolotarev rational functions\u2014introduced by Zolotarev (ZOLO) in 1877\u2014this new PD algorithm ZOLO-PD converges within two iterations even for ill-conditioned matrices, instead of the original six iterations needed for QDWH. ZOLO-PD uses the property of Zolotarev functions that optimality is maintained when two functions are composed in an appropriate manner. The resulting ZOLO-PD has a convergence rate up to 17, in contrast to the cubic convergence rate for QDWH. This comes at the price of higher arithmetic costs and memory footprint. These extra floating-point operations can, however, be processed in an embarrassingly parallel fashion. We demonstrate performance using up to 102,400 cores on two supercomputers. We demonstrate that, in the presence of a large number of processing units, ZOLO-PD is able to outperform QDWH by up to 2.3\u00d7 speedup, especially in situations where QDWH runs out of work, for instance, in the strong scaling mode of operation.<\/jats:p>","DOI":"10.1145\/3328723","type":"journal-article","created":{"date-parts":[[2019,6,10]],"date-time":"2019-06-10T12:10:51Z","timestamp":1560168651000},"page":"1-15","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Massively Parallel Polar Decomposition on Distributed-memory Systems"],"prefix":"10.1145","volume":"6","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6897-1095","authenticated-orcid":false,"given":"Hatem","family":"Ltaief","sequence":"first","affiliation":[{"name":"Extreme Computing Research Center, King Abdullah University of Science and Technology, Thuwal, Jeddah, KSA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dalal","family":"Sukkari","sequence":"additional","affiliation":[{"name":"Extreme Computing Research Center, King Abdullah University of Science and Technology, Thuwal, Jeddah, KSA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aniello","family":"Esposito","sequence":"additional","affiliation":[{"name":"Cray EMEA Research Lab, Bristol, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuji","family":"Nakatsukasa","sequence":"additional","affiliation":[{"name":"Mathematical Institute, University of Oxford, Oxford, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David","family":"Keyes","sequence":"additional","affiliation":[{"name":"Extreme Computing Research Center, King Abdullah University of Science and Technology, Thuwal, Jeddah, KSA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,6,7]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"TOP500 Supercomputing Sites. 2017. Retrieved from http:\/\/www.top500.org\/.  TOP500 Supercomputing Sites. 2017. Retrieved from http:\/\/www.top500.org\/."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.1631"},{"volume-title":"Guide","author":"Blackford L. Suzan","key":"e_1_2_1_3_1"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1137\/070699895"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342010391989"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1080\/00029890.1985.11971554"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/3037568.3037571"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF02163027"},{"volume-title":"Van Loan","year":"1996","author":"Golub Gene H.","key":"e_1_2_1_9_1"},{"volume-title":"Functions of Matrices: Theory and Computation","author":"Higham Nicholas J.","key":"e_1_2_1_10_1"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/0167-8191(94)90073-6"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1137\/0613044"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1024098014869"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10543-006-0053-4"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1137\/090774999"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1137\/140990334"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1137\/110857544"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1137\/120876605"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2427023.2427030"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2427023.2427030"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3309548"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2017.2755655"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2894747"},{"volume":"9833","volume-title":"Proceedings of the 22nd International Conference on Parallel and Distributed Computing (EuroPar\u201916)","author":"Sukkari Dalal","key":"e_1_2_1_24_1"},{"volume-title":"Trefethen and David Bau","year":"1997","author":"Lloyd","key":"e_1_2_1_25_1"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2017.2703149"},{"key":"e_1_2_1_27_1","first-page":"1","volume-title":"Izdat. Akad. Nauk SSSR","author":"Zolotarev E. I.","year":"1877"}],"container-title":["ACM Transactions on Parallel Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3328723","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3328723","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:54:01Z","timestamp":1750204441000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3328723"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,3,31]]},"references-count":27,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2019,3,31]]}},"alternative-id":["10.1145\/3328723"],"URL":"https:\/\/doi.org\/10.1145\/3328723","relation":{},"ISSN":["2329-4949","2329-4957"],"issn-type":[{"type":"print","value":"2329-4949"},{"type":"electronic","value":"2329-4957"}],"subject":[],"published":{"date-parts":[[2019,3,31]]},"assertion":[{"value":"2018-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-03-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-06-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}