{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,1]],"date-time":"2026-02-01T06:09:48Z","timestamp":1769926188521,"version":"3.49.0"},"reference-count":185,"publisher":"Cambridge University Press (CUP)","license":[{"start":{"date-parts":[[2016,5,23]],"date-time":"2016-05-23T00:00:00Z","timestamp":1463961600000},"content-version":"unspecified","delay-in-days":22,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Acta Numerica"],"published-print":{"date-parts":[[2016,5,1]]},"abstract":"<jats:p>Many crucial scientific computing applications, ranging from national security to medical advances, rely on high-performance linear algebra algorithms and technologies, underscoring their importance and broad impact. Here we present the state-of-the-art design and implementation practices for the acceleration of the predominant linear algebra algorithms on large-scale accelerated multicore systems. Examples are given with fundamental dense linear algebra algorithms \u2013 from the LU, QR, Cholesky, and LDLT factorizations needed for solving linear systems of equations, to eigenvalue and singular value decomposition (SVD) problems. The implementations presented are readily available via the open-source PLASMA and MAGMA libraries, which represent the next generation modernization of the popular LAPACK library for accelerated multicore systems.<\/jats:p><jats:p>To generate the extreme level of parallelism needed for the efficient use of these systems, algorithms of interest are redesigned and then split into well-chosen computational tasks. The task execution is scheduled over the computational components of a hybrid system of multicore CPUs with GPU accelerators and\/or Xeon Phi coprocessors, using either static scheduling or light-weight runtime systems. The use of light-weight runtime systems keeps scheduling overheads low, similar to static scheduling, while enabling the expression of parallelism through sequential-like code. This simplifies the development effort and allows exploration of the unique strengths of the various hardware components. Finally, we emphasize the development of innovative linear algebra algorithms using three technologies \u2013 mixed precision arithmetic, batched operations, and asynchronous iterations \u2013 that are currently of high interest for accelerated multicore systems.<\/jats:p>","DOI":"10.1017\/s0962492916000015","type":"journal-article","created":{"date-parts":[[2016,5,27]],"date-time":"2016-05-27T13:37:42Z","timestamp":1464356262000},"page":"1-160","source":"Crossref","is-referenced-by-count":16,"title":["Linear algebra software for large-scale accelerated multicore computing"],"prefix":"10.1017","volume":"25","author":[{"given":"A.","family":"Abdelfattah","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"H.","family":"Anzt","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"J.","family":"Dongarra","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"M.","family":"Gates","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"A.","family":"Haidar","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"J.","family":"Kurzak","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"P.","family":"Luszczek","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"S.","family":"Tomov","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"I.","family":"Yamazaki","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"A.","family":"YarKhan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2016,5,23]]},"reference":[{"key":"S0962492916000015_r130","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.91"},{"key":"S0962492916000015_r164","first-page":"305","article-title":"A comparative study of sparse approximate inverse preconditioners","volume":"30","author":"Tuma","year":"1998","journal-title":"Appl. Numer. Math"},{"key":"S0962492916000015_r182","unstructured":"S. N. Yeralan , T. A. Davis \u00a0and S. Ranka (2013), Sparse multifrontal QR on the GPU. Technical report, University of Florida."},{"key":"S0962492916000015_r27","first-page":"895","volume-title":"Proc. IEEE 27th International Parallel and Distributed Processing Symposium: IPDPS","author":"Ballard","year":"2013"},{"key":"S0962492916000015_r24","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2012.12"},{"key":"S0962492916000015_r88","first-page":"403","volume-title":"Handbook for Automatic Computation II: Linear Algebra","author":"Golub","year":"1971"},{"key":"S0962492916000015_r41","doi-asserted-by":"crossref","unstructured":"L. S. Blackford , J. Choi , A. Cleary , E. D\u2019Azevedo , J. Demmel , I. Dhillon , J. J. Dongarra , S. Hammarling , G. Henry , A. Petitet , K. Stanley , D. Walker \u00a0and R. C. Whaley (1997), ScaLAPACK Users\u2019 Guide, SIAM. www.netlib.org\/scalapack\/slug\/","DOI":"10.1137\/1.9780898719642"},{"key":"S0962492916000015_r83","doi-asserted-by":"publisher","DOI":"10.1111\/j.1365-246X.1988.tb00456.x"},{"key":"S0962492916000015_r35","unstructured":"K. Bergman (2008), Exascale computing study: Technology challenges in achieving exascale systems. DARPA IPTO ExaScale Computing Study."},{"key":"S0962492916000015_r65","unstructured":"J. Demmel , Y. Hida , X. S. Li \u00a0and E. J. Riedy (2007), Extra-precise iterative refinement for overdetermined least squares problems. Technical report EECS-2007-77, UC Berkeley. LAPACK Working Note 188."},{"key":"S0962492916000015_r51","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2008.10.002"},{"key":"S0962492916000015_r17","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-662-48096-0_50"},{"key":"S0962492916000015_r142","doi-asserted-by":"crossref","first-page":"124","DOI":"10.1007\/978-3-319-07518-1_8","volume-title":"Supercomputing","author":"Park","year":"2014"},{"key":"S0962492916000015_r114","unstructured":"Intel (2014) , Intel\u00ae64 and IA-32 architectures software developer\u2019s manual. http:\/\/download.intel.com\/products\/processor\/manual\/"},{"key":"S0962492916000015_r178","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-19328-6_6"},{"key":"S0962492916000015_r117","doi-asserted-by":"publisher","DOI":"10.1007\/BFb0018555"},{"key":"S0962492916000015_r68","first-page":"972","volume-title":"Proc. IEEE 28th International Parallel and Distributed Processing Symposium: IPDPS 2014","author":"Dong","year":"2014"},{"key":"S0962492916000015_r26","doi-asserted-by":"publisher","DOI":"10.1145\/2427023.2427025"},{"key":"S0962492916000015_r177","volume-title":"The Algebraic Eigenvalue Problem","author":"Wilkinson","year":"1988"},{"key":"S0962492916000015_r60","first-page":"#1","article-title":"The University of Florida sparse matrix collection","volume":"38","author":"Davis","year":"1994","journal-title":"ACM Trans. Math. Softw."},{"key":"S0962492916000015_r5","doi-asserted-by":"publisher","DOI":"10.1137\/0914027"},{"key":"S0962492916000015_r149","doi-asserted-by":"publisher","DOI":"10.1145\/1916461.1916462"},{"key":"S0962492916000015_r159","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479896297744"},{"key":"S0962492916000015_r55","doi-asserted-by":"publisher","DOI":"10.1137\/140968896"},{"key":"S0962492916000015_r107","doi-asserted-by":"publisher","DOI":"10.1137\/0708058"},{"key":"S0962492916000015_r168","doi-asserted-by":"publisher","DOI":"10.1016\/j.cam.2004.09.024"},{"key":"S0962492916000015_r146","doi-asserted-by":"publisher","DOI":"10.1137\/0913036"},{"key":"S0962492916000015_r95","doi-asserted-by":"publisher","DOI":"10.1147\/rd.416.0737"},{"key":"S0962492916000015_r179","doi-asserted-by":"publisher","DOI":"10.1137\/14M0973773"},{"key":"S0962492916000015_r78","doi-asserted-by":"publisher","DOI":"10.1016\/0377-0427(89)90367-1"},{"key":"S0962492916000015_r62","unstructured":"J. Demmel , L. Grigori , M. Hoemmen \u00a0and J. Langou (2008a), Implementing communication-optimal parallel and sequential QR factorizations. arXiv:0809.2407"},{"key":"S0962492916000015_r144","unstructured":"D. S. Parker (1995b), A randomizing butterfly transformation useful in block matrix computations. Technical report CSD-950024, Computer Science Department, University of California."},{"key":"S0962492916000015_r36","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-14390-8_40"},{"key":"S0962492916000015_r86","first-page":"182","volume-title":"High Performance Computing for Computational Science: VECPAR 2014","author":"Gates","year":"2014"},{"key":"S0962492916000015_r94","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479892242232"},{"key":"S0962492916000015_r110","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898718027"},{"key":"S0962492916000015_r16","unstructured":"H. Anzt (2012) Asynchronous and multiprecision linear solvers: Scalable and fault-tolerant numerics for energy efficient high performance computing. PhD thesis, Institute for Applied and Numerical Mathematics, Karlsruhe Institute of Technology."},{"key":"S0962492916000015_r124","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-8191(99)00021-6"},{"key":"S0962492916000015_r156","doi-asserted-by":"publisher","DOI":"10.1090\/S0025-5718-1980-0572859-4"},{"key":"S0962492916000015_r97","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2010.5470443"},{"key":"S0962492916000015_r85","doi-asserted-by":"publisher","DOI":"10.1007\/10703040_3"},{"key":"S0962492916000015_r151","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898718003"},{"key":"S0962492916000015_r63","doi-asserted-by":"publisher","DOI":"10.1137\/080731992"},{"key":"S0962492916000015_r166","doi-asserted-by":"publisher","DOI":"10.1137\/0913048"},{"key":"S0962492916000015_r91","doi-asserted-by":"publisher","DOI":"10.1137\/S1064827597323415"},{"key":"S0962492916000015_r33","doi-asserted-by":"publisher","DOI":"10.1016\/j.compstruc.2006.08.029"},{"key":"S0962492916000015_r40","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611971484"},{"key":"S0962492916000015_r154","doi-asserted-by":"publisher","DOI":"10.1137\/S0036142902401074"},{"key":"S0962492916000015_r6","doi-asserted-by":"publisher","DOI":"10.1016\/S0045-7825(99)00242-X"},{"key":"S0962492916000015_r93","doi-asserted-by":"publisher","DOI":"10.1137\/100788926"},{"key":"S0962492916000015_r132","first-page":"92","volume-title":"Proc. PARA 2012: Applied Parallel and Scientific Computing","author":"Messer","year":"2012"},{"key":"S0962492916000015_r49","doi-asserted-by":"publisher","DOI":"10.1177\/1094342007084026"},{"key":"S0962492916000015_r39","doi-asserted-by":"publisher","DOI":"10.1145\/365723.365736"},{"key":"S0962492916000015_r113","doi-asserted-by":"publisher","DOI":"10.1177\/1094342004041296"},{"key":"S0962492916000015_r163","doi-asserted-by":"publisher","DOI":"10.1137\/0611023"},{"key":"S0962492916000015_r89","volume-title":"Matrix Computations","author":"Golub","year":"1996"},{"key":"S0962492916000015_r25","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2008.11.005"},{"key":"S0962492916000015_r180","first-page":"17","volume-title":"11th International Conference on High Performance Computing for Computational Science: VECPAR 2014, Revised Selected Papers (Best Paper Award)","author":"Yamazaki","year":"2014"},{"key":"S0962492916000015_r101","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2012.13"},{"key":"S0962492916000015_r11","unstructured":"E. Anderson \u00a0and J. Dongarra (1990b), Evaluating block algorithm variants in LAPACK. Technical report UT-CS-90-103, Computer Science Department, University of Tennessee. LAPACK Working Note 19."},{"key":"S0962492916000015_r70","unstructured":"J. Dongarra \u00a0and R. C. Whaley (1995), A user\u2019s guide to the BLACS v1.1, Technical report UT-CS-95-281, University of Tennessee, Knoxville. LAPACK Working Note 94, updated 5 May 1997 (version 1.1)."},{"key":"S0962492916000015_r122","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.1164"},{"key":"S0962492916000015_r106","doi-asserted-by":"publisher","DOI":"10.1142\/S0129053392000183"},{"key":"S0962492916000015_r48","doi-asserted-by":"publisher","DOI":"10.1145\/1377596.1377597"},{"key":"S0962492916000015_r165","first-page":"586","volume-title":"IEEE Computer Society Conference on Computer Vision and Pattern Recognition 1991: Proc. CVPR\u201991","author":"Turk","year":"1991"},{"key":"S0962492916000015_r43","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479801384573"},{"key":"S0962492916000015_r145","doi-asserted-by":"publisher","DOI":"10.1137\/0724090"},{"key":"S0962492916000015_r80","doi-asserted-by":"publisher","DOI":"10.1007\/BF01932738"},{"key":"S0962492916000015_r38","doi-asserted-by":"publisher","DOI":"10.1137\/0908009"},{"key":"S0962492916000015_r74","first-page":"429","article-title":"Exploiting fine-grain parallelism in recursive LU factorization","volume":"22","author":"Dongarra","year":"2012","journal-title":"Advances in Parallel Computing Special Issue"},{"key":"S0962492916000015_r96","unstructured":"B. Hadri , H. Ltaief , E. Agullo \u00a0and J. Dongarra (2009), Enhancing parallelism of tile QR factorization for multicore architectures. LAPACK Working Note 222."},{"key":"S0962492916000015_r100","volume-title":"Proc. SC \u201911: International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Haidar","year":"2011"},{"key":"S0962492916000015_r15","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1976.1162766"},{"key":"S0962492916000015_r138","unstructured":"M. Naumov (2011), Parallel solution of sparse triangular linear systems in the preconditioned iterative methods on the GPU. Technical report NVR-2011-001, NVIDIA."},{"key":"S0962492916000015_r1","doi-asserted-by":"publisher","DOI":"10.1007\/BF01931804"},{"key":"S0962492916000015_r76","first-page":"#1","article-title":"Model-driven one-sided factorizations on multicore accelerated systems","volume":"1","author":"Dongarra","year":"2014","journal-title":"Int. J. Supercomputing Frontiers and Innovations"},{"key":"S0962492916000015_r108","doi-asserted-by":"publisher","DOI":"10.1109\/74.80522"},{"key":"S0962492916000015_r31","first-page":"123","volume-title":"Nonlinear Programming","author":"Bartels","year":"1971"},{"key":"S0962492916000015_r109","doi-asserted-by":"publisher","DOI":"10.1146\/annurev.fl.22.010190.001351"},{"key":"S0962492916000015_r45","first-page":"180","volume-title":"Proc. IEEE 18th Euromicro International Conference on Parallel, Distributed and Network-Based Computing: PDP 2010","author":"Broquedis","year":"2010"},{"key":"S0962492916000015_r14","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2012.11"},{"key":"S0962492916000015_r66","doi-asserted-by":"publisher","DOI":"10.1137\/070688778"},{"key":"S0962492916000015_r28","doi-asserted-by":"publisher","DOI":"10.1137\/130929060"},{"key":"S0962492916000015_r81","doi-asserted-by":"publisher","DOI":"10.1145\/356044.356047"},{"key":"S0962492916000015_r141","doi-asserted-by":"publisher","DOI":"10.1007\/BF01386090"},{"key":"S0962492916000015_r112","doi-asserted-by":"publisher","DOI":"10.1007\/BF02287921"},{"key":"S0962492916000015_r158","volume-title":"Introduction to Matrix Computations","author":"Stewart","year":"1973"},{"key":"S0962492916000015_r75","doi-asserted-by":"publisher","DOI":"10.1137\/1026003"},{"key":"S0962492916000015_r174","doi-asserted-by":"publisher","DOI":"10.1016\/0377-0427(94)00067-B"},{"key":"S0962492916000015_r61","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611971446"},{"key":"S0962492916000015_r64","doi-asserted-by":"publisher","DOI":"10.1145\/1141885.1141894"},{"key":"S0962492916000015_r58","doi-asserted-by":"publisher","DOI":"10.1145\/800195.805928"},{"key":"S0962492916000015_r59","volume-title":"Numerical Linear Algebra and Applications","author":"Datta","year":"1995"},{"key":"S0962492916000015_r92","first-page":"782","article-title":"Mathematicians of Gaussian elimination","volume":"58","author":"Grcar","year":"2011","journal-title":"Notices Amer. Math. Soc."},{"key":"S0962492916000015_r148","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2012.12"},{"key":"S0962492916000015_r185","doi-asserted-by":"publisher","DOI":"10.1109\/MAP.2008.4653660"},{"key":"S0962492916000015_r73","unstructured":"J. Dongarra , M. Faverge , H. Ltaief \u00a0and P. Luszczek (2011), Achieving numerical accuracy and high performance using recursive tile LU factorization. Technical report ICL-UT-11-08, Computer Science Department, University of Tennessee, Knoxville."},{"key":"S0962492916000015_r183","unstructured":"M.-C. Yeung \u00a0and T. F. Chan (1995), Probabilistic analysis of Gaussian elimination without pivoting. Technical report CAM95-29, Department of Mathematics, University of California, Los Angeles."},{"key":"S0962492916000015_r21","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479896296921"},{"key":"S0962492916000015_r176","volume-title":"Rounding Errors in Algebraic Processes","author":"Wilkinson","year":"1963"},{"key":"S0962492916000015_r50","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.1301"},{"key":"S0962492916000015_r72","doi-asserted-by":"publisher","DOI":"10.1145\/77626.79170"},{"key":"S0962492916000015_r3","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/180\/1\/012037"},{"key":"S0962492916000015_r125","volume-title":"Proc. 2006 ACM\/IEEE Conference on Supercomputing","author":"Langou","year":"2006"},{"key":"S0962492916000015_r105","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/42288.42291","article-title":"An extended set of Fortran Basic Linear Algebra Subprograms","volume":"14","author":"Hammarling","year":"1988","journal-title":"ACM Trans. Math. Softw."},{"key":"S0962492916000015_r90","doi-asserted-by":"publisher","DOI":"10.1137\/1018113"},{"key":"S0962492916000015_r12","doi-asserted-by":"publisher","DOI":"10.1142\/S0129053389000056"},{"key":"S0962492916000015_r120","first-page":"1","volume-title":"Applied Parallel and Scientific Computing","author":"K\u00e5gstr\u00f6m","year":"2012"},{"key":"S0962492916000015_r82","first-page":"113","article-title":"Large dense numerical linear algebra in 1993: the parallel computing influence","volume":"7","author":"Edelman","year":"1993","journal-title":"Int. J. High Performance Comput. Appl."},{"key":"S0962492916000015_r155","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4757-2312-0"},{"key":"S0962492916000015_r52","first-page":"223","volume-title":"Proc. 15th AGM SIGPLAN Symposium on Principle and Practice of Parallel Programming","author":"Castaldo","year":"2010"},{"key":"S0962492916000015_r135","first-page":"268","article-title":"Matrix computations with Fortran and paging","volume":"15","author":"Moler","year":"1972","journal-title":"Comm. Assoc. Comput. Mach."},{"key":"S0962492916000015_r140","doi-asserted-by":"publisher","DOI":"10.1137\/S1064827599362314"},{"key":"S0962492916000015_r116","first-page":"61","volume-title":"Proc. Third SIAM Conference on Parallel Processing for Scientific Computing","author":"Jessup","year":"1989"},{"key":"S0962492916000015_r103","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-38750-0_6"},{"key":"S0962492916000015_r57","volume-title":"Concepts and Applications of Finite Element Analysis","author":"Cook","year":"2007"},{"key":"S0962492916000015_r10","unstructured":"E. Anderson \u00a0and J. Dongarra (1990a) Implementation guide for LAPACK. Technical report UT-CS-90-101, Computer Science Department, University of Tennessee. LAPACK Working Note 18."},{"key":"S0962492916000015_r172","first-page":"1","volume-title":"2013 IEEE International Conference on Cluster Computing: CLUSTER IEEE","author":"Villa","year":"2013b"},{"key":"S0962492916000015_r133","volume-title":"TOP500 Supercomputer Sites","author":"Meuer","year":"2011"},{"key":"S0962492916000015_r46","doi-asserted-by":"publisher","DOI":"10.1090\/S0025-5718-1977-0428694-0"},{"key":"S0962492916000015_r129","volume-title":"Proc. 2012 IEEE Conference on High Performance Extreme Computing Conference: HPEC 2012","author":"Luszczek","year":"2012"},{"key":"S0962492916000015_r181","unstructured":"A. YarKhan , J. Kurzak \u00a0and J. Dongarra (2011), QUARK Users\u2019 Guide, Innovative Computing Laboratory, Electrical Engineering and Computer Science, University of Tennessee."},{"key":"S0962492916000015_r84","doi-asserted-by":"publisher","DOI":"10.1016\/S0377-0427(00)00409-X"},{"key":"S0962492916000015_r13","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898719604"},{"key":"S0962492916000015_r173","unstructured":"V. Volkov \u00a0and J. W. Demmel (2008), LU, QR and Cholesky factorizations using vector capabilities of GPUs. Technical report UCB\/EECS-2008-49, University of California, Berkeley. LAPACK Working Note 202."},{"key":"S0962492916000015_r47","first-page":"1","volume-title":"Applied Parallel Computing: State of the Art in Scientific Computing, PARA 2006","author":"Buttari","year":"2006"},{"key":"S0962492916000015_r99","volume-title":"Proc. SC \u201913: International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Haidar","year":"2013a"},{"key":"S0962492916000015_r42","first-page":"207","volume-title":"Proceedings of the 5th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming: PPOPP \u201995","author":"Blumofe","year":"1995"},{"key":"S0962492916000015_r22","doi-asserted-by":"publisher","DOI":"10.1177\/109434208700100403"},{"key":"S0962492916000015_r7","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479899358194"},{"key":"S0962492916000015_r143","unstructured":"D. S. Parker (1995a), Random butterfly transformations with applications in computational linear algebra. Technical report CSD-950023, Computer Science Department, University of California."},{"key":"S0962492916000015_r87","first-page":"205","article-title":"Calculating the singular values and pseudoinverse of a matrix","volume":"2","author":"Golub","year":"1965","journal-title":"SIAM J. Numer. Anal."},{"key":"S0962492916000015_r69","doi-asserted-by":"publisher","DOI":"10.1137\/0904049"},{"key":"S0962492916000015_r160","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2010.06.001"},{"key":"S0962492916000015_r153","doi-asserted-by":"publisher","DOI":"10.1137\/0910005"},{"key":"S0962492916000015_r161","unstructured":"S. Tomov , R. Nath , P. Du \u00a0and J. Dongarra (2009), MAGMA version 0.2 Users\u2019 Guide."},{"key":"S0962492916000015_r171","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-40047-6_81"},{"key":"S0962492916000015_r102","first-page":"1150","volume-title":"15th IEEE International Workshop on Parallel and Distributed Scientific and Engineering Computing: PDSEC 2014 (Best Paper)","author":"Haidar","year":"2014b"},{"key":"S0962492916000015_r37","doi-asserted-by":"crossref","unstructured":"C. Bischof (1993), A summary of block schemes for reducing a general matrix to Hessenberg form. Technical report ANL\/MCS-TM-175, Argonne National Laboratory.","DOI":"10.2172\/6698958"},{"key":"S0962492916000015_r136","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.1981.1102568"},{"key":"S0962492916000015_r53","doi-asserted-by":"publisher","DOI":"10.1145\/1248377.1248397"},{"key":"S0962492916000015_r34","first-page":"88","article-title":"Numerical experiments with parallel orderings for ILU preconditioners","volume":"8","author":"Benzi","year":"1999","journal-title":"Electron. Trans. Numer. Anal."},{"key":"S0962492916000015_r128","unstructured":"D. Lukarski (2012), Parallel sparse linear algebra for multi-core and many-core platforms: Parallel solvers and preconditioners. PhD thesis, Karlsruhe Institute of Technology (KIT), Germany."},{"key":"S0962492916000015_r54","unstructured":"J. Choi (1995), A proposal for a set of parallel basic linear algebra subprograms. Technical report UT-CS-95-292, University of Tennessee, Knoxville. LAPACK Working Note 100."},{"key":"S0962492916000015_r77","doi-asserted-by":"publisher","DOI":"10.1137\/0720002"},{"key":"S0962492916000015_r150","doi-asserted-by":"publisher","DOI":"10.1137\/0914028"},{"key":"S0962492916000015_r152","doi-asserted-by":"publisher","DOI":"10.1137\/0911008"},{"key":"S0962492916000015_r111","doi-asserted-by":"publisher","DOI":"10.1037\/h0071325"},{"key":"S0962492916000015_r8","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2005.07.004"},{"key":"S0962492916000015_r20","doi-asserted-by":"publisher","DOI":"10.1137\/0610013"},{"key":"S0962492916000015_r23","doi-asserted-by":"publisher","DOI":"10.1137\/0612048"},{"key":"S0962492916000015_r44","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479801384585"},{"key":"S0962492916000015_r71","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611971811"},{"key":"S0962492916000015_r56","first-page":"1","volume-title":"High Performance Computing","author":"Chow","year":"2015"},{"key":"S0962492916000015_r115","doi-asserted-by":"publisher","DOI":"10.1088\/0029-5515\/46\/7\/S02"},{"key":"S0962492916000015_r79","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1093\/imanum\/12.1.1","article-title":"Stability of methods for matrix inversion","volume":"12","author":"Ducroz","year":"1992","journal-title":"IMA J. Numer. Anal."},{"key":"S0962492916000015_r137","article-title":"Cramming more components onto integrated circuits","volume":"38","author":"Moore","year":"1965","journal-title":"Electronics"},{"key":"S0962492916000015_r127","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-31464-3_67"},{"key":"S0962492916000015_r32","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-31464-3_14"},{"key":"S0962492916000015_r121","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2011.05.001"},{"key":"S0962492916000015_r119","first-page":"135","volume-title":"Proc. Symposium on High Performance Computing","author":"Kabir","year":"2015"},{"key":"S0962492916000015_r4","volume-title":"SC \u201909: Proc. Conference on High Performance Computing Networking, Storage and Analysis","author":"Agullo","year":"2009b"},{"key":"S0962492916000015_r169","doi-asserted-by":"publisher","DOI":"10.1137\/0903021"},{"key":"S0962492916000015_r157","first-page":"1","volume-title":"Proc. International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Solc\u00e0","year":"2015"},{"key":"S0962492916000015_r147","doi-asserted-by":"publisher","DOI":"10.1017\/S0022112009991972"},{"key":"S0962492916000015_r2","volume-title":"GPU Computing Gems","volume":"2","author":"Agullo","year":"2010"},{"key":"S0962492916000015_r18","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2013.05.008"},{"key":"S0962492916000015_r175","unstructured":"I. Wainwright \u00a0and H. P. C. Sweden (2013), Optimized LU-decomposition with full pivot for small batched matrices. In 2013 NVIDIA GPU Tech. Conf."},{"key":"S0962492916000015_r104","doi-asserted-by":"publisher","DOI":"10.1177\/1094342013502097"},{"key":"S0962492916000015_r162","first-page":"1","volume-title":"Proc. IEEE IPDPS \u201910","author":"Tomov","year":"2010b"},{"key":"S0962492916000015_r123","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2007.70813"},{"key":"S0962492916000015_r167","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479897317788"},{"key":"S0962492916000015_r170","doi-asserted-by":"publisher","DOI":"10.1002\/nla.1680010404"},{"key":"S0962492916000015_r131","doi-asserted-by":"publisher","DOI":"10.1007\/s00607-009-0066-3"},{"key":"S0962492916000015_r126","doi-asserted-by":"publisher","DOI":"10.1145\/355841.355847"},{"key":"S0962492916000015_r29","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611971538"},{"key":"S0962492916000015_r134","doi-asserted-by":"publisher","DOI":"10.1145\/321386.321394"},{"key":"S0962492916000015_r67","doi-asserted-by":"publisher","DOI":"10.1016\/0168-9274(91)90011-N"},{"key":"S0962492916000015_r184","doi-asserted-by":"publisher","DOI":"10.1137\/1037125"},{"key":"S0962492916000015_r98","first-page":"491","volume-title":"Proc. IEEE 28th International Parallel and Distributed Processing Symposium: IPDPS","author":"Haidar","year":"2014"},{"key":"S0962492916000015_r30","doi-asserted-by":"crossref","first-page":"573","DOI":"10.1002\/cpe.1476","article-title":"Complex version of high performance computing LINPACK benchmark (HPL)","volume":"22","author":"Barrett","year":"2010","journal-title":"Concurrency Computat. Pract. Exper."},{"key":"S0962492916000015_r118","doi-asserted-by":"publisher","DOI":"10.1007\/11558958_49"},{"key":"S0962492916000015_r19","unstructured":"M. Arioli \u00a0and I. S. Duff (2008) Using FGMRES to obtain backward stability in mixed precision. Technical report RAL-TR-2008-006, Rutherford Appleton Laboratory."},{"key":"S0962492916000015_r9","first-page":"3","volume-title":"Proc. 4th Conference on Parallel Processing for Scientific Computing","author":"Anderson","year":"1989"},{"key":"S0962492916000015_r139","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4757-4393-7"}],"container-title":["Acta Numerica"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S0962492916000015","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,17]],"date-time":"2024-06-17T05:35:11Z","timestamp":1718602511000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S0962492916000015\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,5,1]]},"references-count":185,"alternative-id":["S0962492916000015"],"URL":"https:\/\/doi.org\/10.1017\/s0962492916000015","relation":{},"ISSN":["0962-4929","1474-0508"],"issn-type":[{"value":"0962-4929","type":"print"},{"value":"1474-0508","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,5,1]]}}}