{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T15:14:56Z","timestamp":1784906096501,"version":"3.55.0"},"reference-count":58,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2025,9,25]],"date-time":"2025-09-25T00:00:00Z","timestamp":1758758400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"DOI":"10.13039\/100006228","name":"Oak Ridge National Laboratory","doi-asserted-by":"publisher","award":["DE-AC05-00OR22725"],"award-info":[{"award-number":["DE-AC05-00OR22725"]}],"id":[{"id":"10.13039\/100006228","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000181","name":"Air Force Office of Scientific Research","doi-asserted-by":"publisher","award":["FA8750-19-2-1000"],"award-info":[{"award-number":["FA8750-19-2-1000"]}],"id":[{"id":"10.13039\/100000181","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000083","name":"Directorate for Computer and Information Science and Engineering","doi-asserted-by":"publisher","award":["2004541"],"award-info":[{"award-number":["2004541"]}],"id":[{"id":"10.13039\/100000083","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:p>We present a mixed-precision benchmark called HPL-MxP that uses both a lower-precision LU factorization with a non-stationary iterative refinement based on GMRES. We evaluate the numerical stability of one of the methods of generating the input matrix in a scalable fashion and show how the diagonal scaling affects the solution quality in terms of the backward-error. Some of the performance results at large scale supercomputing installations produced Exascale-level compute throughput numbers thus proving the viability of the proposed benchmark for evaluating such machines. We also present the potential of the benchmark to continue increasing its use with proliferation of hardware accelerators for AI workloads whose reliable evaluation continues to pose a particular challenge for the users.<\/jats:p>","DOI":"10.1177\/10943420251382476","type":"journal-article","created":{"date-parts":[[2025,9,26]],"date-time":"2025-09-26T06:14:42Z","timestamp":1758867282000},"page":"52-62","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":4,"title":["HPL-MxP benchmark: Mixed-precision algorithms, iterative refinement, and scalable data generation"],"prefix":"10.1177","volume":"40","author":[{"given":"Jack","family":"Dongarra","sequence":"first","affiliation":[{"name":"Innovative Computing Laboratory, University of Tennessee, Knoxville, TN, USA"},{"name":"Computer Science and Mathematics, ORNL, Oak Ridge, TN, USA"},{"name":"Applied Mathematics, University of Manchester, Manchester, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0089-6965","authenticated-orcid":false,"given":"Piotr","family":"Luszczek","sequence":"additional","affiliation":[{"name":"Innovative Computing Laboratory, University of Tennessee, Knoxville, TN, USA"},{"name":"MIT Lincoln Laboratory, LLSC, Lexington, MA, US"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2025,9,25]]},"reference":[{"key":"e_1_3_4_2_1","doi-asserted-by":"publisher","DOI":"10.1137\/23M1549079"},{"key":"e_1_3_4_3_1","volume-title":"Standard for binary floating point arithmetic","author":"ANSI\/IEEE Standard 754-1985","year":"1985","unstructured":"ANSI\/IEEE Standard 754-1985 (1985) Standard for binary floating point arithmetic. IEEE."},{"key":"e_1_3_4_4_1","volume-title":"On Improving the Performance of the Linear Solver Restarted GMRES","author":"Baker AH","year":"2003","unstructured":"Baker AH (2003) On Improving the Performance of the Linear Solver Restarted GMRES. PhD Thesis. University of Colorado."},{"key":"e_1_3_4_5_1","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479803422014"},{"key":"e_1_3_4_6_1","doi-asserted-by":"publisher","DOI":"10.1137\/0613015"},{"key":"e_1_3_4_7_1","doi-asserted-by":"publisher","DOI":"10.1137\/17M1122918"},{"key":"e_1_3_4_8_1","doi-asserted-by":"publisher","DOI":"10.1137\/17M1140819"},{"key":"e_1_3_4_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2917663"},{"key":"e_1_3_4_10_1","first-page":"116","article-title":"On the correctness of some bisection-like parallel eigenvalue algorithms in floating point arithmetic","volume":"3","author":"Demmel JW","year":"1995","unstructured":"Demmel JW, Dhillon I, Ren H (1995) On the correctness of some bisection-like parallel eigenvalue algorithms in floating point arithmetic. Electronic Transactions on Numerical Analysis 3: 116\u2013149.","journal-title":"Electronic Transactions on Numerical Analysis"},{"key":"e_1_3_4_11_1","first-page":"13","volume-title":"Contemporary High Performance Computing: From Petascale Toward Exascale","author":"Dongarra J","year":"2013","unstructured":"Dongarra J, Luszczek P (2013) HPC challenge: design, history, and implementation highlights. In: Vetter JS (ed) Contemporary High Performance Computing: From Petascale Toward Exascale. Boca Raton: Taylor & Francis, 13\u201332."},{"key":"e_1_3_4_12_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.728"},{"issue":"1","key":"e_1_3_4_13_1","first-page":"12","article-title":"The high-performance conjugate gradients benchmark","volume":"51","author":"Dongarra J","year":"2018","unstructured":"Dongarra J, Heroux MA, Luszczek P (2018) The high-performance conjugate gradients benchmark. SIAM News 51(1): 12.","journal-title":"SIAM News"},{"key":"e_1_3_4_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF01732607"},{"key":"e_1_3_4_15_1","volume-title":"Eigenvalues and Condition Numbers of Random Matrices","author":"Edelman A","year":"1989","unstructured":"Edelman A (1989) Eigenvalues and Condition Numbers of Random Matrices. PhD Thesis. Massachusetts Institute of Technology."},{"key":"e_1_3_4_16_1","volume-title":"The Annotated C++ Reference Manual","author":"Ellis MA","year":"1990","unstructured":"Ellis MA, Stroustrup B (1990) The Annotated C++ Reference Manual. Reading, Massachusetts: Addison-Wesley."},{"key":"e_1_3_4_17_1","doi-asserted-by":"publisher","DOI":"10.1137\/20M1357238"},{"key":"e_1_3_4_18_1","doi-asserted-by":"publisher","DOI":"10.7717\/peerj-cs.330"},{"key":"e_1_3_4_19_1","volume-title":"Efficient Computation of the Singular Value Decomposition with Applications to Least Squares Problems","author":"Gu M","year":"1994","unstructured":"Gu M, Demmel J, Dhillon I (1994) Efficient Computation of the Singular Value Decomposition with Applications to Least Squares Problems. University of Tennessee. Technical Report CS-94-257, Department of Computer Science."},{"key":"e_1_3_4_20_1","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis","author":"Haidar A","year":"2018","unstructured":"Haidar A, Tomov S, Dongarra J, et al. (2018) Harnessing GPU tensor cores for fast FP16 arithmetic to speed up mixed-precision iterative refinement solvers. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis. IEEE Press."},{"key":"e_1_3_4_21_1","doi-asserted-by":"publisher","DOI":"10.1098\/rspa.2020.0110"},{"key":"e_1_3_4_22_1","doi-asserted-by":"publisher","DOI":"10.1137\/18M1226312"},{"key":"e_1_3_4_23_1","doi-asserted-by":"publisher","DOI":"10.1137\/18M1229511"},{"key":"e_1_3_4_24_1","first-page":"1","article-title":"P1673R12: a free function linear algebra interface based on the BLAS","volume":"14","author":"Hoemmen M","year":"2023","unstructured":"Hoemmen M, Hollman D, Trott C, et al. (2023) P1673R12: a free function linear algebra interface based on the BLAS. ISO JTC21\/SC22\/WG21 Library Evolution Working Group 14: 1\u2013141.","journal-title":"ISO JTC21\/SC22\/WG21 Library Evolution Working Group"},{"key":"e_1_3_4_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1731022.1731027"},{"key":"e_1_3_4_26_1","volume-title":"IEEE standard for floating-point arithmetic. (revision of IEEE 754-2008)","author":"IEEE Std 754-2019","year":"2019","unstructured":"IEEE Std 754-2019 (2019) IEEE standard for floating-point arithmetic. (revision of IEEE 754-2008). IEEE."},{"key":"e_1_3_4_27_1","volume-title":"The C Programming Language","author":"Kernighan BW","year":"1978","unstructured":"Kernighan BW, Ritchie DM (1978) The C Programming Language. Upper Saddle River, New Jersey: Prentice-Hall."},{"key":"e_1_3_4_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ScalA51936.2020.00014"},{"key":"e_1_3_4_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1188455.1188573"},{"key":"e_1_3_4_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.1998.10030"},{"key":"e_1_3_4_31_1","doi-asserted-by":"crossref","unstructured":"Lindquist N Luszczek P Dongarra J (2020) Improving the performance of the GMRES method using mixed-precision techniques. In Proceedings of Smokey Mountain Conference Oak Ridge TN USA 26-28 August-2020.","DOI":"10.1007\/978-3-030-63393-6_4"},{"key":"e_1_3_4_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ScalAH56622.2022.00010"},{"key":"e_1_3_4_33_1","doi-asserted-by":"publisher","unstructured":"Lindquist N Luszczek P Dongara J (2023) Using additive modifications in LU factorization instead of pivoting. In International Conference on Supercomputing (ICS) 2023 Orlando Florida USA June 21-23-2023 14\u201324. Best paper nominee. DOI: 10.1145\/3577193.3593731.","DOI":"10.1145\/3577193.3593731"},{"key":"e_1_3_4_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3699714"},{"key":"e_1_3_4_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41404.2022.00083"},{"issue":"4","key":"e_1_3_4_36_1","first-page":"18","article-title":"Design and implementation of the HPCC benchmark suite","volume":"2","author":"Luszczek P","year":"2006","unstructured":"Luszczek P, Dongarra J, Kepner J (2006) Design and implementation of the HPCC benchmark suite. CT Watch Quarterly 2(4A): 18\u201323.","journal-title":"CT Watch Quarterly"},{"key":"e_1_3_4_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC.2017.8091031"},{"key":"e_1_3_4_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC43674.2020.9286145"},{"key":"e_1_3_4_39_1","doi-asserted-by":"publisher","DOI":"10.1177\/10943420241281050"},{"key":"e_1_3_4_40_1","volume-title":"Cramming more components onto integrated circuits","author":"Moore GE","year":"1965","unstructured":"Moore GE (1965) Cramming more components onto integrated circuits. IEEE."},{"key":"e_1_3_4_41_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cam.2019.112512"},{"key":"e_1_3_4_42_1","doi-asserted-by":"publisher","DOI":"10.1177\/10943420241239588"},{"key":"e_1_3_4_43_1","doi-asserted-by":"publisher","DOI":"10.1515\/crll.1954.193.143"},{"key":"e_1_3_4_44_1","doi-asserted-by":"publisher","DOI":"10.1137\/080725167"},{"key":"e_1_3_4_45_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-95168-3_29"},{"key":"e_1_3_4_46_1","doi-asserted-by":"publisher","DOI":"10.1137\/S1064827500381239"},{"key":"e_1_3_4_47_1","doi-asserted-by":"publisher","DOI":"10.1137\/050630416"},{"key":"e_1_3_4_48_1","volume-title":"34th Conference on Neural Information Processing Systems (Neurips 2020)","author":"Rouhani B","year":"2020","unstructured":"Rouhani B, Lo D, Zhao R, et al. (2020) Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point. In 34th Conference on Neural Information Processing Systems (Neurips 2020). https:\/\/proceedings.neurips.cc\/paper\/2020\/file\/747e32ab0fea7fbd2ad9ec03daa3f840-Paper.pdf"},{"key":"e_1_3_4_49_1","doi-asserted-by":"publisher","DOI":"10.1137\/0914028"},{"key":"e_1_3_4_50_1","doi-asserted-by":"publisher","DOI":"10.1137\/0907058"},{"key":"e_1_3_4_51_1","volume-title":"Numerical Linear Algebra with Applications","author":"\u015awirydowicz K","year":"2019","unstructured":"\u015awirydowicz K, Langou J, Ananthan S, et al. (2019) Low synchronization gram-schmidt and GMRES algorithms. In: Numerical Linear Algebra with Applications."},{"key":"e_1_3_4_52_1","first-page":"1","volume-title":"ScalAH22: 13Th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Heterogeneous Systems","author":"Tsai Y","year":"2022","unstructured":"Tsai Y, Luszczek P, Dongarra J (2022) Mixed-precision algorithm for finding selected eigenvalues and eigenvectors of symmetric and Hermitian matrices. In ScalAH22: 13Th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Heterogeneous Systems. IEEE, 1\u201310."},{"key":"e_1_3_4_53_1","doi-asserted-by":"publisher","DOI":"10.1093\/qjmam\/1.1.287"},{"key":"e_1_3_4_54_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11075-023-01596-9"},{"key":"e_1_3_4_55_1","volume-title":"Rounding Errors in Algebraic Processes London: Notes on Applied Science No. 32, Her Majesty\u2019s Stationery Office","author":"Wilkinson JH","year":"1963","unstructured":"Wilkinson JH (1963) Rounding Errors in Algebraic Processes London: Notes on Applied Science No. 32, Her Majesty\u2019s Stationery Office. Dover."},{"key":"e_1_3_4_56_1","volume-title":"The algebraic eigenvalue problem","author":"Wilkinson JH","year":"1965","unstructured":"Wilkinson JH (1965) The algebraic eigenvalue problem. Oxford University Press."},{"key":"e_1_3_4_57_1","doi-asserted-by":"publisher","DOI":"10.1137\/14M0973773"},{"key":"e_1_3_4_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/PMBS56514.2022.00015"},{"key":"e_1_3_4_59_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF02268390"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/10943420251382476","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/10943420251382476","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/10943420251382476","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/10943420251382476","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:17:50Z","timestamp":1777450670000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/10943420251382476"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,25]]},"references-count":58,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1]]}},"alternative-id":["10.1177\/10943420251382476"],"URL":"https:\/\/doi.org\/10.1177\/10943420251382476","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,25]]}}}