{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T17:39:12Z","timestamp":1781890752141,"version":"3.54.5"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,10,28]],"date-time":"2024-10-28T00:00:00Z","timestamp":1730073600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,10,28]],"date-time":"2024-10-28T00:00:00Z","timestamp":1730073600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100011033","name":"Agencia Estatal de Investigaci\u00f3n","doi-asserted-by":"publisher","award":["RYC2021-033973-I"],"award-info":[{"award-number":["RYC2021-033973-I"]}],"id":[{"id":"10.13039\/501100011033","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004834","name":"Universitat Jaume I","doi-asserted-by":"publisher","award":["UJI-2023-04"],"award-info":[{"award-number":["UJI-2023-04"]}],"id":[{"id":"10.13039\/501100004834","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002809","name":"Generalitat de Catalunya","doi-asserted-by":"publisher","award":["2021-SGR-01007"],"award-info":[{"award-number":["2021-SGR-01007"]}],"id":[{"id":"10.13039\/501100002809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003329","name":"Ministerio de Econom\u00eda y Competitividad","doi-asserted-by":"publisher","award":["PID2019-107255GB"],"award-info":[{"award-number":["PID2019-107255GB"]}],"id":[{"id":"10.13039\/501100003329","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004834","name":"Universitat Jaume I","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004834","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2025,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper investigates the efficient application of half-precision floating-point (FP16) arithmetic on GPUs for boosting LU decompositions in double (FP64) precision. Addressing the motivation to enhance computational efficiency, we introduce two novel algorithms: Pre-Pivoted LU (PRP) and Mixed-precision Panel Factorization (MPF). Deployed in both hybrid CPU-GPU setups and native GPU-only configurations, PRP identifies pivot lists through LU decomposition computed in reduced precision and subsequently reorders matrix rows in FP64 precision before executing LU decomposition without pivoting. Two variants of PRP, namely hPRP and xPRP, are introduced, differing in their computation of pivot lists in full half-precision or mixed half-single precision. The MPF algorithm generates FP64 LU factorization while internally utilizing hPRP for panel factorization, showcasing accuracy on par with standard DGETRF but with superior speed. The study further explores auxiliary functions required for the native mode implementation of PRP variants and MPF.<\/jats:p>","DOI":"10.1007\/s11227-024-06523-w","type":"journal-article","created":{"date-parts":[[2024,10,28]],"date-time":"2024-10-28T10:16:42Z","timestamp":1730110602000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Mixed-precision pre-pivoting strategy for the LU factorization"],"prefix":"10.1007","volume":"81","author":[{"given":"Nima","family":"Sahraneshinsamani","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sandra","family":"Catal\u00e1n","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jos\u00e9 R.","family":"Herrero","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,10,28]]},"reference":[{"issue":"4","key":"6523_CR1","doi-asserted-by":"publisher","first-page":"2536","DOI":"10.1137\/18M1229511","volume":"41","author":"NJ Higham","year":"2019","unstructured":"Higham NJ, Pranesh S, Zounon M (2019) Squeezing a matrix into half precision, with an application to solving linear systems. SIAM Journal on Scientific Computing 41(4):2536\u20132551. https:\/\/doi.org\/10.1137\/18M1229511","journal-title":"SIAM Journal on Scientific Computing"},{"key":"6523_CR2","doi-asserted-by":"crossref","unstructured":"Haidar A, Tomov S, Dongarra J, Higham NJ (2018) Harnessing GPU tensor cores for fast FP16 arithmetic to speed up mixed-precision iterative refinement solvers. In: IEEE (ed.) SC \u201918 Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis, Dallas, Texas, November 11\u201316, 2018, pp. 47\u201314711. IEEE Computer Society Press, pub-IEEE:adr","DOI":"10.1109\/SC.2018.00050"},{"key":"6523_CR3","doi-asserted-by":"publisher","unstructured":"Dongarra J, Grigori L, Higham NJ (2020) Numerical algorithms for high-performance computational science. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 378(2166):20190066 https:\/\/doi.org\/10.1098\/rsta.2019.0066https:\/\/royalsocietypublishing.org\/doi\/pdf\/10.1098\/rsta.2019.0066","DOI":"10.1098\/rsta.2019.0066"},{"issue":"1","key":"6523_CR4","doi-asserted-by":"publisher","first-page":"258","DOI":"10.1137\/19M1298263","volume":"43","author":"NJ Higham","year":"2021","unstructured":"Higham NJ, Pranesh S (2021) Exploiting lower precision arithmetic in solving symmetric positive definite linear systems and least squares problems. SIAM Journal on Scientific Computing 43(1):258\u2013277. https:\/\/doi.org\/10.1137\/19M1298263","journal-title":"SIAM Journal on Scientific Computing"},{"key":"6523_CR5","doi-asserted-by":"publisher","first-page":"237","DOI":"10.1007\/978-3-030-50417-5_18","volume-title":"Computational Sci - ICCS 2020","author":"A Abdelfattah","year":"2020","unstructured":"Abdelfattah A, Tomov S, Dongarra J (2020) Investigating the benefit of FP16-enabled mixed-precision solvers for symmetric positive definite matrices using GPUs. In: Krzhizhanovskaya VV, Z\u00e1vodszky G, Lees MH, Dongarra JJ, Sloot PMA, Brissos S, Teixeira J (eds) Computational Sci - ICCS 2020. Springer, Cham, pp 237\u2013250"},{"key":"6523_CR6","doi-asserted-by":"publisher","unstructured":"Anzt, H., Flegar, G., Gr\u00fctzmacher, T., Quintana-Ort\u00ed, E.S.: Toward a modular precision ecosystem for high-performance computing. Int. J. High Perform. Comput. Appl. 33(6) (2019) https:\/\/doi.org\/10.1177\/1094342019846547","DOI":"10.1177\/1094342019846547"},{"issue":"1","key":"6523_CR7","doi-asserted-by":"publisher","first-page":"501","DOI":"10.1007\/s00440-024-01276-2","volume":"189","author":"H Huang","year":"2024","unstructured":"Huang H, Tikhomirov K (2024) Average-case analysis of the Gaussian elimination with partial pivoting. Probability Theory Related Fields 189(1):501\u2013567. https:\/\/doi.org\/10.1007\/s00440-024-01276-2","journal-title":"Probability Theory Related Fields"},{"key":"6523_CR8","doi-asserted-by":"publisher","unstructured":"Lindquist, N., Luszczek, P., Dongarra, J.: Using additive modifications in lu factorization instead of pivoting. In: Proceedings of the 37th ACM International Conference on Supercomputing. ICS \u201923, pp. 14\u201324. Association for Computing Machinery, New York, NY, USA (2023). https:\/\/doi.org\/10.1145\/3577193.3593731","DOI":"10.1145\/3577193.3593731"},{"issue":"2","key":"6523_CR9","doi-asserted-by":"publisher","first-page":"165","DOI":"10.1177\/10943420221136848","volume":"37","author":"F Lopez","year":"2023","unstructured":"Lopez F, Mary T (2023) Mixed precision LU factorization on GPU tensor cores: reducing data movement and memory footprint. The Int J High Performance Comput Appl 37(2):165\u2013179. https:\/\/doi.org\/10.1177\/10943420221136848","journal-title":"The Int J High Performance Comput Appl"},{"key":"6523_CR10","doi-asserted-by":"publisher","unstructured":"IEEE: IEEE standard for floating-point arithmetic. IEEE Std 754-2008, 1\u201370 (2008) https:\/\/doi.org\/10.1109\/IEEESTD.2008.4610935","DOI":"10.1109\/IEEESTD.2008.4610935"},{"key":"6523_CR11","doi-asserted-by":"publisher","unstructured":"IEEE: IEEE standard for floating-point arithmetic. IEEE Std 754-2019 (Revision of IEEE 754-2008), 1\u201384 (2019) https:\/\/doi.org\/10.1109\/IEEESTD.2019.8766229","DOI":"10.1109\/IEEESTD.2019.8766229"},{"key":"6523_CR12","unstructured":"GNU: The GNU Multiple Precision Arithmetic Library. https:\/\/gmplib.org\/ (2023)"},{"issue":"2","key":"6523_CR13","doi-asserted-by":"publisher","first-page":"13","DOI":"10.1145\/1236463.1236468","volume":"33","author":"L Fousse","year":"2007","unstructured":"Fousse L, Hanrot G, Lef\u00e8vre V, P\u00e9lissier P, Zimmermann P (2007) Mpfr: A multiple-precision binary floating-point library with correct rounding. ACM Transactions on Mathematical Software (TOMS) 33(2):13","journal-title":"ACM Transactions on Mathematical Software (TOMS)"},{"issue":"4","key":"6523_CR14","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1145\/3368086","volume":"45","author":"G Flegar","year":"2019","unstructured":"Flegar G, Scheidegger F, Novakovi\u0107 V, Mariani G, Tom\u00e1s AE, Malossi ACI, Quintana-Ort\u00ed ES (2019) FloatX: A C++ library for customized floating-point arithmetic. ACM Transactions on Mathematical Software 45(4):40\u201314023. https:\/\/doi.org\/10.1145\/3368086","journal-title":"ACM Transactions on Mathematical Software"},{"key":"6523_CR15","doi-asserted-by":"publisher","unstructured":"Van Der\u00a0Hoeven J (2017) Multiple precision floating-point arithmetic on SIMD processors. In: 2017 IEEE 24th Symposium on Computer Arithmetic (ARITH), pp. 2\u20139. https:\/\/doi.org\/10.1109\/ARITH.2017.12","DOI":"10.1109\/ARITH.2017.12"},{"issue":"1","key":"6523_CR16","doi-asserted-by":"publisher","first-page":"26","DOI":"10.1109\/TC.2019.2936192","volume":"69","author":"H Zhang","year":"2020","unstructured":"Zhang H, Chen D, Ko S-B (2020) New flexible multiple-precision multiply-accumulate unit for deep neural network training and inference. IEEE Trans on Computers 69(1):26\u201338. https:\/\/doi.org\/10.1109\/TC.2019.2936192","journal-title":"IEEE Trans on Computers"},{"key":"6523_CR17","doi-asserted-by":"publisher","unstructured":"Durand Y, Guthmuller E, Fuguet C, Fereyre J, Bocco A, Alidori R (2022) Accelerating variants of the conjugate gradient with the variable precision processor. In: 2022 IEEE 29th Symposium on Computer Arithmetic (ARITH), pp. 51\u201357. https:\/\/doi.org\/10.1109\/ARITH54963.2022.00017","DOI":"10.1109\/ARITH54963.2022.00017"},{"key":"6523_CR18","unstructured":"Golub GH, Van Loan CF (2013) Matrix Computations, 4th edn. Johns Hopkins Studies in the Mathematical Sciences, p. 756. The Johns Hopkins University Press, Baltimore, Maryland, USA. https:\/\/jhupbooks.press.jhu.edu\/title\/matrix-computations"},{"key":"6523_CR19","doi-asserted-by":"publisher","first-page":"249","DOI":"10.1016\/0024-3795(91)90337-V","volume":"149","author":"G Poole","year":"1991","unstructured":"Poole G, Neal L (1991) A geometric analysis of gaussian elimination. i. Linear Algebra its Applications 149:249\u2013272. https:\/\/doi.org\/10.1016\/0024-3795(91)90337-V","journal-title":"Linear Algebra its Applications"},{"key":"6523_CR20","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898719604","volume-title":"LAPACK Users\u2019 Guide","author":"E Anderson","year":"1999","unstructured":"Anderson E, Bai Z, Bischof C, Blackford S, Demmel J, Dongarra J, Du Croz J, Greenbaum A, Hammarling S, McKenney A, Sorensen D (1999) LAPACK Users\u2019 Guide, 3rd edn. Society for Industrial and Applied Mathematics, Philadelphia, PA","edition":"3"},{"key":"6523_CR21","unstructured":"Guennebaud G, Jacob B et al (2010) Eigen v3. http:\/\/eigen.tuxfamily.org"},{"key":"6523_CR22","unstructured":"MAGMA: Matrix Algebra on GPU and Multicore Architectures (MAGMA) Project. http:\/\/icl.cs.utk.edu\/magma\/ (2022)"},{"issue":"1","key":"6523_CR23","doi-asserted-by":"publisher","first-page":"417","DOI":"10.1137\/20M1357238","volume":"42","author":"M Fasi","year":"2021","unstructured":"Fasi M, Higham NJ (2021) Matrices with tunable infinity-norm condition number and no need for pivoting in LU factorization. SIAM J. Matrix Anal. Appl 42(1):417\u2013435","journal-title":"SIAM J. Matrix Anal. Appl"},{"issue":"3","key":"6523_CR24","doi-asserted-by":"publisher","first-page":"281","DOI":"10.1145\/321075.321076","volume":"8","author":"JH Wilkinson","year":"1961","unstructured":"Wilkinson JH (1961) Error analysis of direct methods of matrix inversion. J. ACM 8(3):281\u2013330. https:\/\/doi.org\/10.1145\/321075.321076","journal-title":"J. ACM"},{"issue":"4","key":"6523_CR25","doi-asserted-by":"publisher","first-page":"548","DOI":"10.1137\/1013095","volume":"13","author":"JH Wilkinson","year":"1971","unstructured":"Wilkinson JH (1971) Modern error analysis. SIAM Review 13(4):548\u2013568. https:\/\/doi.org\/10.1137\/1013095","journal-title":"SIAM Review"},{"key":"6523_CR26","unstructured":"Higham NJ (1989) How accurate is gaussian elimination? Technical report, Cornell University"},{"key":"6523_CR27","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898718027","volume-title":"Accuracy and Stability of Numerical Algorithms","author":"NJ Higham","year":"2002","unstructured":"Higham NJ (2002) Accuracy and Stability of Numerical Algorithms, 2nd edn. Society for Industrial and Applied Mathematics, USA","edition":"2"},{"issue":"12","key":"6523_CR28","doi-asserted-by":"publisher","first-page":"2700","DOI":"10.1109\/TPDS.2018.2842785","volume":"29","author":"s Abdelfattah","year":"2018","unstructured":"Abdelfattah s, Haidar A, Tomov S, Dongarra J (2018) Analysis and design techniques towards high-performance and energy-efficient dense linear solvers on gpus. IEEE Trans on Parallel Distributed Syst 29(12):2700\u20132712. https:\/\/doi.org\/10.1109\/TPDS.2018.2842785","journal-title":"IEEE Trans on Parallel Distributed Syst"},{"key":"6523_CR29","unstructured":"Strazdins P et al (1998) A comparison of lookahead and algorithmic blocking techniques for parallel matrix factorization. Technical Report TR-CS-98-07, The Australian National University, Department of Computer Science, Canberra 0200 ACT, Australia"},{"key":"6523_CR30","unstructured":"NVIDIA Corporation: Whitepaper: NVIDIA Tesla V100 GPU architecture; the world\u2019s most advanced data center GPU. Technical report, NVIDIA (2017). https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf Accessed 2024-01-27"},{"key":"6523_CR31","unstructured":"NVIDIA Corporation: Whitepaper: NVIDIA A100 tensor core GPU architecture; unprecedented acceleration at every scale. Technical report, NVIDIA (2020). https:\/\/images.nvidia.com\/aem-dam\/en-zz\/Solutions\/data-center\/nvidia-ampere-architecture-whitepaper.pdf Accessed 2024-01-27"},{"key":"6523_CR32","unstructured":"NVIDIA Corporation: Whitepaper: NVIDIA Ampere GA102 GPU architecture; second-generation RTX. Technical report, NVIDIA (2021). https:\/\/www.nvidia.com\/content\/PDF\/nvidia-ampere-ga-102-gpu-architecture-whitepaper-v2.pdf Accessed 2024-01-27"},{"key":"6523_CR33","unstructured":"NVIDIA Corporation: NVIDIA Nsight Systems. https:\/\/docs.nvidia.com\/nsight-systems\/UserGuide\/index.html (2023)"},{"key":"6523_CR34","unstructured":"NVIDIA Corporation: NVIDIA Nsight Compute. https:\/\/docs.nvidia.com\/nsight-compute\/NsightCompute\/index.html (2023)"},{"key":"6523_CR35","volume-title":"The Algebraic Eigenvalue Problem","author":"JH Wilkinson","year":"1988","unstructured":"Wilkinson JH (1988) The Algebraic Eigenvalue Problem. Oxford University Press, Oxford"},{"issue":"1","key":"6523_CR36","doi-asserted-by":"publisher","first-page":"140","DOI":"10.1137\/S0895479894246905","volume":"18","author":"TA Davis","year":"1997","unstructured":"Davis TA, Duff IS (1997) An unsymmetric-pattern multifrontal method for sparse lu factorization. SIAM J Matrix Analysis Appl 18(1):140\u2013158. https:\/\/doi.org\/10.1137\/S0895479894246905","journal-title":"SIAM J Matrix Analysis Appl"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-024-06523-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-024-06523-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-024-06523-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,28]],"date-time":"2024-10-28T10:41:51Z","timestamp":1730112111000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-024-06523-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,28]]},"references-count":36,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1]]}},"alternative-id":["6523"],"URL":"https:\/\/doi.org\/10.1007\/s11227-024-06523-w","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,28]]},"assertion":[{"value":"22 September 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 October 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"87"}}