{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T13:44:40Z","timestamp":1782999880345,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":51,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,7,5]],"date-time":"2026-07-05T00:00:00Z","timestamp":1783209600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,7,6]]},"DOI":"10.1145\/3797905.3800554","type":"proceedings-article","created":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T11:50:37Z","timestamp":1782993037000},"page":"1272-1283","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Non-Delayed Cholesky Factorization"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-9990-230X","authenticated-orcid":false,"given":"Yuchen","family":"Luo","sequence":"first","affiliation":[{"name":"SSSLab, Dept. of CST, China University of Petroleum-Beijing, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-7693-4558","authenticated-orcid":false,"given":"Hongyu","family":"Chen","sequence":"additional","affiliation":[{"name":"SSSLab, Dept. of CST, China University of Petroleum-Beijing, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9525-1659","authenticated-orcid":false,"given":"Shaoshuai","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2150-5759","authenticated-orcid":false,"given":"Weifeng","family":"Liu","sequence":"additional","affiliation":[{"name":"SSSLab, Dept. of CST, China University of Petroleum-Beijing, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,5]]},"reference":[{"key":"e_1_3_3_1_2_2","doi-asserted-by":"crossref","unstructured":"Ahmad Abdelfattah Azzam Haidar Stanimire Tomov and Jack Dongarra. 2016. Performance tuning and optimization techniques of fixed and variable size batched Cholesky factorization on GPUs. Procedia Computer Science 80 (2016) 119\u2013130.","DOI":"10.1016\/j.procs.2016.05.303"},{"key":"e_1_3_3_1_3_2","doi-asserted-by":"crossref","unstructured":"Ahmad Abdelfattah Azzam Haidar Stanimire Tomov and Jack Dongarra. 2017. Fast Cholesky factorization on GPUs for batch and native modes in MAGMA. Journal of Computational Science 20 (2017) 85\u201393.","DOI":"10.1016\/j.jocs.2016.12.009"},{"key":"e_1_3_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC.2018.8547576"},{"key":"e_1_3_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2016.90"},{"key":"e_1_3_3_1_6_2","unstructured":"Emmanuel Agullo Jim Demmel Jack Dongarra Bilel Hadri Jakub Kurzak Julien Langou Hatem Ltaief Piotr Luszczek and Stanimire Tomov. 2009. Numerical linear algebra on emerging architectures: The PLASMA and MAGMA projects. Journal of Physics: Conference Series 180 1 (2009) 012037."},{"key":"e_1_3_3_1_7_2","doi-asserted-by":"crossref","unstructured":"Bjarne\u00a0Stig Andersen Jerzy Wa\u015bniewski and Fred\u00a0G Gustavson. 2001. A recursive formulation of Cholesky factorization of a matrix in packed storage. ACM Transactions on Mathematical Software 27 2 (2001) 214\u2013244.","DOI":"10.1145\/383738.383741"},{"key":"e_1_3_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.5555\/323215"},{"key":"e_1_3_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41404.2022.00034"},{"key":"e_1_3_3_1_10_2","volume-title":"Euro-Par \u201920","author":"Beaumont Olivier","year":"2020","unstructured":"Olivier Beaumont, Julien Langou, Willy Quach, and Alena Shilova. 2020. A makespan lower bound for the scheduling of the tiled cholesky factorization based on alap schedule. In Euro-Par \u201920."},{"key":"e_1_3_3_1_11_2","doi-asserted-by":"crossref","unstructured":"Alfredo Buttari Jack Dongarra Parry Husbands Jakub Kurzak and Katherine Yelick. 2007. Multithreading for synchronization tolerance in matrix factorization. Journal of Physics: Conference Series 78 1 (2007) 012028.","DOI":"10.1088\/1742-6596\/78\/1\/012028"},{"key":"e_1_3_3_1_12_2","doi-asserted-by":"crossref","unstructured":"Alfredo Buttari Julien Langou Jakub Kurzak and Jack Dongarra. 2009. A class of parallel tiled linear algebra algorithms for multicore architectures. Parallel computing 35 1 (2009) 38\u201353.","DOI":"10.1016\/j.parco.2008.10.002"},{"key":"e_1_3_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS53621.2022.00047"},{"key":"e_1_3_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/ProTools49597.2019.00009"},{"key":"e_1_3_3_1_15_2","doi-asserted-by":"crossref","unstructured":"Erin Carson Tom\u00e1\u0161 Gergelits and Ichitaro Yamazaki. 2022. Mixed precision s-step Lanczos and conjugate gradient algorithms. Numerical Linear Algebra with Applications 29 3 (2022) e2425.","DOI":"10.1002\/nla.2425"},{"key":"e_1_3_3_1_16_2","doi-asserted-by":"crossref","unstructured":"Terry Cojean Abdou Guermouche Andra Hugo Raymond Namyst and Pierre-Andr\u00e9 Wacrenier. 2019. Resource aggregation for task-based cholesky factorization on top of modern architectures. Parallel Computing 83 (2019) 73\u201392.","DOI":"10.1016\/j.parco.2018.10.007"},{"key":"e_1_3_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3721145.3725756"},{"key":"e_1_3_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2014.52"},{"key":"e_1_3_3_1_19_2","doi-asserted-by":"crossref","unstructured":"Jack Dongarra Jeremy Du\u00a0Croz Sven Hammarling and Iain\u00a0S Duff. 1990. A set of level 3 basic linear algebra subprograms. ACM Transactions on Mathematical Software 16 1 (1990) 1\u201317.","DOI":"10.1145\/77626.79170"},{"key":"e_1_3_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46079-6_37"},{"key":"e_1_3_3_1_21_2","doi-asserted-by":"crossref","unstructured":"Joseph Dorris Asim YarKhan Jakub Kurzak Piotr Luszczek and Jack Dongarra. 2018. Task based Cholesky decomposition on Xeon Phi architectures using OpenMP. International Journal of Computational Science and Engineering 17 3 (2018) 310\u2013323.","DOI":"10.1504\/IJCSE.2018.095851"},{"key":"e_1_3_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581784.3607050"},{"key":"e_1_3_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356223"},{"key":"e_1_3_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2017.18"},{"key":"e_1_3_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/InPar.2012.6339596"},{"key":"e_1_3_3_1_26_2","doi-asserted-by":"crossref","unstructured":"Fred\u00a0G Gustavson Jerzy Wa\u015bniewski Jack Dongarra Jose\u00a0R Herrero and Julien Langou. 2013. Level-3 Cholesky factorization routines improve performance of many Cholesky algorithms. ACM Transactions on Mathematical Software 39 2 (2013) 1\u201310.","DOI":"10.1145\/2427023.2427026"},{"key":"e_1_3_3_1_27_2","doi-asserted-by":"crossref","unstructured":"Fred\u00a0G Gustavson Jerzy Wa\u015bniewski Jack Dongarra and Julien Langou. 2010. Rectangular full packed format for Cholesky\u2019s algorithm: factorization solution and inversion. ACM Transactions on Mathematical Software 37 2 (2010) 1\u201321.","DOI":"10.1145\/1731022.1731028"},{"key":"e_1_3_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3038228.3038237"},{"key":"e_1_3_3_1_29_2","doi-asserted-by":"crossref","unstructured":"Azzam Haidar Ahmad Abdelfattah Mawussi Zounon Stanimire Tomov and Jack Dongarra. 2017. A guide for achieving high performance with very small matrices on GPU: A case study of batched LU and Cholesky factorizations. IEEE Transactions on Parallel and Distributed Systems 29 5 (2017) 973\u2013984.","DOI":"10.1109\/TPDS.2017.2783929"},{"key":"e_1_3_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC.2016.7761591"},{"key":"e_1_3_3_1_31_2","doi-asserted-by":"crossref","unstructured":"Jakub Kurzak Hartwig Anzt Mark Gates and Jack Dongarra. 2015. Implementation and tuning of batched Cholesky factorization and solve for NVIDIA GPUs. IEEE Transactions on Parallel and Distributed Systems 27 7 (2015) 2036\u20132048.","DOI":"10.1109\/TPDS.2015.2481890"},{"key":"e_1_3_3_1_32_2","doi-asserted-by":"crossref","unstructured":"Jakub Kurzak Alfredo Buttari and Jack Dongarra. 2008. Solving systems of linear equations on the CELL processor using Cholesky factorization. IEEE Transactions on Parallel and Distributed Systems 19 9 (2008) 1175\u20131186.","DOI":"10.1109\/TPDS.2007.70813"},{"key":"e_1_3_3_1_33_2","volume-title":"PARA \u201906","author":"Kurzak Jakub","year":"2006","unstructured":"Jakub Kurzak and Jack Dongarra. 2006. Pipelined Shared Memory Implementation of Linear Algebra Routines with arbitary Lookahead-LU, Cholesky, QR. In PARA \u201906."},{"key":"e_1_3_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3774934.3786442"},{"key":"e_1_3_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-43659-3_45"},{"key":"e_1_3_3_1_36_2","doi-asserted-by":"crossref","unstructured":"Weifeng Liu Ang Li Jonathan\u00a0D Hogg Iain\u00a0S Duff and Brian Vinter. 2017. Fast synchronization-free algorithms for parallel sparse triangular solves with multiple right-hand sides. Concurrency and Computation: Practice and Experience 29 21 (2017) e4244.","DOI":"10.1002\/cpe.4244"},{"key":"e_1_3_3_1_37_2","unstructured":"Hatem Ltaief Stanimire Tomov Rajib Nath and Jack Dongarra. 2010. Hybrid multicore cholesky factorization with multiple gpu accelerators. IEEE Transaction on Parallel and Distributed Systems 48 (2010)."},{"key":"e_1_3_3_1_38_2","doi-asserted-by":"crossref","unstructured":"Yuechen Lu Yuchen Luo Haocheng Lian Zhou Jin and Weifeng Liu. 2021. Implementing LU and Cholesky factorizations on artificial intelligence accelerators. CCF Transactions on High Performance Computing 3 3 (2021) 286\u2013297.","DOI":"10.1007\/s42514-021-00075-8"},{"key":"e_1_3_3_1_39_2","doi-asserted-by":"crossref","unstructured":"Zhengyang Lu and Weifeng Liu. 2023. Tilesptrsv: a tiled algorithm for parallel sparse triangular solve on gpus. CCF Transactions on High Performance Computing 5 2 (2023) 129\u2013143.","DOI":"10.1007\/s42514-023-00151-1"},{"key":"e_1_3_3_1_40_2","doi-asserted-by":"crossref","unstructured":"Rajib Nath Stanimire Tomov and Jack Dongarra. 2010. An improved magma gemm for fermi graphics processing units. The International Journal of High Performance Computing Applications 24 4 (2010) 511\u2013515.","DOI":"10.1177\/1094342010385729"},{"key":"e_1_3_3_1_41_2","unstructured":"NVIDIA. 2024. cuBLAS Library. https:\/\/developer.nvidia.com\/cublas."},{"key":"e_1_3_3_1_42_2","unstructured":"NVIDIA. 2024. cuSOLVER Library. https:\/\/developer.nvidia.com\/cusolver."},{"key":"e_1_3_3_1_43_2","unstructured":"NVIDIA. 2024. CUTLASS: CUDA Templates for Linear Algebra Subroutines. https:\/\/github.com\/NVIDIA\/cutlass."},{"key":"e_1_3_3_1_44_2","unstructured":"Lu Shi Gaoyuan Zou Siqi Wu and Shaoshuai Zhang. 2025. High performance Cholesky factorization on emerging GPU architectures using Tensor Cores. Computer Engineering & Science 47 7 (2025) 1170."},{"key":"e_1_3_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404397.3404400"},{"key":"e_1_3_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3712285.3759895"},{"key":"e_1_3_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC56929.2023.10247767"},{"key":"e_1_3_3_1_48_2","doi-asserted-by":"crossref","unstructured":"Ichitaro Yamazaki Stanimire Tomov and Jack Dongarra. 2015. Mixed-precision Cholesky QR factorization and its case studies on multicore CPU with multiple GPUs. SIAM Journal on Scientific Computing 37 3 (2015) C307\u2013C330.","DOI":"10.1137\/14M0973773"},{"key":"e_1_3_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41406.2024.00064"},{"key":"e_1_3_3_1_50_2","doi-asserted-by":"crossref","unstructured":"Feng Zhang Jiya Su Weifeng Liu Bingsheng He Ruofan Wu Xiaoyong Du and Rujia Wang. 2021. Yuenyeungsptrsv: a thread-level and warp-level fusion synchronization-free sparse triangular solve. IEEE Transactions on Parallel and Distributed Systems 32 9 (2021) 2321\u20132337.","DOI":"10.1109\/TPDS.2021.3066635"},{"key":"e_1_3_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ScalA51936.2020.00011"},{"key":"e_1_3_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18074.2021.9586141"}],"event":{"name":"ICS '26: 2026 International Conference on Supercomputing","location":"Belfast United Kingdom","acronym":"ICS '26","sponsor":["SIGHPC ACM Special Interest Group on High Performance Computing, Special Interest Group on High Performance Computing","SIGARCH ACM Special Interest Group on Computer Architecture"]},"container-title":["Proceedings of the 40th ACM International Conference on Supercomputing"],"original-title":[],"deposited":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T13:02:24Z","timestamp":1782997344000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3797905.3800554"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,5]]},"references-count":51,"alternative-id":["10.1145\/3797905.3800554","10.1145\/3797905"],"URL":"https:\/\/doi.org\/10.1145\/3797905.3800554","relation":{},"subject":[],"published":{"date-parts":[[2026,7,5]]},"assertion":[{"value":"2026-07-05","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}