{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T22:55:31Z","timestamp":1777676131936,"version":"3.51.4"},"reference-count":30,"publisher":"SAGE Publications","issue":"5","license":[{"start":{"date-parts":[[2025,6,3]],"date-time":"2025-06-03T00:00:00Z","timestamp":1748908800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"},{"start":{"date-parts":[[2025,6,3]],"date-time":"2025-06-03T00:00:00Z","timestamp":1748908800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"DOI":"10.13039\/100006151","name":"Basic Energy Sciences","doi-asserted-by":"publisher","award":["17-SC-20-SC"],"award-info":[{"award-number":["17-SC-20-SC"]}],"id":[{"id":"10.13039\/100006151","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002347","name":"Bundesministerium f\u00fcr Bildung und Forschung","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100002347","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2025,9]]},"abstract":"<jats:p>\n                    The direct solution of batches of band linear systems in parallel is important for many applications. In this paper, we elaborate on three new GPU algorithms for the data-parallel direct solution of linear system batches, sharing a band structure. We develop algorithms for three matrix types: tridiagonal, small bandwidth, and wide bandwidth. We exploit the band structure of the matrix, and store it in an efficient fashion (LAPACK band matrix format) and ensure that the SIMD parallelism of the GPUs are maximized. We develop a panel-based factorization for wide-band matrices to ensure coalesced access (with column-major storage) while minimizing main memory traffic. For the tridiagonal solvers, to ensure a high level of concurrency, we adapt a divide-and-conquer approach and utilize co-operative group functionality to efficiently communicate between compute units (in registers) on GPUs. We implement these algorithms for NVIDIA GPUs and study the performance for varying matrix sizes (16 to 1024) and across a range of batch items (upto 1 \u00d7 10\n                    <jats:sup>6<\/jats:sup>\n                    ). We compare the performance of our implementations with the corresponding optimized vendor implementations (cuSPARSE and MKL), with the state-of-the-art GPU library MAGMA, and with the optimized LAPACK implementation provided by Intel MKL on Intel Skylake CPUs. We also showcase the effectiveness of our batched band solvers for matrices originating from XGC, a gyrokinetic Particle-In-Cell (PIC) application optimized for modeling the edge region plasma within a plasma physics application. We show that our implementations are on average \u223c 2\u00d7 (for batched banded solvers, compared to MAGMA and MKL) to \u223c 3\u00d7 (for batched tridiagonal solvers, compared to cuSPARSE) faster than the state-of-the-art and the vendor provided implementations.\n                  <\/jats:p>","DOI":"10.1177\/10943420251347460","type":"journal-article","created":{"date-parts":[[2025,6,25]],"date-time":"2025-06-25T02:33:14Z","timestamp":1750818794000},"page":"615-630","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":0,"title":["Efficient solution of batched band linear systems on GPUs"],"prefix":"10.1177","volume":"39","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7961-1159","authenticated-orcid":false,"given":"Pratik","family":"Nayak","sequence":"first","affiliation":[{"name":"Karlsruhe Institute of Technology, Karlsruhe, Germany"},{"name":"Technical University of Munich, Heilbronn, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Isha","family":"Aggarwal","sequence":"additional","affiliation":[{"name":"Karlsruhe Institute of Technology, Karlsruhe, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2177-952X","authenticated-orcid":false,"given":"Hartwig","family":"Anzt","sequence":"additional","affiliation":[{"name":"Technical University of Munich, Heilbronn, Germany"},{"name":"ICL, University of Tennessee, Knoxville, TN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2025,6,3]]},"reference":[{"key":"e_1_3_5_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3431921"},{"key":"e_1_3_5_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3624062.3624247"},{"key":"e_1_3_5_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ScalA54577.2021.00010"},{"key":"e_1_3_5_5_1","volume-title":"Software, Environments, Tools","author":"Anderson E","year":"1999","unstructured":"Anderson E (ed) (1999) LAPACK users\u2019 guide. In: Software, Environments, Tools. 3rd edition edition. Philadelphia: Society for Industrial and Applied Mathematics.","edition":"3"},{"key":"e_1_3_5_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3480935"},{"key":"e_1_3_5_7_1","unstructured":"Blackford S Dongarra J (1991) LAPACK working note 41 installation guide for LAPACK."},{"key":"e_1_3_5_8_1","unstructured":"Carroll E Gloster A Bustamante MD et al. (2021) A batched GPU methodology for numerical solutions of partial differential equations. arXiv:2107.05395 [physics]."},{"key":"e_1_3_5_9_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2014.07.003"},{"key":"e_1_3_5_10_1","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898718881"},{"key":"e_1_3_5_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-06548-9_1"},{"key":"e_1_3_5_12_1","volume-title":"A Proposed API for Batched Basic Linear Algebra Subprograms","author":"Dongarra J","year":"2016","unstructured":"Dongarra J, Duff I, Gates M, et al. (2016) A Proposed API for Batched Basic Linear Algebra Subprograms. Manchester: The University of Manchester, Vol. 25. Technical Report 2016."},{"key":"e_1_3_5_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2017.05.138"},{"key":"e_1_3_5_14_1","doi-asserted-by":"publisher","DOI":"10.1093\/acprof:oso\/9780198508380.001.0001"},{"key":"e_1_3_5_15_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2019.03.016"},{"key":"e_1_3_5_16_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2016.03.064"},{"key":"e_1_3_5_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1089014.1089020"},{"key":"e_1_3_5_18_1","volume-title":"oneAPI Math Kernel Library","author":"Intel","year":"2023","unstructured":"Intel (2023) oneAPI Math Kernel Library. Santa Clara: Intel Corporation."},{"key":"e_1_3_5_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS53621.2022.00024"},{"key":"e_1_3_5_20_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2023.03.012"},{"key":"e_1_3_5_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3472456.3472484"},{"key":"e_1_3_5_22_1","doi-asserted-by":"publisher","DOI":"10.1088\/0029-5515\/49\/11\/115021"},{"key":"e_1_3_5_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/779359.779361"},{"key":"e_1_3_5_24_1","doi-asserted-by":"publisher","DOI":"10.5445\/IR\/1000165437"},{"key":"e_1_3_5_25_1","doi-asserted-by":"publisher","DOI":"10.5281\/ZENODO.10871244"},{"key":"e_1_3_5_26_1","volume-title":"NVIDIA A100 Tensor Core GPU Architecture","author":"NVIDIA","year":"2020","unstructured":"NVIDIA (2020) NVIDIA A100 Tensor Core GPU Architecture. Santa Clara: NVIDIA Corporation. Technical report."},{"key":"e_1_3_5_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/PDP2018.2018.00123"},{"key":"e_1_3_5_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-71593-9_9"},{"key":"e_1_3_5_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2017.05.145"},{"key":"e_1_3_5_30_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.4909"},{"key":"e_1_3_5_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/SPDP.1991.218237"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/10943420251347460","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/10943420251347460","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/10943420251347460","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/10943420251347460","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:17:45Z","timestamp":1777450665000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/10943420251347460"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,3]]},"references-count":30,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2025,9]]}},"alternative-id":["10.1177\/10943420251347460"],"URL":"https:\/\/doi.org\/10.1177\/10943420251347460","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,3]]}}}