{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,29]],"date-time":"2025-09-29T08:13:39Z","timestamp":1759133619401,"version":"3.41.0"},"reference-count":50,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2019,5,3]],"date-time":"2019-05-03T00:00:00Z","timestamp":1556841600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Math. Softw."],"published-print":{"date-parts":[[2019,6,30]]},"abstract":"<jats:p>Batched dense linear algebra kernels are becoming ubiquitous in scientific applications, ranging from tensor contractions in deep learning to data compression in hierarchical low-rank matrix approximation. Within a single API call, these kernels are capable of simultaneously launching up to thousands of similar matrix computations, removing the expensive overhead of multiple API calls while increasing the occupancy of the underlying hardware. A challenge is that for the existing hardware landscape (x86, GPUs, etc.), only a subset of the required batched operations is implemented by the vendors, with limited support for very small problem sizes. We describe the design and performance of a new class of batched triangular dense linear algebra kernels on very small data sizes (up to 256) using single and multiple GPUs. By deploying recursive formulations, stressing the register usage, maintaining data locality, reducing threads synchronization, and fusing successive kernel calls, the new batched kernels outperform existing state-of-the-art implementations.<\/jats:p>","DOI":"10.1145\/3267101","type":"journal-article","created":{"date-parts":[[2019,5,6]],"date-time":"2019-05-06T17:24:38Z","timestamp":1557163478000},"page":"1-28","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Batched Triangular Dense Linear Algebra Kernels for Very Small Matrix Sizes on GPUs"],"prefix":"10.1145","volume":"45","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9509-7794","authenticated-orcid":false,"given":"Ali","family":"Charara","sequence":"first","affiliation":[{"name":"Extreme Computing Research Center, King Abdullah University of Science and Technology, Thuwal, Kingdom of Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David","family":"Keyes","sequence":"additional","affiliation":[{"name":"Extreme Computing Research Center, King Abdullah University of Science and Technology, Thuwal, Kingdom of Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hatem","family":"Ltaief","sequence":"additional","affiliation":[{"name":"Extreme Computing Research Center, King Abdullah University of Science and Technology, Thuwal, Kingdom of Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,5,3]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2016.05.302"},{"volume-title":"Proceedings of the 10th International Conference on High Performance Computing for Computational Science (VECPAR\u201912)","author":"Abdelfattah A.","key":"e_1_2_1_2_1"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jocs.2016.12.009"},{"volume-title":"Proceedings of the 2016 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW\u201916)","author":"Abdelfattah A.","key":"e_1_2_1_4_1"},{"volume":"9697","volume-title":"Proceedings of the 31st International Conference on High Performance Computing, Julian M. Kunkel, Pavan Balaji, and Jack Dongarra (Eds.). Lecture Notes in Computer Science","author":"Abdelfattah A.","key":"e_1_2_1_5_1"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2016.05.303"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2818311"},{"volume-title":"Proceedings of the International Supercomputing Conference.","author":"Akbudak K.","key":"e_1_2_1_8_1"},{"key":"e_1_2_1_9_1","doi-asserted-by":"crossref","unstructured":"P. Amestoy A. Buttari I. Duff A. Guermouche J. Y. L\u2019Excellent and B. U\u00e7ar. 2011. MUMPS. Springer US Boston MA 1232--1238.  P. Amestoy A. Buttari I. Duff A. Guermouche J. Y. L\u2019Excellent and B. U\u00e7ar. 2011. MUMPS. Springer US Boston MA 1232--1238.","DOI":"10.1007\/978-0-387-09766-4_204"},{"volume-title":"Proceedings of the Applied Parallel Computing: 5th International Workshop on New Paradigms for HPC in Industry and Academia, T. S\u00f8revik, F. Manne, A. H. Gebremedhin, and Randi Moe (Eds.). Springer","author":"Andersen B. S.","key":"e_1_2_1_10_1"},{"key":"e_1_2_1_11_1","doi-asserted-by":"crossref","unstructured":"E. Anderson Z. Bai C. Bischof L. S. Blackford J. Demmel J. Dongarra J. Du Croz A. Greenbaum S. Hammarling A. McKenney and D. Sorensen. 1999. LAPACK User\u2019s Guide (3rd ed.). Society for Industrial and Applied Mathematics Philadelphia.   E. Anderson Z. Bai C. Bischof L. S. Blackford J. Demmel J. Dongarra J. Du Croz A. Greenbaum S. Hammarling A. McKenney and D. Sorensen. 1999. LAPACK User\u2019s Guide (3rd ed.). Society for Industrial and Applied Mathematics Philadelphia.","DOI":"10.1137\/1.9780898719604"},{"volume-title":"Bench-testing Environment for Automated Software Tuning. Innovative Computing Laboratory","author":"ST.","key":"e_1_2_1_12_1"},{"key":"e_1_2_1_13_1","doi-asserted-by":"crossref","unstructured":"W. H. Boukaram G. Turkiyyah H. Ltaief and D. Keyes. 2017. Batched QR and SVD algorithms on GPUs with applications in hierarchical matrix compression. (unpublished).  W. H. Boukaram G. Turkiyyah H. Ltaief and D. Keyes. 2017. Batched QR and SVD algorithms on GPUs with applications in hierarchical matrix compression. (unpublished).","DOI":"10.1016\/j.parco.2017.09.001"},{"key":"e_1_2_1_14_1","unstructured":"A. Charara D. Keyes and H. Ltaief. 2016a. A framework for dense triangular matrix kernels on various manycore architectures. (unpublished). Retrieved from http:\/\/hdl.handle.net\/10754\/622077.  A. Charara D. Keyes and H. Ltaief. 2016a. A framework for dense triangular matrix kernels on various manycore architectures. (unpublished). Retrieved from http:\/\/hdl.handle.net\/10754\/622077."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-43659-3_35"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1137\/S1064827502412887"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCC.2014.30"},{"volume-title":"Proceedings of the 2014 43rd International Conference on Parallel Processing. IEEE Computer Society, 432--440","author":"Dong T.","key":"e_1_2_1_19_1"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.14529\/jsfi150405"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/42288.42291"},{"volume-title":"FRPA: A Framework for Recursive Parallel Algorithms. Master\u2019s thesis. EECS Department","year":"2015","author":"Eliahu D.","key":"e_1_2_1_22_1"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1137\/S0036144503428693"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1147\/rd.444.0605"},{"volume-title":"Proceedings of the 2014 16th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing. 479--486","author":"Falch T. L.","key":"e_1_2_1_25_1"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/1377603.1377607"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/s006070050015"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.5555\/333825.333827"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342014567546"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2716282.2716288"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2688500.2688534"},{"volume-title":"Proceedings of the 30th International Conference on High Performance Computing. Springer, 31--47","author":"Haidar A.","key":"e_1_2_1_32_1"},{"volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201916)","author":"Heinecke A.","key":"e_1_2_1_33_1"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-8191(01)00141-7"},{"volume-title":"AIP Conf. Proc. 1504","year":"2012","author":"Igual F. D.","key":"e_1_2_1_35_1"},{"key":"e_1_2_1_36_1","unstructured":"Intel. 2017. Math Kernel Library (MKL). Retrieved from http:\/\/software.intel.com\/en-us\/articles\/intel-mkl.  Intel. 2017. Math Kernel Library (MKL). Retrieved from http:\/\/software.intel.com\/en-us\/articles\/intel-mkl."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/11558958_3"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/292395.292412"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10915-013-9805-x"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2015.2481890"},{"volume-title":"Matrix Algebra on GPU and Multicore Architectures. Innovative Computing Laboratory","author":"MA.","key":"e_1_2_1_41_1"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-43659-3_48"},{"key":"e_1_2_1_43_1","unstructured":"NVIDIA. 2017a. The CUDA Basic Linear Algebra Subroutines (cuBLAS). Retrieved from http:\/\/developer.nvidia.com\/cublas.  NVIDIA. 2017a. The CUDA Basic Linear Algebra Subroutines (cuBLAS). Retrieved from http:\/\/developer.nvidia.com\/cublas."},{"key":"e_1_2_1_44_1","unstructured":"NVIDIA. 2017b. The CUDA Basic Linear Algebra Subroutines (cuBLAS). Retrieved from http:\/\/docs.nvidia.com\/cuda\/cublas\/index.html#appendix-acknowledgements.  NVIDIA. 2017b. The CUDA Basic Linear Algebra Subroutines (cuBLAS). Retrieved from http:\/\/docs.nvidia.com\/cuda\/cublas\/index.html#appendix-acknowledgements."},{"key":"e_1_2_1_45_1","unstructured":"NVIDIA. 2017c. The CUDA Solver Library (cuSOLVER). Retrieved from http:\/\/developer.nvidia.com\/cusolver.  NVIDIA. 2017c. The CUDA Solver Library (cuSOLVER). Retrieved from http:\/\/developer.nvidia.com\/cusolver."},{"key":"e_1_2_1_46_1","unstructured":"NVIDIA. 2017d. CUDA C Programming Guide. Retrieved from http:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/.  NVIDIA. 2017d. CUDA C Programming Guide. Retrieved from http:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/."},{"volume-title":"Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS\u201914)","author":"Ofenbeck G.","key":"e_1_2_1_47_1"},{"key":"e_1_2_1_48_1","unstructured":"V.\n      Oreste M.\n      Fatica N. A.\n      Gawande and \n      A.\n      Tumeo\n  . \n  2013\n  . Power\/performance trade-offs of small batched LU based solvers on GPUs. In Proceedings of the 19th International Conference on \n  Parallel Processing (Euro-Par\u2019\n  13) Lecture Notes in Computer Science Vol. \n  8097\n  .  V. Oreste M. Fatica N. A. Gawande and A. Tumeo. 2013. Power\/performance trade-offs of small batched LU based solvers on GPUs. In Proceedings of the 19th International Conference on Parallel Processing (Euro-Par\u201913) Lecture Notes in Computer Science Vol. 8097."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3061664"},{"key":"e_1_2_1_50_1","unstructured":"SuiteSparse. 2017. A suite of sparse matrix software. Retrieved from http:\/\/faculty.cse.tamu.edu\/davis\/SuiteSparse\/.  SuiteSparse. 2017. A suite of sparse matrix software. Retrieved from http:\/\/faculty.cse.tamu.edu\/davis\/SuiteSparse\/."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1137\/15M1025785"}],"container-title":["ACM Transactions on Mathematical Software"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3267101","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3267101","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T20:26:52Z","timestamp":1750278412000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3267101"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,5,3]]},"references-count":50,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2019,6,30]]}},"alternative-id":["10.1145\/3267101"],"URL":"https:\/\/doi.org\/10.1145\/3267101","relation":{},"ISSN":["0098-3500","1557-7295"],"issn-type":[{"type":"print","value":"0098-3500"},{"type":"electronic","value":"1557-7295"}],"subject":[],"published":{"date-parts":[[2019,5,3]]},"assertion":[{"value":"2017-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-08-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-05-03","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}