{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T14:53:22Z","timestamp":1781621602849,"version":"3.54.5"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2016,5,10]],"date-time":"2016-05-10T00:00:00Z","timestamp":1462838400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Math. Softw."],"published-print":{"date-parts":[[2016,6,15]]},"abstract":"<jats:p>KBLAS is an open-source, high-performance library that provides optimized kernels for a subset of Level 2 BLAS functionalities on CUDA-enabled GPUs. Since performance of dense matrix-vector multiplication is hindered by the overhead of memory accesses, a double-buffering optimization technique is employed to overlap data motion with computation. After identifying a proper set of tuning parameters, KBLAS efficiently runs on various GPU architectures while avoiding code rewriting and retaining compliance with the standard BLAS API. Another optimization technique allows ensuring coalesced memory access when dealing with submatrices, especially for high-level dense linear algebra algorithms. All KBLAS kernels have been leveraged to a multi-GPU environment, which requires the introduction of new APIs. Considering general matrices, KBLAS is very competitive with existing state-of-the-art kernels and provides a smoother performance across a wide range of matrix dimensions. Considering symmetric and Hermitian matrices, the KBLAS performance outperforms existing state-of-the-art implementations on all matrix sizes and achieves asymptotically up to 50% and 60% speedup against the best competitor on single GPU and multi-GPUs systems, respectively. Performance results also validate our performance model. A subset of KBLAS high-performance kernels have been integrated into NVIDIA's standard BLAS implementation (cuBLAS) for larger dissemination, starting from version 6.0.<\/jats:p>","DOI":"10.1145\/2818311","type":"journal-article","created":{"date-parts":[[2016,5,11]],"date-time":"2016-05-11T12:11:38Z","timestamp":1462968698000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":40,"title":["KBLAS"],"prefix":"10.1145","volume":"42","author":[{"given":"Ahmad","family":"Abdelfattah","sequence":"first","affiliation":[{"name":"Extreme Computing Research Center, KAUST"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"David","family":"Keyes","sequence":"additional","affiliation":[{"name":"Extreme Computing Research Center, KAUST"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hatem","family":"Ltaief","sequence":"additional","affiliation":[{"name":"Extreme Computing Research Center, KAUST"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2016,5,10]]},"reference":[{"key":"e_1_2_1_1_1","series-title":"Lecture Notes in Computer Science","volume-title":"High Performance Computing for Computational Science (VECPAR'12), Michel Dayd, Osni Marques, and Kengo Nakajima (Eds.)","author":"Abdelfattah Ahmad","unstructured":"Ahmad Abdelfattah , Jack Dongarra , David Keyes , and Hatem Ltaief . 2013a. Optimizing memory-bound SYMV kernel on GPU hardware accelerators . In High Performance Computing for Computational Science (VECPAR'12), Michel Dayd, Osni Marques, and Kengo Nakajima (Eds.) . Lecture Notes in Computer Science , Vol. 7851 . Springer , Berlin , 72--79. DOI:http:\/\/dx.doi.org\/10.1007\/978-3-642-38718-0_10 10.1007\/978-3-642-38718-0_10 Ahmad Abdelfattah, Jack Dongarra, David Keyes, and Hatem Ltaief. 2013a. Optimizing memory-bound SYMV kernel on GPU hardware accelerators. In High Performance Computing for Computational Science (VECPAR'12), Michel Dayd, Osni Marques, and Kengo Nakajima (Eds.). Lecture Notes in Computer Science, Vol. 7851. Springer, Berlin, 72--79. DOI:http:\/\/dx.doi.org\/10.1007\/978-3-642-38718-0_10"},{"key":"e_1_2_1_2_1","volume-title":"Euro-Par 2014 Parallel Processing, Fernando Silva","author":"Abdelfattah Ahmad","unstructured":"Ahmad Abdelfattah , Eric Gendron , Damien Gratadour , David Keyes , Hatem Ltaief , Arnaud Sevin , and Fabrice Vidal . 2014. High performance pseudo-analytical simulation of multi-object adaptive optics over multi-GPU systems . In Euro-Par 2014 Parallel Processing, Fernando Silva , I. Dutra, and V. Santos Costa (Eds.). Lecture Notes in Computer Science, Vol. 8632 . Springer International Publishing , 704--715. DOI:http:\/\/dx.doi.org\/10.1007\/978-3-319-09873-9_59 10.1007\/978-3-319-09873-9_59 Ahmad Abdelfattah, Eric Gendron, Damien Gratadour, David Keyes, Hatem Ltaief, Arnaud Sevin, and Fabrice Vidal. 2014. High performance pseudo-analytical simulation of multi-object adaptive optics over multi-GPU systems. In Euro-Par 2014 Parallel Processing, Fernando Silva, I. Dutra, and V. Santos Costa (Eds.). Lecture Notes in Computer Science, Vol. 8632. Springer International Publishing, 704--715. DOI:http:\/\/dx.doi.org\/10.1007\/978-3-319-09873-9_59"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-36949-0_23"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/180\/1\/012037"},{"key":"e_1_2_1_5_1","volume-title":"Sorensen","author":"Anderson E.","year":"1999","unstructured":"E. Anderson , Z. Bai , C. Bischof , Suzan L. Blackford , James W. Demmel , Jack J. Dongarra , J. Du Croz , A. Greenbaum , S. Hammarling , A. McKenney , and Danny C . Sorensen . 1999 . LAPACK User's Guide (3rd ed.). Society for Industrial and Applied Mathematics, Philadelphia . E. Anderson, Z. Bai, C. Bischof, Suzan L. Blackford, James W. Demmel, Jack J. Dongarra, J. Du Croz, A. Greenbaum, S. Hammarling, A. McKenney, and Danny C. Sorensen. 1999. LAPACK User's Guide (3rd ed.). Society for Industrial and Applied Mathematics, Philadelphia."},{"key":"e_1_2_1_6_1","unstructured":"BLAS. 1979. Basic Linear Algebra Subprograms. Retrieved from http:\/\/www.netlib.org\/blas\/.  BLAS. 1979. Basic Linear Algebra Subprograms. Retrieved from http:\/\/www.netlib.org\/blas\/."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1186562.1015800"},{"key":"e_1_2_1_8_1","unstructured":"cuBLAS-XT. 2014. Accelerate BLAS calls with multiple GPUs. Retrieved from https:\/\/developer.nvidia.com\/cublasxt.  cuBLAS-XT. 2014. Accelerate BLAS calls with multiple GPUs. Retrieved from https:\/\/developer.nvidia.com\/cublasxt."},{"key":"e_1_2_1_9_1","volume-title":"CULA: Hybrid GPU accelerated linear algebra routines","author":"Humphrey J. R.","year":"2010","unstructured":"J. R. Humphrey , D. K. Price , K. E. Spagnoli , A. L. Paolini , and E. J. Kelmelis . 2010 . CULA: Hybrid GPU accelerated linear algebra routines . In Society of Photo-Optical Instrumentation Engineers (SPIE) Conference Series , Vol . 7705. 1. J. R. Humphrey, D. K. Price, K. E. Spagnoli, A. L. Paolini, and E. J. Kelmelis. 2010. CULA: Hybrid GPU accelerated linear algebra routines. In Society of Photo-Optical Instrumentation Engineers (SPIE) Conference Series, Vol. 7705. 1."},{"key":"e_1_2_1_10_1","unstructured":"KBLAS. 2014. KAUST Basic Linear Algebra Subprograms. Available at http:\/\/cec.kaust.edu.sa\/Pages\/kblas.aspx. (2014).  KBLAS. 2014. KAUST Basic Linear Algebra Subprograms. Available at http:\/\/cec.kaust.edu.sa\/Pages\/kblas.aspx. (2014)."},{"key":"e_1_2_1_11_1","volume-title":"Hwu","author":"Kirk David B.","year":"2010","unstructured":"David B. Kirk and Wen-mei W . Hwu . 2010 . Programming Massively Parallel Processors: A Hands-on Approach. Morgan Kaufmann Publishers , San Francisco, CA. David B. Kirk and Wen-mei W. Hwu. 2010. Programming Massively Parallel Processors: A Hands-on Approach. Morgan Kaufmann Publishers, San Francisco, CA."},{"key":"e_1_2_1_12_1","volume-title":"Matrix Algebra on GPU and Multicore Architectures. Innovative Computing Laboratory","author":"MAGMA.","unstructured":"MAGMA. 2009. Matrix Algebra on GPU and Multicore Architectures. Innovative Computing Laboratory , University of Tennessee . Retrieved from http:\/\/icl.cs.utk.edu\/magma\/. MAGMA. 2009. Matrix Algebra on GPU and Multicore Architectures. Innovative Computing Laboratory, University of Tennessee. Retrieved from http:\/\/icl.cs.utk.edu\/magma\/."},{"key":"e_1_2_1_14_1","volume-title":"Memory bandwidth and machine balance in current high performance computers","author":"McCalpin John D.","year":"1995","unstructured":"John D. McCalpin . 1995. Memory bandwidth and machine balance in current high performance computers . IEEE Computer Society Technical Committee on Computer Architecture (TCCA) Newsletter ( Dec. 1995 ), 19--25. John D. McCalpin. 1995. Memory bandwidth and machine balance in current high performance computers. IEEE Computer Society Technical Committee on Computer Architecture (TCCA) Newsletter (Dec. 1995), 19--25."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2063384.2063392"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342010385729"},{"key":"e_1_2_1_17_1","volume-title":"BLAS for GPUs","author":"Nath Rajib","unstructured":"Rajib Nath , Stanimire Tomov , and Jack Dongarra . 2010b. BLAS for GPUs . CRC Press , 57--80. DOI:http:\/\/dx.doi.org\/doi:10.1201\/b10376-6 10.1201\/b10376-6 Rajib Nath, Stanimire Tomov, and Jack Dongarra. 2010b. BLAS for GPUs. CRC Press, 57--80. DOI:http:\/\/dx.doi.org\/doi:10.1201\/b10376-6"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/1964238.1964250"},{"key":"e_1_2_1_19_1","unstructured":"NVIDIA. 2009. NVIDIA Fermi Compute Architecture Whitepaper. Retrieved from http:\/\/www.nvidia.com\/content\/PDF\/fermi_white_papers\/NVIDIA_Fermi_Compute_Architecture_Whitepaper.pdf.  NVIDIA. 2009. NVIDIA Fermi Compute Architecture Whitepaper. Retrieved from http:\/\/www.nvidia.com\/content\/PDF\/fermi_white_papers\/NVIDIA_Fermi_Compute_Architecture_Whitepaper.pdf."},{"key":"e_1_2_1_20_1","unstructured":"NVIDIA. 2012. NVIDIA Kepler GK110 Architecture Whitepaper. Retrieved from http:\/\/www.nvidia.com\/content\/PDF\/kepler\/NVIDIA-Kepler-GK110-Architecture-Whitepaper.pdf.  NVIDIA. 2012. NVIDIA Kepler GK110 Architecture Whitepaper. Retrieved from http:\/\/www.nvidia.com\/content\/PDF\/kepler\/NVIDIA-Kepler-GK110-Architecture-Whitepaper.pdf."},{"key":"e_1_2_1_21_1","unstructured":"NVIDIA. 2014a. CUDA C Programming Guide. Retrieved from http:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/.  NVIDIA. 2014a. CUDA C Programming Guide. Retrieved from http:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/."},{"key":"e_1_2_1_22_1","unstructured":"NVIDIA. 2014b. The NVIDIA CUDA Basic Linear Algebra Subroutines. Retrieved from https:\/\/developer.nvidia.com\/cublas\/.  NVIDIA. 2014b. The NVIDIA CUDA Basic Linear Algebra Subroutines. Retrieved from https:\/\/developer.nvidia.com\/cublas\/."},{"key":"e_1_2_1_23_1","unstructured":"NVIDIA. 2014c. cuBLAS::CUDA Toolkit Documentation. http:\/\/docs.nvidia.com\/cuda\/cublas\/#appendix-acknowledgements.  NVIDIA. 2014c. cuBLAS::CUDA Toolkit Documentation. http:\/\/docs.nvidia.com\/cuda\/cublas\/#appendix-acknowledgements."},{"key":"e_1_2_1_24_1","unstructured":"OPENACC. 2011. Directives for Accelerators. Retrieved from http:\/\/www.openacc-standard.org\/.  OPENACC. 2011. Directives for Accelerators. Retrieved from http:\/\/www.openacc-standard.org\/."},{"key":"e_1_2_1_25_1","unstructured":"OPENCL. 2009. The open standard for parallel programming of heterogeneous systems. Retrieved from http:\/\/www.khronos.org\/opencl\/.  OPENCL. 2009. The open standard for parallel programming of heterogeneous systems. Retrieved from http:\/\/www.khronos.org\/opencl\/."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2063384.2063431"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2010.06.001"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.5555\/1413370.1413402"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.3152"}],"container-title":["ACM Transactions on Mathematical Software"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2818311","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2818311","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T05:07:25Z","timestamp":1750223245000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2818311"}},"subtitle":["An Optimized Library for Dense Matrix-Vector Multiplication on GPU Accelerators"],"short-title":[],"issued":{"date-parts":[[2016,5,10]]},"references-count":29,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2016,6,15]]}},"alternative-id":["10.1145\/2818311"],"URL":"https:\/\/doi.org\/10.1145\/2818311","relation":{},"ISSN":["0098-3500","1557-7295"],"issn-type":[{"value":"0098-3500","type":"print"},{"value":"1557-7295","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,5,10]]},"assertion":[{"value":"2014-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-08-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-05-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}