{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T22:46:59Z","timestamp":1777675619622,"version":"3.51.4"},"reference-count":21,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2020,10,9]],"date-time":"2020-10-09T00:00:00Z","timestamp":1602201600000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2021,1]]},"abstract":"<jats:p>General matrix-matrix multiplications with double-precision real and complex entries (DGEMM and ZGEMM) in vendor-supplied BLAS libraries are best optimized for square matrices but often show bad performance for tall &amp; skinny matrices, which are much taller than wide. NVIDIA\u2019s current CUBLAS implementation delivers only a fraction of the potential performance as indicated by the roofline model in this case. We describe the challenges and key characteristics of an implementation that can achieve close to optimal performance. We further evaluate different strategies of parallelization and thread distribution and devise a flexible, configurable mapping scheme. To ensure flexibility and allow for highly tailored implementations we use code generation combined with autotuning. For a large range of matrix sizes in the domain of interest we achieve at least 2\/3 of the roofline performance and often substantially outperform state-of-the art CUBLAS results on an NVIDIA Volta GPGPU.<\/jats:p>","DOI":"10.1177\/1094342020965661","type":"journal-article","created":{"date-parts":[[2020,10,9]],"date-time":"2020-10-09T06:44:38Z","timestamp":1602225878000},"page":"5-19","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":12,"title":["Performance engineering for real and complex tall &amp; skinny matrix multiplication kernels on GPUs"],"prefix":"10.1177","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3547-0611","authenticated-orcid":false,"given":"Dominik","family":"Ernst","sequence":"first","affiliation":[{"name":"Erlangen Regional Computing Center (RRZE), Friedrich-Alexander-Universit\u00e4t Erlangen-N\u00fcrnberg, Erlangen, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Georg","family":"Hager","sequence":"additional","affiliation":[{"name":"Erlangen Regional Computing Center (RRZE), Friedrich-Alexander-Universit\u00e4t Erlangen-N\u00fcrnberg, Erlangen, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jonas","family":"Thies","sequence":"additional","affiliation":[{"name":"German Aerospace Center (DLR), Simulation and Software Technology, K\u00f6ln, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gerhard","family":"Wellein","sequence":"additional","affiliation":[{"name":"Erlangen Regional Computing Center (RRZE), Friedrich-Alexander-Universit\u00e4t Erlangen-N\u00fcrnberg, Erlangen, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2020,10,9]]},"reference":[{"key":"bibr1-1094342020965661","doi-asserted-by":"publisher","DOI":"10.1145\/2858788.2688513"},{"key":"bibr2-1094342020965661","doi-asserted-by":"publisher","DOI":"10.1145\/3330345.3330355"},{"key":"bibr3-1094342020965661","doi-asserted-by":"publisher","DOI":"10.1109\/CDC.1974.270490"},{"key":"bibr4-1094342020965661","unstructured":"Ernst D (2019) CUDA microbenchmarks. Available at: http:\/\/tiny.cc\/cudabench (accessed 2 February 2020)."},{"key":"bibr5-1094342020965661","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-43229-4_43"},{"key":"bibr6-1094342020965661","first-page":"233","volume-title":"Proceedings of parallel CFD \u201899","author":"Gropp WD","year":"1999"},{"key":"bibr7-1094342020965661","unstructured":"Harris M (2013) CUDA pro tip: write flexible kernels with grid-stride loops. Available at: devblogs.nvidia.com\/cuda-pro-tip-write-flexible-kernels-grid-stride-loops\/ (accessed 2 February 2020)."},{"key":"bibr8-1094342020965661","doi-asserted-by":"publisher","DOI":"10.1007\/11751649_84"},{"key":"bibr9-1094342020965661","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2017.56"},{"key":"bibr10-1094342020965661","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2016.58"},{"key":"bibr11-1094342020965661","doi-asserted-by":"publisher","DOI":"10.1145\/3372419"},{"key":"bibr12-1094342020965661","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-92040-5_17"},{"key":"bibr13-1094342020965661","first-page":"1","author":"Kreutzer M","year":"2016","journal-title":"International Journal of Parallel Programming"},{"key":"bibr14-1094342020965661","first-page":"19","volume":"2","author":"McCalpin JD","year":"1995","journal-title":"IEEE Computer Society Technical Committee on Computer Architecture (TCCA) Newsletter"},{"key":"bibr15-1094342020965661","unstructured":"NVIDIA (2019a) CUBLAS reference. Available at: https:\/\/docs.nvidia.com\/cuda\/cublas (accessed 5 May 2019)."},{"key":"bibr16-1094342020965661","unstructured":"NVIDIA (2019b) CUTLASS. Available at: https:\/\/github.com\/NVIDIA\/cutlass (accessed 5 May 2019)."},{"key":"bibr17-1094342020965661","doi-asserted-by":"publisher","DOI":"10.1016\/0024-3795(80)90247-5"},{"key":"bibr18-1094342020965661","doi-asserted-by":"publisher","DOI":"10.1137\/140976017"},{"key":"bibr19-1094342020965661","doi-asserted-by":"publisher","DOI":"10.1007\/BF02165411"},{"key":"bibr20-1094342020965661","doi-asserted-by":"crossref","unstructured":"Thies J, R\u00f6hrig-Z\u00f6llner M, Overmars N, et al. (2019) PHIST: a pipelined, hybrid-parallel iterative solver toolkit. Accepted for ACM Transactions on Mathematical Software. Available at: https:\/\/elib.dlr.de\/123323\/.","DOI":"10.1145\/3402227"},{"key":"bibr21-1094342020965661","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342020965661","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342020965661","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342020965661","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:16:01Z","timestamp":1777450561000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342020965661"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,9]]},"references-count":21,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,1]]}},"alternative-id":["10.1177\/1094342020965661"],"URL":"https:\/\/doi.org\/10.1177\/1094342020965661","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,10,9]]}}}