{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T11:59:20Z","timestamp":1787313560885,"version":"3.56.0"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2008,5,1]],"date-time":"2008-05-01T00:00:00Z","timestamp":1209600000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["EIA-0103642EIA-0079710"],"award-info":[{"award-number":["EIA-0103642EIA-0079710"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Pittsburgh Supercomputing Center","award":["ASC50006P"],"award-info":[{"award-number":["ASC50006P"]}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["1 P20 HG003900-01"],"award-info":[{"award-number":["1 P20 HG003900-01"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Math. Softw."],"published-print":{"date-parts":[[2008,5]]},"abstract":"<jats:p>\n                    On cache based computer architectures using current standard algorithms, Householder bidiagonalization requires a significant portion of the execution time for computing matrix singular values and vectors. In this paper we reorganize the sequence of operations for Householder bidiagonalization of a general\n                    <jats:italic>m<\/jats:italic>\n                    \u00d7\n                    <jats:italic>n<\/jats:italic>\n                    matrix, so that two (_GEMV) vector-matrix multiplications can be done with one pass of the unreduced trailing part of the matrix through cache. Two new BLAS operations approximately cut in half the transfer of data from main memory to cache, reducing execution times by up to 25 per cent. We give detailed algorithm descriptions and compare timings with the current LAPACK bidiagonalization algorithm.\n                  <\/jats:p>","DOI":"10.1145\/1356052.1356055","type":"journal-article","created":{"date-parts":[[2008,5,15]],"date-time":"2008-05-15T14:28:05Z","timestamp":1210861685000},"page":"1-33","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":20,"title":["Cache efficient bidiagonalization using BLAS 2.5 operators"],"prefix":"10.1145","volume":"34","author":[{"given":"Gary W.","family":"Howell","sequence":"first","affiliation":[{"name":"North Carolina State University, Raleigh, NC"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"James W.","family":"Demmel","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, CA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Charles T.","family":"Fulton","sequence":"additional","affiliation":[{"name":"Florida Institute of Technology, Melbourne, FL"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sven","family":"Hammarling","sequence":"additional","affiliation":[{"name":"University of Manchester, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Karen","family":"Marmol","sequence":"additional","affiliation":[{"name":"Harris Corporation, Melbourne, FL"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2008,5,16]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Anderson E. Bai Z. Bischof C. Blackford S. Demmel J. Dongarra J. Greenbaum A. Hammarling S. Mckenney A. and Sorensen D. 1999. LAPACK User's Guide 3rd. Ed. SIAM Philadelphia PA.   Anderson E. Bai Z. Bischof C. Blackford S. Demmel J. Dongarra J. Greenbaum A. Hammarling S. Mckenney A. and Sorensen D. 1999. LAPACK User's Guide 3rd. Ed. SIAM Philadelphia PA.","DOI":"10.1137\/1.9780898719604"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.laa.2004.09.019"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1177\/109434209200600103"},{"key":"e_1_2_1_4_1","unstructured":"Berry M. Do T. O'brien G. Krishna V. and Varadhan S. 1993. SVDPACKC: version 1.0 user's guide Tech. rep. CS-93-194 University of Tennessee Knoxville TN.   Berry M. Do T. O'brien G. Krishna V. and Varadhan S. 1993. SVDPACKC: version 1.0 user's guide Tech. rep. CS-93-194 University of Tennessee Knoxville TN."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1137\/1037127"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1137\/0908009"},{"key":"e_1_2_1_7_1","unstructured":"Blackford S. and Dongarra J. 1999. Installation guide for LAPACK LAPACK Working Note 41.  Blackford S. and Dongarra J. 1999. Installation guide for LAPACK LAPACK Working Note 41."},{"key":"e_1_2_1_8_1","first-page":"1","article-title":"Basic linear algebra subprograms technical (BLAST) forum standard","volume":"12","author":"Blackford L. S.","year":"2002","unstructured":"Blackford , L. S. , Corliss , G. , Demmel , J. , Dongarra , J. , Duff , I. , Hammarling , S. , Henry , G. , Heroux , M. , Hu , C. , Kahan , W. , Kaufmann , L. , Kearfott , B. , Frogh , F. , Li , X. , Maany , Z. , Petitet , A. , Pozo , R. , Remington , K. , Walster , W. , Whaley , C. , Wolff , V. , Gudenberg , J. , and Lumsdaine , A. 2002 . Basic linear algebra subprograms technical (BLAST) forum standard , Int. J. High Perform. Comput. 12 , 1 -- 2 (www.netlib.org\/blas\/blast-forum). Blackford, L. S., Corliss, G., Demmel, J., Dongarra, J., Duff, I., Hammarling, S., Henry, G., Heroux, M., Hu, C., Kahan, W., Kaufmann, L., Kearfott, B., Frogh, F., Li, X., Maany, Z., Petitet, A., Pozo, R., Remington, K., Walster, W., Whaley, C., Wolff, V., Gudenberg, J., and Lumsdaine, A. 2002. Basic linear algebra subprograms technical (BLAST) forum standard, Int. J. High Perform. Comput. 12, 1--2 (www.netlib.org\/blas\/blast-forum).","journal-title":"Int. J. High Perform. Comput."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/567806.567807"},{"key":"e_1_2_1_10_1","unstructured":"Bosner N. and Barlow J. L. 2005. Block and parallel versions of one-sided bidiagonalization. preprint http:\/\/www.cse.psu.edu\/barlow\/block_bidiag.pdf.  Bosner N. and Barlow J. L. 2005. Block and parallel versions of one-sided bidiagonalization. preprint http:\/\/www.cse.psu.edu\/barlow\/block_bidiag.pdf."},{"key":"e_1_2_1_11_1","first-page":"379","article-title":"The design of a parallel dense linear algebra software library: reduction to Hessenberg, tridiagonal, and bidiagonal form. (LAPACK Working Note &num; 92) Num","volume":"10","author":"Choi J.","year":"1995","unstructured":"Choi , J. , Dongarra , J. , and Walker , D. 1995 . The design of a parallel dense linear algebra software library: reduction to Hessenberg, tridiagonal, and bidiagonal form. (LAPACK Working Note &num; 92) Num . Alg. , 10 , 379 -- 399 . Choi, J., Dongarra, J., and Walker, D. 1995. The design of a parallel dense linear algebra software library: reduction to Hessenberg, tridiagonal, and bidiagonal form. (LAPACK Working Note &num; 92) Num. Alg., 10, 379--399.","journal-title":"Alg."},{"key":"e_1_2_1_12_1","unstructured":"Dhillon I. S. 1997. A New O(n2) Algorithm for the symmetric tridiagonal eigenvalue\/eigenvector problem. PhD thesis University of California Berkeley CA.   Dhillon I. S. 1997. A New O(n 2 ) Algorithm for the symmetric tridiagonal eigenvalue\/eigenvector problem. PhD thesis University of California Berkeley CA."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/0377-0427(89)90367-1"},{"key":"e_1_2_1_14_1","volume-title":"Numerical Linear Algebra for High-Performance Computers","author":"Dongarra J.","unstructured":"Dongarra , J. , Duff , I. , Sorensen , D. , and Van Der Vorst , H. 1988. Numerical Linear Algebra for High-Performance Computers . SIAM , Philadelphia, PA . Dongarra, J., Duff, I., Sorensen, D., and Van Der Vorst, H. 1988. Numerical Linear Algebra for High-Performance Computers. SIAM, Philadelphia, PA."},{"key":"e_1_2_1_15_1","first-page":"5","article-title":"Portable memory heirarchy techniques for PDE solvers","volume":"33","author":"Douglas C. C.","year":"2000","unstructured":"Douglas , C. C. , Haase , G. , Hu , J. , Kowarschik , M. , R\u00fcde , U. , and Weiss , C. 2000 . Portable memory heirarchy techniques for PDE solvers : Part I. SIAM News , 33 , 5 . Douglas, C. C., Haase, G., Hu, J., Kowarschik, M., R\u00fcde, U., and Weiss, C. 2000. Portable memory heirarchy techniques for PDE solvers: Part I. SIAM News, 33, 5.","journal-title":"Part I. SIAM News"},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Fernando V. Parlett B. and Dhillon I. 1995. A way to find the most redundant equation in a tridiagonal system. Mathematics Department University of California Berkeley.  Fernando V. Parlett B. and Dhillon I. 1995. A way to find the most redundant equation in a tridiagonal system. Mathematics Department University of California Berkeley.","DOI":"10.21236\/ADA310613"},{"key":"e_1_2_1_17_1","doi-asserted-by":"crossref","unstructured":"Goedecker S. and Hoise A. 2001. Performance Optimization for Numerically Intensive Codes SIAM Philadelphia PA.   Goedecker S. and Hoise A. 2001. Performance Optimization for Numerically Intensive Codes SIAM Philadelphia PA.","DOI":"10.1137\/1.9780898718218"},{"key":"e_1_2_1_18_1","first-page":"205","article-title":"Calculating the singular Values and pseudo-inverse of a matrix","volume":"2","author":"Golub G.","year":"1965","unstructured":"Golub , G. and Kahan , W. 1965 . Calculating the singular Values and pseudo-inverse of a matrix , SIAM J. Num. Anal. , 2 , 205 -- 224 . Golub, G. and Kahan, W. 1965. Calculating the singular Values and pseudo-inverse of a matrix, SIAM J. Num. Anal., 2, 205--224.","journal-title":"SIAM J. Num. Anal."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF02163027"},{"key":"e_1_2_1_20_1","volume-title":"Matrix Computations","author":"Golub G.","unstructured":"Golub , G. and Van Loan , C. F. 1996. Matrix Computations 3 rd Ed. Johns Hopkins University Press , Baltimore, MA . Golub, G. and Van Loan, C. F. 1996. Matrix Computations 3rd Ed. Johns Hopkins University Press, Baltimore, MA.","edition":"3"},{"key":"e_1_2_1_21_1","unstructured":"Gr\u00f6sser B. and Lang B. 1988. Efficient parallel reduction to bidiagonal form Preprint BUGHW-SC 98\/2. http:\/\/www.math.uni-wuppertal\/org\/SciComp\/Preprint\/SC9802ips.gz.  Gr\u00f6sser B. and Lang B. 1988. Efficient parallel reduction to bidiagonal form Preprint BUGHW-SC 98\/2. http:\/\/www.math.uni-wuppertal\/org\/SciComp\/Preprint\/SC9802ips.gz."},{"key":"e_1_2_1_22_1","unstructured":"Howell G. W. 2001. Sparse Householder bidiagonalization. CERFACS Sparse Days. http:\/\/ncsu. edu\/itd\/hpc\/Documents\/Publications\/gary_howell\/cerfacs01.ps  Howell G. W. 2001. Sparse Householder bidiagonalization. CERFACS Sparse Days. http:\/\/ncsu. edu\/itd\/hpc\/Documents\/Publications\/gary_howell\/cerfacs01.ps"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1016\/0167-8191(95)00064-X"},{"key":"e_1_2_1_24_1","volume-title":"A Matlab script for 2-blocking to speed Ralha-Barlow one-sided bidiagonalization. Summer intern project at ERDC","author":"Owens B.","unstructured":"Owens , B. 2003. A Matlab script for 2-blocking to speed Ralha-Barlow one-sided bidiagonalization. Summer intern project at ERDC , MSRC , Vicksburg, MS . http:\/\/ncsu.edu\/itd\/hpc\/Documents\/Publications\/gary_howell\/barlow3.m. Owens, B. 2003. A Matlab script for 2-blocking to speed Ralha-Barlow one-sided bidiagonalization. Summer intern project at ERDC, MSRC, Vicksburg, MS. http:\/\/ncsu.edu\/itd\/hpc\/Documents\/Publications\/gary_howell\/barlow3.m."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/355984.355989"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0024-3795(97)80053-5"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0024-3795(01)00569-9"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1137\/0910005"},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the 9th SIAM Conference on Parallel Processing for Scientific Computing.","author":"Whaley C.","unstructured":"Whaley , C. and Dongarra , J . 1999. Automatically tuned linear algebra in software . In Proceedings of the 9th SIAM Conference on Parallel Processing for Scientific Computing. Whaley, C. and Dongarra, J. 1999. Automatically tuned linear algebra in software. In Proceedings of the 9th SIAM Conference on Parallel Processing for Scientific Computing."}],"container-title":["ACM Transactions on Mathematical Software"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1356052.1356055","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1356052.1356055","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T09:39:09Z","timestamp":1750239549000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1356052.1356055"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,5]]},"references-count":29,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2008,5]]}},"alternative-id":["10.1145\/1356052.1356055"],"URL":"https:\/\/doi.org\/10.1145\/1356052.1356055","relation":{},"ISSN":["0098-3500","1557-7295"],"issn-type":[{"value":"0098-3500","type":"print"},{"value":"1557-7295","type":"electronic"}],"subject":[],"published":{"date-parts":[[2008,5]]},"assertion":[{"value":"2006-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2007-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2008-05-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}