{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T04:33:13Z","timestamp":1759206793665,"version":"3.41.0"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2012,8,1]],"date-time":"2012-08-01T00:00:00Z","timestamp":1343779200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"CICYT","award":["TIN2008-06570-C04-01"],"award-info":[{"award-number":["TIN2008-06570-C04-01"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Math. Softw."],"published-print":{"date-parts":[[2012,8]]},"abstract":"<jats:p>Out-of-core implementations of algorithms for dense matrix computations have traditionally focused on optimal use of memory so as to minimize I\/O, often trading programmability for performance. In this article we show how the current state of hardware and software allows the programmability problem to be addressed without sacrificing performance. This comes from the realizations that memory is cheap and large, making it less necessary to optimally orchestrate I\/O, and that new algorithms view matrices as collections of submatrices and computation as operations with those submatrices. This enables libraries to be coded at a high level of abstraction, leaving the tasks of scheduling the computations and data movement in the hands of a runtime system. This is in sharp contrast to more traditional approaches that leverage optimal use of in-core memory and, at the expense of introducing considerable programming complexity, explicit overlap of I\/O with computation. Performance is demonstrated for this approach on multicore architectures as well as platforms equipped with hardware accelerators.<\/jats:p>","DOI":"10.1145\/2331130.2331133","type":"journal-article","created":{"date-parts":[[2012,9,4]],"date-time":"2012-09-04T12:50:47Z","timestamp":1346763047000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["A Runtime System for Programming Out-of-Core Matrix Algorithms-by-Tiles on Multithreaded Architectures"],"prefix":"10.1145","volume":"38","author":[{"given":"Gregorio","family":"Quintana-Ort\u00ed","sequence":"first","affiliation":[{"name":"Universidad Jaume I"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Francisco D.","family":"Igual","sequence":"additional","affiliation":[{"name":"Universidad Jaume I"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mercedes","family":"Marqu\u00e9s","sequence":"additional","affiliation":[{"name":"Universidad Jaume I"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Enrique S.","family":"Quintana-Ort\u00ed","sequence":"additional","affiliation":[{"name":"Universidad Jaume I"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Robert A.","family":"van de Geijn","sequence":"additional","affiliation":[{"name":"The University of Texas at Austin"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2012,8]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/180\/1\/012037"},{"volume-title":"Proceedings of the 10th IEEE Workshop on Parallel and Distributed Scientific and Engineering Computing.","author":"Barrachina S.","key":"e_1_2_1_3_1","unstructured":"Barrachina , S. , Castillo , M. , Igual , F. D. , Mayo , R. , and Quintana-Ort\u00ed , E. S . 2008. Evaluation and tuning of the level 3 CUBLAS for graphics processors . In Proceedings of the 10th IEEE Workshop on Parallel and Distributed Scientific and Engineering Computing. Barrachina, S., Castillo, M., Igual, F. D., Mayo, R., and Quintana-Ort\u00ed, E. S. 2008. Evaluation and tuning of the level 3 CUBLAS for graphics processors. In Proceedings of the 10th IEEE Workshop on Parallel and Distributed Scientific and Engineering Computing."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.v21:18"},{"key":"e_1_2_1_5_1","unstructured":"Bashe C. J. Johnson L. R. Palmer J. H. and Pugh E. W. 1985. IBM\u2019s Early Computers. MIT Press. Bashe C. J. Johnson L. R. Palmer J. H. and Pugh E. W. 1985. IBM\u2019s Early Computers . MIT Press."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1137\/06067256X"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1055531.1055532"},{"key":"e_1_2_1_8_1","unstructured":"Buttari A. Langou J. Kurzak J. and Dongarra J. 2007a. A class of parallel tiled linear algebra algorithms for multicore architectures. LAPACK Working Note 191 UT-CS-07-600 University of Tennessee. Buttari A. Langou J. Kurzak J. and Dongarra J. 2007a. A class of parallel tiled linear algebra algorithms for multicore architectures. LAPACK Working Note 191 UT-CS-07-600 University of Tennessee."},{"key":"e_1_2_1_9_1","unstructured":"Buttari A. Langou J. Kurzak J. and Dongarra J. 2007b. Parallel tiled QR factorization for multicore architectures. LAPACK Working Note 190 UT-CS-07-598 University of Tennessee. Buttari A. Langou J. Kurzak J. and Dongarra J. 2007b. Parallel tiled QR factorization for multicore architectures. LAPACK Working Note 190 UT-CS-07-598 University of Tennessee."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1248377.1248397"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1006\/jsvi.1996.0111"},{"key":"e_1_2_1_12_1","unstructured":"Golub G. H. and Loan C. F. V. 1996. Matrix Computations 3rd Ed. The Johns Hopkins University Press Baltimore MD. Golub G. H. and Loan C. F. V. 1996. Matrix Computations 3rd Ed. The Johns Hopkins University Press Baltimore MD."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/504210.504213"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1055531.1055534"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC.2010.5434077"},{"key":"e_1_2_1_17_1","unstructured":"Igual F. D. Quintana-Ort\u00ed G. and van de Geijn R. 2009. Level-3 BLAS on a GPU: Picking the low hanging fruit. FLAME Working Note #37 DICC 2009-04-01 Depto. de Ingenier\u00eda y Ciencia de Computadores Universidad Jaime I de Castell\u00f3n. Igual F. D. Quintana-Ort\u00ed G. and van de Geijn R. 2009. Level-3 BLAS on a GPU: Picking the low hanging fruit. FLAME Working Note #37 DICC 2009-04-01 Depto. de Ingenier\u00eda y Ciencia de Computadores Universidad Jaime I de Castell\u00f3n."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/11558958_49"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2008.224"},{"volume-title":"Department of Computer Sciences","author":"Low T. M.","key":"e_1_2_1_20_1","unstructured":"Low , T. M. and van de Geijn , R. 2004. An API for manipulating matrices stored by blocks. FLAME Working Note #12 TR-2004-15 , Department of Computer Sciences , The University of Texas at Austin. Low, T. M. and van de Geijn, R. 2004. An API for manipulating matrices stored by blocks. FLAME Working Note #12 TR-2004-15, Department of Computer Sciences, The University of Texas at Austin."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2009.09.005"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1377612.1377615"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1527286.1527288"},{"volume-title":"Milestones in Computer Science and Information Technology","author":"Reilly E. D.","key":"e_1_2_1_24_1","unstructured":"Reilly , E. D. 2003. Milestones in Computer Science and Information Technology . Greenwood Press , Westport, CT . 164. Reilly, E. D. 2003. Milestones in Computer Science and Information Technology. Greenwood Press, Westport, CT. 164."},{"volume-title":"Proceedings of the ASME International Mechanical Engineering Congress &amp; Exposition.","author":"Schafer N.","key":"e_1_2_1_25_1","unstructured":"Schafer , N. , Serban , R. , and Negrut , D . 2008. Implicit integration in molecular dynamics simulation . In Proceedings of the ASME International Mechanical Engineering Congress &amp; Exposition. Schafer, N., Serban, R., and Negrut, D. 2008. Implicit integration in molecular dynamics simulation. In Proceedings of the ASME International Mechanical Engineering Congress &amp; Exposition."},{"volume-title":"Proceedings of the ACM Special Interest Group on Programming Language Conference. 59--70","author":"Siek J. G.","key":"e_1_2_1_26_1","unstructured":"Siek , J. G. and Lumsdaine , A . 1998a. The matrix template library: A generic programming approach to high performance numerical linear algebra . In Proceedings of the ACM Special Interest Group on Programming Language Conference. 59--70 . Siek, J. G. and Lumsdaine, A. 1998a. The matrix template library: A generic programming approach to high performance numerical linear algebra. In Proceedings of the ACM Special Interest Group on Programming Language Conference. 59--70."},{"volume-title":"Proceedings of the European Conference on Object-Oriented Programming Workshops. 468--469","author":"Siek J. G.","key":"e_1_2_1_27_1","unstructured":"Siek , J. G. and Lumsdaine , A . 1998b. A rational approach to portable high performance: The basic linear algebra instruction set (blais) and the fixed algorithm size template (fast) library . In Proceedings of the European Conference on Object-Oriented Programming Workshops. 468--469 . Siek, J. G. and Lumsdaine, A. 1998b. A rational approach to portable high performance: The basic linear algebra instruction set (blais) and the fixed algorithm size template (fast) library. In Proceedings of the European Conference on Object-Oriented Programming Workshops. 468--469."},{"key":"e_1_2_1_28_1","unstructured":"Silberschatz A. Galvin P. B. and Gagne G. 2008. Operating System Concepts 8th Ed. John Wiley &amp; Sons Inc. Silberschatz A. Galvin P. B. and Gagne G. 2008. Operating System Concepts 8th Ed. John Wiley &amp; Sons Inc."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/236017.236029"},{"key":"e_1_2_1_30_1","unstructured":"YarKhan A. Kurzak J. and Dongarra J. 2011. QUARK Users\u2019 Guide: QUeueing and Runtime for Kernels. Tech. rep. ICL-UT-11-02 Innovative Computing Laboratory University of Tennessee. YarKhan A. Kurzak J. and Dongarra J. 2011. QUARK Users\u2019 Guide: QUeueing and Runtime for Kernels. Tech. rep. ICL-UT-11-02 Innovative Computing Laboratory University of Tennessee."},{"key":"e_1_2_1_31_1","doi-asserted-by":"crossref","unstructured":"Zhang Y. and Sarkar T. K. 2009. Parallel Solution of Integral Equation-Based EM Problems in the Frequency Domain. John Wiley &amp; Sons Hoboken NJ. Zhang Y. and Sarkar T. K. 2009. Parallel Solution of Integral Equation-Based EM Problems in the Frequency Domain . John Wiley &amp; Sons Hoboken NJ.","DOI":"10.1002\/9780470495094"}],"container-title":["ACM Transactions on Mathematical Software"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2331130.2331133","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2331130.2331133","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T08:49:12Z","timestamp":1750236552000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2331130.2331133"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,8]]},"references-count":29,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2012,8]]}},"alternative-id":["10.1145\/2331130.2331133"],"URL":"https:\/\/doi.org\/10.1145\/2331130.2331133","relation":{},"ISSN":["0098-3500","1557-7295"],"issn-type":[{"type":"print","value":"0098-3500"},{"type":"electronic","value":"1557-7295"}],"subject":[],"published":{"date-parts":[[2012,8]]},"assertion":[{"value":"2010-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2012-08-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}