{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:36:35Z","timestamp":1750307795900,"version":"3.41.0"},"reference-count":18,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2008,3,1]],"date-time":"2008-03-01T00:00:00Z","timestamp":1204329600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["CCF-0540926 CCF-0342369 ACI 0305163"],"award-info":[{"award-number":["CCF-0540926 CCF-0342369 ACI 0305163"]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007523","name":"Advanced Cyberinfrastructure","doi-asserted-by":"publisher","award":["CCF-0540926 CCF-0342369 ACI 0305163"],"award-info":[{"award-number":["CCF-0540926 CCF-0342369 ACI 0305163"]}],"id":[{"id":"10.13039\/100007523","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Math. Softw."],"published-print":{"date-parts":[[2008,3]]},"abstract":"<jats:p>We discuss the OpenMP parallelization of linear algebra algorithms that are coded using the Formal Linear Algebra Methods Environment (FLAME) API. This API expresses algorithms at a higher level of abstraction, avoids the use loop and array indices, and represents these algorithms as they are formally derived and presented. We report on two implementations of the workqueuing model, neither of which requires the use of explicit indices to specify parallelism. The first implementation uses the experimental taskq pragma, which may influence the adoption of a similar construct into OpenMP 3.0. The second workqueuing implementation is domain-specific to FLAME but allows us to illustrate the benefits of sorting tasks according to their computational cost prior to parallel execution. In addition, we discuss how scalable parallelization of dense linear algebra algorithms via OpenMP will require a two-dimensional partitioning of operands much like a 2D data distribution is needed on distributed memory architectures. We illustrate the issues and solutions by discussing the parallelization of the symmetric rank-k update and report impressive performance on an SGI system with 14 Itanium2 processors.<\/jats:p>","DOI":"10.1145\/1326548.1326552","type":"journal-article","created":{"date-parts":[[2008,3,19]],"date-time":"2008-03-19T12:58:50Z","timestamp":1205931530000},"page":"1-29","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Scalable parallelization of FLAME code via the workqueuing model"],"prefix":"10.1145","volume":"34","author":[{"given":"Field G. Van","family":"Zee","sequence":"first","affiliation":[{"name":"The University of Texas at Austin, Austin, TX"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Paolo","family":"Bientinesi","sequence":"additional","affiliation":[{"name":"Duke University, Durham, NC"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tze Meng","family":"Low","sequence":"additional","affiliation":[{"name":"The University of Texas at Austin, Austin, TX"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Robert A. van de","family":"Geijn","sequence":"additional","affiliation":[{"name":"The University of Texas at Austin, Austin, TX"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2008,3,19]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Anderson E. Bai Z. Demmel J. Dongarra J. E. DuCroz J. Greenbaum A. Hammarling S. McKenney A. E. Ostrouchov S. and Sorensen D. 1992. LAPACK Users' Guide. SIAM Philadelphia.   Anderson E. Bai Z. Demmel J. Dongarra J. E. DuCroz J. Greenbaum A. Hammarling S. McKenney A. E. Ostrouchov S. and Sorensen D. 1992. LAPACK Users' Guide. SIAM Philadelphia."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1055531.1055532"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1055531.1055533"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/77626.79170"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/42288.42291"},{"key":"e_1_2_1_7_1","unstructured":"Dow E. 2005. Take charge of processor affinity. IBM developerWorks. http:\/\/www.ibm.com\/developerworks\/linux\/library\/l-affinity.html.  Dow E. 2005. Take charge of processor affinity. IBM developerWorks. http:\/\/www.ibm.com\/developerworks\/linux\/library\/l-affinity.html."},{"key":"e_1_2_1_8_1","unstructured":"Garey M. R. Graham R. L. and Ullman J. D. 1973. An analysis of some packing algorithms. In Combinatorial Algorithms R. Rustin Ed. Algorithmics Press New York 39--47.  Garey M. R. Graham R. L. and Ullman J. D. 1973. An analysis of some packing algorithms. In Combinatorial Algorithms R. Rustin Ed. Algorithmics Press New York 39--47."},{"key":"e_1_2_1_9_1","unstructured":"Goto K. 2006. http:\/\/www.cs.utexas.edu\/users\/kgoto.  Goto K. 2006. http:\/\/www.cs.utexas.edu\/users\/kgoto."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1356052.1356053"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/800125.804034"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/355841.355847"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1065944.1065965"},{"key":"e_1_2_1_15_1","unstructured":"OpenMP Architecture Review Board. 2006. http:\/\/www.openmp.org\/.  OpenMP Architecture Review Board. 2006. http:\/\/www.openmp.org\/."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/779359.779365"},{"key":"e_1_2_1_17_1","unstructured":"Shah S. Haab G. Peterson P. and Throop J. 1999. Flexible control structures for parallelism in OpenMP. In P2 6 7 European Workshop on OpenMP (EWOMP).  Shah S. Haab G. Peterson P. and Throop J. 1999. Flexible control structures for parallelism in OpenMP. In P2 6 7 European Workshop on OpenMP (EWOMP)."},{"volume-title":"European Workshop on OpenMP (EWOMP).","author":"Su E.","key":"e_1_2_1_18_1","unstructured":"Su , E. , Tian , X. , Girkar , M. , Haab , G. , Shah , S. , and Peterson , P . 2002. Compiler support of the workqueuing execution model for Intel SMP architectures . In European Workshop on OpenMP (EWOMP). Su, E., Tian, X., Girkar, M., Haab, G., Shah, S., and Peterson, P. 2002. Compiler support of the workqueuing execution model for Intel SMP architectures. In European Workshop on OpenMP (EWOMP)."},{"volume-title":"Using PLAPACK: Parallel Linear Algebra Package","author":"van de Geijn R. A.","key":"e_1_2_1_19_1","unstructured":"van de Geijn , R. A. 1997. Using PLAPACK: Parallel Linear Algebra Package . The MIT Press . van de Geijn, R. A. 1997. Using PLAPACK: Parallel Linear Algebra Package. The MIT Press."},{"key":"e_1_2_1_20_1","unstructured":"Weisstein E. W. 2006. Bin-Packing Problem. From MathWorld---A Wolfram Web Resource. http:\/\/mathworld.wolfram.com\/Bin-PackingProblem.html.  Weisstein E. W. 2006. Bin-Packing Problem. From MathWorld---A Wolfram Web Resource. http:\/\/mathworld.wolfram.com\/Bin-PackingProblem.html."}],"container-title":["ACM Transactions on Mathematical Software"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1326548.1326552","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1326548.1326552","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T13:56:25Z","timestamp":1750254985000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1326548.1326552"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,3]]},"references-count":18,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2008,3]]}},"alternative-id":["10.1145\/1326548.1326552"],"URL":"https:\/\/doi.org\/10.1145\/1326548.1326552","relation":{},"ISSN":["0098-3500","1557-7295"],"issn-type":[{"type":"print","value":"0098-3500"},{"type":"electronic","value":"1557-7295"}],"subject":[],"published":{"date-parts":[[2008,3]]},"assertion":[{"value":"2006-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2007-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2008-03-19","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}