{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,4,3]],"date-time":"2022-04-03T06:06:25Z","timestamp":1648965985277},"reference-count":11,"publisher":"World Scientific Pub Co Pte Lt","issue":"04","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Parallel Process. Lett."],"published-print":{"date-parts":[[2014,12]]},"abstract":"<jats:p> A systolic array provides an alternative computing paradigm to the von Neumann architecture. Though its hardware implementation has failed as a paradigm to design integrated circuits in the past, we are now discovering that the systolic array as a software virtualization layer can lead to an extremely scalable execution paradigm. To demonstrate this scalability, in this paper, we design and implement a 3D virtual systolic array to compute a tile QR decomposition of a tall-and-skinny dense matrix. Our implementation is based on a state-of-the-art algorithm that factorizes a panel based on a tree-reduction. Freed from the constraint of a planar layout, we present a three-dimensional virtual systolic array architecture for this algorithm. Using a runtime developed as a part of the Parallel Ultra Light Systolic Array Runtime (PULSAR) project, we demonstrate on a Cray-XT5 machine how our virtual systolic array can be mapped to a large-scale machine and obtain excellent parallel performance. This is an important contribution since such a QR decomposition is used, for example, to compute a least squares solution of an overdetermined system, which arises in many scientific and engineering problems. <\/jats:p>","DOI":"10.1142\/s0129626414420043","type":"journal-article","created":{"date-parts":[[2014,12,31]],"date-time":"2014-12-31T04:43:13Z","timestamp":1420000993000},"page":"1442004","source":"Crossref","is-referenced-by-count":0,"title":["Design and Implementation of a Large Scale Tree-Based QR Decomposition Using a 3D Virtual Systolic Array and a Lightweight Runtime"],"prefix":"10.1142","volume":"24","author":[{"given":"Ichitaro","family":"Yamazaki","sequence":"first","affiliation":[{"name":"University of Tennessee, Knoxville, Tennessee 37996-3450, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jakub","family":"Kurzak","sequence":"additional","affiliation":[{"name":"University of Tennessee, Knoxville, Tennessee 37996-3450, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Piotr","family":"Luszczek","sequence":"additional","affiliation":[{"name":"University of Tennessee, Knoxville, Tennessee 37996-3450, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jack","family":"Dongarra","sequence":"additional","affiliation":[{"name":"University of Tennessee, Knoxville, Tennessee 37996-3450, USA"},{"name":"Oak Ridge National Laboratory, Oak Ridge, Tennessee 37831, USA"},{"name":"University of Manchester, Manchester, M13 9PL, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2014,12,30]]},"reference":[{"issue":"1","key":"p_3","first-page":"1094","volume":"25","author":"Dongarra J.","year":"2011","journal-title":"Int. J. High Perf. Comput. Applic."},{"key":"p_11","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.1829"},{"key":"p_12","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2012.242"},{"key":"p_14","doi-asserted-by":"publisher","DOI":"10.1023\/B:GRID.0000024072.93701.f3"},{"key":"p_15","doi-asserted-by":"publisher","DOI":"10.1147\/rd.515.0593"},{"key":"p_17","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.1463"},{"key":"p_19","doi-asserted-by":"crossref","unstructured":"C. Augonnet and R. Namyst. A unified runtime system for heterogeneous multicore architectures. In Proceedings of the Euro-Par 2008 Workshops - Parallel Processing, Lecture Notes in Computer Science, pages 174-183, Las Palmas de Gran Canaria, Spain, August 2008. Springer. DOI: 10.1007\/978-3-642-00955-6 22.10.1007\/978-3-642-00955-6","DOI":"10.1007\/978-3-642-00955-6"},{"key":"p_20","doi-asserted-by":"publisher","DOI":"10.1109\/2.214440"},{"key":"p_21","doi-asserted-by":"publisher","DOI":"10.1145\/291889.291893"},{"key":"p_22","doi-asserted-by":"publisher","DOI":"10.1145\/209937.209958"},{"key":"p_23","doi-asserted-by":"publisher","DOI":"10.1109\/99.660313"}],"container-title":["Parallel Processing Letters"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0129626414420043","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,8,6]],"date-time":"2019-08-06T20:00:02Z","timestamp":1565121602000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/abs\/10.1142\/S0129626414420043"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,12]]},"references-count":11,"journal-issue":{"issue":"04","published-online":{"date-parts":[[2014,12,30]]},"published-print":{"date-parts":[[2014,12]]}},"alternative-id":["10.1142\/S0129626414420043"],"URL":"https:\/\/doi.org\/10.1142\/s0129626414420043","relation":{},"ISSN":["0129-6264","1793-642X"],"issn-type":[{"value":"0129-6264","type":"print"},{"value":"1793-642X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2014,12]]}}}