{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T07:38:48Z","timestamp":1740123528940,"version":"3.37.3"},"reference-count":23,"publisher":"Springer Science and Business Media LLC","issue":"11","license":[{"start":{"date-parts":[[2024,4,15]],"date-time":"2024-04-15T00:00:00Z","timestamp":1713139200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,4,15]],"date-time":"2024-04-15T00:00:00Z","timestamp":1713139200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100008431","name":"Consejer\u00eda de Educaci\u00f3n, Junta de Castilla y Le\u00f3n","doi-asserted-by":"publisher","award":["VA226P20","VA226P20","VA226P20","VA226P20","VA226P20"],"award-info":[{"award-number":["VA226P20","VA226P20","VA226P20","VA226P20","VA226P20"]}],"id":[{"id":"10.13039\/501100008431","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004837","name":"Ministerio de Ciencia e Innovaci\u00f3n","doi-asserted-by":"publisher","award":["PID2022-142292NB-I00","PID2022-142292NB-I00","PID2022-142292NB-I00","PID2022-142292NB-I00"],"award-info":[{"award-number":["PID2022-142292NB-I00","PID2022-142292NB-I00","PID2022-142292NB-I00","PID2022-142292NB-I00"]}],"id":[{"id":"10.13039\/501100004837","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Red Espa\u00f1ola de Supercomputaci\u00f3n","award":["IM-2023-3-0020","IM-2023-3-0020"],"award-info":[{"award-number":["IM-2023-3-0020","IM-2023-3-0020"]}]},{"DOI":"10.13039\/501100007515","name":"Universidad de Valladolid","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100007515","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2024,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>There are many works devoted to improving the matrix product computation, as it is used in a wide variety of scientific applications arising from many different fields. In this work, we propose alternative data distribution policies and communication patterns to reduce the elapsed time when computing triangular matrix products in distributed memory environments. In particular, we focus on commodity clusters, where the number of nodes is limited, proposing alternatives to traditional approaches in order to improve this operation\u2019s performance. Our proposal overcomes the performance results associated with the state-of-the-art libraries, such as ScaLAPACK and SLATE, offering execution times that are up to 30% faster.<\/jats:p>","DOI":"10.1007\/s11227-024-06097-7","type":"journal-article","created":{"date-parts":[[2024,4,15]],"date-time":"2024-04-15T18:02:01Z","timestamp":1713204121000},"page":"16630-16653","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Performance improvement of the triangular matrix product in commodity clusters"],"prefix":"10.1007","volume":"80","author":[{"given":"Inmaculada","family":"Santamaria-Valenzuela","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Roc\u00edo","family":"Carratal\u00e1-S\u00e1ez","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuri","family":"Torres","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Diego R.","family":"Llanos","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arturo","family":"Gonzalez-Escribano","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,4,15]]},"reference":[{"key":"6097_CR1","doi-asserted-by":"publisher","unstructured":"Sterling TL (2011). In: Padua D (ed) Clusters. Springer, Boston, pp 289\u2013297. https:\/\/doi.org\/10.1007\/978-0-387-09766-4_18","DOI":"10.1007\/978-0-387-09766-4_18"},{"issue":"1137\/1","key":"6097_CR2","first-page":"9780898719642","volume":"10","author":"LS Blackford","year":"1997","unstructured":"Blackford LS, Choi J, Cleary A, D\u2019Azevedo E, Demmel J, Dhillon I, Dongarra J, Hammarling S, Henry G, Petitet A, Stanley K, Walker D, Whaley RC (1997) ScaLAPACK Users\u2019 Guide. Soc Ind Appl Math. doi 10(1137\/1):9780898719642","journal-title":"Soc Ind Appl Math. doi"},{"key":"6097_CR3","doi-asserted-by":"publisher","unstructured":"Gates M, Kurzak J, Charara A, YarKhan A, Dongarra J (2019) Slate: design of a modern distributed and accelerated linear algebra library. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. SC\u201919. Association for Computing Machinery, New York. https:\/\/doi.org\/10.1145\/3295500.3356223","DOI":"10.1145\/3295500.3356223"},{"issue":"4","key":"6097_CR4","doi-asserted-by":"publisher","first-page":"255","DOI":"10.1002\/(SICI)1096-9128(199704)9:4<255::AID-CPE250>3.0.CO;2-2","volume":"9","author":"RA Geijn","year":"1997","unstructured":"Geijn RA, Watts J (1997) SUMMA: scalable universal matrix multiplication algorithm. Concurr Pract Exp 9(4):255\u2013274","journal-title":"Concurr Pract Exp"},{"key":"6097_CR5","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2019.102597","volume":"91","author":"V Manin","year":"2020","unstructured":"Manin V, Lang B (2020) Cannon-type triangular matrix multiplication for the reduction of generalized HPD eigenproblems to standard form. Parallel Comput 91:102597. https:\/\/doi.org\/10.1016\/j.parco.2019.102597","journal-title":"Parallel Comput"},{"key":"6097_CR6","doi-asserted-by":"publisher","unstructured":"Sankaran A, Alashti NA, Psarras C, Bientinesi P (2022) Benchmarking the linear algebra awareness of tensorflow and pytorch. In: 2022 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), pp 924\u2013933. https:\/\/doi.org\/10.1109\/IPDPSW55747.2022.00150","DOI":"10.1109\/IPDPSW55747.2022.00150"},{"issue":"3","key":"6097_CR7","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/1356052.1356053","volume":"34","author":"K Goto","year":"2008","unstructured":"Goto K, Geijn RA (2008) Anatomy of high-performance matrix multiplication. ACM Trans Math Softw 34(3):1\u201325. https:\/\/doi.org\/10.1145\/1356052.1356053","journal-title":"ACM Trans Math Softw"},{"issue":"3","key":"6097_CR8","first-page":"14","volume":"41","author":"FG Zee","year":"2015","unstructured":"Zee FG, Geijn RA (2015) BLIS: A framework for rapidly instantiating BLAS functionality. ACM Trans Math Softw 41(3):14\u201311433","journal-title":"ACM Trans Math Softw"},{"key":"6097_CR9","doi-asserted-by":"publisher","unstructured":"Wang Q, Zhang X, Zhang Y, Yi Q (2013) Augem: automatically generate high performance dense linear algebra kernels on x86 CPUs. In: SC \u201913: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, pp 1\u201312. https:\/\/doi.org\/10.1145\/2503210.2503219","DOI":"10.1145\/2503210.2503219"},{"key":"6097_CR10","doi-asserted-by":"publisher","first-page":"1276","DOI":"10.4236\/am.2015.68121","volume":"6","author":"A Kalinkin","year":"2015","unstructured":"Kalinkin A, Anders A, Anders R (2015) Intel Math Kernel Library PARDISO for Intel Xeon PhiTM Manycore Coprocessor. Appl Math 6:1276\u20131281. https:\/\/doi.org\/10.4236\/am.2015.68121","journal-title":"Appl Math"},{"key":"6097_CR11","unstructured":"Limited A. ARM Performance Libraries Reference Guide. https:\/\/developer.arm.com\/documentation\/101004\/2030\/"},{"key":"6097_CR12","unstructured":"(AMD) AMD. AMD Optimizing CPU Libraries (AOCL). https:\/\/www.amd.com\/en\/developer\/aocl.html#documentation"},{"key":"6097_CR13","doi-asserted-by":"publisher","unstructured":"Anderson E, Bai Z, Bischof C, Blackford LS, Demmel J, Dongarra J, Du\u00a0Croz J, Greenbaum A, Hammarling S, McKenney A, Sorensen D (1999) LAPACK Users\u2019 Guide, 3rd edn. Society for Industrial and Applied Mathematics. https:\/\/doi.org\/10.1137\/1.9780898719604","DOI":"10.1137\/1.9780898719604"},{"key":"6097_CR14","unstructured":"Geijn RA, Watts J (1995) Summa: Scalable universal matrix multiplication algorithm. Technical report, USA"},{"issue":"1","key":"6097_CR15","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/77626.79170","volume":"16","author":"JJ Dongarra","year":"1990","unstructured":"Dongarra JJ, Du Croz J, Hammarling S, Duff IS (1990) A set of level 3 basic linear algebra subprograms. ACM Trans Math Softw 16(1):1\u201317. https:\/\/doi.org\/10.1145\/77626.79170","journal-title":"ACM Trans Math Softw"},{"key":"6097_CR16","unstructured":"Forum MP (1994) MPI: a message-passing interface standard. Technical report, USA"},{"key":"6097_CR17","unstructured":"Corp I (2023) Intel math kernel library. https:\/\/software.intel.com\/en-us\/mkl"},{"key":"6097_CR18","unstructured":"Du P, Tomov S, Dongarra J (2012) Providing GPU capability to LU and QR within the scalapack framework. Technical Report UT-CS-12-699 (2012)"},{"key":"6097_CR19","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1016\/j.procs.2012.04.008","volume":"9","author":"E D\u2019Azevedo","year":"2012","unstructured":"D\u2019Azevedo E, Hill JC (2012) Parallel LU factorization on GPU cluster. Proc Comput Sci 9:67\u201375. https:\/\/doi.org\/10.1016\/j.procs.2012.04.008","journal-title":"Proc Comput Sci"},{"key":"6097_CR20","unstructured":"Chandra R, Dagum L, Kohr D, Menon R, Maydan D, McDonald J (2001) Parallel programming in OpenMP. Morgan kaufmann"},{"key":"6097_CR21","unstructured":"Whaley RC, Dongarra JJ (1999) Automatically tuned linear algebra software. In: Proceedings of the Ninth SIAM Conference on Parallel Processing for Scientific Computing, PPSC 1999, San Antonio, Texas, USA, March 22\u201324, 1999. SIAM"},{"key":"6097_CR22","doi-asserted-by":"crossref","unstructured":"Tukanov N, Srinivasaraghavan R, Moreira JE, Low TM (2022) Modeling matrix engines for portability and performance. In: 2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS), pp 1173\u20131183","DOI":"10.1109\/IPDPS53621.2022.00117"},{"key":"6097_CR23","doi-asserted-by":"publisher","first-page":"153","DOI":"10.1016\/J.JPDC.2019.12.002","volume":"138","author":"P Valero-Lara","year":"2020","unstructured":"Valero-Lara P, Catal\u00e1n S, Martorell X, Usui T, Labarta J (2020) sLASs: A fully automatic auto-tuned linear algebra library based on OpenMP extensions implemented in OmpSs (LASs library). J Parallel Distrib Comput 138:153\u2013171. https:\/\/doi.org\/10.1016\/J.JPDC.2019.12.002","journal-title":"J Parallel Distrib Comput"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-024-06097-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-024-06097-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-024-06097-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,25]],"date-time":"2024-06-25T11:10:42Z","timestamp":1719313842000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-024-06097-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,15]]},"references-count":23,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2024,7]]}},"alternative-id":["6097"],"URL":"https:\/\/doi.org\/10.1007\/s11227-024-06097-7","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"type":"print","value":"0920-8542"},{"type":"electronic","value":"1573-0484"}],"subject":[],"published":{"date-parts":[[2024,4,15]]},"assertion":[{"value":"21 March 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 April 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no Conflict of interest as defined by Springer, or other interests that might be perceived to influence the results and\/or discussion reported in this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"Not applicable.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}]}}