{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T23:14:39Z","timestamp":1776122079293,"version":"3.50.1"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2018,1,3]],"date-time":"2018-01-03T00:00:00Z","timestamp":1514937600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Math. Softw."],"published-print":{"date-parts":[[2018,9,30]]},"abstract":"<jats:p>\n            The Sparse Matrix-Vector Multiplication (SpMV) kernel ranks among the most important and thoroughly studied linear algebra operations, as it lies at the heart of many iterative methods for the solution of sparse linear systems, and often constitutes a severe performance bottleneck. Its optimization, which is intimately associated with the data structures used to store the sparse matrix, has always been of particular interest to the applied mathematics and computer science communities and has attracted further attention since the advent of multicore architectures. In this article, we present SparseX, an open source software package for SpMV targeting multicore platforms, that employs the state-of-the-art\n            <jats:italic>Compressed Sparse eXtended<\/jats:italic>\n            (CSX) sparse matrix storage format to deliver high efficiency through a highly usable \u201cBLAS-like\u201d interface that requires limited or no tuning. Performance results indicate that our library achieves superior performance over competitive libraries on large-scale problems.\n          <\/jats:p>","DOI":"10.1145\/3134442","type":"journal-article","created":{"date-parts":[[2018,1,4]],"date-time":"2018-01-04T16:27:31Z","timestamp":1515083251000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":25,"title":["SparseX"],"prefix":"10.1145","volume":"44","author":[{"given":"Athena","family":"Elafrou","sequence":"first","affiliation":[{"name":"National Technical University of Athens, Athens, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vasileios","family":"Karakasis","sequence":"additional","affiliation":[{"name":"Swiss National Supercomputing Centre, ETH Zurich, Lugano, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Theodoros","family":"Gkountouvas","sequence":"additional","affiliation":[{"name":"Cornell University, New York, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kornilios","family":"Kourtis","sequence":"additional","affiliation":[{"name":"IBM Reasearch Zurich, Zurich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Georgios","family":"Goumas","sequence":"additional","affiliation":[{"name":"National Technical University of Athens, Athens, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nectarios","family":"Koziris","sequence":"additional","affiliation":[{"name":"National Technical University of Athens, Athens, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,1,3]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"crossref","volume-title":"Performance Evaluation of Scientific Applications on POWER8","author":"Adinetz Andrew V.","DOI":"10.1007\/978-3-319-17248-4_2"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/147877.147901"},{"key":"e_1_2_1_3_1","volume-title":"Master\u2019s thesis","author":"Ankit J."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/125826.125925"},{"key":"e_1_2_1_5_1","volume-title":"Lois Curfman McInnes, and Barry F. Smith","author":"Balay Satish","year":"1997"},{"key":"e_1_2_1_6_1","unstructured":"V. H. F. Batista G. O. Ainsworth Jr. and F. L. B. Ribeiro. 2010. Parallel structurally-symmetric sparse matrix-vector products on multi-core processors. Computing Research Repository (CoRR) abs\/1003.0952.  V. H. F. Batista G. O. Ainsworth Jr. and F. L. B. Ribeiro. 2010. Parallel structurally-symmetric sparse matrix-vector products on multi-core processors. Computing Research Repository (CoRR) abs\/1003.0952."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CDC.2006.376798"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1542275.1542294"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/800195.805928"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2049662.2049663"},{"key":"e_1_2_1_12_1","volume-title":"Technical Report 4744. Sandia National Laboratories.","author":"Dongarra J.","year":"2013"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1155\/2001\/527931"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/355815.355817"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2013.248"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2012.290"},{"key":"e_1_2_1_17_1","volume-title":"Raftery","author":"Gneiting Tilmann","year":"2005"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-008-0251-8"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.68"},{"key":"e_1_2_1_20_1","unstructured":"M. H. Gutknecht. 2007. Block Krylov space methods for linear systems with multiple right-hand sides: An introduction. Modern Mathematical Models Methods and Algorithms for Real World Systems 420--447.  M. H. Gutknecht. 2007. Block Krylov space methods for linear systems with multiple right-hand sides: An introduction. Modern Mathematical Models Methods and Algorithms for Real World Systems 420--447."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1198555.1198768"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1186736.1186737"},{"key":"e_1_2_1_23_1","volume-title":"Technical Report 4744. Sandia National Laboratories.","author":"Heroux M. A.","year":"2013"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1089014.1089021"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.6028\/jres.049.044"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/L-CA.2013.6"},{"key":"e_1_2_1_27_1","volume-title":"Optimizing Sparse Matrix Computations for Register Reuse in SPARSITY","author":"Im Eun-Jin"},{"key":"e_1_2_1_28_1","unstructured":"Intel\u00ae Coorporation. 2013. Intel\u00ae Math Kernel Library. Retrieved from http:\/\/software.intel.com\/en-us\/intel-mkl.  Intel\u00ae Coorporation. 2013. Intel\u00ae Math Kernel Library. Retrieved from http:\/\/software.intel.com\/en-us\/intel-mkl."},{"key":"e_1_2_1_29_1","unstructured":"ITRS. 2011. International Technology Roadmap for Semiconductors: Assembly and Packaging. Retrieved from http:\/\/www.itrs.net\/Links\/2005ITRS\/AP2011.pdf.  ITRS. 2011. International Technology Roadmap for Semiconductors: Assembly and Packaging. Retrieved from http:\/\/www.itrs.net\/Links\/2005ITRS\/AP2011.pdf."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2012.290"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/1366230.1366244"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/1941553.1941587"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1137\/130930352"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/355841.355847"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2491956.2462181"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2751205.2751209"},{"key":"e_1_2_1_37_1","unstructured":"D. Lukarski. 2013. PARALUTION project. Retrieved from http:\/\/www.paralution.com\/.  D. Lukarski. 2013. PARALUTION project. Retrieved from http:\/\/www.paralution.com\/."},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the ISCA 25th International Conference on Computers and Their Applications (CATA\u201910)","author":"Martone M."},{"key":"e_1_2_1_39_1","unstructured":"J. D. McCalpin. 1995. STREAM: Sustainable Memory Bandwidth in High Performance Computing. Retrieved from http:\/\/www.cs.virginia.edu\/stream\/.  J. D. McCalpin. 1995. STREAM: Sustainable Memory Bandwidth in High Performance Computing. Retrieved from http:\/\/www.cs.virginia.edu\/stream\/."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/331532.331562"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/356616.356618"},{"key":"e_1_2_1_42_1","doi-asserted-by":"crossref","volume-title":"Numerical Methods for Large Eigenvalue Problems","author":"Saad Y.","DOI":"10.1137\/1.9781611970739"},{"key":"e_1_2_1_43_1","volume-title":"Iterative Methods for Sparse Linear Systems","author":"Saad Y.","edition":"2"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/PROC.1967.6011"},{"key":"e_1_2_1_45_1","volume-title":"LIKWID: Lightweight performance tools. CoRR abs\/1104.4874","author":"Treibig Jan","year":"2011"},{"key":"e_1_2_1_46_1","volume-title":"OSKI: A library of automatically tuned sparse matrix kernels. J. Phys,: Conf. Ser. 16, 521","author":"Vuduc R.","year":"2005"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/11557654_91"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/1183401.1183444"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"}],"container-title":["ACM Transactions on Mathematical Software"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3134442","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3134442","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:30:10Z","timestamp":1750217410000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3134442"}},"subtitle":["A Library for High-Performance Sparse Matrix-Vector Multiplication on Multicore Platforms"],"short-title":[],"issued":{"date-parts":[[2018,1,3]]},"references-count":48,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2018,9,30]]}},"alternative-id":["10.1145\/3134442"],"URL":"https:\/\/doi.org\/10.1145\/3134442","relation":{},"ISSN":["0098-3500","1557-7295"],"issn-type":[{"value":"0098-3500","type":"print"},{"value":"1557-7295","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,1,3]]},"assertion":[{"value":"2015-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-08-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-01-03","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}