{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T06:02:53Z","timestamp":1783576973405,"version":"3.55.0"},"reference-count":29,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2008,11,25]],"date-time":"2008-11-25T00:00:00Z","timestamp":1227571200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2009,10]]},"DOI":"10.1007\/s11227-008-0251-8","type":"journal-article","created":{"date-parts":[[2008,11,26]],"date-time":"2008-11-26T15:22:00Z","timestamp":1227712920000},"page":"36-77","source":"Crossref","is-referenced-by-count":76,"title":["Performance evaluation of the sparse matrix-vector multiplication on modern architectures"],"prefix":"10.1007","volume":"50","author":[{"given":"Georgios","family":"Goumas","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kornilios","family":"Kourtis","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nikos","family":"Anastopoulos","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Vasileios","family":"Karakasis","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nectarios","family":"Koziris","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2008,11,25]]},"reference":[{"key":"251_CR1","first-page":"32","volume-title":"Supercomputing\u201992","author":"RC Agarwal","year":"1992","unstructured":"Agarwal RC, Gustavson FG, Zubair M (1992) a high performance algorithm using pre-processing for the sparse matrix-vector multiplication. In: Supercomputing\u201992, Minnesota, November 1992. IEEE, New York, pp 32\u201341"},{"key":"251_CR2","unstructured":"Asanovic K, Bodik R, Catanzaro BC, Gebis JJ, Husbands P, Keutzer K, Patterson DA, Plishker WL, Shalf J, Williams SW, Yelick KA (2006) The landscape of parallel computing research: A view from Berkeley. Technical Report UCB\/EECS-2006-183, EECS Department, University of California, Berkeley"},{"issue":"1","key":"251_CR3","doi-asserted-by":"crossref","first-page":"64","DOI":"10.1007\/s11227-007-0149-x","volume":"44","author":"E Athanasaki","year":"2008","unstructured":"Athanasaki E, Anastopoulos N, Kourtis K, Koziris N (2008) Exploring the performance limits of simultaneous multithreading for memory intensive applications. J Supercomput 44(1):64\u201397","journal-title":"J Supercomput"},{"key":"251_CR4","doi-asserted-by":"crossref","DOI":"10.1137\/1.9781611971538","volume-title":"Templates for the solution of linear systems: building blocks for iterative methods","author":"R Barrett","year":"1994","unstructured":"Barrett R, Berry M, Chan TF, Demmel J, Donato JM, Dongarra J, Eijkhout V, Pozo R, Romine C, der Vorst HV (1994) Templates for the solution of linear systems: building blocks for iterative methods. SIAM, Philadelphia"},{"key":"251_CR5","unstructured":"Buttari A, Eijkhout V, Langou J, Filippone S (2005) Performance optimization and modeling of blocked sparse kernels. Technical Report ICL-UT-04-05, Innovative Computing Laboratory, University of Tennessee"},{"key":"251_CR6","unstructured":"Catalyuerek UV, Aykanat C (1996) Decomposing irregularly sparse matrices for parallel matrix-vector multiplication. In: Lecture notes in computer science, vol 1117, pp 75\u201386"},{"key":"251_CR7","unstructured":"Davis T (1997) University of Florida Sparse Matrix Collection. http:\/\/www.cise.ufl.edu\/research\/sparse\/matrices . NA Digest 97(23)"},{"key":"251_CR8","unstructured":"Geus R, R\u00f6llin S (1999) Towards a fast parallel sparse matrix-vector multiplication. In: Parallel computing: fundamentals and applications, international conference ParCo. Imperial College Press, 1999, pp 308\u2013315"},{"key":"251_CR9","volume-title":"Proceedings of parallel CFD\u201999","author":"W Gropp","year":"1999","unstructured":"Gropp W, Kaushik D, Keyes D, Smith B (1999) Toward realistic performance bounds for implicit cfd codes. In: Ecer A et al. (eds) Proceedings of parallel CFD\u201999. Elsevier, Amsterdam"},{"key":"251_CR10","unstructured":"Im E (2000) Optimizing the performance of sparse matrix-vector multiplication. PhD thesis, University of California, Berkeley"},{"key":"251_CR11","unstructured":"Im E, Yelick K (1999) Optimizing sparse matrix-vector multiplication on SMPs. In: 9th SIAM conference on parallel processing for scientific computing, SIAM, March 1999"},{"key":"251_CR12","doi-asserted-by":"crossref","unstructured":"Im E, Yelick K (2001) Optimizing sparse matrix computations for register reuse in SPARSITY. In: Lecture notes in computer science, vol\u00a02073, pp 127\u2013136","DOI":"10.1007\/3-540-45545-0_22"},{"key":"251_CR13","unstructured":"Kotakemori H, Hasegawa H, Kajiyama T, Nukada A, Suda R, Nishida A (2005) Performance evaluation of parallel sparse matrix-vector products on SGI Altix3700. In: 1st International workshop on OpenMP (IWOMP), Eugene, OR, USA, June 2005"},{"issue":"3","key":"251_CR14","doi-asserted-by":"crossref","first-page":"322","DOI":"10.1145\/263326.263382","volume":"15","author":"JL Lo","year":"1997","unstructured":"Lo JL, Eggers SJ, Emer JS, Levy HM, Stamm RL, Tullsen DM (1997) Converting thread-level parallelism to instruction-level parallelism via simultaneous multithreading. ACM Trans Comput Syst 15(3):322\u2013354","journal-title":"ACM Trans Comput Syst"},{"issue":"2","key":"251_CR15","doi-asserted-by":"crossref","first-page":"225","DOI":"10.1177\/1094342004038951","volume":"18","author":"J Mellor-Crummey","year":"2004","unstructured":"Mellor-Crummey J, Garvin J (2004) Optimizing sparse matrix-vector product computations using unroll and jam. Int J High Perform Comput Appl 18(2):225","journal-title":"Int J High Perform Comput Appl"},{"key":"251_CR16","unstructured":"Mitchell N, Carter L, Ferrante J, Tullsen D (1999) Instruction level parallelism vs. thread level parallelism on simultaneous multi-threading processors. In: Proceedings of supercomputing\u201999 (CD-ROM), Portland, OR, November 1999. ACM SIGARCH and IEEE"},{"issue":"4","key":"251_CR17","doi-asserted-by":"crossref","first-page":"703","DOI":"10.1007\/BF01932741","volume":"29","author":"GV Paolini","year":"1989","unstructured":"Paolini GV, Radicati di Brozolo G (1989) Data structures to vectorize CG algorithms for general sparsity patterns. BIT Numer Math 29(4):703\u2013718","journal-title":"BIT Numer Math"},{"key":"251_CR18","doi-asserted-by":"crossref","unstructured":"Pichel JC, Heras DB, Cabaleiro JC, Rivera FF (2004) Improving the locality of the sparse matrix-vector product on shared memory multiprocessors. In: PDP, IEEE Computer Society, 2004, pp 66\u201371","DOI":"10.1109\/EMPDP.2004.1271429"},{"issue":"8\u20139","key":"251_CR19","doi-asserted-by":"crossref","first-page":"858","DOI":"10.1016\/j.parco.2005.04.012","volume":"31","author":"JC Pichel","year":"2005","unstructured":"Pichel JC, Heras DB, Cabaleiro JC, Rivera FF (2005) Performance optimization of irregular codes based on the combination of reordering and blocking techniques. Parallel Comput 31(8\u20139):858\u2013876","journal-title":"Parallel Comput"},{"key":"251_CR20","doi-asserted-by":"crossref","unstructured":"Pinar A, Heath MT (1999) Improving performance of sparse matrix-vector multiplication. In: Supercomputing\u201999, Portland, OR, November 1999. ACM SIGARCH and IEEE","DOI":"10.1145\/331532.331562"},{"key":"251_CR21","unstructured":"Saad Y (1990) Sparskit: A basic tool kit for sparse matrix computation. Technical report, Center for Supercomputing Research and Development, University of Illinois at Urbana Champaign"},{"key":"251_CR22","doi-asserted-by":"crossref","DOI":"10.1137\/1.9780898718003","volume-title":"Iterative methods for sparse linear systems","author":"Y Saad","year":"2003","unstructured":"Saad Y (2003) Iterative methods for sparse linear systems. SIAM, Philadelphia"},{"key":"251_CR23","first-page":"578","volume-title":"Supercomputing\u201992","author":"O Temam","year":"1992","unstructured":"Temam O, Jalby W (1992) Characterizing the behavior of sparse algorithms on caches. In: Supercomputing\u201992, Minnesota, November 1992. IEEE, New York, pp 578\u2013587"},{"issue":"6","key":"251_CR24","doi-asserted-by":"crossref","first-page":"711","DOI":"10.1147\/rd.416.0711","volume":"41","author":"S Toledo","year":"1997","unstructured":"Toledo S (1997) Improving the memory-system performance of sparse-matrix vector multiplication. IBM J Res Dev 41(6):711\u2013725","journal-title":"IBM J Res Dev"},{"key":"251_CR25","doi-asserted-by":"crossref","unstructured":"Vuduc R, Demmel J, Yelick K, Kamil S, Nishtala R, Lee B (2002) Performance optimizations and bounds for sparse matrix-vector multiply. In: Supercomputing, Baltimore, MD, November, 2002","DOI":"10.1109\/SC.2002.10025"},{"key":"251_CR26","series-title":"Lecture notes in computer science","doi-asserted-by":"crossref","first-page":"807","DOI":"10.1007\/11557654_91","volume-title":"High performance computing and communications","author":"RW Vuduc","year":"2005","unstructured":"Vuduc RW, Moon H (2005) Fast sparse matrix-vector multiplication by exploiting variable block structure. In: High performance computing and communications. Lecture notes in computer science, vol 3726. Springer, Berlin, pp 807\u2013816"},{"key":"251_CR27","doi-asserted-by":"crossref","unstructured":"White J, Sadayappan P (1997) On improving the performance of sparse matrix-vector multiplication. In: 4th International conference on high performance computing (HiPC \u201997), 1997","DOI":"10.1109\/HIPC.1997.634472"},{"key":"251_CR28","doi-asserted-by":"crossref","first-page":"307","DOI":"10.1145\/1183401.1183444","volume-title":"ICS \u201906: Proceedings of the 20th annual international conference on supercomputing","author":"J Willcock","year":"2006","unstructured":"Willcock J, Lumsdaine A (2006) Accelerating sparse matrix computations via data compression. In: ICS \u201906: Proceedings of the 20th annual international conference on supercomputing, New York, NY, USA, 2006. ACM Press, New York, pp 307\u2013316"},{"key":"251_CR29","doi-asserted-by":"crossref","unstructured":"Williams S, Oilker L, Vuduc R, Shalf J, Yelick K, Demmel J (2007) Optimization of sparse matrix-vector multiplication on emerging multicore platforms. In: Supercomputing\u201907, Reno, NV, November 2007","DOI":"10.1145\/1362622.1362674"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-008-0251-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11227-008-0251-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-008-0251-8","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,6,1]],"date-time":"2019-06-01T10:23:57Z","timestamp":1559384637000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11227-008-0251-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,11,25]]},"references-count":29,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2009,10]]}},"alternative-id":["251"],"URL":"https:\/\/doi.org\/10.1007\/s11227-008-0251-8","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2008,11,25]]}}}