{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,8,23]],"date-time":"2023-08-23T05:09:28Z","timestamp":1692767368952},"reference-count":55,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2015,3,29]],"date-time":"2015-03-29T00:00:00Z","timestamp":1427587200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2015,7]]},"DOI":"10.1007\/s11227-015-1409-9","type":"journal-article","created":{"date-parts":[[2015,3,29]],"date-time":"2015-03-29T02:24:36Z","timestamp":1427595876000},"page":"2644-2667","update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["A methodology for speeding up matrix vector multiplication for single\/multi-core architectures"],"prefix":"10.1007","volume":"71","author":[{"given":"Vasilios","family":"Kelefouras","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Angeliki","family":"Kritikakou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elissavet","family":"Papadima","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Costas","family":"Goutis","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2015,3,29]]},"reference":[{"issue":"2","key":"1409_CR1","first-page":"101","volume":"35","author":"RC Whaley","year":"2005","unstructured":"Whaley RC, Petitet A (2005) Minimizing development and maintenance costs in supporting persistently optimized BLAS. Softw: Pract Exp 35(2):101\u2013121","journal-title":"Softw: Pract Exp"},{"key":"1409_CR2","unstructured":"OpenBlas (2012). http:\/\/xianyi.github.com\/OpenBLAS"},{"key":"1409_CR3","unstructured":"Krivutsenko A (2008) GotoBLAS\u2014anatomy of a fast matrix multiplication. Technical report"},{"key":"1409_CR4","unstructured":"Guennebaud G, Jacob B et al (2010) Eigen v3. http:\/\/eigen.tuxfamily.org"},{"key":"1409_CR5","unstructured":"Intel: Intel MKL (2012). http:\/\/software.intel.com\/en-us\/intel-mkl"},{"key":"1409_CR6","doi-asserted-by":"crossref","unstructured":"Bilmes J, Asanovi\u0107 K, Chin C, Demmel J (1997) Optimizing matrix multiply using PHiPAC: a portable, high-performance, ANSI C coding methodology. In: Proceedings of the international conference on supercomputing. ACM SIGARC, Vienna, Austria","DOI":"10.1145\/263580.263662"},{"key":"1409_CR7","doi-asserted-by":"crossref","unstructured":"Frigo M, Johnson SG (1997) The fastest Fourier transform in the west. Technical report. Cambridge, MA, USA","DOI":"10.21236\/ADA479065"},{"key":"1409_CR8","doi-asserted-by":"crossref","unstructured":"Milder P, Franchetti F, Hoe JC, P\u00fcschel M (2012) Computer generation of hardware for linear digital signal processing transforms. ACM Trans Des Autom Electron Syst 17(2) 15:1\u201315:33. doi: 10.1145\/2159542.2159547","DOI":"10.1145\/2159542.2159547"},{"key":"1409_CR9","unstructured":"Pinter SS (1996) Register allocation with instruction scheduling: a new approach. J Prog Lang 4(1):21\u201338"},{"key":"1409_CR10","doi-asserted-by":"crossref","unstructured":"Shobaki G, Shawabkeh M, Rmaileh NEA (2008) Preallocation instruction scheduling with register pressure minimization using a combinatorial optimization approach. ACM Trans Archit Code Optim 10(3):14:1\u201314:31. doi: 10.1145\/2512432","DOI":"10.1145\/2512432"},{"key":"1409_CR11","doi-asserted-by":"crossref","unstructured":"Bacon DF, Graham SL, Sharp OJ (1994) Compiler transformations for high-performance computing. ACM Comput Surv 26(4):345\u2013420. doi: 10.1145\/197405.197406","DOI":"10.1145\/197405.197406"},{"key":"1409_CR12","unstructured":"Granston E, Holler A (2001) Automatic recommendation of compiler options. In: Proceedings of the workshop on feedback-directed and dynamic optimization (FDDO)"},{"key":"1409_CR13","doi-asserted-by":"crossref","unstructured":"Triantafyllis S, Vachharajani M, Vachharajani N, August DI (2003) Compiler optimization-space exploration. In: Proceedings of the international symposium on code generation and optimization: feedback-directed and runtime optimization, CGO \u201903, pp 204\u2013215. IEEE Computer Society, Washington, DC, USA. http:\/\/dl.acm.org\/citation.cfm?id=776261.776284","DOI":"10.1109\/CGO.2003.1191546"},{"key":"1409_CR14","doi-asserted-by":"crossref","unstructured":"Cooper KD, Subramanian D, Torczon L (2002) Adaptive optimizing compilers for the 21st century. J Supercomput 23(1):7\u201322. doi: 10.1023\/A:1015729001611","DOI":"10.1023\/A:1015729001611"},{"key":"1409_CR15","doi-asserted-by":"crossref","unstructured":"Kisuki T, Knijnenburg PMW, O\u2019Boyle MFP, Bodin F, Wijshoff HAG (1999) A feasibility study in iterative compilation. In: Proceedings of the 2nd international symposium on high performance computing, ISHPC \u201999, pp 121\u2013132. Springer-Verlag, London, UK. http:\/\/dl.acm.org\/citation.cfm?id=646347.690219","DOI":"10.1007\/BFb0094916"},{"key":"1409_CR16","doi-asserted-by":"crossref","unstructured":"Kulkarni PA, Whalley DB, Tyson GS, Davidson JW (2009) Practical exhaustive optimization phase order exploration and evaluation. ACM Trans Archit Code Optim 6(1):1:1\u20131:36. doi: 10.1145\/1509864.1509865","DOI":"10.1145\/1509864.1509865"},{"issue":"6","key":"1409_CR17","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1145\/996893.996863","volume":"39","author":"P Kulkarni","year":"2004","unstructured":"Kulkarni P, Hines S, Hiser J, Whalley D, Davidson J, Jones D (2004) Fast searches for effective optimization phase sequences. SIGPLAN Not 39(6):171\u2013182. doi: 10.1145\/996893.996863","journal-title":"SIGPLAN Not"},{"key":"1409_CR18","doi-asserted-by":"crossref","unstructured":"Park E, Kulkarni S, Cavazos J (2011) An evaluation of different modeling techniques for iterative compilation. In: Proceedings of the 14th international conference on compilers, architectures and synthesis for embedded systems, CASES \u201911, pp 65\u201374. ACM, New York, NY, USA. doi: 10.1145\/2038698.2038711","DOI":"10.1145\/2038698.2038711"},{"key":"1409_CR19","doi-asserted-by":"crossref","unstructured":"Monsifrot A, Bodin F, Quiniou R (2002) A machine learning approach to automatic production of compiler heuristics. In: Proceedings of the 10th international conference on artificial intelligence: methodology, systems, and applications, AIMSA \u201902, pp 41\u201350. Springer-Verlag, London, UK. http:\/\/dl.acm.org\/citation.cfm?id=646053.677574","DOI":"10.1007\/3-540-46148-5_5"},{"key":"1409_CR20","doi-asserted-by":"crossref","unstructured":"Stephenson M, Amarasinghe S, Martin M, O\u2019Reilly UM (2003) Meta optimization: improving compiler heuristics with machine learning. SIGPLAN Not 38(5):77\u201390 (2003). doi: 10.1145\/780822.781141","DOI":"10.1145\/780822.781141"},{"key":"1409_CR21","doi-asserted-by":"crossref","unstructured":"Tartara M, Crespi Reghizzi S (2013) Continuous learning of compiler heuristics. ACM Trans Archit Code Optim 9(4):46:1\u201346:25. doi: 10.1145\/2400682.2400705","DOI":"10.1145\/2400682.2400705"},{"key":"1409_CR22","doi-asserted-by":"crossref","unstructured":"Agakov F, Bonilla E, Cavazos J, Franke B, Fursin G, O\u2019Boyle MFP, Thomson J, Toussaint M, Williams CKI (2006) Using machine learning to focus iterative optimization. In: Proceedings of the international symposium on code generation and optimization, CGO \u201906, pp 295\u2013305. IEEE Computer Society, Washington, DC, USA. doi: 10.1109\/CGO.2006.37","DOI":"10.1109\/CGO.2006.37"},{"key":"1409_CR23","doi-asserted-by":"crossref","unstructured":"Nethercote N, Seward J (2007) Valgrind: a framework for heavyweight dynamic binary instrumentation. SIGPLAN Not 42(6):89\u2013100. doi: 10.1145\/1273442.1250746","DOI":"10.1145\/1273442.1250746"},{"key":"1409_CR24","unstructured":"Simplescalar CI, Burger D, Austin TM (1997) The SimpleScalar tool set, version 2.0. Technical report"},{"issue":"1\u20132","key":"1409_CR25","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1016\/S0167-8191(00)00087-9","volume":"27","author":"RC Whaley","year":"2001","unstructured":"Whaley RC, Petitet A, Dongarra JJ (2001) Automated empirical optimization of software and the ATLAS project. Parallel Comput 27(1\u20132):3\u201335","journal-title":"Parallel Comput"},{"key":"1409_CR26","doi-asserted-by":"crossref","unstructured":"Whaley RC, Dongarra J (1999) Automatically tuned linear algebra software. In: 9th SIAM conference on parallel processing for scientific computing. CD-ROM Proceedings","DOI":"10.1109\/SC.1998.10004"},{"key":"1409_CR27","doi-asserted-by":"crossref","unstructured":"Whaley RC, Dongarra J (1998) Automatically tuned linear algebra software. In: SuperComputing 1998: high performance networking and computing","DOI":"10.1109\/SC.1998.10004"},{"key":"1409_CR28","doi-asserted-by":"crossref","unstructured":"Whaley RC, Dongarra J (1997) Automatically tuned linear algebra software. Technical report. UT-CS-97-366, University of Tennessee","DOI":"10.1109\/SC.1998.10004"},{"key":"1409_CR29","unstructured":"See homepage for details: ATLAS homepage (2012). http:\/\/math-atlas.sourceforge.net\/"},{"issue":"4","key":"1409_CR30","doi-asserted-by":"crossref","first-page":"511","DOI":"10.1142\/S0129626408003545","volume":"18","author":"N Fujimoto","year":"2008","unstructured":"Fujimoto N (2008) Dense matrix\u2013vector multiplication on the CUDA architecture. Parallel Process Lett 18(4):511\u2013530","journal-title":"Parallel Process Lett"},{"key":"1409_CR31","doi-asserted-by":"crossref","unstructured":"Fujimoto N (2008) Faster matrix\u2013vector multiplication on GeForce 8800GTX. In: IPDPS, pp 1\u20138. IEEE. http:\/\/dblp.uni-trier.de\/db\/conf\/ipps\/ipdps2008.html","DOI":"10.1109\/IPDPS.2008.4536350"},{"key":"1409_CR32","doi-asserted-by":"crossref","first-page":"73","DOI":"10.1142\/S0129053395000051","volume":"7","author":"B Hendrickson","year":"1995","unstructured":"Hendrickson B, Leland R, Plimpton S (1995) An efficient parallel algorithm for matrix\u2013vector multiplication. Int J High Speed Comput 7:73\u201388","journal-title":"Int J High Speed Comput"},{"key":"1409_CR33","doi-asserted-by":"crossref","unstructured":"S\u00f8rensen HHB (2012) High-performance matrix\u2013vector multiplication on the GPU. In: Proceedings of the 2011 international conference on parallel processing, Euro-Par\u201911, pp 377\u2013386. Springer-Verlag, Berlin, Heidelberg","DOI":"10.1007\/978-3-642-29737-3_42"},{"issue":"3","key":"1409_CR34","doi-asserted-by":"crossref","first-page":"397","DOI":"10.1109\/TPDS.2011.174","volume":"23","author":"N Zhang","year":"2012","unstructured":"Zhang N (2012) A novel parallel scan for multicore processors and its application in sparse matrix\u2013vector multiplication. IEEE Trans Parallel Distrib Syst 23(3):397\u2013404. doi: 10.1109\/TPDS.2011.174","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"1409_CR35","doi-asserted-by":"crossref","unstructured":"Williams S, Oliker L, Vuduc R, Shalf J, Yelick K, Demmel J (2007) Optimization of sparse matrix\u2013vector multiplication on emerging multicore platforms. In: Proceedings of the 2007 ACM\/IEEE conference on supercomputing, SC \u201907, pp 38:1\u201338:12. ACM, New York, NY, USA. doi: 10.1145\/1362622.1362674","DOI":"10.1145\/1362622.1362674"},{"key":"1409_CR36","doi-asserted-by":"crossref","unstructured":"Goumas G, Kourtis K, Anastopoulos N, Karakasis V, Koziris N (2009) Performance evaluation of the sparse matrix\u2013vector multiplication on modern architectures. J Supercomput 50(1):36\u201377. doi: 10.1007\/s11227-008-0251-8","DOI":"10.1007\/s11227-008-0251-8"},{"key":"1409_CR37","doi-asserted-by":"crossref","unstructured":"Michailidis PD, Margaritis KG (2010) Performance models for matrix computations on multicore processors using OpenMP. In: Proceedings of the 2010 international conference on parallel and distributed computing. Applications and Technologies, PDCAT \u201910, pp 375\u2013380. IEEE Computer Society, Washington, DC, USA. doi: 10.1109\/PDCAT.2010.52","DOI":"10.1109\/PDCAT.2010.52"},{"key":"1409_CR38","doi-asserted-by":"crossref","unstructured":"Schmollinger M, Kaufmann M (2002) Algorithms for SMP-clusters dense matrix\u2013vector multiplication. In: Proceedings of the 16th international parallel and distributed processing Sysmposium, IPDPS \u201902, pp 57\u2013. IEEE Computer Society, Washington, DC, USA. http:\/\/dl.acm.org\/citation.cfm?id=645610.661893","DOI":"10.1109\/IPDPS.2002.1016540"},{"key":"1409_CR39","doi-asserted-by":"crossref","unstructured":"Waghmare VN, Kendre SV, Chordiya SG (2011) Article: performance analysis of matrix\u2013vector multiplication in hybrid (MPI + OpenMP). Int J Comput Appl 22(5):22\u201325. Published by Foundation of Computer Science","DOI":"10.5120\/2579-3561"},{"key":"1409_CR40","doi-asserted-by":"crossref","unstructured":"Baker AH, Schulz M, Yang UM (2011) On the performance of an algebraic multigrid solver on multicore clusters. In: Proceedings of the 9th international conference on high performance computing for computational science, VECPAR\u201910, pp 102\u2013115. Springer-Verlag, Berlin, Heidelberg. http:\/\/dl.acm.org\/citation.cfm?id=1964238.1964252","DOI":"10.1007\/978-3-642-19328-6_12"},{"key":"1409_CR41","unstructured":"Parallel methods for matrix\u2013vector multiplication. http:\/\/www.hpcc.unn.ru\/mskurs\/ENG\/DOC\/pp07.pdf"},{"issue":"11","key":"1409_CR42","doi-asserted-by":"crossref","first-page":"1783","DOI":"10.1016\/0167-8191(95)00032-9","volume":"21","author":"SM Bhandarkar","year":"1995","unstructured":"Bhandarkar SM, Arabnia HR (1995) The REFINE multiprocessor\u2014theoretical properties and algorithms. Parallel Comput 21(11):1783\u20131805","journal-title":"Parallel Comput"},{"key":"1409_CR43","unstructured":"Arabnia HR, Smith JW (1993) A reconfigurable interconnection network for imaging operations and its implementation using a multi-stage switching box. pp 349\u2013357"},{"issue":"1","key":"1409_CR44","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1023\/A:1022804606389","volume":"25","author":"MA Wani","year":"2003","unstructured":"Wani MA, Arabnia HR (2003) Parallel edge-region-based segmentation algorithm targeted at reconfigurable MultiRing network. J Supercomput 25(1):43\u201362","journal-title":"J Supercomput"},{"issue":"2","key":"1409_CR45","doi-asserted-by":"crossref","first-page":"188","DOI":"10.1016\/0743-7315(90)90028-N","volume":"10","author":"HR Arabnia","year":"1990","unstructured":"Arabnia HR (1990) A parallel algorithm for the arbitrary rotation of digitized images using process-and-data-decomposition approach. J Parallel Distrib Comput 10(2):188\u2013192","journal-title":"J Parallel Distrib Comput"},{"issue":"1","key":"1409_CR46","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1111\/j.1467-8659.1989.tb00448.x","volume":"8","author":"HR Arabnia","year":"1989","unstructured":"Arabnia HR, Oliver MA (1989) A transputer network for fast operations on digitised images. Comput Graph Forum 8(1):3\u201311","journal-title":"Comput Graph Forum"},{"issue":"1","key":"1409_CR47","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1006\/jpdc.1995.1011","volume":"24","author":"SM Bhandarkar","year":"1995","unstructured":"Bhandarkar SM, Arabnia HR (1995) The Hough transform on a reconfigurable multi-ring network. J Parallel Distrib Comput 24(1):107\u2013114","journal-title":"J Parallel Distrib Comput"},{"issue":"5","key":"1409_CR48","doi-asserted-by":"crossref","first-page":"425","DOI":"10.1093\/comjnl\/30.5.425","volume":"30","author":"HR Arabnia","year":"1987","unstructured":"Arabnia HR, Oliver MA (1987) A transputer network for the arbitrary rotation of digitised images. Comput J 30(5):425\u2013432","journal-title":"Comput J"},{"issue":"3","key":"1409_CR49","doi-asserted-by":"crossref","first-page":"243","DOI":"10.1007\/BF00130109","volume":"10","author":"HR Arabnia","year":"1996","unstructured":"Arabnia HR, Bhandarkar SM (1996) Parallel stereocorrelation on a reconfigurable multi-ring network. J Supercomput 10(3):243\u2013269","journal-title":"J Supercomput"},{"issue":"1","key":"1409_CR50","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1111\/j.1467-8659.1987.tb00340.x","volume":"6","author":"HR Arabnia","year":"1987","unstructured":"Arabnia HR, Oliver MA (1987) Arbitrary rotation of raster images with SIMD machine architectures. Comput Graph Forum 6(1):3\u201311","journal-title":"Comput Graph Forum"},{"issue":"02","key":"1409_CR51","doi-asserted-by":"crossref","first-page":"201","DOI":"10.1142\/S0218001495000110","volume":"9","author":"SM Bhandarkar","year":"1995","unstructured":"Bhandarkar SM, Arabnia HR, Smith JW (1995) A reconfigurable architecture for image processing and computer vision. Int J Pattern Recognit Artif Intell 9(02):201\u2013229","journal-title":"Int J Pattern Recognit Artif Intell"},{"key":"1409_CR52","doi-asserted-by":"crossref","unstructured":"Arabnia H (1995) A distributed stereocorrelation algorithm. In: Computer communications and networks, 1995. Proceedings, 4th international conference on, pp 479\u2013482, IEEE","DOI":"10.1109\/ICCCN.1995.540163"},{"key":"1409_CR53","unstructured":"Intel core 2 duo processor E6550. http:\/\/ark.intel.com\/Product.aspx?id=30783"},{"key":"1409_CR54","unstructured":"Intel core 2 duo processor T6600. http:\/\/ark.intel.com\/products\/37255\/Intel-Core2-Duo-Processor-T6600"},{"key":"1409_CR55","unstructured":"Intel i7-2600K Processor. http:\/\/ark.intel.com\/products\/52214"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-015-1409-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11227-015-1409-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-015-1409-9","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,8,22]],"date-time":"2019-08-22T11:37:49Z","timestamp":1566473869000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11227-015-1409-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,3,29]]},"references-count":55,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2015,7]]}},"alternative-id":["1409"],"URL":"https:\/\/doi.org\/10.1007\/s11227-015-1409-9","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2015,3,29]]}}}