{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T10:19:34Z","timestamp":1778840374620,"version":"3.51.4"},"reference-count":32,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2013,7,20]],"date-time":"2013-07-20T00:00:00Z","timestamp":1374278400000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J Sign Process Syst"],"published-print":{"date-parts":[[2014,12]]},"DOI":"10.1007\/s11265-013-0812-9","type":"journal-article","created":{"date-parts":[[2013,7,19]],"date-time":"2013-07-19T04:57:25Z","timestamp":1374209845000},"page":"241-255","source":"Crossref","is-referenced-by-count":4,"title":["A Methodology for Speeding up MVM for Regular, Toeplitz and Bisymmetric Toeplitz Matrices"],"prefix":"10.1007","volume":"77","author":[{"given":"Vasilios I.","family":"Kelefouras","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Angeliki S.","family":"Kritikakou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Konstantinos","family":"Siourounis","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Costas E.","family":"Goutis","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2013,7,20]]},"reference":[{"key":"812_CR1","unstructured":"Intel core 2 duo processor e6550 http:\/\/ark.intel.com\/Product.aspx?id=30783 ."},{"key":"812_CR2","unstructured":"Ubuntu manuals http:\/\/manpages.ubuntu.com\/manpages\/lucid\/man1\/time.1.html ."},{"key":"812_CR3","unstructured":"Xilinx, virtex-5 fpga ml507 evaluation platform @ http:\/\/www.xilinx.com\/products\/boards-and-kits\/hw-v5-ml507-uni-g.htm ."},{"key":"812_CR4","unstructured":"Openblas (2012). Available at http:\/\/xianyi.github.com\/OpenBLAS\/ ."},{"issue":"2","key":"812_CR5","doi-asserted-by":"crossref","first-page":"830","DOI":"10.1007\/s11227-010-0474-3","volume":"59","author":"N Alachiotis","year":"2012","unstructured":"Alachiotis, N., Kelefouras, V.I., Athanasiou, G.S., Michail, H.E., Kritikakou, A.S., Goutis, C.E. (2012). A data locality methodology for matrix\u2014matrix multiplication algorithm. Journal of Supercomputing, 59(2), 830\u2013851. doi: 10.1007\/s11227-010-0474-3 .","journal-title":"Journal of Supercomputing"},{"key":"812_CR6","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1109\/2.982917","volume":"35","author":"T Austin","year":"2002","unstructured":"Austin, T., Larson, E., Ernst, D. (2002). Simplescalar: an infrastructure for computer system modeling. Computer, 35, 59\u201367. doi: 10.1109\/2.982917 . URL http:\/\/dl.acm.org\/citation.cfm?id=619072.621910 .","journal-title":"Computer"},{"key":"812_CR7","doi-asserted-by":"crossref","DOI":"10.1145\/263580.263662","volume-title":"Optimizing matrix multiply using PHiPAC: A portable, high-performance, ANSI C coding methodology. In Proceedings of the international conference on supercomputing","author":"J Bilmes","year":"1997","unstructured":"Bilmes, J., Asanovi\u0107, K., Chin, C., Demmel, J. (1997). Optimizing matrix multiply using PHiPAC: A portable, high-performance, ANSI C coding methodology. In Proceedings of the international conference on supercomputing. Vienna: ACM SIGARC."},{"key":"812_CR8","first-page":"134","volume-title":"Mv-ft: efficient implementation for matrix-vector multiplication on ft64 stream processor. In Proceedings of the second international conference on digital society, ICDS \u201908","author":"J Du","year":"2008","unstructured":"Du, J., Ao, F., Yang, X. (2008). Mv-ft: efficient implementation for matrix-vector multiplication on ft64 stream processor. In Proceedings of the second international conference on digital society, ICDS \u201908 (pp. 134\u2013139). Washington: IEEE Computer Society. doi: 10.1109\/ICDS.2008.16."},{"key":"812_CR9","doi-asserted-by":"crossref","unstructured":"Frigo, M., & Johnson, S.G. (1997). The fastest fourier transform in the west. Tech. rep., Cambridge.","DOI":"10.21236\/ADA479065"},{"key":"812_CR10","doi-asserted-by":"crossref","unstructured":"Fujimoto, N. (2008). Faster matrix-vector multiplication on geforce 8800gtx. In International parallel and distributed processing symposium\/international parallel processing symposium (pp. 1\u20138). doi: 10.1109\/IPDPS.2008.4536350 .","DOI":"10.1109\/IPDPS.2008.4536350"},{"key":"812_CR11","unstructured":"Guennebaud, G., Jacob, B., et al. (2010). Eigen v3. http:\/\/eigen.tuxfamily.org ."},{"key":"812_CR12","doi-asserted-by":"crossref","unstructured":"Garzia, F., Brunelli, C., Rossi, D., Nurmi, J. (2008). Implementation of a floating-point matrix-vector multiplication on a reconfigurable architecture. In IPDPS (pp. 1\u20136).","DOI":"10.1109\/IPDPS.2008.4536538"},{"key":"812_CR13","volume-title":"Matrix computations","author":"GH Golub","year":"1996","unstructured":"Golub, G.H., & Loan, C.F.V. (1996). Matrix computations, 3rd edn. Baltimore: The Johns Hopkins University Press.","edition":"3"},{"key":"812_CR14","doi-asserted-by":"crossref","first-page":"73","DOI":"10.1142\/S0129053395000051","volume":"7","author":"B Hendrickson","year":"1995","unstructured":"Hendrickson, B., Leland, R., Plimpton, S. (1995). An efficient parallel algorithm for matrix-vector multiplication. International Journal of High Speed Computing, 7, 73\u201388.","journal-title":"International Journal of High Speed Computing"},{"key":"812_CR15","unstructured":"Intel (2012). Intel mkl, Available at http:\/\/software.intel.com\/en-us\/intel-mkl ."},{"issue":"12","key":"812_CR16","doi-asserted-by":"crossref","first-page":"6217","DOI":"10.1109\/TSP.2011.2168525","volume":"59","author":"VI Kelefouras","year":"2011","unstructured":"Kelefouras, V.I., Athanasiou, G., Alachiotis, N., Michail, H.E., Kritikakou, A., Goutis, C.E. (2011). A methodology for speeding up fast fourier transform focusing on memory architecture utilization. IEEE Transactions on Signal Processing, 59(12), 6217\u20136226.","journal-title":"IEEE Transactions on Signal Processing"},{"key":"812_CR17","unstructured":"Krivutsenko, A. (2008). Gotoblas-anatomy of a fast matrix multiplication. Tech. rep."},{"issue":"1","key":"812_CR18","doi-asserted-by":"crossref","first-page":"1:1","DOI":"10.1145\/1509864.1509865","volume":"6","author":"PA Kulkarni","year":"2009","unstructured":"Kulkarni, P.A., Whalley, D.B., Tyson, G.S., Davidson, J.W. (2009). Practical exhaustive optimization phase order exploration and evaluation. TACO, 6(1), 1:1\u20131:36","journal-title":"TACO"},{"key":"812_CR19","unstructured":"van de Geijn, L. (1993). Distributed memory matrix-vector multiplication and conjugate gradient algorithms. SC Conference, 484\u2013492. http:\/\/doi.ieeecomputersociety.org\/10.1109\/SUPERC.1993.1263496 ."},{"issue":"2","key":"812_CR20","doi-asserted-by":"crossref","first-page":"15:1","DOI":"10.1145\/2159542.2159547","volume":"17","author":"P Milder","year":"2012","unstructured":"Milder, P., Franchetti, F., Hoe, J.C., P\u00fcschel, M. (2012). Computer generation of hardware for linear digital signal processing transforms. ACM Transactions on Design Automation of Electronic Systems, 17(2), 15:1\u201315:33. doi: 10.1145\/2159542.2159547 .","journal-title":"ACM Transactions on Design Automation of Electronic Systems"},{"issue":"6","key":"812_CR21","doi-asserted-by":"crossref","first-page":"89","DOI":"10.1145\/1273442.1250746","volume":"42","author":"N Nethercote","year":"2007","unstructured":"Nethercote, N., & Seward, J. (2007). Valgrind: a framework for heavyweight dynamic binary instrumentation. SIGPLAN Not, 42(6), 89\u2013100. doi: 10.1145\/1273442.1250746 .","journal-title":"SIGPLAN Not"},{"issue":"6","key":"812_CR22","doi-asserted-by":"crossref","first-page":"248","DOI":"10.1145\/173262.155114","volume":"28","author":"SS Pinter","year":"1996","unstructured":"Pinter, S.S. (1996). Register allocation with instruction scheduling. SIGPLAN Not, 28(6), 248\u2013257 doi: 10.1145\/173262.155114 .","journal-title":"SIGPLAN Not"},{"key":"812_CR23","volume-title":"Mips iv instruction set, revision 3.1.","author":"J PRICE","year":"1995","unstructured":"PRICE, J. (1995). Mips iv instruction set, revision 3.1. Tech. rep. Mountain View: MIPS Technologies, Inc."},{"key":"812_CR24","first-page":"57","volume-title":"Algorithms for smp-clusters dense matrix-vector multiplication. In Proceedings of the 16th international parallel and distributed processing symposium, IPDPS \u201902","author":"M Schmollinger","year":"2002","unstructured":"Schmollinger, M., & Kaufmann, M. (2002). Algorithms for smp-clusters dense matrix-vector multiplication. In Proceedings of the 16th international parallel and distributed processing symposium, IPDPS \u201902 (p. 57). Washington: IEEE Computer Society. http:\/\/portal.acm.org\/citation.cfm?id=645610.661893 ."},{"key":"812_CR25","first-page":"42","volume-title":"kappa numa: a model for clusters of smp machines. In Proceedings of the th international conference on parallel processing and applied mathematics-revised papers, PPAM \u201901","author":"M Schmollinger","year":"2002","unstructured":"Schmollinger, M., & Kaufmann, M. (2002). kappa numa: a model for clusters of smp-machines. In Proceedings of the th international conference on parallel processing and applied mathematics-revised papers, PPAM \u201901 (pp. 42\u201350). London: Springer. http:\/\/portal.acm.org\/citation.cfm?id=645813.668567 ."},{"key":"812_CR26","unstructured":"See homepage for details: Atlas homepage. http:\/\/math-atlas.sourceforge.net\/ ."},{"key":"812_CR27","unstructured":"Thoziyoor, D.T.S., Tarjan, D., Thoziyoor, S. (2006). Cacti 4.0. Tech. rep."},{"key":"812_CR28","unstructured":"Whaley, R.C., & Dongarra, J. (1997). Automatically tuned linear algebra software. Tech. Rep. UT-CS-97-366. University of Tennessee."},{"key":"812_CR29","unstructured":"Whaley, R.C., & Dongarra, J. (1998). Automatically tuned linear algebra software. In SuperComputing 1998: high performance networking and computing."},{"key":"812_CR30","unstructured":"Whaley, R.C., & Dongarra, J. (1999). Automatically tuned linear algebra software. In Ninth SIAM conference on parallel processing for scientific computing. CD-ROM Proceedings."},{"issue":"2","key":"812_CR31","first-page":"101","volume":"35","author":"RC Whaley","year":"2005","unstructured":"Whaley, R.C., & Petitet, A. (2005). Minimizing development and maintenance costs in supporting persistently optimized BLAS. Software: Practice and Experience, 35(2), 101\u2013121. http:\/\/www.cs.utsa.edu\/whaley\/papers\/spercw04.ps .","journal-title":"Software: Practice and Experience"},{"issue":"1\u20132","key":"812_CR32","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1016\/S0167-8191(00)00087-9","volume":"27","author":"RC Whaley","year":"2001","unstructured":"Whaley, R.C., Petitet, A., Dongarra, J.J. (2001). Automated empirical optimization of software and the ATLAS project. Parallel Computing , 27(1\u20132), 3\u201335. Also available as University of Tennessee LAPACK Working Note #147, UT-CS-00-448, 2000 ( www.netlib.org\/lapack\/lawns\/lawn147.ps ).","journal-title":"Parallel Computing"}],"container-title":["Journal of Signal Processing Systems"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11265-013-0812-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11265-013-0812-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11265-013-0812-9","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,7,18]],"date-time":"2019-07-18T21:51:18Z","timestamp":1563486678000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11265-013-0812-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,7,20]]},"references-count":32,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2014,12]]}},"alternative-id":["812"],"URL":"https:\/\/doi.org\/10.1007\/s11265-013-0812-9","relation":{},"ISSN":["1939-8018","1939-8115"],"issn-type":[{"value":"1939-8018","type":"print"},{"value":"1939-8115","type":"electronic"}],"subject":[],"published":{"date-parts":[[2013,7,20]]}}}