{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,10,1]],"date-time":"2023-10-01T17:59:58Z","timestamp":1696183198017},"reference-count":47,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2013,4,5]],"date-time":"2013-04-05T00:00:00Z","timestamp":1365120000000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2013,10]]},"DOI":"10.1007\/s11227-013-0916-9","type":"journal-article","created":{"date-parts":[[2013,4,4]],"date-time":"2013-04-04T05:30:15Z","timestamp":1365053415000},"page":"431-487","source":"Crossref","is-referenced-by-count":7,"title":["High-performance optimizations on tiled many-core embedded systems: a matrix multiplication case study"],"prefix":"10.1007","volume":"66","author":[{"given":"Arslan","family":"Munir","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Farinaz","family":"Koushanfar","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ann","family":"Gordon-Ross","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sanjay","family":"Ranka","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2013,4,5]]},"reference":[{"key":"916_CR1","volume-title":"Proc of the 15th international Euro-Par conference on parallel processing (Euro-Par\u201909)","author":"N Yuan","year":"2009","unstructured":"Yuan\u00a0N, Zhou\u00a0Y, Tan\u00a0G, Zhang\u00a0J, Fan\u00a0D (2009) High performance matrix multiplication on many cores. In: Proc of the 15th international Euro-Par conference on parallel processing (Euro-Par\u201909), Delft, The Netherlands, August 2009"},{"key":"916_CR2","unstructured":"MAXIMUMPC (2007) Fast forward: multicore vs manycore. June. Available online: http:\/\/www.maximumpc.com\/article\/fast_forward_multicore_vs_manycore"},{"key":"916_CR3","unstructured":"Wikipedia (2013) Multi-core processor. February. Available online: http:\/\/en.wikipedia.org\/wiki\/Manycore"},{"key":"916_CR4","unstructured":"Tilera (2013) Tilera cloud computing. February. Available online: http:\/\/www.tilera.com\/solutions\/cloud_computing"},{"key":"916_CR5","unstructured":"Tilera (2013) Tilera TILEmpower platform. February. Available online: http:\/\/www.tilera.com\/sites\/default\/files\/productbriefs\/TILEProEmpower_PB021_v4.pdf"},{"issue":"3","key":"916_CR6","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1109\/MM.2009.41","volume":"29","author":"M Levy","year":"2009","unstructured":"Levy\u00a0M, Conte\u00a0T (2009) Embedded multicore processors and systems. IEEE MICRO 29(3):7\u20139","journal-title":"IEEE MICRO"},{"issue":"10","key":"916_CR7","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1145\/1562764.1562783","volume":"52","author":"K Asanovic","year":"2009","unstructured":"Asanovic K, Bodik R, Demmel J, Keaveny T, Keutzer K, Kubiatowicz J, Morgan N, Patterson K, Sen D, Wawrzynek J, Wessel D, Yelick K (2009) A view of the parallel computing landscape. Commun ACM 52(10):56\u201367","journal-title":"Commun ACM"},{"key":"916_CR8","volume-title":"Proc of ACM 3rd conference on computing frontiers (CF)","author":"Jd Cuvillo","year":"2006","unstructured":"Cuvillo Jd, Zhu\u00a0W, Gao GR (2006) Landing OpenMP on Cyclops-64: an efficient mapping of OpenMP to a many-core system-on-a-chip. In: Proc of ACM 3rd conference on computing frontiers (CF), Ischia, Italy, May 2006"},{"issue":"1","key":"916_CR9","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1109\/JSSC.2007.910957","volume":"43","author":"SR Vangal","year":"2008","unstructured":"Vangal SR, Howard\u00a0J, Ruhl\u00a0G, Dighe\u00a0S, Wilson\u00a0H, Tschanz\u00a0J, Finan\u00a0D, Singh\u00a0A, Jacob\u00a0T, Jain\u00a0S, Erraguntla\u00a0V, Roberts\u00a0C, Hoskote\u00a0Y, Borkar\u00a0N, Borkar\u00a0S (2008) An 80-tile sub-100-W TeraFLOPS processor in 65-nm CMOS. IEEE J Solid-State Circuits 43(1):29\u201341","journal-title":"IEEE J Solid-State Circuits"},{"issue":"3","key":"916_CR10","doi-asserted-by":"crossref","DOI":"10.1145\/1698772.1698782","volume":"9","author":"E Musoll","year":"2010","unstructured":"Musoll\u00a0E (2010) A cost-effective load-balancing policy for tile-based, massive multi-core packet processors. ACM Trans Embedded Comput Syst 9(3):24","journal-title":"ACM Trans Embedded Comput Syst"},{"key":"916_CR11","doi-asserted-by":"crossref","first-page":"274","DOI":"10.1007\/978-3-642-24568-8_14","volume-title":"Transactions on high-performance embedded architectures and compilers IV (HiPEAC IV)","author":"N Wu","year":"2011","unstructured":"Wu N, Yang Q, Wen M, He Y, Ren J, Guan M, Zhang C (2011) Tiled multi-core stream architecture. In: Transactions on high-performance embedded architectures and compilers IV (HiPEAC IV), vol 4, pp 274\u2013293"},{"key":"916_CR12","volume-title":"Proc of IEEE\/ACM conference on supercomputing (SC)","author":"TG Mattson","year":"2008","unstructured":"Mattson TG, Wijngaart RVd, Frumkin\u00a0M (2008) Programming the Intel 80-core network-on-a-chip terascale processor. In: Proc of IEEE\/ACM conference on supercomputing (SC), Austin, Texas, November 2008"},{"key":"916_CR13","unstructured":"Crowell\u00a0T (2011) Will 2011 mark the beginning of manycore? January. Available online: http:\/\/talbottcrowell.wordpress.com\/2011\/01\/01\/manycore\/"},{"key":"916_CR14","unstructured":"Tilera (2012) Manycore without boundaries: TILEPro64 processor. May. Available online: http:\/\/www.tilera.com\/products\/processors\/TILEPRO64"},{"key":"916_CR15","first-page":"349","volume-title":"Lecture notes in computer science","author":"R Brown","year":"2008","unstructured":"Brown R, Sharapov I (2008) Performance and programmability comparison between OpenMP and MPI implementations of a molecular modeling application. In: Lecture notes in computer science, vol 4315. Springer, Berlin, pp 349\u2013360"},{"issue":"11","key":"916_CR16","doi-asserted-by":"crossref","first-page":"1185","DOI":"10.1109\/71.476190","volume":"6","author":"X Sun","year":"1995","unstructured":"Sun\u00a0X, Zhu\u00a0J (1995) Performance considerations of shared virtual memory machines. IEEE Trans Parallel Distrib Syst 6(11):1185\u20131194","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"916_CR17","unstructured":"Cortesi\u00a0D (1998) Origin2000 and Onyx2 performance tuning and optimization guide. Available online: http:\/\/techpubs.sgi.com\/library\/dynaweb_docs\/0640\/SGI_Developer\/books\/OrOn2_PfTune\/sgi_html\/index.html"},{"key":"916_CR18","volume-title":"Proc of the international parallel and distributed processing symposium (IPDPS)","author":"M Krishnan","year":"2004","unstructured":"Krishnan\u00a0M, Nieplocha\u00a0J (2004) SRUMMA: a matrix multiplication algorithm suitable for clusters and scalable shared memory systems. In: Proc of the international parallel and distributed processing symposium (IPDPS), Santa Fe, New Mexico, April 2004"},{"key":"916_CR19","first-page":"44","volume-title":"Proc of the ACM international conference on supercomputing (ICS)","author":"H-J Lee","year":"1997","unstructured":"Lee H-J, Robertson JP, Fortes\u00a0J (1997) Generalized Cannon\u2019s algorithm for parallel matrix multiplication. In: Proc of the ACM international conference on supercomputing (ICS), Vienna, Austria, July 1997, pp 44\u201351"},{"key":"916_CR20","unstructured":"van de Geijn RA, Watts\u00a0J (1995) Summa: scalable universal matrix multiplication algorithm. University of Texas at Austin, Tech rep. Available online: http:\/\/www.ncstrl.org:8900\/ncstrl\/servlet\/search?formname=detail&id=oai%3Ancstrlh%3Autexas_cs%3AUTEXAS_CS%2F%2FCS-TR-95-13"},{"key":"916_CR21","volume-title":"Handbook on multicore computing","author":"J Li","year":"2012","unstructured":"Li\u00a0J, Ranka\u00a0S, Sahni\u00a0S (2012) GPU matrix multiplication. In: Rajasekaran\u00a0S (ed) Handbook on multicore computing. CRC Press, Boca Raton"},{"key":"916_CR22","unstructured":"More\u00a0A (2008) A case study on high performance matrix multiplication. Available online: mm-matrixmultiplicationtool.googlecode.com\/files\/mm.pdf"},{"issue":"4","key":"916_CR23","doi-asserted-by":"crossref","first-page":"352","DOI":"10.1145\/98267.98290","volume":"16","author":"N Higham","year":"1990","unstructured":"Higham\u00a0N (1990) Exploiting fast matrix multiplication within the level 3 BLAS. ACM Trans Math Softw 16(4):352\u2013368","journal-title":"ACM Trans Math Softw"},{"issue":"3","key":"916_CR24","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1356052.1356053","volume":"34","author":"K Goto","year":"2008","unstructured":"Goto\u00a0K, Geijn\u00a0R (2008) Anatomy of high-performance matrix multiplication. ACM Trans Math Softw 34(3):1\u201325","journal-title":"ACM Trans Math Softw"},{"key":"916_CR25","unstructured":"Nishtala\u00a0R, Vuduc RW, Demmel JW, Yelick KA (2004) Performance modeling and analysis of cache blocking in sparse matrix vector multiply. Tech rep UCB\/CSD-04-1335, EECS Department, University of California, Berkeley. Available online: http:\/\/www.eecs.berkeley.edu\/Pubs\/TechRpts\/2004\/5535.html"},{"key":"916_CR26","doi-asserted-by":"crossref","first-page":"63","DOI":"10.1145\/106972.106981","volume-title":"Proc of the fourth ACM international conference on architectural support for programming languages and operating systems (ASPLOS)","author":"MD Lam","year":"1991","unstructured":"Lam MD, Rothberg EE, Wolf ME (1991) The cache performance and optimizations of blocked algorithms. In: Proc of the fourth ACM international conference on architectural support for programming languages and operating systems (ASPLOS), Santa Clara, California, April 1991, pp 63\u201374"},{"key":"916_CR27","volume-title":"Stream processor architecture","author":"S Rixner","year":"2002","unstructured":"Rixner\u00a0S (2002) Stream processor architecture. Kluwer Academic, Norwell"},{"key":"916_CR28","volume-title":"Proc of the 2005 and 2006 international conference on OpenMP shared memory parallel programming (IWOMP\u201905\/IWOMP\u201906)","author":"W Zhu","year":"2005","unstructured":"Zhu\u00a0W, Cuvillo Jd, Gao GR (2005) Performance characteristics of OpenMP language constructs on a many-core-on-a-chip architecture. In: Proc of the 2005 and 2006 international conference on OpenMP shared memory parallel programming (IWOMP\u201905\/IWOMP\u201906), Eugene, Oregon, June 2005"},{"key":"916_CR29","volume-title":"Proc of the ACM Euro-Par conference on parallel processing","author":"E Garcia","year":"2010","unstructured":"Garcia\u00a0E, Venetis\u00a0I, Khan\u00a0R, Gao\u00a0G (2010) Optimized dense matrix multiplication on a many-core architecture. In: Proc of the ACM Euro-Par conference on parallel processing"},{"key":"916_CR30","first-page":"1","volume-title":"Proc of the IEEE aerospace conference","author":"S Safari","year":"2012","unstructured":"Safari\u00a0S, Fijany\u00a0A, Diotalevi\u00a0F, Hosseini\u00a0F (2012) Highly parallel and fast implementation of stereo vision algorithms on MIMD many-core Tilera architecture. In: Proc of the IEEE aerospace conference, Boston, MA, August 2012, pp 1\u201311"},{"key":"916_CR31","volume-title":"Proc of the IEEE international performance computing and communications conference (IPCCC)","author":"A Munir","year":"2012","unstructured":"Munir\u00a0A, Gordon-Ross\u00a0A, Ranka\u00a0S (2012) Parallelized benchmark-driven performance evaluation of SMPs and tiled multi-core architectures for embedded systems. In: Proc of the IEEE international performance computing and communications conference (IPCCC), Austin, Texas, December 2012"},{"key":"916_CR32","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4419-0263-4","volume-title":"Multicore processors and systems","author":"S Keckler","year":"2009","unstructured":"Keckler\u00a0S, Olukotun\u00a0K, Hofstee\u00a0H (2009) Multicore processors and systems. Springer, Berlin"},{"key":"916_CR33","unstructured":"Tilera (2012) Manycore without boundaries: TILE64 processor. April. Available online: http:\/\/www.tilera.com\/products\/processors\/TILE64"},{"key":"916_CR34","unstructured":"Intel (2013) Intel\u2019s teraflops research chip. February. Available online: http:\/\/download.intel.com\/pressroom\/kits\/Teraflops\/Teraflops_Research_Chip_Overview.pdf"},{"issue":"5","key":"916_CR35","doi-asserted-by":"crossref","first-page":"51","DOI":"10.1109\/MM.2007.4378783","volume":"27","author":"Y Hoskote","year":"2007","unstructured":"Hoskote\u00a0Y, Vangal\u00a0S, Singh\u00a0A, Borkar\u00a0N, Borkar\u00a0S (2007) A 5-GHz mesh interconnect for a TeraFLOPS processor. IEEE MICRO 27(5):51\u201361","journal-title":"IEEE MICRO"},{"key":"916_CR36","unstructured":"IBM (2012) Linux and Symmetric Multiprocessing, February. Available online: http:\/\/www.ibm.com\/developerworks\/library\/l-linux-smp\/"},{"key":"916_CR37","volume-title":"Tilera official documentation","author":"Tilera","year":"2009","unstructured":"Tilera (2009) Tile processor architecture overview for the TILEPro series. In: Tilera official documentation. November"},{"key":"916_CR38","volume-title":"Tilera official documentation","author":"Tilera","year":"2010","unstructured":"Tilera (2010) Multicore development environment system programmer\u2019s guide. In: Tilera official documentation. March"},{"key":"916_CR39","volume-title":"Tilera official documentation","author":"Tilera","year":"2009","unstructured":"Tilera (2009) Tile processor architecture overview. In: Tilera official documentation. November"},{"key":"916_CR40","volume-title":"Introduction to parallel computing","author":"V Kumar","year":"1994","unstructured":"Kumar\u00a0V, Grama\u00a0A, Gupta\u00a0A, Karypis\u00a0G (1994) Introduction to parallel computing. Benjamin-Cummings, Redwood City"},{"key":"916_CR41","volume-title":"Tilera official documentation","author":"Tilera","year":"2010","unstructured":"Tilera (2010) Multicore development environment optimization guide. In: Tilera official documentation. March"},{"key":"916_CR42","unstructured":"ARM (2012) Cortex-A15 MPCore: technical reference manual. April. Available online: http:\/\/infocenter.arm.com\/help\/topic\/com.arm.doc.ddi0438e\/DDI0438E_cortex_a15_r3p0_trm.pdf"},{"key":"916_CR43","unstructured":"Oracle (2013) Sun studio 12: Fortran programming guide. February. Available online: http:\/\/docs.oracle.com\/cd\/E19205-01\/819-5262\/aeuic\/index.html"},{"key":"916_CR44","volume-title":"Proc of 20th annual IEEE international conference on parallel processing (ICPP)","author":"S Mahlke","year":"1991","unstructured":"Mahlke\u00a0S, Warter\u00a0N, Chen\u00a0W, Chang\u00a0P, Hwu W-m (1991) The effect of compiler optimizations on available parallelism in scalar programs. In: Proc of 20th annual IEEE international conference on parallel processing (ICPP), Austin, Texas, August 1991"},{"key":"916_CR45","doi-asserted-by":"crossref","unstructured":"Williams\u00a0J, Massie\u00a0C, George\u00a0A, Richardson\u00a0J, Gosrani\u00a0K, Lam\u00a0H (2010) Characterization of fixed and reconfigurable multi-core devices for application acceleration. ACM Trans on Reconfigurable Technology and Systems 3(4)","DOI":"10.1145\/1862648.1862649"},{"key":"916_CR46","volume-title":"Tilera official documentation","author":"Tilera","year":"2010","unstructured":"Tilera (2010) TILEmPower appliance user\u2019s guide. In: Tilera official documentation. January"},{"key":"916_CR47","volume-title":"Tilera official documentation","author":"Tilera","year":"2009","unstructured":"Tilera (2009) Tilera multicore development environment: iLib API reference manual. In: Tilera official documentation. April"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-013-0916-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11227-013-0916-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-013-0916-9","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,7,11]],"date-time":"2019-07-11T16:27:33Z","timestamp":1562862453000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11227-013-0916-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,4,5]]},"references-count":47,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2013,10]]}},"alternative-id":["916"],"URL":"https:\/\/doi.org\/10.1007\/s11227-013-0916-9","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2013,4,5]]}}}