{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,5,27]],"date-time":"2023-05-27T03:40:13Z","timestamp":1685158813862},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2009,10,10]],"date-time":"2009-10-10T00:00:00Z","timestamp":1255132800000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2011,4]]},"DOI":"10.1007\/s11227-009-0340-3","type":"journal-article","created":{"date-parts":[[2009,10,9]],"date-time":"2009-10-09T09:37:27Z","timestamp":1255081047000},"page":"25-55","source":"Crossref","is-referenced-by-count":3,"title":["Region-based parallelization of irregular reductions on\u00a0explicitly managed memory hierarchies"],"prefix":"10.1007","volume":"56","author":[{"given":"Seonggun","family":"Kim","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hwansoo","family":"Han","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kwang-Moo","family":"Choe","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2009,10,10]]},"reference":[{"key":"340_CR1","unstructured":"Ahn JH, Erez\u00a0M, Dally WJ (2005) Scatter-Add in data parallel architectures. In: HPCA\u201905: international symposium on high-performance computer architecture, pp\u00a0132\u2013142"},{"key":"340_CR2","unstructured":"Arevalo\u00a0A, Matinata RM, Pandian\u00a0M, Peri\u00a0E, Ruby\u00a0K, Thomas\u00a0F, Almond\u00a0C. Programming the cell broadband engine architecture: examples and best practices. http:\/\/www.redbooks.ibm.com\/redbooks\/pdfs\/sg247575.pdf"},{"key":"340_CR3","unstructured":"Asanovic\u00a0K, Bodik\u00a0R, Catanzaro BC, Gebis JJ, Husbands\u00a0P, Keutzer\u00a0K, Patterson DA, Plishker WL, Shalf\u00a0J, Williams SW, Yelick KA (2006) The landscape of parallel computing research: a view from Berkeley. Tech Rep UCB\/EECS-2006-183, EECS Department, University of California, Berkeley"},{"key":"340_CR4","unstructured":"Balart\u00a0J, Duran\u00a0A, Gonzalez\u00a0M, Martorell\u00a0X, Ayguade\u00a0E, Labarta\u00a0J (2004) Nanos Mercurium: a research compiler for OpenMP. In: EWOMP\u201904: European workshop on OpenMP, pp\u00a0103\u2013109"},{"key":"340_CR5","unstructured":"Bellens\u00a0P, Perez JM, Badia RM, Labarta\u00a0J (2006) CellSs: a programming model for the Cell BE architecture. In: SC\u201906: ACM\/IEEE conference on supercomputing, p\u00a086"},{"key":"340_CR6","doi-asserted-by":"crossref","first-page":"187","DOI":"10.1002\/jcc.540040211","volume":"4","author":"B Brooks","year":"1983","unstructured":"Brooks\u00a0B, Bruccoleri\u00a0R, Olafson\u00a0D, States\u00a0D, Swaminathan\u00a0S, Karplus\u00a0M (1983) CHARMM: a program for macromolecular energy, minimization, and dynamics calculations. J\u00a0Comput Chem 4:187\u2013217","journal-title":"J\u00a0Comput Chem"},{"key":"340_CR7","volume-title":"LCPC\u201906: international workshop on languages and compilers for parallel computing","author":"T Chen","year":"2006","unstructured":"Chen\u00a0T, Sura\u00a0Z, O\u2019Brien\u00a0K, O\u2019Brien\u00a0J (2006) Optimizing the use of static buffers for DMA on a Cell chip. In: LCPC\u201906: international workshop on languages and compilers for parallel computing. Springer, Berlin"},{"key":"340_CR8","unstructured":"ClearSpeed. ClearSpeed whitepaper: CSX processor architecture. http:\/\/www.clearspeed.com\/docs\/resources\/ClearSpeed_Architecture_Whitepaper_Feb07v2.pdf"},{"key":"340_CR9","doi-asserted-by":"crossref","unstructured":"Ding\u00a0C, Kennedy\u00a0K (1999) Improving cache performance of dynamic applications with computation and data layout transformations. In: PLDI\u201999: ACM SIGPLAN conference on programming language design and implementation","DOI":"10.1145\/301618.301670"},{"key":"340_CR10","doi-asserted-by":"crossref","unstructured":"Eichenberger AE, O\u2019Brien\u00a0K, O\u2019Brien\u00a0K, Wu\u00a0P, Chen\u00a0T, Oden PH, Prener DA, Shepherd JC, So\u00a0B, Sura\u00a0Z, Wang\u00a0A, Zhang\u00a0T, Zhao\u00a0P, Gschwind\u00a0M (2005) Optimizing compiler for the cell processor. In: PACT \u201905: international conference on parallel architectures and compilation techniques, pp\u00a0161\u2013172","DOI":"10.1109\/PACT.2005.33"},{"issue":"1","key":"340_CR11","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1147\/sj.451.0059","volume":"45","author":"AE Eichenberger","year":"2006","unstructured":"Eichenberger AE, O\u2019Brien JK, O\u2019Brien KM, Wu\u00a0P, Chen\u00a0T, Oden PH, Prener DA, Shepherd JC, So\u00a0B, Sura\u00a0Z, Wang\u00a0A, Zhang\u00a0T, Zhao\u00a0P, Gschwind MK, Archambault\u00a0R, Gao\u00a0Y, Koo\u00a0R (2006) Using advanced compiler technology to exploit the performance of the Cell Broadband EngineTM architecture. IBM Syst\u00a0J 45(1):59\u201384","journal-title":"IBM Syst\u00a0J"},{"key":"340_CR12","first-page":"65","volume-title":"LCPC\u201991: workshop on languages and compilers for parallel computing","author":"R Eigenmann","year":"1991","unstructured":"Eigenmann\u00a0R, Hoeflinger\u00a0J, Li\u00a0Z, Padua\u00a0D (1991) Experience in the automatic parallelization of four perfect-benchmark programs. In: LCPC\u201991: workshop on languages and compilers for parallel computing. Springer, Berlin, pp\u00a065\u201383"},{"key":"340_CR13","unstructured":"Fatahalian\u00a0K, Horn DR, Knight TJ, Leem\u00a0L, Houston\u00a0M, Park JY, Erez\u00a0M, Ren\u00a0M, Aiken\u00a0A, Dally WJ, Hanrahan\u00a0P (2006) Sequoia: programming the memory hierarchy. In: SC\u201906: ACM\/IEEE conference on supercomputing, p\u00a083"},{"key":"340_CR14","doi-asserted-by":"crossref","unstructured":"Feautrier\u00a0P (1988) Array expansion. In: ICS\u201988: international conference on supercomputing, pp\u00a0429\u2013441","DOI":"10.1145\/55364.55406"},{"issue":"2\u20133","key":"340_CR15","doi-asserted-by":"crossref","first-page":"155","DOI":"10.1002\/cpe.769","volume":"16","author":"E Gutierrez","year":"2004","unstructured":"Gutierrez\u00a0E, Plata\u00a0O, Zapata\u00a0E (2004) Data partitioning-based parallel irregular reductions. Concurr Comput Pract Exp 16(2\u20133):155\u2013172","journal-title":"Concurr Comput Pract Exp"},{"issue":"3","key":"340_CR16","doi-asserted-by":"crossref","first-page":"133","DOI":"10.1016\/j.parco.2008.01.003","volume":"34","author":"E Guti\u00e9rrez","year":"2008","unstructured":"Guti\u00e9rrez\u00a0E, Plata\u00a0O, Zapata EL (2008) An analytical model of locality-based parallel irregular reductions. Parallel Comput 34(3):133\u2013157","journal-title":"Parallel Comput"},{"key":"340_CR17","doi-asserted-by":"crossref","first-page":"102","DOI":"10.1109\/ISCA.2004.1310767","volume-title":"ISCA\u201904: international symposium on computer architecture","author":"L Hammond","year":"2004","unstructured":"Hammond\u00a0L, Wong\u00a0V, Chen\u00a0M, Carlstrom BD, Davis JD, Hertzberg\u00a0B, Prabhu MK, Wijaya\u00a0H, Kozyrakis\u00a0C, Olukotun\u00a0K (2004) Transactional memory coherence and consistency. In: ISCA\u201904: international symposium on computer architecture. IEEE Computer Society, Los Alamitos, p\u00a0102"},{"issue":"7","key":"340_CR18","doi-asserted-by":"crossref","first-page":"606","DOI":"10.1109\/TPDS.2006.88","volume":"17","author":"H Han","year":"2006","unstructured":"Han\u00a0H, Tseng CW (2006) Exploiting locality for irregular scientific codes. IEEE Trans Parallel Distrib Syst 17(7):606\u2013618","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"340_CR19","unstructured":"Han\u00a0H, Tseng CW (2000) A\u00a0comparison of locality transformations for irregular codes. In: LCR\u201900: international workshop on languages, compilers, and run-time systems for scalable computers, pp\u00a070\u201384"},{"key":"340_CR20","doi-asserted-by":"crossref","unstructured":"Hofstee HP (2005) Power efficient processor architecture and the cell processor. In: HPCA\u201905: international symposium on high-performance computer architecture, pp\u00a0258\u2013262","DOI":"10.1109\/HPCA.2005.26"},{"key":"340_CR21","unstructured":"IBM: cell broadband engine programming handbook version 1.1"},{"key":"340_CR22","unstructured":"IBM:\u2009using\u2009the\u2009IBM\u2009XL\u2009C\/C++\u2009alpha\u2009edition\u2009for\u2009multicore\u2009acceleration\u2009single-source\u2009compiler. https:\/\/www-01.ibm.com\/chips\/techlib\/techlib.nsf\/techdocs\/C609359652E175AF00257353006E8063"},{"key":"340_CR23","unstructured":"IGC at ETH Zurich: Cell\/B.E. technology-based software. http:\/\/www-03.ibm.com\/technology\/cell\/software.html"},{"key":"340_CR24","unstructured":"Kahle\u00a0J. Cell architecture (presentation slides). http:\/\/www.power.org\/resources\/devcorner\/cellcorner\/CellTraining_Track1"},{"key":"340_CR25","doi-asserted-by":"crossref","unstructured":"Kodukula\u00a0I, Ahmed\u00a0N, Pingali\u00a0K (1997) Data-centric multi-level blocking. In: PLDI\u201997: ACM SIGPLAN conference on programming language design and implementation, pp\u00a0346\u2013357","DOI":"10.1145\/258915.258946"},{"key":"340_CR26","unstructured":"Lee SI, Johnson TA, Eigenmann\u00a0R (2003) Cetus\u2014an extensible compiler infrastructure for source-to-source transformation. In: LCPC\u201903: international workshop on languages and compilers for parallel computing, pp\u00a0539\u2013553"},{"key":"340_CR27","unstructured":"Li\u00a0Z (1992) Array privatization for parallel execution of loops. In: ICS\u201992: international conference on supercomputing, pp\u00a0313\u2013322"},{"key":"340_CR28","unstructured":"Lin\u00a0Y, Padua DA (1998) On the automatic parallelization of sparse and irregular Fortran programs. In: LCR\u201998: international workshop on languages, compilers, and run-time systems for scalable computers, pp\u00a041\u201356"},{"key":"340_CR29","doi-asserted-by":"crossref","unstructured":"Mellor-Crummey\u00a0J, Whalley\u00a0D, Kennedy\u00a0K (1999) Improving memory hierarchy performance for irregular applications. In: ICS\u201999: international conference on supercomputing","DOI":"10.1145\/305138.305228"},{"key":"340_CR30","doi-asserted-by":"crossref","unstructured":"Mirchandaney\u00a0R, Saltz JH, Smith RM, Nico DM, Crowley\u00a0K (1988) Principles of runtime support for parallel processors. In: ICS\u201988: international conference on supercomputing, pp\u00a0140\u2013152","DOI":"10.1145\/55364.55378"},{"key":"340_CR31","first-page":"192","volume-title":"PACT\u201999: international conference on parallel architectures and compilation techniques","author":"N Mitchell","year":"1999","unstructured":"Mitchell\u00a0N, Carter\u00a0L, Ferrante\u00a0J (1999) Localizing non-affine array references. In: PACT\u201999: international conference on parallel architectures and compilation techniques. IEEE Computer Society, Washington, p\u00a0192"},{"key":"340_CR32","unstructured":"NVIDIA. NVIDIA GeForce GTX 200 GPU architectural overview. http:\/\/www.nvidia.com\/object\/io_1213615494642.html"},{"key":"340_CR33","unstructured":"Poletto\u00a0M, Engler DR, Kaashoek MF (1996) tcc: a template-based compiler for \u2018C. In: WCSSS\u201996: workshop on compiler support for systems software, pp\u00a01\u20137"},{"key":"340_CR34","doi-asserted-by":"crossref","unstructured":"Schneider\u00a0S, Yeom JS, Rose\u00a0B, Linford JC, Sandu\u00a0A, Nikolopoulos DS (2009) A\u00a0comparison of programming models for multiprocessors with explicitly managed memory hierarchies. In: PPoPP\u201909: ACM SIGPLAN symposium on principles and practice of parallel programming, pp\u00a0131\u2013140","DOI":"10.1145\/1504176.1504197"},{"key":"340_CR35","doi-asserted-by":"crossref","unstructured":"Shavit\u00a0N, Touitou\u00a0D (1995) Software transactional memory. In: PODC\u201995: ACM symposium on principles of distributed computing, pp\u00a0204\u2013213","DOI":"10.1145\/224964.224987"},{"key":"340_CR36","doi-asserted-by":"crossref","unstructured":"Strout MM, Carter\u00a0L, Ferrante\u00a0J (2003) Compile-time composition of run-time data and iteration reorderings. In: PLDI\u201903: ACM SIGPLAN conference on programming language design and implementation","DOI":"10.1145\/781139.781142"},{"key":"340_CR37","unstructured":"The OpenMP architecture review board: the OpenMP API specification for parallel programming. http:\/\/openmp.org"},{"issue":"1","key":"340_CR38","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1145\/216585.216588","volume":"23","author":"WA Wulf","year":"1995","unstructured":"Wulf WA, McKee SA (1995) Hitting the memory wall: implications of the obvious. SIGARCH Comput Arch News 23(1):20\u201324","journal-title":"SIGARCH Comput Arch News"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-009-0340-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11227-009-0340-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-009-0340-3","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,27]],"date-time":"2023-05-27T03:03:22Z","timestamp":1685156602000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11227-009-0340-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,10,10]]},"references-count":38,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2011,4]]}},"alternative-id":["340"],"URL":"https:\/\/doi.org\/10.1007\/s11227-009-0340-3","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2009,10,10]]}}}