{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:24:40Z","timestamp":1750307080154,"version":"3.41.0"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2013,4,1]],"date-time":"2013-04-01T00:00:00Z","timestamp":1364774400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities of China","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61003002"],"award-info":[{"award-number":["61003002"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003399","name":"Science and Technology Commission of Shanghai Municipality","doi-asserted-by":"publisher","award":["12QA1401700"],"award-info":[{"award-number":["12QA1401700"]}],"id":[{"id":"10.13039\/501100003399","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2013,4]]},"abstract":"<jats:p>The prevalence of chip multiprocessors opens opportunities of running data-parallel applications originally in clusters on a single machine with many cores. MapReduce, a simple and elegant programming model to program large-scale clusters, has recently been shown a promising alternative to harness the multicore platform.<\/jats:p>\n          <jats:p>\n            The differences such as memory hierarchy and communication patterns between clusters and multicore platforms raise new challenges to design and implement an efficient MapReduce system on multicore. This article argues that it is more efficient for MapReduce to iteratively process small chunks of data in turn than processing a large chunk of data at a time on shared memory multicore platforms. Based on the argument, we extend the general MapReduce programming model with a \u201ctiling strategy\u201d, called\n            <jats:italic>Tiled<\/jats:italic>\n            -\n            <jats:italic>MapReduce<\/jats:italic>\n            (TMR). TMR partitions a large MapReduce job into a number of small subjobs and iteratively processes one subjob at a time with efficient use of resources; TMR finally merges the results of all subjobs for output. Based on Tiled-MapReduce, we design and implement several optimizing techniques targeting multicore, including the reuse of the input buffer among subjobs, a NUCA\/NUMA-aware scheduler, and pipelining a subjob\u2019s reduce phase with the successive subjob\u2019s map phase, to optimize the memory, cache, and CPU resources accordingly. Further, we demonstrate that Tiled-MapReduce supports fine-grained fault tolerance and enables several usage scenarios such as online and incremental computing on multicore machines.\n          <\/jats:p>\n          <jats:p>Performance evaluation with our prototype system called Ostrich on a 48-core machine shows that Ostrich saves up to 87.6% memory, causes less cache misses, and makes more efficient use of CPU cores, resulting in a speedup ranging from 1.86x to 3.07x over Phoenix. Ostrich also efficiently supports fine-grained fault tolerance, online, and incremental computing with small performance penalty.<\/jats:p>","DOI":"10.1145\/2445572.2445575","type":"journal-article","created":{"date-parts":[[2013,4,9]],"date-time":"2013-04-09T12:17:58Z","timestamp":1365509878000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Tiled-MapReduce"],"prefix":"10.1145","volume":"10","author":[{"given":"Rong","family":"Chen","sequence":"first","affiliation":[{"name":"Shanghai Jiao Tong University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haibo","family":"Chen","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2013,4]]},"reference":[{"unstructured":"Ahmad F. Lee S. Thottethodi M. and Vijaykumar T. N. 2011a. Mapreduce benchmarks. http:\/\/web.ics.purdue.edu\/~fahmad\/benchmarks.htm.  Ahmad F. Lee S. Thottethodi M. and Vijaykumar T. N. 2011a. Mapreduce benchmarks. http:\/\/web.ics.purdue.edu\/~fahmad\/benchmarks.htm.","key":"e_1_2_1_1_1"},{"unstructured":"Ahmad F. Lee S. Thottethodi M. and Vijaykumar T. N. 2011b. MapReduce with communication overlap (MaRCO). Tech. rep. ECE-TR-413 Electrical and Computer Engineering Purdue University.  Ahmad F. Lee S. Thottethodi M. and Vijaykumar T. N. 2011b. MapReduce with communication overlap (MaRCO). Tech. rep. ECE-TR-413 Electrical and Computer Engineering Purdue University.","key":"e_1_2_1_2_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_3_1","DOI":"10.1145\/2150976.2150984"},{"unstructured":"AMD 2013. Codeanalyst performance analyzer. http:\/\/developer.amd.com\/tools\/codeanalyst\/pages\/default.aspx.  AMD 2013. Codeanalyst performance analyzer. http:\/\/developer.amd.com\/tools\/codeanalyst\/pages\/default.aspx.","key":"e_1_2_1_4_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_5_1","DOI":"10.1145\/1966445.1966472"},{"doi-asserted-by":"publisher","key":"e_1_2_1_6_1","DOI":"10.1145\/1807128.1807150"},{"doi-asserted-by":"publisher","key":"e_1_2_1_7_1","DOI":"10.1145\/2038916.2038923"},{"volume-title":"Hadoop: A framework for running applications on large clusters built of commodity hardware","year":"2005","author":"Bialecki A.","key":"e_1_2_1_8_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_9_1","DOI":"10.1145\/227234.227246"},{"doi-asserted-by":"publisher","key":"e_1_2_1_10_1","DOI":"10.1145\/1278480.1278667"},{"unstructured":"Borthakur D. 2013. The hadoop distributed file system: Architecture and design. http:\/\/hadoop.apache.org\/hdfs\/docs\/current\/hdfs_design.html.  Borthakur D. 2013. The hadoop distributed file system: Architecture and design. http:\/\/hadoop.apache.org\/hdfs\/docs\/current\/hdfs_design.html.","key":"e_1_2_1_11_1"},{"volume-title":"Proceedings of the 8th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201908)","author":"Boyd-Wickizer S.","key":"e_1_2_1_12_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_13_1","DOI":"10.1145\/1854273.1854337"},{"volume-title":"Proceedings of Neural Information Processing Systems Conference (NIPS\u201906)","author":"Chu C. T.","key":"e_1_2_1_14_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_15_1","DOI":"10.1145\/207110.207162"},{"volume-title":"Proceedings of the 7th USENIX Conference on Networked Systems Design and Implementation (NSDI\u201910)","author":"Condie T.","key":"e_1_2_1_16_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_17_1","DOI":"10.1145\/1327452.1327492"},{"unstructured":"de Kruijf M. and Sankaralingam K. 2007. MapReduce for the cell be architecture. Tech. rep. CS-TR-2007 Computer Sciences University of Wisconsin.  de Kruijf M. and Sankaralingam K. 2007. MapReduce for the cell be architecture. Tech. rep. CS-TR-2007 Computer Sciences University of Wisconsin.","key":"e_1_2_1_18_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_19_1","DOI":"10.1109\/eScience.2008.59"},{"doi-asserted-by":"publisher","key":"e_1_2_1_20_1","DOI":"10.1145\/1851476.1851593"},{"unstructured":"Evans J. 2013. jemalloc. http:\/\/www.canonware.com\/jemalloc\/.  Evans J. 2013. jemalloc. http:\/\/www.canonware.com\/jemalloc\/.","key":"e_1_2_1_21_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_22_1","DOI":"10.1145\/277650.277725"},{"volume-title":"Matrix Computations","author":"Golub G.","doi-asserted-by":"crossref","key":"e_1_2_1_23_1","DOI":"10.56021\/9781421407944"},{"doi-asserted-by":"publisher","key":"e_1_2_1_24_1","DOI":"10.1109\/69.273032"},{"unstructured":"Gray J. 2013. http:\/\/www.hpl.hp.com\/hosted\/sortbenchmark.  Gray J. 2013. http:\/\/www.hpl.hp.com\/hosted\/sortbenchmark.","key":"e_1_2_1_25_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_26_1","DOI":"10.1145\/1454115.1454152"},{"doi-asserted-by":"publisher","key":"e_1_2_1_27_1","DOI":"10.1145\/253260.253291"},{"doi-asserted-by":"publisher","key":"e_1_2_1_28_1","DOI":"10.1145\/1854273.1854303"},{"doi-asserted-by":"publisher","key":"e_1_2_1_29_1","DOI":"10.1145\/1272996.1273005"},{"unstructured":"Khronos Group. 2009. Opencl overview. http:\/\/www.khronos.org\/opencl\/.  Khronos Group. 2009. Opencl overview. http:\/\/www.khronos.org\/opencl\/.","key":"e_1_2_1_30_1"},{"volume-title":"Proceedings of the 24th International Workshop on Languages and Compilers for Parallel Computing (LCPC\u201911)","author":"Kim J.","key":"e_1_2_1_31_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_32_1","DOI":"10.1145\/1854273.1854301"},{"doi-asserted-by":"publisher","key":"e_1_2_1_33_1","DOI":"10.1109\/ICDCS.2011.26"},{"unstructured":"Levon J. 2004. OProfile Manual. Victoria University of Manchester. http:\/\/oprofile.sourceforge.net\/doc\/.  Levon J. 2004. OProfile Manual . Victoria University of Manchester. http:\/\/oprofile.sourceforge.net\/doc\/.","key":"e_1_2_1_34_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_35_1","DOI":"10.1145\/1346281.1346318"},{"volume-title":"Proceedings of the USENIX Conference on USENIX Annual Technical Conference (USENIX ATC\u201911)","author":"Logothetis D.","key":"e_1_2_1_36_1"},{"unstructured":"Mao Y. Morris R. and Kaashoek F. 2010. Optimizing MapReduce for Multicore Architectures. Tech. rep. MIT-CSAIL-TR-2010-020 Computer Science and Artificial Intelligence Laboratory Massachusetts Institute of Technology.  Mao Y. Morris R. and Kaashoek F. 2010. Optimizing MapReduce for Multicore Architectures. Tech. rep. MIT-CSAIL-TR-2010-020 Computer Science and Artificial Intelligence Laboratory Massachusetts Institute of Technology.","key":"e_1_2_1_37_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_38_1","DOI":"10.1109\/eScience.2008.62"},{"doi-asserted-by":"publisher","key":"e_1_2_1_39_1","DOI":"10.1145\/319151.319160"},{"doi-asserted-by":"publisher","key":"e_1_2_1_40_1","DOI":"10.1145\/1376616.1376726"},{"doi-asserted-by":"publisher","key":"e_1_2_1_41_1","DOI":"10.1109\/ICPADS.2009.143"},{"doi-asserted-by":"publisher","key":"e_1_2_1_42_1","DOI":"10.1155\/2005\/962135"},{"volume-title":"Proceedings of the Conference on Hot Topics in Cloud Computing (HotCloud\u201909)","author":"Popa L.","key":"e_1_2_1_43_1"},{"volume-title":"Proceedings of the 9th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201910)","author":"Power R.","key":"e_1_2_1_44_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_45_1","DOI":"10.1109\/HPCA.2007.346181"},{"doi-asserted-by":"publisher","key":"e_1_2_1_46_1","DOI":"10.1145\/1723112.1723129"},{"doi-asserted-by":"publisher","key":"e_1_2_1_47_1","DOI":"10.1145\/1966445.1966452"},{"doi-asserted-by":"publisher","key":"e_1_2_1_48_1","DOI":"10.1145\/1851476.1851597"},{"doi-asserted-by":"publisher","key":"e_1_2_1_49_1","DOI":"10.1145\/1996092.1996095"},{"doi-asserted-by":"publisher","key":"e_1_2_1_50_1","DOI":"10.1145\/1272996.1273004"},{"doi-asserted-by":"publisher","key":"e_1_2_1_51_1","DOI":"10.14778\/1687553.1687609"},{"doi-asserted-by":"publisher","key":"e_1_2_1_52_1","DOI":"10.1145\/79173.79181"},{"doi-asserted-by":"publisher","key":"e_1_2_1_53_1","DOI":"10.1109\/2.612254"},{"unstructured":"Wikimedia. 2013. Downloads. http:\/\/dumps.wikimedia.org\/.  Wikimedia. 2013. Downloads. http:\/\/dumps.wikimedia.org\/.","key":"e_1_2_1_54_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_55_1","DOI":"10.1109\/PACT.2011.22"},{"volume-title":"Proceedings of the IEEE International Conference on Data Engineering (ICDE\u201910)","author":"Yang C.","key":"e_1_2_1_56_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_57_1","DOI":"10.1145\/1247480.1247602"},{"doi-asserted-by":"publisher","key":"e_1_2_1_58_1","DOI":"10.1109\/IISWC.2009.5306783"},{"volume-title":"Proceedings of the 8th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201908)","author":"Yu Y.","key":"e_1_2_1_59_1"},{"volume-title":"Proceedings of the 2nd USENIX Conference on Hot Topics in Cloud Computing (HotCloud\u201910)","author":"Zaharia M.","key":"e_1_2_1_60_1"},{"volume-title":"Proceedings of the 8th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201908)","author":"Zaharia M.","key":"e_1_2_1_61_1"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2445572.2445575","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2445572.2445575","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T09:34:09Z","timestamp":1750239249000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2445572.2445575"}},"subtitle":["Efficient and Flexible MapReduce Processing on Multicore with Tiling"],"short-title":[],"issued":{"date-parts":[[2013,4]]},"references-count":61,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2013,4]]}},"alternative-id":["10.1145\/2445572.2445575"],"URL":"https:\/\/doi.org\/10.1145\/2445572.2445575","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2013,4]]},"assertion":[{"value":"2012-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2012-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-04-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}