{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,18]],"date-time":"2025-11-18T15:34:15Z","timestamp":1763480055150,"version":"3.41.0"},"reference-count":21,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2015,6,2]],"date-time":"2015-06-02T00:00:00Z","timestamp":1433203200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGMETRICS Perform. Eval. Rev."],"published-print":{"date-parts":[[2015,6,2]]},"abstract":"<jats:p>Cloud computing offers a new, attractive option to customers for quickly provisioning any size Hadoop cluster, consuming resources as a service, executing their MapReduce workload, and then paying for the time these resources were used. One of the open questions in such environments is the right choice of resources (and their amount) a user should lease from the service provider. Typically, there is a variety of different types of VM instances in the Cloud (e.g., small, medium, or large EC2 instances). The capacity differences of the offered VMs are reflected in VM's pricing. Therefore, for the same price a user can get a variety of Hadoop clusters based on different VM instance types. We observe that the performance of MapReduce applications may vary significantly on different platforms. This makes a selection of the best cost\/performance platform for a given workload a non-trivial problem, especially when it contains multiple jobs with different platform preferences. We aim to solve the following problem: given a completion time target for a set of MapReduce jobs, determine a homogeneous or heterogeneous Hadoop cluster configuration (i.e., the number, types of VMs, and the job schedule) for processing these jobs within a given deadline while minimizing the rented infrastructure cost. In this work,1 we design an efficient and fast simulation-based framework for evaluating and selecting the right underlying platform for achieving the desirable Service Level Objectives (SLOs). Our evaluation study with Amazon EC2 platform reveals that for different workload mixes, an optimized platform choice may result in 45-68% cost savings for achieving the same performance objectives when using different (but seemingly equivalent) choices. Moreover, depending on a workload the heterogeneous solution may outperform the homogeneous cluster solution by 26-42%. We provide additional insights explaining the obtained results by profiling the performance characteristics of used applications and underlying EC2 platforms. The results of our simulation study are validated through experiments with Hadoop clusters deployed on different Amazon EC2 instances.<\/jats:p>","DOI":"10.1145\/2788402.2788409","type":"journal-article","created":{"date-parts":[[2015,6,3]],"date-time":"2015-06-03T15:35:55Z","timestamp":1433345755000},"page":"38-50","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":26,"title":["Exploiting Cloud Heterogeneity to Optimize Performance and Cost of MapReduce Processing"],"prefix":"10.1145","volume":"42","author":[{"given":"Zhuoyao","family":"Zhang","sequence":"first","affiliation":[{"name":"Google Inc., Mountain View, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ludmila","family":"Cherkasova","sequence":"additional","affiliation":[{"name":"HewlettPackard Labs, Palo Alto, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Boon Thau","family":"Loo","sequence":"additional","affiliation":[{"name":"University of Pennsylvania, Philadelphia, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2015,6,2]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Apache Rumen: a tool to extract job characterization data from job tracker logs. https:\/\/issues.apache.org\/jira\/browse\/MAPREDUCE-728.  Apache Rumen: a tool to extract job characterization data from job tracker logs. https:\/\/issues.apache.org\/jira\/browse\/MAPREDUCE-728."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2150976.2150984"},{"key":"e_1_2_1_3_1","unstructured":"Apache. Mumak: Map-Reduce Simulator. https:\/\/issues.apache.org\/jira\/browse\/ MAPREDUCE-751.  Apache. Mumak: Map-Reduce Simulator. https:\/\/issues.apache.org\/jira\/browse\/ MAPREDUCE-751."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2038916.2038934"},{"key":"e_1_2_1_5_1","volume-title":"Proc. of 5th Conf. on Innovative Data Systems Research (CIDR)","author":"Herodotou H.","year":"2011","unstructured":"H. Herodotou , H. Lim , G. Luo , N. Borisov , L. Dong , F. Cetin , and S. Babu . Starfish: A Self-tuning System for Big Data Analytics . In Proc. of 5th Conf. on Innovative Data Systems Research (CIDR) , 2011 . H. Herodotou, H. Lim, G. Luo, N. Borisov, L. Dong, F. Cetin, and S. Babu. Starfish: A Self-tuning System for Big Data Analytics. In Proc. of 5th Conf. on Innovative Data Systems Research (CIDR), 2011."},{"key":"e_1_2_1_6_1","doi-asserted-by":"crossref","unstructured":"S. Johnson. Optimal Two- and Three-Stage Production Schedules with Setup Times Included. Naval Res. Log. Quart. 1954.  S. Johnson. Optimal Two- and Three-Stage Production Schedules with Setup Times Included. Naval Res. Log. Quart. 1954.","DOI":"10.1002\/nav.3800010110"},{"key":"e_1_2_1_7_1","volume-title":"Proc. of the First Workshop on Hot Topics in Cloud Computing","author":"Kambatla K.","year":"2009","unstructured":"K. Kambatla , A. Pathak , and H. Pucha . Towards optimizing hadoop provisioning in the cloud . In Proc. of the First Workshop on Hot Topics in Cloud Computing , 2009 . K. Kambatla, A. Pathak, and H. Pucha. Towards optimizing hadoop provisioning in the cloud. In Proc. of the First Workshop on Hot Topics in Cloud Computing, 2009."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1989323.1989355"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2010.73"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1462704.1462713"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLOUD.2011.14"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2011.36"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-25821-3_9"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1998582.1998637"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/MASCOTS.2012.12"},{"key":"e_1_2_1_16_1","volume-title":"Intl. Symposium on Modelling, Analysis and Simulation of Computer and Telecommunication Systems (MASCOTS)","author":"Wang G.","year":"2009","unstructured":"G. Wang , A. Butt , P. Pandey , and K. Gupta . A Simulation Approach to Evaluating Design Decisions in MapReduce Setups . In Intl. Symposium on Modelling, Analysis and Simulation of Computer and Telecommunication Systems (MASCOTS) , 2009 . G. Wang, A. Butt, P. Pandey, and K. Gupta. A Simulation Approach to Evaluating Design Decisions in MapReduce Setups. In Intl. Symposium on Modelling, Analysis and Simulation of Computer and Telecommunication Systems (MASCOTS), 2009."},{"key":"e_1_2_1_18_1","volume-title":"Proc. of the IPDPS Workshops: Heterogeneity in Computing","author":"Xie J.","year":"2010","unstructured":"J. Xie Improving mapreduce performance through data placement in heterogeneous hadoop clusters . In Proc. of the IPDPS Workshops: Heterogeneity in Computing , 2010 . J. Xie et al. Improving mapreduce performance through data placement in heterogeneous hadoop clusters. In Proc. of the IPDPS Workshops: Heterogeneity in Computing, 2010."},{"key":"e_1_2_1_19_1","volume-title":"Proc. of OSDI","author":"Zaharia M.","year":"2008","unstructured":"M. Zaharia Improving mapreduce performance in heterogeneous environments . In Proc. of OSDI , 2008 . M. Zaharia et al. Improving mapreduce performance in heterogeneous environments. In Proc. of OSDI, 2008."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2592784.2592785"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2371536.2371546"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/NOMS.2014.6838231"}],"container-title":["ACM SIGMETRICS Performance Evaluation Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2788402.2788409","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2788402.2788409","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T05:07:44Z","timestamp":1750223264000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2788402.2788409"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,6,2]]},"references-count":21,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2015,6,2]]}},"alternative-id":["10.1145\/2788402.2788409"],"URL":"https:\/\/doi.org\/10.1145\/2788402.2788409","relation":{},"ISSN":["0163-5999"],"issn-type":[{"type":"print","value":"0163-5999"}],"subject":[],"published":{"date-parts":[[2015,6,2]]},"assertion":[{"value":"2015-06-02","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}