{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:51:51Z","timestamp":1750308711454,"version":"3.41.0"},"reference-count":20,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2013,9,1]],"date-time":"2013-09-01T00:00:00Z","timestamp":1377993600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["CNS-1117185, CNS-0845552, IIS-0812270, CCF-0964471"],"award-info":[{"award-number":["CNS-1117185, CNS-0845552, IIS-0812270, CCF-0964471"]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000144","name":"Division of Computer and Network Systems","doi-asserted-by":"publisher","award":["CNS-1117185, CNS-0845552, IIS-0812270, CCF-0964471"],"award-info":[{"award-number":["CNS-1117185, CNS-0845552, IIS-0812270, CCF-0964471"]}],"id":[{"id":"10.13039\/100000144","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000145","name":"Division of Information and Intelligent Systems","doi-asserted-by":"publisher","award":["CNS-1117185, CNS-0845552, IIS-0812270, CCF-0964471"],"award-info":[{"award-number":["CNS-1117185, CNS-0845552, IIS-0812270, CCF-0964471"]}],"id":[{"id":"10.13039\/100000145","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Auton. Adapt. Syst."],"published-print":{"date-parts":[[2013,9]]},"abstract":"<jats:p>\n            Many applications associated with live business intelligence are written as complex data analysis programs defined by directed acyclic graphs of MapReduce jobs, for example, using Pig, Hive, or Scope frameworks. An increasing number of these applications have additional requirements for completion time guarantees. In this article, we consider the popular Pig framework that provides a high-level SQL-like abstraction on top of MapReduce engine for processing large data sets. There is a lack of performance models and analysis tools for automated performance management of such MapReduce jobs. We offer a performance modeling environment for Pig programs that automatically profiles jobs from the past runs and aims to solve the following inter-related problems: (i) estimating the completion time of a Pig program as a function of allocated resources; (ii) estimating the amount of resources (a number of map and reduce slots) required for completing a Pig program with a given (soft) deadline. First, we design a\n            <jats:italic>basic<\/jats:italic>\n            performance model that accurately predicts completion time and required resource allocation for a Pig program that is defined as a sequence of MapReduce jobs: predicted completion times are within 10% of the measured ones. Second, we optimize a Pig program execution by enforcing the\n            <jats:italic>optimal schedule<\/jats:italic>\n            of its concurrent jobs. For DAGs with concurrent jobs, this optimization helps reducing the program completion time: 10%--27% in our experiments. Moreover, it eliminates possible nondeterminism of concurrent jobs\u2019 execution in the Pig program, and therefore, enables a more accurate performance model for Pig programs. Third, based on these optimizations, we propose a\n            <jats:italic>refined<\/jats:italic>\n            performance model for Pig programs with concurrent jobs. The proposed approach leads to significant resource savings (20%--60% in our experiments) compared with the original, unoptimized solution. We validate our solution using a 66-node Hadoop cluster and a diverse set of workloads: PigMix benchmark, TPC-H queries, and customized queries mining a collection of HP Labs\u2019 web proxy logs.\n          <\/jats:p>","DOI":"10.1145\/2518017.2518019","type":"journal-article","created":{"date-parts":[[2013,10,1]],"date-time":"2013-10-01T18:14:28Z","timestamp":1380651268000},"page":"1-28","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Performance Modeling and Optimization of Deadline-Driven Pig Programs"],"prefix":"10.1145","volume":"8","author":[{"given":"Zhuoyao","family":"Zhang","sequence":"first","affiliation":[{"name":"HP Labs and University of Pennsylvania"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ludmila","family":"Cherkasova","sequence":"additional","affiliation":[{"name":"Hewlett-Packard Labs"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Abhishek","family":"Verma","sequence":"additional","affiliation":[{"name":"HP Labs and University of Illinois at Urbana-Champaign"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Boon Thau","family":"Loo","sequence":"additional","affiliation":[{"name":"University of Pennsylvania"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2013,9]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Apache. 2010. PigMix Benchmark. http:\/\/wiki.apache.org\/pig\/PigMix.  Apache. 2010. PigMix Benchmark. http:\/\/wiki.apache.org\/pig\/PigMix."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.14778\/1454159.1454166"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1327452.1327492"},{"volume-title":"Proceedings of the 5th International Workshop on Self Managing Database Systems (SMDB).","author":"Ganapathi A.","key":"e_1_2_1_4_1","unstructured":"Ganapathi , A. , Chen , Y. , Fox , A. , Katz , R. , and Patterson , D . 2010. Statistics-driven workload modeling for the cloud . In Proceedings of the 5th International Workshop on Self Managing Database Systems (SMDB). Ganapathi, A., Chen, Y., Fox, A., Katz, R., and Patterson, D. 2010. Statistics-driven workload modeling for the cloud. In Proceedings of the 5th International Workshop on Self Managing Database Systems (SMDB)."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687553.1687568"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.14778\/3402707.3402746"},{"volume-title":"Proceedings of the 5th Conference on Innovative Data Systems Research (CIDR).","author":"Herodotou H.","key":"e_1_2_1_7_1","unstructured":"Herodotou , H. , Lim , H. , Luo , G. , Borisov , N. , Dong , L. , Cetin , F. , and Babu , S . 2011. Starfish: A self-tuning system for big data analytics . In Proceedings of the 5th Conference on Innovative Data Systems Research (CIDR). Herodotou, H., Lim, H., Luo, G., Borisov, N., Dong, L., Cetin, F., and Babu, S. 2011. Starfish: A self-tuning system for big data analytics. In Proceedings of the 5th Conference on Innovative Data Systems Research (CIDR)."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1272998.1273005"},{"key":"e_1_2_1_9_1","doi-asserted-by":"crossref","unstructured":"Johnson S. M. 1954. Optimal two- and three-stage production schedules with setup times included. Naval Res. Log. Quart.  Johnson S. M. 1954. Optimal two- and three-stage production schedules with setup times included. Naval Res. Log. Quart.","DOI":"10.1002\/nav.3800010110"},{"volume-title":"Proceedings of the 1st Workshop on Hot Topics in Cloud Computing.","author":"Kambatla K.","key":"e_1_2_1_10_1","unstructured":"Kambatla , K. , Pathak , A. , and Pucha , H . 2009. Towards optimizing hadoop provisioning in the cloud . In Proceedings of the 1st Workshop on Hot Topics in Cloud Computing. Kambatla, K., Pathak, A., and Pucha, H. 2009. Towards optimizing hadoop provisioning in the cloud. In Proceedings of the 1st Workshop on Hot Topics in Cloud Computing."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807167.1807223"},{"volume-title":"Proceedings of ICDE.","author":"Morton K.","key":"e_1_2_1_12_1","unstructured":"Morton , K. , Friesen , A. , Balazinska , M. , and Grossman , D . 2010b. Estimating the progress of MapReduce pipelines . In Proceedings of ICDE. Morton, K., Friesen, A., Balazinska, M., and Grossman, D. 2010b. Estimating the progress of MapReduce pipelines. In Proceedings of ICDE."},{"volume-title":"Proceedings of the 12th IEEE\/IFIP Network Operations and Management Symposium.","author":"Polo J.","key":"e_1_2_1_13_1","unstructured":"Polo , J. , Carrera , D. , Becerra , Y. , Torres , J. , Ayguad\u00e9 , E. , Steinder , M. , and Whalley , I . 2010. Performance-driven task co-scheduling for MapReduce environments . In Proceedings of the 12th IEEE\/IFIP Network Operations and Management Symposium. Polo, J., Carrera, D., Becerra, Y., Torres, J., Ayguad\u00e9, E., Steinder, M., and Whalley, I. 2010. Performance-driven task co-scheduling for MapReduce environments. In Proceedings of the 12th IEEE\/IFIP Network Operations and Management Symposium."},{"volume-title":"Proc. of VLDB.","author":"Thusoo A.","key":"e_1_2_1_14_1","unstructured":"Thusoo , A. , Sarma , J. S. , Jain , N. , Shao , Z. , Chakka , P. , Anthony , S. , Liu , H. , Wyckoff , P. , and Murthy , R . 2009. Hive - a warehousing solution over a map-reduce framework . Proc. of VLDB. Thusoo, A., Sarma, J. S., Jain, N., Shao, Z., Chakka, P., Anthony, S., Liu, H., Wyckoff, P., and Murthy, R. 2009. Hive - a warehousing solution over a map-reduce framework. Proc. of VLDB."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLOUD.2011.14"},{"issue":"8","key":"e_1_2_1_16_1","first-page":"0","article-title":"TPC Benchmark H (Decision Support)","volume":"2","author":"Transaction Processing Performance Council (TPC).","year":"2008","unstructured":"Transaction Processing Performance Council (TPC). 2008 . TPC Benchmark H (Decision Support) , Version 2 . 8 . 0 . http:\/\/www.tpc.org\/tpch\/. Transaction Processing Performance Council (TPC). 2008. TPC Benchmark H (Decision Support), Version 2.8.0. http:\/\/www.tpc.org\/tpch\/.","journal-title":"Version"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1998582.1998637"},{"volume-title":"Proceedings of the 5th Workshop on Large Scale Distributed Systems and Middleware (LADIS).","author":"Verma A.","key":"e_1_2_1_18_1","unstructured":"Verma , A. , Cherkasova , L. , and Campbell , R. H . 2011b. SLO-driven right-sizing and resource provisioning of MapReduce jobs . In Proceedings of the 5th Workshop on Large Scale Distributed Systems and Middleware (LADIS). Verma, A., Cherkasova, L., and Campbell, R. H. 2011b. SLO-driven right-sizing and resource provisioning of MapReduce jobs. In Proceedings of the 5th Workshop on Large Scale Distributed Systems and Middleware (LADIS)."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2038916.2038927"},{"volume-title":"Proceedings of the 11th ACM\/IFIP\/USENIX Middleware Conference.","author":"Wolf J.","key":"e_1_2_1_20_1","unstructured":"Wolf , J. , Rajan , D. , Hildrum , K. , Khandekar , R. , Kumar , V. , Parekh , S. , Wu , K.-L. , and Balmin , A . 2010. FLEX: A slot allocation scheduling optimizer for MapReduce workloads . In Proceedings of the 11th ACM\/IFIP\/USENIX Middleware Conference. Wolf, J., Rajan, D., Hildrum, K., Khandekar, R., Kumar, V., Parekh, S., Wu, K.-L., and Balmin, A. 2010. FLEX: A slot allocation scheduling optimizer for MapReduce workloads. In Proceedings of the 11th ACM\/IFIP\/USENIX Middleware Conference."}],"container-title":["ACM Transactions on Autonomous and Adaptive Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2518017.2518019","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2518017.2518019","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T20:14:12Z","timestamp":1750277652000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2518017.2518019"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,9]]},"references-count":20,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2013,9]]}},"alternative-id":["10.1145\/2518017.2518019"],"URL":"https:\/\/doi.org\/10.1145\/2518017.2518019","relation":{},"ISSN":["1556-4665","1556-4703"],"issn-type":[{"type":"print","value":"1556-4665"},{"type":"electronic","value":"1556-4703"}],"subject":[],"published":{"date-parts":[[2013,9]]},"assertion":[{"value":"2013-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-09-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}