{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,30]],"date-time":"2026-05-30T02:09:48Z","timestamp":1780106988434,"version":"3.54.0"},"reference-count":26,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2020,2,7]],"date-time":"2020-02-07T00:00:00Z","timestamp":1581033600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,2,7]],"date-time":"2020-02-07T00:00:00Z","timestamp":1581033600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Cloud Comp"],"published-print":{"date-parts":[[2020,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Task stragglers in MapReduce jobs dramatically impede job execution of data-intensive computing in cloud data centers. This impedance is due to the uneven distribution of input data, heterogeneous data nodes, resource contention situations, and network configurations. Data skew of intermediate data in MapReduce job causes delay failures due to the violation of job completion time. Data-intensive computing frameworks, such as MapReduce or Hadoop YARN, employ HashPartitioner. This partitioner may cause intermediate data skew, which results in straggler reducers. In this paper, we strive to make Hadoop YARN more efficient in cloud environments. We present, a new partitioning scheme, called balanced data clusters partitioner (BDCP), to handle straggler Reduce tasks based on sampling of input data and feedback information about the current processing task. Our extensive experimental results show that BDCP can outperform the default Hadoop HashPartitioner and Range partitioner. BDCP can assist in straggler mitigation during reduce phase and minimize the job completion time in MapReduce jobs within data-intensive cloud computing.<\/jats:p>","DOI":"10.1186\/s13677-019-0139-6","type":"journal-article","created":{"date-parts":[[2020,2,7]],"date-time":"2020-02-07T00:05:09Z","timestamp":1581033909000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Improvement of job completion time in data-intensive cloud computing applications"],"prefix":"10.1186","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0229-3586","authenticated-orcid":false,"given":"Ibrahim Adel","family":"Ibrahim","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mostafa","family":"Bassiouni","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2020,2,7]]},"reference":[{"key":"139_CR1","unstructured":"MapReduce: Official Apache Hadoop Website. http:\/\/hadoop.apache.org. Accessed 14 Feb 2019."},{"key":"139_CR2","doi-asserted-by":"crossref","unstructured":"Wu H (2016) Big data management the mass weather logs In: International Conference on Smart Computing and Communication, 122\u2013132.. Springer.","DOI":"10.1007\/978-3-319-52015-5_13"},{"key":"139_CR3","unstructured":"White T (2009) Hadoop, \u201cThe Definitive Guide (1\u2019st ed.)\u201d"},{"key":"139_CR4","doi-asserted-by":"publisher","unstructured":"Subramanian V, Wang L, Lee E-J, Chen P (2010) Rapid processing of synthetic seismograms using windows azure cloud In: 2010 IEEE Second International Conference on Cloud Computing Technology and Science.. IEEE. https:\/\/doi.org\/10.1109\/cloudcom.2010.110.","DOI":"10.1109\/cloudcom.2010.110"},{"issue":"9","key":"139_CR5","doi-asserted-by":"publisher","first-page":"2520","DOI":"10.1109\/TPDS.2014.2350972","volume":"26","author":"Q Chen","year":"2015","unstructured":"Chen Q, Yao J, Xiao Z (2015) Libra: Lightweight data skew mitigation in mapreduce. IEEE Trans Parallel Distrib Syst 26(9):2520\u20132533.","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"139_CR6","doi-asserted-by":"publisher","first-page":"149","DOI":"10.1016\/j.future.2014.06.009","volume":"43","author":"F Zhang","year":"2015","unstructured":"Zhang F, Cao J, Khan SU, Li K, Hwang K (2015) A task-level adaptive mapreduce framework for real-time streaming data in healthcare applications. Futur Gener Comput Syst 43:149\u2013160.","journal-title":"Futur Gener Comput Syst"},{"key":"139_CR7","unstructured":"MapReduce Job. Word Count. http:\/\/spark.apache.org\/examples.html. Accessed 27 Apr 2019."},{"key":"139_CR8","doi-asserted-by":"publisher","first-page":"66","DOI":"10.1016\/j.future.2013.09.010","volume":"36","author":"D Lee","year":"2014","unstructured":"Lee D, Kim J-S, Maeng S (2014) Large-scale incremental processing with mapreduce. Futur Gener Comput Syst 36:66\u201379.","journal-title":"Futur Gener Comput Syst"},{"key":"139_CR9","unstructured":"Range Partitioner, [EB\/OL]. http:\/\/spark.apache.org\/docs\/1.3.0\/api\/java\/org\/apache\/spark\/RangePartitioner.html. Accessed 11 Apr 2019."},{"key":"139_CR10","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1145\/2213836.2213840","volume-title":"Proceedings of the 2012 ACM SIGMOD International Conference on Management of Data","author":"Y Kwon","year":"2012","unstructured":"Kwon Y, Balazinska M, Howe B, Rolia J (2012) Skewtune: mitigating skew in mapreduce applications In: Proceedings of the 2012 ACM SIGMOD International Conference on Management of Data, 25\u201336.. ACM, Scottsdale."},{"key":"139_CR11","doi-asserted-by":"publisher","first-page":"145","DOI":"10.1016\/j.procs.2014.05.014","volume":"29","author":"MAH Hassan","year":"2014","unstructured":"Hassan MAH, Bamha M, Loulergue F (2014) Handling data-skew effects in join operations using mapreduce. Procedia Comput Sci 29:145\u2013158.","journal-title":"Procedia Comput Sci"},{"issue":"1","key":"139_CR12","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2830544.2830546","volume":"17","author":"D Karapiperis","year":"2015","unstructured":"Karapiperis D, Verykios VS (2015) Load-balancing the distance computations in record linkage. ACM SIGKDD Explor Newsl 17(1):1\u20137.","journal-title":"ACM SIGKDD Explor Newsl"},{"key":"139_CR13","first-page":"49","volume-title":"Proceedings of the Symposium on High Performance Computing","author":"L Vu","year":"2015","unstructured":"Vu L, Alaghband G (2015) A load balancing parallel method for frequent pattern mining on multi-core cluster In: Proceedings of the Symposium on High Performance Computing, 49\u201358.. Society for Computer Simulation International, Alexandria."},{"key":"139_CR14","doi-asserted-by":"publisher","first-page":"993","DOI":"10.1016\/j.future.2017.03.013","volume":"105","author":"Jianjiang Li","year":"2020","unstructured":"Li J, Liu Y, Pan J, Zhang P, Chen W, Wang L (2017) Map-balance-reduce: an improved parallel programming model for load balancing of mapreduce. Futur Gener Comput Syst. https:\/\/doi.org\/10.1016\/j.future.2017.03.013.","journal-title":"Future Generation Computer Systems"},{"key":"139_CR15","doi-asserted-by":"publisher","unstructured":"Xu Y, Zou P, Qu W, Li Z, Li K, Cui X (2012) Sampling-based partitioning in mapreduce for skewed data In: 2012 Seventh ChinaGrid Annual Conference.. IEEE. https:\/\/doi.org\/10.1109\/chinagrid.2012.18.","DOI":"10.1109\/chinagrid.2012.18"},{"key":"139_CR16","doi-asserted-by":"publisher","first-page":"287","DOI":"10.1016\/j.future.2016.06.027","volume":"78","author":"Z Tang","year":"2018","unstructured":"Tang Z, Zhang X, Li K, Li K (2018) An intermediate data placement algorithm for load balancing in spark computing environment. Futur Gener Comput Syst 78:287\u2013301.","journal-title":"Futur Gener Comput Syst"},{"key":"139_CR17","doi-asserted-by":"publisher","unstructured":"Ibrahim IA, Bassiouni M (2017) Improving mapreduce performance with progress and feedback based speculative execution In: 2017 IEEE International Conference on Smart Cloud (SmartCloud).. IEEE. https:\/\/doi.org\/10.1109\/smartcloud.2017.25.","DOI":"10.1109\/smartcloud.2017.25"},{"key":"139_CR18","first-page":"185","volume-title":"Presented as Part of the 10th {USENIX} Symposium on Networked Systems Design and Implementation ({NSDI} 13)","author":"G Ananthanarayanan","year":"2013","unstructured":"Ananthanarayanan G, Ghodsi A, Shenker S, Stoica I (2013) Effective straggler mitigation: Attack of the clones In: Presented as Part of the 10th {USENIX} Symposium on Networked Systems Design and Implementation ({NSDI} 13), 185\u2013198.. USENIX, Lombard."},{"key":"139_CR19","first-page":"7","volume":"8","author":"M Zaharia","year":"2008","unstructured":"Zaharia M, Konwinski A, Joseph AD, Katz RH, Stoica I (2008) Improving mapreduce performance in heterogeneous environments. Osdi 8:7.","journal-title":"Osdi"},{"key":"139_CR20","first-page":"1","volume-title":"2010 IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum (IPDPSW)","author":"J Xie","year":"2010","unstructured":"Xie J, Yin S, Ruan X, Ding Z, Tian Y, Majors J, Manzanares A, Qin X (2010) Improving mapreduce performance through data placement in heterogeneous hadoop clusters In: 2010 IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum (IPDPSW), 1\u20139.. IEEE, Atlanta."},{"key":"139_CR21","doi-asserted-by":"publisher","unstructured":"Lin C, Guo W, Lin C (2013) Self-learning mapreduce scheduler in multi-job environment In: 2013 International Conference on Cloud Computing and Big Data, 610\u2013612.. IEEE. https:\/\/doi.org\/10.1109\/cloudcom-asia.2013.95.","DOI":"10.1109\/cloudcom-asia.2013.95"},{"key":"139_CR22","doi-asserted-by":"publisher","unstructured":"Ibrahim IA, Dai W, Bassiouni M (2016) Intelligent data placement mechanism for replicas distribution in cloud storage systems In: 2016 IEEE International Conference on Smart Cloud (SmartCloud).. IEEE. https:\/\/doi.org\/10.1109\/smartcloud.2016.23.","DOI":"10.1109\/smartcloud.2016.23"},{"issue":"1","key":"139_CR23","doi-asserted-by":"publisher","first-page":"23","DOI":"10.1186\/2192-113X-2-23","volume":"2","author":"W Dai","year":"2013","unstructured":"Dai W, Bassiouni M (2013) An improved task assignment scheme for hadoop running in the clouds. J Cloud Comput Adv Syst Appl 2(1):23.","journal-title":"J Cloud Comput Adv Syst Appl"},{"key":"139_CR24","doi-asserted-by":"publisher","unstructured":"Dai W, Ibrahim I, Bassiouni M (2016) A new replica placement policy for hadoop distributed file system In: 2016 IEEE 2nd International Conference on Big Data Security on Cloud (BigDataSecurity), IEEE International Conference on High Performance and Smart Computing (HPSC), and IEEE International Conference on Intelligent Data and Security (IDS), 262\u2013267.. IEEE. https:\/\/doi.org\/10.1109\/bigdatasecurity-hpsc-ids.2016.30.","DOI":"10.1109\/bigdatasecurity-hpsc-ids.2016.30"},{"key":"139_CR25","doi-asserted-by":"publisher","unstructured":"Dai W, Ibrahim I, Bassiouni M (2016) Improving load balance for data-intensive computing on cloud platforms In: 2016 IEEE International Conference on Smart Cloud (SmartCloud).. IEEE. https:\/\/doi.org\/10.1109\/smartcloud.2016.44.","DOI":"10.1109\/smartcloud.2016.44"},{"key":"139_CR26","doi-asserted-by":"publisher","unstructured":"Khatami Z, Hong S, Lee J, Depner S, Chafi H, Ramanujam J, Kaiser H (2017) A load-balanced parallel and distributed sorting algorithm implemented with PGX.D In: 2017 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW).. IEEE. https:\/\/doi.org\/10.1109\/ipdpsw.2017.30.","DOI":"10.1109\/ipdpsw.2017.30"}],"container-title":["Journal of Cloud Computing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13677-019-0139-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/s13677-019-0139-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13677-019-0139-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,2,8]],"date-time":"2021-02-08T05:19:49Z","timestamp":1612761589000},"score":1,"resource":{"primary":{"URL":"https:\/\/journalofcloudcomputing.springeropen.com\/articles\/10.1186\/s13677-019-0139-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,2,7]]},"references-count":26,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,12]]}},"alternative-id":["139"],"URL":"https:\/\/doi.org\/10.1186\/s13677-019-0139-6","relation":{},"ISSN":["2192-113X"],"issn-type":[{"value":"2192-113X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,2,7]]},"assertion":[{"value":"21 May 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 September 2019","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 February 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare that they have no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"8"}}