{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,8,6]],"date-time":"2024-08-06T11:31:03Z","timestamp":1722943863565},"reference-count":25,"publisher":"Engineering and Technology Publishing","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["jcm"],"published-print":{"date-parts":[[2020]]},"abstract":"<jats:p>In contrast to HPC clusters, when big data is processing in a distributed, particularly dynamic and opportunistic environment, the overall performance must be impaired and even bottlenecked by the dynamics of overlay and the opportunism of computing nodes. The dynamics and opportunism are caused by churn and unreliability of a generic distributed environment, and they cannot be ignored or avoided. Understanding impact factors, their impact strength and the relevance between these impacts is the foundation of potential optimization. This paper derives the research background, methodology and results by reasoning the necessity of distributed environments for big data processing, scrutinizing the dynamics and opportunism of distributed environments, classifying impact factors, proposing evaluation metrics and carrying out a series of intensive experiments. The result analysis of this paper provides important insights to the impact strength of the factors and the relevance of impact across the factors. The production of the results aims at paving a way to future optimization or avoidance of potential bottlenecks for big data processing in distributed environments.<\/jats:p>","DOI":"10.12720\/jcm.15.11.776-789","type":"journal-article","created":{"date-parts":[[2020,12,29]],"date-time":"2020-12-29T07:30:16Z","timestamp":1609227016000},"page":"776-789","source":"Crossref","is-referenced-by-count":2,"title":["The Experimental Study of Performance Impairment of Big Data Processing in Dynamic and Opportunistic Environments"],"prefix":"10.12720","author":[{"name":"School of Engineering & Technology, Central Queensland University, Australia","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Li","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"William W.","family":"Guo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"4977","published-online":{"date-parts":[[2020]]},"reference":[{"key":"ref0","doi-asserted-by":"publisher","unstructured":"[1] E. J. Korpela, \"SETI@home, BOINC, and volunteer distributed computing,\" Annual Review of Earth and Planetary Sciences, vol. 40, pp. 69-87, 2012.","DOI":"10.1146\/annurev-earth-040809-152348"},{"key":"ref1","unstructured":"[2] Oracle 2016. An Enterprise Architect's Guide to Big Data -Reference Architecture Overview, Oracle Enterprise Architecture White Paper. [Online]. Available: http:\/\/www.oracle.com\/technetwork\/topics\/entarch\/articles\/oea-big-data-guide-1522052.pdf"},{"key":"ref2","unstructured":"[3] ATLAS@Home 2020. [Online]. Available: http:\/\/lhcathome.web.cern.ch\/projects\/atlas"},{"key":"ref3","unstructured":"[4] Asteroids@home 2020. [Online]. Available: http:\/\/asteroidsathome.net\/"},{"key":"ref4","unstructured":"[5] Einstein@Home 2020. [Online]. Available: https:\/\/einsteinathome.org\/"},{"key":"ref5","unstructured":"[6] D. P. Anderson, \"BOINC: A system for public-resource computing and storage,\" in Proc. 5th IEEE\/ACM International Conference on Grid Computing, 2004, pp. 4-10."},{"key":"ref6","unstructured":"[7] L. Sarmenta, \"Volunteer computing,\" PhD thesis, Massachusetts Institute of Technology, 2001."},{"key":"ref7","unstructured":"[8] R. Casado. (2013). The Three Generations of Big Data Processing. [Online]. Available: https:\/\/www.slideshare.net\/Datadopter\/the-threegenerations-of-big-data-processing"},{"key":"ref8","doi-asserted-by":"publisher","unstructured":"[9] J. Dean and S. Ghemawat, \"MapReduce: Simplified data processing on large clusters,\" Communications of the ACM, vol. 51, no. 1, pp. 107-113, 2008.","DOI":"10.1145\/1327452.1327492"},{"key":"ref9","doi-asserted-by":"publisher","unstructured":"[10] I. Stoica, R. Morris, D. Liben-Nowell, D. R. Karger, M. F. Kaashoek, F. Dabek, and H. Balakrishnan, \"Chord: A scalable peer-to-peer lookup protocol for internet applications,\" IEEE\/ACM Transactions on Networking, vol. 11, no. 1, pp. 17-32, 2003.","DOI":"10.1109\/TNET.2002.808407"},{"key":"ref10","unstructured":"[11] S. Kaffille and K. Loesing. (2007). Open Chord (1.0.4) User's Manual, The University of Bamberg, Germany. [Online]. Available: https:\/\/sourceforge.net\/projects\/openchord\/"},{"key":"ref11","doi-asserted-by":"publisher","unstructured":"[12] S. Singh, R. Garg, and P. K. Mishra, \"Observations on factors affecting performance of mapreduce based apriori on hadoop cluster,\" in Proc. International Conference on Computing, Communication and Automation, 2016, pp. 87-94.","DOI":"10.1109\/CCAA.2016.7813695"},{"key":"ref12","unstructured":"[13] Hadoop Project. 2020. [Online]. Available: https:\/\/cwiki.apache.org\/confluence\/display\/HADOOP2\/ProjectDescription"},{"key":"ref13","doi-asserted-by":"publisher","unstructured":"[14] D. Cheng, J. Rao, Y. Guo, C. Jiang, and X. Zhou, \"Improving performance of heterogeneous mapreduce clusters with adaptive task tuning,\" IEEE Transactions on Parallel and Distributed Systems, vol. 28, no. 3, pp. 774-786, 2017.","DOI":"10.1109\/TPDS.2016.2594765"},{"key":"ref14","unstructured":"[15] Apache Software Foundation 2020a, Wordcount Example. [Online]. Available: https:\/\/cwiki.apache.org\/confluence\/display\/HADOOP2\/WordCount"},{"key":"ref15","unstructured":"[16] Apache Software Foundation. (2020). Grep Example, [Online]. Available: https:\/\/cwiki.apache.org\/confluence\/display\/HADOOP2\/Grep"},{"key":"ref16","unstructured":"[17] Apache Software Foundation. (2020). Terasort Example, [Online]. Available: http:\/\/hadoop.apache.org\/docs\/current\/api\/org\/apache\/hadoop\/examples\/terasort\/package-summary.html"},{"key":"ref17","doi-asserted-by":"publisher","unstructured":"[18] A. Spivak, and D. Nasonov, \"Data preloading and data placement for MapReduce performance improving,\" Procedia Computer Science, vol. 101, pp. 379-387, 2016.","DOI":"10.1016\/j.procs.2016.11.044"},{"key":"ref18","doi-asserted-by":"publisher","unstructured":"[19] Y. Chen, A. Ganapathi, R. Griffith, and R. Katz, \"The case for evaluating MapReduce performance using workload suites,\" in Proc. 19th Annual International Symposium on Modelling, Analysis, and Simulation of Computer and Telecommunication Systems, 2011, pp. 390-399.","DOI":"10.1109\/MASCOTS.2011.12"},{"key":"ref19","doi-asserted-by":"publisher","unstructured":"[20] R. Han and X. Lu, \"On big data benchmarking,\" Lecture Notes in Computer Science, vol. 8807, pp. 3-18, 2014.","DOI":"10.1007\/978-3-319-13021-7_1"},{"key":"ref20","doi-asserted-by":"publisher","unstructured":"[21] W. H. Lee, H. G. Jun, and H. J. Kim, \"Hadoop mapreduce performance enhancement using in-node combiners,\" International Journal of Computer Science & Information Technology, vol. 7, no. 5, pp. 1-17, 2015.","DOI":"10.5121\/ijcsit.2015.7501"},{"key":"ref21","doi-asserted-by":"publisher","unstructured":"[22] D. Ardagna, S. Bernardi, E. Gianniti, S. Aliabadi, D. Perez-Palacin, and J. Requeno, \"Modeling performance of hadoop applications: A journey from queueing networks to stochastic well formed nets,\" in Proc. International Conference on Algorithms and Architectures for Parallel Processing, 2016, pp. 599-613.","DOI":"10.1007\/978-3-319-49583-5_47"},{"key":"ref22","doi-asserted-by":"publisher","unstructured":"[23] E. Dede, Z. Fadika, M. Govindaraju, and L. Ramakrishnan, \"Benchmarking mapreduce implementations under different application scenarios,\" Future Generation Computer Systems, vol. 36, pp. 389-399, 2014.","DOI":"10.1016\/j.future.2014.01.001"},{"key":"ref23","doi-asserted-by":"publisher","unstructured":"[24] X. Zhang, Y. Wu, and C. Zhao, \"MrHeter: Improving MapReduce performance in heterogeneous environments,\" Cluster Computing, vol. 19, no. 4, pp. 1691-1701, 2016.","DOI":"10.1007\/s10586-016-0625-2"},{"key":"ref24","doi-asserted-by":"publisher","unstructured":"[25] S. Monsalve, F. Carballeira, and A. Calder\u00f3n, \"A new volunteer computing model for data \u2010 intensive applications,\" Concurrency and Computation Practice and Experience, vol. 29, no. 24, 2017.","DOI":"10.1002\/cpe.4198"}],"container-title":["Journal of Communications"],"original-title":[],"link":[{"URL":"http:\/\/www.jocm.us\/uploadfile\/2020\/1013\/20201013053622846.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,11,24]],"date-time":"2021-11-24T07:42:31Z","timestamp":1637739751000},"score":1,"resource":{"primary":{"URL":"http:\/\/www.jocm.us\/show-246-1605-1.html"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020]]},"references-count":25,"URL":"https:\/\/doi.org\/10.12720\/jcm.15.11.776-789","relation":{},"ISSN":["1796-2021"],"issn-type":[{"type":"print","value":"1796-2021"}],"subject":[],"published":{"date-parts":[[2020]]}}}