{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:15:46Z","timestamp":1760235346781,"version":"build-2065373602"},"reference-count":24,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2021,8,21]],"date-time":"2021-08-21T00:00:00Z","timestamp":1629504000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["BDCC"],"abstract":"<jats:p>This paper identifies four common misconceptions about the scalability of volunteer computing on big data problems. The misconceptions are then clarified by analyzing the relationship between scalability and the impact factors including the problem size of big data, the heterogeneity and dynamics of volunteers, and the overlay structure. This paper proposes optimization strategies to find the optimal overlay for the given big data problem. This paper forms multiple overlays to optimize the performance of individual steps in terms of MapReduce paradigm. The optimization is to achieve the maximum overall performance by using a minimum number of volunteers, not overusing resources. This paper has demonstrated that the simulations on the concerned factors can fast find the optimization points. This paper concludes that always welcoming more volunteers is an overuse of available resources because they do not always bring benefit to the overall performance. Finding optimal use of volunteers are possible for the given big data problems even on the dynamics and opportunism of volunteers.<\/jats:p>","DOI":"10.3390\/bdcc5030038","type":"journal-article","created":{"date-parts":[[2021,8,22]],"date-time":"2021-08-22T21:42:12Z","timestamp":1629668532000},"page":"38","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["The Optimization Strategies on Clarification of the Misconceptions of Big Data Processing in Dynamic and Opportunistic Environments"],"prefix":"10.3390","volume":"5","author":[{"given":"Wei","family":"Li","sequence":"first","affiliation":[{"name":"School of Engineering & Technology, Central Queensland University, Rockhampton, QLD 4702, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2416-4101","authenticated-orcid":false,"given":"Maolin","family":"Tang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Queensland University of Technology, Brisbane, QLD 4000, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,8,21]]},"reference":[{"key":"ref_1","unstructured":"Oracle (2021, March 12). An Enterprise Architect\u2019s Guide to Big Data\u2014Reference Architecture Overview. Oracle Enterprise Architecture White Paper. Available online: http:\/\/www.oracle.com\/technetwork\/topics\/entarch\/articles\/oea-big-data-guide-1522052.pdf."},{"key":"ref_2","unstructured":"Sarmenta, L. (2001). Volunteer Computing. [Ph.D. Thesis, Massachusetts Institute of Technology]."},{"key":"ref_3","unstructured":"(2021, March 12). ATLAS@Home. Available online: http:\/\/lhcathome.web.cern.ch\/projects\/atlas."},{"key":"ref_4","unstructured":"(2021, March 12). Asteroids@home. Available online: http:\/\/asteroidsathome.net\/."},{"key":"ref_5","unstructured":"(2021, March 12). Einstein@Home. Available online: https:\/\/einsteinathome.org\/."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Li, W., Guo, W., and Li, M. (2020). The Impact Factors on the Competence of Big Data Processing. Int. J. Comput. Appl.","DOI":"10.1080\/1206212X.2020.1719623"},{"key":"ref_7","unstructured":"Casado, R. (2021, March 12). The Three Generations of Big Data Processing. Available online: https:\/\/www.slideshare.net\/Datadopter\/the-three-generations-of-big-data-processing."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1145\/1327452.1327492","article-title":"MapReduce: Simplified Data Processing on Large Clusters","volume":"51","author":"Dean","year":"2008","journal-title":"Commun. ACM"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1109\/TNET.2002.808407","article-title":"Chord: A Scalable Peer-to-Peer Lookup Protocol for Internet Applications","volume":"11","author":"Stoica","year":"2003","journal-title":"IEEE\/ACM Trans. Netw."},{"key":"ref_10","unstructured":"Kaffille, S., and Loesing, K. (2007). Open Chord (1.0.4) User\u2019s Manual 2007, The University of Bamberg. Available online: https:\/\/sourceforge.net\/projects\/open-chord\/."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Fadika, Z., Govindaraju, M., Canon, R., and Ramakrishnan, L. (2012, January 24\u201329). Evaluating Hadoop for Data-Intensive Scientific Operations. Proceedings of the IEEE 5th International Conference on Cloud Computing, Honolulu, HI, USA.","DOI":"10.1109\/CLOUD.2012.118"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"389","DOI":"10.1016\/j.future.2014.01.001","article-title":"Benchmarking MapReduce Implementations under Different Application Scenarios","volume":"36","author":"Dede","year":"2014","journal-title":"Future Gener. Comput. Syst."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"774","DOI":"10.1109\/TPDS.2016.2594765","article-title":"Improving Performance of Heterogeneous MapReduce Clusters with Adaptive Task Tuning","volume":"28","author":"Cheng","year":"2017","journal-title":"IEEE Trans. Parallel Distrib. Syst."},{"key":"ref_14","unstructured":"(2021, March 12). Hadoop. Available online: https:\/\/cwiki.apache.org\/confluence\/display\/HADOOP2\/ProjectDescription."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Jothi, A., and Indumathy, P. (2015, January 19\u201320). Increasing Performance of Parallel and Distributed Systems in High Performance Computing using Weight Based Approach. Proceedings of the International Conference on Circuits, Power and Computing Technologies, Nagercoil, India.","DOI":"10.1109\/ICCPCT.2015.7159347"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"208","DOI":"10.1016\/j.future.2016.02.015","article-title":"Enabling Fast Failure Recovery in Shared Hadoop Clusters: Towards Failure-aware Scheduling","volume":"74","author":"Yildiz","year":"2017","journal-title":"Future Gener. Comput. Syst."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Singh, S., Garg, R., and Mishra, P.K. (2016, January 29\u201330). Observations on Factors Affecting Performance of MapReduce based Apriori on Hadoop Cluster. Proceedings of the International Conference on Computing, Communication and Automation, Greater Noida, India.","DOI":"10.1109\/CCAA.2016.7813695"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Ardagna, D., Bernardi, S., Gianniti, E., Aliabadi, S., Perez-Palacin, D., and Requeno, J. (2016, January 14\u201316). Modeling Performance of Hadoop Applications: A Journey from Queueing Networks to Stochastic Well-Formed Nets. Proceedings of the International Conference on Algorithms and Architectures for Parallel Processing, Granada, Spain.","DOI":"10.1007\/978-3-319-49583-5_47"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Perarnau, S., and Sato, M. (2014, January 19\u201323). Victim Selection and Distributed Work Stealing Performance: A Case Study. Proceedings of the 28th IEEE International Symposium on Parallel and Distributed Processing, Phoenix, AZ, USA.","DOI":"10.1109\/IPDPS.2014.74"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Vu, T.T., and Derbel, B. (2014, January 26\u201329). Link-Heterogeneous Work Stealing. Proceedings of the 4th IEEE\/ACM International Symposium on Cluster, Cloud and Grid Computing, Chicago, IL, USA.","DOI":"10.1109\/CCGrid.2014.85"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1691","DOI":"10.1007\/s10586-016-0625-2","article-title":"MrHeter: Improving MapReduce Performance in Heterogeneous Environments","volume":"19","author":"Zhang","year":"2016","journal-title":"Clust. Comput."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"971","DOI":"10.12720\/jcm.14.10.971-979","article-title":"The Optimization Potential of Volunteer Computing for Compute or Data Intensive Applications","volume":"14","author":"Li","year":"2019","journal-title":"J. Commun."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"6820","DOI":"10.1109\/TII.2020.3046036","article-title":"Time-Series Regeneration with Convolutional Recurrent Generative Adversarial Network for Remaining Useful Life Estimation","volume":"17","author":"Zhang","year":"2021","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Chen, Y., Ganapathi, A., Griffith, R., and Katz, R. (2011, January 25\u201327). The Case for Evaluating MapReduce Performance Using Workload Suites. Proceedings of the 19th Annual International Symposium on Modelling, Analysis, and Simulation of Computer and Telecommunication Systems, Singapore.","DOI":"10.1109\/MASCOTS.2011.12"}],"container-title":["Big Data and Cognitive Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2504-2289\/5\/3\/38\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:48:49Z","timestamp":1760165329000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2504-2289\/5\/3\/38"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,21]]},"references-count":24,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2021,9]]}},"alternative-id":["bdcc5030038"],"URL":"https:\/\/doi.org\/10.3390\/bdcc5030038","relation":{},"ISSN":["2504-2289"],"issn-type":[{"type":"electronic","value":"2504-2289"}],"subject":[],"published":{"date-parts":[[2021,8,21]]}}}