{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,10]],"date-time":"2026-01-10T23:18:50Z","timestamp":1768087130313,"version":"3.49.0"},"reference-count":34,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2021,1,26]],"date-time":"2021-01-26T00:00:00Z","timestamp":1611619200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>One of the most important tasks of any platform for big data processing is storing the data received. Different systems have different requirements for the storage formats of big data, which raises the problem of choosing the optimal data storage format to solve the current problem. This paper describes the five most popular formats for storing big data, presents an experimental evaluation of these formats and a methodology for choosing the format. The following data storage formats will be considered: avro, CSV, JSON, ORC, parquet. At the first stage, a comparative analysis of the main characteristics of the studied formats was carried out; at the second stage, an experimental evaluation of these formats was prepared and carried out. For the experiment, an experimental stand was deployed with tools for processing big data installed on it. The aim of the experiment was to find out characteristics of data storage formats, such as the volume and processing speed for different operations using the Apache Spark framework. In addition, within the study, an algorithm for choosing the optimal format from the presented alternatives was developed using tropical optimization methods. The result of the study is presented in the form of a technique for obtaining a vector of ratings of data storage formats for the Apache Hadoop system, based on an experimental assessment using Apache Spark.<\/jats:p>","DOI":"10.3390\/sym13020195","type":"journal-article","created":{"date-parts":[[2021,1,26]],"date-time":"2021-01-26T12:03:57Z","timestamp":1611662637000},"page":"195","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Choosing a Data Storage Format in the Apache Hadoop System Based on Experimental Evaluation Using Apache Spark"],"prefix":"10.3390","volume":"13","author":[{"given":"Vladimir","family":"Belov","sequence":"first","affiliation":[{"name":"Department of Intelligent Information Security Systems, MIREA\u2014Russian Technological University, 119454 Moscow, Russia"},{"name":"Data Center, Russian Academy of Education, 119121 Moscow, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrey","family":"Tatarintsev","sequence":"additional","affiliation":[{"name":"Departments of Higher Mathematics 2, MIREA\u2014Russian Technological University, 119454 Moscow, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1254-9132","authenticated-orcid":false,"given":"Evgeny","family":"Nikulchev","sequence":"additional","affiliation":[{"name":"Department of Intelligent Information Security Systems, MIREA\u2014Russian Technological University, 119454 Moscow, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,1,26]]},"reference":[{"key":"ref_1","first-page":"175","article-title":"Big data analytics: A literature review","volume":"2","author":"Chong","year":"2015","journal-title":"J. Manag. Anal."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Moro Visconti, R., and Morea, D. (2019). Big Data for the Sustainability of Healthcare Project Financing. Sustainability, 11.","DOI":"10.3390\/su11133748"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1993","DOI":"10.1108\/MD-07-2018-0754","article-title":"A bibliometric analysis of research on Big Data analytics for business and management","volume":"57","author":"Ardito","year":"2018","journal-title":"Manag. Decis."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Cappa, F., Oriani, R., Peruffo, E., and McCarthy, I.P. (2020). Big Data for Creating and Capturing Value in the Digitalized Environment: Unpacking the Effects of Volume, Variety and Veracity on Firm Performance. J. Prod. Innov. Manag.","DOI":"10.1111\/jpim.12545"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1080\/17538947.2016.1239771","article-title":"Big Data and cloud computing: Innovation opportunities and challenges","volume":"10","author":"Yang","year":"2017","journal-title":"Int. J. Digit. Earth"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"133","DOI":"10.1016\/j.jss.2016.11.037","article-title":"Performance evaluation of cloud-based log file analysis with Apache Hadoop and Apache Spark","volume":"125","author":"Mavridis","year":"2017","journal-title":"J. Syst. Softw."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Lee, S., Jo, J.Y., and Kim, Y. (2019, January 29\u201331). Survey of Data Locality in Apache Hadoop. Proceedings of the 2019 IEEE International Conference on Big Data, Cloud Computing, Data Science & Engineering (BCD), Honolulu, HI, USA.","DOI":"10.1109\/BCD.2019.8885148"},{"key":"ref_8","unstructured":"Garg, K., and Kaur, D. (August, January 29). Sentiment Analysis on Twitter Data using Apache Hadoop and Performance Evaluation on Hadoop MapReduce and Apache Spark. Proceedings of the International Conference on Artificial Intelligence (ICAI), Las Vegas, NV, USA."},{"key":"ref_9","unstructured":"Hive (2021, January 11). 2020 Apache Hive Specification. Available online: https:\/\/cwiki.apache.org\/confluence\/display\/HIVE."},{"key":"ref_10","unstructured":"Impala (2021, January 11). 2020 Apache Impala Specification. Available online: https:\/\/impala.apache.org\/impala-docs.html."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"14","DOI":"10.30699\/fhi.v8i1.180","article-title":"BigData Analysis in Healthcare: Apache Hadoop, Apache spark and Apache Flink","volume":"8","author":"Nazari","year":"2019","journal-title":"Front. Health Inform."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1007\/s41060-016-0027-9","article-title":"Big data analytics on Apache Spark","volume":"1","author":"Salloum","year":"2016","journal-title":"Int. J. Data Sci. Anal."},{"key":"ref_13","first-page":"605","article-title":"A new algebraic solution to multidimensional minimax location problems with Chebyshev distance","volume":"11","author":"Krivulin","year":"2012","journal-title":"WSEAS Trans. Math."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Gusev, A., Ilin, D., and Nikulchev, E. (2020). The Dataset of the Experimental Evaluation of Software Components for Application Design Selection Directed by the Artificial Bee Colony Algorithm. Data, 5.","DOI":"10.3390\/data5030059"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"357","DOI":"10.1016\/j.eswa.2016.10.047","article-title":"Evolutionary composition of QoS-aware web services: A many-objective perspective","volume":"72","author":"Parejo","year":"2017","journal-title":"Expert Syst. Appl."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1002\/spe.2656","article-title":"Software component identification and selection: A research review","volume":"49","author":"Gholamshahi","year":"2019","journal-title":"Softw. Pract. Exp."},{"key":"ref_17","first-page":"420","article-title":"Effective Selection of Software Components Based on Experimental Evaluations of Quality of Operation","volume":"28","author":"Gusev","year":"2020","journal-title":"Eng. Lett."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"19","DOI":"10.32362\/2500-316X-2020-8-5-19-33","article-title":"Life cycle support software components","volume":"8","author":"Kudzh","year":"2020","journal-title":"Russ. Technol. J."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1007\/s10619-019-07271-0","article-title":"A cost-based storage format selector for materialized results in big data frameworks","volume":"38","author":"Munir","year":"2020","journal-title":"Distrib. Parallel Databases"},{"key":"ref_20","unstructured":"Nicholls, B., Adangwa, M., Estes, R., Iradukunda, H.N., Zhang, Q., and Zhu, T. (2020). Benchmarking Resource Usage of Underlying Datatypes of Apache Spark. arXiv, Available online: https:\/\/arxiv.org\/abs\/2012.04192."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Wang, X., and Xie, Z. (2020). The Case for Alternative Web Archival Formats to Expedite The Data-To-Insight Cycle. arXiv.","DOI":"10.1145\/3383583.3398542"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3427478.3427479","article-title":"Proceedings of the ACM\/IEEE Joint Conference on Digital Libraries 2020 in Wuhan virtually","volume":"1","author":"He","year":"2020","journal-title":"ACM Sigweb Newsl."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Ahmed, S., Ali, M.U., Ferzund, J., Sarwar, M.A., Rehman, A., and Mehmood, A. (2017). Modern Data Formats for Big Bioinformatics Data Analytics. Int. J. Adv. Comput. Sci. Appl., 8.","DOI":"10.14569\/IJACSA.2017.080450"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"267","DOI":"10.3846\/mla.2017.1033","article-title":"A Comparison of HDFS Compact Data Formats: Avro Versus Parquet","volume":"9","author":"Plase","year":"2017","journal-title":"Moksl. Liet. Ateitis"},{"key":"ref_25","unstructured":"Khan, S., Liu, X., Ali, S.A., and Alam, M. (2019). Storage Solutions for Big Data Systems: A Qualitative Study and Comparison. arXiv, Available online: https:\/\/arxiv.org\/abs\/1904.11498."},{"key":"ref_26","first-page":"1","article-title":"NoSQL Database: New Era of Databases for Big data Analytics-Classification, Characteristics and Comparison","volume":"6","author":"Moniruzzaman","year":"2013","journal-title":"Int. J. Database Theory Appl."},{"key":"ref_27","unstructured":"Apache (2021, January 11). Avro specification 2012. Available online: http:\/\/avro.apache.org\/docs\/current\/spec.html."},{"key":"ref_28","unstructured":"ORC (2021, January 11). ORC Specification 2020. Available online: https:\/\/orc.apache.org\/specification\/ORCv1\/."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2522968.2522979","article-title":"The family of mapreduce and large-scale data processing systems","volume":"46","author":"Sakr","year":"2013","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"ref_30","unstructured":"Apache (2021, January 11). Parquet Official Documentation 2018. Available online: https:\/\/parquet.apache.org\/documen-tation\/latest\/."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Chellappan, S., and Ganesan, D. (2018). Introduction to Apache Spark and Spark Core. Practical Apache Spark, Apress.","DOI":"10.1007\/978-1-4842-3652-9"},{"key":"ref_32","first-page":"95","article-title":"Spark: Cluster computing with working sets","volume":"10","author":"Zaharia","year":"2010","journal-title":"HotCloud"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Krivulin, N., and Sergeev, S. (2017, January 20\u201321). Tropical optimization techniques in multi-criteria decision making with Analytical Hierarchy Process. Proceedings of the 2017 European Modelling Symposium (EMS), Manchester, UK.","DOI":"10.1109\/EMS.2017.18"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Krivulin, N. (2018). Methods of tropical optimization in rating alternatives based on pairwise comparisons. Operations Research Proceedings 2016, Springer.","DOI":"10.1007\/978-3-319-55702-1_13"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/13\/2\/195\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:15:41Z","timestamp":1760159741000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/13\/2\/195"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,1,26]]},"references-count":34,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2021,2]]}},"alternative-id":["sym13020195"],"URL":"https:\/\/doi.org\/10.3390\/sym13020195","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,1,26]]}}}