{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T22:12:58Z","timestamp":1777500778075,"version":"3.51.4"},"reference-count":19,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2022,9,9]],"date-time":"2022-09-09T00:00:00Z","timestamp":1662681600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,9,9]],"date-time":"2022-09-09T00:00:00Z","timestamp":1662681600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100009473","name":"Universidad de M\u00e1laga","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100009473","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Cluster Comput"],"published-print":{"date-parts":[[2023,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The popularization of Hadoop as the the-facto standard platform for data analytics in the context of Big Data applications has led to the upsurge of SQL-on-Hadoop systems, which provide scalable query execution engines allowing the use of SQL queries on data stored in HDFS. In this context, Kubernetes appears as the leading choice to simplify the deployment and scaling of containerized applications; however, there is a lack of studies about the performance of SQL-on-Hadoop systems deployed on Kubernetes, and this is the gap we intend to fill in this paper. We present an experimental study involving four representative SQL scalable platforms: Apache Drill, Apache Hive, Apache Spark SQL and Trino. Concretely, we analyze the performance of these systems when they are deployed on a Hadoop cluster with Kubernetes by using the TPC-H benchmark. The results of our study can help practitioners and users about what they can expect in terms of performance if they plan to use the advantages of Kubernetes to deploy applications using the analyzed SQL scalable platforms.<\/jats:p>","DOI":"10.1007\/s10586-022-03718-9","type":"journal-article","created":{"date-parts":[[2022,9,9]],"date-time":"2022-09-09T18:28:53Z","timestamp":1662748133000},"page":"1935-1947","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["On the performance of SQL scalable systems on Kubernetes: a comparative study"],"prefix":"10.1007","volume":"26","author":[{"given":"Cristian","family":"Cardas","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jos\u00e9 F.","family":"Aldana-Mart\u00edn","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Antonio M.","family":"Burgue\u00f1o-Romero","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Antonio J.","family":"Nebro","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jose M.","family":"Mateos","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Juan J.","family":"S\u00e1nchez","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,9,9]]},"reference":[{"key":"3718_CR1","volume-title":"Hadoop: The Definitive Guide","author":"T White","year":"2009","unstructured":"White, T.: Hadoop: The Definitive Guide. O\u2019Reilly Media, Sebastopol (2009)"},{"key":"3718_CR2","volume-title":"Programming Hive: Data Warehouse and Query Language for Hadoop","author":"E Capriolo","year":"2012","unstructured":"Capriolo, E., Wampler, D., Rutherglen, J.: Programming Hive: Data Warehouse and Query Language for Hadoop. O\u2019Reilly Media, Sebastopol (2012)"},{"key":"3718_CR3","volume-title":"Getting Started with Impala: Interactive SQL for Apache Hadoop","author":"J Russell","year":"2014","unstructured":"Russell, J.: Getting Started with Impala: Interactive SQL for Apache Hadoop. O\u2019Reilly Media, Sebastopol (2014)"},{"key":"3718_CR4","volume-title":"Trino: The Definitive Guide","author":"M Fuller","year":"2021","unstructured":"Fuller, M., Traverso, M., Moser, M.: Trino: The Definitive Guide. O\u2019Reilly Media, Sebastopol, (2021)"},{"key":"3718_CR5","volume-title":"Learning Apache Drill: Query and Analyze Distributed Data Sources with SQL","author":"C Givre","year":"2018","unstructured":"Givre, C., Rogers, P.: Learning Apache Drill: Query and Analyze Distributed Data Sources with SQL. O\u2019Reilly Media, Sebastopol (2018)"},{"key":"3718_CR6","volume-title":"Spark: The Definitive Guide: Big Data Processing Made Simple","author":"B Chambers","year":"2018","unstructured":"Chambers, B., Zaharia, M.: Spark: The Definitive Guide: Big Data Processing Made Simple. O\u2019Reilly Media, Sebastopol (2018)"},{"key":"3718_CR7","doi-asserted-by":"crossref","unstructured":"Abdollahi\u00a0Vayghan, L., Saied, M.A., Toeroe, M., Khendek, F.: Deploying microservice based applications with kubernetes: Experiments and lessons learned. In: 2018 IEEE 11th International Conference on Cloud Computing (CLOUD), pp. 970\u2013973 (2018)","DOI":"10.1109\/CLOUD.2018.00148"},{"issue":"12","key":"3718_CR8","doi-asserted-by":"publisher","first-page":"1295","DOI":"10.14778\/2732977.2733002","volume":"7","author":"A Floratou","year":"2014","unstructured":"Floratou, A., Minhas, U.F., \u00d6zcan, F.: Sql-on-hadoop: full circle back to shared-nothing database architectures. Proc. VLDB Endow. 7(12), 1295\u20131306 (2014)","journal-title":"Proc. VLDB Endow."},{"issue":"4","key":"3718_CR9","doi-asserted-by":"publisher","first-page":"64","DOI":"10.1145\/369275.369291","volume":"29","author":"M Poess","year":"2000","unstructured":"Poess, M., Floyd, C.: New TPC benchmarks for decision support and web commerce. SIGMOD Rec. 29(4), 64\u201371 (2000)","journal-title":"SIGMOD Rec."},{"key":"3718_CR10","doi-asserted-by":"crossref","unstructured":"Poess, M., Rabl, T., Jacobsen, H.-A.: Analysis of tpc-ds: The first standard benchmark for SQL-based big data systems. In: Proceedings of the 2017 Symposium on Cloud Computing. SoCC \u201917, pp. 573\u2013585. Association for Computing Machinery, New York (2017)","DOI":"10.1145\/3127479.3128603"},{"key":"3718_CR11","doi-asserted-by":"publisher","first-page":"154","DOI":"10.1007\/978-3-319-13021-7_12","volume-title":"Big Data Benchmarks, Performance Optimization, and Emerging Hardware","author":"Y Chen","year":"2014","unstructured":"Chen, Y., Qin, X., Bian, H., Chen, J., Dong, Z., Du, X., Gao, Y., Liu, D., Lu, J., Zhang, H.: A study of SQL-on-Hadoop systems. In: Zhan, J., Han, R., Weng, C. (eds.) Big Data Benchmarks, Performance Optimization, and Emerging Hardware, pp. 154\u2013166. Springer, Cham (2014)"},{"key":"3718_CR12","doi-asserted-by":"crossref","unstructured":"Tapdiya, A., Fabbri, D.: A comparative analysis of state-of-the-art sql-on-hadoop systems for interactive analytics. In: Nie, J., Obradovic, Z., Suzumura, T., Ghosh, R., Nambiar, R., Wang, C., Zang, H., Baeza-Yates, R., Hu, X., Kepner, J., Cuzzocrea, A., Tang, J., Toyoda, M. (eds.) 2017 IEEE International Conference on Big Data, BigData 2017, Boston, MA, USA, December 11\u201314, 2017, pp. 1349\u20131356. IEEE Computer Society, Boston (2017)","DOI":"10.1109\/BigData.2017.8258066"},{"key":"3718_CR13","doi-asserted-by":"crossref","unstructured":"Pavlo, A., Paulson, E., Rasin, A., Abadi, D.J., DeWitt, D.J., Madden, S., Stonebraker, M.: A comparison of approaches to large-scale data analysis. SIGMOD \u201909, pp. 165\u2013178. Association for Computing Machinery, New York (2009)","DOI":"10.1145\/1559845.1559865"},{"issue":"2","key":"3718_CR14","doi-asserted-by":"crossref","first-page":"1297","DOI":"10.1002\/widm.1297","volume":"9","author":"M Rodrigues","year":"2019","unstructured":"Rodrigues, M., Santos, M.Y., Bernardino, J.: Big data processing tools: an experimental performance evaluation. WIREs Data Min. Knowl. Discov. 9(2), 1297 (2019)","journal-title":"WIREs Data Min. Knowl. Discov."},{"key":"3718_CR15","doi-asserted-by":"publisher","first-page":"1347","DOI":"10.1007\/s10586-019-02914-4","volume":"22","author":"V Aluko","year":"2019","unstructured":"Aluko, V., Sakr, S.: Big sql systems: an experimental evaluation. Clust. Comput. 22, 1347\u20131377 (2019)","journal-title":"Clust. Comput."},{"key":"3718_CR16","doi-asserted-by":"crossref","unstructured":"Zhu, C., Han, B., Zhao, Y.: A comparative study of spark on the bare metal and kubernetes. In: 2020 6th International conference on big data and information analytics (BigDIA), pp. 117\u2013124 (2020)","DOI":"10.1109\/BigDIA51454.2020.00027"},{"key":"3718_CR17","doi-asserted-by":"crossref","unstructured":"Raju, A., Ramanathan, R., Hemavathy, R.: A comparative study of spark schedulers\u2019 performance. In: 2019 4th international conference on computational systems and information technology for sustainable solution (CSITSS), vol. 4, pp. 1\u20135 (2019)","DOI":"10.1109\/CSITSS47250.2019.9031028"},{"key":"3718_CR18","unstructured":"Transaction Processing Performance Council TPC: TPC Benchmark H (Decision Support) Standard Specification (2021)"},{"key":"3718_CR19","doi-asserted-by":"publisher","first-page":"61","DOI":"10.1007\/978-3-319-04936-6_5","volume-title":"Performance Characterization and Benchmarking","author":"P Boncz","year":"2014","unstructured":"Boncz, P., Neumann, T., Erling, O.: TPC-H analyzed: hidden messages and lessons learned from an influential benchmark. In: Nambiar, R., Poess, M. (eds.) Performance Characterization and Benchmarking, pp. 61\u201376. Springer, Cham (2014)"}],"container-title":["Cluster Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10586-022-03718-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10586-022-03718-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10586-022-03718-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,26]],"date-time":"2023-11-26T19:58:26Z","timestamp":1701028706000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10586-022-03718-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,9]]},"references-count":19,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,6]]}},"alternative-id":["3718"],"URL":"https:\/\/doi.org\/10.1007\/s10586-022-03718-9","relation":{},"ISSN":["1386-7857","1573-7543"],"issn-type":[{"value":"1386-7857","type":"print"},{"value":"1573-7543","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,9]]},"assertion":[{"value":"21 June 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 July 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 August 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 September 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}