{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,25]],"date-time":"2025-09-25T18:23:20Z","timestamp":1758824600215},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"7","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2024,3]]},"abstract":"<jats:p>The proliferation of big data and analytic workloads has driven the need for cloud compute and cluster-based job processing. With Apache Spark, users can process terabytes of data at ease with hundreds of parallel executors. Providing low latency access to Spark clusters and sessions is a challenging problem due to the large overheads of cluster creation and session startup. In this paper, we introduce Intelligent Pooling, a system for proactively provisioning compute resources to combat the aforementioned overheads. Our system (1) predicts usage patterns using an innovative hybrid Machine Learning (ML) model with low latency and high accuracy; and (2) optimizes the pool size dynamically to meet customer demand while reducing extraneous COGS.<\/jats:p>\n          <jats:p>The proposed system auto-tunes its hyper-parameters to balance between performance and operational cost with minimal to no engineering input. Evaluated using large-scale production data, Intelligent Pooling achieves up to 43% reduction in cluster idle time compared to static pooling when targeting 99% pool hit rate. Currently deployed in production, Intelligent Pooling is on track to save tens of million dollars in COGS per year as compared to traditional pre-provisioned pools.<\/jats:p>","DOI":"10.14778\/3654621.3654629","type":"journal-article","created":{"date-parts":[[2024,5,30]],"date-time":"2024-05-30T22:21:08Z","timestamp":1717107668000},"page":"1618-1627","update-policy":"http:\/\/dx.doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Intelligent Pooling: Proactive Resource Provisioning in Large-scale Cloud Service"],"prefix":"10.14778","volume":"17","author":[{"given":"Deepak","family":"Ravikumar","sequence":"first","affiliation":[{"name":"Purdue University, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alex","family":"Yeo","sequence":"additional","affiliation":[{"name":"Netflix, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yiwen","family":"Zhu","sequence":"additional","affiliation":[{"name":"Microsoft, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aditya","family":"Lakra","sequence":"additional","affiliation":[{"name":"Microsoft, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Harsha","family":"Nagulapalli","sequence":"additional","affiliation":[{"name":"Microsoft, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Santhosh","family":"Ravindran","sequence":"additional","affiliation":[{"name":"Microsoft, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Steve","family":"Suh","sequence":"additional","affiliation":[{"name":"Microsoft, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Niharika","family":"Dutta","sequence":"additional","affiliation":[{"name":"Microsoft, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrew","family":"Fogarty","sequence":"additional","affiliation":[{"name":"Microsoft, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yoonjae","family":"Park","sequence":"additional","affiliation":[{"name":"Microsoft, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sumeet","family":"Khushalani","sequence":"additional","affiliation":[{"name":"Microsoft, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arijit","family":"Tarafdar","sequence":"additional","affiliation":[{"name":"Microsoft, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kunal","family":"Parekh","sequence":"additional","affiliation":[{"name":"Microsoft, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Subru","family":"Krishnan","sequence":"additional","affiliation":[{"name":"Microsoft, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,5,30]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330667"},{"key":"e_1_2_1_2_1","volume-title":"Implementing Apache Spark jobs execution and Apache Spark cluster creation for Openstack Sahara. 27, 5","author":"Aleksiyants A","year":"2015","unstructured":"A Aleksiyants, O Borisenko, D Turdakov, A Sher, and S Kuznetsov. 2015. Implementing Apache Spark jobs execution and Apache Spark cluster creation for Openstack Sahara. 27, 5 (2015), 35--48."},{"key":"e_1_2_1_3_1","unstructured":"Amazon. 2022. Amazon AWS. Retrieved July 2 2022 from https:\/\/aws.amazon.com"},{"key":"e_1_2_1_4_1","unstructured":"Apache. 2022. Apache Flink. Retrieved July 2 2022 from https:\/\/flink.apache.org"},{"key":"e_1_2_1_5_1","unstructured":"Apache. 2022. Apache Storm. Retrieved July 2 2022 from https:\/\/storm.apache.org"},{"key":"e_1_2_1_6_1","unstructured":"AWS. 2022. Azure Synapse. Retrieved July 2 2022 from https:\/\/aws.amazon.com\/emr\/features\/spark\/"},{"key":"e_1_2_1_7_1","volume-title":"GluonTS-Probabilistic Time Series Modeling in Python. Retrieved","author":"Amazon AWS.","year":"2022","unstructured":"Amazon AWS. 2022. GluonTS-Probabilistic Time Series Modeling in Python. Retrieved July 2, 2022 from https:\/\/ts.gluon.ai\/stable\/"},{"key":"e_1_2_1_8_1","doi-asserted-by":"crossref","first-page":"929","DOI":"10.1109\/TCC.2016.2586064","article-title":"A forecasting methodology for workload forecasting in cloud systems","volume":"6","author":"Baldan Francisco J","year":"2016","unstructured":"Francisco J Baldan, Sergio Ramirez-Gallego, Christoph Bergmeir, Francisco Herrera, and Jose M Benitez. 2016. A forecasting methodology for workload forecasting in cloud systems. IEEE Transactions on Cloud Computing 6, 4 (2016), 929--941.","journal-title":"IEEE Transactions on Cloud Computing"},{"key":"e_1_2_1_9_1","volume-title":"European Conference on Parallel Processing. Springer, 106--118","author":"Baresi Luciano","year":"2018","unstructured":"Luciano Baresi and Giovanni Quattrocchi. 2018. Towards vertically scalable spark applications. In European Conference on Parallel Processing. Springer, 106--118."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11036-018-0996-0"},{"key":"e_1_2_1_11_1","volume-title":"2014 IEEE 6th International Conference on Cloud Computing Technology and Science. IEEE, 168--173","author":"Biswas Anshuman","year":"2014","unstructured":"Anshuman Biswas, Shikharesh Majumdar, Biswajit Nandy, and Ali El-Haraki. 2014. Automatic resource provisioning: a machine learning based proactive approach. In 2014 IEEE 6th International Conference on Cloud Computing Technology and Science. IEEE, 168--173."},{"key":"e_1_2_1_12_1","volume-title":"Use of reactive and proactive elasticity to adjust resources provisioning in the cloud provider. In 2016 IEEE 18th International Conference on High Performance Computing and Communications","author":"Bouabdallah Raouia","unstructured":"Raouia Bouabdallah, Soufiene Lajmi, and Khaled Ghedira. 2016. Use of reactive and proactive elasticity to adjust resources provisioning in the cloud provider. In 2016 IEEE 18th International Conference on High Performance Computing and Communications; IEEE 14th International Conference on Smart City; IEEE 2nd International Conference on Data Science and Systems (HPCC\/SmartCity\/DSS). IEEE, 1155--1162."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3358090"},{"key":"e_1_2_1_14_1","volume-title":"Best practices: pools for Databricks. Retrieved","year":"2022","unstructured":"DataBricks. 2022. Best practices: pools for Databricks. Retrieved July 2, 2022 from https:\/\/docs.databricks.com\/clusters\/instance-pools\/pool-best-practices.html"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2408776.2408794"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1287\/msom.1070.0203"},{"key":"e_1_2_1_17_1","volume-title":"Prophet: Forecasting at scale. Retrieved","year":"2022","unstructured":"Facebook. 2022. Prophet: Forecasting at scale. Retrieved July 2, 2022 from https:\/\/facebook.github.io\/prophet\/"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137765.3137786"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/2749482.2749839"},{"key":"e_1_2_1_20_1","volume-title":"Google Cloud Platform. Retrieved","year":"2022","unstructured":"Google. 2022. Google Cloud Platform. Retrieved July 2, 2022 from https:\/\/cloud.google.com"},{"key":"e_1_2_1_21_1","unstructured":"Google. 2022. Serverless Spark. Retrieved July 2 2022 from https:\/\/cloud.google.com\/dataproc-serverless\/docs"},{"key":"e_1_2_1_22_1","volume-title":"Spark through Vertex AI. Retrieved","year":"2022","unstructured":"Google. 2022. Spark through Vertex AI. Retrieved July 2, 2022 from https:\/\/cloud.google.com\/vertex-ai-workbench"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2479871.2479899"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10618-020-00710-y"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1002\/spe.2737"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1137\/S1052623499363220"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.3390\/app7080777"},{"key":"e_1_2_1_28_1","volume-title":"Aggressive resource provisioning for ensuring QoS in virtualized environments","author":"Liu Jinzhao","year":"2014","unstructured":"Jinzhao Liu, Yaoxue Zhang, Yuezhi Zhou, Di Zhang, and Hao Liu. 2014. Aggressive resource provisioning for ensuring QoS in virtualized environments. IEEE transactions on cloud computing 3, 2 (2014), 119--131."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10723-014-9314-7"},{"key":"e_1_2_1_30_1","volume-title":"6th {USENIX} Workshop on Hot Topics in Cloud Computing (HotCloud 14).","author":"Lu Qinghua","unstructured":"Qinghua Lu, Liming Zhu, Xiwei Xu, Len Bass, Shanshan Li, Weishan Zhang, and Ning Wang. 2014. Mechanisms and architectures for tail-tolerant system operations in cloud. In 6th {USENIX} Workshop on Hot Topics in Cloud Computing (HotCloud 14)."},{"key":"e_1_2_1_31_1","volume-title":"Azure Data Explorer - Kusto. Retrieved","year":"2022","unstructured":"Microsoft. 2022. Azure Data Explorer - Kusto. Retrieved July 2, 2022 from https:\/\/docs.microsoft.com\/en-us\/azure\/data-explorer\/kusto\/query\/"},{"key":"e_1_2_1_32_1","volume-title":"Retrieved","year":"2022","unstructured":"Microsoft. 2022. Azure Fabric. Retrieved July 24, 2023 from https:\/\/learn.microsoft.com\/en-us\/fabric\/data-engineering\/spark-compute"},{"key":"e_1_2_1_33_1","unstructured":"Microsoft. 2022. Azure HDInsight. Retrieved July 2 2022 from https:\/\/docs.microsoft.com\/en-us\/azure\/hdinsight\/spark\/apache-spark-overview"},{"key":"e_1_2_1_34_1","unstructured":"Microsoft. 2022. Azure Synapse. Retrieved July 2 2022 from https:\/\/docs.microsoft.com\/en-us\/azure\/synapse-analytics\/spark\/apache-spark-overview"},{"key":"e_1_2_1_35_1","volume-title":"Introduction to Cosmos DB. Retrieved","year":"2022","unstructured":"Microsoft. 2022. Introduction to Cosmos DB. Retrieved July 2, 2022 from https:\/\/docs.microsoft.com\/en-us\/azure\/cosmos-db\/introduction"},{"key":"e_1_2_1_36_1","unstructured":"Microsoft. 2022. Microsoft Azure. Retrieved July 2 2022 from https:\/\/azure.microsoft.com"},{"key":"e_1_2_1_37_1","unstructured":"Microsoft. 2022. NimbusMLbu. Retrieved July 2 2022 from https:\/\/docs.microsoft.com\/en-us\/nimbusml\/overview"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2016.04.001"},{"key":"e_1_2_1_39_1","volume-title":"2016 18th Asia-Pacific Network Operations and Management Symposium (APNOMS). IEEE, 1--4.","author":"Oh Yoori","year":"2016","unstructured":"Yoori Oh, Jieun Choi, Eunjung Song, Moonji Kim, and Yoonhee Kim. 2016. A SLA-based Spark cluster scaling method in cloud environment. In 2016 18th Asia-Pacific Network Operations and Management Symposium (APNOMS). IEEE, 1--4."},{"key":"e_1_2_1_40_1","volume-title":"Qun Guo, Alekh Jindal, Ajay Kalhan, Morgan Oslake, Sonia Parchani, Vijay Ramani, Raj Sellappan, Saikat Sen, Sheetal Shrotri, Soundararajan Srinivasan, Ping Xia, Shize Xu, Alicia Yang, and Yiwen Zhu.","author":"Poppe Olga","year":"2020","unstructured":"Olga Poppe, Tayo Amuneke, Dalitso Banda, Aritra De, Ari Green, Manon Knoertzer, Ehi Nosakhare, Karthik Rajendran, Deepak Shankargouda, Meina Wang, Alan Au, Carlo Curino, Qun Guo, Alekh Jindal, Ajay Kalhan, Morgan Oslake, Sonia Parchani, Vijay Ramani, Raj Sellappan, Saikat Sen, Sheetal Shrotri, Soundararajan Srinivasan, Ping Xia, Shize Xu, Alicia Yang, and Yiwen Zhu. 2020. Seagull: An Infrastructure for Load Prediction and Optimized Resource Allocation. In PVLDB. VLDB Endowment, 154--162."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.14778\/3514061.3514073"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICC.2018.8422788"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2843966.2843972"},{"key":"e_1_2_1_44_1","volume-title":"Heuristics for Optimization and Learning","author":"Thonglek Kundjanasith","unstructured":"Kundjanasith Thonglek, Kohei Ichikawa, Chatchawal Sangkeettrakarn, and Apivadee Piyatumrong. 2021. Auto-scaling system in apache spark cluster using model-based deep reinforcement learning. In Heuristics for Optimization and Learning. Springer, 347--360."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3220060"},{"key":"e_1_2_1_46_1","unstructured":"Wikipedia. 2023. Pareto front. https:\/\/en.wikipedia.org\/wiki\/Pareto_front."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2007477.1952699"},{"key":"e_1_2_1_48_1","volume-title":"2nd USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 10)","author":"Zaharia Matei","year":"2010","unstructured":"Matei Zaharia, Mosharaf Chowdhury, Michael J Franklin, Scott Shenker, and Ion Stoica. 2010. Spark: Cluster computing with working sets. In 2nd USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 10)."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/2934664"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467401"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM.2017.8057118"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457569"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3654621.3654629","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,30]],"date-time":"2024-05-30T22:26:16Z","timestamp":1717107976000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3654621.3654629"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3]]},"references-count":52,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2024,3]]}},"alternative-id":["10.14778\/3654621.3654629"],"URL":"https:\/\/doi.org\/10.14778\/3654621.3654629","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2024,3]]},"assertion":[{"value":"2024-05-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}