{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,10,1]],"date-time":"2026-10-01T05:07:19Z","timestamp":1790831239723,"version":"4.1.0"},"reference-count":40,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2020,8]]},"abstract":"<jats:p>Right-sizing resource allocation for big-data queries, particularly in serverless environments, is critical for improving infrastructure operational efficiency, capacity availability, query performance predictability, and for reducing unnecessary wait times. In this paper, we present AutoToken --- a simple and effective predictor for estimating the peak resource usage of recurring big data queries. It uses multiple query plan identifiers to identify recurring query templates and to learn models with the goal of reducing over-allocation in future instances of those queries. AutoToken is computationally light, for both training and scoring, is easily deployable at scale, and is integrated with the Peregrine workload optimization infrastructure at Microsoft. We extensively evaluate AutoToken on SCOPE jobs from our production clusters and show that it outperforms state-of-the-art solutions for peak resource estimation.<\/jats:p>\n          <jats:p>We also discuss our plans towards supporting repeatable and extensible research on resource prediction for SCOPE jobs, including describing a simulation methodology for generating arbitrary-sized datasets with similar characteristics as the production datasets.<\/jats:p>","DOI":"10.14778\/3415478.3415554","type":"journal-article","created":{"date-parts":[[2020,9,14]],"date-time":"2020-09-14T18:46:46Z","timestamp":1600109206000},"page":"3326-3339","source":"Crossref","is-referenced-by-count":18,"title":["AutoToken"],"prefix":"10.14778","volume":"13","author":[{"given":"Rathijit","family":"Sen","sequence":"first","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alekh","family":"Jindal","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hiren","family":"Patel","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shi","family":"Qiao","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,8]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Amazon Athena. https:\/\/aws.amazon.com\/athena\/."},{"key":"e_1_2_1_2_1","unstructured":"Cholesky decomposition. https:\/\/en.wikipedia.org\/wiki\/Cholesky_decomposition."},{"key":"e_1_2_1_3_1","unstructured":"Kullback-Leibler divergence. https:\/\/en.wikipedia.org\/wiki\/Kullback-Leibler_divergence."},{"key":"e_1_2_1_4_1","unstructured":"oj! Algorithms. https:\/\/www.ojalgo.org\/."},{"key":"e_1_2_1_5_1","first-page":"469","volume-title":"Proceedings of the 14th USENIX Conference on Networked Systems Design and Implementation, NSDI'17","author":"Alipourfard O.","year":"2017","unstructured":"O. Alipourfard, H. H. Liu, J. Chen, S. Venkataraman, M. Yu, and M. Zhang. CherryPick: Adaptively unearthing the best cloud configurations for big data analytics. In Proceedings of the 14th USENIX Conference on Networked Systems Design and Implementation, NSDI'17, pages 469--482, USA, Mar. 2017. USENIX Association."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jnca.2017.01.016"},{"key":"e_1_2_1_7_1","first-page":"192","volume-title":"Proceedings of the 2018 ACM\/SPEC International Conference on Performance Engineering, IKPE '18","author":"Ardagna D.","year":"2018","unstructured":"D. Ardagna, E. Barbierato, A. Evangelinou, E. Gianniti, M. Gribaudo, T. B. M. Pinto, A. Guimar\u00e3es, A. P. Couto da Silva, and J. M. Almeida. Performance prediction of cloud-based big data applications. In Proceedings of the 2018 ACM\/SPEC International Conference on Performance Engineering, IKPE '18, pages 192--199, New York, NY, USA, 2018. Association for Computing Machinery."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CloudCom.2014.147"},{"key":"e_1_2_1_9_1","first-page":"285","volume-title":"Proceedings of the 11th USENIX Conference on Operating Systems Design and Implementation, OSDI'14","author":"Boutin E.","year":"2014","unstructured":"E. Boutin, J. Ekanayake, W. Lin, B. Shi, J. Zhou, Z. Qian, M. Wu, and L. Zhou. Apollo: Scalable and coordinated scheduling for cloud-scale computing. In Proceedings of the 11th USENIX Conference on Operating Systems Design and Implementation, OSDI'14, pages 285--300, USA, 2014. USENIX Association."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3167918.3167948"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.14778\/1454159.1454166"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3358090"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3267809.3267819"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3401071.3401656"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2168836.2168847"},{"key":"e_1_2_1_16_1","unstructured":"Google BigQuery. https:\/\/cloud.google.com\/bigquery."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2038916.2038934"},{"key":"e_1_2_1_18_1","first-page":"261","volume-title":"Fifth Biennial Conference on Innovative Data Systems Research","author":"Herodotou H.","year":"2011","unstructured":"H. Herodotou, H. Lim, G. Luo, N. Borisov, L. Dong, F. B. Cetin, and S. Babu. Starfish: A self-tuning system for big data analytics. In Fifth Biennial Conference on Innovative Data Systems Research, pages 261--272. www.cidrdb.org, 2011."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/SCC.2013.67"},{"key":"e_1_2_1_20_1","unstructured":"V. Jalaparti H. Ballani T. Karagiannis A. Rowstron and P. Costa. Bazaar: Enabling predictable performance in datacenters. Technical Report MSR-TR-2012-38 Feb. 2012."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357223.3362726"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3190656"},{"key":"e_1_2_1_23_1","first-page":"117","volume-title":"Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation, OSDI'16","author":"Jyothi S. A.","year":"2016","unstructured":"S. A. Jyothi, C. Curino, I. Menache, S. M. Narayanamurthy, A. Tumanov, J. Yaniv, R. Mavlyutov, I. n. Goiri, S. Krishnan, J. Kulkarni, and S. Rao. Morpheus: Towards automated SLOs for enterprise clusters. In Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation, OSDI'16, pages 117--134, USA, 2016. USENIX Association."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/NOMS.2012.6212065"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2015.2405552"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLOUD.2016.0011"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3005745.3005750"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1809049.1809052"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2016.04.001"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078195"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2987550.2987566"},{"issue":"12","key":"e_1_2_1_32_1","first-page":"1850","article-title":"SparkCruise: Handsfree computation reuse","volume":"12","author":"Roy A.","year":"2019","unstructured":"A. Roy, A. Jindal, H. Patel, A. Gosalia, S. Krishnan, and C. Curino. SparkCruise: Handsfree computation reuse in Spark. PVLDB, 12(12):1850--1853, 2019.","journal-title":"Spark. PVLDB"},{"issue":"1","key":"e_1_2_1_33_1","first-page":"460","article-title":"Runtime measurements in the cloud: Observing, analyzing, and reducing variance","volume":"3","author":"Schad J.","year":"2010","unstructured":"J. Schad, J. Dittrich, and J.-A. Quian\u00e9-Ruiz. Runtime measurements in the cloud: Observing, analyzing, and reducing variance. PVLDB, 3(1-2):460--471, 2010.","journal-title":"PVLDB"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3380584"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLOUD.2011.14"},{"key":"e_1_2_1_36_1","unstructured":"C. Vazquez R. Krishnan and E. John. Time series forecasting of cloud data center workloads for dynamic resource provisioning. Journal of Wireless Mobile Networks Ubiquitous Computing and Dependable Applications (JoWUA) 6(3):87--110 Sept. 2015."},{"key":"e_1_2_1_37_1","first-page":"363","volume-title":"Proceedings of the 13th Usenix Conference on Networked Systems Design and Implementation, NSDI'16","author":"Venkataraman S.","year":"2016","unstructured":"S. Venkataraman, Z. Yang, M. Franklin, B. Recht, and I. Stoica. Ernest: Efficient performance prediction for large-scale advanced analytics. In Proceedings of the 13th Usenix Conference on Networked Systems Design and Implementation, NSDI'16, pages 363--378, Santa Clara, CA, USA, Mar. 2016. USENIX Association."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-25821-3_9"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/3291264.3291267"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10723-011-9201-4"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3415478.3415554","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,17]],"date-time":"2025-09-17T02:46:52Z","timestamp":1758077212000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3415478.3415554"}},"subtitle":["predicting peak parallelism for big data analytics at Microsoft"],"short-title":[],"issued":{"date-parts":[[2020,8]]},"references-count":40,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2020,8]]}},"alternative-id":["10.14778\/3415478.3415554"],"URL":"https:\/\/doi.org\/10.14778\/3415478.3415554","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2020,8]]}}}