{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T02:10:38Z","timestamp":1784167838806,"version":"3.55.0"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"11","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2024,7]]},"abstract":"<jats:p>Autonomous database management systems (DBMSs) aim to optimize themselves automatically without human guidance. They rely on machine learning (ML) models that predict their run-time behavior to evaluate whether a candidate configuration is beneficial without the expensive execution of queries. However, the high cost of collecting the training data to build these models makes them impractical for real-world deployments. Furthermore, these models are instance-specific and thus require retraining whenever the DBMS's environment changes. State-of-the-art methods spend over 93% of their time running queries for training versus tuning.<\/jats:p>\n          <jats:p>To mitigate this problem, we present the Boot framework for automatically accelerating training data collection in DBMSs. Boot utilizes macro- and micro-acceleration (MMA) techniques that modify query execution semantics with approximate run-time telemetry and skip repetitive parts of the training process. To evaluate Boot, we integrated it into a database gym for PostgreSQL. Our experimental evaluation shows that Boot reduces training collection times by up to 268\u00d7 with modest degradation in model accuracy. These results also indicate that our MMA-based approach scales with dataset size and workload complexity.<\/jats:p>","DOI":"10.14778\/3681954.3682030","type":"journal-article","created":{"date-parts":[[2024,8,30]],"date-time":"2024-08-30T16:23:36Z","timestamp":1725035016000},"page":"3680-3693","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Hit the Gym: Accelerating Query Execution to Efficiently Bootstrap Behavior Models for Self-Driving Database Management Systems"],"prefix":"10.14778","volume":"17","author":[{"given":"Wan Shen","family":"Lim","sequence":"first","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lin","family":"Ma","sequence":"additional","affiliation":[{"name":"University of Michigan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"William","family":"Zhang","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Matthew","family":"Butrovich","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Samuel","family":"Arch","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrew","family":"Pavlo","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,8,30]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2011. TPC-H dbgen. Retrieved 2024-07-04 from https:\/\/github.com\/electrum\/tpch-dbgen"},{"key":"e_1_2_1_2_1","unstructured":"2021. pgmustard: Calculating per-operation times in EXPLAIN ANALYZE. Retrieved 2024-07-04 from https:\/\/www.pgmustard.com\/blog\/calculating-per-operation-times-in-postgres-explain-analyze"},{"key":"e_1_2_1_3_1","unstructured":"2023. Llama 2: Inference code for LLaMA models. Retrieved 2024-07-04 from https:\/\/github.com\/facebookresearch\/llama"},{"key":"e_1_2_1_4_1","unstructured":"2024. pgtune - tuning PostgreSQL config by your hardware. Retrieved 2024-07-04 from https:\/\/github.com\/le0pard\/pgtune"},{"key":"e_1_2_1_5_1","volume-title":"NeurIPS Data-Centric AI Workshop. Retrieved 2024-07-04 from https:\/\/datacentricai.org\/neurips21\/","author":"DCAI","year":"2021","unstructured":"DCAI 2021. 2021. NeurIPS Data-Centric AI Workshop. Retrieved 2024-07-04 from https:\/\/datacentricai.org\/neurips21\/"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1066157.1066171"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3639305"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3517845"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3056097"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1066157.1066223"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1007568.1007659"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the 23rd International Conference on Very Large Data Bases","author":"Chaudhuri Surajit","unstructured":"Surajit Chaudhuri and Vivek R. Narasayya. 1997. An Efficient, Cost-Driven Index Selection Tool for Microsoft SQL Server. In Proceedings of the 23rd International Conference on Very Large Data Bases (Athens, Greece) (VLDB '97). Very Large Data Bases Endowment Inc., 146--155."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3314035"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.14778\/3484224.3484234"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407820"},{"key":"e_1_2_1_17_1","volume-title":"AutoGluon-Tabular: Robust and Accurate AutoML for Structured Data. arXiv preprint arXiv:2003.06505","author":"Erickson Nick","year":"2020","unstructured":"Nick Erickson, Jonas Mueller, Alexander Shirkov, Hang Zhang, Pedro Larroy, Mu Li, and Alexander Smola. 2020. AutoGluon-Tabular: Robust and Accurate AutoML for Structured Data. arXiv preprint arXiv:2003.06505 (2020)."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376732"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/69.273032"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1006\/jcss.1996.0041"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389704"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.14778\/3384345.3384349"},{"key":"e_1_2_1_23_1","volume-title":"Forecasting: Principles and Practice","author":"Hyndman Rob J","year":"2021","unstructured":"Rob J Hyndman and George Athanasopoulos. 2021. Forecasting: Principles and Practice, 3rd edition. Retrieved 2024-07-04 from https:\/\/otexts.com\/fpp3\/","edition":"3"},{"key":"e_1_2_1_24_1","volume-title":"Proceedings of the 18th International Conference on Very Large Data Bases (VLDB '92)","author":"Ioannidis Yannis E.","unstructured":"Yannis E. Ioannidis, Raymond T. Ng, Kyuseok Shim, and Timos K. Sellis. 1992. Parametric Query Optimization. In Proceedings of the 18th International Conference on Very Large Data Bases (VLDB '92). 103--114."},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems","author":"Ke Guolin","year":"2017","unstructured":"Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017. LightGBM: a highly efficient gradient boosting decision tree. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS'17). 3149--3157."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.14778\/2095686.2095696"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2903728"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-017-0480-7"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.14778\/3352063.3352129"},{"key":"e_1_2_1_30_1","volume-title":"Database Gyms. In CIDR 2023, Conference on Innovative Data Systems Research.","author":"Lim Wan Shen","year":"2023","unstructured":"Wan Shen Lim, Matthew Butrovich, William Zhang, Andrew Crotty, Lin Ma, Peijing Xu, Johannes Gehrke, and Andrew Pavlo. 2023. Database Gyms. In CIDR 2023, Conference on Innovative Data Systems Research."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/1007568.1007658"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389768"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196908"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457276"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3452838"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342644"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342646"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3056098"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457568"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.14778\/2556549.2556563"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.14778\/3231751.3231757"},{"key":"e_1_2_1_42_1","volume-title":"Self-Driving Database Management Systems. In CIDR 2017, Conference on Innovative Data Systems Research.","author":"Pavlo Andrew","year":"2017","unstructured":"Andrew Pavlo, Gustavo Angulo, Joy Arulraj, Haibin Lin, Jiexi Lin, Lin Ma, Prashanth Menon, Todd Mowry, Matthew Perron, Ian Quah, Siddharth Santurkar, Anthony Tomasic, Skye Toor, Dana Van Aken, Ziqi Wang, Yingjun Wu, Ran Xian, and Tieying Zhang. 2017. Self-Driving Database Management Systems. In CIDR 2017, Conference on Innovative Data Systems Research."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476311.3476411"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3526152"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.14778\/3547305.3547309"},{"key":"e_1_2_1_46_1","unstructured":"The Transaction Processing Council. 2021. TPC-DS Benchmark (Revision 3.2.0). Retrieved 2024-07-04 from https:\/\/www.tpc.org\/TPC_Documents_Current_Versions\/pdf\/TPC-DS_v3.2.0.pdf"},{"key":"e_1_2_1_47_1","unstructured":"The Transaction Processing Council. 2022. TPC-H Benchmark (Revision 3.0.1). Retrieved 2024-07-04 from https:\/\/www.tpc.org\/TPC_Documents_Current_Versions\/pdf\/TPC-H_v3.0.1.pdf"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3064029"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457286"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.14778\/3484224.3484236"},{"key":"e_1_2_1_51_1","unstructured":"Kaiwen Wang Junxiong Wang Yueying Li Nathan Kallus Immanuel Trummer and Wen Sun. 2023. JoinGym: An Efficient Query Optimization Environment for Reinforcement Learning. arXiv:2307.11704 [cs.LG]"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654985"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1080\/00401706.1962.10490022"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2013.6544899"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3526128"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626246.3653391"},{"key":"e_1_2_1_57_1","volume-title":"CIDR 2022, Conference on Innovative Data Systems Research.","author":"Wu Ziniu","year":"2022","unstructured":"Ziniu Wu, Pei Yu, Peilun Yang, Rong Zhu, Yuxing Han, Yaliang Li, Defu Lian, Kai Zeng, and Jingren Zhou. 2022. A Unified Transferable Model for ML-Enhanced DBMS. In CIDR 2022, Conference on Innovative Data Systems Research."},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3517885"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.14778\/3421424.3421432"},{"key":"e_1_2_1_60_1","volume-title":"Matthew Butrovich, and Andrew Pavlo.","author":"Zhang William","year":"2024","unstructured":"William Zhang, Wan Shen Lim, Matthew Butrovich, and Andrew Pavlo. 2024. The Holon Approach for Simultaneously Tuning Multiple Components in a Self-Driving Database Management System with Machine Learning via Synthesized Proto-Actions. In Under Submission."},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.14778\/3632093.3632114"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.14778\/3636218.3636235"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.14778\/3485450.3485456"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.14778\/3397230.3397238"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3555041.3589674"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3681954.3682030","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,4]],"date-time":"2024-09-04T18:45:12Z","timestamp":1725475512000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3681954.3682030"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7]]},"references-count":65,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2024,7]]}},"alternative-id":["10.14778\/3681954.3682030"],"URL":"https:\/\/doi.org\/10.14778\/3681954.3682030","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2024,7]]},"assertion":[{"value":"2024-08-30","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}