{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,10,1]],"date-time":"2026-10-01T05:07:22Z","timestamp":1790831242309,"version":"4.1.0"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,3,12]],"date-time":"2024-03-12T00:00:00Z","timestamp":1710201600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2024,3,12]]},"abstract":"<jats:p>Database systems often rely on historical query traces to perform workload-based performance tuning. However, real production workloads are time-evolving, making historical queries ineffective for optimizing future workloads. To address this challenge, we propose SIBYL, an end-to-end machine learning-based framework that accurately forecasts a sequence of future queries, with the entire query statements, in various prediction windows. Drawing insights from real-workloads, we propose template-based featurization techniques and develop a stacked-LSTM with an encoder-decoder architecture for accurate forecasting of query workloads. We also develop techniques to improve forecasting accuracy over large prediction windows and achieve high scalability over large workloads with high variability in arrival rates of queries. Finally, we propose techniques to handle workload drifts. Our evaluation on four real workloads demonstrates that SIBYL can forecast workloads with an 87.3% median F1 score, and can result in 1.7\u00d7 and 1.3\u00d7 performance improvement when applied to materialized view selection and index selection applications, respectively.<\/jats:p>","DOI":"10.1145\/3639308","type":"journal-article","created":{"date-parts":[[2024,3,26]],"date-time":"2024-03-26T18:51:32Z","timestamp":1711479092000},"page":"1-27","source":"Crossref","is-referenced-by-count":14,"title":["Sibyl: Forecasting Time-Evolving Query Workloads"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6338-3289","authenticated-orcid":false,"given":"Hanxian","family":"Huang","sequence":"first","affiliation":[{"name":"University of California San Diego, La Jolla, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-0866-7275","authenticated-orcid":false,"given":"Tarique","family":"Siddiqui","sequence":"additional","affiliation":[{"name":"Microsoft Research, Redmond, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0457-8429","authenticated-orcid":false,"given":"Rana","family":"Alotaibi","sequence":"additional","affiliation":[{"name":"Microsoft Gray Systems Lab, Redmond, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3712-7358","authenticated-orcid":false,"given":"Carlo","family":"Curino","sequence":"additional","affiliation":[{"name":"Microsoft Gray Systems Lab, Redmond, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2920-1431","authenticated-orcid":false,"given":"Jyoti","family":"Leeka","sequence":"additional","affiliation":[{"name":"Microsoft, Mountain View, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8844-8165","authenticated-orcid":false,"given":"Alekh","family":"Jindal","sequence":"additional","affiliation":[{"name":"SmartApps, Bellevue, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8766-0946","authenticated-orcid":false,"given":"Jishen","family":"Zhao","sequence":"additional","affiliation":[{"name":"University of California San Diego, La Jolla, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-9151-6024","authenticated-orcid":false,"given":"Jes\u00fas","family":"Camacho-Rodr\u00edguez","sequence":"additional","affiliation":[{"name":"Microsoft Gray Systems Lab, Mountain View, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6835-8434","authenticated-orcid":false,"given":"Yuanyuan","family":"Tian","sequence":"additional","affiliation":[{"name":"Microsoft Gray Systems Lab, Mountain View, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,3,26]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"[n. d.]. Dexter. https:\/\/github.com\/ankane\/dexter."},{"key":"e_1_2_2_2_1","unstructured":"[n. d.]. IBM Db2. https:\/\/www.ibm.com\/analytics\/us\/en\/db2."},{"key":"e_1_2_2_3_1","unstructured":"[n. d.]. IBM Informix. https:\/\/www.ibm.com\/products\/informix."},{"key":"e_1_2_2_4_1","unstructured":"[n. d.]. Microsoft SQL Server. https:\/\/www.microsoft.com\/en-us\/sql-server\/sql-server-2022."},{"key":"e_1_2_2_5_1","unstructured":"[n. d.]. Oracle. https:\/\/www.oracle.com\/database."},{"key":"e_1_2_2_6_1","unstructured":"[n. d.]. SQL Server - Parameter Markers. https:\/\/learn.microsoft.com\/sql\/odbc\/reference\/appendixes\/parameter-markers."},{"key":"e_1_2_2_7_1","volume-title":"Osdi","volume":"16","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al . 2016. Tensorflow: a system for large-scale machine learning.. In Osdi, Vol. 16. Savannah, GA, USA, 265--283."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.14778\/3551793.3551857"},{"key":"e_1_2_2_9_1","volume-title":"Automated Selection of Materialized Views and Indexes in SQL Databases. In VLDB 2000, Proceedings of 26th International Conference on Very Large Data Bases, September 10--14","author":"Agrawal Sanjay","year":"2000","unstructured":"Sanjay Agrawal, Surajit Chaudhuri, and Vivek R. Narasayya. 2000. Automated Selection of Materialized Views and Indexes in SQL Databases. In VLDB 2000, Proceedings of 26th International Conference on Very Large Data Bases, September 10--14, 2000, Cairo, Egypt. Morgan Kaufmann, 496--505."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3190662"},{"key":"e_1_2_2_11_1","first-page":"12","article-title":"AutoAdmin Project at Microsoft Research","volume":"34","author":"Bruno Nicolas","year":"2011","unstructured":"Nicolas Bruno, Surajit Chaudhuri, Arnd Christian K\u00f6nig, Vivek R. Narasayya, Ravishankar Ramamurthy, and Manoj Syamala. 2011. AutoAdmin Project at Microsoft Research: Lessons Learned. IEEE Data Eng. Bull. 34, 4 (2011), 12--19.","journal-title":"Lessons Learned. IEEE Data Eng. Bull."},{"key":"e_1_2_2_12_1","volume-title":"Self-Tuning Database Systems: A Decade of Progress. In VLDB '07","author":"Chaudhuri Surajit","year":"2007","unstructured":"Surajit Chaudhuri and Vivek Narasayya. 2007. Self-Tuning Database Systems: A Decade of Progress. In VLDB '07. 3--14."},{"key":"e_1_2_2_13_1","volume-title":"Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio.","author":"Cho Kyunghyun","year":"2014","unstructured":"Kyunghyun Cho, Bart Van Merri\u00ebnboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078 (2014)."},{"key":"e_1_2_2_14_1","unstructured":"Dineshen Chuckravanen. [n. d.]. Approximate entropy as a measure of cognitive fatigue: an eeg pilot study. ([n. d.])."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19--1423"},{"key":"e_1_2_2_16_1","unstructured":"Edgar Haren. 2017. Oracle Revolutionizes Cloud with the World's First Self-Driving Database. https:\/\/blogs.oracle.com\/database\/post\/oracle-revolutionizes-cloud-with-the-worlds-first-self-driving-database."},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2013.79"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/BigData.2015.7363913"},{"key":"e_1_2_2_19_1","volume-title":"Event Based Forecasting of Database Workloads. In 2018 IEEE 4th International Conference on Computer and Communications (ICCC). 1767--1773","author":"Getta Janusz R.","year":"2018","unstructured":"Janusz R. Getta. 2018. Event Based Forecasting of Database Workloads. In 2018 IEEE 4th International Conference on Computer and Communications (ICCC). 1767--1773."},{"key":"e_1_2_2_20_1","volume-title":"Long short-term memory. Neural computation 9, 8","author":"Hochreiter Sepp","year":"1997","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long short-term memory. Neural computation 9, 8 (1997), 1735--1780."},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDEW.2010.5452738"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-85713-6_10"},{"key":"e_1_2_2_23_1","volume-title":"TEALED: A Multi-Step Workload Forecasting Approach Using Time-Sensitive EMD and Auto LSTM Encoder-Decoder. In Database Systems for Advanced Applications. 706--713.","author":"Huang Xiuqi","year":"2022","unstructured":"Xiuqi Huang, Yunlong Cheng, Xiaofeng Gao, and Guihai Chen. 2022. TEALED: A Multi-Step Workload Forecasting Approach Using Time-Sensitive EMD and Auto LSTM Encoder-Decoder. In Database Systems for Advanced Applications. 706--713."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4612-4380-9_35"},{"key":"e_1_2_2_25_1","volume-title":"Query2vec: An evaluation of NLP techniques for generalized workload analytics. arXiv preprint arXiv:1801.05613","author":"Jain Shrainik","year":"2018","unstructured":"Shrainik Jain, Bill Howe, Jiaqi Yan, and Thierry Cruanes. 2018. Query2vec: An evaluation of NLP techniques for generalized workload analytics. arXiv preprint arXiv:1801.05613 (2018)."},{"key":"e_1_2_2_26_1","first-page":"800","article-title":"Selecting subexpressions to materialize at datacenter scale","volume":"11","author":"Jindal Alekh","year":"2018","unstructured":"Alekh Jindal, Konstantinos Karanasos, Sriram Rao, and Hiren Patel. 2018. Selecting subexpressions to materialize at datacenter scale. VLDB 11, 7 (2018), 800--812.","journal-title":"VLDB"},{"key":"e_1_2_2_27_1","unstructured":"Alekh Jindal Shi Qiao Hiren Patel Abhishek Roy Jyoti Leeka and Brandon Haynes. 2021. Production Experiences from Computation Reuse at Microsoft.. In EDBT. 623--634."},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3190656"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.14778\/1880172.1880175"},{"key":"e_1_2_2_30_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)."},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1142\/S0219519416500779"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2016.7472641"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196908"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3542700.3542703"},{"key":"e_1_2_2_35_1","volume-title":"Neo: A learned query optimizer. arXiv preprint arXiv:1904.03711","author":"Marcus Ryan","year":"2019","unstructured":"Ryan Marcus, Parimarjan Negi, Hongzi Mao, Chi Zhang, Mohammad Alizadeh, Tim Kraska, Olga Papaemmanouil, and Nesime Tatbul. 2019. Neo: A learned query optimizer. arXiv preprint arXiv:1904.03711 (2019)."},{"key":"e_1_2_2_36_1","volume-title":"Knapsack problems: algorithms and computer implementations","author":"Martello Silvano","unstructured":"Silvano Martello and Paolo Toth. 1990. Knapsack problems: algorithms and computer implementations. John Wiley & Sons, Inc."},{"key":"e_1_2_2_37_1","first-page":"64","article-title":"Recurrent neural networks","volume":"5","author":"Medsker Larry R","year":"2001","unstructured":"Larry R Medsker and LC Jain. 2001. Recurrent neural networks. Design and Applications 5 (2001), 64--67.","journal-title":"Design and Applications"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3442338"},{"key":"e_1_2_2_39_1","unstructured":"A.V. Oppenheim. 1999. Discrete-Time Signal Processing. Pearson Education."},{"key":"e_1_2_2_40_1","unstructured":"Oracle. 2006. Oracle Database 10g Release 2: The Self-Managing Database. Technical Report. Oracle."},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/872757.872835"},{"key":"e_1_2_2_42_1","volume-title":"Ziqi Wang, Yingjun Wu, Ran Xian, and Tieying Zhang.","author":"Pavlo Andrew","year":"2017","unstructured":"Andrew Pavlo, Gustavo Angulo, Joy Arulraj, Haibin Lin, Jiexi Lin, Lin Ma, Prashanth Menon, Todd C. Mowry, Matthew Perron, Ian Quah, Siddharth Santurkar, Anthony Tomasic, Skye Toor, Dana Van Aken, Ziqi Wang, Yingjun Wu, Ran Xian, and Tieying Zhang. 2017. Self-Driving Database Management Systems. In CIDR."},{"key":"e_1_2_2_43_1","volume-title":"Scikit-Learn: Machine Learning in Python. J. Mach. Learn. Res. 12, null (nov","author":"Pedregosa Fabian","year":"2011","unstructured":"Fabian Pedregosa, Ga\u00ebl Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and \u00c9douard Duchesnay. 2011. Scikit-Learn: Machine Learning in Python. J. Mach. Learn. Res. 12, null (nov 2011), 2825--2830."},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.88.6.2297"},{"key":"e_1_2_2_45_1","volume-title":"et al","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al . 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9."},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1002\/widm.1249"},{"key":"e_1_2_2_47_1","doi-asserted-by":"crossref","unstructured":"Tarique Siddiqui Alekh Jindal Shi Qiao Hiren Patel and Wangchao Le. 2020. Cost Models for Big Data Query Processing: Learning Retrofitting and Our Findings. In SIGMOD. 99--113.","DOI":"10.1145\/3318464.3380584"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3190650"},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415513"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1080\/00031305.2017.1380080"},{"key":"e_1_2_2_51_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_2_52_1","volume-title":"Machine learning 8","author":"Watkins Christopher JCH","year":"1992","unstructured":"Christopher JCH Watkins and Peter Dayan. 1992. Q-learning. Machine learning 8 (1992), 279--292."},{"key":"e_1_2_2_53_1","volume-title":"Feature Hashing for Large Scale Multitask Learning. In ICML '09","author":"Weinberger Kilian","year":"2009","unstructured":"Kilian Weinberger, Anirban Dasgupta, John Langford, Alex Smola, and Josh Attenberg. 2009. Feature Hashing for Large Scale Multitask Learning. In ICML '09. 1113--1120."},{"key":"e_1_2_2_54_1","first-page":"3","article-title":"Towards a Learning Optimizer for Shared Clouds","volume":"12","author":"Wu Chenggang","year":"2018","unstructured":"Chenggang Wu, Alekh Jindal, Saeed Amizadeh, Hiren Patel, Wangchao Le, Shi Qiao, and Sriram Rao. 2018. Towards a Learning Optimizer for Shared Clouds. PVLDB 12, 3 (nov 2018), 210--222.","journal-title":"PVLDB"},{"key":"e_1_2_2_55_1","doi-asserted-by":"crossref","unstructured":"Haitao Yuan Guoliang Li Ling Feng Ji Sun and Yue Han. 2020. Automatic view generation with deep learning and reinforcement learning. In ICDE. 1501--1512.","DOI":"10.1109\/ICDE48307.2020.00133"},{"key":"e_1_2_2_56_1","volume-title":"VLDB '04","author":"Zilio Daniel C.","year":"2004","unstructured":"Daniel C. Zilio, Jun Rao, Sam Lightstone, Guy Lohman, Adam Storm, Christian Garcia-Arellano, and Scott Fadden. 2004. DB2 Design Advisor: Integrated Automatic Physical Database Design. In VLDB '04. 1087--1097."}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639308","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3639308","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T15:13:16Z","timestamp":1755789196000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639308"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,12]]},"references-count":56,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,3,12]]}},"alternative-id":["10.1145\/3639308"],"URL":"https:\/\/doi.org\/10.1145\/3639308","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,12]]}}}