{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T14:33:31Z","timestamp":1784903611756,"version":"3.55.0"},"reference-count":71,"publisher":"Association for Computing Machinery (ACM)","issue":"10","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2024,6]]},"abstract":"<jats:p>\n            Database administrators (DBAs) play an important role in managing database systems. However, it is hard and tedious for DBAs to manage vast database instances and give timely response (waiting for hours is intolerable in many online cases). In addition, existing empirical methods only support limited diagnosis scenarios, which are also labor-intensive to update the diagnosis rules for database version updates. Recently large language models (LLMs) have shown great potential in various fields. Thus, we propose\n            <jats:italic>D-Bot<\/jats:italic>\n            , an LLM-based database diagnosis system that can automatically acquire knowledge from diagnosis documents, and generate reasonable and well-founded diagnosis report (i.e., identifying the root causes and solutions) within acceptable time (e.g., under 10 minutes compared to hours by a DBA). The techniques in\n            <jats:italic>D-Bot<\/jats:italic>\n            include (\n            <jats:italic>i<\/jats:italic>\n            ) offline knowledge extraction from documents, (\n            <jats:italic>ii<\/jats:italic>\n            ) automatic prompt generation (e.g., knowledge matching, tool retrieval), (\n            <jats:italic>iii<\/jats:italic>\n            ) root cause analysis using tree search algorithm, and (\n            <jats:italic>iv<\/jats:italic>\n            ) collaborative mechanism for complex anomalies with multiple root causes. We verify\n            <jats:italic>D-Bot<\/jats:italic>\n            on real benchmarks (including 539 anomalies of six typical applications), and the results show\n            <jats:italic>D-Bot<\/jats:italic>\n            can effectively identify root causes of unseen anomalies and\n            <jats:italic>significantly outperforms traditional methods and vanilla models like GPT-4.<\/jats:italic>\n          <\/jats:p>","DOI":"10.14778\/3675034.3675043","type":"journal-article","created":{"date-parts":[[2024,8,6]],"date-time":"2024-08-06T22:19:11Z","timestamp":1722982751000},"page":"2514-2527","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":36,"title":["D-Bot: Database Diagnosis System using Large Language Models"],"prefix":"10.14778","volume":"17","author":[{"given":"Xuanhe","family":"Zhou","sequence":"first","affiliation":[{"name":"Department of Computer Science, Tsinghua University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guoliang","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Tsinghua University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhaoyan","family":"Sun","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Tsinghua University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhiyuan","family":"Liu","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Tsinghua University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weize","family":"Chen","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Tsinghua University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianming","family":"Wu","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Tsinghua University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiesi","family":"Liu","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Tsinghua University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ruohang","family":"Feng","sequence":"additional","affiliation":[{"name":"Pigsty"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guoyang","family":"Zeng","sequence":"additional","affiliation":[{"name":"ModelBest"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,8,6]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Retrieved","author":"PG.","year":"2023","unstructured":"2023. HypoPG. Retrieved December 1, 2023 from https:\/\/github.com\/HypoPG\/hypopg"},{"key":"e_1_2_1_2_1","volume-title":"Retrieved","author":"AI.","year":"2023","unstructured":"2023. OpenAI. Retrieved December 1, 2023 from https:\/\/openai.com\/"},{"key":"e_1_2_1_3_1","volume-title":"PGTune - calculate configuration for PostgreSQL based on the maximum performance for a given hardware configuration. Retrieved","year":"2023","unstructured":"2023. PGTune - calculate configuration for PostgreSQL based on the maximum performance for a given hardware configuration. Retrieved December 1, 2023 from https:\/\/pgtune.leopard.in.ua\/"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the 2018 International Conference on Management of Data. 221--230","author":"Begoli Edmon","year":"2018","unstructured":"Edmon Begoli, Jes\u00fas Camacho-Rodr\u00edguez, Julian Hyde, Michael J Mior, and Daniel Lemire. 2018. Apache calcite: A foundational framework for optimized query processing over heterogeneous data sources. In Proceedings of the 2018 International Conference on Management of Data. 221--230."},{"key":"e_1_2_1_5_1","volume-title":"Kolmogorov-smirnov test: Overview","author":"Berger Vance W","year":"2014","unstructured":"Vance W Berger and YanYan Zhou. 2014. Kolmogorov-smirnov test: Overview. Wiley statsref: Statistics reference online (2014)."},{"key":"e_1_2_1_6_1","unstructured":"Harrison Chase. 2022. LangChain. https:\/\/github.com\/hwchase17\/langchain."},{"key":"e_1_2_1_7_1","volume-title":"Narasayya","author":"Chaudhuri Surajit","year":"1997","unstructured":"Surajit Chaudhuri and Vivek R. Narasayya. 1997. An Efficient Cost-Driven Index Selection Tool for Microsoft SQL Server. In VLDB. 146--155."},{"key":"e_1_2_1_8_1","volume-title":"Xgboost: A scalable tree boosting system. In sigkdd. 785--794.","author":"Chen Tianqi","year":"2016","unstructured":"Tianqi Chen and Carlos Guestrin. 2016. Xgboost: A scalable tree boosting system. In sigkdd. 785--794."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2211.12588"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2308.10848"},{"key":"e_1_2_1_11_1","volume-title":"Second Biennial Conference on Innovative Data Systems Research, CIDR","author":"Dias Karl","year":"2005","unstructured":"Karl Dias, Mark Ramacher, Uri Shaft, Venkateshwaran Venkataramani, and Graham Wood. 2005. Automatic Performance Diagnosis and Tuning in Oracle. In Second Biennial Conference on Innovative Data Systems Research, CIDR 2005, Asilomar, CA, USA, January 4-7, 2005, Online Proceedings. www.cidrdb.org, 84--94. http:\/\/cidrdb.org\/cidr2005\/papers\/P07.pdf"},{"key":"e_1_2_1_12_1","volume-title":"Openprompt: An open-source framework for prompt-learning. arXiv preprint arXiv:2111.01998","author":"Ding Ning","year":"2021","unstructured":"Ning Ding, Shengding Hu, Weilin Zhao, Yulin Chen, Zhiyuan Liu, Hai-Tao Zheng, and Maosong Sun. 2021. Openprompt: An open-source framework for prompt-learning. arXiv preprint arXiv:2111.01998 (2021)."},{"key":"e_1_2_1_13_1","volume-title":"Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping. arXiv preprint arXiv:2002.06305","author":"Dodge Jesse","year":"2020","unstructured":"Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020. Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping. arXiv preprint arXiv:2002.06305 (2020)."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2305.14325"},{"key":"e_1_2_1_15_1","volume-title":"PAL: Program-aided Language Models. In International Conference on Machine Learning, ICML 2023","volume":"10799","author":"Gao Luyu","year":"2023","unstructured":"Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. PAL: Program-aided Language Models. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceedings of Machine Learning Research, Vol. 202), Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 10764--10799. https:\/\/proceedings.mlr.press\/v202\/gao23f.html"},{"key":"e_1_2_1_16_1","volume-title":"KNN model-based approach in classification","author":"Guo Gongde","unstructured":"Gongde Guo, Hui Wang, David Bell, Yaxin Bi, and Kieran Greer. 2003. KNN model-based approach in classification. In CoopIS. Springer, 986--996."},{"key":"e_1_2_1_17_1","unstructured":"Li Hai-Xiang Li Xiao-Yan Liu Chang and et al. 2021. Systematic definition and classification of data anomalies in DBMS (English Version). arXiv preprint arXiv:2110.14230 (2021)."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2308.00352"},{"key":"e_1_2_1_19_1","first-page":"1","volume-title":"Proc. ACM Manag. Data 1","author":"Huang Shiyue","year":"2023","unstructured":"Shiyue Huang, Ziwei Wang, Xinyi Zhang, Yaofeng Tu, Zhongliang Li, and Bin Cui. 2023. DBPA: A Benchmark for Transactional Database Performance Anomalies. Proc. ACM Manag. Data 1, 1 (2023), 72:1--72:26. 10.1145\/3588926"},{"key":"e_1_2_1_20_1","first-page":"845","article-title":"AI-based Database Performance Diagnosis","volume":"32","author":"Jin Lianyuan","year":"2021","unstructured":"Lianyuan Jin and Guoliang Li. 2021. AI-based Database Performance Diagnosis. Journal of Software 32, 3 (2021), 845--858.","journal-title":"Journal of Software"},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the 2019 International Conference on Management of Data, SIGMOD Conference 2019","author":"Kalmegh Prajakta","year":"2019","unstructured":"Prajakta Kalmegh, Shivnath Babu, and Sudeepa Roy. 2019. iQCAR: inter-Query Contention Analyzer for Data Analytics Frameworks. In Proceedings of the 2019 International Conference on Management of Data, SIGMOD Conference 2019, Amsterdam, The Netherlands, June 30 - July 5, 2019, Peter A. Boncz, Stefan Manegold, Anastasia Ailamaki, Amol Deshpande, and Tim Kraska (Eds.). ACM, 918--935. 10.1145\/3299869.3319904"},{"key":"e_1_2_1_22_1","volume-title":"ICADIWT","author":"Khan Kamran","year":"2014","unstructured":"Kamran Khan, Saif Ur Rehman, Kamran Aziz, Simon Fong, and Sababady Sarasvady. 2014. DBSCAN: Past, present and future. In ICADIWT 2014. IEEE, 232--238."},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the NetDB","volume":"11","author":"Kreps Jay","year":"2011","unstructured":"Jay Kreps, Neha Narkhede, Jun Rao, et al. 2011. Kafka: A distributed messaging system for log processing. In Proceedings of the NetDB, Vol. 11. Athens, Greece, 1--7."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","unstructured":"Hai Lan Zhifeng Bao and Yuwei Peng. 2020. An Index Advisor Using Deep Reinforcement Learning. In CIKM. ACM 2105--2108. 10.1145\/3340531.3412106","DOI":"10.1145\/3340531.3412106"},{"key":"e_1_2_1_25_1","volume-title":"NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following.","author":"Li Dacheng","year":"2023","unstructured":"Dacheng Li, Rulin Shao, Anze Xie, Ying Sheng, Lianmin Zheng, Joseph Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang. 2023. How Long Can Context Length of Open-Source LLMs truly Promise?. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2303.17760"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.14778\/3352063.3352129"},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 996--1008","author":"Li Zeyan","year":"2022","unstructured":"Zeyan Li, Nengwen Zhao, Mingjie Li, et al. 2022. Actionable and interpretable fault localization for recurring failures in online service systems. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 996--1008."},{"key":"e_1_2_1_29_1","volume-title":"FluxInfer: Automatic Diagnosis of Performance Anomaly for Online Database System. In 39th IEEE International Performance Computing and Communications Conference, IPCCC 2020","author":"Liu Ping","year":"2020","unstructured":"Ping Liu, Shenglin Zhang, Yongqian Sun, Yuan Meng, Jiahai Yang, and Dan Pei. 2020. FluxInfer: Automatic Diagnosis of Performance Anomaly for Online Database System. In 39th IEEE International Performance Computing and Communications Conference, IPCCC 2020, Austin, TX, USA, November 6-8, 2020. IEEE, 1--8. 10.1109\/IPCCC50635.2020.9391550"},{"key":"e_1_2_1_30_1","volume-title":"38th IEEE International Conference on Data Engineering, ICDE 2022","author":"Liu Xiaoze","year":"2022","unstructured":"Xiaoze Liu, Zheng Yin, Chao Zhao, Congcong Ge, Lu Chen, Yunjun Gao, Dimeng Li, Ziting Wang, Gaozhong Liang, Jian Tan, and Feifei Li. 2022. PinSQL: Pinpoint Root Cause SQLs to Resolve Performance Issues in Cloud Databases. In 38th IEEE International Conference on Data Engineering, ICDE 2022, Kuala Lumpur, Malaysia, May 9-12, 2022. IEEE, 2549--2561. 10.1109\/ICDE53745.2022.00236"},{"key":"e_1_2_1_31_1","unstructured":"Yuhe Liu Changhua Pei Longlong Xu Bohan Chen Mingze Sun Zhirui Zhang Yongqian Sun Shenglin Zhang Kun Wang Haiming Zhang et al. 2023. OpsEval: A Comprehensive Task-Oriented AIOps Benchmark for Large Language Models. arXiv preprint arXiv:2310.07637 (2023)."},{"key":"e_1_2_1_32_1","volume-title":"22nd IEEE International Symposium on Cluster, Cloud and Internet Computing, CCGrid 2022","author":"Lu Xianglin","year":"2022","unstructured":"Xianglin Lu, Zhe Xie, Zeyan Li, Mingjie Li, Xiaohui Nie, Nengwen Zhao, Qingyang Yu, Shenglin Zhang, Kaixin Sui, Lin Zhu, and Dan Pei. 2022. Generic and Robust Performance Diagnosis via Causal Inference for OLTP Database Systems. In 22nd IEEE International Symposium on Cluster, Cloud and Internet Computing, CCGrid 2022, Taormina, Italy, May 16-19, 2022. IEEE, 655--664. 10.1109\/CCGrid54584.2022.00075"},{"key":"e_1_2_1_33_1","unstructured":"Killian Lucas. 2023. Open Interpreter. https:\/\/github.com\/charlespwd\/project-title."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.14778\/3389133.3389136"},{"key":"e_1_2_1_35_1","doi-asserted-by":"crossref","first-page":"621","DOI":"10.1016\/j.procs.2016.04.140","article-title":"Towards identifying performance anomalies","volume":"83","author":"Malik Haroon","year":"2016","unstructured":"Haroon Malik and Elhadi M Shakshuki. 2016. Towards identifying performance anomalies. Procedia Computer Science 83 (2016), 621--627.","journal-title":"Procedia Computer Science"},{"key":"e_1_2_1_36_1","volume-title":"WebGPT: Browser-assisted question-answering with human feedback. CoRR abs\/2112.09332","author":"Nakano Reiichiro","year":"2021","unstructured":"Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021. WebGPT: Browser-assisted question-answering with human feedback. CoRR abs\/2112.09332 (2021). arXiv:2112.09332 https:\/\/arxiv.org\/abs\/2112.09332"},{"key":"e_1_2_1_37_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3066156","article-title":"DUCT: An upper confidence bound approach to distributed constraint optimization problems","volume":"8","author":"Ottens Brammert","year":"2017","unstructured":"Brammert Ottens, Christos Dimitrakakis, and Boi Faltings. 2017. DUCT: An upper confidence bound approach to distributed constraint optimization problems. ACM Transactions on Intelligent Systems and Technology (TIST) 8, 5 (2017), 1--27.","journal-title":"ACM Transactions on Intelligent Systems and Technology (TIST)"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2304.03442"},{"key":"e_1_2_1_39_1","unstructured":"Chen Qian Xin Cong Cheng Yang Weize Chen Yusheng Su and et al. 2023. Communicative Agents for Software Development. arXiv preprint arXiv:2307.07924 (2023)."},{"key":"e_1_2_1_40_1","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023","author":"Qin Yujia","year":"2023","unstructured":"Yujia Qin, Zihan Cai, Dian Jin, Lan Yan, Shihao Liang, Kunlun Zhu, Yankai Lin, Xu Han, Ning Ding, Huadong Wang, Ruobing Xie, Fanchao Qi, Zhiyuan Liu, Maosong Sun, and Jie Zhou. 2023. WebCPM: Interactive Web Search for Chinese Long-form Question Answering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, 8968--8988. 10.18653\/v1\/2023.acl-long.499"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","unstructured":"Yujia Qin Shengding Hu Yankai Lin Weize Chen Ning Ding Ganqu Cui Zheni Zeng Yufei Huang Chaojun Xiao Chi Han Yi Ren Fung Yusheng Su Huadong Wang Cheng Qian Runchu Tian Kunlun Zhu Shihao Liang Xingyu Shen Bokai Xu Zhen Zhang Yining Ye Bowen Li Ziwei Tang Jing Yi Yuzhang Zhu Zhenning Dai Lan Yan Xin Cong Yaxi Lu Weilin Zhao Yuxiang Huang Junxi Yan Xu Han Xian Sun Dahai Li Jason Phang Cheng Yang Tongshuang Wu Heng Ji Zhiyuan Liu and Maosong Sun. 2023. Tool Learning with Foundation Models. CoRR abs\/2304.08354 (2023). arXiv:2304.08354 10.48550\/arXiv.2304.08354","DOI":"10.48550\/arXiv.2304.08354"},{"key":"e_1_2_1_42_1","unstructured":"Yujia Qin Shengding Hu Yankai Lin and et al. 2023. Tool learning with foundation models. arXiv preprint arXiv:2304.08354 (2023)."},{"key":"e_1_2_1_43_1","unstructured":"Yujia Qin Shihao Liang Yining Ye Kunlun Zhu Lan Yan Yaxi Lu Yankai Lin Xin Cong Xiangru Tang Bill Qian Sihan Zhao Runchu Tian Ruobing Xie Jie Zhou Mark Gerstein Dahai Li Zhiyuan Liu and Maosong Sun. 2023. Tool-LLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. arXiv:2307.16789 [cs.AI]"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2307.16789"},{"key":"e_1_2_1_45_1","volume-title":"Anomaly detection in role administered relational databases---A novel method","author":"Ramachandran Raji","unstructured":"Raji Ramachandran, R Nidhin, and PP Shogil. 2018. Anomaly detection in role administered relational databases---A novel method. In ICACCI. IEEE, 1017--1021."},{"key":"e_1_2_1_46_1","volume-title":"A Survey of Hallucination in Large Foundation Models. arXiv preprint arXiv:2309.05922","author":"Rawte Vipula","year":"2023","unstructured":"Vipula Rawte, Amit Sheth, and Amitava Das. 2023. A Survey of Hallucination in Large Foundation Models. arXiv preprint arXiv:2309.05922 (2023)."},{"key":"e_1_2_1_47_1","volume-title":"Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084","author":"Reimers Nils","year":"2019","unstructured":"Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084 (2019)."},{"key":"e_1_2_1_48_1","doi-asserted-by":"crossref","unstructured":"Stephen Robertson Hugo Zaragoza et al. 2009. The probabilistic relevance framework: BM25 and beyond. Foundations and Trends\u00ae in Information Retrieval 3 4 (2009) 333--389.","DOI":"10.1561\/1500000019"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2302.04761"},{"key":"e_1_2_1_50_1","volume-title":"Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems 36","author":"Shinn Noah","year":"2024","unstructured":"Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems 36 (2024)."},{"key":"e_1_2_1_51_1","volume-title":"Decision tree methods: applications for classification and prediction. Shanghai archives of psychiatry 27, 2","author":"Song Yan-Yan","year":"2015","unstructured":"Yan-Yan Song and LU Ying. 2015. Decision tree methods: applications for classification and prediction. Shanghai archives of psychiatry 27, 2 (2015), 130."},{"key":"e_1_2_1_52_1","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023","author":"Trivedi Harsh","year":"2023","unstructured":"Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2023. Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, 10014--10037. 10.18653\/v1\/2023.acl-long.557"},{"key":"e_1_2_1_53_1","volume-title":"SIGMOD '22: International Conference on Management of Data","author":"Trummer Immanuel","year":"2022","unstructured":"Immanuel Trummer. 2022. DB-BERT: A Database Tuning Tool that \"Reads the Manual\". In SIGMOD '22: International Conference on Management of Data, Philadelphia, PA, USA, June 12 - 17, 2022, Zachary G. Ives, Angela Bonifati, and Amr El Abbadi (Eds.). ACM, 190--203. 10.1145\/3514221.3517843"},{"key":"e_1_2_1_54_1","unstructured":"Gary Valentin Michael Zuliani Daniel C. Zilio Guy M. Lohman and Alan Skelley. 2000. DB2 Advisor: An Optimizer Smart Enough to Recommend Its Own Indexes. In ICDE. 101--110."},{"key":"e_1_2_1_55_1","volume-title":"Proceedings of the 2022 International Conference on Management of Data. 94--107","author":"Wang Zhaoguo","year":"2022","unstructured":"Zhaoguo Wang, Zhou Zhou, Yicun Yang, Haoran Ding, Gansen Hu, Ding Ding, Chuzhe Tang, Haibo Chen, and Jinyang Li. 2022. Wetune: Automatic discovery and verification of query rewrite rules. In Proceedings of the 2022 International Conference on Management of Data. 94--107."},{"key":"e_1_2_1_56_1","volume-title":"Denny Zhou, et al.","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824--24837."},{"key":"e_1_2_1_57_1","volume-title":"Index Selection in Relational Databases","author":"Whang Kyu-Young","year":"1987","unstructured":"Kyu-Young Whang. 1987. Index Selection in Relational Databases. Foundations of Data Organization (1987), 487--500."},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2308.08155"},{"key":"e_1_2_1_59_1","unstructured":"Shunyu Yao Dian Yu Jeffrey Zhao Izhak Shafran Tom Griffiths Yuan Cao and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In NeurIPS Alice Oh Tristan Naumann Amir Globerson Kate Saenko Moritz Hardt and Sergey Levine (Eds.). http:\/\/papers.nips.cc\/paper_files\/paper\/2023\/hash\/271db9922b8d1f4dd7aaef84ed5ac703-Abstract-Conference.html"},{"key":"e_1_2_1_60_1","volume-title":"React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629","author":"Yao Shunyu","year":"2022","unstructured":"Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629 (2022)."},{"key":"e_1_2_1_61_1","volume-title":"ReAct: Synergizing Reasoning and Acting in Language Models. In The Eleventh International Conference on Learning Representations, ICLR 2023","author":"Yao Shunyu","year":"2023","unstructured":"Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net. https:\/\/openreview.net\/pdf?id=WE_vluYUL-X"},{"key":"e_1_2_1_62_1","volume-title":"Proceedings of the 2016 International Conference on Management of Data, SIGMOD Conference 2016","author":"Yoon Dong Young","year":"2016","unstructured":"Dong Young Yoon, Ning Niu, and Barzan Mozafari. 2016. DBSherlock: A Performance Diagnostic Tool for Transactional Databases. In Proceedings of the 2016 International Conference on Management of Data, SIGMOD Conference 2016, San Francisco, CA, USA, June 26 - July 01, 2016, Fatma \u00d6zcan, Georgia Koutrika, and Sam Madden (Eds.). ACM, 1599--1614. 10.1145\/2882903.2915218"},{"key":"e_1_2_1_63_1","volume-title":"Representation Learning for Natural Language Processing","author":"Zeng Guoyang","unstructured":"Guoyang Zeng, Xu Han, Zhengyan Zhang, Zhiyuan Liu, Yankai Lin, and Maosong Sun. 2023. OpenBMB: Big Model Systems for Large-Scale Representation Learning. In Representation Learning for Natural Language Processing. Springer Nature Singapore Singapore, 463--489."},{"key":"e_1_2_1_64_1","volume-title":"Proc. VLDB Endow.","author":"Zhao Xinyang","year":"2024","unstructured":"Xinyang Zhao, Xuanhe Zhou, and Guoliang Li. 2024. Chat2Data: An Interactive Data Analysis System with RAG, Vector Databases and LLMs. Proc. VLDB Endow. (2024)."},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.14778\/3485450.3485456"},{"key":"e_1_2_1_66_1","volume-title":"Llm as dba. arXiv preprint arXiv:2308.05481","author":"Zhou Xuanhe","year":"2023","unstructured":"Xuanhe Zhou, Guoliang Li, and Zhiyuan Liu. 2023. Llm as dba. arXiv preprint arXiv:2308.05481 (2023)."},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.14778\/3611540.3611633"},{"key":"e_1_2_1_68_1","volume-title":"Autoindex: An incremental index management system for dynamic workloads","author":"Zhou Xuanhe","year":"2022","unstructured":"Xuanhe Zhou, Luyang Liu, and et al. 2022. Autoindex: An incremental index management system for dynamic workloads. In ICDE. IEEE, 2196--2208."},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1007\/S41019-023-00235-6"},{"key":"e_1_2_1_70_1","volume-title":"Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba.","author":"Zhou Yongchao","year":"2022","unstructured":"Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2022. Large Language Models Are Human-Level Prompt Engineers. (2022). arXiv:2211.01910 http:\/\/arxiv.org\/abs\/2211.01910"},{"key":"e_1_2_1_71_1","volume-title":"Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba.","author":"Zhou Yongchao","year":"2022","unstructured":"Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2022. Large language models are human-level prompt engineers. arXiv preprint arXiv:2211.01910 (2022)."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3675034.3675043","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,6]],"date-time":"2024-08-06T22:28:01Z","timestamp":1722983281000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3675034.3675043"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6]]},"references-count":71,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2024,6]]}},"alternative-id":["10.14778\/3675034.3675043"],"URL":"https:\/\/doi.org\/10.14778\/3675034.3675043","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2024,6]]},"assertion":[{"value":"2024-08-06","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}