{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T20:13:21Z","timestamp":1784664801908,"version":"3.55.0"},"reference-count":23,"publisher":"Association for Computing Machinery (ACM)","issue":"8","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2023,4]]},"abstract":"<jats:p>\n            Some recent works have shown the advantages of reinforcement learning (RL) based learned query optimizers. These works often use the cost (i.e., the estimation of cost model) or the latency (i.e., execution time) as guidance signals for training their learned models. However, cost-based learning underperforms in latency and latency-based learning is time-intensive. In order to bypass such a dilemma, researchers attempt to transfer a learned value network from the cost domain to the latency domain. We recognize critical insights in cost\/latency-based training, prompting us to transfer the reward function rather than the value network. Based on this idea, we propose a two-stage RL-based framework,\n            <jats:italic>BASE<\/jats:italic>\n            , to bridge the gap between cost and latency. After learning a policy based on cost signals in its first stage,\n            <jats:italic>BASE<\/jats:italic>\n            formulates transferring the reward function as a variant of inverse reinforcement learning. Intuitively,\n            <jats:italic>BASE<\/jats:italic>\n            learns to calibrate the reward function and updates the policy regarding the calibrated one in a mutually-improved manner. Extensive experiments exhibit the superiority of\n            <jats:italic>BASE<\/jats:italic>\n            on two benchmark datasets: Our optimizer outperforms traditional DBMS, using 30% less training time than SOTA methods. Meanwhile, our approach can enhance the efficiency of other learning-based optimizers.\n          <\/jats:p>","DOI":"10.14778\/3594512.3594525","type":"journal-article","created":{"date-parts":[[2023,6,23]],"date-time":"2023-06-23T00:28:36Z","timestamp":1687480116000},"page":"1958-1966","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":21,"title":["BASE: Bridging the Gap between Cost and Latency for Query Optimization"],"prefix":"10.14778","volume":"16","author":[{"given":"Xu","family":"Chen","sequence":"first","affiliation":[{"name":"University of Electronic Science and Technology of China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhen","family":"Wang","sequence":"additional","affiliation":[{"name":"Alibaba Group"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuncheng","family":"Liu","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yaliang","family":"Li","sequence":"additional","affiliation":[{"name":"Alibaba Group"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kai","family":"Zeng","sequence":"additional","affiliation":[{"name":"Alibaba Group"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bolin","family":"Ding","sequence":"additional","affiliation":[{"name":"Alibaba Group"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jingren","family":"Zhou","sequence":"additional","affiliation":[{"name":"Alibaba Group"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Han","family":"Su","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kai","family":"Zheng","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,6,22]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Zihong Yuan, Pierre Senellart, and St\u00e9phane Bressan.","author":"Basu Debabrota","year":"2015","unstructured":"Debabrota Basu , Qian Lin , Weidong Chen , Hoang Tam Vo , Zihong Yuan, Pierre Senellart, and St\u00e9phane Bressan. 2015 . Cost-model oblivious database tuning with reinforcement learning. In Database and Expert Systems Applications . 253--268. Debabrota Basu, Qian Lin, Weidong Chen, Hoang Tam Vo, Zihong Yuan, Pierre Senellart, and St\u00e9phane Bressan. 2015. Cost-model oblivious database tuning with reinforcement learning. In Database and Expert Systems Applications. 253--268."},{"key":"e_1_2_1_2_1","doi-asserted-by":"crossref","first-page":"313","DOI":"10.1016\/j.ins.2016.10.037","article-title":"Ranked batch-mode active learning","volume":"379","author":"Cardoso Thiago NC","year":"2017","unstructured":"Thiago NC Cardoso , Rodrigo M Silva , S\u00e9rgio Canuto , Mirella M Moro , and Marcos A Gon\u00e7alves . 2017 . Ranked batch-mode active learning . Information Sciences 379 (2017), 313 -- 337 . Thiago NC Cardoso, Rodrigo M Silva, S\u00e9rgio Canuto, Mirella M Moro, and Marcos A Gon\u00e7alves. 2017. Ranked batch-mode active learning. Information Sciences 379 (2017), 313--337.","journal-title":"Information Sciences"},{"key":"e_1_2_1_3_1","volume-title":"Generative adversarial imitation learning. Advances in neural information processing systems 29","author":"Ho Jonathan","year":"2016","unstructured":"Jonathan Ho and Stefano Ermon . 2016. Generative adversarial imitation learning. Advances in neural information processing systems 29 ( 2016 ). Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. Advances in neural information processing systems 29 (2016)."},{"key":"e_1_2_1_4_1","first-page":"4","article-title":"How to train your robot with deep reinforcement learning: lessons we have learned","volume":"40","author":"Ibarz Julian","year":"2021","unstructured":"Julian Ibarz , Jie Tan , Chelsea Finn , Mrinal Kalakrishnan , Peter Pastor , and Sergey Levine . 2021 . How to train your robot with deep reinforcement learning: lessons we have learned . The International Journal of Robotics Research 40 , 4 -- 5 (2021), 698--721. Julian Ibarz, Jie Tan, Chelsea Finn, Mrinal Kalakrishnan, Peter Pastor, and Sergey Levine. 2021. How to train your robot with deep reinforcement learning: lessons we have learned. The International Journal of Robotics Research 40, 4--5 (2021), 698--721.","journal-title":"The International Journal of Robotics Research"},{"key":"e_1_2_1_5_1","unstructured":"Emilie Kaufmann Olivier Capp\u00e9 and Aur\u00e9lien Garivier. 2012. On Bayesian upper confidence bounds for bandit problems. In Artificial intelligence and statistics. PMLR 592--600.  Emilie Kaufmann Olivier Capp\u00e9 and Aur\u00e9lien Garivier. 2012. On Bayesian upper confidence bounds for bandit problems. In Artificial intelligence and statistics. PMLR 592--600."},{"key":"e_1_2_1_6_1","volume-title":"Learning to optimize join queries with deep reinforcement learning. arXiv preprint arXiv:1808.03196","author":"Krishnan Sanjay","year":"2018","unstructured":"Sanjay Krishnan , Zongheng Yang , Ken Goldberg , Joseph Hellerstein , and Ion Stoica . 2018. Learning to optimize join queries with deep reinforcement learning. arXiv preprint arXiv:1808.03196 ( 2018 ). Sanjay Krishnan, Zongheng Yang, Ken Goldberg, Joseph Hellerstein, and Ion Stoica. 2018. Learning to optimize join queries with deep reinforcement learning. arXiv preprint arXiv:1808.03196 (2018)."},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","first-page":"204","DOI":"10.14778\/2850583.2850594","article-title":"How good are query optimizers, really","volume":"9","author":"Leis Viktor","year":"2015","unstructured":"Viktor Leis , Andrey Gubichev , Atanas Mirchev , Peter Boncz , Alfons Kemper , and Thomas Neumann . 2015 . How good are query optimizers, really ? Proceedings of the VLDB Endowment 9 , 3 (2015), 204 -- 215 . Viktor Leis, Andrey Gubichev, Atanas Mirchev, Peter Boncz, Alfons Kemper, and Thomas Neumann. 2015. How good are query optimizers, really? Proceedings of the VLDB Endowment 9, 3 (2015), 204--215.","journal-title":"Proceedings of the VLDB Endowment"},{"key":"e_1_2_1_8_1","volume-title":"Machine learning proceedings","author":"Lewis David D","year":"1994","unstructured":"David D Lewis and Jason Catlett . 1994. Heterogeneous uncertainty sampling for supervised learning . In Machine learning proceedings 1994 . Elsevier , 148--156. David D Lewis and Jason Catlett. 1994. Heterogeneous uncertainty sampling for supervised learning. In Machine learning proceedings 1994. Elsevier, 148--156."},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data. 175--191","author":"Ma Lin","year":"2020","unstructured":"Lin Ma , Bailu Ding , Sudipto Das , and Adith Swaminathan . 2020 . Active learning for ML enhanced database systems . In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data. 175--191 . Lin Ma, Bailu Ding, Sudipto Das, and Adith Swaminathan. 2020. Active learning for ML enhanced database systems. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data. 175--191."},{"key":"e_1_2_1_10_1","doi-asserted-by":"crossref","first-page":"6","DOI":"10.1145\/3542700.3542703","article-title":"Bao: Making learned query optimization practical","volume":"51","author":"Marcus Ryan","year":"2022","unstructured":"Ryan Marcus , Parimarjan Negi , Hongzi Mao , Nesime Tatbul , Mohammad Alizadeh , and Tim Kraska . 2022 . Bao: Making learned query optimization practical . ACM SIGMOD Record 51 , 1 (2022), 6 -- 13 . Ryan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul, Mohammad Alizadeh, and Tim Kraska. 2022. Bao: Making learned query optimization practical. ACM SIGMOD Record 51, 1 (2022), 6--13.","journal-title":"ACM SIGMOD Record"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the VLDB Endowment 12","author":"Marcus Ryan","year":"2019","unstructured":"Ryan Marcus , Parimarjan Negi , Hongzi Mao , Chi Zhang , Mohammad Alizadeh , Tim Kraska , Olga Papaemmanouil , and Nesime Tatbul . 2019 . Neo: A learned query optimizer . Proceedings of the VLDB Endowment 12 , 11 (2019). Ryan Marcus, Parimarjan Negi, Hongzi Mao, Chi Zhang, Mohammad Alizadeh, Tim Kraska, Olga Papaemmanouil, and Nesime Tatbul. 2019. Neo: A learned query optimizer. Proceedings of the VLDB Endowment 12, 11 (2019)."},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the First International Workshop on Exploiting Artificial Intelligence Techniques for Data Management. 1--4.","author":"Marcus Ryan","year":"2018","unstructured":"Ryan Marcus and Olga Papaemmanouil . 2018 . Deep reinforcement learning for join order enumeration . In Proceedings of the First International Workshop on Exploiting Artificial Intelligence Techniques for Data Management. 1--4. Ryan Marcus and Olga Papaemmanouil. 2018. Deep reinforcement learning for join order enumeration. In Proceedings of the First International Workshop on Exploiting Artificial Intelligence Techniques for Data Management. 1--4."},{"key":"e_1_2_1_13_1","volume-title":"2019 IEEE 15th International Conference on Intelligent Computer Communication and Processing (ICCP). IEEE, 509--515","author":"Miok Kristian","year":"2019","unstructured":"Kristian Miok , Dong Nguyen-Doan , Daniela Zaharie , and Marko Robnik-\u0160ikonja . 2019 . Generating data using Monte Carlo dropout . In 2019 IEEE 15th International Conference on Intelligent Computer Communication and Processing (ICCP). IEEE, 509--515 . Kristian Miok, Dong Nguyen-Doan, Daniela Zaharie, and Marko Robnik-\u0160ikonja. 2019. Generating data using Monte Carlo dropout. In 2019 IEEE 15th International Conference on Intelligent Computer Communication and Processing (ICCP). IEEE, 509--515."},{"key":"e_1_2_1_14_1","volume-title":"Icml","volume":"99","author":"Ng Andrew Y","year":"1999","unstructured":"Andrew Y Ng , Daishi Harada , and Stuart Russell . 1999 . Policy invariance under reward transformations: Theory and application to reward shaping . In Icml , Vol. 99 . Citeseer, 278--287. Andrew Y Ng, Daishi Harada, and Stuart Russell. 1999. Policy invariance under reward transformations: Theory and application to reward shaping. In Icml, Vol. 99. Citeseer, 278--287."},{"key":"e_1_2_1_15_1","first-page":"1345","article-title":"A survey on transfer learning","volume":"22","author":"Pan Sinno Jialin","year":"2009","unstructured":"Sinno Jialin Pan and Qiang Yang . 2009 . A survey on transfer learning . IEEE Transactions on knowledge and data engineering 22 , 10 (2009), 1345 -- 1359 . Sinno Jialin Pan and Qiang Yang. 2009. A survey on transfer learning. IEEE Transactions on knowledge and data engineering 22, 10 (2009), 1345--1359.","journal-title":"IEEE Transactions on knowledge and data engineering"},{"key":"e_1_2_1_16_1","unstructured":"RyanMarcus. 2021. BAO source code. https:\/\/github.com\/learnedsystems\/BaoForPostgreSQL.  RyanMarcus. 2021. BAO source code. https:\/\/github.com\/learnedsystems\/BaoForPostgreSQL."},{"key":"e_1_2_1_17_1","volume-title":"Felix Martin Schuhknecht, and Jens Dittrich","author":"Sharma Ankur","year":"2018","unstructured":"Ankur Sharma , Felix Martin Schuhknecht, and Jens Dittrich . 2018 . The case for automatic database administration using deep reinforcement learning. arXiv preprint arXiv:1801.05643 (2018). Ankur Sharma, Felix Martin Schuhknecht, and Jens Dittrich. 2018. The case for automatic database administration using deep reinforcement learning. arXiv preprint arXiv:1801.05643 (2018)."},{"key":"e_1_2_1_18_1","volume-title":"End-to-End Robotic Reinforcement Learning without Reward Engineering. environment (eg, by placing additional sensors) 34","author":"Singh Avi","year":"2019","unstructured":"Avi Singh , Larry Yang , Kristian Hartikainen , Chelsea Finn , and Sergey Levine . 2019. End-to-End Robotic Reinforcement Learning without Reward Engineering. environment (eg, by placing additional sensors) 34 ( 2019 ), 44. Avi Singh, Larry Yang, Kristian Hartikainen, Chelsea Finn, and Sergey Levine. 2019. End-to-End Robotic Reinforcement Learning without Reward Engineering. environment (eg, by placing additional sensors) 34 (2019), 44."},{"key":"e_1_2_1_19_1","unstructured":"The PostgreSQL Global Development Group. 2022. PostgreSQL 10 Documentation. https:\/\/www.postgresql.org\/docs\/10\/index.html.  The PostgreSQL Global Development Group. 2022. PostgreSQL 10 Documentation. https:\/\/www.postgresql.org\/docs\/10\/index.html."},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the 2022 International Conference on Management of Data. 931--944","author":"Yang Zongheng","year":"2022","unstructured":"Zongheng Yang , Wei-Lin Chiang , Sifei Luan , Gautam Mittal , Michael Luo , and Ion Stoica . 2022 . Balsa: Learning a Query Optimizer Without Expert Demonstrations . In Proceedings of the 2022 International Conference on Management of Data. 931--944 . Zongheng Yang, Wei-Lin Chiang, Sifei Luan, Gautam Mittal, Michael Luo, and Ion Stoica. 2022. Balsa: Learning a Query Optimizer Without Expert Demonstrations. In Proceedings of the 2022 International Conference on Management of Data. 931--944."},{"key":"e_1_2_1_21_1","volume-title":"2020 IEEE 36th International Conference on Data Engineering (ICDE). IEEE, 1297--1308","author":"Yu Xiang","year":"2020","unstructured":"Xiang Yu , Guoliang Li , Chengliang Chai , and Nan Tang . 2020 . Reinforcement learning with tree-lstm for join order selection . In 2020 IEEE 36th International Conference on Data Engineering (ICDE). IEEE, 1297--1308 . Xiang Yu, Guoliang Li, Chengliang Chai, and Nan Tang. 2020. Reinforcement learning with tree-lstm for join order selection. In 2020 IEEE 36th International Conference on Data Engineering (ICDE). IEEE, 1297--1308."},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of the 2022 International Conference on Management of Data. 2299--2311","author":"Zhang Wangda","year":"2022","unstructured":"Wangda Zhang , Matteo Interlandi , Paul Mineiro , Shi Qiao , Nasim Ghazanfari , Karlen Lie , Marc Friedman , Rafah Hosn , Hiren Patel , and Alekh Jindal . 2022 . Deploying a steered query optimizer in production at Microsoft . In Proceedings of the 2022 International Conference on Management of Data. 2299--2311 . Wangda Zhang, Matteo Interlandi, Paul Mineiro, Shi Qiao, Nasim Ghazanfari, Karlen Lie, Marc Friedman, Rafah Hosn, Hiren Patel, and Alekh Jindal. 2022. Deploying a steered query optimizer in production at Microsoft. In Proceedings of the 2022 International Conference on Management of Data. 2299--2311."},{"key":"e_1_2_1_23_1","volume-title":"Diverse mini-batch active learning. arXiv preprint arXiv:1901.05954","author":"Zhdanov Fedor","year":"2019","unstructured":"Fedor Zhdanov . 2019. Diverse mini-batch active learning. arXiv preprint arXiv:1901.05954 ( 2019 ). Fedor Zhdanov. 2019. Diverse mini-batch active learning. arXiv preprint arXiv:1901.05954 (2019)."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3594512.3594525","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,23]],"date-time":"2023-06-23T00:34:55Z","timestamp":1687480495000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3594512.3594525"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,4]]},"references-count":23,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2023,4]]}},"alternative-id":["10.14778\/3594512.3594525"],"URL":"https:\/\/doi.org\/10.14778\/3594512.3594525","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2023,4]]},"assertion":[{"value":"2023-06-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}