{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T17:17:06Z","timestamp":1783790226958,"version":"3.55.0"},"reference-count":20,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2023,8]]},"abstract":"<jats:p>In the age of the Digital Revolution, almost all human activities, from industrial and business operations to medical and academic research, are reliant on the constant integration and utilisation of ever-increasing volumes of data. However, the explosive volume and complexity of data makes data querying and exploration challenging even for experts, and makes the need to democratise the access to data, even for non-technical users, all the more evident. It is time to lift all technical barriers, by empowering users to access relational databases through conversation. We consider 3 main research areas that a natural language data interface is based on: Text-to-SQL, SQL-to-Text, and Data-to-Text. The purpose of this tutorial is a deep dive into these areas, covering state-of-the-art techniques and models, and explaining how the progress in the deep learning field has led to impressive advancements. We will present benchmarks that sparked research and competition, and discuss open problems and research opportunities with one of the most important challenges being the integration of these 3 research areas into one conversational system.<\/jats:p>","DOI":"10.14778\/3611540.3611575","type":"journal-article","created":{"date-parts":[[2023,9,15]],"date-time":"2023-09-15T11:32:37Z","timestamp":1694777557000},"page":"3878-3881","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["Natural Language Interfaces for Databases with Deep Learning"],"prefix":"10.14778","volume":"16","author":[{"given":"George","family":"Katsogiannis-Meimarakis","sequence":"first","affiliation":[{"name":"Athena Research Center, Athens, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mike","family":"Xydas","sequence":"additional","affiliation":[{"name":"Athena Research Center, Athens, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Georgia","family":"Koutrika","sequence":"additional","affiliation":[{"name":"Athena Research Center, Athens, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,8]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6243"},{"key":"e_1_2_1_2_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 [cs.CL]","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 [cs.CL]"},{"key":"e_1_2_1_3_1","volume-title":"55th annual meeting of the Association for Computational Linguistics (ACL).","author":"Gardent Claire","unstructured":"Claire Gardent, Anastasia Shimorina, Shashi Narayan, and Laura Perez-Beltrachini. 2017. Creating training corpora for nlg micro-planning. In 55th annual meeting of the Association for Computational Linguistics (ACL)."},{"key":"e_1_2_1_4_1","unstructured":"Wonseok Hwang Jinyeong Yim Seunghyun Park and Minjoon Seo. 2019. A Comprehensive Exploration on WikiSQL with Table-Aware Word Contextualization. arXiv:1902.01069 [cs.CL]"},{"key":"e_1_2_1_5_1","volume-title":"Curriculum-Based Self-Training Makes Better Few-Shot Learners for Data-to-Text Generation. arXiv:2206.02712","author":"Ke Pei","year":"2022","unstructured":"Pei Ke, Haozhe Ji, Zhenyu Yang, Yi Huang, Junlan Feng, Xiaoyan Zhu, and Minlie Huang. 2022. Curriculum-Based Self-Training Makes Better Few-Shot Learners for Data-to-Text Generation. arXiv:2206.02712 (2022)."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213929"},{"key":"e_1_2_1_7_1","volume-title":"Neural text generation from structured data with application to the biography domain. arXiv:1603.07771","author":"Lebret R\u00e9mi","year":"2016","unstructured":"R\u00e9mi Lebret, David Grangier, and Michael Auli. 2016. Neural text generation from structured data with application to the biography domain. arXiv:1603.07771 (2016)."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11925"},{"key":"e_1_2_1_9_1","unstructured":"Qin Lyu Kaushik Chakrabarti Shobhit Hathi Souvik Kundu Jianwen Zhang and Zheng Chen. 2020. Hybrid Ranking Network for Text-to-SQL. arXiv:2008.04759 [cs.CL]"},{"key":"e_1_2_1_10_1","volume-title":"Improving Compositional Generalization with Self-Training for Data-to-Text Generation. arXiv:2110.08467","author":"Mehta Sanket Vaibhav","year":"2021","unstructured":"Sanket Vaibhav Mehta, Jinfeng Rao, Yi Tay, Mihir Kale, Ankur Parikh, Hongtao Zhong, and Emma Strubell. 2021. Improving Compositional Generalization with Self-Training for Data-to-Text Generation. arXiv:2110.08467 (2021)."},{"key":"e_1_2_1_11_1","volume-title":"ToTTo: A controlled table-to-text generation dataset. arXiv:2004.14373","author":"Parikh Ankur P","year":"2020","unstructured":"Ankur P Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das. 2020. ToTTo: A controlled table-to-text generation dataset. arXiv:2004.14373 (2020)."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33016908"},{"key":"e_1_2_1_13_1","volume-title":"PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from Language Models. arXiv:2109.05093 [cs.CL]","author":"Scholak Torsten","year":"2021","unstructured":"Torsten Scholak, Nathan Schucher, and Dzmitry Bahdanau. 2021. PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from Language Models. arXiv:2109.05093 [cs.CL]"},{"key":"e_1_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Bailin Wang Richard Shin Xiaodong Liu Oleksandr Polozov and Matthew Richardson. 2020. RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers. arXiv:1911.04942 [cs.CL]","DOI":"10.18653\/v1\/2020.acl-main.677"},{"key":"e_1_2_1_15_1","volume-title":"Challenges in data-to-document generation. arXiv:1707.08052","author":"Wiseman Sam","year":"2017","unstructured":"Sam Wiseman, Stuart M Shieber, and Alexander M Rush. 2017. Challenges in data-to-document generation. arXiv:1707.08052 (2017)."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1112"},{"key":"e_1_2_1_17_1","unstructured":"Xiaojun Xu Chang Liu and Dawn Song. 2017. SQLNet: Generating Structured Queries From Natural Language Without Reinforcement Learning. arXiv:1711.04436 [cs.CL]"},{"key":"e_1_2_1_18_1","volume-title":"Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task. arXiv:1809.08887 [cs.CL]","author":"Yu Tao","year":"2019","unstructured":"Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2019. Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task. arXiv:1809.08887 [cs.CL]"},{"key":"e_1_2_1_19_1","unstructured":"Victor Zhong Caiming Xiong and Richard Socher. 2017. Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning. arXiv:1709.00103 [cs.CL]"},{"key":"e_1_2_1_20_1","volume-title":"Modeling graph structure in transformer for better AMR-to-text generation. arXiv:1909.00136","author":"Zhu Jie","year":"2019","unstructured":"Jie Zhu, Junhui Li, Muhua Zhu, Longhua Qian, Min Zhang, and Guodong Zhou. 2019. Modeling graph structure in transformer for better AMR-to-text generation. arXiv:1909.00136 (2019)."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3611540.3611575","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,10]],"date-time":"2025-09-10T22:35:35Z","timestamp":1757543735000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3611540.3611575"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8]]},"references-count":20,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2023,8]]}},"alternative-id":["10.14778\/3611540.3611575"],"URL":"https:\/\/doi.org\/10.14778\/3611540.3611575","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2023,8]]},"assertion":[{"value":"2023-08-01","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}