{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T19:15:25Z","timestamp":1776885325170,"version":"3.51.2"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,11]]},"abstract":"<jats:p>\n            Natural Language Interface to Database (NLIDB) translates human utterances into SQL queries and enables database interactions for non-expert users. Recently, neural network models have become a major approach to implementing NLIDB. However, neural NLIDB faces challenges due to variations in natural language and database schema design. For instance, one user intent or database conceptual model can be expressed in\n            <jats:italic>various forms.<\/jats:italic>\n            However, existing benchmarks, using hold-out datasets, cannot provide thorough understanding of how good neural NLIDBs really are in real-world situations and its robustness against such variations. A key difficulty is to annotate SQL queries for inputs under real-world variations, requiring considerable manual effort and expert knowledge.\n          <\/jats:p>\n          <jats:p>To systematically assess the robustness of neural NLIDBs without extensive manual effort, we propose MT-Teql, a unified framework to benchmark NLIDBs against real-world language and schema variations. Inspired by recent advances in DBMS metamorphic testing, MT-Teql implements semantics-preserving transformations on utterances and database schemas to generate their variants. NLIDBs can thus be examined for robustness utilizing utterances\/schemas and their variants without requiring manual intervention.<\/jats:p>\n          <jats:p>We benchmarked nine neural NLIDBs using 62,430 inputs and identified 15,433 defects. We analyzed potential root causes of defects and conducted a user study to show how MT-Teql can assist developers to systematically assess NLIDBs. We further show that the transformed (error-triggering) inputs can be used to augment popular NLIDBs and eliminate 46.5%(\u00b15.0%) errors made by them without compromising their accuracy on standard benchmarks. We summarize lessons from this study that can provide insights to select and design NLIDBs that fit particular usage scenarios.<\/jats:p>","DOI":"10.14778\/3494124.3494139","type":"journal-article","created":{"date-parts":[[2022,2,5]],"date-time":"2022-02-05T00:31:46Z","timestamp":1644021106000},"page":"569-582","source":"Crossref","is-referenced-by-count":27,"title":["MT-teql"],"prefix":"10.14778","volume":"15","author":[{"given":"Pingchuan","family":"Ma","sequence":"first","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong SAR, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuai","family":"Wang","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong SAR, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,2,4]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2019.00041"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1448"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1378"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE51399.2021.00220"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.260"},{"key":"e_1_2_1_7_1","volume-title":"EungGyun Kim, and Dong Ryeol Shin.","author":"Choi DongHyun","year":"2020","unstructured":"DongHyun Choi , Myeong Cheol Shin , EungGyun Kim, and Dong Ryeol Shin. 2020 . RYANSQL : Recursively Applying Sketch-based Slot Fillings for Complex Text-to-SQL in Cross-Domain Databases . arXiv preprint arXiv:2004.03125 (2020). DongHyun Choi, Myeong Cheol Shin, EungGyun Kim, and Dong Ryeol Shin. 2020. RYANSQL: Recursively Applying Sketch-based Slot Fillings for Complex Text-to-SQL in Cross-Domain Databases. arXiv preprint arXiv:2004.03125 (2020)."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.3115\/1075812.1075823"},{"key":"e_1_2_1_9_1","volume-title":"Christopher Meek, Oleksandr Polozov, Huan Sun, and Matthew Richardson.","author":"Deng Xiang","year":"2020","unstructured":"Xiang Deng , Ahmed Hassan Awadallah , Christopher Meek, Oleksandr Polozov, Huan Sun, and Matthew Richardson. 2020 . Structure-Grounded Pretraining for Text-to-SQL. arXiv preprint arXiv:2010.12773 (2020). Xiang Deng, Ahmed Hassan Awadallah, Christopher Meek, Oleksandr Polozov, Huan Sun, and Matthew Richardson. 2020. Structure-Grounded Pretraining for Text-to-SQL. arXiv preprint arXiv:2010.12773 (2020)."},{"key":"e_1_2_1_10_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT (1).","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT (1). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT (1)."},{"key":"e_1_2_1_11_1","volume-title":"Improving text-to-sql evaluation methodology. arXiv preprint arXiv:1806.09029","author":"Finegan-Dollak Catherine","year":"2018","unstructured":"Catherine Finegan-Dollak , Jonathan K Kummerfeld , Li Zhang , Karthik Ramanathan , Sesh Sadasivam , Rui Zhang , and Dragomir Radev . 2018. Improving text-to-sql evaluation methodology. arXiv preprint arXiv:1806.09029 ( 2018 ). Catherine Finegan-Dollak, Jonathan K Kummerfeld, Li Zhang, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang, and Dragomir Radev. 2018. Improving text-to-sql evaluation methodology. arXiv preprint arXiv:1806.09029 (2018)."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.117"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1097"},{"key":"e_1_2_1_14_1","volume-title":"Towards complex text-to-sql in cross-domain database with intermediate representation. arXiv preprint arXiv:1905.08205","author":"Guo Jiaqi","year":"2019","unstructured":"Jiaqi Guo , Zecheng Zhan , Yan Gao , Yan Xiao , Jian-Guang Lou , Ting Liu , and Dongmei Zhang . 2019. Towards complex text-to-sql in cross-domain database with intermediate representation. arXiv preprint arXiv:1905.08205 ( 2019 ). Jiaqi Guo, Zecheng Zhan, Yan Gao, Yan Xiao, Jian-Guang Lou, Ting Liu, and Dongmei Zhang. 2019. Towards complex text-to-sql in cross-domain database with intermediate representation. arXiv preprint arXiv:1905.08205 (2019)."},{"key":"e_1_2_1_15_1","volume-title":"Learning a neural semantic parser from user feedback. arXiv preprint arXiv:1704.08760","author":"Iyer Srinivasan","year":"2017","unstructured":"Srinivasan Iyer , Ioannis Konstas , Alvin Cheung , Jayant Krishnamurthy , and Luke Zettlemoyer . 2017. Learning a neural semantic parser from user feedback. arXiv preprint arXiv:1704.08760 ( 2017 ). Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, Jayant Krishnamurthy, and Luke Zettlemoyer. 2017. Learning a neural semantic parser from user feedback. arXiv preprint arXiv:1704.08760 (2017)."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/3401960.3401970"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735461.2735468"},{"key":"e_1_2_1_18_1","unstructured":"Pingchuan Ma and Shuai Wang. 2021. MT-Teql: Evaluating and Augmenting Neural NLIDB on Real-world Linguistic and Schema Variations. Supplementary Material. https:\/\/bit.ly\/MT-Teql-sm.  Pingchuan Ma and Shuai Wang. 2021. MT-Teql: Evaluating and Augmenting Neural NLIDB on Real-world Linguistic and Schema Variations. Supplementary Material. https:\/\/bit.ly\/MT-Teql-sm."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/3491440.3491504"},{"key":"e_1_2_1_20_1","volume-title":"Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI","author":"Madaan Nishtha","year":"2021","unstructured":"Nishtha Madaan , Inkit Padhi , Naveen Panwar , and Diptikalyan Saha . 2021. Generate Your Counterfactuals: Towards Controlled Counterfactual Generation for Text . In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021 , Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2--9, 2021. AAAI Press , 13516--13524. Nishtha Madaan, Inkit Padhi, Naveen Panwar, and Diptikalyan Saha. 2021. Generate Your Counterfactuals: Towards Controlled Counterfactual Generation for Text. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2--9, 2021. AAAI Press, 13516--13524."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.14778\/2794367.2794377"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.3115\/1220355.1220376"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.3115\/116580.116612"},{"key":"e_1_2_1_25_1","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei Ilya Sutskever etal 2019. Language models are unsupervised multitask learners. OpenAI blog 1 8 (2019) 9.  Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei Ilya Sutskever et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1 8 (2019) 9."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1621"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.442"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409710"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3428279"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/3488766.3488804"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994536"},{"key":"e_1_2_1_32_1","volume-title":"DuoRAT: Towards Simpler Text-to-SQL Models. arXiv preprint arXiv:2010.11119","author":"Scholak Torsten","year":"2020","unstructured":"Torsten Scholak , Raymond Li , Dzmitry Bahdanau , Harm de Vries , and Chris Pal . 2020. DuoRAT: Towards Simpler Text-to-SQL Models. arXiv preprint arXiv:2010.11119 ( 2020 ). Torsten Scholak, Raymond Li, Dzmitry Bahdanau, Harm de Vries, and Chris Pal. 2020. DuoRAT: Towards Simpler Text-to-SQL Models. arXiv preprint arXiv:2010.11119 (2020)."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2016.2532875"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407858"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1010"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2931464"},{"key":"e_1_2_1_37_1","volume-title":"Astraea: Grammar-based Fairness Testing. arXiv preprint arXiv:2010.02542","author":"Soremekun Ezekiel","year":"2020","unstructured":"Ezekiel Soremekun , Sakshi Udeshi , and Sudipta Chattopadhyay . 2020 . Astraea: Grammar-based Fairness Testing. arXiv preprint arXiv:2010.02542 (2020). Ezekiel Soremekun, Sakshi Udeshi, and Sudipta Chattopadhyay. 2020. Astraea: Grammar-based Fairness Testing. arXiv preprint arXiv:2010.02542 (2020)."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.5555\/3298023.3298212"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.742"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3180155.3180220"},{"key":"e_1_2_1_41_1","volume-title":"Meta-Learning for Domain Generalization in Semantic Parsing. arXiv preprint arXiv:2010.11988","author":"Wang Bailin","year":"2020","unstructured":"Bailin Wang , Mirella Lapata , and Ivan Titov . 2020. Meta-Learning for Domain Generalization in Semantic Parsing. arXiv preprint arXiv:2010.11988 ( 2020 ). Bailin Wang, Mirella Lapata, and Ivan Titov. 2020. Meta-Learning for Domain Generalization in Semantic Parsing. arXiv preprint arXiv:2010.11988 (2020)."},{"key":"e_1_2_1_42_1","volume-title":"Rat-sql: Relation-aware schema encoding and linking for text-to-sql parsers. arXiv preprint arXiv:1911.04942","author":"Wang Bailin","year":"2019","unstructured":"Bailin Wang , Richard Shin , Xiaodong Liu , Oleksandr Polozov , and Matthew Richardson . 2019 . Rat-sql: Relation-aware schema encoding and linking for text-to-sql parsers. arXiv preprint arXiv:1911.04942 (2019). Bailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov, and Matthew Richardson. 2019. Rat-sql: Relation-aware schema encoding and linking for text-to-sql parsers. arXiv preprint arXiv:1911.04942 (2019)."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.5555\/972942.972944"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3380589"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.523"},{"key":"e_1_2_1_46_1","volume-title":"Sqlnet: Generating structured queries from natural language without reinforcement learning. arXiv preprint arXiv:1711.04436","author":"Xu Xiaojun","year":"2017","unstructured":"Xiaojun Xu , Chang Liu , and Dawn Song . 2017 . Sqlnet: Generating structured queries from natural language without reinforcement learning. arXiv preprint arXiv:1711.04436 (2017). Xiaojun Xu, Chang Liu, and Dawn Song. 2017. Sqlnet: Generating structured queries from natural language without reinforcement learning. arXiv preprint arXiv:1711.04436 (2017)."},{"key":"e_1_2_1_47_1","volume-title":"Type-sql: Knowledge-based type-aware neural text-to-sql generation. arXiv preprint arXiv:1804.09769","author":"Yu Tao","year":"2018","unstructured":"Tao Yu , Zifan Li , Zilin Zhang , Rui Zhang , and Dragomir Radev . 2018 . Type-sql: Knowledge-based type-aware neural text-to-sql generation. arXiv preprint arXiv:1804.09769 (2018). Tao Yu, Zifan Li, Zilin Zhang, Rui Zhang, and Dragomir Radev. 2018. Type-sql: Knowledge-based type-aware neural text-to-sql generation. arXiv preprint arXiv:1804.09769 (2018)."},{"key":"e_1_2_1_48_1","volume-title":"Bailin Wang, Yi Chern Tan, Xinyi Yang, Dragomir Radev, Richard Socher, and Caiming Xiong.","author":"Yu Tao","year":"2020","unstructured":"Tao Yu , Chien-Sheng Wu , Xi Victoria Lin , Bailin Wang, Yi Chern Tan, Xinyi Yang, Dragomir Radev, Richard Socher, and Caiming Xiong. 2020 . GraPPa: Grammar- Augmented Pre-Training for Table Semantic Parsing . arXiv preprint arXiv:2009.13845 (2020). Tao Yu, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang, Yi Chern Tan, Xinyi Yang, Dragomir Radev, Richard Socher, and Caiming Xiong. 2020. GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing. arXiv preprint arXiv:2009.13845 (2020)."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1193"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1425"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.5555\/1864519.1864543"},{"key":"e_1_2_1_52_1","volume-title":"Caiming Xiong, Richard Socher, Michael R Lyu, Irwin King, and Steven CH Hoi.","author":"Zeng Jichuan","year":"2020","unstructured":"Jichuan Zeng , Xi Victoria Lin , Caiming Xiong, Richard Socher, Michael R Lyu, Irwin King, and Steven CH Hoi. 2020 . Photon : A Robust Cross-Domain Text-to-SQL System . arXiv preprint arXiv:2007.15280 (2020). Jichuan Zeng, Xi Victoria Lin, Caiming Xiong, Richard Socher, Michael R Lyu, Irwin King, and Steven CH Hoi. 2020. Photon: A Robust Cross-Domain Text-to-SQL System. arXiv preprint arXiv:2007.15280 (2020)."},{"key":"e_1_2_1_53_1","volume-title":"Semantic Evaluation for Text-to-SQL with Distilled Test Suites. arXiv preprint arXiv:2010.02840","author":"Zhong Ruiqi","year":"2020","unstructured":"Ruiqi Zhong , Tao Yu , and Dan Klein . 2020. Semantic Evaluation for Text-to-SQL with Distilled Test Suites. arXiv preprint arXiv:2010.02840 ( 2020 ). Ruiqi Zhong, Tao Yu, and Dan Klein. 2020. Semantic Evaluation for Text-to-SQL with Distilled Test Suites. arXiv preprint arXiv:2010.02840 (2020)."},{"key":"e_1_2_1_54_1","volume-title":"Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning. CoRR abs\/1709.00103","author":"Zhong Victor","year":"2017","unstructured":"Victor Zhong , Caiming Xiong , and Richard Socher . 2017. Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning. CoRR abs\/1709.00103 ( 2017 ). Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning. CoRR abs\/1709.00103 (2017)."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3494124.3494139","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:33:36Z","timestamp":1672227216000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3494124.3494139"}},"subtitle":["evaluating and augmenting neural NLIDB on real-world linguistic and schema variations"],"short-title":[],"issued":{"date-parts":[[2021,11]]},"references-count":53,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2021,11]]}},"alternative-id":["10.14778\/3494124.3494139"],"URL":"https:\/\/doi.org\/10.14778\/3494124.3494139","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,11]]}}}