{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T16:18:11Z","timestamp":1781367491954,"version":"3.54.1"},"reference-count":66,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/501100006374","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62422209,62132014,62272304"],"award-info":[{"award-number":["62422209,62132014,62272304"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100006374","name":"National Science Foundation","doi-asserted-by":"publisher","award":["2220407"],"award-info":[{"award-number":["2220407"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2025,6,17]]},"abstract":"<jats:p>State-of-the-art Text-to-SQL models rely on fine-tuning or few-shot prompting to help LLMs learn from training datasets containing mappings from natural language (NL) queries to SQL statements. Consequently, the quality of the dataset can greatly affect the accuracy of these Text-to-SQL models. Unlike other NL tasks, Text-to-SQL datasets are prone to errors despite extensive manual efforts due to the subtle semantics of SQL. Our study has found a non-negligible (&gt;30%) portion of incorrect NL to SQL mapping cases exists in popular datasets Spider and BIRD.<\/jats:p>\n                  <jats:p>This paper aims to improve the quality of Text-to-SQL training datasets and thereby increase the accuracy of the resulting models. To do so, we propose a necessary correctness condition called execution consistency. For a given database instance, an NL to SQL mapping satisfies execution consistency if the execution result of an NL query matches that of the corresponding SQL. We develop SQLDriller to detect incorrect NL to SQL mappings based on execution consistency in a best-effort manner by crafting database instances that likely result in violations of execution consistency. It generates multiple candidate SQL predictions that differ in their syntax structures. Using a SQL equivalence checker, SQLDriller obtains counterexample database instances that can distinguish non-equivalent candidate SQLs. It then checks the execution consistency of an NL to SQL mapping under this set of counterexamples. The evaluation shows SQLDriller effectively detects and fixes incorrect mappings in the Text-to-SQL dataset, and it improves the model accuracy by up to 13.6%.<\/jats:p>","DOI":"10.1145\/3725271","type":"journal-article","created":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T21:23:29Z","timestamp":1750281809000},"page":"1-28","source":"Crossref","is-referenced-by-count":7,"title":["Automated Validating and Fixing of Text-to-SQL Translation with Execution Consistency"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-5303-2599","authenticated-orcid":false,"given":"Yicun","family":"Yang","sequence":"first","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0220-5726","authenticated-orcid":false,"given":"Zhaoguo","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-7486-1391","authenticated-orcid":false,"given":"Yu","family":"Xia","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-3603-2165","authenticated-orcid":false,"given":"Zhuoran","family":"Wei","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-8138-8639","authenticated-orcid":false,"given":"Haoran","family":"Ding","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3267-0776","authenticated-orcid":false,"given":"Ruzica","family":"Piskac","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Yale University, New Haven, Connecticut, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9720-0361","authenticated-orcid":false,"given":"Haibo","family":"Chen","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9574-1746","authenticated-orcid":false,"given":"Jinyang","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Computer Science, New York University, New York, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,18]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"Alibaba DAMO Academy. 2024. BIRD Website Homepage. https:\/\/bird-bench.github.io\/."},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1017\/S135132490000005X"},{"key":"e_1_2_2_3_1","volume-title":"The Eleventh International Conference on Learning Representations, ICLR 2023","author":"Arora Simran","year":"2023","unstructured":"Simran Arora, Avanika Narayan, Mayee F. Chen, Laurel J. Orr, Neel Guha, Kush Bhatia, Ines Chami, and Christopher R\u00e9. 2023. Ask Me Anything: A simple strategy for prompting language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1--5, 2023. OpenReview.net. https:\/\/openreview.net\/forum?id=bhUPJnS2g0X"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-99524-9_24"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.436"},{"key":"e_1_2_2_6_1","volume-title":"Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al.","author":"Chen Mark","year":"2021","unstructured":"Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. CoRR, Vol. abs\/2107.03374 (2021). showeprint[arXiv]2107.03374 https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2409.02038"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.14778\/3236187.3236200"},{"key":"e_1_2_2_9_1","volume-title":"Cosette: An Automated Prover for SQL. In 8th Biennial Conference on Innovative Data Systems Research, CIDR","author":"Chu Shumo","year":"2017","unstructured":"Shumo Chu, Chenglong Wang, Konstantin Weitz, and Alvin Cheung. 2017. Cosette: An Automated Prover for SQL. In 8th Biennial Conference on Innovative Data Systems Research, CIDR 2017, Chaminade, CA, USA, January 8--11, 2017, Online Proceedings. www.cidrdb.org. http:\/\/cidrdb.org\/cidr2017\/papers\/p51-chu-cidr17.pdf"},{"key":"e_1_2_2_10_1","first-page":"179","volume-title":"Proceeding of the IFIP Working Conference Data Base Management, Carg\u00e8se","author":"Codd E. F.","year":"1974","unstructured":"E. F. Codd. 1974. Seven Steps to Rendezvous with the Casual User. In Data Base Management, Proceeding of the IFIP Working Conference Data Base Management, Carg\u00e8se, Corsica, France, April 1--5, 1974. North-Holland, 179-200."},{"key":"e_1_2_2_11_1","unstructured":"The SQLite Consortium. 2024. SQLite Release 3.45.1. https:\/\/www.sqlite.org\/releaselog\/3_45_1.html."},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/978--3--540--78800--3_24"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/n19--1423"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626768"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.64"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2307.07306"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00410"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.187"},{"key":"e_1_2_2_19_1","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018","volume":"360","author":"Finegan-Dollak Catherine","year":"2018","unstructured":"Catherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang, and Dragomir R. Radev. 2018. Improving Text-to-SQL Evaluation Methodology. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15--20, 2018, Volume 1: Long Papers. Association for Computational Linguistics, 351-360. https:\/\/aclanthology.org\/P18--1033\/"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.14778\/3641204.3641221"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.3390\/ai2040043"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3649849"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/320251.320253"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17--1089"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735461.2735468"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i11.26535"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654930"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","unstructured":"Jinyang Li Binyuan Hui Reynold Cheng Bowen Qin Chenhao Ma Nan Huo Fei Huang Wenyu Du Luo Si and Yongbin Li. 2023a. Graphix-T5: mixing pre-trained transformers with graph-aware layers for text-to-SQL parsing. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence and Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence and Thirteenth Symposium on Educational Advances in Artificial Intelligence (AAAI'23\/IAAI'23\/EAAI'23). AAAI Press Article 1467 9 pages. https:\/\/doi.org\/10.1609\/aaai.v37i11.26536","DOI":"10.1609\/aaai.v37i11.26536"},{"key":"e_1_2_2_30_1","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Li Jinyang","year":"2023","unstructured":"Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Ruiying Geng, Nan Huo, Xuanhe Zhou, Chenhao Ma, Guoliang Li, Kevin C.C. Chang, Fei Huang, Reynold Cheng, and Yongbin Li. 2023b. Can LLM already serve as a database interface? a big bench for large-scale database grounded text-to-SQLs. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS '23). Curran Associates Inc., Red Hook, NY, USA, Article 1835, 28 pages."},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654979"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1162\/dint_a_00194"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2408.05109"},{"key":"e_1_2_2_34_1","volume-title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR, Vol. abs\/1907.11692 (2019). showeprint[arXiv]1907.11692 http:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11704-024--40763--6"},{"key":"e_1_2_2_36_1","unstructured":"LILY Lab of Yale University. 2024. Spider Website Homepage. https:\/\/yale-lily.github.io\/spider."},{"key":"e_1_2_2_37_1","unstructured":"OpenAI. 2023. GPT-4 Technical Report. CoRR Vol. abs\/2303.08774 (2023). showeprint[arXiv]2303.08774 https:\/\/doi.org\/10.48550\/arXiv.2303.08774"},{"key":"e_1_2_2_38_1","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Papicchio Simone","year":"2023","unstructured":"Simone Papicchio, Paolo Papotti, and Luca Cagliero. 2023. QATCH: benchmarking SQL-centric tasks with table representation learning models on your data. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS '23). Curran Associates Inc., Red Hook, NY, USA, Article 1348, 20 pages."},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/EDUCON52537.2022.9766617"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/604045.604120"},{"key":"e_1_2_2_41_1","first-page":"326","volume-title":"Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023","author":"Wightman Gwenyth Portillo","year":"2023","unstructured":"Gwenyth Portillo Wightman, Alexandra Delucia, and Mark Dredze. 2023. Strength in Numbers: Estimating Confidence of Large Language Models by Prompt Agreement. In Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). Association for Computational Linguistics, Toronto, Canada, 326-362. https:\/\/aclanthology.org\/2023.trustnlp-1.28"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.52202\/075280-1577"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415557"},{"key":"e_1_2_2_44_1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., Vol. 21, 1, Article 140 (Jan. 2020), 67 pages.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-short.15"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2406.19073"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.779"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2405.16755"},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2302.13971"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.677"},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3452822"},{"key":"e_1_2_2_53_1","volume-title":"Self-Consistency Improves Chain of Thought Reasoning in Language Models. In The Eleventh International Conference on Learning Representations, ICLR 2023","author":"Wang Xuezhi","year":"2023","unstructured":"Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1--5, 2023. OpenReview.net. https:\/\/openreview.net\/forum?id=1PL1NIMMrw"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3526125"},{"key":"e_1_2_2_55_1","unstructured":"Wikipedia. 2024. Jaccard Similarity. Retrieved from https:\/\/en.wikipedia.org\/wiki\/Jaccard_index."},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133887"},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1425"},{"key":"e_1_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.14778\/3659437.3659452"},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2312.10321"},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.29"},{"key":"e_1_2_2_61_1","volume-title":"Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning. CoRR","author":"Zhong Victor","year":"2017","unstructured":"Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning. CoRR, Vol. abs\/1709.00103 (2017). showeprint[arXiv]1709.00103 http:\/\/arxiv.org\/abs\/1709.00103"},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-emnlp.192"},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342267"},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE53745.2022.00250"},{"key":"e_1_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1007\/s41019-023-00235--6"},{"key":"e_1_2_2_66_1","doi-asserted-by":"publisher","DOI":"10.14778\/3685800.3685816"}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3725271","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T18:53:49Z","timestamp":1774983229000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3725271"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,17]]},"references-count":66,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,6,17]]}},"alternative-id":["10.1145\/3725271"],"URL":"https:\/\/doi.org\/10.1145\/3725271","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,17]]}}}