{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,12]],"date-time":"2026-03-12T10:52:13Z","timestamp":1773312733138,"version":"3.50.1"},"reference-count":36,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>\n                    As software systems increasingly rely on natural language interfaces, ensuring the reliability of these systems is crucial. One critical component is the ability to accurately translate natural language queries into corresponding SQL queries, a field known as Text-to-SQL. However, the scarcity of high-quality, large-scale, and domain-specific Text-to-SQL datasets hinders the development of reliable and robust models. To tackle these challenges, we propose\n                    <jats:sc>SelectCraft<\/jats:sc>\n                    , a novel automatic generation approach designed to create realistic Text-to-SQL datasets tailored to specific domains. Our method leverages existing databases and their structures to generate complex text-SQL pairs that mirror real-world usage scenarios. As a proof of concept, we have successfully generated a substantial financial Text-to-SQL dataset, denominated as\n                    <jats:sc>BanQies<\/jats:sc>\n                    , encompassing over 1 million samples utilizing our proposed approach. Moreover, we introduce\n                    <jats:sc>BanQL<\/jats:sc>\n                    , a new large language model (LLM) based on\n                    <jats:italic toggle=\"yes\">StarCoder2<\/jats:italic>\n                    , a state-of-the-art code-based LLM, and fine-tuned on our newly created dataset. We evaluate\n                    <jats:sc>BanQL<\/jats:sc>\n                    performance against several state-of-the-art models, demonstrating significant enhancements in accuracy and generalizability, highlighting the advantages of incorporating domain-specific data in Text-to-SQL tasks. We firmly believe that our contributions have the potential to improve the overall reliability of Text-to-SQL software systems.\n                  <\/jats:p>\n                  <jats:p>CCS Concepts: \u2022 Software and its engineering; \u2022 Computing methodologies \u2192 Natural languageprocessing; \u2022 Information systems \u2192 Structured Query Language;<\/jats:p>","DOI":"10.1145\/3746226","type":"journal-article","created":{"date-parts":[[2025,6,26]],"date-time":"2025-06-26T11:35:08Z","timestamp":1750937708000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Towards Automating Domain-Specific Data Generation for Text-to-SQL: A Comprehensive Approach"],"prefix":"10.1145","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-9387-6847","authenticated-orcid":false,"given":"Salmane","family":"Chafik","sequence":"first","affiliation":[{"name":"College of Computing, Mohammed VI Polytechnic University, Ben Guerir, Morocco"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7657-4738","authenticated-orcid":false,"given":"Saad","family":"Ezzini","sequence":"additional","affiliation":[{"name":"Information and Computer Science Department, King Fahd University of Petroleum &amp; Minerals, Dhahran, Saudi Arabia and Interdisciplinary Research Center for Intelligent Manufacturing and Robotics, King Fahd University of Petroleum &amp; Minerals, Dhahran, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4225-911X","authenticated-orcid":false,"given":"Ismail","family":"Berrada","sequence":"additional","affiliation":[{"name":"College of Computing, Mohammed VI Polytechnic University, Ben Guerir, Morocco"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,3,11]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et al. 2023. Gpt-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","unstructured":"Benoit Courty Victor Schmidt Sasha Luccioni Goyal-Kamal Marion Coutarel Boris Feld J\u00e9r\u00e9my Lecourt Liam Connell Amine Saboni Inimaz et al. 2024. mlco2\/codecarbon: v2.4.1. DOI: 10.5281\/zenodo.11171501","DOI":"10.5281\/zenodo.11171501"},{"key":"e_1_3_2_4_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4171\u20134186."},{"key":"e_1_3_2_5_2","unstructured":"Yujian Gan Xinyun Chen and Matthew Purver. 2021. Exploring underexplored limitations of cross-domain text-to-sql generalization. arXiv:2109.05157. Retrieved from https:\/\/arxiv.org\/abs\/2109.05157"},{"issue":"1","key":"e_1_3_2_6_2","first-page":"9","article-title":"A review of ChatGPT AI\u2019s impact on several business sectors","volume":"1","author":"George A. Shaji","year":"2023","unstructured":"A. Shaji George and A. S. Hovan George. 2023. A review of ChatGPT AI\u2019s impact on several business sectors. Partners Universal International Innovation Journal 1, 1 (2023), 9\u201323.","journal-title":"Partners Universal International Innovation Journal"},{"key":"e_1_3_2_7_2","unstructured":"Moshe Hazoom Vibhor Malik and Ben Bogin. 2021. Text-to-SQL in the wild: A naturally-occurring dataset based on stack exchange data. arXiv:2106.05006. Retrieved from https:\/\/arxiv.org\/abs\/2106.05006"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.3115\/116580.116613"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-022-00776-8"},{"key":"e_1_3_2_10_2","unstructured":"Denis Kocetkov Raymond Li Loubna Ben Allal Jia Li Chenghao Mou Carlos Mu\u00f1oz Ferrandis Yacine Jernite Margaret Mitchell Sean Hughes Thomas Wolf et al. 2022. The stack: 3 TB of permissively licensed source code. arXiv:2211.15533. Retrieved from https:\/\/arxiv.org\/abs\/2211.15533"},{"key":"e_1_3_2_11_2","unstructured":"Numbers Station Labs. 2023. NSText2SQL: An Open Source Text-to-SQL Dataset for Foundation Model Training. Retrieved from https:\/\/github.com\/NumbersStationAI\/NSQL"},{"key":"e_1_3_2_12_2","article-title":"Can LLM already serve as a database interface? A big bench for large-scale database grounded text-to-sqls","volume":"36","author":"Li Jinyang","year":"2024","unstructured":"Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Ruiying Geng, Nan Huo, et al. 2024. Can LLM already serve as a database interface? A big bench for large-scale database grounded text-to-sqls. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems, Vol"},{"key":"e_1_3_2_13_2","unstructured":"Raymond Li Loubna Ben Allal Yangtian Zi Niklas Muennighoff Denis Kocetkov Chenghao Mou Marc Marone Christopher Akiki Jia Li Jenny Chim et al. 2023. StarCoder: May the source be with you!. arXiv:2305.06161. Retrieved from https:\/\/arxiv.org\/abs\/2305.06161"},{"key":"e_1_3_2_14_2","first-page":"74","volume-title":"Text Summarization Branches Out","year":"2004","unstructured":"Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out. Association for Computational Linguistics, 74\u201381. Retrieved from https:\/\/www.aclweb.org\/anthology\/W04-1013"},{"key":"e_1_3_2_15_2","first-page":"1950","article-title":"Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning","volume":"35","author":"Liu Haokun","year":"2022","unstructured":"Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A. Raffel. 2022. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 35, 1950\u20131965.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems, Vol"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3640310.3674091"},{"key":"e_1_3_2_17_2","unstructured":"Anton Lozhkov Raymond Li Loubna Ben Allal Federico Cassano Joel Lamy-Poirier Nouamane Tazi Ao Tang Dmytro Pykhtar Jiawei Liu Yuxiang Wei et al. 2024. Starcoder 2 and the stack v2: The next generation. arXiv:2402.19173. Retrieved from https:\/\/arxiv.org\/abs\/2402.19173"},{"issue":"253","key":"e_1_3_2_18_2","first-page":"1","article-title":"Estimating the carbon footprint of bloom, a 176b parameter language model","volume":"24","author":"Luccioni Alexandra Sasha","year":"2023","unstructured":"Alexandra Sasha Luccioni, Sylvain Viguier, and Anne-Laure Ligozat. 2023. Estimating the carbon footprint of bloom, a 176b parameter language model. Journal of Machine Learning Research 24, 253 (2023), 1\u201315.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3658542"},{"key":"e_1_3_2_20_2","unstructured":"Erik Nijkamp Hiroaki Hayashi Caiming Xiong Silvio Savarese and Yingbo Zhou. 2023. Codegen2: Lessons for training LLMS on programming and natural languages. arXiv:2305.02309. Retrieved from https:\/\/arxiv.org\/abs\/2305.02309"},{"key":"e_1_3_2_21_2","unstructured":"Erik Nijkamp Bo Pang Hiroaki Hayashi Lifu Tu Huan Wang Yingbo Zhou Silvio Savarese and Caiming Xiong. 2022. Codegen: An open large language model for code with multi-turn program synthesis. arXiv:2203.13474. Retrieved from https:\/\/arxiv.org\/abs\/2203.13474"},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","unstructured":"Kishore Papineni Salim Roukos Todd Ward and Wei Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics 311\u2013318. Retrieved from https:\/\/aclanthology.org\/P02-1040.pdf","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_2_23_2","unstructured":"David Patterson Joseph Gonzalez Quoc Le Chen Liang Lluis-Miquel Munguia Daniel Rothchild David So Maud Texier and Jeff Dean. 2021. Carbon emissions and large neural network training. arXiv:2104.10350. Retrieved from https:\/\/arxiv.org\/abs\/2104.10350"},{"issue":"140","key":"e_1_3_2_24_2","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21, 140 (2020), 1\u201367. Retrieved from http:\/\/jmlr.org\/papers\/v21\/20-074.html","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_25_2","unstructured":"Baptiste Roziere Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing Ellen Tan Yossi Adi Jingyu Liu Tal Remez J\u00e9r\u00e9my Rapin et al. 2023. Code llama: Open foundation models for code. arXiv:2308.12950. Retrieved from https:\/\/arxiv.org\/abs\/2308.12950"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.779"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3639477.3639732"},{"key":"e_1_3_2_28_2","unstructured":"Alane Laughlin Suhr Kenton Lee Ming-Wei Chang and Pete Shaw. 2020. Exploring unexplored generalization challenges for cross-database semantic parsing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Retrieved from https:\/\/aclanthology.org\/2020.acl-main.742\/"},{"key":"e_1_3_2_29_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems, Vol"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3366423.3380120"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","unstructured":"Yue Wang Weishi Wang Shafiq Joty and Steven C. H. Hoi. 2021. Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. arXiv:2109.00859. Retrieved from https:\/\/arxiv.org\/abs\/2109.00859","DOI":"10.18653\/v1\/2021.emnlp-main.685"},{"key":"e_1_3_2_32_2","unstructured":"Xiaoyu Yin Dagmar Gromann and Sebastian Rudolph. 2019. Neural machine translating from natural language to SPARQL. arXiv:1906.09302. Retrieved from https:\/\/arxiv.org\/abs\/1906.09302"},{"key":"e_1_3_2_33_2","article-title":"Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task","author":"Yu Tao","year":"2018","unstructured":"Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al. 2018. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP).","journal-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP)"},{"key":"e_1_3_2_34_2","first-page":"1050","volume-title":"Proceedings of the National Conference on Artificial Intelligence","author":"Zelle John M.","year":"1996","unstructured":"John M. Zelle and Raymond J. Mooney. 1996. Learning to parse database queries using inductive logic programming. In Proceedings of the National Conference on Artificial Intelligence, 1050\u20131055."},{"key":"e_1_3_2_35_2","unstructured":"Jingqing Zhang Yao Zhao Mohammad Saleh and Peter J. Liu. 2019. PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization. arXiv:1912.08777. Retrieved from https:\/\/arxiv.org\/abs\/1912.08777"},{"key":"e_1_3_2_36_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Zhang Tianyi","year":"2020","unstructured":"Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating text generation with BERT. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=SkeHuCVFDr"},{"key":"e_1_3_2_37_2","unstructured":"Victor Zhong Caiming Xiong and Richard Socher. 2017. Seq2SQL: Generating structured queries from natural language using reinforcement learning. arXiv:1709.00103. Retrieved from https:\/\/arxiv.org\/abs\/1709.00103"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746226","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T16:27:39Z","timestamp":1773246459000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746226"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,11]]},"references-count":36,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3746226"],"URL":"https:\/\/doi.org\/10.1145\/3746226","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,11]]},"assertion":[{"value":"2024-11-06","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}