{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T23:16:45Z","timestamp":1770938205355,"version":"3.50.1"},"reference-count":43,"publisher":"Association for Computing Machinery (ACM)","issue":"2","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2026,2,28]]},"abstract":"<jats:p>\n                    Recent advances in text-to-SQL models, which translate natural language questions (NLQs) into executable SQL queries, have made interacting with relational databases more accessible, even for those with limited technical ability. This task has seen significant improvement with the release of multiple English datasets and benchmarks such as WikiSQL, SPIDER, and BIRD, each covering different domains and levels of complexity. Non-English high-resource languages, such as Chinese, Russian, and Arabic, have also benefited from these advances, either through the translation of existing datasets or the creation of new ones.\n                    <jats:italic toggle=\"yes\">Dialect2SQL<\/jats:italic>\n                    , a newly released text-to-SQL dataset, is dedicated to the Moroccan dialect (Darija), which is known for its complexity and distinctiveness compared to other Arabic dialects and Modern Standard Arabic. In this article, we conduct a comprehensive study on text-to-SQL for Darija by conducting several experiments mainly on the\n                    <jats:italic toggle=\"yes\">Dialect2SQL<\/jats:italic>\n                    dataset using different approaches and configurations with two code-based large language models, StarCoder2 and Qwen-2.5-Coder. The experiments reveal the performance gap between models fine-tuned on English data and those fine-tuned on Darija. Additionally, the results illustrate the positive impact of incorporating multi-language datasets during training. In particular, the gap decreases from 10.1% to 6.7% in BLEU, and from 12.5% to 5.7% in TSED.\n                  <\/jats:p>","DOI":"10.1145\/3787497","type":"journal-article","created":{"date-parts":[[2026,1,24]],"date-time":"2026-01-24T18:31:24Z","timestamp":1769279484000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["DarijaDB: Unlocking Text-to-SQL for Arabic Dialects"],"prefix":"10.1145","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-9387-6847","authenticated-orcid":false,"given":"Salmane","family":"Chafik","sequence":"first","affiliation":[{"name":"College of Computing, Mohammed VI Polytechnic University","place":["Ben Guerir, Morocco"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7657-4738","authenticated-orcid":false,"given":"Saad","family":"Ezzini","sequence":"additional","affiliation":[{"name":"Information and Computer Science Department, King Fahd University of Petroleum & Minerals","place":["Dhahran, Saudi Arabia"]},{"name":"Interdisciplinary Research Center for Intelligent Manufacturing and Robotics, King Fahd University of Petroleum & Minerals","place":["Dhahran, Saudi Arabia"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4225-911X","authenticated-orcid":false,"given":"Ismail","family":"Berrada","sequence":"additional","affiliation":[{"name":"College of Computing, Mohammed VI Polytechnic University","place":["Ben Guerir, Morocco"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,2,12]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et\u00a0al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3605098.3636065."},{"key":"e_1_3_2_4_2","unstructured":"Jinze Bai Shuai Bai Yunfei Chu Zeyu Cui Kai Dang Xiaodong Deng Yang Fan Wenbin Ge Yu Han Fei Huang et\u00a0al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609. Retrieved from https:\/\/arxiv.org\/abs\/2309.16609"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-emnlp.175"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","unstructured":"Ben Bogin Matt Gardner and Jonathan Berant. 2019. Representing schema structure with graph neural networks for text-to-SQL parsing. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 4560\u20134565.","DOI":"10.18653\/v1\/P19-1448"},{"key":"e_1_3_2_7_2","first-page":"86","volume-title":"Proceedings of the 4th Workshop on Arabic Corpus Linguistics (WACL-4)","author":"Chafik Salmane","year":"2025","unstructured":"Salmane Chafik, Saad Ezzini, and Ismail Berrada. 2025. Dialect2SQL: A novel text-to-SQL dataset for Arabic dialects with a focus on Moroccan Darija. In Proceedings of the 4th Workshop on Arabic Corpus Linguistics (WACL-4). 86\u201392."},{"key":"e_1_3_2_8_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde De Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman et\u00a0al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374. Retrieved from https:\/\/arxiv.org\/abs\/2107.03374"},{"issue":"2","key":"e_1_3_2_9_2","first-page":"309","article-title":"Ryansql: Recursively applying sketch-based slot fillings for complex text-to-sql in cross-domain databases","volume":"47","author":"Choi DongHyun","year":"2021","unstructured":"DongHyun Choi, Myeong Cheol Shin, EungGyun Kim, and Dong Ryeol Shin. 2021. Ryansql: Recursively applying sketch-based slot fillings for complex text-to-sql in cross-domain databases. Computational Linguistics 47, 2 (2021), 309\u2013332.","journal-title":"Computational Linguistics"},{"key":"e_1_3_2_10_2","unstructured":"Marta R. Costa-juss\u00e0 James Cross Onur \u00c7elebi Maha Elbayad Kenneth Heafield Kevin Heffernan Elahe Kalbassi Janice Lam Daniel Licht Jean Maillard et\u00a0al. 2022. No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672. Retrieved from https:\/\/arxiv.org\/abs\/2207.04672"},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Longxu Dou Yan Gao Mingyang Pan Dingzirui Wang Wanxiang Che Dechen Zhan and Jian-Guang Lou. 2023. MultiSpider: Towards benchmarking multilingual text-to-SQL semantic parsing. In Proceedings of the AAAI Conference on Artificial Intelligence Vol. 37. 12745\u201312753.","DOI":"10.1609\/aaai.v37i11.26499"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.14569\/IJACSA.2023.0140347"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-FoSE59343.2023.00008"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-06458-6_13"},{"key":"e_1_3_2_15_2","doi-asserted-by":"crossref","unstructured":"Yujian Gan Xinyun Chen Qiuping Huang Matthew Purver John R. Woodward Jinxia Xie and Pengsheng Huang. 2021. Towards robustness of text-to-SQL models against synonym substitution. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2505\u20132515.","DOI":"10.18653\/v1\/2021.acl-long.195"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","unstructured":"Yujian Gan Xinyun Chen and Matthew Purver. 2021. Exploring underexplored limitations of cross-domain text-to-SQL generalization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 8926\u20138931.","DOI":"10.18653\/v1\/2021.emnlp-main.702"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","unstructured":"Dawei Gao Haibin Wang Yaliang Li Xiuyu Sun Yichen Qian Bolin Ding and Jingren Zhou. 2024. Text-to-sql empowered by large language models: A benchmark evaluation. In Proceedings of the VLDB Endow. 17 5 (Jan. 2024) 1132\u20131145. 10.14778\/3641204.3641221","DOI":"10.14778\/3641204.3641221"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.180"},{"key":"e_1_3_2_19_2","doi-asserted-by":"crossref","unstructured":"Jiaqi Guo Zecheng Zhan Yan Gao Yan Xiao Jian-Guang Lou Ting Liu and Dongmei Zhang. 2019. Towards complex text-to-sql in cross-domain database with intermediate representation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 4524\u20134535.","DOI":"10.18653\/v1\/P19-1444"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","unstructured":"Moshe Hazoom Vibhor Malik and Ben Bogin. 2021. Text-to-SQL in the wild: A naturally-occurring dataset based on stack exchange data. In Proceedings of the 1st Workshop on Natural Language Processing for Programming (NLP4Prog 2021). Association for Computational Linguistics Online 77\u201387. 10.18653\/v1\/2021.nlp4prog-1.9","DOI":"10.18653\/v1\/2021.nlp4prog-1.9"},{"key":"e_1_3_2_21_2","volume-title":"Proceedings of the Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27, 1990","author":"Hemphill Charles T.","year":"1990","unstructured":"Charles T. Hemphill, John J. Godfrey, and George R. Doddington. 1990. The ATIS spoken language systems pilot corpus. In Proceedings of the Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27, 1990."},{"key":"e_1_3_2_22_2","article-title":"Large language models for software engineering: A systematic literature review","author":"Hou Xinyi","year":"2024","unstructured":"Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2024. Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology 33, 8 (2024), 1\u201379.","journal-title":"ACM Transactions on Software Engineering and Methodology"},{"key":"e_1_3_2_23_2","first-page":"2790","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Houlsby Neil","year":"2019","unstructured":"Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the International Conference on Machine Learning. PMLR, 2790\u20132799."},{"key":"e_1_3_2_24_2","unstructured":"Saihao Huang Lijie Wang Zhenghua Li Zeyang Liu Chenhui Dou Fukang Yan Xinyan Xiao Hua Wu and Min Zhang. 2022. SeSQL: Yet another large-scale session-level Chinese text-to-SQL dataset. arXiv:2208.12711. Retrieved from https:\/\/arxiv.org\/abs\/2208.12711"},{"key":"e_1_3_2_25_2","unstructured":"Binyuan Hui Jian Yang Zeyu Cui Jiaxi Yang Dayiheng Liu Lei Zhang Tianyu Liu Jiajun Zhang Bowen Yu Keming Lu et\u00a0al. 2024. Qwen2. 5-coder technical report. arXiv preprint arXiv:2409.12186. Retrieved from https:\/\/arxiv.org\/abs\/2409.12186"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-022-00776-8"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","unstructured":"Chia-Hsuan Lee Oleksandr Polozov and Matthew Richardson. 2021. KaggleDBQA: Realistic evaluation of text-to-SQL parsers. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Association for Computational Linguistics Online 2261\u20132273. 10.18653\/v1\/2021.acl-long.176","DOI":"10.18653\/v1\/2021.acl-long.176"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.14778\/2735461.2735468"},{"key":"e_1_3_2_29_2","unstructured":"Jinyang Li Binyuan Hui Ge Qu Jiaxi Yang Binhua Li Bowen Li Bailin Wang Bowen Qin Rongyu Cao Ruiying Geng et\u00a0al. 2023. Can LLM already serve as a database interface? a big bench for large-scale database grounded text-to-SQLs. Advances in Neural Information Processing Systems 36 (2023) 42330\u201342357."},{"key":"e_1_3_2_30_2","unstructured":"Raymond Li Loubna Ben Allal Yangtian Zi Niklas Muennighoff Denis Kocetkov Chenghao Mou Marc Marone Christopher Akiki Jia Li Jenny Chim et\u00a0al. 2023. Starcoder: May the source be with you!Transactions on Machine Learning Research (2023)."},{"key":"e_1_3_2_31_2","unstructured":"Anton Lozhkov Raymond Li Loubna Ben Allal Federico Cassano Joel Lamy-Poirier Nouamane Tazi Ao Tang Dmytro Pykhtar Jiawei Liu Yuxiang Wei et\u00a0al. 2024. Starcoder 2 and the stack v2: The next generation. arXiv:2402.19173. Retrieved from https:\/\/arxiv.org\/abs\/2402.19173"},{"key":"e_1_3_2_32_2","unstructured":"Erik Nijkamp Bo Pang Hiroaki Hayashi Lifu Tu Huan Wang Yingbo Zhou Silvio Savarese and Caiming Xiong. 2023. CodeGen: An open large language model for code with multi-turn program synthesis. In The Eleventh International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=iaYcJKpY2B_"},{"key":"e_1_3_2_33_2","unstructured":"Aissam Outchakoucht and Hamza Es-Samaali. 2021. Moroccan Dialect -Darija- Open Dataset. arXiv:2103.09687. Retrieved from https:\/\/arxiv.org\/abs\/2103.09687"},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","unstructured":"Kishore Papineni Salim Roukos Todd Ward and Wei jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311\u2013318.","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_2_35_2","unstructured":"Bowen Qin Binyuan Hui Lihan Wang Min Yang Jinyang Li Binhua Li Ruiying Geng Rongyu Cao Jian Sun Luo Si Fei Huang and Yongbin Li. 2022. A survey on text-to-SQL parsing: Concepts methods and future directions. arXiv:2208.13629. Retrieved from https:\/\/arxiv.org\/abs\/2208.13629"},{"key":"e_1_3_2_36_2","unstructured":"Baptiste Roziere Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing Ellen Tan Yossi Adi Jingyu Liu Romain Sauvestre Tal Remez et\u00a0al. 2023. Code llama: Open foundation models for code. arXiv:2308.12950. Retrieved from https:\/\/arxiv.org\/abs\/2308.12950"},{"key":"e_1_3_2_37_2","unstructured":"Guokan Shang Hadi Abdine Yousef Khoubrane Amr Mohamed Yassine Abbahaddou Sofiane Ennadir Imane Momayiz Xuguang Ren Eric Moulines Preslav Nakov et\u00a0al. 2025. Atlas-chat: Adapting large language models for low-resource Moroccan Arabic dialect. In Proceedings of the First Workshop on Language Models for Low-Resource Languages. 9\u201330."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3639477.3639732"},{"key":"e_1_3_2_39_2","first-page":"479","volume-title":"Proceedings of the 22nd Annual Conference of the European Association for Machine Translation","author":"Tiedemann J\u00f6rg","year":"2020","unstructured":"J\u00f6rg Tiedemann and Santhosh Thottingal. 2020. OPUS-MT\u2013building open translation services for the world. In Proceedings of the 22nd Annual Conference of the European Association for Machine Translation. 479\u2013480."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3366423.3380120"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","unstructured":"Yue Wang Weishi Wang Shafiq Joty and Steven C. H. Hoi. 2021. Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics Online and Punta Cana Dominican Republic 8696\u20138708. 10.18653\/v1\/2021.emnlp-main.685","DOI":"10.18653\/v1\/2021.emnlp-main.685"},{"key":"e_1_3_2_42_2","doi-asserted-by":"crossref","unstructured":"Tao Yu Rui Zhang Kai Yang Michihiro Yasunaga Dongxu Wang Zifan Li James Ma Irene Li Qingning Yao Shanelle Roman Zilin Zhang and Dragomir Radev. 2019. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. arXiv:1809.08887. Retrieved from https:\/\/arxiv.org\/abs\/1809.08887","DOI":"10.18653\/v1\/D18-1425"},{"key":"e_1_3_2_43_2","first-page":"1050","volume-title":"Proceedings of the National Conference on Artificial Intelligence","author":"Zelle John M.","year":"1996","unstructured":"John M. Zelle and Raymond J. Mooney. 1996. Learning to parse database queries using inductive logic programming. In Proceedings of the National Conference on Artificial Intelligence. 1050\u20131055."},{"key":"e_1_3_2_44_2","unstructured":"Victor Zhong Caiming Xiong and Richard Socher. 2017. Seq2SQL: Generating structured queries from natural language using reinforcement learning. arXiv:1709.00103. Retrieved from https:\/\/arxiv.org\/abs\/1709.00103"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3787497","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T22:20:38Z","timestamp":1770934838000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3787497"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,12]]},"references-count":43,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,2,28]]}},"alternative-id":["10.1145\/3787497"],"URL":"https:\/\/doi.org\/10.1145\/3787497","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,12]]},"assertion":[{"value":"2025-04-10","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-31","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}