{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,9]],"date-time":"2026-05-09T00:08:08Z","timestamp":1778285288640,"version":"3.51.4"},"reference-count":33,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2025,1,2]],"date-time":"2025-01-02T00:00:00Z","timestamp":1735776000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2019YFB1405000"],"award-info":[{"award-number":["2019YFB1405000"]}]},{"name":"National Key Research and Development Program of China","award":["61873309"],"award-info":[{"award-number":["61873309"]}]},{"name":"National Key Research and Development Program of China","award":["92046024"],"award-info":[{"award-number":["92046024"]}]},{"name":"National Key Research and Development Program of China","award":["92146002"],"award-info":[{"award-number":["92146002"]}]},{"name":"National Natural Science Foundation of China","award":["2019YFB1405000"],"award-info":[{"award-number":["2019YFB1405000"]}]},{"name":"National Natural Science Foundation of China","award":["61873309"],"award-info":[{"award-number":["61873309"]}]},{"name":"National Natural Science Foundation of China","award":["92046024"],"award-info":[{"award-number":["92046024"]}]},{"name":"National Natural Science Foundation of China","award":["92146002"],"award-info":[{"award-number":["92146002"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>With the rapid development of prominent models, NL2SQL has made many breakthroughs, but customers still hope that the accuracy of NL2SQL can be continuously improved through optimization. The method based on large models has brought revolutionary changes to NL2SQL. This paper innovatively proposes a new NL2SQL method based on a large language model (LLM), which could be adapted to an edge-cloud computing platform. First, natural language is converted into Python language, and then SQL is generated through Python. At the same time, considering the traceability characteristics of financial industry regulatory requirements, this paper uses the open-source big model DeepSeek. After testing on the BIRD dataset, compared with most NL2SQL models based on large language models, EX is at least 2.73% higher than the original method, F1 is at least 3.72 higher than the original method, and VES is 6.34% higher than the original method. Through this innovative algorithm, the accuracy of NL2SQL in the financial industry is greatly improved, which can provide business personnel with a robust database access mode.<\/jats:p>","DOI":"10.3390\/fi17010012","type":"journal-article","created":{"date-parts":[[2025,1,2]],"date-time":"2025-01-02T06:05:10Z","timestamp":1735797910000},"page":"12","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["FI-NL2PY2SQL: Financial Industry NL2SQL Innovation Model Based on Python and Large Language Model"],"prefix":"10.3390","volume":"17","author":[{"given":"Xiaozheng","family":"Du","sequence":"first","affiliation":[{"name":"School of Computer Science, Fudan University, Shanghai 200438, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shijing","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Computer Science, Fudan University, Shanghai 200438, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5810-701X","authenticated-orcid":false,"given":"Feng","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, Shanghai Normal University Tianhua College, No. 1661 Shengxin North Road, Shanghai 201815, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cheng","family":"Wang","sequence":"additional","affiliation":[{"name":"Business Analysis BU, GienTech Technology Co., Ltd., Shanghai 200232, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Binh Minh","family":"Nguyen","sequence":"additional","affiliation":[{"name":"School of Information and Communication Technology, Hanoi University of Science and Technology, No. 1 Dai Co Viet, Hai Ba Trung, Hanoi 100000, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,1,2]]},"reference":[{"key":"ref_1","unstructured":"Liu, X., Shen, S., Li, B., Ma, P., Jiang, R., Zhang, Y., Fan, J., Li, G., Tang, N., and Luo, Y. (2024). A Survey of NL2SQL with Large Language Models: Where are we, and where are we going?. arXiv."},{"key":"ref_2","unstructured":"Shi, L., Tang, Z., Zhang, N., Zhang, X., and Yang, Z. (2024). A Survey on Employing Large Language Models for Text-to-SQL Tasks. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Yavuz, S., G\u00fcr, I., Su, Y., and Yan, X. (November, January 31). What it takes to achieve 100% condition accuracy on WikiSQL. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium.","DOI":"10.18653\/v1\/D18-1197"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., and Roman, S. (2018). Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. arXiv.","DOI":"10.18653\/v1\/D18-1425"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"415","DOI":"10.51594\/csitrj.v5i2.791","article-title":"Business intelligence in the era of big data: A review of analytical tools and competitive advantage","volume":"5","author":"Adewusi","year":"2024","journal-title":"Comput. Sci. IT Res. J."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"905","DOI":"10.1007\/s00778-022-00776-8","article-title":"A survey on deep learning approaches for text-to-SQL","volume":"32","author":"Koutrika","year":"2023","journal-title":"VLDB J."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Li, F., and Jagadish, H.V. (2014, January 22\u201327). NaLIR: An interactive natural language interface for querying relational databases. Proceedings of the 2014 ACM SIGMOD International Conference on Management of Data, New York, NY, USA.","DOI":"10.1145\/2588555.2594519"},{"key":"ref_8","unstructured":"Zhao, Y., Jiang, J., Hu, Y., Lan, W., Zhu, H., Chauhan, A., Li, A., Pan, L., Wang, J., and Hang, C.-W. (2022). Importance of synthesizing high-quality data for text-to-sql parsing. arXiv."},{"key":"ref_9","unstructured":"Vaswani, A. (2017, January 4\u20139). Attention is all you need. Advances in Neural Information Processing Systems. Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_10","unstructured":"Talaei, S., Pourreza, M., Chang, Y.C., Mirhoseini, A., and Saberi, A. (2024). Chess: Contextual harnessing for efficient sql synthesis. arXiv."},{"key":"ref_11","first-page":"1","article-title":"Codes: Towards building open-source language models for text-to-sql","volume":"2","author":"Li","year":"2024","journal-title":"Proc. ACM Manag. Data"},{"key":"ref_12","unstructured":"Pourreza, M., and Rafiei, D. (2024, January 9\u201315). Din-sql: Decomposed in-context learning of text-to-sql with self-correction. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_13","unstructured":"Cafero\u011flu, H.A., and Ulusoy, \u00d6. (2024). E-SQL: Direct Schema Linking via Question Enrichment in Text-to-SQL. arXiv."},{"key":"ref_14","first-page":"1","article-title":"Few-shot text-to-sql translation using structure and content prompt learning","volume":"1","author":"Gu","year":"2023","journal-title":"Proc. ACM Manag. Data"},{"key":"ref_15","unstructured":"Pourreza, M., Li, H., Sun, R., Chung, Y., Talaei, S., Kakkar, G.T., Gan, Y., Saberi, A., Ozcan, F., and Arik, S.O. (2024). CHASE-SQL: Multi-Path Reasoning and Preference Optimized Candidate Selection in Text-to-SQL. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Mao, W., Wang, R., Guo, J., Zeng, J., Gao, C., Han, P., and Liu, C. (2024). Enhancing Text-to-SQL Parsing through Question Rewriting and Execution-Guided Refinement. Findings of the Association for Computational Linguistics ACL 2024, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2024.findings-acl.120"},{"key":"ref_17","unstructured":"Hong, Z., Yuan, Z., Zhang, Q., Chen, H., Dong, J., Huang, F., and Huang, X. (2024). Next-Generation Database Interfaces: A Survey of LLM-based Text-to-SQL. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Yang, J., Hui, B., Yang, M., Yang, J., Lin, J., and Zhou, C. (2024). Synthesizing text-to-sql data from weak and strong llms. arXiv.","DOI":"10.18653\/v1\/2024.acl-long.425"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Gao, D., Wang, H., Li, Y., Sun, X., Qian, Y., Ding, B., and Zhou, J. (2023). Text-to-sql empowered by large language models: A benchmark evaluation. arXiv.","DOI":"10.14778\/3641204.3641221"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"255","DOI":"10.1002\/hcs2.61","article-title":"Large language models in health care: Development, applications, and challenges","volume":"2","author":"Yang","year":"2023","journal-title":"Health Care Sci."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Pourreza, M., and Rafiei, D. (2024). Dts-sql: Decomposed text-to-sql with small large language models. arXiv.","DOI":"10.18653\/v1\/2024.findings-emnlp.481"},{"key":"ref_22","unstructured":"Jiang, J., Wang, F., Shen, J., Kim, S., and Kim, S. (2024). A Survey on Large Language Models for Code Generation. arXiv."},{"key":"ref_23","unstructured":"DeepSeek-AI, Q.Z., Zhu, Q., Guo, D., Shao, Z., Yang, D., Wang, P., Xu, R., Wu, Y., Li, Y., and Gao, H. (2024). DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence. arXiv."},{"key":"ref_24","unstructured":"Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q.V., and Zhou, D. (December, January 28). Chain-of-thought prompting elicits reasoning in large language models. Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_25","unstructured":"Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. (2022). Self-consistency improves chain of thought reasoning in language models. arXiv."},{"key":"ref_26","unstructured":"Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K. (2024, January 9\u201315). Tree of thoughts: Deliberate problem solving with large language models. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Hao, S., Gu, Y., Ma, H., Hong, J.J., Wang, Z., Wang, D.Z., and Hu, Z. (2023). Reasoning with language model is planning with world model. arXiv.","DOI":"10.18653\/v1\/2023.emnlp-main.507"},{"key":"ref_28","unstructured":"Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2022). React: Synergizing reasoning and acting in language models. arXiv."},{"key":"ref_29","unstructured":"Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., and Yang, Y. (2024, January 9\u201315). Self-refine: Iterative refinement with self-feedback. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_30","unstructured":"Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S. (2024, January 9\u201315). Reflexion: Language agents with verbal reinforcement learning. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_31","unstructured":"Zhou, A., Yan, K., Shlapentokh-Rothman, M., Wang, H., and Wang, Y.X. (2023). Language agent tree search unifies reasoning acting and planning in language models. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"103193","DOI":"10.1016\/j.sysarc.2024.103193","article-title":"Transformers in source code generation: A comprehensive survey","volume":"153","author":"Ghaemi","year":"2024","journal-title":"J. Syst. Archit."},{"key":"ref_33","unstructured":"Fu, Y., Panda, R., Niu, X., Yue, X., Hajishirzi, H., Kim, Y., and Peng, H. (2024). Data engineering for scaling language models to 128k context. arXiv."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/17\/1\/12\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,7]],"date-time":"2025-10-07T15:23:43Z","timestamp":1759850623000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/17\/1\/12"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,2]]},"references-count":33,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,1]]}},"alternative-id":["fi17010012"],"URL":"https:\/\/doi.org\/10.3390\/fi17010012","relation":{},"ISSN":["1999-5903"],"issn-type":[{"value":"1999-5903","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,2]]}}}