{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T15:58:09Z","timestamp":1785427089575,"version":"3.56.0"},"reference-count":35,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2025,4,29]],"date-time":"2025-04-29T00:00:00Z","timestamp":1745884800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>In the field of databases, Large Language Models (LLMs) have recently been studied for generating SQL queries from textual descriptions, while their use for conceptual or logical data modeling remains less explored. The conceptual design of relational databases commonly relies on the entity-relationship (ER) data model, where translation rules enable mapping an ER schema into corresponding relational tables with their constraints. Our study investigates the capability of LLMs to describe in natural language a database conceptual data model based on the ER schema. Whether for documentation, onboarding, or communication with non-technical stakeholders, LLMs can significantly improve the process of explaining the ER schema by generating accurate descriptions about how the components interact as well as the represented information. To guide the LLM with challenging constructs, specific hints are defined to provide an enriched ER schema. Different LLMs have been explored (ChatGPT 3.5 and 4, Llama2, Gemini, Mistral 7B) and different metrics (F1 score, ROUGE, perplexity) are used to assess the quality of the generated descriptions and compare the different LLMs.<\/jats:p>","DOI":"10.3390\/info16050368","type":"journal-article","created":{"date-parts":[[2025,4,30]],"date-time":"2025-04-30T05:50:17Z","timestamp":1745992217000},"page":"368","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Exploring Large Language Models\u2019 Ability to Describe Entity-Relationship Schema-Based Conceptual Data Models"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7003-129X","authenticated-orcid":false,"given":"Andrea","family":"Avignone","sequence":"first","affiliation":[{"name":"Department of Control and Computer Engineering, Politecnico di Torino, 10129 Torino, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alessia","family":"Tierno","sequence":"additional","affiliation":[{"name":"Department of Control and Computer Engineering, Politecnico di Torino, 10129 Torino, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3819-2696","authenticated-orcid":false,"given":"Alessandro","family":"Fiori","sequence":"additional","affiliation":[{"name":"Department of Control and Computer Engineering, Politecnico di Torino, 10129 Torino, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5740-5004","authenticated-orcid":false,"given":"Silvia","family":"Chiusano","sequence":"additional","affiliation":[{"name":"Department of Control and Computer Engineering, Politecnico di Torino, 10129 Torino, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,4,29]]},"reference":[{"key":"ref_1","unstructured":"Atzeni, P., Ceri, S., Paraboschi, S., and Torlone, R. (1999). Database Systems\u2014Concepts, Languages and Architectures, McGraw-Hill Book Company."},{"key":"ref_2","unstructured":"Ramakrishnan, R., and Gehrke, J. (2003). Database Management Systems, McGraw-Hill. [3rd ed.]."},{"key":"ref_3","unstructured":"Silberschatz, A., Korth, H.F., and Sudarshan, S. (2010). Database System Concepts, McGraw-Hill. [6th ed.]."},{"key":"ref_4","unstructured":"Kimball, R., and Ross, M. (2002). The Data Warehouse Toolkit: The Complete Guide to Dimensional Modeling, Wiley. [2nd ed.]."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3512467","article-title":"A Survey of Knowledge-enhanced Text Generation","volume":"54","author":"Yu","year":"2022","journal-title":"ACM Comput. Surv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Chen, X., Gao, C., Chen, C., Zhang, G., and Liu, Y. (2025). An Empirical Study on Challenges for LLM Application Developers. ACM Trans. Softw. Eng. Methodol., just accepted.","DOI":"10.1145\/3715007"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3695988","article-title":"Large Language Models for Software Engineering: A Systematic Literature Review","volume":"33","author":"Hou","year":"2024","journal-title":"ACM Trans. Softw. Eng. Methodol."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"905","DOI":"10.1007\/s00778-022-00776-8","article-title":"A Survey on Deep Learning Approaches for Text-to-SQL","volume":"32","author":"Koutrika","year":"2023","journal-title":"VLDB J."},{"key":"ref_9","unstructured":"Korhonen, A., Traum, D., and M\u00e0rquez, L. (August, January 28). Towards Complex Text-to-SQL in Cross-Domain Database with Intermediate Representation. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"101983","DOI":"10.1016\/j.wpi.2020.101983","article-title":"Patent claim generation by fine-tuning OpenAI GPT-2","volume":"62","author":"Lee","year":"2020","journal-title":"World Pat. Inf."},{"key":"ref_11","unstructured":"Davis, B., Graham, Y., Kelleher, J., and Sripada, Y. (2020, January 15\u201318). RecipeNLG: A Cooking Recipes Dataset for Semi-Structured Text Generation. Proceedings of the 13th International Conference on Natural Language Generation, Dublin, Ireland."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Avignone, A., Fiori, A., Chiusano, S., and Rizzo, G. (2023, January 18\u201320). Generation of Textual\/Video Descriptions for Technological Products Based on Structured Data. Proceedings of the 2023 IEEE 17th International Conference on Application of Information and Communication Technologies (AICT), Baku, Azerbaijan.","DOI":"10.1109\/AICT59525.2023.10313177"},{"key":"ref_13","unstructured":"Scott, D., Bel, N., and Zong, C. (2020, January 8\u201313). TableGPT: Few-shot Table-to-Text Generation with Table Structure Reconstruction and Content Matching. Proceedings of the International Conference on Computational Linguistics, Barcelona, Spain."},{"key":"ref_14","first-page":"4","article-title":"Application of Large Language Models to Software Engineering Tasks: Opportunities, Risks, and Implications","volume":"40","author":"Ozkaya","year":"2023","journal-title":"IEEE Softw."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zheng, Z., Ning, K., Zhong, Q., Chen, J., Chen, W., Guo, L., Wang, W., and Wang, Y. (2024). Towards an Understanding of Large Language Models in Software Engineering tasks. Empir. Softw. Engg., 30.","DOI":"10.1007\/s10664-024-10602-0"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Marques, N., Silva, R.R., and Bernardino, J. (2024). Using ChatGPT in Software Requirements Engineering: A Comprehensive Review. Future Internet, 16.","DOI":"10.3390\/fi16060180"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Cai, R., Xu, B., Zhang, Z., Yang, X., Li, Z., and Liang, Z. (2018, January 13\u201319). An Encoder-Decoder Framework translating Natural Language to Database Queries. Proceedings of the IJCAI\u201918: International Joint Conference on Artificial Intelligence, Stockholm, Sweden.","DOI":"10.24963\/ijcai.2018\/553"},{"key":"ref_18","unstructured":"Deng, N., Chen, Y., and Zhang, Y. (2022). Recent Advances in Text-to-SQL: A Survey of What We Have and What We Expect. arXiv, abs\/2208.10099."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"C\u00e2mara, V., Mendonca-Neto, R., Silva, A., and Cordovil, L. (2024, January 5\u20138). A Large Language Model approach to SQL-to-Text Generation. Proceedings of the 2024 IEEE International Conference on Consumer Electronics (ICCE), Las Vegas, NV, USA.","DOI":"10.1109\/ICCE59016.2024.10444148"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"43156","DOI":"10.1109\/ACCESS.2020.2977613","article-title":"M-SQL: Multi-Task Representation Learning for Single-Table Text2sql Generation","volume":"8","author":"Zhang","year":"2020","journal-title":"IEEE Access"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1280","DOI":"10.1109\/TR.2023.3336330","article-title":"LI-EMRSQL: Linking Information Enhanced Text2SQL Parsing on Complex Electronic Medical Records","volume":"73","author":"Li","year":"2024","journal-title":"IEEE Trans. Reliab."},{"key":"ref_22","unstructured":"Walker, M., Ji, H., and Stent, A. TypeSQL: Knowledge-Based Type-Aware Neural Text-to-SQL Generation. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), New Orleans, LA, USA."},{"key":"ref_23","unstructured":"Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B.D., and Schockaert, S. (2025, January 19\u201324). Leveraging LLM-Generated Schema Descriptions for Unanswerable Question Detection in Clinical Data. Proceedings of the International Conference on Computational Linguistics, Abu Dhabi, United Arab Emirates."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Ferrari, A., Abualhaija, S., and Arora, C. (2024, January 24\u201325). Model Generation with LLMs: From Requirements to UML Sequence Diagrams. Proceedings of the 2024 IEEE 32nd International Requirements Engineering Conference Workshops (REW), Reykjavik, Iceland.","DOI":"10.1109\/REW61692.2024.00044"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Wang, B., Wang, C., Liang, P., Li, B., and Zeng, C. (2024, January 7\u201313). How LLMs Aid in UML Modeling: An Exploratory Study with Novice Analysts. Proceedings of the 2024 IEEE International Conference on Software Services Engineering (SSE), Shenzhen, China.","DOI":"10.1109\/SSE62657.2024.00046"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Ardimento, P., Bernardi, M.L., and Cimitile, M. (July, January 30). Teaching UML using a RAG-based LLM. Proceedings of the 2024 International Joint Conference on Neural Networks (IJCNN), Yokohama, Japan.","DOI":"10.1109\/IJCNN60899.2024.10651492"},{"key":"ref_27","unstructured":"Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., and Askell, A. (2020, January 6\u201312). Language Models are Few-shot Learners. Proceedings of the NIPS \u201920: 34th International Conference on Neural Information Processing Systems, Red Hook, NY, USA."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Sivarajkumar, S., Kelley, M., Samolyk-Mazzanti, A., Visweswaran, S., and Wang, Y. (2023). An Empirical Evaluation of Prompting Strategies for Large Language Models in Zero-Shot Clinical Natural Language Processing. arXiv, abs\/2309.08008.","DOI":"10.2196\/preprints.55318"},{"key":"ref_29","unstructured":"Kojima, T., Gu, S.S., Reid, M., Matsuo, Y., and Iwasawa, Y. (December, January 28). Large Language Models are Zero-shot Reasoners. Proceedings of the NIPS \u201922: 36th International Conference on Neural Information Processing Systems, Red Hook, NY, USA."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Reynolds, L., and McDonell, K. (2021, January 8\u201313). Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm. Proceedings of the CHI EA \u201921: Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, New York, NY, USA.","DOI":"10.1145\/3411763.3451760"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L. (2022). Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?. arXiv, abs\/2202.12837.","DOI":"10.18653\/v1\/2022.emnlp-main.759"},{"key":"ref_32","unstructured":"Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E.H., Le, Q.V., and Zhou, D. (December, January 28). Chain-of-thought Prompting elicits reasoning in Large Language Models. Proceedings of the NIPS \u201922: 36th International Conference on Neural Information Processing Systems, Red Hook, NY, USA."},{"key":"ref_33","unstructured":"Zhang, Z., Zhang, A., Li, M., and Smola, A. (2022). Automatic Chain of Thought Prompting in Large Language Models. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"S63","DOI":"10.1121\/1.2016299","article-title":"Perplexity\u2014A Measure of the Difficulty of Speech Recognition Tasks","volume":"62","author":"Jelinek","year":"1977","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"102393","DOI":"10.1016\/j.artmed.2022.102393","article-title":"Semantic Coherence Markers: The Contribution of Perplexity Metrics","volume":"134","author":"Colla","year":"2022","journal-title":"Artif. Intell. Med."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/5\/368\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T17:24:31Z","timestamp":1760030671000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/5\/368"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,29]]},"references-count":35,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2025,5]]}},"alternative-id":["info16050368"],"URL":"https:\/\/doi.org\/10.3390\/info16050368","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,29]]}}}