{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T13:32:34Z","timestamp":1785418354615,"version":"3.56.0"},"reference-count":0,"publisher":"European Alliance for Innovation n.o.","issue":"7","license":[{"start":{"date-parts":[[2026,1,26]],"date-time":"2026-01-26T00:00:00Z","timestamp":1769385600000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-sa\/4.0"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["ICST Transactions on Scalable Information Systems"],"abstract":"<jats:p>INTRODUCTION: Large Language Models (LLMs), a major breakthrough in artificial intelligence, have been widely applied across various domains in recent years. Their powerful capabilities in language comprehension and generation enable effective handling of diverse natural language processing tasks, such as text generation, question answering, machine translation, and information retrieval. This paper investigates the application of LLM technology in data integration, a core aspect of data governance. In contrast to end-to-end black-box approaches, we reframe data integration as a problem of discovering interpretable mapping rules through symbolic regression.OBJECTIVES: We begin by defining the fundamental problem of data integration. We then propose a general-purpose large model framework for data governance, built on a deep symbolic regression foundation. The framework comprises a symbolic expression generator and a metadata-enhanced executor, aiming to achieve both high accuracy and interpretability.METHODS: The model is trained using a combination of recurrent neural networks and reinforcement learning techniques, for expression generation and the execution of the discovered rules is structured based on a Transformer encoder architecture enhanced with a dedicated metadata embedding layer. To enhance performance, we incorporate metadata fine-tuning, where the generated symbolic expressions serve as key metadata to guide the integration process.RESULTS: Finally, the proposed model is evaluated on two representative data integration tasks, with experimental results demonstrating its effectiveness.CONCLUSION: The results validate its practical quality and highlight the advantage of the symbolic regression paradigm in enhancing interpretability.<\/jats:p>","DOI":"10.4108\/eetsis.10245","type":"journal-article","created":{"date-parts":[[2026,1,26]],"date-time":"2026-01-26T10:37:09Z","timestamp":1769423829000},"source":"Crossref","is-referenced-by-count":0,"title":["Research on the Application of Large Language Model in Data Integration"],"prefix":"10.4108","volume":"12","author":[{"given":"Zhanfang","family":"Chen","sequence":"first","affiliation":[{"id":[{"id":"https:\/\/ror.org\/007mntk44","id-type":"ROR","asserted-by":"publisher"}],"name":"Changchun University of Science and Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-6040-2938","authenticated-orcid":true,"given":"Yuan","family":"Ren","sequence":"additional","affiliation":[{"id":[{"id":"https:\/\/ror.org\/007mntk44","id-type":"ROR","asserted-by":"publisher"}],"name":"Changchun University of Science and Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaoming","family":"Jiang","sequence":"additional","affiliation":[{"id":[{"id":"https:\/\/ror.org\/007mntk44","id-type":"ROR","asserted-by":"publisher"}],"name":"Changchun University of Science and Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ruipeng","family":"Qi","sequence":"additional","affiliation":[{"id":[{"id":"https:\/\/ror.org\/007mntk44","id-type":"ROR","asserted-by":"publisher"}],"name":"Changchun University of Science and Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"2587","published-online":{"date-parts":[[2026,1,26]]},"container-title":["ICST Transactions on Scalable Information Systems"],"original-title":[],"link":[{"URL":"https:\/\/publications.eai.eu\/index.php\/sis\/article\/download\/10245\/3876","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/publications.eai.eu\/index.php\/sis\/article\/download\/10245\/3876","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,27]],"date-time":"2026-01-27T09:09:29Z","timestamp":1769504969000},"score":1,"resource":{"primary":{"URL":"https:\/\/publications.eai.eu\/index.php\/sis\/article\/view\/10245"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,26]]},"references-count":0,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2026,1,21]]}},"URL":"https:\/\/doi.org\/10.4108\/eetsis.10245","relation":{},"ISSN":["2032-9407"],"issn-type":[{"value":"2032-9407","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,26]]}}}