{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,27]],"date-time":"2026-07-27T19:11:56Z","timestamp":1785179516460,"version":"3.55.0"},"reference-count":34,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2025,6,24]],"date-time":"2025-06-24T00:00:00Z","timestamp":1750723200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Fundamental Research Funds for the Central Universities","award":["2023QNYL24"],"award-info":[{"award-number":["2023QNYL24"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Multimodal knowledge graph entity alignment is a key basic task of knowledge fusion and integration, which is used to identify entities with semantic equivalent but different representation forms in different knowledge graphs. Previous entity alignment research has mostly focused on encoding and utilizing basic features such as entity names and attributes; however, it is difficult to comprehensively capture the rich semantic information of entities by solely relying on these basic features. To effectively overcome this limitation, this paper proposes a fusion-optimized multimodal entity alignment method, FMEA-TD. Compared with previous work, this method makes full use of the textual description information in the knowledge graph to provide rich supplements for entity features, thereby better capturing the entity semantics and solving the problems faced by relying solely on the entity\u2019s own features. FMEA-TD is able to effectively fuse the entity\u2019s own information and text description information through multimodal cooperation confidence, establish the interaction mechanism between them, and thus promote mutual collaboration between different modalities, which enhances the model\u2019s ability to understand the semantic text. Experimentally validated, FMEA-TD outperforms current state-of-the-art baseline methods on public knowledge graph datasets.<\/jats:p>","DOI":"10.3390\/info16070534","type":"journal-article","created":{"date-parts":[[2025,6,24]],"date-time":"2025-06-24T10:44:41Z","timestamp":1750761881000},"page":"534","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Fusion-Optimized Multimodal Entity Alignment with Textual Descriptions"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-5026-4016","authenticated-orcid":false,"given":"Chenchen","family":"Wang","sequence":"first","affiliation":[{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance, Ministry of Education, Minzu University of China, Beijing 100081, China"},{"name":"School of Information Engineering, Minzu University of China, Beijing 100081, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"family":"Chaomurilige","sequence":"additional","affiliation":[{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance, Ministry of Education, Minzu University of China, Beijing 100081, China"},{"name":"School of Chinese Ethnic Minority Languages and Literatures, Minzu University of China, Beijing 100081, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu","family":"Weng","sequence":"additional","affiliation":[{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance, Ministry of Education, Minzu University of China, Beijing 100081, China"},{"name":"School of Information Engineering, Minzu University of China, Beijing 100081, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xuan","family":"Liu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance, Ministry of Education, Minzu University of China, Beijing 100081, China"},{"name":"School of Information Engineering, Minzu University of China, Beijing 100081, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zheng","family":"Liu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance, Ministry of Education, Minzu University of China, Beijing 100081, China"},{"name":"School of Chinese Ethnic Minority Languages and Literatures, Minzu University of China, Beijing 100081, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,6,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"167","DOI":"10.3233\/SW-140134","article-title":"Dbpedia\u2014A large-scale, multilingual knowledge base extracted from wikipedia","volume":"6","author":"Lehmann","year":"2015","journal-title":"Semant. Web"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"203","DOI":"10.1016\/j.websem.2008.06.001","article-title":"Yago: A large ontology from wikipedia and wordnet","volume":"6","author":"Suchanek","year":"2008","journal-title":"J. Web Semant."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Moussallem, D., Ngonga Ngomo, A.C., Buitelaar, P., and Arcan, M. (2019, January 19\u201321). Utilizing knowledge graphs for neural machine translation augmentation. Proceedings of the 10th International Conference on Knowledge Capture, Marina Del Rey, CA, USA.","DOI":"10.1145\/3360901.3364423"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Srivastava, S., Patidar, M., Chowdhury, S., Agarwal, P., Bhattacharya, I., and Shroff, G. (2021, January 19\u201323). Complex question answering on knowledge graphs using machine translation and multi-task learning. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics, Online. Main Volume.","DOI":"10.18653\/v1\/2021.eacl-main.300"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Tao, W., Zhu, H., Tan, K., Wang, J., Liang, Y., Jiang, H., Yuan, P., and Lan, Y. (2024, January 8\u201312). Finqa: A training-free dynamic knowledge graph question answering system in finance with llm-based revision. Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Vilnius, Lithuania.","DOI":"10.1007\/978-3-031-70371-3_32"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Li, M., Zareian, A., Lin, Y., Pan, X., Whitehead, S., Chen, B., Wu, B., Ji, H., Chang, S.F., and Voss, C. (2020, January 5\u201310). Gaia: A fine-grained multimedia knowledge extraction system. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Online.","DOI":"10.18653\/v1\/2020.acl-demos.11"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Wen, H., Lin, Y., Lai, T., Pan, X., Li, S., Lin, X., Zhou, B., Li, M., Wang, H., and Zhang, H. (2021, January 6\u201311). Resin: A dockerized schema-guided cross-document cross-lingual cross-media information extraction and event tracking system. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Demonstrations, Online.","DOI":"10.18653\/v1\/2021.naacl-demos.16"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"715","DOI":"10.1109\/TKDE.2022.3224228","article-title":"Multi-modal knowledge graph construction and application: A survey","volume":"36","author":"Zhu","year":"2022","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Li, J., Luo, R., Sun, J., Xiao, J., and Yang, Y. (2024, January 8\u201312). Prior bilinear-based models for knowledge graph completion. Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Vilnius, Lithuania.","DOI":"10.1007\/978-3-031-70352-2_19"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Yu, M., Zuo, Y., Zhang, W., Zhao, M., Xu, T., Zhao, Y., Guo, J., and Yu, J. (2024, January 8\u201312). Graph attention network with relational dynamic factual fusion for knowledge graph completion. Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Vilnius, Lithuania.","DOI":"10.1007\/978-3-031-70359-1_6"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Liu, F., Chen, M., Roth, D., and Collier, N. (2021, January 19\u201321). Visual pivoting for (unsupervised) entity alignment. Proceedings of the AAAI Conference on Artificial Intelligence, Online.","DOI":"10.1609\/aaai.v35i5.16550"},{"key":"ref_12","unstructured":"Lin, Z., Zhang, Z., Wang, M., Shi, Y., Wu, X., and Zheng, Y. (2022). Multi-modal contrastive representation learning for entity alignment. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Li, Y., Chen, J., Li, Y., Xiang, Y., Chen, X., and Zheng, H.T. (2023, January 4\u201310). Vision, deduction and alignment: An empirical study on multi-modal knowledge graph alignment. Proceedings of the ICASSP 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece.","DOI":"10.1109\/ICASSP49357.2023.10094863"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Mao, X., Wang, W., Wu, Y., and Lan, M. (2021). From alignment to assignment: Frustratingly simple unsupervised entity alignment. arXiv.","DOI":"10.18653\/v1\/2021.emnlp-main.226"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Qi, Z., Zhang, Z., Chen, J., Chen, X., Xiang, Y., Zhang, N., and Zheng, Y. (2021). Unsupervised knowledge graph alignment by probabilistic reasoning and semantic embedding. arXiv.","DOI":"10.24963\/ijcai.2021\/278"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Jiang, C., Qian, Y., Chen, L., Gu, Y., and Xie, X. (2023, January 18\u201322). Unsupervised deep cross-language entity alignment. Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Turin, Italy.","DOI":"10.1007\/978-3-031-43421-1_1"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zeng, W., Tang, J., and Zhao, X. (2019, January 16\u201320). Iterative representation learning for entity alignment leveraging textual information. Proceedings of the Machine Learning and Knowledge Discovery in Databases: International Workshops of ECML PKDD 2019, W\u00fcrzburg, Germany. Part I.","DOI":"10.1007\/978-3-030-43823-4_40"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"134","DOI":"10.1007\/978-3-030-55130-8_12","article-title":"Mmea: Entity alignment for multi-modal knowledge graph","volume":"Volume 13","author":"Chen","year":"2020","journal-title":"Proceedings of the Knowledge Science, Engineering and Management: 13th International Conference, KSEM 2020"},{"key":"ref_19","first-page":"e1","article-title":"Bert-int: A bert-based interaction model for knowledge graph alignment","volume":"100","author":"Tang","year":"2020","journal-title":"Interactions"},{"key":"ref_20","unstructured":"Cao, B., Xia, Y., Ding, Y., Zhang, C., and Hu, Q. (2024). Predictive dynamic fusion. arXiv."},{"key":"ref_21","unstructured":"Veli\u010dkovi\u0107, P., Cucurull, G., Casanova, A., Romero, A., Lio, P., and Bengio, Y. (2017). Graph attention networks. arXiv."},{"key":"ref_22","unstructured":"Kipf, T.N., and Welling, M. (2016). Semi-supervised classification with graph convolutional networks. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C.D. (2014, January 25\u201329). Glove: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Reimers, N. (2019). Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv.","DOI":"10.18653\/v1\/D19-1410"},{"key":"ref_26","unstructured":"Devlin, J. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_27","unstructured":"Hadsell, R., Chopra, S., and LeCun, Y. (2006, January 17\u201322). Dimensionality reduction by learning an invariant mapping. Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201906), New York, NY, USA."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Chen, M., Tian, Y., Yang, M., and Zaniolo, C. (2016). Multilingual knowledge graph embeddings for cross-lingual knowledge alignment. arXiv.","DOI":"10.24963\/ijcai.2017\/209"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Sun, Z., Hu, W., Zhang, Q., and Qu, Y. (2018, January 13\u201319). Bootstrapping entity alignment with knowledge graph embedding. Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), Stockholm, Sweden.","DOI":"10.24963\/ijcai.2018\/611"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Li, C., Cao, Y., Hou, L., Shi, J., Li, J., and Chua, T.S. (2019). Semi-Supervised Entity Alignment via Joint Knowledge Embedding Model and Cross-Graph Mode, Association for Computational Linguistics.","DOI":"10.18653\/v1\/D19-1274"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Yang, H.W., Zou, Y., Shi, P., Lu, W., Lin, J., and Sun, X. (2019). Aligning cross-lingual entities with multi-aspect information. arXiv.","DOI":"10.18653\/v1\/D19-1451"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wu, Y., Liu, X., Feng, Y., Wang, Z., and Zhao, D. (2019). Jointly learning entity and relation representations for entity alignment. arXiv.","DOI":"10.18653\/v1\/D19-1023"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wang, H., Liu, Q., Huang, R., and Zhang, J. (2023). Multi-modal entity alignment method based on feature enhancement. Appl. Sci., 13.","DOI":"10.3390\/app13116747"},{"key":"ref_34","unstructured":"Brody, S., Alon, U., and Yahav, E. (2021). How attentive are graph attention networks?. arXiv."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/7\/534\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T17:57:54Z","timestamp":1760032674000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/7\/534"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,24]]},"references-count":34,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2025,7]]}},"alternative-id":["info16070534"],"URL":"https:\/\/doi.org\/10.3390\/info16070534","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,24]]}}}