{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T10:38:29Z","timestamp":1783420709620,"version":"3.54.6"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"8","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62207011, 62407013, and 62377009"],"award-info":[{"award-number":["62207011, 62407013, and 62377009"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Natural Science Foundation of Hubei Province of China","award":["2025AFB653"],"award-info":[{"award-number":["2025AFB653"]}]},{"name":"the Hubei Provincial Key Laboratory of Artificial Intelligence and Smart Learning","award":["2025AISL001"],"award-info":[{"award-number":["2025AISL001"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>\n            Knowledge graphs (KGs) are frequently confronted with the challenge of incompleteness, a problem that extends to multimodal knowledge graphs (MKGs). The primary goal of multimodal knowledge graph completion (MKGC) is to predict missing entities within MKGs. However, current MKGC methods face difficulties in adequately addressing modal preferences and imbalances in modal information. To overcome these issues, we introduce AdaMKGC, an innovative hybrid model incorporating an adaptive modality interaction transformer. This model employs a dynamic attention interaction strategy and a self-enhancing sampling approach. AdaMKGC achieves a more precise utilization of multimodal information by integrating modal preference information into modal interactions. Additionally, it effectively mitigates the issue of modal imbalance through targeted sampling and adjustment for entities with deficient information. Experimental evaluations demonstrate AdaMKGC\u2019s superior performance in overcoming these prevalent challenges. Compared to existing state-of-the-art MKGC models, AdaMKGC shows a notable enhancement of 28% in MR on the WN18-IMG dataset and an improvement of 2.7% in Hits@1 on the FB15k-237-IMG dataset. Our code is available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/HubuKG\/AdaMKGC\">https:\/\/github.com\/HubuKG\/AdaMKGC<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3760786","type":"journal-article","created":{"date-parts":[[2025,8,18]],"date-time":"2025-08-18T16:00:46Z","timestamp":1755532846000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":16,"title":["Adaptive Modality Interaction Transformer for Multimodal Knowledge Graph Completion"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-1908-8167","authenticated-orcid":false,"given":"Yue","family":"Jian","sequence":"first","affiliation":[{"name":"School of Computer Science, Hubei University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8952-215X","authenticated-orcid":false,"given":"Miao","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Hubei University, Wuhan, China and Hubei Provincial Key Laboratory of Artificial Intelligence and Smart Learning, Central China Normal University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-0382-6045","authenticated-orcid":false,"given":"Ziyue","family":"Qin","sequence":"additional","affiliation":[{"name":"School of Computer Science, Hubei University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-4222-6399","authenticated-orcid":false,"given":"Chuyuan","family":"Xie","sequence":"additional","affiliation":[{"name":"School of Computer Science, Hubei University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7784-7689","authenticated-orcid":false,"given":"Kui","family":"Xiao","sequence":"additional","affiliation":[{"name":"School of Computer Science, Hubei University and Hubei Key Laboratory of Big Data Intelligent Analysis and Application (Hubei University), Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6045-5208","authenticated-orcid":false,"given":"Yan","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Hubei University and Hubei Key Laboratory of Big Data Intelligent Analysis and Application (Hubei University), Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0443-0094","authenticated-orcid":false,"given":"Zhifei","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science, Hubei University and Hubei Key Laboratory of Big Data Intelligent Analysis and Application (Hubei University), Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,17]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3124805"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1068"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1522"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2022.103242"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376746"},{"key":"e_1_3_2_7_2","first-page":"2787","volume-title":"Proceedings of the 26th International Conference on Neural Information Processing Systems","author":"Bordes Antoine","year":"2013","unstructured":"Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko. 2013. Translating embeddings for modeling multi-relational data. In Proceedings of the 26th International Conference on Neural Information Processing Systems, 2787\u20132795."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3531992"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2022.11.042"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11573"},{"key":"e_1_3_2_11_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4171\u20134186."},{"key":"e_1_3_2_12_2","volume-title":"Proceedings of the 9th International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2021. An image is worth 16\u2009\u00d7\u200916 words: Transformers for image recognition at scale. In Proceedings of the 9th International Conference on Learning Representations."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2021.3070843"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2021.107188"},{"key":"e_1_3_2_15_2","first-page":"5583","volume-title":"Proceedings of the 38th International Conference on Machine Learning","volume":"139","author":"Kim Wonjae","year":"2021","unstructured":"Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. ViLT: Vision-and-language transformer without convolution or region supervision. In Proceedings of the 38th International Conference on Machine Learning, Vol. 139, 5583\u20135594."},{"key":"e_1_3_2_16_2","first-page":"1","volume-title":"Proceedings of the 3rd International Conference on Learning Representations","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the 3rd International Conference on Learning Representations, 1\u201315."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.488"},{"key":"e_1_3_2_18_2","unstructured":"Liunian Harold Li Mark Yatskar Da Yin Cho-Jui Hsieh and Kai-Wei Chang. 2019. VisualBERT: A simple and performant baseline for vision and language. arXiv:1908.03557. Retrieved from https:\/\/arxiv.org\/abs\/1908.03557"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i5.20521"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2023.3282989"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v29i1.9491"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3217449"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2014.2347204"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-21348-0_30"},{"key":"e_1_3_2_25_2","first-page":"13","volume-title":"Proceedings of the International Conference on Neural Information Processing Systems","author":"Lu Jiasen","year":"2019","unstructured":"Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Proceedings of the International Conference on Neural Information Processing Systems, 13\u201323."},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/219717.219748"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-2053"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.5555\/3104482.3104584"},{"key":"e_1_3_2_29_2","first-page":"1","article-title":"Automatic differentiation in PyTorch","author":"Paszke Adam","year":"2017","unstructured":"Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017. Automatic differentiation in PyTorch. In Proceedings of the Advances in Neural Information Processing Systems, 1\u20134.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1359"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-93417-4_38"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33013060"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01519"},{"key":"e_1_3_2_34_2","first-page":"1","volume-title":"Proceedings of the 7th International Conference on Learning Representations","author":"Sun Zhiqing","year":"2019","unstructured":"Zhiqing Sun, Zhi-Hong Deng, Jian-Yun Nie, and Jian Tang. 2019. RotatE: Knowledge graph embedding by relational rotation in complex space. In Proceedings of the 7th International Conference on Learning Representations, 1\u201318."},{"key":"e_1_3_2_35_2","first-page":"2071","volume-title":"Proceedings of the 33rd International Conference on Machine Learning","author":"Trouillon Th\u00e9o","year":"2016","unstructured":"Th\u00e9o Trouillon, Johannes Welbl, Sebastian Riedel, \u00c9ric Gaussier, and Guillaume Bouchard. 2016. Complex embeddings for simple link prediction. In Proceedings of the 33rd International Conference on Machine Learning, 2071\u20132080."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i03.5694"},{"key":"e_1_3_2_37_2","first-page":"1","volume-title":"Proceedings of the 8th International Conference on Learning Representations","author":"Vashishth Shikhar","year":"2020","unstructured":"Shikhar Vashishth, Soumya Sanyal, Vikram Nitin, and Partha P. Talukdar. 2020. Composition-based multi-relational graph convolutional networks. In Proceedings of the 8th International Conference on Learning Representations, 1\u201316."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3450043"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2022\/382"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.295"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475470"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2017.2754499"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2019.8852079"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v28i1.8870"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.5555\/3172077.3172327"},{"key":"e_1_3_2_46_2","first-page":"1","volume-title":"Proceedings of the 3rd International Conference on Learning Representations","author":"Yang Bishan","year":"2014","unstructured":"Bishan Yang, Wen-Tau Yih, Xiaodong He, Jianfeng Gao, and Li Deng. 2014. Embedding entities and relations for learning and inference in knowledge bases. In Proceedings of the 3rd International Conference on Learning Representations, 1\u201312."},{"key":"e_1_3_2_47_2","unstructured":"Liang Yao Chengsheng Mao and Yuan Luo. 2019. KG-BERT: BERT for knowledge graph completion. arXiv:1909.03193. Retrieved from https:\/\/arxiv.org\/abs\/1909.03193"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3449879"},{"key":"e_1_3_2_49_2","first-page":"1","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Zhang Ningyu","year":"2023","unstructured":"Ningyu Zhang, Lei Li, Xiang Chen, Xiaozhuan Liang, Shumin Deng, and Huajun Chen. 2023. Multimodal analogical reasoning over knowledge graphs. In Proceedings of the 11th International Conference on Learning Representations, 1\u201320."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE51399.2021.00280"},{"key":"e_1_3_2_51_2","volume-title":"Proceedings of the ACM Knowledge Discovery and Data Mining","author":"Zhang Yichi","year":"2022","unstructured":"Yichi Zhang and Wen Zhang. 2022. Knowledge graph completion with pre-trained multimodal transformer and twins negative sampling. In Proceedings of the ACM Knowledge Discovery and Data Mining."},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2020.3005952"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i4.20360"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3760786","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,19]],"date-time":"2025-09-19T17:08:42Z","timestamp":1758301722000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3760786"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,17]]},"references-count":52,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3760786"],"URL":"https:\/\/doi.org\/10.1145\/3760786","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,17]]},"assertion":[{"value":"2024-10-04","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-10","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-17","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}