{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T06:41:19Z","timestamp":1775544079598,"version":"3.50.1"},"reference-count":47,"publisher":"Association for Computing Machinery (ACM)","issue":"8","license":[{"start":{"date-parts":[[2024,6,29]],"date-time":"2024-06-29T00:00:00Z","timestamp":1719619200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2021ZD0140407"],"award-info":[{"award-number":["2021ZD0140407"]}]},{"name":"National High Level Hospital Clinical Research Funding","award":["2022-PUMCH-C041"],"award-info":[{"award-number":["2022-PUMCH-C041"]}]},{"name":"Beijing Natural Science Foundation","award":["722231"],"award-info":[{"award-number":["722231"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,8,31]]},"abstract":"<jats:p>\n            Recent years have witnessed the successful application of knowledge graph techniques in structured data processing, while how to incorporate knowledge from visual and textual modalities into knowledge graphs has been given less attention. To better organize them, Multimodal Knowledge Graphs (MKGs), comprising the structural triplets of traditional Knowledge Graphs (KGs) together with entity-related multimodal data (e.g., images and texts), have been introduced consecutively. However, it is still a great challenge to explore MKGs due to their inherent incompleteness. Although most existing Multimodal Knowledge Graph Completion (MKGC) approaches can infer missing triplets based on available factual triplets and multimodal information, they almost ignore the modal conflicts and supervisory effect, failing to achieve a more comprehensive understanding of entities. To address these issues, we propose a novel\n            <jats:underline>H<\/jats:underline>\n            ierarchical\n            <jats:underline>K<\/jats:underline>\n            nowledge\n            <jats:underline>A<\/jats:underline>\n            lignment (\n            <jats:bold>HKA<\/jats:bold>\n            ) framework for MKGC. Specifically, a macro-knowledge alignment module is proposed to capture global semantic relevance between modalities for dealing with modal conflicts in MKG. Furthermore, a micro-knowledge alignment module is also developed to reveal the local consistency information through inter- and intra-modality supervisory effects more effectively. By integrating different modal predictions, a final decision can be made. Experimental results on three benchmark MKGC tasks have demonstrated the effectiveness of the proposed HKA framework.\n          <\/jats:p>","DOI":"10.1145\/3664288","type":"journal-article","created":{"date-parts":[[2024,5,11]],"date-time":"2024-05-11T11:37:16Z","timestamp":1715427436000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["HKA: A Hierarchical Knowledge Alignment Framework for Multimodal Knowledge Graph Completion"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-2958-9288","authenticated-orcid":false,"given":"Yunhui","family":"Xu","sequence":"first","affiliation":[{"name":"Institute of Information Science, Beijing Jiaotong University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9326-9863","authenticated-orcid":false,"given":"Youru","family":"Li","sequence":"additional","affiliation":[{"name":"Institute of information Sciecne, Beijing Jiaotong University, Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8773-2789","authenticated-orcid":false,"given":"Muhao","family":"Xu","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Beijing Jiaotong University, Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7315-3276","authenticated-orcid":false,"given":"Zhenfeng","family":"Zhu","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Beijing Jiaotong University, Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8581-9554","authenticated-orcid":false,"given":"Yao","family":"Zhao","sequence":"additional","affiliation":[{"name":"Institute of Information Science; Beijing Key Laboratory of Advanced Information Science and Network Technology, Beijing Jiaotong University, Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,6,29]]},"reference":[{"key":"e_1_3_1_2_2","article-title":"Translating embeddings for modeling multi-relational data","volume":"26","author":"Bordes Antoine","year":"2013","unstructured":"Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko. 2013. Translating embeddings for modeling multi-relational data. Advances in Neural Information Processing Systems 26 (2013), 2787\u20132795.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i03.5665"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3539244"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2019.112948"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3531992"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11573"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210187"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11538"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1207\/s15516709cog1402_1"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2021.03.132"},{"key":"e_1_3_1_12_2","first-page":"2505","volume-title":"International Conference on Machine Learning","author":"Guo Lingbing","year":"2019","unstructured":"Lingbing Guo, Zequn Sun, and Wei Hu. 2019. Learning to exploit long-term relational dependencies in knowledge graphs. In International Conference on Machine Learning. PMLR, 2505\u20132514."},{"key":"e_1_3_1_13_2","first-page":"1","article-title":"Personalized recommendation system based on knowledge embedding and historical behavior","author":"Hui Bei","year":"2022","unstructured":"Bei Hui, Lizong Zhang, Xue Zhou, Xiao Wen, and Yuhui Nian. 2022. Personalized recommendation system based on knowledge embedding and historical behavior. Applied Intelligence 52 (2022), 1\u201313.","journal-title":"Applied Intelligence"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1067"},{"key":"e_1_3_1_15_2","article-title":"Simple embedding for link prediction in knowledge graphs","volume":"31","author":"Kazemi Seyed Mehran","year":"2018","unstructured":"Seyed Mehran Kazemi and David Poole. 2018. Simple embedding for link prediction in knowledge graphs. Advances in Neural Information Processing Systems 31 (2018), 4289\u20134300.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.153"},{"key":"e_1_3_1_17_2","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. Imagenet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems 25 (2012), 84\u201390.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6795"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.469"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3543507.3583554"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3545573"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v29i1.9491"},{"key":"e_1_3_1_23_2","first-page":"2168","volume-title":"International Conference on Machine Learning","author":"Liu Hanxiao","year":"2017","unstructured":"Hanxiao Liu, Yuexin Wu, and Yiming Yang. 2017. Analogical inference for multi-relational embeddings. In International Conference on Machine Learning. PMLR, 2168\u20132178."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-21348-0_30"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1223"},{"key":"e_1_3_1_26_2","article-title":"Recurrent models of visual attention","volume":"27","author":"Mnih Volodymyr","year":"2014","unstructured":"Volodymyr Mnih, Nicolas Heess, Alex Graves, and Koray Kavukcuoglu. 2014. Recurrent models of visual attention. Advances in Neural Information Processing Systems 27 (2014), 2204\u20132212.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611778"},{"key":"e_1_3_1_28_2","first-page":"3104482","volume-title":"International Conference on Machine Learning","volume":"11","author":"Nickel Maximilian","year":"2011","unstructured":"Maximilian Nickel, Volker Tresp, Hans-Peter Kriegel, et\u00a0al. 2011. A three-way model for collective learning on multi-relational data. In International Conference on Machine Learning, Vol. 11. 3104482\u20133104584."},{"key":"e_1_3_1_29_2","first-page":"8748","volume-title":"International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et\u00a0al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_1_30_2","article-title":"Dynamic routing between capsules","volume":"30","author":"Sabour Sara","year":"2017","unstructured":"Sara Sabour, Nicholas Frosst, and Geoffrey E. Hinton. 2017. Dynamic routing between capsules. Advances in Neural Information Processing Systems 30 (2017), 3859\u20133869.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-93417-4_38"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/1242572.1242667"},{"key":"e_1_3_1_33_2","article-title":"RotatE: Knowledge graph embedding by relational rotation in complex space","author":"Sun Zhiqing","year":"2019","unstructured":"Zhiqing Sun, Zhihong Deng, Jian-Yun Nie, and Jian Tang. 2019. RotatE: Knowledge graph embedding by relational rotation in complex space. In International Conference on Learning Representations,International Conference on Learning Representations.","journal-title":"International Conference on Learning Representations,International Conference on Learning Representations"},{"key":"e_1_3_1_34_2","series-title":"JMLR Workshop and Conference Proceedings","first-page":"2071","volume-title":"International Conference on Machine Learning 2016","volume":"48","author":"Trouillon Th\u00e9o","year":"2016","unstructured":"Th\u00e9o Trouillon, Johannes Welbl, Sebastian Riedel, \u00c9ric Gaussier, and Guillaume Bouchard. 2016. Complex embeddings for simple link prediction. In International Conference on Machine Learning 2016(JMLR Workshop and Conference Proceedings, Vol. 48). JMLR.org, 2071\u20132080. http:\/\/proceedings.mlr.press\/v48\/trouillon16.html"},{"key":"e_1_3_1_35_2","first-page":"2180","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long and Short Papers)","author":"Vu Thanh","year":"2019","unstructured":"Thanh Vu, Tu Dinh Nguyen, Dat Quoc Nguyen, Dinh Phung, et\u00a0al. 2019. A capsule network-based embedding model for knowledge graph completion and search personalization. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long and Short Papers). 2180\u20132189."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3450043"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475470"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2019.8852079"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v28i1.8870"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v30i1.10329"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.5555\/3172077.3172327"},{"key":"e_1_3_1_42_2","volume-title":"3rd International Conference on Learning Representations","author":"Yang Bishan","year":"2015","unstructured":"Bishan Yang, Wen-tau Yih, Xiaodong He, Jianfeng Gao, and Li Deng. 2015. Embedding entities and relations for learning and inference in knowledge bases. In 3rd International Conference on Learning Representations."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3532009"},{"key":"e_1_3_1_44_2","article-title":"KG-BERT: BERT for knowledge graph completion","author":"Yao Liang","year":"2019","unstructured":"Liang Yao, Chengsheng Mao, and Yuan Luo. 2019. KG-BERT: BERT for knowledge graph completion. arXiv preprint arXiv:1909.03193 (2019).","journal-title":"arXiv preprint arXiv:1909.03193"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.45"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12057"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.719"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE55515.2023.00015"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3664288","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3664288","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:06:14Z","timestamp":1750291574000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3664288"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,29]]},"references-count":47,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2024,8,31]]}},"alternative-id":["10.1145\/3664288"],"URL":"https:\/\/doi.org\/10.1145\/3664288","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,29]]},"assertion":[{"value":"2023-11-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-24","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-06-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}