{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T02:39:38Z","timestamp":1782873578611,"version":"3.54.5"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"7","funder":[{"DOI":"10.13039\/501100009592","name":"Beijing Municipal Science & Technology Commission","doi-asserted-by":"crossref","award":["Z231100001723002"],"award-info":[{"award-number":["Z231100001723002"]}],"id":[{"id":"10.13039\/501100009592","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Graduate Research and Practice Projects of Minzu University of China","award":["SJCX2024019"],"award-info":[{"award-number":["SJCX2024019"]}]},{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2022YFF0902500"],"award-info":[{"award-number":["2022YFF0902500"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Compilation of the Uyghur Bayangjing and development of its text-image corpus","award":["24VJXG063"],"award-info":[{"award-number":["24VJXG063"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,7,31]]},"abstract":"<jats:p>In the realm of Knowledge-based Visual Question Answering (KB-VQA), the intricacy of the task lies in adeptly retrieving pertinent information from external sources and seamlessly aligning and amalgamating multimodal features. While numerous studies have effectively leveraged external knowledge to enrich factual connections among entities, there exists a tendency to overlook the significant reservoir of implicit information inherent in the visual-textual dimension. This oversight often results in suboptimal alignment and an undue reliance on the knowledge base. To address these challenges, this article introduces a novel strategy called SCAG. This approach aggregates the semantic co-occurring attention from diverse regions within images and various tokens within textual inputs using guidance weights to construct joint probabilistic representations grounded in the visual and textual dimensions, respectively. By employing this alignment strategy, the goal is to substantially mitigate information loss, reinforce inter-feature constraints within the model, reduce reliance on external knowledge sources, and enhance self-reasoning capabilities. The efficacy of our proposed model is comprehensively evaluated on the VQAv2 and OK-VQA datasets, with comparative analyses against multiple models conducted on the Ambiguous Knowledge (AK) dataset. Notably, our model exhibits a noteworthy 4.62% improvement over the state-of-the-art in addressing the knowledge dependency problem.<\/jats:p>","DOI":"10.1145\/3734220","type":"journal-article","created":{"date-parts":[[2025,5,6]],"date-time":"2025-05-06T10:46:21Z","timestamp":1746528381000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["SCAG: Semantic Co-occurring Attention Guided Alignment for Knowledge-based Visual Question Answering"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8522-565X","authenticated-orcid":false,"given":"Zheng","family":"Liu","sequence":"first","affiliation":[{"name":"The Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance, Ministry of Education, Minzu University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-6004-9147","authenticated-orcid":false,"given":"Kunyu","family":"Yang","sequence":"additional","affiliation":[{"name":"The Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance, Ministry of Education, Minzu University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0787-550X","authenticated-orcid":false,"given":"Yu","family":"Weng","sequence":"additional","affiliation":[{"name":"The Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance, Ministry of Education, Minzu University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-8375-5027","authenticated-orcid":false,"given":"Zheng","family":"He","sequence":"additional","affiliation":[{"name":"The Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance, Ministry of Education, Minzu University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5064-6526","authenticated-orcid":false,"given":"Xuan","family":"Liu","sequence":"additional","affiliation":[{"name":"The Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance, Ministry of Education, Minzu University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6861-9684","authenticated-orcid":false,"given":"Honghao","family":"Gao","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,7,18]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3489142"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2022.3211622"},{"issue":"3","key":"e_1_3_2_4_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3618301","article-title":"Cross-modality multiple relations learning for knowledge-based visual question answering","volume":"20","author":"Wang Yan","year":"2023","unstructured":"Yan Wang, Peize Li, Qingyi Si, Hanwen Zhang, Wenyu Zang, Zheng Lin, and Peng Fu. 2023. Cross-modality multiple relations learning for knowledge-based visual question answering. ACM Transactions on Multimedia Computing, Communications, and Applications 20, 3 (Oct. 2023), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612389"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-88361-4_9"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01389"},{"issue":"4","key":"e_1_3_2_8_2","doi-asserted-by":"crossref","first-page":"3471","DOI":"10.1007\/s11760-024-03013-7","article-title":"A focus fusion attention mechanism integrated with image captions for knowledge graph-based visual question answering","volume":"18","author":"Ma Mingyang","year":"2024","unstructured":"Mingyang Ma, Turdi Tohti, Yi Liang, Zicheng Zuo, and Askar Hamdulla. 2024. A focus fusion attention mechanism integrated with image captions for knowledge graph-based visual question answering. Signal, Image and Video Processing 18, 4 (2024), 3471\u20133482.","journal-title":"Signal, Image and Video Processing"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","unstructured":"Qiao Jin Zheng Yuan Guangzhi Xiong Qianlan Yu Huaiyuan Ying Chuanqi Tan Mosha Chen Songfang Huang Xiaozhong Liu and Sheng Yu. 2022. Biomedical question answering: A survey of approaches and challenges. ACM Computing Surveys 55 2 (Jan. 2022) 1\u201336.","DOI":"10.1145\/3490238"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01973"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00503"},{"key":"e_1_3_2_12_2","first-page":"10560","article-title":"REVIVE: Regional visual representation matters in knowledge-based visual question answering","volume":"35","author":"Lin Yuanze","year":"2022","unstructured":"Yuanze Lin, Yujia Xie, Dongdong Chen, Yichong Xu, Chenguang Zhu, and Lu Yuan. 2022. REVIVE: Regional visual representation matters in knowledge-based visual question answering. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 35, 10560\u201310571.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.772"},{"key":"e_1_3_2_14_2","first-page":"956","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Gui Liangke","year":"2022","unstructured":"Liangke Gui, Borui Wang, Qiuyuan Huang, Alexander Hauptmann, Yonatan Bisk, and Jianfeng Gao. 2022. KAT: A knowledge augmented transformer for vision-and-language. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz (Eds.), Association for Computational Linguistics, 956\u2013968."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01438"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3688848"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612516"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","first-page":"12096","DOI":"10.18653\/v1\/2023.findings-emnlp.809","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Ghosal Deepanway","year":"2023","unstructured":"Deepanway Ghosal, Navonil Majumder, Roy Lee, Rada Mihalcea, and Soujanya Poria. 2023. Language guided visual question answering: Elevate your multimodal language model using knowledge-enriched prompts. In Findings of the Association for Computational Linguistics: EMNLP 2023. Houda Bouamor, Juan Pino, and Kalika Bali (Eds.), Association for Computational Linguistics, 12096\u201312102."},{"key":"e_1_3_2_19_2","first-page":"780","volume-title":"Proceedings of the International Conference on Intelligent Computing","author":"Chen Zheng","year":"2023","unstructured":"Zheng Chen and Yaxin Wen. 2023. Exploiting query knowledge embedding and trilinear joint embedding for visual question answering. In Proceedings of the International Conference on Intelligent Computing. Springer, 780\u2013791."},{"key":"e_1_3_2_20_2","first-page":"121","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part XXX","author":"Li Xiujun","year":"2020","unstructured":"Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. 2020. OSCAR: Object-semantics aligned pre-training for vision-language tasks. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part XXX. Springer, 121\u2013137."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3634918"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3607827.3616839"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611891"},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","unstructured":"Hao Tan and Mohit Bansal. 2019. LXMERT: Learning cross-modality encoder representations from transformers. arXiv:1908.07490. Retrieved from https:\/\/arxiv.org\/abs\/1908.07490","DOI":"10.18653\/v1\/D19-1514"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_2_27_2","first-page":"21","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Thauvin Dao","year":"2023","unstructured":"Dao Thauvin and St\u00e9phane Herbin. 2023. Knowledge informed sequential scene graph verification using VQA. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 21\u201331."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018876"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.108153"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","unstructured":"Keyang Cheng Zhou Jiang Yu Si Yi Ding and Hongjian Gu and Humaira Abdul Ghafoor. 2024. Multimodal joint representation by shared knowledge embedding with vision-language view for VQA. Research Square (2024). DOI: 10.21203\/rs.3.rs-2684160\/v1","DOI":"10.21203\/rs.3.rs-2684160\/v1"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","first-page":"421","DOI":"10.1007\/978-3-031-56027-9_26","volume-title":"Advances in Information Retrieval","author":"Lerner Paul","year":"2024","unstructured":"Paul Lerner, Olivier Ferret, and Camille Guinaudeau. 2024. Cross-modal retrieval for knowledge-based visual question answering. In Advances in Information Retrieval. Nazli Goharian, Nicola Tonellotto, Yulan He, Aldo Lipani, Graham McDonald, Craig Macdonald, and Iadh Ounis (Eds.), Springer, Cham, 421\u2013438."},{"key":"e_1_3_2_32_2","first-page":"5067","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Gao Feng","year":"2022","unstructured":"Feng Gao, Qing Ping, Govind Thattai, Aishwarya Reganti, Ying Nian Wu, and Prem Natarajan. 2022. Transform-retrieve-generate: Natural language-centric outside-knowledge visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 5067\u20135077."},{"key":"e_1_3_2_33_2","first-page":"23369","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Hu Ziniu","year":"2023","unstructured":"Ziniu Hu, Ahmet Iscen, Chen Sun, Zirui Wang, Kai-Wei Chang, Yizhou Sun, Cordelia Schmid, David A. Ross, and Alireza Fathi. 2023. REVEAL: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge memory. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 23369\u201323379."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.517"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00277"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","first-page":"489","DOI":"10.18653\/v1\/2020.findings-emnlp.44","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201920)","author":"Gard\u00e8res Fran\u00e7ois","year":"2020","unstructured":"Fran\u00e7ois Gard\u00e8res, Maryam Ziaeefard, Baptiste Abeloos, and Freddy Lecue. 2020. ConceptBert: Concept-aware representation for visual question answering. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201920), 489\u2013498."},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.7717\/peerj-cs.353"},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","unstructured":"Zihao Zhu J. Yu Yujing Wang Yajing Sun Yue Hu and Qi Wu. 2020. Mucko: Multi-layer cross-modal knowledge reasoning for fact-based visual question answering. arXiv:2006.09073. Retrieved from https:\/\/arxiv.org\/abs\/2006.09073","DOI":"10.24963\/ijcai.2020\/153"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2024.3384270"},{"key":"e_1_3_2_40_2","first-page":"5939","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Li Zejun","year":"2023","unstructured":"Zejun Li, Zhihao Fan, Jingjing Chen, Qi Zhang, Xuanjing Huang, and Zhongyu Wei. 2023. Unifying cross-lingual and cross-modal modeling towards weakly supervised multilingual vision-language pre-training. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.), Association for Computational Linguistics, 5939\u20135958."},{"key":"e_1_3_2_41_2","first-page":"3976","volume-title":"Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Ni Minheng","year":"2021","unstructured":"Minheng Ni, Haoyang Huang, Lin Su, Edward Cui, Taroon Bharti, Lijuan Wang, Dongdong Zhang, and Nan Duan. 2021. M3P: Learning universal representations via multitask multilingual multimodal pre-training. In Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3976\u20133985."},{"key":"e_1_3_2_42_2","doi-asserted-by":"crossref","first-page":"5731","DOI":"10.18653\/v1\/2023.acl-long.315","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zeng Yan","year":"2023","unstructured":"Yan Zeng, Wangchunshu Zhou, Ao Luo, Ziming Cheng, and Xinsong Zhang. 2023. Cross-view language modeling: Towards unified cross-lingual cross-modal pre-training. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.), Association for Computational Linguistics, 5731\u20135746."},{"key":"e_1_3_2_43_2","doi-asserted-by":"crossref","first-page":"4153","DOI":"10.1109\/CVPR46437.2021.00414","volume-title":"Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhou Mingyang","year":"2021","unstructured":"Mingyang Zhou, Luowei Zhou, Shuohang Wang, Yu Cheng, Linjie Li, Zhou Yu, and Jingjing Liu. 2021. UC2: Universal cross-lingual cross-modal vision-and-language pre-training. In Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4153\u20134163."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.100"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.74"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3734220","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,18]],"date-time":"2025-07-18T23:54:48Z","timestamp":1752882888000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3734220"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,18]]},"references-count":45,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2025,7,31]]}},"alternative-id":["10.1145\/3734220"],"URL":"https:\/\/doi.org\/10.1145\/3734220","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,18]]},"assertion":[{"value":"2024-10-04","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-18","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}