{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T05:17:09Z","timestamp":1779081429954,"version":"3.51.4"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2025,5,17]],"date-time":"2025-05-17T00:00:00Z","timestamp":1747440000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2022YFB3305500"],"award-info":[{"award-number":["2022YFB3305500"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62273089"],"award-info":[{"award-number":["62273089"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100021171","name":"Guangdong Basic and Applied Basic Research Foundation","doi-asserted-by":"crossref","award":["024A1515010237"],"award-info":[{"award-number":["024A1515010237"]}],"id":[{"id":"10.13039\/501100021171","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Hong Kong Research Grants Council under the General Research","award":["PolyU15200023"],"award-info":[{"award-number":["PolyU15200023"]}]},{"name":"Faculty Research Grants","award":["SDS24A8"],"award-info":[{"award-number":["SDS24A8"]}]},{"name":"Direct Grant","award":["DR25E8"],"award-info":[{"award-number":["DR25E8"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2025,6,30]]},"abstract":"<jats:p>\n            Knowledge-based visual question answering (KB-VQA) requires reasoning about the visual grounding relations between the images and questions by incorporating external knowledge. Existing works typically retrieve knowledge from knowledge graphs by leveraging global multimodal representations of image\u2013text pairs for graph convolution, which neglect contextual clues at hop granularity, resulting in suboptimal spreading and leveraging of contextual information. To this end, we propose a multi-hop graph reasoning network (MGRN) for KB-VQA, which consists of a knowledge graph constructor (KGC) module, a semantic-instructed graph reasoning (SGR) module, and an answering module. MGRN exploits multimodal semantics from given images and questions as instructions for graph reasoning to obtain the knowledge representation from either the scene graph or knowledge base. Specifically, KGC fuses the scene graph with triplets from ConceptNet and Comet to construct a contextual knowledge graph for retrieving knowledge representation. Furthermore, SGR conducts multi-hop graph reasoning to select top-\n            <jats:italic>K<\/jats:italic>\n            knowledge items for answering by passing and filtering interplay messages on contextual knowledge graphs under the guidance of multimodal semantic representation. Extensive experiments conducted on two public datasets show the effectiveness and outperformance of our method.\n          <\/jats:p>","DOI":"10.1145\/3724125","type":"journal-article","created":{"date-parts":[[2025,3,19]],"date-time":"2025-03-19T17:45:03Z","timestamp":1742406303000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["A Multi-Hop Graph Reasoning Network for Knowledge-Based VQA"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-4074-9165","authenticated-orcid":false,"given":"Zihan","family":"Hu","sequence":"first","affiliation":[{"name":"Guangdong University of Technology, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-0093-1128","authenticated-orcid":false,"given":"Jiuxiang","family":"You","sequence":"additional","affiliation":[{"name":"Hiroshima University, Higashihiroshima, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9392-1375","authenticated-orcid":false,"given":"Zhenguo","family":"Yang","sequence":"additional","affiliation":[{"name":"Guangdong University of Technology, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3201-0038","authenticated-orcid":false,"given":"Xiaoping","family":"Li","sequence":"additional","affiliation":[{"name":"Guangdong University of Technology, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0965-3617","authenticated-orcid":false,"given":"Haoran","family":"Xie","sequence":"additional","affiliation":[{"name":"Lingnan University, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3370-471X","authenticated-orcid":false,"given":"Qing","family":"Li","sequence":"additional","affiliation":[{"name":"The Hong Kong Polytechnic University, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6237-6607","authenticated-orcid":false,"given":"Wenyin","family":"Liu","sequence":"additional","affiliation":[{"name":"Guangdong University of Technology, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,5,17]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"2425","volume-title":"International Conference on Computer Vision","author":"Antol Stanislaw","year":"2015","unstructured":"Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: Visual question answering. In International Conference on Computer Vision, 2425\u20132433."},{"key":"e_1_3_1_3_2","first-page":"104","volume-title":"European Conference on Computer Vision","volume":"12375","author":"Chen Yenchun","year":"2020","unstructured":"Yenchun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. UNITER: Universal image-text representation learning. In European Conference on Computer Vision, Vol. 12375, 104\u2013120."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3579051.3579053"},{"key":"e_1_3_1_5_2","first-page":"19","article-title":"S3-Net: A fast scene understanding network by single-shot segmentation for autonomous driving","volume":"12","author":"Cheng Yuan","year":"2021","unstructured":"Yuan Cheng, Yuchao Yang, Hai-Bao Chen, Ngai Wong, and Hao Yu. 2021. S3-Net: A fast scene understanding network by single-shot segmentation for autonomous driving. ACM Transactions on Intelligent Systems and Technology 12 (2021), 19.","journal-title":"ACM Transactions on Intelligent Systems and Technology"},{"key":"e_1_3_1_6_2","first-page":"565","volume-title":"Proceedings of the VLDB Endowment","volume":"10","author":"Cui Wanyun","year":"2017","unstructured":"Wanyun Cui, Yanghua Xiao, Haixun Wang, Yangqiu Song, Seung-won Hwang, and Wei Wang. 2017. KBQA: Learning question answering over QA corpora and knowledge bases. Proceedings of the VLDB Endowment 10 (2017), 565\u2013576."},{"key":"e_1_3_1_7_2","first-page":"3298","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Dai Bo","year":"2017","unstructured":"Bo Dai, Yuqi Zhang, and Dahua Lin. 2017. Detecting visual relationships with deep relational networks. In IEEE Conference on Computer Vision and Pattern Recognition, 3298\u20133308."},{"key":"e_1_3_1_8_2","first-page":"4271","article-title":"Funnel-transformer: Filtering out sequential redundancy for efficient language processing","volume":"33","author":"Dai Zihang","year":"2020","unstructured":"Zihang Dai, Guokun Lai, Yiming Yang, and Quoc Le. 2020. Funnel-transformer: Filtering out sequential redundancy for efficient language processing. Advances in Neural Information Processing Systems 33 (2020), 4271\u20134282.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_9_2","first-page":"4171","volume-title":"The Annual Conference of the North American Chapter of the Association for Computational Linguistics","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In The Annual Conference of the North American Chapter of the Association for Computational Linguistics, 4171\u20134186."},{"key":"e_1_3_1_10_2","first-page":"5057","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Gao Feng","year":"2022","unstructured":"Feng Gao, Qing Ping, Govind Thattai, Aishwarya N. Reganti, Ying Nian Wu, and Prem Natarajan. 2022. Transform-Retrieve-Generate: Natural language-centric outside-knowledge visual question answering. In IEEE Conference on Computer Vision and Pattern Recognition, 5057\u20135067."},{"key":"e_1_3_1_11_2","first-page":"6894","volume-title":"Conference on Empirical Methods in Natural Language Processing","author":"Gao Tianyu","year":"2021","unstructured":"Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. In Conference on Empirical Methods in Natural Language Processing, 6894\u20136910."},{"key":"e_1_3_1_12_2","first-page":"489","volume-title":"Conference on Empirical Methods in Natural Language Processing","author":"Gard\u00e8res Fran\u00e7ois","year":"2020","unstructured":"Fran\u00e7ois Gard\u00e8res, Maryam Ziaeefard, Baptiste Abeloos, and Freddy L\u00e9cu\u00e9. 2020. ConceptBert: Concept-aware representation for visual question answering. In Conference on Empirical Methods in Natural Language Processing, 489\u2013498."},{"key":"e_1_3_1_13_2","first-page":"2061","volume-title":"ACM International Conference on Multimedia","author":"Guo Yangyang","year":"2022","unstructured":"Yangyang Guo, Liqiang Nie, Yongkang Wong, Yibing Liu, Zhiyong Cheng, and Mohan S. Kankanhalli. 2022. A unified end-to-end retriever-reader framework for knowledge-based VQA. In ACM International Conference on Multimedia, 2061\u20132069."},{"key":"e_1_3_1_14_2","first-page":"553","volume-title":"ACM International Conference on Web Search and Data Mining","author":"He Gaole","year":"2021","unstructured":"Gaole He, Yunshi Lan, Jing Jiang, Wayne Xin Zhao, and Ji-Rong Wen. 2021. Improving multi-hop knowledge base question answering by learning intermediate supervision signals. In ACM International Conference on Web Search and Data Mining, 553\u2013561."},{"key":"e_1_3_1_15_2","first-page":"373","article-title":"Hypergraph Transformer: Weakly-supervised multi-hop reasoning for knowledge-based visual question answering","author":"Heo Yu-Jung","year":"2022","unstructured":"Yu-Jung Heo, Eun-Sol Kim, Woo Suk Choi, and Byoung-Tak Zhang. 2022. Hypergraph Transformer: Weakly-supervised multi-hop reasoning for knowledge-based visual question answering. In Annual Meeting of the Association for Computational Linguistics, 373\u2013390.","journal-title":"Annual Meeting of the Association for Computational Linguistics"},{"key":"e_1_3_1_16_2","first-page":"6384","volume-title":"AAAI Conference on Artificial Intelligence","author":"Hwang Jena D.","year":"2021","unstructured":"Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, and Yejin Choi. 2021. COMET-ATOMIC 2020: On symbolic and neural commonsense knowledge graphs. In AAAI Conference on Artificial Intelligence, 6384\u20136392."},{"key":"e_1_3_1_17_2","first-page":"662","volume-title":"European Conference on Computer Vision","volume":"13696","author":"Kamath Amita","year":"2022","unstructured":"Amita Kamath, Christopher Clark, Tanmay Gupta, Eric Kolve, Derek Hoiem, and Aniruddha Kembhavi. 2022. Webly supervised concept expansion for general purpose vision models. In European Conference on Computer Vision, Vol. 13696, 662\u2013681."},{"key":"e_1_3_1_18_2","first-page":"5583","volume-title":"International Conference on Machine Learning","volume":"139","author":"Kim Wonjae","year":"2021","unstructured":"Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. ViLT: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning, Vol. 139, 5583\u20135594."},{"key":"e_1_3_1_19_2","first-page":"12888","volume-title":"International Conference on Machine Learning","volume":"162","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. 2022. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, Vol. 162, 12888\u201312900."},{"key":"e_1_3_1_20_2","first-page":"9694","volume-title":"Annual Conference on Neural Information Processing Systems","author":"Li Junnan","year":"2021","unstructured":"Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven Chu-Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. In Annual Conference on Neural Information Processing Systems, 9694\u20139705."},{"key":"e_1_3_1_21_2","first-page":"2592","article-title":"UNIMO: Towards unified-modal understanding and generation via cross-modal contrastive learning","author":"Li Wei","year":"2021","unstructured":"Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. 2021. UNIMO: Towards unified-modal understanding and generation via cross-modal contrastive learning. In Annual Meeting of the Association for Computational Linguistics, 2592\u20132607.","journal-title":"Annual Meeting of the Association for Computational Linguistics"},{"key":"e_1_3_1_22_2","first-page":"512","article-title":"Commonsense knowledge base completion","author":"Li Xiang","year":"2016","unstructured":"Xiang Li, Aynaz Taheri, Lifu Tu, and Kevin Gimpel. 2016. Commonsense knowledge base completion. In Annual Meeting of the Association for Computational Linguistics, 512\u2013524.","journal-title":"Annual Meeting of the Association for Computational Linguistics"},{"key":"e_1_3_1_23_2","first-page":"121","volume-title":"European Conference on Computer Vision","volume":"12375","author":"Li Xiujun","year":"2020","unstructured":"Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. 2020. Oscar: Object-semantics aligned pre-training for vision-language tasks. In European Conference on Computer Vision, Vol. 12375, 121\u2013137."},{"key":"e_1_3_1_24_2","first-page":"7244","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Li Yikang","year":"2017","unstructured":"Yikang Li, Wanli Ouyang, Xiaogang Wang, and Xiaoou Tang. 2017. ViP-CNN: Visual phrase guided convolutional neural network. In IEEE Conference on Computer Vision and Pattern Recognition, 7244\u20137253."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_1_26_2","first-page":"11238","volume-title":"Conference on Empirical Methods in Natural Language Processing","author":"Lin Weizhe","year":"2022","unstructured":"Weizhe Lin and Bill Byrne. 2022. Retrieval augmented visual question answering with outside knowledge. In Conference on Empirical Methods in Natural Language Processing, 11238\u201311254."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2021.107650"},{"key":"e_1_3_1_28_2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_1_29_2","first-page":"852","volume-title":"European Conference on Computer Vision","volume":"9905","author":"Lu Cewu","year":"2016","unstructured":"Cewu Lu, Ranjay Krishna, Michael S. Bernstein, and Li Fei-Fei. 2016. Visual relationship detection with language priors. In European Conference on Computer Vision, Vol. 9905, 852\u2013869."},{"key":"e_1_3_1_30_2","first-page":"13","volume-title":"Annual Conference on Neural Information Processing Systems 2019","author":"Lu Jiasen","year":"2019","unstructured":"Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Annual Conference on Neural Information Processing Systems 2019, 13\u201323."},{"key":"e_1_3_1_31_2","first-page":"744","volume-title":"International Conference on Learning Representations","author":"Lu Jiasen","year":"2023","unstructured":"Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi. 2023. UNIFIED-IO: A unified model for vision, language, and multi-modal tasks. In International Conference on Learning Representations, 744\u2013751."},{"key":"e_1_3_1_32_2","first-page":"14111","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Marino Kenneth","year":"2021","unstructured":"Kenneth Marino, Xinlei Chen, Devi Parikh, Abhinav Gupta, and Marcus Rohrbach. 2021. KRISP: Integrating implicit and symbolic knowledge for open-domain knowledge-based VQA. In IEEE Conference on Computer Vision and Pattern Recognition, 14111\u201314121."},{"key":"e_1_3_1_33_2","first-page":"3195","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Marino Kenneth","year":"2019","unstructured":"Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019. OK-VQA: A visual question answering benchmark requiring external knowledge. In IEEE Conference on Computer Vision and Pattern Recognition, 3195\u20133204."},{"key":"e_1_3_1_34_2","first-page":"6097","article-title":"Multi-hop reading comprehension through question decomposition and rescoring","author":"Min Sewon","year":"2019","unstructured":"Sewon Min, Victor Zhong, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2019. Multi-hop reading comprehension through question decomposition and rescoring. In Annual Meeting of the Association for Computational Linguistics, 6097\u20136109.","journal-title":"Annual Meeting of the Association for Computational Linguistics"},{"key":"e_1_3_1_35_2","first-page":"2156","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Nam Hyeonseob","year":"2017","unstructured":"Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim. 2017. Dual attention networks for multimodal reasoning and matching. In IEEE Conference on Computer Vision and Pattern Recognition, 2156\u20132164."},{"key":"e_1_3_1_36_2","first-page":"3353","volume-title":"International Conference on Information and Knowledge Management","author":"Nuthalapati Sai Vidyaranya","year":"2021","unstructured":"Sai Vidyaranya Nuthalapati, Ramraj Chandradevan, Eleonora Giunchiglia, Bowen Li, Maxime Kayser, Thomas Lukasiewicz, and Carl Yang. 2021. Lightweight visual question answering using scene graphs. In International Conference on Information and Knowledge Management, 3353\u20133357."},{"key":"e_1_3_1_37_2","first-page":"8748","volume-title":"International Conference on Machine Learning","volume":"139","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, Vol. 139, 8748\u20138763."},{"key":"e_1_3_1_38_2","first-page":"1155","volume-title":"IEEE Winter Conference on Applications of Computer Vision","author":"Ravi Sahithya","year":"2023","unstructured":"Sahithya Ravi, Aditya Chinchure, Leonid Sigal, Renjie Liao, and Vered Shwartz. 2023. VLC-BERT: Visual question answering with contextualized commonsense knowledge. In IEEE Winter Conference on Applications of Computer Vision, 1155\u20131165."},{"key":"e_1_3_1_39_2","first-page":"3980","volume-title":"Conference on Empirical Methods in Natural Language Processing","author":"Reimers Nils","year":"2019","unstructured":"Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using siamese BERT-networks. In Conference on Empirical Methods in Natural Language Processing, 3980\u20133990."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2022.118669"},{"key":"e_1_3_1_42_2","first-page":"146","volume-title":"European Conference on Computer Vision","volume":"13668","author":"Schwenk Dustin","year":"2022","unstructured":"Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. 2022. A-OKVQA: A benchmark for visual question answering using world knowledge. In European Conference on Computer Vision, Vol. 13668, 146\u2013162."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-021-11276-2"},{"key":"e_1_3_1_44_2","first-page":"4444","volume-title":"AAAI Conference on Artificial Intelligence","author":"Speer Robyn","year":"2017","unstructured":"Robyn Speer, Joshua Chin, and Catherine Havasi. 2017. ConceptNet 5.5: An open multilingual graph of general knowledge. In AAAI Conference on Artificial Intelligence, 4444\u20134451."},{"key":"e_1_3_1_45_2","first-page":"21","volume-title":"International Conference on Learning Representations","author":"Su Weijie","year":"2020","unstructured":"Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020. VL-BERT: Pre-training of generic visual-linguistic representations. In International Conference on Learning Representations, 21\u201334."},{"key":"e_1_3_1_46_2","first-page":"5099","volume-title":"Conference on Empirical Methods in Natural Language Processing","author":"Tan Hao","year":"2019","unstructured":"Hao Tan and Mohit Bansal. 2019. LXMERT: Learning cross-modality encoder representations from transformers. In Conference on Empirical Methods in Natural Language Processing, 5099\u20135110."},{"key":"e_1_3_1_47_2","first-page":"3713","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Tang Kaihua","year":"2020","unstructured":"Kaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi, and Hanwang Zhang. 2020. Unbiased scene graph generation from biased training. In IEEE Conference on Computer Vision and Pattern Recognition, 3713\u20133722."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/2629489"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3151957"},{"key":"e_1_3_1_50_2","unstructured":"Yanan Wang Michihiro Yasunaga Hongyu Ren Shinya Wada and Jure Leskovec. 2022. VQA-GNN: Reasoning with multimodal semantic graph for visual question answering. arXiv:2205.11501. Retrieved from https:\/\/arxiv.org\/abs\/2205.11501"},{"key":"e_1_3_1_51_2","first-page":"2712","volume-title":"AAAI Conference on Artificial Intelligence","author":"Wu Jialin","year":"2022","unstructured":"Jialin Wu, Jiasen Lu, Ashish Sabharwal, and Roozbeh Mottaghi. 2022. Multi-modal answer validation for knowledge-based VQA. In AAAI Conference on Artificial Intelligence, 2712\u20132721."},{"key":"e_1_3_1_52_2","unstructured":"Jialian Wu Jianfeng Wang Zhengyuan Yang Zhe Gan Zicheng Liu Junsong Yuan and Lijuan Wang. 2022. GRiT: A generative region-to-text transformer for object understanding. arXiv:2212.00280. Retrieved from https:\/\/arxiv.org\/abs\/2212.00280"},{"key":"e_1_3_1_53_2","first-page":"3097","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Xu Danfei","year":"2017","unstructured":"Danfei Xu, Yuke Zhu, Christopher B. Choy, and Li Fei-Fei. 2017. Scene graph generation by iterative message passing. In IEEE Conference on Computer Vision and Pattern Recognition, 3097\u20133106."},{"key":"e_1_3_1_54_2","first-page":"3081","volume-title":"AAAI Conference on Artificial Intelligence","author":"Yang Zhengyuan","year":"2022","unstructured":"Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Yumao Lu, Zicheng Liu, and Lijuan Wang. 2022. An empirical study of GPT-3 for few-shot knowledge-based VQA. In AAAI Conference on Artificial Intelligence, 3081\u20133089."},{"key":"e_1_3_1_55_2","unstructured":"Liang Yao Chengsheng Mao and Yuan Luo. 2019. KG-BERT: BERT for knowledge graph completion. arXiv:1909.03193. Retrieved from https:\/\/arxiv.org\/abs\/1909.03193"},{"key":"e_1_3_1_56_2","first-page":"535","volume-title":"The Annual Conference of the North American Chapter of the Association for Computational Linguistics","author":"Yasunaga Michihiro","year":"2021","unstructured":"Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec. 2021. QA-GNN: Reasoning with language models and knowledge graphs for question answering. In The Annual Conference of the North American Chapter of the Association for Computational Linguistics, 535\u2013546."},{"key":"e_1_3_1_57_2","first-page":"3208","volume-title":"AAAI Conference on Artificial Intelligence","author":"Yu Fei","year":"2021","unstructured":"Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2021. ERNIE-ViL: Knowledge enhanced vision-language representations through scene graphs. In AAAI Conference on Artificial Intelligence, 3208\u20133216."},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3616542"},{"key":"e_1_3_1_59_2","first-page":"13041","volume-title":"AAAI Conference on Artificial Intelligence","author":"Zhou Luowei","year":"2020","unstructured":"Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao. 2020. Unified vision-language pre-training for image captioning and VQA. In AAAI Conference on Artificial Intelligence, 13041\u201313049."},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3495161"},{"key":"e_1_3_1_61_2","first-page":"1097","volume-title":"International Joint Conferences on Artificial Intelligence","author":"Zhu Zihao","year":"2020","unstructured":"Zihao Zhu, Jing Yu, Yujing Wang, Yajing Sun, Yue Hu, and Qi Wu. 2020. Mucko: Multi-layer cross-modal knowledge reasoning for fact-based visual question answering. In International Joint Conferences on Artificial Intelligence, 1097\u20131103."},{"key":"e_1_3_1_62_2","unstructured":"Mikolov Tomas Kai Chen Greg Corrado and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv:1301.3781. Retrieved from https:\/\/arxiv.org\/abs\/arXiv:1301.3781"},{"key":"e_1_3_1_63_2","unstructured":"\u0158eh\u016f\u0159ek Radim and Petr Sojka. 2011. Gensim statistical semantics in python. Retrieved from http:\/\/nlp.fi.muni.cz\/projekty\/gensim\/"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3724125","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3724125","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:59Z","timestamp":1750295939000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3724125"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,17]]},"references-count":62,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,6,30]]}},"alternative-id":["10.1145\/3724125"],"URL":"https:\/\/doi.org\/10.1145\/3724125","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,17]]},"assertion":[{"value":"2023-11-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-21","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-17","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}