{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T03:41:52Z","timestamp":1784691712871,"version":"3.55.0"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2025,3,10]],"date-time":"2025-03-10T00:00:00Z","timestamp":1741564800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Central Government Guided Local Science and Technology Development","award":["YDZX2022028"],"award-info":[{"award-number":["YDZX2022028"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62172376, U22A2068"],"award-info":[{"award-number":["62172376, U22A2068"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,4,30]]},"abstract":"<jats:p>Visual question answering (VQA) is a challenging task that requires models to understand both visual and linguistic inputs and produce accurate answers. However, VQA models often exploit biases in datasets to make predictions rather than reasoning based on the inputs. Prior approaches to debiasing have suggested the implementation of a supplementary model, deliberately designed to exhibit bias, which subsequently informs the training of a resilient target model. Nevertheless, such techniques merely quantify the model\u2019s divergence based on the statistical distribution of labels within the training dataset or in relation to unimodal branches. In this work, we propose a novel method of generating bias from the target model itself, called LEGO, which aims to combat the language guidance bias. Specifically, LEGO framework employs a generative network that assimilates the biases inherent in the target model by integrating adversarial goals with the principles of knowledge distillation. Then, we use a debiased contrastive learning strategy to model the language guidance bias of caption and question. In the process of modeling, in order to obtain robust semantic coreference, the multimodal representations of two semantic granularity are modeled by mutual information fusion and contrast learning difference modeling. We evaluate our method on various VQA-biased datasets, including VQA-CP2, GQA-OOD, and RSICD, and show that it outperforms similar methods.<\/jats:p>","DOI":"10.1145\/3715141","type":"journal-article","created":{"date-parts":[[2025,1,30]],"date-time":"2025-01-30T15:42:16Z","timestamp":1738251736000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Language-guided Bias Generation Contrastive Strategy for Visual Question Answering"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2909-4317","authenticated-orcid":false,"given":"Enyuan","family":"Zhao","sequence":"first","affiliation":[{"name":"Ocean University of China, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4209-7387","authenticated-orcid":false,"given":"Ning","family":"Song","sequence":"additional","affiliation":[{"name":"Ocean University of China, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5587-3863","authenticated-orcid":false,"given":"Ze","family":"Zhang","sequence":"additional","affiliation":[{"name":"Ocean University of China, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4952-7666","authenticated-orcid":false,"given":"Jie","family":"Nie","sequence":"additional","affiliation":[{"name":"Ocean University of China, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4406-536X","authenticated-orcid":false,"given":"Xinyue","family":"Liang","sequence":"additional","affiliation":[{"name":"Ocean University of China, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6965-0162","authenticated-orcid":false,"given":"Zhiqiang","family":"Wei","sequence":"additional","affiliation":[{"name":"Ocean University of China, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,3,10]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"4971","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Agrawal Aishwarya","year":"2018","unstructured":"Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. 2018. Don\u2019t just assume; look and answer: Overcoming priors for visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4971\u20134980."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_2_4_2","first-page":"2425","volume-title":"Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV)","author":"Antol Stanislaw","year":"2015","unstructured":"Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: Visual question answering. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), 2425\u20132433."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.5555\/3524938.3524988"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3590773"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3382684"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3291379"},{"key":"e_1_3_2_9_2","first-page":"841","volume-title":"Proceedings of the 33rd Conference on Neural Information Processing Systems (NeurIPS)","author":"Cadene Remi","year":"2019","unstructured":"Remi Cadene, Corentin Dancette, Hedi Ben-younes, Matthieu Cord, and Devi Parikh. 2019. RUBi: Reducing unimodal biases in visual question answering. In Proceedings of the 33rd Conference on Neural Information Processing Systems (NeurIPS), 841\u2013852."},{"key":"e_1_3_2_10_2","first-page":"10800","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Chen Long","year":"2020","unstructured":"Long Chen, Xin Yan, Jun Xiao, Hanwang Zhang, Shiliang Pu, and Yueting Zhuang. 2020. Counterfactual samples synthesizing for robust visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10800\u201310809."},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Long Chen Yuhang Zheng and Jun Xiao. 2022. Rethinking data augmentation for robust visual question answering. arXiv:2207.08739. Retrieved from https:\/\/arxiv.org\/abs\/2207.08739","DOI":"10.1007\/978-3-031-20059-5_6"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.5555\/3524938.3525087"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3318220"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01124"},{"key":"e_1_3_2_15_2","first-page":"4069","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Clark Christopher","year":"2019","unstructured":"Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. 2019. Don\u2019t take the easy way out: Ensemble based methods for avoiding known dataset biases. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP), 4069\u20134082."},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","unstructured":"Hongchao Fang and Pengtao Xie. 2020. CERT: Contrastive self-supervised learning for language understanding. arXiv:2005.12766. Retrieved from https:\/\/arxiv.org\/abs\/2005.12766","DOI":"10.36227\/techrxiv.12308378.v1"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","first-page":"878","DOI":"10.18653\/v1\/2020.emnlp-main.63","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Gokhale Tejas","year":"2020","unstructured":"Tejas Gokhale, Pratyay Banerjee, Chitta Baral, and Yezhou Yang. 2020. MUTANT: A training paradigm for out-of-distribution generalization in visual question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 878\u2013892."},{"key":"e_1_3_2_18_2","first-page":"2672","volume-title":"Proceedings of the Conference on Neural Information Processing Systems (NeurIPS)","author":"Goodfellow Ian","year":"2014","unstructured":"Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS), 2672\u20132680."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3128322"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00161"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3240337"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3673902"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_2_24_2","unstructured":"R. Devon Hjelm Alex Fedorov Samuel Lavoie-Marchildon Karan Grewal Phil Bachman Adam Trischler and Yoshua Bengio. 2018. Learning deep representations by mutual information estimation and maximization. arXiv:1808.06670. Retrieved from https:\/\/arxiv.org\/abs\/1808.06670"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6776"},{"key":"e_1_3_2_26_2","first-page":"18","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920)","author":"V. Gouthaman K.","year":"2020","unstructured":"Gouthaman K. V. and Anurag Mittal. 2020. Reducing language biases in visual question answering with visually-grounded question encoder. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920). Springer, 18\u201334."},{"key":"e_1_3_2_27_2","first-page":"9694","article-title":"Align before fuse: Vision and language representation learning with momentum distillation","volume":"34","author":"Li Junnan","year":"2021","unstructured":"Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 34, 9694\u20139705.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3264524"},{"key":"e_1_3_2_29_2","first-page":"18545","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"38","author":"Li Peize","year":"2024","unstructured":"Peize Li, Qingyi Si, Peng Fu, Zheng Lin, and Yan Wang. 2024. Object attribute matters in visual question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 18545\u201318553."},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3300938"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3489142"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.311"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.265"},{"key":"e_1_3_2_34_2","first-page":"22820","article-title":"Fine-grained late-interaction multi-modal retrieval for retrieval augmented visual question answering","volume":"36","author":"Lin Weizhe","year":"2024","unstructured":"Weizhe Lin, Jinghong Chen, Jingbiao Mei, Alexandru Coca, and Bill Byrne. 2024. Fine-grained late-interaction multi-modal retrieval for retrieval augmented visual question answering. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 36, 22820\u201322840.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-short.135"},{"key":"e_1_3_2_36_2","first-page":"12700","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Niu Yulei","year":"2021","unstructured":"Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji-Rong Wen. 2021. Counterfactual VQA: A cause-effect look at language bias. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12700\u201312710."},{"key":"e_1_3_2_37_2","first-page":"16292","article-title":"Introspective distillation for robust question answering","volume":"34","author":"Niu Yulei","year":"2021","unstructured":"Yulei Niu and Hanwang Zhang. 2021. Introspective distillation for robust question answering. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 34, 16292\u201316304.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i5.28253"},{"key":"e_1_3_2_39_2","first-page":"1548","article-title":"Overcoming language priors in visual question answering with adversarial regularization","author":"Ramakrishnan Sainandan","year":"2018","unstructured":"Sainandan Ramakrishnan, Aishwarya Agrawal, and Stefan Lee. 2018. Overcoming language priors in visual question answering with adversarial regularization. In Proceedings of the 32nd Conference on Neural Information Processing Systems (NeurIPS), 1548\u20131558.","journal-title":"Proceedings of the 32nd Conference on Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_3_2_40_2","doi-asserted-by":"crossref","first-page":"2591","DOI":"10.1109\/ICCV.2019.00268","article-title":"Taking a hint: Leveraging explanations to make vision and language models more grounded","author":"Selvaraju Ramprasaath R.","year":"2019","unstructured":"Ramprasaath R. Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Shalini Ghosh, Larry Heck, Dhruv Batra, and Devi Parikh. 2019. Taking a hint: Leveraging explanations to make vision and language models more grounded. In Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), 2591\u20132600.","journal-title":"Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV)"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01438"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.359"},{"key":"e_1_3_2_43_2","unstructured":"Damien Teney Kushal Kafle Robik Shrestha Ehsan Abbasnejad Christopher Kanan and Anton van den Hengel. 2020. On the value of out-of-distribution testing: An example of Goodhart\u2019s law. arXiv:2005.09241. Retrieved from https:\/\/arxiv.org\/abs\/2005.09241"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3618301"},{"key":"e_1_3_2_45_2","first-page":"3784","article-title":"Debiased visual question answering from feature and sample perspectives","volume":"34","author":"Wen Zhiquan","year":"2021","unstructured":"Zhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu, and Qi Wu. 2021. Debiased visual question answering from feature and sample perspectives. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 34, 3784\u20133796.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3455059"},{"key":"e_1_3_2_47_2","unstructured":"Yike Wu Yu Zhao Shiwan Zhao Ying Zhang Xiaojie Yuan Guoqing Zhao and Ning Jiang. 2022. Overcoming language priors in visual question answering via distinguishing superficially similar instances. arXiv:2209.08529. Retrieved from https:\/\/arxiv.org\/abs\/2209.08529"},{"key":"e_1_3_2_48_2","first-page":"21","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Yang Zichao","year":"2016","unstructured":"Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. 2016. Stacked attention networks for image question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 21\u201329."},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3316767"},{"key":"e_1_3_2_50_2","doi-asserted-by":"crossref","unstructured":"Dejiao Zhang Feng Nan Xiaokai Wei Shang-Wen Li Henghui Zhu Kathleen R. McKeown Ramesh Nallapati Andrew O. Arnold and Bing Xiang. 2021. Supporting clustering with contrastive learning. arXiv:2103.12953. Retrieved from https:\/\/arxiv.org\/abs\/2103.12953","DOI":"10.18653\/v1\/2021.naacl-main.427"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i7.28543"},{"key":"e_1_3_2_52_2","first-page":"1083","volume-title":"Proceedings of the 29th International Joint Conference on Artificial Intelligence (IJCAI)","author":"Zhu Xi","year":"2021","unstructured":"Xi Zhu, Zhendong Mao, Chunxiao Liu, Peng Zhang, Bin Wang, and Yongdong Zhang. 2021. Overcoming language priors with self-supervised learning for visual question answering. In Proceedings of the 29th International Joint Conference on Artificial Intelligence (IJCAI), 1083\u20131089."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-acl.190"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715141","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3715141","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:18Z","timestamp":1750295898000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715141"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,10]]},"references-count":52,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,4,30]]}},"alternative-id":["10.1145\/3715141"],"URL":"https:\/\/doi.org\/10.1145\/3715141","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,10]]},"assertion":[{"value":"2024-07-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}