{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T05:02:11Z","timestamp":1750309331840,"version":"3.41.0"},"reference-count":70,"publisher":"Association for Computing Machinery (ACM)","issue":"12","license":[{"start":{"date-parts":[[2024,11,21]],"date-time":"2024-11-21T00:00:00Z","timestamp":1732147200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Key Research Platforms and Projects of the Guangdong Provincial Department of Education","award":["2023ZDZX1034"],"award-info":[{"award-number":["2023ZDZX1034"]}]},{"name":"Natural Science Foundation of Shenzhen","award":["JCYJ20230807142703006"],"award-info":[{"award-number":["JCYJ20230807142703006"]}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"crossref","award":["2023M743003"],"award-info":[{"award-number":["2023M743003"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,12,31]]},"abstract":"<jats:p>Language prior is a major block to improving the generalization of visual question answering (VQA) models. Recent work has revealed that synthesizing extra training samples to balance training sets is a promising way to alleviate language priors. However, most existing methods synthesize extra samples in a manner independent of training processes, which neglect the fact that the language priors memorized by VQA models are changing during training, resulting in insufficient synthesized samples. In this article, we propose an adversarial sample synthesis method, which synthesizes different adversarial samples by adversarial masking at different training epochs to cope with the changing memorized language priors. The basic idea behind our method is to use adversarial masking to synthesize adversarial samples that will cause the model to make wrong answers. To this end, we design a generative module to carry out adversarial masking by attacking the VQA model and introduce a bias-oriented objective to supervise the training of the generative module. We couple the sample synthesis with the training process of the VQA model, which ensures that the synthesized samples at different training epochs are beneficial to the VQA model. We incorporated the proposed method into three VQA models including UpDn, LMH, and LXMERT and conducted experiments on three datasets including VQA-CP v1, VQA-CP v2, and VQA v2. Experimental results demonstrate that a large improvement of our method, such as 16.22% gains on LXMERT in the overall accuracy of VQA-CP v2.<\/jats:p>","DOI":"10.1145\/3688848","type":"journal-article","created":{"date-parts":[[2024,9,16]],"date-time":"2024-09-16T14:59:06Z","timestamp":1726498746000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Adversarial Sample Synthesis for Visual Question Answering"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5769-3739","authenticated-orcid":false,"given":"Chuanhao","family":"Li","sequence":"first","affiliation":[{"name":"Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science &amp; Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-5109-2386","authenticated-orcid":false,"given":"Chenchen","family":"Jing","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China and Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science &amp; Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-6422-6056","authenticated-orcid":false,"given":"Zhen","family":"Li","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science &amp; Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6300-6336","authenticated-orcid":false,"given":"Yuwei","family":"Wu","sequence":"additional","affiliation":[{"name":"Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University, Shenzhen, China and Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science &amp; Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1900-8945","authenticated-orcid":false,"given":"Yunde","family":"Jia","sequence":"additional","affiliation":[{"name":"Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University, Shenzhen, China and Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science &amp; Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,11,21]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"9690","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201920)","author":"Agarwal Vedika","year":"2020","unstructured":"Vedika Agarwal, Rakshith Shetty, and Mario Fritz. 2020. Towards causal VQA: Revealing and reducing spurious correlations by invariant and covariant semantic editing. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201920), 9690\u20139698."},{"key":"e_1_3_2_3_2","first-page":"4971","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201918)","author":"Agrawal Aishwarya","year":"2018","unstructured":"Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. 2018. Don\u2019t just assume; look and answer: Overcoming priors for visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201918), 4971\u20134980."},{"key":"e_1_3_2_4_2","first-page":"6077","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201918)","author":"Anderson Peter","year":"2018","unstructured":"Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201918), 6077\u20136086."},{"key":"e_1_3_2_5_2","first-page":"39","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201916)","author":"Andreas Jacob","year":"2016","unstructured":"Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016. Neural module networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201916), 39\u201348."},{"key":"e_1_3_2_6_2","first-page":"2425","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201915)","author":"Antol Stanislaw","year":"2015","unstructured":"Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: Visual question answering. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201915), 2425\u20132433."},{"key":"e_1_3_2_7_2","first-page":"11671","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201923)","author":"Basu Abhipsa","year":"2023","unstructured":"Abhipsa Basu, Sravanti Addepalli, and R. Venkatesh Babu. 2023. RMLVQA: A margin loss approach for visual question answering with language biases. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201923), 11671\u201311680."},{"unstructured":"Remi Cadene Corentin Dancette Hedi Ben-younes Matthieu Cord and Devi Parikh. 2019. RUBi: Reducing unimodal biases for visual question answering. In Advances in Neural Information Processing Systems (NeurIPS \u201919) Vol. 32 841\u2013852.","key":"e_1_3_2_8_2"},{"doi-asserted-by":"publisher","unstructured":"Jianjian Cao Xiameng Qin Sanyuan Zhao and Jianbing Shen. 2022. Bilateral cross-modality graph matching attention for feature fusion in visual question answering. IEEE Transactions on Neural Networks and Learning Systems (2022) 1\u201312. DOI: 10.1109\/TNNLS.2021.3135655","key":"e_1_3_2_9_2","DOI":"10.1109\/TNNLS.2021.3135655"},{"key":"e_1_3_2_10_2","first-page":"10800","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201920)","author":"Chen Long","year":"2020","unstructured":"Long Chen, Xin Yan, Jun Xiao, Hanwang Zhang, Shiliang Pu, and Yueting Zhuang. 2020. Counterfactual samples synthesizing for robust visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201920), 10800\u201310809."},{"key":"e_1_3_2_11_2","first-page":"11681","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Cho Jae Won","year":"2023","unstructured":"Jae Won Cho, Dong-Jin Kim, Hyeonggon Ryu, and In So Kweon. 2023. Generative bias for robust visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11681\u201311690."},{"doi-asserted-by":"publisher","unstructured":"Kyunghyun Cho Bart Van Merri\u00ebnboer Caglar Gulcehre Dzmitry Bahdanau Fethi Bougares Holger Schwenk and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv:1406.1078. Retrieved from 10.48550\/arXiv.1406.1078","key":"e_1_3_2_12_2","DOI":"10.48550\/arXiv.1406.1078"},{"key":"e_1_3_2_13_2","first-page":"4069","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing (EMNLP-IJCNLP \u201919)","author":"Clark Christopher","year":"2019","unstructured":"Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. 2019. Don\u2019t take the easy way out: Ensemble based methods for avoiding known dataset biases. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing (EMNLP-IJCNLP \u201919), 4069\u20134082."},{"key":"e_1_3_2_14_2","first-page":"4109","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI \u201923)","author":"Peng Xian Ling Mao, Yuanyuan Fu, Dangyang Chen, Daowan","year":"2023","unstructured":"Xian Ling Mao, Yuanyuan Fu, Dangyang Chen, Daowan Peng, and Wei Wei. 2023. An empirical study on the language modal in visual question answering. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI \u201923), 4109\u20134117."},{"unstructured":"Zhe Gan Yen-Chun Chen Linjie Li Chen Zhu Yu Cheng and Jingjing Liu. 2020. Large-scale adversarial training for vision-and-language representation learning. In Advances in Neural Information Processing Systems (NeurIPS \u201920) Vol. 33 6616\u20136628.","key":"e_1_3_2_15_2"},{"doi-asserted-by":"publisher","key":"e_1_3_2_16_2","DOI":"10.18653\/v1\/2020.emnlp-main.63"},{"key":"e_1_3_2_17_2","first-page":"6904","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201917)","author":"Goyal Yash","year":"2017","unstructured":"Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201917), 6904\u20136913."},{"doi-asserted-by":"publisher","key":"e_1_3_2_18_2","DOI":"10.1109\/TIP.2021.3097180"},{"doi-asserted-by":"publisher","key":"e_1_3_2_19_2","DOI":"10.1007\/s10489-022-03559-4"},{"key":"e_1_3_2_20_2","first-page":"3608","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201918)","author":"Gurari Danna","year":"2018","unstructured":"Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham. 2018. Vizwiz grand challenge: Answering visual questions from blind people. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201918), 3608\u20133617."},{"key":"e_1_3_2_21_2","first-page":"1584","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201921)","author":"Han Xinzhe","year":"2021","unstructured":"Xinzhe Han, Shuhui Wang, Chi Su, Qingming Huang, and Qi Tian. 2021. Greedy gradient ensemble for robust visual question answering. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201921), 1584\u20131593."},{"key":"e_1_3_2_22_2","first-page":"441","volume-title":"Proceedings of the International Conference on Robotics and Automation Engineering (ICRAE \u201917)","author":"He Bin","year":"2017","unstructured":"Bin He, Meng Xia, Xinguo Yu, Pengpeng Jian, Hao Meng, and Zhanwen Chen. 2017. An educational robot system of visual question answering for preschoolers. In Proceedings of the International Conference on Robotics and Automation Engineering (ICRAE \u201917). IEEE, 441\u2013445."},{"key":"e_1_3_2_23_2","first-page":"53","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201918)","author":"Hu Ronghang","year":"2018","unstructured":"Ronghang Hu, Jacob Andreas, Trevor Darrell, and Kate Saenko. 2018. Explainable neural computation via stack neural module networks. In Proceedings of the European Conference on Computer Vision (ECCV \u201918), 53\u201369."},{"key":"e_1_3_2_24_2","first-page":"804","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201917)","author":"Hu Ronghang","year":"2017","unstructured":"Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko. 2017. Learning to reason: End-to-end module networks for visual question answering. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201917), 804\u2013813."},{"key":"e_1_3_2_25_2","first-page":"10294","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201919)","author":"Hu Ronghang","year":"2019","unstructured":"Ronghang Hu, Anna Rohrbach, Trevor Darrell, and Kate Saenko. 2019. Language-conditioned graph networks for relational reasoning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201919), 10294\u201310303."},{"doi-asserted-by":"publisher","unstructured":"Drew A. Hudson and Christopher D. Manning. 2018. Compositional attention networks for machine reasoning. arXiv:1803.03067. Retrieved from 10.48550\/arXiv.1803.03067","key":"e_1_3_2_26_2","DOI":"10.48550\/arXiv.1803.03067"},{"key":"e_1_3_2_27_2","first-page":"1122","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI \u201922)","volume":"2","author":"Jing Chenchen","year":"2022","unstructured":"Chenchen Jing, Yunde Jia, Yuwei Wu, Chuanhao Li, and Qi Wu. 2022. Learning the dynamics of visual relational reasoning via reinforced path routing. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI \u201922), Vol. 2, 1122\u20131130."},{"key":"e_1_3_2_28_2","first-page":"11181","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI \u201920)","volume":"34","author":"Jing Chenchen","year":"2020","unstructured":"Chenchen Jing, Yuwei Wu, Xiaoxun Zhang, Yunde Jia, and Qi Wu. 2020. Overcoming language priors in vqa via decomposed linguistic representations. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI \u201920), Vol. 34, 11181\u201311188."},{"key":"e_1_3_2_29_2","first-page":"2901","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201917)","author":"Johnson Justin","year":"2017","unstructured":"Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201917), 2901\u20132910."},{"key":"e_1_3_2_30_2","first-page":"2989","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201917)","author":"Johnson Justin","year":"2017","unstructured":"Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Judy Hoffman, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017. Inferring and executing programs for visual reasoning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201917), 2989\u20132998."},{"key":"e_1_3_2_31_2","first-page":"2989","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201917)","author":"Johnson Justin","year":"2017","unstructured":"Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Judy Hoffman, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017. Inferring and executing programs for visual reasoning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201917), 2989\u20132998."},{"key":"e_1_3_2_32_2","first-page":"198","volume-title":"Proceedings of the 10th International Conference on Natural Language Generation","author":"Kafle Kushal","year":"2017","unstructured":"Kushal Kafle, Mohammed Yousefhussien, and Christopher Kanan. 2017. Data augmentation for visual question answering. In Proceedings of the 10th International Conference on Natural Language Generation, 198\u2013202."},{"doi-asserted-by":"publisher","key":"e_1_3_2_33_2","DOI":"10.48550\/arXiv.1412.6980"},{"key":"e_1_3_2_34_2","first-page":"3001","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV \u201922)","author":"Kolling Camila","year":"2022","unstructured":"Camila Kolling, Martin More, Nathan Gavenski, Eduardo Pooch, Ot\u00e1vio Parraga, and Rodrigo C. Barros. 2022. Efficient counterfactual debiasing for visual question answering. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV \u201922), 3001\u20133010."},{"key":"e_1_3_2_35_2","first-page":"18","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201920)","author":"K. V. Gouthaman","year":"2020","unstructured":"Gouthaman K. V. and Anurag Mittal. 2020. Reducing language biases in visual question answering with visually-grounded question encoder. In Proceedings of the European Conference on Computer Vision (ECCV \u201920). Springer, 18\u201334."},{"doi-asserted-by":"publisher","unstructured":"Linjie Li Zhe Gan and Jingjing Liu. 2020. A closer look at the robustness of vision-and-language pre-trained models. arXiv:2012.08673. Retrieved from 10.48550\/arXiv.2012.08673","key":"e_1_3_2_36_2","DOI":"10.48550\/arXiv.2012.08673"},{"key":"e_1_3_2_37_2","first-page":"4655","article-title":"Visual question answering with question representation update (QRU)","volume":"29","author":"Li Ruiyu","year":"2016","unstructured":"Ruiyu Li and Jiaya Jia. 2016. Visual question answering with question representation update (QRU). In Advances in Neural Information Processing Systems (NeurIPS \u201916), Vol. 29, 4655\u20134663.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS \u201916)"},{"doi-asserted-by":"publisher","key":"e_1_3_2_38_2","DOI":"10.18653\/v1\/2020.emnlp-main.265"},{"key":"e_1_3_2_39_2","first-page":"149","volume-title":"Proceedings of the European Chapter of the Association for Computational Linguistics (EACL \u201923)","author":"Lin Weizhe","year":"2023","unstructured":"Weizhe Lin, Zhilin Wang, and Bill Byrne. 2023. FVQA 2.0: Introducing adversarial samples into fact-based visual question answering. In Proceedings of the European Chapter of the Association for Computational Linguistics (EACL \u201923), 149\u2013157."},{"doi-asserted-by":"publisher","key":"e_1_3_2_40_2","DOI":"10.1145\/3616399"},{"doi-asserted-by":"publisher","unstructured":"Jie Ma Pinghui Wang Zewei Wang Dechen Kong Min Hu Ting Han and Jun Liu. 2023. Adaptive loose optimization for robust question answering. arXiv:2305.03971. Retrieved from 10.48550\/arXiv.2305.03971","key":"e_1_3_2_41_2","DOI":"10.48550\/arXiv.2305.03971"},{"doi-asserted-by":"crossref","unstructured":"Aihua Mao Feng Chen Ziying Ma and Ken Lin. 2024. Overcoming Language Priors in Visual Question Answering with Cumulative Learning Strategy. Available at SSRN 4740502.","key":"e_1_3_2_42_2","DOI":"10.2139\/ssrn.4740502"},{"key":"e_1_3_2_43_2","first-page":"12700","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201921)","author":"Niu Yulei","year":"2021","unstructured":"Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji-Rong Wen. 2021. Counterfactual VQA: A cause-effect look at language bias. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201921), 12700\u201312710."},{"key":"e_1_3_2_44_2","first-page":"16292","article-title":"Introspective distillation for robust question answering","volume":"34","author":"Niu Yulei","year":"2021","unstructured":"Yulei Niu and Hanwang Zhang. 2021. Introspective distillation for robust question answering. In Advances in Neural Information Processing Systems (NeurIPS \u201921), Vol. 34, 16292\u201316304.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS \u201921)"},{"doi-asserted-by":"publisher","key":"e_1_3_2_45_2","DOI":"10.1109\/TMM.2021.3097502"},{"doi-asserted-by":"publisher","key":"e_1_3_2_46_2","DOI":"10.1145\/3534123"},{"doi-asserted-by":"publisher","key":"e_1_3_2_47_2","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_48_2","first-page":"1541","article-title":"Overcoming language priors in visual question answering with adversarial regularization","volume":"31","author":"Ramakrishnan Sainandan","year":"2018","unstructured":"Sainandan Ramakrishnan, Aishwarya Agrawal, and Stefan Lee. 2018. Overcoming language priors in visual question answering with adversarial regularization. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 31. 1541\u20131551.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_3_2_49_2","first-page":"91","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"28","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems (NeurIPS \u201915), Vol. 28, 91\u201399.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS \u201915)"},{"key":"e_1_3_2_50_2","first-page":"3070","article-title":"Multimodal graph networks for compositional generalization in visual question answering","volume":"33","author":"Saqur Raeid","year":"2020","unstructured":"Raeid Saqur and Karthik Narasimhan. 2020. Multimodal graph networks for compositional generalization in visual question answering. In Advances in Neural Information Processing Systems (NeurIPS \u201920), Vol. 33, 3070\u20133081.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS \u201920)"},{"key":"e_1_3_2_51_2","first-page":"618","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201917)","author":"Selvaraju Ramprasaath R.","year":"2017","unstructured":"Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017. Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201917), 618\u2013626."},{"key":"e_1_3_2_52_2","first-page":"2591","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201919)","author":"Selvaraju Ramprasaath R.","year":"2019","unstructured":"Ramprasaath R. Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Shalini Ghosh, Larry Heck, Dhruv Batra, and Devi Parikh. 2019. Taking a hint: Leveraging explanations to make vision and language models more grounded. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201919), 2591\u20132600."},{"key":"e_1_3_2_53_2","first-page":"20346","article-title":"Human-adversarial visual question answering","volume":"34","author":"Sheng Sasha","year":"2021","unstructured":"Sasha Sheng, Amanpreet Singh, Vedanuj Goswami, Jose Magana, Tristan Thrush, Wojciech Galuba, Devi Parikh, and Douwe Kiela. 2021. Human-adversarial visual question answering. In Advances in Neural Information Processing Systems (NeurIPS \u201921), Vol. 34, 20346\u201320359.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS \u201921)"},{"key":"e_1_3_2_54_2","first-page":"7716","article-title":"Adversarial scene editing: Automatic object removal from weak supervision","author":"Shetty Rakshith","year":"2018","unstructured":"Rakshith Shetty, Mario Fritz, and Bernt Schiele. 2018. Adversarial scene editing: Automatic object removal from weak supervision. In Advances in Neural Information Processing Systems (NeurIPS \u201918). Curran Associates, 7716\u20137726.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS \u201918)"},{"key":"e_1_3_2_55_2","first-page":"4101","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL \u201921)","author":"Si Qingyi","year":"2021","unstructured":"Qingyi Si, Zheng Lin, Ming yu Zheng, Peng Fu, and Weiping Wang. 2021. Check it again: Progressive visual question answering via visual entailment. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL \u201921), 4101\u20134110."},{"key":"e_1_3_2_56_2","first-page":"2647","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW \u201923)","author":"Sood Ekta","year":"2023","unstructured":"Ekta Sood, Fabian K\u00f6gel, Philipp M\u00fcller, Dominike Thomas, Mihai B\u00e2ce, and Andreas Bulling. 2023. Multimodal integration of human-Like attention in visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW \u201923), 2647\u20132657."},{"doi-asserted-by":"publisher","unstructured":"Xiangrui Su Qi Zhang Chongyang Shi Jiachang Liu and Liang Hu. 2023. Syntax tree constrained graph network for visual question answering. arXiv:2309.09179. Retrieved from 10.48550\/arXiv.2309.09179","key":"e_1_3_2_57_2","DOI":"10.48550\/arXiv.2309.09179"},{"key":"e_1_3_2_58_2","first-page":"747","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL \u201923)","author":"Subramanian Sanjay","year":"2023","unstructured":"Sanjay Subramanian, Medhini Narasimhan, Kushal Khangaonkar, Kevin Yang, Arsha Nagrani, Cordelia Schmid, Andy Zeng, Trevor Darrell, and Dan Klein. 2023. Modular visual question answering via code generation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL \u201923), 747\u2013761."},{"doi-asserted-by":"publisher","unstructured":"Yuwei Sun Hideya Ochiai and Jun Sakuma. 2023. Instance-level trojan attacks on visual question answering via adversarial learning in neuron activation space. arXiv:2304.00436. Retrieved from 10.48550\/arXiv.2304.00436","key":"e_1_3_2_59_2","DOI":"10.48550\/arXiv.2304.00436"},{"key":"e_1_3_2_60_2","first-page":"5100","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing (EMNLP-IJCNLP \u201919)","author":"Tan Hao","year":"2019","unstructured":"Hao Tan and Mohit Bansal. 2019. LXMERT: Learning cross-modality encoder representations from transformers. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing (EMNLP-IJCNLP \u201919), 5100\u20135111."},{"key":"e_1_3_2_61_2","first-page":"437","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201920)","author":"Tang Ruixue","year":"2020","unstructured":"Ruixue Tang, Chao Ma, Wei Emma Zhang, Qi Wu, and Xiaokang Yang. 2020. Semantic equivalent adversarial data augmentation for visual question answering. In Proceedings of the European Conference on Computer Vision (ECCV \u201920). Springer, 437\u2013453."},{"key":"e_1_3_2_62_2","first-page":"407","article-title":"On the value of out-of-distribution testing: An example of goodhart\u2019s law","volume":"33","author":"Teney Damien","year":"2020","unstructured":"Damien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha, Christopher Kanan, and Anton Van Den Hengel. 2020. On the value of out-of-distribution testing: An example of goodhart\u2019s law. In Advances in Neural Information Processing Systems (NeurIPS \u201920), Vol. 33, 407\u2013417.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS \u201920)"},{"key":"e_1_3_2_63_2","first-page":"5998","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS \u201917), Vol. 30, 5998\u20136008.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS \u201917)"},{"doi-asserted-by":"publisher","unstructured":"Ali Vosoughi Shijian Deng Songyang Zhang Yapeng Tian Chenliang Xu and Jiebo Luo. 2023. Unveiling cross modality bias in visual question answering: A causal view with possible worlds vqa. arXiv:2305.19664. Retrieved from 10.48550\/arXiv.2305.19664","key":"e_1_3_2_64_2","DOI":"10.48550\/arXiv.2305.19664"},{"key":"e_1_3_2_65_2","first-page":"3784","article-title":"Debiased visual question answering from feature and sample perspectives","volume":"34","author":"Wen Zhiquan","year":"2021","unstructured":"Zhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu, and Qi Wu. 2021. Debiased visual question answering from feature and sample perspectives. In Advances in Neural Information Processing Systems (NeurIPS \u201921), Vol. 34, 3784\u20133796.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS \u201921)"},{"key":"e_1_3_2_66_2","first-page":"8604","article-title":"Self-critical reasoning for robust visual question answering","volume":"32","author":"Wu Jialin","year":"2019","unstructured":"Jialin Wu and Raymond Mooney. 2019. Self-critical reasoning for robust visual question answering. In Advances in Neural Information Processing Systems (NeurIPS \u201919), Vol. 32, 8604\u20138614.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS \u201919)"},{"key":"e_1_3_2_67_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo (ICME \u201921)","author":"Yang Chao","year":"2021","unstructured":"Chao Yang, Su Feng, Dongsheng Li, Huawei Shen, Guoqing Wang, and Bin Jiang. 2021. Learning content and context with language bias for Visual Question Answering. In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME \u201921). IEEE, 1\u20136."},{"key":"e_1_3_2_68_2","first-page":"21","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201916)","author":"Yang Zichao","year":"2016","unstructured":"Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. 2016. Stacked attention networks for image question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201916), 21\u201329."},{"doi-asserted-by":"publisher","key":"e_1_3_2_69_2","DOI":"10.1109\/TMM.2020.2995278"},{"key":"e_1_3_2_70_2","first-page":"1083","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI \u201920)","author":"Zhu Xi","year":"2020","unstructured":"Xi Zhu, Zhendong Mao, Chunxiao Liu, Peng Zhang, Bin Wang, and Yongdong Zhang. 2020. Overcoming language priors with self-supervised learning for visual question answering. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI \u201920), 1083\u20131089."},{"key":"e_1_3_2_71_2","first-page":"8217","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201922)","author":"Zhu Zihao","year":"2022","unstructured":"Zihao Zhu. 2022. From shallow to deep: Compositional reasoning over graphs for visual question answering. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201922). IEEE, 8217\u20138221."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3688848","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3688848","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:04:10Z","timestamp":1750291450000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3688848"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,21]]},"references-count":70,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2024,12,31]]}},"alternative-id":["10.1145\/3688848"],"URL":"https:\/\/doi.org\/10.1145\/3688848","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2024,11,21]]},"assertion":[{"value":"2023-10-07","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-08-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-11-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}