{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T03:30:05Z","timestamp":1779334205056,"version":"3.51.4"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"10","license":[{"start":{"date-parts":[[2024,10,29]],"date-time":"2024-10-29T00:00:00Z","timestamp":1730160000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2022ZD0118801"],"award-info":[{"award-number":["2022ZD0118801"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62206279 and U21B2043"],"award-info":[{"award-number":["62206279 and U21B2043"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,10,31]]},"abstract":"<jats:p>\n            Despite most state-of-the-art models having achieved amazing performance in\n            <jats:bold>Visual Question Answering (VQA)<\/jats:bold>\n            , they usually utilize biases to answer the question. Recently, some studies synthesize counterfactual training samples to help the model to mitigate the biases. However, these synthetic samples need extra annotations and often contain noises. Moreover, these methods simply add synthetic samples to the training data to train the model with the cross-entropy loss, which cannot make the best use of synthetic samples to mitigate the biases. In this article, to mitigate the biases in VQA more effectively, we propose a\n            <jats:bold>Hierarchical Counterfactual Contrastive Learning (HCCL)<\/jats:bold>\n            method. Firstly, to avoid introducing noises and extra annotations, our method automatically masks the unimportant features in original pairs to obtain positive samples and create mismatched question-image pairs as negative samples. Then our method uses feature-level and answer-level contrastive learning to make the original sample close to positive samples in the feature space, while away from negative samples in both feature and answer spaces. In this way, the VQA model can learn the robust multimodal features and focus on both visual and language information to produce the answer. Our HCCL method can be adopted in different baselines, and the experimental results on VQA v2, VQA-CP, and GQA-OOD datasets show that our method is effective in mitigating the biases in VQA, which improves the robustness of the VQA model.\n          <\/jats:p>","DOI":"10.1145\/3673902","type":"journal-article","created":{"date-parts":[[2024,6,27]],"date-time":"2024-06-27T19:53:12Z","timestamp":1719517992000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["HCCL: Hierarchical Counterfactual Contrastive Learning for Robust Visual Question Answering"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5504-1799","authenticated-orcid":false,"given":"Dongze","family":"Hao","sequence":"first","affiliation":[{"name":"Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, China and School of Artificial Intelligence, University of Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5607-7854","authenticated-orcid":false,"given":"Qunbo","family":"Wang","sequence":"additional","affiliation":[{"name":"Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2142-5580","authenticated-orcid":false,"given":"Xinxin","family":"Zhu","sequence":"additional","affiliation":[{"name":"Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0903-9131","authenticated-orcid":false,"given":"Jing","family":"Liu","sequence":"additional","affiliation":[{"name":"Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, China and School of Artificial Intelligence, University of Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,10,29]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01006"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00971"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00522"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.12"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018102"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00209"},{"key":"e_1_3_1_10_2","first-page":"830","article-title":"RUBi: Reducing unimodal biases for visual question answering","volume":"32","author":"Cadene Remi","year":"2019","unstructured":"Remi Cadene, Corentin Dancette, Matthieu Cord, Devi Parikh, et al. 2019b. RUBi: Reducing unimodal biases for visual question answering. Advances in Neural Information Processing Systems 32 (2019), 830\u2013850.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01081"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3290012"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20059-5_6"},{"key":"e_1_3_1_14_2","first-page":"1597","volume-title":"International Conference on Machine Learning","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020a. A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning. PMLR, 1597\u20131607."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01124"},{"key":"e_1_3_1_16_2","unstructured":"Christopher Clark Mark Yatskar and Luke Zettlemoyer. 2019. Don\u2019t take the easy way out: Ensemble based methods for avoiding known dataset biases. arXiv:1909.03683. Retrieved from https:\/\/aclanthology.org\/D19-1418"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2017.10.001"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","unstructured":"Akira Fukui Dong Huk Park Daylen Yang Anna Rohrbach Trevor Darrell and Marcus Rohrbach. 2016. Multimodal compact bilinear pooling for visual question answering and visual grounding. arXiv: 1606.01847. DOI: 10.18653\/v1\/D16-1044","DOI":"10.18653\/v1\/D16-1044"},{"key":"e_1_3_1_19_2","first-page":"3197","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"Gat Itai","year":"2020","unstructured":"Itai Gat, Idan Schwartz, Alexander Schwing, and Tamir Hazan. 2020. Removing bias in multi-modal classifiers: Regularization by maximizing functional entropies. In Proceedings of the 34th International Conference on Neural Information Processing Systems. 3197\u20133208."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.63"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.670"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","unstructured":"Gabriel Grand and Yonatan Belinkov. 2019. Adversarial regularization for visual question answering: Strengths shortcomings and side effects. arXiv:1906.08430. DOI: 10.18653\/v1\/W19-1801","DOI":"10.18653\/v1\/W19-1801"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00161"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00686"},{"key":"e_1_3_1_26_2","first-page":"448","volume-title":"International Conference on Machine Learning","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning. PMLR, 448\u2013456."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.217"},{"key":"e_1_3_1_28_2","volume-title":"International Conference on Learning Representations (ICLR)","author":"Kaushik Divyansh","year":"2020","unstructured":"Divyansh Kaushik, Eduard Hovy, and Zachary C. Lipton. 2020. Learning the difference that makes a difference with counterfactually augmented data. In International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00280"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.5555\/3326943.3327087"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV51458.2022.00263"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58601-0_2"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.311"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.265"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.3016083"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01251"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3487042"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00915"},{"key":"e_1_3_1_40_2","first-page":"8026","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. In Proceedings of the 33rd International Conference on Neural Information Processing Systems. 8026\u20138037."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_1_42_2","first-page":"1083","volume-title":"Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence","author":"Ramakrishnan Sainandan","year":"2018","unstructured":"Sainandan Ramakrishnan, Aishwarya Agrawal, and Stefan Lee. 2018. Overcoming language priors in visual question answering with adversarial regularization. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence. 1083\u20131089."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.74"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00268"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","unstructured":"Robik Shrestha Kushal Kafle and Christopher Kanan. 2020. A negative case analysis of visual grounding methods for vqa. arXiv: 2004.05704. DOI: 10.18653\/v1\/2020.acl-main.727","DOI":"10.18653\/v1\/2020.acl-main.727"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","unstructured":"Hao Tan and Mohit Bansal. 2019. LXMERT: Learning cross-modality encoder representations from transformers. arXiv: 1908.07490. DOI: 10.18653\/v1\/D19-1514","DOI":"10.18653\/v1\/D19-1514"},{"key":"e_1_3_1_47_2","first-page":"407","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"Teney Damien","year":"2020","unstructured":"Damien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha, Christopher Kanan, and Anton Van Den Hengel. 2020b. On the value of out-of-distribution testing: An example of goodhart\u2019s law. In Proceedings of the 34th International Conference on Neural Information Processing Systems. 407\u2013417."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58607-2_34"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","unstructured":"Aaron Van den Oord Yazhe Li and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv: 1807.03748. DOI: 10.48550\/arXiv.1807.03748","DOI":"10.48550\/arXiv.1807.03748"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","unstructured":"Zixu Wang Yishu Miao and Lucia Specia. 2021. Cross-modal generative augmentation for visual question answering. arXiv: 2105.04780. DOI: 10.48550\/arXiv.2105.0478","DOI":"10.48550\/arXiv.2105.0478"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.432"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.5555\/3540261.3540550"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3455059"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.10"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2018.2817340"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.5555\/3495724.3497245"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","unstructured":"Xi Zhu Zhendong Mao Chunxiao Liu Peng Zhang Bin Wang and Yongdong Zhang. 2020. Overcoming language priors with self-supervised learning for visual question answering. arXiv: 2012.11528. DOI: 10.5555\/3491440.3491591","DOI":"10.5555\/3491440.3491591"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3673902","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3673902","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:58:23Z","timestamp":1750294703000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3673902"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,29]]},"references-count":56,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2024,10,31]]}},"alternative-id":["10.1145\/3673902"],"URL":"https:\/\/doi.org\/10.1145\/3673902","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,29]]},"assertion":[{"value":"2023-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-06-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-29","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}