{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,21]],"date-time":"2026-04-21T14:50:59Z","timestamp":1776783059213,"version":"3.51.2"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62402484, 62232016, 62072442, and 62272445"],"award-info":[{"award-number":["62402484, 62232016, 62072442, and 62272445"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Key Research and Development Program of China","award":["2024YFF0618800"],"award-info":[{"award-number":["2024YFF0618800"]}]},{"name":"Youth Innovation Promotion Association Chinese Academy of Sciences, Basic Research Program of ISCAS","award":["ISCAS-JCZD-202304"],"award-info":[{"award-number":["ISCAS-JCZD-202304"]}]},{"name":"Major Program of ISCAS","award":["ISCAS-ZD-202302"],"award-info":[{"award-number":["ISCAS-ZD-202302"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2026,1,31]]},"abstract":"<jats:p>Visual entailment (VE) is a multimodal reasoning task consisting of image-sentence pairs whereby a promise is defined by an image, and a sentence describes a hypothesis. The goal is to predict whether the image semantically entails the sentence. VE systems have been widely adopted in many downstream tasks such as image caption and visual question answering. However, the robustness of VE systems still faces significant challenges. One of the reasons is that the VE system suffers object-confusing defect when some similar objects exist. It outputs a positive prediction inferred by an erroneous object relationship, which will result in a fault negative prediction if the noised object does not exist.<\/jats:p>\n                  <jats:p>\n                    Previous approaches generate tests primarily relied on some general perturbations, such as simulating noise or weather interference in images, or substituting synonyms or rewriting sentences in texts. To test the object-confusing defect in VE systems, it requires perceiving and understanding key objects and entities and maintain the semantic relevance between cross-modal inputs, making it challenging to generate effective tests with high quality. Therefore, we propose\n                    <jats:monospace>VEglue<\/jats:monospace>\n                    , an object-aligned joint erasing approach for VE systems testing. It first aligns the object regions in the premise and object descriptions in the hypothesis to identify linked and un-linked objects. Then, based on the alignment information, three metamorphic relations are designed to jointly erase the objects of the two modalities. We evaluate\n                    <jats:monospace>VEglue<\/jats:monospace>\n                    on four widely used VE systems involving two public datasets, and the results demonstrate that\n                    <jats:monospace>VEglue<\/jats:monospace>\n                    could detect 11,609 issues on average with a 52.5% Issue Finding Rate (IFR). Furthermore, we leverage the tests generated by\n                    <jats:monospace>VEglue<\/jats:monospace>\n                    to retrain the VE systems, which largely improves model performance (50.8% increase in accuracy) on newly generated tests without sacrificing the accuracy on the original test set.\n                  <\/jats:p>","DOI":"10.1145\/3731244","type":"journal-article","created":{"date-parts":[[2025,4,18]],"date-time":"2025-04-18T10:41:21Z","timestamp":1744972881000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["VEglue: Testing Visual Entailment Systems via Object-Aligned Joint Erasing"],"prefix":"10.1145","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-7063-3044","authenticated-orcid":false,"given":"Zhiyuan","family":"Chang","sequence":"first","affiliation":[{"name":"State Key Laboratory of Intelligent Game, Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-7936-5593","authenticated-orcid":false,"given":"Mingyang","family":"Li","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Intelligent Game, Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9941-6713","authenticated-orcid":false,"given":"Junjie","family":"Wang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Intelligent Game, Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-8193-8497","authenticated-orcid":false,"given":"Cheng","family":"Li","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Intelligent Game, Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2618-5694","authenticated-orcid":false,"given":"Qing","family":"Wang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Intelligent Game, Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,12,11]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"6077","volume-title":"2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201918)","author":"Anderson Peter","year":"2018","unstructured":"Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201918), 6077\u20136086."},{"key":"e_1_3_2_3_2","first-page":"2425","volume-title":"2015 IEEE International Conference on Computer Vision (ICCV \u201915)","author":"Antol Stanislaw","year":"2015","unstructured":"Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: Visual question answering. In 2015 IEEE International Conference on Computer Vision (ICCV \u201915), 2425\u20132433."},{"key":"e_1_3_2_4_2","unstructured":"Jinze Bai Shuai Bai Shusheng Yang Shijie Wang Sinan Tan Peng Wang Junyang Lin Chang Zhou and Jingren Zhou. 2023. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv:2308.12966. Retrieved from https:\/\/arxiv.org\/abs\/2308.12966"},{"key":"e_1_3_2_5_2","unstructured":"Biwei Cao Jiuxin Cao Jie Gui Jiayun Shen Bo Liu Lei He Yuan Yan Tang and James Tin-Yau Kwok. 2022. AlignVE: Visual entailment recognition based on alignment relations. arXiv:2211.08736. Retrieved from https:\/\/arxiv.org\/abs\/2211.08736"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"2411","DOI":"10.18653\/v1\/2020.findings-emnlp.218","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020.","author":"Cao Yue","year":"2020","unstructured":"Yue Cao and Xiaojun Wan. 2020. DivGAN: Towards diverse paraphrase generation via diversified generative adversarial network. In Findings of the Association for Computational Linguistics: EMNLP 2020. Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, 2411\u20132421."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE51524.2021.9678670"},{"issue":"1","key":"e_1_3_2_8_2","first-page":"4:1","article-title":"Metamorphic testing: A review of challenges and opportunities","volume":"51","author":"Chen Tsong Yueh","year":"2018","unstructured":"Tsong Yueh Chen, Fei-Ching Kuo, Huai Liu, Pak-Lok Poon, Dave Towey, T. H. Tse, and Zhi Quan Zhou. 2018. Metamorphic testing: A review of challenges and opportunities. ACM Comput. Surv. 51, 1 (2018), 4:1\u20134:27.","journal-title":"ACM Comput. Surv"},{"key":"e_1_3_2_9_2","first-page":"4171","volume-title":"2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT \u201919)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT \u201919), 4171\u20134186."},{"key":"e_1_3_2_10_2","unstructured":"Virginie Do Oana-Maria Camburu Zeynep Akata and Thomas Lukasiewicz. 2020. e-SNLI-VE: Corrected visual-textual entailment with natural language explanations. arXiv:2004.03744. Retrieved from https:\/\/arxiv.org\/abs\/2004.03744"},{"key":"e_1_3_2_11_2","first-page":"2006","volume-title":"24th European Conference on Artificial Intelligence (ECAI \u201920), Including 10th Conference on Prestigious Applications of Artificial Intelligence (PAIS \u201920)","author":"Eberts Markus","year":"2020","unstructured":"Markus Eberts and Adrian Ulges. 2020. Span-based joint entity and relation extraction with transformer pre-training. In 24th European Conference on Artificial Intelligence (ECAI \u201920), Including 10th Conference on Prestigious Applications of Artificial Intelligence (PAIS \u201920), 2006\u20132013."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-009-0275-4"},{"key":"e_1_3_2_13_2","volume-title":"7th International Conference on Learning Representations (ICLR \u201919)","author":"Geirhos Robert","year":"2019","unstructured":"Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel. 2019. ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In 7th International Conference on Learning Representations (ICLR \u201919)."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-018-1116-0"},{"key":"e_1_3_2_15_2","first-page":"107","volume-title":"42nd International Conference on Software Engineering (ICSE \u201920)","author":"Gupta Shashij","year":"2020","unstructured":"Shashij Gupta. 2020. Machine translation testing via pathological invariance. In 42nd International Conference on Software Engineering (ICSE \u201920), 107\u2013109."},{"key":"e_1_3_2_16_2","first-page":"961","volume-title":"42nd International Conference on Software Engineering (ICSE \u201920)","author":"He Pinjia","year":"2020","unstructured":"Pinjia He, Clara Meister, and Zhendong Su. 2020. Structure-invariant testing for machine translation. In 42nd International Conference on Software Engineering (ICSE \u201920), 961\u2013973."},{"key":"e_1_3_2_17_2","volume-title":"7th International Conference on Learning Representations (ICLR \u201919)","author":"Hendrycks Dan","year":"2019","unstructured":"Dan Hendrycks and Thomas G. Dietterich. 2019. Benchmarking neural network robustness to common corruptions and perturbations. In 7th International Conference on Learning Representations (ICLR \u201919)."},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_19_2","first-page":"17959","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201922)","author":"Hu Xiaowei","year":"2022","unstructured":"Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, and Lijuan Wang. 2022. Scaling up vision-language pretraining for image captioning. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201922), 17959\u201317968."},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","unstructured":"Kento Kawaharazuka Yoshiki Obinata Naoaki Kanazawa Kei Okada and Masayuki Inaba. 2023. Robotic applications of pre-trained vision-language models to various recognition behaviors. arXiv:2303.05674. Retrieved from https:\/\/arxiv.org\/abs\/2303.05674","DOI":"10.1109\/Humanoids57100.2023.10375211"},{"key":"e_1_3_2_21_2","first-page":"260","volume-title":"The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL HLT \u201916)","author":"Lample Guillaume","year":"2016","unstructured":"Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016. Neural architectures for named entity recognition. In The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL HLT \u201916). Kevin Knight, Ani Nenkova, and Owen Rambow (Eds.), The Association for Computational Linguistics, 260\u2013270."},{"key":"e_1_3_2_22_2","unstructured":"Bo Li Gexiang Fang Yang Yang Quansen Wang Wei Ye Wen Zhao and Shikun Zhang. 2023. Evaluating ChatGPT\u2019s information extraction capabilities: An assessment of performance explainability calibration and faithfulness. arXiv:2304.11633. Retrieved from https:\/\/arxiv.org\/abs\/2304.11633"},{"key":"e_1_3_2_23_2","first-page":"9694","volume-title":"Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021 (NeurIPS \u201921)","author":"Li Junnan","year":"2021","unstructured":"Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven Chu-Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021 (NeurIPS \u201921), 9694\u20139705."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.5555\/3304222.3304355"},{"key":"e_1_3_2_25_2","unstructured":"Haotian Liu Chunyuan Li Yuheng Li and Yong Jae Lee. 2023. Improved baselines with visual instruction tuning. arXiv:2310.03744. Retrieved from https:\/\/arxiv.org\/abs\/2310.03744"},{"key":"e_1_3_2_26_2","unstructured":"Haotian Liu Chunyuan Li Qingyang Wu and Yong Jae Lee. 2023. Visual instruction tuning. arXiv:2304.08485. Retrieved from https:\/\/arxiv.org\/abs\/2304.08485"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","unstructured":"Hanmeng Liu Ruoxi Ning Zhiyang Teng Jian Liu Qiji Zhou and Yue Zhang. 2023. Evaluating the logical reasoning ability of ChatGPT and GPT-4. arXiv:2304.03439. DOI: 10.48550\/arXiv.2304.03439","DOI":"10.48550\/arXiv.2304.03439"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460319.3464829"},{"key":"e_1_3_2_29_2","first-page":"598","volume-title":"44th IEEE\/ACM 44th International Conference on Software Engineering (ICSE \u201922)","author":"Liu Zixi","year":"2022","unstructured":"Zixi Liu, Yang Feng, Yining Yin, and Zhenyu Chen. 2022. DeepState: Selecting test suites to enhance the robustness of recurrent neural networks. In 44th IEEE\/ACM 44th International Conference on Software Engineering (ICSE \u201922), 598\u2013609."},{"key":"e_1_3_2_30_2","unstructured":"Zheheng Luo Qianqian Xie and Sophia Ananiadou. 2023. ChatGPT as a factual inconsistency evaluator for abstractive text summarization. arXiv:2303.15621. Retrieved from https:\/\/arxiv.org\/abs\/2303.15621"},{"key":"e_1_3_2_31_2","unstructured":"Chengqi Lyu Wenwei Zhang Haian Huang Yue Zhou Yudong Wang Yanyi Liu Shilong Zhang and Kai Chen. 2022. RTMDet: An empirical study of designing real-time object detectors. arXiv:2212.07784. Retrieved from https:\/\/arxiv.org\/abs\/2212.07784"},{"key":"e_1_3_2_32_2","first-page":"1682","volume-title":"Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014","author":"Malinowski Mateusz","year":"2014","unstructured":"Mateusz Malinowski and Mario Fritz. 2014. A multi-world approach to question answering about real-world scenes based on uncertain input. In Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, 1682\u20131690."},{"key":"e_1_3_2_33_2","first-page":"11","volume-title":"2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201916)","author":"Mao Junhua","year":"2016","unstructured":"Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L. Yuille, and Kevin Murphy. 2016. Generation and comprehension of unambiguous object descriptions. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201916), 11\u201320."},{"key":"e_1_3_2_34_2","unstructured":"Meta AI. 2024. Introducing meta Llama 3: The most capable openly available LLM to date. Retrieved from https:\/\/ai.meta.com\/blog\/meta-llama-3\/"},{"key":"e_1_3_2_35_2","volume-title":"Workshop on Multi-Modal Fake News and Hate-Speech Detection (DE-FACTIFY \u201922) Co-located with the 36th AAAI Conference on Artificial Intelligence (AAAI \u201922)","author":"Mishra Shreyash","year":"2022","unstructured":"Shreyash Mishra, S. Suryavardan, Amrit Bhaskar, Parul Chopra, Aishwarya N. Reganti, Parth Patwa, Amitava Das, Tanmoy Chakraborty, Amit P. Sheth, and Asif Ekbal. 2022. FACTIFY: A multi-modal fact verification dataset. In Workshop on Multi-Modal Fake News and Hate-Speech Detection (DE-FACTIFY \u201922) Co-located with the 36th AAAI Conference on Artificial Intelligence (AAAI \u201922)."},{"key":"e_1_3_2_36_2","first-page":"717","volume-title":"Advances in Information Retrieval - 45th European Conference on Information Retrieval (ECIR \u201923)","author":"Nguyen Manh-Duy","year":"2023","unstructured":"Manh-Duy Nguyen, Binh T. Nguyen, and Cathal Gurrin. 2023. HADA: A graph-based amalgamation framework in image-text retrieval. In Advances in Information Retrieval - 45th European Conference on Information Retrieval (ECIR \u201923), 717\u2013731."},{"key":"e_1_3_2_37_2","unstructured":". OpenAI. (2023). GPT-4V(ision) System Card. Technical Report OpenAI San Francisco CA. Retrieved September 25 2023 from https:\/\/openai.com\/index\/gpt-4v-system-card"},{"key":"e_1_3_2_38_2","unstructured":"OpenAI. 2024. Hello GPT-4O. Retrieved from https:\/\/openai.com\/index\/hello-gpt-4o\/"},{"key":"e_1_3_2_39_2","first-page":"8748","volume-title":"38th International Conference on Machine Learning (ICML \u201921)\n                  Proceedings of Machine Learning Research, Vol. 139","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In 38th International Conference on Machine Learning (ICML \u201921). Marina Meila and Tong Zhang (Eds.). Proceedings of Machine Learning Research, Vol. 139, PMLR, 8748\u20138763."},{"key":"e_1_3_2_40_2","doi-asserted-by":"crossref","first-page":"779","DOI":"10.1109\/CVPR.2016.91","volume-title":"2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201916)","author":"Redmon Joseph","year":"2016","unstructured":"Joseph Redmon, Santosh Kumar Divvala, Ross B. Girshick, and Ali Farhadi. 2016. You only look once: Unified, real-time object detection. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201916), 779\u2013788."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_2_42_2","unstructured":"S. Suryavardan Shreyash Mishra Parth Patwa Megha Chakraborty Anku Rani Aishwarya Reganti Aman Chadha Amitava Das Amit P. Sheth Manoj Chinnakotla et al. 2023. Factify 2: A multimodal fake news and satire news dataset. arXiv:2304.03897. Retrieved from https:\/\/arxiv.org\/abs\/2304.03897"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3551349.3556953"},{"key":"e_1_3_2_44_2","first-page":"4101","volume-title":"59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL\/IJCNLP \u201921)","author":"Si Qingyi","year":"2021","unstructured":"Qingyi Si, Zheng Lin, Mingyu Zheng, Peng Fu, and Weiping Wang. 2021. Check it again: Progressive visual question answering via visual entailment. In 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL\/IJCNLP \u201921), 4101\u20134110."},{"key":"e_1_3_2_45_2","first-page":"1181","volume-title":"44th IEEE\/ACM 44th International Conference on Software Engineering (ICSE \u201922)","author":"Sun Zeyu","year":"2022","unstructured":"Zeyu Sun, Jie M. Zhang, Yingfei Xiong, Mark Harman, Mike Papadakis, and Lu Zhang. 2022. Improving machine translation systems via isotopic replacement. In 44th IEEE\/ACM 44th International Conference on Software Engineering (ICSE \u201922), 1181\u20131192."},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","first-page":"3172","DOI":"10.1109\/WACV51458.2022.00323","volume-title":"IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV \u201922)","author":"Suvorov Roman","year":"2022","unstructured":"Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, and Victor Lempitsky. 2022. Resolution-robust large mask inpainting with Fourier convolutions. In IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV \u201922), 3172\u20133182."},{"key":"e_1_3_2_47_2","first-page":"5998","volume-title":"Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, 5998\u20136008."},{"key":"e_1_3_2_48_2","unstructured":"Peng Wang Shijie Wang Junyang Lin Shuai Bai Xiaohuan Zhou Jingren Zhou Xinggang Wang and Chang Zhou. 2023. ONE-PEACE: Exploring one general representation model toward unlimited modalities. arXiv:2305.11172. Retrieved from https:\/\/arxiv.org\/abs\/2305.11172"},{"key":"e_1_3_2_49_2","first-page":"23318","volume-title":"International Conference on Machine Learning (ICML \u201922)","author":"Wang Peng","year":"2022","unstructured":"Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022. OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Learning (ICML \u201922), 23318\u201323340."},{"key":"e_1_3_2_50_2","doi-asserted-by":"crossref","first-page":"2387","DOI":"10.1109\/ICSE48619.2023.00200","volume-title":"2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE)","author":"Wang Wenxuan","year":"2023","unstructured":"Wenxuan Wang, Jen-tse Huang, Weibin Wu, Jianping Zhang, Yizhan Huang, Shuqing Li, Pinjia He, and Michael R. Lyu. 2023. MTTM: Metamorphic testing for textual content moderation software. In 2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 2387\u20132399."},{"key":"e_1_3_2_51_2","first-page":"347","volume-title":"Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL \u201921)","author":"Wang Xiao","year":"2021","unstructured":"Xiao Wang, Qin Liu, Tao Gui, Qi Zhang, Yicheng Zou, Xin Zhou, Jiacheng Ye, Yongxin Zhang, Rui Zheng, Zexiong Pang, et al. 2021. TextFlint: Unified multilingual robustness evaluation toolkit for natural language processing. In Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL \u201921), 347\u2013355."},{"key":"e_1_3_2_52_2","unstructured":"Ning Xie Farley Lai Derek Doran and Asim Kadav. 2018. Visual entailment task for visually-grounded language learning. arXiv:1811.10582. Retrieved from https:\/\/arxiv.org\/abs\/1811.10582"},{"key":"e_1_3_2_53_2","unstructured":"Ning Xie Farley Lai Derek Doran and Asim Kadav. 2019. Visual entailment: A novel task for fine-grained image understanding. arXiv:1901.06706. Retrieved from https:\/\/arxiv.org\/abs\/1901.06706"},{"key":"e_1_3_2_54_2","doi-asserted-by":"crossref","first-page":"8","DOI":"10.18653\/v1\/2023.clinicalnlp-1.2","volume-title":"5th Clinical Natural Language Processing Workshop (ClinicalNLP@ACL \u201923)","author":"Yanaka Hitomi","year":"2023","unstructured":"Hitomi Yanaka, Yuta Nakamura, Yuki Chida, and Tomoya Kurosawa. 2023. Medical visual textual entailment for numerical understanding of vision-and-language models. In 5th Clinical Natural Language Processing Workshop (ClinicalNLP@ACL \u201923), 8\u201318."},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548284"},{"key":"e_1_3_2_56_2","unstructured":"Zhengyuan Yang Linjie Li Kevin Lin Jianfeng Wang Chung-Ching Lin Zicheng Liu and Lijuan Wang. 2023. The dawn of LMMs: Preliminary explorations with GPT-4V(ision). arXiv:2309.17421. Retrieved from https:\/\/arxiv.org\/abs\/2309.17421"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598094"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3533767.3534389"},{"key":"e_1_3_2_59_2","first-page":"4470","volume-title":"2019 IEEE\/CVF International Conference on Computer Vision (ICCV \u201919)","author":"Yu Jiahui","year":"2019","unstructured":"Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S. Huang. 2019. Free-form image inpainting with gated convolution. In 2019 IEEE\/CVF International Conference on Computer Vision (ICCV \u201919), IEEE, 4470\u20134479."},{"key":"e_1_3_2_60_2","first-page":"16908","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201921)","author":"Yuan Yuanyuan","year":"2021","unstructured":"Yuanyuan Yuan, Shuai Wang, Mingyue Jiang, and Tsong Yueh Chen. 2021. Perception matters: Detecting perception failures of VQA models using metamorphic testing. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201921), 16908\u201316917."},{"key":"e_1_3_2_61_2","volume-title":"The 11th International Conference on Learning Representations (ICLR \u201923)","author":"Zhang Hao","year":"2023","unstructured":"Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M. Ni, and Heung-Yeung Shum. 2023. DINO: DETR with Improved DeNoising anchor boxes for end-to-end object detection. In The 11th International Conference on Learning Representations (ICLR \u201923)."},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2018.2876865"},{"key":"e_1_3_2_63_2","unstructured":"Qihuang Zhong Liang Ding Juhua Liu Bo Du and Dacheng Tao. 2023. Can ChatGPT understand too? A comparative study on ChatGPT and fine-tuned BERT. arXiv:2302.10198. Retrieved from https:\/\/arxiv.org\/abs\/2302.10198"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3731244","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,11]],"date-time":"2025-12-11T15:55:48Z","timestamp":1765468548000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3731244"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,11]]},"references-count":62,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1,31]]}},"alternative-id":["10.1145\/3731244"],"URL":"https:\/\/doi.org\/10.1145\/3731244","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,11]]},"assertion":[{"value":"2024-05-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-12","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}