{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T15:02:58Z","timestamp":1775228578149,"version":"3.50.1"},"reference-count":39,"publisher":"Association for Computing Machinery (ACM)","issue":"11","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62306342"],"award-info":[{"award-number":["62306342"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Excellent Young Scientists Fund in Hunan Province","award":["2024JJ4070"],"award-info":[{"award-number":["2024JJ4070"]}]},{"name":"Key Laboratory of Computing Power Network and Information Security, Ministry of Education","award":["2023ZD032"],"award-info":[{"award-number":["2023ZD032"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>\n                    Multi-modal sarcasm detection involves determining whether a given multi-modal input conveys sarcastic intent by analyzing the underlying sentiment. Recently, vision large language models have shown remarkable success on various of multi-modal tasks. Inspired by this, we systematically investigate the impact of vision large language models in zero-shot multi-modal sarcasm detection task. Furthermore, to capture different perspectives of sarcastic expressions, we propose a multi-view agent framework, S\n                    <jats:sup>3<\/jats:sup>\n                    Agent, designed to enhance zero-shot multi-modal sarcasm detection by leveraging three critical perspectives:\n                    <jats:italic toggle=\"yes\">superficial expression<\/jats:italic>\n                    ,\n                    <jats:italic toggle=\"yes\">semantic information<\/jats:italic>\n                    , and\n                    <jats:italic toggle=\"yes\">sentiment expression<\/jats:italic>\n                    . Our experiments on the MMSD2.0 dataset, which involves six models and four prompting strategies, demonstrate that our approach achieves state-of-the-art performance. Our method achieves an average improvement of 13.2% in accuracy. Moreover, we evaluate our method on the text-only sarcasm detection task, where it also surpasses baseline approaches.\n                  <\/jats:p>","DOI":"10.1145\/3690642","type":"journal-article","created":{"date-parts":[[2024,8,29]],"date-time":"2024-08-29T12:24:27Z","timestamp":1724934267000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["S\n                    <sup>3<\/sup>\n                    Agent: Unlocking the Power of VLLM for Zero-Shot Multi-Modal Sarcasm Detection"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5374-8931","authenticated-orcid":false,"given":"Peng","family":"Wang","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, Central South University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4339-2855","authenticated-orcid":false,"given":"Yongheng","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Central South University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3026-6347","authenticated-orcid":false,"given":"Hao","family":"Fei","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9154-7858","authenticated-orcid":false,"given":"Qiguang","family":"Chen","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-2426-9711","authenticated-orcid":false,"given":"Yukai","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Central South University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6870-5678","authenticated-orcid":false,"given":"Jiasheng","family":"Si","sequence":"additional","affiliation":[{"name":"Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Qilu University of Technology (Shandong Academy of Sciences), Jinan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1840-3540","authenticated-orcid":false,"given":"Wenpeng","family":"Lu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Qilu University of Technology (Shandong Academy of Sciences), Jinan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0188-1394","authenticated-orcid":false,"given":"Min","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Central South University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3619-675X","authenticated-orcid":false,"given":"Libo","family":"Qin","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Central South University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,11,10]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.semeval-1.111"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.20"},{"key":"e_1_3_1_4_2","unstructured":"Jinze Bai Shuai Bai Shusheng Yang Shijie Wang Sinan Tan Peng Wang Junyang Lin Chang Zhou and Jingren Zhou. 2023. Qwen-VL: A Frontier large vision-language model with versatile abilities.arXiv:2308.12966. Retrieved from https:\/\/arxiv.org\/abs\/2308.12966"},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","unstructured":"David Bamman and Noah Smith. 2015. Contextualized sarcasm detection on Twitter. In Proceedings of the International AAAI Conference on Web and Social Media Vol. 9 574\u2013577.","DOI":"10.1609\/icwsm.v9i1.14655"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1239"},{"key":"e_1_3_1_7_2","doi-asserted-by":"crossref","unstructured":"Qiguang Chen Libo Qin Jin Zhang Zhi Chen Xiao Xu and Wanxiang Che. 2024. M 3 CoT: A novel benchmark for multi-domain multi-step multi-modal chain-of-thought. arXiv:2405.16473. Retrieved from https:\/\/arxiv.org\/abs\/2405.16473","DOI":"10.18653\/v1\/2024.acl-long.446"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2024.127428"},{"key":"e_1_3_1_9_2","unstructured":"DeepSeek-AI. 2024. DeepSeek-V2: A strong economical and efficient mixture-of-experts language model. arXiv:2405.04434 [cs.CL]. Retrieved from https:\/\/arxiv.org\/abs\/2405.04434"},{"key":"e_1_3_1_10_2","first-page":"13109","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Fei Hao","year":"2024","unstructured":"Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, Meishan Zhang, Mong-Li Lee, and Wynne Hsu. 2024. Video-of-thought: Step-by-step video reasoning from perception to cognition. In Proceedings of the International Conference on Machine Learning, Vol. 235, 13109\u201313125."},{"key":"e_1_3_1_11_2","unstructured":"Shengding Hu Yuge Tu Xu Han Chaoqun He Ganqu Cui Xiang Long Zhi Zheng Yewei Fang Yuxiang Huang Weilin Zhao Xinrong Zhang Zheng Leng Thai Kaihuo Zhang Chongyi Wang Yuan Yao Chenyang Zhao Jie Zhou Jie Cai Zhongwu Zhai Ning Ding Chao Jia Guoyang Zeng Dahai Li Zhiyuan Liu and Maosong Sun. 2024. MiniCPM: Unveiling the potential of small language models with scalable training strategies. arXiv:2404.06395."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02644"},{"key":"e_1_3_1_13_2","unstructured":"Douwe Kiela Hamed Firooz Aravind Mohan Vedanuj Goswami Amanpreet Singh Pratik Ringshia and Davide Testuggine. 2020. The hateful memes challenge: Detecting hate speech in multimodal memes. In Proceedings of the Advances in Neural Information Processing Systems Vol. 33 2611\u20132624."},{"key":"e_1_3_1_14_2","first-page":"22199","volume-title":"Proceedings of the Advances in Neural Information Processing SystemsCurran Associates, Inc","volume":"35","author":"Kojima Takeshi","year":"2022","unstructured":"Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Proceedings of the Advances in Neural Information Processing Systems. S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 22199\u201322213. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/file\/8bb0d291acd4acf06ef112099c16f326-Paper-Conference.pdf"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/tmm.2024.3384060"},{"key":"e_1_3_1_16_2","first-page":"27831","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Kuckreja Kartik","year":"2024","unstructured":"Kartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer, Abhijit Das, Salman Khan, and Fahad Shahbaz Khan. 2024. Geochat: Grounded large vision-language model for remote sensing. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 27831\u201327840."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.124"},{"key":"e_1_3_1_18_2","unstructured":"Bin Lin Zhenyu Tang Yang Ye Jiaxi Cui Bin Zhu Peng Jin Junwu Zhang Munan Ning and Li Yuan. 2024. MoE-LLaVA: Mixture of experts for large vision-language models. arXiv:2401.15947. Retrieved from https:\/\/arxiv.org\/abs\/2401.15947"},{"key":"e_1_3_1_19_2","unstructured":"Hongzhan Lin Zixin Chen Ziyang Luo Mingfei Cheng Jing Ma and Guang Chen. 2024. CofiPara: A coarse-to-fine paradigm for multimodal sarcasm target identification with large multimodal models. arXiv:2405.00390. Retrieved from https:\/\/arxiv.org\/abs\/2405.00390"},{"key":"e_1_3_1_20_2","unstructured":"Haotian Liu Chunyuan Li Yuheng Li and Yong Jae Lee. 2023. Improved baselines with visual instruction tuning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201924) 26296\u201326306. Retrieved from https:\/\/openaccess.thecvf.com\/content\/CVPR2024\/html\/Liu_Improved_Baselines_with_Visual_Instruction_Tuning_CVPR_2024_paper.html"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.225"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-emnlp.74"},{"key":"e_1_3_1_23_2","unstructured":"Haoyu Lu Wen Liu Bo Zhang Bingxuan Wang Kai Dong Bo Liu Jingxiang Sun Tongzheng Ren Zhuoshu Li Yaofeng Sun and Chengqi Deng Hanwei Xu Zhenda Xie Chong Ruan. 2024. DeepSeek-VL: Towards real-world vision-language understanding. arXiv:2403.05525. Retrieved from https:\/\/arxiv.org\/abs\/2403.05525"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICACSIS.2013.6761575"},{"key":"e_1_3_1_25_2","first-page":"1601","volume-title":"Proceedings of the 26th International Conference on Computational Linguistics: Technical Papers (COLING \u201916)The COLING 2016 Organizing Committee","author":"Poria Soujanya","year":"2016","unstructured":"Soujanya Poria, Erik Cambria, Devamanyu Hazarika, and Prateek Vij. 2016. A deeper look into sarcastic tweets using deep convolutional neural networks. In Proceedings of the 26th International Conference on Computational Linguistics: Technical Papers (COLING \u201916). Yuji Matsumoto and Rashmi Prasad (Eds.), The COLING 2016 Organizing Committee, 1601\u20131612. DOI: https:\/\/aclanthology.org\/C16-1151"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.689"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2964321"},{"key":"e_1_3_1_28_2","unstructured":"Gemini Team Rohan Anil Sebastian Borgeaud Yonghui Wu Jean-Baptiste Alayrac Jiahui Yu Radu Soricut Johan Schalkwyk Andrew M. Dai Anja Hauth et al. 2023. Gemini: A family of highly capable multimodal models. arXiv:2312.11805."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.139"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.3233\/978-1-60750-606-5-765"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.147"},{"key":"e_1_3_1_32_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Wang Xuezhi","year":"2022","unstructured":"Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. In Proceedings of the 11th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=1PL1NIMMrw"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"key":"e_1_3_1_34_2","first-page":"53366","volume-title":"Proceedings of the International Conference on Machine Learning","volume":"235","author":"Wu Shengqiong","year":"2024","unstructured":"Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. 2024. NExT-GPT: Any-to-any multimodal LLM. In Proceedings of the International Conference on Machine Learning, Vol. 235, 53366\u201353397."},{"key":"e_1_3_1_35_2","doi-asserted-by":"crossref","unstructured":"Tao Xiong Peiran Zhang Hongbo Zhu and Yihui Yang. 2019. Sarcasm detection with self-matching networks and low-rank bilinear pooling. In Proceedings of the World Wide Web Conference 2115\u20132124.","DOI":"10.1145\/3308558.3313735"},{"key":"e_1_3_1_36_2","unstructured":"Shukang Yin Chaoyou Fu Sirui Zhao Ke Li Xing Sun Tong Xu and Enhong Chen. 2023. A survey on multimodal large language models. arXiv:2306.13549. Retrieved from https:\/\/arxiv.org\/abs\/2306.13549"},{"key":"e_1_3_1_37_2","unstructured":"Alex Young Bei Chen Chao Li Chengen Huang Ge Zhang Guanwei Zhang Heng Li Jiangcheng Zhu Jianqun Chen Jing Chang Kaidong Yu Peng Liu Qiang Liu Shawn Yue Senbin Yang Shiming Yang Tao Yu Wen Xie Wenhao Huang Xiaohui Hu Xiaoyi Ren Xinyao Niu Pengcheng Nie Yuchi Xu Yudong Liu Yue Wang Yuxuan Cai Zhenyu Gu Zhiyuan Liu and Zonghong Dai. 2024. Yi: Open foundation models by 01.AI. arXiv:2403.04652. Retrieved from https:\/\/arxiv.org\/abs\/2403.04652"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.aacl-main.32"},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","unstructured":"Yucheng Zhou Xiang Li Qianning Wang and Jianbing Shen. 2024. Visual in-context learning for large vision-language models. arXiv:2402.11574.","DOI":"10.18653\/v1\/2024.findings-acl.940"},{"key":"e_1_3_1_40_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Zhou Yongchao","year":"2023","unstructured":"Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2023. Large language models are human-level prompt engineers. In Proceedings of the 11th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=92gvk82DE-"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3690642","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,10]],"date-time":"2025-11-10T14:50:54Z","timestamp":1762786254000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3690642"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,10]]},"references-count":39,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3690642"],"URL":"https:\/\/doi.org\/10.1145\/3690642","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,10]]},"assertion":[{"value":"2024-06-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-08-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}