{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,14]],"date-time":"2026-06-14T21:31:18Z","timestamp":1781472678267,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":49,"publisher":"ACM","funder":[{"name":"National Science Foundation of China","award":["62272456"],"award-info":[{"award-number":["62272456"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,6,18]]},"DOI":"10.1145\/3733102.3733138","type":"proceedings-article","created":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T11:14:07Z","timestamp":1750158847000},"page":"24-34","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Can GPT tell us why these images are synthesized? Empowering Multimodal Large Language Models for Forensics"],"prefix":"10.1145","author":[{"given":"Yiran","family":"He","sequence":"first","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences., Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yun","family":"Cao","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences., Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-2415-2076","authenticated-orcid":false,"given":"Bowen","family":"Yang","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences., Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-8580-6493","authenticated-orcid":false,"given":"Zeyu","family":"Zhang","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences., Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,17]]},"reference":[{"key":"e_1_3_3_2_2_2","doi-asserted-by":"crossref","unstructured":"Vishal Asnani Xi Yin Tal Hassner and Xiaoming Liu. 2023. Reverse engineering of generative models: Inferring model hyperparameters from generated images. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 12 (2023) 15477\u201315493.","DOI":"10.1109\/TPAMI.2023.3301451"},{"key":"e_1_3_3_2_3_2","unstructured":"Xiao Bi Deli Chen Guanting Chen Shanhuang Chen Damai Dai Chengqi Deng Honghui Ding Kai Dong Qiushi Du Zhe Fu et\u00a0al. 2024. Deepseek llm: Scaling open-source language models with longtermism. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2401.02954 (2024)."},{"key":"e_1_3_3_2_4_2","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared\u00a0D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et\u00a0al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020) 1877\u20131901."},{"key":"e_1_3_3_2_5_2","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1007\/978-3-030-58574-7_7","volume-title":"Computer vision\u2013ECCV 2020: 16th European conference, Glasgow, UK, August 23\u201328, 2020, proceedings, part XXVI 16","author":"Chai Lucy","year":"2020","unstructured":"Lucy Chai, David Bau, Ser-Nam Lim, and Phillip Isola. 2020. What makes fake images detectable? understanding properties that generalize. In Computer vision\u2013ECCV 2020: 16th European conference, Glasgow, UK, August 23\u201328, 2020, proceedings, part XXVI 16. Springer, 103\u2013120."},{"key":"e_1_3_3_2_6_2","doi-asserted-by":"crossref","unstructured":"Beijing Chen Xin Liu Yuhui Zheng Guoying Zhao and Yun-Qing Shi. 2021. A robust GAN-generated face detection method based on dual-color spaces and an improved Xception. IEEE Transactions on Circuits and Systems for Video Technology 32 6 (2021) 3527\u20133538.","DOI":"10.1109\/TCSVT.2021.3116679"},{"key":"e_1_3_3_2_7_2","unstructured":"Wei-Lin Chiang Zhuohan Li Ziqing Lin Ying Sheng Zhanghao Wu Hao Zhang Lianmin Zheng Siyuan Zhuang Yonghao Zhuang Joseph\u00a0E Gonzalez et\u00a0al. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https:\/\/vicuna. lmsys. org (accessed 14 April 2023) 2 3 (2023) 6."},{"key":"e_1_3_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW59228.2023.00104"},{"key":"e_1_3_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10095167"},{"key":"e_1_3_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00791"},{"key":"e_1_3_3_2_11_2","unstructured":"Ricard Durall Margret Keuper Franz-Josef Pfreundt and Janis Keuper. 2019. Unmasking deepfakes with simple features. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/1911.00686 (2019)."},{"key":"e_1_3_3_2_12_2","doi-asserted-by":"publisher","unstructured":"Michael Fink and Pietro Perona. 2022. Caltech 10k Web Faces. 10.22002\/D1.20132","DOI":"10.22002\/D1.20132"},{"key":"e_1_3_3_2_13_2","first-page":"3247","volume-title":"International conference on machine learning","author":"Frank Joel","year":"2020","unstructured":"Joel Frank, Thorsten Eisenhofer, Lea Sch\u00f6nherr, Asja Fischer, Dorothea Kolossa, and Thorsten Holz. 2020. Leveraging frequency analysis for deep fake image recognition. In International conference on machine learning. PMLR, 3247\u20133258."},{"key":"e_1_3_3_2_14_2","unstructured":"Tsu-Jui Fu Wenze Hu Xianzhi Du William\u00a0Yang Wang Yinfei Yang and Zhe Gan. 2023. Guiding instruction-based image editing via multimodal large language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2309.17102 (2023)."},{"key":"e_1_3_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01457"},{"key":"e_1_3_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICME51207.2021.9428429"},{"key":"e_1_3_3_2_17_2","doi-asserted-by":"crossref","unstructured":"Hui Guo Shu Hu Xin Wang Ming-Ching Chang and Siwei Lyu. 2022. Robust attentive deep neural network for detecting gan-generated faces. IEEE Access 10 (2022) 32574\u201332583.","DOI":"10.1109\/ACCESS.2022.3157297"},{"key":"e_1_3_3_2_18_2","first-page":"893","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Jia Shan","year":"2023","unstructured":"Shan Jia, Mingzhen Huang, Zhou Zhou, Yan Ju, Jialing Cai, and Siwei Lyu. 2023. Autosplice: A text-prompt manipulated image dataset for media forensics. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 893\u2013903."},{"key":"e_1_3_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW63382.2024.00436"},{"key":"e_1_3_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01300"},{"key":"e_1_3_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP46576.2022.9897820"},{"key":"e_1_3_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00453"},{"key":"e_1_3_3_2_23_2","unstructured":"Chunyuan Li Cliff Wong Sheng Zhang Naoto Usuyama Haotian Liu Jianwei Yang Tristan Naumann Hoifung Poon and Jianfeng Gao. 2023. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems 36 (2023) 28541\u201328564."},{"key":"e_1_3_3_2_24_2","doi-asserted-by":"publisher","unstructured":"Fei-Fei Li Marco Andreeto Marc\u2019Aurelio Ranzato and Pietro Perona. 2022. Caltech 101. 10.22002\/D1.20086","DOI":"10.22002\/D1.20086"},{"key":"e_1_3_3_2_25_2","unstructured":"Jiawei Li Fanrui Zhang Jiaying Zhu Esther Sun Qiang Zhang and Zheng-Jun Zha. 2024. Forgerygpt: Multimodal large language model for explainable image forgery detection and localization. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.10238 (2024)."},{"key":"e_1_3_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00853"},{"key":"e_1_3_3_2_27_2","unstructured":"Haotian Liu Chunyuan Li Qingyang Wu and Yong\u00a0Jae Lee. 2023. Visual instruction tuning. Advances in neural information processing systems 36 (2023) 34892\u201334916."},{"key":"e_1_3_3_2_28_2","unstructured":"Tingkai Liu Yunzhe Tao Haogeng Liu Qihang Fan Ding Zhou Huaibo Huang Ran He and Hongxia Yang. 2023. DeVAn: Dense Video Annotation for Video-Language Models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2310.05060 (2023)."},{"key":"e_1_3_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681089"},{"key":"e_1_3_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00808"},{"key":"e_1_3_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP46576.2022.9897310"},{"key":"e_1_3_3_2_32_2","doi-asserted-by":"crossref","first-page":"384","DOI":"10.1109\/MIPR.2018.00084","volume-title":"2018 IEEE conference on multimedia information processing and retrieval (MIPR)","author":"Marra Francesco","year":"2018","unstructured":"Francesco Marra, Diego Gragnaniello, Davide Cozzolino, and Luisa Verdoliva. 2018. Detection of gan-generated fake images over social networks. In 2018 IEEE conference on multimedia information processing and retrieval (MIPR). IEEE, 384\u2013389."},{"key":"e_1_3_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58610-2_6"},{"key":"e_1_3_3_2_34_2","unstructured":"Alec Radford and Karthik Narasimhan. 2018. Improving Language Understanding by Generative Pre-Training."},{"key":"e_1_3_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00104"},{"key":"e_1_3_3_2_37_2","unstructured":"Yixuan Su Tian Lan Huayang Li Jialu Xu Yan Wang and Deng Cai. 2023. Pandagpt: One model to instruction-follow them all. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2305.16355 (2023)."},{"key":"e_1_3_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV51458.2022.00323"},{"key":"e_1_3_3_2_39_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et\u00a0al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2302.13971 (2023)."},{"key":"e_1_3_3_2_40_2","unstructured":"Ming Wang Yuanzhong Liu Xiaoyu Liang Songlian Li Yijie Huang Xiaoming Zhang Sijia Shen Chaofeng Guan Daling Wang Shi Feng et\u00a0al. 2024. LangGPT: Rethinking structured reusable prompt design framework for LLMs from the programming language. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2402.16929 (2024)."},{"key":"e_1_3_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00872"},{"key":"e_1_3_3_2_42_2","unstructured":"Xuansheng Wu Jiayi Yuan Wenlin Yao Xiaoming Zhai and Ninghao Liu. 2025. Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2502.15576 (2025)."},{"key":"e_1_3_3_2_43_2","unstructured":"Zhipei Xu Xuanyu Zhang Runyi Li Zecheng Tang Qing Huang and Jian Zhang. 2024. Fakeshield: Explainable image forgery detection and localization via multi-modal large language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.02761 (2024)."},{"key":"e_1_3_3_2_44_2","doi-asserted-by":"crossref","unstructured":"Miaomiao Yu Sigang Ju Jun Zhang Shuohao Li Jun Lei and Xiaofei Li. 2022. Patch-DFD: Patch-based end-to-end DeepFake discriminator. Neurocomputing 501 (2022) 583\u2013595.","DOI":"10.1016\/j.neucom.2022.06.013"},{"key":"e_1_3_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00765"},{"key":"e_1_3_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/WIFS47025.2019.9035107"},{"key":"e_1_3_3_2_47_2","first-page":"337","volume-title":"PRICAI 2021: Trends in Artificial Intelligence: 18th Pacific Rim International Conference on Artificial Intelligence, PRICAI 2021, Hanoi, Vietnam, November 8\u201312, 2021, Proceedings, Part III 18","author":"Zhang Xueqi","year":"2021","unstructured":"Xueqi Zhang, Shuo Wang, Chenyu Liu, Min Zhang, Xiaohan Liu, and Haiyong Xie. 2021. Thinking in patch: Towards generalizable forgery detection with patch transformation. In PRICAI 2021: Trends in Artificial Intelligence: 18th Pacific Rim International Conference on Artificial Intelligence, PRICAI 2021, Hanoi, Vietnam, November 8\u201312, 2021, Proceedings, Part III 18. Springer, 337\u2013352."},{"key":"e_1_3_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00222"},{"key":"e_1_3_3_2_49_2","unstructured":"Deyao Zhu Jun Chen Xiaoqian Shen Xiang Li and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2304.10592 (2023)."},{"key":"e_1_3_3_2_50_2","unstructured":"Haochen Zhu Gang Cao and Xianglin Huang. 2023. Progressive feedback-enhanced transformer for image forgery localization. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2311.08910 (2023)."}],"event":{"name":"IH&MMSEC '25: ACM Workshop on Information Hiding and Multimedia Security","location":"San Jose USA","acronym":"IH&MMSEC '25","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the ACM Workshop on Information Hiding and Multimedia Security"],"original-title":[],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T11:14:50Z","timestamp":1750158890000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3733102.3733138"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,17]]},"references-count":49,"alternative-id":["10.1145\/3733102.3733138","10.1145\/3733102"],"URL":"https:\/\/doi.org\/10.1145\/3733102.3733138","relation":{},"subject":[],"published":{"date-parts":[[2025,6,17]]},"assertion":[{"value":"2025-06-17","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}