{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T13:08:24Z","timestamp":1784034504455,"version":"3.55.0"},"reference-count":264,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272297"],"award-info":[{"award-number":["62272297"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Shanghai Pujiang Program","award":["24PJA056"],"award-info":[{"award-number":["24PJA056"]}]},{"name":"Startup Fund for Young Faculty at SJTU"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,2,28]]},"abstract":"<jats:p>The rapid proliferation of AI-Generated Content (AIGC), spanning text, images, video, and audio, has created a dual-edged sword of unprecedented creativity and significant societal risks, including misinformation and disinformation. This survey provides a comprehensive and structured overview of the current landscape of AIGC detection technologies. We begin by chronicling the evolution of generative models, from foundational GANs to state-of-the-art diffusion and transformer-based architectures. We then systematically review detection methodologies across all modalities, organizing them into a novel taxonomy of External Detection and Internal Detection. For each modality, we trace the technical progression from early feature-based methods to advanced deep learning, while also covering critical tasks like model attribution and tampered region localization. Furthermore, we survey the ecosystem of publicly available detection tools and practical applications. Finally, we distill the primary challenges facing the field\u2013including generalization, robustness, interpretability, and the lack of universal benchmarks\u2013and conclude by outlining key future directions, such as the development of holistic AI Safety Agents, dynamic evaluation standards, and AI-driven governance frameworks. This survey aims to provide researchers and practitioners with a clear, in-depth understanding of the state-of-the-art and critical frontiers in the ongoing endeavor to ensure a safe and trustworthy AIGC ecosystem.<\/jats:p>","DOI":"10.1145\/3760526","type":"journal-article","created":{"date-parts":[[2025,8,13]],"date-time":"2025-08-13T11:15:47Z","timestamp":1755083747000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Advancements in AI-Generated Content Forensics: A Systematic Literature Review"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8750-7036","authenticated-orcid":false,"given":"Qiang","family":"Xu","sequence":"first","affiliation":[{"name":"School of Computer Science, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-3297-6576","authenticated-orcid":false,"given":"Wenpeng","family":"Mu","sequence":"additional","affiliation":[{"name":"School of Computer Science, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-8168-4038","authenticated-orcid":false,"given":"Jianing","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3253-5136","authenticated-orcid":false,"given":"Tanfeng","family":"Sun","sequence":"additional","affiliation":[{"name":"School of Computer Science, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9758-0579","authenticated-orcid":false,"given":"Xinghao","family":"Jiang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,9]]},"reference":[{"key":"e_1_3_1_2_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022","author":"Ouyang Long","year":"2022","unstructured":"Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2024.3440097"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-024-07487-w"},{"key":"e_1_3_1_5_2","unstructured":"Xai. 2024. Grok-2. (2024). Retrieved from https:\/\/grok2.cc\/"},{"key":"e_1_3_1_6_2","unstructured":"Google. 2025. gemini-2.5-pro. (2025). Retrieved from https:\/\/deepmind.google\/models\/gemini\/pro\/"},{"key":"e_1_3_1_7_2","unstructured":"European Commission. 2024. Artificial intelligence act. (2024). Retrieved from https:\/\/artificialintelligenceact.eu\/"},{"key":"e_1_3_1_8_2","unstructured":"The Council of Europe. 2024. The Framework Convention on Artificial Intelligence. (2024). Retrieved from https:\/\/www.coe.int\/en\/web\/artificial-intelligence\/the-framework-convention-on-artificial-intelligence"},{"key":"e_1_3_1_9_2","unstructured":"National Technical Committee. 2024. AI Safety Governance Framework. (2024). Retrieved from https:\/\/www.tc260.org.cn\/front\/postDetail.html?id=20240909102807"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.3389\/fpos.2025.1561776"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00549"},{"key":"e_1_3_1_12_2","unstructured":"Jingyi Deng Chenhao Lin Zhengyu Zhao Shuai Liu Qian Wang and Chao Shen. 2024. A survey of defenses against AI-generated visual media: Detection disruption and authentication. arXiv:2407.10575. Retrieved from https:\/\/arxiv.org\/abs\/2407.10575. (2024)."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2023.102103"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3703626"},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","unstructured":"Li Lin Neeraj Gupta Yue Zhang Hainan Ren Chun-Hao Liu Feng Ding Xin Wang Xin Li Luisa Verdoliva and Shu Hu. 2024. Detecting multimedia generated by large AI models: A survey. arXiv:2402.00045. Retrieved from https:\/\/arxiv.org\/abs\/2402.00045. (2024).","DOI":"10.36227\/techrxiv.170723324.44685515\/v1"},{"key":"e_1_3_1_16_2","unstructured":"Yueying Zou Peipei Li Zekun Li Huaibo Huang Xing Cui Xuannan Liu Chenghanyu Zhang and Ran He. 2025. Survey on AI-generated media detection: From non-MLLM to MLLM. arXiv:2502.05240. Retrieved from https:\/\/arxiv.org\/abs\/2502.05240 (2025)."},{"key":"e_1_3_1_17_2","first-page":"2672","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014","author":"Goodfellow Ian J.","year":"2014","unstructured":"Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In Proceedings of the Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014. 2672\u20132680."},{"key":"e_1_3_1_18_2","first-page":"5998","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017. 5998\u20136008."},{"key":"e_1_3_1_19_2","first-page":"8162","volume-title":"Proceedings of the 38th International Conference on Machine Learning, ICML 2021 (Proceedings of Machine Learning Research)","volume":"139","author":"Nichol Alexander Quinn","year":"2021","unstructured":"Alexander Quinn Nichol and Prafulla Dhariwal. 2021. Improved denoising diffusion probabilistic models. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021 (Proceedings of Machine Learning Research), Vol. 139. PMLR, 8162\u20138171."},{"key":"e_1_3_1_20_2","unstructured":"OpenAI. 2023. GPT-4. (2023). Retrieved from https:\/\/openai.com\/index\/gpt-4\/"},{"key":"e_1_3_1_21_2","unstructured":"Meta. 2025. LLaMA-4. (2025). Retrieved from https:\/\/www.llama.com\/models\/llama-4\/"},{"key":"e_1_3_1_22_2","unstructured":"DeepSeek. 2025. DeepSeek-R1. (2025). Retrieved from https:\/\/www.deepseek.com\/"},{"key":"e_1_3_1_23_2","unstructured":"Biyang Guo Xin Zhang Ziyuan Wang Minqi Jiang Jinran Nie Yuxuan Ding Jianwei Yue and Yupeng Wu. 2023. How close is ChatGPT to human experts? comparison corpus evaluation and detection. arXiv:2301.07597. Retrieved from https:\/\/arxiv.org\/abs\/2301.07597 (2023)."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TBDATA.2025.3536929"},{"key":"e_1_3_1_25_2","unstructured":"Zhenpeng Su Xing Wu Wei Zhou Guangyuan Ma and Songlin Hu. 2023. HC3 plus: A semantic-invariant human ChatGPT comparison corpus. arXiv:2309.02731. Retrieved from https:\/\/arxiv.org\/abs\/2309.02731. (2023)."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.810"},{"key":"e_1_3_1_27_2","unstructured":"Chujie Gao Dongping Chen Qihui Zhang Yue Huang Yao Wan and Lichao Sun. 2024. LLM-as-a-Coauthor: The challenges of detecting LLM-human mixcase. arXiv:2401.05952. Retrieved from https:\/\/arxiv.org\/abs\/2401.05952. (2024)."},{"key":"e_1_3_1_28_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024","author":"Wu Junchao","year":"2024","unstructured":"Junchao Wu, Runzhe Zhan, Derek F. Wong, Shu Yang, Xinyi Yang, Yulin Yuan, and Lidia S. Chao. 2024. DetectRL: Benchmarking LLM-generated text detection in real-world scenarios. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.naacl-long.444"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3658644.3670392"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3696410.3714770"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.51"},{"key":"e_1_3_1_33_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023","author":"Zhu Mingjian","year":"2023","unstructured":"Mingjian Zhu, Hanting Chen, Qiangyu Yan, Xudong Huang, Guanyu Lin, Wei Li, Zhijun Tu, Hailin Hu, Jie Hu, and Yunhe Wang. 2023. GenImage: A million-scale benchmark for detecting AI-generated image. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3576915.3616588"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00239"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2024.3356122"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i4.32363"},{"key":"e_1_3_1_38_2","unstructured":"Shilin Yan Ouxiang Li Jiayin Cai Yanbin Hao Xiaolong Jiang Yao Hu and Weidi Xie. 2024. A sanity check for AI-generated image detection. arXiv:2406.19435. Retrieved from https:\/\/arxiv.org\/abs\/2406.19435. (2024)."},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","unstructured":"Zhaopan Xu Pengfei Zhou Jiaxin Ai Wangbo Zhao Kai Wang Xiaojiang Peng Wenqi Shao Hongxun Yao and Kaipeng Zhang. 2025. MPBench: A comprehensive multimodal reasoning benchmark for process errors identification. arXiv:2503.12505. Retrieved from https:\/\/arxiv.org\/abs\/2503.12505. (2025).","DOI":"10.18653\/v1\/2025.findings-acl.1112"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-024-02255-9"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00009"},{"key":"e_1_3_1_42_2","unstructured":"Brian Dolhansky Russ Howes Ben Pflaum Nicole Baram and Cristian Canton-Ferrer. 2019. The deepfake detection challenge (DFDC) preview dataset. arXiv:1910.08854. Retrieved from https:\/\/arxiv.org\/abs\/1910.08854. (2019)."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00327"},{"key":"e_1_3_1_44_2","unstructured":"Haoxing Chen Yan Hong Zizheng Huang Zhuoer Xu Zhangxuan Gu Yaohui Li Jun Lan Huijia Zhu Jianfu Zhang Weiqiang Wang and Huaxiong Li. 2024. DeMamba: AI-generated video detection on million-scale genvideo benchmark. arXiv:2405.19707. Retrieved from https:\/\/arxiv.org\/abs\/2405.19707. (2024)."},{"key":"e_1_3_1_45_2","volume-title":"Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024","author":"Ju Xuan","year":"2024","unstructured":"Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan, Xintao Wang, Ailing Zeng, Yu Xiong, Qiang Xu, and Ying Shan. 2024. MiraData: A large-scale video dataset with long durations and structured captions. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024."},{"key":"e_1_3_1_46_2","unstructured":"Zhenliang Ni Qiangyu Yan Mouxiao Huang Tianning Yuan Yehui Tang Hailin Hu Xinghao Chen and Yunhe Wang. 2025. GenVidBench: A challenging benchmark for detecting AI-generated video. arXiv:2501.11340. Retrieved from https:\/\/arxiv.org\/abs\/2501.11340. (2025)."},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-2249"},{"key":"e_1_3_1_48_2","doi-asserted-by":"crossref","unstructured":"Junichi Yamagishi Xin Wang Massimiliano Todisco Md. Sahidullah Jose Patino Andreas Nautsch Xuechen Liu Kong Aik Lee Tomi Kinnunen Nicholas W. D. Evans and H\u00e9ctor Delgado. 2021. ASVspoof 2021: Accelerating progress in spoofed and deepfake speech detection. arXiv:2109.00537. Retrieved from https:\/\/arxiv.org\/abs\/2109.00537. (2021).","DOI":"10.21437\/ASVSPOOF.2021-8"},{"key":"e_1_3_1_49_2","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021","author":"Frank Joel","year":"2021","unstructured":"Joel Frank and Lea Sch\u00f6nherr. 2021. WaveFake: A data set to facilitate audio deepfake detection. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021."},{"key":"e_1_3_1_50_2","unstructured":"Xiang Li Pin-Yu Chen and Wenqi Wei. 2024. SONAR: A synthetic AI-audio detection framework and benchmark. arXiv:2410.04324. Retrieved from https:\/\/arxiv.org\/abs\/2410.04324. (2024)."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN60899.2024.10650962"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.dib.2024.110743"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3680795"},{"key":"e_1_3_1_54_2","unstructured":"Luca Comanducci Paolo Bestagini and Stefano Tubaro. 2024. FakeMusicCaps: A dataset for detection and attribution of synthetic music generated via text-to-music models. arXiv:2409.10684. Retrieved from https:\/\/arxiv.org\/abs\/2409.10684. (2024)."},{"key":"e_1_3_1_55_2","unstructured":"Md Awsafur Rahman Zaber Ibn Abdul Hakim Najibul Haque Sarker Bishmoy Paul and Shaikh Anowarul Fattah. 2024. SONICS: Synthetic or not - identifying counterfeit songs. arXiv:2408.14080. Retrieved from https:\/\/arxiv.org\/abs\/2408.14080. (2024)."},{"key":"e_1_3_1_56_2","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021","author":"Khalid Hasam","year":"2021","unstructured":"Hasam Khalid, Shahroz Tariq, Minha Kim, and Simon S. Woo. 2021. FakeAVCeleb: A novel audio-video multimodal deepfake dataset. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021."},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACVW58289.2023.00071"},{"key":"e_1_3_1_58_2","doi-asserted-by":"crossref","unstructured":"Zhixi Cai Kalin Stefanov Abhinav Dhall and Munawar Hayat. 2022. Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization. arXiv:2204.06228. Retrieved from https:\/\/arxiv.org\/abs\/2204.06228. (2022).","DOI":"10.1109\/DICTA56598.2022.10034605"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-emnlp.959"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1609\/icwsm.v18i1.31437"},{"key":"e_1_3_1_61_2","first-page":"180","volume-title":"Proceedings of the Pattern Recognition - 27th International Conference, ICPR 2024 (Lecture Notes in Computer Science)","volume":"15314","author":"Hou Yang","year":"2024","unstructured":"Yang Hou, Haitao Fu, Chunkai Chen, Zida Li, Haoyu Zhang, and Jianjun Zhao. 2024. PolyGlotFake: A novel multilingual and multimodal deepfake dataset. In Proceedings of the Pattern Recognition - 27th International Conference, ICPR 2024 (Lecture Notes in Computer Science), Vol. 15314. Springer, 180\u2013193."},{"key":"e_1_3_1_62_2","volume-title":"Proceedings of the 13th International Conference on Learning Representations, ICLR 2025","author":"Thakral Kartik","year":"2025","unstructured":"Kartik Thakral, Rishabh Ranjan, Akanksha Singh, Akshat Jain, Mayank Vatsa, and Richa Singh. 2025. ILLUSION: Unveiling truth with a comprehensive multi-modal, multi-lingual deepfake dataset. In Proceedings of the 13th International Conference on Learning Representations, ICLR 2025. OpenReview.net."},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3701716.3715306"},{"key":"e_1_3_1_64_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. In Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020."},{"key":"e_1_3_1_65_2","first-page":"11895","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019","author":"Song Yang","year":"2019","unstructured":"Yang Song and Stefano Ermon. 2019. Generative modeling by estimating gradients of the data distribution. In Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019. 11895\u201311907."},{"key":"e_1_3_1_66_2","volume-title":"Proceedings of the 9th International Conference on Learning Representations, ICLR 2021","author":"Song Yang","year":"2021","unstructured":"Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2021. Score-based generative modeling through stochastic differential equations. In Proceedings of the 9th International Conference on Learning Representations, ICLR 2021. OpenReview.net."},{"key":"e_1_3_1_67_2","unstructured":"OpenAI. 2023. DALL \\(\\cdot\\) E 3. (2023). Retrieved from https:\/\/openai.com\/index\/dall-e-3\/"},{"key":"e_1_3_1_68_2","unstructured":"Google. 2024. Imagen. (2024). Retrieved from https:\/\/deepmind.google\/models\/imagen\/"},{"key":"e_1_3_1_69_2","unstructured":"Midjourney. 2022. Midjourney. (2022). Retrieved from https:\/\/www.midjourney.com\/"},{"key":"e_1_3_1_70_2","unstructured":"Amazon. 2022. Stable Diffusion. (2022). Retrieved from https:\/\/aws.amazon.com\/cn\/campaigns\/stable-diffusion\/"},{"key":"e_1_3_1_71_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations, ICLR 2023","author":"Singer Uriel","year":"2023","unstructured":"Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. 2023. Make-a-video: Text-to-video generation without text-video data. In Proceedings of the 11th International Conference on Learning Representations, ICLR 2023. OpenReview.net."},{"key":"e_1_3_1_72_2","doi-asserted-by":"crossref","first-page":"205","DOI":"10.1007\/978-3-031-73033-7_12","volume-title":"Proceedings of the Computer Vision - ECCV 2024 (Lecture Notes in Computer Science)","volume":"15120","author":"Girdhar Rohit","year":"2024","unstructured":"Rohit Girdhar, Mannat Singh, Andrew Brown, Quentin Duval, Samaneh Azadi, Sai Saketh Rambhatla, Akbar Shah, Xi Yin, Devi Parikh, and Ishan Misra. 2024. Factorizing text-to-video generation by explicit image conditioning. In Proceedings of the Computer Vision - ECCV 2024 (Lecture Notes in Computer Science), Vol. 15120. Springer, 205\u2013224."},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-demo.49"},{"key":"e_1_3_1_74_2","unstructured":"Andreas Blattmann Tim Dockhorn Sumith Kulal Daniel Mendelevitch Maciej Kilian Dominik Lorenz Yam Levi Zion English Vikram Voleti Adam Letts Varun Jampani and Robin Rombach. 2023. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv:2311.15127. Retrieved from https:\/\/arxiv.org\/abs\/2311.15127. (2023)."},{"key":"e_1_3_1_75_2","unstructured":"OpenAI. 2024. Sora. (2024). Retrieved from https:\/\/openai.com\/index\/sora\/"},{"key":"e_1_3_1_76_2","unstructured":"Goodle. 2025. Veo. (2025). Retrieved from https:\/\/deepmind.google\/models\/veo\/"},{"key":"e_1_3_1_77_2","unstructured":"Runway. 2025. Gen-4. (2025). Retrieved from https:\/\/runwayml.com\/research\/introducing-runway-gen-4"},{"key":"e_1_3_1_78_2","unstructured":"Vidu. 2025. Vidu. (2025). Retrieved from https:\/\/www.vidu.cn\/"},{"key":"e_1_3_1_79_2","unstructured":"Alibaba. 2025. Wan2.1. (2025). Retrieved from https:\/\/wan2.video\/"},{"key":"e_1_3_1_80_2","unstructured":"Xiaoran Fan Chao Pang Tian Yuan He Bai Renjie Zheng Pengfei Zhu Shuohuan Wang Jun-Kun Chen Zeyu Chen Liang Huang Yu Sun and Hua Wu. 2022. ERNIE-SAT: Speech and text joint pretraining for cross-lingual multi-speaker text-to-speech. arXiv:2211.03545. Retrieved from https:\/\/arxiv.org\/abs\/2211.03545. (2022)."},{"key":"e_1_3_1_81_2","unstructured":"Keyu An Qian Chen Chong Deng Zhihao Du Changfeng Gao Zhifu Gao Yue Gu Ting He Hangrui Hu Kai Hu Shengpeng Ji Yabin Li Zerui Li Heng Lu Haoneng Luo Xiang Lv Bin Ma Ziyang Ma Chongjia Ni Changhe Song Jiaqi Shi Xian Shi Hao Wang Wen Wang Yuxuan Wang Zhangyu Xiao Zhijie Yan Yexin Yang Bin Zhang Qinglin Zhang Shiliang Zhang Nan Zhao and Siqi Zheng. 2024. FunAudioLLM: Voice understanding and generation foundation models for natural interaction between humans and LLMs. arXiv:2407.04051. Retrieved from https:\/\/arxiv.org\/abs\/2407.04051. (2024)."},{"key":"e_1_3_1_82_2","unstructured":"Suno. 2024. Suno. (2024). Retrieved from https:\/\/suno.com\/blog\/v3"},{"key":"e_1_3_1_83_2","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3594095"},{"key":"e_1_3_1_84_2","doi-asserted-by":"publisher","DOI":"10.1145\/3600211.3604722"},{"key":"e_1_3_1_85_2","first-page":"1125","volume-title":"Proceedings of the 37th Annual Conference on Learning Theory, June 30 - July 3, 2023, Edmonton, Canada (Proceedings of Machine Learning Research)","volume":"247","author":"Christ Miranda","year":"2024","unstructured":"Miranda Christ, Sam Gunn, and Or Zamir. 2024. Undetectable watermarks for language models. In Proceedings of the 37th Annual Conference on Learning Theory, June 30 - July 3, 2023, Edmonton, Canada (Proceedings of Machine Learning Research), Vol. 247. PMLR, 1125\u20131139."},{"key":"e_1_3_1_86_2","article-title":"Robust distortion-free watermarks for language models","volume":"2024","author":"Kuditipudi Rohith","year":"2024","unstructured":"Rohith Kuditipudi, John Thickstun, Tatsunori Hashimoto, and Percy Liang. 2024. Robust distortion-free watermarks for language models. Trans. Mach. Learn. Res. 2024 (2024).","journal-title":"Trans. Mach. Learn. Res."},{"key":"e_1_3_1_87_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, ICLR 2024","author":"Zhao Xuandong","year":"2024","unstructured":"Xuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, and Yu-Xiang Wang. 2024. Provable robust watermarking for AI-generated text. In Proceedings of the 12th International Conference on Learning Representations, ICLR 2024. OpenReview.net."},{"key":"e_1_3_1_88_2","unstructured":"Aiwei Liu Leyi Pan Xuming Hu Shu\u2019ang Li Lijie Wen Irwin King and Philip S. Yu. 2023. A private watermark for large language models. arXiv:2307.16230. Retrieved from https:\/\/arxiv.org\/abs\/2307.16230. (2023)."},{"key":"e_1_3_1_89_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11704-024-40751-w"},{"key":"e_1_3_1_90_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-024-08025-4"},{"key":"e_1_3_1_91_2","unstructured":"Xiaojun Xu Yuanshun Yao and Yang Liu. 2024. Learning to watermark LLM-generated text via reinforcement learning. arXiv:2403.10553. Retrieved from https:\/\/arxiv.org\/abs\/2403.10553. (2024)."},{"key":"e_1_3_1_92_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.630"},{"key":"e_1_3_1_93_2","unstructured":"Matthias Gall\u00e9 Jos Rozen Germ\u00e1n Kruszewski and Hady Elsahar. 2021. Unsupervised and distributional detection of machine-generated text. arXiv:2111.02878. Retrieved from https:\/\/arxiv.org\/abs\/2111.02878. (2021)."},{"key":"e_1_3_1_94_2","unstructured":"Tharindu Kumarage Joshua Garland Amrita Bhattacharjee Kirill Trapeznikov Scott W. Ruston and Huan Liu. 2023. Stylometric detection of AI-generated text in twitter timelines. arXiv:2303.03697. Retrieved from https:\/\/arxiv.org\/abs\/2303.03697. (2023)."},{"key":"e_1_3_1_95_2","first-page":"24950","volume-title":"Proceedings of the International Conference on Machine Learning, ICML 2023 (Proceedings of Machine Learning Research)","volume":"202","author":"Mitchell Eric","year":"2023","unstructured":"Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, and Chelsea Finn. 2023. DetectGPT: Zero-shot machine-generated text detection using probability curvature. In Proceedings of the International Conference on Machine Learning, ICML 2023 (Proceedings of Machine Learning Research), Vol. 202. PMLR, 24950\u201324962."},{"key":"e_1_3_1_96_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024","author":"Guo Xun","year":"2024","unstructured":"Xun Guo, Yongxin He, Shan Zhang, Ting Zhang, Wanquan Feng, Haibin Huang, and Chongyang Ma. 2024. DeTeCtive: Detecting AI-generated text via multi-level contrastive learning. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024."},{"key":"e_1_3_1_97_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, ICLR 2024","author":"Soto Rafael A. Rivera","year":"2024","unstructured":"Rafael A. Rivera Soto, Kailin Koch, Aleem Khan, Barry Y. Chen, Marcus Bishop, and Nicholas Andrews. 2024. Few-shot detection of machine-generated text using style representations. In Proceedings of the 12th International Conference on Learning Representations, ICLR 2024. OpenReview.net."},{"key":"e_1_3_1_98_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.495"},{"key":"e_1_3_1_99_2","doi-asserted-by":"crossref","unstructured":"Zhen Tao Zhiyu Li Runyu Chen Dinghao Xi and Wei Xu. 2024. Unveiling large language models generated texts: A multi-level fine-grained detection framework. arXiv:2410.14231. Retrieved from https:\/\/arxiv.org\/abs\/2410.14231. (2024).","DOI":"10.2139\/ssrn.4999806"},{"key":"e_1_3_1_100_2","unstructured":"Ram Mohan Rao Kadiyala Siddartha Pullakhandam Kanwal Mehreen Drishti Sharma Siddhant Gupta Jebish Purbey Ashay Srivastava Subhasya Tippareddy Arvind Reddy Bobbili Suraj Telugara Chandrashekhar Modabbir Adeeb Srinadh Vura Suman Debnath and Hamza Farooq. 2025. Robust and fine-grained detection of AI generated texts. arXiv:2504.11952. Retrieved from https:\/\/arxiv.org\/abs\/2504.11952. (2025)."},{"key":"e_1_3_1_101_2","first-page":"474","volume-title":"Proceedings of the Computer Vision - ECCV 2024 (Lecture Notes in Computer Science)","volume":"15067","author":"Hua Hang","year":"2024","unstructured":"Hang Hua, Jing Shi, Kushal Kafle, Simon Jenni, Daoan Zhang, John P. Collomosse, Scott Cohen, and Jiebo Luo. 2024. FineMatch: Aspect-based fine-grained image and text mismatch detection and correction. In Proceedings of the Computer Vision - ECCV 2024 (Lecture Notes in Computer Science), Vol. 15067. Springer, 474\u2013491."},{"key":"e_1_3_1_102_2","unstructured":"Linyang Li Pengyu Wang Ke Ren Tianxiang Sun and Xipeng Qiu. 2023. Origin tracing and detecting of LLMs. arXiv:2304.14072. Retrieved from https:\/\/arxiv.org\/abs\/2304.14072. (2023)."},{"key":"e_1_3_1_103_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-naacl.8"},{"key":"e_1_3_1_104_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.aacl-main.84"},{"key":"e_1_3_1_105_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.218"},{"key":"e_1_3_1_106_2","doi-asserted-by":"publisher","DOI":"10.1145\/3716846"},{"key":"e_1_3_1_107_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.68"},{"key":"e_1_3_1_108_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.854"},{"key":"e_1_3_1_109_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01274"},{"key":"e_1_3_1_110_2","article-title":"Reducing contextual hallucinations in large language models through attention map optimization","author":"Ainsworth Eloise","year":"2024","unstructured":"Eloise Ainsworth, Justin Wycliffe, and Florence Winslow. 2024. Reducing contextual hallucinations in large language models through attention map optimization. Authorea Preprints (2024).","journal-title":"Authorea Preprints"},{"key":"e_1_3_1_111_2","unstructured":"Neeraj Varshney Wenlin Yao Hongming Zhang Jianshu Chen and Dong Yu. 2023. A stitch in time saves nine: Detecting and mitigating hallucinations of LLMs by validating low-confidence generation. arXiv:2307.03987. Retrieved from https:\/\/arxiv.org\/abs\/2307.03987. (2023)."},{"key":"e_1_3_1_112_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i18.29991"},{"key":"e_1_3_1_113_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.557"},{"key":"e_1_3_1_114_2","volume-title":"Proceedings of the 41st International Conference on Machine Learning, ICML 2024","author":"Du Yilun","year":"2024","unstructured":"Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. 2024. Improving factuality and reasoning in language models through multiagent debate. In Proceedings of the 41st International Conference on Machine Learning, ICML 2024. OpenReview.net."},{"key":"e_1_3_1_115_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.212"},{"key":"e_1_3_1_116_2","doi-asserted-by":"crossref","unstructured":"Xiao Fang Shangkun Che Minjia Mao Hongzhe Zhang Ming Zhao and Xiaohang Zhao. 2023. Bias of AI-generated content: An examination of news produced by large language models. arXiv:2309.09825. Retrieved from https:\/\/arxiv.org\/abs\/2309.09825. (2023).","DOI":"10.21203\/rs.3.rs-3499674\/v1"},{"key":"e_1_3_1_117_2","doi-asserted-by":"publisher","DOI":"10.1038\/s43588-025-00789-7"},{"key":"e_1_3_1_118_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.121542"},{"key":"e_1_3_1_119_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSS.2024.3392469"},{"key":"e_1_3_1_120_2","doi-asserted-by":"publisher","DOI":"10.1007\/s43681-024-00568-6"},{"key":"e_1_3_1_121_2","unstructured":"Tiffany Zhu Iain Weissburg Kexun Zhang and William Yang Wang. 2024. Human bias in the face of AI: The role of human judgement in AI generated text evaluation. arXiv:2410.03723. Retrieved from https:\/\/arxiv.org\/abs\/2410.03723. (2024)."},{"key":"e_1_3_1_122_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCB48548.2020.9304936"},{"key":"e_1_3_1_123_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2024.108341"},{"key":"e_1_3_1_124_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3390945"},{"key":"e_1_3_1_125_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3313503"},{"key":"e_1_3_1_126_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2023.3335359"},{"key":"e_1_3_1_127_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2023.3269152"},{"key":"e_1_3_1_128_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01165"},{"key":"e_1_3_1_129_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10094654"},{"key":"e_1_3_1_130_2","doi-asserted-by":"publisher","DOI":"10.1145\/3558004"},{"key":"e_1_3_1_131_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00858"},{"key":"e_1_3_1_132_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2023.3324739"},{"key":"e_1_3_1_133_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-43153-1_29"},{"key":"e_1_3_1_134_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICICML63543.2024.10957875"},{"key":"e_1_3_1_135_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01974"},{"key":"e_1_3_1_136_2","doi-asserted-by":"publisher","DOI":"10.7717\/peerj-cs.2775"},{"key":"e_1_3_1_137_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00308"},{"key":"e_1_3_1_138_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01137"},{"key":"e_1_3_1_139_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3680776"},{"key":"e_1_3_1_140_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP48485.2024.10446824"},{"key":"e_1_3_1_141_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00661"},{"key":"e_1_3_1_142_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01027"},{"key":"e_1_3_1_143_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.410"},{"key":"e_1_3_1_144_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10095167"},{"key":"e_1_3_1_145_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW59228.2023.00104"},{"key":"e_1_3_1_146_2","doi-asserted-by":"publisher","DOI":"10.1109\/OJSP.2023.3337714"},{"key":"e_1_3_1_147_2","doi-asserted-by":"publisher","DOI":"10.7717\/peerj-cs.2127"},{"key":"e_1_3_1_148_2","first-page":"17006","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024","author":"Luo Yunpeng","year":"2024","unstructured":"Yunpeng Luo, Junlong Du, Ke Yan, and Shouhong Ding. 2024. LaRE\u2303 2: Latent reconstruction error based method for diffusion-generated image detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024. IEEE, 17006\u201317015."},{"key":"e_1_3_1_149_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.02051"},{"key":"e_1_3_1_150_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW63382.2024.00439"},{"key":"e_1_3_1_151_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024","author":"Du Xuefeng","year":"2024","unstructured":"Xuefeng Du, Chaowei Xiao, and Sharon Li. 2024. HaloScope: Harnessing unlabeled LLM generations for hallucination detection. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024."},{"key":"e_1_3_1_152_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-emnlp.262"},{"key":"e_1_3_1_153_2","first-page":"125","volume-title":"Proceedings of the Computer Vision - ECCV 2024 (Lecture Notes in Computer Science)","volume":"15141","author":"Liu Shi","year":"2024","unstructured":"Shi Liu, Kecheng Zheng, and Wei Chen. 2024. Paying more attention to image: A training-free method for alleviating hallucination in LVLMs. In Proceedings of the Computer Vision - ECCV 2024 (Lecture Notes in Computer Science), Vol. 15141. Springer, 125\u2013140."},{"key":"e_1_3_1_154_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00283"},{"key":"e_1_3_1_155_2","doi-asserted-by":"publisher","DOI":"10.1109\/TTS.2024.3365421"},{"key":"e_1_3_1_156_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00459"},{"key":"e_1_3_1_157_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01591"},{"key":"e_1_3_1_158_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.image.2023.117010"},{"key":"e_1_3_1_159_2","doi-asserted-by":"publisher","DOI":"10.1145\/3588574"},{"key":"e_1_3_1_160_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3281448"},{"key":"e_1_3_1_161_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022","author":"Guan Jiazhi","year":"2022","unstructured":"Jiazhi Guan, Hang Zhou, Zhibin Hong, Errui Ding, Jingdong Wang, Chengbin Quan, and Youjian Zhao. 2022. Delving into sequential patches for deepfake detection. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022."},{"key":"e_1_3_1_162_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW63382.2024.00443"},{"key":"e_1_3_1_163_2","unstructured":"Rohit Kundu Hao Xiong Vishal Mohanty Athula Balachandran and Amit K. Roy-Chowdhury. 2024. Towards a universal synthetic video detector: From face or background manipulations to fully AI-generated content. arXiv:2412.12278. Retrieved from https:\/\/arxiv.org\/abs\/2412.12278. (2024)."},{"key":"e_1_3_1_164_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00114"},{"key":"e_1_3_1_165_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.108832"},{"key":"e_1_3_1_166_2","doi-asserted-by":"publisher","DOI":"10.1145\/3512527.3531415"},{"key":"e_1_3_1_167_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2024.3433596"},{"key":"e_1_3_1_168_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICME51207.2021.9428361"},{"key":"e_1_3_1_169_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2023.01.001"},{"key":"e_1_3_1_170_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3217950"},{"key":"e_1_3_1_171_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00471"},{"key":"e_1_3_1_172_2","unstructured":"Zhixuan Chu Lei Zhang Yichen Sun Siqiao Xue Zhibo Wang Zhan Qin and Kui Ren. 2024. Sora detector: A unified hallucination detection for large text-to-video models. arXiv:2405.04180. Retrieved from https:\/\/arxiv.org\/abs\/2405.04180. (2024)."},{"key":"e_1_3_1_173_2","unstructured":"Ruiyang Zhang Hu Zhang and Zhedong Zheng. 2024. VL-uncertainty: Detecting hallucination in large vision-language model via uncertainty estimation. arXiv:2411.11919. Retrieved from https:\/\/arxiv.org\/abs\/2411.11919. (2024)."},{"key":"e_1_3_1_174_2","unstructured":"Jiacheng Zhang Yang Jiao Shaoxiang Chen Jingjing Chen and Yu-Gang Jiang. 2024. EventHallusion: Diagnosing event hallucinations in video LLMs. arXiv:2409.16597. Retrieved from https:\/\/arxiv.org\/abs\/2409.16597. (2024)."},{"key":"e_1_3_1_175_2","unstructured":"Chaoyu Li Eun Woo Im and Pooyan Fazli. 2024. VidHalluc: Evaluating temporal hallucinations in multimodal large language models for video understanding. arXiv:2412.03735. Retrieved from https:\/\/arxiv.org\/abs\/2412.03735. (2024)."},{"key":"e_1_3_1_176_2","unstructured":"Zongxia Li Xiyang Wu Yubin Qin Guangyao Shi Hongyang Du Dinesh Manocha Tianyi Zhou and Jordan Lee Boyd-Graber. 2025. VideoHallu: Evaluating and mitigating multi-modal hallucinations for synthetic videos. arXiv:2505.01481. Retrieved from https:\/\/arxiv.org\/abs\/2505.01481. (2025)."},{"key":"e_1_3_1_177_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cose.2024.103860"},{"key":"e_1_3_1_178_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681673"},{"key":"e_1_3_1_179_2","unstructured":"Leqi Shen Tao He Guoqiang Gong Fan Yang Yifeng Zhang Pengzhang Liu Sicheng Zhao and Guiguang Ding. 2025. LLaVA-MLB: Mitigating and leveraging attention bias for training-free video LLMs. arXiv:2503.11205. Retrieved from https:\/\/arxiv.org\/abs\/2503.11205. (2025)."},{"key":"e_1_3_1_180_2","doi-asserted-by":"publisher","DOI":"10.1145\/3552466.3556526"},{"key":"e_1_3_1_181_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2024.106320"},{"key":"e_1_3_1_182_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW59228.2023.00097"},{"key":"e_1_3_1_183_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10095927"},{"key":"e_1_3_1_184_2","doi-asserted-by":"publisher","DOI":"10.1109\/MMSP59012.2023.10337724"},{"key":"e_1_3_1_185_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-1768"},{"key":"e_1_3_1_186_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2023-1820"},{"key":"e_1_3_1_187_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414234"},{"key":"e_1_3_1_188_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2023-542"},{"key":"e_1_3_1_189_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2023-1206"},{"key":"e_1_3_1_190_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9747766"},{"key":"e_1_3_1_191_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10096741"},{"key":"e_1_3_1_192_2","doi-asserted-by":"publisher","DOI":"10.1145\/3595916.3626406"},{"key":"e_1_3_1_193_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP48485.2024.10447923"},{"key":"e_1_3_1_194_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i17.29929"},{"key":"e_1_3_1_195_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP48485.2024.10446270"},{"key":"e_1_3_1_196_2","unstructured":"Chun-Yi Kuan and Hung-yi Lee. 2024. Can large audio-language models truly hear? tackling hallucinations with multi-task assessment and stepwise audio reasoning. arXiv:2410.16130. Retrieved from https:\/\/arxiv.org\/abs\/2410.16130. (2024)."},{"key":"e_1_3_1_197_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2024-1076"},{"key":"e_1_3_1_198_2","unstructured":"Chun-Yi Kuan and Hung-yi Lee. 2025. Teaching audio-aware large language models what does not hear: Mitigating hallucinations through synthesized negative samples. arXiv:2505.14518. Retrieved from https:\/\/arxiv.org\/abs\/2505.14518. (2025)."},{"key":"e_1_3_1_199_2","unstructured":"Tzu-wen Hsu Ke-Han Lu Cheng-Han Chiang and Hung-yi Lee. 2025. Reducing object hallucination in large audio-language models via audio-aware decoding. arXiv:2506.07233. Retrieved from https:\/\/arxiv.org\/abs\/2506.07233. (2025)."},{"key":"e_1_3_1_200_2","first-page":"4418","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 - Workshops","author":"Yadav Amit Kumar Singh","year":"2024","unstructured":"Amit Kumar Singh Yadav, Kratika Bhagtani, Davide Salvi, Paolo Bestagini, and Edward J. Delp. 2024. FairSSD: Understanding bias in synthetic speech detectors. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 - Workshops. IEEE, 4418\u20134428."},{"key":"e_1_3_1_201_2","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533089"},{"key":"e_1_3_1_202_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024","author":"Ma Jie","year":"2024","unstructured":"Jie Ma, Min Hu, Pinghui Wang, Wangchun Sun, Lingyun Song, Hongbin Pei, Jun Liu, and Youtian Du. 2024. Look, listen, and answer: Overcoming biases for audio-visual question answering. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024."},{"key":"e_1_3_1_203_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2025.113417"},{"key":"e_1_3_1_204_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3229966"},{"key":"e_1_3_1_205_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.37"},{"key":"e_1_3_1_206_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2022.3231338"},{"key":"e_1_3_1_207_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2024.102715"},{"key":"e_1_3_1_208_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681089"},{"key":"e_1_3_1_209_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i1.31977"},{"key":"e_1_3_1_210_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i1.32036"},{"key":"e_1_3_1_211_2","unstructured":"Zhipei Xu Xuanyu Zhang Runyi Li Zecheng Tang Qing Huang and Jian Zhang. 2024. FakeShield: Explainable image forgery detection and localization via multi-modal large language models. arXiv:2410.02761. Retrieved from https:\/\/arxiv.org\/abs\/2410.02761. (2024)."},{"key":"e_1_3_1_212_2","unstructured":"Jiawei Li Fanrui Zhang Jiaying Zhu Esther Sun Qiang Zhang and Zheng-Jun Zha. 2024. ForgeryGPT: Multimodal large language model for explainable image forgery detection and localization. arXiv:2410.10238. Retrieved from https:\/\/arxiv.org\/abs\/2410.10238. (2024)."},{"key":"e_1_3_1_213_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-024-02245-x"},{"key":"e_1_3_1_214_2","unstructured":"Yize Chen Zhiyuan Yan Siwei Lyu and Baoyuan Wu. 2024. X \\({}^{\\mbox{2}}\\) -DFD: A framework for eXplainable and eXtendable Deepfake Detection. arXiv:2410.06126. Retrieved from https:\/\/arxiv.org\/abs\/2410.06126. (2024)."},{"key":"e_1_3_1_215_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01240"},{"key":"e_1_3_1_216_2","unstructured":"Sahibzada Adil Shahzad Ammarah Hashmi Yan-Tsung Peng Yu Tsao and Hsin-Min Wang. 2023. AV-Lip-Sync+: Leveraging AV-HuBERT to exploit multimodal inconsistency for video deepfake detection. arXiv:2311.02733. Retrieved from https:\/\/arxiv.org\/abs\/2311.02733. (2023)."},{"key":"e_1_3_1_217_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2023.3262148"},{"key":"e_1_3_1_218_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3326694"},{"key":"e_1_3_1_219_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10095247"},{"key":"e_1_3_1_220_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024","author":"Liang Yachao","year":"2024","unstructured":"Yachao Liang, Min Yu, Gang Li, Jianguo Jiang, Boquan Li, Feng Yu, Ning Zhang, Xiang Meng, and Weiqing Huang. 2024. SpeechForensics: Audio-visual speech representation learning for face forgery detection. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024."},{"key":"e_1_3_1_221_2","doi-asserted-by":"publisher","DOI":"10.1145\/3591106.3592218"},{"key":"e_1_3_1_222_2","unstructured":"Keyang Xuan Li Yi Fan Yang Ruochen Wu Yi R. Fung and Heng Ji. 2024. LEMMA: Towards LVLM-enhanced multimodal misinformation detection with external knowledge augmentation. arXiv:2402.11943. Retrieved from https:\/\/arxiv.org\/abs\/2402.11943. (2024)."},{"key":"e_1_3_1_223_2","doi-asserted-by":"publisher","DOI":"10.1145\/3627673.3679826"},{"key":"e_1_3_1_224_2","unstructured":"Yiran He Yun Cao Bowen Yang and Zeyu Zhang. 2025. Can GPT tell us why these images are synthesized? empowering multimodal large language models for forensics. arXiv:2504.11686. Retrieved from https:\/\/arxiv.org\/abs\/2504.11686. (2025)."},{"key":"e_1_3_1_225_2","first-page":"160","volume-title":"Proceedings of the Pattern Recognition - 27th International Conference, ICPR 2024 (Lecture Notes in Computer Science)","volume":"15321","author":"Keita Mamadou","year":"2024","unstructured":"Mamadou Keita, Wassim Hamidouche, Hessen Bougueffa Eutamene, Abdelmalik Taleb-Ahmed, and Abdenour Hadid. 2024. FIDAVL: Fake image detection and attribution using vision-language model. In Proceedings of the Pattern Recognition - 27th International Conference, ICPR 2024 (Lecture Notes in Computer Science), Vol. 15321. Springer, 160\u2013176."},{"key":"e_1_3_1_226_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01452"},{"key":"e_1_3_1_227_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01230"},{"key":"e_1_3_1_228_2","unstructured":"Zhiyuan Zhao Bin Wang Linke Ouyang Xiaoyi Dong Jiaqi Wang and Conghui He. 2023. Beyond hallucinations: Enhancing LVLMs through hallucination-aware direct preference optimization. arXiv:2311.16839. Retrieved from https:\/\/arxiv.org\/abs\/2311.16839. (2023)."},{"key":"e_1_3_1_229_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i16.29771"},{"key":"e_1_3_1_230_2","unstructured":"Nick Jiang Anish Kachinthaya Suzie Petryk and Yossi Gandelsman. 2024. Interpreting and editing vision-language representations to mitigate hallucinations. arXiv:2410.02762. Retrieved from https:\/\/arxiv.org\/abs\/2410.02762. (2024)."},{"key":"e_1_3_1_231_2","doi-asserted-by":"crossref","unstructured":"S. Yin C. Fu S. Zhao et\u00a0al. 2024. Woodpecker: Hallucination correction for multimodal large language models. Science China Information Sciences 67 12 (2024) 220105.","DOI":"10.1007\/s11432-024-4251-x"},{"key":"e_1_3_1_232_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-short.30"},{"key":"e_1_3_1_233_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00434"},{"key":"e_1_3_1_234_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.ltedi-1.8"},{"key":"e_1_3_1_235_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.naacl-demo.8"},{"key":"e_1_3_1_236_2","unstructured":"Winston AI. 2023. Winston AI Detector. (2023). Retrieved from https:\/\/gowinston.ai\/"},{"key":"e_1_3_1_237_2","unstructured":"Originality ai. 2022. AI Checker Plagiarism Checker and Fact Checker. (2022). Retrieved from https:\/\/originality.ai\/"},{"key":"e_1_3_1_238_2","unstructured":"GPTZero. 2023. GPTZero. (2023). Retrieved from https:\/\/gptzero.me\/"},{"key":"e_1_3_1_239_2","unstructured":"ZeroGPT. 2023. ZeroGPT. (2023). Retrieved from https:\/\/www.zerogpt.com\/"},{"key":"e_1_3_1_240_2","unstructured":"Crossplag. 2023. Crossplag AI Detector. (2023). Retrieved from https:\/\/crossplag.com\/ai-content-detector\/"},{"key":"e_1_3_1_241_2","unstructured":"Sapling. 2023. AI Content Detector. (2023). Retrieved from https:\/\/sapling.ai\/ai-content-detector"},{"key":"e_1_3_1_242_2","unstructured":"Undetectable AI. 2023. Undetectable.ai. (2023). Retrieved from https:\/\/undetectable.ai\/"},{"key":"e_1_3_1_243_2","unstructured":"AI or Not. 2023. Image AI Detector. (2023). Retrieved from https:\/\/www.aiornot.com\/"},{"key":"e_1_3_1_244_2","unstructured":"Illuminarty.AI. 2023. Illuminarty Detector. (2023). Retrieved from https:\/\/app.illuminarty.ai\/"},{"key":"e_1_3_1_245_2","unstructured":"Google. 2023. SynthID. (2023). Retrieved from https:\/\/deepmind.google\/science\/synthid\/"},{"key":"e_1_3_1_246_2","unstructured":"Tencent. 2023. Zhuque AI Detection Assistant. (2023). Retrieved from https:\/\/matrix.tencent.com\/ai-detect\/"},{"key":"e_1_3_1_247_2","unstructured":"ADVANCE.AI. 2023. Liveness Detector. (2023). Retrieved from https:\/\/www.advanceai.com.cn\/liveness-detection"},{"key":"e_1_3_1_248_2","unstructured":"Hive Moderation. 2024. Hive AI-Generated Content Detector. (2024). Retrieved from https:\/\/hivemoderation.com\/"},{"key":"e_1_3_1_249_2","unstructured":"Attestiv. 2024. Attestiv Deepfake Detector. (2024). Retrieved from https:\/\/attestiv.com\/"},{"key":"e_1_3_1_250_2","unstructured":"Deepware. 2020. Deepfake Videos Scanner and Detector. (2020). Retrieved from https:\/\/deepware.ai\/"},{"key":"e_1_3_1_251_2","unstructured":"DetectIQ. 2024. Advanced AI-generated content Detector. (2024). Retrieved from https:\/\/detectiq.io\/"},{"key":"e_1_3_1_252_2","unstructured":"Pindrop. 2025. Pindrop Pulse. (2025). Retrieved from https:\/\/www.pindrop.com\/"},{"key":"e_1_3_1_253_2","unstructured":"Ircamamplify. 2024. AI Speech Detector. (2024). Retrieved from https:\/\/www.ircamamplify.io\/"},{"key":"e_1_3_1_254_2","unstructured":"Resemble AI. 2023. Free Deepfake Detector. (2023). Retrieved from https:\/\/www.resemble.ai\/free-deepfake-detector\/"},{"key":"e_1_3_1_255_2","unstructured":"AI Voice Detector. 2024. AI Voice Detector. (2024). Retrieved from https:\/\/aivoicedetector.com\/"},{"key":"e_1_3_1_256_2","unstructured":"ScreenApp. 2024. AI Video Detector. (2024). Retrieved from https:\/\/screenapp.io\/features\/ai-video-detector"},{"key":"e_1_3_1_257_2","unstructured":"Reality Defender. 2025. Reality Defender. (2025). Retrieved from https:\/\/www.realitydefender.com\/"},{"key":"e_1_3_1_258_2","unstructured":"Sensity. 2025. All-In-One Deepfake Detector. (2025). Retrieved from https:\/\/sensity.ai\/"},{"key":"e_1_3_1_259_2","unstructured":"DuckDuckGoose AI. 2025. Deepfake Detection Platform. (2025). Retrieved from https:\/\/www.duckduckgoose.ai\/"},{"key":"e_1_3_1_260_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.121641"},{"key":"e_1_3_1_261_2","doi-asserted-by":"publisher","DOI":"10.1145\/3715073.3715076"},{"key":"e_1_3_1_262_2","doi-asserted-by":"publisher","DOI":"10.1080\/08839514.2025.2463722"},{"key":"e_1_3_1_263_2","doi-asserted-by":"publisher","DOI":"10.3389\/fdata.2024.1402745"},{"key":"e_1_3_1_264_2","unstructured":"Ishaan Domkundwar Ishaan Bhola Riddhik Kochhar and Mukunda N S. 2024. Safeguarding AI agents: Developing and analyzing safety architectures. arXiv:2409.03793. Retrieved from https:\/\/arxiv.org\/abs\/2409.03793. (2024)."},{"key":"e_1_3_1_265_2","doi-asserted-by":"publisher","DOI":"10.1057\/s41284-024-00435-3"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3760526","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,9]],"date-time":"2025-09-09T14:31:01Z","timestamp":1757428261000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3760526"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,9]]},"references-count":264,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,2,28]]}},"alternative-id":["10.1145\/3760526"],"URL":"https:\/\/doi.org\/10.1145\/3760526","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,9]]},"assertion":[{"value":"2024-12-02","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-03","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-09","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}