{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,24]],"date-time":"2026-03-24T04:30:35Z","timestamp":1774326635682,"version":"3.50.1"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2024,12,24]],"date-time":"2024-12-24T00:00:00Z","timestamp":1734998400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U20B2047, 62072421, 62002334, 62102386, and 62121002"],"award-info":[{"award-number":["U20B2047, 62072421, 62002334, 62102386, and 62121002"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100017668","name":"Key Research and Development program of Anhui Province","doi-asserted-by":"crossref","award":["2022k07020008"],"award-info":[{"award-number":["2022k07020008"]}],"id":[{"id":"10.13039\/501100017668","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["WK5290000003 and WK2100000011"],"award-info":[{"award-number":["WK5290000003 and WK2100000011"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,2,28]]},"abstract":"<jats:p>The highly realistic avatar in the metaverse may lead to deepfakes of facial identity. Malicious users can more easily obtain the three-dimensional structure of faces, thus using deepfake technology to create counterfeit videos with higher realism. To automatically discern facial videos forged with the advancing generation techniques, deepfake detectors need to achieve stronger generalization abilities. Inspired by transfer learning, neural networks pre-trained on other large-scale face-related tasks would provide fundamental features for deepfake detection. We propose a video-level deepfake detection method based on a temporal transformer with a self-supervised audio\u2013visual contrastive learning approach for pre-training the deepfake detector. The proposed method learns motion representations in the mouth region by encouraging the paired video and audio representations to be close while unpaired ones to be diverse. The deepfake detector adopts the pre-trained weights and partially fine-tunes on deepfake datasets. Extensive experiments show that our self-supervised pre-training method can effectively improve the accuracy and robustness of our deepfake detection model without extra human efforts. Compared with existing deepfake detection methods, our proposed method achieves better generalization ability in cross-dataset evaluations.<\/jats:p>\n          <jats:p\/>","DOI":"10.1145\/3651311","type":"journal-article","created":{"date-parts":[[2024,3,13]],"date-time":"2024-03-13T11:53:58Z","timestamp":1710330838000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Audio-Visual Contrastive Pre-train for Face Forgery Detection"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-0171-9543","authenticated-orcid":false,"given":"Hanqing","family":"Zhao","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4703-4641","authenticated-orcid":false,"given":"Wenbo","family":"Zhou","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4642-4373","authenticated-orcid":false,"given":"Dongdong","family":"Chen","sequence":"additional","affiliation":[{"name":"Microsoft Cloud &amp; AI, Seattle, United States"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5576-6108","authenticated-orcid":false,"given":"Weiming","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6429-9297","authenticated-orcid":false,"given":"Ying","family":"Guo","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-1982-1522","authenticated-orcid":false,"given":"Zhen","family":"Cheng","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3438-8994","authenticated-orcid":false,"given":"Pengfei","family":"Yan","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4417-9316","authenticated-orcid":false,"given":"Nenghai","family":"Yu","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,12,24]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Alexei Baevski Henry Zhou Abdelrahman Mohamed and Michael Auli. 2020. wav2vec 2.0: a framework for self-supervised learning of speech representations. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS\u201920) Curran Associates Inc. Vancouver BC Canada."},{"key":"e_1_3_1_3_2","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Bulat Adrian","year":"2017","unstructured":"Adrian Bulat and Georgios Tzimiropoulos. 2017. How far are we from solving the 2D & 3D Face Alignment problem? (and a dataset of 230,000 3D facial landmarks). In Proceedings of the International Conference on Computer Vision."},{"key":"e_1_3_1_4_2","volume-title":"European Conference on Computer Vision (ECCV\u201920)","author":"Chai Lucy","year":"2020","unstructured":"Lucy Chai, David Bau, Ser-Nam Lim, and Phillip Isola. 2020. What makes fake images detectable? Understanding properties that generalize. In European Conference on Computer Vision (ECCV\u201920)."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413630"},{"key":"e_1_3_1_6_2","unstructured":"Ting Chen Simon Kornblith Mohammad Norouzi and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In Proceedings of the 37th International Conference on Machine Learning (ICML\u201920) JMLR.org."},{"key":"e_1_3_1_7_2","first-page":"9620","article-title":"An empirical study of training self-supervised vision transformers","author":"Chen Xinlei","year":"2021","unstructured":"Xinlei Chen, Saining Xie, and Kaiming He. 2021. An empirical study of training self-supervised vision transformers. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201921), 9620\u20139629.","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201921)"},{"key":"e_1_3_1_8_2","first-page":"7012","article-title":"Distilling audio-visual knowledge by compositional contrastive learning","author":"Chen Yanbei","year":"2021","unstructured":"Yanbei Chen, Yongqin Xian, A. Sophia Koepke, Ying Shan, and Zeynep Akata. 2021. Distilling audio-visual knowledge by compositional contrastive learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921), 7012\u20137021.","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921)"},{"key":"e_1_3_1_9_2","volume-title":"Proceedings of the Asian Conference on Computer Vision","author":"Chung J. S.","year":"2016","unstructured":"J. S. Chung and A. Zisserman. 2016. Lip reading in the wild. In Proceedings of the Asian Conference on Computer Vision."},{"key":"e_1_3_1_10_2","volume-title":"Proceedings of the ACCV Workshops","author":"Chung Joon Son","year":"2016","unstructured":"Joon Son Chung and Andrew Zisserman. 2016. Out of time: Automated lip sync in the wild. In Proceedings of the ACCV Workshops."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","unstructured":"Ishan Dave Rohit Gupta Mamshad Nayeem Rizve and Mubarak Shah. 2022. TCLR: Temporal contrastive learning for video representation. Computer Vision and Image Understanding 219 (2022) 103406. DOI:10.1016\/j.cviu.2022.103406","DOI":"10.1016\/j.cviu.2022.103406"},{"key":"e_1_3_1_12_2","article-title":"Deepfake detection using spatiotemporal convolutional networks","volume":"2006","author":"Lima Oscar de","year":"2020","unstructured":"Oscar de Lima, Sean Franklin, Shreshtha Basu, Blake Karwoski, and Annet George. 2020. Deepfake detection using spatiotemporal convolutional networks. arXiv: 2006.14749. Retrieved from https:\/\/arxiv.org\/abs\/2006.14749","journal-title":"arXiv"},{"key":"e_1_3_1_13_2","first-page":"5202","article-title":"RetinaFace: Single-shot multi-level face localisation in the wild","author":"Deng Jiankang","year":"2020","unstructured":"Jiankang Deng, J. Guo, Evangelos Ververas, Irene Kotsia, Stefanos Zafeiriou, and InsightFace FaceSoft. 2020. RetinaFace: Single-shot multi-level face localisation in the wild. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920), 5202\u20135211.","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)"},{"key":"e_1_3_1_14_2","unstructured":"Brian Dolhansky. 2020. The DeepFake Detection Challenge Dataset."},{"key":"e_1_3_1_15_2","first-page":"3298","article-title":"A large-scale study on unsupervised spatiotemporal representation learning","author":"Feichtenhofer Christoph","year":"2021","unstructured":"Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross B. Girshick, and Kaiming He. 2021. A large-scale study on unsupervised spatiotemporal representation learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921), 3298\u20133308.","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921)"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275043"},{"key":"e_1_3_1_17_2","article-title":"Contrastive audio-visual masked autoencoder","volume":"2210","author":"Gong Yuan","year":"2022","unstructured":"Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David F. Harwath, Leonid Karlinsky, Hilde Kuehne, and James R. Glass. 2022. Contrastive audio-visual masked autoencoder. arXiv 2210.07839. Retrieved from https:\/\/arxiv.org\/abs\/2210.07839","journal-title":"arXiv"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00500"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2916751"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","unstructured":"Simon Jenni Alexander Black and John Collomosse. 2023. Audio-visual contrastive learning with temporal self-supervision. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence and Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence and Thirteenth Symposium on Educational Advances in Artificial Intelligence (AAAI\u201923\/IAAI\u201923\/EAAI\u201923) AAAI Press. DOI:10.1609\/aaai.v37i7.25967","DOI":"10.1609\/aaai.v37i7.25967"},{"key":"e_1_3_1_21_2","first-page":"2886","article-title":"DeeperForensics-1.0: A large-scale dataset for real-world face forgery detection","author":"Jiang Liming","year":"2020","unstructured":"Liming Jiang, Wayne Wu, Ren Li, Chen Qian, and Chen Change Loy. 2020. DeeperForensics-1.0: A large-scale dataset for real-world face forgery detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920), 2886\u20132895.","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","unstructured":"T. Karras S. Laine M. Aittala J. Hellsten J. Lehtinen and T. Aila. 2020. Analyzing and improving the image quality of styleGAN. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) IEEE Computer Society Los Alamitos CA USA 8107\u20138116. DOI:10.1109\/CVPR42600.2020.00813","DOI":"10.1109\/CVPR42600.2020.00813"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00505"},{"key":"e_1_3_1_24_2","article-title":"Exposing deepfake videos by detecting face warping artifacts","volume":"1811","author":"Li Yuezun","year":"2019","unstructured":"Yuezun Li and Siwei Lyu. 2019. Exposing deepfake videos by detecting face warping artifacts. arXiv: 1811.00656. Retrieved from https:\/\/arxiv.org\/abs\/1811.00656","journal-title":"arXiv:"},{"key":"e_1_3_1_25_2","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW\u201919)","author":"Li Yuezun","year":"2019","unstructured":"Yuezun Li and Siwei Lyu. 2019. Exposing deepfake videos by detecting face warping artifacts. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW\u201919)."},{"key":"e_1_3_1_26_2","first-page":"3204","article-title":"Celeb-DF: A large-scale challenging dataset for deepfake forensics","author":"Li Yuezun","year":"2020","unstructured":"Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. 2020. Celeb-DF: A large-scale challenging dataset for deepfake forensics. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920) (2020), 3204\u20133213.","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00083"},{"key":"e_1_3_1_28_2","first-page":"1000","article-title":"Self-supervised contrastive learning for audio-visual action recognition","author":"Liu Yang","year":"2022","unstructured":"Yang Liu, Ying Hua Tan, and Haoyu Lan. 2022. Self-supervised contrastive learning for audio-visual action recognition. In Proceedings of the IEEE International Conference on Image Processing (ICIP\u201922), 1000\u20131004.","journal-title":"Proceedings of the IEEE International Conference on Image Processing (ICIP\u201922)"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9415063"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58571-6_39"},{"key":"e_1_3_1_31_2","first-page":"12470","article-title":"Audio-visual instance discrimination with cross-modal agreement","author":"Morgado Pedro","year":"2021","unstructured":"Pedro Morgado, Nuno Vasconcelos, and Ishan Misra. 2021. Audio-visual instance discrimination with cross-modal agreement. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921), 12470\u201312481.","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921)"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/BTAS46853.2019.9185974"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.108832"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00009"},{"issue":"1","key":"e_1_3_1_35_2","article-title":"Recurrent convolutional strategies for face manipulation detection in videos","volume":"3","author":"Sabir Ekraam","year":"2019","unstructured":"Ekraam Sabir, Jiaxin Cheng, Ayush Jaiswal, Wael AbdAlmageed, Iacopo Masi, and Prem Natarajan. 2019. Recurrent convolutional strategies for face manipulation detection in videos. Interfaces (GUI) 3, 1 (2019).","journal-title":"Interfaces (GUI)"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.107950"},{"key":"e_1_3_1_37_2","article-title":"First order motion model for image animation","volume":"32","author":"Siarohin Aliaksandr","year":"2019","unstructured":"Aliaksandr Siarohin, St\u00e9phane Lathuili\u00e8re, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019. First order motion model for image animation. In Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","unstructured":"Justus Thies Michael Zollh\u00f6fer and Matthias Nie\u00dfner. 2019. Deferred neural rendering: image synthesis using neural textures. ACM Trans. Graph. 38 4 (July 2019). DOI:10.1145\/3306346.3323035","DOI":"10.1145\/3306346.3323035"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV48630.2021.00137"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","unstructured":"Sheng-Yu Wang Oliver Wang Richard Zhang Andrew Owens and Alexei A. Efros. 2020. CNN-generated images are surprisingly easy to spot... for now. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 8692\u20138701. DOI:10.1109\/CVPR42600.2020.00872","DOI":"10.1109\/CVPR42600.2020.00872"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01261-8_41"},{"key":"e_1_3_1_42_2","first-page":"103","volume-title":"Proceedings of the International Conference on Audio, Language and Image Processing (ICALIP\u201918)","author":"Yan Shuqi","year":"2018","unstructured":"Shuqi Yan, Shaorong He, Xue Lei, Guanhua Ye, and Zhifeng Xie. 2018. Video face swap based on autoencoder generation network. In Proceedings of the International Conference on Audio, Language and Image Processing (ICALIP\u201918). IEEE, 103\u2013108."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00222"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01475"},{"key":"e_1_3_1_45_2","first-page":"4834","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhu Yuhao","year":"2021","unstructured":"Yuhao Zhu, Qi Li, Jian Wang, Cheng-Zhong Xu, and Zhenan Sun. 2021. One shot face swapping on megapixels. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 4834\u20134844."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3651311","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3651311","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:49:54Z","timestamp":1750286994000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3651311"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,24]]},"references-count":44,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,2,28]]}},"alternative-id":["10.1145\/3651311"],"URL":"https:\/\/doi.org\/10.1145\/3651311","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,24]]},"assertion":[{"value":"2023-12-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-26","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}