{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T15:12:57Z","timestamp":1784301177530,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":74,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,6,27]],"date-time":"2022-06-27T00:00:00Z","timestamp":1656288000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Natural Science Foundation","award":["62072116"],"award-info":[{"award-number":["62072116"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,6,27]]},"DOI":"10.1145\/3512527.3531415","type":"proceedings-article","created":{"date-parts":[[2022,6,23]],"date-time":"2022-06-23T22:23:32Z","timestamp":1656023012000},"page":"615-623","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":312,"title":["M2TR: Multi-modal Multi-scale Transformers for Deepfake Detection"],"prefix":"10.1145","author":[{"given":"Junke","family":"Wang","sequence":"first","affiliation":[{"name":"Fudan University, shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zuxuan","family":"Wu","sequence":"additional","affiliation":[{"name":"Fudan University, shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenhao","family":"Ouyang","sequence":"additional","affiliation":[{"name":"Fudan University, shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xintong","family":"Han","sequence":"additional","affiliation":[{"name":"Huya Inc, shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jingjing","family":"Chen","sequence":"additional","affiliation":[{"name":"Fudan University, shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu-Gang","family":"Jiang","sequence":"additional","affiliation":[{"name":"Fudan University, shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ser-Nam","family":"Li","sequence":"additional","affiliation":[{"name":"Meta AI, Newyork, NY, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,6,27]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"crossref","unstructured":"Darius Afchar Vincent Nozick Junichi Yamagishi and Isao Echizen. 2018. Mesonet: a compact facial video forgery detection network. In WIFS.  Darius Afchar Vincent Nozick Junichi Yamagishi and Isao Echizen. 2018. Mesonet: a compact facial video forgery detection network. In WIFS.","DOI":"10.1109\/WIFS.2018.8630761"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"crossref","unstructured":"Nicolas Carion Francisco Massa Gabriel Synnaeve Nicolas Usunier Alexander Kirillov and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. In ECCV.  Nicolas Carion Francisco Massa Gabriel Synnaeve Nicolas Usunier Alexander Kirillov and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. In ECCV.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"crossref","unstructured":"Joao Carreira and Andrew Zisserman. 2017. Quo vadis action recognition? a new model and the kinetics dataset. In CVPR.  Joao Carreira and Andrew Zisserman. 2017. Quo vadis action recognition? a new model and the kinetics dataset. In CVPR.","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"crossref","unstructured":"Dongdong Chen Jing Liao Lu Yuan Nenghai Yu and Gang Hua. 2017. Coherent online video style transfer. In ICCV.  Dongdong Chen Jing Liao Lu Yuan Nenghai Yu and Gang Hua. 2017. Coherent online video style transfer. In ICCV.","DOI":"10.1109\/ICCV.2017.126"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"crossref","unstructured":"Mo Chen Vahid Sedighi Mehdi Boroumand and Jessica Fridrich. 2017. JPEG-phase-aware convolutional neural network for steganalysis of JPEG images. In IHMSW.  Mo Chen Vahid Sedighi Mehdi Boroumand and Jessica Fridrich. 2017. JPEG-phase-aware convolutional neural network for steganalysis of JPEG images. In IHMSW.","DOI":"10.1145\/3082031.3083248"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"crossref","unstructured":"Shen Chen Taiping Yao Yang Chen Shouhong Ding Jilin Li and Rongrong Ji. 2021. Local Relation Learning for Face Forgery Detection. In AAAI.  Shen Chen Taiping Yao Yang Chen Shouhong Ding Jilin Li and Rongrong Ji. 2021. Local Relation Learning for Face Forgery Detection. In AAAI.","DOI":"10.1609\/aaai.v35i2.16193"},{"key":"e_1_3_2_2_7_1","volume-title":"Xception: Deep learning with depthwise separable convolutions. In CVPR.","author":"Chollet Fran\u00e7ois","year":"2017","unstructured":"Fran\u00e7ois Chollet . 2017 . Xception: Deep learning with depthwise separable convolutions. In CVPR. Fran\u00e7ois Chollet. 2017. Xception: Deep learning with depthwise separable convolutions. In CVPR."},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Wenyan Cong Jianfu Zhang Li Niu Liu Liu Zhixin Ling Weiyuan Li and Liqing Zhang. 2020. DoveNet: Deep Image Harmonization via Domain Verification. In CVPR.  Wenyan Cong Jianfu Zhang Li Niu Liu Liu Zhixin Ling Weiyuan Li and Liqing Zhang. 2020. DoveNet: Deep Image Harmonization via Domain Verification. In CVPR.","DOI":"10.1109\/CVPR42600.2020.00842"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3082031.3083247"},{"key":"e_1_3_2_2_10_1","unstructured":"DeepFake Detection Dataset. 2019. https:\/\/ai.googleblog.com\/2019\/09\/contributing-data-to-deepfake-detection.html.  DeepFake Detection Dataset. 2019. https:\/\/ai.googleblog.com\/2019\/09\/contributing-data-to-deepfake-detection.html."},{"key":"e_1_3_2_2_11_1","unstructured":"Deepfakes. 2018. github. https:\/\/github.com\/deepfakes\/faceswap.  Deepfakes. 2018. github. https:\/\/github.com\/deepfakes\/faceswap."},{"key":"e_1_3_2_2_12_1","volume-title":"Imagenet: A large-scale hierarchical image database. In CVPR.","author":"Deng Jia","year":"2009","unstructured":"Jia Deng , Wei Dong , Richard Socher , Li-Jia Li , Kai Li , and Li Fei-Fei . 2009 . Imagenet: A large-scale hierarchical image database. In CVPR. Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In CVPR."},{"key":"e_1_3_2_2_13_1","volume-title":"Retinaface: Single-shot multi-level face localisation in the wild. In CVPR.","author":"Deng Jiankang","year":"2020","unstructured":"Jiankang Deng , Jia Guo , Evangelos Ververas , Irene Kotsia , and Stefanos Zafeiriou . 2020 . Retinaface: Single-shot multi-level face localisation in the wild. In CVPR. Jiankang Deng, Jia Guo, Evangelos Ververas, Irene Kotsia, and Stefanos Zafeiriou. 2020. Retinaface: Single-shot multi-level face localisation in the wild. In CVPR."},{"key":"e_1_3_2_2_14_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT.","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT."},{"key":"e_1_3_2_2_15_1","volume-title":"The deepfake detection challenge (dfdc) dataset. arXiv preprint arXiv:2006.07397","author":"Dolhansky Brian","year":"2020","unstructured":"Brian Dolhansky , Russ Howes , Ben Pflaum , Nicole Baram , and Cristian Canton Ferrer . 2020. The deepfake detection challenge (dfdc) dataset. arXiv preprint arXiv:2006.07397 ( 2020 ). Brian Dolhansky, Russ Howes, Ben Pflaum, Nicole Baram, and Cristian Canton Ferrer. 2020. The deepfake detection challenge (dfdc) dataset. arXiv preprint arXiv:2006.07397 (2020)."},{"key":"e_1_3_2_2_16_1","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly Jakob Uszkoreit and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR.  Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly Jakob Uszkoreit and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR."},{"key":"e_1_3_2_2_17_1","volume-title":"Unmasking deepfakes with simple features. arXiv preprint arXiv:1911.00686","author":"Durall Ricard","year":"2019","unstructured":"Ricard Durall , Margret Keuper , Franz-Josef Pfreundt , and Janis Keuper . 2019. Unmasking deepfakes with simple features. arXiv preprint arXiv:1911.00686 ( 2019 ). Ricard Durall, Margret Keuper, Franz-Josef Pfreundt, and Janis Keuper. 2019. Unmasking deepfakes with simple features. arXiv preprint arXiv:1911.00686 (2019)."},{"key":"e_1_3_2_2_18_1","unstructured":"Face-parsing. 2019. github. https:\/\/github.com\/zllrunning\/face-parsing.PyTorch.  Face-parsing. 2019. github. https:\/\/github.com\/zllrunning\/face-parsing.PyTorch."},{"key":"e_1_3_2_2_19_1","unstructured":"FaceShifter. 2020. github. https:\/\/github.com\/mindslab-ai\/faceshifter.  FaceShifter. 2020. github. https:\/\/github.com\/mindslab-ai\/faceshifter."},{"key":"e_1_3_2_2_20_1","unstructured":"Faceswap. 2018. github. https:\/\/github.com\/MarekKowalski\/FaceSwap\/.  Faceswap. 2018. github. https:\/\/github.com\/MarekKowalski\/FaceSwap\/."},{"key":"e_1_3_2_2_21_1","volume-title":"Rich models for steganalysis of digital images. TIFS","author":"Fridrich Jessica","year":"2012","unstructured":"Jessica Fridrich and Jan Kodovsky . 2012. Rich models for steganalysis of digital images. TIFS ( 2012 ). Jessica Fridrich and Jan Kodovsky. 2012. Rich models for steganalysis of digital images. TIFS (2012)."},{"key":"e_1_3_2_2_22_1","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR.  Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR."},{"key":"e_1_3_2_2_23_1","unstructured":"Yinan He Bei Gan Siyu Chen Yichun Zhou Guojun Yin Luchuan Song Lu Sheng Jing Shao and Ziwei Liu. 2021. ForgeryNet: A Versatile Benchmark for Comprehensive Forgery Analysis. In CVPR.  Yinan He Bei Gan Siyu Chen Yichun Zhou Guojun Yin Luchuan Song Lu Sheng Jing Shao and Ziwei Liu. 2021. ForgeryNet: A Versatile Benchmark for Comprehensive Forgery Analysis. In CVPR."},{"key":"e_1_3_2_2_24_1","volume-title":"et al","author":"Howard Andrew","year":"2019","unstructured":"Andrew Howard , Mark Sandler , Grace Chu , Liang-Chieh Chen , Bo Chen , Mingxing Tan , Weijun Wang , Yukun Zhu , Ruoming Pang , Vijay Vasudevan , et al . 2019 . Searching for mobilenetv3. In ICCV. Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, et al . 2019. Searching for mobilenetv3. In ICCV."},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"crossref","unstructured":"Haozhi Huang Hao Wang Wenhan Luo Lin Ma Wenhao Jiang Xiaolong Zhu Zhifeng Li and Wei Liu. 2017. Real-time neural style transfer for videos. In CVPR.  Haozhi Huang Hao Wang Wenhan Luo Lin Ma Wenhao Jiang Xiaolong Zhu Zhifeng Li and Wei Liu. 2017. Real-time neural style transfer for videos. In CVPR.","DOI":"10.1109\/CVPR.2017.745"},{"key":"e_1_3_2_2_26_1","volume-title":"Deep frequent spatial temporal learning for face anti-spoofing. arXiv preprint arXiv:2002.03723","author":"Huang Ying","year":"2020","unstructured":"Ying Huang , Wenwei Zhang , and Jinzhuo Wang . 2020. Deep frequent spatial temporal learning for face anti-spoofing. arXiv preprint arXiv:2002.03723 ( 2020 ). Ying Huang, Wenwei Zhang, and Jinzhuo Wang. 2020. Deep frequent spatial temporal learning for face anti-spoofing. arXiv preprint arXiv:2002.03723 (2020)."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"crossref","unstructured":"Hyeonseong Jeon Youngoh Bang and Simon S Woo. 2020. FDFtNet: Facing off fake images using fake detection fine-tuning network. In ICT Systems Security and Privacy Protection.  Hyeonseong Jeon Youngoh Bang and Simon S Woo. 2020. FDFtNet: Facing off fake images using fake detection fine-tuning network. In ICT Systems Security and Privacy Protection.","DOI":"10.1007\/978-3-030-58201-2_28"},{"key":"e_1_3_2_2_28_1","volume-title":"Transfiguring portraits. ACM TOG","author":"Kemelmacher-Shlizerman Ira","year":"2016","unstructured":"Ira Kemelmacher-Shlizerman . 2016. Transfiguring portraits. ACM TOG ( 2016 ). Ira Kemelmacher-Shlizerman. 2016. Transfiguring portraits. ACM TOG (2016)."},{"key":"e_1_3_2_2_29_1","volume-title":"Dlib-ml: A machine learning toolkit. JMLR","author":"King Davis E","year":"2009","unstructured":"Davis E King . 2009 . Dlib-ml: A machine learning toolkit. JMLR (2009). Davis E King. 2009. Dlib-ml: A machine learning toolkit. JMLR (2009)."},{"key":"e_1_3_2_2_30_1","volume-title":"Deepfakes: a new threat to face recognition? assessment and detection. arXiv preprint arXiv:1812.08685","author":"Korshunov Pavel","year":"2018","unstructured":"Pavel Korshunov and S\u00e9bastien Marcel . 2018. Deepfakes: a new threat to face recognition? assessment and detection. arXiv preprint arXiv:1812.08685 ( 2018 ). Pavel Korshunov and S\u00e9bastien Marcel. 2018. Deepfakes: a new threat to face recognition? assessment and detection. arXiv preprint arXiv:1812.08685 (2018)."},{"key":"e_1_3_2_2_31_1","volume-title":"Anastasios Roussos, and Stefanos Zafeiriou.","author":"Koujan Mohammad Rami","year":"2020","unstructured":"Mohammad Rami Koujan , Michail Christos Doukas , Anastasios Roussos, and Stefanos Zafeiriou. 2020 . Head2head: Video-based neural head synthesis. arXiv preprint arXiv:2005.10954 (2020). Mohammad Rami Koujan, Michail Christos Doukas, Anastasios Roussos, and Stefanos Zafeiriou. 2020. Head2head: Video-based neural head synthesis. arXiv preprint arXiv:2005.10954 (2020)."},{"key":"e_1_3_2_2_32_1","unstructured":"Chenyang Lei Yazhou Xing and Qifeng Chen. 2020. Blind Video Temporal Consistency via Deep Video Prior. In NIPS.  Chenyang Lei Yazhou Xing and Qifeng Chen. 2020. Blind Video Temporal Consistency via Deep Video Prior. In NIPS."},{"key":"e_1_3_2_2_33_1","volume-title":"Faceshifter: Towards high fidelity and occlusion aware face swapping. arXiv preprint arXiv:1912.13457","author":"Li Lingzhi","year":"2019","unstructured":"Lingzhi Li , Jianmin Bao , Hao Yang , Dong Chen , and Fang Wen . 2019 . Faceshifter: Towards high fidelity and occlusion aware face swapping. arXiv preprint arXiv:1912.13457 (2019). Lingzhi Li, Jianmin Bao, Hao Yang, Dong Chen, and Fang Wen. 2019. Faceshifter: Towards high fidelity and occlusion aware face swapping. arXiv preprint arXiv:1912.13457 (2019)."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"crossref","unstructured":"Lingzhi Li Jianmin Bao Ting Zhang Hao Yang Dong Chen Fang Wen and Baining Guo. 2020. Face x-ray for more general face forgery detection. In CVPR.  Lingzhi Li Jianmin Bao Ting Zhang Hao Yang Dong Chen Fang Wen and Baining Guo. 2020. Face x-ray for more general face forgery detection. In CVPR.","DOI":"10.1109\/CVPR42600.2020.00505"},{"key":"e_1_3_2_2_35_1","volume-title":"Exposing deepfake videos by detecting face warping artifacts. arXiv preprint arXiv:1811.00656","author":"Li Yuezun","year":"2018","unstructured":"Yuezun Li and Siwei Lyu . 2018. Exposing deepfake videos by detecting face warping artifacts. arXiv preprint arXiv:1811.00656 ( 2018 ). Yuezun Li and Siwei Lyu. 2018. Exposing deepfake videos by detecting face warping artifacts. arXiv preprint arXiv:1811.00656 (2018)."},{"key":"e_1_3_2_2_36_1","unstructured":"Yuezun Li and Siwei Lyu. 2019. Exposing DeepFake Videos By Detecting Face Warping Artifacts. In CVPRW.  Yuezun Li and Siwei Lyu. 2019. Exposing DeepFake Videos By Detecting Face Warping Artifacts. In CVPRW."},{"key":"e_1_3_2_2_37_1","volume-title":"Celeb-df: A large-scale challenging dataset for deepfake forensics. In CVPR.","author":"Li Yuezun","year":"2020","unstructured":"Yuezun Li , Xin Yang , Pu Sun , Honggang Qi , and Siwei Lyu . 2020 . Celeb-df: A large-scale challenging dataset for deepfake forensics. In CVPR. Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. 2020. Celeb-df: A large-scale challenging dataset for deepfake forensics. In CVPR."},{"key":"e_1_3_2_2_38_1","unstructured":"Zhengzhe Liu Xiaojuan Qi and Philip HS Torr. 2020. Global texture enhancement for fake face detection in the wild. In CVPR.  Zhengzhe Liu Xiaojuan Qi and Philip HS Torr. 2020. Global texture enhancement for fake face detection in the wild. In CVPR."},{"key":"e_1_3_2_2_39_1","volume-title":"Shenoy Pratik Gurudatt","author":"Masi Iacopo","year":"2020","unstructured":"Iacopo Masi , Aditya Killekar , Royston Marian Mascarenhas , Shenoy Pratik Gurudatt , and Wael AbdAlmageed. 2020 . Two-branch recurrent network for isolating deepfakes in videos. In ECCV. Iacopo Masi, Aditya Killekar, Royston Marian Mascarenhas, Shenoy Pratik Gurudatt, and Wael AbdAlmageed. 2020. Two-branch recurrent network for isolating deepfakes in videos. In ECCV."},{"key":"e_1_3_2_2_40_1","volume-title":"Joon Son Chung, and Andrew Zisserman","author":"Nagrani Arsha","year":"2017","unstructured":"Arsha Nagrani , Joon Son Chung, and Andrew Zisserman . 2017 . Voxceleb : a large-scale speaker identification dataset. arXiv preprint arXiv:1706.08612 (2017). Arsha Nagrani, Joon Son Chung, and Andrew Zisserman. 2017. Voxceleb: a large-scale speaker identification dataset. arXiv preprint arXiv:1706.08612 (2017)."},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"crossref","unstructured":"Kamyar Nazeri Eric Ng Tony Joseph Faisal Qureshi and Mehran Ebrahimi. 2019. EdgeConnect: Structure Guided Image Inpainting using Edge Prediction. In ICCVW.  Kamyar Nazeri Eric Ng Tony Joseph Faisal Qureshi and Mehran Ebrahimi. 2019. EdgeConnect: Structure Guided Image Inpainting using Edge Prediction. In ICCVW.","DOI":"10.1109\/ICCVW.2019.00408"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"crossref","unstructured":"Huy H Nguyen Fuming Fang Junichi Yamagishi and Isao Echizen. 2019. Multi- task learning for detecting and segmenting manipulated facial images and videos. In BTAS.  Huy H Nguyen Fuming Fang Junichi Yamagishi and Isao Echizen. 2019. Multi- task learning for detecting and segmenting manipulated facial images and videos. In BTAS.","DOI":"10.1109\/BTAS46853.2019.9185974"},{"key":"e_1_3_2_2_43_1","volume-title":"Use of a capsule network to detect fake images and videos. arXiv preprint arXiv:1910.12467","author":"Nguyen Huy H","year":"2019","unstructured":"Huy H Nguyen , Junichi Yamagishi , and Isao Echizen . 2019. Use of a capsule network to detect fake images and videos. arXiv preprint arXiv:1910.12467 ( 2019 ). Huy H Nguyen, Junichi Yamagishi, and Isao Echizen. 2019. Use of a capsule network to detect fake images and videos. arXiv preprint arXiv:1910.12467 (2019)."},{"key":"e_1_3_2_2_44_1","volume-title":"Fsgan: Subject agnostic face swapping and reenactment. In ICCV.","author":"Nirkin Yuval","year":"2019","unstructured":"Yuval Nirkin , Yosi Keller , and Tal Hassner . 2019 . Fsgan: Subject agnostic face swapping and reenactment. In ICCV. Yuval Nirkin, Yosi Keller, and Tal Hassner. 2019. Fsgan: Subject agnostic face swapping and reenactment. In ICCV."},{"key":"e_1_3_2_2_45_1","volume-title":"Ganimation: Anatomically-aware facial animation from a single image. In ECCV.","author":"Pumarola Albert","year":"2018","unstructured":"Albert Pumarola , Antonio Agudo , Aleix M Martinez , Alberto Sanfeliu , and Francesc Moreno-Noguer . 2018 . Ganimation: Anatomically-aware facial animation from a single image. In ECCV. Albert Pumarola, Antonio Agudo, Aleix M Martinez, Alberto Sanfeliu, and Francesc Moreno-Noguer. 2018. Ganimation: Anatomically-aware facial animation from a single image. In ECCV."},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"crossref","unstructured":"Yuyang Qian Guojun Yin Lu Sheng Zixuan Chen and Jing Shao. 2020. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In ECCV.  Yuyang Qian Guojun Yin Lu Sheng Zixuan Chen and Jing Shao. 2020. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In ECCV.","DOI":"10.1007\/978-3-030-58610-2_6"},{"key":"e_1_3_2_2_47_1","unstructured":"Zhaofan Qiu Ting Yao and Tao Mei. 2017. Learning spatio-temporal representation with pseudo-3d residual networks. In ICCV.  Zhaofan Qiu Ting Yao and Tao Mei. 2017. Learning spatio-temporal representation with pseudo-3d residual networks. In ICCV."},{"key":"e_1_3_2_2_48_1","volume":"202","author":"Raffel Colin","unstructured":"Colin Raffel , Noam Shazeer , Adam Roberts , Katherine Lee , Sharan Narang , Michael Matena , Yanqi Zhou , Wei Li , and Peter J. Liu. 202 0. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. JMLR (2020). Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. JMLR (2020).","journal-title":"Peter J. Liu."},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"crossref","unstructured":"Andreas Rossler Davide Cozzolino Luisa Verdoliva Christian Riess Justus Thies and Matthias Nie\u00dfner. 2019. Faceforensics++: Learning to detect manipulated facial images. In ICCV.  Andreas Rossler Davide Cozzolino Luisa Verdoliva Christian Riess Justus Thies and Matthias Nie\u00dfner. 2019. Faceforensics++: Learning to detect manipulated facial images. In ICCV.","DOI":"10.1109\/ICCV.2019.00009"},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"crossref","unstructured":"Manuel Ruder Alexey Dosovitskiy and Thomas Brox. 2016. Artistic style transfer for videos. In GCPR.  Manuel Ruder Alexey Dosovitskiy and Thomas Brox. 2016. Artistic style transfer for videos. In GCPR.","DOI":"10.1007\/978-3-319-45886-1_3"},{"key":"e_1_3_2_2_51_1","volume-title":"First order motion model for image animation. arXiv preprint arXiv:2003.00196","author":"Siarohin Aliaksandr","year":"2020","unstructured":"Aliaksandr Siarohin , St\u00e9phane Lathuili\u00e8re , Sergey Tulyakov , Elisa Ricci , and Nicu Sebe . 2020. First order motion model for image animation. arXiv preprint arXiv:2003.00196 ( 2020 ). Aliaksandr Siarohin, St\u00e9phane Lathuili\u00e8re, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2020. First order motion model for image animation. arXiv preprint arXiv:2003.00196 (2020)."},{"key":"e_1_3_2_2_52_1","unstructured":"Karen Simonyan and Andrew Zisserman. [n.d.]. Very deep convolutional networks for large-scale image recognition. In ICLR.  Karen Simonyan and Andrew Zisserman. [n.d.]. Very deep convolutional networks for large-scale image recognition. In ICLR."},{"key":"e_1_3_2_2_53_1","volume-title":"Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In CVPR.","author":"Sun Deqing","year":"2018","unstructured":"Deqing Sun , Xiaodong Yang , Ming-Yu Liu , and Jan Kautz . 2018 . Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In CVPR. Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. 2018. Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In CVPR."},{"key":"e_1_3_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073640"},{"key":"e_1_3_2_2_55_1","volume-title":"Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML.","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc Le . 2019 . Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML. Mingxing Tan and Quoc Le. 2019. Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML."},{"key":"e_1_3_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3323035"},{"key":"e_1_3_2_2_57_1","doi-asserted-by":"crossref","unstructured":"Justus Thies Michael Zollhofer Marc Stamminger Christian Theobalt and Matthias Niessner. 2016. Face2Face: Real-Time Face Capture and Reenactment of RGB Videos. In CVPR.  Justus Thies Michael Zollhofer Marc Stamminger Christian Theobalt and Matthias Niessner. 2016. Face2Face: Real-Time Face Capture and Reenactment of RGB Videos. In CVPR.","DOI":"10.1109\/CVPR.2016.262"},{"key":"e_1_3_2_2_58_1","unstructured":"Hugo Touvron Matthieu Cord Matthijs Douze Francisco Massa Alexandre Sablayrolles and Herv\u00e9 J\u00e9gou. 2021. Training data-efficient image transformers & distillation through attention. In ICML.  Hugo Touvron Matthieu Cord Matthijs Douze Francisco Massa Alexandre Sablayrolles and Herv\u00e9 J\u00e9gou. 2021. Training data-efficient image transformers & distillation through attention. In ICML."},{"key":"e_1_3_2_2_59_1","volume-title":"Icface: Interpretable and controllable face reenactment using gans. In WACV.","author":"Tripathy Soumya","year":"2020","unstructured":"Soumya Tripathy , Juho Kannala , and Esa Rahtu . 2020 . Icface: Interpretable and controllable face reenactment using gans. In WACV. Soumya Tripathy, Juho Kannala, and Esa Rahtu. 2020. Icface: Interpretable and controllable face reenactment using gans. In WACV."},{"key":"e_1_3_2_2_60_1","volume-title":"Visualizing data using t-SNE. JMLR","author":"der Maaten Laurens Van","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton . 2008. Visualizing data using t-SNE. JMLR ( 2008 ). Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. JMLR (2008)."},{"key":"e_1_3_2_2_61_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is All you Need. In NIPS.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is All you Need. In NIPS."},{"key":"e_1_3_2_2_62_1","doi-asserted-by":"crossref","unstructured":"Junke Wang Zuxuan Wu Jingjing Chen Xintong Han Abhinav Shrivastava Ser-Nam Lim and Yu-Gang Jiang. 2022. ObjectFormer for Image Manipulation Detection and Localization. In CVPR.  Junke Wang Zuxuan Wu Jingjing Chen Xintong Han Abhinav Shrivastava Ser-Nam Lim and Yu-Gang Jiang. 2022. ObjectFormer for Image Manipulation Detection and Localization. In CVPR.","DOI":"10.1109\/CVPR52688.2022.00240"},{"key":"e_1_3_2_2_63_1","volume-title":"Efficient Video Transformers with Spatial-Temporal Token Selection. arXiv preprint arXiv:2111.11591","author":"Wang Junke","year":"2021","unstructured":"Junke Wang , Xitong Yang , Hengduo Li , Zuxuan Wu , and Yu-Gang Jiang . 2021. Efficient Video Transformers with Spatial-Temporal Token Selection. arXiv preprint arXiv:2111.11591 ( 2021 ). Junke Wang, Xitong Yang, Hengduo Li, Zuxuan Wu, and Yu-Gang Jiang. 2021. Efficient Video Transformers with Spatial-Temporal Token Selection. arXiv preprint arXiv:2111.11591 (2021)."},{"key":"e_1_3_2_2_64_1","doi-asserted-by":"crossref","unstructured":"Sheng-Yu Wang Oliver Wang Richard Zhang Andrew Owens and Alexei A Efros. 2020. CNN-generated images are surprisingly easy to spot... for now. In CVPR.  Sheng-Yu Wang Oliver Wang Richard Zhang Andrew Owens and Alexei A Efros. 2020. CNN-generated images are surprisingly easy to spot... for now. In CVPR.","DOI":"10.1109\/CVPR42600.2020.00872"},{"key":"e_1_3_2_2_65_1","volume-title":"Deepfake Video Detection Using Convolutional Vision Transformer. arXiv preprint arXiv:2102.11126","author":"Wodajo Deressa","year":"2021","unstructured":"Deressa Wodajo and Solomon Atnafu . 2021. Deepfake Video Detection Using Convolutional Vision Transformer. arXiv preprint arXiv:2102.11126 ( 2021 ). Deressa Wodajo and Solomon Atnafu. 2021. Deepfake Video Detection Using Convolutional Vision Transformer. arXiv preprint arXiv:2102.11126 (2021)."},{"key":"e_1_3_2_2_66_1","volume-title":"Mannat Singh, Eric Mintun, Trevor Darrell, and Ross Girshick.","author":"Xiao Tete","year":"2021","unstructured":"Tete Xiao , Piotr Dollar , Mannat Singh, Eric Mintun, Trevor Darrell, and Ross Girshick. 2021 . Early Convolutions Help Transformers See Better. In NIPS. Tete Xiao, Piotr Dollar, Mannat Singh, Eric Mintun, Trevor Darrell, and Ross Girshick. 2021. Early Convolutions Help Transformers See Better. In NIPS."},{"key":"e_1_3_2_2_67_1","unstructured":"Huijuan Xu Abir Das and Kate Saenko. 2017. R-c3d: Region convolutional 3d network for temporal activity detection. In ICCV.  Huijuan Xu Abir Das and Kate Saenko. 2017. R-c3d: Region convolutional 3d network for temporal activity detection. In ICCV."},{"key":"e_1_3_2_2_68_1","doi-asserted-by":"crossref","unstructured":"Xin Yang Yuezun Li and Siwei Lyu. 2019. Exposing deep fakes using inconsistent head poses. In ICASSP.  Xin Yang Yuezun Li and Siwei Lyu. 2019. Exposing deep fakes using inconsistent head poses. In ICASSP.","DOI":"10.1109\/ICASSP.2019.8683164"},{"key":"e_1_3_2_2_69_1","doi-asserted-by":"crossref","unstructured":"Yang Yang and Xiaojie Guo. 2020. Generative Landmark Guided Face Inpainting. In PRCV.  Yang Yang and Xiaojie Guo. 2020. Generative Landmark Guided Face Inpainting. In PRCV.","DOI":"10.1007\/978-3-030-60633-6_2"},{"key":"e_1_3_2_2_70_1","volume-title":"Resnest: Split-attention networks. arXiv preprint arXiv:2004.08955","author":"Zhang Hang","year":"2020","unstructured":"Hang Zhang , Chongruo Wu , Zhongyue Zhang , Yi Zhu , Haibin Lin , Zhi Zhang , Yue Sun , Tong He , Jonas Mueller , R Manmatha , 2020 . Resnest: Split-attention networks. arXiv preprint arXiv:2004.08955 (2020). Hang Zhang, Chongruo Wu, Zhongyue Zhang, Yi Zhu, Haibin Lin, Zhi Zhang, Yue Sun, Tong He, Jonas Mueller, R Manmatha, et al. 2020. Resnest: Split-attention networks. arXiv preprint arXiv:2004.08955 (2020)."},{"key":"e_1_3_2_2_71_1","doi-asserted-by":"crossref","unstructured":"Hanqing Zhao Wenbo Zhou Dongdong Chen Tianyi Wei Weiming Zhang and Nenghai Yu. 2021. Multi-attentional Deepfake Detection. In CVPR.  Hanqing Zhao Wenbo Zhou Dongdong Chen Tianyi Wei Weiming Zhang and Nenghai Yu. 2021. Multi-attentional Deepfake Detection. In CVPR.","DOI":"10.1109\/CVPR46437.2021.00222"},{"key":"e_1_3_2_2_72_1","doi-asserted-by":"crossref","unstructured":"Peng Zhou Xintong Han Vlad I Morariu and Larry S Davis. 2017. Two-stream neural networks for tampered face detection. In CVPRW.  Peng Zhou Xintong Han Vlad I Morariu and Larry S Davis. 2017. Two-stream neural networks for tampered face detection. In CVPRW.","DOI":"10.1109\/CVPRW.2017.229"},{"key":"e_1_3_2_2_73_1","unstructured":"Xizhou Zhu Weijie Su Lewei Lu Bin Li Xiaogang Wang and Jifeng Dai. 2021. Deformable {DETR}: Deformable Transformers for End-to-End Object Detection. In ICLR.  Xizhou Zhu Weijie Su Lewei Lu Bin Li Xiaogang Wang and Jifeng Dai. 2021. Deformable {DETR}: Deformable Transformers for End-to-End Object Detection. In ICLR."},{"key":"e_1_3_2_2_74_1","unstructured":"Bojia Zi Minghao Chang Jingjing Chen Xingjun Ma and Yu-Gang Jiang. 2020. WildDeepfake: A Challenging Real-World Dataset for Deepfake Detection. In ACM MM.  Bojia Zi Minghao Chang Jingjing Chen Xingjun Ma and Yu-Gang Jiang. 2020. WildDeepfake: A Challenging Real-World Dataset for Deepfake Detection. In ACM MM."}],"event":{"name":"ICMR '22: International Conference on Multimedia Retrieval","location":"Newark NJ USA","acronym":"ICMR '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 2022 International Conference on Multimedia Retrieval"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3512527.3531415","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3512527.3531415","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:12Z","timestamp":1750188612000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3512527.3531415"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,27]]},"references-count":74,"alternative-id":["10.1145\/3512527.3531415","10.1145\/3512527"],"URL":"https:\/\/doi.org\/10.1145\/3512527.3531415","relation":{},"subject":[],"published":{"date-parts":[[2022,6,27]]},"assertion":[{"value":"2022-06-27","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}