{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T16:22:30Z","timestamp":1781367750025,"version":"3.54.1"},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272020"],"award-info":[{"award-number":["62272020"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100003787","name":"Hebei Natural Science Foundation","doi-asserted-by":"crossref","award":["No. F2024202047"],"award-info":[{"award-number":["No. F2024202047"]}],"id":[{"id":"10.13039\/501100003787","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100021171","name":"Guangdong Basic and Applied Basic Research Foundation","doi-asserted-by":"crossref","award":["2024A1515012536"],"award-info":[{"award-number":["2024A1515012536"]}],"id":[{"id":"10.13039\/501100021171","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Macau Science and Technology Development Fund","award":["001\/2024\/SKL, 0119\/2024\/RIB2, and 0022\/2022\/A1"],"award-info":[{"award-number":["001\/2024\/SKL, 0119\/2024\/RIB2, and 0022\/2022\/A1"]}]},{"DOI":"10.13039\/501100004733","name":"University of Macau","doi-asserted-by":"crossref","award":["MYRG-CRG2025-00031-FST and MYRG-GRG2025-00086-FST"],"award-info":[{"award-number":["MYRG-CRG2025-00031-FST and MYRG-GRG2025-00086-FST"]}],"id":[{"id":"10.13039\/501100004733","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Fundamental Research Funds for Central Universities","award":["SKLCCSE-2025ZX-23"],"award-info":[{"award-number":["SKLCCSE-2025ZX-23"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>With advances in generation techniques, malicious users can easily generate deepfake videos, which can cause severe social problems and trust issues. Therefore, deepfake video detection has received increasing attention in recent years. Given that forgery clues are often subtle and imperceptible, effective detection relies heavily on multi-grained learning. However, existing approaches fail to systematically incorporate multi-grained learning across the key components of network training\u2014namely, the training data, network structure, and supervision strategy\u2014thus limiting their performance. In this article, we propose a multi-grained parallel spatio-temporal deepfake video detection architecture, which introduces a novel framework to mine more discriminative deepfake cues throughout the training pipeline. Firstly, we design a parallel spatio-temporal network combined with a cross-guided mechanism to concurrently extract frame-level spatial features and patch-level temporal features, while leveraging the relationship between spatial artifacts and temporal inconsistencies to enable multi-grained spatio-temporal synchronous learning. Secondly, we propose segment-level data augmentation strategies, including frame-random consistent self-blending and spatio-temporal data augmentation, which improve training data diversity at both frame and patch levels, thereby improving the model\u2019s ability to learn comprehensive deepfake representations. Finally, we construct a multi-grained supervision, comprising a patch-level temporal loss, a distance-based frame-level spatial loss, and a standard segment-level loss, for subtle deepfake feature learning. Extensive experiments demonstrate that our method possesses strong robustness and the generalization ability outperforms the current state-of-the-art methods across a series of deepfake datasets, including FaceForensics++, CelebDF, DFDC, DeeperForensics, and Faceshifter, on average.<\/jats:p>","DOI":"10.1145\/3789507","type":"journal-article","created":{"date-parts":[[2026,1,28]],"date-time":"2026-01-28T13:16:32Z","timestamp":1769606192000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["A Multi-Grained Parallel Spatio-Temporal Learning Architecture for Deepfake Video Detection"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-0960-0581","authenticated-orcid":false,"given":"Hui","family":"Miao","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, Beihang University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4592-8083","authenticated-orcid":false,"given":"Yuanfang","family":"Guo","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Beihang University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9330-2662","authenticated-orcid":false,"given":"Leo Yu","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Information and Communication Technology, Griffith University, Gold Coast, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6015-2618","authenticated-orcid":false,"given":"Jiantao","family":"Zhou","sequence":"additional","affiliation":[{"name":"Department of Computer and Information Science, University of Macau, Macau, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8001-2703","authenticated-orcid":false,"given":"Yunhong","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Beihang University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,3,23]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"GitHub. 2021. PyTorch Library for CAM Methods. Retrieved May 13 2025 from https:\/\/github.com\/jacobgil\/pytorch-grad-cam"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/WIFS.2018.8630761"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.116"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2014.2336244"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00408"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58574-7_7"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00114"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01789"},{"key":"e_1_3_1_11_2","unstructured":"GitHub. [n.d.]. Deepfakes. Retrieved February 20 2025 from https:\/\/github.com\/deepfakes\/faceswap"},{"key":"e_1_3_1_12_2","unstructured":"Brian Dolhansky Joanna Bitton Ben Pflaum Jikuo Lu Russ Howes Menglin Wang and Cristian Canton Ferrer. 2020. The deepfake detection challenge (DFDC) dataset. arXiv:2006.07397. Retrieved from https:\/\/arxiv.org\/abs\/2006.07397"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19781-9_2"},{"key":"e_1_3_1_14_2","unstructured":"Alexey Dosovitskiy. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_1_15_2","unstructured":"GitHub. [n.d.]. FaceSwap. Retrieved February 20 2025 from https:\/\/github.com\/MarekKowalski\/FaceSwap"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3536426"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i1.19955"},{"key":"e_1_3_1_18_2","first-page":"4517","article-title":"Delving into sequential patches for deepfake detection","volume":"35","author":"Guan Jiazhi","year":"2022","unstructured":"Jiazhi Guan, Hang Zhou, Zhibin Hong, Errui Ding, Jingdong Wang, Chengbin Quan, and Youjian Zhao. 2022. Delving into sequential patches for deepfake detection. In Advances in Neural Information Processing Systems, Vol. 35, 4517\u20134530.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01453"},{"key":"e_1_3_1_20_2","first-page":"5039","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Haliassos Alexandros","year":"2021","unstructured":"Alexandros Haliassos, Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. 2021. Lips don\u2019t lie: A generalisable and robust approach to face forgery detection. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 5039\u20135049."},{"key":"e_1_3_1_21_2","unstructured":"Yue-Hua Han Tai-Ming Huang Shu-Tzu Lo Po-Han Huang Kai-Lung Hua and Jun-Cheng Chen. 2024. Towards more general video-based deepfake detection through facial feature guided adaptation for foundation model. arXiv:2404.05583. Retrieved from https:\/\/arxiv.org\/abs\/2404.05583"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00296"},{"key":"e_1_3_1_23_2","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv:1412.6980. Retrieved from https:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00639"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00512"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00505"},{"key":"e_1_3_1_27_2","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"Li Yuezun","year":"2019","unstructured":"Yuezun Li and Siwei Lyu. 2019. Exposing DeepFake videos by detecting face warping artifacts. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00327"},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","unstructured":"Li Lin Neeraj Gupta Yue Zhang Hainan Ren Chun-Hao Liu Feng Ding Xin Wang Xin Li Luisa Verdoliva and Shu Hu. 2024. Detecting multimedia generated by large AI models: A survey. arXiv:2402.00045. Retrieved from https:\/\/arxiv.org\/abs\/2402.00045","DOI":"10.36227\/techrxiv.170723324.44685515\/v1"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00808"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01647"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/BTAS46853.2019.9185974"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i3.16344"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2024.3441821"},{"key":"e_1_3_1_35_2","first-page":"8748","volume-title":"International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00009"},{"issue":"1","key":"e_1_3_1_37_2","first-page":"80","article-title":"Recurrent convolutional strategies for face manipulation detection in videos","volume":"3","author":"Sabir Ekraam","year":"2019","unstructured":"Ekraam Sabir, Jiaxin Cheng, Ayush Jaiswal, Wael AbdAlmageed, Iacopo Masi, and Prem Natarajan. 2019. Recurrent convolutional strategies for face manipulation detection in videos. Interfaces (GUI) 3, 1 (2019), 80\u201387.","journal-title":"Interfaces (GUI)"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01816"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00502"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i4.25658"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3323035"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/2929464.2929475"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3588574"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00402"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.02071"},{"key":"e_1_3_1_46_2","unstructured":"Zhiyuan Yan Yandan Zhao Shen Chen Mingyi Guo Xinghe Fu Taiping Yao Shouhong Ding and Li Yuan. 2024. Generalizing deepfake video detection with plug-and-play: Video-level blending and spatiotemporal adapter tuning. arXiv:2408.17065. Retrieved from https:\/\/arxiv.org\/abs\/2408.17065"},{"issue":"2","key":"e_1_3_1_47_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3672566","article-title":"Deepfake video detection using facial feature points and Ch-Transformer","volume":"21","author":"Yang Rui","year":"2024","unstructured":"Rui Yang, Rushi Lan, Zhenrong Deng, Xiaonan Luo, and Xiyan Sun. 2024. Deepfake video detection using facial feature points and Ch-Transformer. ACM Transactions on Multimedia Computing, Communications, and Applications 21, 2 (2024), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683164"},{"key":"e_1_3_1_49_2","unstructured":"Andrii Yermakov Jan Cech and Jiri Matas. 2025. Unlocking the hidden potential of CLIP in generalizable deepfake detection. arXiv:2503.19683. Retrieved from https:\/\/arxiv.org\/abs\/2503.19683"},{"key":"e_1_3_1_50_2","unstructured":"ZAO. [n.d.]. ZAO App. Retrieved February 20 2025 from https:\/\/zaodownload.com\/"},{"issue":"2","key":"e_1_3_1_51_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3664654","article-title":"Spatiotemporal inconsistency learning and interactive fusion for deepfake video detection","volume":"21","author":"Zhang Dengyong","year":"2024","unstructured":"Dengyong Zhang, Wenjie Zhu, Xin Liao, Feifan Qi, Gaobo Yang, and Xiangling Ding. 2024. Spatiotemporal inconsistency learning and interactive fusion for deepfake video detection. ACM Transactions on Multimedia Computing, Communications, and Applications 21, 2 (2024), 1\u201324.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01332"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2023.3239223"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00222"},{"issue":"2","key":"e_1_3_1_55_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3651311","article-title":"Audio-visual contrastive pre-train for face forgery detection","volume":"21","author":"Zhao Hanqing","year":"2024","unstructured":"Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Weiming Zhang, Ying Guo, Zhen Cheng, Pengfei Yan, and Nenghai Yu. 2024. Audio-visual contrastive pre-train for face forgery detection. ACM Transactions on Multimedia Computing, Communications, and Applications 21, 2 (2024), 1\u201316.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01477"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3789507","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,23]],"date-time":"2026-03-23T15:51:32Z","timestamp":1774281092000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3789507"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,23]]},"references-count":55,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3789507"],"URL":"https:\/\/doi.org\/10.1145\/3789507","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,23]]},"assertion":[{"value":"2025-06-10","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-23","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}