{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T03:31:56Z","timestamp":1770175916357,"version":"3.49.0"},"reference-count":235,"publisher":"Association for Computing Machinery (ACM)","issue":"7","funder":[{"name":"Postdoctoral Fellowship Program and China Postdoctoral Science Foundation","award":["2024M764093, BX20250485"],"award-info":[{"award-number":["2024M764093, BX20250485"]}]},{"name":"Beijing Natural Science Foundation","award":["4254100"],"award-info":[{"award-number":["4254100"]}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["KG16336301"],"award-info":[{"award-number":["KG16336301"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>Talking head synthesis, an advanced method for generating portrait videos from a still image driven by specific content, has garnered widespread attention in virtual reality, augmented reality and game production. Recently, significant breakthroughs have been made with the introduction of novel models such as the transformer and the diffusion model. Current methods can not only generate new content but also edit the generated material. This survey systematically reviews the technology, categorizing it into three pivotal domains: portrait generation, driving mechanisms, and editing techniques. We summarize milestone studies and critically analyze their innovations and shortcomings within each domain. Additionally, we organize an extensive collection of datasets and provide a thorough performance analysis of current methodologies based on various evaluation metrics, aiming to furnish a clear framework and robust data support for future research. Finally, we explore application scenarios of talking head synthesis, illustrate them with specific cases, and examine potential future directions. Moreover, we present a first\u2010of\u2010its\u2010kind system\u2010level analysis of integration\u2010specific challenges\u2014such as identity consistency drift, interface incompatibility, and real\u2010time synchronization\u2014that emerge only when portrait generation, driving, and editing are jointly deployed.<\/jats:p>","DOI":"10.1145\/3785656","type":"journal-article","created":{"date-parts":[[2025,12,19]],"date-time":"2025-12-19T12:03:12Z","timestamp":1766145792000},"page":"1-43","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["A Survey of Talking Head Synthesis Techniques: Portrait Generation, Driving Mechanisms, and Editing"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3711-3821","authenticated-orcid":false,"given":"Ming","family":"Meng","sequence":"first","affiliation":[{"name":"School of Data Science and Intelligent Media, State Key Laboratory of Media Convergence and Communication, Key Laboratory of Media Audio & Video, Communication University of China","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0832-8697","authenticated-orcid":false,"given":"Yufei","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Computer and Communication Engineering, University of Science and Technology Beijing","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-9495-2907","authenticated-orcid":false,"given":"Bo","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Data Science and Intelligent Media, Communication University of China","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0048-2084","authenticated-orcid":false,"given":"Yonggui","family":"Zhu","sequence":"additional","affiliation":[{"name":"School of Data Science and Intelligent Media, Communication University of China","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-2344-1025","authenticated-orcid":false,"given":"Weimin","family":"Shi","sequence":"additional","affiliation":[{"name":"The State Key Laboratory of Virtual Reality Technology and Systems, Beihang University","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-0894-5824","authenticated-orcid":false,"given":"Maxwell","family":"Wen","sequence":"additional","affiliation":[{"name":"Leibowitz AI, University of Electronic Science and Technology of China","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6324-1712","authenticated-orcid":false,"given":"Zhaoxin","family":"Fan","sequence":"additional","affiliation":[{"name":"Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,2,3]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"crossref","unstructured":"Francisca Adoma Acheampong Henry Nunoo-Mensah and Wenyu Chen. 2021. Transformer models for text-based emotion detection: A review of BERT-based approaches. Artificial Intelligence Review 54 8 (2021) 5789\u20135829.","DOI":"10.1007\/s10462-021-09958-2"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/WIFS.2018.8630761"},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","unstructured":"Triantafyllos Afouras Joon Son Chung Andrew Senior Oriol Vinyals and Andrew Zisserman. 2018. Deep audio-visual speech recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 44 12 (2018) 8717\u20138727.","DOI":"10.1109\/TPAMI.2018.2889052"},{"key":"e_1_3_1_5_2","unstructured":"Triantafyllos Afouras Joon Son Chung and Andrew Zisserman. 2018. LRS3-TED: A large-scale dataset for visual speech recognition. arXiv:1809.00496. Retrieved from https:\/\/arxiv.org\/abs\/1809.00496"},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","unstructured":"Petar S Aleksic and Aggelos K Katsaggelos. 2004. Speech-to-video synthesis using MPEG-4 compliant visual features. IEEE Transactions on Circuits and Systems for Video Technology 14 5 (2004) 682\u2013692.","DOI":"10.1109\/TCSVT.2004.826760"},{"key":"e_1_3_1_7_2","doi-asserted-by":"crossref","unstructured":"Najwa Alghamdi Steve Maddock Ricard Marxer Jon Barker and Guy J Brown. 2018. A corpus of audio-visual lombard speech with frontal and profile views. The Journal of the Acoustical Society of America 143 6 (2018) EL523\u2013EL529.","DOI":"10.1121\/1.5042758"},{"key":"e_1_3_1_8_2","unstructured":"Alia Abdul-Hassan and Ameer Badr. 1970. VoxCeleb1: Speaker age-group classification using probabilistic neural network. The International Arab Journal of Information Technology (IAJIT) 19 06 (1970) 115\u2013121."},{"key":"e_1_3_1_9_2","unstructured":"Wissam Antoun Fady Baly and Hazem Hajj. 2020. Arabert: Transformer-based model for arabic language understanding. arXiv:2003.00104. Retrieved from https:\/\/arxiv.org\/abs\/2003.00104"},{"key":"e_1_3_1_10_2","first-page":"5470","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Barron Jonathan T","year":"2022","unstructured":"Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. 2022. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 5470\u20135479."},{"key":"e_1_3_1_11_2","unstructured":"Alexander Bergman Petr Kellnhofer Wang Yifan Eric Chan David Lindell and Gordon Wetzstein. 2022. Generative neural articulated radiance fields. Advances in Neural Information Processing Systems 35 (2022) 19900\u201319916."},{"key":"e_1_3_1_12_2","first-page":"580","volume-title":"Proceedings of the 16th International Conference on Computational Processing of Portuguese","author":"Bernardo Brayan","year":"2024","unstructured":"Brayan Bernardo and Paula Costa. 2024. A speech-driven talking head based on a two-stage generative framework. In Proceedings of the 16th International Conference on Computational Processing of Portuguese. 580\u2013586."},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Dan Bigioi Shubhajit Basak Micha\u0142 Stypu\u0142kowski Maciej Zieba Hugh Jordan Rachel McDonnell and Peter Corcoran. 2024. Speech driven video editing via an audio-conditioned diffusion model. Image and Vision Computing 142 (2024) 104911.","DOI":"10.1016\/j.imavis.2024.104911"},{"key":"e_1_3_1_14_2","first-page":"157","volume-title":"Proceedings of the Seminal Graphics Papers: Pushing the Boundaries, Volume 2","author":"Blanz Volker","year":"1999","unstructured":"Volker Blanz and Thomas Vetter. 1999. A morphable model for the synthesis of 3D faces. In Proceedings of the Seminal Graphics Papers: Pushing the Boundaries, Volume 2. 157\u2013164."},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","unstructured":"Sam Bond-Taylor Adam Leach Yang Long and Chris G Willcocks. 2021. Deep generative modelling: A comparative review of vaes gans normalizing flows energy-based and autoregressive models. IEEE Transactions on Pattern Analysis and Machine Intelligence 44 11 (2021) 7327\u20137347.","DOI":"10.1109\/TPAMI.2021.3116668"},{"key":"e_1_3_1_16_2","unstructured":"Stella Bounareli Vasileios Argyriou and Georgios Tzimiropoulos. 2022. Finding directions in GAN\u2019s latent space for neural face reenactment. British Machine Vision Conference (BMVC) (2022)."},{"key":"e_1_3_1_17_2","first-page":"7149","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV)","author":"Bounareli Stella","year":"2023","unstructured":"Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras, and Georgios Tzimiropoulos. 2023. HyperReenact: One-shot reenactment via jointly learning to refine and retarget faces. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV). 7149\u20137159."},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","unstructured":"Carlos Busso Srinivas Parthasarathy Alec Burmania Mohammed AbdelWahab Najmeh Sadoughi and Emily Mower Provost. 2016. MSP-IMPROV: An acted corpus of dyadic interactions to study emotion perception. IEEE Transactions on Affective Computing 8 1 (2016) 67\u201380.","DOI":"10.1109\/TAFFC.2016.2515617"},{"key":"e_1_3_1_19_2","first-page":"3981","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Cai Shengqu","year":"2022","unstructured":"Shengqu Cai, Anton Obukhov, Dengxin Dai, and Luc Van Gool. 2022. Pix2nerf: Unsupervised conditional p-gan for single image to neural radiance fields translation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 3981\u20133990."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV56688.2023.00440"},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","unstructured":"Houwei Cao David G Cooper Michael K Keutmann Ruben C Gur Ani Nenkova and Ragini Verma. 2014. Crema-d: Crowd-sourced emotional multimodal actors dataset. IEEE Transactions on Affective Computing 5 4 (2014) 377\u2013390.","DOI":"10.1109\/TAFFC.2014.2336244"},{"key":"e_1_3_1_22_2","first-page":"16123","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Chan Eric R","year":"2022","unstructured":"Eric R Chan, Connor Z Lin, Matthew A Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas J Guibas, Jonathan Tremblay, Sameh Khamis, et\u00a0al. 2022. Efficient geometry-aware 3d generative adversarial networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 16123\u201316133."},{"key":"e_1_3_1_23_2","unstructured":"Lele Chen Guofeng Cui Ziyi Kou Haitian Zheng and Chenliang Xu. 2020. What comprises a good talking-head video generation?: A survey and benchmark. arXiv:2005.03201. Retrieved from https:\/\/arxiv.org\/abs\/2005.03201"},{"key":"e_1_3_1_24_2","first-page":"35","volume-title":"Proceedings of the European Conference on Computer Vision(ECCV)","author":"Chen Lele","year":"2020","unstructured":"Lele Chen, Guofeng Cui, Celong Liu, Zhong Li, Ziyi Kou, Yi Xu, and Chenliang Xu. 2020. Talking-head generation with rhythmic head motion. In Proceedings of the European Conference on Computer Vision(ECCV). Springer, 35\u201351."},{"key":"e_1_3_1_25_2","first-page":"520","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Chen Lele","year":"2018","unstructured":"Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. 2018. Lip movements generation at a glance. In Proceedings of the European Conference on Computer Vision (ECCV). 520\u2013535."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_32"},{"key":"e_1_3_1_27_2","first-page":"7832","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Chen Lele","year":"2019","unstructured":"Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. 2019. Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 7832\u20137841."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i3.32241"},{"key":"e_1_3_1_29_2","first-page":"1","volume-title":"Proceedings of the SIGGRAPH Asia 2022 Conference Papers","author":"Cheng Kun","year":"2022","unstructured":"Kun Cheng, Xiaodong Cun, Yong Zhang, Menghan Xia, Fei Yin, Mingrui Zhu, Xuan Wang, Jue Wang, and Nannan Wang. 2022. Videoretalking: Audio-based lip synchronization for talking head video editing in the wild. In Proceedings of the SIGGRAPH Asia 2022 Conference Papers. 1\u20139."},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","unstructured":"Kyusun Cho Joungbin Lee Heeji Yoon Yeobin Hong Jaehoon Ko Sangjun Ahn and Seungryong Kim. 2024. GaussianTalker: Real-time high-fidelity talking head synthesis with audio-driven 3D gaussian splatting. arXiv:2404.16012. Retrieved from https:\/\/arxiv.org\/abs\/2404.16012","DOI":"10.1145\/3664647.3681627"},{"key":"e_1_3_1_31_2","doi-asserted-by":"crossref","unstructured":"Kyoungho Choi Ying Luo and Jenq-Neng Hwang. 2001. Hidden markov model inversion for audio-to-visual conversion in an MPEG-4 facial animation system. Journal of VLSI Signal Processing Systems for Signal Image and Video Technology 29 (2001) 51\u201361.","DOI":"10.1023\/A:1011171430700"},{"key":"e_1_3_1_32_2","first-page":"175","volume-title":"Proceedings of the 1999 IEEE 3rd Workshop on Multimedia Signal Processing (Cat. No. 99TH8451)","author":"Choi Kyoung Ho","year":"1999","unstructured":"Kyoung Ho Choi and Jenq-Neng Hwang. 1999. Baum-welch hidden markov model inversion for reliable audio-to-visual conversion. In Proceedings of the 1999 IEEE 3rd Workshop on Multimedia Signal Processing (Cat. No. 99TH8451). IEEE, 175\u2013180."},{"key":"e_1_3_1_33_2","unstructured":"Erika Chuang and Chris Bregler. 2002. Performance driven facial animation using blendshape interpolation. Computer Science Technical Report Stanford University 2 2 (2002) 3."},{"key":"e_1_3_1_34_2","doi-asserted-by":"crossref","unstructured":"Joon Son Chung Arsha Nagrani and Andrew Zisserman. 2018. Voxceleb2: Deep speaker recognition. arXiv:1806.05622. Retrieved from https:\/\/arxiv.org\/abs\/1806.05622","DOI":"10.21437\/Interspeech.2018-1929"},{"key":"e_1_3_1_35_2","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1007\/978-3-319-54184-6_6","volume-title":"Computer Vision\u2013ACCV 2016: 13th Asian Conference on Computer Vision, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part II 13","author":"Chung Joon Son","year":"2017","unstructured":"Joon Son Chung and Andrew Zisserman. 2017. Lip reading in the wild. In Computer Vision\u2013ACCV 2016: 13th Asian Conference on Computer Vision, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part II 13. Springer, 87\u2013103."},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","unstructured":"Martin Cooke Jon Barker Stuart Cunningham and Xu Shao. 2006. An audio-visual corpus for speech perception and automatic speech recognition. Journal of the Acoustical Society of America 120 5 (2006) 2421.","DOI":"10.1121\/1.2229005"},{"key":"e_1_3_1_37_2","first-page":"10101","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Cudeiro Daniel","year":"2019","unstructured":"Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black. 2019. Capture, learning, and synthesis of 3D speaking styles. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 10101\u201310111."},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","unstructured":"Andrzej Czyzewski Bozena Kostek Piotr Bratoszewski Jozef Kotus and Marcin Szykulski. 2017. An audio-visual corpus for multimodal automatic speech recognition. Journal of Intelligent Information Systems 49 (2017) 167\u2013192.","DOI":"10.1007\/s10844-016-0438-z"},{"key":"e_1_3_1_39_2","volume-title":"Proceedings of the INTERSPEECH 2019-20th Annual Conference of the International Speech Communication Association","author":"Dahmani Sara","year":"2019","unstructured":"Sara Dahmani, Vincent Colotte, Val\u00e9rian Girard, and Slim Ouni. 2019. Conditional variational auto-encoder for text-driven expressive audiovisual speech synthesis. In Proceedings of the INTERSPEECH 2019-20th Annual Conference of the International Speech Communication Association."},{"key":"e_1_3_1_40_2","first-page":"20311","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Dan\u011b\u010dek Radek","year":"2022","unstructured":"Radek Dan\u011b\u010dek, Michael J Black, and Timo Bolkart. 2022. Emoca: Emotion driven monocular face capture and animation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 20311\u201320322."},{"key":"e_1_3_1_41_2","first-page":"12882","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Deng Kangle","year":"2022","unstructured":"Kangle Deng, Andrew Liu, Jun-Yan Zhu, and Deva Ramanan. 2022. Depth-supervised nerf: Fewer views and faster training for free. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 12882\u201312891."},{"key":"e_1_3_1_42_2","unstructured":"Michail Christos Doukas Evangelos Ververas Viktoriia Sharmanska and Stefanos Zafeiriou. 2023. Free-headgan: Neural talking head synthesis with explicit gaze control. IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)."},{"key":"e_1_3_1_43_2","first-page":"14398","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV)","author":"Doukas Michail Christos","year":"2021","unstructured":"Michail Christos Doukas, Stefanos Zafeiriou, and Viktoriia Sharmanska. 2021. Headgan: One-shot neural head synthesis and editing. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV). 14398\u201314407."},{"key":"e_1_3_1_44_2","first-page":"8498","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Drobyshev Nikita","year":"2024","unstructured":"Nikita Drobyshev, Antoni Bigata Casademunt, Konstantinos Vougioukas, Zoe Landgraf, Stavros Petridis, and Maja Pantic. 2024. Emoportraits: Emotion-enhanced multimodal one-shot head avatars. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 8498\u20138507."},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","first-page":"372","DOI":"10.1007\/978-3-319-93764-9_35","volume-title":"Proceedings of the Latent Variable Analysis and Signal Separation: 14th International Conference, LVA\/ICA 2018, Guildford, UK, July 2\u20135, 2018, Proceedings 14","author":"Eskimez Sefik Emre","year":"2018","unstructured":"Sefik Emre Eskimez, Ross K Maddox, Chenliang Xu, and Zhiyao Duan. 2018. Generating talking face landmarks from speech. In Proceedings of the Latent Variable Analysis and Signal Separation: 14th International Conference, LVA\/ICA 2018, Guildford, UK, July 2\u20135, 2018, Proceedings 14. Springer, 372\u2013381."},{"key":"e_1_3_1_46_2","doi-asserted-by":"crossref","unstructured":"Sefik Emre Eskimez You Zhang and Zhiyao Duan. 2021. Speech driven talking face generation from a single image and an emotion condition. IEEE Transactions on Multimedia 24 (2021) 3480\u20133490.","DOI":"10.1109\/TMM.2021.3099900"},{"key":"e_1_3_1_47_2","unstructured":"Patrick Esser Robin Rombach Andreas Blattmann and Bjorn Ommer. 2021. Imagebart: Bidirectional context with multinomial diffusion for autoregressive image synthesis. Advances in Neural Information Processing Systems 34 (2021) 3518\u20133532."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01268"},{"key":"e_1_3_1_49_2","first-page":"4884","volume-title":"Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Fan Bo","year":"2015","unstructured":"Bo Fan, Lijuan Wang, Frank K Soong, and Lei Xie. 2015. Photo-real talking head with deep bidirectional LSTM. In Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 4884\u20134888."},{"key":"e_1_3_1_50_2","doi-asserted-by":"crossref","unstructured":"Bo Fan Lei Xie Shan Yang Lijuan Wang and Frank K Soong. 2016. A deep bidirectional LSTM approach for video-realistic talking head. Multimedia Tools and Applications 75 (2016) 5287\u20135309.","DOI":"10.1007\/s11042-015-2944-3"},{"key":"e_1_3_1_51_2","first-page":"18770","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Fan Yingruo","year":"2022","unstructured":"Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura. 2022. Faceformer: Speech-driven 3d facial animation with transformers. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 18770\u201318780."},{"key":"e_1_3_1_52_2","unstructured":"Guanwen Feng Zhihao Qian Yunan Li Siyu Jin Qiguang Miao and Chi-Man Pun. 2024. LES-talker: Fine-grained emotion editing for talking head generation in linear emotion space. arXiv:2411.09268. Retrieved from https:\/\/arxiv.org\/abs\/2411.09268"},{"key":"e_1_3_1_53_2","doi-asserted-by":"crossref","unstructured":"Yao Feng Haiwen Feng Michael J Black and Timo Bolkart. 2021. Learning an animatable detailed 3D face model from in-the-wild images. ACM Transactions on Graphics (ToG) 40 4 (2021) 1\u201313.","DOI":"10.1145\/3450626.3459936"},{"key":"e_1_3_1_54_2","doi-asserted-by":"crossref","unstructured":"Ohad Fried Ayush Tewari Michael Zollh\u00f6fer Adam Finkelstein Eli Shechtman Dan B Goldman Kyle Genova Zeyu Jin Christian Theobalt and Maneesh Agrawala. 2019. Text-based editing of talking-head video. ACM Transactions on Graphics (TOG) 38 4 (2019) 1\u201314.","DOI":"10.1145\/3306346.3323028"},{"key":"e_1_3_1_55_2","first-page":"8649","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Gafni Guy","year":"2021","unstructured":"Guy Gafni, Justus Thies, Michael Zollhofer, and Matthias Nie\u00dfner. 2021. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 8649\u20138658."},{"key":"e_1_3_1_56_2","first-page":"5609","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Gao Yue","year":"2023","unstructured":"Yue Gao, Yuan Zhou, Jinglu Wang, Xiao Li, Xiang Ming, and Yan Lu. 2023. High-fidelity and freely controllable talking head video generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 5609\u20135619."},{"key":"e_1_3_1_57_2","doi-asserted-by":"crossref","unstructured":"Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2020. Generative adversarial networks. Communications of the ACM 63 11 (2020) 139\u2013144.","DOI":"10.1145\/3422622"},{"key":"e_1_3_1_58_2","doi-asserted-by":"crossref","unstructured":"Shreyank N Gowda Dheeraj Pandey and Shashank Narayana Gowda. 2023. From pixels to portraits: A comprehensive survey of talking head generation techniques and applications. arXiv:2308.16041. Retrieved from https:\/\/arxiv.org\/abs\/2308.16041","DOI":"10.2139\/ssrn.4573122"},{"key":"e_1_3_1_59_2","first-page":"10696","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Gu Shuyang","year":"2022","unstructured":"Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. 2022. Vector quantized diffusion model for text-to-image synthesis. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 10696\u201310706."},{"key":"e_1_3_1_60_2","first-page":"1","volume-title":"Proceedings of the 2018 15th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS)","author":"G\u00fcera David","year":"2018","unstructured":"David G\u00fcera and Edward J Delp. 2018. Deepfake video detection using recurrent neural networks. In Proceedings of the 2018 15th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS). IEEE, 1\u20136."},{"key":"e_1_3_1_61_2","unstructured":"Jianzhu Guo Dingyun Zhang Xiaoqiang Liu Zhizhou Zhong Yuan Zhang Pengfei Wan and Di Zhang. 2024. Liveportrait: Efficient portrait animation with stitching and retargeting control. arXiv:2407.03168. Retrieved from https:\/\/arxiv.org\/abs\/2407.03168"},{"key":"e_1_3_1_62_2","first-page":"5784","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV)","author":"Guo Yudong","year":"2021","unstructured":"Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, and Juyong Zhang. 2021. Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV). 5784\u20135794."},{"key":"e_1_3_1_63_2","first-page":"2024","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Gupta Sonam","year":"2022","unstructured":"Sonam Gupta, Arti Keshari, and Sukhendu Das. 2022. Rv-gan: Recurrent gan for unconditional video generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 2024\u20132033."},{"key":"e_1_3_1_64_2","unstructured":"Kai Han An Xiao Enhua Wu Jianyuan Guo Chunjing Xu and Yunhe Wang. 2021. Transformer in transformer. Advances in Neural Information Processing Systems 34 (2021) 15908\u201315919."},{"key":"e_1_3_1_65_2","doi-asserted-by":"crossref","unstructured":"Naomi Harte and Eoin Gillen. 2015. TCD-TIMIT: An audio-visual corpus of continuous speech. IEEE Transactions on Multimedia 17 5 (2015) 603\u2013615.","DOI":"10.1109\/TMM.2015.2407694"},{"key":"e_1_3_1_66_2","first-page":"55","volume-title":"Proceedings of the European Conference on Computer Vision(ECCV)","author":"He Qianyun","year":"2024","unstructured":"Qianyun He, Xinya Ji, Yicheng Gong, Yuanxun Lu, Zhengyu Diao, Linjia Huang, Yao Yao, Siyu Zhu, Zhan Ma, Songcen Xu, et\u00a0al. 2024. EmoTalk3D: High-fidelity free-view synthesis of emotional 3D talking head. In Proceedings of the European Conference on Computer Vision(ECCV). Springer, 55\u201372."},{"key":"e_1_3_1_67_2","unstructured":"Martin Heusel Hubert Ramsauer Thomas Unterthiner Bernhard Nessler and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in Neural Information Processing Systems 30 (2017)."},{"key":"e_1_3_1_68_2","unstructured":"Jonathan Ho William Chan Chitwan Saharia Jay Whang Ruiqi Gao Alexey Gritsenko Diederik P Kingma Ben Poole Mohammad Norouzi David J Fleet et\u00a0al. 2022. Imagen video: High definition video generation with diffusion models. arXiv:2210.02303. Retrieved from https:\/\/arxiv.org\/abs\/2210.02303"},{"key":"e_1_3_1_69_2","unstructured":"Jonathan Ho Ajay Jain and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems 33 (2020) 6840\u20136851."},{"key":"e_1_3_1_70_2","first-page":"23062","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(CVPR)","author":"Hong Fa-Ting","year":"2023","unstructured":"Fa-Ting Hong and Dan Xu. 2023. Implicit identity representation conditioned memory compensation network for talking head video generation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(CVPR). 23062\u201323072."},{"key":"e_1_3_1_71_2","first-page":"3397","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Hong Fa-Ting","year":"2022","unstructured":"Fa-Ting Hong, Longhao Zhang, Li Shen, and Dan Xu. 2022. Depth-aware generative adversarial network for talking head video generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 3397\u20133406."},{"key":"e_1_3_1_72_2","first-page":"20374","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Hong Yang","year":"2022","unstructured":"Yang Hong, Bo Peng, Haiyao Xiao, Ligang Liu, and Juyong Zhang. 2022. Headnerf: A real-time nerf-based parametric head model. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 20374\u201320384."},{"key":"e_1_3_1_73_2","first-page":"642","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Hsu Gee-Sern","year":"2022","unstructured":"Gee-Sern Hsu, Chun-Hung Tsai, and Hung-Yi Wu. 2022. Dual-generator face reenactment. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 642\u2013650."},{"key":"e_1_3_1_74_2","unstructured":"Huaibo Huang Ran He Zhenan Sun Tieniu Tan and Zhihang Li2018. Introvae: Introspective variational autoencoders for photographic image synthesis. Advances in Neural Information Processing Systems 31 (2018)."},{"key":"e_1_3_1_75_2","unstructured":"Jiehui Huang Xiao Dong Wenhui Song Zheng Chong Zhenchao Tang Jun Zhou Yuhao Cheng Long Chen Hanhui Li Yiqiang Yan et\u00a0al. 2024. Consistentid: Portrait generation with multimodal fine-grained identity preserving. arXiv:2404.16771. Retrieved from https:\/\/arxiv.org\/abs\/2404.16771"},{"key":"e_1_3_1_76_2","unstructured":"Lianghua Huang Di Chen Yu Liu Yujun Shen Deli Zhao and Jingren Zhou. 2023. Composer: Creative and controllable image synthesis with composable conditions. Article 558 (2023) 21 pages. arXiv preprint arXiv:2302.09778. https:\/\/arxiv.org\/abs\/2302.09778"},{"key":"e_1_3_1_77_2","first-page":"22596","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Huang Mengqi","year":"2023","unstructured":"Mengqi Huang, Zhendong Mao, Zhuowei Chen, and Yongdong Zhang. 2023. Towards accurate image coding: Improved autoregressive image generation with dynamic vector quantization. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 22596\u201322605."},{"key":"e_1_3_1_78_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3551574"},{"key":"e_1_3_1_79_2","first-page":"6080","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Huang Ziqi","year":"2023","unstructured":"Ziqi Huang, Kelvin CK Chan, Yuming Jiang, and Ziwei Liu. 2023. Collaborative diffusion for multi-modal face generation and editing. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 6080\u20136090."},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","unstructured":"Amir Jamaludin Joon Son Chung and Andrew Zisserman. 2019. You said that?: Synthesising talking faces from audio. International Journal of Computer Vision 127 11\u201312 (Dec.2019) 1767\u20131779. DOI:10.1007\/s11263-019-01150-y","DOI":"10.1007\/s11263-019-01150-y"},{"key":"e_1_3_1_81_2","doi-asserted-by":"publisher","DOI":"10.1145\/3528233.3530745"},{"key":"e_1_3_1_82_2","unstructured":"Liming Jiang Bo Dai Wayne Wu and Chen Change Loy. 2021. Deceive d: Adaptive pseudo augmentation for gan training with limited data. Advances in Neural Information Processing Systems 34 (2021) 21655\u201321667."},{"key":"e_1_3_1_83_2","doi-asserted-by":"crossref","unstructured":"Felix Juefei-Xu Run Wang Yihao Huang Qing Guo Lei Ma and Yang Liu. 2022. Countering malicious deepfakes: Survey battleground and horizon. International Journal of Computer Vision 130 7 (2022) 1678\u20131734.","DOI":"10.1007\/s11263-022-01606-8"},{"key":"e_1_3_1_84_2","doi-asserted-by":"crossref","unstructured":"Amina Kammoun Rim Slama Hedi Tabia Tarek Ouni and Mohmed Abid. 2022. Generative adversarial networks for face generation: A survey. ACM Computing Surveys 55 5 (2022) 1\u201337.","DOI":"10.1145\/3527850"},{"key":"e_1_3_1_85_2","doi-asserted-by":"crossref","unstructured":"Tero Karras Timo Aila Samuli Laine Antti Herva and Jaakko Lehtinen. 2017. Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Transactions on Graphics (TOG) 36 4 (2017) 1\u201312.","DOI":"10.1145\/3072959.3073658"},{"key":"e_1_3_1_86_2","unstructured":"Tero Karras Timo Aila Samuli Laine and Jaakko Lehtinen. 2017. Progressive growing of GANs for improved quality stability and variation. arXiv:1710.10196. Retrieved from https:\/\/arxiv.org\/abs\/1710.10196"},{"key":"e_1_3_1_87_2","first-page":"4401","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Karras Tero","year":"2019","unstructured":"Tero Karras, Samuli Laine, and Timo Aila. 2019. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 4401\u20134410."},{"key":"e_1_3_1_88_2","first-page":"8110","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Karras Tero","year":"2020","unstructured":"Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 8110\u20138119."},{"key":"e_1_3_1_89_2","doi-asserted-by":"crossref","unstructured":"Achhardeep Kaur Azadeh Noori Hoshyar Vidya Saikrishna Selena Firmin and Feng Xia. 2024. Deepfake video detection: Challenges and opportunities. Artificial Intelligence Review 57 6 (2024) 159.","DOI":"10.1007\/s10462-024-10810-6"},{"key":"e_1_3_1_90_2","first-page":"345","volume-title":"Proceedings of the European Conference on Computer Vision(ECCV)","author":"Khakhulin Taras","year":"2022","unstructured":"Taras Khakhulin, Vanessa Sklyarova, Victor Lempitsky, and Egor Zakharov. 2022. Realistic one-shot mesh-based head avatars. In Proceedings of the European Conference on Computer Vision(ECCV). Springer, 345\u2013362."},{"key":"e_1_3_1_91_2","unstructured":"Diederik P Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv:1312.6114. Retrieved from https:\/\/arxiv.org\/abs\/1312.6114"},{"key":"e_1_3_1_92_2","first-page":"11523","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Lee Doyup","year":"2022","unstructured":"Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han. 2022. Autoregressive image generation using residual quantization. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 11523\u201311532."},{"key":"e_1_3_1_93_2","unstructured":"Kwangho Lee Patrick Kwon Myung-Ki Lee Namhyuk Ahn and Junsoo Lee. 2023. LPMM: Intuitive pose control for neural talking-head model via landmark-parameter morphable model. arXiv.2305.10456. Retrieved from https:\/\/arxiv.org\/abs\/2305.10456"},{"key":"e_1_3_1_94_2","first-page":"563","volume-title":"Generative adversarial networks for face generation: A survey Pacific Rim International Conference on Artificial Intelligence","author":"Lee Soonkyu","year":"2002","unstructured":"Soonkyu Lee and DongSuk Yook. 2002. Audio-to-visual conversion using hidden markov models. In Generative adversarial networks for face generation: A survey Pacific Rim International Conference on Artificial Intelligence. Springer, 563\u2013570."},{"key":"e_1_3_1_95_2","first-page":"4625","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"39","author":"Li Bonan","year":"2025","unstructured":"Bonan Li, Zicheng Zhang, Xuecheng Nie, Congying Han, Yinhan Hu, Xinmin Qiu, and Tiande Guo. 2025. Styo: Stylize your face in only one-shot. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 4625\u20134633."},{"key":"e_1_3_1_96_2","first-page":"10585","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Hong","year":"2024","unstructured":"Hong Li, Yutang Feng, Song Xue, Xuhui Liu, Bohan Zeng, Shanglin Li, Boyu Liu, Jianzhuang Liu, Shumin Han, and Baochang Zhang. 2024. UV-IDM: Identity-conditioned latent diffusion model for face UV-texture generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10585\u201310595."},{"key":"e_1_3_1_97_2","first-page":"127","volume-title":"Proceedings of the European Conference on Computer Vision(ECCV)","author":"Li Jiahe","year":"2024","unstructured":"Jiahe Li, Jiawei Zhang, Xiao Bai, Jin Zheng, Xin Ning, Jun Zhou, and Lin Gu. 2024. Talkinggaussian: Structure-persistent 3d talking head synthesis via gaussian splatting. In Proceedings of the European Conference on Computer Vision(ECCV). Springer, 127\u2013145."},{"key":"e_1_3_1_98_2","first-page":"7568","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(CVPR)","author":"Li Jiahe","year":"2023","unstructured":"Jiahe Li, Jiawei Zhang, Xiao Bai, Jun Zhou, and Lin Gu. 2023. Efficient region-aware neural radiance fields for high-fidelity talking portrait synthesis. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(CVPR). 7568\u20137578."},{"key":"e_1_3_1_99_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i3.16286"},{"key":"e_1_3_1_100_2","first-page":"17969","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Li Weichuang","year":"2023","unstructured":"Weichuang Li, Longhao Zhang, Dong Wang, Bin Zhao, Zhigang Wang, Mulin Chen, Bang Zhang, Zhongjian Wang, Liefeng Bo, and Xuelong Li. 2023. One-shot high-fidelity talking-head synthesis with deformable neural radiance field. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 17969\u201317978."},{"key":"e_1_3_1_101_2","unstructured":"Xueting Li Shalini De Mello Sifei Liu Koki Nagano Umar Iqbal and Jan Kautz. 2024. Generalizable one-shot 3D neural head avatar. Advances in Neural Information Processing Systems 36 (2024)."},{"key":"e_1_3_1_102_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681644"},{"key":"e_1_3_1_103_2","first-page":"3387","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Liang Borong","year":"2022","unstructured":"Borong Liang, Yan Pan, Zhizhi Guo, Hang Zhou, Zhibin Hong, Xiaoguang Han, Junyu Han, Jingtuo Liu, Errui Ding, and Jingdong Wang. 2022. Expressive talking head generation with granular audio-visual control. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 3387\u20133396."},{"key":"e_1_3_1_104_2","first-page":"3441","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"38","author":"Lin Luoyang","year":"2024","unstructured":"Luoyang Lin, Zutao Jiang, Xiaodan Liang, Liqian Ma, Michael C Kampffmeyer, and Xiaochun Cao. 2024. PTUS: Photo-realistic talking upper-body synthesis via 3D-aware motion decomposition warping. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 3441\u20133449."},{"key":"e_1_3_1_105_2","first-page":"2099","volume-title":"Proceedings of the 2023 IEEE International Conference on Multimedia and Expo (ICME)","author":"Liu Jin","year":"2023","unstructured":"Jin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, and Jizhong Han. 2023. Font: Flow-guided one-shot talking head generation with natural head motions. In Proceedings of the 2023 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 2099\u20132104."},{"key":"e_1_3_1_106_2","first-page":"1","volume-title":"Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Liu Jin","year":"2023","unstructured":"Jin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, and Jizhong Han. 2023. OPT: One-shot pose-controllable talking head generation. In Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1\u20135."},{"key":"e_1_3_1_107_2","unstructured":"Kangning Liu Yu-Chuan Su Ruijin Cang Xuhui Jia and Wei (Alex)Hong. 2023. Controllable one-shot face video synthesis with semantic aware prior. arXiv:2304.14471. Retrieved from https:\/\/arxiv.org\/abs\/2304.14471"},{"key":"e_1_3_1_108_2","doi-asserted-by":"crossref","unstructured":"Meng Liu Da Li Yongqiang Li Xuemeng Song and Liqiang Nie. 2024. Audio-semantic enhanced pose-driven talking head generation. IEEE Transactions on Circuits and Systems for Video Technology 34 11 (2024) 11056\u201311069.","DOI":"10.1109\/TCSVT.2024.3414412"},{"key":"e_1_3_1_109_2","doi-asserted-by":"crossref","unstructured":"Pengfei Liu Wenjin Deng Hengda Li Jintai Wang Yinglin Zheng Yiwei Ding Xiaohu Guo and Ming Zeng. 2024. MusicFace: Music-driven expressive singing face synthesis. Computational Visual Media 10 1 (2024) 119\u2013136.","DOI":"10.1007\/s41095-023-0343-7"},{"key":"e_1_3_1_110_2","doi-asserted-by":"crossref","unstructured":"Shiguang Liu. 2023. Audio-driven talking face generation: A review. Journal of the Audio Engineering Society 71 7\/8 (2023) 408\u2013419.","DOI":"10.17743\/jaes.2022.0081"},{"key":"e_1_3_1_111_2","first-page":"11877","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Liu Yunfan","year":"2019","unstructured":"Yunfan Liu, Qi Li, and Zhenan Sun. 2019. Attribute-aware face aging with wavelet-based generative adversarial networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 11877\u201311886."},{"key":"e_1_3_1_112_2","doi-asserted-by":"crossref","unstructured":"Steven R Livingstone and Frank A Russo. 2018. The ryerson audio-visual database of emotional speech and song (RAVDESS): A dynamic multimodal set of facial and vocal expressions in north american english. PloS one 13 5 (2018) e0196391.","DOI":"10.1371\/journal.pone.0196391"},{"key":"e_1_3_1_113_2","doi-asserted-by":"crossref","unstructured":"Yuanxun Lu Jinxiang Chai and Xun Cao. 2021. Live speech portraits: Real-time photorealistic talking-head animation. ACM Transactions on Graphics (TOG) 40 6 (2021) 1\u201317.","DOI":"10.1145\/3478513.3480484"},{"key":"e_1_3_1_114_2","first-page":"282","volume-title":"Proceedings of the European Conference on Computer Vision(ECCV)","author":"Lu Yongyi","year":"2018","unstructured":"Yongyi Lu, Yu-Wing Tai, and Chi-Keung Tang. 2018. Attribute-guided face generation using conditional cyclegan. In Proceedings of the European Conference on Computer Vision(ECCV). 282\u2013297."},{"key":"e_1_3_1_115_2","unstructured":"Yifeng Ma Suzhen Wang Yu Ding Bowen Ma Tangjie Lv Changjie Fan Zhipeng Hu Zhidong Deng and Xin Yu. 2023. TalkCLIP: Talking head generation with text-guided expressive speaking styles. arXiv:2304.00334. Retrieved from https:\/\/arxiv.org\/abs\/2304.00334"},{"key":"e_1_3_1_116_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i2.25280"},{"key":"e_1_3_1_117_2","unstructured":"Yifeng Ma Shiwei Zhang Jiayu Wang Xiang Wang Yingya Zhang and Zhidong Deng. 2023. DreamTalk: When expressive talking head generation meets diffusion probabilistic models. arXiv.2312.09767. Retrieved from https:\/\/arxiv.org\/abs\/2312.09767"},{"key":"e_1_3_1_118_2","first-page":"16901","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Ma Zhiyuan","year":"2023","unstructured":"Zhiyuan Ma, Xiangyu Zhu, Guo-Jun Qi, Zhen Lei, and Lei Zhang. 2023. Otavatar: One-shot talking face avatar with controllable tri-plane rendering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 16901\u201316910."},{"key":"e_1_3_1_119_2","doi-asserted-by":"crossref","unstructured":"Nadia Magnenat-Thalmann E Primeau and Daniel Thalmann. 1988. Abstract muscle action procedures for human face animation. The Visual Computer 3 (1988) 290\u2013297.","DOI":"10.1007\/BF01914864"},{"key":"e_1_3_1_120_2","unstructured":"Arun Mallya Ting-Chun Wang and Ming-Yu Liu. 2022. Implicit warping for animation with image sets. Advances in Neural Information Processing Systems 35 (2022) 22438\u201322450."},{"key":"e_1_3_1_121_2","unstructured":"Elman Mansimov Emilio Parisotto Lei Jimmy Ba and Ruslan Salakhutdinov. 2016. Generating images from captions with attention. arxiv.1511.02793. Retrieved from https:\/\/arxiv.org\/abs\/1511.02793"},{"key":"e_1_3_1_122_2","doi-asserted-by":"crossref","first-page":"83","DOI":"10.1109\/WACVW.2019.00020","volume-title":"Proceedings of the 2019 IEEE Winter Applications of Computer Vision Workshops (WACVW)","author":"Matern Falko","year":"2019","unstructured":"Falko Matern, Christian Riess, and Marc Stamminger. 2019. Exploiting visual artifacts to expose deepfakes and face manipulations. In Proceedings of the 2019 IEEE Winter Applications of Computer Vision Workshops (WACVW). IEEE, 83\u201392."},{"key":"e_1_3_1_123_2","first-page":"5084","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Men Yifang","year":"2020","unstructured":"Yifang Men, Yiming Mao, Yuning Jiang, Wei-Ying Ma, and Zhouhui Lian. 2020. Controllable person image synthesis with attribute-decomposed gan. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 5084\u20135093."},{"key":"e_1_3_1_124_2","first-page":"13829","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV)","author":"Meshry Moustafa","year":"2021","unstructured":"Moustafa Meshry, Saksham Suri, Larry S Davis, and Abhinav Shrivastava. 2021. Learned spatial representations for few-shot talking-head synthesis. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV). 13829\u201313838."},{"key":"e_1_3_1_125_2","doi-asserted-by":"crossref","unstructured":"Ben Mildenhall Pratul P Srinivasan Matthew Tancik Jonathan T Barron Ravi Ramamoorthi and Ren Ng. 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM 65 1 (2021) 99\u2013106.","DOI":"10.1145\/3503250"},{"key":"e_1_3_1_126_2","first-page":"3290","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Mittal Gaurav","year":"2020","unstructured":"Gaurav Mittal and Baoyuan Wang. 2020. Animating face using disentangled audio representations. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 3290\u20133298."},{"key":"e_1_3_1_127_2","first-page":"2307","volume-title":"Proceedings of the ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Nguyen Huy H","year":"2019","unstructured":"Huy H Nguyen, Junichi Yamagishi, and Isao Echizen. 2019. Capsule-forensics: Using capsule networks to detect forged images and videos. In Proceedings of the ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2307\u20132311."},{"key":"e_1_3_1_128_2","first-page":"4954","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Ni Haomiao","year":"2024","unstructured":"Haomiao Ni, Jiachen Liu, Yuan Xue, and Sharon X Huang. 2024. 3D-aware talking-head video motion transfer. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 4954\u20134964."},{"key":"e_1_3_1_129_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV56688.2023.00049"},{"key":"e_1_3_1_130_2","first-page":"11453","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Niemeyer Michael","year":"2021","unstructured":"Michael Niemeyer and Andreas Geiger. 2021. Giraffe: Representing scenes as compositional generative neural feature fields. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 11453\u201311464."},{"key":"e_1_3_1_131_2","first-page":"2998","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV)","author":"Oorloff Trevine","year":"2023","unstructured":"Trevine Oorloff and Yaser Yacoob. 2023. Expressive talking head video encoding in StyleGAN2 latent space. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV). 2998\u20133007."},{"key":"e_1_3_1_132_2","unstructured":"Dongwei Pan Long Zhuo Jingtan Piao Huiwen Luo Wei Cheng Yuxin Wang Siming Fan Shengqi Liu Lei Yang Bo Dai et\u00a0al. 2023. RenderMe-360: A large digital asset library and benchmarks towards high-fidelity head avatars. Advances in Neural Information Processing Systems 36 (2023) 7993\u20138005."},{"key":"e_1_3_1_133_2","first-page":"427","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Pang Youxin","year":"2023","unstructured":"Youxin Pang, Yong Zhang, Weize Quan, Yanbo Fan, Xiaodong Cun, Ying Shan, and Dong-ming Yan. 2023. Dpe: Disentanglement of pose and expression for general video portrait editing. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 427\u2013436."},{"key":"e_1_3_1_134_2","first-page":"18781","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Papantoniou Foivos Paraperas","year":"2022","unstructured":"Foivos Paraperas Papantoniou, Panagiotis P Filntisis, Petros Maragos, and Anastasios Roussos. 2022. Neural emotion director: Speech-preserving semantic control of facial expressions in\u201d in-the-wild\u201d videos. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 18781\u201318790."},{"key":"e_1_3_1_135_2","doi-asserted-by":"publisher","DOI":"10.1145\/800193.569955"},{"key":"e_1_3_1_136_2","doi-asserted-by":"crossref","unstructured":"Ziqiao Peng Wentao Hu Junyuan Ma Xiangyu Zhu Xiaomei Zhang Hao Zhao Hui Tian Jun He Hongyan Liu and Zhaoxin Fan. 2025. SyncTalk++: High-fidelity and efficient synchronized talking heads synthesis using gaussian splatting. arXiv:2506.14742. Retrieved from https:\/\/arxiv.org\/abs\/2506.14742","DOI":"10.1109\/TPAMI.2025.3630057"},{"key":"e_1_3_1_137_2","first-page":"666","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Peng Ziqiao","year":"2024","unstructured":"Ziqiao Peng, Wentao Hu, Yue Shi, Xiangyu Zhu, Xiaomei Zhang, Hao Zhao, Jun He, Hongyan Liu, and Zhaoxin Fan. 2024. Synctalk: The devil is in the synchronization for talking head synthesis. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 666\u2013676."},{"key":"e_1_3_1_138_2","unstructured":"Soujanya Poria Devamanyu Hazarika Navonil Majumder Gautam Naik Erik Cambria and Rada Mihalcea. 2018. Meld: A multimodal multi-party dataset for emotion recognition in conversations. arXiv:1810.02508. Retrieved from https:\/\/arxiv.org\/abs\/1810.02508"},{"key":"e_1_3_1_139_2","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1145\/3394171.3413532","volume-title":"Proceedings of the 28th ACM International Conference on Multimedia","author":"Prajwal KR","year":"2020","unstructured":"KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar. 2020. A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM International Conference on Multimedia. 484\u2013492."},{"key":"e_1_3_1_140_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00370"},{"key":"e_1_3_1_141_2","first-page":"8821","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Ramesh Aditya","year":"2021","unstructured":"Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In Proceedings of the International Conference on Machine Learning. Pmlr, 8821\u20138831."},{"key":"e_1_3_1_142_2","first-page":"1060","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Reed Scott","year":"2016","unstructured":"Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016. Generative adversarial text to image synthesis. In Proceedings of the International Conference on Machine Learning. PMLR, 1060\u20131069."},{"key":"e_1_3_1_143_2","first-page":"13759","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV)","author":"Ren Yurui","year":"2021","unstructured":"Yurui Ren, Ge Li, Yuanqi Chen, Thomas H Li, and Shan Liu. 2021. Pirenderer: Controllable portrait image generation via semantic neural rendering. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV). 13759\u201313768."},{"key":"e_1_3_1_144_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV48630.2021.00009"},{"key":"e_1_3_1_145_2","first-page":"1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV)","author":"Rossler Andreas","year":"2019","unstructured":"Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nie\u00dfner. 2019. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV). 1\u201311."},{"key":"e_1_3_1_146_2","unstructured":"Ekraam Sabir Jiaxin Cheng Ayush Jaiswal Wael AbdAlmageed Iacopo Masi and Prem Natarajan. 2019. Recurrent convolutional strategies for face manipulation detection in videos. Interfaces (GUI) 3 1 (2019) 80\u201387."},{"key":"e_1_3_1_147_2","doi-asserted-by":"publisher","DOI":"10.1145\/3528233.3530757"},{"key":"e_1_3_1_148_2","unstructured":"Tim Salimans Ian Goodfellow Wojciech Zaremba Vicki Cheung Alec Radford and Xi Chen. 2016. Improved techniques for training gans. Advances in Neural Information Processing Systems 29 (2016)."},{"key":"e_1_3_1_149_2","unstructured":"Katja Schwarz Yiyi Liao Michael Niemeyer and Andreas Geiger. 2020. Graf: Generative radiance fields for 3d-aware image synthesis. Advances in Neural Information Processing Systems 33 (2020) 20154\u201320166."},{"key":"e_1_3_1_150_2","doi-asserted-by":"crossref","unstructured":"Tong Sha Wei Zhang Tong Shen Zhoujun Li and Tao Mei. 2023. Deep person generation: A survey from the perspective of face pose and cloth synthesis. ACM Computing Surveys 55 12 (2023) 1\u201337.","DOI":"10.1145\/3575656"},{"key":"e_1_3_1_151_2","first-page":"666","volume-title":"Proceedings of the European Conference on Computer Vision(ECCV)","author":"Shen Shuai","year":"2022","unstructured":"Shuai Shen, Wanhua Li, Zheng Zhu, Yueqi Duan, Jie Zhou, and Jiwen Lu. 2022. Learning dynamic facial radiance fields for few-shot talking head synthesis. In Proceedings of the European Conference on Computer Vision(ECCV). Springer, 666\u2013682."},{"key":"e_1_3_1_152_2","first-page":"1982","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Shen Shuai","year":"2023","unstructured":"Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li, Zheng Zhu, Jie Zhou, and Jiwen Lu. 2023. Difftalk: Crafting diffusion models for generalized audio-driven portraits animation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 1982\u20131991."},{"key":"e_1_3_1_153_2","first-page":"2377","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Siarohin Aliaksandr","year":"2019","unstructured":"Aliaksandr Siarohin, St\u00e9phane Lathuili\u00e8re, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019. Animating arbitrary objects via deep motion transfer. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 2377\u20132386."},{"key":"e_1_3_1_154_2","unstructured":"Aliaksandr Siarohin St\u00e9phane Lathuili\u00e8re Sergey Tulyakov Elisa Ricci and Nicu Sebe. 2019. First order motion model for image animation. Advances in Neural Information Processing Systems 32 (2019)."},{"key":"e_1_3_1_155_2","first-page":"13653","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Siarohin Aliaksandr","year":"2021","unstructured":"Aliaksandr Siarohin, Oliver J Woodford, Jian Ren, Menglei Chai, and Sergey Tulyakov. 2021. Motion representations for articulated animation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 13653\u201313662."},{"key":"e_1_3_1_156_2","doi-asserted-by":"crossref","unstructured":"Linsen Song Wayne Wu Chen Qian Ran He and Chen Change Loy. 2022. Everybody\u2019s talkin\u2019: Let me talk as you want. IEEE Transactions on Information Forensics and Security 17 (2022) 585\u2013598.","DOI":"10.1109\/TIFS.2022.3146783"},{"key":"e_1_3_1_157_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00502"},{"key":"e_1_3_1_158_2","unstructured":"Kuiyuan Sun Xiaolong Liu Xiaolong Li Yao Zhao and Wei Wang. 2024. Multi-modal driven pose-controllable talking head generation. ACM Transactions on Multimedia Computing Communications and Applications 20 12 (2024) 1\u201323."},{"key":"e_1_3_1_159_2","doi-asserted-by":"crossref","unstructured":"Zhonglin Sun Siyang Song Ioannis Patras and Georgios Tzimiropoulos. 2024. Cemiface: Center-based semi-hard synthetic face generation for face recognition. Advances in Neural Information Processing Systems 37 (2024) 35612\u201335638.","DOI":"10.52202\/079017-1123"},{"key":"e_1_3_1_160_2","first-page":"5043","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"38","author":"Sun Zhaoxu","year":"2024","unstructured":"Zhaoxu Sun, Yuze Xuan, Fang Liu, and Yang Xiang. 2024. FG-EmoTalk: Talking head video generation with fine-grained controllable facial expressions. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 5043\u20135051."},{"key":"e_1_3_1_161_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00628"},{"key":"e_1_3_1_162_2","doi-asserted-by":"crossref","unstructured":"Supasorn Suwajanakorn Steven M Seitz and Ira Kemelmacher-Shlizerman. 2017. Synthesizing obama: Learning lip sync from audio. ACM Transactions on Graphics (ToG) 36 4 (2017) 1\u201313.","DOI":"10.1145\/3072959.3073640"},{"key":"e_1_3_1_163_2","first-page":"2818","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Szegedy Christian","year":"2016","unstructured":"Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 2818\u20132826."},{"key":"e_1_3_1_164_2","doi-asserted-by":"crossref","unstructured":"Shuai Tan Bin Ji Mengxiao Bi and Ye Pan. 2024. EDTalk: Efficient disentanglement for emotional talking head synthesis. arXiv:2404.01647. Retrieved from https:\/\/arxiv.org\/abs\/2404.01647","DOI":"10.1007\/978-3-031-72658-3_23"},{"key":"e_1_3_1_165_2","first-page":"22146","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV)","author":"Tan Shuai","year":"2023","unstructured":"Shuai Tan, Bin Ji, and Ye Pan. 2023. Emmn: Emotional motion memory network for audio-driven emotional talking face generation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV). 22146\u201322156."},{"key":"e_1_3_1_166_2","unstructured":"Junshu Tang Bo Zhang Binxin Yang Ting Zhang Dong Chen Lizhuang Ma and Fang Wen. 2022. 3DFaceShop: Explicitly controllable 3D-aware portrait generation. IEEE Transactions on Visualization and Computer Graphics 29 1 (2022) 1\u201318."},{"key":"e_1_3_1_167_2","first-page":"716","volume-title":"Proceedings of the European Conference on Computer Vision(ECCV)","author":"Thies Justus","year":"2020","unstructured":"Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nie\u00dfner. 2020. Neural voice puppetry: Audio-driven facial reenactment. In Proceedings of the European Conference on Computer Vision(ECCV). Springer, 716\u2013731."},{"key":"e_1_3_1_168_2","doi-asserted-by":"crossref","unstructured":"Justus Thies Michael Zollh\u00f6fer and Matthias Nie\u00dfner. 2019. Deferred neural rendering: Image synthesis using neural textures. Acm Transactions on Graphics (TOG) 38 4 (2019) 1\u201312.","DOI":"10.1145\/3306346.3323035"},{"key":"e_1_3_1_169_2","first-page":"2387","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Thies Justus","year":"2016","unstructured":"Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nie\u00dfner. 2016. Face2face: Real-time face capture and reenactment of rgb videos. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 2387\u20132395."},{"key":"e_1_3_1_170_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093474"},{"key":"e_1_3_1_171_2","first-page":"1329","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Tripathy Soumya","year":"2021","unstructured":"Soumya Tripathy, Juho Kannala, and Esa Rahtu. 2021. Facegan: Facial attribute controllable reenactment gan. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 1329\u20131338."},{"key":"e_1_3_1_172_2","first-page":"1526","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Tulyakov Sergey","year":"2018","unstructured":"Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz. 2018. Mocogan: Decomposing motion and content for video generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 1526\u20131535."},{"key":"e_1_3_1_173_2","doi-asserted-by":"crossref","unstructured":"Konstantinos Vougioukas Stavros Petridis and Maja Pantic. 2020. Realistic speech-driven facial animation with gans. International Journal of Computer Vision 128 5 (2020) 1398\u20131413.","DOI":"10.1007\/s11263-019-01251-8"},{"key":"e_1_3_1_174_2","unstructured":"Baiqin Wang Xiangyu Zhu Fan Shen Hao Xu and Zhen Lei. 2025. PC-Talk: Precise facial animation control for audio-driven talking face generation. arXiv:2503.14295. Retrieved from https:\/\/arxiv.org\/abs\/2503.14295"},{"key":"e_1_3_1_175_2","unstructured":"Cong Wang Kuan Tian Jun Zhang Yonghang Guan Feng Luo Fei Shen Zhiwei Jiang Qing Gu Xiao Han and Wei Yang. 2024. V-Express: Conditional dropout for progressive training of portrait video generation. arXiv:2406.02511. Retrieved from https:\/\/arxiv.org\/abs\/2406.02511"},{"key":"e_1_3_1_176_2","first-page":"17979","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Wang Duomin","year":"2023","unstructured":"Duomin Wang, Yu Deng, Zixin Yin, Heung-Yeung Shum, and Baoyuan Wang. 2023. Progressive disentangled representation learning for fine-grained controllable talking head synthesis. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 17979\u201317989."},{"key":"e_1_3_1_177_2","first-page":"26212","volume-title":"Proceedings of the Computer Vision and Pattern Recognition Conference","author":"Wang Haotian","year":"2025","unstructured":"Haotian Wang, Yuzhe Weng, Yueyan Li, Zilu Guo, Jun Du, Shutong Niu, Jiefeng Ma, Shan He, Xiaoyan Wu, Qiming Hu, et\u00a0al. 2025. Emotivetalk: Expressive talking head generation through audio information decoupling and emotional video diffusion. In Proceedings of the Computer Vision and Pattern Recognition Conference. 26212\u201326221."},{"key":"e_1_3_1_178_2","first-page":"13844","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Wang Jiayu","year":"2023","unstructured":"Jiayu Wang, Kang Zhao, Shiwei Zhang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. 2023. Lipformer: High-fidelity and generalizable talking face generation with a pre-learned facial codebook. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 13844\u201313853."},{"key":"e_1_3_1_179_2","first-page":"700","volume-title":"Proceedings of the European Conference on Computer Vision(ECCV)","author":"Wang Kaisiyuan","year":"2020","unstructured":"Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. 2020. Mead: A large-scale audio-visual dataset for emotional talking-face generation. In Proceedings of the European Conference on Computer Vision(ECCV). Springer, 700\u2013717."},{"key":"e_1_3_1_180_2","doi-asserted-by":"crossref","first-page":"679","DOI":"10.1109\/3DV53792.2021.00077","volume-title":"Proceedings of the 2021 International Conference on 3D Vision (3DV)","author":"Wang Qiulin","year":"2021","unstructured":"Qiulin Wang, Lu Zhang, and Bo Li. 2021. SAFA: Structure aware face animation. In Proceedings of the 2021 International Conference on 3D Vision (3DV). IEEE, 679\u2013688."},{"key":"e_1_3_1_181_2","doi-asserted-by":"crossref","unstructured":"Suzhen Wang Lincheng Li Yu Ding Changjie Fan and Xin Yu. 2021. Audio2head: Audio-driven one-shot talking-head generation with natural head motion. arXiv:2107.09293. Retrieved from https:\/\/arxiv.org\/abs\/2107.09293","DOI":"10.24963\/ijcai.2021\/152"},{"key":"e_1_3_1_182_2","first-page":"2531","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"36","author":"Wang Suzhen","year":"2022","unstructured":"Suzhen Wang, Lincheng Li, Yu Ding, and Xin Yu. 2022. One-shot talking face generation from single-speaker audio-visual correlation learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 2531\u20132539."},{"key":"e_1_3_1_183_2","unstructured":"Ting-Chun Wang Ming-Yu Liu Jun-Yan Zhu Guilin Liu Andrew Tao Jan Kautz and Bryan Catanzaro. 2018. Video-to-video synthesis. arXiv:1808.06601. Retrieved from https:\/\/arxiv.org\/abs\/1808.06601"},{"key":"e_1_3_1_184_2","first-page":"10039","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Wang Ting-Chun","year":"2021","unstructured":"Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu. 2021. One-shot free-view neural talking-head synthesis for video conferencing. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 10039\u201310049."},{"key":"e_1_3_1_185_2","first-page":"10039","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Wang Ting-Chun","year":"2021","unstructured":"Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu. 2021. One-shot free-view neural talking-head synthesis for video conferencing. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 10039\u201310049."},{"key":"e_1_3_1_186_2","first-page":"22851","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV)","author":"Wang Yuhan","year":"2023","unstructured":"Yuhan Wang, Liming Jiang, and Chen Change Loy. 2023. Styleinv: A temporal style modulated inversion network for unconditional video generation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV). 22851\u201322861."},{"key":"e_1_3_1_187_2","first-page":"8105","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"39","author":"Wang Yu","year":"2025","unstructured":"Yu Wang, Yunfei Liu, Fa-Ting Hong, Meng Cao, Lijian Lin, and Yu Li. 2025. AnyTalk: Multi-modal driven multi-domain talking head generation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 8105\u20138113."},{"key":"e_1_3_1_188_2","unstructured":"Yaohui Wang Di Yang Francois Bremond and Antitza Dantcheva. 2022. Latent image animator: Learning to animate images via latent space navigation. (2022). Retrieved from https:\/\/openreview.net\/forum?id=7r6kDq0mK_"},{"key":"e_1_3_1_189_2","doi-asserted-by":"crossref","unstructured":"Zhou Wang Alan C Bovik Hamid R Sheikh and Eero P Simoncelli. 2004. Image quality assessment: From error visibility to structural similarity. IEEE Transactions on Image Processing 13 4 (2004) 600\u2013612.","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_1_190_2","unstructured":"Zirui Wang Shangzhe Wu Weidi Xie Min Chen and Victor Adrian Prisacariu. 2021. NeRF\u2013: Neural radiance fields without known camera parameters. arXiv:2102.07064. Retrieved from https:\/\/arxiv.org\/abs\/2102.07064"},{"key":"e_1_3_1_191_2","doi-asserted-by":"crossref","unstructured":"Keith Waters. 1987. A muscle model for animation three-dimensional facial expression. ACM Siggraph Computer Graphics 21 4 (1987) 17\u201324.","DOI":"10.1145\/37402.37405"},{"key":"e_1_3_1_192_2","first-page":"670","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Wiles Olivia","year":"2018","unstructured":"Olivia Wiles, A Koepke, and Andrew Zisserman. 2018. X2face: A network for controlling face generation using images, audio, and pose codes. In Proceedings of the European Conference on Computer Vision (ECCV). 670\u2013686."},{"key":"e_1_3_1_193_2","first-page":"7623","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV)","author":"Wu Jay Zhangjie","year":"2023","unstructured":"Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. 2023. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV). 7623\u20137633."},{"key":"e_1_3_1_194_2","first-page":"603","volume-title":"Proceedings of the European Conference on Computer Vision(ECCV)","author":"Wu Wayne","year":"2018","unstructured":"Wayne Wu, Yunxuan Zhang, Cheng Li, Chen Qian, and Chen Change Loy. 2018. Reenactgan: Learning to reenact faces via boundary transfer. In Proceedings of the European Conference on Computer Vision(ECCV). 603\u2013619."},{"key":"e_1_3_1_195_2","doi-asserted-by":"publisher","unstructured":"Yiqian Wu Hao Xu Xiangjun Tang Xien Chen Siyu Tang Zhebin Zhang Chen Li and Xiaogang Jin. 2024. Portrait3D: Text-guided high-quality 3D portrait generation using pyramid representation and GANs prior. ACM Transactions on Graphics 43 4 Article 45 (July2024) 12 pages. DOI:10.1145\/3658162","DOI":"10.1145\/3658162"},{"key":"e_1_3_1_196_2","first-page":"8532","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"39","author":"Wu Zhenhua","year":"2025","unstructured":"Zhenhua Wu, Linxuan Jiang, Xiang Li, Chaowei Fang, Yipeng Qin, and Guanbin Li. 2025. Hierarchically controlled deformable 3D gaussians for talking head synthesis. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 8532\u20138540."},{"key":"e_1_3_1_197_2","unstructured":"Cheng-hsin Wuu Ningyuan Zheng Scott Ardisson Rohan Bali Danielle Belko Eric Brockmeyer Lucas Evans Timothy Godisart Hyowon Ha Xuhua Huang et\u00a0al. 2022. Multiface: A dataset for neural face rendering. arXiv:2207.11243. Retrieved from https:\/\/arxiv.org\/abs\/2207.11243"},{"key":"e_1_3_1_198_2","doi-asserted-by":"crossref","unstructured":"Lei Xie and Zhi-Qiang Liu. 2007. A coupled HMM approach to video-realistic speech animation. Pattern Recognition 40 8 (2007) 2325\u20132340.","DOI":"10.1016\/j.patcog.2006.12.001"},{"key":"e_1_3_1_199_2","first-page":"8753","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"39","author":"Xie Yifan","year":"2025","unstructured":"Yifan Xie, Tao Feng, Xin Zhang, Xiangyang Luo, Zixuan Guo, Weijiang Yu, Heng Chang, Fei Ma, and Fei Richard Yu. 2025. Pointtalk: Audio-driven dynamic lip point cloud for 3d gaussian-based talking head synthesis. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 8753\u20138761."},{"key":"e_1_3_1_200_2","doi-asserted-by":"crossref","first-page":"3170","DOI":"10.1145\/3664647.3681108","volume-title":"Proceedings of the 32nd ACM International Conference on Multimedia","author":"Xiong Lingyu","year":"2024","unstructured":"Lingyu Xiong, Xize Cheng, Jintao Tan, Xianjia Wu, Xiandong Li, Lei Zhu, Fei Ma, Minglei Li, Huang Xu, and Zhihui Hu. 2024. SegTalker: Segmentation-based talking face generation with mask-guided local editing. In Proceedings of the 32nd ACM International Conference on Multimedia. 3170\u20133179."},{"key":"e_1_3_1_201_2","unstructured":"Chao Xu Shaoting Zhu Junwei Zhu Tianxin Huang Jiangning Zhang Ying Tai and Yong Liu. 2023. Multimodal-driven talking face generation via a unified diffusion-based generator. 1\u20134. arXiv:2305.02594. Retrieved from https:\/\/arxiv.org\/abs\/2305.02594"},{"key":"e_1_3_1_202_2","first-page":"1316","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Xu Tao","year":"2018","unstructured":"Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. 2018. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 1316\u20131324."},{"key":"e_1_3_1_203_2","unstructured":"Yuelang Xu Benwang Chen Zhe Li Hongwen Zhang Lizhen Wang Zerong Zheng and Yebin Liu. 2023. Gaussian head avatar: Ultra high-fidelity head avatar via dynamic gaussians. arXiv:2312.03029. Retrieved from https:\/\/arxiv.org\/abs\/2312.03029"},{"key":"e_1_3_1_204_2","first-page":"15909","volume-title":"Proceedings of the Computer Vision and Pattern Recognition Conference","author":"Xu Zunnan","year":"2025","unstructured":"Zunnan Xu, Zhentao Yu, Zixiang Zhou, Jun Zhou, Xiaoyu Jin, Fa-Ting Hong, Xiaozhong Ji, Junwei Zhu, Chengfei Cai, Shiyu Tang, et\u00a0al. 2025. Hunyuanportrait: Implicit condition control for enhanced portrait animation. In Proceedings of the Computer Vision and Pattern Recognition Conference. 15909\u201315919."},{"key":"e_1_3_1_205_2","doi-asserted-by":"crossref","unstructured":"Eli Yamamoto Satoshi Nakamura and Kiyohiro Shikano. 1998. Lip movement synthesis from speech based on hidden markov models. Speech Communication 26 1-2 (1998) 105\u2013115.","DOI":"10.1016\/S0167-6393(98)00054-5"},{"key":"e_1_3_1_206_2","unstructured":"Shurong Yang Huadong Li Juhao Wu Minhao Jing Linze Li Renhe Ji Jiajun Liang and Haoqiang Fan. 2024. Megactor: Harness the power of raw video for vivid portrait animation. arXiv:2405.20851. Retrieved from https:\/\/arxiv.org\/abs\/2405.20851"},{"key":"e_1_3_1_207_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i7.28473"},{"key":"e_1_3_1_208_2","first-page":"1","volume-title":"Proceedings of the 2019 14th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2019)","author":"Yang Shuang","year":"2019","unstructured":"Shuang Yang, Yuanhang Zhang, Dalu Feng, Mingmin Yang, Chenhao Wang, Jingyun Xiao, Keyu Long, Shiguang Shan, and Xilin Chen. 2019. LRW-1000: A naturally-distributed large-scale benchmark for lip reading in the wild. In Proceedings of the 2019 14th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2019). IEEE, 1\u20138."},{"key":"e_1_3_1_209_2","first-page":"1773","volume-title":"Proceedings of the 28th ACM International Conference on Multimedia","author":"Yao Guangming","year":"2020","unstructured":"Guangming Yao, Yi Yuan, Tianjia Shao, and Kun Zhou. 2020. Mesh guided one-shot face reenactment using graph convolutional networks. In Proceedings of the 28th ACM International Conference on Multimedia. 1773\u20131781."},{"key":"e_1_3_1_210_2","unstructured":"Shunyu Yao RuiZhe Zhong Yichao Yan Guangtao Zhai and Xiaokang Yang. 2022. Dfa-nerf: Personalized talking head generation via disentangled face attributes neural rendering. arXiv:2201.00791. Retrieved from https:\/\/arxiv.org\/abs\/2201.00791"},{"key":"e_1_3_1_211_2","unstructured":"Zhenhui Ye Ziyue Jiang Yi Ren Jinglin Liu Jinzheng He and Zhou Zhao. 2022. GeneFace: Generalized and high-fidelity audio-driven 3D talking face synthesis. In Proceedings of International Conference on Learning Representation. Retrieved from https:\/\/openreview.net\/forum?id=YfwMIDhPccD"},{"key":"e_1_3_1_212_2","unstructured":"Ran Yi Zipeng Ye Juyong Zhang Hujun Bao and Yong-Jin Liu. 2020. Audio-driven talking face video generation with learning-based personalized head pose. arXiv:2002.10137. Retrieved from https:\/\/arxiv.org\/abs\/2002.10137"},{"key":"e_1_3_1_213_2","first-page":"85","volume-title":"Proceedings of the European Conference on Computer Vision(ECCV)","author":"Yin Fei","year":"2022","unstructured":"Fei Yin, Yong Zhang, Xiaodong Cun, Mingdeng Cao, Yanbo Fan, Xuan Wang, Qingyan Bai, Baoyuan Wu, Jue Wang, and Yujiu Yang. 2022. Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan. In Proceedings of the European Conference on Computer Vision(ECCV). Springer, 85\u2013101."},{"key":"e_1_3_1_214_2","unstructured":"Jiahui Yu Yuanzhong Xu Jing Yu Koh Thang Luong Gunjan Baid Zirui Wang Vijay Vasudevan Alexander Ku Yinfei Yang Burcu Karagol Ayan et\u00a0al. 2022. Scaling autoregressive models for content-rich text-to-image generation. Transactions on Machine Learning Research (2022). arXiv preprint arXiv:2206.10789. Retrieved from https:\/\/arxiv.org\/abs\/2206.10789"},{"key":"e_1_3_1_215_2","first-page":"7645","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV)","author":"Yu Zhentao","year":"2023","unstructured":"Zhentao Yu, Zixin Yin, Deyu Zhou, Duomin Wang, Finn Wong, and Baoyuan Wang. 2023. Talking head generation with probabilistic audio-to-visual diffusion priors. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV). 7645\u20137655."},{"key":"e_1_3_1_216_2","first-page":"524","volume-title":"Proceedings of the European Conference on Computer Vision(ECCV)","author":"Zakharov Egor","year":"2020","unstructured":"Egor Zakharov, Aleksei Ivakhnenko, Aliaksandra Shysheya, and Victor Lempitsky. 2020. Fast bi-layer neural synthesis of one-shot realistic head avatars. In Proceedings of the European Conference on Computer Vision(ECCV). Springer, 524\u2013540."},{"key":"e_1_3_1_217_2","first-page":"9459","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV)","author":"Zakharov Egor","year":"2019","unstructured":"Egor Zakharov, Aliaksandra Shysheya, Egor Burkov, and Victor Lempitsky. 2019. Few-shot adversarial learning of realistic neural talking head models. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV). 9459\u20139468."},{"key":"e_1_3_1_218_2","first-page":"22096","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Zhang Bowen","year":"2023","unstructured":"Bowen Zhang, Chenyang Qi, Pan Zhang, Bo Zhang, HsiangTao Wu, Dong Chen, Qifeng Chen, Yong Wang, and Fang Wen. 2023. Metaportrait: Identity-preserving talking head generation with fast personalized adaptation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 22096\u201322105."},{"key":"e_1_3_1_219_2","doi-asserted-by":"crossref","unstructured":"Chenxu Zhang Saifeng Ni Zhipeng Fan Hongbo Li Ming Zeng Madhukar Budagavi and Xiaohu Guo. 2021. 3d talking face with personalized pose dynamics. IEEE Transactions on Visualization and Computer Graphics 29 2 (2021) 1438\u20131449.","DOI":"10.1109\/TVCG.2021.3117484"},{"key":"e_1_3_1_220_2","unstructured":"Chenshuang Zhang Chaoning Zhang Mengchun Zhang and In So Kweon. 2023. Text-to-image diffusion model in generative ai: A survey. arXiv:2303.07909. Retrieved from https:\/\/arxiv.org\/abs\/2303.07909"},{"key":"e_1_3_1_221_2","first-page":"3867","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV)","author":"Zhang Chenxu","year":"2021","unstructured":"Chenxu Zhang, Yifan Zhao, Yifei Huang, Ming Zeng, Saifeng Ni, Madhukar Budagavi, and Xiaohu Guo. 2021. Facial: Synthesizing dynamic talking face with implicit attribute learning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision(ICCV). 3867\u20133876."},{"key":"e_1_3_1_222_2","first-page":"586","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Zhang Richard","year":"2018","unstructured":"Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 586\u2013595."},{"key":"e_1_3_1_223_2","first-page":"8652","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Zhang Wenxuan","year":"2023","unstructured":"Wenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang, Xi Shen, Yu Guo, Ying Shan, and Fei Wang. 2023. Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 8652\u20138661."},{"key":"e_1_3_1_224_2","first-page":"2747","volume-title":"Proceedings of the Interspeech","volume":"2743","author":"Zhang Xinjian","year":"2013","unstructured":"Xinjian Zhang, Lijuan Wang, Gang Li, Frank Seide, and Frank K Soong. 2013. A new language independent, photo-realistic talking head driven by voice only. In Proceedings of the Interspeech, Vol. 2743. 2747."},{"key":"e_1_3_1_225_2","first-page":"12335","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Zhang Xi","year":"2020","unstructured":"Xi Zhang, Xiaolin Wu, Xinliang Zhai, Xianye Ben, and Chengjie Tu. 2020. Davd-net: Deep audio-aided video decompression of talking heads. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 12335\u201312344."},{"key":"e_1_3_1_226_2","first-page":"3543","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"37","author":"Zhang Zhimeng","year":"2023","unstructured":"Zhimeng Zhang, Zhipeng Hu, Wenjin Deng, Changjie Fan, Tangjie Lv, and Yu Ding. 2023. Dinet: Deformation inpainting network for realistic face visually dubbing on high resolution video. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 3543\u20133551."},{"key":"e_1_3_1_227_2","first-page":"3661","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Zhang Zhimeng","year":"2021","unstructured":"Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan. 2021. Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 3661\u20133670."},{"key":"e_1_3_1_228_2","first-page":"3657","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Zhao Jian","year":"2022","unstructured":"Jian Zhao and Hui Zhang. 2022. Thin-plate spline motion model for image animation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 3657\u20133666."},{"key":"e_1_3_1_229_2","unstructured":"Zhongyuan Zhao Zhenyu Bao Qing Li Guoping Qiu and Kanglin Liu. 2024. PSAvatar: A point-based morphable shape model for real-time head avatar creation with 3D gaussian splatting. arXiv:2401.12900. Retrieved from https:\/\/arxiv.org\/abs\/2401.12900"},{"key":"e_1_3_1_230_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33019299"},{"key":"e_1_3_1_231_2","first-page":"4176","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR)","author":"Zhou Hang","year":"2021","unstructured":"Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu. 2021. Pose-controllable talking face generation by implicitly modularized audio-visual representation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR). 4176\u20134186."},{"key":"e_1_3_1_232_2","doi-asserted-by":"crossref","unstructured":"Yang Zhou Xintong Han Eli Shechtman Jose Echevarria Evangelos Kalogerakis and Dingzeyu Li. 2020. Makelttalk: Speaker-aware talking-head animation. ACM Transactions On Graphics (TOG) 39 6 (2020) 1\u201315.","DOI":"10.1145\/3414685.3417774"},{"key":"e_1_3_1_233_2","doi-asserted-by":"crossref","unstructured":"Zhenglin Zhou Fan Ma Hehe Fan and Yi Yang. 2024. HeadStudio: Text to animatable head avatars with 3D gaussian splatting. arXiv:2402.06149. Retrieved from https:\/\/arxiv.org\/abs\/2402.06149","DOI":"10.1007\/978-3-031-73411-3_9"},{"key":"e_1_3_1_234_2","first-page":"650","volume-title":"Proceedings of the European Conference on Computer Vision(ECCV)","author":"Zhu Hao","year":"2022","unstructured":"Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy. 2022. CelebV-HQ: A large-scale video facial attributes dataset. In Proceedings of the European Conference on Computer Vision(ECCV). Springer, 650\u2013667."},{"key":"e_1_3_1_235_2","doi-asserted-by":"crossref","unstructured":"Hao Zhu Haotian Yang Longwei Guo Yidi Zhang Yanru Wang Mingkai Huang Menghua Wu Qiu Shen Ruigang Yang and Xun Cao. 2023. Facescape: 3d facial dataset and benchmark for single-view 3d face reconstruction. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 12 (2023) 14528\u201314545.","DOI":"10.1109\/TPAMI.2023.3307338"},{"key":"e_1_3_1_236_2","doi-asserted-by":"crossref","unstructured":"Kaifeng Zou Sylvain Faisan Boyang Yu S\u00e9bastien Valette and Hyewon Seo. 2024. 4d facial expression diffusion model. ACM Transactions on Multimedia Computing Communications and Applications 21 1 (2024) 1\u201323.","DOI":"10.1145\/3653455"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3785656","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,3]],"date-time":"2026-02-03T14:20:00Z","timestamp":1770128400000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3785656"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,3]]},"references-count":235,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3785656"],"URL":"https:\/\/doi.org\/10.1145\/3785656","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,3]]},"assertion":[{"value":"2024-06-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-08","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}