{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,21]],"date-time":"2026-02-21T09:16:25Z","timestamp":1771665385182,"version":"3.50.1"},"publisher-location":"Cham","reference-count":120,"publisher":"Springer International Publishing","isbn-type":[{"value":"9783030876630","type":"print"},{"value":"9783030876647","type":"electronic"}],"license":[{"start":{"date-parts":[[2022,1,1]],"date-time":"2022-01-01T00:00:00Z","timestamp":1640995200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,1,31]],"date-time":"2022-01-31T00:00:00Z","timestamp":1643587200000},"content-version":"vor","delay-in-days":30,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Talking face\u00a0generation aims at synthesizing coherent and realistic face sequences given an input speech. The task enjoys a wide spectrum of downstream applications, such as teleconferencing, movie dubbing, and virtual assistant. The emergence of deep learning\u00a0and cross-modality research has led to many interesting works that address talking face\u00a0generation. Despite great research efforts in talking face generation, the problem remains challenging due to the need for fine-grained control of face components and the generalization to arbitrary sentences. In this chapter, we first discuss the definition and underlying challenges of the problem. Then, we present an overview of recent progress in talking face\u00a0generation. In addition, we introduce some widely used datasets and performance metrics. Finally, we discuss open questions, potential future directions, and ethical considerations in this task.<\/jats:p>","DOI":"10.1007\/978-3-030-87664-7_8","type":"book-chapter","created":{"date-parts":[[2022,1,31]],"date-time":"2022-01-31T09:03:06Z","timestamp":1643619786000},"page":"163-188","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Talking Faces: Audio-to-Video Face Generation"],"prefix":"10.1007","author":[{"given":"Yuxin","family":"Wang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Linsen","family":"Song","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wayne","family":"Wu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chen","family":"Qian","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ran","family":"He","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chen Change","family":"Loy","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,1,31]]},"reference":[{"key":"8_CR1","unstructured":"Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y (2014) Generative adversarial nets. In: Proceedings of the advances in neural information processing systems, vol 27"},{"key":"8_CR2","unstructured":"Radford A, Metz L, Chintala S (2016) Unsupervised representation learning with deep convolutional generative adversarial networks. In: Proceedings of the international conference on learning representations"},{"key":"8_CR3","unstructured":"Mirza M, Osindero S (2014) Conditional generative adversarial nets. CoRR arXiv:abs\/1411.1784"},{"key":"8_CR4","doi-asserted-by":"crossref","unstructured":"Chen L, Li Z, Maddox RK, Duan Z, Xu C (2018) Lip movements generation at a glance. In: Proceedings of the European conference on computer vision, pp 520\u2013535","DOI":"10.1007\/978-3-030-01234-2_32"},{"key":"8_CR5","doi-asserted-by":"crossref","unstructured":"Zhou H, Liu Y, Liu Z, Luo P, Wang X (2019) Talking face generation by adversarially disentangled audio-visual representation. In: Proceedings of the AAAI conference on artificial intelligence, vol 33, no 1, pp 9299\u20139306","DOI":"10.1609\/aaai.v33i01.33019299"},{"key":"8_CR6","doi-asserted-by":"crossref","unstructured":"Chen L, Maddox RK, Duan Z, Xu C (2019) Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 7832\u20137841","DOI":"10.1109\/CVPR.2019.00802"},{"key":"8_CR7","doi-asserted-by":"crossref","unstructured":"Song Y, Zhu J, Li D, Wang A, Qi H (2019) Talking face generation by conditional recurrent adversarial network. In: Kraus S (ed) Proceedings of the international joint conference on artificial intelligence, pp 919\u2013925","DOI":"10.24963\/ijcai.2019\/129"},{"key":"8_CR8","doi-asserted-by":"crossref","unstructured":"Zhu H, Huang H, Li Y, Zheng A, He R (2020) Arbitrary talking face generation via attentional audio-visual coherence learning. In: Proceedings of the international joint conference on artificial intelligence, pp 2362\u20132368","DOI":"10.24963\/ijcai.2020\/327"},{"key":"8_CR9","doi-asserted-by":"crossref","unstructured":"Pham HX, Cheung S, Pavlovic V (2017) Speech-driven 3d facial animation with implicit emotional awareness: a deep learning approach. In: Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pp 80\u201388","DOI":"10.1109\/CVPRW.2017.287"},{"issue":"4","key":"8_CR10","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3072959.3073658","volume":"36","author":"T Karras","year":"2017","unstructured":"Karras T, Aila T, Laine S, Herva A, Lehtinen J (2017) Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Trans Graph 36(4):1\u201312","journal-title":"ACM Trans Graph"},{"issue":"4","key":"8_CR11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3072959.3073699","volume":"36","author":"S Taylor","year":"2017","unstructured":"Taylor S, Kim T, Yue Y, Mahler M, Krahe J, Rodriguez AG, Hodgins J, Matthews I (2017) A deep learning approach for generalized speech animation. ACM Trans Graph 36(4):1\u201311","journal-title":"ACM Trans Graph"},{"issue":"4","key":"8_CR12","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3306346.3323028","volume":"38","author":"O Fried","year":"2019","unstructured":"Fried O, Tewari A, Zollh\u00f6fer M, Finkelstein A, Shechtman E, Goldman DB, Genova K, Jin Z, Theobalt C, Agrawala M (2019) Text-based editing of talking-head video. ACM Trans Graph 38(4):1\u201314","journal-title":"ACM Trans Graph"},{"issue":"4","key":"8_CR13","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2897824.2925984","volume":"35","author":"P Edwards","year":"2016","unstructured":"Edwards P, Landreth C, Fiume E, Singh K (2016) Jali: an animator-centric viseme model for expressive lip synchronization. ACM Trans Graph 35(4):1\u201311","journal-title":"ACM Trans Graph"},{"issue":"4","key":"8_CR14","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3197517.3201292","volume":"37","author":"Y Zhou","year":"2018","unstructured":"Zhou Y, Xu Z, Landreth C, Kalogerakis E, Maji S, Singh K (2018) Visemenet: audio-driven animator-centric speech animation. ACM Trans Graph 37(4):1\u201310","journal-title":"ACM Trans Graph"},{"issue":"4","key":"8_CR15","doi-asserted-by":"publisher","first-page":"118","DOI":"10.1002\/vis.4340020404","volume":"2","author":"J Lewis","year":"1991","unstructured":"Lewis J (1991) Automated lip-sync: background and techniques. J Visualization Comput Animat 2(4):118\u2013122","journal-title":"J Visualization Comput Animat"},{"key":"8_CR16","doi-asserted-by":"crossref","unstructured":"Guiard-Marigny T, Tsingos N, Adjoudani A, Benoit C, Gascuel M-P (1996) 3d models of the lips for realistic speech animation. In: Proceedings of the computer animation, pp 80\u201389","DOI":"10.1109\/CA.1996.540490"},{"key":"8_CR17","doi-asserted-by":"crossref","unstructured":"Bregler C, Covell M, Slaney M (1997) Video rewrite: driving visual speech with audio. In: Proceedings of the annual conference on computer graphics and interactive techniques, pp 353\u2013360","DOI":"10.1145\/258734.258880"},{"key":"8_CR18","doi-asserted-by":"crossref","unstructured":"Brand M (1999) Voice puppetry. In: Proceedings of the annual conference on computer graphics and interactive techniques, pp 21\u201328","DOI":"10.1145\/311535.311537"},{"issue":"8","key":"8_CR19","doi-asserted-by":"publisher","first-page":"2325","DOI":"10.1016\/j.patcog.2006.12.001","volume":"40","author":"L Xie","year":"2007","unstructured":"Xie L, Liu Z-Q (2007) A coupled HMM approach to video-realistic speech animation. Pattern Recogn 40(8):2325\u20132340","journal-title":"Pattern Recogn"},{"issue":"2","key":"8_CR20","doi-asserted-by":"publisher","first-page":"193","DOI":"10.1111\/cgf.12552","volume":"34","author":"P Garrido","year":"2015","unstructured":"Garrido P, Valgaerts L, Sarmadi H, Steiner I, Varanasi K, Perez P, Theobalt C (2015) Vdub: modifying face video of actors for plausible visual alignment to a dubbed audio track. Comput Graph Forum 34(2):193\u2013204","journal-title":"Comput Graph Forum"},{"key":"8_CR21","doi-asserted-by":"crossref","unstructured":"Charles J, Magee D, Hogg D (2016) Virtual immortality: reanimating characters from TV shows. In: Proceedings of the European conference on computer vision, pp 879\u2013886","DOI":"10.1007\/978-3-319-49409-8_71"},{"issue":"6","key":"8_CR22","first-page":"1","volume":"39","author":"Y Zhou","year":"2020","unstructured":"Zhou Y, Han X, Shechtman E, Echevarria J, Kalogerakis E, Li D (2020) Makelttalk: speaker-aware talking-head animation. ACM Trans Graph 39(6):1\u201315","journal-title":"ACM Trans Graph"},{"issue":"3","key":"8_CR23","doi-asserted-by":"publisher","first-page":"152","DOI":"10.1109\/6046.865480","volume":"2","author":"E Cosatto","year":"2000","unstructured":"Cosatto E, Graf HP (2000) Photo-realistic talking-heads from image samples. IEEE Trans Multimedia 2(3):152\u2013163","journal-title":"IEEE Trans Multimedia"},{"key":"8_CR24","doi-asserted-by":"crossref","unstructured":"Fan B, Wang L, Soong FK, Xie L (2015) Photo-real talking head with deep bidirectional LSTM. In: Proceedings of the IEEE international conference on acoustics, speech and signal processing, pp 4884\u20134888","DOI":"10.1109\/ICASSP.2015.7178899"},{"key":"8_CR25","doi-asserted-by":"crossref","unstructured":"Vougioukas K, Petridis S, Pantic M (2019) Realistic speech-driven facial animation with gans. Int J Comput Vis 1\u201316","DOI":"10.1007\/s11263-019-01251-8"},{"key":"8_CR26","doi-asserted-by":"crossref","unstructured":"Prajwal KR, Mukhopadhyay R, Namboodiri VP, Jawahar CV (2020) A lip sync expert is all you need for speech to lip generation in the wild. In: Proceedings of the ACM international conference on multimedia, pp 484\u2013492","DOI":"10.1145\/3394171.3413532"},{"key":"8_CR27","doi-asserted-by":"crossref","unstructured":"Das D, Biswas S, Sinha S, Bhowmick B (2020) Speech-driven facial animation using cascaded gans for learning of motion and texture. In: Proceedings of the European conference on computer vision, pp 408\u2013424","DOI":"10.1007\/978-3-030-58577-8_25"},{"key":"8_CR28","doi-asserted-by":"crossref","unstructured":"Yao X, Fried O, Fatahalian K, Agrawala M (2020) Iterative text-based editing of talking-heads using neural retargeting. arXiv preprint arXiv:2011.10688","DOI":"10.1145\/3449063"},{"key":"8_CR29","doi-asserted-by":"crossref","unstructured":"Wu W, Zhang Y, Li C, Qian C, Loy CC (2018) ReenactGAN learning to reenact faces via boundary transfer. In: Proceedings of the European conference on computer vision, pp 603\u2013619","DOI":"10.1007\/978-3-030-01246-5_37"},{"key":"8_CR30","doi-asserted-by":"crossref","unstructured":"Song L, Wu W, Fu C, Qian C, Loy CC, He R (2021) Everything\u2019s talkin\u2019: Pareidolia face reenactment. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR46437.2021.00227"},{"issue":"4","key":"8_CR31","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3072959.3073640","volume":"36","author":"S Suwajanakorn","year":"2017","unstructured":"Suwajanakorn S, Seitz SM, Kemelmacher-Shlizerman I (2017) Synthesizing Obama: learning lip sync from audio. ACM Trans Graph 36(4):1\u201313","journal-title":"ACM Trans Graph"},{"issue":"2","key":"8_CR32","first-page":"5","volume":"3","author":"E Friesen","year":"1978","unstructured":"Friesen E, Ekman P (1978) Facial action coding system: a technique for the measurement of facial movement. Palo Alto 3(2):5","journal-title":"Palo Alto"},{"key":"8_CR33","unstructured":"Song L, Wu W, Qian C, He R, Loy CC (2020) Everybody\u2019s talkin\u2019: let me talk as you want. arXiv arXiv:abs\/2001.05201"},{"issue":"2","key":"8_CR34","doi-asserted-by":"publisher","first-page":"98","DOI":"10.1109\/MRA.2012.2192811","volume":"19","author":"M Mori","year":"2012","unstructured":"Mori M, MacDorman KF, Kageki N (2012) The uncanny valley [from the field]. IEEE Robot Autom Mag 19(2):98\u2013100","journal-title":"IEEE Robot Autom Mag"},{"key":"8_CR35","unstructured":"Hannun AY, Case C, Casper J, Catanzaro B, Diamos G, Elsen E, Prenger R, Satheesh S, Sengupta S, Coates A, Ng AY (2014) Deep speech: scaling up end-to-end speech recognition. CoRR abs\/1412.5567"},{"key":"8_CR36","unstructured":"Amodei D, Ananthanarayanan S, Anubhai R, Bai J, Battenberg E, Case C, Casper J, Catanzaro B, Cheng Q, Chen G et\u00a0al (2016) Deep speech 2: end-to-end speech recognition in English and mandarin. In: Proceedings of the international conference on machine learning, pp 173\u2013182"},{"key":"8_CR37","doi-asserted-by":"crossref","unstructured":"Chen D, Ren S, Wei Y, Cao X, Sun J (2014) Joint cascade face detection and alignment. In: Proceedings of the European conference on computer vision, pp 109\u2013122","DOI":"10.1007\/978-3-319-10599-4_8"},{"key":"8_CR38","doi-asserted-by":"crossref","unstructured":"Chen D, Cao X, Wen F, Sun J (2013) Blessing of dimensionality: high-dimensional feature and its efficient compression for face verification. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 3025\u20133032","DOI":"10.1109\/CVPR.2013.389"},{"key":"8_CR39","doi-asserted-by":"crossref","unstructured":"Seibold C, Samek W, Hilsmann A, Eisert P (2017) Detection of face morphing attacks by deep learning. In: Proceedings of the international workshop on digital watermarking, pp 107\u2013120","DOI":"10.1007\/978-3-319-64185-0_9"},{"key":"8_CR40","doi-asserted-by":"crossref","unstructured":"Lewenberg Y, Bachrach Y, Shankar S, Criminisi A (2016) Predicting personal traits from facial images using convolutional neural networks augmented with facial landmark information. In: Proceedings of the AAAI conference on artificial intelligence, vol\u00a030, no\u00a01","DOI":"10.1609\/aaai.v30i1.9844"},{"key":"8_CR41","doi-asserted-by":"crossref","unstructured":"Di X, Sindagi VA, Patel VM (2018) GP-GAN: gender preserving GAN for synthesizing faces from landmarks. In: Proceedings of the international conference on pattern recognition, pp 1079\u20131084","DOI":"10.1109\/ICPR.2018.8545081"},{"key":"8_CR42","doi-asserted-by":"crossref","unstructured":"Garrido P, Valgaerts L, Rehmsen O, Thormahlen T, Perez P, Theobalt C (2014) Automatic face reenactment. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 4217\u20134224","DOI":"10.1109\/CVPR.2014.537"},{"key":"8_CR43","doi-asserted-by":"crossref","unstructured":"Pumarola A, Agudo A, Martinez AM, Sanfeliu A, Moreno-Noguer F (2018) Ganimation: anatomically-aware facial animation from a single image. In: Proceedings of the European conference on computer vision, pp 818\u2013833","DOI":"10.1007\/978-3-030-01249-6_50"},{"key":"8_CR44","doi-asserted-by":"crossref","unstructured":"Blanz V, Vetter T (1999) A morphable model for the synthesis of 3d faces. In: Proceedings of the annual conference on computer graphics and interactive techniques, pp 187\u2013194","DOI":"10.1145\/311535.311556"},{"key":"8_CR45","doi-asserted-by":"crossref","unstructured":"Paysan P, Knothe R, Amberg B, Romdhani S, Vetter T (2009) A 3d face model for pose and illumination invariant face recognition. In: Proceedings of the IEEE international conference on advanced video and signal-based surveillance, pp 296\u2013301","DOI":"10.1109\/AVSS.2009.58"},{"key":"8_CR46","unstructured":"Besl PJ, McKay ND (1992) Method for registration of 3-d shapes. In: Sensor fusion IV: control paradigms and data structures, vol 1611. International Society for Optics and Photonics, pp 586\u2013606"},{"issue":"4","key":"8_CR47","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/1778765.1778839","volume":"29","author":"E Kalogerakis","year":"2010","unstructured":"Kalogerakis E, Hertzmann A, Singh K (2010) Learning 3d mesh segmentation and labeling. ACM Trans Graph 29(4):1\u201312","journal-title":"ACM Trans Graph"},{"issue":"6","key":"8_CR48","first-page":"1","volume":"36","author":"T Li","year":"2017","unstructured":"Li T, Bolkart T, Black MJ, Li H, Romero J (2017) Learning a model of facial shape and expression from 4d scans. ACM Trans Graph 36(6):1\u201317","journal-title":"ACM Trans Graph"},{"issue":"1","key":"8_CR49","doi-asserted-by":"publisher","first-page":"78","DOI":"10.1109\/TPAMI.2017.2778152","volume":"41","author":"X Zhu","year":"2017","unstructured":"Zhu X, Liu X, Lei Z, Li SZ (2017) Face alignment in full pose range: a 3d total solution. IEEE Trans Pattern Anal Mach Intell 41(1):78\u201392","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"8_CR50","first-page":"152","volume":"12364","author":"J Guo","year":"2020","unstructured":"Guo J, Zhu X, Yang Y, Yang F, Lei Z, Li SZ (2020) Towards fast, accurate and stable 3d dense face alignment. Proceedings of the European conference on computer vision 12364:152\u2013168","journal-title":"Proceedings of the European conference on computer vision"},{"issue":"3","key":"8_CR51","first-page":"413","volume":"20","author":"C Cao","year":"2013","unstructured":"Cao C, Weng Y, Zhou S, Tong Y, Zhou K (2013) Facewarehouse: a 3d facial expression database for visual computing. IEEE Trans Visualization Comput Graph 20(3):413\u2013425","journal-title":"IEEE Trans Visualization Comput Graph"},{"key":"8_CR52","doi-asserted-by":"crossref","unstructured":"Bolkart T, Wuhrer S (2015) A groupwise multilinear correspondence optimization for 3d faces. In: Proceedings of the IEEE international conference on computer vision, pp 3604\u20133612","DOI":"10.1109\/ICCV.2015.411"},{"key":"8_CR53","doi-asserted-by":"crossref","unstructured":"Blanz V, Romdhani S, Vetter T (2002) Face identification across different poses and illuminations with a 3d morphable model. In: Proceedings of the IEEE international conference on automatic face and gesture recognition, pp 202\u2013207","DOI":"10.1109\/AFGR.2002.1004155"},{"key":"8_CR54","doi-asserted-by":"crossref","unstructured":"Gecer B, Ploumpis S, Kotsia I, Zafeiriou S (2019) Ganfit: generative adversarial network fitting for high fidelity 3d face reconstruction. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 1155\u20131164","DOI":"10.1109\/CVPR.2019.00125"},{"key":"8_CR55","doi-asserted-by":"crossref","unstructured":"Zhou H, Liu J, Liu Z, Liu Y, Wang X (2020) Rotate-and-render: unsupervised photorealistic face rotation from single-view images. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 5911\u20135920","DOI":"10.1109\/CVPR42600.2020.00595"},{"issue":"4","key":"8_CR56","first-page":"1","volume":"37","author":"H Kim","year":"2018","unstructured":"Kim H, Garrido P, Tewari A, Xu W, Thies J, Niessner M, P\u00e9rez P, Richardt C, Zollh\u00f6fer M, Theobalt C (2018) Deep video portraits. ACM Trans Graph 37(4):1\u201314","journal-title":"Deep video portraits. ACM Trans Graph"},{"key":"8_CR57","doi-asserted-by":"crossref","unstructured":"Thies J, Elgharib M, Tewari A, Theobalt C, Nie\u00dfner M (2020) Neural voice puppetry: audio-driven facial reenactment. In: Proceedings of the European conference on computer vision, pp 716\u2013731","DOI":"10.1007\/978-3-030-58517-4_42"},{"key":"8_CR58","doi-asserted-by":"crossref","unstructured":"Thies J, Zollhofer M, Stamminger M, Theobalt C, Nie\u00dfner M (2016) Face2face: real-time face capture and reenactment of RGB videos. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 2387\u20132395","DOI":"10.1109\/CVPR.2016.262"},{"key":"8_CR59","doi-asserted-by":"crossref","unstructured":"Booth J, Roussos A, Zafeiriou S, Ponniah A, Dunaway D (2016) A 3d morphable model learnt from 10,000 faces. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 5543\u20135552","DOI":"10.1109\/CVPR.2016.598"},{"key":"8_CR60","unstructured":"Rubin S, Berthouzoz F, Mysore GJ, Li W, Agrawala M, Content-based tools for editing audio stories. In: Proceedings of the ACM symposium on user interface software and technology, pp 113\u2013122"},{"issue":"3","key":"8_CR61","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2890493","volume":"35","author":"P Garrido","year":"2016","unstructured":"Garrido P, Zollh\u00f6fer M, Casas D, Valgaerts L, Varanasi K, P\u00e9rez P, Theobalt C (2016) Reconstruction of personalized 3d face rigs from monocular video. ACM Transactions on Graphics 35(3):1\u201315","journal-title":"ACM Transactions on Graphics"},{"issue":"11","key":"8_CR62","doi-asserted-by":"publisher","first-page":"1767","DOI":"10.1007\/s11263-019-01150-y","volume":"127","author":"A Jamaludin","year":"2019","unstructured":"Jamaludin A, Chung JS, Zisserman A (2019) You said that?: synthesising talking faces from audio. Int J Comput Vision 127(11):1767\u20131779","journal-title":"Int J Comput Vision"},{"key":"8_CR63","doi-asserted-by":"crossref","unstructured":"Wiles O, Koepke A, Zisserman A (2018) X2face: a network for controlling face generation using images, audio, and pose codes. In: Proceedings of the European conference on computer vision, pp 670\u2013686","DOI":"10.1007\/978-3-030-01261-8_41"},{"key":"8_CR64","doi-asserted-by":"crossref","unstructured":"Guo Y, Chen K, Liang S, Liu Y, Bao H, Zhang J (2021) Ad-nerf: audio driven neural radiance fields for talking head synthesis. arXiv preprint arXiv:2103.11078","DOI":"10.1109\/ICCV48922.2021.00573"},{"key":"8_CR65","doi-asserted-by":"crossref","unstructured":"Chung JS, Zisserman A (2016) Out of time: automated lip sync in the wild. In: Proceedings of the Asian conference on computer vision, pp 251\u2013263","DOI":"10.1007\/978-3-319-54427-4_19"},{"key":"8_CR66","unstructured":"Prajwal KR, Mukhopadhyay R, Philip J, Jha A, Namboodiri V, Jawahar CV (2019) Towards automatic face-to-face translation. In: Proceedings of the ACM international conference on multimedia, pp 1428\u20131436"},{"key":"8_CR67","doi-asserted-by":"crossref","unstructured":"Agarwal S, Farid H, Fried O, Agrawala M (2020) Detecting deep-fake videos from phoneme-viseme mismatches. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition workshops, pp 660\u2013661","DOI":"10.1109\/CVPRW50498.2020.00338"},{"issue":"3","key":"8_CR68","doi-asserted-by":"publisher","first-page":"388","DOI":"10.1145\/566654.566594","volume":"21","author":"T Ezzat","year":"2002","unstructured":"Ezzat T, Geiger G, Poggio T (2002) Trainable videorealistic speech animation. ACM Trans Graph 21(3):388\u2013398","journal-title":"ACM Trans Graph"},{"key":"8_CR69","doi-asserted-by":"crossref","unstructured":"Chang Y-J, Ezzat T (2005) Transferable videorealistic speech animation. In: Proceedings of the ACM SIGGRAPH\/Eurographics symposium on computer animation, pp 143\u2013151","DOI":"10.1145\/1073368.1073388"},{"issue":"1","key":"8_CR70","doi-asserted-by":"publisher","first-page":"9","DOI":"10.1109\/79.911195","volume":"18","author":"T Chen","year":"2001","unstructured":"Chen T (2001) Audiovisual speech processing. IEEE Signal Process Mag 18(1):9\u201321","journal-title":"IEEE Signal Process Mag"},{"issue":"1","key":"8_CR71","doi-asserted-by":"publisher","first-page":"51","DOI":"10.1023\/A:1011171430700","volume":"29","author":"K Choi","year":"2001","unstructured":"Choi K, Luo Y, Hwang J-N (2001) Hidden markov model inversion for audio-to-visual conversion in an mpeg-4 facial animation system. J VLSI Signal Process Syst Signal Image Video Technol 29(1):51\u201361","journal-title":"J VLSI Signal Process Syst Signal Image Video Technol"},{"key":"8_CR72","doi-asserted-by":"crossref","unstructured":"Wang L, Qian X, Han W, Soong FK (2010) Synthesizing photo-real talking head via trajectory-guided sample selection. In: Proceedings of the annual conference of the international speech communication association0","DOI":"10.21437\/Interspeech.2010-194"},{"key":"8_CR73","unstructured":"Karras T, Aila T, Laine S, Lehtinen J (2018) Progressive growing of gans for improved quality, stability, and variation. In: Proceedings of the international conference on learning representations"},{"key":"8_CR74","doi-asserted-by":"crossref","unstructured":"Karras T, Laine S, Aila T (2019) A style-based generator architecture for generative adversarial networks. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 4401\u20134410","DOI":"10.1109\/CVPR.2019.00453"},{"key":"8_CR75","doi-asserted-by":"crossref","unstructured":"Karras T, Laine S, Aittala M, Hellsten J, Lehtinen J, Aila T (2020) Analyzing and improving the image quality of stylegan. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 8110\u20138119","DOI":"10.1109\/CVPR42600.2020.00813"},{"issue":"11","key":"8_CR76","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","volume":"86","author":"Y LeCun","year":"1998","unstructured":"LeCun Y, Bottou L, Bengio Y, Haffner P (1998) Gradient-based learning applied to document recognition. Proc IEEE 86(11):2278\u20132324","journal-title":"Proc IEEE"},{"key":"8_CR77","doi-asserted-by":"crossref","unstructured":"Devries T, Biswaranjan K, Taylor GW (2014) Multi-task learning of facial landmarks and expression. In: Proceedings of the Canadian conference on computer and robot vision, pp 98\u2013103","DOI":"10.1109\/CRV.2014.21"},{"key":"8_CR78","unstructured":"Krizhevsky A et\u00a0al (2009) Learning multiple layers of features from tiny images. Master\u2019s thesis, University of Tront"},{"key":"8_CR79","doi-asserted-by":"crossref","unstructured":"Lu Y, Tai Y-W, Tang C-K (2018) Attribute-guided face generation using conditional cyclegan. In: Proceedings of the European conference on computer vision, pp 282\u2013297","DOI":"10.1007\/978-3-030-01258-8_18"},{"issue":"1","key":"8_CR80","first-page":"2506","volume":"33","author":"L Song","year":"2019","unstructured":"Song L, Cao J, Song L, Hu Y, He R (2019) Geometry-aware face completion and editing. Proc AAAI Conf Artif Intell 33(1):2506\u20132513","journal-title":"Proc AAAI Conf Artif Intell"},{"key":"8_CR81","doi-asserted-by":"crossref","unstructured":"Huang X, Belongie S (2017) Arbitrary style transfer in real-time with adaptive instance normalization. In: Proceedings of the IEEE international conference on computer vision, pp 1501\u20131510","DOI":"10.1109\/ICCV.2017.167"},{"key":"8_CR82","doi-asserted-by":"crossref","unstructured":"Shen Z, Huang M, Shi J, Xue X, Huang TS (2019) Towards instance-level image-to-image translation. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 3683\u20133692","DOI":"10.1109\/CVPR.2019.00380"},{"key":"8_CR83","doi-asserted-by":"crossref","unstructured":"Yin Y, Jiang S, Robinson JP, Fu Y (2020) Dual-attention GAN for large-pose face frontalization. In: Proceedings of the IEEE international conference on automatic face and gesture recognition, pp 24\u201331","DOI":"10.1109\/FG47880.2020.00004"},{"key":"8_CR84","unstructured":"Qiao F, Yao N, Jiao Z, Li Z, Chen H, Wang H (2018) Geometry-contrastive GAN for facial expression transfer. arXiv preprint 1802.01822"},{"key":"8_CR85","doi-asserted-by":"crossref","unstructured":"Wang K, Wu Q, Song L, Yang Z, Wu W, Qian C, He R, Qiao Y, Loy CC (2020) Mead: a large-scale audio-visual dataset for emotional talking-face generation. In: Proceedings of the European conference on computer vision, pp 700\u2013717","DOI":"10.1007\/978-3-030-58589-1_42"},{"key":"8_CR86","doi-asserted-by":"crossref","unstructured":"Zhou H, Sun Y, Wu W, Loy CC, Wang X, Liu Z (2021) Pose-controllable talking face generation by implicitly modularized audio-visual representation. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR46437.2021.00416"},{"key":"8_CR87","doi-asserted-by":"crossref","unstructured":"Ji X, Zhou H, Wang K, Wu W, Loy CC, Cao X, Xu F (2021) Audio-driven emotional video portraits. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR46437.2021.01386"},{"key":"8_CR88","doi-asserted-by":"crossref","unstructured":"Wang L, Han W, Soong FK (2012) High quality lip-sync animation for 3d photo-realistic talking head. In: Proceedings of the IEEE international conference on acoustics, speech and signal processing, pp 4529\u20134532","DOI":"10.1109\/ICASSP.2012.6288925"},{"key":"8_CR89","unstructured":"Yi R, Ye Z, Zhang J, Bao H, Liu Y-J (2020) Audio-driven talking face video generation with natural head pose. arXiv preprint arXiv:2002.10137"},{"key":"8_CR90","doi-asserted-by":"crossref","unstructured":"Chen L, Cui G, Liu C, Li Z, Kou Z, Xu Y, Xu C (2020) Talking-head generation with rhythmic head motion. In: Proceedings of the European conference on computer vision, pp 35\u201351","DOI":"10.1007\/978-3-030-58545-7_3"},{"key":"8_CR91","doi-asserted-by":"crossref","unstructured":"Liu K, Ostermann J (2011) Realistic facial expression synthesis for an image-based talking head. In: IEEE international conference on multimedia and expo, pp 1\u20136","DOI":"10.1109\/ICME.2011.6011835"},{"issue":"5","key":"8_CR92","doi-asserted-by":"publisher","first-page":"2421","DOI":"10.1121\/1.2229005","volume":"120","author":"M Cooke","year":"2006","unstructured":"Cooke M, Barker J, Cunningham S, Shao X (2006) An audio-visual corpus for speech perception and automatic speech recognition. J Acoust Soc America 120(5):2421\u20132424","journal-title":"J Acoust Soc America"},{"issue":"5","key":"8_CR93","doi-asserted-by":"publisher","first-page":"603","DOI":"10.1109\/TMM.2015.2407694","volume":"17","author":"N Harte","year":"2015","unstructured":"Harte N, Gillen E (2015) TCD-TIMIT: an audio-visual corpus of continuous speech. IEEE Trans Multimedia 17(5):603\u2013615","journal-title":"IEEE Trans Multimedia"},{"issue":"4","key":"8_CR94","doi-asserted-by":"publisher","first-page":"377","DOI":"10.1109\/TAFFC.2014.2336244","volume":"5","author":"H Cao","year":"2014","unstructured":"Cao H, Cooper DG, Keutmann MK, Gur RC, Nenkova A, Verma R (2014) Crema-d: crowd-sourced emotional multimodal actors dataset. IEEE Trans Affect Comput 5(4):377\u2013390","journal-title":"IEEE Trans Affect Comput"},{"key":"8_CR95","doi-asserted-by":"crossref","unstructured":"Chung JS, Zisserman A (2016) Lip reading in the wild. In: Proceedings of the Asian conference on computer vision, pp 87\u2013103","DOI":"10.1007\/978-3-319-54184-6_6"},{"key":"8_CR96","doi-asserted-by":"crossref","unstructured":"Chung JS, Senior A, Vinyals O, Zisserman A (2017) Lip reading sentences in the wild. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3444\u20133453","DOI":"10.1109\/CVPR.2017.367"},{"key":"8_CR97","doi-asserted-by":"crossref","unstructured":"Chung JS, Zisserman A (2017) Lip reading in profile. In: Proceedings of the British machine vision conference","DOI":"10.1007\/978-3-319-54184-6_6"},{"key":"8_CR98","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2019.101027","volume":"60","author":"A Nagrani","year":"2020","unstructured":"Nagrani A, Chung JS, Xie W, Zisserman A (2020) Voxceleb: large-scale speaker verification in the wild. Comput Speech Lang 60:101027","journal-title":"Comput Speech Lang"},{"key":"8_CR99","doi-asserted-by":"crossref","unstructured":"Chung JS, Nagrani A, Zisserman A (2018) Voxceleb2: deep speaker recognition. In: In proceedings of the annual conference of the international speech communication association, pp 1086\u20131090","DOI":"10.21437\/Interspeech.2018-1929"},{"issue":"4","key":"8_CR100","doi-asserted-by":"publisher","first-page":"600","DOI":"10.1109\/TIP.2003.819861","volume":"13","author":"Z Wang","year":"2004","unstructured":"Wang Z, Bovik AC, Sheikh HR, Simoncelli EP (2004) Image quality assessment: from error visibility to structural similarity. IEEE Trans Image Process 13(4):600\u2013612","journal-title":"IEEE Trans Image Process"},{"key":"8_CR101","unstructured":"Salimans T, Goodfellow IJ, Zaremba W, Cheung V, Radford A, Chen X (2016) Improved techniques for training GANs. In: Proceedings of the neural information processing systems"},{"key":"8_CR102","unstructured":"Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S (2017) GANs trained by a two time-scale update rule converge to a local nash equilibrium. In: Proceedings of the international conference on neural information processing systems, pp 6629\u20136640"},{"issue":"9","key":"8_CR103","doi-asserted-by":"publisher","first-page":"2678","DOI":"10.1109\/TIP.2011.2131660","volume":"20","author":"ND Narvekar","year":"2011","unstructured":"Narvekar ND, Karam LJ (2011) A no-reference image blur metric based on the cumulative probability of blur detection (CPBD). IEEE Trans Image Process 20(9):2678\u20132683","journal-title":"IEEE Trans Image Process"},{"key":"8_CR104","doi-asserted-by":"publisher","first-page":"149","DOI":"10.1016\/j.proeng.2013.09.086","volume":"64","author":"K De","year":"2013","unstructured":"De K, Masilamani V (2013) Image sharpness measure for blurred images in frequency domain. Procedia Eng 64:149\u2013158","journal-title":"Procedia Eng"},{"key":"8_CR105","doi-asserted-by":"crossref","unstructured":"Vougioukas K, Petridis S, Pantic M (2019) End-to-end speech-driven realistic facial animation with temporal GANs. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition workshops, pp 37\u201340","DOI":"10.1007\/s11263-019-01251-8"},{"key":"8_CR106","unstructured":"Assael YM, Shillingford B, Whiteson S, De\u00a0Freitas N (2016) Lipnet: end-to-end sentence-level lipreading. arXiv preprint arXiv:1611.01599"},{"key":"8_CR107","unstructured":"Chen L, Cui G, Kou Z, Zheng H, Xu C (2020) What comprises a good talking-head video generation?: a survey and benchmark. arXiv preprint arXiv:2005.03201"},{"key":"8_CR108","doi-asserted-by":"crossref","unstructured":"Schroff F, Kalenichenko D, Philbin J (2015) Facenet: a unified embedding for face recognition and clustering. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 815\u2013823","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"8_CR109","doi-asserted-by":"crossref","unstructured":"Deng J, Guo J, Xue N, Zafeiriou S (2019) Arcface: additive angular margin loss for deep face recognition. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 4690\u20134699","DOI":"10.1109\/CVPR.2019.00482"},{"key":"8_CR110","unstructured":"Zhang J, Zeng X, Xu C, Chen J, Liu Y, Jiang Y (2020) Apb2facev2: real-time audio-guided multi-face reenactment. arXiv preprint arXiv:2010.13017"},{"issue":"3","key":"8_CR111","doi-asserted-by":"publisher","first-page":"243","DOI":"10.1016\/0165-1781(81)90070-6","volume":"5","author":"CN Karson","year":"1981","unstructured":"Karson CN, Berman KF, Donnelly EF, Mendelson WB, Kleinman JE, Wyatt RJ (1981) Speaking, thinking, and blinking. Psychiatry Res 5(3):243\u2013246","journal-title":"Psychiatry Res"},{"issue":"12","key":"8_CR112","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0208030","volume":"13","author":"P H\u00f6mke","year":"2018","unstructured":"H\u00f6mke P, Holler J, Levinson SC (2018) Eye blinks are perceived as communicative signals in human face-to-face interaction. PloS one 13(12):e0208030","journal-title":"PloS one"},{"issue":"1","key":"8_CR113","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2926713","volume":"36","author":"Z Shu","year":"2016","unstructured":"Shu Z, Shechtman E, Samaras D, Hadap S (2016) Eyeopener: editing eyes in the wild. ACM Trans Graph 36(1):1\u201313","journal-title":"ACM Trans Graph"},{"issue":"6","key":"8_CR114","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2816795.2818056","volume":"34","author":"J Thies","year":"2015","unstructured":"Thies J, Zollh\u00f6fer M, Nie\u00dfner M, Valgaerts L, Stamminger M, Theobalt C (2015) Real-time expression transfer for facial reenactment. ACM Trans Graph 34(6):1\u201314","journal-title":"ACM Trans Graph"},{"key":"8_CR115","doi-asserted-by":"crossref","unstructured":"Velinov Z, Papas M, Bradley D, Gotardo PFU, Mirdehghan P, Marschner S, Nov\u00e1k J, Beeler T (2018) Appearance capture and modeling of human teeth. ACM Trans Graph 37(6): 207:1\u2013207:13","DOI":"10.1145\/3272127.3275098"},{"issue":"6","key":"8_CR116","first-page":"1","volume":"39","author":"L Yang","year":"2020","unstructured":"Yang L, Shi Z, Wu Y, Li X, Zhou K, Fu H, Zheng Y (2020) Iorthopredictor: model-guided deep prediction of teeth alignment. ACM Trans Graph 39(6):1\u201315","journal-title":"ACM Trans Graph"},{"key":"8_CR117","doi-asserted-by":"crossref","unstructured":"Wang T-C, Mallya A, Liu M-Y (2021) One-shot free-view neural talking-head synthesis for video conferencing. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR46437.2021.00991"},{"key":"8_CR118","unstructured":"Sadoughi N, Busso C (2019) Speech-driven expressive talking lips with conditional sequential generative adversarial networks. IEEE Trans Affect Comput"},{"issue":"5","key":"8_CR119","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0196391","volume":"13","author":"SR Livingstone","year":"2018","unstructured":"Livingstone SR, Russo FA (2018) The Ryerson audio-visual database of emotional speech and song (ravdess): a dynamic, multimodal set of facial and vocal expressions in north American English. PloS one 13(5):e0196391","journal-title":"PloS one"},{"key":"8_CR120","doi-asserted-by":"crossref","unstructured":"Zhu H, Luo M-D, Wang R, Zheng A-H, He R (2021) Deep audio-visual learning: a survey. Int J Autom Comput 1\u201326","DOI":"10.1007\/s11633-021-1293-0"}],"container-title":["Advances in Computer Vision and Pattern Recognition","Handbook of Digital Face Manipulation and Detection"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-030-87664-7_8","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,1,27]],"date-time":"2024-01-27T16:04:11Z","timestamp":1706371451000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-030-87664-7_8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022]]},"ISBN":["9783030876630","9783030876647"],"references-count":120,"URL":"https:\/\/doi.org\/10.1007\/978-3-030-87664-7_8","relation":{},"ISSN":["2191-6586","2191-6594"],"issn-type":[{"value":"2191-6586","type":"print"},{"value":"2191-6594","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022]]},"assertion":[{"value":"31 January 2022","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}}]}}