{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,19]],"date-time":"2026-03-19T14:53:33Z","timestamp":1773932013096,"version":"3.50.1"},"reference-count":47,"publisher":"Springer Science and Business Media LLC","issue":"33","license":[{"start":{"date-parts":[[2024,3,1]],"date-time":"2024-03-01T00:00:00Z","timestamp":1709251200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,3,1]],"date-time":"2024-03-01T00:00:00Z","timestamp":1709251200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004721","name":"The University of Tokyo","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004721","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In this work, we present a novel method for simultaneously controlling the head pose and the facial expressions of a given input image using a 3D keypoint-based GAN. Existing methods for controlling head pose and expressions simultaneously are not suitable for real images, or they generate unnatural results because it is not trivial to capture head pose (large changes) and expressions (small changes) simultaneously. In this work, we achieve simultaneous control of head pose and facial expressions by introducing 3D facial keypoints for GAN-based facial image synthesis, unlike the existing 2D landmark-based approach. As a result, our method can handle both large variations due to different head poses and subtle variations due to changing facial expressions faithfully. Furthermore, our model takes audio input as an additional modality for further enhancing the quality of generated images. Our model was evaluated on the VoxCeleb2 dataset to demonstrate its state-of-the-art performance for both facial reenactment and facial image manipulation tasks, and our model tends not to be affected by the driving images.<\/jats:p>","DOI":"10.1007\/s11042-024-18449-9","type":"journal-article","created":{"date-parts":[[2024,3,1]],"date-time":"2024-03-01T09:02:18Z","timestamp":1709283738000},"page":"79861-79878","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Simultaneous control of head pose and expressions in 3D facial keypoint-based GAN"],"prefix":"10.1007","volume":"83","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-7248-0288","authenticated-orcid":false,"given":"Tomoyuki","family":"Hatakeyama","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ryosuke","family":"Furuta","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yoichi","family":"Sato","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,3,1]]},"reference":[{"key":"18449_CR1","doi-asserted-by":"crossref","unstructured":"Zhu JY, Park T, Isola P, Efros AA (2017) Unpaired image-to-image translation using cycle-consistent adversarial networks. In: Proc. of the IEEE\/CVF international conference on computer vision (ICCV 2017), pp 2223\u20132232","DOI":"10.1109\/ICCV.2017.244"},{"key":"18449_CR2","doi-asserted-by":"crossref","unstructured":"Choi Y, Choi M, Kim M, Ha JW, Kim S, Choo J (2018) StarGAN: unified generative adversarial networks for multi-domain image-to-image translation. In: Proc. of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2018), pp 8789\u20138797","DOI":"10.1109\/CVPR.2018.00916"},{"key":"18449_CR3","unstructured":"Li M, Zuo W, Zhang D (2016) Deep identity-aware transfer of facial attributes. In: arXiv preprint arXiv:1610.05586"},{"key":"18449_CR4","unstructured":"Perarnau G, van\u00a0de Weijer J, Raducanu B, \u00c1lvarez JM (2016) Invertible conditional GANs for image editing. In: Advances in neural information processing systems (NeurIPS 2016) workshop on adversarial training"},{"key":"18449_CR5","doi-asserted-by":"crossref","unstructured":"Shen W, Liu R (2017) Learning residual images for face attribute manipulation. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2017), pp 4030\u20134038","DOI":"10.1109\/CVPR.2017.135"},{"key":"18449_CR6","unstructured":"Larsen ABL, S\u00f8nderby SK, Larochelle H, Winther O (2016) Autoencoding beyond pixels using a learned similarity metric. In: Proc of the machine learning research (PMLR 2016), pp 1558\u20131566"},{"key":"18449_CR7","doi-asserted-by":"crossref","unstructured":"Zhou H, Liu J, Liu Z, Liu Y, Wang X (2020) Rotate-and-Render: unsupervised photorealistic face rotation from single-view images. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2020), pp 5910\u20135919","DOI":"10.1109\/CVPR42600.2020.00595"},{"key":"18449_CR8","doi-asserted-by":"crossref","unstructured":"Wang TC, Mallya A, Liu MY (2021) One-shot free-view neural talking-head synthesis for video conferencing. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2021), pp 10039\u201310049","DOI":"10.1109\/CVPR46437.2021.00991"},{"key":"18449_CR9","doi-asserted-by":"crossref","unstructured":"Pumarola A, Agudo A, Martinez AM, Sanfeliu A, Moreno-Noguer F (2018) GANimation: anatomically-aware facial animation from a single image. In: Proc of the European conference on computer vision (ECCV 2018), pp 835\u2013851","DOI":"10.1007\/978-3-030-01249-6_50"},{"key":"18449_CR10","doi-asserted-by":"crossref","unstructured":"Tripathy S, Kannala J, Rahtu E (2021) FACEGAN: facial Attribute controllable reenactment GAN. In: Proc of the IEEE\/CVF winter conference on applications of computer vision (WACV 2021), pp 1329\u20131338","DOI":"10.1109\/WACV48630.2021.00137"},{"key":"18449_CR11","doi-asserted-by":"crossref","unstructured":"Ekman P, Friesen WV (1978) Facial action coding system: a technique for the measurement of facial movement. In: Palo Altom: consulting psychologists press","DOI":"10.1037\/t27734-000"},{"key":"18449_CR12","doi-asserted-by":"crossref","unstructured":"Deng Y, Yang J, Chen D, Wen F, Tong X (2020) Disentangled and controllable face image generation via 3D imitative-contrastive learning. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2020)","DOI":"10.1109\/CVPR42600.2020.00520"},{"key":"18449_CR13","doi-asserted-by":"crossref","unstructured":"Tewari A, Elgharib M, Bharaj G, Bernard F, Seidel HP, P\u00e9rez P, Z\u00f6llhofer M, Theobalt C (2020) StyleRig: rigging StyleGAN for 3D control over portrait images. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2020), pp 6142\u20136151","DOI":"10.1109\/CVPR42600.2020.00618"},{"key":"18449_CR14","doi-asserted-by":"crossref","unstructured":"Karras T, Laine S, Aila T (2019) A style-based generator architecture for generative adversarial networks. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2019), pp 4401\u20134410","DOI":"10.1109\/CVPR.2019.00453"},{"key":"18449_CR15","doi-asserted-by":"crossref","unstructured":"Karras T, Laine S, Aittala M, Hellsten J, Lehtinen J, Aila T (2020) Analyzing and improving the image quality of StyleGAN. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2020), pp 8110\u20138119","DOI":"10.1109\/CVPR42600.2020.00813"},{"key":"18449_CR16","doi-asserted-by":"crossref","unstructured":"Thies J, Zollh\u00f6fer M, tamminger M, Theobalt C, Nie\u00dfner M (2016) Face2Face: real-time face capture and reenactment of RGB videos. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2016) pp 2387\u20132395","DOI":"10.1109\/CVPR.2016.262"},{"issue":"6","key":"18449_CR17","doi-asserted-by":"publisher","first-page":"196:1","DOI":"10.1145\/3130800.3130818","volume":"36","author":"H Averbuch-Elor","year":"2017","unstructured":"Averbuch-Elor H, Cohen-Or D, Kopf J, Cohen MF (2017) Bringing portraits to life. ACM Trans Graph (TOG) 36(6):196:1-196:13","journal-title":"ACM Trans Graph (TOG)"},{"key":"18449_CR18","unstructured":"Wang TC, Liu MY, Zhu JY, Liu G, Tao A, Kautz J, Catanzaro B (2018) Video-to-video synthesis. In: Proc of the advances in neural information processing systems (NeurIPS 2018)"},{"key":"18449_CR19","unstructured":"Wang TC, Liu MY, Tao A, Liu G, Kautz J, Catanzaro B (2019) Few-shot video-to-video synthesis. In: Proc of the advances in neural information processing systems (NeurIPS 2019)"},{"key":"18449_CR20","doi-asserted-by":"crossref","unstructured":"Zakharov E, Shysheya A, Burkov E, Lempitsky V (2019) Few-shot adversarial learning of realistic neural talking head models. In: Proc of the IEEE\/CVF International conference on computer vision (ICCV 2019), pp 9459\u20139468","DOI":"10.1109\/ICCV.2019.00955"},{"key":"18449_CR21","doi-asserted-by":"crossref","unstructured":"Wang TC, Liu MY, Zhu JY, Tao A, Kautz J, Catanzaro B (2018) High-resolution image synthesis and semantic manipulation with conditional GANs. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2018), pp 8798\u20138807","DOI":"10.1109\/CVPR.2018.00917"},{"key":"18449_CR22","doi-asserted-by":"crossref","unstructured":"Wiles O, Koepke AS, Zisserman A (2018) X2Face: A network for controlling face generation by using images, audio, and pose codes. In: Proc of the European conference on computer vision (ECCV 2018), pp 690\u2013706","DOI":"10.1007\/978-3-030-01261-8_41"},{"key":"18449_CR23","doi-asserted-by":"crossref","unstructured":"Nirkin Y, Keller Y, Hassner T (2019) FSGAN: subject agnostic face swapping and reenactment. In: Proc of the IEEE\/CVF international conference on computer vision (ICCV 2019), pp 7184\u20137193","DOI":"10.1109\/ICCV.2019.00728"},{"key":"18449_CR24","unstructured":"Siarohin A, Lathuili\u00e8re S, Tulyakov S, Ricci E, Sebe N (2019) First order motion model for image animation. In: Proc of the advances in neural information processing systems (NeurIPS 2019)"},{"key":"18449_CR25","doi-asserted-by":"crossref","unstructured":"Siarohin A, Menapace W, Skorokhodov I, Olszewski K, Ren J, Lee HY, Chai M, Tulyakov S (2023) Unsupervised volumetric animation. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2023)","DOI":"10.1109\/CVPR52729.2023.00452"},{"key":"18449_CR26","doi-asserted-by":"crossref","unstructured":"Chen L, Maddox RK, Duan Z, Xu C (2019) Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2019), pp 7824\u20137833","DOI":"10.1109\/CVPR.2019.00802"},{"key":"18449_CR27","unstructured":"Song L, Wu W, Qian C, He R, Loy CC (2020) Everybody\u2019s Talkin\u2019: let me talk as you want. In: arXiv preprint arXiv:2001.05201"},{"key":"18449_CR28","doi-asserted-by":"crossref","unstructured":"Thies J, Elgharib M, Tewari A, Theobalt C, Nie\u00dfner M (2020) Neural voice puppetry: audio-driven facial reenactment. In: Proc of the European conference on computer vision (ECCV 2020)","DOI":"10.1007\/978-3-030-58517-4_42"},{"key":"18449_CR29","doi-asserted-by":"crossref","unstructured":"Guo Y, Chen K, Liang S, Liu Y, Bao H, Zhang J (2021) AD-NeRF: audio driven neural radiance fields for talking head synthesis. In: Proc of the IEEE\/CVF international conference on computer vision (ICCV 2021)","DOI":"10.1109\/ICCV48922.2021.00573"},{"key":"18449_CR30","doi-asserted-by":"crossref","unstructured":"Stypulkowski M, Vougioukas K, He S, Zieba M, Petridis S, Pantic M (2024) Diffused heads: diffusion models beat GANs on talking-face generation. In: Proc of the IEEE\/CVF winter conference on applications of computer vision (WACV 2024), pp 5091\u20135100","DOI":"10.1109\/WACV57701.2024.00502"},{"key":"18449_CR31","unstructured":"Amodei D, Anubhai R, Battenberg E, Case C, Casper J, Catanzaro B, JC et\u00a0al (2016) Deep speech 2 : end-to-end speech recognition in English and Mandarin. In: Proc of the international conference on machine learning (ICML 2016), pp 173\u2013182"},{"key":"18449_CR32","doi-asserted-by":"crossref","unstructured":"Panayotov V, Chen G, Povey D, Khudanpur S (2015) LibriSpeech: an ASR corpus based on public domain audio books. In: Proc of the IEEE international conference on acoustics, speech and signal processing (ICASSP 2015), pp 5206\u20135210","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"18449_CR33","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li LJ, Li K, Fei-Fei L (2009) ImageNet: a large-scale hierarchical image database. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2009), pp 248\u2013255","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"18449_CR34","doi-asserted-by":"crossref","unstructured":"Parkhi OM, Vedaldi A, Zisserman A (2015) Deep face recognition. In: Proc of the British machine vision conference (BMVC 2015), pp 41.1\u201341.12","DOI":"10.5244\/C.29.41"},{"key":"18449_CR35","unstructured":"Lim JH, Ye JC (2017) Geometric GAN. In: arXiv preprint arXiv:1705.02894"},{"key":"18449_CR36","doi-asserted-by":"crossref","unstructured":"Ruiz N, Chong E, Rehg JM (2018) Fine-grained head pose estimation without keypoints. In: the IEEE conference on computer vision and pattern recognition (CVPR 2018) workshop, pp 2187\u20132196","DOI":"10.1109\/CVPRW.2018.00281"},{"key":"18449_CR37","doi-asserted-by":"crossref","unstructured":"Zhu X, Lei Z, Liu X, Shi H, Li SZ (2016) Face alignment across large poses: a 3D solution. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2016), pp 146\u2013155","DOI":"10.1109\/CVPR.2016.23"},{"key":"18449_CR38","doi-asserted-by":"crossref","unstructured":"Yu C, Wang J, Peng C, Gao C, Yu G, Sang N (2018) BiSeNet: bilateral segmentation network for real-time semantic segmentation. In: Proc of the European conference on computer vision (ECCV 2018), pp 334\u2013349","DOI":"10.1007\/978-3-030-01261-8_20"},{"key":"18449_CR39","doi-asserted-by":"crossref","unstructured":"Lee CH, Liu Z, Wu L, Luo P (2020) MaskGAN: towards diverse and interactive facial image manipulation. In: Proc of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR 2020), pp 5549\u20135558","DOI":"10.1109\/CVPR42600.2020.00559"},{"key":"18449_CR40","unstructured":"Kingma DP, Ba J (2014) Adam: a method for stochastic optimization. In: arXiv preprint arXiv:1412.6980"},{"key":"18449_CR41","doi-asserted-by":"crossref","unstructured":"Chung JS, Nagrani A, Zisserman A (2018) VoxCeleb2: deep speaker recognition. In: Proc of the interspeech, pp 1086\u20131090","DOI":"10.21437\/Interspeech.2018-1929"},{"key":"18449_CR42","doi-asserted-by":"crossref","unstructured":"Baltrusaitis T, Zadeh A, Lim YC, Morency LP (2018) OpenFace 2.0: facial behavior analysis toolkit. In: Proc of the IEEE international conference on automatic face gesture recognition (FG 2018), pp 59\u201366","DOI":"10.1109\/FG.2018.00019"},{"key":"18449_CR43","doi-asserted-by":"crossref","unstructured":"Baltrusaitis T, Mahmoud M, Robinson P (2015) Cross-dataset learning and person-specific normalisation for automatic action unit detection. In: Proc of the IEEE international conference on automatic face gesture recognition (FG 2015), pp 1\u20136","DOI":"10.1109\/FG.2015.7284869"},{"key":"18449_CR44","doi-asserted-by":"crossref","unstructured":"Zakharov E, Ivakhnenko A, Shysheya A, Lempitsky V (2020) Fast bi-layer neural synthesis of one-shot realistic head avatars. In: Proc of the European conference on computer vision (ECCV 2020), pp 524\u2013540","DOI":"10.1007\/978-3-030-58610-2_31"},{"key":"18449_CR45","doi-asserted-by":"crossref","unstructured":"Wang Z, Bovik AC, Sheikh HR, Member S, Simoncelli EP, Member S (2004) Image quality assessment: from error visibility to structural similarity. In: Proc of the IEEE transactions on image processing (TIP 2004) pp 600\u2013612","DOI":"10.1109\/TIP.2003.819861"},{"key":"18449_CR46","unstructured":"Wang Z, Simoncelli EP, Bovik AC (2003) Multi-scale structural similarity for image quality assessment. In: Proceedings of the IEEE asilomar conference on signals, systems, and computers (Asilomar), vol\u00a02, pp 1398\u20131402"},{"key":"18449_CR47","unstructured":"Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S (2017) GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In: Proc of the advances in neural information processing systems (NeurIPS 2017), pp 6629\u20136640"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-024-18449-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-024-18449-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-024-18449-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,7]],"date-time":"2024-10-07T13:28:26Z","timestamp":1728307706000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-024-18449-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,1]]},"references-count":47,"journal-issue":{"issue":"33","published-online":{"date-parts":[[2024,10]]}},"alternative-id":["18449"],"URL":"https:\/\/doi.org\/10.1007\/s11042-024-18449-9","relation":{},"ISSN":["1573-7721"],"issn-type":[{"value":"1573-7721","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,1]]},"assertion":[{"value":"10 October 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 January 2024","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 January 2024","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 March 2024","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of Interest"}}]}}