{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,11]],"date-time":"2026-04-11T20:37:55Z","timestamp":1775939875180,"version":"3.50.1"},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","license":[{"start":{"date-parts":[[2024,8,10]],"date-time":"2024-08-10T00:00:00Z","timestamp":1723248000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"abstract":"<jats:p>Talking head, driving a source image to generate a talking video using other modality information, has made great progress in recent years. However, there are two main issues: 1) These methods are designed to utilize a single modality of information. 2) Most methods cannot control head pose. To address these problems, we propose a novel framework that can utilize multi-modal information to generate a talking head video, while achieving arbitrary head pose control by a movement sequence. Specifically, first, to extend driving information to multiple modalities, multi-modal information is encoded to a unified semantic latent space to generate expression parameters. Secondly, to disentangle attributes, the 3D Morphable Model (3DMM) is utilized to obtain identity information from the source image, and translation and rotation information from the target image. Thirdly, to control head pose and mouth shape, the source image is warped by a motion field generated by the expression parameter, translation parameter, and angle parameter. Finally, all the above parameters are utilized to render a landmark map, and the warped source image is combined with the landmark map to generate a delicate talking head video. Our experimental results demonstrate that our proposed method is capable of achieving state-of-the-art performance in terms of visual quality, lip-audio synchronization, and head pose control.<\/jats:p>","DOI":"10.1145\/3673901","type":"journal-article","created":{"date-parts":[[2024,8,10]],"date-time":"2024-08-10T11:02:51Z","timestamp":1723287771000},"update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Multi-Modal Driven Pose-Controllable Talking Head Generation"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-0077-7388","authenticated-orcid":false,"given":"Kuiyuan","family":"Sun","sequence":"first","affiliation":[{"name":"Institute of Information Science, Beijing Key Laboratory of Advanced Information Science and Network Technology, Beijing Jiaotong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3940-207X","authenticated-orcid":false,"given":"Xiaolong","family":"Liu","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Beijing Key Laboratory of Advanced Information Science and Network Technology, Beijing Jiaotong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6181-6044","authenticated-orcid":false,"given":"Xiaolong","family":"Li","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Beijing Key Laboratory of Advanced Information Science and Network Technology, Beijing Jiaotong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8581-9554","authenticated-orcid":false,"given":"Yao","family":"Zhao","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Beijing Key Laboratory of Advanced Information Science and Network Technology, Beijing Jiaotong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5477-1017","authenticated-orcid":false,"given":"Wei","family":"wang","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Beijing Key Laboratory of Advanced Information Science and Network Technology, Beijing Jiaotong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,8,10]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Volker Blanz and Thomas Vetter. 1999. A morphable model for the synthesis of 3D faces. In Siggraph. 187\u2013194.","DOI":"10.1145\/311535.311556"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01565"},{"key":"e_1_2_1_3_1","volume-title":"Lip Reading Sentences in the Wild. In IEEE Conference on Computer Vision and Pattern Recognition.","author":"Chung J. S.","unstructured":"J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman. 2017. Lip Reading Sentences in the Wild. In IEEE Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_2_1_4_1","volume-title":"Asian conference on computer vision. Springer, 251\u2013263","author":"Chung Joon Son","year":"2016","unstructured":"Joon Son Chung and Andrew Zisserman. 2016. Out of time: automated lip sync in the wild. In Asian conference on computer vision. Springer, 251\u2013263."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01034"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2019.00038"},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Xuanyi Dong Yan Yan Wanli Ouyang and Yi Yang. 2018. Style aggregated network for facial landmark detection. In CVPR. 379\u2013388.","DOI":"10.1109\/CVPR.2018.00047"},{"key":"e_1_2_1_8_1","volume-title":"et\u00a0al","author":"Dosovitskiy Alexey","year":"2020","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et\u00a0al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","unstructured":"Yao Feng Haiwen Feng Michael J. Black and Timo Bolkart. 2021. Learning an Animatable Detailed 3D Face Model from In-The-Wild Images. ACM Transactions on Graphics (Proc. SIGGRAPH) 40 8. https:\/\/doi.org\/10.1145\/3450626.3459936","DOI":"10.1145\/3450626.3459936"},{"key":"e_1_2_1_10_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3450626.3459936","article-title":"Learning an animatable detailed 3d face model from in-the-wild images","volume":"40","author":"Feng Yao","year":"2021","unstructured":"Yao Feng, Haiwen Feng, Michael J Black, and Timo Bolkart. 2021. Learning an animatable detailed 3d face model from in-the-wild images. ACM Transactions on Graphics (ToG) 40, 4 (2021), 1\u201313.","journal-title":"ACM Transactions on Graphics (ToG)"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528223.3530164"},{"key":"e_1_2_1_12_1","volume-title":"PFLD: A practical facial landmark detector. arXiv preprint arXiv:1902.10859","author":"Guo Xiaojie","year":"2019","unstructured":"Xiaojie Guo, Siyuan Li, Jinke Yu, Jiawan Zhang, Jiayi Ma, Lin Ma, Wei Liu, and Haibin Ling. 2019. PFLD: A practical facial landmark detector. arXiv preprint arXiv:1902.10859 (2019)."},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 20374\u201320384","author":"Hong Yang","year":"2022","unstructured":"Yang Hong, Bo Peng, Haiyao Xiao, Ligang Liu, and Juyong Zhang. 2022. Headnerf: A real-time nerf-based parametric head model. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 20374\u201320384."},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision. 22157\u201322167","author":"Huang Tianyu","year":"2023","unstructured":"Tianyu Huang, Bowen Dong, Yunhan Yang, Xiaoshui Huang, Rynson WH Lau, Wanli Ouyang, and Wangmeng Zuo. 2023. Clip2point: Transfer clip to point cloud classification with image-depth pre-training. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 22157\u201322167."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.167"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01386"},{"key":"e_1_2_1_17_1","doi-asserted-by":"crossref","unstructured":"Liming Jiang Ren Li Wayne Wu Chen Qian and Chen Change Loy. 2020. Deeperforensics-1.0: A large-scale dataset for real-world face forgery detection. In CVPR. 2889\u20132898.","DOI":"10.1109\/CVPR42600.2020.00296"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"e_1_2_1_19_1","volume-title":"Faceshifter: Towards high fidelity and occlusion aware face swapping. arXiv preprint arXiv:1912.13457","author":"Li Lingzhi","year":"2019","unstructured":"Lingzhi Li, Jianmin Bao, Hao Yang, Dong Chen, and Fang Wen. 2019. Faceshifter: Towards high fidelity and occlusion aware face swapping. arXiv preprint arXiv:1912.13457 (2019)."},{"key":"e_1_2_1_20_1","volume-title":"Talking Face Generation via Facial Anatomy. ACM Transactions on Multimedia Computing, Communications and Applications","author":"Liu Shiguang","year":"2022","unstructured":"Shiguang Liu and Huixin Wang. 2022. Talking Face Generation via Facial Anatomy. ACM Transactions on Multimedia Computing, Communications and Applications (2022)."},{"key":"e_1_2_1_21_1","volume-title":"TCSD: Triple Complementary Streams Detector for Comprehensive Deepfake Detection. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)","author":"Liu Xiaolong","year":"2022","unstructured":"Xiaolong Liu, Yang Yu, Xiaolong Li, Yao Zhao, and Guodong Guo. 2022. TCSD: Triple Complementary Streams Detector for Comprehensive Deepfake Detection. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) (2022)."},{"key":"e_1_2_1_22_1","volume-title":"2019 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct). IEEE, 200\u2013202","author":"Liu Zhaoxiang","year":"2019","unstructured":"Zhaoxiang Liu, Huan Hu, Zipeng Wang, Kai Wang, Jinqiang Bai, and Shiguo Lian. 2019. Video synthesis of human upper body with realistic face. In 2019 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct). IEEE, 200\u2013202."},{"key":"e_1_2_1_23_1","first-page":"1","article-title":"ML-CookGAN: Multi-Label Generative Adversarial Network for Food Image Generation","volume":"19","author":"Liu Zhiming","year":"2023","unstructured":"Zhiming Liu, Kai Niu, and Zhiqiang He. 2023. ML-CookGAN: Multi-Label Generative Adversarial Network for Food Image Generation. ACM Transactions on Multimedia Computing, Communications and Applications 19, 2s (2023), 1\u201321.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_2_1_24_1","doi-asserted-by":"crossref","unstructured":"Luming Ma and Zhigang Deng. 2019. Real-time hierarchical facial performance capture. In ACM SIGGRAPH. 1\u201310.","DOI":"10.1145\/3306131.3317016"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503250"},{"key":"e_1_2_1_26_1","volume-title":"Joon Son Chung, and Andrew Zisserman","author":"Nagrani Arsha","year":"2017","unstructured":"Arsha Nagrani, Joon Son Chung, and Andrew Zisserman. 2017. Voxceleb: a large-scale speaker identification dataset. arXiv preprint arXiv:1706.08612 (2017)."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00244"},{"key":"e_1_2_1_28_1","doi-asserted-by":"crossref","unstructured":"Pascal Paysan Reinhard Knothe Brian Amberg Sami Romdhani and Thomas Vetter. 2009. A 3D face model for pose and illumination invariant face recognition. In 2009 sixth IEEE international conference on advanced video and signal based surveillance. Ieee 296\u2013301.","DOI":"10.1109\/AVSS.2009.58"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413532"},{"key":"e_1_2_1_30_1","volume-title":"Ganimation: Anatomically-aware facial animation from a single image. In ECCV. 818\u2013833.","author":"Pumarola Albert","year":"2018","unstructured":"Albert Pumarola, Antonio Agudo, Aleix M Martinez, Alberto Sanfeliu, and Francesc Moreno-Noguer. 2018. Ganimation: Anatomically-aware facial animation from a single image. In ECCV. 818\u2013833."},{"key":"e_1_2_1_31_1","volume-title":"Fastspeech 2: Fast and high-quality end-to-end text to speech. arXiv preprint arXiv:2006.04558","author":"Ren Yi","year":"2020","unstructured":"Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020. Fastspeech 2: Fast and high-quality end-to-end text to speech. arXiv preprint arXiv:2006.04558 (2020)."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01350"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00121"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"e_1_2_1_35_1","volume-title":"First order motion model for image animation. Advances in Neural Information Processing Systems 32","author":"Siarohin Aliaksandr","year":"2019","unstructured":"Aliaksandr Siarohin, St\u00e9phane Lathuili\u00e8re, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019. First order motion model for image animation. Advances in Neural Information Processing Systems 32 (2019)."},{"key":"e_1_2_1_36_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_2_1_37_1","volume-title":"Next3D: Generative Neural Texture Rasterization for 3D-Aware Head Avatars. arXiv preprint arXiv:2211.11208","author":"Sun Jingxiang","year":"2022","unstructured":"Jingxiang Sun, Xuan Wang, Lizhen Wang, Xiaoyu Li, Yong Zhang, Hongwen Zhang, and Yebin Liu. 2022. Next3D: Generative Neural Texture Rasterization for 3D-Aware Head Avatars. arXiv preprint arXiv:2211.11208 (2022)."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073640"},{"key":"e_1_2_1_39_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_1_40_1","volume-title":"Audio2head: Audio-driven one-shot talking-head generation with natural head motion. arXiv preprint arXiv:2107.09293","author":"Wang Suzhen","year":"2021","unstructured":"Suzhen Wang, Lincheng Li, Yu Ding, Changjie Fan, and Xin Yu. 2021. Audio2head: Audio-driven one-shot talking-head generation with natural head motion. arXiv preprint arXiv:2107.09293 (2021)."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i3.20154"},{"key":"e_1_2_1_42_1","volume-title":"Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. Advances in neural information processing systems 29","author":"Wu Jiajun","year":"2016","unstructured":"Jiajun Wu, Chengkai Zhang, Tianfan Xue, Bill Freeman, and Josh Tenenbaum. 2016. Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. Advances in neural information processing systems 29 (2016)."},{"key":"e_1_2_1_43_1","doi-asserted-by":"crossref","unstructured":"Wayne Wu Chen Qian Shuo Yang Quan Wang Yici Cai and Qiang Zhou. 2018. Look at boundary: A boundary-aware face alignment algorithm. In CVPR. 2129\u20132138.","DOI":"10.1109\/CVPR.2018.00227"},{"key":"e_1_2_1_44_1","volume-title":"CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior. arXiv preprint arXiv:2301.02379","author":"Xing Jinbo","year":"2023","unstructured":"Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, and Tien-Tsin Wong. 2023. CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior. arXiv preprint arXiv:2301.02379 (2023)."},{"key":"e_1_2_1_45_1","first-page":"1","article-title":"Progressive Transformer Machine for Natural Character Reenactment","volume":"19","author":"Xu Yongzong","year":"2023","unstructured":"Yongzong Xu, Zhijing Yang, Tianshui Chen, Kai Li, and Chunmei Qing. 2023. Progressive Transformer Machine for Natural Character Reenactment. ACM Transactions on Multimedia Computing, Communications and Applications 19, 2s (2023), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_2_1_46_1","volume-title":"High-Fidelity Face Reenactment via Identity-Matched Correspondence Learning. ACM Transactions on Multimedia Computing, Communications and Applications","author":"Xue Han","year":"2022","unstructured":"Han Xue, Jun Ling, Anni Tang, Li Song, Rong Xie, and Wenjun Zhang. 2022. High-Fidelity Face Reenactment via Identity-Matched Correspondence Learning. ACM Transactions on Multimedia Computing, Communications and Applications (2022)."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1631\/FITEE.2100463"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3499026"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00955"},{"key":"e_1_2_1_50_1","volume-title":"Freenet: Multi-identity face reenactment. In CVPR. 5326\u20135335.","author":"Zhang Jiangning","year":"2020","unstructured":"Jiangning Zhang, Xianfang Zeng, Mengmeng Wang, Yusu Pan, Liang Liu, Yong Liu, Yu Ding, and Changjie Fan. 2020. Freenet: Multi-identity face reenactment. In CVPR. 5326\u20135335."},{"key":"e_1_2_1_51_1","volume-title":"Controlvideo: Training-free controllable text-to-video generation. arXiv preprint arXiv:2305.13077","author":"Zhang Yabo","year":"2023","unstructured":"Yabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang, Wangmeng Zuo, and Qi Tian. 2023. Controlvideo: Training-free controllable text-to-video generation. arXiv preprint arXiv:2305.13077 (2023)."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00416"},{"key":"e_1_2_1_53_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3414685.3417774","article-title":"Makelttalk: speaker-aware talking-head animation","volume":"39","author":"Zhou Yang","year":"2020","unstructured":"Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. 2020. Makelttalk: speaker-aware talking-head animation. ACM Transactions on Graphics (TOG) 39, 6 (2020), 1\u201315.","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_1_54_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision. 22257\u201322267","author":"Zhu Hongguang","year":"2023","unstructured":"Hongguang Zhu, Yunchao Wei, Xiaodan Liang, Chunjie Zhang, and Yao Zhao. 2023. Ctp: Towards vision-language continual pretraining via compatible momentum contrast and topology preservation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 22257\u201322267."},{"key":"e_1_2_1_55_1","volume-title":"CelebV-HQ: A Large-Scale Video Facial Attributes Dataset. arXiv preprint arXiv:2207.12393","author":"Zhu Hao","year":"2022","unstructured":"Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy. 2022. CelebV-HQ: A Large-Scale Video Facial Attributes Dataset. arXiv preprint arXiv:2207.12393 (2022)."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3673901","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3673901","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:58:23Z","timestamp":1750294703000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3673901"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,10]]},"references-count":55,"alternative-id":["10.1145\/3673901"],"URL":"https:\/\/doi.org\/10.1145\/3673901","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,8,10]]},"assertion":[{"value":"2023-04-04","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-06-04","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-08-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"3673901"}}