{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,15]],"date-time":"2026-03-15T04:48:17Z","timestamp":1773550097573,"version":"3.50.1"},"reference-count":69,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62431015 and 62571317"],"award-info":[{"award-number":["62431015 and 62571317"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Fundamental Research Funds for the Central Universities, Shanghai Key Laboratory of Digital Media Processing and Transmission","award":["22DZ2229005"],"award-info":[{"award-number":["22DZ2229005"]}]},{"DOI":"10.13039\/501100013314","name":"111 project","doi-asserted-by":"crossref","award":["BP0719010"],"award-info":[{"award-number":["BP0719010"]}],"id":[{"id":"10.13039\/501100013314","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,3,31]]},"abstract":"<jats:p>\n                    Although audio-driven talking face generation has witnessed significant advances in recent years, two problems remain to be solved. First, the existing models cannot control the long-term head actions as humans expect, because the audio can only provide short-term cues, such as the rhythms and sentiments, for head movements. Second, generating long-term head poses and ensuring accurate lip motions remain challenging due to the difficulty in harnessing the optimization process for large-scale head movements and small-scale mouth motions. In this study, we propose a novel method to address these issues. First, to alleviate the limitations of audio conditions, we propose a Pose Latent Diffusion (PLD) model to generate head motions from two kinds of input modalities: the input audio and user-controlled text prompts. The audio provides short-term rhythm correspondence with the head movements, while the text prompts describe the long-term semantics of head motions. Second, we propose a refinement-based learning strategy to synthesize head movements and accurate lip motions using two cascaded networks, namely CoarseNet and RefineNet. The CoarseNet estimates coarse global motions to produce animated images with changed poses, and the RefineNet progressively estimates finer lip motions from low to high resolutions, yielding improved lip-synchronization performance. Experiments demonstrate that our method can achieve better pose diversity and realness compared to audio-based pose generation baselines, and our video generator model outperforms state-of-the-art methods in synthesizing natural head motions. Projects and demos are available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/junleen.github.io\/projects\/posetalk\">https:\/\/junleen.github.io\/projects\/posetalk<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3788691","type":"journal-article","created":{"date-parts":[[2026,2,14]],"date-time":"2026-02-14T14:27:41Z","timestamp":1771079261000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["PoseTalk: Exploring Text- and Audio-Based Pose Control for One-Shot Talking Face Generation"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7260-7141","authenticated-orcid":false,"given":"Jun","family":"Ling","sequence":"first","affiliation":[{"name":"Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University, Shanghai, China and Shanghai Key Laboratory of Collaborative Computing in Spacial Heterogeneous Networks, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-9646-0720","authenticated-orcid":false,"given":"Yiwen","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4699-4701","authenticated-orcid":false,"given":"Han","family":"Xue","sequence":"additional","affiliation":[{"name":"Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8261-5337","authenticated-orcid":false,"given":"Rong","family":"Xie","sequence":"additional","affiliation":[{"name":"Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7124-5182","authenticated-orcid":false,"given":"Li","family":"Song","sequence":"additional","affiliation":[{"name":"Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University, Shanghai, China and MOE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,2,27]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"12449","article-title":"Wav2Vec 2.0: A framework for self-supervised learning of speech representations","volume":"33","author":"Baevski Alexei","year":"2020","unstructured":"Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. Wav2Vec 2.0: A framework for self-supervised learning of speech representations. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 33, 12449\u201312460.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_3_2","unstructured":"Tim Brooks Bill Peebles Connor Holmes Will DePue Yufei Guo Li Jing David Schnurr Joe Taylor Troy Luhman Eric Luhman Clarence Ng Ricky Wang and Aditya Ramesh. 2024. Video Generation Models as World Simulators. Retrieved from https:\/\/openai.com\/research\/video-generation-models-as-world-simulators"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58545-7_3"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00802"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9747495"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01726"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3550469.3555399"},{"key":"e_1_3_2_9_2","first-page":"251","volume-title":"Proceedings of the Asian Conference on Computer Vision: ACCV 2016 International Workshops (ACCV \u201916 Workshops)","author":"Chung Joon Son","year":"2016","unstructured":"Joon Son Chung and Andrew Zisserman. 2016. Out of time: Automated lip sync in the wild. In Proceedings of the Asian Conference on Computer Vision: ACCV 2016 International Workshops (ACCV \u201916 Workshops), 251\u2013263."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00941"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00482"},{"key":"e_1_3_2_12_2","doi-asserted-by":"crossref","unstructured":"Yao Feng Haiwen Feng Michael J. Black and Timo Bolkart. 2021. Learning an animatable detailed 3D face model from in-the-wild images. ACM Transactions on Graphics 40 4 (2021) 1\u201313.","DOI":"10.1145\/3450626.3459936"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.02069"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-3015"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00509"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00573"},{"key":"e_1_3_2_17_2","volume-title":"Proceedings of the International Conference on Learning Representation","author":"Guo Yuwei","year":"2024","unstructured":"Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. 2024. AnimateDiff: Animate your personalized text-to-image diffusion models without specific tuning. In Proceedings of the International Conference on Learning Representation."},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"e_1_3_2_19_2","volume-title":"Proceedings of the International Conference on Learning Representation","author":"He Tianyu","year":"2024","unstructured":"Tianyu He, Junliang Guo, Runyi Yu, Yuchi Wang, Jialiang Zhu, Kaikai An, Leyi Li, Xu Tan, Chunyu Wang, Han Hu, et al. 2024. GAIA: Zero-shot talking avatar generation. In Proceedings of the International Conference on Learning Representation."},{"key":"e_1_3_2_20_2","first-page":"6626","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local Nash equilibrium. In Proceedings of the Advances in Neural Information Processing Systems, 6626\u20136637."},{"key":"e_1_3_2_21_2","volume-title":"Proceedings of the NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications","author":"Ho Jonathan","year":"2022","unstructured":"Jonathan Ho and Tim Salimans. 2022. Classifier-free diffusion guidance. In Proceedings of the NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications."},{"key":"e_1_3_2_22_2","first-page":"8633","article-title":"Video diffusion models","volume":"35","author":"Ho Jonathan","year":"2022","unstructured":"Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. 2022. Video diffusion models. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 35, 8633\u20138646.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.167"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3528233.3530745"},{"key":"e_1_3_2_25_2","unstructured":"Diederik P. Kingma and Max Welling. 2013. Auto-encoding variational Bayes. arXiv:1312.6114. Retrieved from https:\/\/arxiv.org\/abs\/1312.6114"},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","unstructured":"Haojie Li Hao Chen Yining Huang Tianshui Chen and Shuangping Huang. 2025. Enhancing lip dynamic authenticity: Learning 3D temporal representations for talking head synthesis. ACM Transactions on Multimedia Computing Communications and Applications 21 10 (2025) 1\u201321.","DOI":"10.1145\/3750048"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00696"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00338"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2023.3333552"},{"key":"e_1_3_2_30_2","doi-asserted-by":"crossref","unstructured":"Jun Ling Han Xue Anni Tang Rong Xie and Li Song. 2024. ViCoFace: Learning disentangled latent motion representations for visual-consistent face reenactment. ACM Transactions on Multimedia Computing Communications and Applications 20 12 (2024) 1\u201324.","DOI":"10.1145\/3698769"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","unstructured":"Shiguang Liu and Huixin Wang. 2023. Talking face generation via facial anatomy. ACM Transactions on Multimedia Computing Communications and Applications 19 3 (2023) 1\u201319.","DOI":"10.1145\/3571746"},{"key":"e_1_3_2_32_2","doi-asserted-by":"crossref","unstructured":"Zhilei Liu Xiaoxing Liu Sen Chen Jiaxing Liu Longbiao Wang and Chongke Bi. 2024. Multimodal fusion for talking face generation utilizing speech-related facial action units. ACM Transactions on Multimedia Computing Communications and Applications 20 9 (2024) 1\u201324.","DOI":"10.1145\/3672565"},{"key":"e_1_3_2_33_2","unstructured":"Camillo Lugaresi Jiuqiang Tang Hadon Nash Chris McClanahan Esha Uboweja Michael Hays Fan Zhang Chuo-Ling Chang Ming Guang Yong Juhyun Lee et al. 2019. MediaPipe: A framework for building perception pipelines. arXiv:1906.08172. Retrieved from https:\/\/arxiv.org\/abs\/1906.08172"},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","unstructured":"Yiyang Ma Haowei Kuang Huan Yang Jianlong Fu and Jiaying Liu. 2024. Prompt-based modality bridging for unified text-to-face generation and manipulation. ACM Transactions on Multimedia Computing Communications and Applications 20 12 (2024) 1\u201323.","DOI":"10.1145\/3694974"},{"key":"e_1_3_2_35_2","first-page":"1","article-title":"StyleTalk: One-shot talking head generation with controllable speaking styles","author":"Ma Yifeng","year":"2023","unstructured":"Yifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan, Tangjie Lv, Yu Ding, Zhidong Deng, and Xin Yu. 2023. StyleTalk: One-shot talking head generation with controllable speaking styles. In Proceedings of the AAAI Conference on Artificial Intelligence, 1\u20139.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_2_36_2","unstructured":"Yifeng Ma Shiwei Zhang Jiayu Wang Xiang Wang Yingya Zhang and Zhidong Deng. 2023. DreamTalk: When expressive talking head generation meets diffusion probabilistic models. arXiv:2312.09767. Retrieved from https:\/\/arxiv.org\/abs\/2312.09767"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-950"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-4009"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20047-2_28"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413532"},{"key":"e_1_3_2_41_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01350"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00197"},{"key":"e_1_3_2_45_2","first-page":"1","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"32","author":"Siarohin Aliaksandr","year":"2019","unstructured":"Aliaksandr Siarohin, St\u00e9phane Lathuili\u00e8re, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019. First order motion model for image animation. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 32, 1\u201311."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00502"},{"key":"e_1_3_2_47_2","unstructured":"Kuiyuan Sun Xiaolong Liu Xiaolong Li Yao Zhao and Wei Wang. 2024. Multi-modal driven pose-controllable talking head generation. ACM Transactions on Multimedia Computing Communications and Applications 20 12 (2024) 1\u201323."},{"key":"e_1_3_2_48_2","doi-asserted-by":"crossref","unstructured":"Anni Tang Tianyu He Xu Tan Jun Ling Runnan Li Sheng Zhao Li Song and Jiang Bian. 2024. Memories are one-to-many mapping alleviators in talking face generation. IEEE Transactions on Pattern Analysis and Machine Intelligence 46 12 (2024) 8758\u20138770.","DOI":"10.1109\/TPAMI.2024.3409380"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20047-2_21"},{"key":"e_1_3_2_50_2","volume-title":"Proceedings of the International Conference on Learning Representation","author":"Tevet Guy","year":"2023","unstructured":"Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-Or, and Amit Haim Bermano. 2023. Human motion diffusion model. In Proceedings of the International Conference on Learning Representation."},{"key":"e_1_3_2_51_2","doi-asserted-by":"crossref","unstructured":"Linrui Tian Qi Wang Bang Zhang and Liefeng Bo. 2024. EMO: Emote portrait alive-generating expressive portrait videos with audio2video diffusion model under weak conditions. arXiv:2402.17485. Retrieved from https:\/\/arxiv.org\/abs\/2402.17485","DOI":"10.1007\/978-3-031-73010-8_15"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01724"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58589-1_42"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2021\/152"},{"key":"e_1_3_2_55_2","volume-title":"Proceedings of the International Conference on Learning Representation","author":"Wang Yaohui","year":"2022","unstructured":"Yaohui Wang, Di Yang, Francois Bremond, and Antitza Dantcheva. 2022. Latent image animator: Learning to animate images via latent space navigation. In Proceedings of the International Conference on Learning Representation."},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3449075"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_2_58_2","unstructured":"Mingwang Xu Hui Li Qingkun Su Hanlin Shang Liwei Zhang Ce Liu Jingdong Wang Yao Yao and Siyu Zhu. 2024. 2024. Hallo: Hierarchical audio-driven visual synthesis for portrait image animation. arXiv:2406.08801. Retrieved from https:\/\/arxiv.org\/abs\/2406.08801"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611765"},{"key":"e_1_3_2_60_2","volume-title":"Proceedings of the International Conference on Learning Representation","author":"Ye Zhenhui","year":"2023","unstructured":"Zhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu, Jinzheng He, and Zhou Zhao. 2023. GeneFace: Generalized and high-fidelity audio-driven 3D talking face synthesis. In Proceedings of the International Conference on Learning Representation."},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01422"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00384"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00836"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00366"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00938"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33019299"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00416"},{"key":"e_1_3_2_69_2","doi-asserted-by":"crossref","unstructured":"Yang Zhou Xintong Han Eli Shechtman Jose Echevarria Evangelos Kalogerakis and Dingzeyu Li. 2020. MakeltTalk: Speaker-aware talking-head animation. ACM Transactions on Graphics 39 6 (2020) 1\u201315.","DOI":"10.1145\/3414685.3417774"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00545"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3788691","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,15]],"date-time":"2026-03-15T03:48:44Z","timestamp":1773546524000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3788691"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,27]]},"references-count":69,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,3,31]]}},"alternative-id":["10.1145\/3788691"],"URL":"https:\/\/doi.org\/10.1145\/3788691","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,27]]},"assertion":[{"value":"2025-05-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-03","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}