{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T19:55:15Z","timestamp":1783454115696,"version":"3.55.0"},"reference-count":67,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T00:00:00Z","timestamp":1783382400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2021ZD0112100"],"award-info":[{"award-number":["2021ZD0112100"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National NSF of China","award":["62120106009"],"award-info":[{"award-number":["62120106009"]}]},{"name":"National NSF of China","award":["62372033"],"award-info":[{"award-number":["62372033"]}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["2022XKRC015"],"award-info":[{"award-number":["2022XKRC015"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]},{"name":"China Academy of Traditional Chinese Medicine Science and Technology Innovation Project","award":["CAM-KJHT-2024-0026"],"award-info":[{"award-number":["CAM-KJHT-2024-0026"]}]},{"name":"Shandong Provincial Natural Science Foundation\u2014Excellent Youth Program","award":["2026HWYQ-029"],"award-info":[{"award-number":["2026HWYQ-029"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>Speech-driven 3D facial animation has been studied for a long time, yet many challenges still prevent achieving truly natural results. One major issue is that current methods often overlook head motion during speech. To address this, we conducted a preliminary investigation and identified several key obstacles: the scarcity of 3D facial animation datasets with head motion and the limitations of regular sequence prediction models (e.g., Transformer, LSTM, and GRU), which are designed for discrete dynamic sequences and do not align well with continuous 3D head motion sequences. To solve these issues, we propose a head motion prediction module. This module uses audio and the initial motion state of the mesh to predict head motion. By employing ordinary differential equations (ODE) to model continuous dynamic sequences, it predicts head movements that closely resemble real head motion, making the animation more realistic and natural. Additionally, recognizing that the head motion generation given audio is not a one-to-one mapping problem, we introduce a noise module to help the head motion prediction module generate varied head motions given the same audio input. We also observed that previous facial animation methods primarily focus on generating vertices for the mouth region but use a single model to generate the entire face. This approach wastes some of the model\u2019s fitting capacity on other regions. To solve this problem, we propose a cascaded mesh generation module that uses two modules to separately generate the vertex of mouth region and other facial regions. Extensive experiments and a perceptual user study show that our approach outperforms existing methods and produces relatively natural head motion.<\/jats:p>","DOI":"10.1145\/3811818","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:10:33Z","timestamp":1782839433000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Speech-Driven 3D Facial Animation with Natural Head Movements"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-0077-7388","authenticated-orcid":false,"given":"Kuiyuan","family":"Sun","sequence":"first","affiliation":[{"name":"Institute of Information Science, Beijing Key Laboratory of Advanced Information Science and Network Technology, Beijing Jiaotong University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6189-9521","authenticated-orcid":false,"given":"Jichao","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Ocean University of China, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5477-1017","authenticated-orcid":false,"given":"Wei","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Beijing Key Laboratory of Advanced Information Science and Network Technology, Beijing Jiaotong University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8581-9554","authenticated-orcid":false,"given":"Yao","family":"Zhao","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Beijing Key Laboratory of Advanced Information Science and Network Technology, Beijing Jiaotong University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6597-7248","authenticated-orcid":false,"given":"Nicu","family":"Sebe","sequence":"additional","affiliation":[{"name":"Department of Information and Computer Science, University of Trento, Trento, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,7]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"T. Afouras J. S. Chung A. Senior O. Vinyals and A. Zisserman. 2018. Deep audio-visual speech recognition. arXiv:1809.02108. Retrieved from https:\/\/arxiv.org\/abs\/1809.02108"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02009"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.14778\/3137765.3137775"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01962"},{"key":"e_1_3_2_6_2","volume-title":"Advances in Neural Information Processing Systems","volume":"31","author":"Chen Ricky T. Q.","year":"2018","unstructured":"Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David K. Duvenaud. 2018. Neural ordinary differential equations. In Advances in Neural Information Processing Systems, Vol. 31."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3539234"},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","unstructured":"Joon Son Chung Arsha Nagrani and Andrew Zisserman. 2018. VoxCeleb2: Deep speaker recognition. arXiv:1806.05622. Retrieved from https:\/\/arxiv.org\/abs\/1806.05622","DOI":"10.21437\/Interspeech.2018-1929"},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of the Asian Conference on Computer Vision","author":"Chung J. S.","year":"2016","unstructured":"J. S. Chung and A. Zisserman. 2016. Lip reading in the wild. In Proceedings of the Asian Conference on Computer Vision."},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","unstructured":"Razvan-Gabriel Cirstea Chenjuan Guo Bin Yang Tung Kieu Xuanyi Dong and Shirui Pan. 2022. Triformer: Triangular variable-specific attentions for long sequence multivariate time series forecasting\u2013full version. arXiv:2204.13767. Retrieved from https:\/\/arxiv.org\/abs\/2204.13767","DOI":"10.24963\/ijcai.2022\/277"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01034"},{"key":"e_1_3_2_12_2","first-page":"894","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Cuturi Marco","year":"2017","unstructured":"Marco Cuturi and Mathieu Blondel. 2017. Soft-DTW: A differentiable loss function for time-series. In Proceedings of the International Conference on Machine Learning. PMLR, 894\u2013903."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01967"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01821"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459936"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01022"},{"key":"e_1_3_2_17_2","unstructured":"Ian J. Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial networks. arXiv:1406.2661. Retrieved from https:\/\/arxiv.org\/abs\/1406.2661"},{"key":"e_1_3_2_18_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201922)","author":"Hong Yang","year":"2022","unstructured":"Yang Hong, Bo Peng, Haiyao Xiao, Ligang Liu, and Juyong Zhang. 2022. HeadNeRF: A real-time NeRF-based parametric head model. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201922)."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3160360"},{"key":"e_1_3_2_20_2","unstructured":"Cheng Hsin Wuu Ningyuan Zheng Scott Ardisson Rohan Bali Danielle Belko Eric Brockmeyer Lucas Evans Timothy Godisart Hyowon Ha Xuhua Huang et al. 2023. Multiface: A dataset for neural face rendering. arXiv:2207.11243. Retrieved from https:\/\/arxiv.org\/abs\/2207.11243"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3050672"},{"key":"e_1_3_2_22_2","unstructured":"Diederik P. Kingma and Max Welling. 2022. Auto-Encoding Variational Bayes. arXiv:1312.6114. Retrieved from https:\/\/arxiv.org\/abs\/1312.6114"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3592455"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","unstructured":"Mike Lewis Yinhan Liu Naman Goyal Marjan Ghazvininejad Abdelrahman Mohamed Omer Levy Ves Stoyanov and Luke Zettlemoyer. 2019. BART: Denoising sequence-to-sequence pre-training for natural language generation translation and comprehension. arXiv:1910.13461. Retrieved from https:\/\/arxiv.org\/abs\/1910.13461","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01644"},{"key":"e_1_3_2_26_2","unstructured":"Jiahe Li Jiawei Zhang Xiao Bai Jin Zheng Xin Ning Jun Zhou and Lin Gu. 2024. TalkingGaussian: Structure-persistent 3D talking head synthesis via Gaussian splatting. arXiv:2404.15264. Retrieved from https:\/\/arxiv.org\/abs\/2404.15264"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3130800.3130813"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3141604"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijforecast.2021.03.012"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3271816"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3172548"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0196391"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01621"},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","unstructured":"Ben Mildenhall Pratul P. Srinivasan Matthew Tancik Jonathan T. Barron Ravi Ramamoorthi and Ren Ng. 2020. NeRF: Representing scenes as neural radiance fields for view synthesis. arXiv:2003.08934. Retrieved from https:\/\/arxiv.org\/abs\/2003.08934","DOI":"10.1007\/978-3-030-58452-8_24"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.earscirev.2018.12.005"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2019.101027"},{"key":"e_1_3_2_37_2","unstructured":"Dario Pavllo David Grangier and Michael Auli. 2018. QuaterNet: A quaternion-based recurrent model for human motion. arXiv:1805.06485. Retrieved from https:\/\/arxiv.org\/abs\/1805.06485"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01642"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01891"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01919"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00121"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_2_43_2","unstructured":"Amin Shabani Amir Abdi Lili Meng and Tristan Sylvain. 2023. Scaleformer: Iterative multi-scale refining transformers for time series forecasting. arXiv:2206.04038. Retrieved from https:\/\/arxiv.org\/abs\/2206.04038"},{"key":"e_1_3_2_44_2","volume-title":"Advances in Neural Information Processing Systems","volume":"32","author":"Siarohin Aliaksandr","year":"2019","unstructured":"Aliaksandr Siarohin, St\u00e9phane Lathuili\u00e8re, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019. First order motion model for image animation. In Advances in Neural Information Processing Systems, Vol. 32."},{"key":"e_1_3_2_45_2","doi-asserted-by":"crossref","unstructured":"Stefan Stan Kazi Injamamul Haque and Zerrin Yumak. 2023. FaceDiffuser: Speech-driven 3D facial animation synthesis using diffusion. arXiv:2309.11306. Retrieved from https:\/\/arxiv.org\/abs\/2309.11306","DOI":"10.1145\/3623264.3624447"},{"key":"e_1_3_2_46_2","unstructured":"Fan-Keng Sun and Duane S. Boning. 2022. FreDo: Frequency domain-based long-term time series forecasting. arXiv:220512301. Retrieved from https:\/\/arxiv.org\/abs\/2205.12301"},{"key":"e_1_3_2_47_2","doi-asserted-by":"crossref","unstructured":"Zhiyao Sun Tian Lv Sheng Ye Matthieu Gaetan Lin Jenny Sheng Yu-Hui Wen Minjing Yu and Yong-Jin Liu. 2023. DiffPoseTalk: Speech-driven stylistic 3D facial animation and head pose generation via diffusion models. arXiv:2310.00434. Retrieved from https:\/\/arxiv.org\/abs\/2310.00434","DOI":"10.1145\/3658221"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01885"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41591-018-0300-7"},{"key":"e_1_3_2_50_2","unstructured":"Aaron van den Oord Oriol Vinyals and Koray Kavukcuoglu. 2018. Neural discrete representation learning. arXiv:171100937. Retrieved from https:\/\/arxiv.org\/abs\/1711.00937"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3322326"},{"key":"e_1_3_2_52_2","unstructured":"Qianyun Wang Zhenfeng Fan and Shihong Xia. 2021. 3D-TalkEmo: Learning to synthesize 3D emotional talking head. arXiv:2104.12051. Retrieved from https:\/\/arxiv.org\/abs\/2104.12051"},{"key":"e_1_3_2_53_2","first-page":"6607","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Wang Yuyang","year":"2019","unstructured":"Yuyang Wang, Alex Smola, Danielle Maddix, Jan Gasthaus, Dean Foster, and Tim Januschowski. 2019. Deep factors for forecasting. In Proceedings of the International Conference on Machine Learning. PMLR, 6607\u20136617."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i5.25754"},{"key":"e_1_3_2_55_2","first-page":"22419","volume-title":"Advances in Neural Information Processing Systems","volume":"34","author":"Wu Haixu","year":"2021","unstructured":"Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long. 2021. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. In Advances in Neural Information Processing Systems, Vol. 34, 22419\u201322430."},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3289757"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01229"},{"key":"e_1_3_2_58_2","unstructured":"Mingwang Xu Hui Li Qingkun Su Hanlin Shang Liwei Zhang Ce Liu Jingdong Wang Luc Van Gool Yao Yao and Siyu Zhu. 2024. Hallo: Hierarchical audio-driven visual synthesis for portrait image animation. arXiv:2406.08801. Retrieved from https:\/\/arxiv.org\/abs\/2406.08801"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00189"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3322895"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.2990354"},{"key":"e_1_3_2_62_2","unstructured":"Zhenhui Ye Ziyue Jiang Yi Ren Jinglin Liu JinZheng He and Zhou Zhao. 2023. GeneFace: Generalized and high-fidelity audio-driven 3D talking face synthesis. arXiv:2301.13430. Retrieved from https:\/\/arxiv.org\/abs\/2301.13430"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCI.2018.2840738"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3091863"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00836"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2016.2557063"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i12.17325"},{"key":"e_1_3_2_68_2","first-page":"27268","volume-title":"Proceedings of the 39th International Conference on Machine Learning","author":"Zhou Tian","year":"2022","unstructured":"Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin. 2022. FEDformer: Frequency enhanced decomposed transformer for long-term series forecasting. In Proceedings of the 39th International Conference on Machine Learning, Vol. 162. Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (Eds.), PMLR, 27268\u201327286. Retrieved from https:\/\/proceedings.mlr.press\/v162\/zhou22g.html"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3811818","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T18:39:40Z","timestamp":1783449580000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3811818"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,7]]},"references-count":67,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3811818"],"URL":"https:\/\/doi.org\/10.1145\/3811818","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,7]]},"assertion":[{"value":"2025-05-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-14","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-07-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}