{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,11]],"date-time":"2026-04-11T09:59:10Z","timestamp":1775901550435,"version":"3.50.1"},"reference-count":69,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61902415 and 62421002"],"award-info":[{"award-number":["61902415 and 62421002"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Open Fund of PDL","award":["WDZC20235250106"],"award-info":[{"award-number":["WDZC20235250106"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>Text-driven Talking Head Generation (THG) marks a significant advancement in the video production industry by enabling the creation of realistic talking head videos with minimal data input. While previous research has explored few-shot talking head synthesis, these methods often fall short in terms of lip-sync consistency and expression diversity, both of which are essential for practical applications. In this article, we present a novel multi-modal framework, VTalker, for synthesizing talking heads with specific vocal tones and speaking styles. First, a Text-to-Speech (T2S) model is developed for generating speech with a given voice tone from text. Second, we design a novel multi-modal fusion module to effectively integrate speech and text features. Third, we propose the V-DiT (Video Diffusion Transformer) as the backbone to frame the generation of talking heads as a temporally iterative denoising task. To further enhance performance, the appearance and temporal conditions are incorporated into the backbone as tokens, along with patches, to improve image fidelity and ensure smooth spatio-temporal motion. This innovative design enables our model to generate high-fidelity, text-synchronized talking head videos that generalize smoothly across various identities. Extensive experiments demonstrate the effectiveness of our approach in generating high-quality text-driven talking head videos.<\/jats:p>","DOI":"10.1145\/3793546","type":"journal-article","created":{"date-parts":[[2026,2,14]],"date-time":"2026-02-14T14:27:41Z","timestamp":1771079261000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["VTalker: Text-Driven Synthesis of Talking Head with Vision Diffusion Transformer"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4039-895X","authenticated-orcid":false,"given":"Yali","family":"Cai","sequence":"first","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Processing, College of Computer Science and Technology, National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6752-7892","authenticated-orcid":false,"given":"Peng","family":"Qiao","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Processing, College of Computer Science and Technology, National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9743-2034","authenticated-orcid":false,"given":"Dongsheng","family":"Li","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Processing, College of Computer Science and Technology, National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,11]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/2503385.2503473"},{"key":"e_1_3_1_3_2","doi-asserted-by":"crossref","unstructured":"Fan Bao Shen Nie Kaiwen Xue Yue Cao Chongxuan Li Hang Su and Jun Zhu. 2023. All are worth words: A ViT backbone for diffusion models. arXiv:2209.12152. Retrieved from https:\/\/arxiv.org\/abs\/2209.12152","DOI":"10.1109\/CVPR52729.2023.02171"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02171"},{"key":"e_1_3_1_5_2","first-page":"4","article-title":"Is space-time attention all you need for video understanding","author":"Bertasius Gedas","year":"2021","unstructured":"Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Machine Learning, 4.","journal-title":"Proceedings of the 38th International Conference on Machine Learning"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_32"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_32"},{"key":"e_1_3_1_8_2","unstructured":"Shoufa Chen Mengmeng Xu Jiawei Ren Yuren Cong Sen He Yanping Xie Animesh Sinha Ping Luo Tao Xiang and Juan-Manuel Perez-Rua. 2023. GenTron: Delving deep into diffusion transformers for image and video generation. arXiv:2312.04557. Retrieved from https:\/\/arxiv.org\/abs\/2312.04557"},{"key":"e_1_3_1_9_2","first-page":"6441","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Shoufa","year":"2024","unstructured":"Shoufa Chen, Mengmeng Xu, Jiawei Ren, Yuren Cong, Sen He, Yanping Xie, Animesh Sinha, Ping Luo, Tao Xiang, and Juan-Manuel Perez-Rua. 2024. GenTron: Diffusion transformers for image and video generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 6441\u20136451."},{"key":"e_1_3_1_10_2","unstructured":"Hanbo Cheng Limin Lin Chenyu Liu Pengcheng Xia Pengfei Hu Jiefeng Ma Jun Du and Jia Pan. 2024. DAWN: Dynamic frame avatar with non-autoregressive diffusion framework for talking head video generation. arXiv:2410.13726. Retrieved from https:\/\/arxiv.org\/abs\/2410.13726"},{"key":"e_1_3_1_11_2","first-page":"251","volume-title":"ACCV Workshops,","author":"Chung Joon Son","year":"2016","unstructured":"Joon Son Chung and Andrew Zisserman. 2016. Out of time: Automated lip sync in the wild. In ACCV Workshops, 251\u2013263."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2017.2765202"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3613800"},{"key":"e_1_3_1_14_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold et al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.02069"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00573"},{"key":"e_1_3_1_17_2","unstructured":"Fa-Ting Hong Zunnan Xu Zixiang Zhou Jun Zhou Xiu Li Qin Lin Qinglin Lu and Dan Xu. 2025. Audio-visual controlled video diffusion with masked selective state spaces modeling for natural talking head generation. arXiv:2504.02542. Retrieved from https:\/\/arxiv.org\/abs\/2504.02542"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","first-page":"23802","DOI":"10.1609\/aaai.v38i21.30570","article-title":"Audiogpt: Understanding and generating speech, music, sound, and talking head","author":"Huang Rongjie","year":"2024","unstructured":"Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al. 2024. Audiogpt: Understanding and generating speech, music, sound, and talking head. In Proceedings of the 38th AAAI Conference on Artificial Intelligence, 23802\u201323804.","journal-title":"Proceedings of the 38th AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_1_19_2","unstructured":"Jaehyeon Kim Jungil Kong and Juhee Son. 2021. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. arXiv:2106.06103. Retrieved from https:\/\/arxiv.org\/abs\/2106.06103"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1561\/2200000056"},{"key":"e_1_3_1_21_2","unstructured":"Diederik P. Kingma and Max Welling. 2022. Auto-encoding variational bayes. arXiv:1312.6114. Retrieved from https:\/\/arxiv.org\/abs\/1312.6114"},{"key":"e_1_3_1_22_2","unstructured":"Jungil Kong Jaehyeon Kim and Jaekyoung Bae. 2020. HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis. arXiv:2010.05646. Retrieved from https:\/\/arxiv.org\/abs\/2010.05646"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","unstructured":"Jungil Kong Jihoon Park Beomjeong Kim Jeongmin Kim Dohee Kong and Sangjin Kim. 2023. VITS2: Improving quality and efficiency of single-stage text-to-speech with adversarial learning and architecture design. arXiv:2307.16430. Retrieved from https:\/\/arxiv.org\/abs\/2307.16430","DOI":"10.21437\/Interspeech.2023-534"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2023.3247001"},{"key":"e_1_3_1_25_2","first-page":"3037","article-title":"Ae-nerf: Audio enhanced neural radiance field for few shot talking head synthesis","author":"Li Dongze","year":"2024","unstructured":"Dongze Li, Kang Zhao, Wei Wang, Bo Peng, Yingya Zhang, Jing Dong, and Tieniu Tan. 2024. Ae-nerf: Audio enhanced neural radiance field for few shot talking head synthesis. In Proceedings of the 38th AAAI Conference on Artificial Intelligence, 3037\u20133045.","journal-title":"Proceedings of the 38th AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_1_26_2","first-page":"1","volume-title":"Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Li Jingyi","year":"2023","unstructured":"Jingyi Li, Weiping Tu, and Li Xiao. 2023. Freevc: Towards high-quality text-free one-shot voice conversion. In Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1\u20135."},{"key":"e_1_3_1_27_2","unstructured":"Jiahe Li Jiawei Zhang Xiao Bai Jun Zhou and Lin Gu. 2023. Efficient region-aware neural radiance fields for high-fidelity talking portrait synthesis. arXiv:2307.09323. Retrieved from https:\/\/arxiv.org\/abs\/2307.09323"},{"key":"e_1_3_1_28_2","first-page":"1911","article-title":"Write-a-speaker: Text-based emotional and rhythmic talking-head generation","author":"Li Lincheng","year":"2021","unstructured":"Lincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding, Yixing Zheng, Xin Yu, and Changjie Fan. 2021. Write-a-speaker: Text-based emotional and rhythmic talking-head generation. In Proceedings of the 35th AAAI Conference on Artificial Intelligence, 1911\u20131920.","journal-title":"Proceedings of the 35th AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_1_29_2","unstructured":"Lincheng Li Suzhen Wang Zhimeng Zhang Yu Ding Yixing Zheng Xin Yu and Changjie Fan. 2021. Write-a-speaker: Text-based emotional and rhythmic talking-head generation. arXiv:2104.07995. Retrieved from https:\/\/arxiv.org\/abs\/2104.07995"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00544"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.3390\/e25101440"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3414412"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2021.3049196"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2024.3517748"},{"key":"e_1_3_1_35_2","unstructured":"Ilya Loshchilov and Frank Hutter. 2017. Decoupled Weight Decay Regularization. arXiv:1711.05101. Retrieved from https:\/\/arxiv.org\/abs\/1711.05101"},{"key":"e_1_3_1_36_2","first-page":"1896","article-title":"Styletalk: One-shot talking head generation with controllable speaking styles","author":"Ma Yifeng","year":"2023","unstructured":"Yifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan, Tangjie Lv, Yu Ding, Zhidong Deng, and Xin Yu. 2023. Styletalk: One-shot talking head generation with controllable speaking styles. In Proceedings of the 37th AAAI Conference on Artificial Intelligence, 1896\u20131904.","journal-title":"Proceedings of the 37th AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_1_37_2","unstructured":"Yifeng Ma Suzhen Wang Zhipeng Hu Changjie Fan Tangjie Lv Yu Ding Zhidong Deng and Xin Yu. 2023. StyleTalk: One-shot talking head generation with controllable speaking styles. arXiv:2301.01081. Retrieved from https:\/\/arxiv.org\/abs\/2301.01081"},{"key":"e_1_3_1_38_2","unstructured":"Yifeng Ma Shiwei Zhang Jiayu Wang Xiang Wang Yingya Zhang and Zhidong Deng. 2023. Dreamtalk: When expressive talking head generation meets diffusion probabilistic models. arXiv:2312.09767. Retrieved from https:\/\/arxiv.org\/abs\/2312.09767"},{"key":"e_1_3_1_39_2","unstructured":"OpenAI. 2024. Sora: Creating Video from Text. Retrieved from https:\/\/openai.com\/sora"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00387"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00070"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413532"},{"key":"e_1_3_1_43_2","unstructured":"Zengyi Qin Wenliang Zhao Xumin Yu and Xin Sun. 2023. Openvoice: Versatile instant voice cloning. arXiv:2312.01479. Retrieved from https:\/\/arxiv.org\/abs\/2312.01479"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","unstructured":"Nataniel Ruiz Eunji Chong and James M. Rehg. 2018. Fine-grained head pose estimation without keypoints. arXiv:1710.00925. Retrieved from https:\/\/arxiv.org\/abs\/1710.00925","DOI":"10.1109\/CVPRW.2018.00281"},{"key":"e_1_3_1_46_2","unstructured":"Aliaksandr Siarohin St\u00e9phane Lathuili\u00e8re Sergey Tulyakov Elisa Ricci and Nicu Sebe. 2020. First order motion model for image animation. arXiv:2003.00196. Retrieved from https:\/\/arxiv.org\/abs\/2003.00196"},{"key":"e_1_3_1_47_2","doi-asserted-by":"crossref","unstructured":"Li Siyao Weijiang Yu Tianpei Gu Chunze Lin Quan Wang Chen Qian Chen Change Loy and Ziwei Liu. 2022. Bailando: 3D dance generation by actor-critic GPT with choreographic memory. arXiv:2203.13055. Retrieved from https:\/\/arxiv.org\/abs\/2203.13055","DOI":"10.1109\/CVPR52688.2022.01077"},{"key":"e_1_3_1_48_2","unstructured":"stability. 2024. Stable Diffusion 3.0. Retrieved from https:\/\/stability.ai\/news\/stable-diffusion-3"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00502"},{"issue":"1","key":"e_1_3_1_50_2","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1109\/JSAC.2022.3221953","article-title":"Txt2Vid: Ultra-low bitrate compression of talking-head videos via text","volume":"41","author":"Tandon Pulkit","year":"2022","unstructured":"Pulkit Tandon, Shubham Chandak, Pat Pataranutaporn, Yimeng Liu, Anesu M. Mapuranga, Pattie Maes, Tsachy Weissman, and Misha Sra. 2022. Txt2Vid: Ultra-low bitrate compression of talking-head videos via text. IEEE Journal on Selected Areas in Communications 41, 1 (2022), 107\u2013118.","journal-title":"IEEE Journal on Selected Areas in Communications"},{"key":"e_1_3_1_51_2","doi-asserted-by":"crossref","unstructured":"Yehui Tang Yunhe Wang Jianyuan Guo Zhijun Tu Kai Han Hailin Hu and Dacheng Tao. 2024. A survey on transformer compression. arXiv:2402.05964. Retrieved from https:\/\/arxiv.org\/abs\/2402.05964","DOI":"10.1111\/jan.16422"},{"key":"e_1_3_1_52_2","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2023. Attention is all you need. arXiv:1706.03762. Retrieved from https:\/\/arxiv.org\/abs\/1706.03762"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.3390\/app13053100"},{"key":"e_1_3_1_54_2","unstructured":"Ting-Chun Wang Ming-Yu Liu Jun-Yan Zhu Guilin Liu Andrew Tao Jan Kautz and Bryan Catanzaro. 2018. Video-to-video synthesis. arXiv:1808.06601. Retrieved from https:\/\/arxiv.org\/abs\/1808.06601"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_1_56_2","unstructured":"Zhichao Wang Mengyu Dai and Keld Lundgaard. 2023. Text-to-video: A two-stage framework for zero-shot identity-agnostic talking-head generation. arXiv:2308.06457. Retrieved from https:\/\/arxiv.org\/abs\/2308.06457"},{"key":"e_1_3_1_57_2","unstructured":"Wayne Wu Chen Qian Shuo Yang Quan Wang Yici Cai and Qiang Zhou. 2018. Look at boundary: A boundary-aware face alignment algorithm. arXiv:1805.10483. Retrieved from https:\/\/arxiv.org\/abs\/1805.10483"},{"key":"e_1_3_1_58_2","unstructured":"Zhen Xing Qijun Feng Haoran Chen Qi Dai Han Hu Hang Xu Zuxuan Wu and Yu-Gang Jiang. 2023. A survey on video diffusion models. arXiv:2310.10647. Retrieved from https:\/\/arxiv.org\/abs\/2310.10647"},{"key":"e_1_3_1_59_2","unstructured":"Jingjing Xu Xu Sun Zhiyuan Zhang Guangxiang Zhao and Junyang Lin. 2019. Understanding and improving layer normalization. arXiv:1911.07013. Retrieved from https:\/\/arxiv.org\/abs\/1911.07013"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626235"},{"key":"e_1_3_1_61_2","unstructured":"Zhuoyi Yang Jiayan Teng Wendi Zheng Ming Ding Shiyu Huang Jiazheng Xu Yuanming Yang Wenyi Hong Xiaohan Zhang Guanyu Feng et al. 2024. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv:2408.06072. Retrieved from https:\/\/arxiv.org\/abs\/2408.06072"},{"key":"e_1_3_1_62_2","unstructured":"Zhenhui Ye Tianyun Zhong Yi Ren Jiaqi Yang Weichuang Li Jiawei Huang Ziyue Jiang Jinzheng He Rongjie Huang Jinglin Liu et al. 2024. Real3d-portrait: One-shot realistic 3d talking portrait synthesis. arXiv:2401.08503. Retrieved from https:\/\/arxiv.org\/abs\/2401.08503"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00703"},{"key":"e_1_3_1_64_2","first-page":"421","volume-title":"CCF Conference on Computer Supported Cooperative Work and Social Computing","author":"Yuan Hui","year":"2023","unstructured":"Hui Yuan, Ping Li, Gansen Zhao, and Jun Zhang. 2023. Improving voice style conversion via self-attention VAE with feature disentanglement. In CCF Conference on Computer Supported Cooperative Work and Social Computing. Springer, 421\u2013435."},{"key":"e_1_3_1_65_2","first-page":"2659","volume-title":"Proceedings of the ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Zhang Sibo","year":"2022","unstructured":"Sibo Zhang, Jiahong Yuan, Miao Liao, and Liangjun Zhang. 2022. Text2video: Text-driven talking-head video synthesis with personalized phoneme-pose dictionary. In Proceedings of the ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2659\u20132663."},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00836"},{"key":"e_1_3_1_67_2","doi-asserted-by":"crossref","first-page":"154","DOI":"10.1007\/978-3-031-46674-8_11","volume-title":"Proceedings of the International Conference on Advanced Data Mining and Applications","author":"Zhang Xulong","year":"2023","unstructured":"Xulong Zhang, Jianzong Wang, Ning Cheng, and Jing Xiao. 2023. Voice conversion with denoising diffusion probabilistic gan models. In Proceedings of the International Conference on Advanced Data Mining and Applications. Springer, 154\u2013167."},{"key":"e_1_3_1_68_2","doi-asserted-by":"crossref","first-page":"3660","DOI":"10.1109\/CVPR46437.2021.00366","volume-title":"Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhang Zhimeng","year":"2021","unstructured":"Zhimeng Zhang, Lincheng Li, and Yu Ding. 2021. Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset. In Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3660\u20133669."},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2023.110132"},{"key":"e_1_3_1_70_2","doi-asserted-by":"crossref","unstructured":"Hang Zhou Yasheng Sun Wayne Wu Chen Change Loy Xiaogang Wang and Ziwei Liu. 2021. Pose-controllable talking face generation by implicitly modularized audio-visual representation. arXiv:2104.11116. Retrieved from https:\/\/arxiv.org\/abs\/2104.11116","DOI":"10.1109\/CVPR46437.2021.00416"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3793546","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,11]],"date-time":"2026-04-11T09:19:30Z","timestamp":1775899170000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3793546"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,11]]},"references-count":69,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3793546"],"URL":"https:\/\/doi.org\/10.1145\/3793546","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,11]]},"assertion":[{"value":"2025-04-13","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-17","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}