{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,10]],"date-time":"2025-12-10T04:09:35Z","timestamp":1765339775487,"version":"3.46.0"},"publisher-location":"New York, NY, USA","reference-count":85,"publisher":"ACM","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,10,27]]},"DOI":"10.1145\/3746027.3754734","type":"proceedings-article","created":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T07:26:51Z","timestamp":1761377211000},"page":"905-914","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-1239-8119","authenticated-orcid":false,"given":"Gaoxiang","family":"Cong","sequence":"first","affiliation":[{"name":"Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China and University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1943-8219","authenticated-orcid":false,"given":"Liang","family":"Li","sequence":"additional","affiliation":[{"name":"Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4436-8830","authenticated-orcid":false,"given":"Jiadong","family":"Pan","sequence":"additional","affiliation":[{"name":"Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China and University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-1129-8914","authenticated-orcid":false,"given":"Zhedong","family":"Zhang","sequence":"additional","affiliation":[{"name":"Hangzhou Dianzi University, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5988-5494","authenticated-orcid":false,"given":"Amin","family":"Beheshti","sequence":"additional","affiliation":[{"name":"Macquarie University, Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3027-8364","authenticated-orcid":false,"given":"Anton","family":"van den Hengel","sequence":"additional","affiliation":[{"name":"University of Adelaide, Adelaide, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4312-5682","authenticated-orcid":false,"given":"Yuankai","family":"Qi","sequence":"additional","affiliation":[{"name":"Macquarie University, Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-8793-6953","authenticated-orcid":false,"given":"Qingming","family":"Huang","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,10,27]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Seed-tts: A family of high-quality versatile speech generation models. arXiv preprint arXiv:2406.02430","author":"Anastassiou Philip","year":"2024","unstructured":"Philip Anastassiou, Jiawei Chen, Jitong Chen, Yuanzhe Chen, Zhuo Chen, Ziyi Chen, Jian Cong, Lelai Deng, Chuang Ding, Lu Gao, et al., 2024. Seed-tts: A family of high-quality versatile speech generation models. arXiv preprint arXiv:2406.02430 (2024)."},{"key":"e_1_3_2_1_2_1","unstructured":"Alexei Baevski Yuhao Zhou Abdelrahman Mohamed and Michael Auli. 2020. wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations. In NIPS."},{"key":"e_1_3_2_1_3_1","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. In NeurIPS."},{"key":"e_1_3_2_1_4_1","first-page":"21210","article-title":"V2C","author":"Chen Qi","year":"2022","unstructured":"Qi Chen, Mingkui Tan, Yuankai Qi, Jiaqiu Zhou, Yuanqing Li, and Qi Wu. 2022. V2C: Visual Voice Cloning. In CVPR. 21210-21219.","journal-title":"Visual Voice Cloning. In CVPR."},{"key":"e_1_3_2_1_5_1","volume-title":"Neural ordinary differential equations. Advances in neural information processing systems","author":"Chen Ricky TQ","year":"2018","unstructured":"Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud. 2018. Neural ordinary differential equations. Advances in neural information processing systems, Vol. 31 (2018)."},{"key":"e_1_3_2_1_6_1","volume-title":"Vall-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers. arXiv preprint arXiv:2406.05370","author":"Chen Sanyuan","year":"2024","unstructured":"Sanyuan Chen, Shujie Liu, Long Zhou, Yanqing Liu, Xu Tan, Jinyu Li, Sheng Zhao, Yao Qian, and Furu Wei. 2024a. Vall-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers. arXiv preprint arXiv:2406.05370 (2024)."},{"key":"e_1_3_2_1_7_1","volume-title":"F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching. arXiv preprint arXiv:2410.06885","author":"Chen Yushen","year":"2024","unstructured":"Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Chunhui Wang, Jian Zhao, Kai Yu, and Xie Chen. 2024b. F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching. arXiv preprint arXiv:2410.06885 (2024)."},{"key":"e_1_3_2_1_8_1","first-page":"1","article-title":"V2SFlow","author":"Choi Jeongsoo","year":"2025","unstructured":"Jeongsoo Choi, Ji-Hoon Kim, Jinyu Li, Joon Son Chung, and Shujie Liu. 2025. V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow. In ICASSP. IEEE, 1-5.","journal-title":"Video-to-Speech Generation with Speech Decomposition and Rectified Flow. In ICASSP. IEEE"},{"key":"e_1_3_2_1_9_1","first-page":"27325","article-title":"Av2av: Direct audio-visual speech to audio-visual speech translation with unified audio-visual speech representation","author":"Choi Jeongsoo","year":"2024","unstructured":"Jeongsoo Choi, Se Jin Park, Minsu Kim, and Yong Man Ro. 2024. Av2av: Direct audio-visual speech to audio-visual speech translation with unified audio-visual speech representation. In CVPR. 27325-27337.","journal-title":"CVPR."},{"key":"e_1_3_2_1_10_1","volume-title":"Out of Time: Automated Lip Sync in the Wild. In ACCV Workshop. 251-263","author":"Chung Joon Son","year":"2016","unstructured":"Joon Son Chung and Andrew Zisserman. 2016. Out of Time: Automated Lip Sync in the Wild. In ACCV Workshop. 251-263."},{"key":"e_1_3_2_1_11_1","first-page":"14687","article-title":"Learning to Dub Movies via Hierarchical Prosody Models","author":"Cong Gaoxiang","year":"2023","unstructured":"Gaoxiang Cong, Liang Li, Yuankai Qi, Zheng-Jun Zha, Qi Wu, Wenyu Wang, Bin Jiang, Ming-Hsuan Yang, and Qingming Huang. 2023. Learning to Dub Movies via Hierarchical Prosody Models. In CVPR. 14687-14697.","journal-title":"CVPR."},{"key":"e_1_3_2_1_12_1","volume-title":"EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing. arXiv preprint arXiv:2412.08988","author":"Cong Gaoxiang","year":"2024","unstructured":"Gaoxiang Cong, Jiadong Pan, Liang Li, Yuankai Qi, Yuxin Peng, Anton van den Hengel, Jian Yang, and Qingming Huang. 2024a. EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing. arXiv preprint arXiv:2412.08988 (2024)."},{"key":"e_1_3_2_1_13_1","first-page":"6767","article-title":"StyleDubber","author":"Cong Gaoxiang","year":"2024","unstructured":"Gaoxiang Cong, Yuankai Qi, Liang Li, Amin Beheshti, Zhedong Zhang, Anton van den Hengel, Ming-Hsuan Yang, Chenggang Yan, and Qingming Huang. 2024b. StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing. In Findings of ACL. 6767-6779.","journal-title":"Towards Multi-Scale Style Learning for Movie Dubbing. In Findings of ACL."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.2229005"},{"key":"e_1_3_2_1_15_1","first-page":"1331","article-title":"Stochastic context consistency reasoning for domain adaptive object detection","author":"Cui Yiming","year":"2024","unstructured":"Yiming Cui, Liang Li, Jiehua Zhang, Chenggang Yan, Hongkui Wang, Shuai Wang, Heng Jin, and Li Wu. 2024. Stochastic context consistency reasoning for domain adaptive object detection. In ACM MM. 1331-1340.","journal-title":"ACM MM."},{"key":"e_1_3_2_1_16_1","volume-title":"Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens. arXiv preprint arXiv:2407.05407","author":"Du Zhihao","year":"2024","unstructured":"Zhihao Du, Qian Chen, Shiliang Zhang, Kai Hu, Heng Lu, Yexin Yang, Hangrui Hu, Siqi Zheng, Yue Gu, Ziyang Ma, et al., 2024a. Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens. arXiv preprint arXiv:2407.05407 (2024)."},{"key":"e_1_3_2_1_17_1","unstructured":"Zhihao Du Yuxuan Wang Qian Chen Xian Shi Xiang Lv Tianyu Zhao Zhifu Gao Yexin Yang Changfeng Gao Hui Wang et al. 2024b. Cosyvoice 2: Scalable streaming speech synthesis with large language models. arXiv preprint arXiv:2412.10117 (2024)."},{"key":"e_1_3_2_1_18_1","unstructured":"Daya Guo Dejian Yang Haowei Zhang Junxiao Song Ruoyu Zhang Runxin Xu Qihao Zhu Shirong Ma Peiyi Wang Xiao Bi et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948 (2025)."},{"key":"e_1_3_2_1_19_1","volume-title":"Fireredtts: A foundation text-to-speech framework for industry-level generative speech applications. arXiv preprint arXiv:2409.03283","author":"Guo Hao-Han","year":"2024","unstructured":"Hao-Han Guo, Kun Liu, Fei-Yu Shen, Yi-Chen Wu, Feng-Long Xie, Kun Xie, and Kai-Tuo Xu. 2024b. Fireredtts: A foundation text-to-speech framework for industry-level generative speech applications. arXiv preprint arXiv:2409.03283 (2024)."},{"key":"e_1_3_2_1_20_1","first-page":"11121","article-title":"VoiceFlow","author":"Guo Yiwei","year":"2024","unstructured":"Yiwei Guo, Chenpeng Du, Ziyang Ma, Xie Chen, and Kai Yu. 2024a. VoiceFlow: Efficient Text-To-Speech with Rectified Flow Matching. In ICASSP. 11121-11125.","journal-title":"Efficient Text-To-Speech with Rectified Flow Matching. In ICASSP."},{"key":"e_1_3_2_1_21_1","first-page":"16582","article-title":"Neural Dubber: Dubbing for Videos According to Scripts","author":"Hu Chenxu","year":"2021","unstructured":"Chenxu Hu, Qiao Tian, Tingle Li, Yuping Wang, Yuxuan Wang, and Hang Zhao. 2021. Neural Dubber: Dubbing for Videos According to Scripts. In NeurIPS. 16582-16595.","journal-title":"NeurIPS."},{"key":"e_1_3_2_1_22_1","first-page":"8818","article-title":"Faces that Speak: Jointly Synthesising Talking Face and Speech from Text","author":"Jang Youngjoon","year":"2024","unstructured":"Youngjoon Jang, Ji-Hoon Kim, Junseok Ahn, Doyeop Kwak, Hongsun Yang, Yooncheol Ju, Ilhwan Kim, Byeong-Yeol Kim, and Joon Son Chung. 2024. Faces that Speak: Jointly Synthesising Talking Face and Speech from Text. In CVPR. 8818-8828.","journal-title":"CVPR."},{"key":"e_1_3_2_1_23_1","unstructured":"Shengpeng Ji Ziyue Jiang Wen Wang Yifu Chen Minghui Fang Jialong Zuo Qian Yang Xize Cheng Zehan Wang Ruiqi Li et al. 2024a. Wavtokenizer: an efficient acoustic discrete codec tokenizer for audio language modeling. arXiv preprint arXiv:2408.16532 (2024)."},{"key":"e_1_3_2_1_24_1","unstructured":"Shengpeng Ji Ziyue Jiang Wen Wang Yifu Chen Minghui Fang Jialong Zuo Qian Yang Xize Cheng Zehan Wang Ruiqi Li et al. 2024b. Wavtokenizer: an efficient acoustic discrete codec tokenizer for audio language modeling. arXiv preprint arXiv:2408.16532 (2024)."},{"key":"e_1_3_2_1_25_1","unstructured":"Zeqian Ju Yuancheng Wang Kai Shen Xu Tan Detai Xin Dongchao Yang Eric Liu Yichong Leng Kaitao Song Siliang Tang Zhizheng Wu Tao Qin Xiangyang Li Wei Ye Shikun Zhang Jiang Bian Lei He Jinyu Li and Sheng Zhao. 2024. NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models. In ICML."},{"key":"e_1_3_2_1_26_1","unstructured":"Jaehyeon Kim Sungwon Kim Jungil Kong and Sungroh Yoon. 2020. Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search. In NeurIPS."},{"key":"e_1_3_2_1_27_1","volume-title":"From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech. arXiv preprint arXiv:2503.16956","author":"Kim Ji-Hoon","year":"2025","unstructured":"Ji-Hoon Kim, Jeongsoo Choi, Jaehun Kim, Chaeyoung Jung, and Joon Son Chung. 2025. From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech. arXiv preprint arXiv:2503.16956 (2025)."},{"key":"e_1_3_2_1_28_1","volume-title":"Evelina Bakhturina, Mikyas Desta, Rafael Valle, Sungroh Yoon, and Bryan Catanzaro.","author":"Kim Sungwon","year":"2023","unstructured":"Sungwon Kim, Kevin J. Shih, Rohan Badlani, Jo a, o Felipe Santos, Evelina Bakhturina, Mikyas Desta, Rafael Valle, Sungroh Yoon, and Bryan Catanzaro. 2023. P-Flow: A Fast and Data-Efficient Zero-Shot TTS through Speech Prompting. In NeurIPS."},{"key":"e_1_3_2_1_29_1","first-page":"17022","article-title":"HiFi-GAN","author":"Kong Jungil","year":"2020","unstructured":"Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020. HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis. In NIPS. 17022-17033.","journal-title":"In NIPS."},{"key":"e_1_3_2_1_30_1","unstructured":"Rithesh Kumar Prem Seetharaman Alejandro Luebs Ishaan Kumar and Kundan Kumar. 2023. High-fidelity audio compression with improved rvqgan. In NeurIPS."},{"key":"e_1_3_2_1_31_1","volume-title":"Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale. In NeurIPS.","author":"Le Matthew","year":"2023","unstructured":"Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, and Wei-Ning Hsu. 2023. Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale. In NeurIPS."},{"key":"e_1_3_2_1_32_1","volume-title":"Joon Son Chung, and Soo-Whan Chung","author":"Lee Jiyoung","year":"2023","unstructured":"Jiyoung Lee, Joon Son Chung, and Soo-Whan Chung. 2023a. Imaginary Voice: Face-Styled Diffusion Model for Text-to-Speech. In ICASSP. 1-5."},{"key":"e_1_3_2_1_33_1","unstructured":"Sang-gil Lee Wei Ping Boris Ginsburg Bryan Catanzaro and Sungroh Yoon. 2023b. BigVGAN: A Universal Neural Vocoder with Large-Scale Training. In ICLR."},{"key":"e_1_3_2_1_34_1","first-page":"4626","article-title":"Frame-Level Signal-to-Noise Ratio Estimation Using Deep Learning","author":"Li Hao","year":"2020","unstructured":"Hao Li, DeLiang Wang, Xueliang Zhang, and Guanglai Gao. 2020. Frame-Level Signal-to-Noise Ratio Estimation Using Deep Learning.. In Interspeech. 4626-4630.","journal-title":"Interspeech."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2025.3597267"},{"key":"e_1_3_2_1_36_1","first-page":"1","article-title":"FCConDubber: Fine And Coarse Grained Prosody Alignment For Expressive Video Dubbing via Contrastive Audio-Motion Pretraining","author":"Li Qiulin","year":"2025","unstructured":"Qiulin Li, Zhichao Wu, Hanwei Li, Xin Dong, and Qun Yang. 2025b. FCConDubber: Fine And Coarse Grained Prosody Alignment For Expressive Video Dubbing via Contrastive Audio-Motion Pretraining. In ICASSP. 1-5.","journal-title":"ICASSP."},{"key":"e_1_3_2_1_37_1","unstructured":"Yinghao Aaron Li Cong Han Vinay S. Raghavan Gavin Mischler and Nima Mesgarani. 2023. StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models. In NeurIPS."},{"key":"e_1_3_2_1_38_1","volume-title":"Heli Ben-Hamu, Maximilian Nickel, and Matt Le.","author":"Lipman Yaron","year":"2022","unstructured":"Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2022. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747 (2022)."},{"key":"e_1_3_2_1_39_1","first-page":"3003","volume-title":"IEEE PAMI","volume":"45","author":"Liu Xuejing","year":"2023","unstructured":"Xuejing Liu, Liang Li, Shuhui Wang, Zheng-Jun Zha, Zechao Li, Qi Tian, and Qingming Huang. 2023. Entity-Enhanced Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding. IEEE PAMI, Vol. 45, 3 (2023), 3003-3018."},{"key":"e_1_3_2_1_40_1","first-page":"11966","article-title":"A ConvNet for the 2020s","author":"Liu Zhuang","year":"2022","unstructured":"Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. 2022. A ConvNet for the 2020s. In CVPR. 11966-11976.","journal-title":"CVPR."},{"key":"e_1_3_2_1_41_1","first-page":"8032","article-title":"Visualtts","author":"Lu Junchen","year":"2022","unstructured":"Junchen Lu, Berrak Sisman, Rui Liu, Mingyang Zhang, and Haizhou Li. 2022. Visualtts: TTS with Accurate Lip-Speech Synchronization for Automatic Voice Over. In ICASSP. 8032-8036.","journal-title":"In ICASSP."},{"key":"e_1_3_2_1_42_1","first-page":"7608","article-title":"Towards Practical Lipreading with Distilled and Efficient Models","author":"Ma Pingchuan","year":"2021","unstructured":"Pingchuan Ma, Brais Martinez, Stavros Petridis, and Maja Pantic. 2021. Towards Practical Lipreading with Distilled and Efficient Models. In ICASSP. 7608-7612.","journal-title":"ICASSP."},{"key":"e_1_3_2_1_43_1","first-page":"6319","article-title":"Lipreading Using Temporal Convolutional Networks","author":"Martinez Brais","year":"2020","unstructured":"Brais Martinez, Pingchuan Ma, Stavros Petridis, and Maja Pantic. 2020. Lipreading Using Temporal Convolutional Networks. In ICASSP. 6319-6323.","journal-title":"ICASSP."},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-1386"},{"key":"e_1_3_2_1_45_1","first-page":"11341","article-title":"Matcha-TTS: A fast TTS architecture with conditional flow matching","author":"Mehta Shivam","year":"2024","unstructured":"Shivam Mehta, Ruibo Tu, Jonas Beskow, \u00c9va Sz\u00e9kely, and Gustav Eje Henter. 2024. Matcha-TTS: A fast TTS architecture with conditional flow matching. In ICASSP. 11341-11345.","journal-title":"ICASSP."},{"key":"e_1_3_2_1_46_1","unstructured":"Lingwei Meng Long Zhou Shujie Liu Sanyuan Chen Bing Han Shujie Hu Yanqing Liu Jinyu Li Sheng Zhao Xixin Wu et al. 2024. Autoregressive speech synthesis without vector quantization. arXiv preprint arXiv:2407.08551 (2024)."},{"key":"e_1_3_2_1_47_1","unstructured":"Fabian Mentzer David Minnen Eirikur Agustsson and Michael Tschannen. 2024. Finite Scalar Quantization: VQ-VAE Made Simple. In ICLR."},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2004-668"},{"key":"e_1_3_2_1_49_1","first-page":"13793","article-title":"Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis","author":"Prajwal K. R.","year":"2020","unstructured":"K. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, and C. V. Jawahar. 2020. Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis. In CVPR. 13793-13802.","journal-title":"CVPR."},{"key":"e_1_3_2_1_50_1","first-page":"28492","article-title":"Robust Speech Recognition via Large-Scale Weak Supervision","author":"Radford Alec","year":"2023","unstructured":"Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust Speech Recognition via Large-Scale Weak Supervision. In ICML. 28492-28518.","journal-title":"ICML."},{"key":"e_1_3_2_1_51_1","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei Ilya Sutskever et al. 2019. Language models are unsupervised multitask learners. OpenAI blog Vol. 1 8 (2019) 9."},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414878"},{"key":"e_1_3_2_1_53_1","unstructured":"Yi Ren Chenxu Hu Xu Tan Tao Qin Sheng Zhao Zhou Zhao and Tie-Yan Liu. 2021. FastSpeech 2: Fast and High-Quality End-to-End Text to Speech. In ICLR."},{"key":"e_1_3_2_1_54_1","volume-title":"Utmos: Utokyo-sarulab system for voicemos challenge","author":"Saeki Takaaki","year":"2022","unstructured":"Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari. 2022. Utmos: Utokyo-sarulab system for voicemos challenge 2022. arXiv preprint arXiv:2204.02152 (2022)."},{"key":"e_1_3_2_1_55_1","volume-title":"Maximum likelihood training of score-based diffusion models. Advances in neural information processing systems","author":"Song Yang","year":"2021","unstructured":"Yang Song, Conor Durkan, Iain Murray, and Stefano Ermon. 2021. Maximum likelihood training of score-based diffusion models. Advances in neural information processing systems, Vol. 34 (2021), 1415-1428."},{"key":"e_1_3_2_1_56_1","volume-title":"Tae-Hyun Oh, and David Harwath.","author":"Sung-Bin Kim","year":"2025","unstructured":"Kim Sung-Bin, Jeongsoo Choi, Puyuan Peng, Joon Son Chung, Tae-Hyun Oh, and David Harwath. 2025. VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models. arXiv preprint arXiv:2504.02386 (2025)."},{"key":"e_1_3_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3339628"},{"key":"e_1_3_2_1_58_1","volume-title":"Visual Position Prompt for MLLM based Visual Grounding","author":"Tang Wei","year":"2025","unstructured":"Wei Tang, Yanpeng Sun, Qinying Gu, and Zechao Li. 2025. Visual Position Prompt for MLLM based Visual Grounding. IEEE Trans. Multimedia (2025)."},{"key":"e_1_3_2_1_59_1","unstructured":"Yuandong Tian. 2022. Understanding Deep Contrastive Learning via Coordinate-wise Optimization. In NeurIPS."},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3365104"},{"key":"e_1_3_2_1_61_1","first-page":"6306","article-title":"Neural Discrete Representation Learning","author":"van den Oord A\u00e4ron","year":"2017","unstructured":"A\u00e4ron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017. Neural Discrete Representation Learning. In NeurIPS. 6306-6315.","journal-title":"NeurIPS."},{"key":"e_1_3_2_1_62_1","unstructured":"Chengyi Wang Sanyuan Chen Yu Wu Ziqiang Zhang Long Zhou Shujie Liu Zhuo Chen Yanqing Liu Huaming Wang Jinyu Li et al. 2023a. Neural codec language models are zero-shot text to speech synthesizers. arXiv preprint arXiv:2301.02111 (2023)."},{"key":"e_1_3_2_1_63_1","first-page":"14653","article-title":"Seeing What You Said: Talking Face Generation Guided by a Lip Reading Expert","author":"Wang Jiadong","year":"2023","unstructured":"Jiadong Wang, Xinyuan Qian, Malu Zhang, Robby T. Tan, and Haizhou Li. 2023b. Seeing What You Said: Talking Face Generation Guided by a Lip Reading Expert. In CVPR. 14653-14662.","journal-title":"CVPR."},{"key":"e_1_3_2_1_64_1","volume-title":"Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens. arXiv preprint arXiv:2503.01710","author":"Wang Xinsheng","year":"2025","unstructured":"Xinsheng Wang, Mingqi Jiang, Ziyang Ma, Ziyu Zhang, Songxiang Liu, Linqin Li, Zheng Liang, Qixi Zheng, Rui Wang, Xiaoqin Feng, et al., 2025. Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens. arXiv preprint arXiv:2503.01710 (2025)."},{"key":"e_1_3_2_1_65_1","volume-title":"Maskgct: Zero-shot text-to-speech with masked generative codec transformer. arXiv preprint arXiv:2409.00750","author":"Wang Yuancheng","year":"2024","unstructured":"Yuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng, Haotian Guo, Jiachen Zheng, Qiang Zhang, Xueyao Zhang, Shunsi Zhang, and Zhizheng Wu. 2024. Maskgct: Zero-shot text-to-speech with masked generative codec transformer. arXiv preprint arXiv:2409.00750 (2024)."},{"key":"e_1_3_2_1_66_1","unstructured":"An Yang Baosong Yang Beichen Zhang Binyuan Hui Bo Zheng Bowen Yu Chengyuan Li Dayiheng Liu Fei Huang Haoran Wei et al. 2024. Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115 (2024)."},{"key":"e_1_3_2_1_67_1","volume-title":"Stablevc: Style controllable zero-shot voice conversion with conditional flow matching. arXiv preprint arXiv:2412.04724","author":"Yao Jixun","year":"2024","unstructured":"Jixun Yao, Yuguang Yang, Yu Pan, Ziqian Ning, Jiaohao Ye, Hongbin Zhou, and Lei Xie. 2024. Stablevc: Style controllable zero-shot voice conversion with conditional flow matching. arXiv preprint arXiv:2412.04724 (2024)."},{"key":"e_1_3_2_1_68_1","volume-title":"arXiv preprint arXiv:2502.01046","author":"Ye Jiaxin","year":"2025","unstructured":"Jiaxin Ye, Boyuan Cao, and Hongming Shan. 2025a. Emotional Face-to-Speech. arXiv preprint arXiv:2502.01046 (2025)."},{"key":"e_1_3_2_1_69_1","volume-title":"arXiv preprint arXiv:2503.14928","author":"Ye Jiaxin","year":"2025","unstructured":"Jiaxin Ye and Hongming Shan. 2025. Shushing! Let's Imagine an Authentic Speech from the Silent Video. arXiv preprint arXiv:2503.14928 (2025)."},{"key":"e_1_3_2_1_70_1","first-page":"1","article-title":"Temporal Modeling Matters","author":"Ye Jiaxin","year":"2023","unstructured":"Jiaxin Ye, Xin-Cheng Wen, Yujie Wei, Yong Xu, Kunhong Liu, and Hongming Shan. 2023. Temporal Modeling Matters: A Novel Temporal Emotional Modeling Approach for Speech Emotion Recognition. In ICASSP. 1-5.","journal-title":"In ICASSP."},{"key":"e_1_3_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1049\/cje.2021.00.455"},{"key":"e_1_3_2_1_72_1","unstructured":"Zhen Ye Peiwen Sun Jiahe Lei Hongzhan Lin Xu Tan Zheqi Dai Qiuqiang Kong Jianyi Chen Jiahao Pan Qifeng Liu et al. 2024. Codec does matter: Exploring the semantic shortcoming of codec for audio language model. arXiv preprint arXiv:2408.17175 (2024)."},{"key":"e_1_3_2_1_73_1","volume-title":"Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis. arXiv preprint arXiv:2502.04128","author":"Ye Zhen","year":"2025","unstructured":"Zhen Ye, Xinfa Zhu, Chi-Min Chan, Xinsheng Wang, Xu Tan, Jiahe Lei, Yi Peng, Haohe Liu, Yizhu Jin, Zheqi DAI, et al., 2025b. Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis. arXiv preprint arXiv:2502.04128 (2025)."},{"key":"e_1_3_2_1_74_1","unstructured":"Yochai Yemini Aviv Shamsian Lior Bracha Sharon Gannot and Ethan Fetaya. 2024. LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading. In ICLR."},{"key":"e_1_3_2_1_75_1","first-page":"1526","article-title":"LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech","author":"Zen Heiga","year":"2019","unstructured":"Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. 2019. LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech. In Interspeech. 1526-1530.","journal-title":"Interspeech."},{"key":"e_1_3_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3432099"},{"key":"e_1_3_2_1_77_1","volume-title":"DeepAudio-V1: Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation. arXiv preprint arXiv:2503.22265","author":"Zhang Haomin","year":"2025","unstructured":"Haomin Zhang, Chang Liu, Junjie Zheng, Zihao Chen, Chaofan Ding, and Xinhan Di. 2025b. DeepAudio-V1: Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation. arXiv preprint arXiv:2503.22265 (2025)."},{"key":"e_1_3_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.23919\/cje.2022.00.414"},{"key":"e_1_3_2_1_79_1","first-page":"12251","article-title":"Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment","author":"Zhang Xueyao","year":"2025","unstructured":"Xueyao Zhang, Yuancheng Wang, Chaoren Wang, Ziniu Li, Zhuo Chen, and Zhizheng Wu. 2025c. Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment. In ACL. 12251-12270.","journal-title":"ACL."},{"key":"e_1_3_2_1_80_1","doi-asserted-by":"crossref","unstructured":"Zhedong Zhang Liang Li Gaoxiang Cong YIN Haibing Yuhan Gao Chenggang Yan Anton van den Hengel and Yuankai Qi. 2024b. From Speaker to Dubber: Movie Dubbing with Prosody and Duration Consistency Learning. In ACM MM.","DOI":"10.1145\/3664647.3680777"},{"key":"e_1_3_2_1_81_1","volume-title":"Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing. arXiv preprint arXiv:2503.12042","author":"Zhang Zhedong","year":"2025","unstructured":"Zhedong Zhang, Liang Li, Chenggang Yan, Chunshan Liu, Anton van den Hengel, and Yuankai Qi. 2025a. Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing. arXiv preprint arXiv:2503.12042 (2025)."},{"key":"e_1_3_2_1_82_1","first-page":"332","article-title":"Generating High-Quality Symbolic Music Using Fine-Grained Discriminators","author":"Zhang Zhedong","year":"2024","unstructured":"Zhedong Zhang, Liang Li, Jiehua Zhang, Zhenghui Hu, Hongkui Wang, Chenggang Yan, Jian Yang, and Yuankai Qi. 2024d. Generating High-Quality Symbolic Music Using Fine-Grained Discriminators. In ICPR. 332-344.","journal-title":"ICPR."},{"key":"e_1_3_2_1_83_1","volume-title":"MCDubber: Multimodal Context-Aware Expressive Video Dubbing. arXiv preprint arXiv:2408.11593","author":"Zhao Yuan","year":"2024","unstructured":"Yuan Zhao, Zhenqi Jia, Rui Liu, De Hu, Feilong Bao, and Guanglai Gao. 2024. MCDubber: Multimodal Context-Aware Expressive Video Dubbing. arXiv preprint arXiv:2408.11593 (2024)."},{"key":"e_1_3_2_1_84_1","volume-title":"Narration, Monologue Adaptive Movie Dubbing Via Multi-Modal Chain-of-Thoughts Reasoning Guidance. arXiv preprint arXiv:2503.23660","author":"Zheng Junjie","year":"2025","unstructured":"Junjie Zheng, Zihao Chen, Chaofan Ding, and Xinhan Di. 2025. DeepDubber-V1: Towards High Quality and Dialogue, Narration, Monologue Adaptive Movie Dubbing Via Multi-Modal Chain-of-Thoughts Reasoning Guidance. arXiv preprint arXiv:2503.23660 (2025)."},{"key":"e_1_3_2_1_85_1","volume-title":"Autoregressive Speech Synthesis with Next-Distribution Prediction. arXiv preprint arXiv:2412.16846","author":"Zhu Xinfa","year":"2024","unstructured":"Xinfa Zhu, Wenjie Tian, and Lei Xie. 2024. Autoregressive Speech Synthesis with Next-Distribution Prediction. arXiv preprint arXiv:2412.16846 (2024)."}],"event":{"name":"MM '25: The 33rd ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Dublin Ireland","acronym":"MM '25"},"container-title":["Proceedings of the 33rd ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746027.3754734","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,10]],"date-time":"2025-12-10T04:05:44Z","timestamp":1765339544000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746027.3754734"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,27]]},"references-count":85,"alternative-id":["10.1145\/3746027.3754734","10.1145\/3746027"],"URL":"https:\/\/doi.org\/10.1145\/3746027.3754734","relation":{},"subject":[],"published":{"date-parts":[[2025,10,27]]},"assertion":[{"value":"2025-10-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}