{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,10]],"date-time":"2026-01-10T05:33:01Z","timestamp":1768023181542,"version":"3.49.0"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2025,3,11]],"date-time":"2025-03-11T00:00:00Z","timestamp":1741651200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Nature Science Foundation of China","doi-asserted-by":"crossref","award":["U23B2053 and 62301521"],"award-info":[{"award-number":["U23B2053 and 62301521"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"the Anhui Provincial Natural Science Foundation","award":["2308085QF200"],"award-info":[{"award-number":["2308085QF200"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>\n            Recently, fine-grained prosody representations have emerged and attracted growing attention to address the one-to-many problem in text-to-speech (TTS). In this article, we propose the PhonemeVec, a pre-trained prosody representations with considering the contextual information. To obtain the contextual prosody representations, we improve the data2vec framework according to the characteristics of prosody to extract the PhonemeVec from the low-band mel-spectrogram, and pre-train on a 960 hours Chinese corpus with high quality and diverse pronunciation. PhonemeVec is subsequently integrated into FastSpeech2, supervising the prosody modeling of the text encoder. Experiments conducted on the Blizzard Challenge 2019 dataset show that the integration of PhonemeVec results in the synthesis of more natural speech. Additionally, objective evaluations confirm that the application of PhonemeVec reduces the distortions between the generated speech and original recordings in terms of duration and F0. Audio samples can be found at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"http:\/\/home.ustc.edu.cn\/~wsmzzz\/PhonemeVec\/demo.html\">http:\/\/home.ustc.edu.cn\/~wsmzzz\/PhonemeVec\/demo.html<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3711828","type":"journal-article","created":{"date-parts":[[2025,1,15]],"date-time":"2025-01-15T10:56:17Z","timestamp":1736938577000},"page":"1-15","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["PhonemeVec: A Phoneme-Level Contextual Prosody Representation For Speech Synthesis"],"prefix":"10.1145","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-8004-937X","authenticated-orcid":false,"given":"Shiming","family":"Wang","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9421-1049","authenticated-orcid":false,"given":"Li-Ping","family":"Chen","sequence":"additional","affiliation":[{"name":"Dept. of Electronic Engineering and Information Science, University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6668-022X","authenticated-orcid":false,"given":"Yang","family":"Ai","sequence":"additional","affiliation":[{"name":"Electronic Engineering and Information Science, University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-9626-456X","authenticated-orcid":false,"given":"Yajun","family":"Hu","sequence":"additional","affiliation":[{"name":"iFlytek Co Ltd, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7853-5273","authenticated-orcid":false,"given":"Zhen-Hua","family":"Ling","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,3,11]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33016706"},{"key":"e_1_3_2_3_2","unstructured":"Yi Ren Yangjun Ruan Xu Tan Tao Qin Sheng Zhao Zhou Zhao and Tie-Yan Liu. 2019. Fastspeech: Fast robust and controllable text to speech. In Proc. NeurIPS."},{"key":"e_1_3_2_4_2","volume-title":"Proceedings of the ICLR","author":"Ren Yi","year":"2019","unstructured":"Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2019. FastSpeech 2: Fast and high-quality end-to-end text to speech. In Proceedings of the ICLR."},{"key":"e_1_3_2_5_2","unstructured":"Jaehyeon Kim Sungwon Kim Jungil Kong and Sungroh Yoon. 2020. Glow-TTS: A generative flow for text-to-speech via monotonic alignment search. In Proceedings of the NeurIPS. 8067\u20138077."},{"key":"e_1_3_2_6_2","first-page":"5530","volume-title":"Proceedings of the ICML","author":"Kim Jaehyeon","year":"2021","unstructured":"Jaehyeon Kim, Jungil Kong, and Juhee Son. 2021. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. In Proceedings of the ICML. 5530\u20135540."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053795"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3356232"},{"key":"e_1_3_2_9_2","unstructured":"Jungil Kong Jaehyeon Kim and Jaekyoung Bae. 2020. Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis. In Proceedings of the NeurIPS. 17022\u201317033."},{"key":"e_1_3_2_10_2","unstructured":"Xu Tan Tao Qin Frank Soong and Tie-Yan Liu. 2021. A survey on neural speech synthesis. arXiv:2106.15561. Retrieved from https:\/\/arxiv.org\/abs\/2106.15561"},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Chenpeng Du and Kai Yu. 2021. Phone-level prosody modelling with GMM-based MDN for diverse and controllable speech synthesis. IEEE\/ACM Transactions on Audio Speech and Language Processing 30 (2021) 190\u2013201.","DOI":"10.1109\/TASLP.2021.3133205"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.518"},{"key":"e_1_3_2_13_2","unstructured":"Zack Hodari. 2022. Synthesising Prosody with Insufficient Context. The University of Edinburgh."},{"key":"e_1_3_2_14_2","first-page":"11134","volume-title":"Proceedings of the ICML","author":"Weston Jack","year":"2021","unstructured":"Jack Weston, Raphael Lenain, Udeepa Meepegama, and Emil Fristed. 2021. Learning de-identified representations of prosody from raw audio. In Proceedings of the ICML. 11134\u201311145."},{"key":"e_1_3_2_15_2","doi-asserted-by":"crossref","unstructured":"Jennifer Cole. 2015. Prosody in context: A review. Language Cognition and Neuroscience 30 1\u20132 (2015) 1\u201331.","DOI":"10.1080\/23273798.2014.963130"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9413889"},{"key":"e_1_3_2_17_2","first-page":"5180","volume-title":"Proceedings of the ICML","author":"Wang Yuxuan","year":"2018","unstructured":"Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ-Skerry Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A. Saurous. 2018. Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis. In Proceedings of the ICML. 5180\u20135189."},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053520"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053436"},{"key":"e_1_3_2_20_2","unstructured":"Chenpeng Du and Kai Yu. 2021. Rich prosody diversity modelling with phone-level mixture density network. arXiv:2102.00851. Retrieved from https:\/\/arxiv.org\/abs\/2102.00851"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9746744"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9746883"},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","unstructured":"Zhao-Ci Liu Liping Chen Ya-Jun Hu Zhen-Hua Ling and Jia Pan. 2024. PE-wav2vec: A prosody-enhanced speech model for self-supervised prosody learning in TTS. IEEE\/ACM Transactions on Audio Speech and Language Processing 32 (2024).","DOI":"10.1109\/TASLP.2024.3449148"},{"key":"e_1_3_2_24_2","unstructured":"Alexei Baevski Yuhao Zhou Abdelrahman Mohamed and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Proceedings of the NeurIPS. 12449\u201312460."},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","unstructured":"Wei-Ning Hsu Benjamin Bolte Yao-Hung Hubert Tsai Kushal Lakhotia Ruslan Salakhutdinov and Abdelrahman Mohamed. 2021. HuBERT: Self-supervised speech representation learning by masked prediction of hidden units. IEEE\/ACM Transactions on Audio Speech and Language Processing 29 (2021) 3451\u20133460.","DOI":"10.1109\/TASLP.2021.3122291"},{"key":"e_1_3_2_26_2","first-page":"1298","volume-title":"Proceedings of the ICML","author":"Baevski Alexei","year":"2022","unstructured":"Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. 2022. data2vec: A general framework for self-supervised learning in speech, vision and language. In Proceedings of the ICML. 1298\u20131312."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2021-475"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2023.3288409"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00618"},{"key":"e_1_3_2_30_2","unstructured":"George Papamakarios Theo Pavlakou and Iain Murray. 2017. Masked autoregressive flow for density estimation. In Proceedings of the NeurIPS."},{"key":"e_1_3_2_31_2","first-page":"4693","volume-title":"Proceedings of the ICML","author":"Skerry-Ryan R. J.","year":"2018","unstructured":"R. J. Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron Weiss, Rob Clark, and Rif A. Saurous. 2018. Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron. In Proceedings of the ICML. 4693\u20134702."},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683623"},{"key":"e_1_3_2_33_2","unstructured":"Ziyue Jiang Yi Ren Zhenhui Ye Jinglin Liu Chen Zhang Qian Yang Shengpeng Ji Rongjie Huang Chunfeng Wang Xiang Yin et\u00a0al. 2023. Mega-TTS: Zero-shot text-to-speech at scale with intrinsic inductive bias. arXiv:2306.03509. Retrieved from https:\/\/arxiv.org\/abs\/2306.03509"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-2477"},{"key":"e_1_3_2_35_2","volume-title":"Proceedings of the ICLR","author":"Jiang Ziyue","year":"2024","unstructured":"Ziyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He, Zhenhui Ye, Shengpeng Ji, Qian Yang, Chen Zhang, Pengfei Wei, Chunfeng Wang, et\u00a0al. 2024. Mega-TTS 2: Boosting prompting mechanisms for zero-shot speech synthesis. In Proceedings of the ICLR."},{"key":"e_1_3_2_36_2","first-page":"4171","volume-title":"Proceedings of the NAACL-HLT","author":"Kenton Jacob Devlin Ming-Wei Chang","year":"2019","unstructured":"Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the NAACL-HLT. 4171\u20134186."},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","unstructured":"Steffen Schneider Alexei Baevski Ronan Collobert and Michael Auli. 2019. wav2vec: Unsupervised pre-training for speech recognition. In Proceedings of the Interspeech. 3465\u20133469.","DOI":"10.21437\/Interspeech.2019-1873"},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","unstructured":"Shu-wen Yang Heng-Jui Chang Zili Huang Andy T. Liu Cheng-I. Lai Haibin Wu Jiatong Shi Xuankai Chang Hsiang-Sheng Tsai Wen-Chin Huang and others. 2024. A large-scale evaluation of speech foundation models. IEEE\/ACM Transactions on Audio Speech and Language Processing 3 (2024).","DOI":"10.1109\/TASLP.2024.3389631"},{"key":"e_1_3_2_39_2","doi-asserted-by":"crossref","unstructured":"Abdelrahman Mohamed Hung-yi Lee Lasse Borgholt Jakob D. Havtorn Joakim Edin Christian Igel Katrin Kirchhoff Shang-Wen Li Karen Livescu Lars Maal\u00f8e and others. 2022. Self-supervised speech representation learning: A review. IEEE Journal of Selected Topics in Signal Processing 16 6 (2022) 1179\u20131210.","DOI":"10.1109\/JSTSP.2022.3207050"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2023-1087"},{"key":"e_1_3_2_41_2","unstructured":"Iz Beltagy Matthew E. Peters and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv:2004.05150. Retrieved from https:\/\/arxiv.org\/abs\/2004.05150"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2021-1757"},{"key":"e_1_3_2_43_2","unstructured":"Aaron van den Oord Sander Dieleman Heiga Zen Karen Simonyan Oriol Vinyals Alex Graves Nal Kalchbrenner Andrew Senior and Koray Kavukcuoglu. 2016. Wavenet: A generative model for raw audio. arXiv:1609.03499. Retrieved from https:\/\/arxiv.org\/abs\/1609.03499"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.21437\/Blizzard.2019-1"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-4009"},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","unstructured":"Masanori Morise Fumiya Yokomori and Kenji Ozawa. 2016. WORLD: A vocoder-based high-quality speech synthesis system for real-time applications. IEIC TRANSACTIONS on Information and Systems 99 (2016) 1877\u20131884.","DOI":"10.1587\/transinf.2015EDP7457"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3711828","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3711828","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:19:15Z","timestamp":1750295955000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3711828"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,11]]},"references-count":45,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3711828"],"URL":"https:\/\/doi.org\/10.1145\/3711828","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,11]]},"assertion":[{"value":"2024-06-10","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-16","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}