{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T22:04:23Z","timestamp":1774476263698,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":26,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,2,25]],"date-time":"2022-02-25T00:00:00Z","timestamp":1645747200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key R&D Program of China","award":["2019YFF0303001"],"award-info":[{"award-number":["2019YFF0303001"]}]},{"name":"National Nature Science Foundation of China","award":["61871358"],"award-info":[{"award-number":["61871358"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,2,25]]},"DOI":"10.1145\/3529570.3529602","type":"proceedings-article","created":{"date-parts":[[2022,6,29]],"date-time":"2022-06-29T22:15:36Z","timestamp":1656540936000},"page":"187-193","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Speaker Adaption with Intuitive Prosodic Features for Statistical Parametric Speech Synthesis"],"prefix":"10.1145","author":[{"given":"Pengyu","family":"Cheng","sequence":"first","affiliation":[{"name":"National Engineering Laboratory for Speech and Language Information Processing, University of Science and Technology of China, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhenhua","family":"Ling","sequence":"additional","affiliation":[{"name":"National Engineering Laboratory for Speech and Language Information Processing, University of Science and Technology of China, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,6,29]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"[n. d.]. aidatatang 200zh. https:\/\/openslr.org\/62\/.  [n. d.]. aidatatang 200zh. https:\/\/openslr.org\/62\/."},{"key":"e_1_3_2_1_2_1","volume-title":"d.]. Magic Data Technology Co","unstructured":"[n. d.]. Magic Data Technology Co ., Ltd .https:\/\/openslr.com\/68\/. [n. d.]. Magic Data Technology Co., Ltd.https:\/\/openslr.com\/68\/."},{"key":"e_1_3_2_1_3_1","unstructured":"Sercan\u00a0O Arik Jitong Chen Kainan Peng Wei Ping and Yanqi Zhou. 2018. Neural voice cloning with a few samples. arXiv preprint arXiv:1802.06006(2018).  Sercan\u00a0O Arik Jitong Chen Kainan Peng Wei Ping and Yanqi Zhou. 2018. Neural voice cloning with a few samples. arXiv preprint arXiv:1802.06006(2018)."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2020.2980991"},{"key":"e_1_3_2_1_5_1","volume-title":"Adaspeech: Adaptive text to speech for custom voice. In ICLR.","author":"Chen Mingjian","year":"2021","unstructured":"Mingjian Chen , Xu Tan , Bohan Li , Yanqing Liu , Tao Qin , Sheng Zhao , and Tie-Yan Liu . 2021 . Adaspeech: Adaptive text to speech for custom voice. In ICLR. Mingjian Chen, Xu Tan, Bohan Li, Yanqing Liu, Tao Qin, Sheng Zhao, and Tie-Yan Liu. 2021. Adaspeech: Adaptive text to speech for custom voice. In ICLR."},{"key":"e_1_3_2_1_6_1","unstructured":"Yutian Chen Yannis Assael Brendan Shillingford David Budden Scott Reed Heiga Zen Quan Wang Luis\u00a0C Cobo Andrew Trask Ben Laurie 2018. Sample efficient adaptive text-to-speech. arXiv preprint arXiv:1809.10460(2018).  Yutian Chen Yannis Assael Brendan Shillingford David Budden Scott Reed Heiga Zen Quan Wang Luis\u00a0C Cobo Andrew Trask Ben Laurie 2018. Sample efficient adaptive text-to-speech. arXiv preprint arXiv:1809.10460(2018)."},{"key":"e_1_3_2_1_7_1","volume-title":"Attentron: Few-Shot Text-to-Speech Utilizing Attention-Based Variable-Length Embedding. In Interspeech.","author":"Choi Seungwoo","year":"2020","unstructured":"Seungwoo Choi , Seungju Han , Dongyoung Kim , and Sungjoo Ha . 2020 . Attentron: Few-Shot Text-to-Speech Utilizing Attention-Based Variable-Length Embedding. In Interspeech. Seungwoo Choi, Seungju Han, Dongyoung Kim, and Sungjoo Ha. 2020. Attentron: Few-Shot Text-to-Speech Utilizing Attention-Based Variable-Length Embedding. In Interspeech."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9054535"},{"key":"e_1_3_2_1_9_1","unstructured":"Yan Deng Lei He and Frank Soong. 2018. Modeling multi-speaker latent space to improve neural tts: Quick enrolling new speaker and enhancing premium voice. arXiv preprint arXiv:1812.05253(2018).  Yan Deng Lei He and Frank Soong. 2018. Modeling multi-speaker latent space to improve neural tts: Quick enrolling new speaker and enhancing premium voice. arXiv preprint arXiv:1812.05253(2018)."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.21437\/SSW.2019-5"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Yan Huang Lei He Wenning Wei William Gale Jinyu Li and Yifan Gong. 2020. Using Personalized Speech Synthesis and Neural Language Generator for Rapid Speaker Adaptation. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). 7399\u20137403. https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9053104  Yan Huang Lei He Wenning Wei William Gale Jinyu Li and Yifan Gong. 2020. Using Personalized Speech Synthesis and Neural Language Generator for Rapid Speaker Adaptation. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). 7399\u20137403. https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9053104","DOI":"10.1109\/ICASSP40776.2020.9053104"},{"key":"e_1_3_2_1_12_1","first-page":"4485","article-title":"Transfer learning from speaker verification to multispeaker text-to-speech synthesis","volume":"31","author":"Jia Ye","year":"2018","unstructured":"Ye Jia , Yu Zhang , Ron\u00a0 J Weiss , Quan Wang , Jonathan Shen , Fei Ren , Zhifeng Chen , Patrick Nguyen , Ruoming Pang , Ignacio\u00a0Lopez Moreno , 2018 . Transfer learning from speaker verification to multispeaker text-to-speech synthesis . Advances in Neural Information Processing Systems 31 (2018), 4485 \u2013 4495 . Ye Jia, Yu Zhang, Ron\u00a0J Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio\u00a0Lopez Moreno, 2018. Transfer learning from speaker verification to multispeaker text-to-speech synthesis. Advances in Neural Information Processing Systems 31 (2018), 4485\u20134495.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Zvi Kons Slava Shechtman Alex Sorin Carmel Rabinovitz and Ron Hoory. 2019. High quality lightweight and adaptable TTS using LPCNet. In Interspeech. 176\u2013180.  Zvi Kons Slava Shechtman Alex Sorin Carmel Rabinovitz and Ron Hoory. 2019. High quality lightweight and adaptable TTS using LPCNet. In Interspeech. 176\u2013180.","DOI":"10.21437\/Interspeech.2019-1705"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACRIM.1993.407206"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33016706"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Michael McAuliffe Michaela Socolof Sarah Mihuc Michael Wagner and Morgan Sonderegger. 2017. Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi.. In Interspeech Vol.\u00a02017. 498\u2013502.  Michael McAuliffe Michaela Socolof Sarah Mihuc Michael Wagner and Morgan Sonderegger. 2017. Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi.. In Interspeech Vol.\u00a02017. 498\u2013502.","DOI":"10.21437\/Interspeech.2017-1386"},{"key":"e_1_3_2_1_17_1","volume-title":"International Conference on Machine Learning. PMLR, 3683\u20133691","author":"Nachmani Eliya","year":"2018","unstructured":"Eliya Nachmani , Adam Polyak , Yaniv Taigman , and Lior Wolf . 2018 . Fitting new speakers based on a short untranscribed sample . In International Conference on Machine Learning. PMLR, 3683\u20133691 . Eliya Nachmani, Adam Polyak, Yaniv Taigman, and Lior Wolf. 2018. Fitting new speakers based on a short untranscribed sample. In International Conference on Machine Learning. PMLR, 3683\u20133691."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"crossref","unstructured":"Tuomo Raitio Ramya Rasipuram and Dan Castellani. 2020. Controllable neural text-to-speech synthesis using intuitive prosodic features. In Interspeech.  Tuomo Raitio Ramya Rasipuram and Dan Castellani. 2020. Controllable neural text-to-speech synthesis using intuitive prosodic features. In Interspeech.","DOI":"10.21437\/Interspeech.2020-2861"},{"key":"e_1_3_2_1_19_1","unstructured":"Yi Ren Chenxu Hu Xu Tan Tao Qin Sheng Zhao Zhou Zhao and Tie-Yan Liu. 2021. Fastspeech 2: Fast and high-quality end-to-end text to speech. In ICLR.  Yi Ren Chenxu Hu Xu Tan Tao Qin Sheng Zhao Zhou Zhao and Tie-Yan Liu. 2021. Fastspeech 2: Fast and high-quality end-to-end text to speech. In ICLR."},{"key":"e_1_3_2_1_20_1","volume-title":"Fastspeech: Fast, robust and controllable text to speech. in Advances in Neural Information Processing Systems","author":"Ren Yi","year":"2019","unstructured":"Yi Ren , Yangjun Ruan , Xu Tan , Tao Qin , Sheng Zhao , Zhou Zhao , and Tie-Yan Liu . 2019 . Fastspeech: Fast, robust and controllable text to speech. in Advances in Neural Information Processing Systems (2019), 3171\u20133180. Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2019. Fastspeech: Fast, robust and controllable text to speech. in Advances in Neural Information Processing Systems (2019), 3171\u20133180."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461368"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Yao Shi Hui Bu Xin Xu Shaoji Zhang and Ming Li. 2020. AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines. arXiv preprint arXiv:2010.11567(2020).  Yao Shi Hui Bu Xin Xu Shaoji Zhang and Ming Li. 2020. AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines. arXiv preprint arXiv:2010.11567(2020).","DOI":"10.21437\/Interspeech.2021-755"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"crossref","unstructured":"Andros Tjandra Sakriani Sakti and Satoshi Nakamura. 2018. Machine speech chain with one-shot speaker adaptation. arXiv preprint arXiv:1803.10525(2018).  Andros Tjandra Sakriani Sakti and Satoshi Nakamura. 2018. Machine speech chain with one-shot speaker adaptation. arXiv preprint arXiv:1803.10525(2018).","DOI":"10.21437\/Interspeech.2018-1558"},{"key":"e_1_3_2_1_24_1","volume-title":"International Conference on Machine Learning. PMLR, 5180\u20135189","author":"Wang Yuxuan","year":"2018","unstructured":"Yuxuan Wang , Daisy Stanton , Yu Zhang , RJ- Skerry Ryan , Eric Battenberg , Joel Shor , Ying Xiao , Ye Jia , Fei Ren , and Rif\u00a0 A Saurous . 2018 . Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis . In International Conference on Machine Learning. PMLR, 5180\u20135189 . Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ-Skerry Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif\u00a0A Saurous. 2018. Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis. In International Conference on Machine Learning. PMLR, 5180\u20135189."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053795"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683623"}],"event":{"name":"ICDSP 2022: 2022 6th International Conference on Digital Signal Processing","location":"Chengdu China","acronym":"ICDSP 2022"},"container-title":["Proceedings of the 6th International Conference on Digital Signal Processing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3529570.3529602","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3529570.3529602","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:09:13Z","timestamp":1750183753000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3529570.3529602"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,25]]},"references-count":26,"alternative-id":["10.1145\/3529570.3529602","10.1145\/3529570"],"URL":"https:\/\/doi.org\/10.1145\/3529570.3529602","relation":{},"subject":[],"published":{"date-parts":[[2022,2,25]]},"assertion":[{"value":"2022-06-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}