{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T17:12:58Z","timestamp":1765041178974,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":22,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,2,26]],"date-time":"2021-02-26T00:00:00Z","timestamp":1614297600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61871358"],"award-info":[{"award-number":["61871358"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,2,26]]},"DOI":"10.1145\/3458380.3458405","type":"proceedings-article","created":{"date-parts":[[2021,9,23]],"date-time":"2021-09-23T16:51:57Z","timestamp":1632415917000},"page":"146-150","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Learning Deep and Wide Contextual Representations Using BERT for Statistical Parametric Speech Synthesis"],"prefix":"10.1145","author":[{"given":"Ya-Jie","family":"Zhang","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhen-Hua","family":"Ling","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,9,23]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473(2014). Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473(2014)."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969239.2969304"},{"key":"e_1_3_2_1_3_1","volume-title":"BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805(2018).","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805(2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805(2018)."},{"key":"e_1_3_2_1_4_1","volume-title":"Paragraph-based prosodic cues for speech synthesis applications. Speech prosody","author":"Farrus Mireia","year":"2016","unstructured":"Mireia Farrus , Catherine Lai , and Johanna\u00a0 D Moore . 2016. Paragraph-based prosodic cues for speech synthesis applications. Speech prosody ( 2016 ), 1143\u20131147. Mireia Farrus, Catherine Lai, and Johanna\u00a0D Moore. 2016. Paragraph-based prosodic cues for speech synthesis applications. Speech prosody (2016), 1143\u20131147."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1984.1164317"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"crossref","first-page":"4430","DOI":"10.21437\/Interspeech.2019-3177","article-title":"Pre-Trained Text Embeddings for Enhanced Text-to-Speech Synthesis","volume":"2019","author":"Hayashi Tomoki","year":"2019","unstructured":"Tomoki Hayashi , Shinji Watanabe , Tomoki Toda , Kazuya Takeda , Shubham Toshniwal , and Karen Livescu . 2019 . Pre-Trained Text Embeddings for Enhanced Text-to-Speech Synthesis . Proc. Interspeech 2019 (2019), 4430 \u2013 4434 . Tomoki Hayashi, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Shubham Toshniwal, and Karen Livescu. 2019. Pre-Trained Text Embeddings for Enhanced Text-to-Speech Synthesis. Proc. Interspeech 2019(2019), 4430\u20134434.","journal-title":"Proc. Interspeech"},{"key":"e_1_3_2_1_7_1","unstructured":"Wei-Ning Hsu Yu Zhang Ron\u00a0J Weiss Heiga Zen Yonghui Wu Yuxuan Wang Yuan Cao Ye Jia Zhifeng Chen Jonathan Shen 2018. Hierarchical Generative Modeling for Controllable Speech Synthesis. arXiv preprint arXiv:1810.07217(2018). Wei-Ning Hsu Yu Zhang Ron\u00a0J Weiss Heiga Zen Yonghui Wu Yuxuan Wang Yuan Cao Ye Jia Zhifeng Chen Jonathan Shen 2018. Hierarchical Generative Modeling for Controllable Speech Synthesis. arXiv preprint arXiv:1810.07217(2018)."},{"key":"e_1_3_2_1_8_1","volume-title":"Blizzard Challenge Workshop.","author":"Jiang Yuan","year":"2019","unstructured":"Yuan Jiang , Ya-Jun Hu , Li-Juan Liu , Hong-Chuan Wu , Zhi-Kun Wang , Yang Ai , Zhen-Hua Ling , and Li-Rong Dai . 2019 . The USTC system for Blizzard Challenge 2019 . In Blizzard Challenge Workshop. Yuan Jiang, Ya-Jun Hu, Li-Juan Liu, Hong-Chuan Wu, Zhi-Kun Wang, Yang Ai, Zhen-Hua Ling, and Li-Rong Dai. 2019. The USTC system for Blizzard Challenge 2019. In Blizzard Challenge Workshop."},{"key":"e_1_3_2_1_9_1","volume-title":"Efficient Neural Audio Synthesis. In International Conference on Machine Learning. 2415\u20132424","author":"Kalchbrenner Nal","year":"2018","unstructured":"Nal Kalchbrenner , Erich Elsen , Karen Simonyan , Seb Noury , Norman Casagrande , Edward Lockhart , Florian Stimberg , Aaron Oord , Sander Dieleman , and Koray Kavukcuoglu . 2018 . Efficient Neural Audio Synthesis. In International Conference on Machine Learning. 2415\u20132424 . Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu. 2018. Efficient Neural Audio Synthesis. In International Conference on Machine Learning. 2415\u20132424."},{"key":"e_1_3_2_1_10_1","unstructured":"Naihan Li Shujie Liu Yanqing Liu Sheng Zhao Ming Liu and Ming Zhou. 2018. Close to Human Quality TTS with Transformer. arXiv preprint arXiv:1809.08895(2018). Naihan Li Shujie Liu Yanqing Liu Sheng Zhao Ming Liu and Ming Zhou. 2018. Close to Human Quality TTS with Transformer. arXiv preprint arXiv:1809.08895(2018)."},{"key":"e_1_3_2_1_11_1","unstructured":"Wei Ping Kainan Peng Andrew Gibiansky Sercan\u00a0O Arik Ajay Kannan Sharan Narang Jonathan Raiman and John Miller. 2018. Deep voice 3: Scaling text-to-speech with convolutional sequence learning. (2018). Wei Ping Kainan Peng Andrew Gibiansky Sercan\u00a0O Arik Ajay Kannan Sharan Narang Jonathan Raiman and John Miller. 2018. Deep voice 3: Scaling text-to-speech with convolutional sequence learning. (2018)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/78.650093"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461368"},{"key":"e_1_3_2_1_14_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning,ICML","author":"Skerry-Ryan RJ","year":"2018","unstructured":"RJ Skerry-Ryan , Eric Battenberg , Ying Xiao , Yuxuan Wang , Daisy Stanton , Joel Shor , Ron\u00a0 J Weiss , Rob Clark , and Rif\u00a0 A Saurous . 2018 . Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron . In Proceedings of the 35th International Conference on Machine Learning,ICML 2018. 4700\u20134709. RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron\u00a0J Weiss, Rob Clark, and Rif\u00a0A Saurous. 2018. Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron. In Proceedings of the 35th International Conference on Machine Learning,ICML 2018. 4700\u20134709."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/SLT.2018.8639682"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-32381-3_16"},{"key":"e_1_3_2_1_17_1","unstructured":"A\u00e4ron Van Den\u00a0Oord Sander Dieleman Heiga Zen Karen Simonyan Oriol Vinyals Alex Graves Nal Kalchbrenner Andrew\u00a0W Senior and Koray Kavukcuoglu. 2016. WaveNet: A generative model for raw audio.. In SSW. 125. A\u00e4ron Van Den\u00a0Oord Sander Dieleman Heiga Zen Karen Simonyan Oriol Vinyals Alex Graves Nal Kalchbrenner Andrew\u00a0W Senior and Koray Kavukcuoglu. 2016. WaveNet: A generative model for raw audio.. In SSW. 125."},{"key":"e_1_3_2_1_18_1","volume-title":"Tacotron: Towards End-to-End Speech Synthesis. In INTERSPEECH.","author":"Wang Yuxuan","year":"2017","unstructured":"Yuxuan Wang , R.\u00a0 J. Skerry-Ryan , Daisy Stanton , Yonghui Wu , Ron\u00a0 J. Weiss , Navdeep Jaitly , Zongheng Yang , Ying Xiao , Zhifeng Chen , Samy Bengio , Quoc\u00a0 V. Le , Yannis Agiomyrgiannakis , Rob Clark , and Rif\u00a0 A. Saurous . 2017 . Tacotron: Towards End-to-End Speech Synthesis. In INTERSPEECH. Yuxuan Wang, R.\u00a0J. Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron\u00a0J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc\u00a0V. Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif\u00a0A. Saurous. 2017. Tacotron: Towards End-to-End Speech Synthesis. In INTERSPEECH."},{"key":"e_1_3_2_1_19_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning, ICML 2018,. 5167\u20135176","author":"Wang Yuxuan","year":"2018","unstructured":"Yuxuan Wang , Daisy Stanton , Yu Zhang , RJ Skerry-Ryan , Eric Battenberg , Joel Shor , Ying Xiao , Fei Ren , Ye Jia , and Rif\u00a0 A Saurous . 2018 . Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis . In Proceedings of the 35th International Conference on Machine Learning, ICML 2018,. 5167\u20135176 . Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Fei Ren, Ye Jia, and Rif\u00a0A Saurous. 2018. Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018,. 5167\u20135176."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"crossref","unstructured":"Y. Xiao L. He H. Ming and F.\u00a0K. Soong. 2020. Improving Prosody with Linguistic and BERT Derived Features in Multi-Speaker Based Mandarin Chinese Neural TTS. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). 6704\u20136708. Y. Xiao L. He H. Ming and F.\u00a0K. Soong. 2020. Improving Prosody with Linguistic and BERT Derived Features in Multi-Speaker Based Mandarin Chinese Neural TTS. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). 6704\u20136708.","DOI":"10.1109\/ICASSP40776.2020.9054337"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"crossref","first-page":"4480","DOI":"10.21437\/Interspeech.2019-1418","article-title":"Pre-Trained Text Representations for Improving Front-End Text Processing in Mandarin Text-to-Speech Synthesis","volume":"2019","author":"Yang Bing","year":"2019","unstructured":"Bing Yang , Jiaqi Zhong , and Shan Liu . 2019 . Pre-Trained Text Representations for Improving Front-End Text Processing in Mandarin Text-to-Speech Synthesis . Proc. Interspeech 2019 (2019), 4480 \u2013 4484 . Bing Yang, Jiaqi Zhong, and Shan Liu. 2019. Pre-Trained Text Representations for Improving Front-End Text Processing in Mandarin Text-to-Speech Synthesis. Proc. Interspeech 2019(2019), 4480\u20134484.","journal-title":"Proc. Interspeech"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683623"}],"event":{"name":"ICDSP 2021: 2021 5th International Conference on Digital Signal Processing","acronym":"ICDSP 2021","location":"Chengdu China"},"container-title":["2021 5th International Conference on Digital Signal Processing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3458380.3458405","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3458380.3458405","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:24:43Z","timestamp":1750195483000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3458380.3458405"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,2,26]]},"references-count":22,"alternative-id":["10.1145\/3458380.3458405","10.1145\/3458380"],"URL":"https:\/\/doi.org\/10.1145\/3458380.3458405","relation":{},"subject":[],"published":{"date-parts":[[2021,2,26]]},"assertion":[{"value":"2021-09-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}