{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T20:13:08Z","timestamp":1776888788922,"version":"3.51.2"},"reference-count":77,"publisher":"Association for Computing Machinery (ACM)","issue":"10","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U23B2053"],"award-info":[{"award-number":["U23B2053"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2025,10,31]]},"abstract":"<jats:p>Conventional speech synthesis techniques have made significant strides towards achieving human-like performance. However, the domain of audiobook speech synthesis still presents notable challenges. On one hand, the speech in audiobooks exhibits rich prosodic expressiveness, posing substantial difficulties in prosody modeling. On the other hand, the reader of audiobooks uses different voices to perform dialogues of different characters, which has been inadequately explored in existing speech synthesis methods. To address the first challenge, we integrate discourse-scale prosody modeling into the conventional autoencoder-based framework and introduce generative adversarial networks (GANs) for phoneme-level prosody code prediction. Regarding the second challenge, we further explore a character voice encoder based on the pretrained speaker verification model, integrating it into our proposed method. Experimental results validate that the proposed method enhances the prosodic expressiveness of synthesized audiobook speech. Moreover, it demonstrates the capacity to produce distinctive voices for different audiobook characters without compromising the naturalness of the synthesized speech.<\/jats:p>","DOI":"10.1145\/3749644","type":"journal-article","created":{"date-parts":[[2025,8,11]],"date-time":"2025-08-11T11:26:06Z","timestamp":1754911566000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Enhanced Prosody Modeling and Character Voice Controlling for Audiobook Speech Synthesis"],"prefix":"10.1145","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-5102-1182","authenticated-orcid":false,"given":"Ning-Qian","family":"Wu","sequence":"first","affiliation":[{"name":"National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China","place":["Hefei, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7853-5273","authenticated-orcid":false,"given":"Zhen-Hua","family":"Ling","sequence":"additional","affiliation":[{"name":"National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China","place":["Hefei, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,9,12]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Xu Tan Tao Qin Frank Soong and Tie-Yan Liu. 2021. A survey on neural speech synthesis. arXiv:2106.15561. Retrieved from https:\/\/arxiv.org\/abs\/2106.15561"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461368"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33016706"},{"key":"e_1_3_2_5_2","first-page":"3165","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada","author":"Ren Yi","year":"2019","unstructured":"Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2019. FastSpeech: Fast, robust and controllable text to speech. In Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada. Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d\u2019Alch\u00e9-Buc, Emily B. Fox, and Roman Garnett (Eds.), 3165\u20133174."},{"key":"e_1_3_2_6_2","volume-title":"Proceedings of the 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021","author":"Ren Yi","year":"2021","unstructured":"Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2021. FastSpeech 2: Fast and high-quality end-to-end text to speech. In Proceedings of the 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414718"},{"key":"e_1_3_2_8_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020.","author":"Kim Jaehyeon","year":"2020","unstructured":"Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon. 2020. Glow-TTS: A generative flow for text-to-speech via monotonic alignment search. In Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020.Hugo Larochelle, Marc\u2019Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)."},{"key":"e_1_3_2_9_2","series-title":"Proceedings of Machine Learning Research","first-page":"8599","volume-title":"Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event","volume":"139","author":"Popov Vadim","year":"2021","unstructured":"Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail A. Kudinov. 2021. Grad-TTS: A diffusion probabilistic model for text-to-speech. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event(Proceedings of Machine Learning Research, Vol. 139). Marina Meila and Tong Zhang (Eds.), PMLR, 8599\u20138608."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.21437\/INTERSPEECH.2023-534"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3356232"},{"key":"e_1_3_2_12_2","series-title":"Proceedings of Machine Learning Research","first-page":"4700","volume-title":"Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm\u00e4ssan, Stockholm, Sweden, July 10-15, 2018","volume":"80","author":"Skerry-Ryan R. J.","year":"2018","unstructured":"R. J. Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron J. Weiss, Rob Clark, and Rif A. Saurous. 2018. Towards end-to-end prosody transfer for expressive speech synthesis with tacotron. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm\u00e4ssan, Stockholm, Sweden, July 10-15, 2018(Proceedings of Machine Learning Research, Vol. 80). Jennifer G. Dy and Andreas Krause (Eds.), PMLR, 4700\u20134709."},{"key":"e_1_3_2_13_2","series-title":"Proceedings of Machine Learning Research","first-page":"5167","volume-title":"Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm\u00e4ssan, Stockholm, Sweden, July 10-15, 2018","volume":"80","author":"Wang Yuxuan","year":"2018","unstructured":"Yuxuan Wang, Daisy Stanton, Yu Zhang, R. J. Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A. Saurous. 2018. Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm\u00e4ssan, Stockholm, Sweden, July 10-15, 2018(Proceedings of Machine Learning Research, Vol. 80). Jennifer G. Dy and Andreas Krause (Eds.), PMLR, 5167\u20135176."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414413"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9413696"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-2053"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-1430"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414102"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.21437\/SSW.2021-12"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.21437\/SSW.2021-37"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10096247"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2023.3278184"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.21437\/SSW.2023-22"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU57964.2023.10389629"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2024-1862"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9746238"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2021-465"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10096285"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2024.3402088"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2023-1779"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP48485.2024.10445804"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3133205"},{"key":"e_1_3_2_34_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020.","author":"Kong Jungil","year":"2020","unstructured":"Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020. HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis. In Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020.Hugo Larochelle, Marc\u2019Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053795"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10095105"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-2176"},{"key":"e_1_3_2_38_2","first-page":"13198","volume-title":"Proceedings of the 35th AAAI Conference on Artificial Intelligence, AAAI 2021, 33rd Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The 11th Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021","author":"Lee Sang-Hoon","year":"2021","unstructured":"Sang-Hoon Lee, Hyun-Wook Yoon, Hyeong-Rae Noh, Ji-Hoon Kim, and Seong-Whan Lee. 2021. Multi-SpectroGAN: High-diversity and high-fidelity spectrogram generation with adversarial style combination for speech synthesis. In Proceedings of the 35th AAAI Conference on Artificial Intelligence, AAAI 2021, 33rd Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The 11th Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021. AAAI,13198\u201313206."},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","unstructured":"Yinghao Aaron Li Cong Han and Nima Mesgarani. 2025. StyleTTS: A style-based generative model for natural and diverse text-to-speech synthesis. IEEE Journal of Selected Topics in Signal Processing 19 1 (2025) 283\u2013296. DOI:10.1109\/JSTSP.2025.3530171","DOI":"10.1109\/JSTSP.2025.3530171"},{"key":"e_1_3_2_40_2","first-page":"270","volume-title":"Proceedings of the Speech Communication; 15th ITG Conference","author":"Sani Paolo","year":"2023","unstructured":"Paolo Sani, Judith Bauer, Frank Zalkow, Emanuel AP Habets, and Christian Dittmar. 2023. Improving the naturalness of synthesized spectrograms for TTS using GANBased post-processing. In Proceedings of the Speech Communication; 15th ITG Conference. VDE, 270\u2013274."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10095513"},{"key":"e_1_3_2_42_2","first-page":"4485","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montr\u00e9al, Canada","author":"Jia Ye","year":"2018","unstructured":"Ye Jia, Yu Zhang, Ron J. Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez-Moreno, and Yonghui Wu. 2018. Transfer learning from speaker verification to multispeaker text-to-speech synthesis. In Proceedings of the Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montr\u00e9al, Canada. Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol\u00f2 Cesa-Bianchi, and Roman Garnett (Eds.), 4485\u20134495."},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-1032"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-2650"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461375"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9415078"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9413889"},{"key":"e_1_3_2_48_2","series-title":"Proceedings of Machine Learning Research","first-page":"3331","volume-title":"Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA","volume":"97","author":"Kenter Tom","year":"2019","unstructured":"Tom Kenter, Vincent Wan, Chun-an Chan, Rob Clark, and Jakub Vit. 2019. CHiVE: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA(Proceedings of Machine Learning Research, Vol. 97). Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.), PMLR, 3331\u20133340."},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2018-1113"},{"key":"e_1_3_2_50_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10-16, 2023.","author":"Li Yinghao Aaron","year":"2023","unstructured":"Yinghao Aaron Li, Cong Han, Vinay S. Raghavan, Gavin Mischler, and Nima Mesgarani. 2023. StyleTTS 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10-16, 2023.Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)."},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP48485.2024.10446191"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3074757"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/SLT48900.2021.9383629"},{"key":"e_1_3_2_54_2","first-page":"6306","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA","author":"Oord A\u00e4ron van den","year":"2017","unstructured":"A\u00e4ron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017. Neural discrete representation learning. In Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA. Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (Eds.), 6306\u20136315."},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053436"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-1615"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2024.3395994"},{"key":"e_1_3_2_58_2","doi-asserted-by":"crossref","first-page":"2280","DOI":"10.21437\/Interspeech.2024-715","volume-title":"Proceedings of the Interspeech 2024","author":"Korotkova Yuliya","year":"2024","unstructured":"Yuliya Korotkova, Ilya Kalinovskiy, and Tatiana Vakhrusheva. 2024. Word-level text markup for prosody control in speech synthesis. In Proceedings of the Interspeech 2024. 2280\u20132284."},{"key":"e_1_3_2_59_2","volume-title":"Proceedings of the 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014.","author":"Kingma Diederik P.","year":"2014","unstructured":"Diederik P. Kingma and Max Welling. 2014. Auto-encoding variational bayes. In Proceedings of the 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014.Yoshua Bengio and Yann LeCun (Eds.)."},{"key":"e_1_3_2_60_2","volume-title":"Proceedings of the 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019","author":"Hsu Wei-Ning","year":"2019","unstructured":"Wei-Ning Hsu, Yu Zhang, Ron J. Weiss, Heiga Zen, Yonghui Wu, Yuxuan Wang, Yuan Cao, Ye Jia, Zhifeng Chen, Jonathan Shen, Patrick Nguyen, and Ruoming Pang. 2019. Hierarchical generative modeling for controllable speech synthesis. In Proceedings of the 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net."},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683623"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/SLT.2018.8639682"},{"key":"e_1_3_2_63_2","first-page":"2962","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA","author":"Gibiansky Andrew","year":"2017","unstructured":"Andrew Gibiansky, Sercan \u00d6mer Arik, Gregory Frederick Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou. 2017. Deep voice 2: Multi-speaker neural text-to-speech. In Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA. Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (Eds.), 2962\u20132970."},{"key":"e_1_3_2_64_2","volume-title":"Proceedings of the 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019","author":"Chen Yutian","year":"2019","unstructured":"Yutian Chen, Yannis M. Assael, Brendan Shillingford, David Budden, Scott E. Reed, Heiga Zen, Quan Wang, Luis C. Cobo, Andrew Trask, Ben Laurie, \u00c7aglar G\u00fcl\u00e7ehre, A\u00e4ron van den Oord, Oriol Vinyals, and Nando de Freitas. 2019. Sample efficient adaptive text-to-speech. In Proceedings of the 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net."},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-3139"},{"key":"e_1_3_2_66_2","first-page":"10040","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montr\u00e9al, Canada","author":"Arik Sercan \u00d6mer","year":"2018","unstructured":"Sercan \u00d6mer Arik, Jitong Chen, Kainan Peng, Wei Ping, and Yanqi Zhou. 2018. Neural voice cloning with a few samples. In Proceedings of the Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montr\u00e9al, Canada. Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol\u00f2 Cesa-Bianchi, and Roman Garnett (Eds.), 10040\u201310050."},{"key":"e_1_3_2_67_2","series-title":"Proceedings of Machine Learning Research","first-page":"2709","volume-title":"Proceedings of the International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA.","volume":"162","author":"Casanova Edresson","year":"2022","unstructured":"Edresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo C\u00e2ndido J\u00fanior, Eren G\u00f6lge, and Moacir A. Ponti. 2022. YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone. In Proceedings of the International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA.Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesv\u00e1ri, Gang Niu, and Sivan Sabato (Eds.), Proceedings of Machine Learning Research, Vol. 162, PMLR, 2709\u20132720."},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCSLP57327.2022.10037956"},{"key":"e_1_3_2_69_2","first-page":"426","volume-title":"Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2024 - Volume 2: Short Papers, St. Julian\u2019s, Malta, March 17-22, 2024","author":"Tankala Pavan","year":"2024","unstructured":"Pavan Tankala, Preethi Jyothi, Preeti Rao, and Pushpak Bhattacharyya. 2024. STORiCo: Storytelling TTS for hindi with character voice modulation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2024 - Volume 2: Short Papers, St. Julian\u2019s, Malta, March 17-22, 2024 (2024). Yvette Graham and Matthew Purver (Eds.), Association for Computational Linguistics, 426\u2013431. Retrieved from https:\/\/aclanthology.org\/2024.eacl-short.37"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.21437\/INTERSPEECH.2022-638"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-3015"},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.304"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.400"},{"key":"e_1_3_2_74_2","first-page":"3455","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023","author":"Chen Yue","year":"2023","unstructured":"Yue Chen, Tianwei He, Hongbin Zhou, Jia-Chen Gu, Heng Lu, and Zhen-Hua Ling. 2023. Symbolization, prompt, and classification: A framework for implicit speaker identification in novels. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023. Houda Bouamor, Juan Pino, and Kalika Bali (Eds.), Association for Computational Linguistics, 3455\u20133467."},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.21437\/BLIZZARD.2017-13"},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.1587\/TRANSINF.2015EDP7457"},{"key":"e_1_3_2_77_2","first-page":"498","volume-title":"Proceedings of the Interspeech 2017, 18th Annual Conference of the International Speech Communication Association, Stockholm, Sweden, August 20-24, 2017","author":"McAuliffe Michael","year":"2017","unstructured":"Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger. 2017. Montreal forced aligner: Trainable text-speech alignment using kaldi. In Proceedings of the Interspeech 2017, 18th Annual Conference of the International Speech Communication Association, Stockholm, Sweden, August 20-24, 2017. Francisco Lacerda (Ed.), ISCA, 498\u2013502."},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2018-1929"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3749644","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,12]],"date-time":"2025-09-12T12:54:43Z","timestamp":1757681683000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3749644"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,12]]},"references-count":77,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2025,10,31]]}},"alternative-id":["10.1145\/3749644"],"URL":"https:\/\/doi.org\/10.1145\/3749644","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,12]]},"assertion":[{"value":"2024-02-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-03","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}