{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T21:35:17Z","timestamp":1776116117028,"version":"3.50.1"},"reference-count":71,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2024,10,22]],"date-time":"2024-10-22T00:00:00Z","timestamp":1729555200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,10,22]],"date-time":"2024-10-22T00:00:00Z","timestamp":1729555200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Technische Hochschule N\u00fcrnberg"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Speech Technol"],"published-print":{"date-parts":[[2024,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Generative Adversarial Networks (GANs) have demonstrated promising results as end-to-end models for whispered to voiced speech conversion. Leveraging non-autoregressive systems like GANs capable of performing conditional waveform generation eliminates the need for separate models to estimate voiced speech features, and leads to faster inference compared to autoregressive methods. This study aims to identify the optimal GAN architecture for the whispered to voiced speech conversion task by comparing six state-of-the-art models. Furthermore, we present a method for evaluating the preservation of speaker identity and local accent, using embeddings obtained from speaker- and language identification systems. Our experimental results show that building the speech conversion system based on the HiFi-GAN architecture yields the best objective evaluation scores, outperforming the baseline by <jats:inline-formula><jats:alternatives><jats:tex-math>$$\\sim$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mo>\u223c<\/mml:mo>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>\u00a09% relative using frequency-weighted Signal-to-Noise Ratio and Log Likelihood Ratio, as well as by <jats:inline-formula><jats:alternatives><jats:tex-math>$$\\sim$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mo>\u223c<\/mml:mo>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>\u00a029% relative using Root Mean Squared Error. In subjective tests, HiFi-GAN yielded a mean opinion score of 2.9, significantly outperforming the baseline with a score of 1.4. Furthermore, HiFi-GAN enhanced ASR performance and preserved speaker identity and accent, with correct language detection rates of up to <jats:inline-formula><jats:alternatives><jats:tex-math>$$\\sim$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mo>\u223c<\/mml:mo>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>\u00a098%.<\/jats:p>","DOI":"10.1007\/s10772-024-10161-1","type":"journal-article","created":{"date-parts":[[2024,10,22]],"date-time":"2024-10-22T07:02:17Z","timestamp":1729580537000},"page":"1093-1110","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Generative adversarial networks for whispered to voiced speech conversion: a comparative study"],"prefix":"10.1007","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-7293-9131","authenticated-orcid":false,"given":"Dominik","family":"Wagner","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ilja","family":"Baumann","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tobias","family":"Bocklet","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,10,22]]},"reference":[{"key":"10161_CR1","first-page":"12449","volume":"33","author":"A Baevski","year":"2020","unstructured":"Baevski, A., Zhou, Y., Mohamed, A., & Auli, M. (2020). wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems, 33, 12449\u201312460.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"10161_CR2","unstructured":"Binkowski, M., Donahue, J., Dieleman, S., Clark, A., Elsen, E., Casagrande, N., Cobo, L.C., & Simonyan, K. (2020). High fidelity speech synthesis with adversarial networks. In 8th International conference on learning representations, (ICLR 2020), Addis Ababa, Ethiopia, April 26\u201330, 2020."},{"issue":"3","key":"10161_CR3","doi-asserted-by":"publisher","first-page":"331","DOI":"10.1097\/aud.0b013e3181ff3515","volume":"32","author":"F Chen","year":"2011","unstructured":"Chen, F., & Loizou, P. C. (2011). Predicting the intelligibility of vocoded speech. Ear and Hearing, 32(3), 331\u2013338. https:\/\/doi.org\/10.1097\/aud.0b013e3181ff3515","journal-title":"Ear and Hearing"},{"key":"10161_CR4","doi-asserted-by":"publisher","unstructured":"Desplanques, B., Thienpondt, J., & Demuynck, K. (2020). ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification. In Proceedings Interspeech 2020, (pp. 3830\u20133834). https:\/\/doi.org\/10.21437\/Interspeech.2020-2650","DOI":"10.21437\/Interspeech.2020-2650"},{"key":"10161_CR5","doi-asserted-by":"publisher","first-page":"184","DOI":"10.1109\/JSTSP.2013.2283471","volume":"8","author":"D Erro","year":"2014","unstructured":"Erro, D., Sainz, I., Navas, E., & Hernaez, I. (2014). Harmonics plus noise model based vocoder for statistical parametric speech synthesis. IEEE Journal of Selected Topics in Signal Processing, 8, 184\u2013194.","journal-title":"IEEE Journal of Selected Topics in Signal Processing"},{"issue":"02","key":"10161_CR6","doi-asserted-by":"publisher","first-page":"652","DOI":"10.1109\/TPAMI.2019.2938758","volume":"43","author":"S Gao","year":"2021","unstructured":"Gao, S., Cheng, M., Zhao, K., Zhang, X., Yang, M., & Torr, P. (2021). Res2net: A new multi-scale backbone architecture. IEEE Transactions on Pattern Analysis & Machine Intelligence, 43(02), 652\u2013662. https:\/\/doi.org\/10.1109\/TPAMI.2019.2938758","journal-title":"IEEE Transactions on Pattern Analysis & Machine Intelligence"},{"key":"10161_CR7","unstructured":"Gao, T., Zhou, J., Wang, H., Tao, L., & Kwan, H.K. (2021). Attention-guided generative adversarial network for whisper to normal speech conversion. Preprint at ArXiv abs\/2111.01342"},{"key":"10161_CR8","unstructured":"Garofolo, J., Lamel, L., Fisher, W., Fiscus, J., Pallett, D., Dahlgren, N., & Zue, V. (1993). TIMIT acoustic-phonetic continuous speech corpus LDC93S1. Linguistic Data Consortium"},{"key":"10161_CR9","unstructured":"Goodfellow, I.J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. In Proceedings of the 27th international conference on neural information processing systems, (vol. 2, pp. 2672\u20132680)."},{"key":"10161_CR10","doi-asserted-by":"publisher","first-page":"630","DOI":"10.1007\/978-3-319-46493-0_38","volume-title":"Computer Vision - ECCV 2016","author":"K He","year":"2016","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Identity mappings in deep residual networks. In B. Leibe, J. Matas, N. Sebe, & M. Welling (Eds.), Computer vision - ECCV 2016 (pp. 630\u2013645). Springer."},{"issue":"3","key":"10161_CR11","doi-asserted-by":"publisher","first-page":"1703","DOI":"10.1121\/1.417354","volume":"100","author":"I Holube","year":"1996","unstructured":"Holube, I., & Kollmeier, B. (1996). Speech intelligibility prediction in hearing-impaired listeners based on a psychoacoustically motivated perception model. The Journal of the Acoustical Society of America, 100(3), 1703\u20131716. https:\/\/doi.org\/10.1121\/1.417354","journal-title":"The Journal of the Acoustical Society of America"},{"key":"10161_CR12","doi-asserted-by":"publisher","first-page":"3451","DOI":"10.1109\/TASLP.2021.3122291","volume":"29","author":"W-N Hsu","year":"2021","unstructured":"Hsu, W.-N., Bolte, B., Tsai, Y.-H.H., Lakhotia, K., Salakhutdinov, R., & Mohamed, A. (2021). Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE\/ACM Transactions on Audio, Speech, and Language Processing, 29, 3451\u20133460.","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"10161_CR14","doi-asserted-by":"publisher","unstructured":"Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. In 2018 IEEE\/CVF conference on computer vision and pattern recognition, (pp. 7132\u20137141). https:\/\/doi.org\/10.1109\/CVPR.2018.00745","DOI":"10.1109\/CVPR.2018.00745"},{"issue":"1","key":"10161_CR13","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1109\/TASL.2007.911054","volume":"16","author":"Y Hu","year":"2008","unstructured":"Hu, Y., & Loizou, P. C. (2008). Evaluation of objective quality measures for speech enhancement. IEEE Transactions on Audio, Speech, and Language Processing, 16(1), 229\u2013238. https:\/\/doi.org\/10.1109\/TASL.2007.911054","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"key":"10161_CR15","unstructured":"ITU-T: ITU-T recommendation P.808\u2014subjective evaluation of speech quality with a crowdsourcing approach. Recommendation P.808, International Telecommunication Union, Geneva (2021)"},{"key":"10161_CR16","doi-asserted-by":"publisher","unstructured":"Iwamoto, K., Ochiai, T., Delcroix, M., Ikeshita, R., Sato, H., Araki, S., & Katagiri, S.(2022). How bad are artifacts?: Analyzing the impact of speech enhancement errors on ASR. In Proceedings Interspeech 2022, (pp. 5418\u20135422). https:\/\/doi.org\/10.21437\/Interspeech.2022-318","DOI":"10.21437\/Interspeech.2022-318"},{"key":"10161_CR17","doi-asserted-by":"publisher","unstructured":"Jang, W., Lim, D., Yoon, J., Kim, B., & Kim, J. (2021). UnivNet: A neural vocoder with multi-resolution spectrogram discriminators for high-fidelity waveform generation. In Proceedings Interspeech 2021, (pp. 2207\u20132211). https:\/\/doi.org\/10.21437\/Interspeech.2021-1016 . ISCA.","DOI":"10.21437\/Interspeech.2021-1016"},{"key":"10161_CR18","unstructured":"Kalchbrenner, N., Elsen, E., Simonyan, K., Noury, S., Casagrande, N., Lockhart, E., Stimberg, F., Oord, A., Dieleman, S., & Kavukcuoglu, K.(2018). Efficient neural audio synthesis. In Proceedings of the 35th international conference on machine learning (PMLR), (pp. 2410\u20132419). Proceedings of Machine Learning Research."},{"key":"10161_CR20","doi-asserted-by":"publisher","unstructured":"Kim, J.W., Salamon, J., Li, P., & Bello, J.P. (2018). Crepe: A convolutional representation for pitch estimation. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP), (pp. 161\u2013165). https:\/\/doi.org\/10.1109\/ICASSP.2018.8461329","DOI":"10.1109\/ICASSP.2018.8461329"},{"key":"10161_CR19","unstructured":"Kim, T., Cha, M., Kim, H., Lee, J. K., & Kim, J. (2017) Learning to discover cross-domain relations with generative adversarial networks. In Proceedings of the 34th international conference on machine learning, (vol. 70, pp. 1857\u20131865)."},{"key":"10161_CR21","unstructured":"Kingma, D.P., & Ba, J.(2015). Adam: A method for stochastic optimization. In Bengio, Y., LeCun, Y. (eds.) 3rd International conference on learning representations, (ICLR 2015), May 7\u20139, 2015, Conference Track Proceedings. Preprint at arXiv:1412.6980"},{"key":"10161_CR22","unstructured":"Kong, J., Kim, J., Bae, J. (2020). Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. In Advances in neural information processing systems, (vol. 33, pp. 17022\u201317033)."},{"key":"10161_CR23","unstructured":"Kong, Z., Ping, W., Huang, J., Zhao, K., & Catanzaro, B. (2021). Diffwave: A versatile diffusion model for audio synthesis. In International conference on learning representations."},{"key":"10161_CR24","unstructured":"Kumar, K., Kumar, R., Boissiere, T., Gestin, L., Teoh, W.Z., Sotelo, J., Br\u00e9bisson, A., Bengio, Y., & Courville, A (2019) Melgan: Generative adversarial networks for conditional waveform synthesis. In Advances in neural information processing systems, vol. 32."},{"key":"10161_CR25","doi-asserted-by":"publisher","unstructured":"Leng, Y., Tan, X., Zhao, S., Soong, F., Li, X.-Y., & Qin, T. (2021). MBNET: Mos prediction for synthesized speech with mean-bias network. In 2021 IEEE international conference on acoustics, speech and signal processing (ICASSP 2021), (pp. 391\u2013395). https:\/\/doi.org\/10.1109\/ICASSP39728.2021.9413877","DOI":"10.1109\/ICASSP39728.2021.9413877"},{"key":"10161_CR26","doi-asserted-by":"publisher","first-page":"130495","DOI":"10.1109\/ACCESS.2019.2940700","volume":"7","author":"H Lian","year":"2019","unstructured":"Lian, H., Hu, Y., Yu, W., Zhou, J., & Zheng, W. (2019). Whisper to normal speech conversion using sequence-to-sequence mapping model with auditory attention. IEEE Access, 7, 130495\u2013130504.","journal-title":"IEEE Access"},{"key":"10161_CR27","unstructured":"Lim, B.P.(2010). Computational differences between whispered and non-whispered speech. PhD thesis, University of Illinois."},{"key":"10161_CR28","doi-asserted-by":"publisher","DOI":"10.1201\/b14529","volume-title":"Speech enhancement: Theory and practice","author":"PC Loizou","year":"2013","unstructured":"Loizou, P. C. (2013). Speech enhancement: Theory and practice (2nd ed.). CRC Press Inc.","edition":"2"},{"key":"10161_CR29","doi-asserted-by":"crossref","unstructured":"Lorenzo-Trueba, J., Yamagishi, J., Toda, T., Saito, D., Villavicencio, F., Kinnunen, T., & Ling, Z. (2018). The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods.","DOI":"10.21437\/Odyssey.2018-28"},{"key":"10161_CR30","unstructured":"Loshchilov, I., & Hutter, F. (2019). Decoupled weight decay regularization. In International conference on learning representations. Retrieved from https:\/\/openreview.net\/forum?id=Bkg6RiCqY7"},{"issue":"5","key":"10161_CR31","doi-asserted-by":"publisher","first-page":"3387","DOI":"10.1121\/1.3097493","volume":"125","author":"J Ma","year":"2009","unstructured":"Ma, J., Hu, Y., & Loizou, P. C. (2009). Objective measures for predicting speech intelligibility in noisy conditions based on new band-importance functions. The Journal of the Acoustical Society of America, 125(5), 3387. https:\/\/doi.org\/10.1121\/1.3097493","journal-title":"The Journal of the Acoustical Society of America"},{"key":"10161_CR32","first-page":"2579","volume":"9","author":"L Maaten","year":"2008","unstructured":"Maaten, L., & Hinton, G. (2008). Visualizing data using T-SNE. Journal of Machine Learning Research, 9, 2579\u20132605.","journal-title":"Journal of Machine Learning Research"},{"key":"10161_CR33","doi-asserted-by":"publisher","unstructured":"Malaviya, H., Shah, J., Patel, M., Munshi, J., & Patil, H.A. (2020). Mspec-net: Multi-domain speech conversion network. In 2020 IEEE international conference on acoustics, speech and signal processing (ICASSP 2020), (pp. 7764\u20137768). https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9052966","DOI":"10.1109\/ICASSP40776.2020.9052966"},{"key":"10161_CR34","doi-asserted-by":"publisher","unstructured":"Mao, X., Li, Q., Xie, H., Lau, R.Y.K., Wang, Z., & Smolley, S.P.(2017). Least squares generative adversarial networks. In 2017 IEEE international conference on computer vision (ICCV), (pp. 2813\u20132821). https:\/\/doi.org\/10.1109\/ICCV.2017.304","DOI":"10.1109\/ICCV.2017.304"},{"key":"10161_CR35","doi-asserted-by":"crossref","unstructured":"Mashimo, M., Toda, T., Shikano, K., & Campbell, N. (2001). Evaluation of cross-language voice conversion based on GMM and straight. In Proceedings 7th European conference on speech communication and technology (Eurospeech 2001), (pp. 361\u2013364). ISCA.","DOI":"10.21437\/Eurospeech.2001-111"},{"key":"10161_CR36","doi-asserted-by":"publisher","unstructured":"McAuliffe, M., Socolof, M., Mihuc, S., Wagner, M., & Sonderegger, M. (2017). Montreal forced aligner: Trainable text-speech alignment using Kaldi. In Proceedings Interspeech 2017, (pp. 498\u2013502). https:\/\/doi.org\/10.21437\/Interspeech.2017-1386","DOI":"10.21437\/Interspeech.2017-1386"},{"key":"10161_CR37","doi-asserted-by":"publisher","unstructured":"McLoughlin, I.V., Li, J., & Song, Y. (2013). Reconstruction of continuous voiced speech from whispers. In Proceedings Interspeech 2013, (pp. 1022\u20131026). https:\/\/doi.org\/10.21437\/Interspeech.2013-111","DOI":"10.21437\/Interspeech.2013-111"},{"key":"10161_CR38","doi-asserted-by":"crossref","unstructured":"Meenakshi, G.N., & Ghosh, P.K. (2018). Whispered speech to neutral speech conversion using bidirectional LSTMS. In Proceedings Interspeech 2018, (pp. 491\u2013495). ISCA.","DOI":"10.21437\/Interspeech.2018-1487"},{"key":"10161_CR39","doi-asserted-by":"publisher","DOI":"10.1587\/transinf.2015EDP7457","author":"M Morise","year":"2016","unstructured":"Morise, M., Yokomori, F., & Ozawa, K. (2016). World: A vocoder-based high-quality speech synthesis system for real-time applications. IEICE Transactions on Information and Systems. https:\/\/doi.org\/10.1587\/transinf.2015EDP7457","journal-title":"IEICE Transactions on Information and Systems"},{"key":"10161_CR40","unstructured":"Morrison, M., Kumar, R., Kumar, K., Seetharaman, P., Courville, A., & Bengio, Y. (2022). Chunked autoregressive GAN for conditional waveform synthesis. In International conference on learning representations (ICLR)."},{"issue":"0","key":"10161_CR41","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1080\/02699206.2022.2092425","volume":"0","author":"F Mouret","year":"2022","unstructured":"Mouret, F., Crevier-Buchman, L., & Pillot-Loiseau, C. (2022). Intelligibility of pseudo-whispered speech after total laryngectomy. Clinical Linguistics & Phonetics. Advance online publication\u00a0https:\/\/doi.org\/10.1080\/02699206.2022.2092425","journal-title":"Clinical Linguistics &amp; Phonetics"},{"key":"10161_CR42","doi-asserted-by":"publisher","unstructured":"Naderi, B., & Cutler, R. (2020). An open source implementation of ITU-T Recommendation P.808 with validation. In Proceedings Interspeech 2020, (pp. 2862\u20132866). https:\/\/doi.org\/10.21437\/Interspeech.2020-2665","DOI":"10.21437\/Interspeech.2020-2665"},{"key":"10161_CR43","doi-asserted-by":"publisher","unstructured":"Nagrani, A., Chung, J.S., & Zisserman, A. (2017). VoxCeleb: A large-scale speaker identification dataset. In Proceedings Interspeech 2017, (pp. 2616\u20132620). https:\/\/doi.org\/10.21437\/Interspeech.2017-950","DOI":"10.21437\/Interspeech.2017-950"},{"key":"10161_CR44","unstructured":"Niranjan, A., Sharma, M., Gutha, S.B.C., & Shaik, M.A.B. (2021). End-to-end whisper to natural speech conversion using modified transformer network. Preprint at ArXiv abs\/2004.09347"},{"key":"10161_CR45","unstructured":"Oord, A., Vinyals, O., & Kavukcuoglu, K. (2017). Neural discrete representation learning. In Advances in neural information processing systems, (NeurIPS 2017), vol. 30."},{"key":"10161_CR46","doi-asserted-by":"crossref","unstructured":"Parmar, M., Doshi, S., Shah, N.J., Patel, M., & Patil, H.A. (2019). Effectiveness of cross-domain architectures for whisper-to-normal speech conversion. In 27th European signal processing conference (EUSIPCO), (pp. 1\u20135). EURASIP.","DOI":"10.23919\/EUSIPCO.2019.8902961"},{"key":"10161_CR48","doi-asserted-by":"crossref","unstructured":"Pascual, S., Bonafonte, A., & Serr\u00e0 , J. (2017). SEGAN: Speech enhancement generative adversarial network. In Proceedings Interspeech 2017, (pp. 3642\u20133646). ISCA.","DOI":"10.21437\/Interspeech.2017-1428"},{"key":"10161_CR47","doi-asserted-by":"crossref","unstructured":"Pascual, S., Bonafonte, A., Serr\u00e0 , J., & Gonz\u00e1\u00a1lez L\u00f3pez, J.A.(2018). Whispered-to-voiced alaryngeal speech conversion with generative adversarial networks. In Proceedings IberSPEECH 2018, (pp. 117\u2013121). ISCA.","DOI":"10.21437\/IberSPEECH.2018-25"},{"key":"10161_CR49","doi-asserted-by":"publisher","unstructured":"Patel, M., Parmar, M., Doshi, S., Shah, N., & Patil, H. (2019). Novel inception-GAN for whispered-to-normal speech conversion. In Proceedings 10th ISCA workshop on speech synthesis (SSW 10), (pp. 87\u201392). https:\/\/doi.org\/10.21437\/SSW.2019-16","DOI":"10.21437\/SSW.2019-16"},{"key":"10161_CR50","doi-asserted-by":"publisher","unstructured":"Patel, M., Purohit, M., Shah, J., & Patil, H.A. (2020). CinC-GAN for effective F0 prediction for whisper-to-normal speech conversion. In 2020 28th European signal processing conference (EUSIPCO), (pp. 411\u2013415). Retrieved from https:\/\/doi.org\/10.23919\/Eusipco47968.2020.9287385 . https:\/\/www.eurasip.org\/proceedings\/eusipco\/eusipco2020\/pdfs\/0000411.pdf","DOI":"10.23919\/Eusipco47968.2020.9287385"},{"key":"10161_CR51","doi-asserted-by":"crossref","unstructured":"Prenger, R., Valle, R., & Catanzaro, B. (2019). Waveglow: A flow-based generative network for speech synthesis. In 2019 IEEE international conference on acoustics, speech and signal processing (ICASSP), (pp. 3617\u20133621). IEEE.","DOI":"10.1109\/ICASSP.2019.8683143"},{"key":"10161_CR52","doi-asserted-by":"publisher","unstructured":"Rekimoto, J. (2023). WESPER: Zero-shot and realtime whisper to normal voice conversion for whisper-based speech interactions. In Proceedings of the 2023 CHI conference on human factors in computing systems (CHI \u201923). https:\/\/doi.org\/10.1145\/3544548.3580706","DOI":"10.1145\/3544548.3580706"},{"key":"10161_CR53","unstructured":"Ren, Y., Hu, C., Tan, X., Qin, T., Zhao, S., Zhao, Z., & Liu, T.-Y. (2021). Fastspeech 2: Fast and high-quality end-to-end text to speech. In International conference on learning representations (ICLR). https:\/\/openreview.net\/forum?id=piLPYqxtWuA"},{"key":"10161_CR54","doi-asserted-by":"publisher","unstructured":"Rosenberg, A., & Ramabhadran, B. (2017). Bias and statistical significance in evaluating speech synthesis with mean opinion scores. In Proceedings Interspeech 2017, (pp. 3976\u20133980). https:\/\/doi.org\/10.21437\/Interspeech.2017-479","DOI":"10.21437\/Interspeech.2017-479"},{"key":"10161_CR55","doi-asserted-by":"publisher","unstructured":"Safari, P., India, M., & Hernando, J. (2020). Self-attention encoding and pooling for speaker recognition. In Proceedings Interspeech 2020, (pp. 941\u2013945). https:\/\/doi.org\/10.21437\/Interspeech.2020-1446","DOI":"10.21437\/Interspeech.2020-1446"},{"key":"10161_CR56","unstructured":"Salimans, T., & Kingma, D.P. (2016). Weight normalization: A simple reparameterization to accelerate training of deep neural networks. In Advances in neural information processing systems, vol. 29. Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2016\/file\/ed265bc903a5a097f61d3ec064d96d2e-Paper.pdf"},{"key":"10161_CR57","unstructured":"Shah, N., Parmar, M., Shah, N., & Patil, H.A.(2018). Novel MMSE DiscoGAN for crossdomain whisper-to-speech conversion. In Machine learning in speech and language processing workshop, (MLSLP), (pp. 1\u20133). Google."},{"key":"10161_CR58","doi-asserted-by":"publisher","unstructured":"Snyder, D., Garcia-Romero, D., McCree, A., Sell, G., Povey, D., & Khudanpur, S. (2018). Spoken Language Recognition using X-vectors. In Proceedings the speaker and language recognition workshop (Odyssey 2018), (pp. 105\u2013111). https:\/\/doi.org\/10.21437\/Odyssey.2018-15","DOI":"10.21437\/Odyssey.2018-15"},{"issue":"7","key":"10161_CR59","doi-asserted-by":"publisher","first-page":"2125","DOI":"10.1109\/TASL.2011.2114881","volume":"19","author":"CH Taal","year":"2011","unstructured":"Taal, C. H., Hendriks, R. C., Heusdens, R., & Jensen, J. (2011). An algorithm for intelligibility prediction of time-frequency weighted noisy speech. IEEE Transactions on Audio, Speech, and Language Processing, 19(7), 2125\u20132136. https:\/\/doi.org\/10.1109\/TASL.2011.2114881","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"issue":"8","key":"10161_CR60","doi-asserted-by":"publisher","first-page":"2222","DOI":"10.1109\/TASL.2007.907344","volume":"15","author":"T Toda","year":"2007","unstructured":"Toda, T., Black, A. W., & Tokuda, K. (2007). Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory. IEEE Transactions on Audio, Speech, and Language Processing, 15(8), 2222\u20132235. https:\/\/doi.org\/10.1109\/TASL.2007.907344","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"key":"10161_CR62","doi-asserted-by":"publisher","unstructured":"Toda, T., Chen, L.-H., Saito, D., Villavicencio, F., Wester, M., Wu, Z., & Yamagishi, J. (2016). The voice conversion challenge 2016. In Proceedings Interspeech 2016, (pp. 1632\u20131636). https:\/\/doi.org\/10.21437\/Interspeech.2016-1066","DOI":"10.21437\/Interspeech.2016-1066"},{"key":"10161_CR61","doi-asserted-by":"crossref","unstructured":"Toda, T., & Shikano, K. (2005). Nam-to-speech conversion with gaussian mixture models. In INTERSPEECH 2005\u2014 Eurospeech, 9th European conference on speech communication and technology, (pp. 1957\u20131960). ISCA.","DOI":"10.21437\/Interspeech.2005-611"},{"key":"10161_CR63","doi-asserted-by":"publisher","unstructured":"Tseng, W.-C., Huang, C.-y., Kao, W.-T., Lin, Y.Y., Lee, H.-y.(2021). Utilizing self-supervised representations for MOS prediction. In Proceedings Interspeech 2021, (pp. 2781\u20132785). https:\/\/doi.org\/10.21437\/Interspeech.2021-2013","DOI":"10.21437\/Interspeech.2021-2013"},{"key":"10161_CR64","doi-asserted-by":"crossref","unstructured":"Valk, J., & Alum\u00e4e, T. (2021). VoxLingua107: A dataset for spoken language recognition. In Proceedings IEEE SLT Workshop.","DOI":"10.1109\/SLT48900.2021.9383459"},{"key":"10161_CR71","unstructured":"van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., Kavukcuoglu, K.(2016). WaveNet: A generative model for raw audio. In Proceedings 9th ISCA workshop on speech synthesis workshop (SSW 9), (p. 125)."},{"key":"10161_CR65","doi-asserted-by":"crossref","unstructured":"Wagner, D., Bayerl, S. P., Maruri, H. C., & Bocklet, T. (2022). Generative models for improved naturalness intelligibility and voicing of whispered speech. In 2022 IEEE spoken language technology workshop (SLT), (pp. 943\u2013948).","DOI":"10.1109\/SLT54892.2023.10022796"},{"key":"10161_CR66","doi-asserted-by":"publisher","unstructured":"Yamamoto, R., Song, E., & Kim, J.-M.(2019). Probability density distillation with generative adversarial networks for high-quality parallel waveform generation. In Proceedings Interspeech 2019, (pp. 699\u2013703). https:\/\/doi.org\/10.21437\/Interspeech.2019-1965","DOI":"10.21437\/Interspeech.2019-1965"},{"key":"10161_CR67","doi-asserted-by":"publisher","unstructured":"Yamamoto, R., Song, E., & Kim, J. (2020). Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram. In 2020 IEEE international conference on acoustics, speech and signal processing (ICASSP 2020), (pp. 6199\u20136203). https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9053795","DOI":"10.1109\/ICASSP40776.2020.9053795"},{"key":"10161_CR68","doi-asserted-by":"crossref","unstructured":"Yang, G., Yang, S., Liu, K., Fang, P., Chen, W., & Xie, L. (2021). Multi-band melgan: Faster waveform generation for high-quality text-to-speech. In 2021 IEEE spoken language technology workshop (SLT), (pp. 492\u2013498). IEEE.","DOI":"10.1109\/SLT48900.2021.9383551"},{"key":"10161_CR69","doi-asserted-by":"publisher","unstructured":"Yu, C., Lu, H., Hu, N., Yu, M., Weng, C., Xu, K., Liu, P., Tuo, D., Kang, S., Lei, G., Su, D., & Yu, D. (2020). DurIAN: Duration informed attention network for speech synthesis. In Proceedings Interspeech 2020, (pp. 2027\u20132031). https:\/\/doi.org\/10.21437\/Interspeech.2020-2968 .","DOI":"10.21437\/Interspeech.2020-2968"},{"key":"10161_CR70","doi-asserted-by":"publisher","unstructured":"Zeng, Z., Wang, J., Cheng, N., & Xiao, J. (2021). Lvcnet: Efficient condition-dependent modeling network for waveform generation. In 2021 IEEE international conference on acoustics, speech and signal processing (ICASSP 2021), (pp. 6054\u20136058). https:\/\/doi.org\/10.1109\/ICASSP39728.2021.9414710","DOI":"10.1109\/ICASSP39728.2021.9414710"}],"container-title":["International Journal of Speech Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10772-024-10161-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10772-024-10161-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10772-024-10161-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,12,16]],"date-time":"2024-12-16T10:09:01Z","timestamp":1734343741000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10772-024-10161-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,22]]},"references-count":71,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,12]]}},"alternative-id":["10161"],"URL":"https:\/\/doi.org\/10.1007\/s10772-024-10161-1","relation":{},"ISSN":["1381-2416","1572-8110"],"issn-type":[{"value":"1381-2416","type":"print"},{"value":"1572-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,22]]},"assertion":[{"value":"9 April 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 October 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 October 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}