{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T05:39:55Z","timestamp":1777613995139,"version":"3.51.4"},"reference-count":35,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,10,28]],"date-time":"2023-10-28T00:00:00Z","timestamp":1698451200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,10,28]],"date-time":"2023-10-28T00:00:00Z","timestamp":1698451200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100012165","name":"Key Technologies Research and Development Program","doi-asserted-by":"publisher","award":["2020AAA0107902"],"award-info":[{"award-number":["2020AAA0107902"]}],"id":[{"id":"10.13039\/501100012165","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Non-parallel data voice conversion (VC) has achieved considerable breakthroughs due to self-supervised pre-trained representation (SSPR) being used in recent years. Features extracted by the pre-trained model are expected to contain more content information. However, in common VC with SSPR, there is no special implementation to remove speaker information in the content representation extraction by SSPR, which prevents further purification of the speaker information from SSPR representation. Moreover, in conventional VC, Mel-spectrogram is often selected as the reconstructed acoustic feature, which is not consistent with the input of the content encoder and results in some information lost. Motivated by the above, we proposed W2VC to settle the issues. W2VC consists of three parts: (1) We reconstruct feature from WavLM representation (WLMR) that is more consistent with the input of content encoder; (2) Connectionist temporal classification (CTC) is used to align content representation and text context from phoneme level, content encoder plus gradient reversal layer (GRL) based speaker classifier are used to remove speaker information in the content representation extraction; (3) WLMR-based HiFi-GAN is trained to convert WLMR to waveform speech. VC experimental results show that GRL can purify well the content information of the self-supervised model. The GRL purification and CTC supervision on the content encoder are complementary in improving the VC performance. Moreover, the synthesized speech using the WLMR retrained vocoder achieves better results in both subjective and objective evaluation. The proposed method is evaluated on the VCTK and CMU databases. It is shown the method achieves 8.901 in objective MCD, 4.45 in speech naturalness, and 3.62 in speaker similarity of subjective MOS score, which is superior to the baseline.<\/jats:p>","DOI":"10.1186\/s13636-023-00312-8","type":"journal-article","created":{"date-parts":[[2023,10,28]],"date-time":"2023-10-28T14:02:27Z","timestamp":1698501747000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["W2VC: WavLM representation based one-shot voice conversion with gradient reversal distillation and CTC supervision"],"prefix":"10.1186","volume":"2023","author":[{"given":"Hao","family":"Huang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lin","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4954-8398","authenticated-orcid":false,"given":"Jichen","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ying","family":"Hu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liang","family":"He","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,10,28]]},"reference":[{"key":"312_CR1","doi-asserted-by":"crossref","unstructured":"A. Kain, M.W. Macon, in Proceedings of the 1998 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP\u201998 (Cat. No. 98CH36181). Spectral voice conversion for text-to-speech synthesis, vol. 1 (IEEE, 1998), pp. 285\u2013288","DOI":"10.1109\/ICASSP.1998.674423"},{"key":"312_CR2","doi-asserted-by":"crossref","unstructured":"K. Kobayashi, T. Toda, T. Nakano, M. Goto, G. Neubig, S. Sakti, S. Nakamura, in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Regression approaches to perceptual age control in singing voice conversion (IEEE, 2014), pp. 7904\u20137908","DOI":"10.1109\/ICASSP.2014.6855139"},{"key":"312_CR3","doi-asserted-by":"crossref","unstructured":"Z. Du, B. Sisman, K. Zhou, H. Li, in 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). Expressive voice conversion: A joint framework for speaker identity and emotional style transfer (IEEE, 2021), pp. 594\u2013601","DOI":"10.1109\/ASRU51503.2021.9687906"},{"key":"312_CR4","doi-asserted-by":"crossref","unstructured":"C.C. Hsu, H.T. Hwang, Y.C. Wu, Y. Tsao, H.M. Wang, in 2016 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA). Voice conversion from non-parallel corpora using variational auto-encoder (IEEE, 2016), pp. 1\u20136","DOI":"10.1109\/APSIPA.2016.7820786"},{"key":"312_CR5","doi-asserted-by":"crossref","unstructured":"X. Tian, J. Wang, H. Xu, E.S. Chng, H. Li, in Odyssey. Average modeling approach to voice conversion with non-parallel data, vol. 2018 (2018), pp. 227\u2013232","DOI":"10.21437\/Odyssey.2018-32"},{"issue":"8","key":"312_CR6","doi-asserted-by":"publisher","first-page":"2222","DOI":"10.1109\/TASL.2007.907344","volume":"15","author":"T Toda","year":"2007","unstructured":"T. Toda, A.W. Black, K. Tokuda, Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory. IEEE Trans. Audio Speech. Lang. Process. 15(8), 2222\u20132235 (2007)","journal-title":"IEEE Trans. Audio Speech. Lang. Process."},{"key":"312_CR7","unstructured":"D.Y. Wu, Y.H. Chen, H.Y. Lee, Vqvc+: One-shot voice conversion by vector quantization and u-net architecture. arXiv preprint arXiv:2006.04154 (2020)"},{"key":"312_CR8","unstructured":"D.P. Kingma, M. Welling, Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013)"},{"key":"312_CR9","doi-asserted-by":"crossref","unstructured":"L. Wan, Q. Wang, A. Papir, I.L. Moreno, in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Generalized end-to-end loss for speaker verification (IEEE, 2018), pp. 4879\u20134883","DOI":"10.1109\/ICASSP.2018.8462665"},{"key":"312_CR10","doi-asserted-by":"crossref","unstructured":"L. Sun, K. Li, H. Wang, S. Kang, H. Meng, in 2016 IEEE International Conference on Multimedia and Expo (ICME). Phonetic posteriorgrams for many-to-one voice conversion without parallel data training (IEEE, 2016), pp. 1\u20136","DOI":"10.1109\/ICME.2016.7552917"},{"key":"312_CR11","doi-asserted-by":"crossref","unstructured":"W.C. Huang, S.W. Yang, T. Hayashi, H.Y. Lee, S. Watanabe, T. Toda, in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). S3prl-vc: Open-source voice conversion framework with self-supervised speech representations (IEEE, 2022), pp. 6552\u20136556","DOI":"10.1109\/ICASSP43922.2022.9746430"},{"key":"312_CR12","doi-asserted-by":"crossref","unstructured":"S.W. Yang, P.H. Chi, Y.S. Chuang, C.I.J. Lai, K. Lakhotia, Y.Y. Lin, A.T. Liu, J. Shi, X. Chang, G.T. Lin, et al., Superb: Speech processing universal performance benchmark. arXiv preprint arXiv:2105.01051 (2021)","DOI":"10.21437\/Interspeech.2021-1775"},{"key":"312_CR13","doi-asserted-by":"crossref","unstructured":"Y.Y. Lin, C.M. Chien, J.H. Lin, H.y. Lee, L.s. Lee, in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Fragmentvc: Any-to-any voice conversion by end-to-end extracting and fusing fine-grained voice fragments with attention (IEEE, 2021), pp. 5939\u20135943","DOI":"10.1109\/ICASSP39728.2021.9413699"},{"key":"312_CR14","doi-asserted-by":"crossref","unstructured":"J.H. Lin, Y.Y. Lin, C.M. Chien, H.Y. Lee, S2vc: A framework for any-to-any voice conversion with self-supervised pretrained representations. arXiv preprint arXiv:2104.02901 (2021)","DOI":"10.21437\/Interspeech.2021-1356"},{"key":"312_CR15","doi-asserted-by":"crossref","unstructured":"B. van Niekerk, L. Nortje, H. Kamper, Vector-quantized neural networks for acoustic unit discovery in the zerospeech 2020 challenge. arXiv preprint arXiv preprint arXiv:2005.09409 (2020)","DOI":"10.21437\/Interspeech.2020-1693"},{"key":"312_CR16","unstructured":"A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, K. Kavukcuoglu, Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499 (2016)"},{"key":"312_CR17","unstructured":"K. Kumar, R. Kumar, T. De Boissiere, L. Gestin, W.Z. Teoh, J. Sotelo, A. de\u00a0Br\u00e9bisson, Y. Bengio, A.C. Courville, Melgan: Generative adversarial networks for conditional waveform synthesis. Adv. Neural Inf. Process. Syst. 32 (2019)"},{"key":"312_CR18","first-page":"17022","volume":"33","author":"J Kong","year":"2020","unstructured":"J. Kong, J. Kim, J. Bae, Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. Adv. Neural Inf. Process. Syst. 33, 17022\u201317033 (2020)","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"312_CR19","doi-asserted-by":"crossref","unstructured":"A. Graves, S. Fern\u00e1ndez, F. Gomez, J. Schmidhuber, in Proceedings of the 23rd international conference on Machine learning. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks (2006), pp. 369\u2013376","DOI":"10.1145\/1143844.1143891"},{"key":"312_CR20","unstructured":"Y. Ganin, V. Lempitsky, in International conference on machine learning. Unsupervised domain adaptation by backpropagation (PMLR, 2015), pp. 1180\u20131189"},{"key":"312_CR21","doi-asserted-by":"crossref","unstructured":"X. Zhao, F. Liu, C. Song, Z. Wu, S. Kang, D. Tuo, H. Meng, in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Disentangling content and fine-grained prosody information via hybrid asr bottleneck features for voice conversion (IEEE, 2022), pp. 7022\u20137026","DOI":"10.1109\/ICASSP43922.2022.9747625"},{"issue":"6","key":"312_CR22","doi-asserted-by":"publisher","first-page":"1505","DOI":"10.1109\/JSTSP.2022.3188113","volume":"16","author":"S Chen","year":"2022","unstructured":"S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al., Wavlm: Large-scale self-supervised pre-training for full stack speech processing. IEEE J. Sel. Top. Signal Process. 16(6), 1505\u20131518 (2022)","journal-title":"IEEE J. Sel. Top. Signal Process."},{"key":"312_CR23","doi-asserted-by":"publisher","first-page":"3451","DOI":"10.1109\/TASLP.2021.3122291","volume":"29","author":"WN Hsu","year":"2021","unstructured":"W.N. Hsu, B. Bolte, Y.H.H. Tsai, K. Lakhotia, R. Salakhutdinov, A. Mohamed, Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE\/ACM Trans. Audio Speech Lang. Process. 29, 3451\u20133460 (2021)","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"312_CR24","unstructured":"D. Ulyanov, A. Vedaldi, V. Lempitsky, Instance normalization: The missing ingredient for fast stylization. arXiv preprint arXiv:1607.08022 (2016)"},{"key":"312_CR25","doi-asserted-by":"crossref","unstructured":"X. Huang, S. Belongie, in Proceedings of the IEEE international conference on computer vision. Arbitrary style transfer in real-time with adaptive instance normalization (2017), pp. 1501\u20131510","DOI":"10.1109\/ICCV.2017.167"},{"key":"312_CR26","doi-asserted-by":"crossref","unstructured":"J.C. Chou, C.C. Yeh, H.Y. Lee, One-shot voice conversion by separating speaker and content representations with instance normalization. arXiv preprint arXiv:1904.05742 (2019).","DOI":"10.21437\/Interspeech.2019-2663"},{"key":"312_CR27","doi-asserted-by":"crossref","unstructured":"J. Wang, J. Li, X. Zhao, Z. Wu, S. Kang, H. Meng, Adversarially learning disentangled speech representations for robust multi-factor voice conversion. arXiv preprint (2021)","DOI":"10.21437\/Interspeech.2021-1990"},{"key":"312_CR28","doi-asserted-by":"crossref","unstructured":"A.T. Liu, S.W. Yang, P.H. Chi, P.C. Hsu, H.Y. Lee, in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders (IEEE, 2020), pp. 6419\u20136423","DOI":"10.1109\/ICASSP40776.2020.9054458"},{"key":"312_CR29","unstructured":"Y. Wang, D. Stanton, Y. Zhang, R.S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, R.A. Saurous, in International Conference on Machine Learning. Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis (PMLR, 2018), pp. 5180\u20135189"},{"key":"312_CR30","doi-asserted-by":"crossref","unstructured":"P. Wu, Z. Ling, L. Liu, Y. Jiang, H. Wu, L. Dai, in 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC). End-to-end emotional speech synthesis using style tokens and semi-supervised training (IEEE, 2019), pp. 623\u2013627","DOI":"10.1109\/APSIPAASC47483.2019.9023186"},{"key":"312_CR31","doi-asserted-by":"crossref","unstructured":"R. Valle, J. Li, R. Prenger, B. Catanzaro, in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens (IEEE, 2020), pp. 6189\u20136193","DOI":"10.1109\/ICASSP40776.2020.9054556"},{"key":"312_CR32","unstructured":"C. Veaux, J. Yamagishi, K. MacDonald. CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit[J]. University of Edinburgh. The Centre for Speech Technology Research (CSTR). 6, 15 (2017)."},{"key":"312_CR33","unstructured":"J. Kominek, A.W. Black, in Fifth ISCA workshop on speech synthesis. The cmu arctic speech databases (2004)"},{"key":"312_CR34","doi-asserted-by":"publisher","first-page":"103110","DOI":"10.1016\/j.dsp.2021.103110","volume":"116","author":"X Kang","year":"2021","unstructured":"X. Kang, H. Huang, Y. Hu, Z. Huang, Connectionist temporal classification loss for vector quantized variational autoencoder in zero-shot voice conversion. Digit. Signal Process. 116, 103110 (2021)","journal-title":"Digit. Signal Process."},{"key":"312_CR35","first-page":"12449","volume":"33","author":"A Baevski","year":"2020","unstructured":"A. Baevski, Y. Zhou, A. Mohamed, M. Auli, wav2vec 2.0: A framework for self-supervised learning of speech representations. Adv. Neural Inf. Process. Syst. 33, 12449\u201312460 (2020)","journal-title":"Adv. Neural Inf. Process. Syst."}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-023-00312-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-023-00312-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-023-00312-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,1]],"date-time":"2024-11-01T00:11:43Z","timestamp":1730419903000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-023-00312-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,28]]},"references-count":35,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["312"],"URL":"https:\/\/doi.org\/10.1186\/s13636-023-00312-8","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,28]]},"assertion":[{"value":"3 March 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 October 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 October 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"45"}}