{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,27]],"date-time":"2025-07-27T07:38:10Z","timestamp":1753601890199,"version":"3.28.0"},"reference-count":41,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Inf. &amp; Syst."],"published-print":{"date-parts":[[2024,1,1]]},"DOI":"10.1587\/transinf.2023edp7039","type":"journal-article","created":{"date-parts":[[2023,12,31]],"date-time":"2023-12-31T22:39:14Z","timestamp":1704062354000},"page":"93-104","source":"Crossref","is-referenced-by-count":2,"title":["Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis"],"prefix":"10.1587","volume":"E107.D","author":[{"given":"Kenichi","family":"FUJITA","sequence":"first","affiliation":[{"name":"NTT Human Informatics Laboratories, NTT Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Atsushi","family":"ANDO","sequence":"additional","affiliation":[{"name":"NTT Human Informatics Laboratories, NTT Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yusuke","family":"IJIMA","sequence":"additional","affiliation":[{"name":"NTT Human Informatics Laboratories, NTT Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","doi-asserted-by":"crossref","unstructured":"[1] E. Zetterholm, \u201cIntonation pattern and duration differences in imitated speech,\u201d The International Conference on Speech Prosody, 2002.","DOI":"10.21437\/SpeechProsody.2002-168"},{"key":"2","unstructured":"[2] E. Zetterholm, \u201cThe same but different-three impersonators imitate the same target voices,\u201d The International Congress of Phonetic Sciences (ICPhS), 2003."},{"key":"3","doi-asserted-by":"crossref","unstructured":"[3] D. Gomathi, S.A. Thati, K.V. Sridaran, and B. Yegnanarayana, \u201cAnalysis of mimicry speech,\u201d The International Speech Communication Association, pp.695-698, 2012. 10.21437\/interspeech.2012-218","DOI":"10.21437\/Interspeech.2012-218"},{"key":"4","doi-asserted-by":"publisher","unstructured":"[4] N. Hojo, Y. Ijima, and H. Mizuno, \u201cDNN-based speech synthesis using speaker codes,\u201d IEICE Trans. Inf. &amp; Syst., vol.E101-D, no.2, pp.462-472, 2018. 10.1587\/transinf.2017edp7165","DOI":"10.1587\/transinf.2017EDP7165"},{"key":"5","doi-asserted-by":"crossref","unstructured":"[5] E. Cooper, C.-I. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, and J. Yamagishi, \u201cZero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,\u201d 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.6184-6188, IEEE, 2020. 10.1109\/icassp40776.2020.9054535","DOI":"10.1109\/ICASSP40776.2020.9054535"},{"key":"6","doi-asserted-by":"crossref","unstructured":"[6] Y. Fan, Y. Qian, F.K. Soong, and L. He, \u201cMulti-speaker modeling and speaker adaptation for DNN-based TTS synthesis,\u201d 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.4475-4479, IEEE, 2015. 10.1109\/icassp.2015.7178817","DOI":"10.1109\/ICASSP.2015.7178817"},{"key":"7","doi-asserted-by":"crossref","unstructured":"[7] H.-T. Luong and J. Yamagishi, \u201cScaling and bias codes for modeling speaker-adaptive DNN-based speech synthesis systems,\u201d 2018 IEEE Spoken Language Technology Workshop (SLT), pp.610-617, IEEE, 2018. 10.1109\/slt.2018.8639659","DOI":"10.1109\/SLT.2018.8639659"},{"key":"8","doi-asserted-by":"crossref","unstructured":"[8] H.-T. Luong, X. Wang, J. Yamagishi, and N. Nishizawa, \u201cTraining multi-speaker neural text-to-speech systems using speaker-imbalanced speech corpora,\u201d INTERSPEECH 2019, pp.1303-1307, 2019. 10.21437\/interspeech.2019-1311","DOI":"10.21437\/Interspeech.2019-1311"},{"key":"9","doi-asserted-by":"crossref","unstructured":"[9] N. Dehak, P.J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, \u201cFront-end factor analysis for speaker verification,\u201d IEEE Trans. Audio, Speech, Language Process., vol.19, no.4, pp.788-798, 2010. 10.1109\/tasl.2010.2064307","DOI":"10.1109\/TASL.2010.2064307"},{"key":"10","doi-asserted-by":"crossref","unstructured":"[10] G. Heigold, I. Moreno, S. Bengio, and N. Shazeer, \u201cEnd-to-end text-dependent speaker verification,\u201d 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.5115-5119, IEEE, 2016. 10.1109\/icassp.2016.7472652","DOI":"10.1109\/ICASSP.2016.7472652"},{"key":"11","doi-asserted-by":"crossref","unstructured":"[11] D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, \u201cX-vectors: Robust DNN embeddings for speaker recognition,\u201d 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.5329-5333, IEEE, 2018. 10.1109\/icassp.2018.8461375","DOI":"10.1109\/ICASSP.2018.8461375"},{"key":"12","doi-asserted-by":"crossref","unstructured":"[12] Z. Wu, P. Swietojanski, C. Veaux, S. Renals, and S. King, \u201cA study of speaker adaptation for DNN-based speech synthesis,\u201d INTERSPEECH 2015, pp.879-883, 2015. 10.21437\/interspeech.2015-270","DOI":"10.21437\/Interspeech.2015-270"},{"key":"13","doi-asserted-by":"crossref","unstructured":"[13] R. Doddipatla, N. Braunschweiler, and R. Maia, \u201cSpeaker adaptation in DNN-based speech synthesis using d-vectors,\u201d INTERSPEECH 2017, pp.3404-3408, 2017. 10.21437\/interspeech.2017-1038","DOI":"10.21437\/Interspeech.2017-1038"},{"key":"14","unstructured":"[14] Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. Lopez Moreno, and Y. Wu, \u201cTransfer learning from speaker verification to multispeaker text-to-speech synthesis,\u201d Advances in Neural Information Processing Systems (NeurIPS), 2018."},{"key":"15","doi-asserted-by":"crossref","unstructured":"[15] K. Fujita, A. Ando, and Y. Ijima, \u201cPhoneme duration modeling using speech rhythm-based speaker embeddings for multi-speaker speech synthesis,\u201d INTERSPEECH 2021, pp.3141-3145, 2021. 10.21437\/interspeech.2021-826","DOI":"10.21437\/Interspeech.2021-826"},{"key":"16","unstructured":"[16] A. Gibiansky, S. Arik, G. Diamos, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou, \u201cDeep voice 2: Multi-speaker neural text-to-speech,\u201d Advances in Neural Information Processing Systems (NeurIPS), 2017."},{"key":"17","unstructured":"[17] Y. Wang, D. Stanton, Y. Zhang, R.S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R.A. Saurous, \u201cStyle tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,\u201d International Conference on Machine Learning (ICML), pp.5180-5189, 2018."},{"key":"18","doi-asserted-by":"crossref","unstructured":"[18] M.J. Carey, E.S. Parris, H. Lloyd-Thomas, and S.J. Bennett, \u201cRobust prosodic features for speaker identification,\u201d International Conference on Spoken Language Processing (ICSLP), pp.1800-1803, 1996. 10.21437\/icslp.1996-457","DOI":"10.21437\/ICSLP.1996-457"},{"key":"19","doi-asserted-by":"crossref","unstructured":"[19] C. Miyajima, Y. Hattori, K. Tokuda, T. Masuko, T. Kobayashi, and T. Kitamura, \u201cSpeaker identification using gaussian mixture models based on multi-space probability distribution,\u201d 2001 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.433-436, IEEE, 2001. 10.1109\/icassp.2001.940860","DOI":"10.1109\/ICASSP.2001.940860"},{"key":"20","doi-asserted-by":"crossref","unstructured":"[20] S. Wang, J. Rohdin, L. Burget, O. Plchot, Y. Qian, K. Yu, and J.\u010cernock\u00fd, \u201cOn the usage of phonetic information for text-independent speaker embedding extraction,\u201d INTERSPEECH 2019, pp.1148-1152, 2019. 10.21437\/interspeech.2019-3036","DOI":"10.21437\/Interspeech.2019-3036"},{"key":"21","doi-asserted-by":"crossref","unstructured":"[21] Y. Liu, L. He, J. Liu, and M.T. Johnson, \u201cSpeaker embedding extraction with phonetic information,\u201d INTERSPEECH 2018, pp.2247-2251, 2018. 10.21437\/interspeech.2018-1226","DOI":"10.21437\/Interspeech.2018-1226"},{"key":"22","doi-asserted-by":"publisher","unstructured":"[22] Y. Ijima, N. Miyazaki, H. Mizuno, and S. Sakauchi, \u201cStatistical model training technique based on speaker clustering approach for HMM-based speech synthesis,\u201d Speech Communication, vol.71, pp.50-61, 2015. 10.1016\/j.specom.2015.04.003","DOI":"10.1016\/j.specom.2015.04.003"},{"key":"23","unstructured":"[23] Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.Y. Liu, \u201cFastSpeech 2: Fast and high-quality end-to-end text to speech,\u201d International Conference on Learning Representations, 2021."},{"key":"24","doi-asserted-by":"crossref","unstructured":"[24] I. Elias, H. Zen, J. Shen, Y. Zhang, Y. Jia, R.J. Weiss, and Y. Wu, \u201cParallel tacotron: Non-autoregressive and controllable TTS,\u201d 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.5709-5713, 2021. 10.1109\/icassp39728.2021.9414718","DOI":"10.1109\/ICASSP39728.2021.9414718"},{"key":"25","doi-asserted-by":"crossref","unstructured":"[25] J.S. Chung, J. Huh, S. Mun, M. Lee, H.-S. Heo, S. Choe, C. Ham, S. Jung, B.-J. Lee, and I. Han, \u201cIn defence of metric learning for speaker recognition,\u201d INTERSPEECH 2020, pp.2977-2981, 2020. 10.21437\/interspeech.2020-1064","DOI":"10.21437\/Interspeech.2020-1064"},{"key":"26","doi-asserted-by":"crossref","unstructured":"[26] V. Peddinti, D. Povey, and S. Khudanpur, \u201cA time delay neural network architecture for efficient modeling of long temporal contexts,\u201d INTERSPEECH 2015, pp.3214-3218, 2015. 10.21437\/interspeech.2015-647","DOI":"10.21437\/Interspeech.2015-647"},{"key":"27","doi-asserted-by":"publisher","unstructured":"[27] N.J.M.S. Mary, S. Umesh, and S.V. Katta, \u201cS-vectors and tesa: Speaker embeddings and a speaker authenticator based on transformer encoder,\u201d IEEE\/ACM Trans. Audio, Speech, Language Process., vol.30, pp.404-413, 2021. 10.1109\/taslp.2021.3134566","DOI":"10.1109\/TASLP.2021.3134566"},{"key":"28","doi-asserted-by":"crossref","unstructured":"[28] W. Cai, J. Chen, and M. Li, \u201cExploring the encoding layer and loss function in end-to-end speaker and language recognition system,\u201d Speaker Odyssey 2018, pp.74-81, 2018. 10.21437\/odyssey.2018-11","DOI":"10.21437\/Odyssey.2018-11"},{"key":"29","doi-asserted-by":"crossref","unstructured":"[29] F.A. Rezaur rahman Chowdhury, Q. Wang, I.L. Moreno, and L. Wan, \u201cAttention-based models for text-dependent speaker verification,\u201d 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.5359-5363, 2018. 10.1109\/icassp.2018.8461587","DOI":"10.1109\/ICASSP.2018.8461587"},{"key":"30","unstructured":"[30] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, L.u. Kaiser, and I. Polosukhin, \u201cAttention is all you need,\u201d Advances in Neural Information Processing Systems (NeurIPS), 2017."},{"key":"31","unstructured":"[31] J.D.M.W.C. Kenton and L.K. Toutanova, \u201cBERT: Pre-training of deep bidirectional transformers for language understanding,\u201d NAACL-HLT, pp.4171-4186, 2019."},{"key":"32","unstructured":"[32] T. Brown, B. Mann, N. Ryder, M. Subbiah, J.D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, \u201cLanguage models are few-shot learners,\u201d Advances in Neural Information Processing Systems (NeurIPS), pp.1877-1901, 2020."},{"key":"33","doi-asserted-by":"publisher","unstructured":"[33] W.-N. Hsu, B. Bolte, Y.-H.H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, \u201cHuBERT: Self-supervised speech representation learning by masked prediction of hidden units,\u201d IEEE\/ACM Trans. Audio, Speech, Language Process., vol.29, pp.3451-3460, 2021. 10.1109\/taslp.2021.3122291","DOI":"10.1109\/TASLP.2021.3122291"},{"key":"34","doi-asserted-by":"crossref","unstructured":"[34] H. Zen, A. Senior, and M. Schuster, \u201cStatistical parametric speech synthesis using deep neural networks,\u201d 2013 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.7962-7966, 2013. 10.1109\/icassp.2013.6639215","DOI":"10.1109\/ICASSP.2013.6639215"},{"key":"35","doi-asserted-by":"crossref","unstructured":"[35] J. Shen, R. Pang, R.J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R.A. Saurous, Y. Agiomvrgiannakis, and Y. Wu, \u201cNatural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,\u201d 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.4779-4783, 2018. 10.1109\/icassp.2018.8461368","DOI":"10.1109\/ICASSP.2018.8461368"},{"key":"36","unstructured":"[36] J. Kong, J. Kim, and J. Bae, \u201cHiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,\u201d Advances in Neural Information Processing Systems (NeurIPS), vol.33, pp.17022-17033, 2020."},{"key":"37","doi-asserted-by":"crossref","unstructured":"[37] M.S. Han, \u201cAcoustic manifestations of mora timing in japanese,\u201d The Journal of the Acoustical Society of America, vol.96, no.1, pp.73-82, 1994. 10.1121\/1.410376","DOI":"10.1121\/1.410376"},{"key":"38","doi-asserted-by":"publisher","unstructured":"[38] Y. Saito, S. Takamichi, and H. Saruwatari, \u201cPerceptual-similarity-aware deep speaker representation learning for multi-speaker generative modeling,\u201d IEEE\/ACM Trans. Audio, Speech, Language Process., vol.29, pp.1033-1048, 2021. 10.1109\/taslp.2021.3059114","DOI":"10.1109\/TASLP.2021.3059114"},{"key":"39","unstructured":"[39] L. Van der Maaten and G. Hinton, \u201cVisualizing data using t-SNE,\u201d Journal of machine learning research, vol.9, no.11, pp.2579-2605, 2008."},{"key":"40","doi-asserted-by":"crossref","unstructured":"[40] D.N. Reshef, Y.A. Reshef, H.K. Finucane, S.R. Grossman, G. McVean, P.J. Turnbaugh, E.S. Lander, M. Mitzenmacher, and P.C. Sabeti, \u201cDetecting novel associations in large data sets,\u201d Science, vol.334, no.6062, pp.1518-1524, 2011. 10.1126\/science.1205438","DOI":"10.1126\/science.1205438"},{"key":"41","doi-asserted-by":"crossref","unstructured":"[41] Z. Zeng, J. Wang, N. Cheng, T. Xia, and J. Xiao, \u201cAlignTTS: Efficient feed-forward text-to-speech system without explicit alignment,\u201d 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.6714-6718, 2020. 10.1109\/icassp40776.2020.9054119","DOI":"10.1109\/ICASSP40776.2020.9054119"}],"container-title":["IEICE Transactions on Information and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E107.D\/1\/E107.D_2023EDP7039\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,7]],"date-time":"2024-11-07T05:15:06Z","timestamp":1730956506000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E107.D\/1\/E107.D_2023EDP7039\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,1]]},"references-count":41,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024]]}},"URL":"https:\/\/doi.org\/10.1587\/transinf.2023edp7039","relation":{},"ISSN":["0916-8532","1745-1361"],"issn-type":[{"type":"print","value":"0916-8532"},{"type":"electronic","value":"1745-1361"}],"subject":[],"published":{"date-parts":[[2024,1,1]]},"article-number":"2023EDP7039"}}