{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T10:28:33Z","timestamp":1781087313138,"version":"3.54.1"},"reference-count":27,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,2,18]],"date-time":"2021-02-18T00:00:00Z","timestamp":1613606400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,2,18]],"date-time":"2021-02-18T00:00:00Z","timestamp":1613606400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Nowadays automatic speech recognition (ASR) systems can achieve higher and higher accuracy rates depending on the methodology applied and datasets used. The rate decreases significantly when the ASR system is being used with a non-native speaker of the language to be recognized. The main reason for this is specific pronunciation and accent features related to the mother tongue of that speaker, which influence the pronunciation. At the same time, an extremely limited volume of labeled non-native speech datasets makes it difficult to train, from the ground up, sufficiently accurate ASR systems for non-native speakers.In this research, we address the problem and its influence on the accuracy of ASR systems, using the style transfer methodology. We designed a pipeline for modifying the speech of a non-native speaker so that it more closely resembles the native speech. This paper covers experiments for accent modification using different setups and different approaches, including neural style transfer and autoencoder. The experiments were conducted on English language pronounced by Japanese speakers (<jats:italic>UME-ERJ<\/jats:italic>dataset). The results show that there is a significant relative improvement in terms of the speech recognition accuracy. Our methodology reduces the necessity of training new algorithms for non-native speech (thus overcoming the obstacle related to the data scarcity) and can be used as a wrapper for any existing ASR system. The modification can be performed in real time, before a sample is passed into the speech recognition system itself.<\/jats:p>","DOI":"10.1186\/s13636-021-00199-3","type":"journal-article","created":{"date-parts":[[2021,2,20]],"date-time":"2021-02-20T16:02:26Z","timestamp":1613836946000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":36,"title":["Accent modification for speech recognition of non-native speakers using neural style transfer"],"prefix":"10.1186","volume":"2021","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0051-2762","authenticated-orcid":false,"given":"Kacper","family":"Radzikowski","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Le","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Osamu","family":"Yoshie","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Robert","family":"Nowak","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,2,18]]},"reference":[{"key":"199_CR1","unstructured":"W. Xiong, L. Wu, F. Alleva, J. Droppo, X. Huang, A. Stolcke, The Microsoft 2017 Conversational Speech Recognition System (2017). https:\/\/arxiv.org\/abs\/1708.06073."},{"key":"199_CR2","unstructured":"T. T. Ping, Automatic speech recognition for non-native speakers. PhD thesis, Universit\u00e9 Joseph-Fourier - Grenoble (2008)."},{"key":"199_CR3","doi-asserted-by":"crossref","unstructured":"A. Metallinou, J. Cheng, in Fifteenth Annual Conference of the International Speech Communication Association. Using deep neural networks to improve proficiency assessment for children English language learners, (2014).","DOI":"10.21437\/Interspeech.2014-358"},{"key":"199_CR4","doi-asserted-by":"crossref","unstructured":"T. Drugman, T. Dutoit, in INTERSPEECH 2009, 10th Annual Conference of the International Speech Communication Association, Brighton, United Kingdom, September 6-10, 2009. Glottal closure and opening instant detection from speech signals (ISCA, 2009), pp. 2891\u20132894. http:\/\/www.isca-speech.org\/archive\/interspeech_2009\/i09_2891.html.","DOI":"10.21437\/Interspeech.2009-47"},{"key":"199_CR5","doi-asserted-by":"publisher","unstructured":"K. Livescu, J. Glass, in 2000 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.00CH37100), vol. 3. Lexical modeling of non-native speech for automatic speech recognition, (2000), pp. 1683\u20131686. https:\/\/doi.org\/10.1109\/ICASSP.2000.862074.","DOI":"10.1109\/ICASSP.2000.862074"},{"key":"199_CR6","doi-asserted-by":"publisher","unstructured":"T. Tan, L. Besacier, in 2007 IEEE International Conference on Acoustics, Speech and Signal Processing - ICASSP \u201907, vol. 4. Acoustic Model Interpolation for Non-Native Speech Recognition, (2007), pp. IV-1009\u2013IV-1012. https:\/\/doi.org\/10.1109\/ICASSP.2007.367243.","DOI":"10.1109\/ICASSP.2007.367243"},{"key":"199_CR7","unstructured":"L. M. Tomokiyo, Recognizing non-native speech: characterizing and adapting to non-native usage in LVCSR. PhD thesis, Carnegie Mellon University (2001)."},{"issue":"10","key":"199_CR8","doi-asserted-by":"publisher","first-page":"1533","DOI":"10.1109\/TASLP.2014.2339736","volume":"22","author":"O. Abdel-Hamid","year":"2014","unstructured":"O. Abdel-Hamid, A. -R. Mohamed, H. Jiang, L. Deng, G. Penn, D. Yu, Convolutional neural networks for speech recognition. IEEE\/ACM Trans. Audio Speech Lang. Proc.22(10), 1533\u20131545 (2014).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Proc."},{"key":"199_CR9","first-page":"1","volume":"1","author":"N. Dave","year":"2013","unstructured":"N. Dave, Feature extraction methods LPC, PLP and MFCC in speech recognition. Int. J. Adv. Res. Eng. Technol.1:, 1\u20135 (2013).","journal-title":"Int. J. Adv. Res. Eng. Technol."},{"issue":"4","key":"199_CR10","doi-asserted-by":"publisher","first-page":"788","DOI":"10.1109\/TASL.2010.2064307","volume":"19","author":"N. Dehak","year":"2011","unstructured":"N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, P. Ouellet, Front-end factor analysis for speaker verification. Trans. Audio Speech Lang. Proc.19(4), 788\u2013798 (2011).","journal-title":"Trans. Audio Speech Lang. Proc."},{"issue":"1","key":"199_CR11","doi-asserted-by":"publisher","first-page":"151","DOI":"10.1016\/j.csl.2012.01.008","volume":"27","author":"M. Li","year":"2013","unstructured":"M. Li, Han K.J., Narayanan S., Automatic speaker age and gender recognition using acoustic and prosodic level information fusion. Comput. Speech Lang.27(1), 151\u2013167 (2013). https:\/\/doi.org\/10.1016\/j.csl.2012.01.008.","journal-title":"Comput. Speech Lang."},{"key":"199_CR12","doi-asserted-by":"publisher","unstructured":"A. Graves, A. Mohamed, G. Hinton, in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, vol. 38. Speech recognition with deep recurrent neural networks, (2013), pp. 6645\u20136649. https:\/\/doi.org\/10.1109\/ICASSP.2013.6638947.","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"199_CR13","unstructured":"D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos, E. Elsen, J. Engel, L. Fan, C. Fougner, T. Han, A. Hannun, B. Jun, P. LeGresley, L. Lin, S. Narang, A. Ng, S. Ozair, R. Prenger, J. Raiman, S. Satheesh, D. Seetapun, S. Sengupta, Y. Wang, Z. Wang, C. Wang, B. Xiao, D. Yogatama, J. Zhan, Z. Zhu, Deep speech 2: end-to-end speech recognition in English and Mandarin (2015). https:\/\/arxiv.org\/abs\/1512.02595."},{"issue":"6","key":"199_CR14","doi-asserted-by":"publisher","first-page":"82","DOI":"10.1109\/MSP.2012.2205597","volume":"29","author":"G. E. Hinton","year":"2012","unstructured":"G. E. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, B. Kingsbury, Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups. IEEE Signal Process. Mag.29(6), 82\u201397 (2012). https:\/\/doi.org\/10.1109\/MSP.2012.2205597.","journal-title":"IEEE Signal Process. Mag."},{"key":"199_CR15","unstructured":"K. Simonyan, A. Zisserman, Very Deep Convolutional Networks for Large-Scale Image Recognition. CoRR. abs\/1409.1556: (2015)."},{"key":"199_CR16","doi-asserted-by":"publisher","unstructured":"K. Radzikowski, L. Wang, O. Yoshie, R. Nowak, Dual supervised learning for non-native speech recognition. EURASIP J. Audio Speech Music Process.2019(3), 1\u201310 (2019). https:\/\/doi.org\/10.1186\/s13636-018-0146-4. https:\/\/rdcu.be\/bgUxy.","DOI":"10.1186\/s13636-018-0146-4"},{"key":"199_CR17","unstructured":"R. Kacper, W. Le, Y. Osamu, in Proceedings of the Conference of Institute of Electrical Engineers of Japan, Electronics and Information Systems Division. Non-native english speaker\u2019s speech correction, based on domain focused document, (2016)."},{"key":"199_CR18","first-page":"276","volume-title":"Proceedings of the 18th International Conference on Information Integration and Web-based Applications and Services (iiWAS \u201916)","author":"R. Kacper","year":"2016","unstructured":"R. Kacper, W. Le, Y. Osamu, in Proceedings of the 18th International Conference on Information Integration and Web-based Applications and Services (iiWAS \u201916). Non-native english speakers\u2019 speech correction, based on domain focused document (ACMNew York, 2016), pp. 276\u2013281."},{"key":"199_CR19","unstructured":"R. Kacper, W. Le, Y. Osamu, in Proceedings of the conference of institute of electrical engineers of japan, electronics and information systems division. Non-native speech recognition using characteristic speech features, with respect to nationality, (2017)."},{"key":"199_CR20","unstructured":"L. A. Gatys, A. S. Ecker, M. Bethge, A Neural Algorithm of Artistic Style (2015). https:\/\/arxiv.org\/abs\/1508.06576."},{"key":"199_CR21","unstructured":"J. Johnson, A. Alahi, L. Fei-Fei, Perceptual losses for real-time style transfer and super-resolution (2016). https:\/\/arxiv.org\/abs\/1603.08155."},{"key":"199_CR22","doi-asserted-by":"crossref","unstructured":"E. Grinstein, N. Q. K. Duong, A. Ozerov, P. P\u00e9rez, in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Audio Style Transfer (IEEE, 2018). doi:10.1109\/icassp.2018.8461711.","DOI":"10.1109\/ICASSP.2018.8461711"},{"key":"199_CR23","unstructured":"P. Verma, J. O. Smith, neural style transfer for audio spectograms (2018). https:\/\/arxiv.org\/abs\/1801.01589."},{"key":"199_CR24","unstructured":"K. Simonyan, A. Zisserman, Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"199_CR25","unstructured":"F. Chollet, et al., Keras. GitHub (2015). https:\/\/github.com\/fchollet\/keras."},{"key":"199_CR26","doi-asserted-by":"publisher","first-page":"369","DOI":"10.1145\/1143844.1143891","volume-title":"Proceedings of the 23rd International Conference on Machine Learning (ICML \u201906)","author":"A. Graves","year":"2006","unstructured":"A. Graves, S. Fern\u00e1ndez, F. Gomez, J. Schmidhuber, in Proceedings of the 23rd International Conference on Machine Learning (ICML \u201906). Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks (Association for Computing MachineryNew York, 2006), pp. 369\u2013376. https:\/\/doi.org\/10.1145\/1143844.1143891."},{"key":"199_CR27","doi-asserted-by":"crossref","unstructured":"B. Hixon, E. Schneider, S. Epstein, in INTERSPEECH. Phonemic Similarity Metrics to Compare Pronunciation Methods, (2011).","DOI":"10.21437\/Interspeech.2011-305"}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-021-00199-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/s13636-021-00199-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-021-00199-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,18]],"date-time":"2022-12-18T09:10:51Z","timestamp":1671354651000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-021-00199-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,2,18]]},"references-count":27,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["199"],"URL":"https:\/\/doi.org\/10.1186\/s13636-021-00199-3","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,2,18]]},"assertion":[{"value":"18 November 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 January 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 February 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Not applicable.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"11"}}