{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,19]],"date-time":"2026-02-19T15:31:03Z","timestamp":1771515063766,"version":"3.50.1"},"reference-count":30,"publisher":"MDPI AG","issue":"18","license":[{"start":{"date-parts":[[2021,9,18]],"date-time":"2021-09-18T00:00:00Z","timestamp":1631923200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The performance of voice-controlled systems is usually influenced by accented speech. To make these systems more robust, frontend accent recognition (AR) technologies have received increased attention in recent years. As accent is a high-level abstract feature that has a profound relationship with language knowledge, AR is more challenging than other language-agnostic audio classification tasks. In this paper, we use an auxiliary automatic speech recognition (ASR) task to extract language-related phonetic features. Furthermore, we propose a hybrid structure that incorporates the embeddings of both a fixed acoustic model and a trainable acoustic model, making the language-related acoustic feature more robust. We conduct several experiments on the AESRC dataset. The results demonstrate that our approach can obtain an 8.02% relative improvement compared with the Transformer baseline, showing the merits of the proposed method.<\/jats:p>","DOI":"10.3390\/s21186258","type":"journal-article","created":{"date-parts":[[2021,9,21]],"date-time":"2021-09-21T22:35:20Z","timestamp":1632263720000},"page":"6258","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":14,"title":["Accent Recognition with Hybrid Phonetic Features"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7166-757X","authenticated-orcid":false,"given":"Zhan","family":"Zhang","sequence":"first","affiliation":[{"name":"Department of Information and Electronic Engineering, Zhejiang University, Hangzhou 310007, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuehai","family":"Wang","sequence":"additional","affiliation":[{"name":"Department of Information and Electronic Engineering, Zhejiang University, Hangzhou 310007, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianyi","family":"Yang","sequence":"additional","affiliation":[{"name":"Department of Information and Electronic Engineering, Zhejiang University, Hangzhou 310007, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,9,18]]},"reference":[{"key":"ref_1","unstructured":"Chu, X., Combs, E., Wang, A., and Picheny, M. (2021, July 30). Accented Speech Recognition Inspired by Human Perception. Available online: https:\/\/arxiv.org\/pdf\/2104.04627.pdf."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Huang, C., Chen, T., Li, S., Chang, E., and Zhou, J. (2001, January 3\u20137). Analysis of speaker variability. Proceedings of the 7th European Conference on Speech Communication and Technology, Aalborg, Denmark.","DOI":"10.21437\/Eurospeech.2001-356"},{"key":"ref_3","unstructured":"Levis, J., and Barriuso, T. (2011, January 16\u201317). Nonnative speakers\u2019 pronunciation errors in spoken and read English. Proceedings of the 3rd Pronunciation in Second Language Learning and Teaching (PSLLTP), Ames, IA, USA."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Viglino, T., Motlicek, P., and Cernak, M. (2019, January 15\u201319). End-to-end accented speech recognition. Proceedings of the 20th Annual Conference of the International Speech Communication Association, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-2122"},{"key":"ref_5","unstructured":"Crawshaw, M. (2021, July 29). Multi-Task Learning with Deep Neural Networks: A Survey. Available online: https:\/\/arxiv.org\/abs\/2009.09796."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Shi, X., Yu, F., Lu, Y., Liang, Y., Feng, Q., Wang, D., Qian, Y., and Xie, L. (2021, January 6\u201311). The accented english speech recognition challenge 2020: Open datasets, tracks, baselines, results and methods. Proceedings of the 2021 IEEE International Conference on Acoustics, Speech and Signal Processing, Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9413386"},{"key":"ref_7","unstructured":"Naranjo-Alcazar, J., Perez-Castanos, S., Zuccarello, P., and Cobos, M. (2021, July 28). CNN Depth Analysis with Different Channel Inputs for Acoustic Scene Classification. Available online: https:\/\/arxiv.org\/abs\/1906.04591."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Wu, Y., and Lee, T. (2019, January 12\u201317). Enhancing Sound Texture in CNN-based Acoustic Scene Classification. Proceedings of the 2019 IEEE International Conference on Acoustics, Speech and Signal Processing, Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683490"},{"key":"ref_9","unstructured":"Zheng, W., Mo, Z., Xing, X., and Zhao, G. (2021, July 30). CNNs-Based Acoustic Scene Classification Using Multi-Spectrogram Fusion and Label Expansions. Available online: https:\/\/arxiv.org\/abs\/1809.01543."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Behravan, H., Hautama, V., Siniscalchi, S.M., Kinnunen, T., and Lee, C.H. (2014, January 4\u20139). Introducing attribute features to foreign accent recognition. Proceedings of the 2014 IEEE International Conference on Acoustics, Speech and Signal Processing, Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6854621"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1007\/s10772-015-9328-y","article-title":"MFCC-GMM based accent recognition system for Telugu speech signals","volume":"19","author":"Mannepalli","year":"2016","journal-title":"Int. J. Speech Technol."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1109\/TASLP.2015.2489558","article-title":"I-Vector Modeling of Speech Attributes for Automatic Foreign Accent Recognition","volume":"24","author":"Behravan","year":"2016","journal-title":"Ieee\/Acm Trans. Audio Speech Lang. Process."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Chen, J., Cai, W., Cai, D., Cai, Z., Zhong, H., and Li, M. (2018, January 26\u201329). End-to-end language identification using NetFV and NetVLAD. Proceedings of the 2018 11th International Symposium on Chinese Spoken Language Processing, Taipei, Taiwan.","DOI":"10.1109\/ISCSLP.2018.8706687"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Cai, W., Chen, J., and Li, M. (2018, January 26\u201329). Exploring the Encoding Layer and Loss Function in End-to-End Speaker and Language Recognition System. Proceedings of the Odyssey 2018 The Speaker and Language Recognition Workshop, Les Sables d\u2019Olonne, France.","DOI":"10.21437\/Odyssey.2018-11"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1016\/j.specom.2020.05.003","article-title":"Automatic accent identification as an analytical tool for accent robust automatic speech recognition","volume":"122","author":"Najafian","year":"2020","journal-title":"Speech Commun."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Weninger, F., Sun, Y., Park, J., Willett, D., and Zhan, P. (2019, January 15\u201319). Deep learning based Mandarin accent identification for accent robust ASR. Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-2737"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Li, J., Lavrukhin, V., Ginsburg, B., Leary, R., Kuchaiev, O., Cohen, J.M., Nguyen, H., and Gadde, R.T. (2019, January 15\u201319). Jasper An End-to-End Convolutional Neural Acoustic Model. Proceedings of the Annual Conference of the International Speech Communication Association, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-1819"},{"key":"ref_18","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Snyder, D., Garcia-Romero, D., Sell, G., Povey, D., and Khudanpur, S. (2018, January 15\u201320). X-vectors: Robust dnn embeddings for speaker recognition. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461375"},{"key":"ref_20","first-page":"369","article-title":"Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks","volume":"148","author":"Graves","year":"2006","journal-title":"Acm Int. Conf. Proc. Ser."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201322). Squeeze-and-Excitation Networks. Proceedings of the IEEE conference on computer vision and pattern recognition, Piscataway, NJ, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European conference on computer vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. (2015, January 19\u201324). Librispeech: An ASR corpus based on public domain audio books. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, South Brisbane, QLD, Australia.","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"ref_24","first-page":"8026","article-title":"Pytorch: An imperative style, high-performance deep learning library","volume":"32","author":"Paszke","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_25","unstructured":"Povey, D., Ghoshal, A., Boulianne, G., Burget, L., Glembek, O., Goel, N., Hannemann, M., Motl\u00ed\u010dek, P., Qian, Y., and Schwarz, P. (2011, January 11\u201315). The Kaldi speech recognition toolkit. Proceedings of the IEEE 2011 Workshop on Automatic Speech Recognition and Understanding, Big Island, HI, USA."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Park, D.S., Chan, W., Zhang, Y., Chiu, C.C., Zoph, B., Cubuk, E.D., and Le, Q.V. (2019, January 15\u201319). SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. Proceedings of the Annual Conference of the International Speech Communication Association, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-2680"},{"key":"ref_27","first-page":"2579","article-title":"Visualizing Data using t-SNE","volume":"9","author":"Hinton","year":"2008","journal-title":"J. Mach. Learn. Res."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Hamooni, H., and Mueen, A. (2014, January 14\u201317). Dual-Domain Hierarchical Classification of Phonetic Time Series. Proceedings of the 2014 IEEE International Conference on Data Mining, Shenzhen, China.","DOI":"10.1109\/ICDM.2014.92"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Dekel, O., Keshet, J., and Singer, Y. (2005, January 11\u201313). An Online Algorithm for Hierarchical Phoneme Classification. Proceedings of the International Workshop Machine Learning for Multimodal Interaction, Edinburgh, UK.","DOI":"10.1007\/978-3-540-30568-2_13"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1641","DOI":"10.1109\/29.46546","article-title":"Speaker-independent phone recognition using hidden Markov models","volume":"37","author":"Lee","year":"1989","journal-title":"IEEE Trans. Acoust. Speech Signal Process."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/18\/6258\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:01:41Z","timestamp":1760166101000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/18\/6258"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,18]]},"references-count":30,"journal-issue":{"issue":"18","published-online":{"date-parts":[[2021,9]]}},"alternative-id":["s21186258"],"URL":"https:\/\/doi.org\/10.3390\/s21186258","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,9,18]]}}}