{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T14:12:40Z","timestamp":1774879960485,"version":"3.50.1"},"reference-count":44,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T00:00:00Z","timestamp":1768780800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T00:00:00Z","timestamp":1768780800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Manipal Academy of Higher Education, Manipal"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Speech Technol"],"published-print":{"date-parts":[[2026,3]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Spoken dialect identification (DID) refers to the automatic identification of dialects in a speech sample. Since different dialects of a given language have very high similarity in terms of vocabulary, grammar, pronunciation, etc., accurate identification of the dialect is very challenging. In addition to this, lack of sufficient training data, which is common in low-resource languages, further complicates the issue. Due to these reasons, state-of-the-art DID systems often perform unsatisfactorily. Using a feature representation of the speech that can efficiently encode DID-specific contents in the input even in low-resource conditions, and using a suitable model architecture that can efficiently capture the DID-specific contents in the given input feature representation, including subtle differences between the dialects, can help address this issue. Motivated by this, in this work, we explore the usage of different frame-level and segment-level features, and explore different model architectures for accurate identification of the dialect. Specifically, apart from commonly used Mel-frequency cepstral coefficients (MFCC) features, we explore the usage of bottleneck features and wav2vec2.0 features (both base and large variants) obtained using pre-trained networks as representation of the speech. Following this, we experiment with state-of-the-art x-vector, BLSTM-based-u-vector, and transformer-based-u-vector which uses attention in a hierarchical manner. This helps to efficiently model the DID-specific contents in the input representation of speech. Experiments conducted on Kannada, Konkani, Tamil and Marathi, which are some low-resource languages of India, indicate that wav2vec2.0 large features, when combined with transformer-based model, are better as compared to other combinations of features and model architectures and provided around 10% improvement compared to baselines in all cases.<\/jats:p>","DOI":"10.1007\/s10772-025-10243-8","type":"journal-article","created":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T12:20:48Z","timestamp":1768825248000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Building spoken dialect identification system in low-resource conditions"],"prefix":"10.1007","volume":"29","author":[{"given":"Ananya","family":"Angra","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"H.","family":"Muralikrishna","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"A. D.","family":"Dileep","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Veena","family":"Thenkanidiyoor","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,1,19]]},"reference":[{"key":"10243_CR1","doi-asserted-by":"crossref","unstructured":"Aizat, K., Mohamed, O., Orken, M., Ainur, A., & Zhumazhanov, B. (2020). Identification and authentication of user voice using DNN features and i-vector. Cogent Engineering, 7(1), 1751557","DOI":"10.1080\/23311916.2020.1751557"},{"key":"10243_CR2","doi-asserted-by":"crossref","unstructured":"Ali, A., Dehak, N., Cardinal, P., Khurana, S., Yella, S. H., & Glass, J., et al. (2015). Automatic dialect detection in arabic broadcast speech. arXiv preprint arXiv:150906928","DOI":"10.21437\/Interspeech.2016-1297"},{"key":"10243_CR3","doi-asserted-by":"crossref","unstructured":"Angra, A., Muralikrishna, H., Dileep, A., & Veena, T. (2024). Exploring aggregated wav2vec 2.0 features and dual-stream TDNN for efficient spoken dialect identification. IEEE Access","DOI":"10.1109\/ACCESS.2024.3523951"},{"key":"10243_CR4","unstructured":"Baevski, A., Zhou, Y., Mohamed, A., & Auli, M. wav2vec 2020, 2.0). A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems, 33, 12449\u201312460"},{"key":"10243_CR5","doi-asserted-by":"crossref","unstructured":"Chambers, J. K., & Trudgill, P. (1998). Dialectology. Cambridge University Press","DOI":"10.1017\/CBO9780511805103"},{"key":"10243_CR6","doi-asserted-by":"crossref","unstructured":"Chittaragi, N. B., & Koolagudi, S. G. (2019). Acoustic-phonetic feature based Kannada dialect identification from vowel sounds. International Journal of Speech Technology, 22, 1099\u20131113","DOI":"10.1007\/s10772-019-09646-1"},{"key":"10243_CR7","doi-asserted-by":"crossref","unstructured":"Chittaragi, N. B., & Koolagudi, S. G. (2020). Automatic dialect identification system for Kannada language using single and ensemble SVM algorithms. Language Resources and Evaluation, 54, 553\u2013585","DOI":"10.1007\/s10579-019-09481-5"},{"key":"10243_CR8","doi-asserted-by":"publisher","first-page":"101230","DOI":"10.1016\/j.csl.2021.101230","volume":"70","author":"N. B. Chittaragi","year":"2021","unstructured":"Chittaragi, N. B., & Koolagudi, S. G. (2021). Dialect identification using chroma-spectral shape features with ensemble technique. Computer Speech & Language, 70, 101230","journal-title":"Computer Speech & Language"},{"key":"10243_CR9","doi-asserted-by":"publisher","first-page":"4289","DOI":"10.1007\/s13369-017-2941-0","volume":"43","author":"N. B. Chittaragi","year":"2018","unstructured":"Chittaragi, N. B., Prakash, A., & Koolagudi, S. G. (2018). Dialect identification using spectral and prosodic features on single and ensemble classifiers. Arabian Journal for Science and Engineering, 43, 4289\u20134302","journal-title":"Arabian Journal for Science and Engineering"},{"key":"10243_CR10","first-page":"160","volume-title":"Linguistic resources for AI\/NLP in Indian languages.","author":"N. Choudhary","year":"2019","unstructured":"Choudhary, N., Rajesha, N., Manasa, G., & Ramamoorthy, L. (2019). LDC-IL raw speech corpora: An overview. In Linguistic resources for AI\/NLP in Indian languages. (pp. 160\u2013174). Mysore: Central Institute of Indian Languages"},{"issue":"1","key":"10243_CR11","doi-asserted-by":"publisher","first-page":"111","DOI":"10.1016\/S0095-4470(03)00009-3","volume":"32","author":"C. G. Clopper","year":"2004","unstructured":"Clopper, C. G., & Pisoni, D. B. (2004). Some acoustic cues for the perceptual categorization of American English regional dialects. Journal of Phonetics, 32(1), 111\u2013140","journal-title":"Journal of Phonetics"},{"issue":"6","key":"10243_CR12","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3523179","volume":"21","author":"S. Dey","year":"2022","unstructured":"Dey, S., Sahidullah, M., & Saha, G. (2022). An overview of Indian spoken language recognition from Machine learning perspective. ACM Transactions on Asian and Low-Resource Language Information Processing, 21(6), 1\u201345","journal-title":"ACM Transactions on Asian and Low-Resource Language Information Processing"},{"key":"10243_CR13","doi-asserted-by":"crossref","unstructured":"Feng, K., & Chaspari, T. (2019). Low-resource language identification from speech using transfer learning. In 2019 IEEE 29th international workshop on machine learning for signal processing (MLSP) (pp. 1\u20136). IEEE","DOI":"10.1109\/MLSP.2019.8918833"},{"key":"10243_CR14","doi-asserted-by":"publisher","first-page":"252","DOI":"10.1016\/j.csl.2017.06.008","volume":"46","author":"R. Fer","year":"2017","unstructured":"Fer, R., Mat\u011bjka, P., Gr\u00e9zl, F., Plchot, O., Vesel\u00fd, K., & \u010cernock\u00fd, J. H. (2017). Multilingually trained bottleneck features in spoken language recognition. Computer Speech & Language, 46, 252\u2013267","journal-title":"Computer Speech & Language"},{"key":"10243_CR15","doi-asserted-by":"crossref","unstructured":"Gothi, R., & Rao, P. (2023). Improving Automatic speech recognition with dialect-specific language models. In International conference on speech and computer (pp. 57\u201367). Springer.","DOI":"10.1007\/978-3-031-48309-7_5"},{"key":"10243_CR16","doi-asserted-by":"crossref","unstructured":"He, L. M., Yang, X. B., & Lu, H. J. (2007). A comparison of support vector machines ensemble for classification. In International conference on machine learning and cybernetics (pp. 3613\u20133617). IEEE.","DOI":"10.1109\/ICMLC.2007.4370773"},{"issue":"6","key":"10243_CR17","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3712060","volume":"57","author":"A. Joshi","year":"2025","unstructured":"Joshi, A., Dabre, R., Kanojia, D., Li, Z., Zhan, H., Haffari, G., & Dippold, D. (2025). Natural language processing for dialects of a language: A survey. ACM Computing Surveys, 57(6), 1\u201337","journal-title":"ACM Computing Surveys"},{"key":"10243_CR18","unstructured":"Kannada, W. (2024) [Online; accessed 2024]. Available from: https:\/\/en.wikipedia.org\/wiki\/Kannada"},{"key":"10243_CR19","unstructured":"Language, W. K. (2024). [Online; accessed 2024]. Available from: https:\/\/en.wikipedia.org\/wiki\/Konkani_language"},{"key":"10243_CR20","doi-asserted-by":"crossref","unstructured":"Lee, J. W., Kim, E., Koo, J., & Lee, K. (2022. Available from: Representation selective self-distillation and wav2vec 2.0 feature exploration for spoof-aware speaker verification. Interspeech. https:\/\/api.semanticscholar.org\/CorpusID:247996736","DOI":"10.21437\/Interspeech.2022-11460"},{"key":"10243_CR21","doi-asserted-by":"crossref","unstructured":"Lee, Y., Greenberg, C., Mason, L., & Singer, E. (2022) The 2022 NIST language recognition evaluation plan (LRE22). NIST.","DOI":"10.21437\/Interspeech.2023-241"},{"key":"10243_CR22","doi-asserted-by":"crossref","unstructured":"Lin, W., Madhavi, M., Das, R. K., & Li, H. (2020). Transformer-based Arabic dialect identification. In 2020 International conference on Asian language processing (IALP) (pp. 192\u2013196). IEEE.","DOI":"10.1109\/IALP51396.2020.9310504"},{"key":"10243_CR23","unstructured":"Marathi, W. (2024) [Online; accessed 2024]. Available from: https:\/\/en.wikipedia.org\/wiki\/Marathi_language"},{"key":"10243_CR24","doi-asserted-by":"crossref","unstructured":"Monteiro, S., Angra, A., M, H., Thenkanidiyoor, V., & Dileep, A. D. (2023). Exploring the impact of different approaches for spoken dialect identification of Konkani language. In International conference on speech and computer (pp. 461\u2013474). Springer.","DOI":"10.1007\/978-3-031-48312-7_37"},{"key":"10243_CR25","doi-asserted-by":"crossref","unstructured":"Mothukuri, S. K. P., Hegde, P., Chittaragi, N. B., & Koolagudi, S. G. (2020). Kannada dialect classification using artificial neural networks. In 2020 International conference on artificial intelligence and signal processing (AISP) (pp. 1\u20135). IEEE.","DOI":"10.1109\/AISP48273.2020.9073178"},{"key":"10243_CR26","doi-asserted-by":"crossref","unstructured":"Muralikrishna, H., Gupta, S., Dileep, A. D., & Rajan, P. (2021). Noise-robust spoken language identification using language relevance factor based embedding. In 2021 IEEE spoken language technology workshop (SLT) (pp. 644\u2013651). IEEE.","DOI":"10.1109\/SLT48900.2021.9383503"},{"key":"10243_CR27","doi-asserted-by":"crossref","unstructured":"Muralikrishna, H., Kapoor, S., Dileep, A. D., & Rajan, P. (2021). Spoken language identification in unseen target domain using within-sample similarity loss. In 2021 IEEE international conference on acoustics, speech and signal processing (ICASSP 2021) (pp. 7223\u20137227). IEEE.","DOI":"10.1109\/ICASSP39728.2021.9414090"},{"key":"10243_CR28","doi-asserted-by":"crossref","unstructured":"Narayanan, A., Misra, A., Sim, K. C., Pundak, G., Tripathi, A., Elfeky, M., Haghani, P., Strohman, T., & Bacchiani, M. (2018). Toward domain-invariant speech recognition via large scale training. In 2018 IEEE spoken language technology workshop (SLT) (pp. 441\u2013447). IEEE.","DOI":"10.1109\/SLT.2018.8639610"},{"key":"10243_CR29","unstructured":"Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, E. (2012). Scikit-learn: Machine learning in python. Journal of Machine Learning Research, 12(85), 2825-2830."},{"key":"10243_CR30","volume-title":"Marathi raw speech corpus","author":"L. Ramamoorthy","year":"2019","unstructured":"Ramamoorthy, L., Choudhary, N., Apine, G. R., & Betkekar, A. P. (2019a). Marathi raw speech corpus. Central Institute of Indian Languages, Mysore"},{"key":"10243_CR31","unstructured":"Ramamoorthy, L., Choudhary, N., Patil, V. F., Baji, C. S., Abhyankar, M. N., Rajesha, N., & Manasa, G. (2019b). Kannada raw speech corpus. Mysore: Central Institute of Indian Languages. Dataset"},{"key":"10243_CR32","volume-title":"Konkani raw speech corpus","author":"L. Ramamoorthy","year":"2019","unstructured":"Ramamoorthy, L., Choudhary, N., Varik, S., & Tanawade, R. S. (2019c). Konkani raw speech corpus. Central Institute of Indian Languages, Mysore."},{"key":"10243_CR33","unstructured":"Ramamoorthy, L., Narayan, C., Thennarasu, S., Prem Kumar, L. R., Amudha, R., & Prabagaran, R., & Shrikanth, D. (2021). Tamil Raw Speech Corpus"},{"key":"10243_CR34","unstructured":"Rao, K. S., & Koolagudi, S. G. (2011). Identification of hindi dialects and emotions using spectral and prosodic features of speech. IJSCI: International Journal of Systemics, Cybernetics and Informatics, 9(4), 24\u201333"},{"key":"10243_CR35","doi-asserted-by":"crossref","unstructured":"Schneider, S., Baevski, A., Collobert, R., & Auli, M. (2019). wav2vec: Unsupervised pre-training for speech recognition. In Interspeech 2019 (pp. 3465\u20133469).","DOI":"10.21437\/Interspeech.2019-1873"},{"key":"10243_CR36","doi-asserted-by":"crossref","unstructured":"Sharma, M. (2022). Multi-lingual multi-task speech emotion recognition using wav2vec 2.0. In 2022 IEEE international conference on acoustics, speech and signal processing (ICASSP 2022) (pp. 6907\u20136911). IEEE.","DOI":"10.1109\/ICASSP43922.2022.9747417"},{"key":"10243_CR37","doi-asserted-by":"crossref","unstructured":"Shon, S., Ali, A., & Glass, J. (2018). Convolutional neural networks and language embeddings for end-to-end dialect recognition. arXiv preprint arXiv:180304567","DOI":"10.21437\/Odyssey.2018-14"},{"key":"10243_CR38","doi-asserted-by":"crossref","unstructured":"Shon, S., Ali, A., & Glass, J. (2019). Domain attentive fusion for end-to-end dialect identification with unknown target domain. In 2019 IEEE international conference on acoustics, speech and signal processing (ICASSP 2019) (pp. 5951\u20135955). IEEE.","DOI":"10.1109\/ICASSP.2019.8682533"},{"key":"10243_CR39","doi-asserted-by":"crossref","unstructured":"Singla, Y. K., Shah, J., Chen, C., & Shah, R. R. (2022). What do audio transformers hear? Probing their representations for language delivery & structure. In 2022 IEEE international conference on data mining workshops (ICDMW) (pp. 910\u2013925).","DOI":"10.1109\/ICDMW58026.2022.00120"},{"key":"10243_CR40","doi-asserted-by":"crossref","unstructured":"Snyder, D., Garcia-Romero, D., McCree, A., Sell, G., Povey, D., & Khudanpur, S. (2018a). Spoken language recognition using x-vectors. Odyssey, 105\u2013111","DOI":"10.21437\/Odyssey.2018-15"},{"key":"10243_CR41","doi-asserted-by":"crossref","unstructured":"Snyder, D., Garcia-Romero, D., Sell, G., Povey, D., & Khudanpur, S. (2018b). X-vectors: Robust DNN embeddings for speaker recognition. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP 2018) (pp. 5329\u20135333). IEEE.","DOI":"10.1109\/ICASSP.2018.8461375"},{"key":"10243_CR42","unstructured":"Tamil, W.;. (2024) [Online; accessed 2024]. Available from: https:\/\/en.wikipedia.org\/wiki\/Tamil_language#:text=Tamil%20is%20the%20official%20language,the%20official%20languages%20of%20Singapore"},{"key":"10243_CR43","unstructured":"Torres-Carrasquillo, P. A., Gleason, T. P., & Reynolds, D. A. (2004). Dialect identification using gaussian mixture models. Odyssey. Citeseer, 297\u2013300"},{"key":"10243_CR44","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30"}],"container-title":["International Journal of Speech Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10772-025-10243-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10772-025-10243-8","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10772-025-10243-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T13:23:24Z","timestamp":1774877004000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10772-025-10243-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,19]]},"references-count":44,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,3]]}},"alternative-id":["10243"],"URL":"https:\/\/doi.org\/10.1007\/s10772-025-10243-8","relation":{},"ISSN":["1381-2416","1572-8110"],"issn-type":[{"value":"1381-2416","type":"print"},{"value":"1572-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,19]]},"assertion":[{"value":"25 August 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 December 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 January 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"23"}}