{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T07:11:14Z","timestamp":1777705874351,"version":"3.51.4"},"reference-count":49,"publisher":"SAGE Publications","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IFS"],"published-print":{"date-parts":[[2023,6,1]]},"abstract":"<jats:p>\u00a0Speech Recognition and its potential applications in terms of \u201ctalking devices\u201d have become indispensable in today\u2019s world. Technological advances like mobiles, smart home assistants or tablets extensively use the techniques of automatic speech recognition that works good for adults but cannot always follow and understand children\u2019s speech. The primary goal of this paper is to bridge the gap of communication between voice assistants and Indian children speaking English as secondary language. The issue of lack of children\u2019s speech corpora with English as non-native language, is addressed by creating a dataset of children in the age group of 5-15 years, speaking Hindi or Marathi as their mother tongue and English as their second language. The analysis and implementation of the proposed work shows the accuracy of approximately 96% and potential for further scope by increasing the size of dataset in lower age group. The key contributions of our work are (i) creating speech dataset of Indian children whose mother-tongue is Hindi or Marathi, (ii) employing and evaluating hybrid Convolutional Neural Network (CNN) as an age classifier, (iii) language modeling to customize children vocabulary, (iv) checking accuracy and performance of the system.<\/jats:p>","DOI":"10.3233\/jifs-224472","type":"journal-article","created":{"date-parts":[[2023,5,5]],"date-time":"2023-05-05T12:16:14Z","timestamp":1683288974000},"page":"10799-10813","source":"Crossref","is-referenced-by-count":0,"title":["An approach for Correcting the Word-level Mispronunciations for non-native English-speaking Indian Children"],"prefix":"10.1177","volume":"44","author":[{"given":"Neha","family":"Kasture","sequence":"first","affiliation":[{"name":"Department of Computer Science and Engineering, Indian Institute of Information Technology, Nagpur, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pooja","family":"Jain","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Indian Institute of Information Technology, Nagpur, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","reference":[{"key":"10.3233\/JIFS-224472_ref1","doi-asserted-by":"publisher","first-page":"473","DOI":"10.15388\/infedu.2020.21","article-title":"Voice Assistants and SmartSpeakers in Everyday Life and in Education","volume":"19","author":"Terzopoulos","year":"2020","journal-title":"Informatics inEducation"},{"key":"10.3233\/JIFS-224472_ref2","unstructured":"Purington A. , Taft J.G. , Sannon S. , Bazarova N.N. , Taylor S. Hardman Alexa is my new BFF\u2019: Social Roles, User Satisfaction, and Personification of the Amazon Echo."},{"key":"10.3233\/JIFS-224472_ref3","unstructured":"Potamianos A. , Narayanan S. , Lee S. Automatic speech recognition for children. [Online]. Available: http:\/\/www.isca-speech.org\/archive"},{"key":"10.3233\/JIFS-224472_ref4","doi-asserted-by":"crossref","unstructured":"Lee S. , Potamianos A. , Narayanan S. Acoustics of children\u2019sspeech: Developmental changes of temporal and spectral parameters a), 1999. [Online]. Available: http:\/\/acousticalsociety.org\/content\/terms.","DOI":"10.1121\/1.426686"},{"key":"10.3233\/JIFS-224472_ref5","first-page":"1","article-title":"Development of speech sounds in children","volume":"257","author":"Eguchi","year":"1969","journal-title":"Acta Otolaryngol Suppl"},{"key":"10.3233\/JIFS-224472_ref6","unstructured":"Shivakumar P. Gurunath , Potamianos A., Lee S. and Narayanan S., Improving Speech Recognition for Children using Acoustic Adaptationand Pronunciation Modeling."},{"key":"10.3233\/JIFS-224472_ref7","doi-asserted-by":"crossref","unstructured":"Liao H. et al., Large Vocabulary Automatic Speech Recognition forChildren, 2015.","DOI":"10.21437\/Interspeech.2015-373"},{"key":"10.3233\/JIFS-224472_ref8","doi-asserted-by":"crossref","unstructured":"Palaz D. , Collobert R. , Magimai-Doss M. Estimating Phoneme ClassConditional Probabilities from Raw Speech Signal using ConvolutionalNeural Networks, Apr. 2013, [Online]. Available: http:\/\/arxiv.org\/abs\/1304.1018.","DOI":"10.21437\/Interspeech.2013-438"},{"key":"10.3233\/JIFS-224472_ref9","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2015-6"},{"issue":"4","key":"10.3233\/JIFS-224472_ref10","doi-asserted-by":"publisher","first-page":"215","DOI":"10.1250\/ast.33.215","article-title":"Development of vocal tract and acousticfeatures in children","volume":"33","author":"Mugitani","year":"2012","journal-title":"Acoustical Science and Technology"},{"key":"10.3233\/JIFS-224472_ref11","unstructured":"Gerosa M. , Lee S. , Giuliani D. , Narayanan S. Analyzing children\u2019s speech: an acoustic study of consonants and consonant-vowel transition."},{"key":"10.3233\/JIFS-224472_ref12","doi-asserted-by":"publisher","first-page":"521S","DOI":"10.1093\/ajcn\/72.2.521S","article-title":"Growth and pubertal developmentin children and adolescents: Effects of diet and physical activity","volume":"72","author":"Rogol","year":"2000","journal-title":"Am J Clin Nutr"},{"key":"10.3233\/JIFS-224472_ref13","doi-asserted-by":"crossref","first-page":"349","DOI":"10.1109\/ICASSP.1996.541104","article-title":"A study of speech recognition forchildren and the elderly","volume":"1","author":"Wilpon","year":"1996","journal-title":"1996 IEEE International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings"},{"key":"10.3233\/JIFS-224472_ref14","doi-asserted-by":"crossref","unstructured":"Narayanan S. , Potamianos A. Creating Conversational Interfacesfor Children, 2002.","DOI":"10.1109\/89.985544"},{"key":"10.3233\/JIFS-224472_ref15","doi-asserted-by":"crossref","unstructured":"Li Q. , Russell M. An analysis of the causes of increased errorrates in children2s speech recognition., Jul. 2002.","DOI":"10.21437\/ICSLP.2002-221"},{"key":"10.3233\/JIFS-224472_ref16","doi-asserted-by":"publisher","DOI":"10.1145\/1640377.1640384"},{"key":"10.3233\/JIFS-224472_ref17","unstructured":"Gray S.S. , Willett D. , Lu J. , Pinto J. , Maergner P. and Bodenstab N., Child automatic speech recognition for US English: childinteraction with living-room-electronic-devices, 2014."},{"key":"10.3233\/JIFS-224472_ref18","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-45243-0_76"},{"key":"10.3233\/JIFS-224472_ref19","unstructured":"Booth E. , Carns J. , Kennington C. , Rafla N. Evaluating and Improving Child-Directed Automatic Speech Recognition, 2020 [Online]. Available: https:\/\/cloud.google.com\/speech-to-text\/."},{"key":"10.3233\/JIFS-224472_ref20","unstructured":"Amodei D. et al., Deep Speech 2: End-to-End Speech Recognition inEnglish and Mandarin, Dec. 2015, [Online]. Available: http:\/\/arxiv.org\/abs\/1512.02595"},{"key":"10.3233\/JIFS-224472_ref21","doi-asserted-by":"crossref","unstructured":"Matassoni M. , Gretter R. , Falavigna D. , Giuliani D. Non-native children speech recognition through transfer learning, Sep. 2018, [Online]. Available: http:\/\/arxiv.org\/abs\/1809.09658.","DOI":"10.1109\/ICASSP.2018.8462059"},{"key":"10.3233\/JIFS-224472_ref22","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2020.101077"},{"issue":"3","key":"10.3233\/JIFS-224472_ref23","doi-asserted-by":"publisher","first-page":"489","DOI":"10.1007\/s10772-015-9291-7","article-title":"Pitch Adaptive MFCC Features for Improving Children\u2019s Mismatched ASR","volume":"18","author":"Ghai","year":"2015","journal-title":"Int. J. Speech Technol."},{"key":"10.3233\/JIFS-224472_ref24","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2016-1020"},{"issue":"2","key":"10.3233\/JIFS-224472_ref25","doi-asserted-by":"crossref","first-page":"161","DOI":"10.1006\/csla.2000.0139","article-title":"TheSTAR system: an interactive pronunciation tutor for young children","volume":"14","author":"Russell","year":"2000","journal-title":"Computer Speech & Language"},{"key":"10.3233\/JIFS-224472_ref26","doi-asserted-by":"crossref","unstructured":"Mostow J. , Roth S.F. , Hauptmann A.6 and Kane M., A PrototypeReading Coach that Listens, 1994. [Online]. Available: www.aaai.org","DOI":"10.1145\/215585.215665"},{"key":"10.3233\/JIFS-224472_ref27","unstructured":"Serizel R. , Giuliani D. Deep neural network adaptation forchildren\u2019s and adults\u2019 speech recognition. [Online]. Available: https:\/\/hal.archives-ouvertes.fr\/hal-01393975"},{"key":"10.3233\/JIFS-224472_ref28","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2015-378"},{"key":"10.3233\/JIFS-224472_ref29","unstructured":"Ossama A.-H. , Mohamed A.-R. , Jiang H. , Penn G. Applying convolutional neural networks concepts to hybrid NN-HMM model for speech recognition. [Online]. Available: http:\/\/www.ldc.upenn.edu\/Catalog\/CatalogEntry.jsp?catalogId=LDC93S1"},{"key":"10.3233\/JIFS-224472_ref30","unstructured":"Booth E. , Carns J. , Kennington C. , Rafla N. Evaluating andImproving Child-Directed Automatic Speech Recognition, [Online]. 2020. Available: https:\/\/cloud.google.com\/speech-to-text\/"},{"issue":"1","key":"10.3233\/JIFS-224472_ref31","doi-asserted-by":"publisher","first-page":"184","DOI":"10.1080\/02564602.2020.1824623","article-title":"A particular character speechsynthesis system based on deep learning","volume":"38","author":"Mei","year":"2021","journal-title":"IETE Technical Review"},{"issue":"2","key":"10.3233\/JIFS-224472_ref32","doi-asserted-by":"publisher","first-page":"348","DOI":"10.1109\/TASL.2010.2047812","article-title":"A GenerativeStudent Model for Scoring Word Reading Skills","volume":"19","author":"Tepperman","year":"2011","journal-title":"IEEETransactions on Audio, Speech, and Language Processing"},{"key":"10.3233\/JIFS-224472_ref33","doi-asserted-by":"publisher","DOI":"10.1109\/NCC.2017.8077101"},{"key":"10.3233\/JIFS-224472_ref34","doi-asserted-by":"crossref","unstructured":"Lovato S. , Piper A.M. , Wartella E. Hey Google, Do UnicornsExist?: Conversational Agents as a Path to Answers to Children\u2019sQuestions, Proc 18th ACM Int Conf Interact Des Child, (2019), 2019.","DOI":"10.1145\/3311927.3323150"},{"key":"10.3233\/JIFS-224472_ref35","doi-asserted-by":"crossref","unstructured":"Shobaki K. , Hosom J.-P. , Cole R. The OGI kids\u2019 speech corpus andrecognizers, Jul. (2000), pp. 258\u2013261.","DOI":"10.21437\/ICSLP.2000-800"},{"key":"10.3233\/JIFS-224472_ref36","unstructured":"and K.J.-P.H. , Shobaki R.C. CSLU: Kids\u2019 Speech Version 1.1 \u2013 Linguistic Data Consortiumhttps."},{"key":"10.3233\/JIFS-224472_ref37","doi-asserted-by":"publisher","DOI":"10.21437\/interspeech.2009-209"},{"key":"10.3233\/JIFS-224472_ref38","first-page":"1559","article-title":"Support vector machines versus fast scoring in thelow-dimensional total variability space for speaker verification, in Proceedings of the Annual Conference of the International Speech Communication Association,\u2013","volume":"1","author":"Dehak","year":"2009","journal-title":"INTER SPEECH"},{"key":"10.3233\/JIFS-224472_ref39","doi-asserted-by":"publisher","DOI":"10.1002\/9781118142882"},{"key":"10.3233\/JIFS-224472_ref40","unstructured":"Dehak N. , Kenny P. , Dehak R. , Dumouchel P. , Ouellet P. IEEE transactions on audio, speech and language processing 1 Front-End Factor Analysis For Speaker Verification."},{"key":"10.3233\/JIFS-224472_ref41","doi-asserted-by":"publisher","first-page":"328","DOI":"10.1109\/29.21701","article-title":"Phoneme recognition using time-delay neural networks, Acoustics, Speech and Signal Processing","volume":"37","author":"Waibel","year":"1989","journal-title":"IEEE Transactions on"},{"key":"10.3233\/JIFS-224472_ref42","unstructured":"Krizhevsky A. , Sutskever I. , Hinton G.E. Image Net Classification with Deep Convolutional Neural Networks. [Online]. Available: http:\/\/code.google.com\/p\/cuda-convnet\/."},{"key":"10.3233\/JIFS-224472_ref43","unstructured":"Dubagunta S. Pavankumar , Kabil S.H. and Doss M.M., Improving children speech recognition through feature learning from raw speech signal. [Online]. Available: http:\/\/www.idiap.ch\/en\/people\/directory."},{"issue":"10","key":"10.3233\/JIFS-224472_ref44","doi-asserted-by":"publisher","first-page":"1533","DOI":"10.1109\/TASLP.2014.2339736","article-title":"neural networks for speech recognition","volume":"22","author":"Abdel-Hamid","year":"2014","journal-title":"IEEE Transactions on Audio, Speech and Language Processing"},{"key":"10.3233\/JIFS-224472_ref45","unstructured":"Goodfellow I. , Bengio Y. , Courville A. Deep Learning. MIT Press. 2016."},{"key":"10.3233\/JIFS-224472_ref46","unstructured":"Simonyan K. , Zisserman A. Very Deep Convolutional Networks forLarge-Scale Image Recognition, Sep. 2014, [Online]. Available: http:\/\/arxiv.org\/abs\/1409.1556"},{"issue":"2","key":"10.3233\/JIFS-224472_ref47","doi-asserted-by":"publisher","first-page":"291","DOI":"10.1109\/89.279278","article-title":"Maximum a posteriori estimation for multivariate Gaussian mixture observations of Markov chains","volume":"2","author":"Gauvain","year":"1994","journal-title":"IEEE Transactions on Speech and Audio Processing"},{"issue":"6","key":"10.3233\/JIFS-224472_ref48","doi-asserted-by":"publisher","first-page":"1129","DOI":"10.1109\/TASLP.2016.2544660","article-title":"Improving Short Utterance Speaker Recognition by Modeling Speech Unit Classes","volume":"24","author":"Li","year":"2016","journal-title":"IEEE\/ACMTransactions on Audio, Speech, and Language Processing"},{"issue":"4","key":"10.3233\/JIFS-224472_ref49","doi-asserted-by":"crossref","first-page":"1015","DOI":"10.1109\/TASL.2010.2076389","article-title":"Automatic prediction of children\u2019s reading ability for high-level literacy assessment","volume":"19","author":"Black","year":"2011","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"}],"container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"original-title":[],"link":[{"URL":"https:\/\/content.iospress.com\/download?id=10.3233\/JIFS-224472","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:45:29Z","timestamp":1777455929000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/full\/10.3233\/JIFS-224472"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,1]]},"references-count":49,"journal-issue":{"issue":"6"},"URL":"https:\/\/doi.org\/10.3233\/jifs-224472","relation":{},"ISSN":["1064-1246","1875-8967"],"issn-type":[{"value":"1064-1246","type":"print"},{"value":"1875-8967","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,1]]}}}