{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,5]],"date-time":"2025-12-05T23:55:25Z","timestamp":1764978925093,"version":"3.46.0"},"reference-count":54,"publisher":"Walter de Gruyter GmbH","issue":"1","license":[{"start":{"date-parts":[[2020,7,28]],"date-time":"2020-07-28T00:00:00Z","timestamp":1595894400000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,7,28]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>This paper implements the continuous Hindi Automatic Speech Recognition (ASR) system using the proposed integrated features vector with Recurrent Neural Network (RNN) based Language Modeling (LM). The proposed system also implements the speaker adaptation using Maximum-Likelihood Linear Regression (MLLR) and Constrained Maximum likelihood Linear Regression (C-MLLR). This system is discriminatively trained by Maximum Mutual Information (MMI) and Minimum Phone Error (MPE) techniques with 256 Gaussian mixture per Hidden Markov Model(HMM) state. The training of the baseline system has been done using a phonetically rich Hindi dataset. The results show that discriminative training enhances the baseline system performance by up to 3%. Further improvement of ~7% has been recorded by applying RNN LM. The proposed Hindi ASR system shows significant performance improvement over other current state-of-the-art techniques.<\/jats:p>","DOI":"10.1515\/jisys-2018-0417","type":"journal-article","created":{"date-parts":[[2020,7,28]],"date-time":"2020-07-28T06:17:49Z","timestamp":1595917069000},"page":"165-179","source":"Crossref","is-referenced-by-count":24,"title":["Discriminatively trained continuous Hindi speech recognition using integrated acoustic features and recurrent neural network language modeling"],"prefix":"10.1515","volume":"30","author":[{"given":"A.","family":"Kumar","sequence":"first","affiliation":[{"name":"Computer Engineering Department, National Institute of Technology , Kurukshetra , Haryana , India"},{"name":"Computer Science and Engineering Department, Galgotias University , Greater Noida , Uttar Pradesh , India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"R.K.","family":"Aggarwal","sequence":"additional","affiliation":[{"name":"Computer Engineering Department, National Institute of Technology , Kurukshetra , Haryana , India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"374","published-online":{"date-parts":[[2020,7,28]]},"reference":[{"doi-asserted-by":"crossref","unstructured":"R. K. Aggarwal and M. Dave, Performance evaluation of sequentially combined heterogeneous feature streams for Hindi speech recognition system, Telecommun. Syst. 52 (2013), 1-10.","key":"2025120523494761638_j_jisys-2018-0417_ref_001","DOI":"10.1007\/s11235-011-9623-0"},{"doi-asserted-by":"crossref","unstructured":"M. A. Anusuya and S. K. Katti, Front end analysis of speech recognition: a review, International Journal of Speech Technology 14.2 (2011): 99-145.","key":"2025120523494761638_j_jisys-2018-0417_ref_002","DOI":"10.1007\/s10772-010-9088-7"},{"doi-asserted-by":"crossref","unstructured":"A. Biswas et al., Feature extraction technique using ERB like wavelet sub-band periodic and aperiodic decomposition for TIMIT phoneme recognition, International Journal of Speech Technology 17.4 (2014): 389-399.","key":"2025120523494761638_j_jisys-2018-0417_ref_003","DOI":"10.1007\/s10772-014-9236-6"},{"doi-asserted-by":"crossref","unstructured":"A. Biswas et al., Hindi phoneme classification using Wiener filtered wavelet packet decomposed periodic and aperiodic acoustic feature, Computers & Electrical Engineering 42 (2015): 12-22.","key":"2025120523494761638_j_jisys-2018-0417_ref_004","DOI":"10.1016\/j.compeleceng.2014.12.017"},{"unstructured":"W. Burgos, Gammatone and MFCC Features in Speaker Recognition Dissertation, 2014.","key":"2025120523494761638_j_jisys-2018-0417_ref_005"},{"doi-asserted-by":"crossref","unstructured":"X. Chen et al., Improving the training and evaluation efficiency of recurrent neural network language models, in: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) IEEE, 2015.","key":"2025120523494761638_j_jisys-2018-0417_ref_006","DOI":"10.1109\/ICASSP.2015.7179003"},{"doi-asserted-by":"crossref","unstructured":"X. Chen et al.,Eflcient training and evaluation of recurrent neural network language models for automatic speech recognition, IEEE\/ACM Transactions on Audio, Speech, and Language Processing 24.11 (2016): 2146-2157.","key":"2025120523494761638_j_jisys-2018-0417_ref_007","DOI":"10.1109\/TASLP.2016.2598304"},{"doi-asserted-by":"crossref","unstructured":"X. Chen et al., CUED\u2013RNNLM\u2013An open-source toolkit for eflcient training and evaluation of recurrent neural network language models, in:2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) IEEE, 2016","key":"2025120523494761638_j_jisys-2018-0417_ref_008","DOI":"10.1109\/ICASSP.2016.7472829"},{"doi-asserted-by":"crossref","unstructured":"A. D. Cheveign\u00e9 et al., Concurrent vowel identification. II. Effects of phase, harmonicity, and task, The Journal of the Acoustical Society of America 101.5 (1997): 2848-2856.","key":"2025120523494761638_j_jisys-2018-0417_ref_009","DOI":"10.1121\/1.419476"},{"unstructured":"H. P. Combrinck and E. C. Botha, On the Mel-scaled cepstrum, in:Department of Electrical and Electronic Engineering, University of Pretoria Pretoria, South Africa, 1996.","key":"2025120523494761638_j_jisys-2018-0417_ref_010"},{"doi-asserted-by":"crossref","unstructured":"A. Currey et al., Dynamic adjustment of language models for automatic speech recognition using word similarity, in: Spoken Language Technology Workshop (SLT) IEEE, 2016.","key":"2025120523494761638_j_jisys-2018-0417_ref_011","DOI":"10.1109\/SLT.2016.7846299"},{"doi-asserted-by":"crossref","unstructured":"S. B. Davis, and P. Mermelstein, Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences, Readings in speech recognition (1990): 65-74.","key":"2025120523494761638_j_jisys-2018-0417_ref_012","DOI":"10.1016\/B978-0-08-051584-7.50010-3"},{"doi-asserted-by":"crossref","unstructured":"L. Deng et al., Distributed speech processing in MiPad\u015b multimodal user interface, IEEE Transactions on Speech and Audio Processing 10.8 (2002): 605-619.","key":"2025120523494761638_j_jisys-2018-0417_ref_013","DOI":"10.1109\/TSA.2002.804538"},{"doi-asserted-by":"crossref","unstructured":"M. Dua, R. K. Aggarwal, and M. Biswas, Discriminatively trained continuous Hindi speech recognition system using interpolated recurrent neural network language modeling, In:Neural Computing and Applications (2018): 1-9.","key":"2025120523494761638_j_jisys-2018-0417_ref_014","DOI":"10.1007\/s00521-018-3499-9"},{"doi-asserted-by":"crossref","unstructured":"M. Dua, R. K. Aggarwal, and M. Biswas, Discriminative training using noise robust integrated features and refined HMM modeling, Journal of Intelligent Systems (2018).","key":"2025120523494761638_j_jisys-2018-0417_ref_015","DOI":"10.1515\/jisys-2017-0618"},{"doi-asserted-by":"crossref","unstructured":"O. Farooq O, S. Datta, M.C. Shrotriya, Wavelet sub-band based temporal features for robust Hindi phoneme recognition, International Journal of Wavelets, Multiresolution and Information Processing 8.6 (2010):847-59.","key":"2025120523494761638_j_jisys-2018-0417_ref_016","DOI":"10.1142\/S0219691310003845"},{"doi-asserted-by":"crossref","unstructured":"M. Ferras et al., Comparison of speaker adaptation methods as feature extraction for SVM-based speaker recognition, IEEE Transactions on Audio, Speech, and Language Processing 18.6 (2010): 1366-1378.","key":"2025120523494761638_j_jisys-2018-0417_ref_017","DOI":"10.1109\/TASL.2009.2034187"},{"doi-asserted-by":"crossref","unstructured":"D. Gillick, S. Wegmann, and L. Gillick, Discriminative training for speech recognition is compensating for statistical dependence in the HMM framework, in:2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) IEEE, 2012.","key":"2025120523494761638_j_jisys-2018-0417_ref_018","DOI":"10.1109\/ICASSP.2012.6288979"},{"doi-asserted-by":"crossref","unstructured":"G. Heigold, N. Hermann, S. Ralph, and W. Simon, Discriminative training for automatic speech recognition: Modeling, criteria, optimization, implementation, and performance,In: IEEE Signal Processing Magazine 29.6 (2012): 58-69.","key":"2025120523494761638_j_jisys-2018-0417_ref_019","DOI":"10.1109\/MSP.2012.2197232"},{"doi-asserted-by":"crossref","unstructured":"H. Hermansky, Perceptual linear predictive (PLP) analysis of speech, the Journal of the Acoustical Society of America 87.4 (1990): 1738-1752.","key":"2025120523494761638_j_jisys-2018-0417_ref_020","DOI":"10.1121\/1.399423"},{"doi-asserted-by":"crossref","unstructured":"K. Ishizuka, and T. Nakatani, A feature extraction method using subband based periodicity and aperiodicity decomposition with noise robust frontend processing for automatic speech recognition, Speech communication48.11 (2006): 1447-1457.","key":"2025120523494761638_j_jisys-2018-0417_ref_021","DOI":"10.1016\/j.specom.2006.06.008"},{"unstructured":"R. Jozefowicz et al., Exploring the limits of language modelingpreprint(2016), http:\/\/arxiv.org\/abs\/1602.02410","key":"2025120523494761638_j_jisys-2018-0417_ref_022"},{"doi-asserted-by":"crossref","unstructured":"V. Kadyan, A. Mantri and R. K. Aggarwal, A heterogeneous speech feature vectors generation approach with hybrid hmm classifiers, Int. J. Speech Technology 20.4 (2017): 761-769.","key":"2025120523494761638_j_jisys-2018-0417_ref_023","DOI":"10.1007\/s10772-017-9446-9"},{"unstructured":"S. Kapadia, Discriminative training of hidden Markov models Doctoral dissertation, University of Cambridge, 1998.","key":"2025120523494761638_j_jisys-2018-0417_ref_024"},{"unstructured":"J. Koehler, N. Morgan, H. Hermansky, H. G. Hirsch and G. Tong, Integrating RASTA-PLP into Speech Recognition, in: 1994 IEEE International Conference on Acoustics, Speech, and Signal Processing1 Adelaide, SA, Australia, 1994.","key":"2025120523494761638_j_jisys-2018-0417_ref_025"},{"doi-asserted-by":"crossref","unstructured":"R. Kumar, A. Kumar, and R. K. Pandey, Beta wavelet based ECG signal compression using lossless encoding with modified thresholding, Computers & Electrical Engineering 39.1 (2013): 130-140.","key":"2025120523494761638_j_jisys-2018-0417_ref_026","DOI":"10.1016\/j.compeleceng.2012.04.008"},{"doi-asserted-by":"crossref","unstructured":"N. Kumar and A. G. Andreou, Heteroscedastic discriminant analysis and reduced rank HMMs for improved speech recognition, Speech Commun 26.4 (1998), 283-297.","key":"2025120523494761638_j_jisys-2018-0417_ref_027","DOI":"10.1016\/S0167-6393(98)00061-2"},{"unstructured":"A. G. Kunkle, Sequence scoring experiments using the TIMIT corpus and the HTK recognition framework Dissertation, Florida Institute of Technology, Florida, USA, 2010.","key":"2025120523494761638_j_jisys-2018-0417_ref_028"},{"doi-asserted-by":"crossref","unstructured":"J. Li et al., An overview of noise-robust automatic speech recognition, IEEE\/ACM Transactions on Audio, Speech, and Language Processing 22.4 (2014): 745-777.","key":"2025120523494761638_j_jisys-2018-0417_ref_029","DOI":"10.1109\/TASLP.2014.2304637"},{"doi-asserted-by":"crossref","unstructured":"K. Li et al., Recurrent neural network language model adaptation for conversational speech recognition, In: INTERSPEECH Hyderabad (2018): 1-5.","key":"2025120523494761638_j_jisys-2018-0417_ref_030","DOI":"10.21437\/Interspeech.2018-1413"},{"doi-asserted-by":"crossref","unstructured":"T. Mikolov et al., Recurrent neural network based language model, In:Eleventh annual conference of the international speech communication association(2010).","key":"2025120523494761638_j_jisys-2018-0417_ref_031","DOI":"10.21437\/Interspeech.2010-343"},{"doi-asserted-by":"crossref","unstructured":"T. Mikolov et al., Extensions of recurrent neural network language model, In:2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)IEEE, (2011):5528-5531.","key":"2025120523494761638_j_jisys-2018-0417_ref_032","DOI":"10.1109\/ICASSP.2011.5947611"},{"doi-asserted-by":"crossref","unstructured":"T. Mikolov et al., Context dependent recurrent neural network language model, In:2012 IEEE Spoken Language Technology Workshop (SLT)IEEE, (2012):234-239.","key":"2025120523494761638_j_jisys-2018-0417_ref_033","DOI":"10.1109\/SLT.2012.6424228"},{"doi-asserted-by":"crossref","unstructured":"A. Mohan et al., Acoustic modelling for speech recognition in Indian languages in an agricultural commodities task domain, Speech Communication 56 (2014): 167-180.","key":"2025120523494761638_j_jisys-2018-0417_ref_034","DOI":"10.1016\/j.specom.2013.07.005"},{"unstructured":"J. M. Naik, L. P. Netsch, and G. R. Doddington, Speaker verification over long distance telephone lines, in:International Conference on Acoustics, Speech, and Signal Processing ICASSP-89 IEEE, 1989.","key":"2025120523494761638_j_jisys-2018-0417_ref_035"},{"unstructured":"D. Povey, Discriminative training for large vocabulary speech recognition PhD Diss. University of Cambridge, 2005.","key":"2025120523494761638_j_jisys-2018-0417_ref_036"},{"doi-asserted-by":"crossref","unstructured":"D. Povey et al.,Purely Sequence-Trained Neural Networks for ASR Based on Lattice-Free MMI, In: Interspeech (2016): 2751-2755.","key":"2025120523494761638_j_jisys-2018-0417_ref_037","DOI":"10.21437\/Interspeech.2016-595"},{"doi-asserted-by":"crossref","unstructured":"S. Ranjan. A discrete wavelet transform based approach to Hindi speech recognition, In:2010 International Conference on Signal Acquisition and Processing Bangalore(2010):345-348.","key":"2025120523494761638_j_jisys-2018-0417_ref_038","DOI":"10.1109\/ICSAP.2010.21"},{"doi-asserted-by":"crossref","unstructured":"S. Ranjan, Exploring the discrete wavelet transform as a tool for Hindi speech recognition, International Journal of Computer Theory and Engineering 2.4 (2010): 642.","key":"2025120523494761638_j_jisys-2018-0417_ref_039","DOI":"10.7763\/IJCTE.2010.V2.216"},{"doi-asserted-by":"crossref","unstructured":"D. E. Rumelhart, G. E. Hinton, and R. J. Williams, Learning representations by back-propagating errors, nature 323.6088 (1986): 533.","key":"2025120523494761638_j_jisys-2018-0417_ref_040","DOI":"10.1038\/323533a0"},{"unstructured":"K. Samudravijaya, P. V. S. Rao and S. S. Agrawal, Hindi speech database, in: International Conference on spoken Language Processing Beijing, China, 2002, pp. 456\u2013464.","key":"2025120523494761638_j_jisys-2018-0417_ref_041"},{"doi-asserted-by":"crossref","unstructured":"R. Schluter et al., Gammatone features and feature combination for large vocabulary speech recognition, in:IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2007) 4 IEEE, 2007.","key":"2025120523494761638_j_jisys-2018-0417_ref_042","DOI":"10.1109\/ICASSP.2007.366996"},{"doi-asserted-by":"crossref","unstructured":"Y. Shao, and W. DeLiang, Robust speaker identification using auditory features and computational auditory scene analysis, In: 2008 IEEE International Conference on Acoustics, Speech and Signal ProcessingIEEE, (2008):1589-1592.","key":"2025120523494761638_j_jisys-2018-0417_ref_043","DOI":"10.1109\/ICASSP.2008.4517928"},{"doi-asserted-by":"crossref","unstructured":"Y. Shao et al., A computational auditory scene analysis system for speech segregation and robust speech recognition, Computer Speech & Language 24.1 (2010): 77-93.","key":"2025120523494761638_j_jisys-2018-0417_ref_044","DOI":"10.1016\/j.csl.2008.03.004"},{"doi-asserted-by":"crossref","unstructured":"A. Sharma et al., Hybrid wavelet based LPC features for Hindi speech recognition, International Journal of Information and Communication Technology 1.3-4 (2008): 373-381.","key":"2025120523494761638_j_jisys-2018-0417_ref_045","DOI":"10.1504\/IJICT.2008.024008"},{"doi-asserted-by":"crossref","unstructured":"N. Singh-Miller, Natasha, M. Collins, and T. J. Hazen, Dimensionality reduction for speech recognition using neighborhood components analysis, in: Eighth Annual Conference of the International Speech Communication Association 2007.","key":"2025120523494761638_j_jisys-2018-0417_ref_046","DOI":"10.21437\/Interspeech.2007-376"},{"doi-asserted-by":"crossref","unstructured":"A. Stolcke et al., MLLR transforms as features in speaker recognition, in:Ninth European Conference on Speech Communication and Technology 2005.","key":"2025120523494761638_j_jisys-2018-0417_ref_047","DOI":"10.21437\/Interspeech.2005-647"},{"unstructured":"K. Vertanen, An overview of discriminative training for speech recognition, University of Cambridge Cambridge, UK (2004).","key":"2025120523494761638_j_jisys-2018-0417_ref_048"},{"unstructured":"E. Wong, and S. Sridharan, Comparison of linear prediction cepstrum coefficients and mel-frequency cepstrum coefficients for language identification, in: Proceedings of 2001 International Symposium on Intelligent Multimedia, Video and Speech Processing IEEE, 2001.","key":"2025120523494761638_j_jisys-2018-0417_ref_049"},{"doi-asserted-by":"crossref","unstructured":"Z. Wu et al., A study of speaker adaptation for DNN-based speech synthesis, in:Sixteenth Annual Conference of the International Speech Communication Association 2015.","key":"2025120523494761638_j_jisys-2018-0417_ref_050","DOI":"10.21437\/Interspeech.2015-270"},{"unstructured":"D. Yogatama et al., Memory architectures in recurrent neural network language models, in: Seventh International Conference on Learning Representations (2018).","key":"2025120523494761638_j_jisys-2018-0417_ref_051"},{"unstructured":"S. Young, G. Evermann, M. Gales,T. Hain, D. Kershaw,X. Liu,V. Valtchev, The HTK book Cambridge University Engineering Department, vol 3, pp 1\u2013285,2002.","key":"2025120523494761638_j_jisys-2018-0417_ref_052"},{"doi-asserted-by":"crossref","unstructured":"X. Zhao and D. L. Wang, Analyzing noise robustness of MFCC and GFCC features in speaker identification, in: 2013 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) IEEE, 2013","key":"2025120523494761638_j_jisys-2018-0417_ref_053","DOI":"10.1109\/ICASSP.2013.6639061"},{"doi-asserted-by":"crossref","unstructured":"A. Zolnay, R. Schluter, and H. Ney., Acoustic feature combination for robust speech recognition, in: IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201905)1 IEEE, 2005.","key":"2025120523494761638_j_jisys-2018-0417_ref_054","DOI":"10.1109\/ICASSP.2005.1415149"}],"container-title":["Journal of Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.degruyter.com\/view\/journals\/jisys\/30\/1\/article-p165.xml","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/jisys-2018-0417\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/jisys-2018-0417\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,5]],"date-time":"2025-12-05T23:50:35Z","timestamp":1764978635000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/jisys-2018-0417\/html"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,7,28]]},"references-count":54,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2020,8,15]]},"published-print":{"date-parts":[[2020,8,15]]}},"alternative-id":["10.1515\/jisys-2018-0417"],"URL":"https:\/\/doi.org\/10.1515\/jisys-2018-0417","relation":{},"ISSN":["2191-026X"],"issn-type":[{"type":"electronic","value":"2191-026X"}],"subject":[],"published":{"date-parts":[[2020,7,28]]}}}