{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T11:01:54Z","timestamp":1761562914167,"version":"3.37.3"},"reference-count":39,"publisher":"Wiley","license":[{"start":{"date-parts":[[2022,9,16]],"date-time":"2022-09-16T00:00:00Z","timestamp":1663286400000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100015767","name":"Universitas Syiah Kuala","doi-asserted-by":"publisher","award":["268\/UN11\/SPK\/PNBP\/2020"],"award-info":[{"award-number":["268\/UN11\/SPK\/PNBP\/2020"]}],"id":[{"id":"10.13039\/501100015767","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Applied Computational Intelligence and Soft Computing"],"published-print":{"date-parts":[[2022,9,16]]},"abstract":"<jats:p>Currently, speech recognition datasets are increasingly available freely in various languages. However, speech recognition datasets in the Indonesian language are still challenging to obtain. Consequently, research focusing on speech recognition is challenging to carry out. This research creates Indonesian speech recognition datasets from YouTube channels with subtitles by validating all utterances of downloaded audio to improve the data quality. The quality of the dataset was evaluated using a deep neural network. The time delay neural network (TDNN) was used to build the acoustic model by applying the alignment data from the Gaussian mixture model-hidden Markov model (GMM-HMM). Data augmentation was used to increase the number of validated datasets and enhance the performance of the acoustic model. The results show that the acoustic model built using the validated datasets is better than the unvalidated datasets for all types of lexicons. Utilizing the four lexicon types and increasing the data through augmentation to train the acoustic models can lower the word error rate percentage in the GMM-HMM, TDNN factorization (TDNNF), and CNN-TDNNF-augmented models to 40.85%, 24.96%, and 19.03%, respectively.<\/jats:p>","DOI":"10.1155\/2022\/3227828","type":"journal-article","created":{"date-parts":[[2022,9,16]],"date-time":"2022-09-16T22:05:33Z","timestamp":1663365933000},"page":"1-16","source":"Crossref","is-referenced-by-count":1,"title":["Acoustic Model with Multiple Lexicon Types for Indonesian Speech Recognition"],"prefix":"10.1155","volume":"2022","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3859-6706","authenticated-orcid":true,"given":"Taufik Fuadi","family":"Abidin","sequence":"first","affiliation":[{"name":"Department of Informatics, Universitas Syiah Kuala, Banda Aceh, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0132-5961","authenticated-orcid":true,"given":"Alim","family":"Misbullah","sequence":"additional","affiliation":[{"name":"Department of Informatics, Universitas Syiah Kuala, Banda Aceh, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0370-3620","authenticated-orcid":true,"given":"Ridha","family":"Ferdhiana","sequence":"additional","affiliation":[{"name":"Department of Statistics, Universitas Syiah Kuala, Banda Aceh, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8806-3476","authenticated-orcid":true,"given":"Laina","family":"Farsiah","sequence":"additional","affiliation":[{"name":"Department of Informatics, Universitas Syiah Kuala, Banda Aceh, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8472-1061","authenticated-orcid":true,"given":"Muammar Zikri","family":"Aksana","sequence":"additional","affiliation":[{"name":"Department of Informatics, Universitas Syiah Kuala, Banda Aceh, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5449-6828","authenticated-orcid":true,"given":"Hammam","family":"Riza","sequence":"additional","affiliation":[{"name":"National Research and Innovation Agency, Jakarta, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","reference":[{"key":"1","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2004.828699"},{"key":"2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2003.813274"},{"key":"3","article-title":"Text indexing","volume-title":"Text Mining. Studies in Big Data","author":"T. Jo","year":"2019"},{"author":"K. A. Lee","key":"4","article-title":"Joint application of speech and speaker recognition for automation and security in smart home"},{"key":"5","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2012.2205597"},{"article-title":"Fast and accurate recurrent neural network acoustic models for speech recognition","year":"2015","author":"H. Sak","key":"6"},{"key":"7","doi-asserted-by":"crossref","DOI":"10.1109\/ICASSP.2015.7179001","article-title":"Scaling recurrent neural network language models","author":"W. Williams","year":"2015"},{"key":"8","doi-asserted-by":"crossref","first-page":"167","DOI":"10.1016\/j.procs.2016.04.045","article-title":"Towards robust Indonesian speech recognition with spontaneous-speech adapted acoustic models","volume":"81","author":"D. Hoesen","year":"2016","journal-title":"Procedia Computer Science"},{"key":"9","doi-asserted-by":"publisher","DOI":"10.1109\/ICOIACT.2018.8350748"},{"article-title":"Feature extraction analysis on Indonesian speech recognition system","author":"U. N. Wisesty","key":"10","doi-asserted-by":"crossref","DOI":"10.1109\/ICoICT.2015.7231396"},{"key":"11","doi-asserted-by":"publisher","DOI":"10.1109\/ICSDA.2017.8384448"},{"author":"S. Sakti","key":"12","article-title":"Development of indonesian large vocabulary continuous speech recognition system within a-star project"},{"article-title":"JTubeSpeech: corpus of Japanese speech collected from YouTube for speech recognition and speaker verification","year":"2021","author":"S. Takamichi","key":"13"},{"key":"14","doi-asserted-by":"publisher","DOI":"10.1016\/j.apacoust.2019.107175"},{"key":"15","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2988365"},{"author":"B. Oyo","key":"16","article-title":"A preliminary speech learning tool for improvement of African English accents"},{"article-title":"Urdu speech recognition system for district names of Pakistan: development, challenges and solutions","author":"M. Qasim","key":"17","doi-asserted-by":"crossref","DOI":"10.1109\/ICSDA.2016.7918979"},{"key":"18","doi-asserted-by":"publisher","DOI":"10.1121\/1.5147823"},{"key":"19","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053808"},{"key":"20","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2008.4517531"},{"key":"21","first-page":"270","article-title":"Automatic detection of syllable repetition in read speech for objective assessment of stuttered disfluencies","volume":"22","author":"K. M. Ravikumar","year":"2008","journal-title":"Proceedings of World Academy of Science, Engineering and Technology"},{"key":"22","doi-asserted-by":"crossref","DOI":"10.21437\/Interspeech.2016-1495","article-title":"Acoustic modelling from the signal domain using CNNs","author":"P. Ghahremani","year":"2016"},{"key":"23","doi-asserted-by":"publisher","DOI":"10.1017\/S135132491600005X"},{"key":"24","doi-asserted-by":"publisher","DOI":"10.21437\/interspeech.2014-80"},{"key":"25","doi-asserted-by":"publisher","DOI":"10.1007\/s10772-018-09573-7"},{"key":"26","doi-asserted-by":"publisher","DOI":"10.21533\/pen.v9i4.2450"},{"key":"27","doi-asserted-by":"publisher","DOI":"10.18280\/TS.380212"},{"key":"28","doi-asserted-by":"publisher","DOI":"10.1016\/J.APACOUST.2021.107918"},{"key":"29","doi-asserted-by":"publisher","DOI":"10.3390\/S21217025"},{"key":"30","doi-asserted-by":"crossref","DOI":"10.1109\/SLT48900.2021.9383624","article-title":"Dual application of speech enhancement for automatic speech recognition","author":"A. Pandey","year":"2021"},{"key":"31","doi-asserted-by":"crossref","DOI":"10.21437\/Interspeech.2015-647","article-title":"A time delay neural network architecture for efficient modeling of long temporal contexts","author":"V. Peddinti","year":"2015"},{"key":"32","doi-asserted-by":"crossref","DOI":"10.21437\/Interspeech.2015-527","article-title":"Reverberation robust acoustic modeling using i-vectors with time delay neural networks","author":"V. Peddinti","year":"2015"},{"article-title":"Multistream CNN for Robust Acoustic Modeling","year":"2021","author":"K. J. Han","key":"33"},{"key":"34","doi-asserted-by":"publisher","DOI":"10.3390\/sym11091185"},{"key":"35","doi-asserted-by":"publisher","DOI":"10.3966\/199115992017102805001"},{"key":"36","doi-asserted-by":"publisher","DOI":"10.1016\/j.jksuci.2021.04.002"},{"article-title":"The kaldi speech recognition toolkit","year":"2011","author":"D. Povey","key":"37"},{"author":"T. F. Abidin","key":"38","article-title":"Deep neural network for automatic speech recognition from Indonesian audio using several lexicon types"},{"article-title":"Semi-orthogonal low-rank matrix factorization for deep neural networks","year":"2018","author":"D. Povey","key":"39"}],"container-title":["Applied Computational Intelligence and Soft Computing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/downloads.hindawi.com\/journals\/acisc\/2022\/3227828.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/acisc\/2022\/3227828.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/acisc\/2022\/3227828.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,9,16]],"date-time":"2022-09-16T22:05:47Z","timestamp":1663365947000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.hindawi.com\/journals\/acisc\/2022\/3227828\/"}},"subtitle":[],"editor":[{"given":"Bhargav","family":"Appasani","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2022,9,16]]},"references-count":39,"alternative-id":["3227828","3227828"],"URL":"https:\/\/doi.org\/10.1155\/2022\/3227828","relation":{},"ISSN":["1687-9732","1687-9724"],"issn-type":[{"type":"electronic","value":"1687-9732"},{"type":"print","value":"1687-9724"}],"subject":[],"published":{"date-parts":[[2022,9,16]]}}}