{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,16]],"date-time":"2025-11-16T15:49:57Z","timestamp":1763308197490,"version":"3.41.0"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"10","license":[{"start":{"date-parts":[[2023,10,13]],"date-time":"2023-10-13T00:00:00Z","timestamp":1697155200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"China National Natural Science Foundation","doi-asserted-by":"crossref","award":["62366037, 62066033"],"award-info":[{"award-number":["62366037, 62066033"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Research Foundation for Young Scholars of Inner Mongolia University","award":["10000-23112101\/052"],"award-info":[{"award-number":["10000-23112101\/052"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,10,31]]},"abstract":"<jats:p>Traditional weighted finite-state transducer\u2013 (WFST) based Mongolian automatic speech recognition (ASR) systems use phonemes as pronunciation lexicon modeling units. However, Mongolian is an agglutinative, low-resource language, and building an ASR system based on the phoneme pronunciation lexicon remains a challenge for various reasons. First, the phoneme pronunciation lexicon manually constructed by Mongolian linguists is finite, which is usually used to build a grapheme-to-phoneme conversion (G2P) model to frequently expand new words. However, the data sparsity decreases the robustness of the G2P model and affects the performance of the final ASR system. Second, homophones and polysyllabic words are common in Mongolian, which has a certain impact on the construction of the Mongolian acoustic model. To address these problems, in this work, we first propose a grapheme-to-phoneme alignment model to obtain the mapping relationship between phonemes and subword units. Then, we construct an acoustic subword segmentation set to segment words directly instead of using the traditional G2P method to predict phoneme sequences to expand the pronunciation lexicon. Further, by analyzing the Mongolian encoding form, we also propose an acoustic subword modeling units construction method that removes control characters. Finally, we investigate various acoustic subword modeling units for pronunciation lexicon construction for the Mongolian ASR system. Experiments on a Mongolian dataset with 325 hours of training show that the pronunciation lexicon based on the acoustic subword modeling unit can effectively construct the WFST-based Mongolian ASR system. Further, removing the control characters when building the acoustic subword modeling unit can further improve the ASR system performance.<\/jats:p>","DOI":"10.1145\/3617830","type":"journal-article","created":{"date-parts":[[2023,8,29]],"date-time":"2023-08-29T11:22:44Z","timestamp":1693308164000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["A Comparative Study on Selecting Acoustic Modeling Units for WFST-based Mongolian Speech Recognition"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1647-1539","authenticated-orcid":false,"given":"Wang","family":"Yonghe","sequence":"first","affiliation":[{"name":"College of Computer Science, Inner Mongolia University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7312-1629","authenticated-orcid":false,"given":"Feilong","family":"Bao","sequence":"additional","affiliation":[{"name":"College of Computer Science, Inner Mongolia University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-5513-1192","authenticated-orcid":false,"given":"Gaunglai","family":"Gao","sequence":"additional","affiliation":[{"name":"College of Computer Science, Inner Mongolia University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,10,13]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"State Administration for Market Regulation.","author":"Administration China National Standards","year":"2010","unstructured":"China National Standards Administration. 2010. GB\/T26226-2010 information technology mongolian variation display character set and rules for the use of control characters. In State Administration for Market Regulation."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2013.6639250"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8462105"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1609.03193"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/1187415.1187418"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9746748"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9747842"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2017.2743344"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414082"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143891"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2012.2205597"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ITAIC54216.2022.9836706"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.4324\/9780203987919"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.21437\/Eurospeech.2003-785"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1007"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-1569"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1712.09444"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-99495-6_4"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414509"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9746652"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414928"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2015-647"},{"key":"e_1_3_2_24_2","volume-title":"Proceedings of the IEEE 2011 Workshop on Automatic Speech Recognition and Understanding","author":"Povey Daniel","year":"2011","unstructured":"Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et\u00a0al. 2011. The Kaldi speech recognition toolkit. In Proceedings of the IEEE 2011 Workshop on Automatic Speech Recognition and Understanding. IEEE Signal Processing Society."},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2016-595"},{"key":"e_1_3_2_26_2","first-page":"77","article-title":"Mongolian syntax","year":"1991","unstructured":"Qingge\u2019ertai. 1991. Mongolian syntax. Inner Mongolia People\u2019s Publishing House (1991), 77\u2013133.","journal-title":"Inner Mongolia People\u2019s Publishing House"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9413966"},{"issue":"3","key":"e_1_3_2_28_2","first-page":"1","article-title":"Improving deep learning based automatic speech recognition for Gujarati","volume":"21","author":"Raval Deepang","year":"2021","unstructured":"Deepang Raval, Vyom Pathak, Muktan Patel, and Brijesh Bhatt. 2021. Improving deep learning based automatic speech recognition for Gujarati. Trans. Asian Low-Resourc. Lang. Inf. Process. 21, 3 (2021), 1\u201318.","journal-title":"Trans. Asian Low-Resourc. Lang. Inf. Process."},{"key":"e_1_3_2_29_2","first-page":"157","article-title":"Long short-term memory recurrent neural network architectures for large scale acoustic modeling","author":"Sak Hasim","year":"2014","unstructured":"Hasim Sak, Andrew W. Senior, and Fran\u00e7oise Beaufays. 2014. Long short-term memory recurrent neural network architectures for large scale acoustic modeling. IEEE Trans. Neural Netw. (2014), 157\u2013166.","journal-title":"IEEE Trans. Neural Netw."},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9747713"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/89.985546"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2020.101158"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-103"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.21437\/ICSLP.2002-303"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/AICT52784.2021.9620466"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2020.101141"},{"key":"e_1_3_2_37_2","first-page":"383","volume-title":"Proceedings of the National Conference on Man-Machine Speech Communication","author":"Wang Yonghe","year":"2017","unstructured":"Yonghe Wang, Feilong Bao, Guanglai Gao, and Rui Liu. 2017. Research on Mongolian speech recognition based on TDNN-LSTM. In Proceedings of the National Conference on Man-Machine Speech Communication. 383\u2013391."},{"key":"e_1_3_2_38_2","first-page":"243","volume-title":"International Conference on Natural Language Processing and Chinese Computing","author":"Wang Yonghe","year":"2017","unstructured":"Yonghe Wang, Feilong Bao, Hongwei Zhang, and Guanglai Gao. 2017. Research on Mongolian speech recognition based on FSMN. In International Conference on Natural Language Processing and Chinese Computing. Springer, 243\u2013254."},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9413679"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8462353"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8682494"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-1995"},{"key":"e_1_3_2_43_2","doi-asserted-by":"crossref","first-page":"180","DOI":"10.1007\/978-3-319-25816-4_15","volume-title":"Chinese Computational Linguistics and Natural Language Processing Based on Naturally Annotated Big Data","author":"Zhang Hui","year":"2015","unstructured":"Hui Zhang, Feilong Bao, and Guanglai Gao. 2015. Mongolian speech recognition based on deep neural networks. In Chinese Computational Linguistics and Natural Language Processing Based on Naturally Annotated Big Data. Springer, 180\u2013188."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9413648"},{"key":"e_1_3_2_45_2","doi-asserted-by":"crossref","unstructured":"Wei Zhou Mohammad Zeineldeen Zuoyun Zheng Ralf Schl\u00fcter and Hermann Ney. 2021. Acoustic data-driven subword modeling for end-to-end speech recognition. In Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH\u201921) . 2886\u20132890.","DOI":"10.21437\/Interspeech.2021-1623"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3617830","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3617830","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:57Z","timestamp":1750178277000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3617830"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,13]]},"references-count":44,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2023,10,31]]}},"alternative-id":["10.1145\/3617830"],"URL":"https:\/\/doi.org\/10.1145\/3617830","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2023,10,13]]},"assertion":[{"value":"2022-07-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-08-24","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}