{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,18]],"date-time":"2026-03-18T04:57:11Z","timestamp":1773809831175,"version":"3.50.1"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2023,7,20]],"date-time":"2023-07-20T00:00:00Z","timestamp":1689811200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004543","name":"China Scholarship Council","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004543","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,7,31]]},"abstract":"<jats:p>Speech dataset is an essential component in building commercial speech applications. However, low-resource languages such as Swahili lack such a resource that is vital for spoken digit recognition. For languages where such resources exist, they are usually insufficient. Thus, pre-training methods have been used with external resources to improve continuous speech recognition. However, to the best of our knowledge, no study has investigated the effect of pre-training methods specifically for spoken digit recognition. This study aimed at addressing these problems. First, we developed a Swahili spoken digit dataset for Swahili spoken digit recognition. Then, we investigated the effect of cross-lingual and multi-lingual pre-training methods on spoken digit recognition. Finally, we proposed an effective language-independent pre-training method for spoken digit recognition. The proposed method has the advantage of incorporating target language data during the pre-training stage that leads to an optimal solution when using less training data. Experiments on Swahili (being developed), English, and Gujarati datasets show that our method achieves better performance compared with all the baselines listed in this study.<\/jats:p>","DOI":"10.1145\/3597494","type":"journal-article","created":{"date-parts":[[2023,5,20]],"date-time":"2023-05-20T09:46:36Z","timestamp":1684575996000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Swahili Speech Dataset Development and Improved Pre-training Method for Spoken Digit Recognition"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9184-3078","authenticated-orcid":false,"given":"Alexander R.","family":"Kivaisi","sequence":"first","affiliation":[{"name":"Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science and Technology, Beijing Institute of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6955-4170","authenticated-orcid":false,"given":"Qingjie","family":"Zhao","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science and Technology, Beijing Institute of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2263-3218","authenticated-orcid":false,"given":"Jimmy T.","family":"Mbelwa","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, University of Dar es salaam"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,7,20]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dan Mane Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Viegas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2016. TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. arXiv preprint arXiv:1603.04467."},{"key":"e_1_3_2_3_2","first-page":"96","article-title":"Massively multilingual adversarial speech recognition","volume":"1","author":"Adams Oliver","year":"2019","unstructured":"Oliver Adams, Matthew Wiesner, Shinji Watanabe, and David Yarowsky. 2019. Massively multilingual adversarial speech recognition. NAACL HLT 2019 - 2019 Conf. North Am. Chapter Assoc. Comput. Linguist. Hum. Lang. Technol. - Proc. Conf. 1, 96\u2013108.","journal-title":"NAACL HLT 2019 - 2019 Conf. North Am. Chapter Assoc. Comput. Linguist. Hum. Lang. Technol. - Proc. Conf"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10772-014-9267-z"},{"key":"e_1_3_2_5_2","first-page":"1545","volume-title":"Proceedings of the 2008 IEEE International Conference on Acoustics, Speech and Signal Processing.","author":"de Andrade Bresolin A.","year":"2008","unstructured":"A. de Andrade Bresolin, A. D. D. Neto, and P. J. Alsina. 2008. Digit recognition using wavelet and SVM in Brazilian Portuguese. In Proceedings of the 2008 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 1545\u20131548."},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2005.851998"},{"key":"e_1_3_2_7_2","article-title":"Cross-lingual language model pretraining","volume":"32","author":"Conneau Alexis","year":"2019","unstructured":"Alexis Conneau and Guillaume Lample. 2019. Cross-lingual language model pretraining. Advances in Neural Information Processing Systems 32, (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","first-page":"208","DOI":"10.1007\/978-981-15-4828-4_18","article-title":"Development of a novel database in gujarati language for spoken digits classification","volume":"5","author":"Dalsaniya Nikunj","year":"2020","unstructured":"Nikunj Dalsaniya, Sapan H. Mankad, Sanjay Garg, and Dhuri Shrivastava. 2020. Development of a novel database in gujarati language for spoken digits classification. Advances in Signal Processing and Intelligent Recognition Systems: 5th International Symposium, SIRS 2019, Trivandrum, India, December 18\u201321, 2019, Revised Selected Papers 5, 1209 (2020), 208\u2013219.","journal-title":"Advances in Signal Processing and Intelligent Recognition Systems: 5th International Symposium, SIRS 2019, Trivandrum, India, December 18\u201321, 2019, Revised Selected Papers"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383772"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1980.1163420"},{"key":"e_1_3_2_11_2","first-page":"4818","volume-title":"Proceedings of the 2014 IEEE International Conference on Acoustics, Speech and Signal Processing.","author":"Deng Jun","year":"2014","unstructured":"Jun Deng, Rui Xia, Zixing Zhang, Yang Liu, and Bjorn Schuller. 2014. Introducing shared-hidden-layer autoencoders for transfer learning and their application in acoustic emotion recognition. In Proceedings of the 2014 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 4818\u20134822."},{"issue":"5","key":"e_1_3_2_12_2","first-page":"3406","article-title":"Voice and speech recognition for tamil words and numerals","volume":"2","author":"Dharun V. S.","year":"2012","unstructured":"V. S. Dharun and M. Karnan. 2012. Voice and speech recognition for tamil words and numerals. International Journal of Modern Engineering Research 2, 5 (2012), 3406\u20133414.","journal-title":"International Journal of Modern Engineering Research"},{"key":"e_1_3_2_13_2","first-page":"3755","volume-title":"Proceedings of the 2016 International Conference on Electrical, Electronics, and Optimization Techniques.","author":"Dixit Abhishek","year":"2016","unstructured":"Abhishek Dixit, Abhinav Vidwans, and Pankaj Sharma. 2016. Improved MFCC and LPC algorithm for bundelkhandi isolated digit speech recognition. In Proceedings of the 2016 International Conference on Electrical, Electronics, and Optimization Techniques. IEEE, 3755\u20133759."},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","first-page":"361","DOI":"10.1109\/SLT.2018.8639541","volume-title":"Proceedings of the 2018 IEEE Spoken Language Technology Workshop","author":"Drexler Jennifer","year":"2018","unstructured":"Jennifer Drexler and James Glass. 2018. Combining End-to-End and adversarial training for low-resource speech recognition. In Proceedings of the 2018 IEEE Spoken Language Technology Workshop. IEEE, 361\u2013368."},{"key":"e_1_3_2_15_2","first-page":"5099","volume-title":"Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing.","author":"Emre Eskimez Sefik","year":"2018","unstructured":"Sefik Emre Eskimez, Zhiyao Duan, and Wendi Heinzelman. 2018. Unsupervised learning approach to feature analysis for automatic speech emotion recognition. In Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 5099\u20135103."},{"key":"e_1_3_2_16_2","first-page":"94","article-title":"Developments of Swahili resources for an automatic speech recognition system","author":"Gelas Hadrien","year":"2012","unstructured":"Hadrien Gelas, Laurent Besacier, and Francois Pellegrino. 2012. Developments of Swahili resources for an automatic speech recognition system. Spoken Language Technologies For Under-Resourced Languages. (2012), 94\u2013101.","journal-title":"Spoken Language Technologies For Under-Resourced Languages"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1109\/ASRU.2017.8268947","volume-title":"Proceedings of the 2017 IEEE Automatic Speech Recognition and Understanding Workshop","author":"Ghahremani Pegah","year":"2017","unstructured":"Pegah Ghahremani, Vimal Manohar, Hossein Hadian, Daniel Povey, and Sanjeev Khudanpur. 2017. Investigation of transfer learning for ASR using LF-MMI trained neural networks. In Proceedings of the 2017 IEEE Automatic Speech Recognition and Understanding Workshop. IEEE, 279\u2013286."},{"issue":"2","key":"e_1_3_2_18_2","first-page":"263","article-title":"Recognition of spoken bengali numerals using mlp, svm, rf based models with pca based feature summarization","volume":"15","author":"Gupta Avisek","year":"2018","unstructured":"Avisek Gupta and Kamal Sarkar. 2018. Recognition of spoken bengali numerals using mlp, svm, rf based models with pca based feature summarization. Int. Arab J. Inf. Technol 15, 2 (2018), 263\u2013269.","journal-title":"Int. Arab J. Inf. Technol"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/2500887"},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","first-page":"822","DOI":"10.1109\/ACSSC.1993.342636","volume-title":"Proceedings of 27th Asilomar Conference on Signals, Systems and Computers.","author":"Kim Ki-Chul","year":"1993","unstructured":"Ki-Chul Kim, Il-Song Han, Jun-Hee Lee, Hwang-Soo Lee, and Y. N. Yi. 1993. Spoken digit recognition using URAN (universally reconstructable artificial neural-network) VLSI chip. In Proceedings of 27th Asilomar Conference on Signals, Systems and Computers. IEEE Comput. Soc. Press, 822\u2013825."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-2620"},{"key":"e_1_3_2_22_2","unstructured":"Interspeech 2019 20th Annual Conference of the International Speech Communication Association Graz Austria 15-19 September 2019"},{"key":"e_1_3_2_23_2","article-title":"Multilingual spoken words corpus","volume":"1","author":"Mazumder Mark","year":"2021","unstructured":"Mark Mazumder, Sharad Chitlangia, Colby Banbury, Yiping Kang, Juan Ciro, Keith Achorn, Daniel Galvez, Mark Sabini, Peter Mattson, David Kanter, Greg Diamos, Pete Warden, Josh Meyer, and Vijay Janapa Reddi. 2021. Multilingual spoken words corpus. 35th Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2). 1.","journal-title":"35th Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.25080\/Majora-7b98e3ed-003"},{"issue":"6","key":"e_1_3_2_25_2","first-page":"666","article-title":"Multiclass svm based spoken hindi numerals recognition","volume":"12","author":"Mittal Teena","year":"2015","unstructured":"Teena Mittal and Rajendra Kumar Sharma. 2015. Multiclass svm based spoken hindi numerals recognition. International Arab Journal of Information Technology 12, 6 (2015), 666\u2013671.","journal-title":"International Arab Journal of Information Technology"},{"key":"e_1_3_2_26_2","first-page":"379","volume-title":"Proceedings of the 2009 12th International Conference on Computers and Information Technology.","author":"Muhammad Ghulam","year":"2009","unstructured":"Ghulam Muhammad, Yousef A. Alotaibi, and Mohammad Nurul Huda. 2009. Automatic speech recognition for Bangla digits. In Proceedings of the 2009 12th International Conference on Computers and Information Technology. IEEE, 379\u2013383."},{"key":"e_1_3_2_27_2","first-page":"300","volume-title":"2017 International Conference on Signal Processing and Communication","author":"Mukherjee Himadri","year":"2018","unstructured":"Himadri Mukherjee, Ankita Dhar, Santanu Phadikar, and Kaushik Roy. 2018. RECAL - A language identification system. 2017 International Conference on Signal Processing and Communication. 300\u2013304."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-021-10775-6"},{"key":"e_1_3_2_29_2","first-page":"74","volume-title":"Proceedings of the 2017 9th International Conference on Advanced Computational Intelligence.","author":"Nisar Shibli","year":"2017","unstructured":"Shibli Nisar, Ibrahim Shahzad, Muhammad Adnan Khan, and Muhammad Tariq. 2017. Pashto spoken digits recognition using spectral and prosodic based feature extraction. In Proceedings of the 2017 9th International Conference on Advanced Computational Intelligence. IEEE, 74\u201378."},{"key":"e_1_3_2_30_2","volume-title":"Swahili Language Handbook","author":"Polom\u00e9 Edgar C.","year":"1967","unstructured":"Edgar C. Polom\u00e9. 1967. Swahili Language Handbook. Washington, DC."},{"key":"e_1_3_2_31_2","first-page":"190","volume-title":"Proceedings of the 2013 International Conference on Control Communication and Computing","author":"Renjith S.","year":"2013","unstructured":"S. Renjith, Aju Joseph, and K. K. Anish Babu. 2013. Isolated digit recognition for Malayalam- An application perspective. In Proceedings of the 2013 International Conference on Control Communication and Computing. IEEE, 190\u2013193."},{"key":"e_1_3_2_32_2","first-page":"14","volume-title":"Proceedings of the International Conference on Computer Science, Engineering and Information Technology.","author":"Saxena Babita","year":"2015","unstructured":"Babita Saxena and Charu Wahi. 2015. Hindi digits recognition system on speech data collected in different natural noise environments. In Proceedings of the International Conference on Computer Science, Engineering and Information Technology. 14\u201315."},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","first-page":"3465","DOI":"10.21437\/Interspeech.2019-1873","volume-title":"Proceedings of the Interspeech 2019","author":"Schneider Steffen","year":"2019","unstructured":"Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019. wav2vec: Unsupervised Pre-Training for speech recognition. In Proceedings of the Interspeech 2019, ISCA, ISCA, 3465\u20133469."},{"issue":"2019","key":"e_1_3_2_34_2","doi-asserted-by":"crossref","first-page":"1381","DOI":"10.1016\/j.procs.2020.04.148","article-title":"Bengali spoken digit classification: A deep learning approach using convolutional neural network","volume":"171","author":"Sharmin Riffat","year":"2020","unstructured":"Riffat Sharmin, Shantanu Kumar Rahut, and Mohammad Rezwanul Huq. 2020. Bengali spoken digit classification: A deep learning approach using convolutional neural network. Procedia Computer Science 171, 2019 (2020), 1381\u20131388.","journal-title":"Procedia Computer Science"},{"key":"e_1_3_2_35_2","first-page":"241","volume-title":"Proceedings of the Ibero-American Conference on Artificial Intelligence","author":"Silva Diego F.","year":"2012","unstructured":"Diego F. Silva, Vin\u00edcius M. A. de Souza, Gustavo E. A. P. A. Batista, and Rafael Giusti. 2012. Spoken digit recognition in portuguese using line spectral frequencies. In Proceedings of the Ibero-American Conference on Artificial Intelligence. 241\u2013250."},{"key":"e_1_3_2_36_2","first-page":"54","volume-title":"Proceedings of the 2010 International Conference on Computer Information Systems and Industrial Management Applications","author":"Ghanty Sumit Kumar","year":"2010","unstructured":"Sumit Kumar Ghanty, Soharab Hossain Shaikh, and Nabendu Chaki. 2010. On recognition of spoken Bengali numerals. In Proceedings of the 2010 International Conference on Computer Information Systems and Industrial Management Applications. IEEE, 54\u201359."},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","first-page":"246","DOI":"10.1109\/SLT.2012.6424230","article-title":"Unsupervised cross-lingual knowledge transfer in DNN-based LVCSR","author":"Swietojanski Pawel","year":"2012","unstructured":"Pawel Swietojanski, Arnab Ghoshal, and Steve Renals. 2012. Unsupervised cross-lingual knowledge transfer in DNN-based LVCSR. In Proceedings of the 2012 IEEE Spoken Language Technology Workshop. IEEE, 246\u2013251.","journal-title":"Proceedings of the 2012 IEEE Spoken Language Technology Workshop."},{"key":"e_1_3_2_38_2","first-page":"515","volume-title":"Proceedings of the Interspeech 2013","author":"Vu Ngoc Thang","year":"2013","unstructured":"Ngoc Thang Vu and Tanja Schultz. 2013. Multilingual multilayer perceptron for rapid language adaptation between and across language families. In Proceedings of the Interspeech 2013. ISCA, 515\u2013519."},{"key":"e_1_3_2_39_2","first-page":"1225","volume-title":"Proceedings of the 2015 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference","author":"Wang Dong","year":"2015","unstructured":"Dong Wang and Thomas Fang Zheng. 2015. Transfer learning for speech and language processing. In Proceedings of the 2015 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference. IEEE, 1225\u20131237."},{"issue":"1","key":"e_1_3_2_40_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3321128","article-title":"From genesis to creole language","volume":"19","author":"Wang Hongmin","year":"2020","unstructured":"Hongmin Wang, Jie Yang, and Yue Zhang. 2020. From genesis to creole language. ACM Trans. Asian Low-Resource Lang. Inf. Process. 19, 1 (2020), 1\u201329.","journal-title":"ACM Trans. Asian Low-Resource Lang. Inf. Process."},{"key":"e_1_3_2_41_2","doi-asserted-by":"crossref","first-page":"4375","DOI":"10.21437\/Interspeech.2019-3254","volume-title":"Proceedings of the Interspeech 2019","author":"Wiesner Matthew","year":"2019","unstructured":"Matthew Wiesner, Adithya Renduchintala, Shinji Watanabe, Chunxi Liu, Najim Dehak, and Sanjeev Khudanpur. 2019. Pretraining by backtranslation for end-to-end ASR in low-resource settings. In Proceedings of the Interspeech 2019. ISCA, 4375\u20134379."},{"key":"e_1_3_2_42_2","first-page":"1","volume-title":"Proceedings of the 2018 2nd International Conference on Natural Language and Speech Processing.","author":"Zerari Naima","year":"2018","unstructured":"Naima Zerari, Samir Abdelhamid, Hassen Bouzgou, and Christian Raymond. 2018. Bi-directional recurrent end-to-end neural network classifier for spoken Arab digit recognition. In Proceedings of the 2018 2nd International Conference on Natural Language and Speech Processing. IEEE, 1\u20136."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3597494","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3597494","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:48:45Z","timestamp":1750182525000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3597494"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,20]]},"references-count":41,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2023,7,31]]}},"alternative-id":["10.1145\/3597494"],"URL":"https:\/\/doi.org\/10.1145\/3597494","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,20]]},"assertion":[{"value":"2021-10-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-05-13","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}