{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,7]],"date-time":"2026-05-07T15:54:47Z","timestamp":1778169287302,"version":"3.51.4"},"reference-count":49,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2023,1,11]],"date-time":"2023-01-11T00:00:00Z","timestamp":1673395200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100007421","name":"Li Ka Shing Foundation","doi-asserted-by":"publisher","award":["2020LKSFG04C"],"award-info":[{"award-number":["2020LKSFG04C"]}],"id":[{"id":"10.13039\/100007421","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The rehabilitation of aphasics is fundamentally based on the assessment of speech impairment. Developing methods for assessing speech impairment automatically is important due to the growing number of stroke cases each year. Traditionally, aphasia is assessed manually using one of the well-known assessment batteries, such as the Western Aphasia Battery (WAB), the Chinese Rehabilitation Research Center Aphasia Examination (CRRCAE), and the Boston Diagnostic Aphasia Examination (BDAE). In aphasia testing, a speech-language pathologist (SLP) administers multiple subtests to assess people with aphasia (PWA). The traditional assessment is a resource-intensive process that requires the presence of an SLP. Thus, automating the assessment of aphasia is essential. This paper evaluated and compared custom machine learning (ML) speech recognition algorithms against off-the-shelf platforms using healthy and aphasic speech datasets on the naming and repetition subtests of the aphasia battery. Convolutional neural networks (CNN) and linear discriminant analysis (LDA) are the customized ML algorithms, while Microsoft Azure and Google speech recognition are off-the-shelf platforms. The results of this study demonstrated that CNN-based speech recognition algorithms outperform LDA and off-the-shelf platforms. The ResNet-50 architecture of CNN yielded an accuracy of 99.64 \u00b1 0.26% on the healthy dataset. Even though Microsoft Azure was not trained on the same healthy dataset, it still generated comparable results to the LDA and superior results to Google\u2019s speech recognition platform.<\/jats:p>","DOI":"10.3390\/s23020857","type":"journal-article","created":{"date-parts":[[2023,1,12]],"date-time":"2023-01-12T03:47:01Z","timestamp":1673495221000},"page":"857","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":19,"title":["A Comparative Investigation of Automatic Speech Recognition Platforms for Aphasia Assessment Batteries"],"prefix":"10.3390","volume":"23","author":[{"given":"Seedahmed S.","family":"Mahmoud","sequence":"first","affiliation":[{"name":"Department of Biomedical Engineering, Shantou University, Shantou 515063, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Raphael F.","family":"Pallaud","sequence":"additional","affiliation":[{"name":"Computer and Information Technology Department, IT Institute @ Phoenix College, Phoenix, AZ 85013, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3684-8498","authenticated-orcid":false,"given":"Akshay","family":"Kumar","sequence":"additional","affiliation":[{"name":"Department of Biomedical Engineering, Shantou University, Shantou 515063, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7371-9516","authenticated-orcid":false,"given":"Serri","family":"Faisal","sequence":"additional","affiliation":[{"name":"Computer and Information Technology Department, IT Institute @ Phoenix College, Phoenix, AZ 85013, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yin","family":"Wang","sequence":"additional","affiliation":[{"name":"Department of Biomedical Engineering, Shantou University, Shantou 515063, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qiang","family":"Fang","sequence":"additional","affiliation":[{"name":"Department of Biomedical Engineering, Shantou University, Shantou 515063, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,1,11]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"439","DOI":"10.1161\/CIRCRESAHA.116.308413","article-title":"Global Burden of Stroke","volume":"120","author":"Feigin","year":"2017","journal-title":"Circ. Res."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"417","DOI":"10.1016\/S1474-4422(19)30030-4","article-title":"The global burden of stroke. Persistent and disabling","volume":"18","author":"Gorelick","year":"2019","journal-title":"Lancet Neurol."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"172","DOI":"10.1044\/jslhr.4101.172","article-title":"A meta-analysis of clinical outcomes in the treatment of aphasia","volume":"41","author":"Robey","year":"1998","journal-title":"J. Speech Lang. Hear. Res. JSLHR"},{"key":"ref_4","unstructured":"Chinese Rehabilitation Research Center (2022, February 02). Chinese Rehabilitation Research Center Aphasia Examination (CRRCAE). Available online: https:\/\/wenku.baidu.com\/view\/e209482cbd64783e09122bb5.html."},{"key":"ref_5","unstructured":"Rose, F.C. (1984). The Aachen aphasia test. Advances in Neurology. Progress in Aphasiology, Raven Press."},{"key":"ref_6","unstructured":"Goodglass, H., and Kaplan, E. (1983). The Assessment of Aphasia and Related Disorders, Williams & Wilkins."},{"key":"ref_7","first-page":"323","article-title":"Reliability and validity of bedside version of arabic diagnostic aphasia battery (A-DAB-1) for Lebanese Individuals","volume":"32","author":"Nilipour","year":"2017","journal-title":"Aphasiology"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"819","DOI":"10.1007\/s11265-019-01511-3","article-title":"An End-to-End Approach to Automatic Speech Assessment for Cantonese-speaking People with Aphasia","volume":"92","author":"Qin","year":"2020","journal-title":"J. Signal Process. Syst."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1016\/j.compeleceng.2016.08.021","article-title":"An incremental method combining density clustering and support vector machines for voice pathology detection","volume":"57","author":"Amami","year":"2017","journal-title":"Comput. Electr. Eng."},{"key":"ref_10","first-page":"39","article-title":"Jisuanji fuzhu yanyu jiaozhi xitong. [A computer aided speech correction system]","volume":"14","author":"Ding","year":"1995","journal-title":"Zhongguo Shengwu Yixue Gongcheng Xuebao"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"3191","DOI":"10.1109\/JBHI.2020.3011104","article-title":"An Efficient Deep Learning Based Method for Speech Assessment of Mandarin-Speaking Aphasic Patients","volume":"24","author":"Mahmoud","year":"2020","journal-title":"IEEE J. Biomed. Health Inform."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"2187","DOI":"10.1109\/TASLP.2016.2598428","article-title":"Automatic Assessment of Speech Intelligibility for Individuals with Aphasia","volume":"24","author":"Le","year":"2016","journal-title":"IEEE ACM Trans. Audio Speech Lang. Process"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"349","DOI":"10.1016\/j.cmpb.2011.02.015","article-title":"Comparison of machine learning methods for classifying aphasic and non-aphasic speakers","volume":"104","author":"Juhola","year":"2011","journal-title":"Comput. Methods Programs Biomed."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1682","DOI":"10.1002\/hbm.25321","article-title":"Machine learning-based multimodal prediction of language outcomes in chronic aphasia","volume":"42","author":"Kristinsson","year":"2021","journal-title":"Hum. Brain Mapp."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Qin, Y., Lee, T., Feng, S., and Kong, A.-H. (2018, January 2\u20136). Automatic Speech Assessment for People with Aphasia Using TDNN-BLSTM with Multi-Task Learning. Proceedings of the Interspeech, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-1630"},{"key":"ref_16","unstructured":"Le, D. (2017). Towards Automatic Speech-Language Assessment for Aphasia Rehabilitation. [Ph.D. Thesis, University of Michigan]."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1264","DOI":"10.1109\/TBME.2012.2183367","article-title":"Novel Speech Signal Processing Algorithms for High-Accuracy Classification of Parkinson\u2019s Disease","volume":"59","author":"Tsanas","year":"2012","journal-title":"IEEE Trans. Biomed. Eng."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Shahin, M., Ahmed, B., McKechnie, J., Ballard, K., and Gutierrez-Osuna, R. (2014, January 14\u201318). A comparison of GMM-HMM and DNN-HMM based pronunciation verification techniques for use in the assessment of childhood apraxia of speech. Proceedings of the 15th Annual Conference of the International Speech Communication Association: Celebrating the Diversity of Spoken Languages, Interspeech, Singapore.","DOI":"10.21437\/Interspeech.2014-377"},{"key":"ref_19","unstructured":"Li, A.N. (2010). Shiyuzheng Huanzhe Yuyin Xinhao de Shibie Yanjiu. An Investigation on Speech Recognition of Aphasia Patients. [Master\u2019s Thesis, Xi\u2019an University of Science and Technology]."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Day, M., Dey, R., and Khojandi, A. (2021, January 1\u20135). Predicting Severity in People with Aphasia. A Natural Language Processing and Machine Learning Approach. Computer Science. Proceedings of the 2021 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), Guadalajara, Mexico.","DOI":"10.1109\/EMBC46164.2021.9630694"},{"key":"ref_21","unstructured":"Simplilearn Solutions (2022, July 23). What Is Microsoft Azure: How Does It Work and Services. Available online: https:\/\/www.simplilearn.com\/tutorials\/azure-tutorial\/what-is-azure."},{"key":"ref_22","unstructured":"Microsoft (2022, July 23). Browse Azure Products-AI + Machine Learning. Available online: https:\/\/docs.microsoft.com\/en-us\/azure\/?product=ai-machine-learning."},{"key":"ref_23","unstructured":"GeeksforGeeks (2022, July 23). REST API (Introduction). Available online: https:\/\/www.geeksforgeeks.org\/rest-api-introduction\/."},{"key":"ref_24","unstructured":"Microsoft (2022, July 23). What Is Speech-to-Text. Available online: https:\/\/docs.microsoft.com\/EN-US\/azure\/cognitive-services\/speech-service\/speech-to-text."},{"key":"ref_25","unstructured":"Microsoft (2022, July 23). Language and Voice Support for the Speech Service. Available online: https:\/\/docs.microsoft.com\/en-us\/azure\/cognitive-services\/speech-service\/language-support?tabs=speechtotext."},{"key":"ref_26","unstructured":"Microsoft (2022, July 23). Speech-to-Text Documentation. Available online: https:\/\/docs.microsoft.com\/en-us\/azure\/cognitive-services\/speech-service\/index-speech-to-text."},{"key":"ref_27","unstructured":"Microsoft (2022, July 23). Real-Time Speech-to-Text. Available online: https:\/\/speech.microsoft.com\/portal\/speechtotexttool."},{"key":"ref_28","unstructured":"Google (2022, July 23). Language Support. Available online: https:\/\/cloud.google.com\/speech-to-text\/docs\/languages."},{"key":"ref_29","unstructured":"Google (2022, July 23). Unveiling a New Visual User Interface for Google Cloud\u2019s Speech-to-Text API. Available online: https:\/\/cloud.google.com\/blog\/products\/ai-machine-learning\/google-clouds-new-visual-interface-for-speech-to-text-api."},{"key":"ref_30","unstructured":"Google (2022, July 23). Speech-to-Text Basics. Available online: https:\/\/cloud.google.com\/speech-to-text\/docs\/basics."},{"key":"ref_31","unstructured":"Google (2022, July 23). Speech-to-Text. Available online: https:\/\/cloud.google.com\/speech-to-text\/."},{"key":"ref_32","unstructured":"Schalkwyk, J. (2022, July 23). An All-Neural On-Device Speech Recognizer. Available online: https:\/\/ai.googleblog.com\/2019\/03\/an-all-neural-on-device-speech.html."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1866","DOI":"10.1109\/TSP.2002.800406","article-title":"Adaptive instantaneous frequency estimation of multicomponent FM signals using quadratic time-frequency distributions","volume":"50","author":"Hussain","year":"2002","journal-title":"IEEE Trans. Signal Process."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"133","DOI":"10.1016\/j.bspc.2006.02.001","article-title":"Time-frequency analysis of normal and abnormal biological signals","volume":"1","author":"Mahmoud","year":"2006","journal-title":"Biomed. Signal Process. Control"},{"key":"ref_35","unstructured":"Hussain, Z., and Boashash, B. (2001, January 13\u201316). Design of time-frequency distributions for amplitude and IF estimation of multicomponent signals. Proceedings of the Sixth International Symposium on Signal Processing and Its Applications, Kuala Lumpur, Malaysia."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Dodge, S., and Karam, L. (2016, January 6\u20138). Understanding how image quality affects deep neural networks. Proceedings of the 2016 Eighth International Conference on Quality of Multimedia Experience (QoMEX), Lisbon, Portugal.","DOI":"10.1109\/QoMEX.2016.7498955"},{"key":"ref_37","first-page":"451","article-title":"Effects of Varying Resolution on Performance of CNN based Image Classification: An Experimental Study","volume":"6","author":"Kannojia","year":"2018","journal-title":"Int. J. Comput. Sci. Eng."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Smith, L. (2017, January 24\u201331). Cyclical learning rates for training neural networks. Proceedings of the IEEE Winter Conference on Applications of Computer Vision (ACV), Santa Rosa, CA, USA.","DOI":"10.1109\/WACV.2017.58"},{"key":"ref_40","unstructured":"Kingma, D., and Adam, J.B. (2014). A Method for Stochastic Optimization. arXiv."},{"key":"ref_41","unstructured":"Krogh, A., and Hertz, J. (1991, January 2\u20135). A simple weight decay can improve generalization. Proceedings of the 4th International Conference on Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_42","unstructured":"Howard, J. (2021, December 07). Fastai. Available online: https:\/\/github.com\/fastai\/fastai."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Kohlschein, C., Schmitt, M., Schuller, B., Jeschke, S., and Werner, C. (2017, January 12\u201315). A machine learning based system for the automatic evaluation of aphasia speech. Proceedings of the 2017 IEEE 19th International Conference on e-Health Networking, Applications and Services (Healthcom), Dalian, China.","DOI":"10.1109\/HealthCom.2017.8210766"},{"key":"ref_44","unstructured":"Stack Overflow (2022, July 23). How to Enable Word Level Confidence for MS Azure Speech to Text Service. Available online: https:\/\/stackoverflow.com\/questions\/60229786\/how-to-enable-word-level-confidence-for-ms-azure-speech-to-text-service."},{"key":"ref_45","unstructured":"Google (2022, July 23). Enable Word-Level Confidence. Available online: https:\/\/cloud.google.com\/speech-to-text\/docs\/word-confidence."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"16246","DOI":"10.1109\/ACCESS.2018.2816338","article-title":"Voice Disorder Identification by Using Machine Learning Techniques","volume":"6","author":"Verde","year":"2018","journal-title":"IEEE Access"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Christensen, H., Cunningham, S., Fox, C., Green, P., and Hain, T. (2012, January 9\u201313). A comparative study of adaptive, automatic recognition of disordered speech. Proceedings of the Thirteenth Annual Conference of the International Speech Communication Association, Portland, OR, USA.","DOI":"10.21437\/Interspeech.2012-484"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Mengistu, K., and Rudzicz, F. (2011, January 5\u20138). Comparing Humans and Automatic Speech Recognition Systems in Recognizing Dysarthric Speech. Proceedings of the Advances in Artificial Intelligence, Perth, Australia.","DOI":"10.1007\/978-3-642-21043-3_36"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"3924","DOI":"10.1016\/j.eswa.2015.01.033","article-title":"Exploring the influence of general and specific factors on the recognition accuracy of an ASR system for dysarthric speaker","volume":"42","author":"Mustafa","year":"2015","journal-title":"Expert Syst. Appl."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/2\/857\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:03:33Z","timestamp":1760119413000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/2\/857"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,11]]},"references-count":49,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2023,1]]}},"alternative-id":["s23020857"],"URL":"https:\/\/doi.org\/10.3390\/s23020857","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,11]]}}}