{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T08:11:48Z","timestamp":1783411908218,"version":"3.54.6"},"reference-count":27,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2025,2,12]],"date-time":"2025-02-12T00:00:00Z","timestamp":1739318400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Parkinson\u2019s disease (PD) is a progressive degenerative brain disease that worsens with age, causing areas of the brain to weaken. Vocal dysfunction often emerges as one of the earliest and most prominent indicators of Parkinson\u2019s disease, with a significant number of patients exhibiting vocal impairments during the initial stages of the illness. In view of this, to facilitate the diagnosis of Parkinson\u2019s disease through the analysis of these vocal characteristics, this study focuses on exerting a combination of mel spectrogram and MFCC as spectral features. This study adopts Italian raw audio data to establish an efficient detection framework specifically designed to classify the vocal data into two distinct categories: healthy individuals and patients diagnosed with Parkinson\u2019s disease. To this end, the study proposes a hybrid model that integrates Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs) for the detection of Parkinson\u2019s disease. Certainly, CNNs are employed to extract spatial features from the extracted spectro-temporal characteristics of vocal data, while LSTMs capture temporal dependencies, accelerating a comprehensive analysis of the development of vocal patterns over time. Additionally, the merging of a multi-head attention mechanism significantly enhances the model\u2019s ability to concentrate on essential details, hence improving its overall performance. This unified method aims to enhance the detection of subtle vocal changes associated with Parkinson\u2019s, enhancing overall diagnostic accuracy. The findings declare that this model achieves a noteworthy accuracy of 99.00% for the Parkinson\u2019s disease detection process.<\/jats:p>","DOI":"10.3390\/info16020135","type":"journal-article","created":{"date-parts":[[2025,2,12]],"date-time":"2025-02-12T03:41:57Z","timestamp":1739331717000},"page":"135","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["PD-Net: Parkinson\u2019s Disease Detection Through Fusion of Two Spectral Features Using Attention-Based Hybrid Deep Neural Network"],"prefix":"10.3390","volume":"16","author":[{"given":"Munira","family":"Islam","sequence":"first","affiliation":[{"name":"Department of Electronics and Telecommunication Engineering, Chittagong University of Engineering & Technology, Chattogram 4349, Bangladesh"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2948-6368","authenticated-orcid":false,"given":"Khadija","family":"Akter","sequence":"additional","affiliation":[{"name":"Department of Electronics and Telecommunication Engineering, Chittagong University of Engineering & Technology, Chattogram 4349, Bangladesh"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8251-5168","authenticated-orcid":false,"given":"Md. Azad","family":"Hossain","sequence":"additional","affiliation":[{"name":"Department of Electronics and Telecommunication Engineering, Chittagong University of Engineering & Technology, Chattogram 4349, Bangladesh"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6347-7509","authenticated-orcid":false,"given":"M. Ali Akber","family":"Dewan","sequence":"additional","affiliation":[{"name":"School of Computing and Information Systems, Faculty of Science and Technology, Athabasca University, Athabasca, AB T9S 3A3, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,2,12]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"368","DOI":"10.1136\/jnnp.2007.131045","article-title":"Parkinson\u2019s disease: Clinical features and diagnosis","volume":"79","author":"Jankovic","year":"2008","journal-title":"J. Neurol. Neurosurg. Psychiatry"},{"key":"ref_2","first-page":"S3","article-title":"The Emerging Evidence of the Parkinson Pandemic","volume":"8","author":"Dorsey","year":"2018","journal-title":"J. Park. Dis."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Azadi, H., Akbarzadeh-T, M.R., Shoeibi, A., and Kobravi, H.R. (2021). Evaluating the Effect of Parkinson\u2019s Disease on Jitter and Shimmer Speech Features. Adv. Biomed. Res., 10.","DOI":"10.4103\/abr.abr_254_21"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Movement Disorder Society Task Force on Rating Scales for Parkinson\u2019s Disease (2003). The Unified Parkinson\u2019s Disease Rating Scale (UPDRS): Status and recommendations. Mov. Disord. Off. J. Mov. Disord. Soc., 18, 738\u2013750.","DOI":"10.1002\/mds.10473"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"457","DOI":"10.1097\/00019052-200208000-00009","article-title":"Clinical aspects of parkinson disease","volume":"15","author":"Sethi","year":"2022","journal-title":"Curr. Opin. Neurol."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Alshammri, R., Alharbi, G., Alharbi, E., and Almubark, I. (2023). Machine learning approaches to identify parkinson\u2019s disease using voice signal features. Front. Artif. Intell., 6.","DOI":"10.3389\/frai.2023.1084001"},{"key":"ref_7","unstructured":"Zi\u00f3lko, M. (2025, January 27). Speech Analysis as a Tool for Detection and Monitoring of Medical Conditions: A Review (Preprint). Available online: https:\/\/journals.pan.pl\/dlibra\/show-content?id=128239."},{"key":"ref_8","first-page":"1129","article-title":"Parkinson\u2019s disease","volume":"5510","year":"2022","journal-title":"Rev. Prat."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"548","DOI":"10.1001\/jama.2019.22360","article-title":"Diagnosis and treatment of parkinson disease: A review","volume":"323","author":"Armstrong","year":"2020","journal-title":"JAMA"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Wu, H., Soraghan, J.J., Lowit, A., and Caterina, G.D. (2018, January 2\u20136). A deep learning method for pathological voice detection using convolutional deep belief networks. Proceedings of the Interspeech 2018, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-1351"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Sakar, B.E., Serbes, G., and Sakar, C.O. (2017). Analyzing the effectiveness of vocal features in early telediagnosis of Parkinson\u2019s disease. PLoS ONE, 12.","DOI":"10.1371\/journal.pone.0182428"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"e26305","DOI":"10.2196\/26305","article-title":"Detecting Parkinson\u2019s disease using a web-based speech task: Observational study","volume":"23","author":"Rahman","year":"2021","journal-title":"J. Med. Internet Res."},{"key":"ref_13","first-page":"611","article-title":"Magnetic resonance imaging markers for early diagnosis of parkinson\u2019s diseaseff","volume":"7","author":"Marino","year":"2012","journal-title":"Neural Regen. Res."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Maffia, M., Micco, R.D., Pettorino, M., Siciliano, M., Tessitore, A., and Meo, A.D. (2021). Speech rhythm variation in early-stage Parkinson\u2019s disease: A study on different speaking tasks. Front. Psychol., 12.","DOI":"10.3389\/fpsyg.2021.668291"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"100181","DOI":"10.1016\/j.health.2023.100181","article-title":"An ensemble nearest neighbor boosting technique for prediction of Parkinson\u2019s disease","volume":"3","author":"Shastry","year":"2023","journal-title":"Healthc. Anal."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"95","DOI":"10.3233\/JPD-230002","article-title":"Machine learning-based classification of parkinson\u2019s disease patients using speech biomarkers","volume":"14","author":"Hossain","year":"2023","journal-title":"J. Parkinson\u2019s Dis."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"249","DOI":"10.1016\/j.procs.2023.01.007","article-title":"Early detection of parkinson\u2019s disease using machine learning","volume":"218","author":"Govindu","year":"2023","journal-title":"Procedia Comput. Sci."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Lilhore, U.K., Dalal, S., Faujdar, N., Margala, M., Chakrabarti, P., Chakrabarti, T., Simaiya, S., Kumar, P., Thangaraju, P., and Velmurugan, H. (2023). Hybrid cnn-lstm model with efficient hyperparameter tuning for prediction of parkinson\u2019s disease. Sci. Rep., 13.","DOI":"10.1038\/s41598-023-41314-y"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"4849","DOI":"10.1007\/s00521-020-05233-7","article-title":"Multi-variate vocal data analysis for detection of parkinson disease using deep learning","volume":"33","author":"Nagasubramanian","year":"2020","journal-title":"Neural Comput. Appl."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"22199","DOI":"10.1109\/ACCESS.2017.2762475","article-title":"Assessment of speech intelligibility in parkinson\u2019s disease using a speech-to-text system","volume":"5","author":"Dimauro","year":"2017","journal-title":"IEEE Access"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Abayomi-Alli, O.O., Dama\u0161evi\u010dius, R., Qazi, A., Adedoyin-Olowe, M., and Misra, S. (2022). Data augmentation and deep learning methods in sound classification: A systematic review. Electronics, 11.","DOI":"10.3390\/electronics11223795"},{"key":"ref_22","unstructured":"Surina, S. (2025, January 27). Time and Pitch Scaling in Audio Processing. Available online: https:\/\/www.surina.net\/article\/time-and-pitch-scaling.html."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wongvorachan, T., He, S., and Bulut, O. (2023). A comparison of undersampling, oversampling, and smote methods for dealing with imbalanced classification in educational data mining. Information, 14.","DOI":"10.3390\/info14010054"},{"key":"ref_24","unstructured":"Oppenheim, A.V., and Schafer, R.W. (2009). Discrete-Time Signal Processing, Pearson. [3rd ed.]."},{"key":"ref_25","unstructured":"M\u00fcller, M. (2015). Fundamentals of Audio and Music Processing: With Applications to Signal Processing and Music Information Retrieval, Springer."},{"key":"ref_26","unstructured":"Jurafsky, D., and Martin, J.H. (2020). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition, Pearson. [3rd ed.]."},{"key":"ref_27","unstructured":"Deora, P., Ghaderi, R., Taheri, H., and Thrampoulidis, C. (2023). On the optimization and generalization of multi-head attention. arXiv."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/2\/135\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T16:31:51Z","timestamp":1760027511000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/2\/135"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,12]]},"references-count":27,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,2]]}},"alternative-id":["info16020135"],"URL":"https:\/\/doi.org\/10.3390\/info16020135","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,12]]}}}