{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T22:27:59Z","timestamp":1784672879279,"version":"3.55.0"},"reference-count":49,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2022,3,18]],"date-time":"2022-03-18T00:00:00Z","timestamp":1647561600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Every human being experiences emotions daily, e.g., joy, sadness, fear, anger. These might be revealed through speech\u2014words are often accompanied by our emotional states when we talk. Different acoustic emotional databases are freely available for solving the Emotional Speech Recognition (ESR) task. Unfortunately, many of them were generated under non-real-world conditions, i.e., actors played emotions, and recorded emotions were under fictitious circumstances where noise is non-existent. Another weakness in the design of emotion recognition systems is the scarcity of enough patterns in the available databases, causing generalization problems and leading to overfitting. This paper examines how different recording environmental elements impact system performance using a simple logistic regression algorithm. Specifically, we conducted experiments simulating different scenarios, using different levels of Gaussian white noise, real-world noise, and reverberation. The results from this research show a performance deterioration in all scenarios, increasing the error probability from 25.57% to 79.13% in the worst case. Additionally, a virtual enlargement method and a robust multi-scenario speech-based emotion recognition system are proposed. Our system\u2019s average error probability of 34.57% is comparable to the best-case scenario with 31.55%. The findings support the prediction that simulated emotional speech databases do not offer sufficient closeness to real scenarios.<\/jats:p>","DOI":"10.3390\/s22062343","type":"journal-article","created":{"date-parts":[[2022,3,20]],"date-time":"2022-03-20T21:37:17Z","timestamp":1647812237000},"page":"2343","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":25,"title":["Robust Multi-Scenario Speech-Based Emotion Recognition System"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2797-6073","authenticated-orcid":false,"given":"Fangfang","family":"Zhu-Zhou","sequence":"first","affiliation":[{"name":"Department of Signal Theory and Communications, University of Alcal\u00e1, 28805 Alcal\u00e1 de Henares, Madrid, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1790-3834","authenticated-orcid":false,"given":"Roberto","family":"Gil-Pita","sequence":"additional","affiliation":[{"name":"Department of Signal Theory and Communications, University of Alcal\u00e1, 28805 Alcal\u00e1 de Henares, Madrid, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0890-789X","authenticated-orcid":false,"given":"Joaqu\u00edn","family":"Garc\u00eda-G\u00f3mez","sequence":"additional","affiliation":[{"name":"Department of Signal Theory and Communications, University of Alcal\u00e1, 28805 Alcal\u00e1 de Henares, Madrid, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3073-3278","authenticated-orcid":false,"given":"Manuel","family":"Rosa-Zurera","sequence":"additional","affiliation":[{"name":"Department of Signal Theory and Communications, University of Alcal\u00e1, 28805 Alcal\u00e1 de Henares, Madrid, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,3,18]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Fant, G. (1970). Acoustic Theory of Speech Production, Walter de Gruyter. Number 2.","DOI":"10.1515\/9783110873429"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"390.e21","DOI":"10.1016\/j.jvoice.2012.12.010","article-title":"Vocal indices of stress: A review","volume":"27","author":"Giddens","year":"2013","journal-title":"J. Voice"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1007\/s10919-008-0055-9","article-title":"In a nervous voice: Acoustic analysis and perception of anxiety in social phobics\u2019 speech","volume":"32","author":"Laukka","year":"2008","journal-title":"J. Nonverbal Behav."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1037\/0033-2909.99.2.143","article-title":"Vocal affect expression: A review and a model for future research","volume":"99","author":"Scherer","year":"1986","journal-title":"Psychol. Bull."},{"key":"ref_5","first-page":"137","article-title":"Affective Automotive User Interfaces\u2013Reviewing the State of Driver Affect Research and Emotion Regulation in the Car","volume":"54","author":"Braun","year":"2021","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Vidrascu, L., and Devillers, L. (2005, January 4\u20138). Detection of real-life emotions in call centers. Proceedings of the Ninth European Conference on Speech Communication and Technology, Lisbon, Portugal.","DOI":"10.21437\/Interspeech.2005-582"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1016\/S0167-6393(02)00070-5","article-title":"Emotional speech: Towards a new generation of databases","volume":"40","author":"Campbell","year":"2003","journal-title":"Speech Commun."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"14","DOI":"10.3389\/fcomp.2020.00014","article-title":"Real-time speech emotion recognition using a pre-trained image classification network: Effects of bandwidth reduction and companding","volume":"2","author":"Lech","year":"2020","journal-title":"Front. Comput. Sci."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Heracleous, P., Yasuda, K., Sugaya, F., Yoneyama, A., and Hashimoto, M. (2017, January 23\u201326). Speech emotion recognition in noisy and reverberant environments. Proceedings of the 2017 Seventh International Conference on Affective Computing and Intelligent Interaction (ACII), San Antonio, TX, USA.","DOI":"10.1109\/ACII.2017.8273610"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"457","DOI":"10.2478\/aoa-2013-0054","article-title":"Speech emotion recognition under white noise","volume":"38","author":"Huang","year":"2013","journal-title":"Arch. Acoust."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Schuller, B., Arsic, D., Wallhoff, F., and Rigoll, G. (2006, January 2\u20135). Emotion recognition in the noise applying large acoustic feature sets. Proceedings of the 3rd International Conference\u2014Speech Prosody 2006, Dresden, Germany.","DOI":"10.21437\/SpeechProsody.2006-150"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Zhu-Zhou, F., Gil-Pita, R., Garcia-Gomez, J., and Rosa-Zurera, M. (2021, January 14\u201317). Noise and Codification Effect on Emotional Speech Classification Systems. Proceedings of the IEEE\/WIC\/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology, WI\/IAT 2021, Melbourne, Australia.","DOI":"10.1145\/3498851.3499022"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1016\/j.specom.2019.12.001","article-title":"Speech emotion recognition: Emotional models, databases, features, preprocessing methods, supporting modalities, and classifiers","volume":"116","year":"2020","journal-title":"Speech Commun."},{"key":"ref_14","first-page":"400","article-title":"A review: Speech emotion recognition","volume":"6","author":"Peerzade","year":"2018","journal-title":"Int. J. Comput. Sci. Eng."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"93","DOI":"10.1007\/s10772-018-9491-z","article-title":"Databases, features and classifiers for speech emotion recognition: A review","volume":"21","author":"Swain","year":"2018","journal-title":"Int. J. Speech Technol."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1162","DOI":"10.1016\/j.specom.2006.04.003","article-title":"Emotional speech recognition: Resources, features, and methods","volume":"48","author":"Ververidis","year":"2006","journal-title":"Speech Commun."},{"key":"ref_17","first-page":"31","article-title":"Emotion recognition through the voice","volume":"9","author":"Carrera","year":"1988","journal-title":"Stud. Psychol."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Dupr\u00e9, D., Krumhuber, E.G., K\u00fcster, D., and McKeown, G.J. (2020). A performance comparison of eight commercially available automatic classifiers for facial affect recognition. PLoS ONE, 15.","DOI":"10.1371\/journal.pone.0231968"},{"key":"ref_19","unstructured":"Curumsing, M.K. (2017). Emotion-Oriented Requirements Engineering. [Ph.D. Thesis, Swinburne University of Technology]."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Ekman, P. (1992). Facial expressions of emotion: New findings, new questions. Psychol. Sci.","DOI":"10.1111\/j.1467-9280.1992.tb00253.x"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Dellaert, F., Polzin, T., and Waibel, A. (1996, January 3\u20136). Recognizing emotion in speech. Proceedings of the Fourth International Conference on Spoken Language Processing. ICSLP\u201996, Philadelphia, PA, USA.","DOI":"10.21437\/ICSLP.1996-462"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Kwon, O.W., Chan, K., Hao, J., and Lee, T.W. (2003, January 1\u20134). Emotion recognition by speech signals. Proceedings of the Eighth European Conference on Speech Communication and Technology, Geneva, Switzerland.","DOI":"10.21437\/Eurospeech.2003-80"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Schuller, B., Rigoll, G., and Lang, M. (1003, January 6\u201310). Hidden Markov model-based speech emotion recognition. Proceedings of the 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings.(ICASSP\u201903), Hong Kong, China.","DOI":"10.1109\/ICME.2003.1220939"},{"key":"ref_24","unstructured":"Ververidis, D., and Kotropoulos, C. (2005, January 6). Emotional speech classification using Gaussian mixture models and the sequential floating forward selection algorithm. Proceedings of the 2005 IEEE International Conference on Multimedia and Expo, Amsterdam, The Netherlands."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Busso, C., Deng, Z., Yildirim, S., Bulut, M., Lee, C.M., Kazemzadeh, A., Lee, S., Neumann, U., and Narayanan, S. (2004, January 13\u201315). Analysis of emotion recognition using facial expressions, speech and multimodal information. Proceedings of the 6th International Conference on Multimodal Interfaces, State College, PA, USA.","DOI":"10.1145\/1027933.1027968"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Luengo, I., Navas, E., Hern\u00e1ez, I., and S\u00e1nchez, J. (2005, January 4\u20138). Automatic emotion recognition using prosodic parameters. Proceedings of the Ninth European Conference on Speech Communication and Technology, Lisbon, Portugal.","DOI":"10.21437\/Interspeech.2005-324"},{"key":"ref_27","unstructured":"Borchert, M., and Dusterhoft, A. (November, January 30). Emotions in speech-experiments with prosody and quality features in speech for use in categorical and dimensional emotion recognition environments. Proceedings of the 2005 International Conference on Natural Language Processing and Knowledge Engineering, Wuhan, China."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Mirsamadi, S., Barsoum, E., and Zhang, C. (2017, January 5\u20139). Automatic speech emotion recognition using recurrent neural networks with local attention. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7952552"},{"key":"ref_29","unstructured":"Ruder, S. (2017). An overview of multi-task learning in deep neural networks. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Trigeorgis, G., Ringeval, F., Brueckner, R., Marchi, E., Nicolaou, M.A., Schuller, B., and Zafeiriou, S. (2016, January 20\u201325). Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network. Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472669"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Atcheson, M., Sethu, V., and Epps, J. (2018, January 2\u20136). Demonstrating and Modelling Systematic Time-varying Annotator Disagreement in Continuous Emotion Annotation. Proceedings of the Interspeech, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-1933"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Aouani, H., and Ben Ayed, Y. (2018, January 21\u201324). Emotion recognition in speech using MFCC with SVM, DSVM and auto-encoder. Proceedings of the 2018 4th International Conference on Advanced Technologies for Signal and Image Processing (ATSIP), Sousse, Tunisia.","DOI":"10.1109\/ATSIP.2018.8364518"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Krishna Kishore, K., and Krishna Satish, P. (2013, January 22\u201323). Emotion recognition in speech using MFCC and wavelet features. Proceedings of the 2013 3rd IEEE International Advance Computing Conference (IACC), Ghaziabad, India.","DOI":"10.1109\/IAdCC.2013.6514336"},{"key":"ref_34","unstructured":"Jain, M., Narayan, S., Balaji, P., Bhowmick, A., and Muthu, R.K. (2020). Speech emotion recognition using support vector machine. arXiv."},{"key":"ref_35","first-page":"11","article-title":"Body size projection by voice quality in emotional speech\u2014Evidence from Mandarin Chinese","volume":"5","author":"Liu","year":"2014","journal-title":"Perception"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1340","DOI":"10.3389\/fpsyg.2015.01340","article-title":"Sound frequency affects speech emotion perception: Results from congenital amusia","volume":"6","author":"Lolli","year":"2015","journal-title":"Front. Psychol."},{"key":"ref_37","unstructured":"Gonzalez, S., and Brookes, M. (Spetember, January 29). A pitch estimation filter robust to high levels of noise (PEFAC). Proceedings of the 2011 19th European Signal Processing Conference, Barcelona, Spain."},{"key":"ref_38","unstructured":"Lovric, M. (2011). Bootstrap Methods, Springer."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"235","DOI":"10.1007\/s10994-007-5019-5","article-title":"Active learning for logistic regression: An evaluation","volume":"68","author":"Schein","year":"2007","journal-title":"Mach. Learn."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"19","DOI":"10.1007\/s10462-009-9114-9","article-title":"A review on the combination of binary classifiers in multiclass problems","volume":"30","author":"Lorena","year":"2008","journal-title":"Artif. Intell. Rev."},{"key":"ref_41","unstructured":"Mohino-Herranz, I., S\u00e1nchez-Hevia, H.A., Gil-Pita, R., and Rosa-Zurera, M. (2014). Creation of New Virtual Patterns for Emotion Recognition through PSOLA, Audio Engineering Society. Audio Engineering Society Convention 136."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Burkhardt, F., Paeschke, A., Rolfes, M., Sendlmeier, W.F., and Weiss, B. (2005, January 4\u20138). A database of German emotional speech. Proceedings of the Ninth European Conference on Speech Communication and Technology, Lisbon, Portugal.","DOI":"10.21437\/Interspeech.2005-446"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"190","DOI":"10.1109\/TAFFC.2015.2457417","article-title":"The Geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing","volume":"7","author":"Eyben","year":"2015","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"2203","DOI":"10.1109\/TMM.2014.2360798","article-title":"Learning salient features for speech emotion recognition using convolutional neural networks","volume":"16","author":"Mao","year":"2014","journal-title":"IEEE Trans. Multimed."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Mohino, I., Goni, M., Alvarez, L., Llerena, C., and Gil-Pita, R. (2013, January 12\u201314). Detection of emotions and stress through speech analysis. Proceedings of the Signal Processing, Pattern Recognition and Application-2013, Innsbruck, Austria.","DOI":"10.2316\/P.2013.798-071"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"768","DOI":"10.1016\/j.specom.2010.08.013","article-title":"Automatic speech emotion recognition using modulation spectral features","volume":"53","author":"Wu","year":"2011","journal-title":"Speech Commun."},{"key":"ref_47","unstructured":"Snyder, D., Chen, G., and Povey, D. (2015). Musan: A music, speech, and noise corpus. arXiv."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"182","DOI":"10.1016\/j.apacoust.2008.02.003","article-title":"Fast image method for impulse response calculations of box-shaped rooms","volume":"70","author":"McGovern","year":"2009","journal-title":"Appl. Acoust."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"377","DOI":"10.1109\/TAFFC.2014.2336244","article-title":"Crema-d: Crowd-sourced emotional multimodal actors dataset","volume":"5","author":"Cao","year":"2014","journal-title":"IEEE Trans. Affect. Comput."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/6\/2343\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:38:43Z","timestamp":1760135923000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/6\/2343"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,18]]},"references-count":49,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2022,3]]}},"alternative-id":["s22062343"],"URL":"https:\/\/doi.org\/10.3390\/s22062343","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,3,18]]}}}