{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,28]],"date-time":"2026-02-28T04:30:48Z","timestamp":1772253048329,"version":"3.50.1"},"reference-count":29,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2022,2,22]],"date-time":"2022-02-22T00:00:00Z","timestamp":1645488000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The real challenge in Human-Robot Interaction (HRI) is to build machines capable of perceiving human emotions so that robots can interact with humans in a proper manner. Emotion varies accordingly to many factors, and gender represents one of the most influential ones: an appropriate gender-dependent emotion recognition system is recommended indeed. In this article, we propose a Gender Recognition (GR) module for the gender identification of the speaker, as a preliminary step for the final development of a Speech Emotion Recognition (SER) system. The system was designed to be installed on social robots for hospitalized and living at home patients monitoring. Hence, the importance of reducing the software computational effort of the architecture also minimizing the hardware bulkiness, in order for the system to be suitable for social robots. The algorithm was executed on the Raspberry Pi hardware. For the training, the Italian emotional database EMOVO was used. Results show a GR accuracy value of 97.8%, comparable with the ones found in the literature.<\/jats:p>","DOI":"10.3390\/s22051714","type":"journal-article","created":{"date-parts":[[2022,2,22]],"date-time":"2022-02-22T22:35:00Z","timestamp":1645569300000},"page":"1714","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Gender Identification in a Two-Level Hierarchical Speech Emotion Recognition System for an Italian Social Robot"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3523-4312","authenticated-orcid":false,"given":"Antonio","family":"Guerrieri","sequence":"first","affiliation":[{"name":"Fondazione Neurone Onlus, Viale Regina Margherita 169, 00198 Roma, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eleonora","family":"Braccili","sequence":"additional","affiliation":[{"name":"Fondazione Neurone Onlus, Viale Regina Margherita 169, 00198 Roma, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Federica","family":"Sgr\u00f2","sequence":"additional","affiliation":[{"name":"Fondazione Neurone Onlus, Viale Regina Margherita 169, 00198 Roma, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Giulio Nicol\u00f2","family":"Meldolesi","sequence":"additional","affiliation":[{"name":"Fondazione Neurone Onlus, Viale Regina Margherita 169, 00198 Roma, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,2,22]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Beer, J., Liles, K., Wu, X., and Pakala, S. (2017). Affective Human\u2013Robot Interaction. Emotions and Affect in Human Factors and Human-Computer Interaction, Academic Press.","DOI":"10.1016\/B978-0-12-801851-4.00015-X"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Bartneck, C., Belpaeme, T., Eyssel, F., Kanda, T., Keijsers, M., and Sabanovic, S. (2020). Human-Robot Interaction\u2014An Introduction, Cambridge University Press. Chapter 2.","DOI":"10.1017\/9781108676649"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"630","DOI":"10.1177\/00100002030004008","article-title":"Sex Differences in Emotion: A Critical Review of the Literature and Implications for Counseling Psychology","volume":"30","author":"Wester","year":"2002","journal-title":"Couns. Psychol.\u2014Couns Psychol."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Zhang, L., Wang, L., Dang, J., Guo, L., and Yu, Q. (2018, January 4\u20137). Gender-Aware CNN-BLSTM for Speech Emotion Recognition. Proceedings of the 27th International Conference on Artificial Neural Networks, Rhodes, Greece.","DOI":"10.1007\/978-3-030-01418-6_76"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"5115","DOI":"10.1016\/j.eswa.2011.11.028","article-title":"Cultural dependency analysis for understanding speech emotion","volume":"39","author":"Kamaruddin","year":"2012","journal-title":"Expert Syst. Appl."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Verma, D., Mukhopadhyay, D., and Mark, E. (2016, January 12\u201313). Role of gender influence in vocal Hindi conversations: A study on speech emotion recognition. Proceedings of the 2016 International Conference on Computing Communication Control and automation (ICCUBEA), Pune, India.","DOI":"10.1109\/ICCUBEA.2016.7860021"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Fu, L., Wang, C., and Zhang, Y. (2010, January 5\u20137). A study on influence of gender on speech emotion classification. Proceedings of the 2010 2nd International Conference on Signal Processing Systems, Dalian, China.","DOI":"10.1109\/ICSPS.2010.5555556"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"766","DOI":"10.1016\/j.chb.2007.04.004","article-title":"The role of emotion in computer-mediated communication: A review","volume":"24","author":"Derks","year":"2008","journal-title":"Comput. Hum. Behav."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"22","DOI":"10.4018\/IJIIT.2019100102","article-title":"Speech Emotion Recognition Based on Gender Influence in Emotional Expression","volume":"15","author":"Vasuki","year":"2019","journal-title":"Int. J. Intell. Inf. Technol."},{"key":"ref_10","unstructured":"Vogt, T., and Andr\u00e9, E. (2006, January 24\u201326). Improving Automatic Emotion Recognition from Speech via Gender Differentiation. Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC 2006), Genoa, Italy."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Vinay, S., Gupta, S., and Mehra, A. (2014, January 20\u201321). Gender specific emotion recognition through speech signals. Proceedings of the 2014 International Conference on Signal Processing and Integrated Networks (SPIN), Noida, India.","DOI":"10.1109\/SPIN.2014.6777050"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1016\/j.procs.2019.04.009","article-title":"Recognizing Emotion from Speech Based on Age and Gender Using Hierarchical Models","volume":"151","author":"Shaqra","year":"2019","journal-title":"Procedia Comput. Sci."},{"key":"ref_13","unstructured":"Titze, I. (1994). Principles of Voice Production, Prentice Hall (Currently Published by NCVS.org)."},{"key":"ref_14","unstructured":"Baken, R.J. (1987). Clinical Measurement of Speech and Voice, Taylor and Francis Ltd."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"244","DOI":"10.1109\/TETC.2013.2274797","article-title":"Gender-Driven Emotion Recognition Through Speech Signals For Ambient Intelligence Applications","volume":"1","author":"Bisio","year":"2013","journal-title":"IEEE Trans. Emerg. Top. Comput."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Ramdinmawii, E., and Mittal, V. (2016, January 26\u201328). Gender identification from speech signal by examining the speech production characteristics. Proceedings of the 2016 International Conference on Signal Processing and Communication (ICSC), Noida, India.","DOI":"10.1109\/ICSPCom.2016.7980584"},{"key":"ref_17","first-page":"7213717","article-title":"DGR: Gender Recognition of Human Speech Using One-Dimensional Conventional Neural Network","volume":"2019","author":"Alkhawaldeh","year":"2019","journal-title":"Sci. Program."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Kabil, S., Muckenhirn, H., and Magimai-Doss, M. (2018). On Learning to Identify Genders from Raw Speech Signal Using CNNs, Interspeech.","DOI":"10.21437\/Interspeech.2018-1240"},{"key":"ref_19","unstructured":"Costantini, G., Iadarola, I., Paoloni, A., and Todisco, M. (2014, January 26\u201331). EMOVO Corpus: An Italian Emotional Speech Database. Proceedings of the International Conference on Language Resources and Evaluation (LREC 2014), Reykjavik, Iceland."},{"key":"ref_20","first-page":"51","article-title":"The COST 2102 Italian Audio and Video Emotional Database","volume":"Volume 204","author":"Esposito","year":"2009","journal-title":"Neural Nets WIRN09, Proceedings of the 19th Italian Workshop on Neural Nets, Salerno, Italy, 28\u201330 May 2009"},{"key":"ref_21","first-page":"406","article-title":"The New Italian Audio and Video Emotional Database","volume":"Volume 5967","author":"Esposito","year":"2009","journal-title":"Development of Multimodal Interfaces: Active Listening and Synchrony"},{"key":"ref_22","first-page":"255","article-title":"Emotional Vocal Expressions Recognition Using the COST 2102 Italian Database of Emotional Speech","volume":"Volume 5967","author":"Atassi","year":"2010","journal-title":"Development of Multimodal Interfaces: Active Listening and Synchrony 2010"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1016\/j.knosys.2014.03.019","article-title":"Speech emotion recognition using amplitude modulation parameters and a combined feature selection procedure","volume":"63","author":"Mencattini","year":"2014","journal-title":"Knowl.-Based Syst."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Paliwal, K.K., Lyons, J.G., and W\u00f3jcicki, K.K. (2010, January 13\u201315). Preference for 20\u201340 ms window duration in speech analysis. Proceedings of the 2010 4th International Conference on Signal Processing and Communication Systems, Gold Coast, QLD, Australia.","DOI":"10.1109\/ICSPCS.2010.5709770"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TSA.2005.855834","article-title":"Robust speech recognition in noisy environments based on subband spectral centroid histograms","volume":"14","author":"Gajic","year":"2006","journal-title":"Audio Speech Lang. Process. IEEE Trans."},{"key":"ref_26","first-page":"58","article-title":"Speaker Verification with Adaptive Spectral Subband Centroids","volume":"Volume 4642","author":"Kinnunen","year":"2007","journal-title":"International Conference on Biometrics"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"12","DOI":"10.18178\/ijsps.6.1.12-16","article-title":"Spectral Subband Centroids for Robust Speaker Identification Using Marginalization-based Missing Feature Theory","volume":"6","author":"Nicolson","year":"2019","journal-title":"Int. J. Signal Process. Syst."},{"key":"ref_28","first-page":"1","article-title":"Spectral Subband Centroids as Complementary Features for Speaker Authentication","volume":"Volume 3072","author":"Poh","year":"2004","journal-title":"International Conference on Biometric Authentication"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1091","DOI":"10.1016\/j.sigpro.2007.11.017","article-title":"Speaker segmentation and clustering","volume":"88","author":"Kotti","year":"2008","journal-title":"Signal Process."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/5\/1714\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:25:02Z","timestamp":1760135102000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/5\/1714"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,22]]},"references-count":29,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2022,3]]}},"alternative-id":["s22051714"],"URL":"https:\/\/doi.org\/10.3390\/s22051714","relation":{"has-preprint":[{"id-type":"doi","id":"10.20944\/preprints202112.0134.v1","asserted-by":"object"}]},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,22]]}}}