{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T16:10:21Z","timestamp":1782835821229,"version":"3.54.5"},"reference-count":52,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2023,3,10]],"date-time":"2023-03-10T00:00:00Z","timestamp":1678406400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"University of Auckland Postgraduate Research Student Support fund"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>A low-resource emotional speech synthesis system for empathetic speech synthesis based on modelling prosody features is presented here. Secondary emotions, identified to be needed for empathetic speech, are modelled and synthesised in this investigation. As secondary emotions are subtle in nature, they are difficult to model compared to primary emotions. This study is one of the few to model secondary emotions in speech as they have not been extensively studied so far. Current speech synthesis research uses large databases and deep learning techniques to develop emotion models. There are many secondary emotions, and hence, developing large databases for each of the secondary emotions is expensive. Hence, this research presents a proof of concept using handcrafted feature extraction and modelling of these features using a low-resource-intensive machine learning approach, thus creating synthetic speech with secondary emotions. Here, a quantitative-model-based transformation is used to shape the emotional speech\u2019s fundamental frequency contour. Speech rate and mean intensity are modelled via rule-based approaches. Using these models, an emotional text-to-speech synthesis system to synthesise five secondary emotions-anxious, apologetic, confident, enthusiastic and worried-is developed. A perception test to evaluate the synthesised emotional speech is also conducted. The participants could identify the correct emotion in a forced response test with a hit rate greater than 65%.<\/jats:p>","DOI":"10.3390\/s23062999","type":"journal-article","created":{"date-parts":[[2023,3,10]],"date-time":"2023-03-10T02:05:54Z","timestamp":1678413954000},"page":"2999","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Exploring Prosodic Features Modelling for Secondary Emotions Needed for Empathetic Speech Synthesis"],"prefix":"10.3390","volume":"23","author":[{"given":"Jesin","family":"James","sequence":"first","affiliation":[{"name":"Department of Electrical, Computer, and Software Engineering, The University of Auckland, Auckland 1010, New Zealand"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8523-6607","authenticated-orcid":false,"given":"Balamurali","family":"B.T.","sequence":"additional","affiliation":[{"name":"Science, Maths and Technology, Singapore University of Technology and Design, Singapore 487372, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Catherine","family":"Watson","sequence":"additional","affiliation":[{"name":"Department of Electrical, Computer, and Software Engineering, The University of Auckland, Auckland 1010, New Zealand"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hansj\u00f6rg","family":"Mixdorff","sequence":"additional","affiliation":[{"name":"Computer Science and Media, Berliner Hochschule f\u00fcr Technik, 13353 Berlin, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Eyssel, F., Ruiter, L.D., Kuchenbrandt, D., Bobinger, S., and Hegel, F. (2012, January 5\u20138). \u2018If you sound like me, you must be more human\u2019: On the interplay of robot and user features on human-robot acceptance and anthropomorphism. Proceedings of the International Conference on Human-Robot Interaction, Boston, MA, USA.","DOI":"10.1145\/2157689.2157717"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"James, J., Watson, C.I., and MacDonald, B. (2018, January 27\u201331). Artificial empathy in social robots: An analysis of emotions in speech. Proceedings of the IEEE International Symposium on Robot & Human Interactive Communication, Nanjing, China.","DOI":"10.1109\/ROMAN.2018.8525652"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"2119","DOI":"10.1007\/s12369-020-00691-4","article-title":"Empathetic Speech Synthesis and Testing for Healthcare Robots","volume":"13","author":"James","year":"2020","journal-title":"Int. J. Soc. Robot."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"169","DOI":"10.1080\/02699939208411068","article-title":"An argument for basic emotions","volume":"6","author":"Ekman","year":"1992","journal-title":"Cogn. Emot."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Schr\u00f6der, M. (2001, January 3\u20137). Emotional Speech Synthesis: A Review. Proceedings of the Eurospeech, Alborg, Denmark.","DOI":"10.21437\/Eurospeech.2001-150"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1007\/978-3-540-85483-8_2","article-title":"Affect Simulation with Primary and Secondary Emotions","volume":"Volume 5208","author":"Wachsmuth","year":"2008","journal-title":"Proceedings of the Intelligent Virtual Agents"},{"key":"ref_7","unstructured":"Damasio, A. (1994). Descartes\u2019 Error, Emotion Reason and the Human Brain, Avon Books."},{"key":"ref_8","unstructured":"James, J., Watson, C., and Stoakes, H. (2019, January 5\u20139). Influence of Prosodic features and semantics on secondary emotion production and perception. Proceedings of the International Congress of Phonetic Sciences, Melbourne, Australia."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"369","DOI":"10.1016\/0167-6393(95)00005-9","article-title":"Implementation and testing of a system for producing emotion-by-rule in synthetic speech","volume":"16","author":"Murray","year":"1995","journal-title":"Speech Commun."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1145","DOI":"10.1109\/TASL.2006.876113","article-title":"Prosody conversion from neutral speech to emotional speech","volume":"14","author":"Tao","year":"2006","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_11","unstructured":"Skerry-Ryan, R., Battenberg, E., Xiao, Y., Wang, Y., Stanton, D., Shor, J., Weiss, R.J., Clark, R., and Saurous, R.A. (2018, January 10\u201315). Towards end-to-end prosody transfer for expressive speech synthesis with tacotron. Proceedings of the International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"An, S., Ling, Z., and Dai, L. (2017, January 12\u201315). Emotional statistical parametric speech synthesis using LSTM-RNNs. Proceedings of the APSIPA Conference, Kuala Lumpur, Malaysia.","DOI":"10.1109\/APSIPA.2017.8282282"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Vroomen, J., Collier, R., and Mozziconacci, S. (1993, January 19\u201323). Duration and intonation in emotional speech. Proceedings of the Third European Conference on Speech Communication and Technology, Berlin, Germany.","DOI":"10.21437\/Eurospeech.1993-136"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Masuko, T., Kobayashi, T., and Miyanaga, K. (2004, January 8). A style control technique for HMM-based speech synthesis. Proceedings of the International Conference on Spoken Language Processing, Jeju, Republic of Korea.","DOI":"10.21437\/Interspeech.2004-551"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1099","DOI":"10.1109\/TASL.2006.876123","article-title":"The IBM expressive text-to-speech synthesis system for American English","volume":"14","author":"Pitrelli","year":"2006","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Yamagishi, J., Kobayashi, T., Tachibana, M., Ogata, K., and Nakano, Y. (2007, January 16\u201320). Model adaptation approach to speech synthesis with diverse voices and styles. Proceedings of the International Conference on Acoustics, Speech and Signal Processing, Honolulu, HI, USA.","DOI":"10.1109\/ICASSP.2007.367299"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"394","DOI":"10.1016\/j.specom.2009.12.007","article-title":"Analysis of statistical parametric and unit selection speech synthesis systems applied to emotional speech","volume":"52","author":"Yamagishi","year":"2010","journal-title":"Speech Commun."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Tits, N. (2019, January 3\u20136). A Methodology for Controlling the Emotional Expressiveness in Synthetic Speech-a Deep Learning approach. Proceedings of the International Conference on Affective Computing and Intelligent Interaction, Cambridge, UK.","DOI":"10.1109\/ACIIW.2019.8925241"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhang, M., Tao, J., Jia, H., and Wang, X. (2008, January 11\u201314). Improving HMM based speech synthesis by reducing over-smoothing problems. Proceedings of the International Symposium on Chinese Spoken Language Processing, Singapore.","DOI":"10.1109\/CHINSL.2008.ECP.16"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1097","DOI":"10.1121\/1.405558","article-title":"Toward the simulation of emotion in synthetic speech: A review of the literature on human vocal emotion","volume":"93","author":"Murray","year":"1993","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"James, J., Tian, L., and Watson, C. (2018, January 2\u20136). An open source emotional speech corpus for human robot interaction applications. Proceedings of the Interspeech, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-1349"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"695","DOI":"10.1177\/0539018405058216","article-title":"What are emotions? And how can they be measured?","volume":"44","author":"Scherer","year":"2005","journal-title":"Soc. Sci. Inf."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"116","DOI":"10.1109\/T-AFFC.2012.36","article-title":"Seeing Stars of Valence and Arousal in Blog Posts","volume":"4","author":"Paltoglou","year":"2013","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1853","DOI":"10.1109\/TASLP.2022.3178225","article-title":"Neonatal Bowel Sound Detection Using Convolutional Neural Network and Laplace Hidden Semi-Markov Model","volume":"30","author":"Sitaula","year":"2022","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Burne, L., Sitaula, C., Priyadarshi, A., Tracy, M., Kavehei, O., Hinder, M., Withana, A., McEwan, A., and Marzbanrad, F. (2022). Ensemble Approach on Deep and Handcrafted Features for Neonatal Bowel Sound Detection. IEEE J. Biomed. Health Inform., 1\u201311.","DOI":"10.1109\/JBHI.2022.3217559"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"392","DOI":"10.1016\/j.csl.2017.01.002","article-title":"EMU-SDMS: Advanced speech database management and analysis in R","volume":"45","author":"Winkelmann","year":"2017","journal-title":"Comput. Speech Lang."},{"key":"ref_27","unstructured":"R Core Team (2017). R: A Language and Environment for Statistical Computing, R Foundation for Statistical Computing."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Wickham, H. (2016). ggplot2: Elegant Graphics for Data Analysis, Springer.","DOI":"10.1007\/978-3-319-24277-4"},{"key":"ref_29","unstructured":"James, J., Mixdorff, H., and Watson, C. (2019, January 5\u20139). Quantitative model-based analysis of F0 contours of emotional speech. Proceedings of the International Congress of Phonetic Sciences, Melbourne, Australia."},{"key":"ref_30","unstructured":"Hui, C.T.J., Chin, T.J., and Watson, C. (2014, January 2\u20135). Automatic detection of speech truncation and speech rate. Proceedings of the SST, Christchurch, New Zealand."},{"key":"ref_31","unstructured":"Hirose, K., Fujisaki, H., and Yamaguchi, M. (1984, January 15\u201320). Synthesis by rule of voice fundamental frequency contours of spoken Japanese from linguistic information. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, Calgary, AB, Canada."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Nguyen, D.T., Luong, M.C., Vu, B.K., Mixdorff, H., and Ngo, H.H. (2004, January 8). Fujisaki Model based f0 contours in Vietnamese TTS. Proceedings of the International Conference on Spoken Language Processing, Jeju, Republic of Korea.","DOI":"10.21437\/Interspeech.2004-549"},{"key":"ref_33","unstructured":"Gu, W., and Lee, T. (September, January 31). Quantitative analysis of f0 contours of emotional speech of Mandarin. Proceedings of the 8th ISCA Spee Synthesis Workshop, Barselona, Spain."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Amir, N., Mixdorff, H., Amir, O., Rochman, D., Diamond, G.M., Pfitzinger, H.R., Levi-Isserlish, T., and Abramson, S. (2010, January 10\u201314). Unresolved anger: Prosodic analysis and classification of speech from a therapeutic setting. Proceedings of the Speech Prosody, Chicago, IL, USA.","DOI":"10.21437\/SpeechProsody.2010-88"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Mixdorff, H., Cossio-Mercado, C., H\u00f6nemann, A., Gurlekian, J., Evin, D., and Torres, H. (2015, January 6\u201310). Acoustic correlates of perceived syllable prominence in German. Proceedings of the Annual Conference of the International Speech Communication Association, Dresden, Germany.","DOI":"10.21437\/Interspeech.2015-11"},{"key":"ref_36","unstructured":"Boersma, P., and Weenink, D. (2022, February 01). Praat: Doing Phonetics by Computer [Computer Program]. Version 6.0.46. Available online: https:\/\/www.fon.hum.uva.nl\/praat\/."},{"key":"ref_37","unstructured":"Mixdorff, H. (2000, January 5\u20139). A novel approach to the fully automatic extraction of Fujisaki model parameters. Proceedings of the IEEE Int. Conf. on Acoustics, Speech, and Signal Processing, Istanbul, Turkey."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Mixdorff, H., and Fujisaki, H. (2000, January 16\u201320). A quantitative description of German prosody offering symbolic labels as a by-product. Proceedings of the International Conference on Spoken Language Processing, Beijing, China.","DOI":"10.21437\/ICSLP.2000-218"},{"key":"ref_39","unstructured":"Watson, C.I., and Marchi, A. (2014, January 6\u20139). Resources created for building New Zealand English voices. Proceedings of the Australasian International Conference of Speech Science and Technology, Parramatta, New Zealand."},{"key":"ref_40","unstructured":"Jain, S. (2015). Towards the Creation of Customised Synthetic Voices using Hidden Markov Models on a Healthcare Robot. [Master\u2019s Thesis, The University of Auckland]."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"365","DOI":"10.1023\/A:1025708916924","article-title":"The German text-to-speech synthesis system MARY: A tool for research, development and teaching","volume":"6","author":"Trouvain","year":"2003","journal-title":"Int. J. Speech Technol."},{"key":"ref_42","first-page":"97","article-title":"Accurate short-term analysis of the fundamental frequency and the harmonics-to-noise ratio of a sampled sound","volume":"17","author":"Boersma","year":"1993","journal-title":"Inst. Phon. Sci."},{"key":"ref_43","unstructured":"Kisler, T., Schiel, F., and Sloetjes, H. (2012, January 16\u201322). Signal processing via web services: The use case WebMAUS. Proceedings of the Digital Humanities Conference, Sheffield."},{"key":"ref_44","first-page":"18","article-title":"Classification and Regression by Random Forest","volume":"23","author":"Liaw","year":"2002","journal-title":"R News 2.3"},{"key":"ref_45","unstructured":"Yoav, F., and Robert E, S. (1996, January 3\u20136). Experiments with a new boosting algorithm. Proceedings of the International Conference on Machine Learning, Bari, Italy."},{"key":"ref_46","first-page":"2825","article-title":"Scikit-learn: Machine Learning in Python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2379776.2379786","article-title":"Ensemble approaches for regression: A survey","volume":"45","author":"Soares","year":"2012","journal-title":"ACM Comput. Surv."},{"key":"ref_48","unstructured":"Eide, E., Aaron, A., Bakis, R., Hamza, W., Picheny, M., and Pitrelli, J. (2004, January 14\u201316). A corpus-based approach to expressive speech synthesis. Proceedings of the ISCA ITRW on Speech Synthesis, Pittsburgh, PA, USA."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Ming, H., Huang, D.Y., Dong, M., Li, H., Xie, L., and Zhang, S. (2015, January 21\u201324). Fundamental Frequency Modeling Using Wavelets for Emotional Voice Conversion. Proceedings of the International Conference on Affective Computing and Intelligent Interaction, Xi\u2019an, China.","DOI":"10.1109\/ACII.2015.7344665"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Lu, X., and Pan, T. (2016, January 11\u201312). Research On Prosody Conversion of Affective Speech Based on LIBSVM and PAD Three Dimensional Emotion Model. Proceedings of the Wkhp on Advanced Research & Tech in Industry Applications, Tianjin, China.","DOI":"10.2991\/wartia-16.2016.1"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Robinson, C., Obin, N., and Roebel, A. (2019, January 12\u201317). Sequence-To-Sequence Modelling of F0 for Speech Emotion Conversion. Proceedings of the International Conference on Acoustics, Speech, and Signal Processing, Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683865"},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1016\/j.csl.2003.12.001","article-title":"Measuring speech quality for text-to-speech systems: Development and assessment of a modified mean opinion score (MOS) scale","volume":"19","author":"Viswanathan","year":"2005","journal-title":"Comput. Speech Lang."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/6\/2999\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:51:46Z","timestamp":1760122306000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/6\/2999"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,10]]},"references-count":52,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["s23062999"],"URL":"https:\/\/doi.org\/10.3390\/s23062999","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,10]]}}}