{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T04:52:39Z","timestamp":1780635159963,"version":"3.54.1"},"reference-count":50,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2024,1,12]],"date-time":"2024-01-12T00:00:00Z","timestamp":1705017600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,1,12]],"date-time":"2024-01-12T00:00:00Z","timestamp":1705017600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Federal Ministry for Economic Affairs and Climate Action and the German Aerospace Center","award":["50RP2260A"],"award-info":[{"award-number":["50RP2260A"]}]},{"DOI":"10.13039\/501100008349","name":"Universit\u00e4t Duisburg-Essen","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100008349","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["K\u00fcnstl Intell"],"published-print":{"date-parts":[[2024,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper briefly introduces the Project \u201cAudEeKA\u201d, whose aim is to use speech and other bio signals for emotion recognition to improve remote, but also direct, healthcare. This article takes a look at use cases, goals and challenges, of researching and implementing a possible solution. To gain additional insights, the main-goal of the project is divided into multiple sub-goals, namely speech emotion recognition, stress detection and classification and emotion detection from physiological signals. Also, similar projects are considered and project-specific requirements stemming from use-cases introduced. Possible pitfalls and difficulties are outlined, which are mostly associated with datasets. They also emerge out of the requirements, their accompanying restrictions and first analyses in the area of speech emotion recognition, which are shortly presented and discussed. At the same time, first approaches to solutions for every sub-goal, which include the use of continual learning, and finally a draft of the planned architecture for the envisioned system, is presented. This draft presents a possible solution for combining all sub-goals, while reaching the main goal of a multimodal emotion recognition system.<\/jats:p>","DOI":"10.1007\/s13218-023-00828-3","type":"journal-article","created":{"date-parts":[[2024,1,12]],"date-time":"2024-01-12T09:02:33Z","timestamp":1705050153000},"page":"151-156","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Auditive Emotion Recognition for Empathic AI-Assistants"],"prefix":"10.1007","volume":"38","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-7410-8504","authenticated-orcid":false,"given":"Roswitha","family":"Duwenbeck","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Elsa Andrea","family":"Kirchner","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,1,12]]},"reference":[{"key":"828_CR1","unstructured":"Winnat C (2017) Deutsche aerzte nehmen sich rund sieben minuten zeit pro patient"},{"key":"828_CR2","unstructured":"Stewart MA (1995) Effective physician-patient communication and health outcomes: a review. CMAJ 152(9):1423"},{"key":"828_CR3","doi-asserted-by":"crossref","unstructured":"Nitschke JP, Bartz JA (2022) The association between acute stress & empathy: a systematic literature review. Neurosci Biobehav Rev 144:105003","DOI":"10.1016\/j.neubiorev.2022.105003"},{"key":"828_CR4","doi-asserted-by":"publisher","first-page":"S34","DOI":"10.1046\/j.1525-1497.1999.00263.x","volume":"14","author":"DC Dugdale","year":"1999","unstructured":"Dugdale DC, Epstein R, Pantilat SZ (1999) Time and the patient\u2013physician relationship. J Gen Intern Med 14:S34","journal-title":"J Gen Intern Med"},{"key":"828_CR5","unstructured":"Budde K, Dasch T, Kirchner E, Ohliger U, Schapranow M, Schmidt T, Schwerk A, Thoms J, Zahn T, Hiltawsky K (2020) K\u00fcnstliche intelligenz: Patienten im fokus. Dtsch Arztebl 117(49):A\u20132407"},{"key":"828_CR6","volume-title":"Lernende systeme im gesundheitswesen: Grundlagen, anwendungsszenarien und gestaltungsoptionen","author":"LS-DPL Systeme","year":"2019","unstructured":"Systeme LS-DPL (2019) Lernende systeme im gesundheitswesen: Grundlagen, anwendungsszenarien und gestaltungsoptionen. Bericht der AG Gesundheit, Medizintechnik, Pflege"},{"key":"828_CR7","doi-asserted-by":"crossref","unstructured":"Kim J, Andr\u00e9 E (2006) Emotion recognition using physiological and speech signal in short-term observation. In: Perception and interactive technologies: international tutorial and research workshop, PIT 2006 Kloster Irsee, Germany, June 19\u201321, 2006. Proceedings. Springer, pp\u00a053\u201364","DOI":"10.1007\/11768029"},{"key":"828_CR8","doi-asserted-by":"crossref","unstructured":"Chao L, Tao J, Yang M, Li Y, Wen Z (2015) Long short term memory recurrent neural network based multimodal dimensional emotion recognition. In: Proceedings of the 5th international workshop on audio\/visual emotion challenge, pp\u00a065\u201372","DOI":"10.1145\/2808196.2811634"},{"key":"828_CR9","doi-asserted-by":"crossref","unstructured":"Ranganathan H, Chakraborty S, Panchanathan S (2016) Multimodal emotion recognition using deep learning architectures. In: 2016 IEEE winter conference on applications of computer vision (WACV), pp\u00a01\u20139","DOI":"10.1109\/WACV.2016.7477679"},{"key":"828_CR10","doi-asserted-by":"crossref","unstructured":"Guo H, Jiang N, Shao D (2020) Research on multi-modal emotion recognition based on speech, eeg and ecg signals. In: Robotics and rehabilitation intelligence: first international conference, ICRRI 2020, Fushun, China, September 9\u201311, 2020, Proceedings, Part I 1. Springer, pp\u00a0272\u2013288","DOI":"10.1007\/978-981-33-4929-2_19"},{"key":"828_CR11","doi-asserted-by":"crossref","unstructured":"Bakhshi A, Chalup S (2021) Multimodal emotion recognition based on speech and physiological signals using deep neural networks. In: Pattern recognition. ICPR international workshops and challenges: virtual event, January 10\u201315, 2021, Proceedings, Part VI. Springer, pp\u00a0289\u2013300","DOI":"10.1007\/978-3-030-68780-9_25"},{"key":"828_CR12","doi-asserted-by":"publisher","DOI":"10.1016\/j.compbiomed.2022.105907","volume":"149","author":"Q Wang","year":"2022","unstructured":"Wang Q, Wang M, Yang Y, Zhang X (2022) Multi-modal emotion recognition using EEG and speech signals. Comput Biol Med 149:105907","journal-title":"Comput Biol Med"},{"issue":"1","key":"828_CR13","doi-asserted-by":"publisher","first-page":"32","DOI":"10.1109\/79.911197","volume":"18","author":"R Cowie","year":"2001","unstructured":"Cowie R, Douglas-Cowie E, Tsapatsoulis N, Votsis G, Kollias S, Fellenz W, Taylor JG (2001) Emotion recognition in human\u2013computer interaction. IEEE Signal Process Mag 18(1):32\u201380","journal-title":"IEEE Signal Process Mag"},{"key":"828_CR14","doi-asserted-by":"crossref","unstructured":"Austermann A, Esau N, Kleinjohann L, Kleinjohann B (2005) Prosody based emotion recognition for mexi. In 2005 IEEE\/RSJ international conference on intelligent robots and systems. IEEE, pp\u00a01138\u20131144","DOI":"10.1109\/IROS.2005.1545341"},{"key":"828_CR15","unstructured":"Altun H (2005) Integrating learner\u2019s affective state in intelligent tutoring systems to enhance e-learning applications. GETS 2005 3(1)"},{"key":"828_CR16","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1155\/S1110865704406192","volume":"2004","author":"CL Lisetti","year":"2004","unstructured":"Lisetti CL, Nasoz F (2004) Using noninvasive wearable computers to recognize human emotions from physiological signals. EURASIP J Adv Signal Process 2004:1\u201316","journal-title":"EURASIP J Adv Signal Process"},{"key":"828_CR17","doi-asserted-by":"crossref","unstructured":"Devillers L, Lamel L, Vasilescu I (2003) motion detection in task-oriented spoken dialogues. In: 2003 International conference on multimedia and expo. ICME\u201903. Proceedings (Cat. No. 03TH8698), vol\u00a03, pp\u00a0III\u2013549. IEEE","DOI":"10.1109\/ICME.2003.1221370"},{"key":"828_CR18","doi-asserted-by":"crossref","unstructured":"Tacconi D, Mayora O, Lukowicz P, Arnrich B, Setz C, Troster G, Haring C (2008) Activity and emotion recognition to support early diagnosis of psychiatric diseases. In: 2008 second international conference on pervasive computing technologies for healthcare, pp\u00a0100\u2013102. IEEE","DOI":"10.1109\/PCTHEALTH.2008.4571041"},{"issue":"1","key":"828_CR19","first-page":"53","volume":"2","author":"A Saxena","year":"2020","unstructured":"Saxena A, Khanna A, Gupta D (2020) Emotion recognition and detection methods: a comprehensive survey. J Artif Intell Syst 2(1):53\u201379","journal-title":"J Artif Intell Syst"},{"key":"828_CR20","doi-asserted-by":"crossref","unstructured":"Makiuchi MR, Uto K, Shinoda K (2021) Multimodal emotion recognition with high-level speech and text features. In: 2021 IEEE automatic speech recognition and understanding workshop (ASRU), pp\u00a0350\u2013357","DOI":"10.1109\/ASRU51503.2021.9688036"},{"key":"828_CR21","doi-asserted-by":"crossref","unstructured":"Pepino L, Riera P, Ferrer L, Gravano A (2020) Fusion approaches for emotion recognition from speech using acoustic and text-based features. In: ICASSP 2020\u20142020 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp\u00a06484\u20136488","DOI":"10.1109\/ICASSP40776.2020.9054709"},{"key":"828_CR22","doi-asserted-by":"publisher","first-page":"61672","DOI":"10.1109\/ACCESS.2020.2984368","volume":"8","author":"N-H Ho","year":"2020","unstructured":"Ho N-H, Yang H-J, Kim S-H, Lee G (2020) Multimodal approach of speech emotion recognition using multi-level multi-head fusion attention-based recurrent neural network. IEEE Access 8:61672\u201361686","journal-title":"IEEE Access"},{"key":"828_CR23","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.patrec.2021.03.007","volume":"146","author":"L Schoneveld","year":"2021","unstructured":"Schoneveld L, Othmani A, Abdelkawy H (2021) Leveraging recent advances in deep learning for audio\u2013visual emotion recognition. Pattern Recogn Lett 146:1\u20137","journal-title":"Pattern Recogn Lett"},{"key":"828_CR24","doi-asserted-by":"publisher","first-page":"42","DOI":"10.1016\/j.eswa.2016.08.047","volume":"66","author":"L-A Perez-Gaspar","year":"2016","unstructured":"Perez-Gaspar L-A, Caballero-Morales S-O, Trujillo-Romero F (2016) Multimodal emotion recognition with evolutionary computation for human-robot interaction. Expert Syst Appl 66:42\u201361","journal-title":"Expert Syst Appl"},{"key":"828_CR25","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2022.108580","volume":"244","author":"AI Middya","year":"2022","unstructured":"Middya AI, Nag B, Roy S (2022) Deep learning based multimodal emotion recognition using model-level fusion of audio\u2013visual modalities. Knowl-Based Syst 244:108580","journal-title":"Knowl-Based Syst"},{"key":"828_CR26","doi-asserted-by":"publisher","DOI":"10.1016\/j.jnca.2019.102423","volume":"147","author":"M Imani","year":"2019","unstructured":"Imani M, Montazer GA (2019) A survey of emotion recognition methods with emphasis on e-learning environments. J Netw Comput Appl 147:102423","journal-title":"J Netw Comput Appl"},{"key":"828_CR27","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1007\/s10772-011-9125-1","volume":"15","author":"SG Koolagudi","year":"2012","unstructured":"Koolagudi SG, Rao KS (2012) Emotion recognition from speech: a review. Int J Speech Technol 15:99\u2013117","journal-title":"Int J Speech Technol"},{"key":"828_CR28","doi-asserted-by":"publisher","first-page":"47795","DOI":"10.1109\/ACCESS.2021.3068045","volume":"9","author":"TM Wani","year":"2021","unstructured":"Wani TM, Gunawan TS, Qadri SAA, Kartiwi M, Ambikairajah E (2021) A comprehensive review of speech emotion recognition systems. IEEE Access 9:47795\u201347814","journal-title":"IEEE Access"},{"key":"828_CR29","unstructured":"Muenchen TU, \u201cEight emotional speech databases used - tum.\u201d"},{"issue":"7","key":"828_CR30","doi-asserted-by":"publisher","first-page":"2074","DOI":"10.3390\/s18072074","volume":"18","author":"L Shu","year":"2018","unstructured":"Shu L, Xie J, Yang M, Li Z, Li Z, Liao D, Xu X, Yang X (2018) A review of emotion recognition using physiological signals. Sensors 18(7):2074","journal-title":"Sensors"},{"key":"828_CR31","doi-asserted-by":"publisher","first-page":"1111","DOI":"10.3389\/fpsyg.2020.01111","volume":"11","author":"F Larradet","year":"2020","unstructured":"Larradet F, Niewiadomski R, Barresi G, Caldwell DG, Mattos LS (2020) Toward emotion recognition from physiological signals in the wild: approaching the methodological issues in real-life data collection. Front Psychol 11:1111","journal-title":"Front Psychol"},{"issue":"39\u201358","key":"828_CR32","first-page":"3","volume":"1","author":"PJ Lang","year":"1997","unstructured":"Lang PJ, Bradley MM, Cuthbert BN et al (1997) International affective picture system (IAPS): technical manual and affective ratings. NIMH Center Study Emotion Attent 1(39\u201358):3","journal-title":"NIMH Center Study Emotion Attent"},{"key":"828_CR33","unstructured":"Merkx P, Truong KP, Neerincx MA (2007) Inducing and measuring emotion through a multiplayer first-person shooter computer game. In: Proceedings of the computer games workshop"},{"key":"828_CR34","doi-asserted-by":"crossref","unstructured":"Zhang W, Shu L, Xu X, Liao D (2017) Affective virtual reality system (AVRS): design and ratings of affective VR scenes. In: 2017 international conference on virtual reality and visualization (ICVRV). IEEE, pp\u00a0311\u2013314","DOI":"10.1109\/ICVRV.2017.00072"},{"key":"828_CR35","doi-asserted-by":"crossref","unstructured":"Kim J, Andr\u00e9 E (2009) Fusion of multichannel biosignals towards automatic emotion recognition. Multisensor Fusion Integr Intell Syst 35(Part 1):55\u201368","DOI":"10.1007\/978-3-540-89859-7_5"},{"issue":"2","key":"828_CR36","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1007\/BF00995188","volume":"17","author":"D Matsumoto","year":"1993","unstructured":"Matsumoto D (1993) Ethnic differences in affect intensity, emotion judgments, display rule attitudes, and self-reported emotional expression in an American sample. Motiv Emotion 17(2):107\u2013123","journal-title":"Motiv Emotion"},{"key":"828_CR37","unstructured":"Brody LR (1993) On understanding gender differences in the expression of emotion. Hum Feel Explor Affect Dev Mean, pp\u00a087\u2013121"},{"issue":"1","key":"828_CR38","doi-asserted-by":"publisher","first-page":"28","DOI":"10.1037\/0882-7974.6.1.28","volume":"6","author":"RW Levenson","year":"1991","unstructured":"Levenson RW, Carstensen LL, Friesen WV, Ekman P (1991) Emotion, physiology, and expression in old age. Psychol Aging 6(1):28","journal-title":"Psychol Aging"},{"key":"828_CR39","first-page":"1517","volume":"5","author":"F Burkhardt","year":"2005","unstructured":"Burkhardt F, Paeschke A, Rolfes M, Sendlmeier WF, Weiss B et al (2005) A database of German emotional speech. Interspeech 5:1517\u20131520","journal-title":"Interspeech"},{"key":"828_CR40","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, Vanderplas J, Passos A, Cournapeau D, Brucher M, Perrot M, Duchesnay E (2011) Scikit-learn: machine learning in Python. J Mach Learn Res 12:2825\u20132830","journal-title":"J Mach Learn Res"},{"key":"828_CR41","doi-asserted-by":"crossref","unstructured":"Eyben F, W\u00f6llmer M, Schuller B (2010) Opensmile: the munich versatile and fast open-source audio feature extractor. In: Proceedings of the 18th ACM international conference on multimedia, pp\u00a01459\u20131462","DOI":"10.1145\/1873951.1874246"},{"issue":"4","key":"828_CR42","doi-asserted-by":"publisher","first-page":"397","DOI":"10.1177\/1754073911410747","volume":"3","author":"JL Tracy","year":"2011","unstructured":"Tracy JL, Randles D (2011) Four models of basic emotions: a review of Ekman and Cordaro, Izard, Levenson, and Panksepp and Watt. Emot Rev 3(4):397\u2013405","journal-title":"Emot Rev"},{"issue":"6","key":"828_CR43","doi-asserted-by":"publisher","first-page":"1161","DOI":"10.1037\/h0077714","volume":"39","author":"JA Russell","year":"1980","unstructured":"Russell JA (1980) A circumplex model of affect. J Pers Soc Psychol 39(6):1161","journal-title":"J Pers Soc Psychol"},{"key":"828_CR44","doi-asserted-by":"crossref","unstructured":"Mariotti A (2015) The effects of chronic stress on health: new insights into the molecular mechanisms of brain\u2013body communication. Future Sci OA 1(3):FSO23","DOI":"10.4155\/fso.15.21"},{"key":"828_CR45","doi-asserted-by":"publisher","first-page":"135","DOI":"10.1007\/s11276-015-0960-x","volume":"22","author":"T Gao","year":"2016","unstructured":"Gao T, Song J-Y, Zou J-Y, Ding J-H, Wang D-Q, Jin R-C (2016) An overview of performance trade-off mechanisms in routing protocol for green wireless sensor networks. Wireless Netw 22:135\u2013157","journal-title":"Wireless Netw"},{"key":"828_CR46","doi-asserted-by":"crossref","unstructured":"Gunes H, Piccardi M (2005) Affect recognition from face and body: early fusion vs. late fusion. In: 2005 IEEE international conference on systems, man and cybernetics, vol 4, pp 3437\u20133443","DOI":"10.1109\/ICSMC.2005.1571679"},{"key":"828_CR47","doi-asserted-by":"crossref","unstructured":"Hazarika D, Gorantla S, Poria S, Zimmermann R (2018) Self-attentive feature-level fusion for multimodal emotion detection. In: 2018 IEEE conference on multimedia information processing and retrieval (MIPR), pp\u00a0196\u2013201","DOI":"10.1109\/MIPR.2018.00043"},{"key":"828_CR48","unstructured":"Zheng W-L, Dong B-N, Lu B-L (2014) Multimodal emotion recognition using EEG and eye tracking data. In: 2014 36th annual international conference of the IEEE engineering in medicine and biology society, pp\u00a05040\u20135043"},{"key":"828_CR49","doi-asserted-by":"crossref","unstructured":"Sahoo S, Routray A (2016) Emotion recognition from audio-visual data using rule based decision level fusion. In: 2016 IEEE students? Technology symposium (TechSym), pp\u00a07\u201312","DOI":"10.1109\/TechSym.2016.7872646"},{"key":"828_CR50","doi-asserted-by":"crossref","unstructured":"Song K-S, Nho Y-H, Seo J-H, Kwon D-S (2018) Decision-level fusion method for emotion recognition using multimodal emotion recognition information. In: 2018 15th international conference on ubiquitous robots (UR), pp\u00a0472\u2013476","DOI":"10.1109\/URAI.2018.8441795"}],"container-title":["KI - K\u00fcnstliche Intelligenz"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13218-023-00828-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s13218-023-00828-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13218-023-00828-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,12,5]],"date-time":"2024-12-05T22:20:17Z","timestamp":1733437217000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s13218-023-00828-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,12]]},"references-count":50,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,11]]}},"alternative-id":["828"],"URL":"https:\/\/doi.org\/10.1007\/s13218-023-00828-3","relation":{},"ISSN":["0933-1875","1610-1987"],"issn-type":[{"value":"0933-1875","type":"print"},{"value":"1610-1987","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,12]]},"assertion":[{"value":"11 May 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 December 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 January 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}