{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,11,3]],"date-time":"2024-11-03T04:02:55Z","timestamp":1730606575932,"version":"3.28.0"},"reference-count":59,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"11","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Fundamentals"],"published-print":{"date-parts":[[2024,11,1]]},"DOI":"10.1587\/transfun.2024eap1034","type":"journal-article","created":{"date-parts":[[2024,7,4]],"date-time":"2024-07-04T22:10:41Z","timestamp":1720131041000},"page":"1641-1649","source":"Crossref","is-referenced-by-count":0,"title":["Speech Emotion Detection Using Fusion on Multi-Source Low-Level Information Based Recurrent Branches"],"prefix":"10.1587","volume":"E107.A","author":[{"given":"Jiaxin","family":"WU","sequence":"first","affiliation":[{"name":"School of Integrated Circuits, Southeast University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bing","family":"LI","sequence":"additional","affiliation":[{"name":"School of Integrated Circuits, Southeast University"},{"name":"School of Cyber Science and Engineering, Southeast University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Li","family":"ZHAO","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Southeast University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinzhou","family":"XU","sequence":"additional","affiliation":[{"name":"School of Internet of Things, Nanjing University of Posts and Telecommunications"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","doi-asserted-by":"publisher","unstructured":"[1] S. Li, X. Xing, W. Fan, B. Cai, P. Fordson, and X. Xu, \u201cSpatiotemporal and frequential cascaded attention networks for speech emotion recognition,\u201d Neurocomputing, vol.448, pp.238-248, 2021. 10.1016\/j.neucom.2021.02.094","DOI":"10.1016\/j.neucom.2021.02.094"},{"key":"2","doi-asserted-by":"crossref","unstructured":"[2] R.S. Sudhakar and M.C. Anil, \u201cAnalysis of speech features for emotion detection: A review,\u201d Proc. International Conference on Computing Communication Control and Automation (ICCUBEA), Pune, India, pp.661-664, IEEE, 2015. 10.1109\/ICCUBEA.2015.135","DOI":"10.1109\/ICCUBEA.2015.135"},{"key":"3","doi-asserted-by":"publisher","unstructured":"[3] A. Koduru, H.B. Valiveti, and A.K. Budati, \u201cFeature extraction algorithms to improve the speech emotion recognition rate,\u201d Int. J. Speech Technol., vol.23, no.1, pp.45-55, 2020. 10.1007\/s10772-020-09672-4","DOI":"10.1007\/s10772-020-09672-4"},{"key":"4","doi-asserted-by":"crossref","unstructured":"[4] A. Satt, S. Rozenberg, and R. Hoory, \u201cEfficient emotion recognition from speech using deep learning on spectrograms,\u201d Proc. Annual Conference of the International Speech Communication Association (INTERSPEECH), Stockholm, Sweden, pp.1089-1093, ISCA, 2017. 10.21437\/interspeech.2017-200","DOI":"10.21437\/Interspeech.2017-200"},{"key":"5","doi-asserted-by":"publisher","unstructured":"[5] K. Hartmann, I. Siegert, D. Philippou-H\u00fcbner, and A. Wendemuth, \u201cEmotion detection in HCI: From speech features to emotion space,\u201d IFAC Symposium on Analysis, Design, and Evaluation of Human-Machine Systems, vol.46, no.15, pp.288-295, 2013. 10.3182\/20130811-5-us-2037.00049","DOI":"10.3182\/20130811-5-US-2037.00049"},{"key":"6","doi-asserted-by":"crossref","unstructured":"[6] S. Li, W. Deng, and J. Du, \u201cReliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild,\u201d Proc. IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Hawaii State, USA, pp.2852-2861, IEEE, 2017. 10.1109\/cvpr.2017.277","DOI":"10.1109\/CVPR.2017.277"},{"key":"7","doi-asserted-by":"crossref","unstructured":"[7] S. Bedoya-Jaramillo, E. Belalcazar-Bola\u00f1os, T. Villa-Ca\u00f1as, J. Orozco-Arroyave, J. Arias-Londo\u00f1o, and J. Vargas-Bonilla, \u201cAutomatic emotion detection in speech using mel frequency cesptral coefficients,\u201d Proc. Symposium of Image, Signal Processing, and Artificial Vision (STSIVA), Medellin, Antioquia, Colombia, pp.62-65, IEEE, 2012. 10.1109\/stsiva.2012.6340558","DOI":"10.1109\/STSIVA.2012.6340558"},{"key":"8","doi-asserted-by":"publisher","unstructured":"[8] S. Lalitha, D. Geyasruti, R. Narayanan, and M. Shravani, \u201cEmotion detection using MFCC and cepstrum features,\u201d Procedia Computer Science, vol.70, pp.29-35, 2015. 10.1016\/j.procs.2015.10.020","DOI":"10.1016\/j.procs.2015.10.020"},{"key":"9","doi-asserted-by":"publisher","unstructured":"[9] I. Shahin, O.A. Alomari, A.B. Nassif, I. Afyouni, I.A. Hashem, and A. Elnagar, \u201cAn efficient feature selection method for arabic and english speech emotion recognition using Grey Wolf Optimizer,\u201d Applied Acoustics, vol.205, p.109279, 2023. 10.1016\/j.apacoust.2023.109279","DOI":"10.1016\/j.apacoust.2023.109279"},{"key":"10","doi-asserted-by":"publisher","unstructured":"[10] Mustaqeem, M. Sajjad, and S. Kwon, \u201cClustering-based speech emotion Recognition by incorporating learned features and Deep BiLSTM,\u201d IEEE Access, vol.8, pp.79861-79875, 2020. 10.1109\/access.2020.2990405","DOI":"10.1109\/ACCESS.2020.2990405"},{"key":"11","doi-asserted-by":"crossref","unstructured":"[11] X. Ma, Z. Wu, J. Jia, M. Xu, H. Meng, and L. Cai, \u201cEmotion recognition from variable-length speech segments using deep learning on spectrograms,\u201d Proc. Annual Conference of the International Speech Communication Association (INTERSPEECH), Hyderabad, India, pp.3683-3687, ISCA, 2018. 10.21437\/interspeech.2018-2228","DOI":"10.21437\/Interspeech.2018-2228"},{"key":"12","doi-asserted-by":"publisher","unstructured":"[12] S.P. Mishra, P. Warule, and S. Deb, \u201cVariational mode decomposition based acoustic and entropy features for speech emotion recognition,\u201d Applied Acoustics, vol.212, p.109578, 2023. 10.1016\/j.apacoust.2023.109578","DOI":"10.1016\/j.apacoust.2023.109578"},{"key":"13","doi-asserted-by":"crossref","unstructured":"[13] N. Scheidwasser-Clow, M. Kegler, P. Beckmann, and M. Cernak, \u201cSERAB: A multi-lingual benchmark for speech emotion recognition,\u201d Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Virtual and Singapore, pp.7697-7701, IEEE, 2022. 10.1109\/icassp43922.2022.9747348","DOI":"10.1109\/ICASSP43922.2022.9747348"},{"key":"14","unstructured":"[14] A.S. Tehrani, N. Faridani, and R. Toosi, \u201cUnsupervised representations improve supervised learning in speech emotion recognition,\u201d ArXiv Preprint, ArXiv:2309.12714, 2023. 10.48550\/arXiv.2309.12714"},{"key":"15","doi-asserted-by":"crossref","unstructured":"[15] M. Baruah and B. Banerjee, \u201cSpeech emotion recognition via generation using an attention-based variational recurrent neural network,\u201d Proc. Annual Conference of the International Speech Communication Association (INTERSPEECH), Brno, Czechia, pp.4710-4714, ISCA, 2022. 10.21437\/interspeech.2022-753","DOI":"10.21437\/Interspeech.2022-753"},{"key":"16","doi-asserted-by":"publisher","unstructured":"[16] G.A. Prabhakar, B. Basel, A. Dutta, and C.V.R. Rao, \u201cMultichannel CNN-BLSTM architecture for speech emotion recognition system by fusion of magnitude and phase spectral features using DCCA for consumer applications,\u201d IEEE Trans. Consum. Electron., vol.69, no.2, pp.226-235, 2023. 10.1109\/tce.2023.3236972","DOI":"10.1109\/TCE.2023.3236972"},{"key":"17","doi-asserted-by":"crossref","unstructured":"[17] S. Sarker, K. Akter, and N. Mamun, \u201cA text independent speech emotion recognition based on convolutional neural network,\u201d Proc. International Conference on Electrical, Computer and Communication Engineering (ECCE), Swansea, UK, pp.1-4, IEEE, 2023. 10.1109\/ecce57851.2023.10101666","DOI":"10.1109\/ECCE57851.2023.10101666"},{"key":"18","doi-asserted-by":"publisher","unstructured":"[18] M. Chen, X. He, J. Yang, and H. Zhang, \u201c3-D convolutional recurrent neural networks with attention model for speech emotion recognition,\u201d IEEE Signal Process. Lett., vol.25, no.10, pp.1440-1444, 2018. 10.1109\/lsp.2018.2860246","DOI":"10.1109\/LSP.2018.2860246"},{"key":"19","doi-asserted-by":"publisher","unstructured":"[19] D.M. Schuller and B.W. Schuller, \u201cA review on five recent and near-future developments in computational processing of emotion in the human voice,\u201d Emotion Review, vol.13, no.1, pp.44-50, 2021. 10.1177\/1754073919898526","DOI":"10.1177\/1754073919898526"},{"key":"20","doi-asserted-by":"crossref","unstructured":"[20] C. Marechal, D. Miko\u0142ajewski, K. Tyburek, P. Prokopowicz, L. Bougueroua, C. Ancourt, and K. W\u0229grzyn-Wolska, \u201cSurvey on AI-based multimodal methods for emotion detection,\u201d High-performance Modelling and Simulation for Big Data Applications, LNTCS, vol.11400, pp.307-324, 2019. 10.1007\/978-3-030-16272-6_11","DOI":"10.1007\/978-3-030-16272-6_11"},{"key":"21","doi-asserted-by":"crossref","unstructured":"[21] A. Triantafyllopoulos, S. Liu, and B.W. Schuller, \u201cDeep speaker conditioning for speech emotion recognition,\u201d Proc. International Conference on Multimedia and Expo (ICME), Shenzhen, China, pp.1-6, IEEE, 2021. 10.1109\/icme51207.2021.9428217","DOI":"10.1109\/ICME51207.2021.9428217"},{"key":"22","doi-asserted-by":"crossref","unstructured":"[22] H. Zhou and K. Liu, \u201cSpeech emotion recognition with discriminative feature learning,\u201d Proc. Annual Conference of the International Speech Communication Association (INTERSPEECH), Shanghai, China, pp.4094-4097, ISCA, 2020. 10.21437\/interspeech.2020-2237","DOI":"10.21437\/Interspeech.2020-2237"},{"key":"23","doi-asserted-by":"crossref","unstructured":"[23] D. Dai, Z. Wu, R. Li, X. Wu, J. Jia, and H. Meng, \u201cLearning discriminative features from spectrograms using center loss for speech emotion recognition,\u201d Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Graz, Austria, pp.7405-7409, ISCA, 2019. 10.1109\/icassp.2019.8683765","DOI":"10.1109\/ICASSP.2019.8683765"},{"key":"24","doi-asserted-by":"crossref","unstructured":"[24] P. Kumar, S. Jain, B. Raman, P.P. Roy, and M. Iwamura, \u201cEnd-to-end Triplet loss based emotion embedding system for speech emotion recognition,\u201d Proc. International Conference on Pattern Recognition (ICPR), Virtual Event\/Milano, Italy, pp.8766-8773, Springer, 2021. 10.1109\/icpr48806.2021.9413144","DOI":"10.1109\/ICPR48806.2021.9413144"},{"key":"25","doi-asserted-by":"crossref","unstructured":"[25] T.Y. Lin, P. Goyal, R. Girshick, K. He, and P. Doll\u00e1r, \u201cFocal loss for dense object detection,\u201d Proc. International Conference on Computer Vision (ICCV), Venice, Italy, pp.2980-2988, IEEE, 2017. 10.1109\/iccv.2017.324","DOI":"10.1109\/ICCV.2017.324"},{"key":"26","doi-asserted-by":"crossref","unstructured":"[26] J. Cai, Z. Meng, A.S. Khan, Z. Li, J. O&apos;Reilly, and Y. Tong, \u201cIsland loss for learning discriminative features in facial expression recognition,\u201d Proc. International Conference on Automatic Face &amp; Gesture Recognition (FG), Xi&apos;an, China, pp.302-309, IEEE, 2018. 10.1109\/fg.2018.00051","DOI":"10.1109\/FG.2018.00051"},{"key":"27","doi-asserted-by":"publisher","unstructured":"[27] X.Y. Jing, X. Zhang, X. Zhu, F. Wu, X. You, Y. Gao, S. Shan, and J.Y. Yang, \u201cMultiset feature learning for highly imbalanced data classification,\u201d IEEE Trans. Pattern Anal. Mach. Intell., vol.43, no.1, pp.139-156, 2019. 10.1109\/TPAMI.2019.2929166","DOI":"10.1109\/TPAMI.2019.2929166"},{"key":"28","doi-asserted-by":"crossref","unstructured":"[28] Y. Chang, Z. Ren, T.T. Nguyen, K. Qian, and B.W. Schuller, \u201cKnowledge transfer for on-device speech emotion recognition with neural structured learning,\u201d Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, IEEE, 2023. 10.1109\/ICASSP49357.2023.10096757","DOI":"10.1109\/ICASSP49357.2023.10096757"},{"key":"29","doi-asserted-by":"crossref","unstructured":"[29] P. P\u00e9rez-Toro, D. Rodr\u00edguez-Salas, T. Arias-Vergara, S. Bayerl, P. Klumpp, K. Riedhammer, M. Schuster, E. N\u00f6th, A. Maier, and J. Orozco-Arroyave, \u201cTransferring quantified emotion knowledge for the detection of depression in alzheimer&apos;s disease using forestnets,\u201d Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, IEEE, 2023. 10.1109\/icassp49357.2023.10095219","DOI":"10.1109\/ICASSP49357.2023.10095219"},{"key":"30","doi-asserted-by":"publisher","unstructured":"[30] S. Lalitha, S. Tripathi, and D. Gupta, \u201cEnhanced speech emotion detection using deep neural networks,\u201d Int. J. Speech Technol., vol.22, no.3, pp.497-510, 2019. 10.1007\/s10772-018-09572-8","DOI":"10.1007\/s10772-018-09572-8"},{"key":"31","doi-asserted-by":"crossref","unstructured":"[31] Y. Shen, H. Yang, and L. Lin, \u201cAutomatic depression detection: an emotional audio-textual corpus and a Gru\/Bilstm-based model,\u201d Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Virtual and Singapore, pp.6247-6251, IEEE, 2022. 10.1109\/icassp43922.2022.9746569","DOI":"10.1109\/ICASSP43922.2022.9746569"},{"key":"32","doi-asserted-by":"crossref","unstructured":"[32] W. Wu, M. Wu, and K. Yu, \u201cClimate and weather: Inspecting depression detection via emotion recognition,\u201d Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Virtual and Singapore, pp.6262-6266, IEEE, 2022. 10.1109\/icassp43922.2022.9746634","DOI":"10.1109\/ICASSP43922.2022.9746634"},{"key":"33","doi-asserted-by":"crossref","unstructured":"[33] Y. Feng and L. Devillers, \u201cEnd-to-end continuous speech emotion recognition in real-life customer service call center conversations,\u201d Proc. International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW), pp.1-8, IEEE, 2023. 10.1109\/aciiw59127.2023.10388120","DOI":"10.1109\/ACIIW59127.2023.10388120"},{"key":"34","doi-asserted-by":"publisher","unstructured":"[34] B.T. Atmaja and M. Akagi, \u201cDimensional speech emotion recognition from speech features and word embeddings by using multitask learning,\u201d APSIPA Transactions on Signal and Information Processing, vol.9, no.1, p.e17, 2020. 10.1017\/atsip.2020.14","DOI":"10.1017\/ATSIP.2020.14"},{"key":"35","doi-asserted-by":"publisher","unstructured":"[35] F. Wang, H. Sahli, J. Gao, D. Jiang, and W. Verhelst, \u201cRelevance units machine based dimensional and continuous speech emotion prediction,\u201d Multimed. Tools Appl., vol.74, pp.9983-10000, 2015. 10.1007\/s11042-014-2319-1","DOI":"10.1007\/s11042-014-2319-1"},{"key":"36","doi-asserted-by":"crossref","unstructured":"[36] B. Mirheidari, A. Bittar, N. Cummins, J. Downs, H.L. Fisher, and H. Christensen, \u201cAutomatic detection of expressed emotion from five-minute speech samples: Challenges and opportunities,\u201d Proc. Annual Conference of the International Speech Communication Association (INTERSPEECH), Incheon, Korea, pp.2458-2462, ISCA, 2022. 10.21437\/interspeech.2022-10188","DOI":"10.21437\/Interspeech.2022-10188"},{"key":"37","doi-asserted-by":"crossref","unstructured":"[37] H. Zou, Y. Si, C. Chen, D. Rajan, and E.S. Chng, \u201cSpeech emotion recognition with co-attention based multi-level acoustic information,\u201d Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Virtual and Singapore, pp.7367-7371, IEEE, 2022. 10.1109\/icassp43922.2022.9747095","DOI":"10.1109\/ICASSP43922.2022.9747095"},{"key":"38","doi-asserted-by":"publisher","unstructured":"[38] Z. Yao, Z. Wang, W. Liu, Y. Liu, and J. Pan, \u201cSpeech emotion recognition using fusion of three multi-task learning-based classifiers: HSF-DNN, MS-CNN and LLD-RNN,\u201d Speech Communication, vol.120, pp.11-19, 2020. 10.1016\/j.specom.2020.03.005","DOI":"10.1016\/j.specom.2020.03.005"},{"key":"39","doi-asserted-by":"crossref","unstructured":"[39] M. Luo, H. Phan, and J. Reiss, \u201cCross-modal fusion techniques for utterance-level emotion recognition from text and speech,\u201d Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, IEEE, 2023. 10.1109\/icassp49357.2023.10096885","DOI":"10.1109\/ICASSP49357.2023.10096885"},{"key":"40","doi-asserted-by":"publisher","unstructured":"[40] Y. Xie, R. Liang, Z. Liang, C. Huang, C. Zou, and B. Schuller, \u201cSpeech emotion classification using attention-based LSTM,\u201d IEEE\/ACM Trans. Audio, Speech, Language Process., vol.27, no.11, pp.1675-1685, 2019. 10.1109\/taslp.2019.2925934","DOI":"10.1109\/TASLP.2019.2925934"},{"key":"41","doi-asserted-by":"crossref","unstructured":"[41] L. Tarantino, P.N. Garner, A. Lazaridis, \u201cSelf-attention for speech emotion recognition,\u201d Proc. Annual Conference of the International Speech Communication Association (INTERSPEECH), Graz, Austria, pp.2578-2582, ISCA, 2019. 10.21437\/interspeech.2019-2822","DOI":"10.21437\/Interspeech.2019-2822"},{"key":"42","doi-asserted-by":"crossref","unstructured":"[42] Z. Zhao, H. Wang, H. Wang, and B. Schuller, \u201cHierarchical network with decoupled knowledge distillation for speech emotion recognition,\u201d Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, IEEE, 2023. 10.1109\/icassp49357.2023.10095045","DOI":"10.1109\/ICASSP49357.2023.10095045"},{"key":"43","doi-asserted-by":"crossref","unstructured":"[43] S. Kakouros, T. Stafylakis, L. Mo\u0161ner, and L. Burget, \u201cSpeech-based emotion recognition with self-supervised models using attentive channel-wise correlations and label smoothing,\u201d Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, IEEE, 2023. 10.1109\/ICASSP49357.2023.10094673","DOI":"10.1109\/ICASSP49357.2023.10094673"},{"key":"44","doi-asserted-by":"crossref","unstructured":"[44] K. Liu, D. Wang, D. Wu, and J. Feng, \u201cSpeech emotion recognition via two-stream pooling attention with discriminative channel weighting,\u201d Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, IEEE, 2023. 10.1109\/icassp49357.2023.10095588","DOI":"10.1109\/ICASSP49357.2023.10095588"},{"key":"45","doi-asserted-by":"publisher","unstructured":"[45] M. Rayhan Ahmed, S. Islam, A. Muzahidul Islam, and S. Shatabda, \u201cAn ensemble 1D-CNN-LSTM-GRU model with data augmentation for speech emotion recognition,\u201d Expert Systems with Applications, vol.218, p.119633, 2023. 10.1016\/j.eswa.2023.119633","DOI":"10.1016\/j.eswa.2023.119633"},{"key":"46","doi-asserted-by":"crossref","unstructured":"[46] D. Bertero and P. Fung, \u201cA first look into a convolutional neural network for speech emotion detection,\u201d Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA, pp.5115-5119, IEEE, 2017. 10.1109\/icassp.2017.7953131","DOI":"10.1109\/ICASSP.2017.7953131"},{"key":"47","doi-asserted-by":"publisher","unstructured":"[47] N. Majumder, S. Poria, D. Hazarika, R. Mihalcea, A. Gelbukh, and E. Cambria, \u201cDialogueRNN: An attentive RNN for emotion detection in conversations,\u201d Proc. AAAI Conference on Artificial Intelligence, Hawaii, USA, pp.6818-6825, AAAI Press, 2019. 10.1609\/aaai.v33i01.33016818","DOI":"10.1609\/aaai.v33i01.33016818"},{"key":"48","doi-asserted-by":"crossref","unstructured":"[48] J. Santoso, T. Yamada, K. Ishizuka, T. Hashimoto, and S. Makino, \u201cPerformance improvement of speech emotion recognition by neutral speech detection using autoencoder and intermediate representation,\u201d Proc. Annual Conference of the International Speech Communication Association (INTERSPEECH), Incheon, Korea, pp.4700-4704, ISCA, 2022. 10.21437\/interspeech.2022-584","DOI":"10.21437\/Interspeech.2022-584"},{"key":"49","doi-asserted-by":"publisher","unstructured":"[49] W. Li, J. Xue, R. Tan, C. Wang, Z. Deng, S. Li, G. Guo, and D. Cao, \u201cGlobal-local-feature-fused driver speech emotion detection for intelligent cockpit in automated driving,\u201d IEEE Trans. Intell. Veh., vol.8, no.4, pp.2684-2697, 2023. 10.1109\/tiv.2023.3259988","DOI":"10.1109\/TIV.2023.3259988"},{"key":"50","doi-asserted-by":"publisher","unstructured":"[50] X. Qin, Z. Wu, T. Zhang, Y. Li, J. Luan, B. Wang, L. Wang, and J. Cui, \u201cBERT-ERC: Fine-tuning BERT is enough for emotion recognition in conversation,\u201d Proc. AAAI Conference on Artificial Intelligence, Washington, DC, USA, pp.13492-13500, 2023. 10.1609\/aaai.v37i11.26582","DOI":"10.1609\/aaai.v37i11.26582"},{"key":"51","doi-asserted-by":"crossref","unstructured":"[51] Y. Wang, J. Wang, and X. Zhang, \u201cYNU-HPCC at WASSA-2023 shared task 1: Large-scale language model with LoRA fine-tuning for empathy detection and emotion classification,\u201d Proc. Workshop on Computational Approaches to Subjectivity, Sentiment, &amp; Social Media Analysis (WASSA), Toronto, Canada, pp.526-530, Association for Computational Linguistics, 2023. 10.18653\/v1\/2023.wassa-1.45","DOI":"10.18653\/v1\/2023.wassa-1.45"},{"key":"52","doi-asserted-by":"crossref","unstructured":"[52] F. Eyben, F. Weninger, F. Gross, and B. Schuller, \u201cRecent developments in openSMILE, the munich open-source multimedia feature extractor,\u201d Proc. ACM International Conference on Multimedia, Barcelona, Spain, pp.835-838, ACM, 2013. 10.1145\/2502081.2502224","DOI":"10.1145\/2502081.2502224"},{"key":"53","unstructured":"[53] Y. Wang, A. Boumadane, and A. Heba, \u201cA fine-tuned wav2vec 2.0\/HuBERT benchmark for speech emotion recognition, speaker verification and spoken language understanding,\u201d ArXiv Preprint, ArXiv:2111.02735, 2021. 10.48550\/arXiv.2111.02735"},{"key":"54","doi-asserted-by":"publisher","unstructured":"[54] C. Busso, M. Bulut, C.C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J.N. Chang, S. Lee, and S.S. Narayanan, \u201cIEMOCAP: Interactive emotional dyadic motion capture database,\u201d Lang. Resources &amp; Evaluation, vol.42, no.4, pp.335-359, 2008. 10.1007\/s10579-008-9076-6","DOI":"10.1007\/s10579-008-9076-6"},{"key":"55","doi-asserted-by":"crossref","unstructured":"[55] B. Schuller, S. Steidl, A. Batliner, A. Vinciarelli, K. Scherer, F. Ringeval, M. Chetouani, F. Weninger, F. Eyben, E. Marchi, M. Mortillaro, H. Salamin, A. Polychroniou, F. Valente, and S. Kim, \u201cThe INTERSPEECH 2013 computational paralinguistics challenge: Social signals, conflict, emotion, autism,\u201d Proc. Annual Conference of the International Speech Communication Association (INTERSPEECH), Lyon, France, pp.148-152, ISCA, 2013. 10.21437\/interspeech.2013-56","DOI":"10.21437\/Interspeech.2013-56"},{"key":"56","doi-asserted-by":"crossref","unstructured":"[56] M. Macary, M. Tahon, Y. Est\u00e9ve, and A. Rousseau, \u201cOn the use of self-supervised pre-trained acoustic and linguistic features for continuous speech emotion recognition,\u201d 2021 IEEE Spoken Language Technology Workshop (SLT), pp.373-380, 2021. 10.1109\/slt48900.2021.9383456","DOI":"10.1109\/SLT48900.2021.9383456"},{"key":"57","doi-asserted-by":"publisher","unstructured":"[57] S. Li, P. Song, and W. Zheng, \u201cMulti-source discriminant subspace alignment for cross-domain speech emotion recognition,\u201d IEEE\/ACM Trans. Audio, Speech, Language Process., vol.31, pp.2448-2460, 2023. 10.1109\/taslp.2023.3288415","DOI":"10.1109\/TASLP.2023.3288415"},{"key":"58","doi-asserted-by":"publisher","unstructured":"[58] W. Zhang, P. Song, D. Chen, C. Sheng, and W. Zhang, \u201cCross-corpus speech emotion recognition based on joint transfer subspace learning and regression,\u201d IEEE Trans. Cogn. Develop. Syst., vol.14, no.2, pp.588-598, 2021. 10.1109\/tcds.2021.3055524","DOI":"10.1109\/TCDS.2021.3055524"},{"key":"59","doi-asserted-by":"publisher","unstructured":"[59] W. Zhang and P. Song, \u201cTransfer sparse discriminant subspace learning for cross-corpus speech emotion recognition,\u201d IEEE\/ACM Trans. Audio, Speech, Language Process., vol.28, pp.307-318, 2019. 10.1109\/taslp.2019.2955252","DOI":"10.1109\/TASLP.2019.2955252"}],"container-title":["IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transfun\/E107.A\/11\/E107.A_2024EAP1034\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,2]],"date-time":"2024-11-02T03:24:08Z","timestamp":1730517848000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transfun\/E107.A\/11\/E107.A_2024EAP1034\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,1]]},"references-count":59,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2024]]}},"URL":"https:\/\/doi.org\/10.1587\/transfun.2024eap1034","relation":{},"ISSN":["0916-8508","1745-1337"],"issn-type":[{"type":"print","value":"0916-8508"},{"type":"electronic","value":"1745-1337"}],"subject":[],"published":{"date-parts":[[2024,11,1]]},"article-number":"2024EAP1034"}}