{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,27]],"date-time":"2026-07-27T23:10:10Z","timestamp":1785193810160,"version":"3.55.0"},"reference-count":59,"publisher":"MDPI AG","issue":"22","license":[{"start":{"date-parts":[[2020,11,23]],"date-time":"2020-11-23T00:00:00Z","timestamp":1606089600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Research Foundation (NRF) grant funded by the MSIP of Korea","award":["2019R1A2C2009480"],"award-info":[{"award-number":["2019R1A2C2009480"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Speech emotion recognition predicts the emotional state of a speaker based on the person\u2019s speech. It brings an additional element for creating more natural human\u2013computer interactions. Earlier studies on emotional recognition have been primarily based on handcrafted features and manual labels. With the advent of deep learning, there have been some efforts in applying the deep-network-based approach to the problem of emotion recognition. As deep learning automatically extracts salient features correlated to speaker emotion, it brings certain advantages over the handcrafted-feature-based methods. There are, however, some challenges in applying them to the emotion recognition problem, because data required for properly training deep networks are often lacking. Therefore, there is a need for a new deep-learning-based approach which can exploit available information from given speech signals to the maximum extent possible. Our proposed method, called \u201cFusion-ConvBERT\u201d, is a parallel fusion model consisting of bidirectional encoder representations from transformers and convolutional neural networks. Extensive experiments were conducted on the proposed model using the EMO-DB and Interactive Emotional Dyadic Motion Capture Database emotion corpus, and it was shown that the proposed method outperformed state-of-the-art techniques in most of the test configurations.<\/jats:p>","DOI":"10.3390\/s20226688","type":"journal-article","created":{"date-parts":[[2020,11,23]],"date-time":"2020-11-23T08:18:23Z","timestamp":1606119503000},"page":"6688","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":36,"title":["Fusion-ConvBERT: Parallel Convolution and BERT Fusion for Speech Emotion Recognition"],"prefix":"10.3390","volume":"20","author":[{"given":"Sanghyun","family":"Lee","sequence":"first","affiliation":[{"name":"Department of Electronics and Electrical Engineering, Korea University, Seoul 136-713, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"David K.","family":"Han","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, Drexel University, Philadelphia, PA 19104, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8744-4514","authenticated-orcid":false,"given":"Hanseok","family":"Ko","sequence":"additional","affiliation":[{"name":"Department of Electronics and Electrical Engineering, Korea University, Seoul 136-713, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2020,11,23]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Gupta, R., Malandrakis, N., Xiao, B., Guha, T., Van Segbroeck, M., Black, M., Potamianos, A., and Narayanan, S. (2014, January 3\u20137). Multimodal prediction of affective dimensions and depression in human-computer interactions. Proceedings of the 4th International Workshop on Audio\/Visual Emotion Challenge, Orlando, FL, USA.","DOI":"10.1145\/2661806.2661810"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"90","DOI":"10.1145\/3129340","article-title":"Speech emotion recognition: Two decades in a nutshell, benchmarks, and ongoing trends","volume":"61","author":"Schuller","year":"2018","journal-title":"Commun. ACM"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1880","DOI":"10.1109\/TMM.2013.2269314","article-title":"Two-level hierarchical alignment for semi-coupled HMM-based audiovisual emotion recognition with temporal course","volume":"15","author":"Wu","year":"2013","journal-title":"IEEE Trans. Multimed."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"142","DOI":"10.1109\/TMM.2011.2171334","article-title":"Error weighted semi-coupled hidden Markov model for audio-visual emotion recognition","volume":"14","author":"Lin","year":"2011","journal-title":"IEEE Trans. Multimed."},{"key":"ref_5","first-page":"11","article-title":"The influence of language and culture on the understanding of vocal emotions","volume":"6","author":"Altrov","year":"2015","journal-title":"J. Est. Finno Ugric Linguist."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"572","DOI":"10.1016\/j.patcog.2010.09.020","article-title":"Survey on speech emotion recognition: Features, classification schemes, and databases","volume":"44","author":"Kamel","year":"2011","journal-title":"Pattern Recognit."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Tzinis, E., and Potamianos, A. (2017, January 23\u201326). Segment-based speech emotion recognition using recurrent neural networks. Proceedings of the 2017 Seventh International Conference on Affective Computing and Intelligent Interaction (ACII), San Antonio, TX, USA.","DOI":"10.1109\/ACII.2017.8273599"},{"key":"ref_8","unstructured":"Anand, N., and Verma, P. (2015). Convoluted feelings convolutional and recurrent nets for detecting emotion from audio data. Technical Report, Stanford University."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Trigeorgis, G., Ringeval, F., Brueckner, R., Marchi, E., Nicolaou, M.A., Schuller, B., and Zafeiriou, S. (2016, January 20\u201325). Adieu features?. end-to-end speech emotion recognition using a deep convolutional recurrent network. In Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472669"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1694","DOI":"10.3390\/s17071694","article-title":"Emotion recognition from Chinese speech for smart affective services using a combination of SVM and DBN","volume":"17","author":"Zhu","year":"2017","journal-title":"Sensors"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Kim, Y., Lee, H., and Provost, E.M. (2013, January 26\u201331). Deep learning for robust feature generation in audiovisual emotion recognition. Proceedings of the 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6638346"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Stuhlsatz, A., Meyer, C., Eyben, F., Zielke, T., Meier, G., and Schuller, B. (2011, January 22\u201327). Deep neural networks for acoustic emotion recognition: Raising the benchmarks. Proceedings of the 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Prague, Czech Republic.","DOI":"10.1109\/ICASSP.2011.5947651"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Huang, Z., Dong, M., Mao, Q., and Zhan, Y. (2014, January 3\u20137). Speech emotion recognition using CNN. Proceedings of the 22nd ACM International Conference on Multimedia, Orlando, FL, USA.","DOI":"10.1145\/2647868.2654984"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"2203","DOI":"10.1109\/TMM.2014.2360798","article-title":"Learning salient features for speech emotion recognition using convolutional neural networks","volume":"16","author":"Mao","year":"2014","journal-title":"IEEE Trans. Multimed."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1440","DOI":"10.1109\/LSP.2018.2860246","article-title":"3-D convolutional recurrent neural networks with attention model for speech emotion recognition","volume":"25","author":"Chen","year":"2018","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_16","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_17","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141, and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Peters, M.E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. (2018). Deep contextualized word representations. arXiv.","DOI":"10.18653\/v1\/N18-1202"},{"key":"ref_19","unstructured":"Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (2020, November 23). Improving Language Understanding by Generative Pre-Training. Available online: https:\/\/s3-us-west-2.amazonaws.com\/openaiassets\/research-covers\/languageunsupervised\/languageunderstandingpaper.pdf."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Ling, S., Liu, Y., Salazar, J., and Kirchhoff, K. (2020, January 4\u20138). Deep contextualized acoustic representations for semi-supervised speech recognition. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9053176"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Liu, A.T., Yang, S.W., Chi, P.H., Hsu, P.C., and Lee, H.Y. (2020, January 4\u20138). Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9054458"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Schneider, S., Baevski, A., Collobert, R., and Auli, M. (2019). Wav2vec: Unsupervised pre-training for speech recognition. arXiv.","DOI":"10.21437\/Interspeech.2019-1873"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Gideon, J., Khorram, S., Aldeneh, Z., Dimitriadis, D., and Provost, E.M. (2017). Progressive neural networks for transfer learning in emotion recognition. arXiv.","DOI":"10.21437\/Interspeech.2017-1637"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1007\/s10579-008-9076-6","article-title":"IEMOCAP: Interactive emotional dyadic motion capture database","volume":"42","author":"Busso","year":"2008","journal-title":"Lang. Resour. Eval."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Burkhardt, F., Paeschke, A., Rolfes, M., Sendlmeier, W.F., and Weiss, B. (2005, January 4\u20138). A database of German emotional speech. Proceedings of the Ninth European Conference on Speech Communication and Technology, Lisbon, Portugal.","DOI":"10.21437\/Interspeech.2005-446"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"603","DOI":"10.1016\/S0167-6393(03)00099-2","article-title":"Speech emotion recognition using hidden Markov models","volume":"41","author":"Nwe","year":"2003","journal-title":"Speech Commun."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Schuller, B., Rigoll, G., and Lang, M. (2003, January 6\u201310). Hidden Markov model-based speech emotion recognition. Proceedings of the 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201903), Hong Kong, China.","DOI":"10.1109\/ICME.2003.1220939"},{"key":"ref_28","unstructured":"Ververidis, D., and Kotropoulos, C. (2005, January 6). Emotional speech classification using Gaussian mixture models and the sequential floating forward selection algorithm. Proceedings of the 2005 IEEE International Conference on Multimedia and Expo, Amsterdam, The Netherlands."},{"key":"ref_29","unstructured":"Schuller, B., Rigoll, G., and Lang, M. (2004, January 17\u201321). Speech emotion recognition combining acoustic features and linguistic information in a hybrid support vector machine-belief network architecture. Proceedings of the 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing, Montreal, QC, Canada."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Seehapoch, T., and Wongthanavasu, S. (February, January 31). Speech emotion recognition using support vector machines. Proceedings of the 2013 5th international conference on Knowledge and smart technology (KST), Chonburi, Thailand.","DOI":"10.1109\/KST.2013.6512793"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.eij.2015.05.004","article-title":"Music emotion recognition: The combined evidence of MFCC and residual phase","volume":"17","author":"Nalini","year":"2016","journal-title":"Egypt. Inform. J."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wen, G., Li, H., Huang, J., Li, D., and Xun, E. (2017). Random deep belief networks for recognizing emotions from speech signals. Comput. Intell. Neurosci., 2017.","DOI":"10.1155\/2017\/1945630"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"7717","DOI":"10.1109\/ACCESS.2018.2888882","article-title":"Sound classification using convolutional neural network and tensor deep stacking network","volume":"7","author":"Khamparia","year":"2019","journal-title":"IEEE Access"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"92871","DOI":"10.1109\/ACCESS.2019.2928017","article-title":"ECG arrhythmia classification using STFT-based spectrogram and convolutional neural network","volume":"7","author":"Huang","year":"2019","journal-title":"IEEE Access"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Cummins, N., Amiriparian, S., Hagerer, G., Batliner, A., Steidl, S., and Schuller, B.W. (2017, January 23\u201327). An image-based deep spectrum feature representation for the recognition of emotional speech. Proceedings of the 25th ACM International Conference on Multimedia, Mountain View, CA, USA.","DOI":"10.1145\/3123266.3123371"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"93847","DOI":"10.1109\/ACCESS.2019.2924597","article-title":"Dual exclusive attentive transfer for unsupervised deep convolutional domain adaptation in speech emotion recognition","volume":"7","author":"Ocquaye","year":"2019","journal-title":"IEEE Access"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Keren, G., and Schuller, B. (2016, January 24\u201329). Convolutional RNN: An enhanced model for extracting features from sequential data. Proceedings of the 2016 International Joint Conference on Neural Networks (IJCNN), Vancouver, BC, Canada.","DOI":"10.1109\/IJCNN.2016.7727636"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"125868","DOI":"10.1109\/ACCESS.2019.2938007","article-title":"Speech emotion recognition from 3D log-mel spectrograms with deep learning network","volume":"7","author":"Meng","year":"2019","journal-title":"IEEE Access"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Mirsamadi, S., Barsoum, E., and Zhang, C. (2017, January 5\u20139). Automatic speech emotion recognition using recurrent neural networks with local attention. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7952552"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"2730","DOI":"10.3390\/s19122730","article-title":"Speech emotion recognition with heterogeneous feature unification of deep neural network","volume":"19","author":"Jiang","year":"2019","journal-title":"Sensors"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"75798","DOI":"10.1109\/ACCESS.2019.2921390","article-title":"Exploration of complementary features for speech emotion recognition based on kernel extreme learning machine","volume":"7","author":"Guo","year":"2019","journal-title":"IEEE Access"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Lim, W., Jang, D., and Lee, T. (2016, January 13\u201315). Speech emotion recognition using convolutional and recurrent neural networks. Proceedings of the 2016 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA), Jeju, Korea.","DOI":"10.1109\/APSIPA.2016.7820699"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"McFee, B., Raffel, C., Liang, D., Ellis, D.P., McVicar, M., Battenberg, E., and Nieto, O. (2015, January 6\u201312). Librosa: Audio and music signal analysis in python. Proceedings of the 14th Python in Science Conference, Austin, TX, USA.","DOI":"10.25080\/Majora-7b98e3ed-003"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"153","DOI":"10.1016\/j.dsp.2007.12.004","article-title":"Time\u2013frequency feature representation using energy concentration: An overview of recent advances","volume":"19","author":"Jiang","year":"2009","journal-title":"Digit. Signal Process."},{"key":"ref_45","unstructured":"Douglas, O., and Shaughnessy, O. (2000). Speech Communications: Human and Machine, IEEE Press."},{"key":"ref_46","unstructured":"Ba, J.L., Kiros, J.R., and Hinton, G.E. (2016). Layer normalization. arXiv."},{"key":"ref_47","unstructured":"Xie, C., Tan, M., Gong, B., Yuille, A., and Le, Q.V. (2020). Smooth adversarial training. arXiv."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Palaz, D., Collobert, R., and Doss, M.M. (2013). Estimating phoneme class conditional probabilities from raw speech signal using convolutional neural networks. arXiv.","DOI":"10.21437\/Interspeech.2013-438"},{"key":"ref_49","unstructured":"Behnke, S. (2003, January 20\u201324). Discovering hierarchical speech features using convolutional non-negative matrix factorization. Proceedings of the International Joint Conference on Neural Networks, Portland, OR, USA."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Xia, R., and Liu, Y. (2016, January 8\u201312). DBN-ivector Framework for Acoustic Emotion Recognition. Proceedings of the INTERSPEECH, San Francisco, CA, USA.","DOI":"10.21437\/Interspeech.2016-488"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Zhao, Z., Bao, Z., Zhang, Z., Cummins, N., Wang, H., and Schuller, B.W. (2019, January 15\u201319). Attention-Enhanced Connectionist Temporal Classification for Discrete Speech Emotion Recognition. Proceedings of the INTERSPEECH, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-1649"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. (2015, January 19\u201324). Librispeech: An asr corpus based on public domain audio books. Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, Australia.","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"ref_53","first-page":"1929","article-title":"Dropout: A simple way to prevent neural networks from overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref_54","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Park, D.S., Chan, W., Zhang, Y., Chiu, C.C., Zoph, B., Cubuk, E.D., and Le, Q.V. (2019). Specaugment: A simple data augmentation method for automatic speech recognition. arXiv.","DOI":"10.21437\/Interspeech.2019-2680"},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"5571","DOI":"10.1007\/s11042-017-5292-7","article-title":"Deep features-based speech emotion recognition for smart affective services","volume":"78","author":"Badshah","year":"2019","journal-title":"Multimed. Tools Appl."},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"90368","DOI":"10.1109\/ACCESS.2019.2927384","article-title":"Parallelized convolutional recurrent neural network with spectral features for speech emotion recognition","volume":"7","author":"Jiang","year":"2019","journal-title":"IEEE Access"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Zheng, W., Yu, J., and Zou, Y. (2015, January 21\u201324). An experimental study of speech emotion recognition based on deep convolutional neural networks. Proceedings of the 2015 International Conference on Affective Computing and Intelligent Interaction (ACII), Xi\u2019an, China.","DOI":"10.1109\/ACII.2015.7344669"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Luo, D., Zou, Y., and Huang, D. (2018, January 2\u20136). Investigation on Joint Representation Learning for Robust Feature Extraction in Speech Emotion Recognition. Proceedings of the INTERSPEECH, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-1832"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/22\/6688\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:35:55Z","timestamp":1760178955000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/22\/6688"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,11,23]]},"references-count":59,"journal-issue":{"issue":"22","published-online":{"date-parts":[[2020,11]]}},"alternative-id":["s20226688"],"URL":"https:\/\/doi.org\/10.3390\/s20226688","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,11,23]]}}}