{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T02:42:35Z","timestamp":1787020955975,"version":"build-2736575974"},"reference-count":39,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,5,15]],"date-time":"2023-05-15T00:00:00Z","timestamp":1684108800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,5,15]],"date-time":"2023-05-15T00:00:00Z","timestamp":1684108800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Research project of the Macao Polytechnic University","award":["RP\/ESCA-03\/2021"],"award-info":[{"award-number":["RP\/ESCA-03\/2021"]}]},{"name":"Research project of the Macao Polytechnic University","award":["RP\/FCA-12\/2022"],"award-info":[{"award-number":["RP\/FCA-12\/2022"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Emotion plays a dominant role in speech. The same utterance with different emotions can lead to a completely different meaning. The ability to perform various of emotion during speaking is also one of the typical characters of human. In this case, technology trends to develop advanced speech emotion classification algorithms in the demand of enhancing the interaction between computer and human beings. This paper proposes a speech emotion classification approach based on the paralinguistic and spectral features extraction. The Mel-frequency cepstral coefficients (MFCC) are extracted as spectral feature, and openSMILE is employed to extract the paralinguistic feature. The machine learning techniques multi-layer perceptron classifier and support vector machines are respectively applied into the extracted features for the classification of the speech emotions. We have conducted experiments on the Berlin database to evaluate the performance of the proposed approach. Experimental results show that the proposed approach achieves satisfied performances. Comparisons are conducted in clean condition and noisy condition respectively, and the results indicate better performance of the proposed scheme.<\/jats:p>","DOI":"10.1186\/s13636-023-00290-x","type":"journal-article","created":{"date-parts":[[2023,5,15]],"date-time":"2023-05-15T10:29:59Z","timestamp":1684146599000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["Paralinguistic and spectral feature extraction for speech emotion classification using machine learning techniques"],"prefix":"10.1186","volume":"2023","author":[{"given":"Tong","family":"Liu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7490-6695","authenticated-orcid":false,"given":"Xiaochen","family":"Yuan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,5,15]]},"reference":[{"key":"290_CR1","doi-asserted-by":"crossref","unstructured":"X.\u00a0Cao, M.\u00a0Jia, J.\u00a0Ru, T.w. Pai, Cross-corpus speech emotion recognition using subspace learning and domain adaption. EURASIP J. Audio Speech Music Process. 2022(1), 32 (2022)","DOI":"10.1186\/s13636-022-00264-5"},{"issue":"1","key":"290_CR2","doi-asserted-by":"publisher","first-page":"69","DOI":"10.1109\/TAFFC.2015.2392101","volume":"6","author":"K Wang","year":"2015","unstructured":"K. Wang, N. An, B.N. Li, Y. Zhang, L. Li, Speech emotion recognition using fourier parameters. IEEE Trans. Affect. Comput. 6(1), 69\u201375 (2015)","journal-title":"IEEE Trans. Affect. Comput."},{"issue":"1","key":"290_CR3","doi-asserted-by":"publisher","first-page":"18","DOI":"10.1186\/s13636-021-00208-5","volume":"2021","author":"D Tang","year":"2021","unstructured":"D. Tang, P. Kuppens, L. Geurts, T. van Waterschoot, End-to-end speech emotion recognition using a novel context-stacking dilated convolution neural network. EURASIP J. Audio Speech Music Process. 2021(1), 18 (2021)","journal-title":"EURASIP J. Audio Speech Music Process."},{"issue":"1","key":"290_CR4","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s13636-018-0145-5","volume":"2019","author":"L Sun","year":"2019","unstructured":"L. Sun, S. Fu, F. Wang, Decision tree svm model with fisher feature selection for speech emotion recognition. EURASIP J. Audio Speech Music Process. 2019(1), 1\u201314 (2019)","journal-title":"EURASIP J. Audio Speech Music Process."},{"issue":"3\u20134","key":"290_CR5","doi-asserted-by":"publisher","first-page":"169","DOI":"10.1080\/02699939208411068","volume":"6","author":"P Ekman","year":"1992","unstructured":"P. Ekman, An argument for basic emotions. Cogn. Emot. 6(3\u20134), 169\u2013200 (1992)","journal-title":"Cogn. Emot."},{"issue":"6","key":"290_CR6","doi-asserted-by":"publisher","first-page":"1161","DOI":"10.1037\/h0077714","volume":"39","author":"JA Russell","year":"1980","unstructured":"J.A. Russell, A circumplex model of affect. J. Pers. Soc. Psychol. 39(6), 1161 (1980)","journal-title":"J. Pers. Soc. Psychol."},{"issue":"1","key":"290_CR7","doi-asserted-by":"publisher","first-page":"23","DOI":"10.2991\/ijcis.d.201019.002","volume":"14","author":"A Cabri","year":"2020","unstructured":"A. Cabri, F. Masulli, Z. Mnasri, S. Rovetta et al., Emotion recognition from speech: an unsupervised learning approach. Int. J. Comput. Intell. Syst. 14(1), 23 (2020)","journal-title":"Int. J. Comput. Intell. Syst."},{"issue":"317","key":"290_CR8","first-page":"31","volume":"2293","author":"KR Scherer","year":"1984","unstructured":"K.R. Scherer et al., On the nature and function of emotion: a component process approach. Approaches Emot. 2293(317), 31 (1984)","journal-title":"Approaches Emot."},{"key":"290_CR9","doi-asserted-by":"crossref","unstructured":"Rao, K.S., Koolagudi, S.G. Robust emotion recognition using spectral and prosodic features. In: Springer Science & Business Media, Springer, New York (2013)","DOI":"10.1007\/978-1-4614-6360-3"},{"issue":"11","key":"290_CR10","doi-asserted-by":"publisher","first-page":"1675","DOI":"10.1109\/TASLP.2019.2925934","volume":"27","author":"Y Xie","year":"2019","unstructured":"Y. Xie, R. Liang, Z. Liang, C. Huang, C. Zou, B. Schuller, Speech emotion classification using attention-based lstm. IEEE\/ACM Trans. Audio Speech Lang. Process. 27(11), 1675\u20131685 (2019)","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"1","key":"290_CR11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s13636-022-00240-z","volume":"2022","author":"Y Xu","year":"2022","unstructured":"Y. Xu, W. Wang, H. Cui, M. Xu, M. Li, Paralinguistic singing attribute recognition using supervised machine learning for describing the classical tenor solo singing voice in vocal pedagogy. EURASIP J. Audio Speech Music Process. 2022(1), 1\u201316 (2022)","journal-title":"EURASIP J. Audio Speech Music Process."},{"issue":"3\u20134","key":"290_CR12","doi-asserted-by":"publisher","first-page":"455","DOI":"10.1016\/j.specom.2005.02.018","volume":"46","author":"E Shriberg","year":"2005","unstructured":"E. Shriberg, L. Ferrer, S. Kajarekar, A. Venkataraman, A. Stolcke, Modeling prosodic feature sequences for speaker recognition. Speech Commun. 46(3\u20134), 455\u2013472 (2005)","journal-title":"Speech Commun."},{"issue":"4","key":"290_CR13","doi-asserted-by":"publisher","first-page":"1892","DOI":"10.1109\/TAFFC.2022.3188223","volume":"13","author":"SR Kshirsagar","year":"2022","unstructured":"S.R. Kshirsagar, T.H. Falk, Quality-aware bag of modulation spectrum features for robust speech emotion recognition. IEEE Trans. Affect. Comput. 13(4), 1892\u20131905 (2022)","journal-title":"IEEE Trans. Affect. Comput."},{"key":"290_CR14","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s13636-021-00216-5","volume":"2021","author":"M Geravanchizadeh","year":"2021","unstructured":"M. Geravanchizadeh, E. Forouhandeh, M. Bashirpour, Feature compensation based on the normalization of vocal tract length for the improvement of emotion-affected speech recognition. EURASIP J. Audio Speech Music Process. 2021, 1\u201319 (2021)","journal-title":"EURASIP J. Audio Speech Music Process."},{"key":"290_CR15","doi-asserted-by":"crossref","unstructured":"J.L. Jacobson, D.C. Boersma, R.B. Fields, K.L. Olson, Paralinguistic features of adult speech to infants and small children. Child Dev. 54(2), 436\u2013442 (1983)","DOI":"10.2307\/1129704"},{"key":"290_CR16","doi-asserted-by":"crossref","unstructured":"S.M. Tsai, in 2013 1st International Conference on Orange Technologies (ICOT), A robust zero-watermarking algorithm for audio based on LPCC (IEEE, 2013), pp. 63\u201366","DOI":"10.1109\/ICOT.2013.6521158"},{"key":"290_CR17","unstructured":"C. Ittichaichareon, S. Suksri, T. Yingthawornsuk, in International conference on computer graphics, simulation and modeling (ICGSM'2012), vol. 9, Speech recognition using mfcc, Pattaya, Thailand (2012)"},{"issue":"4","key":"290_CR18","doi-asserted-by":"publisher","first-page":"603","DOI":"10.1016\/S0167-6393(03)00099-2","volume":"41","author":"TL Nwe","year":"2003","unstructured":"T.L. Nwe, S.W. Foo, L.C. De Silva, Speech emotion recognition using hidden markov models. Speech Commun. 41(4), 603\u2013623 (2003)","journal-title":"Speech Commun."},{"key":"290_CR19","unstructured":"F.\u00a0Albu, D.\u00a0Hagiescu, L.\u00a0Vladutu, M.A. Puica, in EDULEARN15 Proceedings, Neural network approaches for children\u2019s emotion recognition in intelligent learning applications (IATED, 2015), pp. 3229\u20133239"},{"issue":"1\u20133","key":"290_CR20","doi-asserted-by":"publisher","first-page":"489","DOI":"10.1016\/j.neucom.2005.12.126","volume":"70","author":"GB Huang","year":"2006","unstructured":"G.B. Huang, Q.Y. Zhu, C.K. Siew, Extreme learning machine: theory and applications. Neurocomputing 70(1\u20133), 489\u2013501 (2006)","journal-title":"Neurocomputing"},{"issue":"2","key":"290_CR21","doi-asserted-by":"publisher","first-page":"1883","DOI":"10.4249\/scholarpedia.1883","volume":"4","author":"LE Peterson","year":"2009","unstructured":"L.E. Peterson, K-nearest neighbor. Scholarpedia 4(2), 1883 (2009)","journal-title":"K-nearest neighbor. Scholarpedia"},{"key":"290_CR22","doi-asserted-by":"crossref","unstructured":"B.\u00a0Schuller, G.\u00a0Rigoll, M.\u00a0Lang, in 2004 IEEE international conference on acoustics, speech, and signal processing, vol.\u00a01, Speech emotion recognition combining acoustic features and linguistic information in a hybrid support vector machine-belief network architecture (IEEE, 2004), pp. I\u2013577","DOI":"10.1109\/ICASSP.2004.1326051"},{"key":"290_CR23","doi-asserted-by":"crossref","unstructured":"K. Han, D. Yu, I. Tashev, in INTERSPEECH 2014, Speech emotion recognition using deep neural network and extreme learning machine, ISCA, Singapore (2014)","DOI":"10.21437\/Interspeech.2014-57"},{"key":"290_CR24","doi-asserted-by":"crossref","unstructured":"B. Schuller, S. Steidl, A. Batliner, F. Burkhardt, L. Devillers, C. M\u00fcller, S. Narayanan, in Proc. INTERSPEECH 2010, The INTERSPEECH 2010 paralinguistic challenge, ISCA, Makuhari, Japan, (2010), pp. 2794\u20132797","DOI":"10.21437\/Interspeech.2010-739"},{"key":"290_CR25","unstructured":"F. Eyben, M. W\u00f6llmer, B. Schuller, in Proceedings of the 18th ACM international conference on Multimedia, Opensmile: the munich versatile and fast open-source audio feature extractor, ACM, New York, United States (2010), pp. 1459\u20131462"},{"key":"290_CR26","unstructured":"J. Joy, A. Kannan, S. Ram, S. Rama, Speech emotion recognition using neural network and MLP classifier, International Journal of Engineering Science and Computing, Pearl Media Publications PVT LTD, 10(4), pp. 25170\u201325172 (2020)"},{"key":"290_CR27","doi-asserted-by":"crossref","unstructured":"B.\u00a0McFee, C.\u00a0Raffel, D.\u00a0Liang, D.P. Ellis, M.\u00a0McVicar, E.\u00a0Battenberg, O.\u00a0Nieto, in Proceedings of the 14th python in science conference, vol.\u00a08, librosa: audio and music signal analysis in python (Citeseer, 2015), pp. 18\u201325","DOI":"10.25080\/Majora-7b98e3ed-003"},{"key":"290_CR28","unstructured":"F.\u00a0Albu, A.\u00a0Mateescu, N.\u00a0Dumitriu, in International Conference on Microelectronics and Computer Science, Architecture selection for a multilayer feedforward network (Citeseer, 1997), pp. 131\u2013134"},{"issue":"1","key":"290_CR29","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1109\/TNN.2004.836197","volume":"16","author":"C Xiang","year":"2005","unstructured":"C. Xiang, S.Q. Ding, T.H. Lee, Geometrical interpretation and architecture selection of MLP. IEEE Trans. Neural Netw. 16(1), 84\u201396 (2005)","journal-title":"IEEE Trans. Neural Netw."},{"key":"290_CR30","doi-asserted-by":"crossref","unstructured":"T.\u00a0Andersen, T.\u00a0Martinez, in IJCNN\u201999. International Joint Conference on Neural Networks. Proceedings (Cat. No. 99CH36339), vol.\u00a03, Cross validation and MLP architecture selection (IEEE, 1999), pp. 1614\u20131619","DOI":"10.1109\/IJCNN.1999.832613"},{"key":"290_CR31","unstructured":"G.\u00a0Roffo, Feature selection library (matlab toolbox). arXiv preprint arXiv:1607.01327 (2016)"},{"key":"290_CR32","unstructured":"S. Russell, P. Norvig, Artificial intelligence: a modern approach, Prentice Hall, London, United Kingdom (2003)"},{"issue":"8","key":"290_CR33","doi-asserted-by":"publisher","first-page":"1226","DOI":"10.1109\/TPAMI.2005.159","volume":"27","author":"H Peng","year":"2005","unstructured":"H. Peng, F. Long, C. Ding, Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy. IEEE Trans. Pattern Anal. Mach. Intell. 27(8), 1226\u20131238 (2005)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"290_CR34","unstructured":"H. Liu, H. Motoda, in Chapman & Hall\/CRC, Computational methods of feature selection (Chapman & Hall\/CRC data mining and knowledge discovery series), Chapman and Hall\/CRC, Florida, United States (2007)"},{"key":"290_CR35","doi-asserted-by":"crossref","unstructured":"H.\u00a0Zeng, Y.m. Cheung, Feature selection and kernel learning for local learning-based clustering. IEEE Trans. Pattern Anal. Mach. Intell. 33(8), 1532\u20131547 (2010)","DOI":"10.1109\/TPAMI.2010.215"},{"key":"290_CR36","unstructured":"Y. Yang, H.T. Shen, Z. Ma, Z. Huang, X. Zhou, in Twenty-second international joint conference on artificial intelligence, L2, 1-norm regularized discriminative feature selection for unsupervised, AAAI Press, Washington, United States (2011)"},{"key":"290_CR37","doi-asserted-by":"crossref","unstructured":"F. Burkhardt, A. Paeschke, M. Rolfes, W.F. Sendlmeier, B. Weiss, et al., in INTERSPEECH, vol. 5, A database of German emotional speech. ISCA, Lisbon, Portugal (2005), pp. 1517\u20131520","DOI":"10.21437\/Interspeech.2005-446"},{"key":"290_CR38","doi-asserted-by":"crossref","unstructured":"Y.\u00a0Fu, X.\u00a0Yuan, in 2020 IEEE 23rd International Conference on Computational Science and Engineering (CSE), Composite feature extraction for speech emotion recognition (IEEE, 2020), pp. 72\u201377","DOI":"10.1109\/CSE50738.2020.00018"},{"issue":"5","key":"290_CR39","doi-asserted-by":"publisher","first-page":"587","DOI":"10.1049\/iet-spr.2016.0336","volume":"11","author":"WA Jassim","year":"2017","unstructured":"W.A. Jassim, R. Paramesran, N. Harte, Speech emotion classification using combined neurogram and INTERSPEECH 2010 paralinguistic challenge features. IET Signal Proc. 11(5), 587\u2013595 (2017)","journal-title":"IET Signal Proc."}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-023-00290-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-023-00290-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-023-00290-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,20]],"date-time":"2024-10-20T07:21:24Z","timestamp":1729408884000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-023-00290-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,15]]},"references-count":39,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["290"],"URL":"https:\/\/doi.org\/10.1186\/s13636-023-00290-x","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,15]]},"assertion":[{"value":"27 February 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 April 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 May 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"23"}}