{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T19:09:31Z","timestamp":1783537771587,"version":"3.55.0"},"reference-count":26,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[2020,5,6]],"date-time":"2020-05-06T00:00:00Z","timestamp":1588723200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"published-print":{"date-parts":[[2020,10,7]]},"abstract":"<jats:p>In speech emotion recognition, most emotional corpora generally have problems such as inconsistent sample length and imbalance of sample categories. Considering these problems, in this paper, a variable length input CRNN deep learning model based on Focal Loss is proposed for speech emotion recognition of anger, happiness, neutrality and sadness in IEMOCAP emotional corpus. In this model, Firstly, a variable-length strategy is introduced to input the speech spectra of the filled speech samples into CNN. Then the effective part of the input sequence is preserved and output by masking matrix and convolution layer. Thirdly, the effective output of input sequence is input into BiGRU network for learning. Finally, the focal loss is used for network training to control and adjust the contribution of various samples to the total loss. Compared with the traditional speech emotion recognition model, simulations show that our method can effectively improve the accuracy and performance of emotion recognition.<\/jats:p>","DOI":"10.3233\/jifs-191129","type":"journal-article","created":{"date-parts":[[2020,5,8]],"date-time":"2020-05-08T14:48:12Z","timestamp":1588949292000},"page":"2791-2796","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":6,"title":["Emotion recognition from speech using deep learning on spectrograms"],"prefix":"10.1177","volume":"39","author":[{"given":"Xingguang","family":"Li","sequence":"first","affiliation":[{"name":"Electronic Information Engineering, Changchun University of Science and Technology, Changchun, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenjun","family":"Song","sequence":"additional","affiliation":[{"name":"Electronic Information Engineering, Changchun University of Science and Technology, Changchun, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zonglin","family":"Liang","sequence":"additional","affiliation":[{"name":"Electronic Information Engineering, Changchun University of Science and Technology, Changchun, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2020,5,6]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3129340"},{"key":"e_1_3_2_3_2","doi-asserted-by":"crossref","unstructured":"SchmidtE.M. and KimY.E. Learning emotion-based acoustic features with deep belief networks [C] (2011) 65\u201368.","DOI":"10.1109\/ASPAA.2011.6082328"},{"key":"e_1_3_2_4_2","doi-asserted-by":"crossref","unstructured":"HanK. YuD. and TasheyI. Speech emotion recognition using deep neural network and extreme learning machine [C]\/\/Fifteenth annual conference of the international speech communication association 2014.","DOI":"10.21437\/Interspeech.2014-57"},{"key":"e_1_3_2_5_2","first-page":"827","article-title":"An experimental study of speech emotion recognition based on deep convolutional neural networks [C]\/\/2015 international conference on affective computing and intelligent interaction (ACII)","author":"Zheng W.Q.","year":"2015","unstructured":"ZhengW.Q., YuJ.S. and ZouY.X., An experimental study of speech emotion recognition based on deep convolutional neural networks [C]\/\/2015 international conference on affective computing and intelligent interaction (ACII), IEEE (2015), 827\u2013831.","journal-title":"IEEE"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1155\/2014\/749604"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2014.2360798"},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","unstructured":"SattA. RozenbergS. and HooryR. Efficient Emotion Recognition from Speech Using Deep Learning on Spectrograms [C]\/\/INTERSPEECH (2017) 1089\u20131093.","DOI":"10.21437\/Interspeech.2017-200"},{"key":"e_1_3_2_9_2","first-page":"1","article-title":"Speech emotion recognition using convolutional and recurrent neural networks [C]\/\/2016 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA)","author":"Lim W.","year":"2016","unstructured":"LimW., JangD. and LeeT., Speech emotion recognition using convolutional and recurrent neural networks [C]\/\/2016 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA), IEEE (2016), 1\u20134.","journal-title":"IEEE"},{"key":"e_1_3_2_10_2","first-page":"2227","article-title":"Automatic speech emotion recognition using recurrent neural networks with local attention [C]\/\/2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Mirsamadi S.","year":"2017","unstructured":"MirsamadiS., BarsoumE. and ZhangC., Automatic speech emotion recognition using recurrent neural networks with local attention [C]\/\/2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE (2017), 2227\u20132231.","journal-title":"IEEE"},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"AytarY. VondrickC. and TorralbaA. Soundnet: Learning sound representations from unlabeled video [C]\/\/Advances in neural information processing systems (2016) 892\u2013900.","DOI":"10.1109\/CVPR.2016.18"},{"key":"e_1_3_2_12_2","first-page":"4580","article-title":"\u201dConvolutional, long short-term memory,fully connected deep neural networks\u201d","author":"Sainath T.N.","year":"2015","unstructured":"SainathT.N., VinyalO. and SakH., \u201dConvolutional, long short-term memory,fully connected deep neural networks\u201d, in IEEE International Conference on Acoustics, Speech and Signal Processing, (2015), 4580\u20134584.","journal-title":"IEEE International Conference on Acoustics, Speech and Signal Processing"},{"key":"e_1_3_2_13_2","article-title":"\u201cHigh-level feature representation using recurrent neural network for speech emotion recognition,\u201d","author":"Lee J.","year":"2015","unstructured":"LeeJ. and TashevI., \u201cHigh-level feature representation using recurrent neural network for speech emotion recognition,\u201d in INTERSPEECH, 2015.","journal-title":"INTERSPEECH"},{"key":"e_1_3_2_14_2","article-title":"\u201dEmotion classification from noisy speech \u2013 a deep learning approach,\u201d","author":"Rana R.","year":"2016","unstructured":"RanaR., \u201dEmotion classification from noisy speech \u2013 a deep learning approach,\u201d arXiv preprint arXiv:1603.05901, 2016.","journal-title":"arXiv preprint arXiv:1603.05901"},{"key":"e_1_3_2_15_2","article-title":"\u201cEmotion recognition from speech with recurrent neural networks,\u201d","author":"Chernykh V.","year":"2017","unstructured":"ChernykhV., SterlingG. and PrihodkoP., \u201cEmotion recognition from speech with recurrent neural networks,\u201d arXiv preprint arXiv:1701.08071, 2017.","journal-title":"arXiv preprint arXiv:1701.08071"},{"key":"e_1_3_2_16_2","first-page":"1","article-title":"\u201dLearning the speech front-end with raw waveform cldnns\u201d","author":"Sainath T.N.","year":"2015","unstructured":"SainathT.N., WeissR.J., SeniorA.W., WilsonK.W. and VinyalsO., \u201dLearning the speech front-end with raw waveform cldnns\u201d, in INTERSPEECH, 2015, pp. 1\u20135.","journal-title":"INTERSPEECH"},{"key":"e_1_3_2_17_2","article-title":"Deep speech: Scaling up end-to-end speech recognition,\u201d","author":"Hunnun A.","year":"2014","unstructured":"HunnunA., CaseC., CasperJ., CatanzaroB., DiamosG., ElsenE., PrengerR., SatheeshS., SenguptaS. and Coates\u201dA., Deep speech: Scaling up end-to-end speech recognition,\u201d Computer Science, 2014.","journal-title":"Computer Science"},{"key":"e_1_3_2_18_2","article-title":"\u201dDeep speech 2: End-to-end speech recognition in english and mandarin,\u201d","author":"Amodei D.","year":"2015","unstructured":"AmodeiD., AnubhaiR., BattenbergE., CaseC., CasperJ., CatanzaroB., ChenJ., ChrzanowskiM., CoatesA. and DiamosG., \u201dDeep speech 2: End-to-end speech recognition in english and mandarin,\u201d, Computer Science 2015.","journal-title":"Computer Science"},{"key":"e_1_3_2_19_2","first-page":"4052","article-title":"\u201cDeep neural networks for small footprint text-dependent speaker verification,\u201d","author":"Variani E.","year":"2014","unstructured":"VarianiE., LeiX., McdermottE., MorenoI.L. and Gonzalez-DominguezJ., \u201cDeep neural networks for small footprint text-dependent speaker verification,\u201d in IEEE International Conference on Acoustics, Speech and Signal Processing, 2014, pp. 4052\u20134056.","journal-title":"IEEE International Conference on Acoustics, Speech and Signal Processing"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.bspc.2018.08.035"},{"key":"e_1_3_2_21_2","article-title":"\u201dAdieu features? end- to-end speech emotion recognition using a deep convolutional recurrent network,\u201d","author":"Trigeorrgis G.","year":"2016","unstructured":"TrigeorrgisG., RingevalF., BruecknerR., MarchiE., NicolaouM.A., SchullerB. and ZafeiriouS., \u201dAdieu features? end- to-end speech emotion recognition using a deep convolutional recurrent network,\u201d in IEEE International Conference on Acoustics, Speech and Signal Processing, 2016.","journal-title":"IEEE International Conference on Acoustics, Speech and Signal Processing"},{"key":"e_1_3_2_22_2","article-title":"You only look once: Unified, real-time object detection","author":"Redmon J.","year":"2016","unstructured":"RedmonJ., DivvalaS., GirshickR. and FarhadiA., You only look once: Unified, real-time object detection. In CVPR, 2016.","journal-title":"CVPR"},{"key":"e_1_3_2_23_2","article-title":"YOLO9000: Better, faster, stronger","author":"Redmon J.","year":"2017","unstructured":"RedmonJ. and FarhadiA., YOLO9000: Better, faster, stronger. In CVPR, 2017.","journal-title":"CVPR"},{"key":"e_1_3_2_24_2","article-title":"SSD: Single shot multibox detector","author":"Liu W.","year":"2016","unstructured":"LiuW., AnguelovD., ErhanD., SzegedyC. and ReedS., SSD: Single shot multibox detector. In ECCV, 2016.","journal-title":"ECCV"},{"key":"e_1_3_2_25_2","article-title":"DSSD: Deconvolutional single shot detector","author":"Fu C.-Y.","year":"2016","unstructured":"FuC.-Y., LiuW., RangaA., TyagiA. and BergA.C., DSSD: Deconvolutional single shot detector. arXiv: 1701. 06659, 2016.","journal-title":"arXiv: 1701. 06659"},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","unstructured":"LinT.Y. GoyalP. and GirshickR. et al. Focal loss for dense object detection [C]\/\/Proceedings of the IEEE International Conference on Computer Vision (2017) 2980\u20132988.","DOI":"10.1109\/ICCV.2017.324"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2018-2228"}],"container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/JIFS-191129","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.3233\/JIFS-191129","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/JIFS-191129","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:40:22Z","timestamp":1777455622000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.3233\/JIFS-191129"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,5,6]]},"references-count":26,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2020,10,7]]}},"alternative-id":["10.3233\/JIFS-191129"],"URL":"https:\/\/doi.org\/10.3233\/jifs-191129","relation":{},"ISSN":["1064-1246","1875-8967"],"issn-type":[{"value":"1064-1246","type":"print"},{"value":"1875-8967","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,5,6]]}}}