{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,17]],"date-time":"2026-01-17T19:35:30Z","timestamp":1768678530506,"version":"3.49.0"},"reference-count":58,"publisher":"Wiley","issue":"1","license":[{"start":{"date-parts":[[2024,10,18]],"date-time":"2024-10-18T00:00:00Z","timestamp":1729209600000},"content-version":"vor","delay-in-days":291,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2023YFC3304903"],"award-info":[{"award-number":["2023YFC3304903"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2023YFF0905404"],"award-info":[{"award-number":["2023YFF0905404"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["International Journal of Intelligent Systems"],"published-print":{"date-parts":[[2024,1]]},"abstract":"<jats:p>Speech emotion recognition plays a crucial role in analyzing psychological disorders, behavioral decision\u2010making, and human\u2010machine interaction applications. However, the majority of current methods for speech emotion recognition heavily rely on data\u2010driven approaches, and the scarcity of emotion speech datasets limits the progress in research and development of emotion analysis and recognition. To address this issue, this study introduces a new English speech dataset specifically designed for emotion analysis and recognition. This dataset consists of 5503 voices from over 60 English speakers in different emotional states. Furthermore, to enhance emotion analysis and recognition, fast Fourier transform (FFT), short\u2010time Fourier transform (STFT), mel\u2010frequency cepstral coefficients (MFCCs), and continuous wavelet transform (CWT) are employed for feature extraction from the speech data. Utilizing these algorithms, the spectrum images of the speeches are obtained, forming four datasets consisting of different speech feature images. Furthermore, to evaluate the dataset, 16 classification models and 19 detection algorithms are selected. The experimental results demonstrate that the majority of classification and detection models achieve exceptionally high recognition accuracy on this dataset, confirming its effectiveness and utility. The dataset proves to be valuable in advancing research and development in the field of emotion recognition.<\/jats:p>","DOI":"10.1155\/2024\/5410080","type":"journal-article","created":{"date-parts":[[2024,10,18]],"date-time":"2024-10-18T13:50:21Z","timestamp":1729259421000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["E\u2010Speech: Development of a Dataset for Speech Emotion Recognition and Analysis"],"prefix":"10.1155","volume":"2024","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2855-0277","authenticated-orcid":false,"given":"Wenjin","family":"Liu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-9385-742X","authenticated-orcid":false,"given":"Jiaqi","family":"Shi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-8668-1669","authenticated-orcid":false,"given":"Shudong","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-3697-9983","authenticated-orcid":false,"given":"Lijuan","family":"Zhou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-9434-1478","authenticated-orcid":false,"given":"Haoming","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2024,10,18]]},"reference":[{"key":"e_1_2_12_1_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.04.028"},{"key":"e_1_2_12_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2022.03.002"},{"key":"e_1_2_12_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.seta.2021.101858"},{"key":"e_1_2_12_4_2","doi-asserted-by":"crossref","unstructured":"JanbakhshiP.andKodrasiI. Experimental Investigation on STFT Phase Representations for Deep Learning-Based Dysarthric Speech Detection ICASSP 2022-2022 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP) May 2022 Singapore IEEE 6477\u20136481.","DOI":"10.1109\/ICASSP43922.2022.9747205"},{"key":"e_1_2_12_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11760-023-02716-7"},{"key":"e_1_2_12_6_2","article-title":"A Multi-Task Learning Speech Synthesis Optimization Method Based on CWT: A Case Study of Tacotron2","volume":"4","author":"Hu G.","year":"2024","journal-title":"EURASIP Journal on Applied Signal Processing"},{"key":"e_1_2_12_7_2","article-title":"Imagenet Classification With Deep Convolutional Neural Networks","volume":"25","author":"Krizhevsky A.","year":"2012","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_12_8_2","unstructured":"SimonyanK.andZissermanA. Very Deep Convolutional Networks for Large-Scale Image Recognition 2014 https:\/\/arxiv.org\/abs\/1409.1556."},{"key":"e_1_2_12_9_2","doi-asserted-by":"crossref","unstructured":"HeK. ZhangX. RenS. andSunJ. Deep Residual Learning for Image Recognition Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition June 2016 Las Vegas NV 770\u2013778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_12_10_2","doi-asserted-by":"crossref","unstructured":"HuangG. LiuZ. MaatenL. V. D. andWeinbergerK. Q. Densely Connected Convolutional Networks Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition July 2017 Honolulu HI 4700\u20134708.","DOI":"10.1109\/CVPR.2017.243"},{"key":"e_1_2_12_11_2","doi-asserted-by":"crossref","unstructured":"XieS. GirshickR. Doll\u00e1rP. TuZ. andHeK. Aggregated Residual Transformations for Deep Neural Networks Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition July 2017 Honolulu HI 1492\u20131500.","DOI":"10.1109\/CVPR.2017.634"},{"key":"e_1_2_12_12_2","doi-asserted-by":"crossref","unstructured":"HuJ. ShenL. andSunG. Squeeze-and-Excitation Networks Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition June 2018 Salt Lake City UT 7132\u20137141.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_2_12_13_2","doi-asserted-by":"crossref","unstructured":"SandlerM. HowardA. ZhuM. ZhmoginovA. andChenL.-C. Mobilenetv2: Inverted Residuals and Linear Bottlenecks Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition June 2018 Salt Lake City UT 4510\u20134520.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_2_12_14_2","doi-asserted-by":"crossref","unstructured":"ZhangX. ZhouX. LinM. andSunJ. Shufflenet: An Extremely Efficient Convolutional Neural Network for Mobile Devices Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition June 2018 Salt Lake City UT 6848\u20136856.","DOI":"10.1109\/CVPR.2018.00716"},{"key":"e_1_2_12_15_2","doi-asserted-by":"crossref","unstructured":"MaN. ZhangX. ZhengH.-T. andSunJ. Shufflenet V2: Practical Guidelines for Efficient CNN Architecture Design Proceedings of the European Conference on Computer Vision June 2018 Munich Germany ECCV 116\u2013131.","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"e_1_2_12_16_2","unstructured":"DosovitskiyA. BeyerL. KolesnikovA.et al. An Image is Worth 16\u2009\u00d7\u200916 Words: Transformers for Image Recognition at Scale International Conference on Learning Representations May 2021 Vienna Austria."},{"key":"e_1_2_12_17_2","doi-asserted-by":"crossref","unstructured":"LiuZ. LinY. CaoY.et al. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows Proceedings of the IEEE\/CVF International Conference on Computer Vision October 2021 Montreal Canada 10012\u201310022.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_2_12_18_2","doi-asserted-by":"crossref","unstructured":"LinT.-Y. GoyalP. GirshickR. HeK. andDoll\u00e1rP. Focal Loss for Dense Object Detection Proceedings of the IEEE International Conference on Computer Vision June 2017 Cambridge MA 2980\u20132988.","DOI":"10.1109\/ICCV.2017.324"},{"key":"e_1_2_12_19_2","doi-asserted-by":"crossref","unstructured":"LiuW. AnguelovD. ErhanD.et al. SSD: Single Shot Multibox Detector Computer Vision--ECCV 2016: 14th European Conference October 2016 Amsterdam Netherlands Springer 21\u201337.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"e_1_2_12_20_2","doi-asserted-by":"crossref","unstructured":"TanM. PangR. andLeQ. V. Efficientdet: Scalable and Efficient Object Detection Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition June 2020 Seattle WA 10781\u201310790.","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"e_1_2_12_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/tpami.2016.2577031"},{"key":"e_1_2_12_22_2","unstructured":"BochkovskiyA. WangC.-Y. andLiaoH.-Y. M. Yolov4: Optimal Speed and Accuracy of Object Detection 2020 https:\/\/arxiv.org\/abs\/2004.10934."},{"key":"e_1_2_12_23_2","doi-asserted-by":"crossref","unstructured":"WangC.-Y. BochkovskiyA. andLiaoH.-Y. M. YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition June 2023 Vancouver Canada 7464\u20137475.","DOI":"10.1109\/CVPR52729.2023.00721"},{"key":"e_1_2_12_24_2","first-page":"1517","article-title":"A Database of German Emotional Speech","volume":"5","author":"Burkhardt F.","year":"2005","journal-title":"Interspeech"},{"key":"e_1_2_12_25_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0196391"},{"key":"e_1_2_12_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-008-9076-6"},{"key":"e_1_2_12_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/taffc.2014.2336244"},{"key":"e_1_2_12_28_2","doi-asserted-by":"crossref","unstructured":"MartinO. KotsiaI. MacqB. andPitasI. The eNTERFACE\u201905 Audio-Visual Emotion Database 22nd International Conference on Data Engineering Workshops (ICDEW\u203206) April 2006 Atlanta GA IEEE.","DOI":"10.1109\/ICDEW.2006.145"},{"key":"e_1_2_12_29_2","first-page":"53","article-title":"Speaker-Dependent Audio-Visual Emotion Recognition","author":"Haq S.","year":"2009","journal-title":"Auditory-Visual Speech Processing"},{"key":"e_1_2_12_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-019-09450-y"},{"key":"e_1_2_12_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/tsa.2004.838534"},{"key":"e_1_2_12_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/s0167-6393(03)00099-2"},{"key":"e_1_2_12_33_2","doi-asserted-by":"crossref","unstructured":"YunS.andYooC. D. Speech Emotion Recognition via a Max-Margin Framework Incorporating a Loss Function Based on the Watson and Tellegen\u2032s Emotion Model 2009 IEEE International Conference on Acoustics Speech and Signal Processing April 2009 Taipei Taiwan IEEE 4169\u20134172.","DOI":"10.1109\/ICASSP.2009.4960547"},{"key":"e_1_2_12_34_2","doi-asserted-by":"crossref","unstructured":"WangY. Research and Implementation of Emotional Feature Classification and Recognition in Speech Signal 2008 International Symposium on Intelligent Information Technology Application Workshops December 2008 Shanghai China IEEE 471\u2013474.","DOI":"10.1109\/IITA.Workshops.2008.219"},{"key":"e_1_2_12_35_2","doi-asserted-by":"crossref","unstructured":"IjimaY. TachibanaM. NoseT. andKobayashiT. Emotional Speech Recognition Based on Style Estimation and Adaptation With Multiple-Regression HMM 2009 IEEE International Conference on Acoustics Speech and Signal Processing April 2009 Taipei Taiwan IEEE 4157\u20134160.","DOI":"10.1109\/ICASSP.2009.4960544"},{"key":"e_1_2_12_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2024.109103"},{"key":"e_1_2_12_37_2","doi-asserted-by":"crossref","unstructured":"AtassiH.andEspositoA. A Speaker Independent Approach to the Classification of Emotional Vocal Expressions 2008 20th IEEE International Conference on Tools With Artificial Intelligence November 2008 Dayton OH IEEE 147\u2013152.","DOI":"10.1109\/ICTAI.2008.158"},{"key":"e_1_2_12_38_2","doi-asserted-by":"crossref","unstructured":"Kumar MishraH.andChandra SekharC. Variational Gaussian Mixture Models for Speech Emotion Recognition 2009 Seventh International Conference on Advances in Pattern Recognition February 2009 Kolkata India IEEE 183\u2013186.","DOI":"10.1109\/ICAPR.2009.89"},{"key":"e_1_2_12_39_2","unstructured":"GraciarenaM. ShribergE. StolckeA. EnosF. HirschbergJ. andKajarekarS. Combining Prosodic Lexical and Cepstral Systems for Deceptive Speech Detection 2006 IEEE International Conference on Acoustics Speech and Signal Processing May 2006 Toulouse France IEEE."},{"key":"e_1_2_12_40_2","doi-asserted-by":"crossref","unstructured":"SeehapochT.andWongthanavasuS. Speech Emotion Recognition Using Support Vector Machines 2013 5th International Conference on Knowledge and Smart Technology (KST) May 2013 Chonburi Thailand IEEE 86\u201391.","DOI":"10.1109\/KST.2013.6512793"},{"key":"e_1_2_12_41_2","doi-asserted-by":"publisher","DOI":"10.5120\/431-636"},{"key":"e_1_2_12_42_2","doi-asserted-by":"crossref","unstructured":"GrimmM. KroschelK. andNarayananS. Support Vector Regression for Automatic Recognition of Spontaneous Emotions in Speech 2007 IEEE International Conference on Acoustics Speech and Signal Processing April 2007 Honolulu HI IEEE.","DOI":"10.1109\/ICASSP.2007.367262"},{"key":"e_1_2_12_43_2","doi-asserted-by":"crossref","unstructured":"PaoT.-L. LiaoW.-Y. ChenY.-T. YehJ.-H. ChengY.-M. andChienC. S. Comparison of Several Classifiers for Emotion Recognition From Noisy Mandarin Speech Third International Conference on Intelligent Information Hiding and Multimedia Signal Processing (IIH-MSP 2007) November 2007 Kaohsiung Taiwan IEEE 23\u201326.","DOI":"10.1109\/IIHMSP.2007.4457484"},{"key":"e_1_2_12_44_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-024-52989-2"},{"key":"e_1_2_12_45_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-023-16036-y"},{"key":"e_1_2_12_46_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00034-023-02571-4"},{"key":"e_1_2_12_47_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.122905"},{"key":"e_1_2_12_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/tmm.2017.2766843"},{"key":"e_1_2_12_49_2","doi-asserted-by":"crossref","unstructured":"HeracleousP. MohammadY. andYoneyamaA. Deep Convolutional Neural Networks for Feature Extraction in Speech Emotion Recognition Human-Computer Interaction. Recognition and Interaction Technologies: Thematic Area HCI 2019 Held as Part of the 21st HCI International Conference HCII 2019 July 2019 Orlando FL Springer 117\u2013132.","DOI":"10.1007\/978-3-030-22643-5_9"},{"key":"e_1_2_12_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/taslp.2021.3076364"},{"key":"e_1_2_12_51_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.bspc.2018.08.035"},{"key":"e_1_2_12_52_2","doi-asserted-by":"crossref","unstructured":"WangJ. XueM. CulhaneR. DiaoE. DingJ. andTarokhV. Speech Emotion Recognition With Dual-Sequence LSTM Architecture ICASSP 2020-2020 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP) May 2020 Barcelona Spain IEEE 6474\u20136478.","DOI":"10.1109\/ICASSP40776.2020.9054629"},{"key":"e_1_2_12_53_2","article-title":"Transformer-Like Model With Linear Attention for Speech Emotion Recognition","volume":"37","author":"Du J.","year":"2021","journal-title":"Journal of Southeast University"},{"key":"e_1_2_12_54_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2024.103102"},{"key":"e_1_2_12_55_2","doi-asserted-by":"crossref","unstructured":"YiZ. ZhaoZ. ShenZ. andZhangT. Multimodal Fusion via Hypergraph Autoencoder and Contrastive Learning for Emotion Recognition in Conversation 2024 https:\/\/arxiv.org\/abs\/2408.00970.","DOI":"10.1145\/3664647.3681633"},{"key":"e_1_2_12_56_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.122946"},{"key":"e_1_2_12_57_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2024.128177"},{"key":"e_1_2_12_58_2","doi-asserted-by":"crossref","unstructured":"ZouH. SiY. ChenC. RajanD. andChngE. S. Speech Emotion Recognition With Co-Attention Based Multi-Level Acoustic Information ICASSP 2022-2022 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP) May 2022 Singapore IEEE 7367\u20137371.","DOI":"10.1109\/ICASSP43922.2022.9747095"}],"container-title":["International Journal of Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1155\/2024\/5410080","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,29]],"date-time":"2024-11-29T21:49:41Z","timestamp":1732916981000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1155\/2024\/5410080"}},"subtitle":[],"editor":[{"given":"Yu-an","family":"Tan","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2024,1]]},"references-count":58,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,1]]}},"alternative-id":["10.1155\/2024\/5410080"],"URL":"https:\/\/doi.org\/10.1155\/2024\/5410080","archive":["Portico"],"relation":{},"ISSN":["0884-8173","1098-111X"],"issn-type":[{"value":"0884-8173","type":"print"},{"value":"1098-111X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1]]},"assertion":[{"value":"2024-07-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-09-23","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-18","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"5410080"}}