{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,28]],"date-time":"2026-02-28T04:30:23Z","timestamp":1772253023664,"version":"3.50.1"},"reference-count":16,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2018,11,7]],"date-time":"2018-11-07T00:00:00Z","timestamp":1541548800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>The competition of speech recognition technology related to smartphones is now getting into full swing with the widespread internet of thing (IoT) devices. For robust speech recognition, it is necessary to detect speech signals in various acoustic environments. Speech\/music classification that facilitates optimized signal processing from classification results has been extensively adapted as an essential part of various electronics applications, such as multi-rate audio codecs, automatic speech recognition, and multimedia document indexing. In this paper, we propose a new technique to improve robustness of a speech\/music classifier for an enhanced voice service (EVS) codec adopted as a voice-over-LTE (VoLTE) speech codec using long short-term memory (LSTM). For effective speech\/music classification, feature vectors implemented with the LSTM are chosen from the features of the EVS. To overcome the diversity of music data, a large scale of data is used for learning. Experiments show that LSTM-based speech\/music classification provides better results than the conventional EVS speech\/music classification algorithm in various conditions and types of speech\/music data, especially at lower signal-to-noise ratio (SNR) than conventional EVS algorithm.<\/jats:p>","DOI":"10.3390\/sym10110605","type":"journal-article","created":{"date-parts":[[2018,11,7]],"date-time":"2018-11-07T10:32:07Z","timestamp":1541586727000},"page":"605","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Improvement of Speech\/Music Classification for 3GPP EVS Based on LSTM"],"prefix":"10.3390","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1833-2041","authenticated-orcid":false,"given":"Sang-Ick","family":"Kang","sequence":"first","affiliation":[{"name":"Department of Electronic Engineering, Inha University, Incheon 22212, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sangmin","family":"Lee","sequence":"additional","affiliation":[{"name":"Department of Electronic Engineering, Inha University, Incheon 22212, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2018,11,7]]},"reference":[{"key":"ref_1","unstructured":"Gao, Y., Shlomot, E., Benyassine, A., Thyssen, J., Su, H.-Y., and Murgia, C. (2001, January 7\u201311). The SMV algorithm selected by TIA and 3GPP2 for CDMA application. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, Salt Lake City, UT, USA."},{"key":"ref_2","unstructured":"3GPP2 Spec (2018, November 06). Source-Controlled Variable-Rate Multimedia Wideband Speech Codec (VMR-WB), Service Option 62 and 63 for Spread Spectrum Systems, 3GPP2-C.S0052-A, v.1.0. Available online: https:\/\/www.3gpp2.org\/Public_html\/Specs\/C.S0052-0_v1.0_040617.pdf."},{"key":"ref_3","unstructured":"3GPP Spec (2018, November 06). Codec for Enhanced Voice Services (EVS), Detailed Algorithm Description, TS 26.445, v.12.0.0. Available online: https:\/\/www.etsi.org\/deliver\/etsi_ts\/126400_126499\/126445\/12.00.00_60\/ts_126445v120000p.pdf."},{"key":"ref_4","unstructured":"Saunders, J. (1996, January 7\u201310). Real-time discrimination of broadcast speech\/music. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, Atlanta, GA, USA."},{"key":"ref_5","unstructured":"Fuchs, G. (September, January 31). A robust speech\/music discriminator for switched audio coding. Proceedings of the Signal Processing Conference (EUSIPCO), Nice, France."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"888","DOI":"10.1587\/transinf.E95.D.888","article-title":"Improvement of SVM-Based Speech\/Music Classification Using Adaptive Kernel Technique","volume":"9","author":"Lim","year":"2012","journal-title":"IEICE Trans. Inf. Syst."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"5375","DOI":"10.1007\/s11042-014-1859-8","article-title":"Efficient implementation techniques of an svm-based speech\/music classifier in smv","volume":"74","author":"Lim","year":"2015","journal-title":"Multimed. Tools Appl."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1109\/LSP.2007.911184","article-title":"Analysis and Improvement of Speech\/Music Classification for 3GPP2 SMV Based on GMM","volume":"15","author":"Song","year":"2008","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"661","DOI":"10.1587\/transfun.E97.A.661","article-title":"Speech\/music classification enhancement for 3gpp2 smv codec based on deep belief networks","volume":"97","author":"Song","year":"2014","journal-title":"IEICE Trans. Fundam. Electron. Commun. Comput. Sci."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Malenovsky, V., Vaillancourt, T., Zhe, W., Choo, K., and Atti, V. (2015, January 19\u201324). Two-stage speech\/music classifier with decision smoothing and sharpening in the EVS codec. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, Brisbane, Australia.","DOI":"10.1109\/ICASSP.2015.7179067"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_12","unstructured":"Jozefowicz, R., Zaremba, W., and Sutskever, I. (2015, January 6\u201311). An empirical exploration of recurrent network architectures. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1109\/72.279181","article-title":"Learning long-term dependencies with gradient descent is difficult","volume":"5","author":"Bengio","year":"1994","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Karneb\u00e4ck, S. (2001, January 3\u20137). Discrimination between speech and music based on a low frequency modulation feature. Proceedings of the Eurospeech, Aalborg, Denmark.","DOI":"10.21437\/Eurospeech.2001-447"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1111\/j.2517-6161.1977.tb01600.x","article-title":"Maximum likelihood from incomplete data via the EM algorithm","volume":"39","author":"Dempster","year":"1977","journal-title":"J. R. Stat. Soc. Ser. B Methodol."},{"key":"ref_16","unstructured":"Fisher, W.M., Doddington, G.R., and Goudie-Marshall, K.M. (1986, January 19\u201320). The DARPA speech recognition research database: Specification and status. Proceedings of the DARPA Workshop Speech Recognition, Palo Alto, CA, USA."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/10\/11\/605\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T15:28:24Z","timestamp":1760196504000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/10\/11\/605"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,11,7]]},"references-count":16,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2018,11]]}},"alternative-id":["sym10110605"],"URL":"https:\/\/doi.org\/10.3390\/sym10110605","relation":{"has-preprint":[{"id-type":"doi","id":"10.20944\/preprints201811.0126.v1","asserted-by":"object"}]},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,11,7]]}}}