{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,7]],"date-time":"2026-05-07T15:50:11Z","timestamp":1778169011470,"version":"3.51.4"},"reference-count":37,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,2,5]],"date-time":"2021-02-05T00:00:00Z","timestamp":1612483200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,2,5]],"date-time":"2021-02-05T00:00:00Z","timestamp":1612483200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100011958","name":"Danmarks Frie Forskningsfond","doi-asserted-by":"publisher","award":["DFF 4184-00056"],"award-info":[{"award-number":["DFF 4184-00056"]}],"id":[{"id":"10.13039\/501100011958","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The presence of degradations in speech signals, which causes acoustic mismatch between training and operating conditions, deteriorates the performance of many speech-based systems. A variety of enhancement techniques have been developed to compensate the acoustic mismatch in speech-based applications. To apply these signal enhancement techniques, however, it is necessary to know prior information about the presence and the type of degradations in speech signals. In this paper, we propose a new convolutional neural network (CNN)-based approach to automatically identify the major types of degradations commonly encountered in speech-based applications, namely additive noise, nonlinear distortion, and reverberation. In this approach, a set of parallel CNNs, each detecting a certain degradation type, is applied to the log-mel spectrogram of audio signals. Experimental results using two different speech types, namely pathological voice and normal running speech, show the effectiveness of the proposed method in detecting the presence and the type of degradations in speech signals which outperforms the state-of-the-art method. Using the score weighted class activation mapping, we provide a visual analysis of how the network makes decision for identifying different types of degradation in speech signals by highlighting the regions of the log-mel spectrogram which are more influential to the target degradation.<\/jats:p>","DOI":"10.1186\/s13636-021-00198-4","type":"journal-article","created":{"date-parts":[[2021,2,5]],"date-time":"2021-02-05T13:08:27Z","timestamp":1612530507000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["A CNN-based approach to identification of degradations in speech signals"],"prefix":"10.1186","volume":"2021","author":[{"given":"Yuki","family":"Saishu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6882-4618","authenticated-orcid":false,"given":"Amir Hossein","family":"Poorjam","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mads Gr\u00e6sb\u00f8ll","family":"Christensen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,2,5]]},"reference":[{"key":"198_CR1","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1017\/ATSIP.2016.16","volume":"5","author":"S. Ghai","year":"2016","unstructured":"S. Ghai, R. Sinha, Adaptive feature truncation to address acoustic mismatch in automatic recognition of children\u2019s speech. APSIPA Trans. Signal Inf. Process.5:, 1\u201313 (2016).","journal-title":"APSIPA Trans. Signal Inf. Process."},{"key":"198_CR2","doi-asserted-by":"publisher","first-page":"95","DOI":"10.1016\/j.forsciint.2004.09.078","volume":"146","author":"A. Alexander","year":"2004","unstructured":"A. Alexander, F. Botti, D. Dessimoz, A. Drygajlo, The effect of mismatched recording conditions on human and automatic speaker recognition in forensic applications. Forensic Sci. Int.146:, 95\u201399 (2004).","journal-title":"Forensic Sci. Int."},{"key":"198_CR3","unstructured":"V. Mitra, A. Tsiartas, E. Shriberg, in International Conference on Acoustics, Speech and Signal Processing (ICASSP). Noise and reverberation effects on depression detection from speech, (2016), pp. 5795\u20135799."},{"key":"198_CR4","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.specom.2020.12.007","volume":"127","author":"A. H. Poorjam","year":"2021","unstructured":"A. H. Poorjam, M. S. Kavalekalam, L. Shi, J. P. Raykov, J. R. Jensen, M. A. Little, M. G. Christensen, Automatic quality control and enhancement for voice-based remote Parkinson\u2019s disease detection. Speech Commun.127:, 1\u201316 (2021).","journal-title":"Speech Commun."},{"key":"198_CR5","unstructured":"M. Fakhry, A. H. Poorjam, M. G. Christensen, in European Signal Processing Conference (EUSIPCO). Speech enhancement by classification of noisy signals decomposed using NMF and Wiener filtering, (2018), pp. 16\u201320."},{"issue":"4","key":"198_CR6","doi-asserted-by":"publisher","first-page":"353","DOI":"10.1007\/s10772-014-9233-9","volume":"17","author":"J. H. L. Hansen","year":"2014","unstructured":"J. H. L. Hansen, A. Kumar, P. Angkititrakul, Environment mismatch compensation using average eigenspace-based methods for robust speech recognition. Int. J. Speech Technol.17(4), 353\u2013364 (2014).","journal-title":"Int. J. Speech Technol."},{"key":"198_CR7","unstructured":"B. W. Gillespie, H. S. Malvar, D. A. F. Florencio, in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6. Speech dereverberation via maximum-kurtosis subband adaptive filtering, (2001), pp. 3701\u20133704."},{"issue":"6","key":"198_CR8","doi-asserted-by":"publisher","first-page":"982","DOI":"10.1109\/TASLP.2015.2416653","volume":"23","author":"H. Kun","year":"2015","unstructured":"H. Kun, W. Yuxuan, W. DeLiang, S. W. William, M. Ivo, Z. Tao, Learning spectral mapping for speech dereverberation and denoising. IEEE Trans. Audio Speech Lang. Process.23(6), 982\u2013992 (2015).","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"198_CR9","unstructured":"J. S. Abel, in IEEE International Conference On Acoustics, Speech, and Signal Processing. Restoring a clipped signalIEEE Computer Society, (1991), pp. 1745\u20131748."},{"issue":"12","key":"198_CR10","doi-asserted-by":"publisher","first-page":"2627","DOI":"10.1109\/TASL.2013.2281570","volume":"21","author":"B. Defraene","year":"2013","unstructured":"B. Defraene, N. Mansour, S. De Hertogh, T. Van Waterschoot, M. Diehl, M. Moonen, Declipping of audio signals using perceptual compressed sensing. IEEE Trans. Audio Speech Lang. Process.21(12), 2627\u20132637 (2013).","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"issue":"7","key":"198_CR11","doi-asserted-by":"publisher","first-page":"1492","DOI":"10.1109\/TASLP.2017.2696307","volume":"25","author":"D. S. Williamson","year":"2017","unstructured":"D. S. Williamson, D. Wang, Time-frequency masking in the complex domain for speech dereverberation and denoising. IEEE\/ACM Trans. Audio Speech Lang. Process.25(7), 1492\u20131501 (2017).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"198_CR12","doi-asserted-by":"publisher","first-page":"740","DOI":"10.1109\/TASLP.2020.2966869","volume":"28","author":"T. Dietzen","year":"2020","unstructured":"T. Dietzen, S. Doclo, M. Moonen, T. van Waterschoot, Integrated sidelobe cancellation and linear prediction Kalman filter for joint multi-microphone speech dereverberation, interfering speech cancellation, and noise reduction. IEEE\/ACM Trans. Audio Speech Lang. Process.28:, 740\u2013754 (2020).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"4","key":"198_CR13","doi-asserted-by":"publisher","first-page":"680","DOI":"10.1109\/TASLP.2016.2518804","volume":"24","author":"I. Kodrasi","year":"2016","unstructured":"I. Kodrasi, S. Doclo, Joint dereverberation and noise reduction based on acoustic multi-channel equalization. IEEE\/ACM Trans. Audio Speech Lang. Process.24(4), 680\u2013693 (2016).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"198_CR14","unstructured":"L. Ma, D. J. Smith, B. P. Milner, in Eighth European Conference on Speech Communication and Technology. Context awareness using environmental noise classification, (2003), pp. 1\u20134."},{"issue":"2","key":"198_CR15","doi-asserted-by":"publisher","first-page":"1112","DOI":"10.1121\/1.4812273","volume":"134","author":"J. M. Desmond","year":"2013","unstructured":"J. M. Desmond, L. M. Collins, C. S. Throckmorton, Using channel-specific statistical models to detect reverberation in cochlear implant stimuli. J. Acoust. Soc. Am.134(2), 1112\u20131120 (2013).","journal-title":"J. Acoust. Soc. Am."},{"issue":"4","key":"198_CR16","first-page":"91","volume":"92","author":"S. V. Aleinik","year":"2014","unstructured":"S. V. Aleinik, M. Y. Nikolaevich, S. A. Vladimirovich, Detection of clipped fragments in acoustic signals. J. Sci. Tech. Inf. Technol. Mech. Opt.92(4), 91\u201397 (2014).","journal-title":"J. Sci. Tech. Inf. Technol. Mech. Opt."},{"key":"198_CR17","doi-asserted-by":"publisher","first-page":"218","DOI":"10.1016\/j.specom.2015.06.008","volume":"72","author":"F. Bie","year":"2015","unstructured":"F. Bie, D. Wang, J. Wang, T. F. Zheng, Detection and reconstruction of clipped speech for speaker recognition. Speech Commun.72:, 218\u2013231 (2015).","journal-title":"Speech Commun."},{"key":"198_CR18","unstructured":"A. H. Poorjam, J. R. Jensen, M. A. Little, M. G. Christensen, in Proceedings of the Annual Conference of the International Speech Communication Association, InterSpeech. Dominant distortion classification for pre-processing of vowels in remote biomedical voice analysis, (2017), pp. 289\u2013293."},{"key":"198_CR19","unstructured":"A. H. Poorjam, M. A. Little, J. R. Jensen, M. G. Christensen, in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). A parametric approach for classification of distortions in pathological voices, (2018), pp. 286\u2013290."},{"key":"198_CR20","unstructured":"H. Wang, M. Du, F. Yang, Z. Zhang, Score-CAM: improved visual explanations via score-weighted class activation mapping. 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 111\u2013119 (2020)."},{"key":"198_CR21","unstructured":"K. Simonyan, A. Zisserman, in 3rd International Conference on Learning Representations, ICLR 2015 - Conference Track Proceedings. Very deep convolutional networks for large-scale image recognition, (2015), pp. 1\u201314."},{"key":"198_CR22","unstructured":"J. Bjorck, C. Gomes, B. Selman, K. Q. Weinberger, in Advances in Neural Information Processing Systems (NeurIPS). Understanding batch normalization, (2018), pp. 7694\u20137705."},{"key":"198_CR23","volume-title":"An Introduction to the Psychology of Hearing","author":"B. C. Moore","year":"2012","unstructured":"B. C. Moore, An Introduction to the Psychology of Hearing (Emerald, Bingley, 2012)."},{"key":"198_CR24","doi-asserted-by":"crossref","unstructured":"K. Choi, G. Fazekas, M. Sandler, K. Cho, in 2018 26th European Signal Processing Conference (EUSIPCO). A comparison of audiosignal preprocessing methods for deep neural networks on musictagging (Rome, 2018), pp. 1870\u20131874.","DOI":"10.23919\/EUSIPCO.2018.8553106"},{"key":"198_CR25","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/sdata.2016.11","volume":"3","author":"B. M. Bot","year":"2016","unstructured":"B. M. Bot, C. Suver, E. C. Neto, M. Kellen, A. Klein, C. Bare, M. Doerr, A. Pratap, J. Wilbanks, E. R. Dorsey, S. H. Friend, A. D. Trister, The mPower study, Parkinson disease mobile data collected using ResearchKit. Sci Data. 3:, 1\u20139 (2016).","journal-title":"Sci Data"},{"issue":"3","key":"198_CR26","doi-asserted-by":"publisher","first-page":"131","DOI":"10.1155\/1999\/327643","volume":"11","author":"A. K. Ho","year":"1998","unstructured":"A. K. Ho, R. Iansek, C. Marigliani, J. L. Bradshaw, S. Gates, Speech impairment in a large sample of patients with Parkinson\u2019s disease. Behav. Neurol.11(3), 131\u2013137 (1998).","journal-title":"Behav. Neurol."},{"issue":"3","key":"198_CR27","doi-asserted-by":"publisher","first-page":"1148","DOI":"10.1121\/1.424266","volume":"104","author":"I. R. Titze","year":"1998","unstructured":"I. R. Titze, D. W. Martin, Principles of voice production. Acoust. Soc. Am.104(3), 1148 (1998).","journal-title":"Acoust. Soc. Am."},{"key":"198_CR28","unstructured":"C. Valentini-Botinhao, Noisy speech database for training speech enhancement algorithms and tts models [online] (2017). Available: http:\/\/dx.doi.org\/10.7488\/ds\/2117."},{"key":"198_CR29","unstructured":"M. Jeub, M. Sch\u00e4fer, H. Kr\u00fcger, C. Nelke, C. Beaugeant, P. Vary, in Proc. Int. Congress on Acoustics (ICA), Sydney, Australia. Do we need dereverberation for hand-held telephony? (2010), pp. 1\u20137."},{"issue":"6","key":"198_CR30","doi-asserted-by":"publisher","first-page":"1187","DOI":"10.1121\/1.1939454","volume":"37","author":"M. R. Schroeder","year":"1965","unstructured":"M. R. Schroeder, New method of measuring reverberation time. J. Acoust. Soc. Am.37(6), 1187\u20131188 (1965).","journal-title":"J. Acoust. Soc. Am."},{"key":"198_CR31","unstructured":"M. R. Schroeder, B. S. Atal, in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 10. Code-excited linear prediction (CELP): high-quality speech at very low bit rates, (1985), pp. 937\u2013940."},{"key":"198_CR32","unstructured":"C. Valentini-Botinhao, Noisy reverberant speech database for training speech enhancement algorithms and tts models [online] (2017). Available: https:\/\/dx.doi.org\/10.7488\/ds\/2139."},{"issue":"5","key":"198_CR33","doi-asserted-by":"publisher","first-page":"3591","DOI":"10.1121\/1.4806631","volume":"133","author":"J. Thiemann","year":"2013","unstructured":"J. Thiemann, N. Ito, E. Vincent, The diverse environments multi-channel acoustic noise database: a database of multichannel environmental noise recordings. J. Acoust. Soc. Am.133(5), 3591\u20133591 (2013).","journal-title":"J. Acoust. Soc. Am."},{"key":"198_CR34","unstructured":"S. L. Smith, P. J. Kindermans, C. Ying, Q. V. Le, in 6th International Conference on Learning Representations. Don\u2019t decay the learning rate, increase the batch size, (2018), pp. 1\u201311."},{"issue":"1","key":"198_CR35","doi-asserted-by":"publisher","first-page":"68","DOI":"10.1145\/3359786","volume":"63","author":"M. Du","year":"2019","unstructured":"M. Du, N. Liu, X. Hu, Techniques for interpretable machine learning. Commun. ACM. 63(1), 68\u201377 (2019).","journal-title":"Commun. ACM"},{"issue":"4","key":"198_CR36","doi-asserted-by":"publisher","first-page":"312","DOI":"10.1016\/j.specom.2007.10.005","volume":"50","author":"X. Lu","year":"2008","unstructured":"X. Lu, J. Dang, An investigation of dependencies between frequency components and speaker characteristics for text-independent speaker identification. Speech Commun.50(4), 312\u2013322 (2008).","journal-title":"Speech Commun."},{"key":"198_CR37","volume-title":"Techniques in speech acoustics","author":"J. Harrington","year":"2012","unstructured":"J. Harrington, S. Cassidy, Techniques in speech acoustics, vol. 8 (Springer Science & Business Media, Netherlands, 2012)."}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-021-00198-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/s13636-021-00198-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-021-00198-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,2,5]],"date-time":"2021-02-05T13:22:12Z","timestamp":1612531332000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-021-00198-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,2,5]]},"references-count":37,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["198"],"URL":"https:\/\/doi.org\/10.1186\/s13636-021-00198-4","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,2,5]]},"assertion":[{"value":"12 October 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 January 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 February 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare that they have no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"9"}}