{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T15:10:14Z","timestamp":1784301014238,"version":"3.55.0"},"reference-count":60,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,7,2]],"date-time":"2021-07-02T00:00:00Z","timestamp":1625184000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,7,2]],"date-time":"2021-07-02T00:00:00Z","timestamp":1625184000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004871","name":"Technische Universit\u00e4t Braunschweig","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004871","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Estimating time-frequency domain masks for single-channel speech enhancement using deep learning methods has recently become a popular research field with promising results. In this paper, we propose a novel <jats:italic>components loss<\/jats:italic> (CL) for the training of neural networks for mask-based speech enhancement. During the training process, the proposed CL offers separate control over preservation of the speech component quality, suppression of the noise component, and preservation of a naturally sounding residual noise component. We illustrate the potential of the proposed CL by evaluating a standard convolutional neural network (CNN) for mask-based speech enhancement. The new CL is compared to several baseline losses, comprising the conventional mean squared error (MSE) loss w.r.t. speech spectral amplitudes or w.r.t. an ideal-ratio mask, auditory-related loss functions, such as the perceptual evaluation of speech quality (PESQ) loss and the perceptual weighting filter loss, and also the recently proposed SNR loss with two masks. Detailed analysis suggests that the proposed CL obtains a better or at least a more balanced performance across all employed instrumental quality metrics, including SNR improvement, speech component quality, enhanced total speech quality, and particularly also delivers a natural sounding residual noise component. For unseen noise types, we excel even perceptually motivated losses by an about 0.2 points higher PESQ score. The recently proposed so-called SNR loss with two masks not only requires a network with more parameters due to the two decoder heads, but also falls behind on PESQ and POLQA and particularly w.r.t. residual noise quality. Note that the proposed CL shows significantly more 1st ranks among the evaluation metrics than any other baseline. It is easy to implement, and code is provided at <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/ifnspaml\/Components-Loss\">https:\/\/github.com\/ifnspaml\/Components-Loss<\/jats:ext-link>.<\/jats:p>","DOI":"10.1186\/s13636-021-00207-6","type":"journal-article","created":{"date-parts":[[2021,7,2]],"date-time":"2021-07-02T11:03:29Z","timestamp":1625223809000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["Components loss for neural networks in mask-based speech enhancement"],"prefix":"10.1186","volume":"2021","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3046-1425","authenticated-orcid":false,"given":"Ziyi","family":"Xu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Samy","family":"Elshamy","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ziyue","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tim","family":"Fingscheidt","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,7,2]]},"reference":[{"issue":"6","key":"207_CR1","doi-asserted-by":"publisher","first-page":"1109","DOI":"10.1109\/TASSP.1984.1164453","volume":"32","author":"Y. Ephraim","year":"1984","unstructured":"Y. Ephraim, D. Malah, Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator. IEEE Trans. Acoust. Speech Signal Process.32(6), 1109\u20131121 (1984).","journal-title":"IEEE Trans. Acoust. Speech Signal Process."},{"issue":"2","key":"207_CR2","doi-asserted-by":"publisher","first-page":"443","DOI":"10.1109\/TASSP.1985.1164550","volume":"33","author":"Y. Ephraim","year":"1985","unstructured":"Y. Ephraim, D. Malah, Speech enhancement using a minimum mean-square error log-spectral amplitude estimator. IEEE Trans. Acoust. Speech Signal Process.33(2), 443\u2013445 (1985).","journal-title":"IEEE Trans. Acoust. Speech Signal Process."},{"key":"207_CR3","doi-asserted-by":"publisher","first-page":"629","DOI":"10.1109\/ICASSP.1996.543199","volume-title":"1996 IEEE International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings","author":"P. Scalart","year":"1996","unstructured":"P. Scalart, J. V. Filho, in 1996 IEEE International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings. Speech enhancement based on a priori signal to noise estimations (IEEEAtlanta, GA, USA, 1996), pp. 629\u2013632."},{"issue":"7","key":"207_CR4","first-page":"1110","volume":"2005","author":"T. Lotter","year":"2005","unstructured":"T. Lotter, P. Vary, Speech enhancement by MAP spectral amplitude estimation using a super-Gaussian speech model. EURASIP J. Appl. Signal Process.2005(7), 1110\u20131126 (2005).","journal-title":"EURASIP J. Appl. Signal Process."},{"key":"207_CR5","doi-asserted-by":"publisher","first-page":"4768","DOI":"10.1109\/ICASSP.2011.5947421","volume-title":"2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"B. Fodor","year":"2011","unstructured":"B. Fodor, T. Fingscheidt, in 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Speech enhancement using a joint map estimator with Gaussian mixture model for (non-) stationary noise (IEEEPrague, Czech Republic, 2011), pp. 4768\u20134771."},{"issue":"3","key":"207_CR6","doi-asserted-by":"publisher","first-page":"336","DOI":"10.1016\/j.specom.2005.02.011","volume":"47","author":"I. Cohen","year":"2005","unstructured":"I. Cohen, Speech enhancement using super-Gaussian speech models and noncausal a priori SNR estimation. Speech Commun.47(3), 336\u2013350 (2005).","journal-title":"Speech Commun."},{"issue":"5","key":"207_CR7","doi-asserted-by":"publisher","first-page":"910","DOI":"10.1109\/TASL.2008.921764","volume":"16","author":"T. Gerkmann","year":"2008","unstructured":"T. Gerkmann, C. Breithaupt, R. Martin, Improved a posteriori speech presence probability estimation based on a likelihood ratio with fixed priors. IEEE Trans. Audio Speech Lang. Process.16(5), 910\u2013919 (2008).","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"issue":"1","key":"207_CR8","doi-asserted-by":"publisher","first-page":"186","DOI":"10.1109\/TASL.2010.2045799","volume":"19","author":"S. Suhadi","year":"2011","unstructured":"S. Suhadi, C. Last, T. Fingscheidt, A data-driven approach to a priori SNR estimation. IEEE Trans. Audio Speech Lang. Process.19(1), 186\u2013195 (2011).","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"207_CR9","first-page":"1740","volume-title":"Sixteenth Annual Conference of the International Speech Communication Association","author":"S. Elshamy","year":"2015","unstructured":"S. Elshamy, N. Madhu, W. J. Tirry, T. Fingscheidt, in Sixteenth Annual Conference of the International Speech Communication Association. An iterative speech model-based a priori SNR estimator (ISCADresden, Germany, 2015), pp. 1740\u20131744."},{"issue":"8","key":"207_CR10","doi-asserted-by":"publisher","first-page":"1592","DOI":"10.1109\/TASLP.2017.2702385","volume":"25","author":"S. Elshamy","year":"2017","unstructured":"S. Elshamy, N. Madhu, W. Tirry, T. Fingscheidt, Instantaneous a priori SNR estimation by cepstral excitation manipulation. IEEE\/ACM Trans. Audio Speech Lang. Process.25(8), 1592\u20131605 (2017).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"207_CR11","first-page":"789","volume-title":"1999 IEEE International Conference on Acoustics, Speech, and Signal Processing","author":"D. Malah","year":"1999","unstructured":"D. Malah, R. V. Cox, A. J. Accardi, in 1999 IEEE International Conference on Acoustics, Speech, and Signal Processing. Tracking speech-presence uncertainty to improve speech enhancement in non-stationary noise environments (IEEEPhoenix, AZ, USA, 1999), pp. 789\u2013792."},{"key":"207_CR12","first-page":"1","volume-title":"Proc. of ITG Conf. on Speech Communication","author":"T. Fingscheidt","year":"2006","unstructured":"T. Fingscheidt, S. Suhadi, in Proc. of ITG Conf. on Speech Communication. Data-driven speech enhancement (ITGKiel, Germany, 2006), pp. 1\u20134."},{"issue":"4","key":"207_CR13","doi-asserted-by":"publisher","first-page":"825","DOI":"10.1109\/TASL.2008.920062","volume":"16","author":"T. Fingscheidt","year":"2008","unstructured":"T. Fingscheidt, S. Suhadi, S. Stan, Environment-optimized speech enhancement. IEEE Trans. Audio Speech Lang. Process. 16(4), 825\u2013834 (2008).","journal-title":"IEEE Trans. Audio Speech Lang. Process"},{"key":"207_CR14","first-page":"1","volume-title":"2006 14th European Signal Processing Conference","author":"J. Erkelens","year":"2006","unstructured":"J. Erkelens, J. Jensen, R. Heusdens, in 2006 14th European Signal Processing Conference. A general optimization procedure for spectral speech enhancement methods (EURASIPFlorence, Italy, 2006), pp. 1\u20135."},{"issue":"7-8","key":"207_CR15","doi-asserted-by":"publisher","first-page":"530","DOI":"10.1016\/j.specom.2006.06.012","volume":"49","author":"J. Erkelens","year":"2007","unstructured":"J. Erkelens, J. Jensen, R. Heusdens, A data-driven approach to optimizing spectral speech enhancement methods for various error criteria. Speech Commun.49(7-8), 530\u2013541 (2007).","journal-title":"Speech Commun."},{"issue":"12","key":"207_CR16","doi-asserted-by":"publisher","first-page":"1849","DOI":"10.1109\/TASLP.2014.2352935","volume":"22","author":"Y. Wang","year":"2014","unstructured":"Y. Wang, A. Narayanan, D. L. Wang, On training targets for supervised speech separation. IEEE\/ACM Trans. Audio Speech Lang. Process.22(12), 1849\u20131858 (2014).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"207_CR17","doi-asserted-by":"publisher","first-page":"577","DOI":"10.1109\/GlobalSIP.2014.7032183","volume-title":"2014 IEEE Global Conference on Signal and Information Processing (GlobalSIP)","author":"F. Weninger","year":"2014","unstructured":"F. Weninger, J. R. Hershey, J. Le Roux, B. Schuller, in 2014 IEEE Global Conference on Signal and Information Processing (GlobalSIP). Discriminatively trained recurrent neural networks for single-channel speech separation (IEEEAtlanta, GA, USA, 2014), pp. 577\u2013581."},{"key":"207_CR18","doi-asserted-by":"publisher","first-page":"1562","DOI":"10.1109\/ICASSP.2014.6853860","volume-title":"2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"P. S. Huang","year":"2014","unstructured":"P. S. Huang, M. Kim, M. H. Johnson, P. Smaragdis, in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Deep learning for monaural speech separation (IEEEFlorence, Italy, 2014), pp. 1562\u20131566."},{"key":"207_CR19","doi-asserted-by":"publisher","first-page":"708","DOI":"10.1109\/ICASSP.2015.7178061","volume-title":"2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"H. Erdogan","year":"2015","unstructured":"H. Erdogan, J. R. Hershey, S. Watanabe, J. Le Roux, in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks (IEEEBrisbane, QLD, Australia, 2015), pp. 708\u2013712."},{"key":"207_CR20","doi-asserted-by":"publisher","first-page":"4390","DOI":"10.1109\/ICASSP.2015.7178800","volume-title":"2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Y. Wang","year":"2015","unstructured":"Y. Wang, D. L. Wang, in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). A deep neural network for time-domain signal reconstruction (IEEEBrisbane, QLD, Australia, 2015), pp. 4390\u20134394."},{"issue":"3","key":"207_CR21","doi-asserted-by":"publisher","first-page":"483","DOI":"10.1109\/TASLP.2015.2512042","volume":"24","author":"D. S. Williamson","year":"2016","unstructured":"D. S. Williamson, Y. Wang, D. L. Wang, Complex ratio masking for monaural speech separation. IEEE\/ACM Trans. Audio Speech Lang. Process.24(3), 483\u2013492 (2016).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"1","key":"207_CR22","doi-asserted-by":"publisher","first-page":"7","DOI":"10.1109\/TASLP.2018.2868407","volume":"27","author":"F. Bao","year":"2019","unstructured":"F. Bao, W. H. Abdulla, A new ratio mask representation for CASA-based speech enhancement. IEEE\/ACM Trans. Audio Speech Lang. Process.27(1), 7\u201319 (2019).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"207_CR23","doi-asserted-by":"publisher","first-page":"239","DOI":"10.1109\/WASPAA.2019.8937222","volume-title":"2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)","author":"M. Strake","year":"2019","unstructured":"M. Strake, B. Defraene, K. Fluyt, W. Tirry, T. Fingscheidt, in 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). Separated noise suppression and speech restoration: LSTM-based speech enhancement in two stages (IEEENew Paltz, NY, USA, 2019), pp. 239\u2013243."},{"key":"207_CR24","doi-asserted-by":"publisher","first-page":"6674","DOI":"10.1109\/ICASSP40776.2020.9054230","volume-title":"ICASSP 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"M. Strake","year":"2020","unstructured":"M. Strake, B. Defraene, K. Fluyt, W. Tirry, T. Fingscheidt, in ICASSP 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Fully convolutional recurrent networks for speech enhancement (IEEEBarcelona, Spain, 2020), pp. 6674\u20136678."},{"issue":"10","key":"207_CR25","doi-asserted-by":"publisher","first-page":"1702","DOI":"10.1109\/TASLP.2018.2842159","volume":"26","author":"D. L. Wang","year":"2018","unstructured":"D. L. Wang, J. T. Chen, Supervised speech separation based on deep learning: an overview. IEEE\/ACM Trans. Audio Speech Lang. Process.26(10), 1702\u20131726 (2018).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"8","key":"207_CR26","doi-asserted-by":"publisher","first-page":"1424","DOI":"10.1109\/TASLP.2016.2558822","volume":"24","author":"J. Du","year":"2016","unstructured":"J. Du, Y. Tu, L. R. Dai, C. H. Lee, A regression approach to single-channel speech separation via high-resolution deep neural networks. IEEE\/ACM Trans. Audio Speech Lang. Process.24(8), 1424\u20131437 (2016).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"207_CR27","doi-asserted-by":"publisher","first-page":"3743","DOI":"10.21437\/Interspeech.2016-1284","volume-title":"Interspeech","author":"P. G. Shivakumar","year":"2016","unstructured":"P. G. Shivakumar, P. G. Georgiou, in Interspeech. Perception optimized deep denoising autoencoders for speech enhancement (ISCASan Francisco, CA, USA, 2016), pp. 3743\u20133747."},{"key":"207_CR28","doi-asserted-by":"publisher","first-page":"1270","DOI":"10.23919\/EUSIPCO.2017.8081412","volume-title":"2017 25th European Signal Processing Conference (EUSIPCO)","author":"Q. J. Liu","year":"2017","unstructured":"Q. J. Liu, W. Wang, P. J. B. Jackson, Y. Tang, in 2017 25th European Signal Processing Conference (EUSIPCO). A perceptually-weighted deep neural network for monaural speech enhancement in various background noise conditions (EURASIPKos, Greece, 2017), pp. 1270\u20131274."},{"key":"207_CR29","doi-asserted-by":"crossref","unstructured":"Z. Zhao, S. Elshamy, T. Fingscheidt, in 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). A perceptual weighting filter loss for DNN training in speech enhancement, (2019), pp. 229\u2013233.","DOI":"10.1109\/WASPAA.2019.8937189"},{"key":"207_CR30","doi-asserted-by":"publisher","first-page":"7000","DOI":"10.1109\/ICASSP.2019.8683341","volume-title":"ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"C. Brauer","year":"2019","unstructured":"C. Brauer, Z. Zhao, D. Lorenz, T. Fingscheidt, in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Learning to dequantize speech signals by primal-dual networks: an approach for acoustic sensor networks (IEEEBrighton, UK, 2019), pp. 7000\u20137004."},{"issue":"11","key":"207_CR31","doi-asserted-by":"publisher","first-page":"1680","DOI":"10.1109\/LSP.2018.2871419","volume":"25","author":"J. M. Mart\u00edn-Do\u00f1as","year":"2018","unstructured":"J. M. Mart\u00edn-Do\u00f1as, A. M. Gomez, J. A. Gonzalez, A. M. Peinado, A deep learning loss function based on the perceptual evaluation of the speech quality. IEEE Signal Process. Lett.25(11), 1680\u20131684 (2018).","journal-title":"IEEE Signal Process. Lett."},{"issue":"10","key":"207_CR32","doi-asserted-by":"publisher","first-page":"1780","DOI":"10.1109\/TASLP.2018.2842156","volume":"26","author":"Y. Koizumi","year":"2018","unstructured":"Y. Koizumi, K. Niwa, Y. Hioka, K. Kobayashi, Y. Haneda, DNN-based source enhancement to increase objective sound quality assessment score. IEEE\/ACM Trans. Audio Speech Lang. Process.26(10), 1780\u20131792 (2018).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"207_CR33","doi-asserted-by":"publisher","first-page":"5059","DOI":"10.1109\/ICASSP.2018.8462040","volume-title":"2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"M. Kolbcek","year":"2018","unstructured":"M. Kolbcek, Z. H. Tan, J. Jensen, in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Monaural speech enhancement using deep neural networks by maximizing a short-time objective intelligibility measure (IEEECalgary, AB, Canada, 2018), pp. 5059\u20135063."},{"key":"207_CR34","doi-asserted-by":"publisher","first-page":"386","DOI":"10.1109\/IWAENC.2018.8521379","volume-title":"2018 16th International Workshop on Acoustic Signal Enhancement (IWAENC)","author":"G. Naithani","year":"2018","unstructured":"G. Naithani, J. Nikunen, L. Bramslow, T. Virtanen, in 2018 16th International Workshop on Acoustic Signal Enhancement (IWAENC). Deep neural network based speech separation optimizing an objective estimator of intelligibility for low latency applications (IEEETokyo, Japan, 2018), pp. 386\u2013390."},{"key":"207_CR35","doi-asserted-by":"publisher","first-page":"5374","DOI":"10.1109\/ICASSP.2018.8461965","volume-title":"2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"H. Zhang","year":"2018","unstructured":"H. Zhang, X. L. Zhang, G. L. Gao, in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Training supervised speech separation system to improve STOI and PESQ directly (IEEECalgary, AB, Canada, 2018), pp. 5374\u20135378."},{"issue":"9","key":"207_CR36","doi-asserted-by":"publisher","first-page":"1570","DOI":"10.1109\/TASLP.2018.2821903","volume":"26","author":"S. W. Fu","year":"2018","unstructured":"S. W. Fu, T. W. Wang, Y. Tsao, X. Lu, H. Kawai, End-to-end waveform utterance enhancement for direct evaluation metrics optimization by fully convolutional neural networks. IEEE\/ACM Trans. Audio Speech Lang. Process.26(9), 1570\u20131584 (2018).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"207_CR37","first-page":"7","volume-title":"Proc. of ITG-Fachtagung \u201cSprachkommunikation\u201d","author":"T. Fingscheidt","year":"1996","unstructured":"T. Fingscheidt, P. Vary, in Proc. of ITG-Fachtagung \u201cSprachkommunikation\u201d. Error concealment by softbit speech decoding (ITGFrankfurt a.M., Germany, 1996), pp. 7\u201310."},{"issue":"3","key":"207_CR38","doi-asserted-by":"publisher","first-page":"240","DOI":"10.1109\/89.905998","volume":"9","author":"T. Fingscheidt","year":"2001","unstructured":"T. Fingscheidt, P. Vary, Softbit speech decoding: a new approach to error concealment. IEEE Trans. Speech Audio Process.9(3), 240\u2013251 (2001).","journal-title":"IEEE Trans. Speech Audio Process."},{"key":"207_CR39","doi-asserted-by":"publisher","first-page":"3499","DOI":"10.21437\/Interspeech.2018-2441","volume-title":"Interspeech","author":"H. Erdogan","year":"2018","unstructured":"H. Erdogan, T. Yoshioka, in Interspeech. Investigations on data augmentation and loss functions for deep learning based speech-background separation (ISCAHyderabad, India, 2018), pp. 3499\u20133503."},{"key":"207_CR40","doi-asserted-by":"publisher","first-page":"4214","DOI":"10.1109\/ICASSP.2010.5495701","volume-title":"2010 IEEE International Conference on Acoustics, Speech and Signal Processing","author":"C. H. Taal","year":"2010","unstructured":"C. H. Taal, R. C. Hendriks, R. Heusdens, J. Jensen, in 2010 IEEE International Conference on Acoustics, Speech and Signal Processing. A short-time objective intelligibility measure for time-frequency weighted noisy speech (IEEEDallas, TX, USA, 2010), pp. 4214\u20134217."},{"key":"207_CR41","unstructured":"ITU, Rec. P.862: perceptual evaluation of speech quality (PESQ): an objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs. International Telecommunication Standardization Sector (ITU-T) (ITU-T, 2001)."},{"key":"207_CR42","first-page":"36","volume-title":"Proc. Workshop on Quality Assessment in Speech, Audio, and Image Communication","author":"S. Gustafsson","year":"1996","unstructured":"S. Gustafsson, R. Martin, P. Vary, in Proc. Workshop on Quality Assessment in Speech, Audio, and Image Communication. On the optimization of speech enhancement systems using instrumental measures (ITG\/EURASIPDarmstadt, Germany, 1996), pp. 36\u201340."},{"key":"207_CR43","first-page":"818","volume-title":"Interspeech","author":"T. Fingscheidt","year":"2007","unstructured":"T. Fingscheidt, S. Suhadi, in Interspeech. Quality assessment of speech enhancement systems by separation of enhanced speech, noise, and echo (ISCAAntwerp, Belgium, 2007), pp. 818\u2013821."},{"key":"207_CR44","first-page":"1","volume-title":"Proc. of 5th Biennial Workshop on DSP for In-Vehicle Systems","author":"H. Yu","year":"2011","unstructured":"H. Yu, T. Fingscheidt, in Proc. of 5th Biennial Workshop on DSP for In-Vehicle Systems. A figure of merit for instrumental optimization of noise reduction algorithms (SpringerKiel, Germany, 2011), pp. 1\u20138."},{"key":"207_CR45","unstructured":"ITU, Rec. P.1100: narrowband hands-free communication in motor vehicles. International Telecommunication Standardization Sector (ITU-T) (2019)."},{"key":"207_CR46","unstructured":"ITU, Rec. P.1110: wideband hands-free communication in motor vehicles. International Telecommunication Standardization Sector (ITU-T) (2015)."},{"key":"207_CR47","unstructured":"ITU, Rec. P.1130: subsystem requirements for automotive speech services. International Telecommunication Standardization Sector (ITU-T) (2015)."},{"issue":"4","key":"207_CR48","doi-asserted-by":"publisher","first-page":"663","DOI":"10.1109\/TASLP.2018.2887337","volume":"27","author":"Z. Zhao","year":"2019","unstructured":"Z. Zhao, H. J. Liu, T. Fingscheidt, Convolutional neural networks to enhance coded speech. IEEE\/ACM Trans. Audio Speech Lang. Process.27(4), 663\u2013678 (2019).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"207_CR49","doi-asserted-by":"publisher","first-page":"7519","DOI":"10.1109\/ICASSP40776.2020.9052968","volume-title":"ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Z. Xu","year":"2020","unstructured":"Z. Xu, S. Elshamy, T. Fingscheidt, in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Using separate losses for speech and noise in mask-based speech enhancement (IEEEBarcelona, Spain, 2020), pp. 7519\u20137523."},{"key":"207_CR50","doi-asserted-by":"publisher","first-page":"871","DOI":"10.1109\/ICASSP40776.2020.9054254","volume-title":"ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Y. Xia","year":"2020","unstructured":"Y. Xia, S. Braun, C. K. A. Reddy, H. Dubey, R. Cutler, I. Tashev, in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Weighted speech distortion losses for neural-network-based real-time speech enhancement (IEEEBarcelona, Spain, 2020), pp. 871\u2013875."},{"key":"207_CR51","doi-asserted-by":"publisher","first-page":"2467","DOI":"10.21437\/Interspeech.2020-2439","volume-title":"Proc. Interspeech","author":"M. Strake","year":"2020","unstructured":"M. Strake, B. Defraene, K. Fluyt, W. Tirry, T. Fingscheidt, in Proc. Interspeech. INTERSPEECH 2020 deep noise suppression challenge: a fully convolutional recurrent network (FCRN) for joint dereverberation and denoising (ISCAShanghai, China, 2020), pp. 2467\u20132471."},{"key":"207_CR52","unstructured":"S. Z. Fu, C. F. Liao, Y. Tsao, S. D. Lin, in International Conference on Machine Learning. MetricGAN: generative adversarial networks based black-box metric scores optimization for speech enhancement, (2019), pp. 2031\u20132041."},{"key":"207_CR53","first-page":"550","volume-title":"Proc. of NIPS","author":"A. Veit","year":"2016","unstructured":"A. Veit, M. J. Wilber, S. Belongie, in Proc. of NIPS. Residual networks behave like ensembles of relatively shallow networks (Curran Associates Inc.Barcelona, Spain, 2016), pp. 550\u2013558."},{"issue":"5","key":"207_CR54","doi-asserted-by":"publisher","first-page":"2421","DOI":"10.1121\/1.2229005","volume":"120","author":"M. Cooke","year":"2006","unstructured":"M. Cooke, J. Barker, S. Cunningham, X. Shao, An audio-visual corpus for speech perception and automatic speech recognition. J. Acoust. Soc. Am.120(5), 2421\u20132424 (2006).","journal-title":"J. Acoust. Soc. Am."},{"key":"207_CR55","doi-asserted-by":"publisher","first-page":"504","DOI":"10.1109\/ASRU.2015.7404837","volume-title":"2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU)","author":"J. Barker","year":"2015","unstructured":"J. Barker, R. Marxer, E. Vincent, S. Watanabe, in 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU). The third \u2018CHiME\u2019speech separation and recognition challenge: dataset, task and baselines (IEEEScottsdale, AZ, USA, 2015), pp. 504\u2013511."},{"key":"207_CR56","unstructured":"ITU, Rec. P.56: objective measurement of active speech level. International Telecommunication Standardization Sector (ITU-T) (2011)."},{"key":"207_CR57","unstructured":"ITU, Rec. P.862.2: corrigendum 1, wideband extension to recommendation P.862 for the assessment of wideband telephone networks and speech codecs. International Telecommunication Standardization Sector (ITU-T) (2017)."},{"key":"207_CR58","unstructured":"ITU, Rec. P.863: perceptual objective listening quality prediction (POLQA). International Telecommunication Union, Telecommunication Standardization Sector (ITU-T) (2018)."},{"key":"207_CR59","doi-asserted-by":"publisher","first-page":"4573","DOI":"10.1109\/ICASSP.2012.6288936","volume-title":"2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"H. Yu","year":"2012","unstructured":"H. Yu, T. Fingscheidt, in 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Black box measurement of musical tones produced by noise reduction systems (IEEEKyoto, Japan, 2012), pp. 4573\u20134576."},{"key":"207_CR60","unstructured":"3GPP, Mandatory speech codec speech processing functions; adaptive multi-rate (AMR) speech codec; transcoding functions (3GPP TS 26.090, Rel. 14). 3GPP; TSG SA (2017)."}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-021-00207-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-021-00207-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-021-00207-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,7,2]],"date-time":"2021-07-02T11:06:57Z","timestamp":1625224017000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-021-00207-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,2]]},"references-count":60,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["207"],"URL":"https:\/\/doi.org\/10.1186\/s13636-021-00207-6","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,7,2]]},"assertion":[{"value":"20 November 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 March 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 July 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"24"}}