{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T05:26:34Z","timestamp":1781587594573,"version":"3.54.5"},"reference-count":64,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2020,12,1]],"date-time":"2020-12-01T00:00:00Z","timestamp":1606780800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,12,10]],"date-time":"2020-12-10T00:00:00Z","timestamp":1607558400000},"content-version":"vor","delay-in-days":9,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"NXP Semiconductors, Product Line Voice and Audio Solutions, Belgium"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["EURASIP J. Adv. Signal Process."],"published-print":{"date-parts":[[2020,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Single-channel speech enhancement in highly non-stationary noise conditions is a very challenging task, especially when interfering speech is included in the noise. Deep learning-based approaches have notably improved the performance of speech enhancement algorithms under such conditions, but still introduce speech distortions if strong noise suppression shall be achieved. We propose to address this problem by using a two-stage approach, first performing noise suppression and subsequently restoring natural sounding speech, using specifically chosen neural network topologies and loss functions for each task. A mask-based long short-term memory (LSTM) network is employed for noise suppression and speech restoration is performed via spectral mapping with a convolutional encoder-decoder network (CED). The proposed method improves speech quality (PESQ) over state-of-the-art single-stage methods by about 0.1 points for unseen highly non-stationary noise types including interfering speech. Furthermore, it is able to increase intelligibility in low-SNR conditions and consistently outperforms all reference methods.<\/jats:p>","DOI":"10.1186\/s13634-020-00707-1","type":"journal-article","created":{"date-parts":[[2020,12,10]],"date-time":"2020-12-10T10:04:47Z","timestamp":1607594687000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":43,"title":["Speech enhancement by LSTM-based noise suppression followed by CNN-based speech restoration"],"prefix":"10.1186","volume":"2020","author":[{"given":"Maximilian","family":"Strake","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bruno","family":"Defraene","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kristoff","family":"Fluyt","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wouter","family":"Tirry","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tim","family":"Fingscheidt","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2020,12,10]]},"reference":[{"issue":"6","key":"707_CR1","doi-asserted-by":"publisher","first-page":"1109","DOI":"10.1109\/TASSP.1984.1164453","volume":"32","author":"Y. Ephraim","year":"1984","unstructured":"Y. Ephraim, D. Malah, Speech enhancement using a minimum mean-square error short-time spectral amplitude estimator. IEEE Trans. Acoust. Speech, Signal Process.32(6), 1109\u20131121 (1984).","journal-title":"IEEE Trans. Acoust. Speech, Signal Process."},{"issue":"2","key":"707_CR2","doi-asserted-by":"publisher","first-page":"443","DOI":"10.1109\/TASSP.1985.1164550","volume":"33","author":"Y. Ephraim","year":"1985","unstructured":"Y. Ephraim, D. Malah, Speech enhancement using a minimum mean-square error log-spectral amplitude estimator. IEEE Trans. Acoust. Speech Sig. Process.33(2), 443\u2013445 (1985).","journal-title":"IEEE Trans. Acoust. Speech Sig. Process."},{"key":"707_CR3","first-page":"629","volume-title":"Proc. of ICASSP","author":"P. Scalart","year":"1996","unstructured":"P. Scalart, J. V. Filho, in Proc. of ICASSP. Speech enhancement based on a priori signal to noise estimation (IEEEAtlanta, 1996), pp. 629\u2013632."},{"issue":"7","key":"707_CR4","first-page":"1110","volume":"2005","author":"T. Lotter","year":"2005","unstructured":"T. Lotter, P. Vary, Speech enhancement by map spectral amplitude estimation using a super-Gaussian speech model. EURASIP J. Adv. Sig. Process.2005(7), 1110\u20131126 (2005).","journal-title":"EURASIP J. Adv. Sig. Process."},{"key":"707_CR5","first-page":"4897","volume-title":"Proc. of ICASSP","author":"C. Breithaupt","year":"2008","unstructured":"C. Breithaupt, T. Gerkmann, R. Martin, in Proc. of ICASSP. A novel a priori SNR estimation approach based on selective cepstro-temporal smoothing (IEEELas Vegas, 2008), pp. 4897\u20134900."},{"issue":"8","key":"707_CR6","doi-asserted-by":"publisher","first-page":"1592","DOI":"10.1109\/TASLP.2017.2702385","volume":"25","author":"S. Elshamy","year":"2017","unstructured":"S. Elshamy, N. Madhu, W. Tirry, T. Fingscheidt, Instantaneous a priori SNR estimation by cepstral excitation manipulation. IEEE\/ACM Trans. Audio, Speech, Lang. Process.25(8), 1592\u20131605 (2017).","journal-title":"IEEE\/ACM Trans. Audio, Speech, Lang. Process."},{"issue":"5","key":"707_CR7","doi-asserted-by":"publisher","first-page":"504","DOI":"10.1109\/89.928915","volume":"9","author":"R. Martin","year":"2001","unstructured":"R. Martin, Noise power spectral density estimation based on optimal smoothing and minimum statistics. IEEE Trans. Speech Audio Process.9(5), 504\u2013512 (2001).","journal-title":"IEEE Trans. Speech Audio Process."},{"issue":"5","key":"707_CR8","doi-asserted-by":"publisher","first-page":"466","DOI":"10.1109\/TSA.2003.811544","volume":"11","author":"I. Cohen","year":"2003","unstructured":"I. Cohen, Noise spectrum estimation in adverse environments: improved minima controlled recursive averaging. IEEE Trans. Speech Audio Process.11(5), 466\u2013475 (2003).","journal-title":"IEEE Trans. Speech Audio Process."},{"issue":"4","key":"707_CR9","doi-asserted-by":"publisher","first-page":"1383","DOI":"10.1109\/TASL.2011.2180896","volume":"20","author":"T. Gerkmann","year":"2012","unstructured":"T. Gerkmann, R. C. Hendriks, Unbiased MMSE-based noise power estimation with low complexity and low tracking delay. IEEE\/ACM Trans. Audio Speech Lang. Process.20(4), 1383\u20131393 (2012).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"2","key":"707_CR10","doi-asserted-by":"publisher","first-page":"220","DOI":"10.1016\/j.specom.2005.08.005","volume":"48","author":"S. Rangachari","year":"2006","unstructured":"S. Rangachari, P. C. Loizou, A noise-estimation algorithm for highly non-stationary environments. Speech Commun.48(2), 220\u2013231 (2006).","journal-title":"Speech Commun."},{"key":"707_CR11","volume-title":"Speech enhancement: theory and practice","author":"C. Loizou Philipos","year":"2007","unstructured":"C. Loizou Philipos, Speech enhancement: theory and practice (CRC Press, Boca Raton, 2007)."},{"issue":"7","key":"707_CR12","doi-asserted-by":"publisher","first-page":"1381","DOI":"10.1109\/TASL.2013.2250961","volume":"21","author":"Y. Wang","year":"2013","unstructured":"Y. Wang, D. L. Wang, Towards scaling up classification-based speech separation. IEEE\/ACM Trans. Audio Speech Lang. Process.21(7), 1381\u20131390 (2013).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"1","key":"707_CR13","doi-asserted-by":"publisher","first-page":"65","DOI":"10.1109\/LSP.2013.2291240","volume":"21","author":"Y. Xu","year":"2014","unstructured":"Y. Xu, J. Du, L. R. Dai, C. H. Lee, An experimental study on speech enhancement based on deep neural networks. IEEE Sig. Process. Lett.21(1), 65\u201368 (2014).","journal-title":"IEEE Sig. Process. Lett."},{"issue":"1","key":"707_CR14","doi-asserted-by":"publisher","first-page":"7","DOI":"10.1109\/TASLP.2014.2364452","volume":"23","author":"Y. Xu","year":"2015","unstructured":"Y. Xu, J. Du, L. R. Dai, C. H. Lee, A regression approach to speech enhancement based on deep neural networks. IEEE\/ACM Trans. Audio Speech Lang. Process.23(1), 7\u201319 (2015).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"12","key":"707_CR15","doi-asserted-by":"publisher","first-page":"1849","DOI":"10.1109\/TASLP.2014.2352935","volume":"22","author":"Y. Wang","year":"2014","unstructured":"Y. Wang, A. Narayanan, D. L. Wang, On training targets for supervised speech separation. IEEE\/ACM Trans. Audio Speech Lang. Process. 22(12), 1849\u20131858 (2014).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process"},{"key":"707_CR16","first-page":"577","volume-title":"Proc. of GlobalSIP Machine Learning Applications in Speech Processing Symposium","author":"F. Weninger","year":"2014","unstructured":"F. Weninger, J. R. Hershey, J. Le Roux, B. Schuller, in Proc. of GlobalSIP Machine Learning Applications in Speech Processing Symposium. Discriminatively trained recurrent neural networks for single-channel speech separation (IEEEAtlanta, 2014), pp. 577\u2013581."},{"issue":"3","key":"707_CR17","doi-asserted-by":"publisher","first-page":"483","DOI":"10.1109\/TASLP.2015.2512042","volume":"24","author":"D. S. Williamson","year":"2016","unstructured":"D. S. Williamson, Y. Wang, D. L. Wang, Complex ratio masking for monaural speech separation. IEEE\/ACM Trans. Audio Speech Lang. Process.24(3), 483\u2013492 (2016).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"4","key":"707_CR18","doi-asserted-by":"publisher","first-page":"663","DOI":"10.1109\/TASLP.2018.2887337","volume":"27","author":"Z. Zhao","year":"2019","unstructured":"Z. Zhao, H. Liu, T. Fingscheidt, Convolutional neural networks to enhance coded speech. IEEE\/ACM Trans. Audio Speech Lang. Process.27(4), 663\u2013678 (2019).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"12","key":"707_CR19","doi-asserted-by":"publisher","first-page":"2460","DOI":"10.1109\/TASLP.2018.2867947","volume":"26","author":"S. Elshamy","year":"2018","unstructured":"S. Elshamy, N. Madhu, W. Tirry, T. Fingscheidt, DNN-supported speech enhancement with cepstral estimation of both excitation and envelope. IEEE\/ACM Trans. Audio Speech Lang. Process.26(12), 2460\u20132474 (2018).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"707_CR20","first-page":"106","volume-title":"Proc. of IWAENC","author":"N. Takahashi","year":"2018","unstructured":"N. Takahashi, N. Goswami, Y. Mitsufuji, in Proc. of IWAENC. MMdenseLSTM: an efficient combination of convolutional and recurrent neural networks for audio source separation (IEEETokyo, 2018), pp. 106\u2013110."},{"key":"707_CR21","first-page":"5054","volume-title":"Proc. of ICASSP","author":"T. Gao","year":"2018","unstructured":"T. Gao, J. Du, L. -R. Dai, C. -H. Lee, in Proc. of ICASSP. Densely connected progressive learning for LSTM-based speech enhancement (IEEECalgary, 2018), pp. 5054\u20135058."},{"key":"707_CR22","first-page":"708","volume-title":"Proc. of ICASSP","author":"H. Erdogan","year":"2015","unstructured":"H. Erdogan, J. R. Hershey, S. Watanabe, J. Le Roux, in Proc. of ICASSP. Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks (IEEEBrisbane, 2015), pp. 708\u2013712."},{"key":"707_CR23","volume-title":"Proc. of INTERSPEECH","author":"T. Fingscheidt","year":"2007","unstructured":"T. Fingscheidt, S. Suhadi, in Proc. of INTERSPEECH. Quality assessment of speech enhancement systems by separation of enhanced speech, noise, and echo (ISCAAntwerpen, 2007)."},{"key":"707_CR24","unstructured":"ITU-T Rec P.1100, Narrow-band hands-free communication in motor vehicles (2015)."},{"issue":"6","key":"707_CR25","doi-asserted-by":"publisher","first-page":"4705","DOI":"10.1121\/1.4986931","volume":"141","author":"J. Chen","year":"2017","unstructured":"J. Chen, D. L. Wang, Long short-term memory for speaker generalization in supervised speech separation. J. Acoust. Soc. Am.141(6), 4705\u20134714 (2017).","journal-title":"J. Acoust. Soc. Am."},{"key":"707_CR26","first-page":"1","volume-title":"Proc. of MLSP","author":"S. -W. Fu","year":"2017","unstructured":"S. -W. Fu, T. Hu, Y. Tsao, X. Lu, in Proc. of MLSP. Complex spectrogram enhancement by convolutional neural network with multi-metrics learning (IEEETokyo, 2017), pp. 1\u20136."},{"key":"707_CR27","doi-asserted-by":"publisher","first-page":"1993","DOI":"10.21437\/Interspeech.2017-1465","volume-title":"Proc. of INTERSPEECH","author":"S. R. Park","year":"2017","unstructured":"S. R. Park, J. Lee, in Proc. of INTERSPEECH. A fully convolutional neural network for speech enhancement (ISCAStockholm, 2017), pp. 1993\u20131997."},{"key":"707_CR28","first-page":"2802","volume-title":"Proc. of NIPS","author":"X. Mao","year":"2016","unstructured":"X. Mao, C. Shen, Y. -B. Yang, in Proc. of NIPS. Image restoration using very deep convolutional encoder-decoder networks with symmetric skip connections (Curran Associates, Inc.Barcelona, 2016), pp. 2802\u20132810."},{"issue":"12","key":"707_CR29","doi-asserted-by":"publisher","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","volume":"39","author":"V. Badrinarayanan","year":"2017","unstructured":"V. Badrinarayanan, A. Kendall, R. Cipolla, SegNet: a deep convolutional encoder-decoder architecture for image segmentation. IEEE Trans. Pattern Anal. Mach. Intell.39(12), 2481\u20132495 (2017).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"707_CR30","first-page":"1520","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"H. Noh","year":"2015","unstructured":"H. Noh, S. Hong, B. Han, in Proceedings of the IEEE International Conference on Computer Vision. Learning deconvolution network for semantic segmentation (IEEESantiago, 2015), pp. 1520\u20131528."},{"key":"707_CR31","first-page":"2401","volume-title":"Proc. of ICASSP","author":"H. Zhao","year":"2018","unstructured":"H. Zhao, S. Zarar, I. Tashev, C. Lee, in Proc. of ICASSP. Convolutional-recurrent neural networks for speech enhancement (IEEECalgary, 2018), pp. 2401\u20132405."},{"key":"707_CR32","doi-asserted-by":"publisher","first-page":"3229","DOI":"10.21437\/Interspeech.2018-1405","volume-title":"Proc. of INTERSPEECH","author":"K. Tan","year":"2018","unstructured":"K. Tan, D. L. Wang, in Proc. of INTERSPEECH. A convolutional recurrent neural network for real-time speech enhancement (ISCAHyderabad, 2018), pp. 3229\u20133233."},{"key":"707_CR33","doi-asserted-by":"crossref","unstructured":"Z. Xu, M. Strake, T. Fingscheidt, Concatenated identical DNN (CI-DNN) to reduce noise-type dependence in DNN-based speech enhancement. arXiv:1810.11217 (2018).","DOI":"10.23919\/EUSIPCO.2019.8903066"},{"key":"707_CR34","first-page":"55","volume-title":"Proc. of CISS","author":"M. Tinston","year":"2009","unstructured":"M. Tinston, Y. Ephraim, in Proc. of CISS. Speech enhancement using the multistage wiener filter (IEEEBaltimore, 2009), pp. 55\u201360."},{"issue":"2","key":"707_CR35","doi-asserted-by":"publisher","first-page":"892","DOI":"10.1121\/1.4884759","volume":"136","author":"D. S. Williamson","year":"2014","unstructured":"D. S. Williamson, Y. Wang, D. L. Wang, Reconstruction techniques for improving the perceptual quality of binary masked speech. J. Acoust. Soc. Am.136(2), 892\u2013902 (2014).","journal-title":"J. Acoust. Soc. Am."},{"key":"707_CR36","volume-title":"Proc. of INTERSPEECH","author":"E. M. Grais","year":"2013","unstructured":"E. M. Grais, H. Erdogan, in Proc. of INTERSPEECH. Spectro-temporal post-enhancement using MMSE estimation in NMF based single-channel source separation (ISCALyon, 2013)."},{"issue":"9","key":"707_CR37","doi-asserted-by":"publisher","first-page":"1773","DOI":"10.1109\/TASLP.2017.2716443","volume":"25","author":"E. M. Grais","year":"2017","unstructured":"E. M. Grais, G. Roma, A. J. R. Simpson, M. D. Plumbley, Two-stage single-channel audio source separation using deep neural networks. IEEE\/ACM Trans. Audio Speech Lang. Process.25(9), 1773\u20131783 (2017).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"4","key":"707_CR38","doi-asserted-by":"publisher","first-page":"663","DOI":"10.1109\/TASLP.2018.2887337","volume":"27","author":"Z. Zhao","year":"2019","unstructured":"Z. Zhao, H. Liu, T. Fingscheidt, Convolutional neural networks to enhance coded speech. ACM Trans. Audio Speech Lang. Process.27(4), 663\u2013678 (2019).","journal-title":"ACM Trans. Audio Speech Lang. Process."},{"key":"707_CR39","first-page":"234","volume-title":"Proc. of WASPAA","author":"M. Strake","year":"2019","unstructured":"M. Strake, B. Defraene, K. Fluyt, W. Tirry, T. Fingscheidt, in Proc. of WASPAA. Separated noise suppression and speech restoration: LSTM-based speech enhancement in two stages (IEEENew Paltz, 2019), pp. 234\u2013238."},{"key":"707_CR40","volume-title":"ITG-Fachtagung Sprachkommunikation","author":"T. Fingscheidt","year":"2006","unstructured":"T. Fingscheidt, S. Suhadi, in ITG-Fachtagung Sprachkommunikation. Data-driven speech enhancement (ITGKiel, 2006)."},{"issue":"4","key":"707_CR41","doi-asserted-by":"publisher","first-page":"825","DOI":"10.1109\/TASL.2008.920062","volume":"16","author":"T. Fingscheidt","year":"2008","unstructured":"T. Fingscheidt, S. Suhadi, S. Stan, Environment-optimized speech enhancement. IEEE\/ACM Trans. Audio Speech Lang. Process.16(4), 825\u2013834 (2008).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"707_CR42","unstructured":"R. Pascanu, C. Gulcehre, K. Cho, Y. Bengio, How to construct deep recurrent neural networks. arXiv:1312.6026 (2013)."},{"key":"707_CR43","first-page":"807","volume-title":"Proc. of ICML","author":"V. Nair","year":"2010","unstructured":"V. Nair, G. E. Hinton, in Proc. of ICML. Rectified linear units improve restricted boltzmann machines (OmnipressHaifa, 2010), pp. 807\u2013814."},{"issue":"8","key":"707_CR44","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S. Hochreiter","year":"1997","unstructured":"S. Hochreiter, J. Schmidhuber, Long short-term memory. Neural Comput.9(8), 1735\u20131780 (1997).","journal-title":"Neural Comput."},{"key":"707_CR45","first-page":"2018","volume-title":"Proc. of ICCV","author":"M. D. Zeiler","year":"2011","unstructured":"M. D. Zeiler, G. W. Taylor, R. Fergus, in Proc. of ICCV. Adaptive deconvolutional networks for mid and high level feature learning (IEEEBarcelona, 2011), pp. 2018\u20132025."},{"key":"707_CR46","volume-title":"Proc. of ICML Workshop on Deep Learning for Audio, Speech, and Language Processing","author":"A. L. Maas","year":"2013","unstructured":"A. L. Maas, A. Y. Hannun, A. Y. Ng, in Proc. of ICML Workshop on Deep Learning for Audio, Speech, and Language Processing. Rectifier nonlinearities improve neural network acoustic models (OmnipressAtlanta, 2013)."},{"key":"707_CR47","unstructured":"V. Dumoulin, F. Visin, A guide to convolution arithmetic for deep learning. arXiv:1603.07285 (2016)."},{"key":"707_CR48","volume-title":"TIMIT acoustic-phonetic continuous speech corpus","author":"J. S. Garofolo","year":"1993","unstructured":"J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, TIMIT acoustic-phonetic continuous speech corpus (Linguistic Data Consortium, Philadelpia, 1993). Linguistic Data Consortium."},{"key":"707_CR49","unstructured":"NTT Advanced Technology Corporation, Super wideband stereo speech database. San Jose, CA, USA. NTT Advanced Technology Corporation."},{"key":"707_CR50","doi-asserted-by":"crossref","first-page":"3110","DOI":"10.21437\/Interspeech.2010-774","volume-title":"Proc. of INTERSPEECH","author":"D. B. Dean","year":"2010","unstructured":"D. B. Dean, S. Sridharan, R. J. Vogt, M. W. Mason, in Proc. of INTERSPEECH. The QUT-NOISE-TIMIT corpus for the evaluation of voice activity detection algorithms (ISCAMakuhari, 2010), pp. 3110\u20133113."},{"key":"707_CR51","first-page":"181","volume-title":"Proc. of ASR2000-Automatic Speech Recognition: Challenges for the New Millenium ISCA Tutorial and Research Workshop","author":"H. -G. Hirsch","year":"2000","unstructured":"H. -G. Hirsch, D. Pearce, in Proc. of ASR2000-Automatic Speech Recognition: Challenges for the New Millenium ISCA Tutorial and Research Workshop. The aurora experimental framework for the performance evaluation of speech recognition systems under noisy conditions (ISCAParis, 2000), pp. 181\u2013188."},{"key":"707_CR52","unstructured":"EG 202 396-1, Speech Processing, ETSI, Transmission and Quality Aspects (STQ); Speech Quality Performance in the Presence of Background Noise; Part 1: Background Noise Simulation Technique and Background Noise Database (2008)."},{"issue":"4","key":"707_CR53","doi-asserted-by":"publisher","first-page":"339","DOI":"10.1016\/0893-6080(88)90007-X","volume":"1","author":"P. J. Werbos","year":"1988","unstructured":"P. J. Werbos, Generalization of backpropagation with application to a recurrent gas market model. Neural Netw.1(4), 339\u2013356 (1988).","journal-title":"Neural Netw."},{"key":"707_CR54","unstructured":"D. P. Kingma, J. Ba, Adam: a method for stochastic optimization. arXiv:1412.6980 (2014)."},{"issue":"6088","key":"707_CR55","doi-asserted-by":"publisher","first-page":"533","DOI":"10.1038\/323533a0","volume":"323","author":"D. E. Rumelhart","year":"1986","unstructured":"D. E. Rumelhart, G. E. Hinton, R. J. Williams, Learning representations by back-propagating errors. Nature. 323(6088), 533\u2013536 (1986).","journal-title":"Nature"},{"key":"707_CR56","unstructured":"H. Yu, Post-filter optimization for multichannel automotive speech enhancement. PhD thesis, Technische Universit\u00e4t Braunschweig (2013)."},{"key":"707_CR57","first-page":"802","volume-title":"Proc. of NIPS","author":"X. Shi","year":"2015","unstructured":"X. Shi, Z. Chen, H. Wang, D. Y. Yeung, W. Wong, W. Woo, in Proc. of NIPS. Convolutional LSTM network: a machine learning approach for precipitation nowcasting (Curran Associates, Inc.Montreal, 2015), pp. 802\u2013810."},{"key":"707_CR58","unstructured":"ITU-T Rec. G.160 Appendix II, Objective measures for the characterization of the basic functioning of noise reduction algorithms (2012)."},{"issue":"4","key":"707_CR59","doi-asserted-by":"publisher","first-page":"670","DOI":"10.1109\/TASLP.2015.2401426","volume":"23","author":"V. Mai","year":"2015","unstructured":"V. Mai, D. Pastor, A. A\u00efssa-El-Bey, R. Le-Bidan, Robust estimation of non-stationary noise power spectrum for speech enhancement. IEEE\/ACM Trans. Audio Speech Lang. Process.23(4), 670\u2013682 (2015).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"5","key":"707_CR60","doi-asserted-by":"publisher","first-page":"703","DOI":"10.1016\/j.sigpro.2008.10.020","volume":"89","author":"M. Rahmani","year":"2009","unstructured":"M. Rahmani, A. Akbari, B. Ayad, B. Lithgow, Noise cross psd estimation using phase information in diffuse noise field. Sig. Process.89(5), 703\u2013709 (2009).","journal-title":"Sig. Process."},{"key":"707_CR61","first-page":"524","volume-title":"Proc. of ICASSP","author":"A. Sugiyama","year":"2015","unstructured":"A. Sugiyama, R. Miyahara, in Proc. of ICASSP. A directional noise suppressor with a specified beamwidth (IEEEBrisbane, 2015), pp. 524\u2013528."},{"key":"707_CR62","unstructured":"ITU-T Rec. P.862, Perceptual evaluation of speech quality (PESQ): an objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs (2001)."},{"key":"707_CR63","first-page":"4214","volume-title":"Proc. of ICASSP","author":"C. H. Taal","year":"2010","unstructured":"C. H. Taal, R. C. Hendriks, R. Heusdens, J. Jensen, in Proc. of ICASSP. A short-time objective intelligibility measure for time-frequency weighted noisy speech (IEEEDallas, 2010), pp. 4214\u20134217."},{"key":"707_CR64","first-page":"36","volume-title":"Proc. of Workshop on Quality Assessment in Speech, Audio, and Image Communication","author":"S. Gustafsson","year":"1996","unstructured":"S. Gustafsson, R. Martin, P. Vary, in Proc. of Workshop on Quality Assessment in Speech, Audio, and Image Communication. On the optimization of speech enhancement systems using instrumental measures (ITG\/EURASIPDarmstadt, 1996), pp. 36\u201340."}],"container-title":["EURASIP Journal on Advances in Signal Processing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13634-020-00707-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/s13634-020-00707-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13634-020-00707-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,3]],"date-time":"2022-12-03T22:58:46Z","timestamp":1670108326000},"score":1,"resource":{"primary":{"URL":"https:\/\/asp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13634-020-00707-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,12]]},"references-count":64,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,12]]}},"alternative-id":["707"],"URL":"https:\/\/doi.org\/10.1186\/s13634-020-00707-1","relation":{},"ISSN":["1687-6180"],"issn-type":[{"value":"1687-6180","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,12]]},"assertion":[{"value":"10 January 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 November 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 December 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors would like to disclose that NXP Semiconductors has filed a patent comprising parts of this work. The authors declare that this did not in any way affect the interpretation or presentation of results in this work and that they have no other competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"49"}}