{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T03:14:09Z","timestamp":1783480449969,"version":"3.55.0"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,9,7]],"date-time":"2021-09-07T00:00:00Z","timestamp":1630972800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,9,7]],"date-time":"2021-09-07T00:00:00Z","timestamp":1630972800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["11874219"],"award-info":[{"award-number":["11874219"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The acoustic echo cannot be entirely removed by linear adaptive filters due to the nonlinear relationship between the echo and the far-end signal. Usually, a post-processing module is required to further suppress the echo. In this paper, we propose a residual echo suppression method based on the modification of dual-path recurrent neural network (DPRNN) to improve the quality of speech communication. Both the residual signal and the auxiliary signal, the far-end signal or the output of the adaptive filter, obtained from the linear acoustic echo cancelation are adopted to form a dual-stream for the DPRNN. We validate the efficacy of the proposed method in the notoriously difficult double-talk situations and discuss the impact of different auxiliary signals on performance. We also compare the performance of the time domain and the time-frequency domain processing. Furthermore, we propose an efficient and applicable way to deploy our method to off-the-shelf loudspeakers by fine-tuning the pre-trained model with little recorded-echo data.<\/jats:p>","DOI":"10.1186\/s13636-021-00221-8","type":"journal-article","created":{"date-parts":[[2021,9,7]],"date-time":"2021-09-07T20:02:42Z","timestamp":1631044962000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Nonlinear residual echo suppression based on dual-stream DPRNN"],"prefix":"10.1186","volume":"2021","author":[{"given":"Hongsheng","family":"Chen","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guoliang","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kai","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jing","family":"Lu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,9,7]]},"reference":[{"issue":"11","key":"221_CR1","doi-asserted-by":"publisher","first-page":"2295","DOI":"10.1016\/S0165-1684(00)00118-3","volume":"80","author":"E. H\u00e4nsler","year":"2000","unstructured":"E. H\u00e4nsler, G. U. Schmidt, Hands-free telephones\u2013joint control of echo cancellation and postfiltering. Signal Process.80(11), 2295\u20132305 (2000).","journal-title":"Signal Process."},{"key":"221_CR2","volume-title":"Adaptive Filter Theory","author":"S. S. Haykin","year":"2002","unstructured":"S. S. Haykin, Adaptive Filter Theory (Prentice Hall, New Jersey, 2002)."},{"key":"221_CR3","unstructured":"F. Albu, H. K. Kwan, in 2004 IEEE International Symposium on Circuits and Systems (IEEE Cat. No.04CH37512), 3. Combined echo and noise cancellation based on Gauss-Seidel pseudo affine projection algorithm, (2004), p. 505."},{"issue":"6","key":"221_CR4","doi-asserted-by":"publisher","first-page":"1140","DOI":"10.1016\/j.sigpro.2005.09.013","volume":"86","author":"G. Enzner","year":"2006","unstructured":"G. Enzner, P. Vary, Frequency-domain adaptive kalman filter for acoustic echo control in hands-free telephones. Signal Process.86(6), 1140\u20131156 (2006).","journal-title":"Signal Process."},{"issue":"12","key":"221_CR5","doi-asserted-by":"publisher","first-page":"1778","DOI":"10.1109\/LSP.2017.2718564","volume":"24","author":"F. Yang","year":"2017","unstructured":"F. Yang, G. Enzner, J. Yang, Frequency-domain adaptive Kalman filter with fast recovery of abrupt echo-path changes. IEEE Signal Process. Lett.24(12), 1778\u20131782 (2017).","journal-title":"IEEE Signal Process. Lett."},{"issue":"2","key":"221_CR6","doi-asserted-by":"publisher","first-page":"342","DOI":"10.1109\/LSP.2019.2890965","volume":"26","author":"W. Fan","year":"2019","unstructured":"W. Fan, K. Chen, J. Lu, J. Tao, Effective improvement of under-modeling frequency-domain Kalman filter. IEEE Signal Process. Lett.26(2), 342\u2013346 (2019).","journal-title":"IEEE Signal Process. Lett."},{"key":"221_CR7","unstructured":"A. N. Birkett, R. A. Goubran, in Proceedings of 1995 Workshop on Applications of Signal Processing to Audio and Accoustics. Limitations of handsfree acoustic echo cancellers due to nonlinear loudspeaker distortion and enclosure vibration effects, (1995), pp. 103\u2013106."},{"issue":"1","key":"221_CR8","doi-asserted-by":"publisher","first-page":"21","DOI":"10.1016\/S0165-1684(97)00173-4","volume":"64","author":"S. Gustafsson","year":"1998","unstructured":"S. Gustafsson, R. Martin, P. Vary, Combined acoustic echo control and noise reduction for hands-free telephony. Signal Process.64(1), 21\u201332 (1998).","journal-title":"Signal Process."},{"issue":"8","key":"221_CR9","doi-asserted-by":"publisher","first-page":"1433","DOI":"10.1109\/TASL.2008.2002071","volume":"16","author":"E. A. P. Habets","year":"2008","unstructured":"E. A. P. Habets, S. Gannot, I. Cohen, P. C. W. Sommen, Joint dereverberation and residual echo suppression of speech signals in noisy environments. IEEE Trans. Audio Speech Lang. Process.16(8), 1433\u20131451 (2008).","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"221_CR10","doi-asserted-by":"publisher","first-page":"77","DOI":"10.1109\/TASLP.2019.2948765","volume":"28","author":"N. K. Desiraju","year":"2020","unstructured":"N. K. Desiraju, S. Doclo, M. Buck, T. Wolff, Online estimation of reverberation parameters for late residual echo suppression. IEEE Trans. Audio Speech Lang. Process.28:, 77\u201391 (2020).","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"issue":"5","key":"221_CR11","doi-asserted-by":"publisher","first-page":"245","DOI":"10.1109\/TSA.2002.800553","volume":"10","author":"S. Gustafsson","year":"2002","unstructured":"S. Gustafsson, R. Martin, P. Jax, P. Vary, A psychoacoustic approach to combined acoustic echo cancellation and noise reduction. IEEE Trans. Speech Audio Process.10(5), 245\u2013256 (2002).","journal-title":"IEEE Trans. Speech Audio Process."},{"key":"221_CR12","unstructured":"A. S. Chhetri, A. C. Surendran, J. W. Stokes, J. C. Platt, in Proc. IWAENC, 5. Regression-based residual acoustic echo suppression, (2005)."},{"key":"221_CR13","doi-asserted-by":"crossref","unstructured":"M. L. Valero, E. Mabande, E. A. P. Habets, in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Signal-based late residual echo spectral variance estimation, (2014), pp. 5914\u20135918.","DOI":"10.1109\/ICASSP.2014.6854738"},{"key":"221_CR14","doi-asserted-by":"crossref","unstructured":"G. Carbajal, R. Serizel, E. Vincent, E. Humbert, in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Multiple-input neural network-based residual echo suppression, (2018), pp. 231\u2013235.","DOI":"10.1109\/ICASSP.2018.8461476"},{"issue":"2","key":"221_CR15","first-page":"322","volume":"161","author":"H. Zhang","year":"2018","unstructured":"H. Zhang, D. Wang, Deep learning for acoustic echo cancellation in noisy and double-talk scenarios. Training. 161(2), 322 (2018).","journal-title":"Training"},{"key":"221_CR16","doi-asserted-by":"crossref","unstructured":"F. Kuech, W. Kellermann, in 2007 IEEE International Conference on Acoustics, Speech and Signal Processing, 1. Nonlinear residual echo suppression using a power filter model of the acoustic echo path, (2007), pp. 73\u201376.","DOI":"10.1109\/ICASSP.2007.366619"},{"key":"221_CR17","doi-asserted-by":"crossref","unstructured":"C. Zhang, X. Zhang, in Proc. Interspeech. A robust and cascaded acoustic echo cancellation based on deep learning, (2020), pp. 3940\u20133944.","DOI":"10.21437\/Interspeech.2020-1260"},{"issue":"8","key":"221_CR18","doi-asserted-by":"publisher","first-page":"1256","DOI":"10.1109\/TASLP.2019.2915167","volume":"27","author":"Y. Luo","year":"2019","unstructured":"Y. Luo, N. Mesgarani, Conv-tasnet: surpassing ideal time-frequency magnitude masking for speech separation. IEEE\/ACM Trans Audio Speech Lang. Process.27(8), 1256\u20131266 (2019).","journal-title":"IEEE\/ACM Trans Audio Speech Lang. Process."},{"key":"221_CR19","doi-asserted-by":"crossref","unstructured":"H. Chen, T. Xiang, K. Chen, J. Lu, in Proc. Interspeech. Nonlinear residual echo suppression based on multi-stream Conv-TasNET, (2020), pp. 3959\u20133963.","DOI":"10.21437\/Interspeech.2020-2234"},{"key":"221_CR20","doi-asserted-by":"crossref","unstructured":"Y. Luo, Z. Chen, T. Yoshioka, in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation (IEEE, 2020), pp. 46\u201350.","DOI":"10.1109\/ICASSP40776.2020.9054266"},{"key":"221_CR21","doi-asserted-by":"crossref","unstructured":"K. He, X. Zhang, S. Ren, J. Sun, in Proceedings of the IEEE International Conference on Computer Vision. Delving deep into rectifiers: surpassing human-level performance on imagenet classification, (2015), pp. 1026\u20131034.","DOI":"10.1109\/ICCV.2015.123"},{"issue":"2","key":"221_CR22","doi-asserted-by":"publisher","first-page":"145","DOI":"10.1109\/MSP.2013.2297439","volume":"32","author":"A. Cichocki","year":"2015","unstructured":"A. Cichocki, D. Mandic, L. De Lathauwer, G. Zhou, Q. Zhao, C. Caiafa, H. A. PHAN, Tensor decompositions for signal processing applications: from two-way to multiway component analysis. IEEE Signal Process. Mag.32(2), 145\u2013163 (2015). https:\/\/doi.org\/10.1109\/MSP.2013.2297439.","journal-title":"IEEE Signal Process. Mag."},{"key":"221_CR23","doi-asserted-by":"crossref","unstructured":"D. Yin, C. Luo, Z. Xiong, W. Zeng, Phasen: a phase-and-harmonics-aware speech enhancement network. arXiv preprint arXiv:1911.04697 (2019).","DOI":"10.1609\/aaai.v34i05.6489"},{"key":"221_CR24","doi-asserted-by":"crossref","unstructured":"Y. Wu, K. He, in Proceedings of the European Conference on Computer Vision (ECCV). Group normalization, (2018), pp. 3\u201319.","DOI":"10.1007\/978-3-030-01261-8_1"},{"key":"221_CR25","doi-asserted-by":"crossref","unstructured":"V. Panayotov, G. Chen, D. Povey, S. Khudanpur, in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Librispeech: an ASR corpus based on public domain audio books (IEEE, 2015), pp. 5206\u20135210.","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"221_CR26","unstructured":"D. Snyder, G. Chen, D. Povey, Musan: a music, speech, and noise corpus. arXiv preprint arXiv:1510.08484 (2015)."},{"issue":"7","key":"221_CR27","doi-asserted-by":"publisher","first-page":"2065","DOI":"10.1109\/TASL.2012.2196512","volume":"20","author":"S. Malik","year":"2012","unstructured":"S. Malik, G. Enzner, State-space frequency-domain adaptive filtering for nonlinear acoustic echo cancellation. IEEE Trans. Audio Speech Lang. Process.20(7), 2065\u20132079 (2012). https:\/\/doi.org\/10.1109\/TASL.2012.2196512.","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"221_CR28","doi-asserted-by":"crossref","unstructured":"D. Comminiello, M. Scarpiniti, L. A. Azpicueta-Ruiz, J. Arenas-Garc\u00eda, A. Uncini, in 2017 25th European Signal Processing Conference (EUSIPCO). Full proportionate functional link adaptive filters for nonlinear acoustic echo cancellation, (2017), pp. 1145\u20131149.","DOI":"10.23919\/EUSIPCO.2017.8081387"},{"issue":"6","key":"221_CR29","doi-asserted-by":"publisher","first-page":"1429","DOI":"10.1109\/TASL.2009.2035038","volume":"18","author":"E. A. Lehmann","year":"2009","unstructured":"E. A. Lehmann, A. M. Johansson, Diffuse reverberation model for efficient image-source simulation of room impulse responses. IEEE Trans. Audio Speech Lang. Process.18(6), 1429\u20131439 (2009).","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"issue":"5","key":"221_CR30","doi-asserted-by":"publisher","first-page":"802","DOI":"10.1109\/5.381848","volume":"83","author":"N. J. Kasdin","year":"1995","unstructured":"N. J. Kasdin, Discrete simulation of colored noise and stochastic processes and 1\/f\/sup\/spl alpha\/\/power law noise generation. Proc. IEEE. 83(5), 802\u2013827 (1995).","journal-title":"Proc. IEEE"},{"key":"221_CR31","doi-asserted-by":"crossref","unstructured":"K. Cho, B. Van Merri\u00ebnboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, Y. Bengio, Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078 (2014).","DOI":"10.3115\/v1\/D14-1179"},{"key":"221_CR32","unstructured":"D. P. Kingma, J. Ba, Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)."},{"key":"221_CR33","unstructured":"A. W. Rix, J. G. Beerends, M. P. Hollier, A. P. Hekstra, in 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.01CH37221), 2. Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs, (2001), pp. 749\u20137522."},{"issue":"4","key":"221_CR34","doi-asserted-by":"publisher","first-page":"1462","DOI":"10.1109\/TSA.2005.858005","volume":"14","author":"E. Vincent","year":"2006","unstructured":"E. Vincent, R. Gribonval, C. F\u00e9votte, Performance measurement in blind audio source separation. IEEE Trans. Audio Speech Lang. Process.14(4), 1462\u20131469 (2006).","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"221_CR35","unstructured":"C. Raffel, B. McFee, E. J. Humphrey, J. Salamon, O. Nieto, D. Liang, D. P. Ellis, C. C. Raffel, in In Proceedings of the 15th International Society for Music Information Retrieval Conference, ISMIR. mir_eval: a transparent implementation of common MIR metrics (Citeseer, 2014)."},{"key":"221_CR36","doi-asserted-by":"crossref","unstructured":"C. H. Taal, R. C. Hendriks, R. Heusdens, J. Jensen, in 2010 IEEE International Conference on Acoustics, Speech and Signal Processing. A short-time objective intelligibility measure for time-frequency weighted noisy speech, (2010), pp. 4214\u20134217.","DOI":"10.1109\/ICASSP.2010.5495701"}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-021-00221-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-021-00221-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-021-00221-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,8]],"date-time":"2023-01-08T20:12:33Z","timestamp":1673208753000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-021-00221-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,7]]},"references-count":36,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["221"],"URL":"https:\/\/doi.org\/10.1186\/s13636-021-00221-8","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,9,7]]},"assertion":[{"value":"5 April 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 August 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 September 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"35"}}