{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T02:20:09Z","timestamp":1760149209155,"version":"build-2065373602"},"reference-count":28,"publisher":"MDPI AG","issue":"14","license":[{"start":{"date-parts":[[2023,7,16]],"date-time":"2023-07-16T00:00:00Z","timestamp":1689465600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Using the source-filter model of speech production, clean speech signals can be decomposed into an excitation component and an envelope component that is related to the phoneme being uttered. Therefore, restoring the envelope of degraded speech during speech enhancement can improve the intelligibility and quality of output. As the number of phonemes in spoken speech is limited, they can be adequately represented by a correspondingly limited number of envelopes. This can be exploited to improve the estimation of speech envelopes from a degraded signal in a data-driven manner. The improved envelopes are then used in a second stage to refine the final speech estimate. Envelopes are typically derived from the linear prediction coefficients (LPCs) or from the cepstral coefficients (CCs). The improved envelope is obtained either by mapping the degraded envelope onto pre-trained codebooks (classification approach) or by directly estimating it from the degraded envelope (regression approach). In this work, we first investigate the optimal features for envelope representation and codebook generation by a series of oracle tests. We demonstrate that CCs provide better envelope representation compared to using the LPCs. Further, we demonstrate that a unified speech codebook is advantageous compared to the typical codebook that manually splits speech and silence as separate entries. Next, we investigate low-complexity neural network architectures to map degraded envelopes to the optimal codebook entry in practical systems. We confirm that simple recurrent neural networks yield good performance with a low complexity and number of parameters. We also demonstrate that with a careful choice of the feature and architecture, a regression approach can further improve the performance at a lower computational cost. However, as also seen from the oracle tests, the benefit of the two-stage framework is now chiefly limited by the statistical noise floor estimate, leading to only a limited improvement in extremely adverse conditions. This highlights the need for further research on joint estimation of speech and noise for optimum enhancement.<\/jats:p>","DOI":"10.3390\/s23146438","type":"journal-article","created":{"date-parts":[[2023,7,17]],"date-time":"2023-07-17T01:06:36Z","timestamp":1689555996000},"page":"6438","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Investigations on the Optimal Estimation of Speech Envelopes for the Two-Stage Speech Enhancement"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9684-6611","authenticated-orcid":false,"given":"Yanjue","family":"Song","sequence":"first","affiliation":[{"name":"IDLab, Ghent University\u2014imec, 9000 Gent, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9131-3309","authenticated-orcid":false,"given":"Nilesh","family":"Madhu","sequence":"additional","affiliation":[{"name":"IDLab, Ghent University\u2014imec, 9000 Gent, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,7,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1109","DOI":"10.1109\/TASSP.1984.1164453","article-title":"Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator","volume":"32","author":"Ephraim","year":"1984","journal-title":"IEEE Trans. Acoust. Speech Signal Process."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"443","DOI":"10.1109\/TASSP.1985.1164550","article-title":"Speech enhancement using a minimum mean-square error log-spectral amplitude estimator","volume":"33","author":"Ephraim","year":"1985","journal-title":"IEEE Trans. Acoust. Speech Signal Process."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"2098","DOI":"10.1109\/TASL.2006.872621","article-title":"Improved signal-to-noise ratio estimation for speech enhancement","volume":"14","author":"Plapous","year":"2006","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"2492","DOI":"10.1109\/TASLP.2022.3190725","article-title":"Improved CEM for speech harmonic enhancement in single channel noise suppression","volume":"30","author":"Song","year":"2022","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1592","DOI":"10.1109\/TASLP.2017.2702385","article-title":"Instantaneous a priori SNR estimation by cepstral excitation manipulation","volume":"25","author":"Elshamy","year":"2017","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"101329","DOI":"10.1016\/j.csl.2021.101329","article-title":"Prediction of speech intelligibility with DNN-based performance measures","volume":"74","author":"Martinez","year":"2022","journal-title":"Comput. Speech Lang."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Arai, K., Araki, S., Ogawa, A., Kinoshita, K., Nakatani, T., Yamamoto, K., and Irino, T. (2019, January 15\u201319). Predicting Speech Intelligibility of Enhanced Speech Using Phone Accuracy of DNN-Based ASR System. Proceedings of the Interspeech, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-1381"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1016\/j.specom.2012.07.003","article-title":"Artificial bandwidth extension of spectral envelope along a Viterbi path","volume":"55","author":"Turan","year":"2013","journal-title":"Speech Commun."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Taal, C.H., Hendriks, R.C., Heusdens, R., and Jensen, J. (2010, January 14\u201319). A short-time objective intelligibility measure for time-frequency weighted noisy speech. Proceedings of the 2010 IEEE International Conference on Acoustics, Speech and Signal Processing, Dallas, TX, USA.","DOI":"10.1109\/ICASSP.2010.5495701"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"163","DOI":"10.1109\/TSA.2005.854113","article-title":"Codebook driven short-term predictor parameter estimation for speech enhancement","volume":"14","author":"Srinivasan","year":"2005","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"2460","DOI":"10.1109\/TASLP.2018.2867947","article-title":"DNN-supported speech enhancement with cepstral estimation of both excitation and envelope","volume":"26","author":"Elshamy","year":"2018","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1324","DOI":"10.1109\/TASL.2011.2177821","article-title":"Model-based speech enhancement with improved spectral envelope estimation via dynamics tracking","volume":"20","author":"Chen","year":"2011","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"469","DOI":"10.1049\/iet-spr.2016.0477","article-title":"Deep neural network-based linear predictive parameter estimations for speech enhancement","volume":"11","author":"Li","year":"2017","journal-title":"IET Signal Process."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Kolb\u00e6k, M., Tan, Z.H., and Jensen, J. (2018, January 15\u201320). Monaural speech enhancement using deep neural networks by maximizing a short-time objective intelligibility measure. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8462040"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Valin, J.M., Isik, U., Phansalkar, N., Giri, R., Helwani, K., and Krishnaswamy, A. (2020). A perceptually-motivated approach for low-complexity, real-time enhancement of fullband speech. arXiv.","DOI":"10.21437\/Interspeech.2020-2730"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"563","DOI":"10.1109\/LSP.2013.2255125","article-title":"Speech spectral envelope enhancement by HMM-based analysis\/resynthesis","volume":"20","author":"Carmona","year":"2013","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"561","DOI":"10.1109\/PROC.1975.9792","article-title":"Linear prediction: A tutorial review","volume":"63","author":"Makhoul","year":"1975","journal-title":"Proc. IEEE"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1036","DOI":"10.1109\/LSP.2007.906208","article-title":"Cepstral smoothing of spectral filter gains for speech enhancement without musical noise","volume":"14","author":"Breithaupt","year":"2007","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Breithaupt, C., Gerkmann, T., and Martin, R. (April, January 31). A novel a priori SNR estimation approach based on selective cepstro-temporal smoothing. Proceedings of the 2008 IEEE International Conference on Acoustics, Speech and Signal Processing, Las Vegas, NV, USA.","DOI":"10.1109\/ICASSP.2008.4518755"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"983","DOI":"10.1109\/TASL.2013.2243434","article-title":"Corpus-based speech enhancement with uncertainty modeling and cepstral smoothing","volume":"21","author":"Nickel","year":"2013","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_21","unstructured":"Chung, J., Gulcehre, C., Cho, K., and Bengio, Y. (2014). Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Braun, S., Gamper, H., Reddy, C.K., and Tashev, I. (2021, January 6\u201312). Towards efficient models for real-time deep noise suppression. Proceedings of the 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2021), Virtual.","DOI":"10.1109\/ICASSP39728.2021.9413580"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Song, Y., and Madhu, N. Aiding speech harmonic recovery in DNN-based single channel noise reduction using cepstral excitation manipulation (CEM) components (in press). Proceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece.","DOI":"10.1109\/ICASSP49357.2023.10096868"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1109\/TCOM.1980.1094577","article-title":"An algorithm for vector quantizer design","volume":"28","author":"Linde","year":"1980","journal-title":"IEEE Trans. Commun."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Vary, P., and Martin, R. (2006). Digital Speech Transmission: Enhancement, Coding and Error Concealment, John Wiley & Sons.","DOI":"10.1002\/0470031743"},{"key":"ref_26","unstructured":"(2011). Objective Measurement of Active Speech Level, International Telecommunication Union-Telecommunication Standardisation Sector (ITU). Number Rec. P.56."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1383","DOI":"10.1109\/TASL.2011.2180896","article-title":"Unbiased MMSE-based noise power estimation with low complexity and low tracking delay","volume":"20","author":"Gerkmann","year":"2011","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_28","unstructured":"(2017). Corrigendum 1, Wideband Extension to Recommendation P.862 for the Assessment of Wideband Telephone Networks and Speech Codecs, International Telecommunication Union-Telecommunication Standardisation Sector (ITU). Number Rec. P.862.2."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/14\/6438\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:12:54Z","timestamp":1760127174000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/14\/6438"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,16]]},"references-count":28,"journal-issue":{"issue":"14","published-online":{"date-parts":[[2023,7]]}},"alternative-id":["s23146438"],"URL":"https:\/\/doi.org\/10.3390\/s23146438","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2023,7,16]]}}}