{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,11]],"date-time":"2025-09-11T19:50:28Z","timestamp":1757620228151,"version":"3.44.0"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,7,29]],"date-time":"2025-07-29T00:00:00Z","timestamp":1753747200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,7,29]],"date-time":"2025-07-29T00:00:00Z","timestamp":1753747200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Numerous post-processing methods have been proposed to improve coded speech quality and intelligibility. However, achieving state-of-the-art enhancement and generalisation across varying distortion levels remains a challenge. To bridge this gap, we propose a Lightweight Causal-Transformer-based Coded Speech Enhancement (LCT-CSE) model employing a causal frequency-time-frequency (FTF) transformer block. This block facilitates temporal and spectral sequential modelling using transformers, efficiently exploiting global dependency across causal-context TF bins while minimising computational overhead. Experimental results indicate that the proposed LCT-CSE model outperforms the considered baselines across mainstream lossy audio codecs, including Opus, AMR-WB, EVS and LC3+, with less footprint and complexity. To further utilise auxiliary, utterance-level information such as bitrate and other general distortion characteristics, building upon the LCT-CSE model, we propose two information incorporation methods. One employs one-hot vector representations and feature fusions, referred to as 1-hot vector-based modulation, while the other dynamically switches information-dependent network paths, termed dynamic linear modulation (DLM). These methods can be used to improve performance in bitrate-information utilisation, with negligible additional computational overhead. The DLM model even achieves comparable performance to bitrate-specific trained (BST) models. We further extend the proposed information incorporation method, DLM, to a generalised scenario, tandem coding. Compared to the two practically used approaches, the DLM-based LCT-CSE model consistently exhibits improved generalisability across varying tandem encoding conditions, based on derivative distortion information. Specifically, it achieves gains up to 0.74 in PESQ, 7% in STOI, and 0.18 in MOS-SIG under various bitrate conditions. This indicates significant potential for further applications where auxiliary information can be utilised.<\/jats:p>","DOI":"10.1186\/s13636-025-00420-7","type":"journal-article","created":{"date-parts":[[2025,7,29]],"date-time":"2025-07-29T12:44:22Z","timestamp":1753793062000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Coded speech enhancement using auxiliary utterance-level information"],"prefix":"10.1186","volume":"2025","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-3957-8859","authenticated-orcid":false,"given":"Haixin","family":"Zhao","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nilesh","family":"Madhu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,7,29]]},"reference":[{"key":"420_CR1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4419-1754-6","author":"Y You","year":"2010","unstructured":"Y. You, Audio coding: theory and applications (Springer. New York (2010). https:\/\/doi.org\/10.1007\/978-1-4419-1754-6","journal-title":"New York"},{"key":"420_CR2","doi-asserted-by":"publisher","unstructured":"J.M. Valin, K. Vos, T.B. Terriberry. Definition of the Opus Audio Codec. RFC 6716 (2012). https:\/\/doi.org\/10.17487\/RFC6716","DOI":"10.17487\/RFC6716"},{"issue":"8","key":"420_CR3","doi-asserted-by":"publisher","first-page":"620","DOI":"10.1109\/TSA.2002.804299","volume":"10","author":"B Bessette","year":"2002","unstructured":"B. Bessette, R. Salami, R. Lefebvre, M. Jelinek, J. Rotola-Pukkila, J. Vainio, H. Mikkola, K. Jarvinen, The adaptive multirate wideband speech codec (amr-wb). IEEE Trans. Speech Audio Process. 10(8), 620\u2013636 (2002). https:\/\/doi.org\/10.1109\/TSA.2002.804299","journal-title":"IEEE Trans. Speech Audio Process."},{"key":"420_CR4","doi-asserted-by":"publisher","unstructured":"S. Bruhn, H. Pobloth, M. Schnell, B. Grill, J. Gibbs, L. Miao, K. J\u00e4rvinen, L. Laaksonen, N. Harada, N. Naka, S. Ragot, S. Proust, T. Sanda, I. Varga, C. Greer, M. Jel\u00ednek, M. Xie, P. Usai, in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Standardization of the new 3GPP EVS codec (2015), pp. 5703\u20135707. https:\/\/doi.org\/10.1109\/ICASSP.2015.7179064","DOI":"10.1109\/ICASSP.2015.7179064"},{"key":"420_CR5","unstructured":"M. Schnell, E. Ravelli, J. B\u00fcthe, M. Schlegel, A. Tomasek, A. Tschekalinskij, J. Svedberg, M. Sehlstedt, in Audio Engineering Society Convention 150. LC3 and LC3plus: The new audio transmission standards for wireless communication (Audio Engineering Society, New York, 2021)"},{"key":"420_CR6","doi-asserted-by":"publisher","unstructured":"J.N. Antons, R. Schleicher, S. Arndt, S. M\u00f6ller, G. Curio, in 2012 Fourth International Workshop on Quality of Multimedia Experience. Too tired for calling? A physiological measure of fatigue caused by bandwidth limitations (2012), pp. 63\u201367. https:\/\/doi.org\/10.1109\/QoMEX.2012.6263840","DOI":"10.1109\/QoMEX.2012.6263840"},{"key":"420_CR7","doi-asserted-by":"publisher","unstructured":"J. Skoglund, J.M. Valin, in Interspeech 2020. Improving opus low bit rate quality with neural speech synthesis (2020), pp. 2847\u20132851. https:\/\/doi.org\/10.21437\/Interspeech.2020-2939","DOI":"10.21437\/Interspeech.2020-2939"},{"key":"420_CR8","doi-asserted-by":"publisher","first-page":"121532","DOI":"10.1109\/ACCESS.2021.3108784","volume":"9","author":"S Hwang","year":"2021","unstructured":"S. Hwang, Y. Cheon, S. Han, I. Jang, J.W. Shin, Enhancement of coded speech using neural network-based side information. IEEE Access 9, 121532\u2013121540 (2021). https:\/\/doi.org\/10.1109\/ACCESS.2021.3108784","journal-title":"IEEE Access"},{"key":"420_CR9","doi-asserted-by":"crossref","unstructured":"A. Mustafa, J. B\u00fcthe, S. Korse, K. Gupta, G. Fuchs, N. Pia, in 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). A streamwise gan vocoder for wideband speech coding at very low bit rate (IEEE, New York, 2021), pp. 66\u201370","DOI":"10.1109\/WASPAA52581.2021.9632750"},{"key":"420_CR10","doi-asserted-by":"crossref","unstructured":"J. B\u00fcthe, J.M. Valin, A. Mustafa, in 2023 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). LACE: A light-weight, causal model for enhancing coded speech through adaptive convolutions (IEEE, New York, 2023), pp. 1\u20135","DOI":"10.1109\/WASPAA58266.2023.10248150"},{"key":"420_CR11","doi-asserted-by":"publisher","unstructured":"J. B\u00fcthe, A. Mustafa, J.M. Valin, K. Helwani, M.M. Goodwin, in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). NOLACE: Improving low-complexity speech codec enhancement through adaptive temporal shaping (2024), pp. 476\u2013480. https:\/\/doi.org\/10.1109\/ICASSP48485.2024.10448332","DOI":"10.1109\/ICASSP48485.2024.10448332"},{"issue":"1","key":"420_CR12","doi-asserted-by":"publisher","first-page":"59","DOI":"10.1109\/89.365380","volume":"3","author":"JH Chen","year":"1995","unstructured":"J.H. Chen, A. Gersho, Adaptive postfiltering for quality enhancement of coded speech. IEEE Trans. Speech Audio Process. 3(1), 59\u201371 (1995)","journal-title":"IEEE Trans. Speech Audio Process."},{"key":"420_CR13","doi-asserted-by":"crossref","unstructured":"G. Fuchs, C.R. Helmrich, G. Markovi\u0107, M. Neusinger, E. Ravelli, T. Moriya, in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Low delay LPC and MDCT-based audio coding in the EVS codec (IEEE, New York, 2015), pp. 5723\u20135727","DOI":"10.1109\/ICASSP.2015.7179068"},{"key":"420_CR14","doi-asserted-by":"crossref","unstructured":"T. Vaillancourt, R. Salami, M. Jelinek, in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). New post-processing techniques for low bit rate celp codecs (IEEE, New York, 2015), pp. 5908\u20135912","DOI":"10.1109\/ICASSP.2015.7179105"},{"key":"420_CR15","doi-asserted-by":"crossref","unstructured":"Z. Zhao, S. Elshamy, H. Liu, T. Fingscheidt, in 2018 16th International Workshop on Acoustic Signal Enhancement (IWAENC). A CNN postprocessor to enhance coded speech (IEEE, New York, 2018), pp. 406\u2013410","DOI":"10.1109\/IWAENC.2018.8521389"},{"issue":"4","key":"420_CR16","doi-asserted-by":"publisher","first-page":"663","DOI":"10.1109\/TASLP.2018.2887337","volume":"27","author":"Z Zhao","year":"2018","unstructured":"Z. Zhao, H. Liu, T. Fingscheidt, Convolutional neural networks to enhance coded speech. IEEE\/ACM Trans. Audio Speech Lang. Process. 27(4), 663\u2013678 (2018)","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"420_CR17","doi-asserted-by":"crossref","unstructured":"S. Korse, K. Gupta, G. Fuchs, in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Enhancement of coded speech using a mask-based post-filter (IEEE, New York, 2020), pp. 6764\u20136768","DOI":"10.1109\/ICASSP40776.2020.9053283"},{"key":"420_CR18","doi-asserted-by":"crossref","unstructured":"K. Gupta, S. Korse, B. Edler, G. Fuchs, in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). A DNN based post-filter to enhance the quality of coded speech in mdct domain (IEEE, New York, 2022), pp. 836\u2013840","DOI":"10.1109\/ICASSP43922.2022.9747410"},{"key":"420_CR19","doi-asserted-by":"crossref","unstructured":"S. Korse, N. Pia, K. Gupta, G. Fuchs, in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Postgan: A GAN-based post-processor to enhance the quality of coded speech (IEEE, New York, 2022), pp. 831\u2013835","DOI":"10.1109\/ICASSP43922.2022.9747733"},{"key":"420_CR20","doi-asserted-by":"crossref","unstructured":"H. Zhao, N. Madhu, in IEEE International Conference on Electronics, Information, and Communication (ICEIC). Bitrate-informed coded speech enhancement model (IEEE, New York, 2024), pp. 666\u2013669","DOI":"10.1109\/ICEIC61013.2024.10457098"},{"key":"420_CR21","doi-asserted-by":"publisher","unstructured":"Y.X. Lu, Y. Ai, Z.H. Ling, in INTERSPEECH 2023. MP-SENet: A speech enhancement model with parallel denoising of magnitude and phase spectra (2023), pp. 3834\u20133838. https:\/\/doi.org\/10.21437\/Interspeech.2023-1441","DOI":"10.21437\/Interspeech.2023-1441"},{"key":"420_CR22","doi-asserted-by":"publisher","unstructured":"J. Chen, Q. Mao, D. Liu, in Interspeech 2020. Dual-path transformer network: Direct context-aware modeling for end-to-end monaural speech separation (2020), pp. 2642\u20132646. https:\/\/doi.org\/10.21437\/Interspeech.2020-2205","DOI":"10.21437\/Interspeech.2020-2205"},{"key":"420_CR23","doi-asserted-by":"crossref","unstructured":"K. Wang, B. He, W.P. Zhu, in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Tstnn: Two-stage transformer based neural network for speech enhancement in the time domain (IEEE, New York, 2021), pp. 7098\u20137102","DOI":"10.1109\/ICASSP39728.2021.9413740"},{"key":"420_CR24","doi-asserted-by":"crossref","unstructured":"A. Biswas, D. Jia, in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Audio codec enhancement with generative adversarial networks (IEEE, New York, 2020), pp. 356\u2013360","DOI":"10.1109\/ICASSP40776.2020.9053113"},{"key":"420_CR25","unstructured":"H. Zhao, N. Madhu, in 2025 33rd European Signal Processing Conference (EUSIPCO). Study of lightweight transformer architectures for single-channel speech enhancement (2025). Available:  https:\/\/arxiv.org\/abs\/2505.21057. Accessed\u00a023 July 2025."},{"key":"420_CR26","doi-asserted-by":"publisher","unstructured":"P. Motlicek, V. Ullal, H. Hermansky, in 2007 IEEE International Conference on Acoustics, Speech and Signal Processing - ICASSP \u201907. Wide-band perceptual audio coding based on frequency-domain linear prediction, vol. 1 (2007), pp. I\u2013265\u2013I\u2013268. https:\/\/doi.org\/10.1109\/ICASSP.2007.366667","DOI":"10.1109\/ICASSP.2007.366667"},{"key":"420_CR27","doi-asserted-by":"publisher","unstructured":"H. Wu, K. Tan, B. Xu, A. Kumar, D. Wong, in INTERSPEECH 2023. Rethinking complex-valued deep neural networks for monaural speech enhancement (2023), pp. 3889\u20133893. https:\/\/doi.org\/10.21437\/Interspeech.2023-686","DOI":"10.21437\/Interspeech.2023-686"},{"key":"420_CR28","doi-asserted-by":"publisher","first-page":"2018","DOI":"10.1109\/LSP.2021.3116502","volume":"28","author":"ZQ Wang","year":"2021","unstructured":"Z.Q. Wang, G. Wichern, J. Le Roux, On the compensation between magnitude and phase in speech separation. IEEE Signal Process. Lett. 28, 2018\u20132022 (2021). https:\/\/doi.org\/10.1109\/LSP.2021.3116502","journal-title":"IEEE Signal Process. Lett."},{"key":"420_CR29","doi-asserted-by":"crossref","unstructured":"S. Braun, H. Gamper, C.K. Reddy, I. Tashev, in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Towards efficient models for real-time deep noise suppression (IEEE, New York, 2021), pp. 656\u2013660","DOI":"10.1109\/ICASSP39728.2021.9413580"},{"key":"420_CR30","doi-asserted-by":"publisher","unstructured":"K. Tesch, N.H. Mohrmann, T. Gerkmann, in Interspeech 2022. On the role of spatial, spectral, and temporal processing for DNN-based non-linear multi-channel speech enhancement (2022), pp. 2908\u20132912. https:\/\/doi.org\/10.21437\/Interspeech.2022-162","DOI":"10.21437\/Interspeech.2022-162"},{"issue":"8","key":"420_CR31","doi-asserted-by":"publisher","first-page":"1276","DOI":"10.1109\/LSP.2018.2849578","volume":"25","author":"J Lee","year":"2018","unstructured":"J. Lee, J. Skoglund, T. Shabestary, H.G. Kang, Phase-sensitive joint learning algorithms for deep learning-based speech enhancement. IEEE Signal Process. Lett. 25(8), 1276\u20131280 (2018)","journal-title":"IEEE Signal Process. Lett."},{"key":"420_CR32","doi-asserted-by":"crossref","unstructured":"K. Wilson, M. Chinen, J. Thorpe, B. Patton, J. Hershey, R.A. Saurous, J. Skoglund, R.F. Lyon, in 2018 16th International Workshop on Acoustic Signal Enhancement (IWAENC). Exploring tradeoffs in models for low-latency speech enhancement (IEEE, 2018), pp. 366\u2013370","DOI":"10.1109\/IWAENC.2018.8521347"},{"key":"420_CR33","doi-asserted-by":"publisher","unstructured":"S. Wisdom, J.R. Hershey, K. Wilson, J. Thorpe, M. Chinen, B. Patton, R.A. Saurous, in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Differentiable consistency constraints for improved deep speech enhancement (2019), pp. 900\u2013904. https:\/\/doi.org\/10.1109\/ICASSP.2019.8682783","DOI":"10.1109\/ICASSP.2019.8682783"},{"key":"420_CR34","doi-asserted-by":"publisher","unstructured":"C.K. Reddy, H. Dubey, K. Koishida, A. Nair, V. Gopal, R. Cutler, S. Braun, H. Gamper, R. Aichner, S. Srinivasan, in Interspeech 2021. Interspeech 2021 deep noise suppression challenge (2021), pp. 2796\u20132800. https:\/\/doi.org\/10.21437\/Interspeech.2021-1609","DOI":"10.21437\/Interspeech.2021-1609"},{"key":"420_CR35","unstructured":"International Telecommunication Union, Wideband extension to Recommendation P.862 for the assessment of wideband telephone networks and speech codecs (ITU, Geneva, 2007), ITU-T Recommendation P.862.2"},{"key":"420_CR36","doi-asserted-by":"crossref","unstructured":"C.H. Taal, R.C. Hendriks, R. Heusdens, J. Jensen, in 2010 IEEE international conference on acoustics, speech and signal processing. A short-time objective intelligibility measure for time-frequency weighted noisy speech (IEEE, New York, 2010), pp. 4214\u20134217","DOI":"10.1109\/ICASSP.2010.5495701"},{"key":"420_CR37","doi-asserted-by":"crossref","unstructured":"C.K. Reddy, V. Gopal, R. Cutler, in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). DNSMOS p. 835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors (IEEE, New York, 2022), pp. 886\u2013890","DOI":"10.1109\/ICASSP43922.2022.9746108"},{"key":"420_CR38","unstructured":"ITU-R, Method for the subjective assessment of intermediate quality level of audio systems. Technical Report BS.1534-3, International Telecommunication Union, Radiocommunication Sector (ITU, Geneva, 2015)"}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-025-00420-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-025-00420-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-025-00420-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,8]],"date-time":"2025-09-08T06:38:33Z","timestamp":1757313513000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-025-00420-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,29]]},"references-count":38,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["420"],"URL":"https:\/\/doi.org\/10.1186\/s13636-025-00420-7","relation":{},"ISSN":["1687-4722"],"issn-type":[{"type":"electronic","value":"1687-4722"}],"subject":[],"published":{"date-parts":[[2025,7,29]]},"assertion":[{"value":"7 February 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 July 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 July 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"29"}}