{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,25]],"date-time":"2025-12-25T07:23:30Z","timestamp":1766647410876,"version":"3.37.3"},"reference-count":50,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T00:00:00Z","timestamp":1701734400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T00:00:00Z","timestamp":1701734400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"RFI OIC"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The task of bandwidth extension addresses the generation of missing high frequencies of audio signals based on knowledge of the low-frequency part of the sound. This task applies to various problems, such as audio coding or audio restoration. In this article, we focus on efficient bandwidth extension of monophonic and polyphonic musical signals using a differentiable digital signal processing (DDSP) model. Such a model is composed of a neural network part with relatively few parameters trained to infer the parameters of a differentiable digital signal processing model, which efficiently generates the output full-band audio signal. <\/jats:p><jats:p>We first address bandwidth extension of monophonic signals, and then propose two methods to explicitly handle polyphonic signals. The benefits of the proposed models are first demonstrated on monophonic and polyphonic synthetic data against a baseline and a deep-learning-based ResNet model. The models are next evaluated on recorded monophonic and polyphonic data, for a wide variety of instruments and musical genres. We show that all proposed models surpass a higher complexity deep learning model for an objective metric computed in the frequency domain. A MUSHRA listening test confirms the superiority of the proposed approach in terms of perceptual quality.<\/jats:p>","DOI":"10.1186\/s13636-023-00315-5","type":"journal-article","created":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T15:01:50Z","timestamp":1701788510000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Efficient bandwidth extension of musical signals using a differentiable harmonic plus noise model"],"prefix":"10.1186","volume":"2023","author":[{"given":"Pierre-Amaury","family":"Grumiaux","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1253-4427","authenticated-orcid":false,"given":"Mathieu","family":"Lagrange","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,12,5]]},"reference":[{"key":"315_CR1","doi-asserted-by":"crossref","unstructured":"E. Vincent, T. Virtanen, S. Gannot (eds.), Audio Source Separation and Speech Enhancement (Wiley, 2018)","DOI":"10.1002\/9781119279860"},{"key":"315_CR2","doi-asserted-by":"publisher","first-page":"922","DOI":"10.1109\/TASL.2011.2168211","volume":"20","author":"A Adler","year":"2012","unstructured":"A. Adler, V. Emiya, M.G. Jafari, M. Elad, R. Gribonval, M.D. Plumbley, Audio inpainting. IEEE Trans. Audio Speech Lang. Process 20, 922\u2013932 (2012)","journal-title":"IEEE Trans. Audio Speech Lang. Process"},{"key":"315_CR3","doi-asserted-by":"crossref","unstructured":"N.R. French, J.C. Steinberg, Factors governing the intelligibility of speech sounds. J. Acoust. Soc. Am.\u00a019, 90\u2013119\u00a0(1947)","DOI":"10.1121\/1.1916407"},{"key":"315_CR4","unstructured":"S.V. Vaseghi, R. Frayling-Cork, Restoration of old gramophone recordings. J. Audio Eng. Soc.\u00a040, 791\u2013801 (1992)"},{"key":"315_CR5","doi-asserted-by":"crossref","unstructured":"C. Gaultier, S. Kiti\u0107, R. Gribonval, N. Bertin, Sparsity-based audio declipping methods: selected overview, new algorithms, and large-scale evaluation. IEEE\/ACM Trans. Audio Speech Lang. Process.\u00a029, 1174\u20131187 (2021)","DOI":"10.1109\/TASLP.2021.3059264"},{"key":"315_CR6","unstructured":"M. Dietz, L. Liljeryd, K. Kjorling, O. Kunz, in Audio Engineering Society Convention, Spectral band replication, a novel approach in audio coding (2002)"},{"key":"315_CR7","doi-asserted-by":"crossref","unstructured":"Y. Ning, S. He, Z. Wu, C. Xing, L.-J. Zhang, A review of deep learning based speech synthesis. Appl. Sci.\u00a09, 4050\u20134066 (2019)","DOI":"10.3390\/app9194050"},{"key":"315_CR8","doi-asserted-by":"crossref","unstructured":"S.H. Mohammadi, A. Kain, in Speech Communication, An overview of voice conversion systems (2017)","DOI":"10.1016\/j.specom.2017.01.008"},{"key":"315_CR9","unstructured":"A.v.d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, K. Kavukcuoglu, Wavenet: a generative model for raw audio. Proc. ISCA (2016)"},{"key":"315_CR10","doi-asserted-by":"crossref","unstructured":"E. Moliner, J. Lehtinen, V. V\u00e4lim\u00e4ki, in International Conference on Acoustics, Speech and Signal Processing, Solving audio inverse problems with a diffusion model (2023)","DOI":"10.1109\/ICASSP49357.2023.10095637"},{"key":"315_CR11","unstructured":"J. Engel, L. Hantrakul, C. Gu, A. Roberts, DDSP: Differentiable Digital Signal Processing. (International Conference on Learning Representations, 2020)"},{"key":"315_CR12","doi-asserted-by":"crossref","unstructured":"S. Sulun, M.E.P. Davies, On filter generalization for music bandwidth extension using deep neural networks. J. Sel. Top. Signal Process.\u00a015, 132\u2013142 (2021)","DOI":"10.1109\/JSTSP.2020.3037485"},{"key":"315_CR13","unstructured":"S. Meltzer, R. Bohm, F. Henn, in Audio Engineering Society Convention, SBR enhanced audio codecs for digital broadcasting such as \u201cDigital Radio Mondiale\u201d (DRM) (2002)"},{"key":"315_CR14","doi-asserted-by":"crossref","unstructured":"F. Nagel, S. Disch, in International Conference on Acoustics, Speech and Signal Processing, A harmonic bandwidth extension method for audio codecs (2009)","DOI":"10.1109\/ICASSP.2009.4959541"},{"key":"315_CR15","unstructured":"S. Chennoukh, A. Gerrits, G. Miet, R. Sluijter, in International Conference on Acoustics, Speech, and Signal Processing. Proceedings, Speech enhancement via frequency bandwidth extension using line spectral frequencies (2001)"},{"key":"315_CR16","doi-asserted-by":"crossref","unstructured":"J. Sadasivan, S. Mukherjee, C.S. Seelamantula, in International Conference on Acoustics, Speech and Signal Processing, Joint dictionary training for bandwidth extension of speech signals (2016)","DOI":"10.1109\/ICASSP.2016.7472814"},{"key":"315_CR17","doi-asserted-by":"crossref","unstructured":"Y. Yoshida, M. Abe, in International Conference on Spoken Language Processing, An algorithm to reconstruct wideband speech from narrowband speech based on codebook mapping (1994)","DOI":"10.21437\/ICSLP.1994-412"},{"key":"315_CR18","unstructured":"K.-Y. Park, H.S. Kim, in International Conference on Acoustics, Speech, and Signal Processing. Proceedings, Narrowband to wideband conversion of speech using GMM based transformation (2000)"},{"key":"315_CR19","doi-asserted-by":"crossref","unstructured":"P. Bauer, T. Fingscheidt, in International Conference on Acoustics, Speech and Signal Processing, An HMM-based artificial bandwidth extension evaluated by cross-language training and test (2008)","DOI":"10.1109\/ICASSP.2008.4518678"},{"key":"315_CR20","doi-asserted-by":"crossref","unstructured":"G.-B. Song, P. Martynovich, A study of HMM-based bandwidth extension of speech signals. Signal Process.\u00a089, 2036\u20132044 (2009)","DOI":"10.1016\/j.sigpro.2009.03.037"},{"key":"315_CR21","doi-asserted-by":"crossref","unstructured":"D. Bansal, B. Raj, P. Smaragdis, in Interspeech, Bandwidth expansion of narrowband speech using non-negative matrix factorization (2005)","DOI":"10.21437\/Interspeech.2005-528"},{"key":"315_CR22","doi-asserted-by":"crossref","unstructured":"D.L. Sun, R. Mazumder, in International Workshop on Machine Learning for Signal Processing, Non-negative matrix completion for bandwidth extension: A convex optimization approach (2013)","DOI":"10.1109\/MLSP.2013.6661924"},{"key":"315_CR23","doi-asserted-by":"crossref","unstructured":"K. Li, C.-H. Lee, in International Conference on Acoustics, Speech and Signal Processing, A deep neural network approach to speech bandwidth expansion (2015)","DOI":"10.1109\/ICASSP.2015.7178801"},{"key":"315_CR24","doi-asserted-by":"crossref","unstructured":"K. Li, Z. Huang, Y. Xu, C.-H. Lee, in Interspeech, DNN-based speech bandwidth expansion and its application to adding high-frequency missing features for automatic speech recognition of narrowband speech (2015)","DOI":"10.21437\/Interspeech.2015-555"},{"key":"315_CR25","doi-asserted-by":"crossref","unstructured":"Y. Wang, S. Zhao, W. Liu, M. Li, J. Kuang, in Interspeech, Speech bandwidth expansion based on deep neural networks (2015)","DOI":"10.21437\/Interspeech.2015-558"},{"key":"315_CR26","doi-asserted-by":"crossref","unstructured":"Y. Gu, Z.-H. Ling, in Interspeech, Waveform modeling using stacked dilated convolutional neural networks for speech bandwidth extension (2017)","DOI":"10.21437\/Interspeech.2017-336"},{"key":"315_CR27","unstructured":"V. Kuleshov, S.Z. Enam, S. Ermon, Audio super resolution using neural networks. (International Conference on Learning Representations, 2017)"},{"key":"315_CR28","doi-asserted-by":"crossref","unstructured":"H. Wang, D. Wang, in International Conference on Acoustics, Speech and Signal Processing, Time-frequency loss for CNN based speech super-resolution (2020)","DOI":"10.1109\/ICASSP40776.2020.9053712"},{"key":"315_CR29","unstructured":"G. Campos, N. Fonseca, A. Ferreira, M. Davies, in Conference on Digital Audio Effects, High frequency magnitude spectrogram reconstruction for music mixtures using convolutional autoencoders (2018)"},{"key":"315_CR30","doi-asserted-by":"crossref","unstructured":"M. Lagrange, F. Gontier, in International Conference on Acoustics, Speech and Signal Processing, Bandwidth extension of musical audio signals with no side information using dilated convolutional neural networks (2020)","DOI":"10.1109\/ICASSP40776.2020.9054194"},{"key":"315_CR31","doi-asserted-by":"crossref","unstructured":"S. Li, S. Villette, P. Ramadas, D.J. Sinder, in International Conference on Acoustics, Speech and Signal Processing, Speech bandwidth extension using generative adversarial networks (2018)","DOI":"10.1109\/ICASSP.2018.8462588"},{"key":"315_CR32","doi-asserted-by":"crossref","unstructured":"X. Li, V. Chebiyyam, K. Kirchhoff, in Interspeech, Speech audio super-resolution for speech recognition (2019)","DOI":"10.21437\/Interspeech.2019-3043"},{"key":"315_CR33","doi-asserted-by":"crossref","unstructured":"J. Su, Y. Wang, A. Finkelstein, Z. Jin, in International Conference on Acoustics, Speech and Signal Processing, Bandwidth extension is all you need (2021)","DOI":"10.1109\/ICASSP39728.2021.9413575"},{"key":"315_CR34","doi-asserted-by":"crossref","unstructured":"E. Moliner, V. V\u00e4lim\u00e4ki, BEHM-GAN: Bandwidth Extension of Historical Music using Generative Adversarial Networks. IEEE\/ACM Trans. Audio Speech Lang. Process\u00a031, 943\u2013956\u00a0(2022)","DOI":"10.1109\/TASLP.2022.3190726"},{"key":"315_CR35","unstructured":"J. Song, C. Meng, S. Ermon, Denoising diffusion implicit models. International Conference on Learning Representations (2020)"},{"key":"315_CR36","unstructured":"E. Moliner, F. Elvander, V. V\u00e4lim\u00e4ki, Zero-shot blind audio bandwidth extension. arXiv preprint arXiv:2306.01433 (2023)"},{"key":"315_CR37","unstructured":"B. Hayes, C. Saitis, G. Fazekas, Neural waveshaping synthesis (International Conference on Music Information Retrieval, 2021)"},{"key":"315_CR38","doi-asserted-by":"crossref","unstructured":"S. Shan, L. Hantrakul, J. Chen, M. Avent, D. in International Conference on Acoustics, Speech and Signal Processing, Trevelyan, Differentiable wavetable synthesis (2021)","DOI":"10.1109\/ICASSP43922.2022.9746940"},{"key":"315_CR39","doi-asserted-by":"crossref","unstructured":"S. Lee, H.-S. Choi, K. Lee, Differentiable artificial reverberation. IEEE\/ACM Trans. Audio Speech Lang. Process\u00a030, 2541\u20132556 (2022)","DOI":"10.1109\/TASLP.2022.3193298"},{"key":"315_CR40","doi-asserted-by":"crossref","unstructured":"C.J. Steinmetz, N.J. Bryan, J.D. Reiss, Style transfer of audio effects with differentiable signal processing. J. Audio Eng. Soc\u00a070, 708\u2013721 (2022)","DOI":"10.17743\/jaes.2022.0025"},{"key":"315_CR41","unstructured":"N. Masuda, D. Saito, in International Society for Music Information Retrieval Conference, Synthesizer sound matching with differentiable DSP (2021)"},{"key":"315_CR42","doi-asserted-by":"crossref","unstructured":"F. Esqueda, B. Kuznetsov, J.D. Parker, in International Conference on Digital Audio Effects, Differentiable white-box virtual analog modeling (2021)","DOI":"10.23919\/DAFx51585.2021.9768272"},{"key":"315_CR43","doi-asserted-by":"crossref","unstructured":"J.W. Kim, J. Salamon, P. Li, J.P. Bello, in International Conference on Acoustics, Speech and Signal Processing, Crepe: a Convolutional Representation for Pitch Estimation (2018)","DOI":"10.1109\/ICASSP.2018.8461329"},{"key":"315_CR44","unstructured":"L. Hantrakul, J. Engel, A. Roberts, C. Gu, in International Society for Music Information Retrieval Conference, Fast and flexible neural audio synthesis (2019)"},{"key":"315_CR45","doi-asserted-by":"crossref","unstructured":"R.M. Bittner, J.J. Bosch, D. Rubinstein, G. Meseguer-Brocal, S. Ewert, in International Conference on Acoustics, Speech and Signal Processing, A lightweight instrument-agnostic Model for polyphonic note transcription and multipitch estimation (2022)","DOI":"10.1109\/ICASSP43922.2022.9746549"},{"key":"315_CR46","unstructured":"C.E. Cella, D. Ghisi, V. Lostanlen, F. L\u00e9vy, J. Fineberg, Y. Maresz, OrchideaSOL: a dataset of extended instrumental techniques for computer-aided orchestration (International Computer Music Conference, 2020)"},{"key":"315_CR47","unstructured":"V. Lostanlen, C.-E. Cella, Deep convolutional networks on the pitch spiral for musical instrument recognition (International Conference on Music Information Retrieval, 2017)"},{"key":"315_CR48","doi-asserted-by":"crossref","unstructured":"B.L. Sturm, An analysis of the GTZAN music genre dataset. In Proceedings of the second international ACM workshop on Music information retrieval with user-centered and multimodal strategies 7\u201312\u00a0(2012)","DOI":"10.1145\/2390848.2390851"},{"key":"315_CR49","unstructured":"R.M. Bittner, J. Wilkins, H. Yip, J.P. Bello, in International Conference on Music Information Retrieval, MedleyDB 2.0: New data and a system for sustainable data collection (2016)"},{"key":"315_CR50","doi-asserted-by":"crossref","unstructured":"M. Schoeffler, S. Bartoschek, F.-R. St\u00f6ter, M. Roess, S. Westphal, B. Edler, J. Herre, webMUSHRA - A Comprehensive Framework for Web-based Listening Tests. J. Open Res. Softw\u00a06, 1-8 (2018)","DOI":"10.5334\/jors.187"}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-023-00315-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-023-00315-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-023-00315-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T15:06:29Z","timestamp":1701788789000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-023-00315-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,5]]},"references-count":50,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["315"],"URL":"https:\/\/doi.org\/10.1186\/s13636-023-00315-5","relation":{},"ISSN":["1687-4722"],"issn-type":[{"type":"electronic","value":"1687-4722"}],"subject":[],"published":{"date-parts":[[2023,12,5]]},"assertion":[{"value":"16 June 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 November 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 December 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"51"}}