{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,27]],"date-time":"2025-12-27T18:35:21Z","timestamp":1766860521218,"version":"3.48.0"},"reference-count":72,"publisher":"Springer Science and Business Media LLC","issue":"42","license":[{"start":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T00:00:00Z","timestamp":1760313600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T00:00:00Z","timestamp":1760313600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100018967","name":"Universitat Pompeu Fabra","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100018967","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Music identification is crucial for distributing royalties in the music industry. This problem is solved using Audio fingerprinting (AFP) algorithms. However, these methods often struggle in real-world scenarios such as TV broadcasting, when music is in the background, masked by other sounds such as speech. While prior research has focused on improving AFP robustness to pitch and tempo variations, less attention has been given to enhancing robustness for background music identification. In this work, we assess whether source separation systems improve background music identification by recovering the music signal in these recordings. We present the first extensive study comprising 13 source separation algorithms and five AFP models. We evaluate them on a public dataset of TV recordings, assessing both music identification performance and computational cost. Our results show that source separation substantially improves peak-based AFP identifications, particularly when music is in the background. Additionally, this finding extends to foreground music, making the approach versatile for various music identification tasks, such as query-by-example. Deep learning-based model NeuralFP* (tailored for background music identification) shows no substantial benefit from adding a separation model as preprocessing. This reproducible study provides a comprehensive evaluation framework, offering valuable insights into using source separation methods to improve music identification in real-world contexts.<\/jats:p>","DOI":"10.1007\/s11042-025-21080-x","type":"journal-article","created":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T14:33:08Z","timestamp":1760365988000},"page":"50595-50628","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Enhanced television broadcast monitoring with source separation-assisted audio fingerprinting: A case study"],"prefix":"10.1007","volume":"84","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2827-8955","authenticated-orcid":false,"given":"Guillem","family":"Cort\u00e8s-Sebasti\u00e0","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2563-075X","authenticated-orcid":false,"given":"Marius","family":"Miron","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8251-9911","authenticated-orcid":false,"given":"Emilio","family":"Molina","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2878-4834","authenticated-orcid":false,"given":"Alex","family":"Ciurana","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1395-2345","authenticated-orcid":false,"given":"Xavier","family":"Serra","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,10,13]]},"reference":[{"key":"21080_CR1","unstructured":"Wang A (2003) An industrial strength audio search algorithm. In: ISMIR 2003, 4th International Conference on Music Information Retrieval, Baltimore, Maryland, USA, October 27-30, 2003, Proceedings, Baltimore, Maryland, USA, pp. 7\u201313. https:\/\/www.ee.columbia.edu\/~dpwe\/papers\/Wang03-shazam.pdf"},{"key":"21080_CR2","doi-asserted-by":"crossref","unstructured":"Gomez E, Cano P, Gomes, L, Batlle E, Bonnet, M (2002) Mixed watermarking-fingerprinting approach for integrity verification of audio recordings. In: Proceedings of the International Telecommunications Symposium, Natal, Brazil. https:\/\/mtg.upf.edu\/files\/publications\/its2002-egomez.pdf","DOI":"10.14209\/its.2002.303"},{"key":"21080_CR3","doi-asserted-by":"publisher","unstructured":"Ouali C, Dumouchel P, Gupta V (2014) A robust audio fingerprinting method for content-based copy detection. In: 12th International Workshop on Content-Based Multimedia Indexing, CBMI 2014. IEEE, Klagenfurt, Austria, pp. 1\u20136. https:\/\/doi.org\/10.1109\/CBMI.2014.6849814","DOI":"10.1109\/CBMI.2014.6849814"},{"key":"21080_CR4","unstructured":"Sonnleitner R, Arzt A, Widmer G (2016) Landmark-based audio fingerprinting for DJ mix monitoring. In: Proceedings of the 17th International Society for Music Information Retrieval Conference, ISMIR 2016, New York City, United States, August 7-11, 2016, New York City, New York, USA, pp. 185\u2013191. http:\/\/m.mr-pc.org\/ismir16\/website\/articles\/187_Paper.pdf"},{"key":"21080_CR5","doi-asserted-by":"publisher","unstructured":"Chang S, Lee D, Park J, Lim H, Lee K, Ko K, Han Y (2021) Neural audio fingerprint for high-specific audio retrieval based on contrastive learning. In: IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2021. IEEE, Toronto, Ontario, Canada, pp. 3025\u2013302https:\/\/doi.org\/10.1109\/ICASSP39728.2021.9414337","DOI":"10.1109\/ICASSP39728.2021.9414337"},{"key":"21080_CR6","doi-asserted-by":"publisher","unstructured":"Cort\u00e8s G, Ciurana A, Molina E, Miron M, Meyers O, Six J, Serra X (2022) Baf: An audio fingerprinting dataset for broadcast monitoring. In: Proceedings of the 23rd International Society for Music Information Retrieval Conference, ISMIR, Bengaluru, India, pp. 908\u20139https:\/\/doi.org\/10.5281\/zenodo.7316812","DOI":"10.5281\/zenodo.7316812"},{"key":"21080_CR7","unstructured":"Haitsma J, Kalker T (2002) A highly robust audio fingerprinting system. In: Proceedings of the 3rd International Conference on Music Information Retrieval (ISMIR 2002), Paris, France, pp. 107\u2013115. http:\/\/ismir2002.ismir.net\/proceedings\/02-FP04-2.pdf"},{"issue":"2008","key":"21080_CR8","doi-asserted-by":"publisher","first-page":"3467","DOI":"10.1016\/j.patcog.2008.05.006","volume":"41","author":"S Baluja","year":"2008","unstructured":"Baluja S, Covell M (2008) Waveprint: Efficient wavelet-based audio fingerprinting. Pattern Recogn. 41(2008):3467\u20133480. https:\/\/doi.org\/10.1016\/j.patcog.2008.05.006","journal-title":"Pattern Recogn."},{"key":"21080_CR9","doi-asserted-by":"publisher","unstructured":"Fenet S, Richard G, Grenier Y (2011) A scalable audio fingerprint method with robustness to pitch-shifting. In: Proceedings of the 12th International Society for Music Information Retrieval Conference, ISMIR 2011, Miami, Florida, USA, October 24\u201328, 2011, Miami, Florida, USA, pp. 121\u2013126https:\/\/doi.org\/10.5281\/zenodo.1417593","DOI":"10.5281\/zenodo.1417593"},{"issue":"2020","key":"21080_CR10","doi-asserted-by":"publisher","first-page":"172343","DOI":"10.1109\/ACCESS.2020.3024951","volume":"8","author":"H Son","year":"2020","unstructured":"Son H, Byun S, Lee S (2020) A robust audio fingerprinting using a new hashing method. IEEE Access 8(2020):172343\u201317235. https:\/\/doi.org\/10.1109\/ACCESS.2020.3024951","journal-title":"IEEE Access"},{"key":"21080_CR11","doi-asserted-by":"publisher","unstructured":"Zhang X, Zhu B, Li L, Li W, Li X, Wang W, Lu P, Zhang W (2015) Sift-based local spectrogram image descriptor: a novel feature for robust music identification. EURASIP Journal on Audio Speech Music Processing (2015):https:\/\/doi.org\/10.1186\/s13636-015-0050-0","DOI":"10.1186\/s13636-015-0050-0"},{"key":"21080_CR12","doi-asserted-by":"publisher","unstructured":"Six J, Leman M (2014) Panako: A scalable acoustic fingerprinting system handling time-scale and pitch modification. In: 15th International Society for Music Information Retrieval Conference (ISMIR-2014), Taipei, Taiwan, pp. 259\u201326https:\/\/doi.org\/10.5281\/zenodo.1416190","DOI":"10.5281\/zenodo.1416190"},{"issue":"2014","key":"21080_CR13","doi-asserted-by":"publisher","first-page":"308","DOI":"10.1016\/j.sigpro.2013.11.023","volume":"98","author":"M Malekesmaeili","year":"2014","unstructured":"Malekesmaeili M, Ward RK (2014) A local fingerprinting approach for audio copy detection. Signal Process. 98(2014):308\u2013321. https:\/\/doi.org\/10.1016\/j.sigpro.2013.11.023","journal-title":"Signal Process."},{"key":"21080_CR14","unstructured":"Sonnleitner R, Widmer G (2014) Quad-based audio fingerprinting robust to time and frequency scaling. In: Proceedings of the 17th International Conference on Digital Audio Effects, DAFx-14, Erlangen, Germany, pp. 173\u2013180. http:\/\/www.dafx14.fau.de\/papers\/dafx14_reinhard_sonnleitner_quad_based_audio_fingerpr.pdf"},{"issue":"2016","key":"21080_CR15","doi-asserted-by":"publisher","first-page":"409","DOI":"10.1109\/TASLP.2015.2509248","volume":"24","author":"R Sonnleitner","year":"2016","unstructured":"Sonnleitner R, Widmer G (2016) Robust quad-based audio fingerprinting. IEEE ACM Transactions on Audio Speech Language Processing 24(2016):409\u201342. https:\/\/doi.org\/10.1109\/TASLP.2015.2509248","journal-title":"IEEE ACM Transactions on Audio Speech Language Processing"},{"key":"21080_CR16","doi-asserted-by":"publisher","unstructured":"Dupraz E, Richard G (2010) Robust frequency-based audio fingerprinting. In: Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2010. IEEE, Dallas, Texas, USA, pp. 281\u201328https:\/\/doi.org\/10.1109\/ICASSP.2010.5495944","DOI":"10.1109\/ICASSP.2010.5495944"},{"key":"21080_CR17","unstructured":"MIREX (2018) Audio fingerprinting results, 14th music information retrieval evaluation exchange (mirex 2018). https:\/\/www.music-ir.org\/mirex\/wiki\/2018:Audio_Fingerprinting_Results. [Accessed October 2024]"},{"key":"21080_CR18","unstructured":"MIREX (2019) Audio fingerprinting results, 15th music information retrieval evaluation exchange (mirex 2019). https:\/\/www.music-ir.org\/mirex\/wiki\/2019:Audio_Fingerprinting. [Accessed October 2024]"},{"key":"21080_CR19","unstructured":"MIREX (2020) Audio fingerprinting results, 16th music information retrieval evaluation exchange (mirex 2020). https:\/\/www.music-ir.org\/mirex\/wiki\/2020:Audio_Fingerprinting_Results. [Accessed October 2024]"},{"key":"21080_CR20","doi-asserted-by":"publisher","unstructured":"Mel\u00e9ndez-Catal\u00e1n B, Molina E, G\u00f3mez E (2019) Open broadcast media audio from TV: A dataset of TV broadcast audio with relative music loudness annotations. Transactions of the International Society for Music Information Retrieval 2(2019):43\u2013https:\/\/doi.org\/10.5334\/tismir.29","DOI":"10.5334\/tismir.29"},{"key":"21080_CR21","unstructured":"IFPI, Global music report 2023, https:\/\/ifpi-website-cms.s3.eu-west-2.amazonaws.com\/GMR_2023_State_of_the_Industry_ee2ea600e2.pdf. [Accessed October 2024]"},{"key":"21080_CR22","doi-asserted-by":"publisher","unstructured":"Han W, Zhou S, Li C, Liu Y, Liu Z (2015) Blind source separation for a robust audio recognition scheme in multiple sound-sources environment. In: Proceedings of the 2015 International Conference on Mechatronics, Electronic, Industrial and Control Engineering, Atlantis Press, Shenyang, China, pp. 1564\u20131568https:\/\/doi.org\/10.2991\/meic-15.2015.358","DOI":"10.2991\/meic-15.2015.358"},{"key":"21080_CR23","unstructured":"Jansson A, Humphrey EJ, Montecchio N, Bittner RM, Kumar A, Weyde T (2017) Singing voice separation with deep u-net convolutional networks. In: Proceedings of the 18th International Society for Music Information Retrieval Conference, ISMIR 2017, Suzhou, China, 2017, pp. 745\u2013751. https:\/\/ejhumphrey.com\/assets\/pdf\/jansson2017singing.pdf"},{"key":"21080_CR24","unstructured":"Stoller D, Ewert S, Dixon S (2018) Wave-u-net: A multi-scale neural network for end-to-end audio source separation, in: Proceedings of the 19th International Society for Music Information Retrieval Conference, ISMIR 2018, Paris, France, 2018, pp. 334\u2013340. http:\/\/ismir2018.ircam.fr\/doc\/pdfs\/205_Paper.pdf"},{"key":"21080_CR25","unstructured":"Yu CY, Chuek KW (2021) Danna-sep: Unite to separate them all, The ISMIR 2021 Workshop on Music Source Separation"},{"key":"21080_CR26","unstructured":"ITU-R BS.1534, Method for the subjective assessment of intermediate quality level of audio systems, International Telecommunication Union Radiocommunication Assembly (2014)"},{"key":"21080_CR27","doi-asserted-by":"publisher","unstructured":"Neuschmied H, Mayer H, Batlle E (2001) Content-based identification of audio titles on the internet, in: Proceedings First International Conference on WEB Delivering of Music. WEDELMUSIC 2001, pp. 96\u2013100. https:\/\/doi.org\/10.1109\/WDM.2001.990163","DOI":"10.1109\/WDM.2001.990163"},{"key":"21080_CR28","unstructured":"Cano P, Batlle E, Mayer H, Neuschmied H (2002) Robust sound modeling for song detection in broadcast audio. In: Proceedings of the 112th AES Convention, Munich, Germany, pp. 1\u20137. https:\/\/mtg.upf.edu\/files\/publications\/aes2002-pcano.pdf"},{"key":"21080_CR29","unstructured":"Allamanche E (2001) Audioid: Towards content-based identification of audio material. In: 100th AES Convention, May, 2001"},{"key":"21080_CR30","unstructured":"Kastner T, Allamanche E, Herre J, Hellmuth O, Cremer M, Grossmann H (2002) Mpeg-7 scalable robust audio fingerprinting. In: Audio Engineering Society Convention 112, Audio Engineering Society.https:\/\/www.iis.fraunhofer.de\/content\/dam\/iis\/de\/doc\/ame\/conference\/AES-112-Convention_MPEG-7_ScalableRobustAudioFingerprinting_AES5511.pdf"},{"key":"21080_CR31","unstructured":"Bardeli R (2004) Robust identification of time-scaled audio, in: Audio Engineering Society Conference: 25th International Conference: Metadata for Audio, Audio Engineering Society. https:\/\/www.researchgate.net\/publication\/200038389_Robust_Identification_of_Time-Scaled_Audio"},{"key":"21080_CR32","unstructured":"Fenet S, Moussallam M, Grenier Y, Richard G, Daudet L (2012) A framework for fingerprint-based detection of repeating objects in multimedia streams. In: 2012 Proceedings of the 20th European Signal Processing Conference (EUSIPCO), pp. 1464\u20131468. https:\/\/www.eurasip.org\/Proceedings\/Eusipco\/Eusipco2012\/Conference\/papers\/1569582247.pdf"},{"key":"21080_CR33","doi-asserted-by":"publisher","unstructured":"Ramona M, Peeters G (2011) Audio identification based on spectral modeling of bark-bands energy and synchronization through onset detection, in: Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2011, May 22-27, 2011, Prague Congress Center, Prague, Czech Republic, IEEE, pp. 477\u2013480. https:\/\/doi.org\/10.1109\/ICASSP.2011.5946444","DOI":"10.1109\/ICASSP.2011.5946444"},{"key":"21080_CR34","doi-asserted-by":"publisher","unstructured":"Ramona M, Peeters G (2013) Audioprint: An efficient audio fingerprint system based on a novel cost-less synchronization scheme, in: IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2013, IEEE, Vancouver, BC, Canada, pp. 818\u2013822. https:\/\/doi.org\/10.1109\/ICASSP.2013.6637762","DOI":"10.1109\/ICASSP.2013.6637762"},{"key":"21080_CR35","unstructured":"Six J (2020) Olaf: Overly lightweight acoustic fingerprinting, in: Demo \/ late-breaking abstracts of 21st International Society for Music Information Retrieval Conference, ISMIR 2020, Montr\u00e9al, Canada. https:\/\/archives.ismir.net\/ismir2020\/latebreaking\/000001.pdf"},{"key":"21080_CR36","doi-asserted-by":"publisher","unstructured":"Six J (20223) Olaf: a lightweight, portable audio search system, Journal of Open Source Software 8(2023):5459. https:\/\/doi.org\/10.21105\/joss.05459.https:\/\/doi.org\/10.21105\/joss.05459","DOI":"10.21105\/joss.05459"},{"key":"21080_CR37","doi-asserted-by":"publisher","first-page":"135","DOI":"10.1016\/j.patrec.2022.06.009","volume":"160","author":"S Serrano","year":"2022","unstructured":"Serrano S, Sahbudin MAB, Chaouch C, Scarpa M (2022) A new fingerprint definition for effective song recognition. Pattern Recogn. Lett. 160:135\u2013141. https:\/\/doi.org\/10.1016\/j.patrec.2022.06.009","journal-title":"Pattern Recogn. Lett."},{"key":"21080_CR38","doi-asserted-by":"publisher","first-page":"31591","DOI":"10.1007\/s11042-023-14787-2","volume":"82","author":"S Serrano","year":"2023","unstructured":"Serrano S, Scarpa M (2023) Accuracy comparisons of fingerprint based song recognition approaches using very high granularity. Multimedia Tools and Applications 82:31591\u201331606. https:\/\/doi.org\/10.1007\/s11042-023-14787-2","journal-title":"Multimedia Tools and Applications"},{"key":"21080_CR39","doi-asserted-by":"publisher","unstructured":"Sahbudin MAB, Chaouch C, Scarpa M, Serrano S (2019) IoT based song recognition for fm radio station broadcasting. In: 2019 7th International Conference on Information and Communication Technology (ICoICT). IEEE, pp. 1\u20136https:\/\/doi.org\/10.1109\/ICoICT.2019.8835190","DOI":"10.1109\/ICoICT.2019.8835190"},{"key":"21080_CR40","unstructured":"Agarwaal A, Kanaujia P, Roy SS, Ghose S (2023) Robust and lightweight audio fingerprint for automatic content recognition. arXiv:2305.09559"},{"key":"21080_CR41","doi-asserted-by":"publisher","unstructured":"D\u00e9fossez A (2021) Hybrid spectrogram and waveform source separation, in: Proceedings of the ISMIR 2021 Workshop on Music Source Separation.https:\/\/doi.org\/10.48550\/arXiv.2111.03600","DOI":"10.48550\/arXiv.2111.03600"},{"key":"21080_CR42","doi-asserted-by":"publisher","unstructured":"Hennequin R, Khlif A, Voituret F, Moussallam M (2020) Spleeter: a fast and efficient music source separation tool with pre-trained models. Journal of Open Source Software 5(2020):2154.https:\/\/doi.org\/10.21105\/joss.02154, Deezer Research","DOI":"10.21105\/joss.02154"},{"key":"21080_CR43","doi-asserted-by":"publisher","unstructured":"Kong Q, Chen K, Liu H, Du X, Berg-Kirkpatrick T, Dubnov S, Plumbley MD (2023) Universal source separation with weakly labelled data. https:\/\/doi.org\/10.48550\/arXiv.2305.07447","DOI":"10.48550\/arXiv.2305.07447"},{"key":"21080_CR44","doi-asserted-by":"publisher","unstructured":"Petermann D, Wichern G, Wang ZQ, Le Roux J (2022) The cocktail fork problem: Three-stem audio separation for real-world soundtracks. In: 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, 2022, pp. 526\u2013530. https:\/\/doi.org\/10.1109\/ICASSP43922.2022.9746005","DOI":"10.1109\/ICASSP43922.2022.9746005"},{"key":"21080_CR45","doi-asserted-by":"publisher","unstructured":"Sawata R, Uhlich S, Takahashi S, Mitsufuji Y (2021) All for one and one for all: Improving music separation by bridging networks. In: 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, Toronto, Ontario, Canada, pp. 51\u201355.https:\/\/doi.org\/10.1109\/ICASSP39728.2021.9414044","DOI":"10.1109\/ICASSP39728.2021.9414044"},{"key":"21080_CR46","doi-asserted-by":"publisher","unstructured":"Plaja-Roglans G, Miron M, Serra X (2022) A diffusion-inspired training strategy for singing voice extraction in the waveform domain, in: Proceedings of the 23rd International Society for Music Information Retrieval Conference, ISMIR 2022, Bengaluru, India, December 4-8, 2022, pp. 685\u2013693. https:\/\/doi.org\/10.5281\/zenodo.7316754","DOI":"10.5281\/zenodo.7316754"},{"key":"21080_CR47","doi-asserted-by":"crossref","unstructured":"Defossez A, Synnaeve G, Adi Y (2020) Real time speech enhancement in the waveform domain, in: Proceedings of the 21st Annual Conference of the International Speech Communication Association (INTERSPEECH 2020), Shanghai, China, p. 3291\u20133295. arXiv:2006.12847","DOI":"10.21437\/Interspeech.2020-2409"},{"key":"21080_CR48","doi-asserted-by":"publisher","unstructured":"Plaja-Roglans G, Miron M, Shankar A, Serra X (2023) Carnatic singing voice separation using cold diffusion on training data with bleeding. In: Proceedings of the 24th International Society for Music Information Retrieval Conference, ISMIR 2023, Milan, Italy, November 5-9, 2023, pp. 553\u2013560. https:\/\/doi.org\/10.5281\/ZENODO.10265347","DOI":"10.5281\/ZENODO.10265347"},{"key":"21080_CR49","doi-asserted-by":"publisher","first-page":"21155","DOI":"10.1007\/s11042-022-11994-1","volume":"81","author":"M Mirbeygi","year":"2022","unstructured":"Mirbeygi M, Mahabadi A, Ranjbar A (2022) Speech and music separation approaches-a survey. Multimedia Tools and Applications 81:21155\u201321197. https:\/\/doi.org\/10.1007\/s11042-022-11994-1","journal-title":"Multimedia Tools and Applications"},{"key":"21080_CR50","doi-asserted-by":"publisher","first-page":"958","DOI":"10.1109\/TNN.2007.915115","volume":"19","author":"K-K Shyu","year":"2008","unstructured":"Shyu K-K, Lee M-H, Wu Y-T, Lee P-L (2008) Implementation of pipelined fastica on fpga for real-time blind source separation. IEEE Trans. Neural Networks 19:958\u2013970. https:\/\/doi.org\/10.1109\/TNN.2007.915115","journal-title":"IEEE Trans. Neural Networks"},{"key":"21080_CR51","doi-asserted-by":"publisher","DOI":"10.4218\/etrij.2023-0249","author":"H Kim","year":"2024","unstructured":"Kim H, Kim J, Park J, Kim S, Park C, Yoo W (2024) Background music monitoring framework and dataset for tv broadcast audio. ETRI J. https:\/\/doi.org\/10.4218\/etrij.2023-0249","journal-title":"ETRI J."},{"key":"21080_CR52","unstructured":"Tensorflow, Sound classification with yamnet, https:\/\/www.tensorflow.org\/hub\/tutorials\/yamnet, 2022. [Accessed October 2024]"},{"key":"21080_CR53","doi-asserted-by":"publisher","unstructured":"Rouard S, Massa F, D\u00e9fossez A (2023) Hybrid transformers for music source separation. In: 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, pp. 1\u20135. https:\/\/doi.org\/10.1109\/ICASSP49357.2023.10096956","DOI":"10.1109\/ICASSP49357.2023.10096956"},{"key":"21080_CR54","doi-asserted-by":"publisher","unstructured":"St\u00f6ter FR, Uhlich S, Liutkus A, Mitsufuji Y (2019) Open-unmix - a reference implementation for music source separation. Journal of Open Source Software. https:\/\/doi.org\/10.21105\/joss.01667","DOI":"10.21105\/joss.01667"},{"key":"21080_CR55","doi-asserted-by":"publisher","unstructured":"Jansson A, Humphrey EJ, Montecchio N, Bittner RM, Kumar A, Weyde T (2017) Singing voice separation with deep u-net convolutional networks, in: Proceedings of the 18th International Society for Music Information Retrieval Conference, ISMIR 2017, ISMIR, Suzhou, China, pp. 745\u2013751. https:\/\/doi.org\/10.5281\/zenodo.1414934","DOI":"10.5281\/zenodo.1414934"},{"key":"21080_CR56","doi-asserted-by":"publisher","unstructured":"Pr\u00e9tet L, Hennequin R, Royo-Letelier J, Vaglio A (2019) Singing voice separation: A study on training data. In: 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, United Kingdom, pp. 506\u2013510.https:\/\/doi.org\/10.1109\/ICASSP.2019.8683555","DOI":"10.1109\/ICASSP.2019.8683555"},{"key":"21080_CR57","doi-asserted-by":"crossref","unstructured":"Zhang L, Li C, Deng F, Wang X (2021) Multi-task audio source separation, 2021 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) 671\u2013678. arXiv:2107.06467","DOI":"10.1109\/ASRU51503.2021.9687922"},{"key":"21080_CR58","doi-asserted-by":"publisher","unstructured":"Schmidt N, Pons J, Miron M (2022) Podcastmix: A dataset for separating music and speech in podcasts. In: Interspeech 2022, 23rd Annual Conference of the International Speech Communication Association, Incheon, Korea, 18-22 September 2022, ISCA, pp. 231\u2013235. https:\/\/doi.org\/10.21437\/Interspeech.2022-41","DOI":"10.21437\/Interspeech.2022-41"},{"key":"21080_CR59","doi-asserted-by":"publisher","unstructured":"Panayotov V, Chen G, Povey D, Khudanpur S (2015) Librispeech: An ASR corpus based on public domain audio books. In: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2015, IEEE, South Brisbane, Queensland, Australia, pp. 5206\u20135210. https:\/\/doi.org\/10.1109\/ICASSP.2015.7178964","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"21080_CR60","unstructured":"Bertin-Mahieux T, Ellis DP, Whitman B, Lamere P (2011) The million song dataset. In: Proceedings of the 12th International Conference on Music Information Retrieval (ISMIR 2011), Miami, Florida, USA. https:\/\/ismir2011.ismir.net\/papers\/OS6-1.pdf"},{"key":"21080_CR61","doi-asserted-by":"crossref","unstructured":"Fonseca E, Favory X, Pons J, Font F, Serra X (2022) FSD50K: an open dataset of human-labeled sound events, IEEE\/ACM Transactions on Audio, Speech, and Language Processing 30 (2022) 829\u2013852. https:\/\/repositori.upf.edu\/bitstream\/handle\/10230\/56072\/Font_tra_fsd5.pdf","DOI":"10.1109\/TASLP.2021.3133208"},{"key":"21080_CR62","doi-asserted-by":"publisher","unstructured":"Gemmeke JF, Ellis DPW, Freedman D, Jansen A, Lawrence W, Moore RC, Plakal M, Ritter M (2017) Audio set: An ontology and human-labeled dataset for audio events. In: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2017, IEEE, New Orleans, Louisiana, USA, pp. 776\u2013780. https:\/\/doi.org\/10.1109\/ICASSP.2017.7952261","DOI":"10.1109\/ICASSP.2017.7952261"},{"key":"21080_CR63","unstructured":"Takahashi N, Mitsufuji Y (2020) D3net: Densely connected multidilated densenet for music source separation. arXiv:2010.01733"},{"key":"21080_CR64","unstructured":"Parmar N, Vaswani A, Uszkoreit J, Kaiser L, Shazeer N, Ku A, Tran D (2018) Image transformer, in: Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm\u00e4ssan, volume\u00a080 of Proceedings of Machine Learning Research, PMLR, Stockholm, Sweden, pp. 4052\u20134061. http:\/\/proceedings.mlr.press\/v80\/parmar18a.html"},{"key":"21080_CR65","doi-asserted-by":"publisher","unstructured":"Mitsufuji Y, Fabbro G, Uhlich S, St\u00f6ter FR, D\u00e9fossez A, Kim M, Choi W, Yu CY, Cheuk KW (2021) Music demixing challenge 2021. Frontiers in Signal Processing 1. https:\/\/doi.org\/10.3389\/frsip.2021.808395","DOI":"10.3389\/frsip.2021.808395"},{"key":"21080_CR66","unstructured":"Ellis D (2014) The 2014 labrosa audio fingerprint system, in: Proceedings of the 15th International Society for Music Information Retrieval Conference, ISMIR 2014, Taipei, Taiwan. https:\/\/www.music-ir.org\/mirex\/abstracts\/2014\/DP1.pdf"},{"key":"21080_CR67","unstructured":"Six J (2021) Panako 2.0 - updates for an acoustic fingerprinting system, in: Demo \/ late-breaking abstracts of 22st International Society for Music Information Retrieval Conference, ISMIR 2021, Online.https:\/\/biblio.ugent.be\/publication\/8726851\/file\/8726856.pdf"},{"key":"21080_CR68","doi-asserted-by":"publisher","first-page":"95","DOI":"10.1016\/S1018-3639(18)30850-X","volume":"19","author":"AI Al-Shoshan","year":"2006","unstructured":"Al-Shoshan AI (2006) Speech and music classification and separation: A review, Journal of King Saud University -. Eng. Sci. 19:95\u2013132. https:\/\/doi.org\/10.1016\/S1018-3639(18)30850-X","journal-title":"Eng. Sci."},{"key":"21080_CR69","doi-asserted-by":"crossref","unstructured":"Serrano S, Scarpa M, Serghini O et\u00a0al (2024) Vggish for music\/speech classification in radio broadcasting, Proceedings European Council For Modelling and Simulation 38 (2024) 550\u2013557. https:\/\/hdl.handle.net\/11570\/3300940","DOI":"10.7148\/2024-0550"},{"key":"21080_CR70","doi-asserted-by":"publisher","unstructured":"Acosta-Ceja JA, Coto-Jim\u00e9nez M, S\u00e1nchez-Guti\u00e9rrez ME, Sagaceta-Mej\u00eda AR, Fres\u00e1n-Figueroa J (2024) Feature engineering for music\/speech detection in costa rica radio broadcast, in: Pattern Recognition - 16th Mexican Conference, MCPR 2024, Xalapa, Mexico, June 19-22, 2024, Proceedings, volume 14755 of Lecture Notes in Computer Science, Springer, pp. 84\u201395. https:\/\/doi.org\/10.1007\/978-3-031-62836-8_9","DOI":"10.1007\/978-3-031-62836-8_9"},{"key":"21080_CR71","unstructured":"Mel\u00e9ndez-Catal\u00e1n B (2020) Relative music loudness estimation using temporal convolutional networks and a cnn feature extraction front-end, in: Proceedings of the 23rd International Conference on Digital Audio Effects (DAFx-20), volume\u00a05, Vienna, Austria, 2020, pp. 273\u2013280. https:\/\/www.dafx.de\/paper-archive\/2020\/proceedings\/papers\/DAFx2020_paper_13.pdf"},{"key":"21080_CR72","doi-asserted-by":"crossref","unstructured":"Schoeffler M, St\u00f6ter FR, Edler B, Herre J (2015) Towards the next generation of web-based experiments: A case study assessing basic audio quality following the itu-r recommendation bs. 1534 (mushra), in: 1st Web Audio Conference, pp. 1\u20136. https:\/\/wac.ircam.fr\/pdf\/wac15_submission_8.pdf","DOI":"10.5334\/jors.187"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-025-21080-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-025-21080-x","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-025-21080-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,27]],"date-time":"2025-12-27T18:34:55Z","timestamp":1766860495000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-025-21080-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,13]]},"references-count":72,"journal-issue":{"issue":"42","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["21080"],"URL":"https:\/\/doi.org\/10.1007\/s11042-025-21080-x","relation":{},"ISSN":["1573-7721"],"issn-type":[{"type":"electronic","value":"1573-7721"}],"subject":[],"published":{"date-parts":[[2025,10,13]]},"assertion":[{"value":"17 May 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 July 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 July 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 October 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no competing interests to declare that are relevant to the content of this article.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}},{"value":"This work does not require ethics approval.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical standard"}},{"value":"This work does not require consent to participate, because it does not involve human subjects.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}}]}}