{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T11:21:19Z","timestamp":1782386479838,"version":"3.54.5"},"reference-count":101,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2022,10,9]],"date-time":"2022-10-09T00:00:00Z","timestamp":1665273600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,10,9]],"date-time":"2022-10-09T00:00:00Z","timestamp":1665273600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100000266","name":"Engineering and Physical Sciences Research Council","doi-asserted-by":"publisher","award":["EP\/T019751\/1"],"award-info":[{"award-number":["EP\/T019751\/1"]}],"id":[{"id":"10.13039\/501100000266","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100010897","name":"Newton Fund","doi-asserted-by":"publisher","award":["623805725"],"award-info":[{"award-number":["623805725"]}],"id":[{"id":"10.13039\/100010897","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002860","name":"China Sponsorship Council","doi-asserted-by":"publisher","award":["202006470010"],"award-info":[{"award-number":["202006470010"]}],"id":[{"id":"10.13039\/501100002860","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Automated audio captioning is a cross-modal translation task that aims to generate natural language descriptions for given audio clips. This task has received increasing attention with the release of freely available datasets in recent years. The problem has been addressed predominantly with deep learning techniques. Numerous approaches have been proposed, such as investigating different neural network architectures, exploiting auxiliary information such as keywords or sentence information to guide caption generation, and employing different training strategies, which have greatly facilitated the development of this field. In this paper, we present a comprehensive review of the published contributions in automated audio captioning, from a variety of existing approaches to evaluation metrics and datasets. We also discuss open challenges and envisage possible future research directions.<\/jats:p>","DOI":"10.1186\/s13636-022-00259-2","type":"journal-article","created":{"date-parts":[[2022,10,9]],"date-time":"2022-10-09T22:02:30Z","timestamp":1665352950000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":44,"title":["Automated audio captioning: an overview of recent progress and new challenges"],"prefix":"10.1186","volume":"2022","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6079-5130","authenticated-orcid":false,"given":"Xinhao","family":"Mei","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xubo","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mark D.","family":"Plumbley","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenwu","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,10,9]]},"reference":[{"key":"259_CR1","doi-asserted-by":"crossref","unstructured":"J.F. Gemmeke, D.P.W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R.C. Moore, M. Plakal, M. Ritter, in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Audio Set: an ontology and human-labeled dataset for audio events (New Orleans, 2017)","DOI":"10.1109\/ICASSP.2017.7952261"},{"issue":"6","key":"259_CR2","doi-asserted-by":"publisher","first-page":"1230","DOI":"10.1109\/TASLP.2017.2690563","volume":"25","author":"Y Xu","year":"2017","unstructured":"Y. Xu, Q. Huang, W. Wang, P. Foster, S. Sigtia, P.J.B. Jackson, M.D. Plumbley, Unsupervised feature learning based on deep models for environmental audio tagging. IEEE\/ACM Trans. Audio Speech Lang. Process. 25(6), 1230 (2017). https:\/\/doi.org\/10.1109\/TASLP.2017.2690563","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"11","key":"259_CR3","doi-asserted-by":"publisher","first-page":"1791","DOI":"10.1109\/TASLP.2019.2930913","volume":"27","author":"Q Kong","year":"2019","unstructured":"Q. Kong, C. Yu, Y. Xu, T. Iqbal, W. Wang, M.D. Plumbley, Weakly labelled AudioSet tagging with attention neural networks. IEEE\/ACM Trans. Audio Speech Lang. Process. 27(11), 1791 (2019). https:\/\/doi.org\/10.1109\/TASLP.2019.2930913","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"259_CR4","doi-asserted-by":"publisher","first-page":"1560","DOI":"10.1109\/LSP.2020.3019702","volume":"27","author":"H Wang","year":"2020","unstructured":"H. Wang, Y. Zou, D. Chong, W. Wang, Modeling label dependencies for audio tagging with graph convolutional network. IEEE Signal Process. Lett. 27, 1560 (2020). https:\/\/doi.org\/10.1109\/LSP.2020.3019702","journal-title":"IEEE Signal Process. Lett."},{"key":"259_CR5","doi-asserted-by":"crossref","unstructured":"Y. Xu, Q. Kong, Q. Huang, W. Wang, M.D. Plumbley, in Proc. IEEE International Joint Conference on Neural Networks (IJCNN). Convolutional gated recurrent neural network incorporating spatial features for audio tagging (2017)","DOI":"10.1109\/IJCNN.2017.7966291"},{"issue":"4","key":"259_CR6","doi-asserted-by":"publisher","first-page":"777","DOI":"10.1109\/TASLP.2019.2895254","volume":"27","author":"Q Kong","year":"2019","unstructured":"Q. Kong, Y. Xu, I. Sobieraj, W. Wang, M.D. Plumbley, Sound event detection and time-frequency segmentation from weakly labelled data. IEEE\/ACM Trans. Audio Speech Lang. Process. 27(4), 777 (2019). https:\/\/doi.org\/10.1109\/TASLP.2019.2895254","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"259_CR7","doi-asserted-by":"publisher","first-page":"2450","DOI":"10.1109\/TASLP.2020.3014737","volume":"28","author":"Q Kong","year":"2020","unstructured":"Q. Kong, Y. Xu, W. Wang, M.D. Plumbley, Sound event detection of weakly labelled data with CNN-transformer and automatic threshold optimization. IEEE\/ACM Trans. Audio Speech Lang. Process. 28, 2450 (2020). https:\/\/doi.org\/10.1109\/TASLP.2020.3014737","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"5","key":"259_CR8","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1109\/MSP.2021.3090678","volume":"38","author":"A Mesaros","year":"2021","unstructured":"A. Mesaros, T. Heittola, T. Virtanen, M.D. Plumbley, Sound event detection: a tutorial. IEEE Signal Process. Mag. 38(5), 67 (2021)","journal-title":"IEEE Signal Process. Mag."},{"issue":"3","key":"259_CR9","doi-asserted-by":"publisher","first-page":"16","DOI":"10.1109\/MSP.2014.2326181","volume":"32","author":"D Barchiesi","year":"2015","unstructured":"D. Barchiesi, D. Giannoulis, D. Stowell, M.D. Plumbley, Acoustic scene classification: classifying environments from the sounds they produce. IEEE Signal Process. Mag. 32(3), 16 (2015)","journal-title":"IEEE Signal Process. Mag."},{"key":"259_CR10","doi-asserted-by":"crossref","unstructured":"H. Wang, Y. Zou, W. Wang, in Interspeech. SpecAugment++: a hidden space data augmentation method for acoustic scene classification (ISCA, 2021), pp. 551-555","DOI":"10.31219\/osf.io\/3mwa7"},{"key":"259_CR11","doi-asserted-by":"crossref","unstructured":"K. Drossos, S. Adavanne, T. Virtanen, in 2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). Automated audio captioning with recurrent neural networks (IEEE, 2017), pp. 374-378","DOI":"10.1109\/WASPAA.2017.8170058"},{"key":"259_CR12","volume-title":"Audio Retrieval with Natural Language Queries: A Benchmark Study","author":"AS Koepke","year":"2022","unstructured":"A.S. Koepke, A.M. Oncescu, J. Henriques, Z. Akata, S. Albanie, Audio retrieval with natural language queries: a benchmark study (IEEE Trans, Multimed, 2022)"},{"key":"259_CR13","doi-asserted-by":"crossref","unstructured":"X. Mei, X. Liu, J. Sun, M.D. Plumbley, W. Wang, On metric learning for audio-text cross-modal retrieval. arXiv preprint arXiv:2203.15537 (2022)","DOI":"10.21437\/Interspeech.2022-11115"},{"key":"259_CR14","unstructured":"I. Sutskever, O. Vinyals, Q.V. Le, in Proceedings of the 27th International Conference on Neural Information Processing Systems, NIPS\u201914. Sequence to sequence learning with neural networks (MIT Press, Cambridge, 2014), p. 3104-3112"},{"issue":"6088","key":"259_CR15","doi-asserted-by":"publisher","first-page":"533","DOI":"10.1038\/323533a0","volume":"323","author":"DE Rumelhart","year":"1986","unstructured":"D.E. Rumelhart, G.E. Hinton, R.J. Williams, Learning representations by back-propagating errors. Nature 323(6088), 533 (1986)","journal-title":"Nature"},{"issue":"7553","key":"259_CR16","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Y LeCun","year":"2015","unstructured":"Y. LeCun, Y. Bengio, G. Hinton, Deep learning. Nature 521(7553), 436 (2015)","journal-title":"Nature"},{"key":"259_CR17","unstructured":"A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, \u0141. Kaiser, I. Polosukhin, in Advances in Neural Information Processing Systems. Attention is all you need (2017), pp. 5998-6008"},{"key":"259_CR18","doi-asserted-by":"crossref","unstructured":"Y. Koizumi, R. Masumura, K. Nishida, M. Yasuda, S. Saito, in INTERSPEECH. A transformer-based audio captioning model with keyword estimation (ISCA, 2020), pp. 1977-1981","DOI":"10.21437\/Interspeech.2020-2087"},{"key":"259_CR19","doi-asserted-by":"crossref","unstructured":"X. Xu, H. Dinkel, M. Wu, K. Yu, in 2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP). Audio caption in a car setting with a sentence-level loss (IEEE, 2021), pp. 1-5","DOI":"10.1109\/ISCSLP49672.2021.9362117"},{"key":"259_CR20","unstructured":"C.D. Kim, B. Kim, H. Lee, G. Kim, in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. AudioCaps: generating captions for audios in the wild (2019), pp. 119-32"},{"key":"259_CR21","unstructured":"X. Xu, Z. Xie, M. Wu, K. Yu, The SJTU system for DCASE2021 Challenge Task 6: audio captioning based on encoder pre-training and reinforcement learning. Tech. rep., DCASE2021 Challenge (2021)"},{"key":"259_CR22","unstructured":"J. Berg, K. Drossos, in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021). Continual learning for automated audio captioning using the learning without forgetting approach (Barcelona, 2021), pp. 140-144"},{"key":"259_CR23","unstructured":"X. Liu, Q. Huang, X. Mei, T. Ko, H. Tang, M.D. Plumbley, W. Wang, in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021). CL4AC: a contrastive loss for audio captioning (Barcelona, 2021), pp. 196-200"},{"key":"259_CR24","doi-asserted-by":"crossref","unstructured":"A. Graves, Sequence transduction with recurrent neural networks. arXiv preprint arXiv:1211.3711 (2012)","DOI":"10.1007\/978-3-642-24797-2"},{"key":"259_CR25","doi-asserted-by":"crossref","unstructured":"T. Virtanen, M.D. Plumbley, D. Ellis, Computational analysis of sound scenes and events (Springer, 2018)","DOI":"10.1007\/978-3-319-63450-0"},{"issue":"1","key":"259_CR26","doi-asserted-by":"publisher","first-page":"51","DOI":"10.1109\/PROC.1978.10837","volume":"66","author":"F Harris","year":"1978","unstructured":"F. Harris, On the use of windows for harmonic analysis with the discrete Fourier transform. Proc. IEEE. 66(1), 51 (1978). https:\/\/doi.org\/10.1109\/PROC.1978.10837","journal-title":"Proc. IEEE."},{"key":"259_CR27","doi-asserted-by":"publisher","first-page":"2880","DOI":"10.1109\/TASLP.2020.3030497","volume":"28","author":"Q Kong","year":"2020","unstructured":"Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, M.D. Plumbley, PANNs: large-scale pretrained audio neural networks for audio pattern recognition. IEEE\/ACM Trans. Audio Speech Lang. Process. 28, 2880 (2020)","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"259_CR28","doi-asserted-by":"crossref","unstructured":"S. Hershey, S. Chaudhuri, D.P. Ellis, J.F. Gemmeke, A. Jansen, R.C. Moore, M. Plakal, D. Platt, R.A. Saurous, B. Seybold, et\u00a0al., in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). CNN architectures for large-scale audio classification (IEEE, 2017), pp. 131-135","DOI":"10.1109\/ICASSP.2017.7952132"},{"key":"259_CR29","doi-asserted-by":"crossref","unstructured":"M. Wu, H. Dinkel, K. Yu, in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Audio caption: listen and tell (IEEE, 2019), pp. 830-834","DOI":"10.1109\/ICASSP.2019.8682377"},{"key":"259_CR30","doi-asserted-by":"crossref","unstructured":"S. Ikawa, K. Kashino, in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2019 Workshop (DCASE2019). Neural audio captioning based on conditional sequence-to-sequence model (New York University, New York, 2019), pp. 99-103","DOI":"10.33682\/7bay-bj41"},{"key":"259_CR31","unstructured":"J. Chung, C. Gulcehre, K. Cho, Y. Bengio, Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555 (2014)"},{"issue":"8","key":"259_CR32","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"S. Hochreiter, J. Schmidhuber, Long short-term memory. Neural Comput. 9(8), 1735 (1997)","journal-title":"Neural Comput."},{"key":"259_CR33","unstructured":"K. Nguyen, K. Drossos, T. Virtanen, in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020). Temporal sub-sampling of audio feature sequences for automated audio captioning (Tokyo, 2020), pp. 110-114"},{"key":"259_CR34","unstructured":"K. Chen, Y. Wu, Z. Wang, X. Zhang, F. Nian, S. Li, X. Shao, in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020). Audio captioning based on transformer and pre-trained CNN (Tokyo, 2020), pp. 21-25"},{"key":"259_CR35","unstructured":"X. Mei, Q. Huang, X. Liu, G. Chen, J. Wu, Y. Wu, J. ZHAO, S. Li, T. Ko, H. Tang, X. Shao, M.D. Plumbley, W. Wang, in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021). An encoder-decoder based audio captioning system with transfer and reinforcement learning (Barcelona, 2021), pp. 206-210"},{"key":"259_CR36","unstructured":"Z. Ye, H. Wang, D. Yang, Y. Zou, in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021). Improving the performance of automated audio captioning via integrating the acoustic and semantic information (Barcelona, 2021), pp. 40-44"},{"key":"259_CR37","unstructured":"Q. Han, W. Yuan, D. Liu, X. Li, Z. Yang, in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021). Automated audio captioning with weakly supervised pre-training and word selection methods (Barcelona, 2021), pp. 6-10"},{"key":"259_CR38","unstructured":"S. Perez-Castanos, J. Naranjo-Alcazar, P. Zuccarello, M. Cobos, in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020). Listen carefully and tell: an audio captioning system based on residual learning and gammatone audio representation (Tokyo, 2020), pp. 150-154"},{"key":"259_CR39","doi-asserted-by":"crossref","unstructured":"A.\u00d6. Eren, M. Sert, in 2020 IEEE International Symposium on Multimedia (ISM). Audio captioning based on combined audio and semantic embeddings (IEEE, 2020), pp. 41-48","DOI":"10.1109\/ISM.2020.00014"},{"key":"259_CR40","doi-asserted-by":"crossref","unstructured":"A. Tran, K. Drossos, T. Virtanen, WaveTransformer: a novel architecture for audio captioning based on learning temporal and time-frequency information. arXiv preprint arXiv:2010.11098 (2020)","DOI":"10.23919\/EUSIPCO54536.2021.9616340"},{"issue":"11","key":"259_CR41","doi-asserted-by":"publisher","first-page":"2298","DOI":"10.1109\/TPAMI.2016.2646371","volume":"39","author":"B Shi","year":"2016","unstructured":"B. Shi, X. Bai, C. Yao, An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE Trans. Pattern. Anal. Mach. Intell. 39(11), 2298 (2016)","journal-title":"IEEE Trans. Pattern. Anal. Mach. Intell."},{"key":"259_CR42","unstructured":"D. Takeuchi, Y. Koizumi, Y. Ohishi, N. Harada, K. Kashino, in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020). Effects of word-frequency based pre- and post-processings for audio captioning (Tokyo, 2020), pp. 190-194"},{"key":"259_CR43","unstructured":"X. Xu, H. Dinkel, M. Wu, K. Yu, in Proceedings of the Detection and Classification of Acoustic Scenes and Events Workshop (DCASE). A CRNN-GRU based reinforcement learning approach to audio captioning (2020), pp. 225-229"},{"key":"259_CR44","doi-asserted-by":"crossref","unstructured":"X. Xu, H. Dinkel, M. Wu, Z. Xie, K. Yu, in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). (IEEE, 2021), pp. 905-909","DOI":"10.1109\/ICASSP39728.2021.9413982"},{"key":"259_CR45","unstructured":"A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, N. Houlsby, in International Conference on Learning Representations. An image is worth 16x16 words: Transformers for image recognition at scale (2021)"},{"key":"259_CR46","unstructured":"H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, H. J\u00e9gou, in International Conference on Machine Learning. Training data-efficient image Transformers & distillation through attention (PMLR, 2021), pp. 10,347-10,357"},{"key":"259_CR47","unstructured":"X. Mei, X. Liu, Q. Huang, M.D. Plumbley, W. Wang, in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021). Audio captioning transformer (Barcelona, 2021), pp. 211-215"},{"key":"259_CR48","unstructured":"C.P. Narisetty, T. Hayashi, R. Ishizaki, S. Watanabe, K. Takeda, in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021). Leveraging state-of-the-art ASR techniques to audio captioning (Barcelona, 2021), pp. 160-164"},{"key":"259_CR49","doi-asserted-by":"crossref","unstructured":"A. Gulati, J. Qin, C.C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, R. Pang, in INTERSPEECH. Conformer: convolution-augmented Transformer for speech recognition (ISCA, 2020), pp. 5036\u20135040","DOI":"10.21437\/Interspeech.2020-3015"},{"key":"259_CR50","volume-title":"Speech and Language Processing","author":"D Jurafsky","year":"2009","unstructured":"D. Jurafsky, J.H. Martin, Speech and Language Processing, 2nd edn. (Prentice-Hall Inc, USA, 2009)","edition":"2"},{"key":"259_CR51","unstructured":"T. Mikolov, I. Sutskever, K. Chen, G.S. Corrado, J. Dean, in Advances in Neural Information Processing Systems. Distributed representations of words and phrases and their compositionality (2013), pp. 3111\u20133119"},{"key":"259_CR52","doi-asserted-by":"crossref","unstructured":"J. Pennington, R. Socher, C.D. Manning, in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). GloVe: global vectors for word representation (2014), pp. 1532-1543","DOI":"10.3115\/v1\/D14-1162"},{"key":"259_CR53","unstructured":"T. Mikolov, \u00c9. Grave, P. Bojanowski, C. Puhrsch, A. Joulin, in Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018). Advances in pre-training distributed word representations (2018)"},{"key":"259_CR54","unstructured":"J. Devlin, M.W. Chang, K. Lee, K. Toutanova, in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). BERT: pre-training of deep bidirectional transformers for language understanding (2019), pp. 4171-4186"},{"issue":"8","key":"259_CR55","first-page":"9","volume":"1","author":"A Radford","year":"2019","unstructured":"A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al., Language models are unsupervised multitask learners. OpenAI blog. 1(8), 9 (2019)","journal-title":"OpenAI blog."},{"key":"259_CR56","unstructured":"B. Weck, X. Favory, K. Drossos, X. Serra, in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021). Evaluating off-the-shelf machine listening and natural language models for automated audio captioning (Barcelona, 2021), pp. 60-64"},{"key":"259_CR57","unstructured":"E. \u00c7ak\u0131r, K. Drossos, T. Virtanen, in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020). Multi-task regularization based on infrequent classes for audio captioning (Tokyo, 2020), pp. 6-10"},{"key":"259_CR58","doi-asserted-by":"crossref","unstructured":"M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, L. Zettlemoyer, in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension (2020), pp. 7871-7880","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"259_CR59","doi-asserted-by":"publisher","first-page":"1604","DOI":"10.1109\/LSP.2022.3189536","volume":"29","author":"F Xiao","year":"2022","unstructured":"F. Xiao, J. Guan, H. Lan, Q. Zhu, W. Wang, Local information assisted attention-free decoder for audio captioning. IEEE Signal Process. Lett. 29, 1604 (2022). https:\/\/doi.org\/10.1109\/LSP.2022.3189536","journal-title":"IEEE Signal Process. Lett."},{"key":"259_CR60","doi-asserted-by":"crossref","unstructured":"X. Mei, X. Liu, H. Liu, J. Sun, M.D. Plumbley, W. Wang, Automated audio captioning with keywords guidance. Tech. rep., DCASE2022 Challenge (2022)","DOI":"10.1186\/s13636-022-00259-2"},{"key":"259_CR61","doi-asserted-by":"crossref","unstructured":"S.J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, V. Goel, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Self-critical sequence training for image captioning (2017), pp. 7008-7024","DOI":"10.1109\/CVPR.2017.131"},{"key":"259_CR62","doi-asserted-by":"publisher","unstructured":"X. Mei, X. Liu, J. Sun, M.D. Plumbley, W. Wang, in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2022), Diverse audio captioning via adversarial training. pp. 8882-8886. https:\/\/doi.org\/10.1109\/ICASSP43922.2022.9746894","DOI":"10.1109\/ICASSP43922.2022.9746894"},{"key":"259_CR63","volume-title":"Reinforcement learning: An introduction","author":"RS Sutton","year":"2018","unstructured":"R.S. Sutton, A.G. Barto, Reinforcement Learning: An Introduction (A Bradford Book, Cambridge, 2018)"},{"key":"259_CR64","doi-asserted-by":"crossref","unstructured":"K. Drossos, S. Lipping, T. Virtanen, in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Clotho: an audio captioning dataset (IEEE, 2020), pp. 736-740","DOI":"10.1109\/ICASSP40776.2020.9052990"},{"key":"259_CR65","unstructured":"I. Martin, A. Mesaros, in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021). Diversity and bias in audio captioning datasets (Barcelona, 2021), pp. 90-94"},{"key":"259_CR66","doi-asserted-by":"publisher","unstructured":"A. Koh, X. Fuzhao, C.E. Siong, in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Automated audio captioning using transfer learning and reconstruction latent space similarity regularization (2022), pp. 7722-7726. https:\/\/doi.org\/10.1109\/ICASSP43922.2022.9747676","DOI":"10.1109\/ICASSP43922.2022.9747676"},{"key":"259_CR67","doi-asserted-by":"publisher","unstructured":"H.H. Wu, P. Seetharaman, K. Kumar, J.P. Bello, in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Wav2CLIP: learning robust audio representations from clip (2022), pp. 4563-4567. https:\/\/doi.org\/10.1109\/ICASSP43922.2022.9747669","DOI":"10.1109\/ICASSP43922.2022.9747669"},{"key":"259_CR68","unstructured":"Y. Koizumi, Y. Ohishi, D. Niizumi, D. Takeuchi, M. Yasuda, Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval. arXiv preprint arXiv:2012.07331 (2020)"},{"key":"259_CR69","unstructured":"F. Gontier, R. Serizel, C. Cerisara, in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021). Automated audio captioning by fine-tuning BART with audioset tags (Barcelona, 2021), pp. 170-174"},{"key":"259_CR70","doi-asserted-by":"crossref","unstructured":"X. Liu, X. Mei, Q. Huang, J. Sun, J. Zhao, H. Liu, M.D. Plumbley, V. K\u0131l\u0131\u00e7, W. Wang, Leveraging pre-trained BERT for audio captioning. arXiv preprint arXiv:2203.02838 (2022)","DOI":"10.23919\/EUSIPCO55093.2022.9909761"},{"key":"259_CR71","unstructured":"T. Chen, S. Kornblith, M. Norouzi, G. Hinton, in International Conference on Machine Learning. A simple framework for contrastive learning of visual representations (PMLR, 2020), pp. 1597-1607"},{"key":"259_CR72","unstructured":"K. He, H. Fan, Y. Wu, S. Xie, R. Girshick, in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020)"},{"key":"259_CR73","doi-asserted-by":"crossref","unstructured":"C. Chen, N. Hou, Y. Hu, H. Zou, X. Qi, E.S. Chng, Interactive audio-text representation for automated audio captioning with contrastive learning. arXiv preprint arXiv:2203.15526 (2022)","DOI":"10.21437\/Interspeech.2022-10510"},{"key":"259_CR74","doi-asserted-by":"crossref","unstructured":"L. Yu, W. Zhang, J. Wang, Y. Yu, in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 31. SeqGAN: sequence generative adversarial nets with policy gradient (2017)","DOI":"10.1609\/aaai.v31i1.10804"},{"key":"259_CR75","doi-asserted-by":"publisher","unstructured":"C. Narisetty, E. Tsunoo, X. Chang, Y. Kashiwagi, M. Hentschel, S. Watanabe, in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Joint speech recognition and audio captioning (2022), pp. 7892-7896. https:\/\/doi.org\/10.1109\/ICASSP43922.2022.9746601","DOI":"10.1109\/ICASSP43922.2022.9746601"},{"key":"259_CR76","unstructured":"S. Perez-Castanos, J. Naranjo-Alcazar, P. Zuccarello, M. Cobos, in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020) (Tokyo, 2020), pp. 150-154"},{"key":"259_CR77","unstructured":"H. Won, B. Kim, I.Y. Kwak, C. Lim, in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021). Transfer learning followed by transformer for automated audio captioning (Barcelona, 2021), pp. 221-225"},{"key":"259_CR78","doi-asserted-by":"crossref","unstructured":"W. Zhu, X. Wang, P. Narayana, K. Sone, S. Basu, W.Y. Wang, in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Towards understanding sample variance in visually grounded language generation: evaluations and observations (2020), pp. 8806-8811","DOI":"10.18653\/v1\/2020.emnlp-main.708"},{"key":"259_CR79","doi-asserted-by":"crossref","unstructured":"K. Papineni, S. Roukos, T. Ward, W.J. Zhu, in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. BLEU: a method for automatic evaluation of machine translation (2002), pp. 311-318","DOI":"10.3115\/1073083.1073135"},{"key":"259_CR80","unstructured":"C.Y. Lin, in Text Summarization Branches Out. Rouge: a package for automatic evaluation of summaries (Association for Computational Linguistics, 2004), pp. 74-81"},{"key":"259_CR81","unstructured":"S. Banerjee, A. Lavie, METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. Intrinsic and extrinsic evaluation measures for machine translation and\/or summarization (2005) pp. 65-72"},{"key":"259_CR82","doi-asserted-by":"crossref","unstructured":"R. Vedantam, C. Lawrence Zitnick, D. Parikh, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. CIDEr: consensus-based image description evaluation (2015), pp. 4566-4575","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"259_CR83","doi-asserted-by":"crossref","unstructured":"P. Anderson, B. Fernando, M. Johnson, S. Gould, in European Conference on Computer Vision. SPICE: semantic propositional image caption evaluation (Springer, 2016), pp. 382-398","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"259_CR84","doi-asserted-by":"crossref","unstructured":"S. Liu, Z. Zhu, N. Ye, S. Guadarrama, K. Murphy, in Proceedings of the IEEE International Conference on Computer Vision. Improved image captioning via policy gradient optimization of SPIDEr (2017), pp. 873-881","DOI":"10.1109\/ICCV.2017.100"},{"key":"259_CR85","unstructured":"T. Zhang, V. Kishore, F. Wu, K.Q. Weinberger, Y. Artzi, in International Conference on Learning Representations (2020)"},{"key":"259_CR86","doi-asserted-by":"crossref","unstructured":"N. Reimers, I. Gurevych, N. Reimers, I. Gurevych, N. Thakur, N. Reimers, J. Daxenberger, I. Gurevych, N. Reimers, I. Gurevych, et\u00a0al., in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Sentence-BERT: sentence embeddings using Siamese BERT-networks (Association for Computational Linguistics, 2019)","DOI":"10.18653\/v1\/D19-1410"},{"key":"259_CR87","doi-asserted-by":"crossref","unstructured":"Z. Zhou, Z. Zhang, X. Xu, Z. Xie, M. Wu, K.Q. Zhu, in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Can audio captions be evaluated with image caption metrics? (2022), pp. 981-985","DOI":"10.1109\/ICASSP43922.2022.9746427"},{"key":"259_CR88","doi-asserted-by":"publisher","unstructured":"J. Novikova, O. Du\u0161ek, A. Cercas Curry, V. Rieser, in Proceedings of the Conference on Empirical Methods in Natural Language Processing. Why we need new evaluation metrics for NLG (Association for Computational Linguistics, Copenhagen, 2017), pp. 2241-2252. https:\/\/doi.org\/10.18653\/v1\/D17-1238","DOI":"10.18653\/v1\/D17-1238"},{"key":"259_CR89","doi-asserted-by":"crossref","unstructured":"F. Font, G. Roma, X. Serra, in Proceedings of the 21st ACM International Conference on Multimedia. Freesound technical demo (2013), pp. 411-412","DOI":"10.1145\/2502081.2502245"},{"key":"259_CR90","unstructured":"A. Mesaros, T. Heittola, T. Virtanen, in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2018 Workshop (DCASE2018). A multi-device dataset for urban acoustic scene classification (2018), pp. 9-13"},{"key":"259_CR91","unstructured":"A. Radford, J.W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et\u00a0al., in International Conference on Machine Learning. Learning transferable visual models from natural language supervision (PMLR, 2021), pp. 8748-8763"},{"key":"259_CR92","doi-asserted-by":"crossref","unstructured":"L. Zhou, H. Palangi, L. Zhang, H. Hu, J. Corso, J. Gao, in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34. Unified vision-language pre-training for image captioning and VQA (2020), pp. 13,041-13,049","DOI":"10.1609\/aaai.v34i07.7005"},{"key":"259_CR93","doi-asserted-by":"crossref","unstructured":"X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei, et\u00a0al., in European Conference on Computer Vision. OSCAR: object-semantics aligned pre-training for vision-language tasks (Springer, 2020), pp. 121-137","DOI":"10.1007\/978-3-030-58577-8_8"},{"key":"259_CR94","unstructured":"J. Gao, X. Meng, S. Wang, X. Li, S. Wang, S. Ma, W. Gao, Masked non-autoregressive image captioning. arXiv preprint arXiv:1906.00717 (2019)"},{"key":"259_CR95","doi-asserted-by":"crossref","unstructured":"B. Dai, S. Fidler, R. Urtasun, D. Lin, in Proceedings of the IEEE International Conference on Computer Vision. Towards diverse and natural image descriptions via a conditional gan (2017), pp. 2970-2979","DOI":"10.1109\/ICCV.2017.323"},{"key":"259_CR96","unstructured":"Y. Tian, C. Guan, J. Goodman, M. Moore, C. Xu, An attempt towards interpretable audio-visual video captioning. arXiv:1812.02872 (2018)"},{"key":"259_CR97","doi-asserted-by":"crossref","unstructured":"V. Iashin, E. Rahtu, in British Machine Vision Conference (BMVC). A better use of audio-visual cues: dense video captioning with bi-modal Transformer. ArXiv abs\/2005.08271 (2020)","DOI":"10.1109\/CVPRW50498.2020.00487"},{"key":"259_CR98","unstructured":"X. Mei, X. Liu, H. Liu, J. Sun, M.D. Plumbley, W. Wang, Language-based audio retrieval with pre-trained models. Tech. rep., DCASE2022 Challenge (2022)"},{"key":"259_CR99","doi-asserted-by":"publisher","first-page":"2283","DOI":"10.1109\/TASLP.2020.3010650","volume":"28","author":"HM Fayek","year":"2020","unstructured":"H.M. Fayek, J. Johnson, Temporal reasoning via audio question answering. IEEE\/ACM Trans. Audio Speech Lang. Process. 28, 2283 (2020)","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"259_CR100","doi-asserted-by":"crossref","unstructured":"X. Liu, T. Iqbal, J. Zhao, Q. Huang, M.D. Plumbley, W. Wang, in IEEE 31st International Workshop on Machine Learning for Signal Processing (MLSP). Conditional sound generation using neural discrete time-frequency representation learning (2021) p. 1\u20136","DOI":"10.1109\/MLSP52302.2021.9596430"},{"key":"259_CR101","doi-asserted-by":"crossref","unstructured":"X. Liu, H. Liu, Q. Kong, X. Mei, J. Zhao, Q. Huang, M.D. Plumbley, W. Wang, Separate what you describe: language-queried audio source separation. arXiv:2203.15147 (2022)","DOI":"10.21437\/Interspeech.2022-10894"}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-022-00259-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-022-00259-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-022-00259-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,5]],"date-time":"2024-10-05T12:45:25Z","timestamp":1728132325000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-022-00259-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,9]]},"references-count":101,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2022,12]]}},"alternative-id":["259"],"URL":"https:\/\/doi.org\/10.1186\/s13636-022-00259-2","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,10,9]]},"assertion":[{"value":"14 April 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 September 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 October 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"WW is an editorial board member of the\u00a0<i>EURASIP Journal on Audio Speech and Music Processing<\/i> and also a guest editor of the special issue \u201cRecent Advances in Computational Sound Scene Analysis\u201d; other authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"26"}}