{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T01:48:05Z","timestamp":1782784085685,"version":"3.54.5"},"reference-count":44,"publisher":"Springer Science and Business Media LLC","issue":"35","license":[{"start":{"date-parts":[[2025,5,23]],"date-time":"2025-05-23T00:00:00Z","timestamp":1747958400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,5,23]],"date-time":"2025-05-23T00:00:00Z","timestamp":1747958400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100006302","name":"Universidad de Alcal\u00e1","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100006302","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Dense video captioning involves detecting and describing events within video sequences. Traditional methods operate in an offline setting, assuming the entire video is available for analysis. In contrast, in this work we introduce a groundbreaking paradigm: Live Video Captioning (LVC), where captions must be generated for video streams in an online manner. This shift brings unique challenges, including processing partial observations of the events and the need for a temporal anticipation of the actions. We formally define the novel problem of LVC and propose innovative evaluation metrics specifically designed for this online scenario, demonstrating their advantages over traditional metrics. To address the novel complexities of LVC, we present a new model that combines deformable transformers with temporal filtering, enabling effective captioning over video streams. Extensive experiments on the ActivityNet Captions dataset validate the proposed approach, showcasing its superior performance in the LVC setting compared to state-of-the-art offline methods. To foster further research, we provide the results of our model and an evaluation toolkit with the new metrics integrated at:\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/gramuah\/lvc\" ext-link-type=\"uri\">https:\/\/github.com\/gramuah\/lvc<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/s11042-025-20908-w","type":"journal-article","created":{"date-parts":[[2025,5,23]],"date-time":"2025-05-23T06:36:58Z","timestamp":1747982218000},"page":"44863-44895","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Live video captioning"],"prefix":"10.1007","volume":"84","author":[{"given":"Eduardo","family":"Blanco-Fern\u00e1ndez","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Carlos","family":"Guti\u00e9rrez-\u00c1lvarez","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nadia","family":"Nasri","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Saturnino","family":"Maldonado-Basc\u00f3n","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2477-0152","authenticated-orcid":false,"given":"Roberto J.","family":"L\u00f3pez-Sastre","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,5,23]]},"reference":[{"key":"20908_CR1","doi-asserted-by":"crossref","unstructured":"Vaishnavi J, Narmatha V (2024) Video captioning \u2013 a survey. Multimed Tools Appl 1\u201332","DOI":"10.1007\/s11042-024-18886-6"},{"key":"20908_CR2","doi-asserted-by":"crossref","unstructured":"Babavalian MR, Kiani K (2024) Video captioning using transformer-based GAN. Multimed Tools Appl 1\u201323","DOI":"10.2139\/ssrn.4511115"},{"issue":"10","key":"20908_CR3","doi-asserted-by":"publisher","first-page":"9718","DOI":"10.1109\/TCSVT.2024.3399933","volume":"34","author":"S Liu","year":"2024","unstructured":"Liu S, Li A, Zhao Y, Wang J, Wang Y (2024) EvCap: Element-aware video captioning. IEEE Trans Circ Syst Video Technol 34(10):9718\u20139731","journal-title":"IEEE Trans Circ Syst Video Technol"},{"issue":"28","key":"20908_CR4","doi-asserted-by":"publisher","first-page":"72113","DOI":"10.1007\/s11042-024-18372-z","volume":"83","author":"X Yao","year":"2024","unstructured":"Yao X, Zeng Y, Gu M, Yuan R, Li J, Ge J (2024) Multi-level video captioning method based on semantic space. Multimed Tools Appl 83(28):72113\u201372130","journal-title":"Multimed Tools Appl"},{"issue":"7","key":"20908_CR5","doi-asserted-by":"publisher","first-page":"3319","DOI":"10.1109\/TCSVT.2022.3232634","volume":"33","author":"M Tang","year":"2023","unstructured":"Tang M, Wang Z, Zeng Z, Li X, Zhou L (2023) Stay in grid: Improving video captioning via fully grid-level representation. IEEE Trans Circ Syst Video Technol 33(7):3319\u20133332","journal-title":"IEEE Trans Circ Syst Video Technol"},{"key":"20908_CR6","doi-asserted-by":"crossref","unstructured":"Xu J, Mei T, Yao T, Rui Y (2016) Msr-vtt: A large video description dataset for bridging video and language. In: Proceedings of the IEEE\/CVF international conference on computer vision (CVPR), pp 5288\u20135296","DOI":"10.1109\/CVPR.2016.571"},{"key":"20908_CR7","doi-asserted-by":"crossref","unstructured":"Wang X, Wu J, Chen J, Li L, Wang YF, Wang WY (2019) Vatex: A large-scale, high-quality multilingual dataset for video-and-language research. In: Proceedings of the IEEE\/CVF international conference on computer vision (ICCV), pp 4580\u20134590","DOI":"10.1109\/ICCV.2019.00468"},{"issue":"23","key":"20908_CR8","doi-asserted-by":"publisher","first-page":"64037","DOI":"10.1007\/s11042-023-17809-1","volume":"83","author":"S Artham","year":"2024","unstructured":"Artham S, Shaikh SH (2024) A neural ODE and transformer-based model for temporal understanding and dense video captioning. Multimedx Tools Appl 83(23):64037\u201364056","journal-title":"Multimedx Tools Appl"},{"key":"20908_CR9","doi-asserted-by":"crossref","unstructured":"Yang A, Nagrani A, Seo PH, Miech A, Pont-Tuset J, Laptev I, Sivic J, Schmid C (2023) Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. In: Proceedings of the IEEE\/CVF international conference on computer vision (CVPR), pp 10714\u201310726","DOI":"10.1109\/CVPR52729.2023.01032"},{"key":"20908_CR10","doi-asserted-by":"crossref","unstructured":"Wang T, Zhang R, Lu Z, Zheng F, Cheng R, Luo P (2021) End-to-end dense video captioning with parallel decoding. In: Proceedings of the IEEE\/CVF international conference on computer vision (CVPR), pp 6847\u20136857","DOI":"10.1109\/ICCV48922.2021.00677"},{"issue":"9","key":"20908_CR11","doi-asserted-by":"publisher","first-page":"3130","DOI":"10.1109\/TCSVT.2019.2936526","volume":"30","author":"Z Zhang","year":"2020","unstructured":"Zhang Z, Xu D, Ouyang W, Tan C (2020) Show, tell and summarize: Dense video captioning using visual cue aided sentence summarization. IEEE Trans Circuits Syst Video Technol 30(9):3130\u20133139. https:\/\/doi.org\/10.1109\/TCSVT.2019.2936526","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"issue":"5","key":"20908_CR12","doi-asserted-by":"publisher","first-page":"1890","DOI":"10.1109\/TCSVT.2020.3014606","volume":"31","author":"T Wang","year":"2021","unstructured":"Wang T, Zheng H, Yu M, Tian Q, Hu H (2021) Event-centric hierarchical representation for dense video captioning. IEEE Trans Circuits Syst Video Technol 31(5):1890\u20131900","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"20908_CR13","doi-asserted-by":"crossref","unstructured":"Krishna R, Hata K, Ren F, Fei-Fei L, Niebles JC (2017) Dense-captioning events in videos. In: International conference on computer vision (ICCV)","DOI":"10.1109\/ICCV.2017.83"},{"key":"20908_CR14","doi-asserted-by":"crossref","unstructured":"Zhou L, Zhou Y, Corso JJ, Socher R, Xiong C (2018) End-to-end dense video captioning with masked transformer. In: Proceedings of the IEEE\/CVF international conference on computer vision (CVPR), pp 8739\u20138748","DOI":"10.1109\/CVPR.2018.00911"},{"key":"20908_CR15","doi-asserted-by":"crossref","unstructured":"Wang J, Jiang W, Ma L, Liu W, Xu Y (2018) Bidirectional attentive fusion with context gating for dense video captioning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR)","DOI":"10.1109\/CVPR.2018.00751"},{"key":"20908_CR16","doi-asserted-by":"crossref","unstructured":"Yang D, Yuan C (2018) Hierarchical context encoding for events captioning in videos. In: 25th IEEE international conference on image processing (ICIP), pp 1288\u20131292","DOI":"10.1109\/ICIP.2018.8451740"},{"key":"20908_CR17","doi-asserted-by":"crossref","unstructured":"Iashin V, Rahtu E (2020) Multi-modal dense video captioning. In: The IEEE\/CVF conference on computer vision and pattern recognition (CVPR) Workshops","DOI":"10.1109\/CVPRW50498.2020.00487"},{"key":"20908_CR18","doi-asserted-by":"publisher","unstructured":"Li Y, Yao T, Pan Y, Chao H, Mei T (2018) Jointly localizing and describing events for dense video captioning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR). https:\/\/doi.org\/10.1109\/CVPR.2018.00782","DOI":"10.1109\/CVPR.2018.00782"},{"key":"20908_CR19","doi-asserted-by":"crossref","unstructured":"Zhou L, Zhou Y, Corso JJ, Socher R, Xiong C (2018) End-to-end dense video captioning with masked transformer. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR), pp 8739\u20138748","DOI":"10.1109\/CVPR.2018.00911"},{"key":"20908_CR20","doi-asserted-by":"crossref","unstructured":"De Geest R, Gavves E, Ghodrati A, Li C Z.and Snoek Tuytelaars T (2016) Online action detection. In: European conference on computer vision (ECCV)","DOI":"10.1007\/978-3-319-46454-1_17"},{"key":"20908_CR21","doi-asserted-by":"crossref","unstructured":"Gao J, Yang Z, Nevatia R (2017) RED: Reinforced encoder-decoder networks for action anticipation. In: British machine vision conference (BMVC)","DOI":"10.5244\/C.31.92"},{"key":"20908_CR22","doi-asserted-by":"crossref","unstructured":"De Geest R, Tuytelaars T (2018) Modeling temporal structure with LSTM for online action detection. In: IEEE Winter conference on applications of computer vision (WACV), pp 1549\u20131557","DOI":"10.1109\/WACV.2018.00173"},{"key":"20908_CR23","doi-asserted-by":"crossref","unstructured":"Li Y, Lan C, Xing J, Zeng W, Yuan C, Liu J (2016) Online human action detection using joint classification-regression recurrent neural networks. In: European conference on computer vision (ECCV)","DOI":"10.1007\/978-3-319-46478-7_13"},{"key":"20908_CR24","doi-asserted-by":"publisher","first-page":"5139","DOI":"10.1109\/ACCESS.2019.2961789","volume":"8","author":"M Baptista-R\u00edos","year":"2020","unstructured":"Baptista-R\u00edos M, L\u00f3pez-Sastre RJ, Caba Heilbron F, Van Gemert JC, Acevedo-Rodr\u00edguez FJ, Maldonado-Basc\u00f3n S (2020) Rethinking online action detection in untrimmed videos: A novel online evaluation protocol. IEEE Access 8:5139\u20135146","journal-title":"IEEE Access"},{"key":"20908_CR25","doi-asserted-by":"crossref","unstructured":"L\u00f3pez-Sastre RJ, Baptista-R\u00edos M, Acevedo-Rodr\u00edguez FJ, Mart\u00edn-Mart\u00edn P, Maldonado-Basc\u00f3n S (2021) Live video action recognition from unsupervised action proposals. In: 17th international conference on machine vision and applications (MVA)","DOI":"10.23919\/MVA51890.2021.9511355"},{"key":"20908_CR26","doi-asserted-by":"publisher","first-page":"395","DOI":"10.1016\/j.neucom.2022.03.069","volume":"491","author":"X Hu","year":"2022","unstructured":"Hu X, Dai J, Li M, Peng C, Li Y, Du S (2022) Online human action detection and anticipation in videos: A survey. Neurocomputing 491:395\u2013413","journal-title":"Neurocomputing"},{"key":"20908_CR27","doi-asserted-by":"crossref","unstructured":"Hori C, Hori T, Le Roux J (2021) Optimizing latency for online video captioning using audio-visualtransformers. In: Interspeech, pp 586\u2013590","DOI":"10.21437\/Interspeech.2021-1975"},{"key":"20908_CR28","doi-asserted-by":"crossref","unstructured":"Hoai M, Torre F (2012) Max-margin early event detectors. In: Proceedings of the IEEE international conference on computer vision (CVPR), pp 2863\u20132870","DOI":"10.1109\/CVPR.2012.6248012"},{"key":"20908_CR29","doi-asserted-by":"crossref","unstructured":"Huang D, Yao S, Wang Y, De La Torre F (2014) Sequential max-margin event detectors. In: European conference on computer vision (ECCV), pp 410\u2013424","DOI":"10.1007\/978-3-319-10578-9_27"},{"key":"20908_CR30","doi-asserted-by":"crossref","unstructured":"Kong Y, Kit D, Fu Y (2014) A discriminative model with multiple temporal scales for action prediction. In: European conference on computer vision (ECCV), pp 596\u2013611","DOI":"10.1007\/978-3-319-10602-1_39"},{"key":"20908_CR31","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I (2017) Attention is all you need. In: NeurIPS, vol 30"},{"key":"20908_CR32","unstructured":"Zhu X, Su W, Lu L, Li B, Wang X, Dai J (2020) Deformable DETR: deformable transformers for end-to-end object detection. In: Proceedings of the international conference on learning representations (ICLR)"},{"key":"20908_CR33","doi-asserted-by":"publisher","unstructured":"Suin M, Rajagopalan A (2020) An efficient framework for dense video captioning. In: Proceedings of the AAAI conference on artificial intelligence, vol 34, pp 12039\u201312046. https:\/\/doi.org\/10.1609\/aaai.v34i07.6881","DOI":"10.1609\/aaai.v34i07.6881"},{"key":"20908_CR34","doi-asserted-by":"crossref","unstructured":"Alwassel H, Giancola S, Ghanem B (2021) Tsp: Temporally-sensitive pretraining of video encoders for localization tasks. In: Proceedings of the IEEE\/CVF international conference on computer vision (ICCV) workshops","DOI":"10.1109\/ICCVW54120.2021.00356"},{"key":"20908_CR35","doi-asserted-by":"crossref","unstructured":"Rezatofighi H, Tsoi N, Gwak J, Sadeghian A, Reid I, Savarese S (2019) Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression . In: Proceedings of the IEEE\/CVF international conference on computer vision (CVPR), pp 658\u2013666","DOI":"10.1109\/CVPR.2019.00075"},{"key":"20908_CR36","doi-asserted-by":"crossref","unstructured":"Lin TY, Goyal P, Girshick R, He K, Dollar P (2017) Focal loss for dense object detection . In: Proceedings of the IEEE\/CVF international conference on computer vision (ICCV), pp 2999\u20133007","DOI":"10.1109\/ICCV.2017.324"},{"key":"20908_CR37","doi-asserted-by":"crossref","unstructured":"Denkowski M, Lavie A (2014) Meteor universal: Language specific translation evaluation for any target language. In: Proceedings of the 9th workshop on statistical machine translation, pp 376\u2013380","DOI":"10.3115\/v1\/W14-3348"},{"key":"20908_CR38","doi-asserted-by":"publisher","unstructured":"Papineni K, Roukos S, Ward T, Zhu WJ (2002) BLEU: a method for automatic evaluation of machine translation. In: Proceedings of the 40th annual meeting on association for computational linguistics, pp 311\u2013318. https:\/\/doi.org\/10.3115\/1073083.1073135","DOI":"10.3115\/1073083.1073135"},{"key":"20908_CR39","unstructured":"Lin CY (2004) ROUGE: A package for automatic evaluation of summaries. In: Text summarization branches out, pp 74\u201381"},{"key":"20908_CR40","doi-asserted-by":"crossref","unstructured":"Caba Heilbron F, Victor Escorcia BG, Niebles JC (2015) Activitynet: A large-scale video benchmark for human activity understanding. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 961\u2013970","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"20908_CR41","unstructured":"Ghanem B, Niebles JC, Snoek C, Caba Heilbron F, Alwassel H, Escorcia V, Khrisna R, Buch S, Duc Dao C (2018) The activity net large-scale activity recognition challenge 2018 summary"},{"key":"20908_CR42","doi-asserted-by":"crossref","unstructured":"Xiong Y, Dai B, Lin D (2018) Move forward and tell: A progressive generator of video descriptions. In: European conference on computer vision (ECCV), pp 489\u2013505","DOI":"10.1007\/978-3-030-01252-6_29"},{"key":"20908_CR43","doi-asserted-by":"crossref","unstructured":"Mun J, Yang L, Ren Z, Xu N, Han B (2019) Streamlined dense video captioning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR)","DOI":"10.1109\/CVPR.2019.00675"},{"issue":"32","key":"20908_CR44","doi-asserted-by":"publisher","first-page":"20011","DOI":"10.1007\/s00521-024-10208-z","volume":"36","author":"Y Akkem","year":"2024","unstructured":"Akkem Y, Biswas SK, Varanasi A (2024) Streamlit-based enhancing crop recommendation systems with advanced explainable artificial intelligence for smart farming. Neural Comput Appl 36(32):20011\u201320025","journal-title":"Neural Comput Appl"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-025-20908-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-025-20908-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-025-20908-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T10:50:02Z","timestamp":1761389402000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-025-20908-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,23]]},"references-count":44,"journal-issue":{"issue":"35","published-online":{"date-parts":[[2025,10]]}},"alternative-id":["20908"],"URL":"https:\/\/doi.org\/10.1007\/s11042-025-20908-w","relation":{},"ISSN":["1573-7721"],"issn-type":[{"value":"1573-7721","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,23]]},"assertion":[{"value":"20 September 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 April 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 May 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 May 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors certify that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest\/Competing interests"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}]}}