{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,28]],"date-time":"2026-08-28T13:58:55Z","timestamp":1787925535869,"version":"build-2784847793"},"reference-count":196,"publisher":"Springer Science and Business Media LLC","issue":"11","license":[{"start":{"date-parts":[[2023,4,11]],"date-time":"2023-04-11T00:00:00Z","timestamp":1681171200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,4,11]],"date-time":"2023-04-11T00:00:00Z","timestamp":1681171200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"publisher","award":["NRF-2019R1A2C1006159"],"award-info":[{"award-number":["NRF-2019R1A2C1006159"]}],"id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"publisher","award":["NRF-2021R1A6A1A03039493"],"award-info":[{"award-number":["NRF-2021R1A6A1A03039493"]}],"id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"publisher"}]},{"name":"2022 Yeungnam University Research Grant"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Intell Rev"],"published-print":{"date-parts":[[2023,11]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Video description refers to understanding visual content and transforming that acquired understanding into automatic textual narration. It bridges the key AI fields of computer vision and natural language processing in conjunction with real-time and practical applications. Deep learning-based approaches employed for video description have demonstrated enhanced results compared to conventional approaches. The current literature lacks a thorough interpretation of the recently developed and employed sequence to sequence techniques for video description. This paper fills that gap by focusing mainly on deep learning-enabled approaches to automatic caption generation. Sequence to sequence models follow an Encoder\u2013Decoder architecture employing a specific composition of CNN, RNN, or the variants LSTM or GRU as an encoder and decoder block. This standard-architecture can be fused with an attention mechanism to focus on a specific distinctiveness, achieving high quality results. Reinforcement learning employed within the Encoder\u2013Decoder structure can progressively deliver state-of-the-art captions by following exploration and exploitation strategies. The transformer mechanism is a modern and efficient transductive architecture for robust output. Free from recurrence, and solely based on self-attention, it allows parallelization along with training on a massive amount of data. It can fully utilize the available GPUs for most NLP tasks. Recently, with the emergence of several versions of transformers, long term dependency handling is not an issue anymore for researchers engaged in video processing for summarization and description, or for autonomous-vehicle, surveillance, and instructional purposes. They can get auspicious directions from this research.<\/jats:p>","DOI":"10.1007\/s10462-023-10414-6","type":"journal-article","created":{"date-parts":[[2023,4,11]],"date-time":"2023-04-11T04:03:34Z","timestamp":1681185814000},"page":"13293-13372","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":44,"title":["Video description: A\u00a0comprehensive survey of deep learning approaches"],"prefix":"10.1007","volume":"56","author":[{"given":"Ghazala","family":"Rafiq","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6713-8766","authenticated-orcid":false,"given":"Muhammad","family":"Rafiq","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gyu Sang","family":"Choi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,4,11]]},"reference":[{"key":"10414_CR1","unstructured":"Aafaq N, Akhtar N, Liu W, Mian A (2019a) Empirical autopsy of deep video captioning frameworks. arXiv:1911.09345"},{"key":"10414_CR2","unstructured":"Aafaq N, Akhtar N, Liu W, Mian A (2019b) Empirical autopsy of deep video captioning frameworks. arXiv:1911.09345"},{"key":"10414_CR3","doi-asserted-by":"publisher","unstructured":"Aafaq N, Mian A, Liu W, Gilani SZ, Sha M (2019c) Video description: a survey of methods, datasets, and evaluation metrics 52(6). https:\/\/doi.org\/10.1145\/3355390","DOI":"10.1145\/3355390"},{"key":"10414_CR4","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3146005","author":"N Aafaq","year":"2022","unstructured":"Aafaq N, Mian AS, Akhtar N, Liu W, Shah M (2022) Dense video captioning with early linguistic information fusion. IEEE Trans Multimedia. https:\/\/doi.org\/10.1109\/TMM.2022.3146005","journal-title":"IEEE Trans Multimedia"},{"key":"10414_CR5","doi-asserted-by":"publisher","first-page":"70797","DOI":"10.1109\/access.2021.3078295","volume":"9","author":"R Agyeman","year":"2021","unstructured":"Agyeman R, Rafiq M, Shin HK, Rinner B, Choi GS (2021) Optimizing spatiotemporal feature learning in 3D convolutional neural networks with pooling blocks. IEEE Access 9:70797\u201370805. https:\/\/doi.org\/10.1109\/access.2021.3078295","journal-title":"IEEE Access"},{"key":"10414_CR6","doi-asserted-by":"publisher","unstructured":"Al-Rfou R, Choe D, Constant N, Guo M, Jones L (2019) Character-level language modeling with deeper self-attention. Proc AAAI Conf Artif Intell 33 , 3159\u20133166. https:\/\/doi.org\/10.1609\/aaai.v33i01.33013159arxiv.org\/abs\/1808.04444","DOI":"10.1609\/aaai.v33i01.33013159"},{"key":"10414_CR7","doi-asserted-by":"publisher","unstructured":"Alzubaidi L, Zhang J, Humaidi AJ, Al-Dujaili A, Duan Y, Al-Shamma O, et al (2021) Review of deep learning: concepts, CNN architectures, challenges, applications, future directions 8(1). https:\/\/doi.org\/10.1186\/s40537-021-00444-8","DOI":"10.1186\/s40537-021-00444-8"},{"key":"10414_CR8","doi-asserted-by":"publisher","unstructured":"Amaresh M, Chitrakala S (2019) Video captioning using deep learning: an overview of methods, datasets and metrics. Proceedings of the 2019 IEEE international conference on communication and signal processing, ICCSP 2019 (pp. 656\u2013661). https:\/\/doi.org\/10.1109\/ICCSP.2019.8698097","DOI":"10.1109\/ICCSP.2019.8698097"},{"key":"10414_CR9","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279","author":"S Antol","year":"2015","unstructured":"Antol S, Agrawal A, Lu J, Mitchell M, Batra D, Zitnick CL, Parikh D (2015) VQA: visual question answering. Proc IEEE Int Conf Comput Vis. https:\/\/doi.org\/10.1109\/ICCV.2015.279","journal-title":"Proc IEEE Int Conf Comput Vis"},{"key":"10414_CR10","doi-asserted-by":"publisher","unstructured":"Arnab A, Dehghani M, Heigold G, Sun C, Lu\u010di\u0107 M, Schmid C (2021) ViViT: a video vision transformer. Proceedings of the IEEE international conference on computer vision, 6816\u20136826. https:\/\/doi.org\/10.1109\/ICCV48922.2021.00676arXiv:2103.15691","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"10414_CR11","doi-asserted-by":"crossref","unstructured":"Babariya RJ, Tamaki T (2020) Meaning guided video captioning. In: Pattern Recognition: 5th Asian Conference, ACPR 2019, Auckland, New Zealand, November 26\u201329, 2019, Revised Selected Papers, Part II 5, pp 478\u2013488. Springer International Publishing","DOI":"10.1007\/978-3-030-41299-9_37"},{"key":"10414_CR12","unstructured":"Bahdanau D, Cho KH, Bengio Y (2015) Neural machine translation by jointly learning to align and translate. 3rd International Conference on Learning Representations, ICLR 2015 -Conference Track Proceedings, 1\u201315. arXiv:1409.0473"},{"key":"10414_CR13","first-page":"102","volume":"2012","author":"A Barbu","year":"2012","unstructured":"Barbu A, Bridge A, Burchill Z, Coroian D, Dickinson S, Fidler S, Zhang Z (2012) Video in sentences out. Uncertainty Artif Intell\u2013Proc 28th Conf\u2013UAI 2012:102\u2013112 arXiv:1204.2742","journal-title":"Uncertainty Artif Intell\u2013Proc 28th Conf\u2013UAI"},{"key":"10414_CR14","doi-asserted-by":"publisher","DOI":"10.1145\/1553374.1553380","author":"Y Bengio","year":"2009","unstructured":"Bengio Y, Louradour J, Collobert R, Weston J (2009) Curriculum learning. ACM Int Conf Proc Ser. https:\/\/doi.org\/10.1145\/1553374.1553380","journal-title":"ACM Int Conf Proc Ser"},{"key":"10414_CR15","doi-asserted-by":"publisher","first-page":"328","DOI":"10.1109\/CIC.2017.00050","volume":"2017","author":"S Bhatt","year":"2017","unstructured":"Bhatt S, Patwa F, Sandhu R (2017) Natural language processing (almost) from scratch. Proc IEEE 3rd Int Conf Collaboration Internet Comput CIC 2017 2017:328\u2013338. https:\/\/doi.org\/10.1109\/CIC.2017.00050","journal-title":"Proc IEEE 3rd Int Conf Collaboration Internet Comput CIC 2017"},{"key":"10414_CR16","unstructured":"Bilkhu M, Wang S, Dobhal T (2019) Attention is all you need for videos: self-attention based video summarization using universal Transformers. arXiv:1906.02792"},{"issue":"7","key":"10414_CR17","doi-asserted-by":"publisher","first-page":"2631","DOI":"10.1109\/TCYB.2018.2831447","volume":"49","author":"Y Bin","year":"2019","unstructured":"Bin Y, Yang Y, Shen F, Xie N, Shen HT, Li X (2019) Describing video with attention-based bidirectional LSTM. IEEE Trans Cybern 49(7):2631\u20132641. https:\/\/doi.org\/10.1109\/TCYB.2018.2831447","journal-title":"IEEE Trans Cybern"},{"key":"10414_CR18","doi-asserted-by":"publisher","unstructured":"Blohm M, Jagfeld G, Sood E, Yu X, Vu NT (2018) Comparing attention-based convolutional and recurrent neural networks: success and limitations in machine reading comprehension. CoNLL 2018\u201322nd Conference on Computational Natural Language Learning, Proceedings, 108\u2013118. https:\/\/doi.org\/10.18653\/v1\/k18-1011arXiv:1808.08744","DOI":"10.18653\/v1\/k18-1011"},{"issue":"May","key":"10414_CR19","first-page":"25","volume":"3024","author":"T Brox","year":"2014","unstructured":"Brox T, Papenberg N, Weickert J (2014) High accuracy optical flow estimation based on warping\u2013presentation. Lecture Notes Comput Sci (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 3024(May):25\u201336","journal-title":"Lecture Notes Comput Sci (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)"},{"issue":"8","key":"10414_CR20","first-page":"1735","volume":"9","author":"R Cascade-correlation","year":"1997","unstructured":"Cascade-correlation R, Chunking NS (1997) Long Short\u2013Term Memory 9(8):1735\u20131780","journal-title":"Long Short\u2013Term Memory"},{"key":"10414_CR21","unstructured":"Chen DL, Dolan WB (2011) Collecting highly parallel data for paraphrase evaluation. Aclhlt 2011\u2013Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies 1 (pp. 190\u2013200)"},{"key":"10414_CR22","doi-asserted-by":"publisher","unstructured":"Chen DZ, Gholami A, Niesner M, Chang AX (2021) Scan2Cap: context-aware dense captioning in RGB-D scans. 3192\u20133202. https:\/\/doi.org\/10.1109\/cvpr46437.2021.00321arXiv:2012.02206","DOI":"10.1109\/cvpr46437.2021.00321"},{"key":"10414_CR23","unstructured":"Chen H, Li J, Hu X (2020) Delving deeper into the decoder for video captioning. arXiv:2001.05614"},{"key":"10414_CR24","unstructured":"Chen H, Lin K, Maye A, Li J, Hu X (2019a) A semantics-assisted video captioning model trained with scheduled sampling. https:\/\/zhuanzhi.ai\/paper\/f88d29f09d1a56a1b1cf719dfc55ea61arXiv:1909.00121"},{"key":"10414_CR25","doi-asserted-by":"publisher","unstructured":"Chen J, Pan Y, Li Y, Yao T, Chao H, Mei T (2019b) Temporal deformable convolutional encoder\u2013decoder networks for video captioning. Proc AAAI Conf Artif Intell 33 , 8167\u20138174. https:\/\/doi.org\/10.1609\/aaai.v33i01.33018167arXiv:1905.01077","DOI":"10.1609\/aaai.v33i01.33018167"},{"issue":"1997","key":"10414_CR26","first-page":"847","volume":"95","author":"M Chen","year":"2018","unstructured":"Chen M, Li Y, Zhang Z, Huang S (2018) TVT: two-view transformer network for video captioning. Proc Mach Learn Res 95(1997):847\u2013862","journal-title":"Proc Mach Learn Res"},{"key":"10414_CR27","doi-asserted-by":"publisher","first-page":"8191","DOI":"10.1609\/aaai.v33i01.33018191","volume":"33","author":"S Chen","year":"2019","unstructured":"Chen S, Jiang Y-G (2019) Motion guided spatial attention for video captioning. Proc AAAI Conf Artif Intel 33:8191\u20138198. https:\/\/doi.org\/10.1609\/aaai.v33i01.33018191","journal-title":"Proc AAAI Conf Artif Intel"},{"key":"10414_CR28","doi-asserted-by":"publisher","first-page":"8421","DOI":"10.1109\/CVPR46437.2021.00832","volume":"1","author":"S Chen","year":"2021","unstructured":"Chen S, Jiang YG (2021c) Towards bridging event captioner and sentence localizer for weakly supervised dense event captioning. Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn 1:8421\u20138431. https:\/\/doi.org\/10.1109\/CVPR46437.2021.00832","journal-title":"Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn"},{"key":"10414_CR29","doi-asserted-by":"publisher","first-page":"6283","DOI":"10.24963\/ijcai.2019\/877","volume":"2019","author":"S Chen","year":"2019","unstructured":"Chen S, Yao T, Jiang YG (2019b) Deep learning for video captioning: a review. IJCAI Int Joint Conf Artif Intell 2019:6283\u20136290. https:\/\/doi.org\/10.24963\/ijcai.2019\/877","journal-title":"IJCAI Int Joint Conf Artif Intell"},{"key":"10414_CR30","doi-asserted-by":"publisher","first-page":"367","DOI":"10.1007\/978-3-030-01261-8_22","volume":"11217","author":"Y Chen","year":"2018","unstructured":"Chen Y, Wang S, Zhang W, Huang Q (2018) Less is more: picking informative frames for video captioning. Lecture Notes Comput Sci (including subseries lecture notes in artificial intelligence and lecture notes in bioinformatics) 11217:367\u2013384. https:\/\/doi.org\/10.1007\/978-3-030-01261-8_22","journal-title":"Lecture Notes Comput Sci (including subseries lecture notes in artificial intelligence and lecture notes in bioinformatics)"},{"key":"10414_CR31","first-page":"1","volume":"2018","author":"Y Chen","year":"2018","unstructured":"Chen Y, Zhang W, Wang S, Li L, Huang Q (2018) Saliency-based spatiotemporal attention for video captioning. 2018 IEEE 4th Int Conf Multimedia Big Data BigMM 2018:1\u20138","journal-title":"2018 IEEE 4th Int Conf Multimedia Big Data BigMM"},{"key":"10414_CR32","unstructured":"Child R, Gray S, Radford A, Sutskever I (2019) Generating Long Sequences with Sparse Transformers. arXiv:1904.10509"},{"key":"10414_CR33","doi-asserted-by":"publisher","unstructured":"Cho K, Van Merri\u00ebnboer B, Gulcehre C, Bahdanau D, Bougares F, Schwenk H, Bengio Y (2014) Learning phrase representations using RNN encoder\u2013decoder for statistical machine translation. EMNLP 2014 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, 1724\u20131734. https:\/\/doi.org\/10.3115\/v1\/d14-1179arXiv:1406.1078","DOI":"10.3115\/v1\/d14-1179"},{"key":"10414_CR34","doi-asserted-by":"publisher","unstructured":"Dai Z, Yang Z, Yang Y, Carbonell J, Le QV, Salakhutdinov R (2020) Transformer-XL: Attentive language models beyond a fixed-length context. ACL 2019 -57th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference, 2978\u20132988. https:\/\/doi.org\/10.18653\/v1\/p19-1285arXiv:1901.02860","DOI":"10.18653\/v1\/p19-1285"},{"key":"10414_CR35","doi-asserted-by":"publisher","unstructured":"Das P, Xu C, Doell RF, Corso JJ (2013) A thousand frames in just a few words: lingual description of videos through latent topics and sparse object stitching. Proceedings of the IEEE computer society conference on computer vision and pattern recognition (pp. 2634\u20132641). https:\/\/doi.org\/10.1109\/CVPR.2013.340","DOI":"10.1109\/CVPR.2013.340"},{"key":"10414_CR36","doi-asserted-by":"publisher","unstructured":"Demeester T, Rockt\u00e4schel T, Riedel S (2016) Lifted rule injection for relation embeddings. Emnlp 2016\u2014conference on empirical methods in natural language processing, proceedings (pp. 1389\u20131399). https:\/\/doi.org\/10.18653\/v1\/d16-1146","DOI":"10.18653\/v1\/d16-1146"},{"key":"10414_CR37","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00030","author":"C Deng","year":"2021","unstructured":"Deng C, Chen S, Chen D, He Y, Wu Q (2021) Sketch, ground, and refine: top-down dense video captioning. Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn. https:\/\/doi.org\/10.1109\/CVPR46437.2021.00030","journal-title":"Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn"},{"key":"10414_CR38","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li LJ, Li K, Fei-Fei L (2009, June). Imagenet: a large-scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition, pp 248\u2013255. IEEE","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"10414_CR39","doi-asserted-by":"publisher","unstructured":"Doddington G (2002) Automatic evaluation of machine translation quality using n-gram co-occurrence statistics, 138. https:\/\/doi.org\/10.3115\/1289189.1289273","DOI":"10.3115\/1289189.1289273"},{"issue":"4","key":"10414_CR40","doi-asserted-by":"publisher","first-page":"677","DOI":"10.1109\/TPAMI.2016.2599174","volume":"39","author":"J Donahue","year":"2017","unstructured":"Donahue J, Hendricks LA, Rohrbach M, Venugopalan S, Guadarrama S, Saenko K, Darrell T (2017) Long-term recurrent convolutional networks for visual recognition and description. IEEE Trans Pattern Analys Mach Intell 39(4):677\u2013691. https:\/\/doi.org\/10.1109\/TPAMI.2016.2599174","journal-title":"IEEE Trans Pattern Analys Mach Intell"},{"key":"10414_CR41","doi-asserted-by":"publisher","first-page":"452","DOI":"10.3115\/v1\/p14-2074","volume":"2","author":"D Elliott","year":"2014","unstructured":"Elliott D, Keller F (2014) Comparing automatic evaluation measures for image description. 52nd Annu Meet Assoc Comput Linguistics ACL 2014\u2013Proc Conf 2:452\u2013457. https:\/\/doi.org\/10.3115\/v1\/p14-2074","journal-title":"52nd Annu Meet Assoc Comput Linguistics ACL 2014\u2013Proc Conf"},{"key":"10414_CR42","unstructured":"Estevam V, Laroca R, Pedrini H, Menotti D (2021) Dense video captioning using unsupervised semantic information. arXiv:2112.08455v1"},{"key":"10414_CR43","doi-asserted-by":"crossref","unstructured":"Fang Z, Gokhale T, Banerjee P, Baral C, Yang Y (2020) Video2Commonsense: generating commonsense descriptions to enrich video captioning. arXiv:2003.05162","DOI":"10.18653\/v1\/2020.emnlp-main.61"},{"issue":"9","key":"10414_CR44","doi-asserted-by":"publisher","first-page":"2045","DOI":"10.1109\/TMM.2017.2729019","volume":"19","author":"L Gao","year":"2017","unstructured":"Gao L, Guo Z, Zhang H, Xu X, Shen HT (2017) Video captioning with attention-based lstm and semantic consistency. IEEE Trans Multimedia 19(9):2045\u20132055. https:\/\/doi.org\/10.1109\/TMM.2017.2729019","journal-title":"IEEE Trans Multimedia"},{"key":"10414_CR45","doi-asserted-by":"publisher","first-page":"202","DOI":"10.1109\/TIP.2021.3120867","volume":"31","author":"L Gao","year":"2022","unstructured":"Gao L, Lei Y, Zeng P, Song J, Wang M, Shen HT (2022) Hierarchical representation network with auxiliary tasks for video captioning and video question answering. IEEE Trans Image Process 31:202\u2013215. https:\/\/doi.org\/10.1109\/TIP.2021.3120867","journal-title":"IEEE Trans Image Process"},{"issue":"8","key":"10414_CR46","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/tpami.2019.2894139","volume":"14","author":"L Gao","year":"2019","unstructured":"Gao L, Li X, Song J, Shen HT (2019) Hierarchical LSTMs with adaptive attention for visual captioning. IEEE Trans Pattern Analys Mach Intell 14(8):1\u20131. https:\/\/doi.org\/10.1109\/tpami.2019.2894139","journal-title":"IEEE Trans Pattern Analys Mach Intell"},{"key":"10414_CR47","doi-asserted-by":"publisher","first-page":"222","DOI":"10.1016\/j.neucom.2018.06.096","volume":"395","author":"L Gao","year":"2020","unstructured":"Gao L, Wang X, Song J, Liu Y (2020) Fused GRU with semantic-temporal attention for video captioning. Neurocomputing 395:222\u2013228. https:\/\/doi.org\/10.1016\/j.neucom.2018.06.096","journal-title":"Neurocomputing"},{"key":"10414_CR48","unstructured":"Gehring J, Dauphin YN (2016) Convolutional Sequence to Sequence Learning. https:\/\/proceedings.mlr.press\/v70\/gehring17a\/gehring17a.pdf"},{"key":"10414_CR49","doi-asserted-by":"crossref","unstructured":"Gella S, Lewis M, Rohrbach M (2020) A dataset for telling the stories of social media videos. Proceedings of the 2018 conference on empirical methods in natural language processing, EMNLP 2018:968\u2013974","DOI":"10.18653\/v1\/D18-1117"},{"key":"10414_CR50","unstructured":"Ging S, Zolfaghari M, Pirsiavash H, Brox T (2020) COOT: cooperative hierarchical transformer for video-text representation learning. (NeurIPS):1\u201327. arXiv:2011.00597"},{"key":"10414_CR51","unstructured":"Gomez AN, Ren M, Urtasun R, Grosse RB (2017) The reversible resid-ual network: backpropagation without storing activations. Adv Neural Inform Process Syst 2017:2215\u20132225. arXiv:1707.04585"},{"key":"10414_CR52","unstructured":"Goodfellow I, Bengio Y, Courville A (2016) Deep learning. MIT Press. (http:\/\/www.deeplearningbook.org)"},{"key":"10414_CR53","unstructured":"Goyal A, Lamb A, Zhang Y, Zhang S, Courville A, Bengio Y (2016) Professor forcing: anew algorithm for training recurrent networks. Adv Neural Inform Process Syst (Nips):4608\u20134616. arXiv:1610.09038"},{"key":"10414_CR54","unstructured":"Hakeem A, Sheikh Y, Shah M (2004) CASE E: a hierarchical event representation for the analysis of videos. Proc Natl Conf Artif Intell:263\u2013268"},{"key":"10414_CR55","doi-asserted-by":"crossref","unstructured":"Hammad M, Hammad M, Elshenawy M (2019) Characterizing the impact of using features extracted from pretrained models on the quality of video captioning sequence-to-sequence models. arXiv:1911.09989","DOI":"10.1007\/978-3-030-59830-3_21"},{"key":"10414_CR56","unstructured":"Hammoudeh A, Vanderplaetse B, Dupont S (2022) Deep soccer captioning with transformer: dataset, semantics-related losses, and multi-level evaluation:1\u201315. arXiv:2202.05728"},{"key":"10414_CR57","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TPAMI.2022.3152247","volume":"8828","author":"K Han","year":"2022","unstructured":"Han K, Wang Y, Chen H, Chen X, Guo J, Liu Z, Tao D (2022) A survey on vision transformer. IEEE Trans Pattern Analys Mach Intel 8828:1\u201320. https:\/\/doi.org\/10.1109\/TPAMI.2022.3152247","journal-title":"IEEE Trans Pattern Analys Mach Intel"},{"key":"10414_CR58","doi-asserted-by":"publisher","first-page":"8393","DOI":"10.1609\/aaai.v33i01.33018393","volume":"33","author":"D He","year":"2019","unstructured":"He D, Zhao X, Huang J, Li F, Liu X, Wen S (2019) Read, watch, and move: reinforcement learning for temporally grounding natural language descriptions in videos. Proceed AAAI Conf Artif Intel 33:8393\u20138400. https:\/\/doi.org\/10.1609\/aaai.v33i01.33018393. arXiv:1901.06829","journal-title":"Proceed AAAI Conf Artif Intel"},{"key":"10414_CR59","doi-asserted-by":"publisher","first-page":"770","DOI":"10.1109\/CVPR.2016.90","volume":"2016","author":"K He","year":"2016","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn 2016:770\u2013778. https:\/\/doi.org\/10.1109\/CVPR.2016.90","journal-title":"Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn"},{"key":"10414_CR60","doi-asserted-by":"publisher","first-page":"4203","DOI":"10.1109\/ICCV.2017.450","volume":"2017","author":"C Hori","year":"2017","unstructured":"Hori C, Hori T, Lee TY, Zhang Z, Harsham B, Hershey JR et al (2017) Attention-based multimodal fusion for video description. Proc IEEE Int Conf Comput Vis 2017:4203\u20134212. https:\/\/doi.org\/10.1109\/ICCV.2017.450","journal-title":"Proc IEEE Int Conf Comput Vis"},{"key":"10414_CR61","doi-asserted-by":"crossref","unstructured":"Hosseinzadeh M, Wang Y, Canada HT (2021) Video captioning of future frames. Winter Conf App Comput Vis:980\u2013989","DOI":"10.1109\/WACV48630.2021.00102"},{"key":"10414_CR62","doi-asserted-by":"publisher","first-page":"8917","DOI":"10.1109\/ICCV.2019.00901","volume":"2019","author":"J Hou","year":"2019","unstructured":"Hou J, Wu X, Zhao W, Luo J, Jia Y (2019) Joint syntax representation learning and visual cue translation for video captioning. IEEE Int Conf Comput Vis 2019:8917\u20138926. https:\/\/doi.org\/10.1109\/ICCV.2019.00901","journal-title":"IEEE Int Conf Comput Vis"},{"key":"10414_CR63","doi-asserted-by":"publisher","DOI":"10.1155\/2022\/3454167","author":"A Hussain","year":"2022","unstructured":"Hussain A, Hussain T, Ullah W, Baik SW (2022) Vision transformer and deep sequence learning for human activity recognition in surveillance videos. Comput Intel Neurosci. https:\/\/doi.org\/10.1155\/2022\/3454167","journal-title":"Comput Intel Neurosci"},{"key":"10414_CR64","unstructured":"Husz\u00e1r F (2015) How (not) to train your generative model: scheduled sampling, likelihood, adversary?:1\u20139. arXiv:1511.05101"},{"key":"10414_CR65","doi-asserted-by":"crossref","unstructured":"Iashin V, Rahtu E (2020) Multi-modal dense video captioning. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pp 958\u2013959","DOI":"10.1109\/CVPRW50498.2020.00487"},{"issue":"13","key":"10414_CR66","doi-asserted-by":"publisher","first-page":"4817","DOI":"10.3390\/s22134817","volume":"22","author":"H Im","year":"2022","unstructured":"Im H, Choi Y-S (2022) UAT: universal attention transformer for video captioning. Sensors 22(13):4817. https:\/\/doi.org\/10.3390\/s22134817","journal-title":"Sensors"},{"key":"10414_CR67","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2021.108332","volume":"117","author":"W Ji","year":"2022","unstructured":"Ji W, Wang R, Tian Y, Wang X (2022) An attention based dual learning approach for video captioning. Appl Soft Comput 117:108332. https:\/\/doi.org\/10.1016\/j.asoc.2021.108332","journal-title":"Appl Soft Comput"},{"key":"10414_CR68","doi-asserted-by":"publisher","unstructured":"Jia Y, Shelhamer E, Donahue J, Karayev S, Long J, Girshick R, et al. (2014) Caffe: convolutional architecture for fast feature embedding. Mm 2014\u2013proceedings of the 2014 ACM conference on multimedia (pp. 675-678). https:\/\/doi.org\/10.1145\/2647868.2654889","DOI":"10.1145\/2647868.2654889"},{"key":"10414_CR69","doi-asserted-by":"publisher","unstructured":"Jin T, Huang S, Chen M, Li Y, Zhang Z (2020) SBAT: Video captioning with sparse boundary-aware transformer. IJCAI Int Joint Conf Artif Intel 2021:630\u2013636. https:\/\/doi.org\/10.24963\/ijcai.2020.88","DOI":"10.24963\/ijcai.2020.88"},{"key":"10414_CR70","doi-asserted-by":"publisher","unstructured":"Karpathy A, Toderici G, Shetty S, Leung T, Sukthankar R, Li FF (2014) Large-scale video classification with convolutional neural net-works. Proceedings of the IEEE computer society conference on computer vision and pattern recognition (pp. 1725\u20131732). https:\/\/doi.org\/10.1109\/CVPR.2014.223","DOI":"10.1109\/CVPR.2014.223"},{"key":"10414_CR71","unstructured":"Kay W, Carreira J, Simonyan K, Zhang B, Hillier C, Vijayanarasimhan S, et al. (2017) The kinetics human action video dataset. arXiv:1705.06950"},{"key":"10414_CR72","doi-asserted-by":"crossref","unstructured":"Kazemzadeh S, Ordonez V, Matten M, Berg TL (2014) ReferItGame: referring to objects in photographs of natural scenes:787\u2013798","DOI":"10.3115\/v1\/D14-1086"},{"key":"10414_CR73","unstructured":"Kenton M-wC, Kristina L, Devlin J (1953) BERT: pre-training of deep bidirectional transformers for language understanding. (Mlm). arXiv:1810.04805v2"},{"key":"10414_CR74","unstructured":"Khan M, Gotoh Y (2012) Describing video contents in natural language. Proceedings of the workshop on innovative hybrid (pp. 27\u201335)"},{"key":"10414_CR75","doi-asserted-by":"publisher","unstructured":"Kilickaya M, Erdem A, Ikizler-Cinbis N, Erdem E (2017) Re-evaluating automatic metrics for image captioning. 15th conference of the european chapter of the association for computational linguistics, EACL 2017\u2013proceedings of conference (Vol. 1, pp. 199-209). Association for Computational Linguistics (ACL). https:\/\/doi.org\/10.18653\/v1\/e17-1019","DOI":"10.18653\/v1\/e17-1019"},{"key":"10414_CR76","unstructured":"Kitaev N, Kaiser L, Levskaya A (2020) Reformer: the efficient transformer, 1\u201312. arXiv:2001.04451"},{"issue":"2","key":"10414_CR77","doi-asserted-by":"publisher","first-page":"171","DOI":"10.1023\/A:1020346032608","volume":"50","author":"A Kojima","year":"2002","unstructured":"Kojima A, Tamura T, Fukunaga K (2002) Natural language description of human activities from video images based on concept hierarchy of actions. Int J Comput Vis 50(2):171\u2013184. https:\/\/doi.org\/10.1023\/A:1020346032608","journal-title":"Int J Comput Vis"},{"key":"10414_CR78","doi-asserted-by":"publisher","first-page":"706","DOI":"10.1109\/ICCV.2017.83","volume":"2017","author":"R Krishna","year":"2017","unstructured":"Krishna R, Hata K, Ren F, Fei-Fei L, Niebles JC (2017) Dense-captioning events in videos. Proc Int Conf Comput Vis 2017:706\u2013715. https:\/\/doi.org\/10.1109\/ICCV.2017.83","journal-title":"Proc Int Conf Comput Vis"},{"key":"10414_CR79","unstructured":"Langkilde-geary I, Knight K (2002) HALogen statistical sentence generator. (July):102\u2013103"},{"key":"10414_CR80","first-page":"44","volume":"2015","author":"N Laokulrat","year":"2016","unstructured":"Laokulrat N, Phan S, Nishida N, Shu R, Ehara Y, Okazaki N, Nakayama H (2016) Generating video description using sequence-to-sequence model with temporal attention. Coling 2015:44\u201352","journal-title":"Coling"},{"key":"10414_CR81","doi-asserted-by":"crossref","unstructured":"Lavie A, Agarwal A (2007) METEOR: an automatic metric for mt evaluation with improved correlation with human judgments. Proceedings of the Second Workshop on Statistical Machine Translation (June):228\u2013231. http:\/\/acl.ldc.upenn.edu\/W\/W05\/W05-09.pdf","DOI":"10.3115\/1626355.1626389"},{"key":"10414_CR82","doi-asserted-by":"publisher","first-page":"134","DOI":"10.1007\/978-3-540-30194-3-16","volume":"3265","author":"A Lavie","year":"2004","unstructured":"Lavie A, Sagae K, Jayaraman S (2004) The significance of recall in automatic metrics for MT evaluation. Lecture Notes Comput Sci (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 3265:134\u2013143. https:\/\/doi.org\/10.1007\/978-3-540-30194-3-16","journal-title":"Lecture Notes Comput Sci (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)"},{"key":"10414_CR83","doi-asserted-by":"publisher","unstructured":"Lee J, Lee Y, Seong S, Kim K, Kim S, Kim J (2019) Capturing long-range dependencies in video captioning. Proc Int Conf Image Process, ICIP, 2019:1880\u20131884. https:\/\/doi.org\/10.1109\/ICIP.2019.8803143","DOI":"10.1109\/ICIP.2019.8803143"},{"key":"10414_CR84","doi-asserted-by":"publisher","DOI":"10.1155\/2018\/3125879","author":"S Lee","year":"2018","unstructured":"Lee S, Kim I (2018) Multimodal feature learning for video captioning. Math Prob Eng. https:\/\/doi.org\/10.1155\/2018\/3125879","journal-title":"Math Prob Eng"},{"key":"10414_CR85","doi-asserted-by":"publisher","unstructured":"Lei J, Wang L, Shen Y, Yu D, Berg T, Bansal M (2020) MART: memory-augmented recurrent transformer for coherent video paragraph captioning:2603\u20132614. https:\/\/doi.org\/10.18653\/v1\/2020.acl-main.233arXiv:2005.05402","DOI":"10.18653\/v1\/2020.acl-main.233"},{"key":"10414_CR86","doi-asserted-by":"publisher","first-page":"447","DOI":"10.1007\/978-3-030-58589-1_27","volume":"12366","author":"J Lei","year":"2020","unstructured":"Lei J, Yu L, Berg TL, Bansal M (2020) TVR: a large-scale dataset for video-subtitle moment retrieval. Lecture Notes Comput Sci (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 12366:447\u2013463. https:\/\/doi.org\/10.1007\/978-3-030-58589-1_27","journal-title":"Lecture Notes Comput Sci (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)"},{"key":"10414_CR87","doi-asserted-by":"crossref","unstructured":"Levine R, Meurers D (2006) Head-driven phrase structure grammar linguistic approach , formal head-driven phrase structure grammar linguistic approach , formal foundations , and computational realization (January)","DOI":"10.1002\/0470018860.s00225"},{"key":"10414_CR88","unstructured":"Li J, Qiu H (2020) Comparing attention-based neural architectures for video captioning,  vol 1194.  Available on: https:\/\/web.stanford.edu\/class\/archive\/cs\/cs224n\/cs224n"},{"key":"10414_CR90","doi-asserted-by":"publisher","unstructured":"Li L, Chen Y-C, Cheng Y, Gan Z, Yu L, Liu J (2020) HERO: hierarchical encoder for video+language omni-representation pre-training, 2046\u20132065. https:\/\/doi.org\/10.18653\/v1\/2020.emnlp-main.161arXiv:2005.00200","DOI":"10.18653\/v1\/2020.emnlp-main.161"},{"issue":"4","key":"10414_CR91","doi-asserted-by":"publisher","first-page":"297","DOI":"10.1109\/tetci.2019.2892755","volume":"3","author":"S Li","year":"2019","unstructured":"Li S, Tao Z, Li K, Fu Y (2019) Visual to text: survey of image and video captioning. IEEE Trans Emerg Top Comput Intel 3(4):297\u2013312. https:\/\/doi.org\/10.1109\/tetci.2019.2892755","journal-title":"IEEE Trans Emerg Top Comput Intel"},{"key":"10414_CR92","doi-asserted-by":"publisher","unstructured":"Li X, Zhao B, Lu X (2017) MAM-RNN: Multi-level attention model based RNN for video captioning. IJCAI International Joint Conference on Artificial Intelligence, 2208\u20132214. https:\/\/doi.org\/10.24963\/ijcai.2017\/307","DOI":"10.24963\/ijcai.2017\/307"},{"issue":"2","key":"10414_CR93","doi-asserted-by":"publisher","first-page":"621","DOI":"10.1007\/s11280-018-0531-z","volume":"22","author":"X Li","year":"2019","unstructured":"Li X, Zhou Z, Chen L, Gao L (2019) Residual attention-based LSTM for video captioning. World Wide Web 22(2):621\u2013636. https:\/\/doi.org\/10.1007\/s11280-018-0531-z","journal-title":"World Wide Web"},{"key":"10414_CR94","doi-asserted-by":"publisher","unstructured":"Li Y, Yao T, Pan Y, Chao H, Mei T (2018) Jointly localizing and describing events for dense video captioning. Proceedings of the IEEE computer society conference on computer vision and pattern recognition (pp. 7492\u20137500). https:\/\/doi.org\/10.1109\/CVPR.2018.00782","DOI":"10.1109\/CVPR.2018.00782"},{"key":"10414_CR95","unstructured":"Lin C-Y (2004) ROUGE: A Package for Automatic Evaluation of Summaries. In: Text summarization branches out. Association for Computational Linguistics. Barcelona, Spain, pp 74\u201381. https:\/\/aclanthology.org\/W04-1013"},{"key":"10414_CR96","unstructured":"Lin K, Gan Z, Wang L (2020) Multi-modal feature fusion with feature attention for vatex captioning challenge 2020:2\u20135. arXiv:2006.03315"},{"key":"10414_CR97","doi-asserted-by":"publisher","unstructured":"Liu F, Ren X, Wu X, Yang B, Ge S, Sun X (2021) O2NA: an object-oriented non-autoregressive approach for controllable video captioning. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021:281\u2013292. https:\/\/doi.org\/10.18653\/v1\/2021.findings-acl.24arXiv:2108.02359","DOI":"10.18653\/v1\/2021.findings-acl.24"},{"key":"10414_CR98","doi-asserted-by":"publisher","unstructured":"Liu S, Ren Z, Yuan J (2018) SibNet: Sibling convolutional encoder for video captioning. MM 2018 -Proceedings of the 2018 ACM Multimedia Conference, 1425\u20131434. https:\/\/doi.org\/10.1145\/3240508.3240667","DOI":"10.1145\/3240508.3240667"},{"key":"10414_CR99","doi-asserted-by":"publisher","unstructured":"Liu S, Ren Z, Yuan J (2020) SibNet: sibling convolutional encoder for video captioning. IEEE Trans Pattern Analys Mach Intel, 1\u20131. https:\/\/doi.org\/10.1109\/tpami.2019.2940007","DOI":"10.1109\/tpami.2019.2940007"},{"key":"10414_CR100","doi-asserted-by":"publisher","unstructured":"Lowe DG (1999) Object recognition from local scale-invariant features. In: Proceedings of the Seventh IEEE International Conference on Computer Vision, Kerkyra, Greece, 1999, pp 1150\u20131157, vol 2.  https:\/\/doi.org\/10.1109\/ICCV.1999.790410","DOI":"10.1109\/ICCV.1999.790410"},{"key":"10414_CR101","unstructured":"Lowell U, Donahue J, Berkeley UC, Rohrbach M, Berkeley UC, Mooney R (2014) Translating videos to natural language using deep recurrent neural networks. arXiv:1412.4729v3"},{"key":"10414_CR102","unstructured":"Lu J, Batra D, Parikh D, Lee S (2019) ViLBERT: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. (NeurIPS), 1\u201311. arXiv:1908.02265"},{"key":"10414_CR103","doi-asserted-by":"publisher","unstructured":"Lu J, Xiong C, Parikh D, Socher R (2017) Knowing when to look: adaptive attention via a visual sentinel for image captioning. Proc 30th IEEE Conf Comput Vis Pattern Recogn CVPR, 2017:3242\u20133250. https:\/\/doi.org\/10.1109\/CVPR.2017.345arXiv:1612.01887","DOI":"10.1109\/CVPR.2017.345"},{"key":"10414_CR104","unstructured":"Luo H, Ji L, Shi B, Huang H, Duan N, Li T, et al. (2020) UniVL: a unified video and language pre-training model for multimodal understanding and generation. arXiv:2002.06353"},{"key":"10414_CR105","doi-asserted-by":"crossref","unstructured":"Madake J (2022) Dense video captioning using BiLSTM encoder, 1\u20136","DOI":"10.1109\/INCET54531.2022.9824569"},{"key":"10414_CR106","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Graves A, Antonoglou I, Wierstra D, Riedmiller M (2013) Playing atari with deep reinforcement learning, 1\u20139. arXiv:1312.5602"},{"issue":"7540","key":"10414_CR107","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG, Hassabis D (2015) Human-level control through deep reinforcement learning. Nature 518(7540):529\u2013533. https:\/\/doi.org\/10.1038\/nature14236","journal-title":"Nature"},{"key":"10414_CR108","doi-asserted-by":"publisher","unstructured":"Montague P (1999) Reinforcement learning: an introduction, by Sutton RS and Barto AG trends in cognitive sciences 3(9): 360. https:\/\/doi.org\/10.1016\/s1364-6613(99)01331-5","DOI":"10.1016\/s1364-6613(99)01331-5"},{"key":"10414_CR109","unstructured":"Olivastri S, Singh G, Cuzzolin F (2019) End-to-end video captioning. International conference on computer vision workshop. https:\/\/zhuanzhi.ai\/paper\/004e3568315600ed58e6a699bef3cbba"},{"key":"10414_CR110","unstructured":"Pan Y, Li Y, Luo J, Xu J, Yao T, Mei T (2020) Auto-captions on GIF: a large-scale video-sentence dataset for vision-language pre-training. arXiv:2007.02375"},{"key":"10414_CR111","doi-asserted-by":"publisher","unstructured":"Pan Y, Mei T, Yao T, Li H, Rui Y (2016) Jointly modeling embedding and translation to bridge video and language. Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn 2016:4594\u20134602. https:\/\/doi.org\/10.1109\/CVPR.2016.497arXiv:1505.01861","DOI":"10.1109\/CVPR.2016.497"},{"key":"10414_CR112","doi-asserted-by":"publisher","unstructured":"Pan Y, Yao T, Li H, Mei T (2017) Video captioning with transferred semantic attributes. Proc 30th IEEE Conf Comput Vis Pattern Recogn CVPR 2017:984\u2013992. https:\/\/doi.org\/10.1109\/CVPR.2017.111arXiv:1611.07675","DOI":"10.1109\/CVPR.2017.111"},{"key":"10414_CR113","doi-asserted-by":"publisher","unstructured":"Pan Y, Yao T, Li Y, Mei T (2020) X-linear attention networks for image captioning. Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn, 10968\u201310977. https:\/\/doi.org\/10.1109\/CVPR42600.2020.01098arXiv:2003.14080","DOI":"10.1109\/CVPR42600.2020.01098"},{"key":"10414_CR114","doi-asserted-by":"publisher","unstructured":"Park J, Song C, Han JH (2018) A study of evaluation metrics and datasets for video captioning. ICIIBMS 2017 -2nd Int Conf Intel Inform Biomed Sci 2018:172\u2013175. https:\/\/doi.org\/10.1109\/ICIIBMS.2017.8279760","DOI":"10.1109\/ICIIBMS.2017.8279760"},{"key":"10414_CR115","doi-asserted-by":"publisher","unstructured":"Pasunuru R, Bansal M (2017) Reinforced video captioning with entailment rewards. Emnlp 2017\u2014conference on empirical methods in natural language processing, proceedings (pp. 979\u2013985). https:\/\/doi.org\/10.18653\/v1\/d17-1103","DOI":"10.18653\/v1\/d17-1103"},{"key":"10414_CR116","doi-asserted-by":"publisher","unstructured":"Peng Y, Wang C, Pei Y, Li Y (2021) Video captioning with global and local text attention. Visual Computer (0123456789). https:\/\/doi.org\/10.1007\/s00371-021-02294-0","DOI":"10.1007\/s00371-021-02294-0"},{"key":"10414_CR117","doi-asserted-by":"publisher","unstructured":"Perez-Martin J, Bustos B, Perez J (2021) Attentive visual semantic specialized network for video captioning, 5767\u20135774. https:\/\/doi.org\/10.1109\/icpr48806.2021.9412898","DOI":"10.1109\/icpr48806.2021.9412898"},{"key":"10414_CR118","doi-asserted-by":"crossref","unstructured":"Perez-Martin J, Bustos B, P\u00e9rez J (2021) Improving video captioning with temporal composition of a visual-syntactic embedding. Winter Conference on Applications of Computer Vision, 3039\u20133049","DOI":"10.1109\/WACV48630.2021.00308"},{"key":"10414_CR119","unstructured":"Phan S, Henter GE, Miyao Y, Satoh S (2017) Consensus-based sequence training for video captioning. arXiv:1712.09532"},{"key":"10414_CR120","unstructured":"Pramanik S, Agrawal P, Hussain A (2019) OmniNet: a unified architecture for multi-modal multi-task learning, 1\u201316. arXiv:1907.07804"},{"key":"10414_CR121","unstructured":"Raffel C, Ellis DPW (2015) Feed-forward networks with attention can solve some long-term memory problems, 1\u20136. arXiv:1512.08756"},{"key":"10414_CR122","doi-asserted-by":"publisher","unstructured":"Rafiq M, Rafiq G, Agyeman R, Jin S-I, Choi G (2020) Scene classification for sports video summarization using transfer learning. Sensors (Switzerland) 20(6). https:\/\/doi.org\/10.3390\/s20061702","DOI":"10.3390\/s20061702"},{"key":"10414_CR123","doi-asserted-by":"publisher","first-page":"121665","DOI":"10.1109\/ACCESS.2021.3108565","volume":"9","author":"M Rafiq","year":"2021","unstructured":"Rafiq M, Rafiq G, Choi GS (2021) Video description: datasets evaluation metrics. IEEE Access 9:121665\u2013121685. https:\/\/doi.org\/10.1109\/ACCESS.2021.3108565","journal-title":"IEEE Access"},{"key":"10414_CR124","doi-asserted-by":"publisher","unstructured":"Ramanishka V, Das A, Park DH, Venugopalan S, Hendricks LA, Rohrbach M, Saenko K (2016) Multimodal video description. MM 2016 -Proceedings of the 2016 ACM Multimedia Conference, 1092\u20131096. https:\/\/doi.org\/10.1145\/2964284.2984066","DOI":"10.1145\/2964284.2984066"},{"key":"10414_CR125","unstructured":"Ranzato M, Chopra S, Auli M, Zaremba W (2016) Sequence level training with recurrent neural networks. 4th international conference on learning representations, ICLR 2016\u2014conference track proceedings (pp. 1\u201316)"},{"key":"10414_CR126","unstructured":"Redmon J, Farhadi A (2018) YOLOv3: an incremental improvement. arXiv:1804.02767"},{"key":"10414_CR127","doi-asserted-by":"publisher","first-page":"1151","DOI":"10.1109\/CVPR.2017.128","volume":"2017","author":"Z Ren","year":"2017","unstructured":"Ren Z, Wang X, Zhang N, Lv X, Li LJ (2017) Deep reinforcement learning-based image captioning with embedding reward. Proc 30th IEEE Conf Comput Vis Pattern Recogn CVPR 2017:1151\u20131159. https:\/\/doi.org\/10.1109\/CVPR.2017.128","journal-title":"Proc 30th IEEE Conf Comput Vis Pattern Recogn CVPR"},{"key":"10414_CR128","doi-asserted-by":"publisher","first-page":"1179","DOI":"10.1109\/CVPR.2017.131","volume":"2017","author":"SJ Rennie","year":"2017","unstructured":"Rennie SJ, Marcheret E, Mroueh Y, Ross J, Goel V (2017) Self-critical sequence training for image captioning. Proc 30th IEEE Conf Comput Vis Pattern Recogn CVPR  2017:1179\u20131195. https:\/\/doi.org\/10.1109\/CVPR.2017.131","journal-title":"Proc 30th IEEE Conf Comput Vis Pattern Recogn CVPR"},{"key":"10414_CR129","unstructured":"Rivera-soto RA, Ord\u00f3\u00f1ez J (2013) Sequence to sequence models for generating video captions. http:\/\/cs231n.stanford.edu\/reports\/2017\/pdfs\/31.pdf"},{"key":"10414_CR130","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.61","author":"M Rohrbach","year":"2013","unstructured":"Rohrbach M, Qiu W, Titov I, Thater S, Pinkal M, Schiele B (2013) Translating video content to natural language descriptions. Proc IEEE Int Conf Comput Vis. https:\/\/doi.org\/10.1109\/ICCV.2013.61","journal-title":"Proc IEEE Int Conf Comput Vis"},{"key":"10414_CR131","doi-asserted-by":"crossref","unstructured":"Ryu H, Kang S, Kang H, Yoo CD (2021) Semantic grouping network for video captioning. arXiv:2102.00831","DOI":"10.1609\/aaai.v35i3.16353"},{"issue":"11","key":"10414_CR132","first-page":"2673","volume":"45","author":"M Schuster","year":"1997","unstructured":"Schuster M, Paliwal KK (1997) Bidirectional recurrent. Neural Netw 45(11):2673\u20132681","journal-title":"Neural Netw"},{"key":"10414_CR133","doi-asserted-by":"crossref","unstructured":"Seo PH, Nagrani A, Arnab A, Schmid C (2022) End-to-end generative pretraining for multimodal video captioning, 17959\u201317968. arXiv:2201.08264","DOI":"10.1109\/CVPR52688.2022.01743"},{"key":"10414_CR134","doi-asserted-by":"publisher","unstructured":"Sharif N, White L, Bennamoun M, Shah SAA (2018) Learning-based composite metrics for improved caption evaluation. ACL 2018 56th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Student Research Workshop, 14\u201320. https:\/\/doi.org\/10.18653\/v1\/p18-3003","DOI":"10.18653\/v1\/p18-3003"},{"key":"10414_CR135","doi-asserted-by":"publisher","first-page":"5159","DOI":"10.1109\/CVPR.2017.548c","volume":"2017","author":"Z Shen","year":"2017","unstructured":"Shen Z, Li J, Su Z, Li M, Chen Y, Jiang YG, Xue X (2017) Weakly supervised dense video captioning. Proc 30th IEEE Conf Comput Vis Pattern Recogn, CVPR 2017 2017:5159\u20135167. https:\/\/doi.org\/10.1109\/CVPR.2017.548c","journal-title":"Proc 30th IEEE Conf Comput Vis Pattern Recogn, CVPR 2017"},{"key":"10414_CR136","doi-asserted-by":"crossref","unstructured":"Song J, Gao L, Guo Z, Liu W, Zhang D, Shen HT (2017) Hierarchical LSTM with adjusted temporal attention for video captioning, 2737\u20132743","DOI":"10.24963\/ijcai.2017\/381"},{"key":"10414_CR137","doi-asserted-by":"publisher","unstructured":"Song Y, Chen S, Jin Q (2021) Towards diverse paragraph captioning for untrimmed videos. Proceedings of the IEEE Comput Soc Conf Comput Vis Pattern Recogn, 11240\u201311249. https:\/\/doi.org\/10.1109\/CVPR46437.2021.01109arXiv:2105.14477","DOI":"10.1109\/CVPR46437.2021.01109"},{"key":"10414_CR138","unstructured":"Su J (2018) Study of Video Captioning Problem. https:\/\/www.semanticscholar.org\/paper\/Study-of-Video-Captioning-Problem-Su\/511f0041124d8d14bbcdc7f0e57f3bfe13a58e99"},{"key":"10414_CR139","doi-asserted-by":"publisher","first-page":"7463","DOI":"10.1109\/ICCV.2019.00756","volume":"2019","author":"C Sun","year":"2019","unstructured":"Sun C, Myers A, Vondrick C, Murphy K, Schmid C (2019) VideoBERT: a joint model for video and language representation learning. Proc IEEE Int Conf Comput Vis 2019:7463\u20137472. https:\/\/doi.org\/10.1109\/ICCV.2019.00756","journal-title":"Proc IEEE Int Conf Comput Vis"},{"key":"10414_CR140","doi-asserted-by":"publisher","first-page":"1300","DOI":"10.1109\/ICME.2019.00226","volume":"2019","author":"L Sun","year":"2019","unstructured":"Sun L, Li B, Yuan C, Zha Z, Hu W (2019) Multimodal semantic attention network for video captioning. Proc IEEE Int Conf Multimedia Expo 2019:1300\u20131305. https:\/\/doi.org\/10.1109\/ICME.2019.00226. arxiv.org\/abs\/1905.02963","journal-title":"Proc IEEE Int Conf Multimedia Expo"},{"key":"10414_CR141","first-page":"4278","volume":"2017","author":"C Szegedy","year":"2017","unstructured":"Szegedy C, Ioffe S, Vanhoucke V, Alemi AA (2017) Inception-v4, inception-ResNet and the impact of residual connections on learning. 31st AAAI Conf Artif Intel AAAI 2017:4278\u20134284","journal-title":"31st AAAI Conf Artif Intel AAAI"},{"key":"10414_CR142","doi-asserted-by":"publisher","unstructured":"Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, et al. (2015) Going deeper with convolutions. Proceedings of the IEEE computer society conference on computer vision and pattern recognition (07-12-June, pp. 1-9). https:\/\/doi.org\/10.1109\/CVPR.2015.7298594","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"10414_CR144","doi-asserted-by":"publisher","unstructured":"Torralba A, Murphy KP, Freeman WT,  Rubin MA (2003) Context-based vision system for place and object recognition. In: Proceedings of the Ninth IEEE International Conference on Computer Vision, ICCV'03, vol 2, pp 273. IEEE Computer Society. https:\/\/doi.org\/10.5555\/946247.946665","DOI":"10.5555\/946247.946665"},{"key":"10414_CR145","doi-asserted-by":"publisher","first-page":"4489","DOI":"10.1109\/ICCV.2015.510","volume":"2015","author":"D Tran","year":"2015","unstructured":"Tran D, Bourdev L, Fergus R, Torresani L, Paluri M (2015) Learning spatiotemporal features with 3D convolutional networks. Proc IEEE Int Conf Comput Vis 2015:4489\u20134497. https:\/\/doi.org\/10.1109\/ICCV.2015.510","journal-title":"Proc IEEE Int Conf Comput Vis"},{"key":"10414_CR146","unstructured":"Uszkoreit J, Kaiser L (2019) Universal transformers, 1-23. arxiv.org\/abs\/arXiv:1807.03819v3"},{"key":"10414_CR147","unstructured":"Vaswani A, Brain G, Shazeer N, Parmar N, Uszkoreit J, Jones L, et al. (2017) Attention is all you need. Adv Neural Inform Process Syst (Nips), 5998\u20136008. http:\/\/papers.nips.cc\/paper\/7181-attention-is-all-you-need.pdf"},{"key":"10414_CR143","doi-asserted-by":"crossref","unstructured":"Vedantam R, Lawrence Zitnick C, Parikh D (2015) Cider: consensus-based image description evaluation. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4566\u20134575","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"10414_CR148","doi-asserted-by":"publisher","first-page":"4534","DOI":"10.1109\/ICCV.2015.515","volume":"2015","author":"S Venugopalan","year":"2015","unstructured":"Venugopalan S, Rohrbach M, Donahue J, Mooney R, Darrell T, Saenko K (2015) Sequence to sequence -video to text. Proceedings IEEE Int Conf Comput Vis 2015:4534\u20134542. https:\/\/doi.org\/10.1109\/ICCV.2015.515","journal-title":"Proceedings IEEE Int Conf Comput Vis"},{"key":"10414_CR149","doi-asserted-by":"publisher","unstructured":"Vo DM, Chen H, Sugimoto A, Nakayama H (2022) NOC-REK: Novel object captioning with retrieved vocabulary from external knowledge. In: 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 2022, pp 17979\u201317987. https:\/\/doi.org\/10.1109\/CVPR52688.2022.01747","DOI":"10.1109\/CVPR52688.2022.01747"},{"key":"10414_CR150","doi-asserted-by":"publisher","unstructured":"Wallach B (2017) Developing: a world made for money (pp. 241\u2013294). https:\/\/doi.org\/10.2307\/j.ctt1d98bxx.10","DOI":"10.2307\/j.ctt1d98bxx.10"},{"key":"10414_CR153","doi-asserted-by":"publisher","unstructured":"Wang D, Song D (2017) Video Captioning with Semantic Information from the Knowledge Base. Proceedings -2017 IEEE International Conference on Big Knowledge, ICBK 2017 , 224\u2013229. https:\/\/doi.org\/10.1109\/ICBK.2017.26","DOI":"10.1109\/ICBK.2017.26"},{"key":"10414_CR152","doi-asserted-by":"publisher","unstructured":"Wang B, Ma L, Zhang W,  Liu W (2018a) Reconstruction network for video captioning. In: 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp 7622\u20137631. https:\/\/doi.org\/10.1109\/CVPR.2018.00795","DOI":"10.1109\/CVPR.2018.00795"},{"key":"10414_CR156","doi-asserted-by":"publisher","unstructured":"Wang X, Chen W, Wu J, Wang YF, Wang WY (2018b) Video captioning via hierarchical reinforcement learning. Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn, 4213\u20134222. https:\/\/doi.org\/10.1109\/CVPR.2018.00443arXiv:1711.11135","DOI":"10.1109\/CVPR.2018.00443"},{"key":"10414_CR157","doi-asserted-by":"crossref","unstructured":"Wang X, Wang, Y-f, Wang WY (2018c) Watch , listen , and describe: globally and locally aligned cross-modal attentions for video captioning, 795\u2013801","DOI":"10.18653\/v1\/N18-2125"},{"key":"10414_CR151","doi-asserted-by":"publisher","first-page":"2641","DOI":"10.1109\/ICCV.2019.00273","volume":"2019","author":"B Wang","year":"2019","unstructured":"Wang B, Ma L, Zhang W, Jiang W, Wang J, Liu W (2019a) Controllable video captioning with pos sequence guidance based on gated fusion network. Proc IEEE Int Conf Comput Vis 2019:2641\u20132650. https:\/\/doi.org\/10.1109\/ICCV.2019.00273. arXiv:1908.10072","journal-title":"Proc IEEE Int Conf Comput Vis"},{"key":"10414_CR158","doi-asserted-by":"publisher","unstructured":"Wang X, Wu J, Chen J, Li L, Wang Y-F, Wang WY (2019b) VATEX: a large-scale, high-quality multilingual dataset for video-and-language research. In: 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), pp 4580\u20134590. https:\/\/doi.org\/10.1109\/ICCV.2019.00468","DOI":"10.1109\/ICCV.2019.00468"},{"key":"10414_CR154","doi-asserted-by":"publisher","unstructured":"Wang H, Zhang Y, Yu X (2020) An overview of image caption generation methods. Computational Intelligence and Neuroscience 2020. https:\/\/doi.org\/10.1155\/2020\/3062706","DOI":"10.1155\/2020\/3062706"},{"key":"10414_CR155","doi-asserted-by":"publisher","unstructured":"Wang T, Zhang R, Lu Z, Zheng F, Cheng R, Luo P (2021) Endto-End Dense Video Captioning with Parallel Decoding. Proceedings of the IEEE International Conference on Computer Vision, 6827\u20136837. https:\/\/doi.org\/10.1109\/ICCV48922.2021.00677arXiv:2108.07781","DOI":"10.1109\/ICCV48922.2021.00677"},{"issue":"2","key":"10414_CR159","doi-asserted-by":"publisher","first-page":"270","DOI":"10.1162\/neco.1989.1.2.270","volume":"1","author":"RJ Williams","year":"1989","unstructured":"Williams RJ, Zipser D (1989) A learning algorithm for continually running fully recurrent neural networks. Neural Comput 1(2):270\u2013280. https:\/\/doi.org\/10.1162\/neco.1989.1.2.270","journal-title":"Neural Comput"},{"key":"10414_CR160","doi-asserted-by":"crossref","unstructured":"Wu D, Zhao H, Bao X, Wildes RP (2022) Sports video analysis on large-scale data (1). arXiv:2208.04897","DOI":"10.1007\/978-3-031-19836-6_2"},{"key":"10414_CR161","doi-asserted-by":"publisher","unstructured":"Wu Z, Yao T, Fu Y, Jiang, Y-G (2017) Deep learning for video classification and captioning. Front Multimedia Res, 3\u201329. https:\/\/doi.org\/10.1145\/3122865.3122867arXiv:1609.06782","DOI":"10.1145\/3122865.3122867"},{"key":"10414_CR162","unstructured":"Xiao H, Shi J (2019a) Diverse video captioning through latent variable expansion with conditional GAN. https:\/\/zhuanzhi.ai\/paper\/943af2926865564d7a84286c23fa2c63 arXiv:1910.12019"},{"key":"10414_CR163","unstructured":"Xiao H, Shi J (2019b) Huanhou Xiao, Jinglun Shi South China University of Technology, Guangzhou China, 619\u2013623"},{"key":"10414_CR164","doi-asserted-by":"publisher","first-page":"318","DOI":"10.1007\/978-3-030-01267-0_19","volume":"11219","author":"S Xie","year":"2018","unstructured":"Xie S, Sun C, Huang J, Tu Z, Murphy K (2018) Rethinking spatiotem-poral feature learning: speed-accuracy trade-offs in video classification. Lecture Notes Comput Sci (including subseries lecture notes in artificial intelligence and lecture notes in bioinformatics) 11219:318\u2013335. https:\/\/doi.org\/10.1007\/978-3-030-01267-0_19","journal-title":"Lecture Notes Comput Sci (including subseries lecture notes in artificial intelligence and lecture notes in bioinformatics)"},{"key":"10414_CR165","doi-asserted-by":"publisher","first-page":"396","DOI":"10.1109\/WACV.2019.00048","volume":"2019","author":"H Xu","year":"2019","unstructured":"Xu H, Li B, Ramanishka V, Sigal L, Saenko K (2019) Joint event detection and description in continuous video streams. Proc 2019 IEEE Winter Conf App Comput Vis, WACV 2019:396\u2013405. https:\/\/doi.org\/10.1109\/WACV.2019.00048. arXiv:1802.10250","journal-title":"Proc 2019 IEEE Winter Conf App Comput Vis, WACV"},{"key":"10414_CR166","doi-asserted-by":"publisher","first-page":"5288","DOI":"10.1109\/CVPR.2016.571","volume":"2016","author":"J Xu","year":"2016","unstructured":"Xu J, Mei T, Yao T, Rui Y (2016) MSR-VTT: a large video description dataset for bridging video and language. Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn 2016:5288\u20135296. https:\/\/doi.org\/10.1109\/CVPR.2016.571","journal-title":"Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn"},{"key":"10414_CR167","doi-asserted-by":"publisher","DOI":"10.3390\/app10124312","author":"J Xu","year":"2020","unstructured":"Xu J, Wei H, Li L, Fu Q, Guo J (2020) Video description model based on temporal-spatial and channel multi-attention mechanisms. Appl Sci (Switzerland). https:\/\/doi.org\/10.3390\/app10124312","journal-title":"Appl Sci (Switzerland)"},{"key":"10414_CR168","doi-asserted-by":"publisher","unstructured":"Xu J, Yao T, Zhang Y, Mei T (2017) Learning multimodal attention LSTM networks for video captioning. MM 2017 -Proceedings of the 2017 ACM Multimedia Conference, 537\u2013545. https:\/\/doi.org\/10.1145\/3123266.3123448","DOI":"10.1145\/3123266.3123448"},{"key":"10414_CR169","unstructured":"Xu K, Ba JL, Kiros R, Cho K, Courville A, Salakhutdinov R, et al. (2015) Show, attend and tell: neural image caption gener-ation with visual attention. 32nd International Conference on Machine Learning, ICML 2015 3:2048\u20132057. arXiv:1502.03044"},{"key":"10414_CR170","doi-asserted-by":"publisher","first-page":"1772","DOI":"10.1109\/TMM.2020.3002669","volume":"23","author":"W Xu","year":"2021","unstructured":"Xu W, Yu J, Miao Z, Wan L, Tian Y, Ji Q (2021) Deep reinforcement polishing network for video captioning. IEEE Trans Multimedia 23:1772\u20131784. https:\/\/doi.org\/10.1109\/TMM.2020.3002669","journal-title":"IEEE Trans Multimedia"},{"issue":"1","key":"10414_CR171","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1109\/TMM.2019.2924576","volume":"22","author":"C Yan","year":"2020","unstructured":"Yan C, Tu Y, Wang X, Zhang Y, Hao X, Zhang Y, Dai Q (2020) STAT: spatial-temporal attention mechanism for video captioning. IEEE Trans Multimedia 22(1):229\u2013241. https:\/\/doi.org\/10.1109\/TMM.2019.2924576","journal-title":"IEEE Trans Multimedia"},{"key":"10414_CR172","unstructured":"Yan L, Zhu M, Yu C (2010) Crowd video captioning. arXiv:1911.05449v1"},{"key":"10414_CR173","doi-asserted-by":"publisher","unstructured":"Yan Y, Zhuang N, Bingbing Ni, Zhang J, Xu M, Zhang Q, et al (2019) Fine-grained video captioning via graph-based multi-granularity interaction learning. IEEE Trans Pattern Analys Mach Intel. https:\/\/doi.org\/10.1109\/TPAMI.2019.2946823","DOI":"10.1109\/TPAMI.2019.2946823"},{"key":"10414_CR174","doi-asserted-by":"publisher","unstructured":"Yang B, Liu F, Zhang C, Zou Y (2019) Non-autoregressive coarse-to-fine video captioning.  In: AAAI Conference on Artificial Intelligence. https:\/\/doi.org\/10.1609\/aaai.v35i4.16421","DOI":"10.1609\/aaai.v35i4.16421"},{"key":"10414_CR175","unstructured":"Yang Z, Yuan Y, Wu Y, Salakhutdinov R, Cohen WW (2016) Review networks for caption generation. Adv Neural Inform Process Syst (Nips), 2369\u20132377. arXiv:1605.07912"},{"key":"10414_CR176","unstructured":"Yin W, Kann K, Yu M, Sch\u00fctze H (2017) Comparative study of CNN and RNN for natural language processing. arXiv:1702.01923"},{"key":"10414_CR177","doi-asserted-by":"publisher","first-page":"4651","DOI":"10.1109\/CVPR.2016.503","volume":"2016","author":"Q You","year":"2016","unstructured":"You Q, Jin H, Wang Z, Fang C, Luo J (2016) Image captioning with semantic attention. Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn 2016:4651\u20134659. https:\/\/doi.org\/10.1109\/CVPR.2016.503. arXiv:1603.03925","journal-title":"Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn"},{"key":"10414_CR178","doi-asserted-by":"crossref","unstructured":"Young P, Lai A, Hodosh M, Hockenmaier J (2014) From image descriptions to visual denotations? New similarity metrics for semantic inference over event descriptions 2:67\u201378","DOI":"10.1162\/tacl_a_00166"},{"key":"10414_CR179","doi-asserted-by":"publisher","first-page":"6119","DOI":"10.1109\/CVPR.2017.648","volume":"2017","author":"Y Yu","year":"2017","unstructured":"Yu Y, Choi J, Kim Y, Yoo K, Lee SH, Kim G (2017) Supervising neural attention models for video captioning by human gaze data. Proc 30th IEEE Conf Comput Vis Pattern Recogn 2017:6119\u20136127. https:\/\/doi.org\/10.1109\/CVPR.2017.648. arXiv:1707.06029","journal-title":"Proc 30th IEEE Conf Comput Vis Pattern Recogn"},{"key":"10414_CR180","doi-asserted-by":"crossref","unstructured":"Yuan Z, Yan X, Liao Y, Guo Y, Li G, Li Z, Cui S (2022) X-Trans2Cap: cross-modal knowledge transfer using transformer for 3D dense captioning, 3\u20134. arXiv:2203.00843","DOI":"10.1109\/CVPR52688.2022.00837"},{"key":"10414_CR181","doi-asserted-by":"publisher","first-page":"6713","DOI":"10.1109\/CVPR.2019.00688","volume":"2019","author":"R Zellers","year":"2019","unstructured":"Zellers R, Bisk Y, Farhadi A, Choi Y, (2019) From recognition to cognition: visual commonsense reasoning. Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn 2019:6713\u20136724. https:\/\/doi.org\/10.1109\/CVPR.2019.00688","journal-title":"Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn"},{"key":"10414_CR182","doi-asserted-by":"crossref","unstructured":"Zhang J, Peng Y (2019) Object-aware aggregation with bidirectional temporal graph for video captioning. https:\/\/zhuanzhi.ai\/paper\/237b5837832fb600d4269cacdb0286e3 arXiv:1906.04375","DOI":"10.1109\/CVPR.2019.00852"},{"key":"10414_CR183","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1016\/j.neucom.2018.09.038","volume":"323","author":"Q Zhang","year":"2019","unstructured":"Zhang Q, Zhang M, Chen T, Sun Z, Ma Y, Yu B (2019) Recent advances in convolutional neural network acceleration. Neurocomputing 323:37\u201351. https:\/\/doi.org\/10.1016\/j.neucom.2018.09.038. arXiv:1807.08596","journal-title":"Neurocomputing"},{"key":"10414_CR184","doi-asserted-by":"publisher","unstructured":"Zhang W, Wang B, Ma L, Liu W (2019) Reconstruct and represent video contents for captioning via reinforcement learning. IEEE Trans Pattern Analys Mach Intel, 1\u20131. https:\/\/doi.org\/10.1109\/tpami.2019.2920899arxiv.org\/abs\/1906.01452","DOI":"10.1109\/tpami.2019.2920899"},{"key":"10414_CR185","doi-asserted-by":"publisher","first-page":"6250","DOI":"10.1109\/CVPR.2017.662","volume":"2017","author":"X Zhang","year":"2017","unstructured":"Zhang X, Gao K, Zhang Y, Zhang D, Li J, Tian Q (2017) Task-driven dynamic fusion: reducing ambiguity in video description. Proc 30th IEEE Conf Comput Vis Pattern Recogn CVPR 2017:6250\u20136258. https:\/\/doi.org\/10.1109\/CVPR.2017.662","journal-title":"Proc 30th IEEE Conf Comput Vis Pattern Recogn CVPR"},{"key":"10414_CR186","doi-asserted-by":"publisher","first-page":"15460","DOI":"10.1109\/CVPR46437.2021.01521","volume":"1","author":"X Zhang","year":"2021","unstructured":"Zhang X, Sun X, Luo Y, Ji J, Zhou Y, Wu Y, Ji R (2021) RSTnet: captioning with adaptive attention on visual and non-visual words. Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn 1:15460\u201315469. https:\/\/doi.org\/10.1109\/CVPR46437.2021.01521","journal-title":"Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn"},{"issue":"1","key":"10414_CR187","doi-asserted-by":"publisher","first-page":"51","DOI":"10.1007\/s10590-010-9073-6","volume":"24","author":"Y Zhang","year":"2010","unstructured":"Zhang Y, Vogel S (2010) Significance tests of automatic machine translation evaluation metrics. Machine Transl 24(1):51\u201365. https:\/\/doi.org\/10.1007\/s10590-010-9073-6","journal-title":"Machine Transl"},{"key":"10414_CR188","doi-asserted-by":"publisher","unstructured":"Zhang Z, Qi Z, Yuan C, Shan Y, Li B, Deng Y, Hu W (2021) Open-book video captioning with retrieve-copy-generate network. Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn, 9832\u20139841. https:\/\/doi.org\/10.1109\/CVPR46437.2021.00971arXiv:2103.05284","DOI":"10.1109\/CVPR46437.2021.00971"},{"key":"10414_CR189","doi-asserted-by":"crossref","unstructured":"Zhang Z, Shi Y, Yuan C, Li B, Wang P, Hu W, Zha Z (2020) Object relational graph with teacher-recommended learning for video captioning. arXiv:2002.11566","DOI":"10.1109\/CVPR42600.2020.01329"},{"key":"10414_CR190","doi-asserted-by":"publisher","unstructured":"Zhao B, Li X, Lu X (2018) Video captioning with tube features. IICAI Int Joint Conf Artif Intel 2018:1177\u20131183. https:\/\/doi.org\/10.24963\/ijcai.2018\/164","DOI":"10.24963\/ijcai.2018\/164"},{"issue":"2002","key":"10414_CR191","doi-asserted-by":"publisher","first-page":"1","DOI":"10.7717\/PEERJ-CS.916","volume":"8","author":"H Zhao","year":"2022","unstructured":"Zhao H, Chen Z, Guo L, Han Z (2022) Video captioning based on vision transformer and reinforcement learning. Peer J Comput Sci 8(2002):1\u201316. https:\/\/doi.org\/10.7717\/PEERJ-CS.916","journal-title":"Peer J Comput Sci"},{"key":"10414_CR192","doi-asserted-by":"publisher","unstructured":"Zheng Q, Wang C, Tao D (2020) Syntax-Aware Action Targeting for Video Captioning. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 13093\u201313102. https:\/\/doi.org\/10.1109\/CVPR42600.2020.01311","DOI":"10.1109\/CVPR42600.2020.01311"},{"key":"10414_CR193","unstructured":"Zhou L, Corso JJ (2016) Towards automatic learning of procedures from web instructional videos. arXiv:1703.09788v3"},{"key":"10414_CR194","doi-asserted-by":"publisher","first-page":"6571","DOI":"10.1109\/CVPR.2019.00674","volume":"2019","author":"L Zhou","year":"2019","unstructured":"Zhou L, Kalantidis Y, Chen X, Corso JJ, Rohrbach M (2019) Grounded video description. Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn 2019:6571\u20136580. https:\/\/doi.org\/10.1109\/CVPR.2019.00674. arXiv:1812.06587","journal-title":"Proc IEEE Comput Soc Conf Comput Vis Pattern Recogn"},{"key":"10414_CR195","doi-asserted-by":"publisher","unstructured":"Zhou L, Zhou Y, Corso JJ, Socher R, Xiong C (2018) End-to-End Dense Video Captioning with Masked Transformer. Proceedings of the IEEE computer society conference on computer vision and pattern recognition (pp. 8739\u20138748). https:\/\/doi.org\/10.1109\/CVPR.2018.00911","DOI":"10.1109\/CVPR.2018.00911"},{"key":"10414_CR196","unstructured":"Zhu X, Guo L, Yao P, Lu S, Liu W, Liu J (2019) Vatex video captioning challenge 2020: multi-view features and hybrid reward strategies for video captioning. arXiv:1910.11102"},{"key":"10414_CR197","doi-asserted-by":"publisher","first-page":"713","DOI":"10.1007\/978-3-030-01216-8-43","volume":"11206","author":"M Zolfaghari","year":"2018","unstructured":"Zolfaghari M, Singh K, Brox T (2018) ECO: efficient convolutional network for online video understanding. Lecture Notes Comput Sci (including subseries lecture notes in artificial intelligence and lecture notes in bioinformatics) 11206:713\u2013730. https:\/\/doi.org\/10.1007\/978-3-030-01216-8-43","journal-title":"Lecture Notes Comput Sci (including subseries lecture notes in artificial intelligence and lecture notes in bioinformatics)"}],"container-title":["Artificial Intelligence Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-023-10414-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10462-023-10414-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-023-10414-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,22]],"date-time":"2023-09-22T23:53:19Z","timestamp":1695426799000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10462-023-10414-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,4,11]]},"references-count":196,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2023,11]]}},"alternative-id":["10414"],"URL":"https:\/\/doi.org\/10.1007\/s10462-023-10414-6","relation":{},"ISSN":["0269-2821","1573-7462"],"issn-type":[{"value":"0269-2821","type":"print"},{"value":"1573-7462","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,4,11]]},"assertion":[{"value":"11 April 2023","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}