{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:10:55Z","timestamp":1750219855440,"version":"3.41.0"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,5,30]],"date-time":"2023-05-30T00:00:00Z","timestamp":1685404800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62172256, 62202278, and 62202272"],"award-info":[{"award-number":["62172256, 62202278, and 62202272"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100007129","name":"Natural Science Foundation of Shandong Province","doi-asserted-by":"crossref","award":["ZR2019ZD06"],"award-info":[{"award-number":["ZR2019ZD06"]}],"id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Major Program of the National Natural Science Foundation of China","award":["61991411"],"award-info":[{"award-number":["61991411"]}]},{"name":"Quan Cheng Laboratory, Jinan, China"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,11,30]]},"abstract":"<jats:p>Video captioning aims to automatically describe a video clip with informative sentences. At present, deep learning-based models have become the mainstream for this task and achieved competitive results on public datasets. Usually, these methods leverage different types of features to generate sentences, e.g., semantic information, 2D or 3D features. However, some methods only treat semantic information as a complement of visual representations and cannot fully exploit it; some of them ignore the relationship between different types of features. In addition, most of them select multiple frames of a video with an equally spaced sampling scheme, resulting in much redundant information. To address these issues, we present a novel video-captioning framework, Semantic Enhanced video captioning with Multi-feature Fusion, SEMF for short. It optimizes the use of different types of features from three aspects. First, a semantic encoder is designed to enhance meaningful semantic features through a semantic dictionary to boost performance. Second, a discrete selection module pays attention to important features and obtains different contexts at different steps to reduce feature redundancy. Finally, a multi-feature fusion module uses a novel relation-aware attention mechanism to separate the common and complementary components of different features to provide more effective visual features for the next step. Moreover, the entire framework can be trained in an end-to-end manner. Extensive experiments are conducted on Microsoft Research Video Description Corpus (MSVD) and MSR-Video to Text (MSR-VTT) datasets. The results demonstrate that SEMF is able to achieve state-of-the-art results.<\/jats:p>","DOI":"10.1145\/3588572","type":"journal-article","created":{"date-parts":[[2023,3,20]],"date-time":"2023-03-20T12:02:31Z","timestamp":1679313751000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Semantic Enhanced Video Captioning with Multi-feature Fusion"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7389-5883","authenticated-orcid":false,"given":"Tian-Zi","family":"Niu","sequence":"first","affiliation":[{"name":"School of Software, Shandong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2500-9488","authenticated-orcid":false,"given":"Shan-Shan","family":"Dong","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3481-4892","authenticated-orcid":false,"given":"Zhen-Duo","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6901-5476","authenticated-orcid":false,"given":"Xin","family":"Luo","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3367-0951","authenticated-orcid":false,"given":"Shanqing","family":"Guo","sequence":"additional","affiliation":[{"name":"School of Cyber Science and Technology, Shandong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9738-4949","authenticated-orcid":false,"given":"Zi","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Information Technology and Electrical Engineering, The University of Queensland, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9972-7370","authenticated-orcid":false,"given":"Xin-Shun","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,5,30]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"12487","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Aafaq Nayyer","year":"2019","unstructured":"Nayyer Aafaq, Naveed Akhtar, Wei Liu, Syed Zulqarnain Gilani, and Ajmal Mian. 2019. Spatio-temporal dynamics and semantic attribute enriched visual encoding for video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.12487\u201312496."},{"key":"e_1_3_2_3_2","first-page":"65","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Banerjee Satanjeev","year":"2005","unstructured":"Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. 65\u201372."},{"key":"e_1_3_2_4_2","first-page":"4724","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Carreira Jo\u00e3o","year":"2017","unstructured":"Jo\u00e3o Carreira and Andrew Zisserman. 2017. Quo Vadis, action recognition? A new model and the kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.4724\u20134733."},{"key":"e_1_3_2_5_2","first-page":"190","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Chen David L.","year":"2011","unstructured":"David L. Chen and William B. Dolan. 2011. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. 190\u2013200."},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"475767","DOI":"10.3389\/frobt.2020.475767","article-title":"A semantics-assisted video captioning model trained with scheduled sampling","volume":"7","author":"Chen Haoran","year":"2020","unstructured":"Haoran Chen, Ke Lin, Alexander Maye, Jianmin Li, and Xiaolin Hu. 2020. A semantics-assisted video captioning model trained with scheduled sampling. Front. Robot. AI 7 (2020), 475767.","journal-title":"Front. Robot. AI"},{"key":"e_1_3_2_7_2","first-page":"333","volume-title":"Proceedings of the European Conference on Computer Vision.","author":"Chen Shaoxiang","year":"2020","unstructured":"Shaoxiang Chen, Wenhao Jiang, Wei Liu, and Yu-Gang Jiang. 2020. Learning modality interaction for temporal sentence localization and event captioning in videos. In Proceedings of the European Conference on Computer Vision.333\u2013351."},{"key":"e_1_3_2_8_2","first-page":"8191","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence.","author":"Chen Shaoxiang","year":"2019","unstructured":"Shaoxiang Chen and Yu-Gang Jiang. 2019. Motion guided spatial attention for video captioning. In Proceedings of the AAAI Conference on Artificial Intelligence.8191\u20138198."},{"key":"e_1_3_2_9_2","first-page":"367","volume-title":"Proceedings of the European Conference on Computer Vision.","author":"Chen Yangyu","year":"2018","unstructured":"Yangyu Chen, Shuhui Wang, Weigang Zhang, and Qingming Huang. 2018. Less is more: Picking informative frames for video captioning. In Proceedings of the European Conference on Computer Vision.367\u2013384."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1179"},{"key":"e_1_3_2_11_2","first-page":"2634","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Das Pradipto","year":"2013","unstructured":"Pradipto Das, Chenliang Xu, Richard F. Doell, and Jason J. Corso. 2013. A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.2634\u20132641."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3063423"},{"issue":"2","key":"e_1_3_2_13_2","article-title":"Semantic embedding guided attention with explicit visual feature fusion for video captioning","volume":"19","author":"Dong Shanshan","year":"2023","unstructured":"Shanshan Dong, Tianzi Niu, Xin Luo, Wu Liu, and Xinshun Xu. 2023. Semantic embedding guided attention with explicit visual feature fusion for video captioning. ACM Trans. Multim. Comput. Commun. Appl. 19, 2 (2023).","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298754"},{"key":"e_1_3_2_15_2","first-page":"1141","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Gan Zhe","year":"2017","unstructured":"Zhe Gan, Chuang Gan, Xiaodong He, Yunchen Pu, Kenneth Tran, Jianfeng Gao, Lawrence Carin, and Li Deng. 2017. Semantic compositional networks for visual captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.1141\u20131150."},{"key":"e_1_3_2_16_2","first-page":"6639","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Gao Peng","year":"2019","unstructured":"Peng Gao, Zhengkai Jiang, Haoxuan You, Pan Lu, Steven C. H. Hoi, Xiaogang Wang, and Hongsheng Li. 2019. Dynamic fusion with intra- and inter-modality attention flow for visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.6639\u20136648."},{"key":"e_1_3_2_17_2","first-page":"2712","volume-title":"Proceedings of the IEEE International Conference on Computer Vision.","author":"Guadarrama Sergio","year":"2013","unstructured":"Sergio Guadarrama, Niveda Krishnamoorthy, Girish Malkarnenkar, Subhashini Venugopalan, Raymond J. Mooney, Trevor Darrell, and Kate Saenko. 2013. YouTube2Text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition. In Proceedings of the IEEE International Conference on Computer Vision.2712\u20132719."},{"key":"e_1_3_2_18_2","first-page":"6546","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Hara Kensho","year":"2018","unstructured":"Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. 2018. Can spatiotemporal 3D CNNs retrace the history of 2D CNNs and ImageNet? In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.6546\u20136555."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_20_2","first-page":"4203","volume-title":"Proceedings of the IEEE International Conference on Computer Vision.","author":"Hori Chiori","year":"2017","unstructured":"Chiori Hori, Takaaki Hori, Teng-Yok Lee, Ziming Zhang, Bret Harsham, John R. Hershey, Tim K. Marks, and Kazuhiro Sumi. 2017. Attention-based multimodal fusion for video description. In Proceedings of the IEEE International Conference on Computer Vision.4203\u20134212."},{"key":"e_1_3_2_21_2","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Jang Eric","year":"2017","unstructured":"Eric Jang, Shixiang Gu, and Ben Poole. 2017. Categorical reparameterization with Gumbel-Softmax. In Proceedings of the International Conference on Learning Representations."},{"issue":"4","key":"e_1_3_2_22_2","first-page":"125:1\u2013125:20","article-title":"Bi-directional co-attention network for image captioning","volume":"17","author":"Jiang Weitao","year":"2021","unstructured":"Weitao Jiang, Weixuan Wang, and Haifeng Hu. 2021. Bi-directional co-attention network for image captioning. ACM Trans. Multim. Comput. Commun. Appl. 17, 4 (2021), 125:1\u2013125:20.","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1020346032608"},{"key":"e_1_3_2_24_2","first-page":"74","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. 74\u201381."},{"key":"e_1_3_2_25_2","first-page":"2047","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence.","author":"Lin Ke","year":"2021","unstructured":"Ke Lin, Zhuoxin Gan, and Liwei Wang. 2021. Augmented partial mutual learning with frame masking for video captioning. In Proceedings of the AAAI Conference on Artificial Intelligence.2047\u20132055."},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2940007"},{"issue":"4","key":"e_1_3_2_27_2","first-page":"128:1\u2013128:22","article-title":"Adaptive attention-based high-level semantic introduction for image caption","volume":"16","author":"Liu Xiaoxiao","year":"2021","unstructured":"Xiaoxiao Liu and Qingyang Xu. 2021. Adaptive attention-based high-level semantic introduction for image caption. ACM Trans. Multim. Comput. Commun. Appl. 16, 4 (2021), 128:1\u2013128:22.","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503927"},{"key":"e_1_3_2_29_2","first-page":"10867","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Pan Boxiao","year":"2020","unstructured":"Boxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee, Adrien Gaidon, Ehsan Adeli, and Juan Carlos Niebles. 2020. Spatio-temporal graph for video captioning with knowledge distillation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.10867\u201310876."},{"key":"e_1_3_2_30_2","first-page":"984","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Pan Yingwei","year":"2017","unstructured":"Yingwei Pan, Ting Yao, Houqiang Li, and Tao Mei. 2017. Video captioning with transferred semantic attributes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.984\u2013992."},{"key":"e_1_3_2_31_2","first-page":"311","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. 311\u2013318."},{"key":"e_1_3_2_32_2","first-page":"8347","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Pei Wenjie","year":"2019","unstructured":"Wenjie Pei, Jiyuan Zhang, Xiangrong Wang, Lei Ke, Xiaoyong Shen, and Yu-Wing Tai. 2019. Memory-attended recurrent network for video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.8347\u20138356."},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_2_35_2","first-page":"2514","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence.","author":"Ryu Hobin","year":"2021","unstructured":"Hobin Ryu, Sunghun Kang, Haeyong Kang, Haeyong Kang, and Chang D. Yoo. 2021. Semantic grouping network for video captioning. In Proceedings of the AAAI Conference on Artificial Intelligence.2514\u20132522."},{"key":"e_1_3_2_36_2","first-page":"1161","volume-title":"Proceedings of the International Conference on Information Knowledge Management.","author":"Song Weiping","year":"2019","unstructured":"Weiping Song, Chence Shi, Zhiping Xiao, Zhijian Duan, Yewen Xu, Ming Zhang, and Jian Tang. 2019. AutoInt: Automatic feature interaction learning via self-attentive neural networks. In Proceedings of the International Conference on Information Knowledge Management.1161\u20131170."},{"key":"e_1_3_2_37_2","first-page":"11245","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Song Yuqing","year":"2021","unstructured":"Yuqing Song, Shizhe Chen, and Qin Jin. 2021. Towards diverse paragraph captioning for untrimmed videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.11245\u201311254."},{"key":"e_1_3_2_38_2","first-page":"4278","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence.","author":"Szegedy Christian","year":"2017","unstructured":"Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A. Alemi. 2017. Inception-v4, inception-ResNet and the impact of residual connections on learning. In Proceedings of the AAAI Conference on Artificial Intelligence.4278\u20134284."},{"key":"e_1_3_2_39_2","first-page":"745","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence.","author":"Tan Ganchao","year":"2020","unstructured":"Ganchao Tan, Daqing Liu, Meng Wang, and Zheng-Jun Zha. 2020. Learning to discretely compose reasoning module networks for video captioning. In Proceedings of the International Joint Conference on Artificial Intelligence.745\u2013752."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3303083"},{"key":"e_1_3_2_41_2","first-page":"2415","volume-title":"Proceedings of the Conference of the North American Chapter of the Association of Computational Linguistics: Human Language Technologies.","author":"Tang Zineng","year":"2021","unstructured":"Zineng Tang, Jie Lei, and Mohit Bansal. 2021. DeCEMBERT: Learning from noisy instructional videos via dense captions and entropy minimization. In Proceedings of the Conference of the North American Chapter of the Association of Computational Linguistics: Human Language Technologies.2415\u20132426."},{"key":"e_1_3_2_42_2","first-page":"1218","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Thomason Jesse","year":"2014","unstructured":"Jesse Thomason, Subhashini Venugopalan, Sergio Guadarrama, Kate Saenko, and Raymond J. Mooney. 2014. Integrating language and vision to generate natural language descriptions of videos in the wild. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. 1218\u20131227."},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_2_44_2","first-page":"2442","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications in Computer Vision.","author":"Vaidya Jayesh","year":"2022","unstructured":"Jayesh Vaidya, Arulkumar Subramaniam, and Anurag Mittal. 2022. Co-segmentation aided two-stream architecture for video captioning. In Proceedings of the IEEE\/CVF Winter Conference on Applications in Computer Vision.2442\u20132452."},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_3_2_46_2","first-page":"4534","volume-title":"Proceedings of the IEEE International Conference on Computer Vision.","author":"Venugopalan Subhashini","year":"2015","unstructured":"Subhashini Venugopalan, Marcus RohrbachJeffrey Donahue, Raymond J. Mooney, Trevor Darrell, and Kate Saenko. 2015. Sequence to sequence\u2014Video to text. In Proceedings of the IEEE International Conference on Computer Vision.4534\u20134542."},{"key":"e_1_3_2_47_2","first-page":"2641","volume-title":"Proceedings of the IEEE International Conference on Computer Vision.","author":"Wang Bairui","year":"2019","unstructured":"Bairui Wang, Lin Ma, Wei Zhang, Wenhao Jiang, Jingwen Wang, and Wei Liu. 2019. Controllable video captioning with POS sequence guidance based on gated fusion network. In Proceedings of the IEEE International Conference on Computer Vision.2641\u20132650."},{"key":"e_1_3_2_48_2","first-page":"7622","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Wang Bairui","year":"2018","unstructured":"Bairui Wang, Lin Ma, Wei Zhang, and Wei Liu. 2018. Reconstruction network for video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.7622\u20137631."},{"key":"e_1_3_2_49_2","first-page":"7512","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Wang Junbo","year":"2018","unstructured":"Junbo Wang, Wei Wang, Yan Huang, Liang Wang, and Tieniu Tan. 2018. M3: Multimodal memory modelling for video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.7512\u20137520."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2923608"},{"issue":"2","key":"e_1_3_2_51_2","first-page":"57:1\u201357:18","article-title":"Learning transferable perturbations for image captioning","volume":"18","author":"Wu Hanjie","year":"2022","unstructured":"Hanjie Wu, Yongtuo Liu, Hongmin Cai, and Shengfeng He. 2022. Learning transferable perturbations for image captioning. ACM Trans. Multim. Comput. Commun. Appl. 18, 2 (2022), 57:1\u201357:18.","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"e_1_3_2_52_2","first-page":"203","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Wu Qi","year":"2016","unstructured":"Qi Wu, Chunhua Shen, Lingqiao Liu, Anthony R. Dick, and Anton van den Hengel. 2016. What value do explicit high level concepts have in vision to language problems? In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.203\u2013212."},{"key":"e_1_3_2_53_2","first-page":"5288","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Xu Jun","year":"2016","unstructured":"Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. MSR-VTT: A large video description dataset for bridging video and language. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.5288\u20135296."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3002669"},{"key":"e_1_3_2_55_2","first-page":"3119","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence.","author":"Yang Bang","year":"2021","unstructured":"Bang Yang, Yuexian Zou, Fenglin Liu, and Can Zhang. 2021. Non-autoregressive coarse-to-fine video captioning. In Proceedings of the AAAI Conference on Artificial Intelligence.3119\u20133127."},{"key":"e_1_3_2_56_2","first-page":"4507","volume-title":"Proceedings of the IEEE International Conference on Computer Vision.","author":"Yao Li","year":"2015","unstructured":"Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher J. Pal, Hugo Larochelle, and Aaron C. Courville. 2015. Describing videos by exploiting temporal structure. In Proceedings of the IEEE International Conference on Computer Vision.4507\u20134515."},{"key":"e_1_3_2_57_2","first-page":"4651","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"You Quanzeng","year":"2016","unstructured":"Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo. 2016. Image captioning with semantic attention. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.4651\u20134659."},{"key":"e_1_3_2_58_2","first-page":"4584","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Yu Haonan","year":"2016","unstructured":"Haonan Yu, Jiang Wang, Zhiheng Huang, Yi Yang, and Wei Xu. 2016. Video paragraph captioning using hierarchical recurrent neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.4584\u20134593."},{"key":"e_1_3_2_59_2","first-page":"1","volume-title":"Proceedings of the International Conference on Multimedia and Big Data","author":"Yuan Jin","year":"2018","unstructured":"Jin Yuan, Chunna Tian, Xiangnan Zhang, Yuxuan Ding, and Wei Wei. 2018. Video captioning with semantic guiding. In Proceedings of the International Conference on Multimedia and Big Data. 1\u20135."},{"key":"e_1_3_2_60_2","first-page":"8327","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Zhang Junchao","year":"2019","unstructured":"Junchao Zhang and Yuxin Peng. 2019. Object-aware aggregation with bidirectional temporal graph for video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.8327\u20138336."},{"key":"e_1_3_2_61_2","first-page":"13275","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Zhang Ziqi","year":"2020","unstructured":"Ziqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li, Peijin Wang, Weiming Hu, and Zheng-Jun Zha. 2020. Object relational graph with teacher-recommended learning for video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.13275\u201313285."},{"key":"e_1_3_2_62_2","first-page":"13093","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Zheng Qi","year":"2020","unstructured":"Qi Zheng, Chaoyue Wang, and Dacheng Tao. 2020. Syntax-aware action targeting for video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.13093\u201313102."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3588572","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3588572","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:47:13Z","timestamp":1750178833000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3588572"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,30]]},"references-count":61,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,11,30]]}},"alternative-id":["10.1145\/3588572"],"URL":"https:\/\/doi.org\/10.1145\/3588572","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2023,5,30]]},"assertion":[{"value":"2022-08-24","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-03-15","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-05-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}