{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T01:47:30Z","timestamp":1782784050752,"version":"3.54.5"},"reference-count":67,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,2,6]],"date-time":"2023-02-06T00:00:00Z","timestamp":1675641600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D program of China","doi-asserted-by":"crossref","award":["2020AAA0105702"],"award-info":[{"award-number":["2020AAA0105702"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Beijing Natural Science Foundation","award":["JQ21017"],"award-info":[{"award-number":["JQ21017"]}]},{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61972397, 62036011, 62192782, 61721004, U19B2038, 61906192"],"award-info":[{"award-number":["61972397, 62036011, 62192782, 61721004, U19B2038, 61906192"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key Research Program of Frontier Sciences, CAS","award":["QYZDJ-SSW-JSC040"],"award-info":[{"award-number":["QYZDJ-SSW-JSC040"]}]},{"name":"University Synergy Innovation Program of Anhui Province","award":["GXXT-2019-025"],"award-info":[{"award-number":["GXXT-2019-025"]}]},{"name":"Science and Technology Service Network Initiative, CAS","award":["KFJ-STS-SCYD-317"],"award-info":[{"award-number":["KFJ-STS-SCYD-317"]}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["WK2100000024"],"award-info":[{"award-number":["WK2100000024"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,5,31]]},"abstract":"<jats:p>Video captioning requires that the model has the abilities of video understanding, video-text alignment, and text generation. Due to the semantic gap between vision and language, conducting video-text alignment is a crucial step to reduce the semantic gap, which maps the representations from the visual to the language domain. However, the existing methods often overlook this step, so the decoder has to directly take the visual representations as input, which increases the decoder\u2019s workload and limits its ability to generate semantically correct captions. In this paper, we propose a video-text alignment module with a retrieval unit and an alignment unit to learn video-text aligned representations for video captioning. Specifically, we firstly propose a retrieval unit to retrieve sentences as additional input which is used as the semantic anchor between visual scene and language description. Then, we employ an alignment unit with the input of the video and retrieved sentences to conduct the video-text alignment. The representations of two modal inputs are aligned in a shared semantic space. The obtained video-text aligned representations are used to generate semantically correct captions. Moreover, retrieved sentences provide rich semantic concepts which are helpful for generating distinctive captions. Experiments on two public benchmarks, i.e., VATEX and MSR-VTT, demonstrate that our method outperforms state-of-the-art performances by a large margin. The qualitative analysis shows that our method generates correct and distinctive captions.<\/jats:p>","DOI":"10.1145\/3546828","type":"journal-article","created":{"date-parts":[[2022,7,7]],"date-time":"2022-07-07T11:29:12Z","timestamp":1657193352000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":25,"title":["Learning Video-Text Aligned Representations for Video Captioning"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0465-6712","authenticated-orcid":false,"given":"Yaya","family":"Shi","sequence":"first","affiliation":[{"name":"School of Information Science and Technology, University of Science and Technology of China, Hefei, Anhui, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9442-5912","authenticated-orcid":false,"given":"Haiyang","family":"Xu","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2219-4961","authenticated-orcid":false,"given":"Chunfeng","family":"Yuan","sequence":"additional","affiliation":[{"name":"NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6114-1411","authenticated-orcid":false,"given":"Bing","family":"Li","sequence":"additional","affiliation":[{"name":"NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9237-8825","authenticated-orcid":false,"given":"Weiming","family":"Hu","sequence":"additional","affiliation":[{"name":"NLPR, Institute of Automation, Chinese Academy of Sciences, China and School of ArtificialIntelligence, University of Chinese Academy of Sciences, China and CAS Center for Excellence in BrainScience and Intelligence Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2510-8993","authenticated-orcid":false,"given":"Zheng-Jun","family":"Zha","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, University of Science and Technology of China, Hefei, Anhui, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,2,6]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"12487","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Aafaq Nayyer","year":"2019","unstructured":"Nayyer Aafaq, Naveed Akhtar, Wei Liu, Syed Zulqarnain Gilani, and Ajmal Mian. 2019. Spatio-temporal dynamics and semantic attribute enriched visual encoding for video captioning. In IEEE Conference on Computer Vision and Pattern Recognition. 12487\u201312496."},{"key":"e_1_3_1_3_2","first-page":"1728","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Bain Max","year":"2021","unstructured":"Max Bain, Arsha Nagrani, G\u00fcl Varol, and Andrew Zisserman. 2021. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 1728\u20131738."},{"key":"e_1_3_1_4_2","first-page":"65","volume-title":"Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization","author":"Banerjee Satanjeev","year":"2005","unstructured":"Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization. 65\u201372."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.5555\/3020652.3020667"},{"key":"e_1_3_1_6_2","first-page":"190","volume-title":"The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference, 19-24 June, 2011, Portland, Oregon, USA","author":"Chen David L.","year":"2011","unstructured":"David L. Chen and William B. Dolan. 2011. Collecting highly parallel data for paraphrase evaluation. In The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference, 19-24 June, 2011, Portland, Oregon, USA, Dekang Lin, Yuji Matsumoto, and Rada Mihalcea (Eds.). The Association for Computer Linguistics, 190\u2013200."},{"key":"e_1_3_1_7_2","first-page":"8191","volume-title":"The Thirty-Third AAAI Conference on Artificial Intelligence","author":"Chen Shaoxiang","year":"2019","unstructured":"Shaoxiang Chen and Yu-Gang Jiang. 2019. Motion guided spatial attention for video captioning. In The Thirty-Third AAAI Conference on Artificial Intelligence. 8191\u20138198."},{"key":"e_1_3_1_8_2","first-page":"10635","volume-title":"2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Shizhe","year":"2020","unstructured":"Shizhe Chen, Yida Zhao, Qin Jin, and Qi Wu. 2020. Fine-grained video-text retrieval with hierarchical graph reasoning. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10635\u201310644."},{"key":"e_1_3_1_9_2","article-title":"Microsoft COCO captions: Data collection and evaluation server","volume":"1504","author":"Chen Xinlei","year":"2015","unstructured":"Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2015. Microsoft COCO captions: Data collection and evaluation server. CoRR abs\/1504.00325 (2015). arxiv:1504.00325.","journal-title":"CoRR"},{"key":"e_1_3_1_10_2","first-page":"898","volume-title":"Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems","author":"Dai Bo","year":"2017","unstructured":"Bo Dai and Dahua Lin. 2017. Contrastive learning for image captioning. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems. 898\u2013907."},{"key":"e_1_3_1_11_2","first-page":"2634","volume-title":"2013 IEEE Conference on Computer Vision and Pattern Recognition","author":"Das Pradipto","year":"2013","unstructured":"Pradipto Das, Chenliang Xu, Richard F. Doell, and Jason J. Corso. 2013. A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching. In 2013 IEEE Conference on Computer Vision and Pattern Recognition. 2634\u20132641."},{"key":"e_1_3_1_12_2","first-page":"248","volume-title":"2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"Deng Jia","year":"2009","unstructured":"Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li. 2009. ImageNet: A large-scale hierarchical image database. In 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. 248\u2013255."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3063423"},{"key":"e_1_3_1_14_2","first-page":"58","volume-title":"British Machine Vision Conference","author":"Dong Jiarong","year":"2018","unstructured":"Jiarong Dong, Ke Gao, Xiaokai Chen, Junbo Guo, Juan Cao, and Yongdong Zhang. 2018. Not all words are equal: Video-specific information loss for video captioning. In British Machine Vision Conference. 58."},{"key":"e_1_3_1_15_2","first-page":"9346","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Dong Jianfeng","year":"2019","unstructured":"Jianfeng Dong, Xirong Li, Chaoxi Xu, Shouling Ji, Yuan He, Gang Yang, and Xun Wang. 2019. Dual encoding for zero-example video retrieval. In IEEE Conference on Computer Vision and Pattern Recognition. 9346\u20139355."},{"key":"e_1_3_1_16_2","article-title":"CLIP2Video: Mastering video-text retrieval via image CLIP","volume":"2106","author":"Fang Han","year":"2021","unstructured":"Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen. 2021. CLIP2Video: Mastering video-text retrieval via image CLIP. CoRR abs\/2106.11097 (2021). arxiv:2106.11097","journal-title":"CoRR"},{"key":"e_1_3_1_17_2","first-page":"10324","volume-title":"2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Guo Longteng","year":"2020","unstructured":"Longteng Guo, Jing Liu, Xinxin Zhu, Peng Yao, Shichen Lu, and Hanqing Lu. 2020. Normalized and geometry-aware self-attention network for image captioning. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10324\u201310333."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.5555\/2900728.2900815"},{"key":"e_1_3_1_19_2","series-title":"Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event","first-page":"3929","volume":"119","author":"Guu Kelvin","year":"2020","unstructured":"Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Retrieval augmented language model pre-training. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event(Proceedings of Machine Learning Research, Vol. 119). PMLR, 3929\u20133938."},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","first-page":"6546","DOI":"10.1109\/CVPR.2018.00685","volume-title":"2018 IEEE Conference on Computer Vision and Pattern Recognition","author":"Hara Kensho","year":"2018","unstructured":"Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. 2018. Can spatiotemporal 3D CNNs retrace the history of 2D CNNs and ImageNet? In 2018 IEEE Conference on Computer Vision and Pattern Recognition. 6546\u20136555."},{"issue":"1","key":"e_1_3_1_21_2","first-page":"26","article-title":"Image captioning with visual-semantic double attention","volume":"15","author":"He Chen","year":"2019","unstructured":"Chen He and Haifeng Hu. 2019. Image captioning with visual-semantic double attention. ACM Trans. Multimedia Comput. Commun. Appl. 15, 1, Article 26 (Jan.2019), 16 pages.","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_23_2","first-page":"8917","volume-title":"2019 IEEE\/CVF International Conference on Computer Vision","author":"Hou Jingyi","year":"2019","unstructured":"Jingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo, and Yunde Jia. 2019. Joint syntax representation learning and visual cue translation for video captioning. In 2019 IEEE\/CVF International Conference on Computer Vision. 8917\u20138926."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3351072"},{"issue":"2","key":"e_1_3_1_25_2","first-page":"72","article-title":"A multi-instance multi-label dual learning approach for video captioning","volume":"17","author":"Ji Wanting","year":"2021","unstructured":"Wanting Ji and Ruili Wang. 2021. A multi-instance multi-label dual learning approach for video captioning. ACM Trans. Multimedia Comput. Commun. Appl. 17, 2s, Article 72 (Jun.2021), 18 pages.","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"issue":"4","key":"e_1_3_1_26_2","first-page":"125","article-title":"Bi-directional Co-attention network for image captioning","volume":"17","author":"Jiang Weitao","year":"2021","unstructured":"Weitao Jiang, Weixuan Wang, and Haifeng Hu. 2021. Bi-directional Co-attention network for image captioning. ACM Trans. Multimedia Comput. Commun. Appl. 17, 4, Article 125 (Nov.2021), 20 pages.","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_1_27_2","article-title":"The kinetics human action video dataset","author":"Kay Will","year":"2017","unstructured":"Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et\u00a0al. 2017. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950 (2017).","journal-title":"arXiv preprint arXiv:1705.06950"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1020346032608"},{"key":"e_1_3_1_29_2","first-page":"706","volume-title":"IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22\u201329, 2017","author":"Krishna Ranjay","year":"2017","unstructured":"Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017. Dense-captioning events in videos. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22\u201329, 2017. IEEE Computer Society, 706\u2013715."},{"key":"e_1_3_1_30_2","first-page":"359","volume-title":"The 50th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference","author":"Kuznetsova Polina","year":"2012","unstructured":"Polina Kuznetsova, Vicente Ordonez, Alexander C. Berg, Tamara L. Berg, and Yejin Choi. 2012. Collective generation of natural image descriptions. In The 50th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference. 359\u2013368."},{"key":"e_1_3_1_31_2","volume-title":"Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6\u201312, 2020, virtual","author":"Lewis Mike","year":"2020","unstructured":"Mike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan, Sida I. Wang, and Luke Zettlemoyer. 2020. Pre-training via paraphrasing. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6\u201312, 2020, virtual, Hugo Larochelle, Marc\u2019Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)."},{"key":"e_1_3_1_32_2","unstructured":"NIPS\u201920 Proceedings of the 34th International Conference on Neural Information Processing Systems Patrick Lewis Ethan Perez Aleksandra Piktus Fabio Petroni Vladimir Karpukhin Naman Goyal Heinrich K\u00fcttler Mike Lewis Wen-tau Yih Tim Rockt\u00e4schel Sebastian Riedel Douwe Kiela Retrieval-augmented generation for knowledge-intensive NLP tasks 2020 793"},{"key":"e_1_3_1_33_2","first-page":"2046","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing","author":"Li Linjie","year":"2020","unstructured":"Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. 2020. HERO: Hierarchical encoder for video+language Omni-representation pre-training. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing. 2046\u20132065."},{"key":"e_1_3_1_34_2","first-page":"74","volume-title":"Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out. 74\u201381."},{"key":"e_1_3_1_35_2","first-page":"1416","volume-title":"2018 ACM Multimedia Conference on Multimedia Conference, MM 2018","author":"Liu Daqing","unstructured":"Daqing Liu, Zheng-Jun Zha, Hanwang Zhang, Yongdong Zhang, and Feng Wu. [n.d.]. Context-aware visual policy network for sequence-level image captioning. In 2018 ACM Multimedia Conference on Multimedia Conference, MM 2018. 1416\u20131424."},{"key":"e_1_3_1_36_2","first-page":"4239","volume-title":"2019 IEEE\/CVF International Conference on Computer Vision","author":"Liu Lixin","year":"2019","unstructured":"Lixin Liu, Jiajun Tang, Xiaojun Wan, and Zongming Guo. 2019. Generating diverse and descriptive image captions using visual paraphrases. In 2019 IEEE\/CVF International Conference on Computer Vision. 4239\u20134248."},{"key":"e_1_3_1_37_2","doi-asserted-by":"crossref","first-page":"353","DOI":"10.1007\/978-3-030-01267-0_21","volume-title":"Computer Vision - ECCV 2018-15th European Conference, Munich, Germany, September 8\u201314, 2018, Proceedings, Part XV","volume":"11219","author":"Liu Xihui","year":"2018","unstructured":"Xihui Liu, Hongsheng Li, Jing Shao, Dapeng Chen, and Xiaogang Wang. 2018. Show, tell and discriminate: Image captioning by self-retrieval with partially labeled data. In Computer Vision - ECCV 2018-15th European Conference, Munich, Germany, September 8\u201314, 2018, Proceedings, Part XV, Vol. 11219. 353\u2013369."},{"key":"e_1_3_1_38_2","volume-title":"7th International Conference on Learning Representations","author":"Loshchilov Ilya","year":"2019","unstructured":"Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations."},{"key":"e_1_3_1_39_2","article-title":"CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval","volume":"2104","author":"Luo Huaishao","year":"2021","unstructured":"Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. 2021. CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval. CoRR abs\/2104.08860 (2021). arxiv:2104.08860.","journal-title":"CoRR"},{"key":"e_1_3_1_40_2","first-page":"6964","volume-title":"2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018","author":"Luo Ruotian","year":"2018","unstructured":"Ruotian Luo, Brian L. Price, Scott Cohen, and Gregory Shakhnarovich. 2018. Discriminability objective for training descriptive captions. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018. IEEE Computer Society, 6964\u20136974."},{"key":"e_1_3_1_41_2","first-page":"9876","volume-title":"2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Miech Antoine","year":"2020","unstructured":"Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. 2020. End-to-end learning of visual representations from uncurated instructional videos. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 9876\u20139886."},{"key":"e_1_3_1_42_2","first-page":"2630","volume-title":"2019 IEEE\/CVF International Conference on Computer Vision","author":"Miech Antoine","year":"2019","unstructured":"Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019. HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips. In 2019 IEEE\/CVF International Conference on Computer Vision. 2630\u20132640."},{"key":"e_1_3_1_43_2","first-page":"1143","volume-title":"Advances in Neural Information Processing Systems 24: 25th Annual Conference on Neural Information Processing Systems","author":"Ordonez Vicente","year":"2011","unstructured":"Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. 2011. Im2Text: Describing images using 1 million captioned photographs. In Advances in Neural Information Processing Systems 24: 25th Annual Conference on Neural Information Processing Systems. 1143\u20131151."},{"key":"e_1_3_1_44_2","first-page":"10867","volume-title":"2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Pan Boxiao","year":"2020","unstructured":"Boxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee, Adrien Gaidon, Ehsan Adeli, and Juan Carlos Niebles. 2020. Spatio-temporal graph for video captioning with knowledge distillation. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10867\u201310876."},{"key":"e_1_3_1_45_2","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. 311\u2013318."},{"key":"e_1_3_1_46_2","first-page":"8347","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Pei Wenjie","year":"2019","unstructured":"Wenjie Pei, Jiyuan Zhang, Xiangrong Wang, Lei Ke, Xiaoyong Shen, and Yu-Wing Tai. 2019. Memory-attended recurrent network for video captioning. In IEEE Conference on Computer Vision and Pattern Recognition. 8347\u20138356."},{"key":"e_1_3_1_47_2","article-title":"A straightforward framework for video retrieval using CLIP","volume":"2102","author":"Portillo-Quintero Jes\u00fas Andr\u00e9s","year":"2021","unstructured":"Jes\u00fas Andr\u00e9s Portillo-Quintero, Jos\u00e9 Carlos Ortiz-Bayliss, and Hugo Terashima-Mar\u00edn. 2021. A straightforward framework for video retrieval using CLIP. CoRR abs\/2102.12443 (2021). arxiv:2102.12443.","journal-title":"CoRR"},{"key":"e_1_3_1_48_2","article-title":"Learning transferable visual models from natural language supervision","volume":"2103","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. CoRR abs\/2103.00020 (2021). arxiv:2103.00020.","journal-title":"CoRR"},{"key":"e_1_3_1_49_2","first-page":"2514","volume-title":"Thirty-Fifth AAAI Conference on Artificial Intelligence","author":"Ryu Hobin","year":"2021","unstructured":"Hobin Ryu, Sunghun Kang, Haeyong Kang, and Chang D. Yoo. 2021. Semantic grouping network for video captioning. In Thirty-Fifth AAAI Conference on Artificial Intelligence. 2514\u20132522."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.5555\/3298023.3298188"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2020\/104"},{"issue":"2","key":"e_1_3_1_52_2","first-page":"31","article-title":"Rich visual and language representation with complementary semantics for video captioning","volume":"15","author":"Tang Pengjie","year":"2019","unstructured":"Pengjie Tang, Hanli Wang, and Qinyu Li. 2019. Rich visual and language representation with complementary semantics for video captioning. ACM Trans. Multimedia Comput. Commun. Appl. 15, 2, Article 31 (Jun.2019), 23 pages.","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_1_53_2","first-page":"5998","volume-title":"Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_1_54_2","doi-asserted-by":"crossref","first-page":"4566","DOI":"10.1109\/CVPR.2015.7299087","volume-title":"2015 IEEE Conference on Computer Vision and Pattern Recognition","author":"Vedantam R.","year":"2015","unstructured":"R. Vedantam, C. L. Zitnick, and D. Parikh. 2015. CIDEr: Consensus-based image description evaluation. In 2015 IEEE Conference on Computer Vision and Pattern Recognition. 4566\u20134575."},{"key":"e_1_3_1_55_2","first-page":"2641","volume-title":"2019 IEEE\/CVF International Conference on Computer Vision","author":"Wang Bairui","year":"2019","unstructured":"Bairui Wang, Lin Ma, Wei Zhang, Wenhao Jiang, Jingwen Wang, and Wei Liu. 2019. Controllable video captioning with POS sequence guidance based on gated fusion network. In 2019 IEEE\/CVF International Conference on Computer Vision. 2641\u20132650."},{"key":"e_1_3_1_56_2","first-page":"4580","volume-title":"2019 IEEE\/CVF International Conference on Computer Vision","author":"Wang Xin","year":"2019","unstructured":"Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. 2019. VaTeX: A large-scale, high-quality multilingual dataset for video-and-language research. In 2019 IEEE\/CVF International Conference on Computer Vision. 4580\u20134590."},{"key":"e_1_3_1_57_2","article-title":"Memorizing transformers","volume":"2203","author":"Wu Yuhuai","year":"2022","unstructured":"Yuhuai Wu, Markus N. Rabe, DeLesley Hutchins, and Christian Szegedy. 2022. Memorizing transformers. CoRR abs\/2203.08913 (2022). arxiv:2203.08913.","journal-title":"CoRR"},{"key":"e_1_3_1_58_2","article-title":"Google\u2019s neural machine translation system: Bridging the gap between human and machine translation","volume":"1609","author":"Wu Yonghui","year":"2016","unstructured":"Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016. Google\u2019s neural machine translation system: Bridging the gap between human and machine translation. CoRR abs\/1609.08144 (2016). arxiv:1609.08144.","journal-title":"CoRR"},{"key":"e_1_3_1_59_2","first-page":"5288","volume-title":"2016 IEEE Conference on Computer Vision and Pattern Recognition","author":"Xu Jun","year":"2016","unstructured":"Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. MSR-VTT: A large video description dataset for bridging video and language. In 2016 IEEE Conference on Computer Vision and Pattern Recognition. 5288\u20135296."},{"key":"e_1_3_1_60_2","first-page":"3119","volume-title":"Thirty-Fifth AAAI Conference on Artificial Intelligence","author":"Yang Bang","year":"2021","unstructured":"Bang Yang, Yuexian Zou, Fenglin Liu, and Can Zhang. 2021. Non-autoregressive coarse-to-fine video captioning. In Thirty-Fifth AAAI Conference on Artificial Intelligence. 3119\u20133127."},{"issue":"2","key":"e_1_3_1_61_2","first-page":"55","article-title":"Image captioning by asking questions","volume":"15","author":"Yang Xiaoshan","year":"2019","unstructured":"Xiaoshan Yang and Changsheng Xu. 2019. Image captioning by asking questions. ACM Trans. Multimedia Comput. Commun. Appl. 15, 2s, Article 55 (Jul.2019), 19 pages.","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_1_62_2","first-page":"4507","volume-title":"2015 IEEE International Conference on Computer Vision","author":"Yao Li","year":"2015","unstructured":"Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher J. Pal, Hugo Larochelle, and Aaron C. Courville. 2015. Describing videos by exploiting temporal structure. In 2015 IEEE International Conference on Computer Vision. 4507\u20134515."},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00371"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2909864"},{"key":"e_1_3_1_65_2","first-page":"8327","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhang Junchao","year":"2019","unstructured":"Junchao Zhang and Yuxin Peng. 2019. Object-aware aggregation with bidirectional temporal graph for video captioning. In IEEE Conference on Computer Vision and Pattern Recognition. 8327\u20138336."},{"key":"e_1_3_1_66_2","first-page":"9837","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Ziqi","year":"2021","unstructured":"Ziqi Zhang, Zhongang Qi, Chunfeng Yuan, Ying Shan, Bing Li, Ying Deng, and Weiming Hu. 2021. Open-book video captioning with retrieve-copy-generate network. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 9837\u20139846."},{"key":"e_1_3_1_67_2","first-page":"13275","volume-title":"2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Ziqi","year":"2020","unstructured":"Ziqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li, Peijin Wang, Weiming Hu, and Zheng-Jun Zha. 2020. Object relational graph with teacher-recommended learning for video captioning. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 13275\u201313285."},{"key":"e_1_3_1_68_2","first-page":"13093","volume-title":"2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zheng Qi","year":"2020","unstructured":"Qi Zheng, Chaoyue Wang, and Dacheng Tao. 2020. Syntax-aware action targeting for video captioning. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 13093\u201313102."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3546828","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3546828","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:41Z","timestamp":1750186841000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3546828"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,6]]},"references-count":67,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,5,31]]}},"alternative-id":["10.1145\/3546828"],"URL":"https:\/\/doi.org\/10.1145\/3546828","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,6]]},"assertion":[{"value":"2021-12-02","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-06-23","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}