{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,28]],"date-time":"2026-03-28T15:55:18Z","timestamp":1774713318507,"version":"3.50.1"},"reference-count":73,"publisher":"Association for Computing Machinery (ACM)","issue":"11","funder":[{"name":"\u201cPioneer\u201d and \u201cLeading Goose\u201d R&D Program of Zhejiang","award":["2023C01038"],"award-info":[{"award-number":["2023C01038"]}]},{"name":"Key R&D Program of Xinjiang, China","award":["2022B01006"],"award-info":[{"award-number":["2022B01006"]}]},{"name":"Zhejiang Provincial Natural Science Foundation of China","award":["LDT23F01013F01"],"award-info":[{"award-number":["LDT23F01013F01"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>Text-video retrieval, a fundamental task for associating textual descriptions with video content, has become increasingly important in the video domain. Most existing methods focus on the single-modality features only considering the knowledge within individual video or text modalities, often neglecting cross-modal interactions. However, a text description corresponds to a specific spatio-temporal content within a video, involving a certain segment of a frame sequence and distinct sub-regions within these frames. Therefore, we focus on the text-conditioned video features to bridge the modality gap. In this article, we propose Spatio-Temporal Attention for video-text retrieval, termed STAttn, which utilizes textual information to focus on the spatio-temporal video content. Our final text-conditioned video features are generated from the text-related video frames and the text-related regions within these frames. First, we propose the Spatial Text-Attention Module (STAM) to learn the spatial information within video frames. STAM introduces the text-related salient patches to capture more fine-grained details. Second, we propose the Temporal Text-Attention Module (TTAM) to learn the temporal relationships between video frames. Temporal Triplet loss is proposed in TTAM to enhance the attention toward the text-related frames. Thus, the two modules learn the text-related spatio-temporal content from both intra-frame and inter-frame aspects. Extensive experiments on three benchmark datasets, MSRVTT, ActivityNet, and DiDeMo, demonstrate that our STAttn outperforms the state-of-the-art methods.<\/jats:p>","DOI":"10.1145\/3715137","type":"journal-article","created":{"date-parts":[[2025,1,28]],"date-time":"2025-01-28T10:53:42Z","timestamp":1738061622000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Spatio-Temporal Attention for Text-Video Retrieval"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7742-9142","authenticated-orcid":false,"given":"Leqi","family":"Shen","sequence":"first","affiliation":[{"name":"School of Software and BNRist, Tsinghua University, Beijing, China and Zhuoxi Institute of Brain and Intelligence, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5843-6411","authenticated-orcid":false,"given":"Sicheng","family":"Zhao","sequence":"additional","affiliation":[{"name":"BNRist, Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-5023-9288","authenticated-orcid":false,"given":"Yifeng","family":"Zhang","sequence":"additional","affiliation":[{"name":"JD.com Inc, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6031-5245","authenticated-orcid":false,"given":"Pengzhang","family":"Liu","sequence":"additional","affiliation":[{"name":"JD.com Inc, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7816-0587","authenticated-orcid":false,"given":"Yongjun","family":"Bao","sequence":"additional","affiliation":[{"name":"JD.com Inc, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0137-9975","authenticated-orcid":false,"given":"Guiguang","family":"Ding","sequence":"additional","affiliation":[{"name":"School of Software and BNRist, Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,11,7]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"5803","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Hendricks Lisa Anne","year":"2017","unstructured":"Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017. Localizing moments in video with natural language. In Proceedings of the International Conference on Computer Vision, 5803\u20135812."},{"key":"e_1_3_1_3_2","first-page":"2425","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Antol Stanislaw","year":"2015","unstructured":"Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the International Conference on Computer Vision, 2425\u20132433."},{"key":"e_1_3_1_4_2","first-page":"1728","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Bain Max","year":"2021","unstructured":"Max Bain, Arsha Nagrani, G\u00fcl Varol, and Andrew Zisserman. 2021. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the International Conference on Computer Vision, 1728\u20131738."},{"key":"e_1_3_1_5_2","first-page":"3041","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Bin Yi","year":"2023","unstructured":"Yi Bin, Haoxuan Li, Yahui Xu, Xing Xu, Yang Yang, and Heng Tao Shen. 2023. Unifying two-stream encoders with transformers for cross-modal retrieval. In Proceedings of the ACM International Conference on Multimedia, 3041\u20133050."},{"key":"e_1_3_1_6_2","first-page":"5194","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Bogolin Simion-Vlad","year":"2022","unstructured":"Simion-Vlad Bogolin, Ioana Croitoru, Hailin Jin, Yang Liu, and Samuel Albanie. 2022. Cross modal retrieval with querybank normalisation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5194\u20135205."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01065"},{"key":"e_1_3_1_8_2","unstructured":"Xing Cheng Hezheng Lin Xiangyu Wu Fan Yang and Dong Shen. 2021. Improving video-text retrieval by multi-stream corpus alignment and dual softmax loss. arXiv:2109.04290. Retrieved from https:\/\/arxiv.org\/abs\/2109.04290"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3499027"},{"key":"e_1_3_1_10_2","first-page":"11583","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Croitoru Ioana","year":"2021","unstructured":"Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin, Andrew Zisserman, Samuel Albanie, and Yang Liu. 2021. Teachtext: Crossmodal generalized distillation for text-video retrieval. In Proceedings of the International Conference on Computer Vision, 11583\u201311593."},{"key":"e_1_3_1_11_2","first-page":"15648","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Deng Chaorui","year":"2023","unstructured":"Chaorui Deng, Qi Chen, Pengda Qin, Da Chen, and Qi Wu. 2023. Prompt switch: Efficient clip adaptation for text-video retrieval. In Proceedings of the International Conference on Computer Vision, 15648\u201315658."},{"key":"e_1_3_1_12_2","first-page":"3354","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Dzabraev Maksim","year":"2021","unstructured":"Maksim Dzabraev, Maksim Kalashnikov, Stepan Komkov, and Aleksandr Petiushko. 2021. Mdmmt: Multidomain multimodal transformer for video retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3354\u20133363."},{"key":"e_1_3_1_13_2","first-page":"13723","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Fang Bo","year":"2023","unstructured":"Bo Fang, Wenhao Wu, Chang Liu, Yu Zhou, Yuxin Song, Weiping Wang, Xiangbo Shu, Xiangyang Ji, and Jingdong Wang. 2023. UATVR: Uncertainty-adaptive text-video retrieval. In Proceedings of the International Conference on Computer Vision, 13723\u201313733."},{"key":"e_1_3_1_14_2","unstructured":"Han Fang Pengfei Xiong Luhui Xu and Yu Chen. 2021. Clip2video: Mastering video-text retrieval via image clip. arXiv:2106.11097. Retrieved from https:\/\/arxiv.org\/abs\/2106.11097"},{"key":"e_1_3_1_15_2","first-page":"13109","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Fei Hao","year":"2024","unstructured":"Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, Meishan Zhang, Mong-Li Lee, and Wynne Hsu. 2024. Video-of-thought: Step-by-step video reasoning from perception to cognition. In Proceedings of the International Conference on Machine Learning, 13109\u201313125."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i2.27943"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58548-8_13"},{"key":"e_1_3_1_18_2","first-page":"5267","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Gao Jiyang","year":"2017","unstructured":"Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. 2017. Tall: Temporal activity localization via language query. In Proceedings of the International Conference on Computer Vision, 5267\u20135275."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV56688.2023.00108"},{"key":"e_1_3_1_20_2","first-page":"5006","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Gorti Satya Krishna","year":"2022","unstructured":"Satya Krishna Gorti, No\u00ebl Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, and Guangwei Yu. 2022. X-pool: Cross-modal language-video attention for text-video retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5006\u20135015."},{"key":"e_1_3_1_21_2","unstructured":"Priya Goyal Piotr Doll\u00e1r Ross Girshick Pieter Noordhuis Lukasz Wesolowski Aapo Kyrola Andrew Tulloch Yangqing Jia and Kaiming He. 2017. Accurate large minibatch sgd: Training imagenet in 1 hour. arXiv:1706.02677. Retrieved from https:\/\/arxiv.org\/abs\/1706.02677"},{"key":"e_1_3_1_22_2","first-page":"11164","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Guan Peiyan","year":"2023","unstructured":"Peiyan Guan, Renjing Pei, Bin Shao, Jianzhuang Liu, Weimian Li, Jiaxi Gu, Hang Xu, Songcen Xu, Youliang Yan, and Edmund Y. Lam. 2023. PIDRo: Parallel isomeric attention with dynamic routing for text-video retrieval. In Proceedings of the International Conference on Computer Vision, 11164\u201311173."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3483381"},{"key":"e_1_3_1_24_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo","author":"He Tao","year":"2024","unstructured":"Tao He, Leqi Shen, Guiguang Ding, Zhiheng Zhou, Tianshi Xu, Xiaofeng Jin, and Yuheng Huang. 2024. Balanced active sampling for person re-identification. In Proceedings of the IEEE International Conference on Multimedia and Expo. IEEE, 1\u20136."},{"key":"e_1_3_1_25_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo","author":"He Tao","year":"2024","unstructured":"Tao He, Leqi Shen, Guiguang Ding, Zhiheng Zhou, Tianshi Xu, Xiaofeng Jin, and Yuheng Huang. 2024. Camera bias regularization for person re-identification. In Proceedings of the IEEE International Conference on Multimedia and Expo. IEEE, 1\u20136."},{"key":"e_1_3_1_26_2","first-page":"879","volume-title":"Proceedings of the 36th AAAI Conference on Artificial Intelligence","author":"He Tao","year":"2022","unstructured":"Tao He, Leqi Shen, Yuchen Guo, Guiguang Ding, and Zhenhua Guo. 2022. Secret: Self-consistent pseudo label refinement for unsupervised domain adaptive person re-identification. In Proceedings of the 36th AAAI Conference on Artificial Intelligence, 879\u2013887."},{"key":"e_1_3_1_27_2","first-page":"12054","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Ibrahimi Sarah","year":"2023","unstructured":"Sarah Ibrahimi, Xiaohang Sun, Pichao Wang, Amanmeet Garg, Ashutosh Sanan, and Mohamed Omar. 2023. Audio-enhanced text-to-video retrieval using text-conditioned feature alignment. In Proceedings of the International Conference on Computer Vision, 12054\u201312064."},{"key":"e_1_3_1_28_2","first-page":"1595","volume-title":"Companion Proceedings of the ACM on Web Conference","author":"Wei Ji","year":"2024","unstructured":"Wei Ji, Ruiqi Shi, Yinwei Wei, Shanshan Zhao, and Roger Zimmermann. 2024. Weakly supervised video moment retrieval via location-irrelevant proposal learning. In Companion Proceedings of the ACM on Web Conference, 1595\u20131603."},{"key":"e_1_3_1_29_2","first-page":"30291","article-title":"Expectation-maximization contrastive learning for compact video-and-language representations","volume":"35","author":"Jin Peng","year":"2022","unstructured":"Peng Jin, Jinfa Huang, Fenglin Liu, Xian Wu, Shen Ge, Guoli Song, David Clifton, and Jie Chen. 2022. Expectation-maximization contrastive learning for compact video-and-language representations. Advances in Neural Information Processing Systems 35 (2022), 30291\u201330306.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_30_2","first-page":"2472","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Jin Peng","year":"2023","unstructured":"Peng Jin, Jinfa Huang, Pengfei Xiong, Shangxuan Tian, Chang Liu, Xiangyang Ji, Li Yuan, and Jie Chen. 2023. Video-text as game players: Hierarchical banzhaf interaction for cross-modal representation learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2472\u20132482."},{"key":"e_1_3_1_31_2","first-page":"938","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence","author":"Jin Peng","year":"2023","unstructured":"Peng Jin, Hao Li, Zesen Cheng, Jinfa Huang, Zhennan Wang, Li Yuan, Chang Liu, and Jie Chen. 2023. Text-video retrieval with disentangled conceptualization and set-to-set alignment. In Proceedings of the International Joint Conference on Artificial Intelligence, 938\u2013946."},{"key":"e_1_3_1_32_2","first-page":"2470","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Jin Peng","year":"2023","unstructured":"Peng Jin, Hao Li, Zesen Cheng, Kehan Li, Xiangyang Ji, Chang Liu, Li Yuan, and Jie Chen. 2023. DiffusionRet: Generative text-video retrieval with diffusion model. In Proceedings of the International Conference on Computer Vision, 2470\u20132481."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_3_1_34_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=8gmWwjFyLj"},{"key":"e_1_3_1_35_2","first-page":"706","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Krishna Ranjay","year":"2017","unstructured":"Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017. Dense-captioning events in videos. In Proceedings of the International Conference on Computer Vision, 706\u2013715."},{"key":"e_1_3_1_36_2","first-page":"7331","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Lei Jie","year":"2021","unstructured":"Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L. Berg, Mohit Bansal, and Jingjing Liu. 2021. Less is more: Clipbert for video-and-language learning via sparse sampling. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 7331\u20137341."},{"key":"e_1_3_1_37_2","first-page":"924","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Li Haoxuan","year":"2023","unstructured":"Haoxuan Li, Yi Bin, Junrong Liao, Yang Yang, and Heng Tao Shen. 2023. Your negative may not be true negative: Boosting image-text matching with false negative elimination. In Proceedings of the ACM International Conference on Multimedia, 924\u2013934."},{"key":"e_1_3_1_38_2","first-page":"9694","article-title":"Align before fuse: Vision and language representation learning with momentum distillation","volume":"34","author":"Li Junnan","year":"2021","unstructured":"Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. Advances in Neural Information Processing Systems 34 (2021), 9694\u20139705.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_39_2","first-page":"4100","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Li Pandeng","year":"2023","unstructured":"Pandeng Li, Chen-Wei Xie, Liming Zhao, Hongtao Xie, Jiannan Ge, Yun Zheng, Deli Zhao, and Yongdong Zhang. 2023. Progressive spatio-temporal prototype matching for text-video retrieval. In Proceedings of the International Conference on Computer Vision, 4100\u20134110."},{"key":"e_1_3_1_40_2","first-page":"6555","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Liu Ruyang","year":"2023","unstructured":"Ruyang Liu, Jingjia Huang, Ge Li, Jiashi Feng, Xinglong Wu, and Thomas H. Li. 2023. Revisiting temporal modeling for clip-based image-to-video knowledge transferring. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6555\u20136564."},{"issue":"3","key":"e_1_3_1_41_2","first-page":"3003","article-title":"Entity-enhanced adaptive reconstruction network for weakly supervised referring expression grounding","volume":"45","author":"Liu Xuejing","year":"2022","unstructured":"Xuejing Liu, Liang Li, Shuhui Wang, Zheng-Jun Zha, Zechao Li, Qi Tian, and Qingming Huang. 2022. Entity-enhanced adaptive reconstruction network for weakly supervised referring expression grounding. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 3 (2022), 3003\u20133018.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_42_2","unstructured":"Yang Liu Samuel Albanie Arsha Nagrani and Andrew Zisserman. 2019. Use what you have: Video retrieval using representations from collaborative experts. arXiv:1907.13487. Retrieved from https:\/\/arxiv.org\/abs\/1907.13487"},{"key":"e_1_3_1_43_2","first-page":"319","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Liu Yuqi","year":"2022","unstructured":"Yuqi Liu, Pengfei Xiong, Luhui Xu, Shengming Cao, and Qin Jin.2022. Ts2-net: Token shift and selection transformer for text-video retrieval. In Proceedings of the European Conference on Computer Vision, 319\u2013335."},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.07.028"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547910"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377875"},{"key":"e_1_3_1_47_2","first-page":"2630","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Miech Antoine","year":"2019","unstructured":"Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the International Conference on Computer Vision, 2630\u20132640."},{"key":"e_1_3_1_48_2","first-page":"19","volume-title":"Proceedings of the ACM International Conference on Multimedia Retrieval","author":"Mithun Niluthpol Chowdhury","year":"2018","unstructured":"Niluthpol Chowdhury Mithun, Juncheng Li, Florian Metze, and Amit K. Roy-Chowdhury. 2018. Learning joint embedding with multimodal cues for cross-modal video-text retrieval. In Proceedings of the ACM International Conference on Multimedia Retrieval, 19\u201327."},{"key":"e_1_3_1_49_2","unstructured":"Aaron van den Oord Yazhe Li and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv:1807.03748. Retrieved from https:\/\/arxiv.org\/abs\/1807.03748"},{"key":"e_1_3_1_50_2","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/978-3-030-77004-4_1","volume-title":"Proceedings of the Mexican Conference on Pattern Recognition","author":"Portillo-Quintero Jes\u00fas Andr\u00e9s","year":"2021","unstructured":"Jes\u00fas Andr\u00e9s Portillo-Quintero, Jos\u00e9 Carlos Ortiz-Bayliss, and Hugo Terashima-Mar\u00edn. 2021. A straightforward framework for video retrieval using clip. In Proceedings of the Mexican Conference on Pattern Recognition, 3\u201312."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462829"},{"key":"e_1_3_1_52_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, 8748\u20138763."},{"key":"e_1_3_1_53_2","first-page":"13756","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Shao Bin","year":"2023","unstructured":"Bin Shao, Jianzhuang Liu, Renjing Pei, Songcen Xu, Peng Dai, Juwei Lu, Weimian Li, and Youliang Yan. 2023. HiVLP: Hierarchical interactive video-language pre-training. In Proceedings of the International Conference on Computer Vision, 13756\u201313766."},{"key":"e_1_3_1_54_2","unstructured":"Leqi Shen Tianxiang Hao Sicheng Zhao Yifeng Zhang Pengzhang Liu Yongjun Bao and Guiguang Ding. 2024. Tempme: Video temporal token merging for efficient text-video retrieval. arXiv:2409.01156. Retrieved from https:\/\/arxiv.org\/abs\/2409.01156"},{"key":"e_1_3_1_55_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo","author":"Shen Leqi","year":"2024","unstructured":"Leqi Shen, Tao He, Sicheng Zhao, Zhelun Shen, Yuchen Guo, Tianshi Xu, and Guiguang Ding. 2024. X-ReID: Cross-instance transformer for identity-level person re-identification. In Proceedings of the IEEE International Conference on Multimedia and Expo, 1\u20136."},{"key":"e_1_3_1_56_2","first-page":"4832","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Shen Leqi","year":"2024","unstructured":"Leqi Shen, Sicheng Zhao, Yifeng Zhang, Hui Chen, Jundong Zhou, Pengzhang Liu, Yongjun Bao, and Guiguang Ding. 2024. Multi-label learning with block diagonal labels. In Proceedings of the ACM International Conference on Multimedia, 4832\u20134840."},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3648368"},{"key":"e_1_3_1_58_2","first-page":"17138","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Tian Kaibin","year":"2024","unstructured":"Kaibin Tian, Ruixiang Zhao, Zijie Xin, Bangxiang Lan, and Xirong Li. 2024. Holistic features are almost sufficient for text-to-video retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 17138\u201317147."},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3365104"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2022.3226328"},{"key":"e_1_3_1_61_2","first-page":"5696","article-title":"Omnivl: One foundation model for image-language and video-language tasks","volume":"35","author":"Wang Junke","year":"2022","unstructured":"Junke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo, Luowei Zhou, Yucheng Zhao, Yujia Xie, Ce Liu, Yu-Gang Jiang, and Lu Yuan. 2022. Omnivl: One foundation model for image-language and video-language tasks. Advances in Neural Information Processing Systems 35 (2022), 5696\u20135710.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_62_2","first-page":"6598","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang Jinpeng","year":"2023","unstructured":"Jinpeng Wang, Yixiao Ge, Rui Yan, Yuying Ge, Kevin Qinghong Lin, Satoshi Tsutsui, Xudong Lin, Guanyu Cai, Jianping Wu, Ying Shan, et al. 2023. All in one: Exploring unified video-language pre-Training. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6598\u20136608."},{"key":"e_1_3_1_63_2","first-page":"2816","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Wang Ziyang","year":"2023","unstructured":"Ziyang Wang, Yi-Lin Sung, Feng Cheng, Gedas Bertasius, and Mohit Bansal. 2023. Unified coarse-to-fine alignment for video-text retrieval. In Proceedings of the International Conference on Computer Vision, 2816\u20132827."},{"key":"e_1_3_1_64_2","first-page":"10941","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wei Xi","year":"2020","unstructured":"Xi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang, and Feng Wu. 2020. Multi-modality cross attention network for image and sentence matching. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 10941\u201310950."},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1049\/ipr2.12265"},{"key":"e_1_3_1_66_2","first-page":"10704","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wu Wenhao","year":"2023","unstructured":"Wenhao Wu, Haipeng Luo, Bo Fang, Jingdong Wang, and Wanli Ouyang. 2023. Cap4Video: What can auxiliary captions do for Text-Video retrieval? In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 10704\u201310713."},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.571"},{"key":"e_1_3_1_68_2","first-page":"2048","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In Proceedings of the International Conference on Learning Representations, 2048\u20132057."},{"key":"e_1_3_1_69_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Xue Hongwei","year":"2022","unstructured":"Hongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu, Ruihua Song, Houqiang Li, and Jiebo Luo. 2022. CLIP-ViP: Adapting pre-trained Image-Text model to video-language alignment. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=GNjzMAgawq"},{"key":"e_1_3_1_70_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Yang Taojiannan","year":"2022","unstructured":"Taojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang, Chen Chen, and Mu Li. 2022. AIM: Adapting image models for efficient video action recognition. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=CIoSZ_HKHS7"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/3478025"},{"key":"e_1_3_1_72_2","first-page":"374","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Zhang Bowen","year":"2018","unstructured":"Bowen Zhang, Hexiang Hu, and Fei Sha. 2018. Cross-modal and hierarchical modeling of video and text. In Proceedings of the European Conference on Computer Vision, 374\u2013390."},{"issue":"12","key":"e_1_3_1_73_2","doi-asserted-by":"crossref","first-page":"9780","DOI":"10.1109\/TPAMI.2024.3432099","article-title":"Inductive state-relabeling adversarial active learning with heuristic clique rescaling","volume":"46","author":"Zhang Beichen","year":"2024","unstructured":"Beichen Zhang, Liang Li, Shuhui Wang, Shaofei Cai, Zheng-Jun Zha, Qi Tian, and Qingming Huang. 2024. Inductive state-relabeling adversarial active learning with heuristic clique rescaling. IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 12 (2024), 9780\u20139796.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_74_2","first-page":"4778","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Zhang Haonan","year":"2023","unstructured":"Haonan Zhang, Lianli Gao, Pengpeng Zeng, Alan Hanjalic, and Heng Tao Shen. 2023. Depth-aware sparse transformer for video-language learning. In Proceedings of the ACM International Conference on Multimedia, 4778\u20134787."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3715137","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,7]],"date-time":"2025-11-07T15:09:39Z","timestamp":1762528179000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715137"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,7]]},"references-count":73,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3715137"],"URL":"https:\/\/doi.org\/10.1145\/3715137","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,7]]},"assertion":[{"value":"2024-06-24","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-08","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}