{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T16:17:54Z","timestamp":1782836274913,"version":"3.54.5"},"reference-count":73,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,12,9]],"date-time":"2023-12-09T00:00:00Z","timestamp":1702080000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272156"],"award-info":[{"award-number":["62272156"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Shanghai Pujiang Program","award":["20PJ1401900"],"award-info":[{"award-number":["20PJ1401900"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,3,31]]},"abstract":"<jats:p>\n            The task of text-video retrieval aims to understand the correspondence between language and vision and has gained increasing attention in recent years. Recent works have demonstrated the superiority of local spatio-temporal relation learning with graph-based models. However, most existing graph-based models are handcrafted and depend heavily on expert knowledge and empirical feedback, which may be unable to mine the high-level fine-grained visual relations effectively. These limitations result in their inability to distinguish videos with the same visual components but different relations. To solve this problem, we propose a novel cross-modal retrieval framework, Bi-Branch Complementary Network (BiC-Net), which modifies Transformer architecture to effectively bridge text-video modalities in a complementary manner via combining local spatio-temporal relation and global temporal information. Specifically, local video representations are encoded using multiple Transformer blocks and additional residual blocks to learn fine-grained spatio-temporal relations and long-term temporal dependency, calling the module a Fine-grained Spatio-temporal Transformer (FST). Global video representations are encoded using a multi-layer Transformer block to learn global temporal features. Finally, we align the spatio-temporal relation and global temporal features with the text feature on two embedding spaces for cross-modal text-video retrieval. Extensive experiments are conducted on MSR-VTT, MSVD, and YouCook2 datasets. The results demonstrate the effectiveness of our proposed model. Our code is public at\n            <jats:italic>\n              <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/lionel-hing\/BiC-Net\">https:\/\/github.com\/lionel-hing\/BiC-Net<\/jats:ext-link>\n            <\/jats:italic>\n            .\n          <\/jats:p>","DOI":"10.1145\/3627103","type":"journal-article","created":{"date-parts":[[2023,10,13]],"date-time":"2023-10-13T15:26:17Z","timestamp":1697210777000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["BiC-Net: Learning Efficient Spatio-temporal Relation for Text-Video Retrieval"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0654-6026","authenticated-orcid":false,"given":"Ning","family":"Han","sequence":"first","affiliation":[{"name":"School of Computer Science, Xiangtan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1908-1157","authenticated-orcid":false,"given":"Yawen","family":"Zeng","sequence":"additional","affiliation":[{"name":"Bytedance AI Lab, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-2468-6560","authenticated-orcid":false,"given":"Chuhao","family":"Shi","sequence":"additional","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0625-0804","authenticated-orcid":false,"given":"Guangyi","family":"Xiao","sequence":"additional","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5902-1824","authenticated-orcid":false,"given":"Hao","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3148-264X","authenticated-orcid":false,"given":"Jingjing","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Computer Sciences, Fudan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,12,9]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.572"},{"key":"e_1_3_2_4_2","unstructured":"Jimmy Lei Ba Jamie Ryan Kiros and Geoffrey E. Hinton. 2016. Layer normalization. arXiv:1607.06450 (2016)."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00175"},{"key":"e_1_3_2_6_2","first-page":"813","volume-title":"Proceedings of the 38th International Conference on Machine Learning","author":"Bertasius Gedas","year":"2021","unstructured":"Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding. In Proceedings of the 38th International Conference on Machine Learning. 813\u2013824."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.5555\/2002472.2002497"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2964315"},{"key":"e_1_3_2_10_2","first-page":"10635","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Shizhe","year":"2020","unstructured":"Shizhe Chen, Yida Zhao, Qin Jin, and Qi Wu. 2020. Fine-grained video-text retrieval with hierarchical graph reasoning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10635\u201310644."},{"key":"e_1_3_2_11_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171\u20134186."},{"key":"e_1_3_2_12_2","doi-asserted-by":"crossref","unstructured":"Jianfeng Dong Xirong Li and Cees G. M. Snoek. 2018. Predicting visual features from text for image and video caption retrieval. IEEE Transactions on Multimedia 20 12 (2018) 3377\u20133388.","DOI":"10.1109\/TMM.2018.2832602"},{"key":"e_1_3_2_13_2","first-page":"9346","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Dong Jianfeng","year":"2019","unstructured":"Jianfeng Dong, Xirong Li, Chaoxi Xu, Shouling Ji, Yuan He, Gang Yang, and Xun Wang. 2019. Dual encoding for zero-example video retrieval. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 9346\u20139355."},{"key":"e_1_3_2_14_2","unstructured":"Jianfeng Dong Xirong Li Chaoxi Xu Xun Yang Gang Yang Xun Wang and Meng Wang. 2021. Dual encoding for video retrieval by text. IEEE Transactions on Pattern Analysis and Machine Intelligence 44 8 (2021) 4065\u20134080."},{"key":"e_1_3_2_15_2","unstructured":"Han Fang Pengfei Xiong Luhui Xu and Yu Chen. 2021. Clip2video: Mastering video-text retrieval via image clip. arXiv:2106.11097 (2021)."},{"key":"e_1_3_2_16_2","first-page":"1005","volume-title":"International Joint Conference on Artificial Intelligence","author":"Feng Zerun","year":"2020","unstructured":"Zerun Feng, Zhimin Zeng, Caili Guo, and Zheng Li. 2020. Exploiting visual semantic reasoning for video-text retrieval. In International Joint Conference on Artificial Intelligence. 1005\u20131011."},{"key":"e_1_3_2_17_2","first-page":"214","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Gabeur Valentin","year":"2020","unstructured":"Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid. 2020. Multi-modal transformer for video retrieval. In Proceedings of the European Conference on Computer Vision. 214\u2013229."},{"key":"e_1_3_2_18_2","unstructured":"Simon Ging Mohammadreza Zolfaghari Hamed Pirsiavash and Thomas Brox. 2020. COOT: Cooperative hierarchical transformer for video-text representation learning. Advances in Neural Information Processing Systems 33 (2020) 22605\u201322618."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3372278.3390709"},{"key":"e_1_3_2_20_2","first-page":"3826","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Han Ning","year":"2021","unstructured":"Ning Han, Jingjing Chen, Guangyi Xiao, Hao Zhang, Yawen Zeng, and Hao Chen. 2021. Fine-grained cross-modal alignment network for text-video retrieval. In Proceedings of the ACM International Conference on Multimedia. 3826\u20133834."},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","unstructured":"Ning Han Jingjing Chen Hao Zhang Huanwen Wang and Hao Chen. 2022. Adversarial multi-grained embedding network for cross-modal text-video retrieval. ACM Transactions on Multimedia Computing Communications and Applications 18 2 (2022) 1\u201323.","DOI":"10.1145\/3483381"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00100"},{"key":"e_1_3_2_23_2","unstructured":"Dan Hendrycks and Kevin Gimpel. 2016. Gaussian error linear units (GELUs). arXiv:1606.08415 (2016)."},{"key":"e_1_3_2_24_2","first-page":"321","volume-title":"Breakthroughs in statistics: methodology and distribution","author":"Hotelling Harold","year":"1936","unstructured":"Harold Hotelling. 1936. Relations between two sets of variates. In Breakthroughs in statistics: methodology and distribution. 321\u2013377."},{"key":"e_1_3_2_25_2","unstructured":"Will Kay Joao Carreira Karen Simonyan Brian Zhang Chloe Hillier Sudheendra Vijayanarasimhan Fabio Viola Tim Green Trevor Back Paul Natsev et\u00a0al. 2017. The kinetics human action video dataset. arXiv:1705.06950 (2017)."},{"key":"e_1_3_2_26_2","first-page":"1","volume-title":"International Conference on Learning Representations","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In International Conference on Learning Representations. 1\u201315."},{"key":"e_1_3_2_27_2","first-page":"1","volume-title":"International Conference on Learning Representations","author":"Kipf Thomas N.","year":"2017","unstructured":"Thomas N. Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations. 1\u201314."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00725"},{"key":"e_1_3_2_29_2","doi-asserted-by":"crossref","unstructured":"Xirong Li Fangming Zhou Chaoxi Xu Jiaqi Ji and Gang Yang. 2020. SEA: Sentence encoder assembly for video retrieval by textual queries. IEEE Transactions on Multimedia 23 (2020) 4351\u20134362.","DOI":"10.1109\/TMM.2020.3042067"},{"key":"e_1_3_2_30_2","first-page":"11915","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Liu Song","year":"2021","unstructured":"Song Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen, Wenkui Ding, and Zhongyuan Wang. 2021. HiT: Hierarchical Transformer with momentum contrast for video-text retrieval. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 11915\u201311925."},{"key":"e_1_3_2_31_2","unstructured":"Xuejing Liu Liang Li Shuhui Wang Zheng-Jun Zha Zechao Li Qi Tian and Qingming Huang. 2022. Entity-enhanced adaptive reconstruction network for weakly supervised referring expression grounding. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 3 (2022) 3003\u20133018."},{"key":"e_1_3_2_32_2","first-page":"279","volume-title":"British Machine Vision Conference","author":"Liu Yang","year":"2019","unstructured":"Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman. 2019. Use what you have: Video retrieval using representations from collaborative experts. In British Machine Vision Conference. 279."},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","unstructured":"Wei Lu Desheng Li Liqiang Nie Peiguang Jing and Yuting Su. 2021. Learning dual low-rank representation for multi-label micro-video classification. IEEE Transactions on Multimedia 25 (2021) 77\u201389.","DOI":"10.1109\/TMM.2021.3121567"},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","unstructured":"Huaishao Luo Lei Ji Ming Zhong Yang Chen Wen Lei Nan Duan and Tianrui Li. 2022. CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning. Neurocomputing 508 (2022) 293\u2013304.","DOI":"10.1016\/j.neucom.2022.07.028"},{"key":"e_1_3_2_35_2","first-page":"9876","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Miech Antoine","year":"2020","unstructured":"Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. 2020. End-to-end learning of visual representations from uncurated instructional videos. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 9876\u20139886."},{"key":"e_1_3_2_36_2","unstructured":"Antoine Miech Ivan Laptev and Josef Sivic. 2018. Learning a text-video embedding from incomplete and heterogeneous data. arXiv:1804.02516 (2018)."},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00272"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3206025.3206064"},{"key":"e_1_3_2_39_2","first-page":"1","volume-title":"International Conference on Learning Representations","author":"Patrick Mandela","year":"2021","unstructured":"Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander G. Hauptmann, Joao F. Henriques, and Andrea Vedaldi. 2021. Support-set bottlenecks for video-text representation learning. In International Conference on Learning Representations. 1\u201318."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2007.383266"},{"key":"e_1_3_2_41_2","doi-asserted-by":"crossref","unstructured":"Mengshi Qi Jie Qin Yi Yang Yunhong Wang and Jiebo Luo. 2021. Semantics-aware spatial-temporal binaries for cross-modal video retrieval. IEEE Transactions on Image Processing 30 (2021) 2989\u20133004.","DOI":"10.1109\/TIP.2020.3048680"},{"key":"e_1_3_2_42_2","first-page":"84","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Qian Xufeng","year":"2019","unstructured":"Xufeng Qian, Yueting Zhuang, Yimeng Li, Shaoning Xiao, Shiliang Pu, and Jun Xiao. 2019. Video relation detection with spatio-temporal graph. In Proceedings of the ACM International Conference on Multimedia. 84\u201393."},{"key":"e_1_3_2_43_2","first-page":"8748","volume-title":"International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et\u00a0al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning. 8748\u20138763."},{"key":"e_1_3_2_44_2","unstructured":"Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems 28 (2015)."},{"key":"e_1_3_2_45_2","first-page":"1584","volume-title":"Annual Conference of the International Speech Communication Association","author":"Rouditchenko Andrew","year":"2021","unstructured":"Andrew Rouditchenko, Angie Boggust, David Harwath, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Rogerio Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, et\u00a0al. 2021. AVLnet: Learning audio-visual language representations from instructional videos. In Annual Conference of the International Speech Communication Association. 1584\u20131588."},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","unstructured":"Olga Russakovsky Jia Deng Hao Su Jonathan Krause Sanjeev Satheesh Sean Ma Zhiheng Huang Andrej Karpathy Aditya Khosla Michael Bernstein et\u00a0al. 2015. ImageNet large scale visual recognition challenge. International Journal of Computer Vision 115 3 (2015) 211\u2013252.","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_2_47_2","doi-asserted-by":"crossref","unstructured":"Xue Song Jingjing Chen Zuxuan Wu and Yu-Gang Jiang. 2021. Spatial-temporal graphs for cross-modal text2video retrieval. IEEE Transactions on Multimedia 24 (2021) 2914\u20132923.","DOI":"10.1109\/TMM.2021.3090595"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413764"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"e_1_3_2_50_2","first-page":"6000","volume-title":"Advances on Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances on Neural Information Processing Systems. 6000\u20136010."},{"key":"e_1_3_2_51_2","first-page":"4534","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Venugopalan Subhashini","year":"2015","unstructured":"Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko. 2015. Sequence to sequence \u2014 video to text. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 4534\u20134542."},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413975"},{"key":"e_1_3_2_53_2","doi-asserted-by":"crossref","unstructured":"Hao Wang Zheng-Jun Zha Liang Li Xuejin Chen and Jiebo Luo. 2023. Semantic and relation modulation for audio-visual event localization. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 6 (2023) 7711\u20137725.","DOI":"10.1109\/TPAMI.2022.3226328"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00695"},{"key":"e_1_3_2_55_2","doi-asserted-by":"crossref","unstructured":"Wei Wang Junyu Gao Xiaoshan Yang and Changsheng Xu. 2022. Many hands make light work: Transferring knowledge from auxiliary tasks for video-text retrieval. IEEE Transactions on Multimedia 25 (2022) 2661\u20132674.","DOI":"10.1109\/TMM.2022.3149716"},{"key":"e_1_3_2_56_2","first-page":"399","volume-title":"European Conference on Computer Vision","author":"Wang Xiaolong","year":"2018","unstructured":"Xiaolong Wang and Abhinav Gupta. 2018. Videos as space-time region graphs. In European Conference on Computer Vision. 399\u2013417."},{"key":"e_1_3_2_57_2","first-page":"5079","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang Xiaohan","year":"2021","unstructured":"Xiaohan Wang, Linchao Zhu, and Yi Yang. 2021. T2VLAD: Global-local sequence alignment for text-video retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5079\u20135088."},{"key":"e_1_3_2_58_2","first-page":"450","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Wray Michael","year":"2019","unstructured":"Michael Wray, Diane Larlus, Gabriela Csurka, and Dima Damen. 2019. Fine-grained action retrieval through multiple parts-of-speech embeddings. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 450\u2013459."},{"key":"e_1_3_2_59_2","first-page":"9964","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wu Jianchao","year":"2019","unstructured":"Jianchao Wu, Limin Wang, Li Wang, Jie Guo, and Gangshan Wu. 2019. Learning actor relation graphs for group activity recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 9964\u20139974."},{"key":"e_1_3_2_60_2","first-page":"3518","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"Wu Peng","year":"2021","unstructured":"Peng Wu, Xiangteng He, Mingqian Tang, Yiliang Lv, and Jing Liu. 2021. HANet: Hierarchical alignment networks for video-text retrieval. In Proceedings of the 29th ACM International Conference on Multimedia. 3518\u20133527."},{"key":"e_1_3_2_61_2","first-page":"447","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Xiao Junbin","year":"2020","unstructured":"Junbin Xiao, Xindi Shang, Xun Yang, Sheng Tang, and Tat-Seng Chua. 2020. Visual relation grounding in videos. In Proceedings of the European Conference on Computer Vision. 447\u2013464."},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.571"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401151"},{"key":"e_1_3_2_64_2","doi-asserted-by":"crossref","unstructured":"Xun Yang Meng Wang Richang Hong Qi Tian and Yong Rui. 2017. Enhancing person re-identification in a self-trained subspace. ACM Transactions on Multimedia Computing Communications and Applications 13 3 (2017) 1\u201323.","DOI":"10.1145\/3089249"},{"key":"e_1_3_2_65_2","first-page":"471","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Yu Youngjae","year":"2018","unstructured":"Youngjae Yu, Jongseok Kim, and Gunhee Kim. 2018. A joint sequence fusion model for video question answering and retrieval. In Proceedings of the European Conference on Computer Vision. 471\u2013487."},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","unstructured":"Yawen Zeng Ning Han Keyu Pan and Qin Jin. 2023. Temporally language grounding with multi-modal multi-prompt tuning. IEEE Transactions on Multimedia (2023) 1\u201312.","DOI":"10.1109\/TMM.2023.3310282"},{"key":"e_1_3_2_67_2","first-page":"3376","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Zeng Yawen","year":"2023","unstructured":"Yawen Zeng, Qin Jin, Tengfei Bao, and Wenfeng Li. 2023. Multi-modal knowledge hypergraph for diverse image retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence. 3376\u20133383."},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3592054"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547908"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475272"},{"key":"e_1_3_2_71_2","doi-asserted-by":"crossref","unstructured":"Yanchao Zhang Weiqing Min Liqiang Nie and Shuqiang Jiang. 2020. Hybrid-attention enhanced two-stream fusion network for video venue prediction. IEEE Transactions on Multimedia 23 (2020) 2917\u20132929.","DOI":"10.1109\/TMM.2020.3019714"},{"key":"e_1_3_2_72_2","first-page":"1","volume-title":"2020 IEEE International Conference on Multimedia and Expo","author":"Zhao Rui","year":"2020","unstructured":"Rui Zhao, Kecheng Zheng, and Zheng-jun Zha. 2020. Stacked convolutional deep encoding network for video-text retrieval. In 2020 IEEE International Conference on Multimedia and Expo. 1\u20136."},{"key":"e_1_3_2_73_2","first-page":"7590","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Zhou Luowei","year":"2018","unstructured":"Luowei Zhou, Chenliang Xu, and Jason J. Corso. 2018. Towards automatic learning of procedures from web instructional videos. In Proceedings of the AAAI Conference on Artificial Intelligence. 7590\u20137598."},{"key":"e_1_3_2_74_2","doi-asserted-by":"crossref","unstructured":"Xiaofeng Zou Kenli Li and Cen Chen. 2020. Multilevel attention based U-shape graph neural network for point clouds learning. IEEE Transactions on Industrial Informatics 18 1 (2020) 448\u2013456.","DOI":"10.1109\/TII.2020.3046627"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3627103","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3627103","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T23:57:04Z","timestamp":1750291024000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3627103"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,9]]},"references-count":73,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,3,31]]}},"alternative-id":["10.1145\/3627103"],"URL":"https:\/\/doi.org\/10.1145\/3627103","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,9]]},"assertion":[{"value":"2022-09-11","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-24","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-12-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}