{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T00:15:16Z","timestamp":1783037716011,"version":"3.54.6"},"reference-count":88,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2024,1,11]],"date-time":"2024-01-11T00:00:00Z","timestamp":1704931200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,4,30]]},"abstract":"<jats:p>This paper addresses the temporal sentence grounding (TSG). Although existing methods have made decent achievements in this task, they not only severely rely on abundant video-query paired data for training, but also easily fail into the dataset distribution bias. To alleviate these limitations, we introduce a novel Equivariant Consistency Regulation Learning (ECRL) framework to learn more discriminative query-related frame-wise representations for each video, in a self-supervised manner. Our motivation comes from that the temporal boundary of the query-guided activity should be consistently predicted under various video-level transformations. Concretely, we first design a series of spatio-temporal augmentations on both foreground and background video segments to generate a set of synthetic video samples. In particular, we devise a self-refine module to enhance the completeness and smoothness of the augmented video. Then, we present a novel self-supervised consistency loss (SSCL) applied on the original and augmented videos to capture their invariant query-related semantic by minimizing the KL-divergence between the sequence similarity of two videos and a prior Gaussian distribution of timestamp distance. At last, a shared grounding head is introduced to predict the transform-equivariant query-guided segment boundaries for both the original and augmented videos. Extensive experiments on three challenging datasets (ActivityNet, TACoS, and Charades-STA) demonstrate both effectiveness and efficiency of our proposed ECRL framework.<\/jats:p>","DOI":"10.1145\/3634749","type":"journal-article","created":{"date-parts":[[2023,11,27]],"date-time":"2023-11-27T16:02:17Z","timestamp":1701100937000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Transform-Equivariant Consistency Learning for Temporal Sentence Grounding"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8179-4508","authenticated-orcid":false,"given":"Daizong","family":"Liu","sequence":"first","affiliation":[{"name":"Peking University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4907-3978","authenticated-orcid":false,"given":"Xiaoye","family":"Qu","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5244-3274","authenticated-orcid":false,"given":"Jianfeng","family":"Dong","sequence":"additional","affiliation":[{"name":"Zhejiang Gongshang University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8629-4622","authenticated-orcid":false,"given":"Pan","family":"Zhou","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5438-1468","authenticated-orcid":false,"given":"Zichuan","family":"Xu","sequence":"additional","affiliation":[{"name":"Dalian University of Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7591-5315","authenticated-orcid":false,"given":"Haozhao","family":"Wang","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7232-2330","authenticated-orcid":false,"given":"Xing","family":"Di","sequence":"additional","affiliation":[{"name":"Protagolabs Inc., USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0927-1259","authenticated-orcid":false,"given":"Weining","family":"Lu","sequence":"additional","affiliation":[{"name":"Tsinghua University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9830-0081","authenticated-orcid":false,"given":"Yu","family":"Cheng","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,1,11]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.618"},{"key":"e_1_3_2_3_2","first-page":"9922","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Benaim Sagie","year":"2020","unstructured":"Sagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri, William T. Freeman, Michael Rubinstein, Michal Irani, and Tali Dekel. 2020. SpeedNet: Learning the speediness in videos. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 9922\u20139931."},{"key":"e_1_3_2_4_2","first-page":"9810","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Cao Meng","year":"2021","unstructured":"Meng Cao, Long Chen, Mike Zheng Shou, Can Zhang, and Yuexian Zou. 2021. On pursuit of designing multi-modal transformer for video grounding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 9810\u20139823."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1015"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6627"},{"key":"e_1_3_2_8_2","first-page":"333","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Chen Shaoxiang","year":"2020","unstructured":"Shaoxiang Chen, Wenhao Jiang, Wei Liu, and Yu-Gang Jiang. 2020. Learning modality interaction for temporal sentence localization and event captioning in videos. In Proceedings of the European Conference on Computer Vision (ECCV). Springer, 333\u2013351."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018199"},{"key":"e_1_3_2_10_2","first-page":"1597","volume-title":"International Conference on Machine Learning","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning. PMLR, 1597\u20131607."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298981"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.167"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2984064"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2832602"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3059295"},{"key":"e_1_3_2_16_2","article-title":"Hierarchical contrast for unsupervised skeleton-based action representation learning","author":"Dong Jianfeng","year":"2022","unstructured":"Jianfeng Dong, Shengkai Sun, Zhonglin Liu, Shujie Chen, Baolong Liu, and Xun Wang. 2022. Hierarchical contrast for unsupervised skeleton-based action representation learning. arXiv preprint arXiv:2212.02082 (2022).","journal-title":"arXiv preprint arXiv:2212.02082"},{"key":"e_1_3_2_17_2","first-page":"3299","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Feichtenhofer Christoph","year":"2021","unstructured":"Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, and Kaiming He. 2021. A large-scale study on unsupervised spatiotemporal representation learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3299\u20133309."},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.563"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33016391"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2019.00032"},{"key":"e_1_3_2_21_2","article-title":"Unsupervised representation learning by predicting image rotations","author":"Gidaris Spyros","year":"2018","unstructured":"Spyros Gidaris, Praveer Singh, and Nikos Komodakis. 2018. Unsupervised representation learning by predicting image rotations. arXiv (2018).","journal-title":"arXiv"},{"key":"e_1_3_2_22_2","first-page":"487","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","volume":"3","author":"Goldberger Jacob","year":"2003","unstructured":"Jacob Goldberger, Shiri Gordon, and Hayit Greenspan. 2003. An efficient image similarity measure based on approximations of KL-divergence between two Gaussian mixtures. In Proceedings of the IEEE International Conference on Computer Vision, Vol. 3. 487\u2013493."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2019.00186"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3073867"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3090521"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01216-8_31"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.83"},{"key":"e_1_3_2_28_2","first-page":"3195","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Kuang Haofei","year":"2021","unstructured":"Haofei Kuang, Yi Zhu, Zhi Zhang, Xinyu Li, Joseph Tighe, S\u00f6ren Schwertfeger, Cyrill Stachniss, and Mu Li. 2021. Video contrastive learning with global context. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 3195\u20133204."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3565573"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3532626"},{"key":"e_1_3_2_31_2","first-page":"9972","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Le Thao Minh","year":"2020","unstructured":"Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran. 2020. Hierarchical conditional relation networks for video question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 9972\u20139981."},{"key":"e_1_3_2_32_2","article-title":"Exploiting prompt caption for video grounding","author":"Li Hongxiang","year":"2023","unstructured":"Hongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li, Zhihong Zhu, and Yuexian Zou. 2023. Exploiting prompt caption for video grounding. arXiv preprint arXiv:2301.05997 (2023).","journal-title":"arXiv preprint arXiv:2301.05997"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00399"},{"key":"e_1_3_2_34_2","first-page":"988","volume-title":"ACM MM","author":"Lin Tianwei","year":"2017","unstructured":"Tianwei Lin, Xu Zhao, and Zheng Shou. 2017. Single shot temporal action detection. In ACM MM. 988\u2013996."},{"key":"e_1_3_2_35_2","first-page":"1665","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"36","author":"Liu Daizong","year":"2022","unstructured":"Daizong Liu, Xiaoye Qu, Xing Di, Yu Cheng, Zichuan Xu, and Pan Zhou. 2022. Memory-guided semantic learning network for temporal sentence grounding. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 1665\u20131673."},{"key":"e_1_3_2_36_2","first-page":"9292","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Liu Daizong","year":"2021","unstructured":"Daizong Liu, Xiaoye Qu, Jianfeng Dong, and Pan Zhou. 2021. Adaptive proposal generation network for temporal sentence localization in videos. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP). 9292\u20139301."},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01108"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3414026"},{"key":"e_1_3_2_39_2","first-page":"1683","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"36","author":"Liu Daizong","year":"2022","unstructured":"Daizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di, Kai Zou, Yu Cheng, Zichuan Xu, and Pan Zhou. 2022. Unsupervised temporal video grounding with deep semantic clustering. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 1683\u20131691."},{"key":"e_1_3_2_40_2","first-page":"9302","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Liu Daizong","year":"2021","unstructured":"Daizong Liu, Xiaoye Qu, and Pan Zhou. 2021. Progressively guide to attend: An iterative alignment framework for temporal sentence grounding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 9302\u20139311."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210003"},{"key":"e_1_3_2_42_2","first-page":"289","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Lu Jiasen","year":"2016","unstructured":"Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016. Hierarchical question-image co-attention for visual question answering. In Advances in Neural Information Processing Systems (NIPS). 289\u2013297."},{"key":"e_1_3_2_43_2","first-page":"1","volume-title":"ACM Multimedia Asia","author":"Ma Ziyang","year":"2021","unstructured":"Ziyang Ma, Xianjing Han, Xuemeng Song, Yiran Cui, and Liqiang Nie. 2021. Hierarchical deep residual reasoning for temporal moment localization. In ACM Multimedia Asia. 1\u20137."},{"key":"e_1_3_2_44_2","doi-asserted-by":"crossref","first-page":"527","DOI":"10.1007\/978-3-319-46448-0_32","volume-title":"Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11\u201314, 2016, Proceedings, Part I 14","author":"Misra Ishan","year":"2016","unstructured":"Ishan Misra, C. Lawrence Zitnick, and Martial Hebert. 2016. Shuffle and learn: Unsupervised learning using temporal order verification. In Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11\u201314, 2016, Proceedings, Part I 14. Springer, 527\u2013544."},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01082"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00279"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.628"},{"key":"e_1_3_2_48_2","article-title":"Uncovering hidden challenges in query-based video moment retrieval","author":"Otani Mayu","year":"2020","unstructured":"Mayu Otani, Yuta Nakashima, Esa Rahtu, and Janne Heikkil\u00e4. 2020. Uncovering hidden challenges in query-based video moment retrieval. arXiv (2020).","journal-title":"arXiv"},{"key":"e_1_3_2_49_2","first-page":"1532","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Pennington Jeffrey","year":"2014","unstructured":"Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 1532\u20131543."},{"key":"e_1_3_2_50_2","first-page":"5152","volume-title":"International Conference on Machine Learning","author":"Piergiovanni A. J.","year":"2019","unstructured":"A. J. Piergiovanni and Michael Ryoo. 2019. Temporal Gaussian mixture layer for videos. In International Conference on Machine Learning. PMLR, 5152\u20135161."},{"key":"e_1_3_2_51_2","doi-asserted-by":"crossref","first-page":"625","DOI":"10.1109\/TMM.2021.3056892","article-title":"Spatial-temporal action localization with hierarchical self-attention","volume":"24","author":"Pramono Rizard Renanda Adhi","year":"2021","unstructured":"Rizard Renanda Adhi Pramono, Yie-Tarng Chen, and Wen-Hsien Fang. 2021. Spatial-temporal action localization with hierarchical self-attention. IEEE Transactions on Multimedia 24 (2021), 625\u2013639.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00689"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12272"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00207"},{"key":"e_1_3_2_55_2","first-page":"2464","volume-title":"IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Rodriguez Cristian","year":"2020","unstructured":"Cristian Rodriguez, Edison Marrese-Taylor, Fatemeh Sadat Saleh, Hongdong Li, and Stephen Gould. 2020. Proposal-free temporal moment localization of a natural-language query in video using guided attention. In IEEE Winter Conference on Applications of Computer Vision (WACV). 2464\u20132473."},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/78.650093"},{"key":"e_1_3_2_57_2","first-page":"1049","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Shou Zheng","year":"2016","unstructured":"Zheng Shou, Dongang Wang, and Shih-Fu Chang. 2016. Temporal action localization in untrimmed videos via multi-stage CNNs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 1049\u20131058."},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_31"},{"key":"e_1_3_2_59_2","first-page":"5179","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Song Yale","year":"2015","unstructured":"Yale Song, Jordi Vallmitjana, Amanda Stent, and Alejandro Jaimes. 2015. TVSum: Summarizing web videos using titles. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5179\u20135187."},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3050067"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_2_62_2","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems (NIPS)."},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6897"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i3.20163"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i4.16406"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2921539"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33019062"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01017"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3016486"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401151"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3140611"},{"key":"e_1_3_2_72_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Yao Ting","year":"2021","unstructured":"Ting Yao, Yiheng Zhang, Zhaofan Qiu, Yingwei Pan, and Tao Mei. 2021. SeCo: Exploring sequence supervision for unsupervised representation learning. In Proceedings of the AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1145\/2534409"},{"key":"e_1_3_2_74_2","volume-title":"Human-centric Multimedia Analysis","author":"Yuan Yitian","year":"2021","unstructured":"Yitian Yuan, Xiaohan Lan, Xin Wang, Long Chen, Zhi Wang, and Wenwu Zhu. 2021. A closer look at temporal sentence grounding in videos: Dataset and metric. In Human-centric Multimedia Analysis."},{"key":"e_1_3_2_75_2","first-page":"534","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Yuan Yitian","year":"2019","unstructured":"Yitian Yuan, Lin Ma, Jingwen Wang, Wei Liu, and Wenwu Zhu. 2019. Semantic conditioned dynamic modulation for temporal sentence grounding in videos. In Advances in Neural Information Processing Systems (NIPS). 534\u2013544."},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33019159"},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01030"},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.1145\/3478025"},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3073235"},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00134"},{"key":"e_1_3_2_81_2","volume-title":"The Annual Meeting of the Association for Computational Linguistics","author":"Zhang Hao","year":"2020","unstructured":"Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou. 2020. Span-based localizing network for natural language video localization. In The Annual Meeting of the Association for Computational Linguistics."},{"key":"e_1_3_2_82_2","article-title":"Towards debiasing temporal sentence grounding in video","author":"Zhang Hao","year":"2021","unstructured":"Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou. 2021. Towards debiasing temporal sentence grounding in video. arXiv (2021).","journal-title":"arXiv"},{"key":"e_1_3_2_83_2","doi-asserted-by":"crossref","first-page":"649","DOI":"10.1007\/978-3-319-46487-9_40","volume-title":"Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part III 14","author":"Zhang Richard","year":"2016","unstructured":"Richard Zhang, Phillip Isola, and Alexei A. Efros. 2016. Colorful image colorization. In Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part III 14. Springer, 649\u2013666."},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6984"},{"key":"e_1_3_2_85_2","first-page":"9235","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Zhang Yaqing","year":"2019","unstructured":"Yaqing Zhang, Xi Li, and Zhongfei Zhang. 2019. Learning a key-value memory co-attention matching network for person re-identification. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 9235\u20139242."},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3113791"},{"key":"e_1_3_2_87_2","doi-asserted-by":"publisher","DOI":"10.1145\/3331184.3331235"},{"key":"e_1_3_2_88_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.317"},{"key":"e_1_3_2_89_2","doi-asserted-by":"publisher","DOI":"10.1145\/3544493"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3634749","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3634749","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T23:44:07Z","timestamp":1750290247000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3634749"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,11]]},"references-count":88,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,4,30]]}},"alternative-id":["10.1145\/3634749"],"URL":"https:\/\/doi.org\/10.1145\/3634749","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,11]]},"assertion":[{"value":"2023-05-06","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-23","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-11","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}