{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T04:45:31Z","timestamp":1781585131052,"version":"3.54.5"},"reference-count":80,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,2,6]],"date-time":"2023-02-06T00:00:00Z","timestamp":1675641600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2018YFB1404102"],"award-info":[{"award-number":["2018YFB1404102"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["61972448, 61902347, 61976188, 62002323"],"award-info":[{"award-number":["61972448, 61902347, 61976188, 62002323"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Public Welfare Technology Research Project of Zhejiang Province","award":["LGF21F020010"],"award-info":[{"award-number":["LGF21F020010"]}]},{"name":"Fundamental Research Funds for the Provincial Universities of Zhejiang"},{"name":"Open Projects Program of the National Laboratory of Pattern Recognition"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,5,31]]},"abstract":"<jats:p>This article targets the task of language-based video moment localization. The language-based setting of this task allows for an open set of target activities, resulting in a large variation of the temporal lengths of video moments. Most existing methods prefer to first sample sufficient candidate moments with various temporal lengths, then match them with the given query to determine the target moment. However, candidate moments generated with a fixed temporal granularity may be suboptimal to handle the large variation in moment lengths. To this end, we propose a novel multi-stage Progressive Localization Network (PLN) that progressively localizes the target moment in a coarse-to-fine manner. Specifically, each stage of PLN has a localization branch and focuses on candidate moments that are generated with a specific temporal granularity. The temporal granularities of candidate moments are different across the stages. Moreover, we devise a conditional feature manipulation module and an upsampling connection to bridge the multiple localization branches. In this fashion, the later stages are able to absorb the previously learned information, thus facilitating the more fine-grained localization. Extensive experiments on three public datasets demonstrate the effectiveness of our proposed PLN for language-based moment localization, especially for localizing short moments in long videos.<\/jats:p>","DOI":"10.1145\/3543857","type":"journal-article","created":{"date-parts":[[2022,6,11]],"date-time":"2022-06-11T22:49:22Z","timestamp":1654987762000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":22,"title":["Progressive Localization Networks for Language-Based Moment Localization"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7875-4388","authenticated-orcid":false,"given":"Qi","family":"Zheng","sequence":"first","affiliation":[{"name":"Zhejiang Gongshang University, Hangzhou, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5244-3274","authenticated-orcid":false,"given":"Jianfeng","family":"Dong","sequence":"additional","affiliation":[{"name":"Zhejiang Gongshang University, Hangzhou, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4907-3978","authenticated-orcid":false,"given":"Xiaoye","family":"Qu","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, Hubei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0201-1638","authenticated-orcid":false,"given":"Xun","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, Anhui, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7231-1260","authenticated-orcid":false,"given":"Yabing","family":"Wang","sequence":"additional","affiliation":[{"name":"Zhejiang Gongshang University, Hangzhou, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9295-1660","authenticated-orcid":false,"given":"Pan","family":"Zhou","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, Hubei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7718-5443","authenticated-orcid":false,"given":"Baolong","family":"Liu","sequence":"additional","affiliation":[{"name":"Zhejiang Gongshang University, Hangzhou, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5566-4689","authenticated-orcid":false,"given":"Xun","family":"Wang","sequence":"additional","affiliation":[{"name":"Zhejiang Gongshang University, Hangzhou, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,2,6]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.675"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413841"},{"key":"e_1_3_1_5_2","first-page":"1130","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Chao Yu-Wei","year":"2018","unstructured":"Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A. Ross, Jia Deng, and Rahul Sukthankar. 2018. Rethinking the faster R-CNN architecture for temporal action localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1130\u20131139."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1015"},{"key":"e_1_3_1_7_2","first-page":"10551","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Chen Long","year":"2020","unstructured":"Long Chen, Chujie Lu, Siliang Tang, Jun Xiao, Dong Zhang, Chilie Tan, and Xiaolin Li. 2020. Rethinking the bottom-up framework for query-based video localization. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 10551\u201310558."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2959977"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58548-8_20"},{"key":"e_1_3_1_10_2","first-page":"8199","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Chen Shaoxiang","year":"2019","unstructured":"Shaoxiang Chen and Yu-Gang Jiang. 2019. Semantic proposal for activity localization in videos via sentence query. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 8199\u20138206."},{"key":"e_1_3_1_11_2","first-page":"601","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Chen Shaoxiang","year":"2020","unstructured":"Shaoxiang Chen and Yu-Gang Jiang. 2020. Hierarchical visual-textual graph for temporal activity localization via language. In Proceedings of the European Conference on Computer Vision. 601\u2013618."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/tmm.2018.2832602"},{"key":"e_1_3_1_13_2","article-title":"Dual encoding for video retrieval by text","author":"Dong Jianfeng","year":"2021","unstructured":"Jianfeng Dong, Xirong Li, Chaoxi Xu, Xun Yang, Gang Yang, Xun Wang, and Meng Wang. 2021. Dual encoding for video retrieval by text. IEEE Transactions on Pattern Analysis and Machine Intelligence. Early access, February 15, 2021.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2796248"},{"key":"e_1_3_1_15_2","article-title":"Reading-strategy inspired visual representation learning for text-to-video retrieval","author":"Dong Jianfeng","year":"2022","unstructured":"Jianfeng Dong, Yabing Wang, Xianke Chen, Xiaoye Qu, Xirong Li, Yuan He, and Xun Wang. 2022. Reading-strategy inspired visual representation learning for text-to-video retrieval. IEEE Transactions on Circuits and Systems for Video Technology. Early access, January 23, 2022.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00630"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.563"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00155"},{"key":"e_1_3_1_19_2","article-title":"Learning video moment retrieval without a single annotated video","author":"Gao Junyu","year":"2021","unstructured":"Junyu Gao and Changsheng Xu. 2021. Learning video moment retrieval without a single annotated video. IEEE Transactions on Circuits and Systems for Video Technology 32, 3 (2021), 1646\u20131657.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.392"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2019.00032"},{"key":"e_1_3_1_22_2","first-page":"549.1\u2013549.14","volume-title":"Proceedings of the British Machine Vision Conference","author":"Hahn Meera","year":"2020","unstructured":"Meera Hahn, Asim Kadav, James M. Rehg, and Hans Peter Graf. 2020. Tripping through time: Efficient localization of activities in videos. In Proceedings of the British Machine Vision Conference. 549.1\u2013549.14."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.618"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3073867"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3323873.3325019"},{"key":"e_1_3_1_26_2","first-page":"1","volume-title":"Proceedings of the International Conference for Learning Representations","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the International Conference for Learning Representations. 1\u201315."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.83"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123343"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2962063"},{"key":"e_1_3_1_30_2","first-page":"11539","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Lin Zhijie","year":"2020","unstructured":"Zhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang, and Huasheng Liu. 2020. Weakly-supervised video moment retrieval via semantic completion network. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 11539\u201311546."},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.2965987"},{"key":"e_1_3_1_32_2","first-page":"1841","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"Liu Daizong","year":"2020","unstructured":"Daizong Liu, Xiaoye Qu, Jianfeng Dong, and Pan Zhou. 2020. Reasoning step-by-step: Temporal sentence localization in videos via deep rectification-modulation network. In Proceedings of the 28th International Conference on Computational Linguistics. 1841\u20131851."},{"key":"e_1_3_1_33_2","first-page":"11235","article-title":"Context-aware Biaffine Localizing Network for temporal sentence grounding","author":"Liu Daizong","year":"2021","unstructured":"Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Yu Cheng, Wei Wei, Zichuan Xu, and Yulai Xie. 2021. Context-aware Biaffine Localizing Network for temporal sentence grounding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 11235\u201311244.","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210003"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240549"},{"key":"e_1_3_1_36_2","first-page":"11612","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Liu Qinying","year":"2020","unstructured":"Qinying Liu and Zilei Wang. 2020. Progressive boundary refinement network for temporal action detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 11612\u201311619."},{"issue":"3","key":"e_1_3_1_37_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3441577","article-title":"Single-shot semantic matching network for moment localization in videos","volume":"17","author":"Liu Xinfang","year":"2021","unstructured":"Xinfang Liu, Xiushan Nie, Junya Teng, Li Lian, and Yilong Yin. 2021. Single-shot semantic matching network for moment localization in videos. ACM Transactions on Multimedia Computing, Communications, and Applications 17, 3 (2021), 1\u201314.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_38_2","article-title":"Centerness-aware network for temporal action proposal","author":"Liu Yuan","year":"2021","unstructured":"Yuan Liu, Jingyuan Chen, Xinpeng Chena, Bing Deng, Jianqiang Huang, and Xiansheng Hua. 2021. Centerness-aware network for temporal action proposal. IEEE Transactions on Circuits and Systems for Video Technology 32, 1 (2021), 5\u201316.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_39_2","first-page":"5147","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing","author":"Lu Chujie","year":"2019","unstructured":"Chujie Lu, Long Chen, Chilie Tan, Xiaolin Li, and Jun Xiao. 2019. DEBUG: A dense bottom-up grounding approach for natural language video localization. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing. 5147\u20135156."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/2487268.2487269"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01186"},{"key":"e_1_3_1_42_2","first-page":"10810","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Mun Jonghwan","year":"2020","unstructured":"Jonghwan Mun, Minsu Cho, and Bohyung Han. 2020. Local-global video-text interactions for temporal grounding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 10810\u201310819."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3052086"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3284750"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_1_46_2","first-page":"3942","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"32","author":"Perez Ethan","year":"2018","unstructured":"Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. 2018. FiLM: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 32. 3942\u20133951."},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3414053"},{"key":"e_1_3_1_48_2","article-title":"Yolov3: An incremental improvement","author":"Redmon Joseph","year":"2018","unstructured":"Joseph Redmon and Ali Farhadi. 2018. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767 (2018).","journal-title":"arXiv preprint arXiv:1804.02767"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00207"},{"key":"e_1_3_1_50_2","first-page":"91","volume-title":"Advances in Neural Information Processing Systems","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems. 91\u201399."},{"key":"e_1_3_1_51_2","first-page":"2464","volume-title":"Proceedings of the IEEE Winter Conference on Applications of Computer Vision","author":"Rodriguez Cristian","year":"2020","unstructured":"Cristian Rodriguez, Edison Marrese-Taylor, Fatemeh Sadat Saleh, Hongdong Li, and Stephen Gould. 2020. Proposal-free temporal moment localization of a natural-language query in video using guided attention. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision. 2464\u20132473."},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33718-5_11"},{"key":"e_1_3_1_53_2","first-page":"5734","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Shou Zheng","year":"2017","unstructured":"Zheng Shou, Jonathan Chan, Alireza Zareian, Kazuyuki Miyazawa, and Shih-Fu Chang. 2017. CDC: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5734\u20135743."},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.119"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_31"},{"key":"e_1_3_1_56_2","first-page":"1","volume-title":"Proceedings of the International Conference for Learning Representations","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In Proceedings of the International Conference for Learning Representations. 1\u201314."},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3086591"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV48630.2021.00213"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00972"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_1_61_2","first-page":"5998","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_1_62_2","first-page":"12168","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Wang Jingwen","year":"2020","unstructured":"Jingwen Wang, Lin Ma, and Wenhao Jiang. 2020. Temporally grounding language queries in videos by contextual boundary-aware prediction. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 12168\u201312175."},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00042"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413862"},{"key":"e_1_3_1_65_2","first-page":"12386","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Wu Jie","year":"2020","unstructured":"Jie Wu, Guanbin Li, Si Liu, and Liang Lin. 2020. Tree-structured policy based progressive reinforcement learning for temporally language grounding in video. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 12386\u201312393."},{"issue":"04","key":"e_1_3_1_66_2","doi-asserted-by":"crossref","first-page":"2986","DOI":"10.1609\/aaai.v35i4.16406","article-title":"Boundary proposal network for two-stage natural language video localization","volume":"35","author":"Xiao Shaoning","year":"2021","unstructured":"Shaoning Xiao, Long Chen, Songyang Zhang, Wei Ji, Jian Shao, Lu Ye, and Jun Xiao. 2021. Boundary proposal network for two-stage natural language video localization. Proceedings of the AAAI Conference on Artificial Intelligence 35, 04, 2986\u20132994.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_1_67_2","first-page":"9062","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Xu Huijuan","year":"2019","unstructured":"Huijuan Xu, Kun He, Bryan A. Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko. 2019. Multilevel language and vision integration for text-to-clip retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 9062\u20139069."},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3140611"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/2962719"},{"key":"e_1_3_1_70_2","first-page":"536","volume-title":"Advances in Neural Information Processing Systems","author":"Yuan Yitian","year":"2019","unstructured":"Yitian Yuan, Lin Ma, Jingwen Wang, Wei Liu, and Wenwu Zhu. 2019. Semantic conditioned dynamic modulation for temporal sentence grounding in videos. In Advances in Neural Information Processing Systems. 536\u2013546."},{"key":"e_1_3_1_71_2","article-title":"Semantic conditioned dynamic modulation for temporal sentence grounding in videos","author":"Yuan Yitian","year":"2020","unstructured":"Yitian Yuan, Lin Ma, Jingwen Wang, Wei Liu, and Wenwu Zhu. 2020. Semantic conditioned dynamic modulation for temporal sentence grounding in videos. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 5 (2020), 2725\u20132741.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_72_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33019159"},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00719"},{"key":"e_1_3_1_74_2","first-page":"10287","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zeng Runhao","year":"2020","unstructured":"Runhao Zeng, Haoming Xu, Wenbing Huang, Peihao Chen, Mingkui Tan, and Chuang Gan. 2020. Dense regression network for video grounding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 10287\u201310296."},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00134"},{"key":"e_1_3_1_76_2","first-page":"3882","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhang Da","year":"2020","unstructured":"Da Zhang, Xiyang Dai, and Yuan-Fang Wang. 2020. METAL: Minimum effort temporal activity localization in untrimmed videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3882\u20133892."},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.585"},{"key":"e_1_3_1_78_2","first-page":"12870","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Zhang Songyang","year":"2020","unstructured":"Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo. 2020. Learning 2D temporal adjacent networks for moment localization with natural language. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 12870\u201312877."},{"key":"e_1_3_1_79_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350879"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1145\/3331184.3331235"},{"key":"e_1_3_1_81_2","article-title":"Temporal textual localization in video via adversarial bi-directional interaction networks","author":"Zhang Zijian","year":"2020","unstructured":"Zijian Zhang, Zhou Zhao, Zhu Zhang, Zhijie Lin, Qi Wang, and Richang Hong. 2020. Temporal textual localization in video via adversarial bi-directional interaction networks. IEEE Transactions on Multimedia 23 (2020), 3306\u20133317.","journal-title":"IEEE Transactions on Multimedia"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3543857","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3543857","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T17:49:40Z","timestamp":1750268980000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3543857"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,6]]},"references-count":80,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,5,31]]}},"alternative-id":["10.1145\/3543857"],"URL":"https:\/\/doi.org\/10.1145\/3543857","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,6]]},"assertion":[{"value":"2021-12-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-05-31","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}