{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,20]],"date-time":"2025-08-20T12:58:28Z","timestamp":1755694708870,"version":"3.41.0"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,2,25]],"date-time":"2023-02-25T00:00:00Z","timestamp":1677283200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,8,31]]},"abstract":"<jats:p>Temporal action proposal generation aims to localize temporal segments of human activities in videos. Current boundary-based proposal generation methods can generate proposals with precise boundary but often suffer from the inferior quality of confidence scores used for proposal retrieving. In this article, we propose an effective and end-to-end action proposal generation method, named ProposalVLAD, with Proposal-Intra Exploring Network (PVPI-Net). We first propose a ProposalVLAD module to dynamically generate global features of the entire video, then we combine the global features and proposal local features to generate the final feature representations for all candidate proposals. Then, we design a novel Proposal-Intra Loss function (PI-Loss) to generate more reliable proposal confidence scores. Extensive experiments on large-scale and challenging datasets demonstrate the effectiveness of our proposed method. Experimental results show that our PVPI-Net achieves significant improvements on two benchmark datasets (i.e., THUMOS\u201914 and ActivityNet-1.3) and sets new records for temporal action detection task.<\/jats:p>","DOI":"10.1145\/3571747","type":"journal-article","created":{"date-parts":[[2022,11,24]],"date-time":"2022-11-24T11:47:35Z","timestamp":1669290455000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["ProposalVLAD with Proposal-Intra Exploring for Temporal Action Proposal Generation"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3706-5321","authenticated-orcid":false,"given":"Kai","family":"Xing","sequence":"first","affiliation":[{"name":"University of Electronic Science and Technology of China, Sichuan Province, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5873-0952","authenticated-orcid":false,"given":"Tao","family":"Li","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Sichuan Province, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3881-9658","authenticated-orcid":false,"given":"Xuanhan","family":"Wang","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Sichuan Province, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,2,25]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"12487","volume-title":"Proceedings of the CVPR","author":"Aafaq Nayyer","year":"2019","unstructured":"Nayyer Aafaq, Naveed Akhtar, Wei Liu, Syed Zulqarnain Gilani, and Ajmal Mian. 2019. Spatio-temporal dynamics and semantic attribute enriched visual encoding for video captioning. In Proceedings of the CVPR. 12487\u201312496."},{"key":"e_1_3_1_3_2","first-page":"2911","volume-title":"Proceedings of the CVPR","author":"Buch Shyamal","year":"2017","unstructured":"Shyamal Buch, Victor Escorcia, Chuanqi Shen, Bernard Ghanem, and Juan Carlos Niebles. 2017. SST: Single-stream temporal action proposals. In Proceedings of the CVPR. 2911\u20132920."},{"key":"e_1_3_1_4_2","first-page":"961","volume-title":"Proceedings of the CVPR","author":"Heilbron Fabian Caba","year":"2015","unstructured":"Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the CVPR. 961\u2013970."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_1_6_2","first-page":"1130","volume-title":"Proceedings of the CVPR","author":"Chao Yu-Wei","year":"2018","unstructured":"Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A. Ross, Jia Deng, and Rahul Sukthankar. 2018. Rethinking the faster R-CNN architecture for temporal action localization. In Proceedings of the CVPR. 1130\u20131139."},{"key":"e_1_3_1_7_2","first-page":"5793","volume-title":"Proceedings of the ICCV","author":"Dai Xiyang","year":"2017","unstructured":"Xiyang Dai, Bharat Singh, Guyue Zhang, Larry S. Davis, and Yan Qiu Chen. 2017. Temporal context network for activity localization in videos. In Proceedings of the ICCV. 5793\u20135802."},{"key":"e_1_3_1_8_2","first-page":"1933","volume-title":"Proceedings of the CVPR","author":"Feichtenhofer Christoph","year":"2016","unstructured":"Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman. 2016. Convolutional two-stream network fusion for video action recognition. In Proceedings of the CVPR. 1933\u20131941."},{"key":"e_1_3_1_9_2","first-page":"68","volume-title":"Proceedings of the ECCV","author":"Gao Jiyang","year":"2018","unstructured":"Jiyang Gao, Kan Chen, and Ram Nevatia. 2018. CTAP: Complementary temporal action proposal generation. In Proceedings of the ECCV. 68\u201383."},{"key":"e_1_3_1_10_2","first-page":"10810","volume-title":"Proceedings of the AAAI","volume":"34","author":"Gao Jialin","year":"2020","unstructured":"Jialin Gao, Zhixiang Shi, Guanshuo Wang, Jiani Li, Yufeng Yuan, Shiming Ge, and Xi Zhou. 2020. Accurate temporal action proposal generation with relation-aware pyramid network. In Proceedings of the AAAI, Vol. 34. 10810\u201310817."},{"key":"e_1_3_1_11_2","first-page":"3628","volume-title":"Proceedings of the ICCV","author":"Gao Jiyang","year":"2017","unstructured":"Jiyang Gao, Zhenheng Yang, Kan Chen, Chen Sun, and Ram Nevatia. 2017. Turn tap: Temporal unit regression network for temporal action proposals. In Proceedings of the ICCV. 3628\u20133636."},{"key":"e_1_3_1_12_2","volume-title":"Proceedings of the BMVC","author":"Gao Jiyang","year":"2017","unstructured":"Jiyang Gao, Zhenheng Yang, and Ram Nevatia. 2017. Cascaded boundary regression for temporal action detection. In Proceedings of the BMVC."},{"key":"e_1_3_1_13_2","unstructured":"Bernard Ghanem Juan Carlos Niebles Cees Snoek Fabian Caba Heilbron Humam Alwassel Ranjay Krishna Victor Escorcia Kenji Hata and Shyamal Buch. 2017. ActivityNet challenge 2017 summary. Retrieved from https:\/\/arxiv.org\/abs\/1710.08011."},{"key":"e_1_3_1_14_2","first-page":"580","volume-title":"Proceedings of the CVPR","author":"Girshick Ross","year":"2014","unstructured":"Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. 2014. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the CVPR. 580\u2013587."},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","DOI":"10.1145\/3495211","article-title":"When did it happen? Duration-informed temporal localization of narrated actions in vlogs","author":"Ignat Oana","year":"2022","unstructured":"Oana Ignat, Santiago Castro, Yuhang Zhou, Jiajun Bao, Dandan Shan, and Rada Mihalcea. 2022. When did it happen? Duration-informed temporal localization of narrated actions in vlogs. ACM Trans. Multimedia Comput. Commun. Appl. (2022).","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_1_16_2","unstructured":"Y.-G. Jiang J. Liu A. Roshan Zamir G. Toderici I. Laptev M. Shah and R. Sukthankar. 2014. THUMOS Challenge: Action Recognition with a Large Number of Classes. http:\/\/crcv.ucf.edu\/THUMOS14\/."},{"key":"e_1_3_1_17_2","first-page":"4405","volume-title":"Proceedings of the ICCV","author":"Kalogeiton Vicky","year":"2017","unstructured":"Vicky Kalogeiton, Philippe Weinzaepfel, Vittorio Ferrari, and Cordelia Schmid. 2017. Action tubelet detector for spatio-temporal action localization. In Proceedings of the ICCV. 4405\u20134413."},{"key":"e_1_3_1_18_2","first-page":"68","volume-title":"Proceedings of the ECCV","author":"Li Yixuan","year":"2020","unstructured":"Yixuan Li, Zixu Wang, Limin Wang, and Gangshan Wu. 2020. Actions as moving points. In Proceedings of the ECCV. 68\u201384."},{"key":"e_1_3_1_19_2","first-page":"11499","volume-title":"Proceedings of the AAAI","volume":"34","author":"Lin Chuming","year":"2020","unstructured":"Chuming Lin, Jian Li, Yabiao Wang, Ying Tai, Donghao Luo, Zhipeng Cui, Chengjie Wang, Jilin Li, Feiyue Huang, and Rongrong Ji. 2020. Fast learning of temporal action proposal via dense boundary generator. In Proceedings of the AAAI, Vol. 34. 11499\u201311506."},{"key":"e_1_3_1_20_2","first-page":"7083","volume-title":"Proceedings of the ICCV","author":"Lin Ji","year":"2019","unstructured":"Ji Lin, Chuang Gan, and Song Han. 2019. TSM: Temporal shift module for efficient video understanding. In Proceedings of the ICCV. 7083\u20137093."},{"key":"e_1_3_1_21_2","first-page":"3889","volume-title":"Proceedings of the ICCV","author":"Lin Tianwei","year":"2019","unstructured":"Tianwei Lin, Xiao Liu, Xin Li, Errui Ding, and Shilei Wen. 2019. BMN: Boundary-matching network for temporal action proposal generation. In Proceedings of the ICCV. 3889\u20133898."},{"key":"e_1_3_1_22_2","first-page":"988","volume-title":"Proceedings of the ACM MM","author":"Lin Tianwei","year":"2017","unstructured":"Tianwei Lin, Xu Zhao, and Zheng Shou. 2017. Single shot temporal action detection. In Proceedings of the ACM MM. 988\u2013996."},{"key":"e_1_3_1_23_2","unstructured":"Tianwei Lin Xu Zhao and Zheng Shou. 2017. Temporal convolution-based action proposal: Submission to ActivityNet 2017. Retrieved from https:\/\/arxiv.org\/abs\/1707.06750."},{"key":"e_1_3_1_24_2","first-page":"3","volume-title":"Proceedings of the ECCV","author":"Lin Tianwei","year":"2018","unstructured":"Tianwei Lin, Xu Zhao, Haisheng Su, Chongjing Wang, and Ming Yang. 2018. BSN: Boundary sensitive network for temporal action proposal generation. In Proceedings of the ECCV. 3\u201319."},{"key":"e_1_3_1_25_2","first-page":"202","article-title":"Interpretable deep generative recommendation models.","volume":"22","author":"Liu Huafeng","year":"2021","unstructured":"Huafeng Liu, Liping Jing, Jingxuan Wen, Pengyu Xu, Jiaqi Wang, Jian Yu, and Michael K Ng. 2021. Interpretable deep generative recommendation models.J. Mach. Learn. Res. 22 (2021), 202\u20131.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_1_26_2","first-page":"3604","volume-title":"Proceedings of the CVPR","author":"Liu Yuan","year":"2019","unstructured":"Yuan Liu, Lin Ma, Yifeng Zhang, Wei Liu, and Shih-Fu Chang. 2019. Multi-granularity generator for temporal action proposal. In Proceedings of the CVPR. 3604\u20133613."},{"key":"e_1_3_1_27_2","first-page":"344","volume-title":"Proceedings of the CVPR","author":"Long Fuchen","year":"2019","unstructured":"Fuchen Long, Ting Yao, Zhaofan Qiu, Xinmei Tian, Jiebo Luo, and Tao Mei. 2019. Gaussian temporal awareness networks for action localization. In Proceedings of the CVPR. 344\u2013353."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3350840"},{"key":"e_1_3_1_29_2","first-page":"485","volume-title":"Proceedings of the CVPR","author":"Qing Zhiwu","year":"2021","unstructured":"Zhiwu Qing, Haisheng Su, Weihao Gan, Dongliang Wang, Wei Wu, Xiang Wang, Yu Qiao, Junjie Yan, Changxin Gao, and Nong Sang. 2021. Temporal context aggregation network for temporal action proposal refinement. In Proceedings of the CVPR. 485\u2013494."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.590"},{"key":"e_1_3_1_31_2","first-page":"1049","volume-title":"Proceedings of the CVPR","author":"Shou Zheng","year":"2016","unstructured":"Zheng Shou, Dongang Wang, and Shih-Fu Chang. 2016. Temporal action localization in untrimmed videos via multi-stage cnns. In Proceedings of the CVPR. 1049\u20131058."},{"key":"e_1_3_1_32_2","first-page":"568","volume-title":"Proceedings of the NeurIPS","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In Proceedings of the NeurIPS. 568\u2013576."},{"key":"e_1_3_1_33_2","unstructured":"Gurkirt Singh and Fabio Cuzzolin. 2016. Untrimmed video classification for activity detection: submission to activitynet challenge. Retrieved from https:\/\/arxiv.org\/abs\/1607.01979."},{"key":"e_1_3_1_34_2","first-page":"2602","volume-title":"Proceedings of the AAAI","volume":"35","author":"Su Haisheng","year":"2021","unstructured":"Haisheng Su, Weihao Gan, Wei Wu, Yu Qiao, and Junjie Yan. 2021. BSN++: Complementary boundary regressor with scale-balanced relation modeling for temporal action proposal generation. In Proceedings of the AAAI, Vol. 35. 2602\u20132610."},{"key":"e_1_3_1_35_2","first-page":"13526","volume-title":"Proceedings of the ICCV","author":"Tan Jing","year":"2021","unstructured":"Jing Tan, Jiaqi Tang, Limin Wang, and Gangshan Wu. 2021. Relaxed transformer decoders for direct action proposal generation. In Proceedings of the ICCV. 13526\u201313535."},{"key":"e_1_3_1_36_2","first-page":"4489","volume-title":"Proceedings of the ICCV","author":"Tran Du","year":"2015","unstructured":"Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015. Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the ICCV. 4489\u20134497."},{"key":"e_1_3_1_37_2","first-page":"6450","volume-title":"Proceedings of the CVPR","author":"Tran Du","year":"2018","unstructured":"Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. 2018. A closer look at spatiotemporal convolutions for action recognition. In Proceedings of the CVPR. 6450\u20136459."},{"key":"e_1_3_1_38_2","first-page":"3169","volume-title":"Proceedings of the CVPR","author":"Wang Heng","year":"2011","unstructured":"Heng Wang, Alexander Kl\u00e4ser, Cordelia Schmid, and Cheng-Lin Liu. 2011. Action recognition by dense trajectories. In Proceedings of the CVPR. 3169\u20133176."},{"key":"e_1_3_1_39_2","first-page":"3551","volume-title":"Proceedings of the ICCV","author":"Wang Heng","year":"2013","unstructured":"Heng Wang and Cordelia Schmid. 2013. Action recognition with improved trajectories. In Proceedings of the ICCV. 3551\u20133558."},{"key":"e_1_3_1_40_2","first-page":"4325","volume-title":"Proceedings of the CVPR","author":"Wang Limin","year":"2017","unstructured":"Limin Wang, Yuanjun Xiong, Dahua Lin, and Luc Van Gool. 2017. Untrimmednets for weakly supervised action recognition and detection. In Proceedings of the CVPR. 4325\u20134334."},{"key":"e_1_3_1_41_2","unstructured":"Limin Wang Yuanjun Xiong Zhe Wang and Yu Qiao. 2015. Towards good practices for very deep two-stream ConvNets. Retrieved from https:\/\/arxiv.org\/abs\/1507.02159."},{"key":"e_1_3_1_42_2","first-page":"1670","volume-title":"Proceedings of the ACM MM","author":"Wang Xuanhan","year":"2022","unstructured":"Xuanhan Wang, Yan Dai, Lianli Gao, and Jingkuan Song. 2022. Skeleton-based action recognition via adaptive cross-form learning. In Proceedings of the ACM MM. 1670\u20131678."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3135145"},{"key":"e_1_3_1_44_2","unstructured":"Saining Xie Chen Sun Jonathan Huang Zhuowen Tu and Kevin Murphy. 2017. Rethinking spatiotemporal feature learning for video understanding. Retrieved from https:\/\/arxiv.org\/abs\/1712.04851."},{"key":"e_1_3_1_45_2","unstructured":"Yuanjun Xiong Limin Wang Zhe Wang Bowen Zhang Hang Song Wei Li Dahua Lin Yu Qiao Luc Van Gool and Xiaoou Tang. 2016. CUHK & ETHZ & SIAT submission to activitynet challenge 2016. Retrieved from https:\/\/arxiv.org\/pdf\/1608.00797.pdf."},{"key":"e_1_3_1_46_2","unstructured":"Yuanjun Xiong Yue Zhao Limin Wang Dahua Lin and Xiaoou Tang. 2017. A pursuit of temporal accuracy in general activity detection. Retrieved from https:\/\/arxiv.org\/abs\/1703.02716."},{"key":"e_1_3_1_47_2","first-page":"10156","volume-title":"Proceedings of the CVPR","author":"Xu Mengmeng","year":"2020","unstructured":"Mengmeng Xu, Chen Zhao, David S Rojas, Ali Thabet, and Bernard Ghanem. 2020. G-tad: Sub-graph localization for temporal action detection. In Proceedings of the CVPR. 10156\u201310165."},{"key":"e_1_3_1_48_2","first-page":"264","volume-title":"Proceedings of the CVPR","author":"Yang Xitong","year":"2019","unstructured":"Xitong Yang, Xiaodong Yang, Ming-Yu Liu, Fanyi Xiao, Larry S. Davis, and Jan Kautz. 2019. STEP: Spatio-temporal progressive learning for video action detection. In Proceedings of the CVPR. 264\u2013272."},{"key":"e_1_3_1_49_2","first-page":"7094","volume-title":"Proceedings of the ICCV","author":"Zeng Runhao","year":"2019","unstructured":"Runhao Zeng, Wenbing Huang, Mingkui Tan, Yu Rong, Peilin Zhao, Junzhou Huang, and Chuang Gan. 2019. Graph convolutional networks for temporal action localization. In Proceedings of the ICCV. 7094\u20137103."},{"issue":"3","key":"e_1_3_1_50_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3321511","article-title":"Moving foreground-aware visual attention and key volume mining for human action recognition","volume":"15","author":"Zhang Junxuan","year":"2019","unstructured":"Junxuan Zhang, Haifeng Hu, and Xinlong Lu. 2019. Moving foreground-aware visual attention and key volume mining for human action recognition. ACM Trans. Multimedia Comput. Commun. Appl. 15, 3 (2019), 1\u201316.","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3164190"},{"key":"e_1_3_1_52_2","first-page":"2914","volume-title":"Proceedings of the ICCV","author":"Zhao Yue","year":"2017","unstructured":"Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin. 2017. Temporal action detection with structured segment networks. In Proceedings of the ICCV. 2914\u20132923."},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3361845"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3571747","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3571747","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:33Z","timestamp":1750182573000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3571747"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,25]]},"references-count":52,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,8,31]]}},"alternative-id":["10.1145\/3571747"],"URL":"https:\/\/doi.org\/10.1145\/3571747","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2023,2,25]]},"assertion":[{"value":"2022-07-11","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-11-06","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}