{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T21:05:49Z","timestamp":1768338349290,"version":"3.49.0"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,2,25]],"date-time":"2023-02-25T00:00:00Z","timestamp":1677283200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62076262, 61673402, 61273270, 60802069"],"award-info":[{"award-number":["62076262, 61673402, 61273270, 60802069"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,8,31]]},"abstract":"<jats:p>\n            Recent action localization works learn in a weakly supervised manner to avoid the expensive cost of human labeling. Those works are mostly based on the Multiple Instance Learning framework, where temporal pooling is an indispensable part that usually relies on the guidance of snippet-level\n            <jats:bold>Class Activation Sequences (CAS)<\/jats:bold>\n            . However, we observe that previous works only leverage a simple convolutional neural network for the generation of CAS, which ignores the weak discriminative foreground action segments and the background ones, and meanwhile, the relationship between different actions has not been considered. To solve this problem, we propose\n            <jats:bold>multiple temporal pooling mechanisms (MTP)<\/jats:bold>\n            for a more sufficient information utilization. Specifically, with the design of the Foreground Variance Branch, Dual Foreground Attention Branch and Hybrid Attention Fine-tuning Branch, MTP can leverage more effective information from different aspects and generate different CASs to guide the learning of temporal pooling. Moreover, different loss functions are designed for a better optimization of individual branches, aiming to effectively distinguish the action from the background. Our method shows excellent results on the THUMOS14 and ActivityNet1.2 datasets.\n          <\/jats:p>","DOI":"10.1145\/3567828","type":"journal-article","created":{"date-parts":[[2022,10,13]],"date-time":"2022-10-13T10:16:59Z","timestamp":1665656219000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Multiple Temporal Pooling Mechanisms for Weakly Supervised Temporal Action Localization"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9802-1226","authenticated-orcid":false,"given":"Peng","family":"Dou","sequence":"first","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-sen University, Guangzhou, Guangdong, People\u2019s Republic of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8842-2045","authenticated-orcid":false,"given":"Ying","family":"Zeng","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-sen University, Guangzhou, Guangdong, People\u2019s Republic of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8633-0235","authenticated-orcid":false,"given":"Zhuoqun","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-sen University, Guangzhou, Guangdong, People\u2019s Republic of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4884-323X","authenticated-orcid":false,"given":"Haifeng","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-sen University, Guangzhou, Guangdong, People\u2019s Republic of China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,2,25]]},"reference":[{"key":"e_1_3_1_2_2","volume-title":"Proceedings of the British Machine Vision Conference 2017","author":"Buch Shyamal","year":"2019","unstructured":"Shyamal Buch, Victor Escorcia, Bernard Ghanem, Li Fei-Fei, and Juan Carlos Niebles. 2019. End-to-end, single-stream temporal action detection in untrimmed videos. In Proceedings of the British Machine Vision Conference 2017. British Machine Vision Association."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.675"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00124"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2959977"},{"key":"e_1_3_1_8_2","first-page":"459","volume-title":"Pattern Recognition and Computer Vision","author":"Dou Peng","year":"2021","unstructured":"Peng Dou, Wei Zhou, Zhongke Liao, and Haifeng Hu. 2021. Feature matching network for weakly-supervised temporal action localization. In Pattern Recognition and Computer Vision. 459\u2013471."},{"key":"e_1_3_1_9_2","volume-title":"2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Gong G.","year":"2020","unstructured":"G. Gong, X. Wang, Y. Mu, and Q. Tian. 2020. Learning temporal co-attention models for unsupervised video action localization. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)."},{"key":"e_1_3_1_10_2","doi-asserted-by":"crossref","unstructured":"C. Gu C. Sun D. A. Ross C. Vondrick C. Pantofaru Y. Li S. Vijayanarasimhan G. Toderici S. Ricco and R. Sukthankar. 2017. AVA: A video dataset of spatio-temporally localized atomic visual actions. (2017).","DOI":"10.1109\/CVPR.2018.00633"},{"key":"e_1_3_1_11_2","volume-title":"Techniques and Applications of Image Understanding","author":"Horn B. K. P.","year":"1981","unstructured":"B. K. P. Horn and B. G. Schunck. 1981. Determining optical flow. In Techniques and Applications of Image Understanding."},{"key":"e_1_3_1_12_2","first-page":"11053","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Huang Linjiang","year":"2020","unstructured":"Linjiang Huang, Yan Huang, Wanli Ouyang, and Liang Wang. 2020. Relational prototypical network for weakly supervised temporal action localization. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 11053\u201311060."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3078324"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2016.10.018"},{"key":"e_1_3_1_15_2","first-page":"1637","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Islam Ashraful","year":"2021","unstructured":"Ashraful Islam, Chengjiang Long, and Richard Radke. 2021. A hybrid attention mechanism for weakly-supervised temporal action localization. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 1637\u20131645."},{"key":"e_1_3_1_16_2","volume-title":"2020 IEEE Winter Conference on Applications of Computer Vision (WACV\u201920)","author":"Islam A.","year":"2020","unstructured":"A. Islam and R. J. Radke. 2020. Weakly supervised temporal action localization using deep metric learning. In 2020 IEEE Winter Conference on Applications of Computer Vision (WACV\u201920)."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475261"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.223"},{"key":"e_1_3_1_19_2","unstructured":"W. Kay J. Carreira K. Simonyan B. Zhang and A. Zisserman. 2017. The kinetics human action video dataset. (2017)."},{"key":"e_1_3_1_20_2","first-page":"3524","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Singh Krishna Kumar","year":"2017","unstructured":"Krishna Kumar Singh and Yong Jae Lee. 2017. Hide-and-seek: Forcing a network to be meticulous for weakly-supervised object and action localization. In Proceedings of the IEEE International Conference on Computer Vision. 3524\u20133533."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.113"},{"key":"e_1_3_1_22_2","first-page":"11320","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Lee Pilhyeon","year":"2020","unstructured":"Pilhyeon Lee, Youngjung Uh, and Hyeran Byun. 2020. Background suppression network for weakly-supervised temporal action localization. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 11320\u201311327."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01225-0_1"},{"key":"e_1_3_1_24_2","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Liu D.","year":"2019","unstructured":"D. Liu, T. Jiang, and Y. Wang. 2019. Completeness modeling and context separation for weakly supervised temporal action localization. In IEEE Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_1_25_2","volume-title":"2019 IEEE\/CVF International Conference on Computer Vision (ICCV\u201920)","author":"Liu Z.","year":"2020","unstructured":"Z. Liu, L. Wang, Q. Zhang, Z. Gao, Z. Niu, N Zheng, and G. Hua. 2020. Weakly supervised temporal action localization through contrast based evaluation networks. In 2019 IEEE\/CVF International Conference on Computer Vision (ICCV\u201920)."},{"key":"e_1_3_1_26_2","doi-asserted-by":"crossref","unstructured":"Z. Liu L. Wang Q. Zhang W. Tang J. Yuan N. Zheng and G. Hua. 2021. ACSNet: Action-context separation network for weakly supervised temporal action localization. (2021).","DOI":"10.1609\/aaai.v35i3.16322"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00043"},{"key":"e_1_3_1_28_2","first-page":"9969","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Luo Wang","year":"2021","unstructured":"Wang Luo, Tianzhu Zhang, Wenfei Yang, Jingen Liu, Tao Mei, Feng Wu, and Yongdong Zhang. 2021. Action unit memory network for weakly supervised temporal action localization. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 9969\u20139979."},{"key":"e_1_3_1_29_2","first-page":"729","volume-title":"European Conference on Computer Vision","author":"Luo Zhekun","year":"2020","unstructured":"Zhekun Luo, Devin Guillory, Baifeng Shi, Wei Ke, Fang Wan, Trevor Darrell, and Huijuan Xu. 2020. Weakly-supervised action localization with expectation-maximization multi-instance learning. In European Conference on Computer Vision. Springer, 729\u2013745."},{"key":"e_1_3_1_30_2","first-page":"283","volume-title":"European Conference on Computer Vision","author":"Min Kyle","year":"2020","unstructured":"Kyle Min and Jason J. Corso. 2020. Adversarial background-aware loss for weakly-supervised temporal activity localization. In European Conference on Computer Vision. Springer, 283\u2013299."},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413687"},{"key":"e_1_3_1_32_2","first-page":"8679","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Narayan Sanath","year":"2019","unstructured":"Sanath Narayan, Hisham Cholakkal, Fahad Shahbaz Khan, and Ling Shao. 2019. 3C-Net: Category count and center loss for weakly-supervised action localization. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 8679\u20138687."},{"key":"e_1_3_1_33_2","first-page":"6752","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Nguyen Phuc","year":"2018","unstructured":"Phuc Nguyen, Ting Liu, Gautam Prasad, and Bohyung Han. 2018. Weakly supervised action localization by sparse temporal pooling network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 6752\u20136761."},{"key":"e_1_3_1_34_2","first-page":"5502","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Nguyen Phuc Xuan","year":"2019","unstructured":"Phuc Xuan Nguyen, Deva Ramanan, and Charless C. Fowlkes. 2019. Weakly-supervised action localization with background modeling. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 5502\u20135511."},{"key":"e_1_3_1_35_2","first-page":"3319","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Pardo Alejandro","year":"2021","unstructured":"Alejandro Pardo, Humam Alwassel, Fabian Caba, Ali Thabet, and Bernard Ghanem. 2021. RefineLoc: Iterative refinement for weakly-supervised action localization. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 3319\u20133328."},{"key":"e_1_3_1_36_2","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201918)","author":"Paul Sujoy","year":"2018","unstructured":"Sujoy Paul, Sourya Roy, and Amit K. Roy-Chowdhury. 2018. W-TALC: Weakly-supervised temporal activity localization and classification. In Proceedings of the European Conference on Computer Vision (ECCV\u201918)."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2020.3018914"},{"key":"e_1_3_1_38_2","first-page":"1009","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Shi Baifeng","year":"2020","unstructured":"Baifeng Shi, Qi Dai, Yadong Mu, and Jingdong Wang. 2020. Weakly-supervised action localization by generative attention modeling. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 1009\u20131019."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.155"},{"key":"e_1_3_1_40_2","doi-asserted-by":"crossref","unstructured":"Z. Shou H. Gao L. Zhang Kazuyuki Mayazawa and Shih-Fu Chang. 2018. AutoLoc: Weakly-supervised temporal action localization. In Proceedings of the European Conference on Computer Vision (ECCV\u201918) .","DOI":"10.1007\/978-3-030-01270-0_10"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.119"},{"key":"e_1_3_1_42_2","volume-title":"Springer, Cham","author":"Sigurdsson G. A.","year":"2016","unstructured":"G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta. 2016. Hollywood in homes: Crowdsourcing data collection for activity understanding. In Springer, Cham."},{"key":"e_1_3_1_43_2","article-title":"Two-stream convolutional networks for action recognition in videos","volume":"27","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. Advances in Neural Information Processing Systems 27 (2014).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00675"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2021.3061289"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.678"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.617"},{"key":"e_1_3_1_49_2","first-page":"10156","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Xu Mengmeng","year":"2020","unstructured":"Mengmeng Xu, Chen Zhao, David S. Rojas, Ali Thabet, and Bernard Ghanem. 2020. G-TAD: Sub-graph localization for temporal action detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10156\u201310165."},{"key":"e_1_3_1_50_2","first-page":"264","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yang Xitong","year":"2019","unstructured":"Xitong Yang, Xiaodong Yang, Ming-Yu Liu, Fanyi Xiao, Larry S. Davis, and Jan Kautz. 2019. Step: Spatio-temporal progressive learning for video action detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 264\u2013272."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.293"},{"issue":"21","key":"e_1_3_1_52_2","doi-asserted-by":"crossref","first-page":"1126","DOI":"10.1049\/el.2019.2088","article-title":"Weakly supervised video action localisation via two-stream action activation network","volume":"55","author":"Yin C.","year":"2019","unstructured":"C. Yin, Z. Liao, H. Hu, and D. Chen. 2019. Weakly supervised video action localisation via two-stream action activation network. Electronics Letters 55, 21 (2019), 1126\u20131127.","journal-title":"Electronics Letters"},{"key":"e_1_3_1_53_2","first-page":"5522","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Yu Tan","year":"2019","unstructured":"Tan Yu, Zhou Ren, Yuncheng Li, Enxu Yan, Ning Xu, and Junsong Yuan. 2019. Temporal structure mining for weakly supervised action detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 5522\u20135531."},{"key":"e_1_3_1_54_2","volume-title":"ICLR 2019-Seventh International Conference on Learning Representations","author":"Yuan Yuan","year":"2019","unstructured":"Yuan Yuan, Yueming Lyu, Xi Shen, Ivor Tsang, and Dit-Yan Yeung. 2019. Marginalized average attentional network for weakly-supervised learning. In ICLR 2019-Seventh International Conference on Learning Representations."},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2922108"},{"key":"e_1_3_1_56_2","first-page":"7094","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zeng Runhao","year":"2019","unstructured":"Runhao Zeng, Wenbing Huang, Mingkui Tan, Yu Rong, Peilin Zhao, Junzhou Huang, and Chuang Gan. 2019. Graph convolutional networks for temporal action localization. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 7094\u20137103."},{"key":"e_1_3_1_57_2","article-title":"Graph convolutional module for temporal action localization in videos","author":"Zeng Runhao","year":"2021","unstructured":"Runhao Zeng, Wenbing Huang, Mingkui Tan, Yu Rong, Peilin Zhao, Junzhou Huang, and Chuang Gan. 2021. Graph convolutional module for temporal action localization in videos. IEEE Transactions on Pattern Analysis and Machine Intelligence (2021).","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_58_2","first-page":"37","volume-title":"European Conference on Computer Vision","author":"Zhai Yuanhao","year":"2020","unstructured":"Yuanhao Zhai, Le Wang, Wei Tang, Qilin Zhang, Junsong Yuan, and Gang Hua. 2020. Two-stream consensus network for weakly-supervised temporal action localization. In European Conference on Computer Vision. Springer, 37\u201354."},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2021.3132287"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.317"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3567828","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3567828","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:51:13Z","timestamp":1750182673000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3567828"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,25]]},"references-count":59,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,8,31]]}},"alternative-id":["10.1145\/3567828"],"URL":"https:\/\/doi.org\/10.1145\/3567828","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,25]]},"assertion":[{"value":"2022-04-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-10-03","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}