{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T22:34:17Z","timestamp":1783636457006,"version":"3.55.0"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T00:00:00Z","timestamp":1782259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation (NSF) of China","doi-asserted-by":"crossref","award":["62576195, 62276155, 62206156"],"award-info":[{"award-number":["62576195, 62276155, 62206156"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key R&D Program of Shandong Province (Major scientific and technological innovation projects), China","award":["2025CXGC020101"],"award-info":[{"award-number":["2025CXGC020101"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>Recently, Temporal Action Localization (TAL) has garnered significant interest in information retrieval community. However, existing supervised\/weakly supervised methods are heavily dependent on extensive labeled temporal boundaries and action categories, which is labor-intensive and time-consuming. Although some unsupervised methods have utilized the \u201citeratively clustering and localization\u201d paradigm for TAL, they still suffer from two pivotal impediments: (1) unsatisfactory video clustering confidence, and (2) unreliable video pseudolabels for model training. To address these limitations, we present a novel self-paced iterative learning model to enhance clustering and localization training simultaneously, thereby facilitating more effective unsupervised TAL. Concretely, we improve the clustering confidence through exploring the contextual feature-robust visual information. Thereafter, we design two (constant- and variable-speed) incremental instance learning strategies for easy-to-hard model training, thus ensuring the reliability of these video pseudolabels and further improving overall localization performance. Extensive experiments on two public datasets demonstrate the superiority of our model over several state-of-the-art competitors.<\/jats:p>","DOI":"10.1145\/3796708","type":"journal-article","created":{"date-parts":[[2026,2,10]],"date-time":"2026-02-10T15:42:37Z","timestamp":1770738157000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Visual Self-paced Iterative Learning for Unsupervised Temporal Action Localization"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5653-8286","authenticated-orcid":false,"given":"Yupeng","family":"Hu","sequence":"first","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-2014-7906","authenticated-orcid":false,"given":"Han","family":"Jiang","sequence":"additional","affiliation":[{"name":"Xi\u2019an Jiaotong University, Xi\u2019an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-1026-4499","authenticated-orcid":false,"given":"Hao","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-4856-8806","authenticated-orcid":false,"given":"Kun","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1344-2513","authenticated-orcid":false,"given":"Haoyu","family":"Tang","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1476-0273","authenticated-orcid":false,"given":"Liqiang","family":"Nie","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3356316"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3663571"},{"key":"e_1_3_2_4_2","first-page":"3985","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Liu Meng","year":"2023","unstructured":"Meng Liu, Fenglei Zhang, Xin Luo, Fan Liu, Yinwei Wei, and Liqiang Nie. 2023. Advancing video question answering with a multi-modal and multi-layer question enhancement network. In Proceedings of the ACM International Conference on Multimedia, 3985\u20133993."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3449986"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"1338","DOI":"10.1109\/TMM.2021.3063631","article-title":"Frame-wise cross-modal matching for video moment retrieval","volume":"24","author":"Tang Haoyu","year":"2021","unstructured":"Haoyu Tang, Jihua Zhu, Meng Liu, Zan Gao, and Zhiyong Cheng. 2021. Frame-wise cross-modal matching for video moment retrieval. IEEE Transactions on Multimedia 24 (2021), 1338\u20131349.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_7_2","first-page":"1","article-title":"Cross-modal representation shift refinement for point-supervised video moment retrieval","author":"Wang Kun","year":"2025","unstructured":"Kun Wang, Yupeng Hu, Hao Liu, Jiang Shao, and Liqiang Nie. 2025. Cross-modal representation shift refinement for point-supervised video moment retrieval. ACM Transactions on Information Systems 44 (2025), 1\u201330.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_2_8_2","first-page":"1","article-title":"Redundancy mitigation: Towards accurate and efficient image-text retrieval","author":"Wang Kun","year":"2025","unstructured":"Kun Wang, Yupeng Hu, Hao Liu, Lirong Jie, and Liqiang Nie. 2025. Redundancy mitigation: Towards accurate and efficient image-text retrieval. IEEE Transactions on Circuits and Systems for Video Technology 1 (2025), 1\u201313.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_2_9_2","first-page":"1476","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Hu Jingjing","year":"2024","unstructured":"Jingjing Hu, Dan Guo, Kun Li, Zhan Si, Xun Yang, and Meng Wang. 2024. Maskable retentive network for video moment retrieval. In Proceedings of the ACM International Conference on Multimedia, 1476\u20131485."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i3.16285"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-022-3783-3"},{"key":"e_1_3_2_12_2","first-page":"1009","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Shi Baifeng","year":"2020","unstructured":"Baifeng Shi, Qi Dai, Yadong Mu, and Jingdong Wang. 2020. Weakly-supervised action localization by generative attention modeling. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1009\u20131019."},{"key":"e_1_3_2_13_2","first-page":"16010","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhang Can","year":"2021","unstructured":"Can Zhang, Meng Cao, Dongming Yang, Jie Chen, and Yuexian Zou. 2021. Cola: Weakly-supervised temporal action localization with snippet contrastive learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 16010\u201316019."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3090521"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3073867"},{"key":"e_1_3_2_16_2","unstructured":"Kun Li Dan Guo Guoliang Chen Chunxiao Fan Jingyuan Xu Zhiliang Wu Hehe Fan and Meng Wang. 2024. Prototypical calibrating ambiguous samples for micro-action recognition. arXiv:2412.14719. Retrieved from https:\/\/arxiv.org\/abs\/2412.14719"},{"key":"e_1_3_2_17_2","first-page":"1","article-title":"Neural multimodal cooperative learning toward micro-video understanding","volume":"29","author":"Wei Yinwei","year":"2019","unstructured":"Yinwei Wei, Xiang Wang, Weili Guan, Liqiang Nie, Zhouchen Lin, and Baoquan Chen. 2019. Neural multimodal cooperative learning toward micro-video understanding. IEEE Transactions on Image Processing 29 (2019), 1\u201314.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3620669"},{"key":"e_1_3_2_19_2","first-page":"9214","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Wang Kun","year":"2024","unstructured":"Kun Wang, Hao Liu, Lirong Jie, Zixu Li, Yupeng Hu, and Liqiang Nie. 2024. Explicit granularity and implicit scale correspondence learning for point-supervised video moment localization. In Proceedings of the ACM International Conference on Multimedia, 9214\u20139223."},{"key":"e_1_3_2_20_2","first-page":"917","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Liu Hao","year":"2025","unstructured":"Hao Liu, Yupeng Hu, Kun Wang, Yinwei Wei, and Liqiang Nie. 2025. Gaming for boundary: Elastic localization for frame-supervised video moment retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval, 917\u2013926."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-024-02205-5"},{"key":"e_1_3_2_22_2","first-page":"5306","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Li Ming","year":"2025","unstructured":"Ming Li, Yupeng Hu, Yinwei Wei, Hao Liu, Haocong Wang, and Weili Guan. 2025. DCount: Decoupled spatial perception and attribute discrimination for referring expression counting. In Proceedings of the ACM International Conference on Multimedia, 5306\u20135315."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3358415"},{"key":"e_1_3_2_24_2","first-page":"13225","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Li Kun","year":"2025","unstructured":"Kun Li, Pengyu Liu, Dan Guo, Fei Wang, Zhiliang Wu, Hehe Fan, and Meng Wang. 2025. Mmad: Multi-label micro-action detection in videos. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 13225\u201313236."},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-023-08703-w"},{"key":"e_1_3_2_26_2","first-page":"11320","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Lee Pilhyeon","year":"2020","unstructured":"Pilhyeon Lee, Youngjung Uh, and Hyeran Byun. 2020. Background suppression network for weakly-supervised temporal action localization. In Proceedings of the AAAI Conference on Artificial Intelligence, 11320\u201311327."},{"issue":"4","key":"e_1_3_2_27_2","first-page":"5252","article-title":"Uncertainty guided collaborative training for weakly supervised and unsupervised temporal action localization","volume":"45","author":"Yang Wenfei","year":"2022","unstructured":"Wenfei Yang, Tianzhu Zhang, Yongdong Zhang, and Feng Wu. 2022. Uncertainty guided collaborative training for weakly supervised and unsupervised temporal action localization. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 4 (2022), 5252\u20135267.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00984"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.389"},{"key":"e_1_3_2_30_2","first-page":"1591","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Hong Fa-Ting","year":"2021","unstructured":"Fa-Ting Hong, Jia-Chang Feng, Dan Xu, Ying Shan, and Wei-Shi Zheng. 2021. Cross-modal consensus network for weakly supervised temporal action localization. In Proceedings of the ACM International Conference on Multimedia, 1591\u20131599."},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01225-0_35"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00560"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.678"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00139"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58539-6_3"},{"key":"e_1_3_2_36_2","first-page":"1189","volume-title":"Proceedings of Advances in Neural Information Processing Systems","author":"Kumar M.","year":"2010","unstructured":"M. Kumar, Benjamin Packer, and Daphne Koller. 2010. Self-paced learning for latent variable models. In Proceedings of Advances in Neural Information Processing Systems, 1189\u20131197."},{"key":"e_1_3_2_37_2","first-page":"2275","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Ma Fan","year":"2017","unstructured":"Fan Ma, Deyu Meng, Qi Xie, Zina Li, and Xuanyi Dong. 2017. Self-paced co-training. In Proceedings of the International Conference on Machine Learning, 2275\u20132284."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_9"},{"key":"e_1_3_2_39_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Liu Hao","year":"2025","unstructured":"Hao Liu, Kun Wang, Yudong Han, Haocong Wang, Yupeng Hu, Chunxiao Wang, and Liqiang Nie. 2025. Curmim: Curriculum masked image modeling. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, 1\u20135."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3243316"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2891895"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00706"},{"key":"e_1_3_2_44_2","first-page":"3854","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zeng Xiangrui","year":"2021","unstructured":"Xiangrui Zeng, Gregory Howe, and Min Xu. 2021. End-to-end robust joint unsupervised image alignment and clustering. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 3854\u20133866."},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00996"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01210"},{"key":"e_1_3_2_47_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.cviu.2016.10.018","article-title":"The THUMOS challenge on action recognition for videos \u201cin the wild\u201d","volume":"155","author":"Idrees Haroon","year":"2017","unstructured":"Haroon Idrees, Amir R. Zamir, Yu-Gang Jiang, Alex Gorban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah. 2017. The THUMOS challenge on action recognition for videos \u201cin the wild\u201d. Computer Vision and Image Understanding 155 (2017), 1\u201323.","journal-title":"Computer Vision and Image Understanding"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_50_2","first-page":"3899","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Liu Ziyi","year":"2019","unstructured":"Ziyi Liu, Le Wang, Qilin Zhang, Zhanning Gao, Zhenxing Niu, Nanning Zheng, and Gang Hua. 2019. Weakly supervised temporal action localization through contrast based evaluation networks. In Proceedings of the IEEE International Conference on Computer Vision, 3899\u20133908."},{"key":"e_1_3_2_51_2","first-page":"2233","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Liu Ziyi","year":"2021","unstructured":"Ziyi Liu, Le Wang, Qilin Zhang, Wei Tang, Junsong Yuan, Nanning Zheng, and Gang Hua. 2021. Acsnet: Action-context separation network for weakly supervised temporal action localization. In Proceedings of the AAAI Conference on Artificial Intelligence, 2233\u20132241."},{"key":"e_1_3_2_52_2","first-page":"1637","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Islam Ashraful","year":"2021","unstructured":"Ashraful Islam, Chengjiang Long, and Richard Radke. 2021. A hybrid attention mechanism for weakly-supervised temporal action localization. In Proceedings of the AAAI Conference on Artificial Intelligence, 1637\u20131645."},{"key":"e_1_3_2_53_2","first-page":"9969","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Luo Wang","year":"2021","unstructured":"Wang Luo, Tianzhu Zhang, Wenfei Yang, Jingen Liu, Tao Mei, Feng Wu, and Yongdong Zhang. 2021. Action unit memory network for weakly supervised temporal action localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 9969\u20139979."},{"key":"e_1_3_2_54_2","first-page":"1854","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Lee Pilhyeon","year":"2021","unstructured":"Pilhyeon Lee, Jinglu Wang, Yan Lu, and Hyeran Byun. 2021. Weakly-supervised temporal action localization by uncertainty modeling. In Proceedings of the AAAI Conference on Artificial Intelligence, 1854\u20131862."},{"key":"e_1_3_2_55_2","first-page":"3319","volume-title":"Proceedings of the IEEE Winter Conference on Applications of Computer Vision","author":"Pardo Alejandro","year":"2021","unstructured":"Alejandro Pardo, Humam Alwassel, Fabian Caba, Ali Thabet, and Bernard Ghanem. 2021. Refineloc: Iterative refinement for weakly-supervised action localization. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision, 3319\u20133328."},{"key":"e_1_3_2_56_2","first-page":"6622","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Tang Xiaojun","year":"2023","unstructured":"Xiaojun Tang, Junsong Fan, Chuanchen Luo, Zhaoxiang Zhang, Man Zhang, and Zongyuan Yang. 2023. DDG-Net: Discriminability-driven graph network for weakly-supervised temporal action localization. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 6622\u20136632."},{"key":"e_1_3_2_57_2","first-page":"1513","volume-title":"Proceedings of the American Association for Artificial Intelligence","author":"Li Zhilin","year":"2023","unstructured":"Zhilin Li, Zilei Wang, and Qinying Liu. 2023. Actionness inconsistency-guided contrastive learning for weakly-supervised temporal action localization. In Proceedings of the American Association for Artificial Intelligence, 1513\u20131521."},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00957"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00237"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i7.28516"},{"issue":"11","key":"e_1_3_2_61_2","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Van der Maaten Laurens","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9, 11 (2008), 2579\u20132605.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/1553374.1553380"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3796708","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:48:39Z","timestamp":1782312519000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3796708"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":61,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3796708"],"URL":"https:\/\/doi.org\/10.1145\/3796708","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-04-06","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}