{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T14:11:23Z","timestamp":1779372683638,"version":"3.53.1"},"reference-count":50,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T00:00:00Z","timestamp":1779321600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2021ZD0111902"],"award-info":[{"award-number":["2021ZD0111902"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["62172022, U21B2038, 62476179"],"award-info":[{"award-number":["62172022, U21B2038, 62476179"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Weakly supervised temporal action localization (WTAL) aims to detect action segments in untrimmed videos only using video-level labels. Existing methods typically follow the multi-instance learning (MIL) paradigm with a top-k strategy, often resulting in incomplete action localization. Moreover, the local and discontinuous nature of actions causes action segments to be isolated and lack sufficient temporal interaction. To address these issues, this article introduces semantic prototypes to enrich video representations, enabling the model to aggregate category-level action cues across videos and recover semantically relevant but weakly activated segments, thereby improving action completeness. A prototype contrastive loss is further employed to improve feature discriminability. Moreover, a sparse temporal interaction unit is designed to jointly model short-term context and long-range dependencies. The boundary-guided loss utilizes the temporal interaction outputs to explicitly constrain semantic responses around action boundaries, promoting sharp and temporally consistent transitions. Based on these, this article proposes a semantic prototype guided sparse temporal interaction network (S2Net), achieving a unified video modeling from full semantic understanding to fine-grained boundary perception. Extensive experiments on THUMOS14 and ActivityNet1.3 demonstrate that S2Net achieves more accurate and complete action localization.<\/jats:p>","DOI":"10.1145\/3807956","type":"journal-article","created":{"date-parts":[[2026,4,15]],"date-time":"2026-04-15T14:52:32Z","timestamp":1776264752000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Semantic Prototype Guided Sparse Temporal Interaction for Weakly Supervised Temporal Action Localization"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7962-4091","authenticated-orcid":false,"given":"Jing","family":"Wang","sequence":"first","affiliation":[{"name":"Beijing Artificial Intelligence Institute, School of Information Science and Technology, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7722-7172","authenticated-orcid":false,"given":"Dehui","family":"Kong","sequence":"additional","affiliation":[{"name":"Beijing Artificial Intelligence Institute, School of Information Science and Technology, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5583-8260","authenticated-orcid":false,"given":"Jinghua","family":"Li","sequence":"additional","affiliation":[{"name":"Beijing Artificial Intelligence Institute, School of Information Science and Technology, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-4355-3169","authenticated-orcid":false,"given":"Xiaochen","family":"Wang","sequence":"additional","affiliation":[{"name":"Beijing Artificial Intelligence Institute, School of Information Science and Technology, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3121-1823","authenticated-orcid":false,"given":"Baocai","family":"Yin","sequence":"additional","affiliation":[{"name":"Beijing Artificial Intelligence Institute, School of Information Science and Technology, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,5,21]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2016.10.018"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2753232"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3321508"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3456795"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1044"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3311447"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2025.3546312"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2916887"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475298"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58601-0_21"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW63382.2024.00276"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV61041.2025.00911"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3213478"},{"key":"e_1_3_1_17_2","unstructured":"Will Kay Joao Carreira Karen Simonyan Brian Zhang Chloe Hillier Sudheendra Vijayanarasimhan Fabio Viola Tim Green Trevor Back Paul Natsev et al. 2017. The kinetics human action video dataset. arXiv:1705.06950. Retrieved from https:\/\/arxiv.org\/abs\/1705.06950"},{"key":"e_1_3_1_18_2","volume-title":"Proceedings of International Conference on Learning Representations","author":"Lee Jun-Tae","year":"2021","unstructured":"Jun-Tae Lee, Mihir Jain, Hyoungwoo Park, and Sungrack Yun. 2021. Cross-attentional audio-visual fusion for weakly-supervised action localization. In Proceedings of International Conference on Learning Representations."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6793"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01026"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3367599"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2024.106905"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3395778"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01771"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-024-01445-2"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00957"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.122000"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3272891"},{"key":"e_1_3_1_29_2","first-page":"689","volume-title":"Proceedings of the 28th International Conference on Machine Learning","author":"Ngiam Jiquan","year":"2011","unstructured":"Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y. Ng. 2011. Multimodal deep learning. In Proceedings of the 28th International Conference on Machine Learning, 689\u2013696."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2024.106788"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3323878"},{"key":"e_1_3_1_32_2","first-page":"8748","volume-title":"Proceedings of the 38th International Conference on Machine Learning","volume":"139","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, Vol. 139. PMLR, 8748\u20138763."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSPW59220.2023.10193217"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2025.121986"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00109"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2024.128688"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2023.109426"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jvcir.2024.104090"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/TBDATA.2025.3528727"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2024.3377468"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3567827"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i3.20216"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i7.28516"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.5555\/1771530.1771554"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1115"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58539-6_3"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-024-06810-6"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3379887"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3234362"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02203"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2024.111207"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3807956","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T13:32:33Z","timestamp":1779370353000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3807956"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,21]]},"references-count":50,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3807956"],"URL":"https:\/\/doi.org\/10.1145\/3807956","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,21]]},"assertion":[{"value":"2025-10-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-16","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}