{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T14:22:01Z","timestamp":1780928521555,"version":"3.54.1"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T00:00:00Z","timestamp":1780876800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"Shenzhen Science and Technology Program","award":["SYSPG20241211173951079"],"award-info":[{"award-number":["SYSPG20241211173951079"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    Weakly supervised video anomaly detection (WVAD) aims to locate events or behaviors that deviate from normal patterns in untrimmed videos using video-level labels. Recent studies typically utilize supplementary modalities to assist anomaly detection. However, these methods suffer from two main issues: (1) The limitations of long-duration anomaly event temporal modeling. The model struggles to consistently maintain key information, resulting in the forgetting phenomenon, which affects the tracking of the event\u2019s overall dynamic evolution and complicates anomaly event analysis and understanding. (2) The multi-modal fusion strategy is insufficient, particularly when there is temporal inconsistency between visual and audio information, causing the model to overlook key information, directly affecting the accurate detection and recognition of anomalous events. To address these issues, we propose a visual-guided long-term temporal context learning network (LTCLNet). The network consists of three key components: a cross-modal interaction module, a multi-modal fusion module, and a visual-guided parameter optimization strategy. First, to address the forgetting issue in long-duration anomaly detection, we designed a cross-modal interaction module. The key part of this module is the establishment of a cross-matrix mechanism. This mechanism achieves bidirectional temporal guidance across modalities. It allows the temporal modeling of each modality to dynamically integrate information from the other modality. This enables the model to continuously track the dynamic evolution of the event. The tracking is facilitated through shared temporal information between the visual and audio modalities. Secondly, to fully exploit the complementary characteristics between different modalities, we introduced a novel temporal reversal integration method in the multi-modal fusion module. This method reverses the feature sequences of each modality to enhance the model\u2019s perception of temporal dynamic changes. By fusing the modality features before and after reversal, the shared temporal structure between modalities is strengthened, improving the model\u2019s ability to capture anomalous information. Additionally, our proposed visual-guided parameter optimization strategy trains a parallel visual modality network as a semantic anchor, ensuring that the model stays aligned with a semantically stable and structurally clear visual flow during the learning process, thus ensuring stability and semantic coherence in the training. Extensive experiments on datasets such as XD-Violence demonstrate that our method significantly outperforms existing approaches, particularly achieving notable improvements in the accuracy and stability of long-term anomaly detection. Our code is publicly available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/ibliever\/LTCLNet\">https:\/\/github.com\/ibliever\/LTCLNet<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3810186","type":"journal-article","created":{"date-parts":[[2026,5,16]],"date-time":"2026-05-16T14:19:31Z","timestamp":1778941171000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Visual-Guided Long Temporal Context Learning Network for Weakly Supervised Video Anomaly Detection"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7233-5318","authenticated-orcid":false,"given":"Hao","family":"Zhou","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, National Engineering Research Center of Digital Life, Sun Yat-sen University, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2712-4412","authenticated-orcid":false,"given":"Ruomei","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Software Engineering, Sun Yat-sen University, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-7788-9589","authenticated-orcid":false,"given":"Linxuan","family":"Han","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, National Engineering Research Center of Digital Life, Sun Yat-sen University, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0031-7721","authenticated-orcid":false,"given":"Baoquan","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, Sun Yat-sen University, Guangzhou, China and Research Institute of Sun Yat-sen University in Shenzhen, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0400-9366","authenticated-orcid":false,"given":"Fan","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, National Engineering Research Center of Digital Life, Sun Yat-sen University, Guangzhou, China and Research Institute of Sun Yat-sen University in Shenzhen, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,8]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.10.044"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3112814"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01734"},{"key":"e_1_3_1_6_2","first-page":"428","volume-title":"Proceedings of the 9th European Conference on Computer Vision","author":"Dalal Navneet","year":"2006","unstructured":"Navneet Dalal, Bill Triggs, and Cordelia Schmid. 2006. Human detection using oriented histograms of flow and appearance. In Proceedings of the 9th European Conference on Computer Vision, 428\u2013441."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3350084"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2025.126497"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCE.2024.3482560"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952261"},{"key":"e_1_3_1_11_2","first-page":"1965","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"Ghadiya Ayush","year":"2024","unstructured":"Ayush Ghadiya, Purbayan Kar, Vishal Chudasama, and Pankaj Wasnik. 2024. Cross-modal fusion and attention mechanism for weakly supervised video anomaly detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, 1965\u20131974."},{"key":"e_1_3_1_12_2","volume-title":"Proceedings of the 1st Conference on Language Modeling","author":"Gu Albert","year":"2024","unstructured":"Albert Gu and Tri Dao. 2024. Mamba: Linear-time sequence modeling with selective state spaces. In Proceedings of the 1st Conference on Language Modeling. Retrieved from https:\/\/openreview.net\/forum?id=tEYskw1VY2"},{"issue":"6","key":"e_1_3_1_13_2","doi-asserted-by":"crossref","first-page":"317","DOI":"10.1007\/s00530-024-01525-3","article-title":"Hierarchical bi-directional conceptual interaction for text-video retrieval","volume":"30","author":"Han Wenpeng","year":"2024","unstructured":"Wenpeng Han, Guanglin Niu, Mingliang Zhou, and Xiaowei Zhang. 2024. Hierarchical bi-directional conceptual interaction for text-video retrieval. Multimedia Systems 30, 6 (2024), 317.","journal-title":"Multimedia Systems"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2026.114765"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.86"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952132"},{"key":"e_1_3_1_17_2","first-page":"6848","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Karim Hamza","year":"2024","unstructured":"Hamza Karim, Keval Doshi, and Yasin Yilmaz. 2024. Real-time weakly supervised video anomaly detection. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, 6848\u20136856."},{"key":"e_1_3_1_18_2","first-page":"1012","volume-title":"Proceedings of the 2022 26th International Conference on Pattern Recognition (ICPR)","author":"Lee Jooyeon","year":"2022","unstructured":"Jooyeon Lee, Woo-Jeoung Nam, and Seong-Whan Lee. 2022. Multi-contextual predictions with vision transformer for video anomaly detection. In Proceedings of the 2022 26th International Conference on Pattern Recognition (ICPR). IEEE, 1012\u20131018."},{"issue":"5","key":"e_1_3_1_19_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3721981","article-title":"Towards energy-efficient audio-visual classification via multimodal interactive spiking neural network","volume":"21","author":"Liu Xu","year":"2025","unstructured":"Xu Liu, Na Xia, Jinxing Zhou, Zhangbin Li, and Dan Guo. 2025. Towards energy-efficient audio-visual classification via multimodal interactive spiking neural network. ACM Transactions on Multimedia Computing, Communications, and Applications 21, 5, Article 144 (May 2025), 1\u201324.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/tcsii.2022.3161049"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/iccv.2013.338"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00775"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3072863"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr.2009.5206641"},{"key":"e_1_3_1_25_2","first-page":"2260","volume-title":"Proceedings of the ICASSP 2021\u20132021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Pang Wen-Feng","year":"2021","unstructured":"Wen-Feng Pang, Qian-Hua He, Yong-Jian Hu, and Yan-Xiong Li. 2021. Violence detection in videos based on fusing visual and audio information. In Proceedings of the ICASSP 2021\u20132021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2260\u20132264."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV56688.2023.00269"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3662185"},{"key":"e_1_3_1_28_2","first-page":"219","volume-title":"Proceedings of the 2nd International Conference on Consumer Electronics and Computer Engineering (ICCECE)","author":"Pu Yujiang","year":"2022","unstructured":"Yujiang Pu and Xiaoyu Wu. 2022. Audio-guided attention network for weakly supervised violence detection. In Proceedings of the 2nd International Conference on Consumer Electronics and Computer Engineering (ICCECE). IEEE, 219\u2013223."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2024.3451935"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr.2012.6247917"},{"key":"e_1_3_1_31_2","volume-title":"Advances in Neural Information Processing Systems","author":"Sch\u00f6lkopf Bernhard","year":"1999","unstructured":"Bernhard Sch\u00f6lkopf, Robert C. Williamson, Alex Smola, John Shawe-Taylor, and John Platt. 1999. Support vector method for novelty detection. In Advances in Neural Information Processing Systems, Vol. 12."},{"issue":"12","key":"e_1_3_1_32_2","doi-asserted-by":"crossref","first-page":"12638","DOI":"10.1109\/TCSVT.2024.3435003","article-title":"Matching multi-scale feature sets in vision transformer for few-shot classification","volume":"34","author":"Song Mingchen","year":"2024","unstructured":"Mingchen Song, Fengqin Yao, Guoqiang Zhong, Zhong Ji, and Xiaowei Zhang. 2024. Matching multi-scale feature sets in vision transformer for few-shot classification. IEEE Transactions on Circuits and Systems for Video Technology 34, 12 (2024), 12638\u201312651.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00678"},{"key":"e_1_3_1_34_2","unstructured":"Shengyang Sun and Xiaojin Gong. 2024. Multi-scale bottleneck transformer for weakly supervised multimodal violence detection. arXiv:2405.05130. Retrieved from https:\/\/arxiv.org\/abs\/2405.05130"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00493"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.120599"},{"issue":"11","key":"e_1_3_1_37_2","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Maaten Laurens Van der","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9, 11 (2008), 2579\u20132605.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-012-0594-8"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2022.3216500"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2636150"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19778-9_42"},{"key":"e_1_3_1_42_2","first-page":"322","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201920)","author":"Wu Peng","year":"2020","unstructured":"Peng Wu, Jing Liu, Yujia Shi, Yujia Sun, Fangtao Shao, Zhaoyang Wu, and Zhiwei Yang. 2020. Not only look, but also listen: Learning multimodal violence detection under weak supervision. In Proceedings of the European Conference on Computer Vision (ECCV \u201920). Springer, 322\u2013339."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i6.28423"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3419842"},{"key":"e_1_3_1_45_2","first-page":"18899","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yang Zhiwei","year":"2024","unstructured":"Zhiwei Yang, Jing Liu, and Peng Wu. 2024. Text prompt with normality guidance for weakly supervised video anomaly detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 18899\u201318908."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547868"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01561"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2025.3553556"},{"key":"e_1_3_1_49_2","first-page":"17385","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Menghao","year":"2024","unstructured":"Menghao Zhang, Jingyu Wang, Qi Qi, Haifeng Sun, Zirui Zhuang, Pengfei Ren, Ruilong Ma, and Jianxin Liao. 2024. Multi-scale video anomaly detection by multi-grained spatio-temporal representation learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 17385\u201317394."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2025.3612592"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr.2011.5995524"},{"key":"e_1_3_1_52_2","first-page":"3769","volume-title":"Proceedings of the 37th AAAI Conference on Artificial Intelligence","volume":"37","author":"Zhou Hang","year":"2023","unstructured":"Hang Zhou, Junqing Yu, and Wei Yang. 2023. Dual memory units with uncertainty regulation for weakly supervised video anomaly detection. In Proceedings of the 37th AAAI Conference on Artificial Intelligence, Vol. 37, 3769\u20133777."},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2024.105286"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3450734"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3810186","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T13:34:51Z","timestamp":1780925691000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3810186"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,8]]},"references-count":53,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3810186"],"URL":"https:\/\/doi.org\/10.1145\/3810186","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,8]]},"assertion":[{"value":"2025-04-02","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-05","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}