{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,23]],"date-time":"2026-03-23T16:50:27Z","timestamp":1774284627282,"version":"3.50.1"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"name":"Open Project of the Application Research Center of Smart Energy Technology of Mianyang City College","award":["Grant ZHNY-2025-06"],"award-info":[{"award-number":["Grant ZHNY-2025-06"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62172417 and 62272461"],"award-info":[{"award-number":["62172417 and 62272461"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>\n                    Video Shadow Detection (VSD) is critical yet challenging, primarily due to ambiguous shadow boundaries and the presence of confusing shadow-like non-shadow regions, which existing methods struggle to resolve effectively by limited temporal modeling. We propose the Dual Sparse Long-Short Term Transformer Network (DSLSTT-Net), a novel framework designed to enhance feature learning by integrating robust temporal consistency and detailed local context. DSLSTT-Net utilizes a dual-stream architecture to concurrently process global temporal information and local shadow feature refinement, enabling effective discrimination between true shadows and confusing areas. At its core, the Sparse Long-Short Term Attention Module (Sparse LSTAM) is introduced to efficiently propagate only high-confidence shadow features from memory, significantly enhancing feature discriminability and computational efficiency. Furthermore, an Adaptive Fusion Module (AFM) dynamically merges purified long-term features with short-term details, optimizing final segmentation. Experimental results confirm that DSLSTT-Net significantly outperforms state-of-the-art methods on VSD benchmarks, validating our approach of dual-stream architecture and sparse temporal modeling. The source code is available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/rayyao\/DSLSTTNet\">https:\/\/github.com\/rayyao\/DSLSTTNet<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3796725","type":"journal-article","created":{"date-parts":[[2026,2,23]],"date-time":"2026-02-23T14:05:27Z","timestamp":1771855527000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Dual Sparse Long-Short Term Transformer for Video Shadow Detection"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-8206-8654","authenticated-orcid":false,"given":"Shuo","family":"Han","sequence":"first","affiliation":[{"name":"Application Research Center of Smart Energy Technology of Mianyang City College, Mianyang, China and School of Computer Sciences and Technology\/School of Artificial Intelligence, China University of Mining and Technology, Xuzhou, China, and Mine Digitization Engineering Research Center of the Ministry of Education, Xuzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2734-915X","authenticated-orcid":false,"given":"Rui","family":"Yao","sequence":"additional","affiliation":[{"name":"Application Research Center of Smart Energy Technology of Mianyang City College, Mianyang, China and School of Computer Sciences and Technology\/School of Artificial Intelligence, China University of Mining and Technology, Xuzhou, China, and Mine Digitization Engineering Research Center of the Ministry of Education, Xuzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-8913-5457","authenticated-orcid":false,"given":"Huili","family":"Hao","sequence":"additional","affiliation":[{"name":"Application Research Center of Smart Energy Technology of Mianyang City College, Mianyang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-7148-6690","authenticated-orcid":false,"given":"Qian","family":"Feng","sequence":"additional","affiliation":[{"name":"School of Computer Sciences and Technology\/School of Artificial Intelligence, China University of Mining and Technology, Xuzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5418-9879","authenticated-orcid":false,"given":"Hancheng","family":"Zhu","sequence":"additional","affiliation":[{"name":"School of Computer Sciences and Technology\/School of Artificial Intelligence, China University of Mining and Technology, Xuzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3564-5090","authenticated-orcid":false,"given":"Jiaqi","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Computer Sciences and Technology\/School of Artificial Intelligence, China University of Mining and Technology, Xuzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6207-0299","authenticated-orcid":false,"given":"Yong","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Computer Sciences and Technology\/School of Artificial Intelligence, China University of Mining and Technology, Xuzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,3,23]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3688803"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3571745"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2024.111126"},{"key":"e_1_3_1_5_2","first-page":"14049","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Guo Lanqing","year":"2023","unstructured":"Lanqing Guo, Chong Wang, Wenhan Yang, Siyu Huang, Yufei Wang, Hanspeter Pfister, and Bihan Wen. 2023. ShadowDiffusion: When degradation prior meets diffusion model for shadow removal. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 14049\u201314058."},{"key":"e_1_3_1_6_2","first-page":"5198","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"38","author":"Tao Xinhao","year":"2024","unstructured":"Xinhao Tao, Junyan Cao, Yan Hong, and Li Niu. 2024. Shadow generation with decomposed mask prediction and attentive shadow filling. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 5198\u20135206"},{"key":"e_1_3_1_7_2","first-page":"8121","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu Qingyang","year":"2024","unstructured":"Qingyang Liu, Junqi You, Jianting Wang, Xinhao Tao, Bo Zhang, and Li Niu. 2024. Shadow generation for composite image using diffusion model. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 8121\u20138130."},{"key":"e_1_3_1_8_2","first-page":"23217","volume-title":"Proceedings of the Computer Vision and Pattern Recognition Conference","author":"Wang Xinrui","year":"2025","unstructured":"Xinrui Wang, Lanqing Guo, Xiyu Wang, Siyu Huang, and Bihan Wen. 2025. SoftShadow: Leveraging soft masks for penumbra-aware shadow removal. In Proceedings of the Computer Vision and Pattern Recognition Conference, 23217\u201323226."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2025.3528347"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2023.110085"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2024.3468877"},{"issue":"3","key":"e_1_3_1_12_2","first-page":"3259","article-title":"Instance shadow detection with a single-stage detector","volume":"45","author":"Wang Tianyu","year":"2022","unstructured":"Tianyu Wang, Xiaowei Hu, Pheng-Ann Heng, and Chi-Wing Fu. 2022. Instance shadow detection with a single-stage detector. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 3 (2022), 3259\u20133273.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00007"},{"key":"e_1_3_1_14_2","first-page":"914","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"36","author":"Hong Yan","year":"2022","unstructured":"Yan Hong, Li Niu, and Jianfu Zhang. 2022. Shadow generation for composite image in real-world scenes. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, 914\u2013922."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2024.3447907"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3049331"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2023.3332257"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612482"},{"key":"e_1_3_1_19_2","first-page":"2715","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Chen Zhihao","year":"2021","unstructured":"Zhihao Chen, Liang Wan, Lei Zhu, Jia Shen, Huazhu Fu, Wennan Liu, and Jing Qin. 2021. Triple-cooperative video shadow detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2715\u20132724."},{"key":"e_1_3_1_20_2","unstructured":"Shilin Hu Hieu Le and Dimitris Samaras. 2021. Temporal feature warping for video shadow detection. arXiv:2107.14287. Retrieved from https:\/\/arxiv.org\/abs\/2107.14287"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00312"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548074"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01007"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00932"},{"key":"e_1_3_1_25_2","first-page":"121","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201918)","author":"Zhu Lei","year":"2018","unstructured":"Lei Zhu, Zijun Deng, Xiaowei Hu, Chi-Wing Fu, Xuemiao Xu, Jing Qin, and Pheng-Ann Heng. 2018. Bidirectional feature pyramid network with recurrent attention residual modules for shadow detection. In Proceedings of the European Conference on Computer Vision (ECCV \u201918), 121\u2013136."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00531"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3283416"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3571745"},{"key":"e_1_3_1_29_2","first-page":"5611","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Zhihao","year":"2020","unstructured":"Zhihao Chen, Lei Zhu, Liang Wan, Song Wang, Wei Feng, and Pheng-Ann Heng. 2020. A multi-task mean teacher for semi-supervised shadow detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 5611\u20135620."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2025.3559891"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2022.3148707"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3104166"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2867733"},{"issue":"5","key":"e_1_3_1_34_2","doi-asserted-by":"crossref","first-page":"3755","DOI":"10.1109\/TCSVT.2023.3319330","article-title":"MC-Blur: A comprehensive benchmark for image deblurring","volume":"34","author":"Zhang Kaihao","year":"2023","unstructured":"Kaihao Zhang, Tao Wang, Wenhan Luo, Wenqi Ren, Bj\u00f6rn Stenger, Wei Liu, Hongdong Li, and Ming-Hsuan Yang. 2023. MC-Blur: A comprehensive benchmark for image deblurring. IEEE Transactions on Circuits and Systems for Video Technology 34, 5 (2023), 3755\u20133767.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2008.916989"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2009.2012924"},{"key":"e_1_3_1_37_2","first-page":"705","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Ding Xinpeng","year":"2022","unstructured":"Xinpeng Ding, Jingwen Yang, Xiaowei Hu, and Xiaomeng Li. 2022. Learning shadow correspondence for video shadow detection. In Proceedings of the European Conference on Computer Vision. Springer, 705\u2013722."},{"key":"e_1_3_1_38_2","first-page":"1","volume-title":"Proceedings of the 18th ACM SIGGRAPH International Conference on Virtual-Reality Continuum and Its Applications in Industry","author":"Lin Junhao","year":"2022","unstructured":"Junhao Lin and Liansheng Wang. 2022. Spatial-temporal fusion network for fast video shadow detection. In Proceedings of the 18th ACM SIGGRAPH International Conference on Virtual-Reality Continuum and Its Applications in Industry, 1\u20135."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3320688"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681236"},{"key":"e_1_3_1_41_2","first-page":"196","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Duan Xin","year":"2025","unstructured":"Xin Duan, Yu Cao, Lei Zhu, Gang Fu, Xin Wang, Renjie Zhang, and Ping Li. 2025. Two-stage video shadow detection via temporal-spatial adaption. In Proceedings of the European Conference on Computer Vision, 196\u2013214."},{"key":"e_1_3_1_42_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017), 5998\u20136008.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_43_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01212"},{"key":"e_1_3_1_45_2","first-page":"16794","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Peng Yansong","year":"2024","unstructured":"Yansong Peng, Hebei Li, Yueyi Zhang, Xiaoyan Sun, and Feng Wu. 2024. Scene adaptive sparse transformer for event-based object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 16794\u201316804."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02279"},{"key":"e_1_3_1_47_2","doi-asserted-by":"crossref","first-page":"111029","DOI":"10.1016\/j.patcog.2024.111029","article-title":"Distilling efficient vision transformers from CNNs for semantic segmentation","volume":"158","author":"Zheng Xu","year":"2025","unstructured":"Xu Zheng, Yunhao Luo, Pengyuan Zhou, and Lin Wang. 2025. Distilling efficient vision transformers from CNNs for semantic segmentation. Pattern Recognition 158, (2025), 111029.","journal-title":"Pattern Recognition"},{"key":"e_1_3_1_48_2","first-page":"12077","article-title":"SegFormer: Simple and efficient design for semantic segmentation with transformers","volume":"34","author":"Xie Enze","year":"2021","unstructured":"Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M. Alvarez, and Ping Luo. 2021. SegFormer: Simple and efficient design for semantic segmentation with transformers. Advances in Neural Information Processing Systems 34 (2021), 12077\u201312090.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_49_2","first-page":"15773","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Norouzi Narges","year":"2024","unstructured":"Narges Norouzi, Svetlana Orlova, Daan De Geus, and Gijs Dubbelman. 2024. ALGM: Adaptive local-then-global token merging for efficient semantic segmentation with plain vision transformers. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 15773\u201315782."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01172"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00061"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00352"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00353"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2024.111637"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.106"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.563"},{"key":"e_1_3_1_62_2","first-page":"684","volume-title":"Proceedings of the 27th International Joint Conference on Artificial Intelligence","volume":"684690","author":"Deng Zijun","year":"2018","unstructured":"Zijun Deng, Xiaowei Hu, Lei Zhu, Xuemiao Xu, Jing Qin, Guoqiang Han, and Pheng-Ann Heng. 2018. R3Net: Recurrent residual refinement network for saliency detection. In Proceedings of the 27th International Joint Conference on Artificial Intelligence, Vol. 684690. AAAI Press, 684\u2013690."},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00374"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00971"},{"key":"e_1_3_1_65_2","first-page":"11781","article-title":"Rethinking space-time networks with improved memory coverage for efficient video object segmentation","volume":"34","author":"Cheng Ho Kei","year":"2021","unstructured":"Ho Kei Cheng, Yu-Wing Tai, and Chi-Keung Tang. 2021. Rethinking space-time networks with improved memory coverage for efficient video object segmentation. Advances in Neural Information Processing Systems 34, (2021), 11781\u201311794.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00090"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3796725","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,23]],"date-time":"2026-03-23T15:51:04Z","timestamp":1774281064000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3796725"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,23]]},"references-count":65,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3796725"],"URL":"https:\/\/doi.org\/10.1145\/3796725","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,23]]},"assertion":[{"value":"2025-09-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-17","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-23","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}