{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T17:35:56Z","timestamp":1783013756248,"version":"3.54.6"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"7","funder":[{"name":"Beijing Natural Science Foundation","award":["4252026"],"award-info":[{"award-number":["4252026"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62203024, 62573369, 62261160576"],"award-info":[{"award-number":["62203024, 62573369, 62261160576"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Research and Development Program of Beijing Municipal Education Commission","award":["KM202310005027"],"award-info":[{"award-number":["KM202310005027"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>\n                    The general-purpose\n                    <jats:bold>Segment Anything Model (SAM)<\/jats:bold>\n                    is limited by the inherent constraints of RGB sensors, which render it inadequate for challenging real-world scenarios such as adverse lighting conditions and rapid motion. In contrast, event cameras, a novel type of bio-inspired visual sensor, offer distinct imaging advantages, including high temporal resolution and a high dynamic range. The event streams generated by these cameras provide spatiotemporal dynamic cues that are often absent in conventional image frames. To overcome the limitations of RGB-based models, we propose\n                    <jats:bold>SAM with Event-based Assistance (EvSAM)<\/jats:bold>\n                    , a novel RGB-event multi-modal semantic segmentation framework. EvSAM leverages the strong generalization capabilities of SAM while incorporating the complementary characteristics of event data to enhance scene comprehension, particularly under adverse conditions. To address the challenges of fusing two modals (image and event) with large data format discrepancy, we introduce two core components: the\n                    <jats:bold>\n                      Multi-spatiotemporal-scale Patch Alignment Block (MS\n                      <jats:sup>2<\/jats:sup>\n                      PAB)\n                    <\/jats:bold>\n                    and the\n                    <jats:bold>Event-based Feature Injector (EFInj)<\/jats:bold>\n                    for SAM. Specifically, the MS\n                    <jats:inline-formula content-type=\"math\/tex\">\n                      <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\({}^{2}\\)<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    PAB captures spatiotemporal semantic coherence from the event stream and transforms it into a frame-based complementary representation using a multi-spatiotemporal alignment strategy. The EFInj introduces a dynamic event feature update mechanism, wherein the fused features at a given layer guide the adaptive generation of deeper event representations. This process facilitates the integration of RGB spatial semantics with event-based motion cues. Owing to these core designs, EvSAM demonstrates superior performance on event-based semantic segmentation datasets, thereby fully validating its distinct advantages in handling extreme visual scenarios. Furthermore, we extend our model to the task of depth estimation, which further demonstrates its strong generalization ability and scalability for various downstream applications.\n                  <\/jats:p>","DOI":"10.1145\/3786794","type":"journal-article","created":{"date-parts":[[2026,1,3]],"date-time":"2026-01-03T13:39:08Z","timestamp":1767447548000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["EvSAM: Segment Anything Model with Event-based Assistance"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-7221-6453","authenticated-orcid":false,"given":"Yi","family":"Ding","sequence":"first","affiliation":[{"name":"College of Computer Science, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-4819-9976","authenticated-orcid":false,"given":"Bowen","family":"Yao","sequence":"additional","affiliation":[{"name":"College of Computer Science, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-4748-1639","authenticated-orcid":false,"given":"Yuhan","family":"Liu","sequence":"additional","affiliation":[{"name":"Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3138-505X","authenticated-orcid":false,"given":"Hao","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Southeast University, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6597-3725","authenticated-orcid":false,"given":"Ding","family":"Ding","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Southeast University, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6058-0217","authenticated-orcid":false,"given":"Zhen","family":"Yang","sequence":"additional","affiliation":[{"name":"College of Computer Science, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5227-1326","authenticated-orcid":false,"given":"Youfu","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Mechanical Engineering, City University of Hong Kong, Hong Kong, SAR"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6253-3564","authenticated-orcid":false,"given":"Yongjian","family":"Deng","sequence":"additional","affiliation":[{"name":"Computer Science, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"1624","volume-title":"Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)","author":"Alonso I\u00f1igo","year":"2018","unstructured":"I\u00f1igo Alonso and Ana Cristina Murillo. 2018. EV-SegNet: Semantic segmentation for event-based cameras. In Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 1624\u20131633. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:54063435"},{"key":"e_1_3_1_3_2","unstructured":"Hangbo Bao Li Dong and Furu Wei. 2021. BEiT: BERT pre-training of image transformers. arXiv:2106.08254. Retrieved from https:\/\/arxiv.org\/abs\/2106.08254"},{"key":"e_1_3_1_4_2","unstructured":"Jonathan Binas Daniel Neil Shih-Chii Liu and Tobi Delbr\u00fcck. 2017. DDD17: End-to-end DAVIS driving dataset. arXiv:1711.01458. Retrieved from https:\/\/arxiv.org\/abs\/1711.01458"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2699184"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","unstructured":"Shoufa Chen Chongjian Ge Zhan Tong Jiangliu Wang Yibing Song Jue Wang and Ping Luo.2022. AdaptFormer: Adapting vision transformers for scalable visual recognition. arXiv:2205.13535. Retrieved from https:\/\/arxiv.org\/abs\/2205.13535","DOI":"10.52202\/068431-1212"},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","unstructured":"Tianrun Chen Ankang Lu Lanyun Zhu Chaotao Ding Chunan Yu Deyi Ji Zejian Li Lingyun Sun Papa Mao and Ying Zang. 2024. SAM2-Adapter: Evaluating & adapting segment anything 2 in downstream tasks: Camouflage shadow medical image segmentation and more. arXiv:2408.04579. Retrieved from https:\/\/arxiv.org\/abs\/2408.04579","DOI":"10.21203\/rs.3.rs-4876632\/v1"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW60793.2023.00361"},{"key":"e_1_3_1_11_2","unstructured":"Zhe Chen Yuchen Duan Wenhai Wang Junjun He Tong Lu Jifeng Dai and Yu Qiao. 2022. 2022. Vision transformer adapter for dense predictions. arXiv:2205.08534. Retrieved from https:\/\/arxiv.org\/abs\/2205.08534"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2021.3060707"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0925-2312(99)00095-8"},{"key":"e_1_3_1_14_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2021. An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations, 2021. Retrieved from https:\/\/openreview.net\/pdf?id=YicbFdNTTy"},{"key":"e_1_3_1_15_2","first-page":"21056","article-title":"Deep residual learning in spiking neural networks","volume":"34","author":"Fang Wei","year":"2021","unstructured":"Wei Fang, Zhaofei Yu, Yanqing Chen, Tiejun Huang, Timoth\u00e9e Masquelier, and Yonghong Tian. 2021. Deep residual learning in spiking neural networks. In Advance Neural Information Processing Systems, Vol. 34, 21056\u201321069. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:235359262","journal-title":"Advance Neural Information Processing Systems"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00364"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.5555\/2354409.2354978"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-0084"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2025.103652"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48891.2023.10161009"},{"key":"e_1_3_1_22_2","volume-title":"Proceedings of the 36th International Conference on Machine Learning","author":"Houlsby Neil","year":"2019","unstructured":"Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:59599816"},{"key":"e_1_3_1_23_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Hu Edward J.","year":"2022","unstructured":"Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2022. LoRA: Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=nZeVKeeFYf9"},{"key":"e_1_3_1_24_2","volume-title":"Proceedings of the IEEE International Conference on 3D Vision (3DV","author":"Gehrig Javier Hidalgo-Carrio Daniel","year":"2020","unstructured":"Daniel Gehrig Javier Hidalgo-Carrio and Davide Scaramuzza. 2020. Learning monocular dense depth from events. In Proceedings of the IEEE International Conference on 3D Vision (3DV). Retrieved from http:\/\/rpg.ifi.uzh.ch\/docs\/3DV20_Hidalgo.pdf"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3249579"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA46639.2022.9811767"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681148"},{"key":"e_1_3_1_29_2","unstructured":"Patrick Lichtsteiner Christoph Posch and Tobi Delbruck. 2006. A 128 \u00d7 128 120 dB 15 \u03bcs Latency Asynchronous Temporal Contrast Vision Sensor. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:2497402"},{"key":"e_1_3_1_30_2","first-page":"10672","volume-title":"Proceedings of the 2022 International Conference on Robotics and Automation (ICRA)","author":"Lin Jiarong","year":"2021","unstructured":"Jiarong Lin and Fu Zhang. 2021. R3LIVE: A robust, real-time, RGB-colored, LiDAR-inertial-visual tightly-coupled state estimation and mapping package. In Proceedings of the 2022 International Conference on Robotics and Automation (ICRA), 10672\u201310678. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:237532664"},{"issue":"2","key":"e_1_3_1_31_2","first-page":"1","article-title":"Dynamic multimodal fusion via meta-learning towards micro-video recommendation","volume":"42","author":"Liu Han","year":"2023","unstructured":"Han Liu, Yinwei Wei, Fan Liu, Wenjie Wang, Liqiang Nie, and Tat-Seng Chua.2023. Dynamic multimodal fusion via meta-learning towards micro-video recommendation. ACM Transactions on Information Systems 42, 2 (2023), 1\u201326.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01862"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3378742"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA46639.2022.9812447"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00070"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA57147.2024.10610700"},{"key":"e_1_3_1_37_2","unstructured":"Xinyang Pu Hecheng Jia Linghao Zheng Feng Wang and Feng Xu. 2024. ClassWise-SAM-Adapter: Parameter efficient fine-tuning adapts segment anything to SAR domain for semantic segmentation. arXiv:2401.02326. Retrieved from https:\/\/arxiv.org\/abs\/2401.02326"},{"key":"e_1_3_1_38_2","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:231591445"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01196"},{"key":"e_1_3_1_40_2","unstructured":"Nikhila Ravi Valentin Gabeur Yuan-Ting Hu Ronghang Hu Chaitanya Ryali Tengyu Ma Haitham Khedr Roman R\u00e4dle Chloe Rolland Laura Gustafson et al. 2024. SAM 2: Segment anything in images and videos. arXiv:2408.00714. Retrieved from https:\/\/arxiv.org\/abs\/2408.00714"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2963386"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3311336"},{"key":"e_1_3_1_43_2","first-page":"3431","volume-title":"Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Shelhamer Evan","year":"2014","unstructured":"Evan Shelhamer, Jonathan Long, and Trevor Darrell. 2014. Fully convolutional networks for semantic segmentation. In Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3431\u20133440. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:1629541"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19830-4_20"},{"key":"e_1_3_1_45_2","unstructured":"Andrew Tao Karan Sapra and Bryan Catanzaro. 2020. Hierarchical multi-scale attention for semantic segmentation. arXiv:2005.10821. Retrieved from https:\/\/arxiv.org\/abs\/2005.10821"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01723"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2025.103547"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/3DV57658.2022.00052"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3380255"},{"key":"e_1_3_1_50_2","first-page":"12077","article-title":"SegFormer: Simple and efficient design for semantic segmentation with transformers","volume":"34","author":"Xie Enze","year":"2021","unstructured":"Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M. Alvarez, and Ping Luo. 2021. SegFormer: Simple and efficient design for semantic segmentation with transformers. In Advance Neural Information Processing Systems, Vol. 34, 12077\u201312090.","journal-title":"Advance Neural Information Processing Systems"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA57147.2024.10611127"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3746027.3755566"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2023.3300537"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00116"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2021.3134828"},{"key":"e_1_3_1_56_2","doi-asserted-by":"crossref","unstructured":"Kaidong Zhang and Dong Liu. 2023. Customized segment anything model for medical image segmentation. arXiv:2304.13785. Retrieved from https:\/\/arxiv.org\/abs\/2304.13785","DOI":"10.2139\/ssrn.4495221"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9747832"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i10.33141"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8794167"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48891.2023.10161563"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00108"},{"key":"e_1_3_1_62_2","doi-asserted-by":"crossref","unstructured":"Zhiyu Zhu Junhui Hou and Dapeng Oliver Wu. 2023. Cross-modal orthogonal high-rank augmentation for RGB-event transformer-trackers. arXiv:2307.04129. Retrieved from https:\/\/arxiv.org\/abs\/2307.04129","DOI":"10.1109\/ICCV51070.2023.02015"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3786794","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:42:13Z","timestamp":1782312133000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3786794"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":61,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3786794"],"URL":"https:\/\/doi.org\/10.1145\/3786794","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-06-20","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}