{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T21:15:39Z","timestamp":1779311739538,"version":"3.51.4"},"reference-count":60,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:00:00Z","timestamp":1750291200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:00:00Z","timestamp":1750291200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Intell."],"published-print":{"date-parts":[[2025,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>This study investigates the application and performance of the Segment Anything Model\u00a02 (SAM2) in the challenging task of video camouflaged object segmentation (VCOS). VCOS involves detecting objects that blend seamlessly in the surroundings for videos due to similar colors and textures and poor light conditions. Compared to the objects in normal scenes, camouflaged objects are much more difficult to detect. SAM2, a video foundation model, has shown potential in various tasks. However, its effectiveness in dynamic camouflaged scenarios remains under-explored. This study presents a comprehensive study on SAM2\u2019s ability in VCOS. First, we assess SAM2\u2019s performance on camouflaged video datasets using different models and prompts (click, box, and mask). Second, we explore the integration of SAM2 with existing multimodal large language models (MLLMs) and VCOS methods. Third, we specifically adapt SAM2 by fine-tuning it on the video camouflaged dataset. Our comprehensive experiments demonstrate that SAM2 has the excellent zero-shot ability to detect camouflaged objects in videos. We also show that this ability could be further improved by specifically adjusting SAM2\u2019s parameters for VCOS.<\/jats:p>","DOI":"10.1007\/s44267-025-00082-1","type":"journal-article","created":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T08:22:45Z","timestamp":1750321365000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["When SAM2 meets video camouflaged object segmentation: a comprehensive evaluation and adaptation"],"prefix":"10.1007","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4584-1893","authenticated-orcid":false,"given":"Yuli","family":"Zhou","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8667-9656","authenticated-orcid":false,"given":"Guolei","family":"Sun","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8948-7892","authenticated-orcid":false,"given":"Yawei","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5487-9845","authenticated-orcid":false,"given":"Guo-Sen","family":"Xie","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8068-3806","authenticated-orcid":false,"given":"Luca","family":"Benini","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2542-3611","authenticated-orcid":false,"given":"Ender","family":"Konukoglu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,6,19]]},"reference":[{"key":"82_CR1","first-page":"3431","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"J. Long","year":"2015","unstructured":"Long, J., Shelhamer, E., & Darrell, T. (2015). Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3431\u20133440). Piscataway: IEEE."},{"key":"82_CR2","first-page":"2961","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"K. He","year":"2017","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., & Girshick, R. (2017). Mask R-CNN. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 2961\u20132969). Piscataway: IEEE."},{"issue":"10","key":"82_CR3","doi-asserted-by":"publisher","first-page":"4651","DOI":"10.1007\/s11263-024-02112-9","volume":"132","author":"P. Huang","year":"2024","unstructured":"Huang, P., Zhang, D., Cheng, D., Han, L., Zhu, P., & Han, J. (2024). M-RRFS: a memory-based robust region feature synthesizer for zero-shot object detection. International Journal of Computer Vision, 132(10), 4651\u20134672.","journal-title":"International Journal of Computer Vision"},{"issue":"10","key":"82_CR4","doi-asserted-by":"publisher","first-page":"6919","DOI":"10.1109\/TPAMI.2024.3387326","volume":"46","author":"G. Sun","year":"2024","unstructured":"Sun, G., Liu, Y., Ding, H., Wu, M., & Van Gool, L. (2024). Learning local and global temporal contexts for video semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(10), 6919\u20136934.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"82_CR5","first-page":"2774","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"D.-P. Fan","year":"2020","unstructured":"Fan, D.-P., Ji, G.-P., Sun, G., Cheng, M.-M., Shen, J., & Shao, L. (2020). Camouflaged object detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 2774\u20132784). Piscataway: IEEE."},{"key":"82_CR6","first-page":"13864","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"X. Cheng","year":"2022","unstructured":"Cheng, X., Xiong, H., Fan, D.-P., Zhong, Y., Harandi, M., Drummond, T., & Ge, Z. (2022). Implicit motion handling for video camouflaged object detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 13864\u201313873). Piscataway: IEEE."},{"issue":"12","key":"82_CR7","doi-asserted-by":"publisher","first-page":"9205","DOI":"10.1109\/TPAMI.2024.3417329","volume":"46","author":"Y. Pang","year":"2024","unstructured":"Pang, Y., Zhao, X., Xiang, T.-Z., Zhang, L., & Lu, H. (2024). ZoomNeXt: a unified collaborative pyramid network for camouflaged object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12), 9205\u20139220.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"82_CR8","first-page":"162","volume-title":"Proceedings of the Asian conference on computer vision","author":"J. Xie","year":"2024","unstructured":"Xie, J., Yang, C., Xie, W., & Zisserman, A. (2024). Moving object segmentation: all you need is SAM (and flow). In Proceedings of the Asian conference on computer vision (pp. 162\u2013178). Cham: Springer."},{"key":"82_CR9","unstructured":"Ravi, N., Gabeur, V., Hu, Y.-T., Hu, R., Ryali, C., Ma, T., Khedr, H., R\u00e4dle, R., Rolland, C., Gustafson, L., et\u00a0al. (2024). SAM 2: segment anything in images and videos. arXiv preprint. arXiv:2408.00714."},{"key":"82_CR10","first-page":"4015","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"A. Kirillov","year":"2023","unstructured":"Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al. (2023). Segment anything. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 4015\u20134026). Piscataway: IEEE."},{"key":"82_CR11","first-page":"433","volume-title":"Proceedings of the 14th European conference on computer vision","author":"P. Bideau","year":"2016","unstructured":"Bideau, P., & Learned-Miller, E. (2016). It\u2019s moving! A probabilistic model for causal motion segmentation in moving camera videos. In B. Leibe, J. Matas, N. Sebe, & M. Welling (Eds.), Proceedings of the 14th European conference on computer vision (pp. 433\u2013449). Cham: Springer."},{"key":"82_CR12","first-page":"11591","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Lv","year":"2021","unstructured":"Lv, Y., Zhang, J., Dai, Y., Li, A., Liu, B., Barnes, N., & Fan, D.-P. (2021). Simultaneously localize, segment and rank the camouflaged objects. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 11591\u201311601). Piscataway: IEEE."},{"key":"82_CR13","first-page":"13791","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"G. Sun","year":"2023","unstructured":"Sun, G., An, Z., Liu, Y., Liu, C., Sakaridis, C., Fan, D.-P., & Van Gool, L. (2023). Indiscernible object counting in underwater scenes. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 13791\u201313801). Piscataway: IEEE."},{"key":"82_CR14","first-page":"19","volume-title":"Proceedings of the 17th European conference on computer vision","author":"J. Pei","year":"2022","unstructured":"Pei, J., Cheng, T., Fan, D.-P., Tang, H., Chen, C., & Van Gool, L. (2022). OSFormer: one-stage camouflaged instance segmentation with transformers. In S. Avidan, G. Brostow, M. Ciss\u00e9, G. M. Farinella, & T. Hassner (Eds.), Proceedings of the 17th European conference on computer vision (pp. 19\u201337). Cham: Springer."},{"key":"82_CR15","first-page":"17169","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Z. Luo","year":"2024","unstructured":"Luo, Z., Liu, N., Zhao, W., Yang, X., Zhang, D., Fan, D.-P., Khan, F., & Han, J. (2024). VSCode: general visual salient and camouflaged object detection with 2D prompt learning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 17169\u201317180). Piscataway: IEEE."},{"key":"82_CR16","unstructured":"Qin, X., Fan, D.-P., Huang, C., Diagne, C., Zhang, Z., Sant\u2019Anna, A. C., Suarez, A., Jagersand, M., & Shao, L. (2021). Boundary-aware segmentation network for mobile and web applications. arXiv preprint. arXiv:2101.04704."},{"key":"82_CR17","doi-asserted-by":"publisher","first-page":"5936","DOI":"10.1109\/TIP.2024.3475219","volume":"33","author":"S. Yao","year":"2024","unstructured":"Yao, S., Sun, H., Xiang, T.-Z., Wang, X., & Cao, X. (2024). Hierarchical graph interaction transformer with dynamic token clustering for camouflaged object detection. IEEE Transactions on Image Processing, 33, 5936\u20135948.","journal-title":"IEEE Transactions on Image Processing"},{"key":"82_CR18","unstructured":"Chen, Z., Zhang, X., Xiang, T.-Z., & Tai, Y. (2024). Adaptive guidance learning for camouflaged object detection. arXiv preprint. arXiv:2405.02824."},{"key":"82_CR19","doi-asserted-by":"crossref","unstructured":"Xing, Y., Kong, D., Zhang, S., Chen, G., Ran, L., Wang, P., & Zhang, Y. (2023). Pre-train, adapt and detect: multi-task adapter tuning for camouflaged object detection. arXiv preprint. arXiv:2307.10685.","DOI":"10.2139\/ssrn.4790957"},{"key":"82_CR20","unstructured":"Sun, W., Liu, C., Zhang, L., Li, Y., Wei, P., Liu, C., Zou, J., Jiao, J., & Ye, Q. (2022). DQnet: cross-model detail querying for camouflaged object detection. arXiv preprint. arXiv:2212.08296."},{"issue":"5","key":"82_CR21","doi-asserted-by":"publisher","first-page":"3597","DOI":"10.1109\/TPAMI.2025.3532440","volume":"47","author":"X. Zhang","year":"2025","unstructured":"Zhang, X., Yin, B., Lin, Z., Hou, Q., Fan, D.-P., & Cheng, M.-M. (2025). Referring camouflaged object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(5), 3597\u20133610.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"82_CR22","first-page":"158","volume-title":"Proceedings of the 18th European conference on computer vision","author":"J. Zhang","year":"2024","unstructured":"Zhang, J., Zhang, R., Shi, Y., Cao, Z., Liu, N., & Khan, F. S. (2024). Learning camouflaged object detection from noisy pseudo label. In A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, & G. Varol (Eds.), Proceedings of the 18th European conference on computer vision (pp. 158\u2013174). Cham: Springer."},{"key":"82_CR23","first-page":"438","volume-title":"Proceedings of the 18th European conference on computer vision","author":"X. Lai","year":"2024","unstructured":"Lai, X., Yang, Z., Hu, J., Zhang, S., Cao, L., Jiang, G., Wang, Z., Zhang, S., & Ji, R. (2024). CamoTeacher: dual-rotation consistency learning for semi-supervised camouflaged object detection. In A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, & G. Varol (Eds.), Proceedings of the 18th European conference on computer vision (pp. 438\u2013455). Cham: Springer."},{"key":"82_CR24","doi-asserted-by":"publisher","first-page":"5154","DOI":"10.1109\/TIFS.2021.3124734","volume":"16","author":"Y. Liu","year":"2021","unstructured":"Liu, Y., Zhang, D., Zhang, Q., & Han, J. (2021). Integrating part-object relationship and contrast for camouflaged object detection. IEEE Transactions on Information Forensics and Security, 16, 5154\u20135166.","journal-title":"IEEE Transactions on Information Forensics and Security"},{"key":"82_CR25","first-page":"201","volume-title":"Proceedings of the 7th Chinese conference on pattern recognition and computer vision","author":"Y. Liu","year":"2024","unstructured":"Liu, Y., & Meng, H. (2024). Camouflaged object detection via scale-feature attention and type-feature attention. In Z. Lin, M.-M. Cheng, R. He, K. Ubul, W. Silamu, H. Zha, J. Zhou, & C.-L. Liu (Eds.), Proceedings of the 7th Chinese conference on pattern recognition and computer vision (pp. 201\u2013213). Cham: Springer."},{"key":"82_CR26","unstructured":"Zhang, D., Cheng, L., Liu, Y., Wang, X., & Han, J. (2024). Mamba capsule routing towards part-whole relational camouflaged object detection. arXiv preprint. arXiv:2410.03987."},{"key":"82_CR27","doi-asserted-by":"publisher","first-page":"7188","DOI":"10.1109\/TMM.2024.3361170","volume":"26","author":"W. Hui","year":"2024","unstructured":"Hui, W., Zhu, Z., Gu, G., Liu, M., & Zhao, Y. (2024). Implicit-explicit motion learning for video camouflaged object detection. IEEE Transactions on Multimedia, 26, 7188\u20137196.","journal-title":"IEEE Transactions on Multimedia"},{"key":"82_CR28","first-page":"19058","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"W. Hui","year":"2024","unstructured":"Hui, W., Zhu, Z., Zheng, S., & Zhao, Y. (2024). Endow SAM with keen eyes: temporal-spatial prompt learning for video camouflaged object detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 19058\u201319067). Piscataway: IEEE."},{"key":"82_CR29","first-page":"1857","volume-title":"Proceedings of the IEEE\/CVF conference on Computer Vision and Pattern Recognition workshops","author":"M. N. Meeran","year":"2024","unstructured":"Meeran, M. N., T, G. A., & Mantha, B. P. (2024). SAM-PM: enhancing video camouflaged object detection using spatio-temporal attention. In Proceedings of the IEEE\/CVF conference on Computer Vision and Pattern Recognition workshops (pp. 1857\u20131866). Piscataway: IEEE."},{"key":"82_CR30","first-page":"832","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"H. Lamdouar","year":"2023","unstructured":"Lamdouar, H., Xie, W., & Zisserman, A. (2023). The making and breaking of camouflage. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 832\u2013842). Piscataway: IEEE."},{"key":"82_CR31","first-page":"7177","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"C. Yang","year":"2021","unstructured":"Yang, C., Lamdouar, H., Lu, E., Zisserman, A., & Xie, W. (2021). Self-supervised video object segmentation by motion grouping. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 7177\u20137188). Piscataway: IEEE."},{"key":"82_CR32","unstructured":"Mansoori, M., Shahabodini, S., Abouei, J., Plataniotis, K. N., & Mohammadi, A. (2024). Polyp SAM 2: advancing zero shot polyp segmentation in colorectal cancer detection. arXiv preprint. arXiv:2408.05892."},{"key":"82_CR33","doi-asserted-by":"crossref","unstructured":"Shen, Y., Ding, H., Shao, X., & Unberath, M. (2024). Performance and non-adversarial robustness of the segment anything model 2 in surgical video segmentation. arXiv preprint. arXiv:2408.04098.","DOI":"10.1117\/12.3047383"},{"key":"82_CR34","unstructured":"Yu, J., Wang, A., Dong, W., Xu, M., Islam, M., Wang, J., Bai, L., & Ren, H. (2024). SAM 2 in robotic surgery: an empirical evaluation for robustness and generalization in surgical video segmentation. arXiv preprint. arXiv:2408.04593."},{"key":"82_CR35","unstructured":"Liu, H., Zhang, E., Wu, J., Hong, M., & Jin, Y. (2024). Surgical SAM 2: real-time segment anything in surgical video by efficient frame pruning. arXiv preprint. arXiv:2408.07931."},{"key":"82_CR36","unstructured":"He, Y., Guo, P., Tang, Y., Myronenko, A., Nath, V., Xu, Z., Yang, D., Zhao, C., Xu, D., & Li, W. (2024). A short review and evaluation of SAM2\u2019s performance in 3D CT image segmentation. arXiv preprint. arXiv:2408.11210."},{"key":"82_CR37","doi-asserted-by":"crossref","unstructured":"Chen, T., Lu, A., Zhu, L., Ding, C., Yu, C., Ji, D., Li, Z., Sun, L., Mao, P., & Zang, Y. (2024). SAM2-Adapter: evaluating & adapting segment anything 2 in downstream tasks: camouflage, shadow, medical image segmentation, and more. arXiv preprint. arXiv:2408.04579.","DOI":"10.21203\/rs.3.rs-4876632\/v1"},{"key":"82_CR38","unstructured":"Xiong, X., Wu, Z., Tan, S., Li, W., Tang, F., Chen, Y., Li, S., Ma, J., & Li, G. (2024). SAM2-UNet: segment anything 2 makes strong encoder for natural and medical image segmentation. arXiv preprint. arXiv:2408.08870."},{"key":"82_CR39","unstructured":"Tang, G., Zhao, W., Ford, L., Benhaim, D., & Zhang, P. (2024). Segment any mesh: zero-shot mesh part segmentation via lifting segment anything 2 to 3D. arXiv preprint. arXiv:2408.13679."},{"key":"82_CR40","unstructured":"Rafaeli, O., Svoray, T., Blushtein-Livnon, R., & Nahlieli, A. (2024). Prompt-based segmentation at multiple resolutions and lighting conditions using segment anything model 2. arXiv preprint. arXiv:2408.06970."},{"key":"82_CR41","unstructured":"Tang, L., & Li, B. (2024). Evaluating SAM2\u2019s role in camouflaged object detection: from SAM to SAM2. arXiv preprint. arXiv:2408.21596."},{"key":"82_CR42","first-page":"488","volume-title":"Proceedings of the Asian conference on computer vision","author":"H. Lamdouar","year":"2020","unstructured":"Lamdouar, H., Yang, C., Xie, W., & Zisserman, A. (2020). Betrayed by motion: camouflaged object discovery via motion segmentation. In Proceedings of the Asian conference on computer vision (pp. 488\u2013503). Cham: Springer."},{"key":"82_CR43","first-page":"4548","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"D.-P. Fan","year":"2017","unstructured":"Fan, D.-P., Cheng, M.-M., Liu, Y., Li, T., & Borji, A. (2017). Structure-measure: a new way to evaluate foreground maps. In Proceedings of the IEEE international conference on computer vision (pp. 4548\u20134557). Piscataway: IEEE."},{"key":"82_CR44","first-page":"248","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"R. Margolin","year":"2014","unstructured":"Margolin, R., Zelnik-Manor, L., & Tal, A. (2014). How to evaluate foreground maps? In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 248\u2013255). Piscataway: IEEE."},{"key":"82_CR45","first-page":"733","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"F. Perazzi","year":"2012","unstructured":"Perazzi, F., Kr\u00e4henb\u00fchl, P., Pritch, Y., & Hornung, A. (2012). Saliency filters: contrast based filtering for salient region detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 733\u2013740). Piscataway: IEEE."},{"key":"82_CR46","first-page":"1597","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"R. Achanta","year":"2009","unstructured":"Achanta, R., Hemami, S., Estrada, F., & S\u00fcsstrunk, S. (2009). Frequency-tuned salient region detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 1597\u20131604). Piscataway: IEEE."},{"key":"82_CR47","first-page":"698","volume-title":"Proceedings of the 27th international joint conference on artificial intelligence","author":"D.-P. Fan","year":"2018","unstructured":"Fan, D.-P., Gong, C., Cao, Y., Ren, B., Cheng, M.-M., & Borji, A. (2018). Enhanced-alignment measure for binary foreground map evaluation. In Proceedings of the 27th international joint conference on artificial intelligence (pp. 698\u2013704). Cham: Springer."},{"key":"82_CR48","unstructured":"Zhang, T., Zhou, Z., & Pei, J. (2024). Evaluation study on SAM 2 for class-agnostic instance-level segmentation. arXiv preprint. arXiv:2409.02567."},{"key":"82_CR49","first-page":"34892","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"H. Liu","year":"2023","unstructured":"Liu, H., Li, C., Wu, Q., & Lee, Y. J. (2023). Visual instruction tuning. In Proceedings of the 37th international conference on neural information processing systems (pp. 34892\u201334916). Red Hook: Curran Associates."},{"key":"82_CR50","unstructured":"Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., & Zhao, R. (2023). Shikra: unleashing multimodal LLM\u2019s referential dialogue magic. arXiv preprint. arXiv:2306.15195."},{"key":"82_CR51","doi-asserted-by":"publisher","first-page":"8805","DOI":"10.1145\/3664647.3680730","volume-title":"Proceedings of the 32nd ACM international conference on multimedia","author":"L. Tang","year":"2024","unstructured":"Tang, L., Jiang, P.-T., Shen, Z.-H., Zhang, H., Chen, J.-W., & Li, B. (2024). Chain of visual perception: harnessing multimodal large language models for zero-shot camouflaged object detection. In Proceedings of the 32nd ACM international conference on multimedia (pp. 8805\u20138814). New York: ACM."},{"key":"82_CR52","unstructured":"Ma, J., Kim, S., Li, F., Baharoon, M., Asakereh, R., Lyu, H., & Wang, B. (2024). Segment anything in medical images and videos: benchmark and deployment. arXiv preprint. arXiv:2408.03322."},{"key":"82_CR53","volume-title":"Proceedings of the 7th international conference on learning representations","author":"I. Loshchilov","year":"2019","unstructured":"Loshchilov, I., & Hutter, F. (2019). Decoupled weight decay regularization. In Proceedings of the 7th international conference on learning representations (pp. 1-18). Retrieved May 5, 2025, from https:\/\/openreview.net\/forum?id=Bkg6RiCqY7."},{"key":"82_CR54","first-page":"8779","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"J.-X. Zhao","year":"2019","unstructured":"Zhao, J.-X., Liu, J.-J., Fan, D.-P., Cao, Y., Yang, J., & Cheng, M.-M. (2019). EGNet: edge guidance network for salient object detection. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 8779\u20138788). Piscataway: IEEE."},{"key":"82_CR55","first-page":"7479","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"X. Qin","year":"2019","unstructured":"Qin, X., Zhang, Z., Huang, C., Gao, C., Dehghan, M., & Jagersand, M. (2019). BASNet: boundary-aware salient object detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 7479\u20137489). Piscataway: IEEE."},{"key":"82_CR56","first-page":"3907","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Z. Wu","year":"2019","unstructured":"Wu, Z., Su, L., & Huang, Q. (2019). Cascaded partial decoder for fast and accurate salient object detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 3907\u20133916). Piscataway: IEEE."},{"key":"82_CR57","first-page":"263","volume-title":"International conference on medical image computing and computer-assisted intervention","author":"D.-P. Fan","year":"2020","unstructured":"Fan, D.-P., Ji, G.-P., Zhou, T., Chen, G., Fu, H., Shen, J., & Shao, L. (2020). PraNet: parallel reverse attention network for polyp segmentation. In International conference on medical image computing and computer-assisted intervention (pp. 263\u2013273). Cham: Springer."},{"issue":"10","key":"82_CR58","doi-asserted-by":"publisher","first-page":"6024","DOI":"10.1109\/TPAMI.2021.3085766","volume":"44","author":"D.-P. Fan","year":"2021","unstructured":"Fan, D.-P., Ji, G.-P., Cheng, M.-M., & Shao, L. (2021). Concealed object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10), 6024\u20136042.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"82_CR59","first-page":"142","volume-title":"Proceedings of the 24th international conference on medical image computing and computer-assisted intervention","author":"G.-P. Ji","year":"2021","unstructured":"Ji, G.-P., Chou, Y.-C., Fan, D.-P., Chen, G., Fu, H., Jha, D., & Shao, L. (2021). Progressively normalized self-attention network for video polyp segmentation. In M. de Bruijne, P. C. Cattin, S. Cotin, N. Padoy, S. Speidel, Y. Zheng, & C. Essert (Eds.), Proceedings of the 24th international conference on medical image computing and computer-assisted intervention (pp. 142\u2013152). Cham: Springer."},{"key":"82_CR60","first-page":"7284","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"P. Yan","year":"2019","unstructured":"Yan, P., Li, G., Xie, Y., Li, Z., Wang, C., Chen, T., & Lin, L. (2019). Semi-supervised video salient object detection using pseudo-labels. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 7284\u20137293). Piscataway: IEEE."}],"container-title":["Visual Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-025-00082-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44267-025-00082-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-025-00082-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T09:03:36Z","timestamp":1750323816000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44267-025-00082-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,19]]},"references-count":60,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,12]]}},"alternative-id":["82"],"URL":"https:\/\/doi.org\/10.1007\/s44267-025-00082-1","relation":{},"ISSN":["2097-3330","2731-9008"],"issn-type":[{"value":"2097-3330","type":"print"},{"value":"2731-9008","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,19]]},"assertion":[{"value":"14 January 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 May 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 May 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 June 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest or competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"10"}}