{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T19:55:17Z","timestamp":1781812517587,"version":"3.54.5"},"reference-count":100,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T00:00:00Z","timestamp":1775520000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T00:00:00Z","timestamp":1775520000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100020870","name":"Office of the Vice-President for Research and Development, Hong Kong University of Science and Technology","doi-asserted-by":"publisher","award":["R9429"],"award-info":[{"award-number":["R9429"]}],"id":[{"id":"10.13039\/501100020870","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001445","name":"DSO National Laboratories - Singapore","doi-asserted-by":"publisher","award":["AISG2-GC-2023-008"],"award-info":[{"award-number":["AISG2-GC-2023-008"]}],"id":[{"id":"10.13039\/501100001445","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001348","name":"Agency for Science, Technology and Research","doi-asserted-by":"publisher","award":["C233312028"],"award-info":[{"award-number":["C233312028"]}],"id":[{"id":"10.13039\/501100001348","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100008629","name":"Info-communications Media Development Authority","doi-asserted-by":"publisher","award":["DTC-RGC-04"],"award-info":[{"award-number":["DTC-RGC-04"]}],"id":[{"id":"10.13039\/501100008629","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001778","name":"Deakin University","doi-asserted-by":"publisher","award":["MAAP Discovery funding (2022-2025)"],"award-info":[{"award-number":["MAAP Discovery funding (2022-2025)"]}],"id":[{"id":"10.13039\/501100001778","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Science Foundation Ireland","award":["22\/FFP-P\/11522"],"award-info":[{"award-number":["22\/FFP-P\/11522"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2026,5]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Text-to-image diffusion techniques have shown exceptional capabilities in producing high-quality, dense visual predictions from open-vocabulary text. This indicates a strong correlation between visual and textual domains in open concepts and that diffusion-based text-to-image models can capture rich and diverse information for computer vision tasks. However, we found that those advantages do not hold for learning of features of camouflaged individuals because of the significant blending between their visual boundaries and their surroundings. In this paper, while leveraging the benefits of diffusion-based techniques and text-image models in open-vocabulary settings, we aim to address a challenging problem in computer vision: open-vocabulary camouflaged instance segmentation (OVCIS). Specifically, we propose a method built upon state-of-the-art diffusion empowered by open-vocabulary to learn multi-scale textual-visual features for camouflaged object representation learning. Such cross-domain representations are desirable in segmenting camouflaged objects where visual cues subtly distinguish the objects from the background, and in segmenting novel object classes which are not seen in training. To enable such powerful representations, we devise complementary modules to effectively fuse cross-domain features, and to engage relevant features towards respective foreground objects. We validate and compare our method with existing ones on several benchmark datasets of camouflaged and generic open-vocabulary instance segmentation. The experimental results confirm the advances of our method over existing ones. We believe that our proposed method would open a new avenue for handling camouflages such as computer vision-based surveillance systems, wildlife monitoring, and military reconnaissance.<\/jats:p>","DOI":"10.1007\/s11263-026-02804-4","type":"journal-article","created":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T03:22:27Z","timestamp":1775532147000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion"],"prefix":"10.1007","volume":"134","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8872-0875","authenticated-orcid":false,"given":"Tuan-Anh","family":"Vu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Duc Thanh","family":"Nguyen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qing","family":"Guo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nhat","family":"Chung","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Binh-Son","family":"Hua","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ivor W.","family":"Tsang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sai-Kit","family":"Yeung","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,4,7]]},"reference":[{"key":"2804_CR1","unstructured":"Baranchuk, D., Voynov, A., Rubachev, I., Khrulkov, V., & Babenko, A. (2022). Label-efficient semantic segmentation with diffusion models. Proceedings of the International Conference on Learning Representations."},{"key":"2804_CR2","doi-asserted-by":"crossref","unstructured":"Beery, S., Van\u00a0Horn, G., & Perona, P. (2018). Recognition in terra incognita. Eccv (pp. 456\u2013473).","DOI":"10.1007\/978-3-030-01270-0_28"},{"key":"2804_CR3","doi-asserted-by":"crossref","unstructured":"Bolya, D., Zhou, C., Xiao, F., & Lee, Y.J. (2019). Yolact: Real-time instance segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 9157\u20139166).","DOI":"10.1109\/ICCV.2019.00925"},{"issue":"5","key":"2804_CR4","doi-asserted-by":"publisher","first-page":"1483","DOI":"10.1109\/TPAMI.2019.2956516","volume":"43","author":"Z Cai","year":"2019","unstructured":"Cai, Z., & Nuno, V. (2019). Cascade r-cnn: High quality object detection and instance segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(5), 1483\u20131498.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2804_CR5","doi-asserted-by":"crossref","unstructured":"Chen, H., Sun, K., Tian, Z., Shen, C., Huang, Y., & Yan, Y. (2020). Blendmask: Top-down meets bottom-up for instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 8573\u20138581).","DOI":"10.1109\/CVPR42600.2020.00860"},{"key":"2804_CR6","doi-asserted-by":"crossref","unstructured":"Chen, K., Pang, J., Wang, J., Xiong, Y., Li, X., Sun, S.. others (2019). Hybrid task cascade for instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 4974\u20134983).","DOI":"10.1109\/CVPR.2019.00511"},{"key":"2804_CR7","doi-asserted-by":"crossref","unstructured":"Cheng, B., Misra, I., Schwing, A.G., Kirillov, A., & Girdhar, R. (2022). Masked-attention mask transformer for universal image segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 1290\u20131299).","DOI":"10.1109\/CVPR52688.2022.00135"},{"key":"2804_CR8","first-page":"17864","volume":"34","author":"B Cheng","year":"2021","unstructured":"Cheng, B., Schwing, A., & Kirillov, A. (2021). Per-pixel classification is not all you need for semantic segmentation. Advances in Neural Information Processing Systems, 34, 17864\u201317875.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2804_CR9","doi-asserted-by":"crossref","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R. & Schiele, B. (2016). The cityscapes dataset for semantic urban scene understanding. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 3213\u20133223).","DOI":"10.1109\/CVPR.2016.350"},{"key":"2804_CR10","doi-asserted-by":"crossref","unstructured":"Desai, K., & Johnson, J. (2021). VirTex: Learning visual representations from textual annotations. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 11162\u201311173).","DOI":"10.1109\/CVPR46437.2021.01101"},{"key":"2804_CR11","unstructured":"Dhariwal, P., & Nichol, A.Q. (2021). Diffusion models beat gans on image synthesis. Proceedings of the Advances in Neural Information Processing Systems (pp. 8780\u20138794)."},{"key":"2804_CR12","unstructured":"Ding, Z., Wang, J., & Tu, Z. (2023). Open-vocabulary universal image segmentation with maskclip. Proceedings of the International Conference on Machine Learning."},{"key":"2804_CR13","doi-asserted-by":"crossref","unstructured":"Dong, B., Pei, J., Gao, R., Xiang, T- Z., Wang, S., & Xiong, H. (2024). A unified query-based paradigm for camouflaged instance segmentation. Proceedings of the acm international conference on multimedia (pp. 2131\u20132138).","DOI":"10.1145\/3581783.3612185"},{"key":"2804_CR14","doi-asserted-by":"crossref","unstructured":"Du, Y., Wei, F., Zhang, Z., Shi, M., Gao, Y., & Li, G. (2022). Learning to prompt for open-vocabulary object detection with vision-language model. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 14064\u201314073).","DOI":"10.1109\/CVPR52688.2022.01369"},{"key":"2804_CR15","doi-asserted-by":"crossref","unstructured":"Esser, P., Rombach, R., & Ommer, B. (2021). Taming transformers for high-resolution image synthesis. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 12873\u201312883).","DOI":"10.1109\/CVPR46437.2021.01268"},{"key":"2804_CR16","doi-asserted-by":"crossref","unstructured":"Fan, D- P., Ji, G- P., Cheng, M- M., & Shao, L. (2022). Concealed object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 6024\u20146042,","DOI":"10.1109\/TPAMI.2021.3085766"},{"key":"2804_CR17","doi-asserted-by":"crossref","unstructured":"Fan, D- P., Ji, G- P., Sun, G., Cheng, M- M., Shen, J., & Shao, L. (2020). Camouflaged object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 2777\u20132787).","DOI":"10.1109\/CVPR42600.2020.00285"},{"key":"2804_CR18","doi-asserted-by":"crossref","unstructured":"Fang, Y., Yang, S., Wang, X., Li, Y., Fang, C., Shan, Y. & Liu, W. (2021). Instances as queries. Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 6910\u20136919).","DOI":"10.1109\/ICCV48922.2021.00683"},{"key":"2804_CR19","doi-asserted-by":"crossref","unstructured":"Fleming, P.J.S., Meek, P.D., Ballard, G., Banks, P.B., Claridge, A.W., Sanderson, J.G., & Swann, D.E. (2014). Camera trapping: Wildlife management and research. CSIRO Publishing.","DOI":"10.1071\/9781486300402"},{"key":"2804_CR20","doi-asserted-by":"crossref","unstructured":"Fu, R., Guo, J., Qin, B., Che, W., Wang, H., & Liu, T. (2014). Learning semantic hierarchies via word embeddings. Proceedings of the 52nd annual meeting of the association for computational linguistics (volume 1: Long papers) (pp. 1199\u20131209).","DOI":"10.3115\/v1\/P14-1113"},{"issue":"4","key":"2804_CR21","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3592133","volume":"42","author":"R Gal","year":"2023","unstructured":"Gal, R., Arar, M., Atzmon, Y., Bermano, A. H., Chechik, G., & Cohen-Or, D. (2023). Encoder-based domain tuning for fast personalization of text-to-image models. ACM Transactions on Graphics, 42(4), 1\u201313.","journal-title":"ACM Transactions on Graphics"},{"key":"2804_CR22","doi-asserted-by":"crossref","unstructured":"Gao, M., Xing, C., Niebles, J.C., Li, J., Xu, R., Liu, W., & Xiong, C. (2022). Open vocabulary object detection with pseudo bounding-box labels. Proceedings of the European Conference on Computer Vision (pp. 266\u2013282).","DOI":"10.1007\/978-3-031-20080-9_16"},{"key":"2804_CR23","doi-asserted-by":"crossref","unstructured":"Ghiasi, G., Gu, X., Cui, Y., & Lin, T. (2022). Scaling open-vocabulary image segmentation with image-level labels. Proceedings of the European Conference on Computer Vision (pp. 540\u2013557).","DOI":"10.1007\/978-3-031-20059-5_31"},{"key":"2804_CR24","unstructured":"Gu, X., Lin, T., Kuo, W., & Cui, Y. (2022). Open-vocabulary object detection via vision and language knowledge distillation. Proceedings of the International Conference on Learning Representations."},{"key":"2804_CR25","doi-asserted-by":"crossref","unstructured":"Guo, R., Niu, D., Qu, L., & Li, Z. (2021). Sotr: Segmenting objects with transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 7157\u20137166).","DOI":"10.1109\/ICCV48922.2021.00707"},{"key":"2804_CR26","doi-asserted-by":"crossref","unstructured":"He, C., Li, K., Zhang, Y., Tang, L., Zhang, Y., Guo, Z., & Li, X. (2023). Camouflaged object detection with feature decomposition and edge reconstruction. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 22046\u201322055).","DOI":"10.1109\/CVPR52729.2023.02111"},{"key":"2804_CR27","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., & Girshick, R.B. (2017). Mask R-CNN. Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 2980\u20132988).","DOI":"10.1109\/ICCV.2017.322"},{"key":"2804_CR28","doi-asserted-by":"crossref","unstructured":"He, Z., Xia, C., Qiao, S., & Li, J. (2024). Text-prompt camouflaged instance segmentation with graduated camouflage learning. Proceedings of the acm international conference on multimedia (pp. 5584\u20135593).","DOI":"10.1145\/3664647.3681132"},{"key":"2804_CR29","unstructured":"Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., & Cohen-Or, D. (2023). Prompt-to-prompt image editing with cross attention control. Proceedings of the International Conference on Learning Representations."},{"key":"2804_CR30","unstructured":"Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Proceedings of the Advances in Neural Information Processing Systems (pp. 6840\u20136851)."},{"key":"2804_CR31","doi-asserted-by":"crossref","unstructured":"Huang, X., & Belongie, S.J. (2017). Arbitrary style transfer in real-time with adaptive instance normalization. Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 1510\u20131519).","DOI":"10.1109\/ICCV.2017.167"},{"key":"2804_CR32","doi-asserted-by":"crossref","unstructured":"Huang, Z., Huang, L., Gong, Y., Huang, C., & Wang, X. (2019). Mask scoring r-cnn. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 6409\u20136418).","DOI":"10.1109\/CVPR.2019.00657"},{"key":"2804_CR33","doi-asserted-by":"publisher","DOI":"10.3389\/frai.2024.1347898","author":"CS Ike","year":"2024","unstructured":"Ike, C. S., Muhammad, N., Bibi, N., Alhazmi, S., & Eoghan, F. (2024). Discriminative context-aware network for camouflaged object detection. Frontiers in Artificial Intelligence. https:\/\/doi.org\/10.3389\/frai.2024.1347898","journal-title":"Frontiers in Artificial Intelligence"},{"key":"2804_CR34","doi-asserted-by":"crossref","unstructured":"Jamali, M., Davidsson, P., Khoshkangini, R., Ljungqvist, M.G., & Mihailescu, R- C. (2025). Context in object detection: a systematic literature review. Artificial Intelligence Review","DOI":"10.1007\/s10462-025-11186-x"},{"key":"2804_CR35","unstructured":"Jia, C., Yang, Y., Xia, Y., Chen, Y., Parekh, Z., Pham, H. & Duerig, T. (2021). Scaling up visual and vision-language representation learning with noisy text supervision. Proceedings of the International Conference on Machine Learning (pp. 4904\u20134916)."},{"key":"2804_CR36","doi-asserted-by":"crossref","unstructured":"Karras, T., Laine, S., Aittala, M., Hellsten, J., Lehtinen, J., & Aila, T. (2020). Analyzing and improving the image quality of stylegan. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 8107\u20138116).","DOI":"10.1109\/CVPR42600.2020.00813"},{"key":"2804_CR37","doi-asserted-by":"crossref","unstructured":"Ke, L., Danelljan, M., Li, X., Tai, Y- W., Tang, C- K., & Yu, F. (2022). Mask transfiner for high-quality instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 4412\u20134421).","DOI":"10.1109\/CVPR52688.2022.00437"},{"key":"2804_CR38","doi-asserted-by":"crossref","unstructured":"Khan, A., Khan, M., Gueaieb, W., El\u00a0Saddik, A., De\u00a0Masi, G., & Karray, F. (2024). Camofocus: Enhancing camouflage object detection with split-feature focal modulation and context refinement. Proceedings of the ieee\/cvf winter conference on applications of computer vision (WACV) (pp. 1434\u20131443).","DOI":"10.1109\/WACV57701.2024.00146"},{"issue":"1","key":"2804_CR39","doi-asserted-by":"publisher","first-page":"83","DOI":"10.1002\/nav.3800020109","volume":"2","author":"HW Kuhn","year":"1955","unstructured":"Kuhn, H. W. (1955). The Hungarian Method for the Assignment Problem. Naval Research Logistics Quarterly, 2(1), 83\u201397.","journal-title":"Naval Research Logistics Quarterly"},{"key":"2804_CR40","doi-asserted-by":"crossref","unstructured":"Kumari, N., Zhang, B., Zhang, R., Shechtman, E., & Zhu, J- Y. (2023). Multi-concept customization of text-to-image diffusion. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 1931\u20131941).","DOI":"10.1109\/CVPR52729.2023.00192"},{"key":"2804_CR41","unstructured":"Kuo, W., Cui, Y., Gu, X., Piergiovanni, A., & Angelova, A. (2023). F-vlm: Open-vocabulary object detection upon frozen vision and language models. Proceedings of the International Conference on Learning Representations."},{"key":"2804_CR42","doi-asserted-by":"crossref","unstructured":"Le, M- Q., Tran, M- T., Le, T- N., Nguyen, T.V., & Do, T- T. (2025). CamoFA: A Learnable Fourier-Based Augmentation for Camouflage Segmentation . 2025 ieee\/cvf winter conference on applications of computer vision (wacv) (pp. 3427\u20133436).","DOI":"10.1109\/WACV61041.2025.00338"},{"key":"2804_CR43","doi-asserted-by":"crossref","unstructured":"Le, T-N., Nguyen, T.V., Nie, Z., Tran, M- T., & Sugimoto, A. (2019). Anabranch network for camouflaged object segmentation. Computer Vision and Image Understanding,184, 45\u201356.","DOI":"10.1016\/j.cviu.2019.04.006"},{"issue":"5","key":"2804_CR44","doi-asserted-by":"publisher","first-page":"4062","DOI":"10.1007\/s10489-024-05369-2","volume":"54","author":"C Li","year":"2024","unstructured":"Li, C., Jiao, G., Yue, G., He, R., & Huang, J. (2024). Multi-scale pooling learning for camouflaged instance segmentation. Applied Intelligence, 54(5), 4062\u20134076.","journal-title":"Applied Intelligence"},{"key":"2804_CR45","doi-asserted-by":"crossref","unstructured":"Li, D., Ling, H., Kim, S.W., Kreis, K., Fidler, S., & Torralba, A. (2022). Bigdatasetgan: Synthesizing imagenet with pixel-wise annotations. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 21298\u201321308).","DOI":"10.1109\/CVPR52688.2022.02064"},{"key":"2804_CR46","doi-asserted-by":"crossref","unstructured":"Lin, T., Maire, M., Belongie, S.J., Hays, J., Perona, P., Ramanan, D.. Zitnick, C.L. (2014). Microsoft COCO: common objects in context. Proceedings of the European Conference on Computer Vision (pp. 740\u2013755).","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"2804_CR47","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.126466","volume":"549","author":"M Liu","year":"2023","unstructured":"Liu, M., & Di, X. (2023). Extraordinary MHNet: Military high-level camouflage object detection network and dataset. Neurocomputing, 549, Article 126466.","journal-title":"Neurocomputing"},{"key":"2804_CR48","unstructured":"Loshchilov, I., & Hutter, F. (2019). Decoupled weight decay regularization. Proceedings of the International Conference on Learning Representations."},{"key":"2804_CR49","doi-asserted-by":"crossref","unstructured":"Luo, N., Pan, Y., Sun, R., Zhang, T., Xiong, Z., & Wu, F. (2023). Camouflaged instance segmentation via explicit de-camouflaging. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 17918\u201317927).","DOI":"10.1109\/CVPR52729.2023.01718"},{"key":"2804_CR50","doi-asserted-by":"crossref","unstructured":"Lyu, Y., Zhang, J., Dai, Y., Li, A., Liu, B., Barnes, N., & Fan, D- P. (2021). Simultaneously localize, segment and rank the camouflaged objects. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 11586\u201311596).","DOI":"10.1109\/CVPR46437.2021.01142"},{"key":"2804_CR51","doi-asserted-by":"crossref","unstructured":"Milletari, F., Navab, N., & Ahmadi, S- A. (2016). V-Net: Fully convolutional neural networks for volumetric medical image segmentation. Proceedings of the International Conference on 3D Vision (pp. 565\u2013571).","DOI":"10.1109\/3DV.2016.79"},{"key":"2804_CR52","doi-asserted-by":"crossref","unstructured":"Minderer, M., Gritsenko, A.A., Stone, A., Neumann, M., Weissenborn, D., Dosovitskiy, A. & Houlsby, N. (2022). Simple open-vocabulary object detection with vision transformers. Proceedings of the European Conference on Computer Vision (pp. 728\u2013755).","DOI":"10.1007\/978-3-031-20080-9_42"},{"key":"2804_CR53","unstructured":"Mokady, R., Hertz, A., & Bermano, A.H. (2021). Clipcap: Clip prefix for image captioning. arXiv preprintarXiv:2111.09734, ,"},{"key":"2804_CR54","doi-asserted-by":"crossref","unstructured":"Nguyen, T.T.T., Eichholtzer, A.C., Driscoll, D.A., Semianiw, N.I., Corva, D.M., Kouzani, A.Z. & Nguyen, D.T. (2023). Sawit: A small-sized animal wild image dataset with annotations. Multimedia Tools and Applications, 1\u201326,","DOI":"10.2139\/ssrn.4313792"},{"key":"2804_CR55","unstructured":"Nichol, A.Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B. & Chen, M. (2022). GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. Proceedings of the International Conference on Machine Learning (pp. 16784\u201316804)."},{"issue":"25","key":"2804_CR56","doi-asserted-by":"publisher","first-page":"E5716","DOI":"10.1073\/pnas.1719367115","volume":"115","author":"MS Norouzzadeh","year":"2018","unstructured":"Norouzzadeh, M. S., Nguyen, A., Kosmala, M., Swanson, A., Palmer, M. S., Packer, C., & Clune, J. (2018). Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning. PNAS, 115(25), E5716\u2013E5725.","journal-title":"PNAS"},{"key":"2804_CR57","doi-asserted-by":"crossref","unstructured":"Pang, Y., Zhao, X., Zuo, J., Zhang, L., & Lu, H. (2024). Open-vocabulary camouflaged object segmentation. Proceedings of the European Conference on Computer Vision (eccv).","DOI":"10.1007\/978-3-031-72970-6_27"},{"key":"2804_CR58","doi-asserted-by":"crossref","unstructured":"Parmar, G., Kumar\u00a0Singh, K., Zhang, R., Li, Y., Lu, J., & Zhu, J- Y. (2023). Zero-shot image-to-image translation. Proceedings of the ACM SIGGRAPH (pp. 1\u201311).","DOI":"10.1145\/3588432.3591513"},{"key":"2804_CR59","doi-asserted-by":"crossref","unstructured":"Pei, J., Cheng, T., Fan, D- P., Tang, H., Chen, C., & Van\u00a0Gool, L. (2022). Osformer: One-stage camouflaged instance segmentation with transformers. Proceedings of the European Conference on Computer Vision (pp. 19\u201337).","DOI":"10.1007\/978-3-031-19797-0_2"},{"key":"2804_CR60","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S. & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning (pp. 8748\u20138763)."},{"key":"2804_CR61","first-page":"1","volume":"21","author":"C Raffel","year":"2020","unstructured":"Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21, 1\u201367.","journal-title":"Journal of Machine Learning Research"},{"key":"2804_CR62","unstructured":"Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., & Chen, M. (2022). Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1\u201327,"},{"key":"2804_CR63","doi-asserted-by":"crossref","unstructured":"Rasheed, H.A., Maaz, M., Khattak, M.U., Khan, S.H., & Khan, F.S. (2022). Bridging the gap between object and image-level representations for open-vocabulary detection. Proceedings of the Advances in Neural Information Processing Systems (pp. 33781\u201333794).","DOI":"10.52202\/068431-2448"},{"key":"2804_CR64","unstructured":"Ren, S., He, K., Girshick, R.B., & Sun, J. (2015). Faster R-CNN: towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems (pp. 91\u201399)."},{"issue":"4","key":"2804_CR65","doi-asserted-by":"publisher","first-page":"5114","DOI":"10.1109\/TPAMI.2022.3201285","volume":"45","author":"P Rewatbowornwong","year":"2023","unstructured":"Rewatbowornwong, P., Tritrong, N., & Suwajanakorn, S. (2023). Repurposing gans for one-shot semantic part segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(4), 5114\u20135125.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2804_CR66","doi-asserted-by":"crossref","unstructured":"Robin, R., Andreas, B., Dominik, L., Patrick, E., & Bj\u00f6rn, O. (2022). High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 10674\u201310685).","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"2804_CR67","doi-asserted-by":"crossref","unstructured":"Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E.. others (2022). Photorealistic text-to-image diffusion models with deep language understanding. Proceedings of the Advances in Neural Information Processing Systems (pp. 36479\u201336494).","DOI":"10.52202\/068431-2643"},{"key":"2804_CR68","doi-asserted-by":"crossref","unstructured":"Sariyildiz, M.B., Perez, J., & Larlus, D. (2020). Learning visual representations with caption annotations. Proceedings of the European Conference on Computer Vision (eccv) (pp. 1\u201317).","DOI":"10.1007\/978-3-030-58598-3_10"},{"key":"2804_CR69","doi-asserted-by":"crossref","unstructured":"Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M.. others (2022). LAION-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 25278\u201325294,","DOI":"10.52202\/068431-1833"},{"key":"2804_CR70","doi-asserted-by":"publisher","DOI":"10.1016\/j.ecoinf.2023.102095","volume":"75","author":"F Sim\u00f5es","year":"2023","unstructured":"Sim\u00f5es, F., Bouveyron, C., & Precioso, F. (2023). Deepwild: Wildlife identification, localisation and estimation on camera trap videos using deep learning. Ecological Informatics, 75, Article 102095.","journal-title":"Ecological Informatics"},{"key":"2804_CR71","unstructured":"Song, J., Meng, C., & Ermon, S. (2021). Denoising diffusion implicit models. Proceedings of the International Conference on Learning Representations."},{"key":"2804_CR72","doi-asserted-by":"crossref","unstructured":"Song, Z., Kang, X., Wei, X., & Li, S. (2023). Pixel-centric context perception network for camouflaged object detection. IEEE Transactions on Neural Networks and Learning Systems, ,","DOI":"10.1109\/TNNLS.2023.3319323"},{"key":"2804_CR73","doi-asserted-by":"crossref","unstructured":"Sun, G., An, Z., Liu, Y., Liu, C., Sakaridis, C., Fan, D- P., & Van\u00a0Gool, L. (2023). Indiscernible object counting in underwater scenes. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 13791\u201313801).","DOI":"10.1109\/CVPR52729.2023.01325"},{"key":"2804_CR74","doi-asserted-by":"crossref","unstructured":"Sun, Y., Chen, G., Zhou, T., Zhang, Y., & Liu, N. (2021). Context-aware cross-level fusion network for camouflaged object detection. Ijcai (pp. 1025\u20131031).","DOI":"10.24963\/ijcai.2021\/142"},{"key":"2804_CR75","doi-asserted-by":"crossref","unstructured":"Tian, Z., Shen, C., & Chen, H. (2020). Conditional convolutions for instance segmentation. Proceedings of the European Conference on Computer Vision (pp. 282\u2013298).","DOI":"10.1007\/978-3-030-58452-8_17"},{"key":"2804_CR76","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s12862-016-0854-2","volume":"17","author":"J Troscianko","year":"2017","unstructured":"Troscianko, J., Skelhorn, J., & Stevens, M. (2017). Quantifying camouflage: how to predict detectability from appearance. BMC Evolutionary Biology, 17, 1\u201313.","journal-title":"BMC Evolutionary Biology"},{"key":"2804_CR77","doi-asserted-by":"publisher","DOI":"10.1016\/j.compbiomed.2024.108186","volume":"171","author":"H Wang","year":"2024","unstructured":"Wang, H., Hu, T., Zhang, Y., Zhang, H., Qi, Y., Wang, L., & Du, M. (2024). Unveiling camouflaged and partially occluded colorectal polyps: Introducing CPSNet for accurate colon polyp segmentation. Computers in Biology and Medicine, 171, Article 108186.","journal-title":"Computers in Biology and Medicine"},{"key":"2804_CR78","first-page":"17721","volume":"33","author":"X Wang","year":"2020","unstructured":"Wang, X., Zhang, R., Kong, T., Li, L., & Shen, C. (2020). Solov2: Dynamic and fast instance segmentation. Advances in Neural Information Processing Systems, 33, 17721\u201317732.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2804_CR79","doi-asserted-by":"publisher","DOI":"10.3390\/app14062494","author":"Y Wen","year":"2024","unstructured":"Wen, Y., Ke, W., & Sheng, H. (2024). Camouflaged object detection based on deep learning with attention-guided edge detection and multi-scale context fusion. Applied Sciences. https:\/\/doi.org\/10.3390\/app14062494","journal-title":"Applied Sciences"},{"issue":"07","key":"2804_CR80","doi-asserted-by":"publisher","first-page":"5092","DOI":"10.1109\/TPAMI.2024.3361862","volume":"46","author":"J Wu","year":"2024","unstructured":"Wu, J., Li, X., Xu, S., Yuan, H., Ding, H., Yang, Y., & Tao, D. (2024). Towards open vocabulary learning: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(07), 5092\u20135113.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2804_CR81","unstructured":"Wu, Y., Kirillov, A., Massa, F., Lo, W- Y., & Girshick, R. (2019). Detectron2. https:\/\/github.com\/facebookresearch\/detectron2."},{"key":"2804_CR82","doi-asserted-by":"crossref","unstructured":"Xiao, J., Chen, T., Hu, X., Zhang, G., & Wang, S. (2023). Boundary-guided context-aware network for camouflaged object detection. Neural Computing and Applications, ,","DOI":"10.1007\/s00521-023-08502-3"},{"key":"2804_CR83","doi-asserted-by":"crossref","unstructured":"Xie, E., Wang, W., Wang, W., Sun, P., Xu, H., Liang, D., & Luo, P. (2021). Segmenting transparent objects in the wild with transformer. Proceedings of the International Joint Conferences on Artificial Intelligence (pp. 1194\u20131200).","DOI":"10.24963\/ijcai.2021\/165"},{"key":"2804_CR84","doi-asserted-by":"crossref","unstructured":"Xu, J., Liu, S., Vahdat, A., Byeon, W., Wang, X., & Mello, S.D. (2023). Open-vocabulary panoptic segmentation with text-to-image diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 2955\u20132966).","DOI":"10.1109\/CVPR52729.2023.00289"},{"key":"2804_CR85","doi-asserted-by":"crossref","unstructured":"Xu, X., Xiong, T., Ding, Z., & Tu, Z. (2023). Masqclip for open-vocabulary universal image segmentation. Proceedings of the ieee\/cvf international conference on computer vision (iccv) (pp. 887\u2013898).","DOI":"10.1109\/ICCV51070.2023.00088"},{"key":"2804_CR86","doi-asserted-by":"publisher","first-page":"43290","DOI":"10.1109\/ACCESS.2021.3064443","volume":"9","author":"J Yan","year":"2021","unstructured":"Yan, J., Le, T., Nguyen, K., Tran, M., Do, T., & Nguyen, T. V. (2021). Mirrornet: Bio-inspired camouflaged object segmentation. IEEE Access, 9, 43290\u201343300.","journal-title":"IEEE Access"},{"key":"2804_CR87","doi-asserted-by":"crossref","unstructured":"Zang, Y., Li, W., Zhou, K., Huang, C., & Loy, C.C. (2022). Open-vocabulary detr with conditional matching. Proceedings of the European Conference on Computer Vision (pp. 106\u2013122).","DOI":"10.1007\/978-3-031-20077-9_7"},{"key":"2804_CR88","doi-asserted-by":"crossref","unstructured":"Zareian, A., Rosa, K.D., Hu, D.H., & Chang, S. (2021). Open-vocabulary object detection using captions. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 14393\u201314402).","DOI":"10.1109\/CVPR46437.2021.01416"},{"key":"2804_CR89","doi-asserted-by":"crossref","unstructured":"Zhang, H., Li, F., Zou, X., Liu, S., Li, C., Yang, J., & Zhang, L. (2023). A simple framework for open-vocabulary segmentation and detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 1020\u20131031).","DOI":"10.1109\/ICCV51070.2023.00100"},{"key":"2804_CR90","unstructured":"Zhang, J., Huang, J., Jin, S., & Lu, S. (2023). Vision-language models for vision tasks: A survey. arXiv preprintarXiv:2304.00685, 1\u201323,"},{"key":"2804_CR91","first-page":"1","volume":"182","author":"Y Zhang","year":"2022","unstructured":"Zhang, Y., Jiang, H., Miura, Y., Manning, C. D., & Langlotz, C. P. (2022). Contrastive learning of medical visual representations from paired images and text. Proceedings of Machine Learning Research, 182, 1\u201324.","journal-title":"Proceedings of Machine Learning Research"},{"key":"2804_CR92","doi-asserted-by":"crossref","unstructured":"Zhao, W., Rao, Y., Liu, Z., Liu, B., Zhou, J., & Lu, J. (2023). Unleashing text-to-image diffusion models for visual perception. Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 5729\u20135739).","DOI":"10.1109\/ICCV51070.2023.00527"},{"key":"2804_CR93","doi-asserted-by":"crossref","unstructured":"Zheng, Y., Wu, J., Qin, Y., Zhang, F., & Cui, L. (2021). Zero-shot instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 2593\u20132602).","DOI":"10.1109\/CVPR46437.2021.00262"},{"key":"2804_CR94","doi-asserted-by":"crossref","unstructured":"Zhong, Y., Yang, J., Zhang, P., Li, C., Codella, N., Li, L.H. & Gao, J. (2022). Regionclip: Region-based language-image pretraining. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 16772\u201316782).","DOI":"10.1109\/CVPR52688.2022.01629"},{"key":"2804_CR95","doi-asserted-by":"crossref","unstructured":"Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., & Torralba, A. (2017). Scene parsing through ade20k dataset. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 5122\u20135130).","DOI":"10.1109\/CVPR.2017.544"},{"issue":"3","key":"2804_CR96","doi-asserted-by":"publisher","first-page":"302","DOI":"10.1007\/s11263-018-1140-0","volume":"127","author":"B Zhou","year":"2019","unstructured":"Zhou, B., Zhao, H., Puig, X., Xiao, T., Fidler, S., Barriuso, A., & Torralba, A. (2019). Semantic understanding of scenes through the ade20k dataset. International Journal of Computer Vision, 127(3), 302\u2013321.","journal-title":"International Journal of Computer Vision"},{"key":"2804_CR97","doi-asserted-by":"crossref","unstructured":"Zhou, H., Qi, L., Shen, T., Huang, H., Yang, X., Li, X. & Yang, M- H. (2025). Rethinking evaluation metrics of open-vocabulary segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence","DOI":"10.1109\/TPAMI.2025.3562930"},{"key":"2804_CR98","doi-asserted-by":"crossref","unstructured":"Zhou, X., Girdhar, R., Joulin, A., Kr\u00e4henb\u00fchl, P., & Misra, I. (2022). Detecting twenty-thousand classes using image-level supervision. Proceedings of the European Conference on Computer Vision (pp. 350\u2013368).","DOI":"10.1007\/978-3-031-20077-9_21"},{"key":"2804_CR99","doi-asserted-by":"crossref","unstructured":"Zou, X., Dou, Z- Y., Yang, J., Gan, Z., Li, L., Li, C. & Gao, J. (2023). Generalized decoding for pixel, image, and language. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 15116\u201315127).","DOI":"10.1109\/CVPR52729.2023.01451"},{"key":"2804_CR100","doi-asserted-by":"crossref","unstructured":"Zou, X., Yang, J., Zhang, H., Li, F., Li, L., Wang, J. & Lee, Y.J. (2023). Segment everything everywhere all at once. Thirty-seventh conference on neural information processing systems.","DOI":"10.52202\/075280-0868"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02804-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-026-02804-4","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02804-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T19:03:49Z","timestamp":1781809429000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-026-02804-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,7]]},"references-count":100,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,5]]}},"alternative-id":["2804"],"URL":"https:\/\/doi.org\/10.1007\/s11263-026-02804-4","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,7]]},"assertion":[{"value":"21 September 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 March 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 April 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 April 2026","order":5,"name":"change_date","label":"Change Date","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Update","order":6,"name":"change_type","label":"Change Type","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The original version of this article is revised due to update in affiliation.","order":7,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"210"}}