{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,18]],"date-time":"2026-04-18T06:41:57Z","timestamp":1776494517090,"version":"3.51.2"},"reference-count":67,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2026,3,9]],"date-time":"2026-03-09T00:00:00Z","timestamp":1773014400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,3,9]],"date-time":"2026-03-09T00:00:00Z","timestamp":1773014400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001665","name":"Agence Nationale de la Recherche","doi-asserted-by":"publisher","award":["ANR-24-CE23-1895-01"],"award-info":[{"award-number":["ANR-24-CE23-1895-01"]}],"id":[{"id":"10.13039\/501100001665","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100030672","name":"NVIDIA AI Technology Center, University of Florida","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100030672","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100005278","name":"Universidad de Antioquia","doi-asserted-by":"publisher","award":["2024-73410"],"award-info":[{"award-number":["2024-73410"]}],"id":[{"id":"10.13039\/501100005278","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2026,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Wildlife monitoring is crucial for studying biodiversity loss and climate change. Camera trap images provide a non-intrusive method for analyzing animal populations and identifying ecological patterns over time. However, manual analysis is time-consuming and resource-intensive. Deep learning, particularly foundation models, has been applied to automate wildlife identification, achieving strong performance when tested on data from the same geographical locations as their training sets. Yet, despite their promise, these models struggle to generalize to new geographical areas, leading to significant performance drops. For example, training an advanced vision-language model, such as CLIP with an adapter, on an African dataset achieves an accuracy of 84.77%. However, this performance drops significantly to 16.17% when the model is tested on an American dataset. This limitation partly arises because existing models rely predominantly on image-based representations, making them sensitive to geographical data distribution shifts, such as variation in background, lighting, and environmental conditions. To address this, we introduce WildIng, a\n                    <jats:bold>Wild<\/jats:bold>\n                    life image\n                    <jats:bold>In<\/jats:bold>\n                    variant representation model for\n                    <jats:bold>g<\/jats:bold>\n                    eographical domain shift. WildIng\u00a0integrates text descriptions with image features, creating a more robust representation to geographical domain shifts. By leveraging textual descriptions, our approach captures consistent semantic information, such as detailed descriptions of the appearance of the species, improving generalization across different geographical locations. Experiments show that WildIng\u00a0enhances the accuracy of foundation models such as BioCLIP by 30% under geographical domain shift conditions. We evaluate WildIng\u00a0on two datasets collected from different regions, namely America and Africa. The code and models are publicly available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/Julian075\/CATALOG\/tree\/WildIng\" ext-link-type=\"uri\">https:\/\/github.com\/Julian075\/CATALOG\/tree\/WildIng<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/s11263-026-02739-w","type":"journal-article","created":{"date-parts":[[2026,3,9]],"date-time":"2026-03-09T17:33:39Z","timestamp":1773077619000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["WildIng: A Wildlife Image Invariant Representation Model for Geographical Domain Shift"],"prefix":"10.1007","volume":"134","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-7287-5761","authenticated-orcid":false,"given":"Julian D.","family":"Santamaria","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Claudia","family":"Isaza","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jhony H.","family":"Giraldo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,3,9]]},"reference":[{"key":"2739_CR1","unstructured":"Abdin, M., Aneja, J., Awadalla, H., Awadallah, A., Awan, A.A., Bach, N., Bahree, A., Bakhtiari, A., Bao, J., & Behl, H., et al. (2024). Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219."},{"key":"2739_CR2","unstructured":"Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.\u00a0L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et\u00a0al. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774."},{"issue":"3","key":"2739_CR3","doi-asserted-by":"publisher","first-page":"23","DOI":"10.2111\/1551-501X(2008)30[23:CCAEOT]2.0.CO;2","volume":"30","author":"SR Archer","year":"2008","unstructured":"Archer, S. R., & Predick, K. I. (2008). Climate change and ecosystems of the southwestern united states. Rangelands, 30(3), 23\u201328.","journal-title":"Rangelands"},{"issue":"4","key":"2739_CR4","doi-asserted-by":"publisher","first-page":"2245","DOI":"10.1109\/TPAMI.2024.3506283","volume":"47","author":"M Awais","year":"2025","unstructured":"Awais, M., Naseer, M., Khan, S., Anwer, R. M., Cholakkal, H., Shah, M., Yang, M.-H., & Khan, F. S. (2025). Foundation models defining a new era in vision: a survey and outlook. IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(4), 2245\u20132264.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2739_CR5","unstructured":"Beery, S., Morris, D., & Yang, S. (2019). Efficient pipeline for camera trap image review. arXiv preprint arXiv:1907.06772."},{"key":"2739_CR6","doi-asserted-by":"crossref","unstructured":"Beery, S., Van Horn, G., & Perona, P. (2018). Recognition in Terra Incognita. In IEEE\/CVF European Conference on Computer Vision, 456\u2013473.","DOI":"10.1007\/978-3-030-01270-0_28"},{"key":"2739_CR7","doi-asserted-by":"crossref","unstructured":"Bendou, Y., Ouasfi, A., Gripon, V., & Boukhayma, A. (2025). Proker: A kernel perspective on few-shot adaptation of large vision-language models. In Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference, 25092\u201325102.","DOI":"10.1109\/CVPR52734.2025.02336"},{"key":"2739_CR8","doi-asserted-by":"crossref","unstructured":"Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Conference of the North American Chapter of the Association for Computational Linguistics, 4171\u20134186.","DOI":"10.18653\/v1\/N19-1423"},{"key":"2739_CR9","doi-asserted-by":"crossref","unstructured":"Duan, J., Chen, L., Tran, S., Yang, J., Xu, Y., Zeng, B., & Chilimbi, T. (2022). Multi-modal alignment using representation codebook. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 15651\u201315660.","DOI":"10.1109\/CVPR52688.2022.01520"},{"key":"2739_CR10","unstructured":"Fabian, Z., Miao, Z., Li, C., Zhang, Y., Liu, Z., Hernandez, A., Arbelaez, P., Link, A., Montes-Rojas, A., & Escucha, R., et al. (2023). Knowledge augmented instruction tuning for zero-shot animal species recognition. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following."},{"key":"2739_CR11","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2025.121989","volume":"705","author":"Z Fang","year":"2025","unstructured":"Fang, Z., Lu, J., & Zhang, G. (2025). Out-of-distribution detection with non-semantic exploration. Information Sciences, 705, Article 121989.","journal-title":"Information Sciences"},{"issue":"9","key":"2739_CR12","doi-asserted-by":"publisher","first-page":"3770","DOI":"10.1007\/s11263-024-02026-6","volume":"132","author":"V Gabeff","year":"2024","unstructured":"Gabeff, V., Ru\u00dfwurm, M., Tuia, D., & Mathis, A. (2024). WildCLIP: Scene and animal attribute retrieval from camera trap data with domain-adapted vision-language models. International Journal of Computer Vision, 132(9), 3770\u20133786.","journal-title":"International Journal of Computer Vision"},{"issue":"8","key":"2739_CR13","doi-asserted-by":"publisher","first-page":"1193","DOI":"10.1049\/cvi2.12318","volume":"18","author":"T Gadot","year":"2024","unstructured":"Gadot, T., Istrate, S., Kim, H., Morris, D., Beery, S., Birch, T., & Ahumada, J. (2024). To crop or not to crop: Comparing whole-image and cropped classification on a large dataset of camera trap images. IET Computer Vision, 18(8), 1193\u20131208.","journal-title":"IET Computer Vision"},{"key":"2739_CR14","doi-asserted-by":"publisher","first-page":"581","DOI":"10.1007\/s11263-023-01891-x","volume":"132","author":"P Gao","year":"2024","unstructured":"Gao, P., Geng, S., Zhang, R., Ma, T., Fang, R., Zhang, Y., Li, H., & Qiao, Y. (2024). Clip-adapter: Better vision-language models with feature adapters. International Journal of Computer Vision, 132, 581\u2013595.","journal-title":"International Journal of Computer Vision"},{"key":"2739_CR15","first-page":"78431","volume":"37","author":"JJ Garau-Luis","year":"2025","unstructured":"Garau-Luis, J. J., Bordes, P., Gonzalez, L., Roller, M., de Almeida, B., Blum, C., Hexemer, L., Laurent, S., Lang, M., Pierrot, T., et al. (2025). Multi-modal transfer learning between biological foundation models. In Advances in Neural Information Processing Systems, 37, 78431\u201378450.","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"2739_CR16","doi-asserted-by":"publisher","first-page":"335","DOI":"10.1007\/s00371-017-1463-9","volume":"35","author":"JH Giraldo","year":"2019","unstructured":"Giraldo, J. H., Salazar, A., Gomez-Villa, A., & Diaz-Pulido, A. (2019). Camera-trap images segmentation using multi-layer robust principal component analysis. The Visual Computer, 35, 335\u2013347.","journal-title":"The Visual Computer"},{"key":"2739_CR17","doi-asserted-by":"publisher","first-page":"24","DOI":"10.1016\/j.ecoinf.2017.07.004","volume":"41","author":"A Gomez-Villa","year":"2017","unstructured":"Gomez-Villa, A., Salazar, A., & Vargas, F. (2017). Towards automatic wild animal monitoring: Identification of animal species in camera-trap images using very deep convolutional neural networks. Ecological informatics, 41, 24\u201332.","journal-title":"Ecological informatics"},{"key":"2739_CR18","unstructured":"Hernandez, A., Miao, Z., Vargas, L., Dodhia, R., & Lavista, J. (2024). Pytorch-Wildlife: A collaborative deep learning framework for conservation. arXiv preprint arXiv:2405.12930."},{"key":"2739_CR19","doi-asserted-by":"crossref","unstructured":"Hogeweg, L.\u00a0E., Gangireddy, R., Brunink, D., Kalkman, V.J., Cornelissen, L., & Kamminga, J.W. (2024). Cood: Combined out-of-distribution detection using multiple measures for anomaly & novel class detection in large-scale hierarchical classification. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pages 3971\u20133980.","DOI":"10.1109\/CVPRW63382.2024.00401"},{"key":"2739_CR20","first-page":"72096","volume":"36","author":"S Huang","year":"2024","unstructured":"Huang, S., Dong, L., Wang, W., Hao, Y., Singhal, S., Ma, S., Lv, T., Cui, L., Mohammed, O. K., Patra, B., et al. (2024). Language is not all you need: Aligning perception with language models. In Advances in Neural Information Processing Systems, 36, 72096\u201372109.","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"2739_CR21","unstructured":"Jiang, A.Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D.S., Casas, D. d.l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. (2023). Mistral 7b. arXiv preprint arXiv:2310.06825."},{"key":"2739_CR22","doi-asserted-by":"publisher","first-page":"583","DOI":"10.1038\/s41586-021-03819-2","volume":"596","author":"J Jumper","year":"2021","unstructured":"Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., \u017d\u00eddek, A., Potapenko, A., et al. (2021). Highly accurate protein structure prediction with alphafold. Nature, 596, 583\u2013589.","journal-title":"Nature"},{"key":"2739_CR23","doi-asserted-by":"crossref","unstructured":"Jung, S.J., Kim, H., & Jang, K.S. (2024). Llm based biological named entity recognition from scientific literature. In IEEE International Conference on Big Data and Smart Computing (BigComp), 433\u2013435.","DOI":"10.1109\/BigComp60711.2024.00095"},{"key":"2739_CR24","unstructured":"Kempf, E., Schrodi, S., Argus, M., & Brox, T. (2025). When and how does clip enable domain and compositional generalization? arXiv preprint arXiv:2502.09507."},{"key":"2739_CR25","doi-asserted-by":"crossref","unstructured":"Khan, Z., & Fu, Y. (2024). Consistency and uncertainty: Identifying unreliable responses from black-box vision-language models for selective visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pages 10854\u201310863.","DOI":"10.1109\/CVPR52733.2024.01032"},{"key":"2739_CR26","doi-asserted-by":"publisher","first-page":"1145","DOI":"10.1016\/j.tplants.2024.04.013","volume":"29","author":"HYI Lam","year":"2024","unstructured":"Lam, H. Y. I., Ong, X. E., & Mutwil, M. (2024). Large language models in plant biology. Trends in Plant Science, 29, 1145\u20131155.","journal-title":"Trends in Plant Science"},{"key":"2739_CR27","unstructured":"Li, J., Li, D., Savarese, S., & Hoi, S. (2023). BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning, 19730\u201319742."},{"key":"2739_CR28","doi-asserted-by":"crossref","unstructured":"Li, J., Li, Y., Fu, Y., Liu, J., Liu, Y., Yang, M., & King, I. (2025a). Clip-powered domain generalization and domain adaptation: A comprehensive survey. arXiv preprint arXiv:2504.14280.","DOI":"10.1109\/TPAMI.2026.3651700"},{"key":"2739_CR29","doi-asserted-by":"publisher","first-page":"7380","DOI":"10.1109\/TMM.2025.3599076","volume":"27","author":"S Li","year":"2025","unstructured":"Li, S., Xu, X., Meng, W., Song, J., Peng, C., & Shen, H. T. (2025). Mitigating hallucinations in large vision-language models via reasoning uncertainty-guided refinement. IEEE Transactions on Multimedia, 27, 7380\u20137391.","journal-title":"IEEE Transactions on Multimedia"},{"key":"2739_CR30","doi-asserted-by":"publisher","DOI":"10.1016\/j.ecoinf.2022.101597","volume":"69","author":"X Li","year":"2022","unstructured":"Li, X., Tian, H., Piao, Z., Wang, G., Xiao, Z., Sun, Y., Gao, E., & Holyoak, M. (2022). cameratrapr: An r package for estimating animal density using camera trapping data. Ecological Informatics, 69, Article 101597.","journal-title":"Ecological Informatics"},{"key":"2739_CR31","unstructured":"Liang, W., Mao, Y., Kwon, Y., Yang, X., & Zou, J. (2023). Accuracy on the curve: On the nonlinear correlation of ml performance between data subpopulations. In IEEE\/CVF International Conference on Machine Learning, 20706\u201320724."},{"key":"2739_CR32","first-page":"34892","volume":"36","author":"H Liu","year":"2024","unstructured":"Liu, H., Li, C., Wu, Q., & Lee, Y. J. (2024). Visual instruction tuning. In Advances in Neural Information Processing Systems, 36, 34892\u201334916.","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"2739_CR33","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s11263-023-01871-1","volume":"132","author":"G Luo","year":"2024","unstructured":"Luo, G., Zhou, Y., Sun, X., Wu, Y., Gao, Y., & Ji, R. (2024). Towards language-guided visual recognition via dynamic convolutions. International Journal of Computer Vision, 132, 1\u201319.","journal-title":"International Journal of Computer Vision"},{"key":"2739_CR34","doi-asserted-by":"crossref","unstructured":"Mushtaq, E., Fabian, Z., Bakman, Y.F., Ramakrishna, A., Soltanolkotabi, M., & Avestimehr, S. (2025). Harmony: Hidden activation representations and model output-aware uncertainty estimation for vision-language models. In Proceedings of the Computer Vision and Pattern Recognition Conference, 1663\u20131668.","DOI":"10.1109\/CVPRW67362.2025.00154"},{"key":"2739_CR35","doi-asserted-by":"crossref","unstructured":"Nguyen, T., Lyu, B., Ishwar, P., Scheutz, M., & Aeron, S. (2022). Trade-off between reconstruction loss and feature alignment for domain generalization. In 2022 21st IEEE International Conference on Machine Learning and Applications, 794\u2013801.","DOI":"10.1109\/ICMLA55696.2022.00132"},{"issue":"1","key":"2739_CR36","doi-asserted-by":"publisher","first-page":"242","DOI":"10.1111\/2041-210X.14031","volume":"14","author":"DL Norman","year":"2023","unstructured":"Norman, D. L., Bischoff, P. H., Wearn, O. R., Ewers, R. M., Rowcliffe, J. M., Evans, B., Sethi, S., Chapman, P. M., & Freeman, R. (2023). Can CNN-based species classification generalise across variation in habitat within a camera trap survey? Methods in Ecology and Evolution, 14(1), 242\u2013251.","journal-title":"Methods in Ecology and Evolution"},{"key":"2739_CR37","unstructured":"Pantazis, O., Brostow, G., Jones, K., & Mac Aodha, O. (2022). SVL-Adapter: Self-Supervised Adapter for Vision-Language Pretrained Models. In British Machine Vision Conference."},{"key":"2739_CR38","doi-asserted-by":"publisher","first-page":"166","DOI":"10.1038\/s44358-025-00022-3","volume":"1","author":"LJ Pollock","year":"2025","unstructured":"Pollock, L. J., Kitzes, J., Beery, S., Gaynor, K. M., Jarzyna, M. A., Mac Aodha, O., Meyer, B., Rolnick, D., Taylor, G. W., Tuia, D., et al. (2025). Harnessing artificial intelligence to fill global shortfalls in biodiversity knowledge. Nature Reviews Biodiversity, 1, 166\u2013182.","journal-title":"Nature Reviews Biodiversity"},{"key":"2739_CR39","doi-asserted-by":"crossref","unstructured":"Pratt, S., Covert, I., Liu, R., & Farhadi, A. (2023). What does a platypus look like? generating customized prompts for zero-shot image classification. In IEEE\/CVF International Conference on Computer Vision, 15691\u201315701.","DOI":"10.1109\/ICCV51070.2023.01438"},{"key":"2739_CR40","first-page":"8748","volume":"139","author":"A Radford","year":"2021","unstructured":"Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021). Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, 139, 8748\u20138763.","journal-title":"In International Conference on Machine Learning"},{"key":"2739_CR41","doi-asserted-by":"publisher","first-page":"191","DOI":"10.1016\/j.tree.2024.11.013","volume":"40","author":"SA Reynolds","year":"2024","unstructured":"Reynolds, S. A., Beery, S., Burgess, N., Burgman, M., Butchart, S. H., Cooke, S. J., Coomes, D., Danielsen, F., Di Minin, E., Dur\u00e1n, A. P., et al. (2024). The potential for ai to revolutionize conservation: a horizon scan. Trends in Ecology & Evolution, 40, 191\u2013207.","journal-title":"Trends in Ecology & Evolution"},{"key":"2739_CR42","doi-asserted-by":"publisher","first-page":"527","DOI":"10.1007\/s11263-024-02180-x","volume":"133","author":"L Riz","year":"2024","unstructured":"Riz, L., Saltori, C., Wang, Y., Ricci, E., & Poiesi, F. (2024). Novel class discovery meets foundation models for 3d semantic segmentation. International Journal of Computer Vision, 133, 527\u2013548.","journal-title":"International Journal of Computer Vision"},{"key":"2739_CR43","doi-asserted-by":"crossref","unstructured":"Santamaria, J.D., Isaza, C., & Giraldo, J.H. (2025). CATALOG: A camera trap language-guided contrastive learning model. In IEEE\/CVF Winter Conference on Applications of Computer Vision, 1197\u20131206,.","DOI":"10.1109\/WACV61041.2025.00124"},{"key":"2739_CR44","unstructured":"Santamaria P, J.D., Giraldo, J.H., Diaz-Pulido, A., & Isaza, C. (2024). Audio vs. visual approach to monitor the critically endangered species atlapetes blancae: Developing deep learning models with limited data. In IARIA Annual Congress on Frontiers in Science, Technology, Services, and Applications, 72\u201380."},{"issue":"7","key":"2739_CR45","doi-asserted-by":"publisher","first-page":"3503","DOI":"10.1002\/ece3.6147","volume":"10","author":"S Schneider","year":"2020","unstructured":"Schneider, S., Greenberg, S., Taylor, G. W., & Kremer, S. C. (2020). Three critical factors affecting automated image species recognition performance for camera traps. Ecology and Evolution, 10(7), 3503\u20133517.","journal-title":"Ecology and Evolution"},{"key":"2739_CR46","doi-asserted-by":"publisher","DOI":"10.1016\/j.ecoinf.2023.102095","volume":"75","author":"F Sim\u00f5es","year":"2023","unstructured":"Sim\u00f5es, F., Bouveyron, C., & Precioso, F. (2023). DeepWILD: Wildlife identification, localisation and estimation on camera trap videos using deep learning. Ecological Informatics, 75, Article 102095.","journal-title":"Ecological Informatics"},{"key":"2739_CR47","doi-asserted-by":"crossref","unstructured":"Stevens, S., Wu, J., Thompson, M.J., Campolongo, E.G., Song, C.H., Carlyn, D.E., Dong, L., Dahdul, W.M., Stewart, C., Berger-Wolf, T., et al. (2024). BioCLIP: A vision foundation model for the tree of life. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pages 19412\u201319424.","DOI":"10.1109\/CVPR52733.2024.01836"},{"key":"2739_CR48","doi-asserted-by":"crossref","unstructured":"Swanson, A., Kosmala, M., Lintott, C., Simpson, R., Smith, A., & Packer, C. (2015). Data from: Snapshot serengeti, high-frequency annotated camera trap images of 40 mammalian species in an african savanna.","DOI":"10.1038\/sdata.2015.26"},{"key":"2739_CR49","unstructured":"Tan, M. & Le, Q. (2021). Efficientnetv2: Smaller models and faster training. In IEEE\/CVF International Conference on Machine Learning, 10096\u201310106."},{"key":"2739_CR50","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s11263-024-02185-6","volume":"133","author":"L Tang","year":"2025","unstructured":"Tang, L., Jiang, P.-T., Xiao, H., & Li, B. (2025). Towards training-free open-world segmentation via image prompt foundation models. International Journal of Computer Vision, 133, 1\u201315.","journal-title":"International Journal of Computer Vision"},{"key":"2739_CR51","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozi\u00e8re, B., Goyal, N., Hambro, E., Azhar, F., & Rodriguez, A. (2023). LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971."},{"issue":"1","key":"2739_CR52","doi-asserted-by":"publisher","first-page":"792","DOI":"10.1038\/s41467-022-27980-y","volume":"13","author":"D Tuia","year":"2022","unstructured":"Tuia, D., Kellenberger, B., Beery, S., Costelloe, B. R., Zuffi, S., Risse, B., Mathis, A., Mathis, M. W., van Langevelde, F., Burghardt, T., et al. (2022). Perspectives in machine learning for wildlife conservation. Nature Communications, 13(1), 792.","journal-title":"Nature Communications"},{"key":"2739_CR53","first-page":"2215","volume":"34","author":"Y Wald","year":"2021","unstructured":"Wald, Y., Feder, A., Greenfeld, D., & Shalit, U. (2021). On calibration and out-of-domain generalization. In Advances in Neural Information Processing Systems, 34, 2215\u20132227.","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"2739_CR54","first-page":"122484","volume":"37","author":"Q Wang","year":"2024","unstructured":"Wang, Q., Lin, Y., Chen, Y., Schmidt, L., Han, B., & Zhang, T. (2024). A sober look at the robustness of clips to spurious features. In Advances in Neural Information Processing Systems, 37, 122484\u2013122523.","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"2739_CR55","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2025.105511","volume":"157","author":"Y Wang","year":"2025","unstructured":"Wang, Y., & Kang, G. (2025). Attention head purification: A new perspective to harness clip for domain generalization. Image and Vision Computing, 157, Article 105511.","journal-title":"Image and Vision Computing"},{"key":"2739_CR56","doi-asserted-by":"publisher","first-page":"392","DOI":"10.1007\/s11263-023-01876-w","volume":"132","author":"W Wu","year":"2024","unstructured":"Wu, W., Sun, Z., Song, Y., Wang, J., & Ouyang, W. (2024). Transferring vision-language models for visual recognition: A classifier perspective. International Journal of Computer Vision, 132, 392\u2013409.","journal-title":"International Journal of Computer Vision"},{"key":"2739_CR57","unstructured":"Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al. (2024). Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115."},{"key":"2739_CR58","doi-asserted-by":"publisher","first-page":"3746","DOI":"10.1007\/s11263-025-02354-1","volume":"133","author":"L Yang","year":"2025","unstructured":"Yang, L., Zhang, R.-Y., Chen, Q., & Xie, X. (2025). Learning with enriched inductive biases for vision-language models. International Journal of Computer Vision, 133, 3746\u20133761.","journal-title":"International Journal of Computer Vision"},{"key":"2739_CR59","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2025.129826","volume":"634","author":"Z Yang","year":"2025","unstructured":"Yang, Z., Tian, Y., Wang, L., & Zhang, J. (2025). Enhancing generalization in camera trap image recognition: Fine-tuning visual language models. Neurocomputing, 634, Article 129826.","journal-title":"Neurocomputing"},{"key":"2739_CR60","doi-asserted-by":"crossref","unstructured":"Yu, H., Zhang, X., Xu, R., Liu, J., He, Y., & Cui, P. (2024). Rethinking the evaluation protocol of domain generalization. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 21897\u201321908.","DOI":"10.1109\/CVPR52733.2024.02068"},{"key":"2739_CR61","doi-asserted-by":"crossref","unstructured":"Yu, R., Liu, S., Yang, X., & Wang, X. (2023). Distribution shift inversion for out-of-distribution prediction. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 3592\u20133602.","DOI":"10.1109\/CVPR52729.2023.00350"},{"key":"2739_CR62","doi-asserted-by":"crossref","unstructured":"Zanella, M., & Ben Ayed, I. (2024). Low-rank few-shot adaptation of vision-language models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1593\u20131603.","DOI":"10.1109\/CVPRW63382.2024.00166"},{"key":"2739_CR63","doi-asserted-by":"publisher","first-page":"825","DOI":"10.1007\/s11263-024-02214-4","volume":"133","author":"Y Zang","year":"2024","unstructured":"Zang, Y., Li, W., Han, J., Zhou, K., & Loy, C. C. (2024). Contextual object detection with multimodal large language models. International Journal of Computer Vision, 133, 825\u2013843.","journal-title":"International Journal of Computer Vision"},{"key":"2739_CR64","doi-asserted-by":"crossref","unstructured":"Zhang, B., Zhang, P., Dong, X., Zang, Y., & Wang, J. (2024a). Long-CLIP: Unlocking the long-text capability of CLIP. In IEEE\/CVF European Conference on Computer Vision, 310\u2013325.","DOI":"10.1007\/978-3-031-72983-6_18"},{"issue":"8","key":"2739_CR65","doi-asserted-by":"publisher","first-page":"5625","DOI":"10.1109\/TPAMI.2024.3369699","volume":"46","author":"J Zhang","year":"2024","unstructured":"Zhang, J., Huang, J., Jin, S., & Lu, S. (2024). Vision-Language Models for Vision Tasks: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(8), 5625\u20135644.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2739_CR66","doi-asserted-by":"crossref","unstructured":"Zhang, R., Zhang, W., Fang, R., Gao, P., Li, K., Dai, J., Qiao, Y., & Li, H. (2022). Tip-adapter: Training-free adaption of clip for few-shot classification. In Proceedings of the IEEE\/CVF European conference on computer vision, 493\u2013510.","DOI":"10.1007\/978-3-031-19833-5_29"},{"key":"2739_CR67","doi-asserted-by":"publisher","first-page":"3375","DOI":"10.1007\/s11263-024-02036-4","volume":"132","author":"L Zhu","year":"2024","unstructured":"Zhu, L., Yin, W., Yang, Y., Wu, F., Zeng, Z., Gu, Q., Wang, X., Zhou, C., & Ye, N. (2024). Vision-language alignment learning under affinity and divergence principles for few-shot out-of-distribution generalization. International Journal of Computer Vision, 132, 3375\u20133407.","journal-title":"International Journal of Computer Vision"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02739-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-026-02739-w","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02739-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,18]],"date-time":"2026-04-18T05:47:55Z","timestamp":1776491275000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-026-02739-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,9]]},"references-count":67,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4]]}},"alternative-id":["2739"],"URL":"https:\/\/doi.org\/10.1007\/s11263-026-02739-w","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,9]]},"assertion":[{"value":"21 March 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 January 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 March 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"183"}}