{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T19:44:46Z","timestamp":1776887086797,"version":"3.51.2"},"reference-count":63,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2024,5,14]],"date-time":"2024-05-14T00:00:00Z","timestamp":1715644800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Natural Science Foundation of Heilongjiang Province","award":["LH2020F040"],"award-info":[{"award-number":["LH2020F040"]}]},{"name":"Natural Science Foundation of Heilongjiang Province","award":["HUDF2022110"],"award-info":[{"award-number":["HUDF2022110"]}]},{"name":"Natural Science Foundation of Heilongjiang Province","award":["ZC2022ZJ010027"],"award-info":[{"award-number":["ZC2022ZJ010027"]}]},{"name":"Natural Science Foundation of Heilongjiang Province","award":["2572017PZ10"],"award-info":[{"award-number":["2572017PZ10"]}]},{"name":"Young Doctoral Research Initiation Fund Project of Harbin University \u201cResearch on Wood Recognition Methods Based on Deep Learning Fusion Model\u201d","award":["LH2020F040"],"award-info":[{"award-number":["LH2020F040"]}]},{"name":"Young Doctoral Research Initiation Fund Project of Harbin University \u201cResearch on Wood Recognition Methods Based on Deep Learning Fusion Model\u201d","award":["HUDF2022110"],"award-info":[{"award-number":["HUDF2022110"]}]},{"name":"Young Doctoral Research Initiation Fund Project of Harbin University \u201cResearch on Wood Recognition Methods Based on Deep Learning Fusion Model\u201d","award":["ZC2022ZJ010027"],"award-info":[{"award-number":["ZC2022ZJ010027"]}]},{"name":"Young Doctoral Research Initiation Fund Project of Harbin University \u201cResearch on Wood Recognition Methods Based on Deep Learning Fusion Model\u201d","award":["2572017PZ10"],"award-info":[{"award-number":["2572017PZ10"]}]},{"name":"Self-funded project of Harbin Science and Technology Plan Research on Computer Vision Recognition Technology of Wood Species Based on transfer learning Fusion Model","award":["LH2020F040"],"award-info":[{"award-number":["LH2020F040"]}]},{"name":"Self-funded project of Harbin Science and Technology Plan Research on Computer Vision Recognition Technology of Wood Species Based on transfer learning Fusion Model","award":["HUDF2022110"],"award-info":[{"award-number":["HUDF2022110"]}]},{"name":"Self-funded project of Harbin Science and Technology Plan Research on Computer Vision Recognition Technology of Wood Species Based on transfer learning Fusion Model","award":["ZC2022ZJ010027"],"award-info":[{"award-number":["ZC2022ZJ010027"]}]},{"name":"Self-funded project of Harbin Science and Technology Plan Research on Computer Vision Recognition Technology of Wood Species Based on transfer learning Fusion Model","award":["2572017PZ10"],"award-info":[{"award-number":["2572017PZ10"]}]},{"name":"Fundamental Research Funds for the Central Universities","award":["LH2020F040"],"award-info":[{"award-number":["LH2020F040"]}]},{"name":"Fundamental Research Funds for the Central Universities","award":["HUDF2022110"],"award-info":[{"award-number":["HUDF2022110"]}]},{"name":"Fundamental Research Funds for the Central Universities","award":["ZC2022ZJ010027"],"award-info":[{"award-number":["ZC2022ZJ010027"]}]},{"name":"Fundamental Research Funds for the Central Universities","award":["2572017PZ10"],"award-info":[{"award-number":["2572017PZ10"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Fine-grained representation is fundamental to species classification based on deep learning, and in this context, cross-modal contrastive learning is an effective method. The diversity of species coupled with the inherent contextual ambiguity of natural language poses a primary challenge in the cross-modal representation alignment of conservation area image data. Integrating cross-modal retrieval tasks with generation tasks contributes to cross-modal representation alignment based on contextual understanding. However, during the contrastive learning process, apart from learning the differences in the data itself, a pair of encoders inevitably learns the differences caused by encoder fluctuations. The latter leads to convergence shortcuts, resulting in poor representation quality and an inaccurate reflection of the similarity relationships between samples in the original dataset within the shared space of features. To achieve fine-grained cross-modal representation alignment, we first propose a residual attention network to enhance consistency during momentum updates in cross-modal encoders. Building upon this, we propose momentum encoding from a multi-task perspective as a bridge for cross-modal information, effectively improving cross-modal mutual information, representation quality, and optimizing the distribution of feature points within the cross-modal shared semantic space. By acquiring momentum encoding queues for cross-modal semantic understanding through multi-tasking, we align ambiguous natural language representations around the invariant image features of factual information, alleviating contextual ambiguity and enhancing model robustness. Experimental validation shows that our proposed multi-task perspective of cross-modal momentum encoders outperforms similar models on standardized image classification tasks and image\u2013text cross-modal retrieval tasks on public datasets by up to 8% on the leaderboard, demonstrating the effectiveness of the proposed method. Qualitative experiments on our self-built conservation area image\u2013text paired dataset show that our proposed method accurately performs cross-modal retrieval and generation tasks among 8142 species, proving its effectiveness on fine-grained cross-modal image\u2013text conservation area image datasets.<\/jats:p>","DOI":"10.3390\/s24103130","type":"journal-article","created":{"date-parts":[[2024,5,15]],"date-time":"2024-05-15T03:35:55Z","timestamp":1715744155000},"page":"3130","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Fine-Grained Cross-Modal Semantic Consistency in Natural Conservation Image Data from a Multi-Task Perspective"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7150-6584","authenticated-orcid":false,"given":"Rui","family":"Tao","sequence":"first","affiliation":[{"name":"College of Computer and Control Engineering, Northeast Forestry University, Harbin 150040, China"},{"name":"College of Artificial Intelligence and Big Data, Hulunbuir University, Hulunbuir 021008, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Meng","family":"Zhu","sequence":"additional","affiliation":[{"name":"College of Information Engineering, Harbin University, Harbin 150076, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haiyan","family":"Cao","sequence":"additional","affiliation":[{"name":"College of Artificial Intelligence and Big Data, Hulunbuir University, Hulunbuir 021008, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Honge","family":"Ren","sequence":"additional","affiliation":[{"name":"College of Computer and Control Engineering, Northeast Forestry University, Harbin 150040, China"},{"name":"Heilongjiang Forestry Intelligent Equipment Engineering Research Center, Harbin 150040, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,5,14]]},"reference":[{"key":"ref_1","unstructured":"Matin, M., Shrestha, T., Chitale, V., and Thomas, S. (2021, January 13\u201317). Exploring the potential of deep learning for classifying camera trap data of wildlife: A case study from Nepal. Proceedings of the AGU Fall Meeting Abstracts, New Orleans, LA, USA."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"E5716","DOI":"10.1073\/pnas.1719367115","article-title":"Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning","volume":"115","author":"Norouzzadeh","year":"2018","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"3019","DOI":"10.1007\/s10531-022-02472-z","article-title":"Inter-observer variance and agreement of wildlife information extracted from camera trap images","volume":"31","author":"Zett","year":"2022","journal-title":"Biodivers. Conserv."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1038\/sdata.2015.26","article-title":"Snapshot Serengeti, high-frequency annotated camera trap images of 40 mammalian species in an African savanna","volume":"2","author":"Swanson","year":"2015","journal-title":"Sci. Data"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1007\/s10980-015-0262-9","article-title":"Volunteer-run cameras as distributed sensors for macrosystem mammal research","volume":"31","author":"McShea","year":"2016","journal-title":"Landsc. Ecol."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"831","DOI":"10.1111\/aje.12540","article-title":"The spotted ghost: Density and distribution of serval Leptailurus serval in Namibia","volume":"56","author":"Edwards","year":"2018","journal-title":"Afr. J. Ecol."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"182","DOI":"10.1111\/aje.12641","article-title":"Dyadic associations reveal clan size and social network structure in the fission\u2013fusion society of spotted hyaenas","volume":"58","author":"Stratford","year":"2020","journal-title":"Afr. J. Ecol."},{"key":"ref_8","unstructured":"Zhang, Y., Jiang, H., Miura, Y., Manning, C.D., and Langlotz, C.P. (2020). Contrastive learning of medical visual representations from paired images and text (2020). arXiv."},{"key":"ref_9","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning, Proceedings of Machine Learning Research, Virtual."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"293","DOI":"10.1016\/j.neucom.2022.07.028","article-title":"CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning","volume":"508","author":"Luo","year":"2022","journal-title":"Neurocomputing"},{"key":"ref_11","unstructured":"Jia, C., Yang, Y., Xia, Y., Chen, Y.T., Parekh, Z., Pham, H., Le, Q., Sung, Y.H., Li, Z., and Duerig, T. (2021, January 18\u201324). Scaling up visual and vision-language representation learning with noisy text supervision. Proceedings of the International Conference on Machine Learning, Proceedings of Machine Learning Research, Virtual."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020, January 16\u201318). Momentum contrast for unsupervised visual representation learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"ref_13","unstructured":"Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. (2020, January 13\u201318). A simple framework for contrastive learning of visual representations. Proceedings of the International Conference on Machine Learning, PMLR, Virtual."},{"key":"ref_14","unstructured":"Li, J., Zhou, P., Xiong, C., and Hoi, S.C. (2021, January 3\u20137). Prototypical Contrastive Learning of Unsupervised Representation. Proceedings of the International Conference on Learning Representations, ICLR2021, Virtual."},{"key":"ref_15","unstructured":"Li, J., Xiong, C., and Hoi, S. (2021, January 3\u20137). MoPro: Webly Supervised Learning with Momentum Prototypes. Proceedings of the International Conference on Learning Representations, ICLR2021, Virtual."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Chen, Y.C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J. (2020, January 23\u201328). Uniter: Universal image-text representation learning. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK.","DOI":"10.1007\/978-3-030-58577-8_7"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"5412","DOI":"10.1109\/TNNLS.2020.2967597","article-title":"Cross-modal attention with semantic consistence for image\u2013text matching","volume":"31","author":"Xu","year":"2020","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Diao, H., Zhang, Y., Ma, L., and Lu, H. (2021, January 2\u20139). Similarity reasoning and filtration for image-text matching. Proceedings of the AAAI Conference on Artificial Intelligence, AAAI, Virtual.","DOI":"10.1609\/aaai.v35i2.16209"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., and Wei, F. (2020, January 23\u201328). Oscar: Object-semantics aligned pre-training for vision-language tasks. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK.","DOI":"10.1007\/978-3-030-58577-8_8"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster r-cnn: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","unstructured":"Gu, X., Lin, T.Y., Kuo, W., and Cui, Y. (2021). Open-vocabulary object detection via vision and language knowledge distillation. arXiv."},{"key":"ref_22","unstructured":"Li, B., Weinberger, K.Q., Belongie, S., Koltun, V., and Ranftl, R. (2022). Language-driven semantic segmentation. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Xu, J., De Mello, S., Liu, S., Byeon, W., Breuel, T., Kautz, J., and Wang, X. (2022, January 18\u201324). Groupvit: Semantic segmentation emerges from text supervision. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01760"},{"key":"ref_24","unstructured":"Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (2017). Proceedings of the Advances in Neural Information Processing Systems, Curran Associates, Inc."},{"key":"ref_25","unstructured":"Kim, W., Son, B., and Kim, I. (2021, January 18\u201324). Vilt: Vision-and-language transformer without convolution or region supervision. Proceedings of the International Conference on Machine Learning, PMLR, Virtual."},{"key":"ref_26","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_27","unstructured":"Bao, H., Wang, W., Dong, L., and Wei, F. (2022). Vl-beit: Generative vision-language pretraining. arXiv."},{"key":"ref_28","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"He, K., Chen, X., Xie, S., Li, Y., Doll\u00e1r, P., and Girshick, R. (2022, January 18\u201324). Masked autoencoders are scalable vision learners. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"ref_30","unstructured":"Li, J., Li, D., Xiong, C., and Hoi, S. (2022, January 17\u201323). Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. Proceedings of the International Conference on Machine Learning, PMLR, Baltimore, MD, USA."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O.K., Singhal, S., and Som, S. (2022). Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv.","DOI":"10.1109\/CVPR52729.2023.01838"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Li, Y., Fan, H., Hu, R., Feichtenhofer, C., and He, K. (2022). Scaling Language-Image Pre-training via Masking. arXiv.","DOI":"10.1109\/CVPR52729.2023.02240"},{"key":"ref_33","unstructured":"Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (2022). Proceedings of the Advances in Neural Information Processing Systems, Curran Associates, Inc."},{"key":"ref_34","unstructured":"Yu, J., Wang, Z., Vasudevan, V., Yeung, L., Seyedhosseini, M., and Wu, Y. (2022). Coca: Contrastive captioners are image-text foundation models. arXiv."},{"key":"ref_35","unstructured":"Wang, Z., Yu, J., Yu, A.W., Dai, Z., Tsvetkov, Y., and Cao, Y. (2021). Simvlm: Simple visual language model pretraining with weak supervision. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Li, C., Xu, H., Tian, J., Wang, W., Yan, M., Bi, B., Ye, J., Chen, H., Xu, G., and Cao, Z. (2022). mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections. arXiv.","DOI":"10.18653\/v1\/2022.emnlp-main.488"},{"key":"ref_37","unstructured":"Oord, A.v.d., Li, Y., and Vinyals, O. (2018). Representation learning with contrastive predictive coding. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Karpathy, A., and Fei-Fei, L. (2015, January 7\u201312). Deep visual-semantic alignments for generating image descriptions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Fleet, D., Pajdla, T., Schiele, B., and Tuytelaars, T. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the Computer Vision\u2014ECCV 2014, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10599-4"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2002, January 6\u201312). Bleu: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Philadelphia, PA, USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Denkowski, M., and Lavie, A. (2014, January 26\u201327). Meteor universal: Language specific translation evaluation for any target language. Proceedings of the Ninth Workshop on Statistical Machine Translation, Baltimore, ND, USA.","DOI":"10.3115\/v1\/W14-3348"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Lawrence Zitnick, C., and Parikh, D. (2015, January 7\u201312). Cider: Consensus-based image description evaluation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Anderson, P., Fernando, B., Johnson, M., and Gould, S. (2016, January 11\u201314). Spice: Semantic propositional image caption evaluation. Proceedings of the Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L. (2018, January 18\u201322). Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Zhou, L., Palangi, H., Zhang, L., Hu, H., Corso, J., and Gao, J. (2020, January 7\u201312). Unified vision-language pre-training for image captioning and vqa. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.7005"},{"key":"ref_46","unstructured":"Mokady, R., Hertz, A., and Bermano, A.H. (2021). Clipcap: Clip prefix for image captioning. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Dou, Z.Y., Xu, Y., Gan, Z., Wang, J., Wang, S., Wang, L., Zhu, C., Zhang, P., Yuan, L., and Peng, N. (2022, January 18\u201324). An empirical study of training end-to-end vision-and-language transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01763"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Cheng, M., Sun, Y., Wang, L., Zhu, X., Yao, K., Chen, J., Song, G., Han, J., Liu, J., and Ding, E. (2022, January 18\u201324). ViSTA: Vision and scene text aggregation for cross-modal retrieval. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00512"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Messina, N., Stefanini, M., Cornia, M., Baraldi, L., Falchi, F., Amato, G., and Cucchiara, R. (2022, January 14\u201316). ALADIN: Distilling Fine-grained Alignment Scores for Efficient Image-Text Matching and Retrieval. Proceedings of the 19th International Conference on Content-Based Multimedia Indexing, Graz, Austria.","DOI":"10.1145\/3549555.3549576"},{"key":"ref_50","unstructured":"Diao, Q., Jiang, Y., Wen, B., Sun, J., and Yuan, Z. (2022). Metaformer: A unified meta framework for fine-grained recognition. arXiv."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Girdhar, R., Singh, M., Ravi, N., van der Maaten, L., Joulin, A., and Misra, I. (2022, January 18\u201324). Omnivore: A single model for many visual modalities. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01563"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Touvron, H., Sablayrolles, A., Douze, M., Cord, M., and J\u00e9gou, H. (2021, January 11\u201317). Grafit: Learning fine-grained image representations with coarse labels. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00091"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Tian, C., Wang, W., Zhu, X., Dai, J., and Qiao, Y. (2022, January 23\u201327). Vl-ltr: Learning class-wise visual-linguistic representation for long-tailed visual recognition. Proceedings of the Computer Vision\u2013ECCV 2022: 17th European Conference, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19806-9_5"},{"key":"ref_54","unstructured":"Gesmundo, A. (2022). A Continual Development Methodology for Large-scale Multitask Dynamic ML Systems. arXiv."},{"key":"ref_55","unstructured":"Liu, J., Huang, X., Liu, Y., and Li, H. (2022). Mixmim: Mixed and masked image modeling for efficient visual representation learning. arXiv."},{"key":"ref_56","unstructured":"Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and J\u00e9gou, H. (2021, January 18\u201324). Training data-efficient image transformers & distillation through attention. Proceedings of the International Conference on Machine Learning. PMLR, Virtual."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Yuan, K., Guo, S., Liu, Z., Zhou, A., Yu, F., and Wu, W. (2021, January 10\u201317). Incorporating convolution designs into visual transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, ICCV, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00062"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Cui, J., Zhong, Z., Tian, Z., Liu, S., Yu, B., and Jia, J. (2022). Generalized Parametric Contrastive Learning. arXiv.","DOI":"10.1109\/ICCV48922.2021.00075"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Van Horn, G., Mac Aodha, O., Song, Y., Cui, Y., Sun, C., Shepard, A., Adam, H., Perona, P., and Belongie, S. (2018, January 18\u201323). The iNaturalist Species Classification and Detection Dataset. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00914"},{"key":"ref_60","unstructured":"Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., and Wortsman, M. (2022). Laion-5b: An open large-scale dataset for training next generation image-text models. arXiv."},{"key":"ref_61","unstructured":"Sanh, V., Webson, A., Raffel, C., Bach, S.H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T.L., and Raja, A. (2021). Multitask prompted training enables zero-shot task generalization. arXiv."},{"key":"ref_62","unstructured":"Honnibal, M., Montani, I., Van Landeghem, S., and Boyd, A. (2020). SpaCy: INDUSTRIAL-Strength Natural Language Processing in Python, Zenodo."},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Yan, J., Xiao, Y., Mukherjee, S., Lin, B.Y., Jia, R., and Ren, X. (2021). On the Robustness of Reading Comprehension Models to Entity Renaming. arXiv.","DOI":"10.18653\/v1\/2022.naacl-main.37"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/10\/3130\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:42:28Z","timestamp":1760107348000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/10\/3130"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,14]]},"references-count":63,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2024,5]]}},"alternative-id":["s24103130"],"URL":"https:\/\/doi.org\/10.3390\/s24103130","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5,14]]}}}