{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T16:13:58Z","timestamp":1783527238532,"version":"3.55.0"},"reference-count":69,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2023,9,18]],"date-time":"2023-09-18T00:00:00Z","timestamp":1694995200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["11974373"],"award-info":[{"award-number":["11974373"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61932005"],"award-info":[{"award-number":["61932005"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Key Project of National Natural Science Foundation of China","award":["11974373"],"award-info":[{"award-number":["11974373"]}]},{"name":"Key Project of National Natural Science Foundation of China","award":["61932005"],"award-info":[{"award-number":["61932005"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Recent research has shown that visual\u2013text pretrained models perform well in traditional vision tasks. CLIP, as the most influential work, has garnered significant attention from researchers. Thanks to its excellent visual representation capabilities, many recent studies have used CLIP for pixel-level tasks. We explore the potential abilities of CLIP in the field of few-shot segmentation. The current mainstream approach is to utilize support and query features to generate class prototypes and then use the prototype features to match image features. We propose a new method that utilizes CLIP to extract text features for a specific class. These text features are then used as training samples to participate in the model\u2019s training process. The addition of text features enables model to extract features that contain richer semantic information, thus making it easier to capture potential class information. To better match the query image features, we also propose a new prototype generation method that incorporates multi-modal fusion features of text and images in the prototype generation process. Adaptive query prototypes were generated by combining foreground and background information from the images with the multi-modal support prototype, thereby allowing for a better matching of image features and improved segmentation accuracy. We provide a new perspective to the task of few-shot segmentation in multi-modal scenarios. Experiments demonstrate that our proposed method achieves excellent results on two common datasets, PASCAL-5i and COCO-20i.<\/jats:p>","DOI":"10.3390\/e25091353","type":"journal-article","created":{"date-parts":[[2023,9,18]],"date-time":"2023-09-18T05:59:06Z","timestamp":1695016746000},"page":"1353","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":19,"title":["CLIP-Driven Prototype Network for Few-Shot Semantic Segmentation"],"prefix":"10.3390","volume":"25","author":[{"given":"Shi-Cheng","family":"Guo","sequence":"first","affiliation":[{"name":"College of Computer Science and Engineering, Shandong University of Science and Technology, Qingdao 266590, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5728-5092","authenticated-orcid":false,"given":"Shang-Kun","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Engineering, Shandong University of Science and Technology, Qingdao 266590, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jing-Yu","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Engineering, Shandong University of Science and Technology, Qingdao 266590, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wei-Min","family":"Zheng","sequence":"additional","affiliation":[{"name":"College of Computer Science and Engineering, Shandong University of Science and Technology, Qingdao 266590, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cheng-Yu","family":"Jiang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Engineering, Shandong University of Science and Technology, Qingdao 266590, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,9,18]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the 13th European Conference of the Computer Vision (ECCV 2014), Zurich, Switzerland. Proceedings\u2014Part V 13.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_3","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20136). Imagenet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA."},{"key":"ref_4","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_5","unstructured":"Siam, M., Oreshkin, B.N., and Jagersand, M. (November, January 27). Amp: Adaptive masked proxies for few-shot segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Liu, L., Cao, J., Liu, M., Guo, Y., Chen, Q., and Tan, M. (2020, January 12\u201316). Dynamic extension nets for few-shot semantic segmentation. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3413915"},{"key":"ref_7","unstructured":"Nguyen, K., and Todorovic, S. (November, January 27). Feature weighting and boosting for few-shot segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_8","unstructured":"Wang, K., Liew, J.H., Zou, Y., Zhou, D., and Feng, J. (November, January 27). Panet: Few-shot image semantic segmentation with prototype alignment. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Liu, Y., Zhang, X., Zhang, S., and He, X. (2020, January 23\u201328). Part-aware prototype network for few-shot semantic segmentation. Proceedings of the 16th European Conference of the Computer Vision (ECCV 2020), Glasgow, UK. Proceedings\u2014Part IX 16.","DOI":"10.1007\/978-3-030-58545-7_9"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Lin, Z., Yu, S., Kuang, Z., Pathak, D., and Ramanan, D. (2023, January 18\u201322). Multimodality helps unimodality: Cross-modal few-shot learning with multimodal models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01852"},{"key":"ref_11","unstructured":"Li, J., Li, D., Xiong, C., and Hoi, S. (2022, January 17\u201323). Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. Proceedings of the International Conference on Machine Learning, Baltimore, MD, USA."},{"key":"ref_12","unstructured":"Lu, J., Batra, D., Parikh, D., and Lee, S. (2019, January 8\u201314). Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O.K., Singhal, S., and Som, S. (2022). Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv.","DOI":"10.1109\/CVPR52729.2023.01838"},{"key":"ref_14","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning, Online."},{"key":"ref_15","unstructured":"Gao, P., Geng, S., Zhang, R., Ma, T., Fang, R., Zhang, Y., Li, H., and Qiao, Y. (2021). Clip-adapter: Better vision-language models with feature adapters. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2337","DOI":"10.1007\/s11263-022-01653-1","article-title":"Learning to prompt for vision-language models","volume":"130","author":"Zhou","year":"2022","journal-title":"Int. J. Comput. Vis."},{"key":"ref_17","unstructured":"Zhang, R., Fang, R., Zhang, W., Gao, P., Li, K., Dai, J., Qiao, Y., and Li, H. (2021). Tip-adapter: Training-free clip-adapter for better vision-language modeling. arXiv."},{"key":"ref_18","unstructured":"Li, B., Weinberger, K.Q., Belongie, S., Koltun, V., and Ranftl, R. (2021, January 3\u20137). Language-driven Semantic Segmentation. Proceedings of the International Conference on Learning Representations, Online."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Rao, Y., Zhao, W., Chen, G., Tang, Y., Zhu, Z., Huang, G., Zhou, J., and Lu, J. (2022, January 18\u201324). Denseclip: Language-guided dense prediction with context-aware prompting. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01755"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Xu, J., De Mello, S., Liu, S., Byeon, W., Breuel, T., Kautz, J., and Wang, X. (2022, January 18\u201324). Groupvit: Semantic segmentation emerges from text supervision. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01760"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhou, K., Yang, J., Loy, C.C., and Liu, Z. (2022, January 18\u201324). Conditional prompt learning for vision-language models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01631"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Khattak, M.U., Rasheed, H., Maaz, M., Khan, S., and Khan, F.S. (2023, January 18\u201322). Maple: Multi-modal prompt learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01832"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Liu, W., Zhang, C., Lin, G., and Liu, F. (2020, January 14\u201319). Crnet: Cross-reference networks for few-shot segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00422"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1145\/1015706.1015720","article-title":"\u201cGrabCut\u201d interactive foreground extraction using iterated graph cuts","volume":"23","author":"Rother","year":"2004","journal-title":"ACM Trans. Graph. (TOG)"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"187","DOI":"10.3233\/FI-2000-411207","article-title":"The watershed transform: Definitions, algorithms and parallelization strategies","volume":"41","author":"Roerdink","year":"2000","journal-title":"Fundam. Inform."},{"key":"ref_27","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the 18th International Conference of the Medical Image Computing and Computer-Assisted Intervention (MICCAI 2015), Munich, Germany. Proceedings\u2014Part III 18."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. (2018, January 8\u201314). Encoder-decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"ref_30","unstructured":"Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A.L. (2014). Semantic image segmentation with deep convolutional nets and fully connected crfs. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_32","unstructured":"Chen, L.C., Papandreou, G., Schroff, F., and Adam, H. (2017). Rethinking atrous convolution for semantic image segmentation. arXiv."},{"key":"ref_33","unstructured":"Oktay, O., Schlemper, J., Folgoc, L.L., Lee, M., Heinrich, M., Misawa, K., Mori, K., McDonagh, S., Hammerla, N.Y., and Kainz, B. (2018). Attention u-net: Learning where to look for the pancreas. arXiv."},{"key":"ref_34","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020, January 26\u201330). An Image is Worth 16 \u00d7 16 Words: Transformers for Image Recognition at Scale. Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., Fu, Y., Feng, J., Xiang, T., and Torr, P.H. (2021, January 19\u201325). Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00681"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Strudel, R., Garcia, R., Laptev, I., and Schmid, C. (2021, January 11\u201317). Segmenter: Transformer for semantic segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00717"},{"key":"ref_37","unstructured":"Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J.M., and Luo, P. (2021, January 6\u201314). SegFormer: Simple and efficient design for semantic segmentation with transformers. Proceedings of the Advances in Neural Information Processing Systems, Online."},{"key":"ref_38","unstructured":"Chen, W.Y., Liu, Y.C., Kira, Z., Wang, Y.C.F., and Huang, J.B. (2019, January 6\u20139). A Closer Look at Few-shot Classification. Proceedings of the International Conference on Learning Representations, New Orleans, LA, USA."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Gidaris, S., and Komodakis, N. (2018, January 18\u201323). Dynamic few-shot visual learning without forgetting. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00459"},{"key":"ref_40","unstructured":"Dhillon, G.S., Chaudhari, P., Ravichandran, A., and Soatto, S. (2019). A baseline for few-shot image classification. arXiv."},{"key":"ref_41","unstructured":"Lake, B., Lee, C.y., Glass, J., and Tenenbaum, J. (2014, January 23\u201326). One-shot learning of generative speech concepts. Proceedings of the Annual Meeting of the Cognitive Science Society, Quebec City, QC, Canada."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Hariharan, B., and Girshick, R. (2017, January 22\u201329). Low-shot visual recognition by shrinking and hallucinating features. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.328"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Cubuk, E.D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q.V. (2018). Autoaugment: Learning augmentation policies from data. arXiv.","DOI":"10.1109\/CVPR.2019.00020"},{"key":"ref_44","unstructured":"Schwartz, E., Karlinsky, L., Shtok, J., Harary, S., Marder, M., Kumar, A., Feris, R., Giryes, R., and Bronstein, A. (2018, January 3\u20138). \u0394-encoder: An effective sample synthesis method for few-shot object recognition. Proceedings of the Annual Conference on Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_45","unstructured":"Allen, K., Shelhamer, E., Shin, H., and Tenenbaum, J. (2019, January 9\u201315). Infinite mixture prototypes for few-shot learning. Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA."},{"key":"ref_46","unstructured":"Koch, G., Zemel, R., and Salakhutdinov, R. (2015, January 6\u201311). Siamese neural networks for one-shot image recognition. Proceedings of the International Conference on Machine Learning (ICML 2015), Lille, France."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Li, W., Wang, L., Xu, J., Huo, J., Gao, Y., and Luo, J. (2019, January 9\u201315). Revisiting local descriptor based image-to-class measure for few-shot learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00743"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Shaban, A., Bansal, S., Liu, Z., Essa, I., and Boots, B. (2017). One-shot learning for semantic segmentation. arXiv.","DOI":"10.5244\/C.31.167"},{"key":"ref_49","unstructured":"Dong, N., and Xing, E.P. (2018, January 3\u20136). Few-shot semantic segmentation with prototype learning. Proceedings of the 2018 British Machine Vision Conference (BMVC 2018), Newcastle, UK."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"3855","DOI":"10.1109\/TCYB.2020.2992433","article-title":"Sg-one: Similarity guidance network for one-shot semantic segmentation","volume":"50","author":"Zhang","year":"2020","journal-title":"IEEE Trans. Cybern."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Fan, Q., Pei, W., Tai, Y.W., and Tang, C.K. (2022, January 23\u201324). Self-support few-shot semantic segmentation. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19800-7_41"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Zhang, C., Lin, G., Liu, F., Yao, R., and Shen, C. (2019, January 15\u201320). Canet: Class-agnostic segmentation networks with iterative refinement and attentive few-shot learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00536"},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"1050","DOI":"10.1109\/TPAMI.2020.3013717","article-title":"Prior guided feature enrichment network for few-shot segmentation","volume":"44","author":"Tian","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Zhao, Q., Liu, B., Lyu, S., and Chen, H. (2023). A self-distillation embedded supervised affinity attention model for few-shot segmentation. IEEE Trans. Cogn. Dev. Syst.","DOI":"10.1109\/TCDS.2023.3251371"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Min, J., Kang, D., and Cho, M. (2021, January 11\u201317). Hypercorrelation squeeze for few-shot segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00686"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Wang, H., Liu, L., Zhang, W., Zhang, J., Gan, Z., Wang, Y., Wang, C., and Wang, H. (2023). Iterative Few-shot Semantic Segmentation from Image Label Text. arXiv.","DOI":"10.24963\/ijcai.2022\/193"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Zhou, C., Loy, C.C., and Dai, B. (2022, January 23\u201324). Extract free dense labels from clip. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19815-1_40"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"L\u00fcddecke, T., and Ecker, A. (2022, January 18\u201324). Image segmentation using text and image prompts. Proceedings of the CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00695"},{"key":"ref_59","unstructured":"Han, M., Zheng, H., Wang, C., Luo, Y., Hu, H., Zhang, J., and Wen, Y. (2023). PartSeg: Few-shot Part Segmentation via Part-aware Prompt Learning. arXiv."},{"key":"ref_60","unstructured":"Shuai, C., Fanman, M., Runtong, Z., Heqian, Q., Hongliang, L., Qingbo, W., and Linfeng, X. (2023). Visual and Textual Prior Guided Mask Assemble for Few-Shot Segmentation and Beyond. arXiv."},{"key":"ref_61","unstructured":"Vinyals, O., Blundell, C., Lillicrap, T., and Wierstra, D. (2016, January 5\u201310). Matching networks for one shot learning. Proceedings of the Advances in Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_62","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","article-title":"The pascal visual object classes (voc) challenge","volume":"88","author":"Everingham","year":"2010","journal-title":"Int. J. Comput. Vis."},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Lu, Z., He, S., Zhu, X., Zhang, L., Song, Y.Z., and Xiang, T. (2021, January 11\u201317). Simpler is better: Few-shot semantic segmentation with classifier weight transformer. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00862"},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"Yang, L., Zhuo, W., Qi, L., Shi, Y., and Gao, Y. (2021, January 11\u201317). Mining latent classes for few-shot segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00860"},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Wu, Z., Shi, X., Lin, G., and Cai, J. (2021, January 11\u201317). Learning meta-class memory for few-shot semantic segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00056"},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"Lang, C., Cheng, G., Tu, B., and Han, J. (2022, January 18\u201324). Learning what not to segment: A new perspective on few-shot segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00789"},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Peng, B., Tian, Z., Wu, X., Wang, C., Liu, S., Su, J., and Jia, J. (2023, January 18\u201322). Hierarchical Dense Correlation Distillation for Few-Shot Segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02264"},{"key":"ref_68","doi-asserted-by":"crossref","unstructured":"Yang, B., Liu, C., Li, B., Jiao, J., and Ye, Q. (2020, January 23\u201328). Prototype mixture models for few-shot semantic segmentation. Proceedings of the 16th European Conference of the Computer Vision (ECCV 2020), Glasgow, UK. Proceedings\u2014Part VIII 16.","DOI":"10.1007\/978-3-030-58598-3_45"},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Zhang, B., Xiao, J., and Qin, T. (2021, January 19\u201325). Self-guided and cross-guided learning for few-shot segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00821"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/25\/9\/1353\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:52:47Z","timestamp":1760129567000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/25\/9\/1353"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,18]]},"references-count":69,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2023,9]]}},"alternative-id":["e25091353"],"URL":"https:\/\/doi.org\/10.3390\/e25091353","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,18]]}}}