{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T15:02:25Z","timestamp":1782313345645,"version":"3.54.5"},"reference-count":23,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,6,10]],"date-time":"2025-06-10T00:00:00Z","timestamp":1749513600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Artif. Intell."],"abstract":"<jats:p>As powerful pre-trained vision-language models (VLMs) like CLIP gain prominence, numerous studies have attempted to combine VLMs for downstream tasks. Among these, prompt learning has been validated as an effective method for adapting to new tasks, which only requires a small number of parameters. However, current prompt learning methods face two challenges: first, a single soft prompt struggles to capture the diverse styles and patterns within a dataset; second, fine-tuning soft prompts is prone to overfitting. To address these challenges, we propose a mixture-of-prompts learning method incorporating a routing module. This module is able to capture a dataset's varied styles and dynamically select the most suitable prompts for each instance. Additionally, we introduce a novel gating mechanism to ensure the router selects prompts based on their similarity to hard prompt templates, which both retains knowledge from hard prompts and improves selection accuracy. We also implement semantically grouped text-level supervision, initializing each soft prompt with the token embeddings of manually designed templates from its group and applying a contrastive loss between the resulted text feature and hard prompt encoded text feature. This supervision ensures that the text features derived from soft prompts remain close to those from their corresponding hard prompts, preserving initial knowledge and mitigating overfitting. Our method has been validated on 11 datasets, demonstrating evident improvements in few-shot learning, domain generalization, and base-to-new generalization scenarios compared to existing baselines. Our approach establishes that multi-prompt specialization with knowledge-preserving routing effectively bridges the adaptability-generalization tradeoff in VLM deployment. The code will be available at <jats:ext-link>https:\/\/github.com\/dyabel\/mocoop<\/jats:ext-link>.<\/jats:p>","DOI":"10.3389\/frai.2025.1580973","type":"journal-article","created":{"date-parts":[[2025,6,10]],"date-time":"2025-06-10T05:23:21Z","timestamp":1749533001000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Mixture of prompts learning for vision-language models"],"prefix":"10.3389","volume":"8","author":[{"given":"Yu","family":"Du","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tong","family":"Niu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rong","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2025,6,10]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2303.08774","article-title":"Gpt-4 technical report","author":"Achiam","year":"2023","journal-title":"arXiv preprint arXiv"},{"key":"B2","doi-asserted-by":"crossref","first-page":"446","DOI":"10.1007\/978-3-319-10599-4_29","article-title":"\u201cFood-101-mining discriminative components with random forests,\u201d","volume-title":"Computer vision-ECCV 2014: 13th European conference, zurich, Switzerland, September 6-12, 2014, proceedings, part VI 13","author":"Bossard","year":"2014"},{"key":"B3","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02225","article-title":"Lasp: text-to-text optimization for language-aware soft prompting of vision and language models","author":"Bulat","year":"2022","journal-title":"arXiv preprint arXiv"},{"key":"B4","first-page":"3606","article-title":"\u201cDescribing textures in the wild,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Cimpoi","year":"2014"},{"key":"B5","first-page":"88","article-title":"\u201cVqgan-clip: open domain image generation and editing with natural language guidance,\u201d","volume-title":"European Conference on Computer Vision","author":"Crowson","year":"2022"},{"key":"B6","doi-asserted-by":"crossref","first-page":"248","DOI":"10.1109\/CVPR.2009.5206848","article-title":"\u201cImagenet: a large-scale hierarchical image database,\u201d","volume-title":"2009 IEEE Conference on Computer Vision and Pattern Recognition","author":"Deng","year":"2009"},{"key":"B7","doi-asserted-by":"publisher","author":"Eslami","year":"2021","DOI":"10.48550\/arXiv.2112.13906"},{"key":"B8","doi-asserted-by":"crossref","first-page":"178","DOI":"10.1109\/CVPR.2004.383","article-title":"\u201cLearning generative visual models from few training examples: an incremental bayesian approach tested on 101 object categories,\u201d","volume-title":"2004 Conference on Computer Vision and Pattern Recognition Workshop","author":"Fei-Fei","year":"2004"},{"key":"B9","doi-asserted-by":"publisher","first-page":"2217","DOI":"10.1109\/JSTARS.2019.2918242","article-title":"Eurosat: a novel dataset and deep learning benchmark for land use and land cover classification","volume":"12","author":"Helber","year":"2019","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens"},{"key":"B10","first-page":"19113","article-title":"\u201cMaple: multi-modal prompt learning,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Khattak","year":"2023"},{"key":"B11","first-page":"554","article-title":"\u201c3D object representations for fine-grained categorization,\u201d","volume-title":"Proceedings of the IEEE International Conference on Computer Vision Workshops","author":"Krause","year":"2013"},{"key":"B12","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1306.5151","article-title":"Fine-grained visual classification of aircraft","author":"Maji","year":"2013","journal-title":"arXiv preprint arXiv:1306.5151"},{"key":"B13","doi-asserted-by":"crossref","first-page":"722","DOI":"10.1109\/ICVGIP.2008.47","article-title":"\u201cAutomated flower classification over a large number of classes,\u201d","volume-title":"2008 Sixth Indian Conference on Computer Vision, Graphics and Image Processing","author":"Nilsback","year":"2008"},{"key":"B14","doi-asserted-by":"crossref","first-page":"3498","DOI":"10.1109\/CVPR.2012.6248092","article-title":"\u201cCats and dogs,\u201d","volume-title":"2012 IEEE Conference on Computer Vision and Pattern Recognition","author":"Parkhi","year":"2012"},{"key":"B15","first-page":"8748","article-title":"\u201cLearning transferable visual models from natural language supervision,\u201d","volume-title":"International Conference on Machine Learning","author":"Radford","year":"2021"},{"key":"B16","doi-asserted-by":"publisher","first-page":"36479","DOI":"10.48550\/arXiv.2205.11487","article-title":"Photorealistic text-to-image diffusion models with deep language understanding","volume":"35","author":"Saharia","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst"},{"key":"B17","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1212.0402","article-title":"Ucf101: a dataset of 101 human actions classes from videos in the wild","author":"Soomro","year":"2012","journal-title":"arXiv preprint arXiv:1212.0402"},{"key":"B18","doi-asserted-by":"crossref","first-page":"3485","DOI":"10.1109\/CVPR.2010.5539970","article-title":"\u201cSun database: large-scale scene recognition from abbey to zoo,\u201d","volume-title":"2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"Xiao","year":"2010"},{"key":"B19","first-page":"6757","article-title":"\u201cVisual-language prompt tuning with knowledge-guided context optimization,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yao","year":"2023"},{"key":"B20","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2210.07225","article-title":"Unified vision and language prompt learning","author":"Zang","year":"2022","journal-title":"arXiv preprint arXiv:2210.07225"},{"key":"B21","first-page":"16816","article-title":"\u201cConditional prompt learning for vision-language models,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhou","year":""},{"key":"B22","doi-asserted-by":"publisher","first-page":"2337","DOI":"10.1007\/s11263-022-01653-1","article-title":"Learning to prompt for vision-language models","volume":"130","author":"Zhou","year":"","journal-title":"Int. J. Comput. Vis"},{"key":"B23","first-page":"15659","article-title":"\u201cPrompt-aligned gradient for prompt tuning,\u201d","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhu","year":"2023"}],"container-title":["Frontiers in Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2025.1580973\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,10]],"date-time":"2025-06-10T05:23:27Z","timestamp":1749533007000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2025.1580973\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,10]]},"references-count":23,"alternative-id":["10.3389\/frai.2025.1580973"],"URL":"https:\/\/doi.org\/10.3389\/frai.2025.1580973","relation":{},"ISSN":["2624-8212"],"issn-type":[{"value":"2624-8212","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,10]]},"article-number":"1580973"}}