{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,7]],"date-time":"2025-11-07T15:49:19Z","timestamp":1762530559744,"version":"build-2065373602"},"reference-count":50,"publisher":"Springer Science and Business Media LLC","issue":"15","license":[{"start":{"date-parts":[[2025,10,1]],"date-time":"2025-10-01T00:00:00Z","timestamp":1759276800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2025,10,1]],"date-time":"2025-10-01T00:00:00Z","timestamp":1759276800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"funder":[{"DOI":"10.13039\/501100002615","name":"Yong In University","doi-asserted-by":"publisher","award":["KY22C215"],"award-info":[{"award-number":["KY22C215"]}],"id":[{"id":"10.13039\/501100002615","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Appl Intell"],"published-print":{"date-parts":[[2025,10]]},"DOI":"10.1007\/s10489-025-06919-y","type":"journal-article","created":{"date-parts":[[2025,10,3]],"date-time":"2025-10-03T12:29:27Z","timestamp":1759494567000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Multimodal prompt learning with selective feature fusion: towards robust cross-modal alignment"],"prefix":"10.1007","volume":"55","author":[{"given":"Jiabao","family":"Han","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yahui","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Zhong","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ying","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xichao","family":"Yuan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,10,3]]},"reference":[{"key":"6919_CR1","doi-asserted-by":"publisher","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 770\u2013778. https:\/\/doi.org\/10.1109\/CVPR.2016.90","DOI":"10.1109\/CVPR.2016.90"},{"key":"6919_CR2","unstructured":"Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I (2021) Learning transferable visual models from natural language supervision. In: Meila M, Zhang T (eds) Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event. Proceedings of Machine Learning Research, vol 139. PMLR, ???, pp 8748\u20138763. http:\/\/proceedings.mlr.press\/v139\/radford21a.html"},{"key":"6919_CR3","unstructured":"Howard AG, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, Andreetto M, Adam H (2017) Mobilenets: efficient convolutional neural networks for mobile vision applications. arXiv:1704.04861"},{"key":"6919_CR4","first-page":"11960","volume":"34","author":"Y Wang","year":"2021","unstructured":"Wang Y, Huang R, Song S, Huang Z, Huang G (2021) Not all images are worth 16x16 words: dynamic transformers for efficient image recognition. Adv Neural Inf Process Syst 34:11960\u201311973","journal-title":"Adv Neural Inf Process Syst"},{"key":"6919_CR5","unstructured":"Touvron H, Cord M, Douze M, Massa F, Sablayrolles A, J\u00e9gou H (2021) Training data-efficient image transformers & distillation through attention. In: Meila M, Zhang T (eds) Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event. Proceedings of Machine Learning Research, vol 139. PMLR, ???, pp 10347\u201310357. http:\/\/proceedings.mlr.press\/v139\/touvron21a.html"},{"key":"6919_CR6","unstructured":"Gao P, Lu J, Li H, Mottaghi R, Kembhavi A (2021) Container: context aggregation network. arXiv:2106.01401"},{"key":"6919_CR7","first-page":"25346","volume":"34","author":"M Mao","year":"2021","unstructured":"Mao M, Zhang R, Zheng H, Ma T, Peng Y, Ding E, Zhang B, Han S et al (2021) Dual-stream network for visual recognition. Adv Neural Inf Process Syst 34:25346\u201325358","journal-title":"Adv Neural Inf Process Syst"},{"key":"6919_CR8","unstructured":"Bahng H, Jahanian A, Sankaranarayanan S, Isola P (2022) Exploring visual prompts for adapting large-scale models. arXiv:2203.17274"},{"key":"6919_CR9","doi-asserted-by":"publisher","unstructured":"Jia M, Tang L, Chen B, Cardie C, Belongie SJ, Hariharan B, Lim S (2022) Visual prompt tuning. In: Avidan S, Brostow GJ, Ciss\u00e9 M, Farinella GM, Hassner T (eds) Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part XXXIII. Lecture Notes in Computer Science, vol 13693. Springer, ???, pp 709\u2013727. https:\/\/doi.org\/10.1007\/978-3-031-19827-4_41","DOI":"10.1007\/978-3-031-19827-4_41"},{"key":"6919_CR10","doi-asserted-by":"crossref","unstructured":"Khattak MU, Rasheed H, Maaz M, Khan S, Khan FS (2023) Maple: multi-modal prompt learning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 19113\u201319122","DOI":"10.1109\/CVPR52729.2023.01832"},{"key":"6919_CR11","doi-asserted-by":"crossref","unstructured":"Xing Y, Wu Q, Cheng D, Zhang S, Liang G, Wang P, Zhang Y (2023) Dual modality prompt tuning for vision-language pre-trained model. IEEE Trans Multimed","DOI":"10.1109\/TMM.2023.3291588"},{"key":"6919_CR12","doi-asserted-by":"crossref","unstructured":"Zhou K, Yang J, Loy CC, Liu Z (2022) Conditional prompt learning for vision-language models. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 16816\u201316825","DOI":"10.1109\/CVPR52688.2022.01631"},{"key":"6919_CR13","unstructured":"Ren S, He K, Girshick R, Sun J (2015) Faster r-cnn: towards real-time object detection with region proposal networks. Adv Neural Inf Process Syst 28"},{"key":"6919_CR14","doi-asserted-by":"crossref","unstructured":"Carion N, Massa F, Synnaeve G, Usunier N, Kirillov A, Zagoruyko S (2020) End-to-end object detection with transformers. In: European conference on computer vision. Springer, pp 213\u2013229","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"6919_CR15","doi-asserted-by":"crossref","unstructured":"Gao P, Zheng M, Wang X, Dai J, Li H (2021) Fast convergence of detr with spatially modulated co-attention. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 3621\u20133630","DOI":"10.1109\/ICCV48922.2021.00360"},{"key":"6919_CR16","unstructured":"Simonyan K, Zisserman A (2014) Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556"},{"key":"6919_CR17","doi-asserted-by":"crossref","unstructured":"Long J, Shelhamer E, Darrell T (2015) Fully convolutional networks for semantic segmentation. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3431\u20133440","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"6919_CR18","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1023\/A:1007379606734","volume":"28","author":"R Caruana","year":"1997","unstructured":"Caruana R (1997) Multitask learning. Mach Learn 28:41\u201375","journal-title":"Mach Learn"},{"key":"6919_CR19","doi-asserted-by":"crossref","unstructured":"Parkhi OM, Vedaldi A, Zisserman A, Jawahar C (2012) Cats and dogs. In: 2012 IEEE conference on computer vision and pattern recognition. IEEE, pp 3498\u20133505","DOI":"10.1109\/CVPR.2012.6248092"},{"key":"6919_CR20","doi-asserted-by":"crossref","unstructured":"Nilsback M-E, Zisserman A (2008) Automated flower classification over a large number of classes. In: 2008 sixth indian conference on computer vision, graphics & image processing. IEEE, pp 722\u2013729","DOI":"10.1109\/ICVGIP.2008.47"},{"issue":"11","key":"6919_CR21","first-page":"1","volume":"2","author":"K Soomro","year":"2012","unstructured":"Soomro K, Zamir AR, Shah M (2012) A dataset of 101 human action classes from videos in the wild. Center Res Comput Vision 2(11):1\u20137","journal-title":"Center Res Comput Vision"},{"key":"6919_CR22","unstructured":"Maji S, Rahtu E, Kannala J, Blaschko M, Vedaldi A (2013) Fine-grained visual classification of aircraft. arXiv:1306.5151"},{"key":"6919_CR23","doi-asserted-by":"crossref","unstructured":"Fei-Fei L, Fergus R, Perona P (2004) Learning generative visual models from few training examples: an incremental bayesian approach tested on 101 object categories. In: 2004 conference on computer vision and pattern recognition workshop. IEEE, pp 178\u2013178","DOI":"10.1109\/CVPR.2004.383"},{"issue":"7","key":"6919_CR24","doi-asserted-by":"publisher","first-page":"2217","DOI":"10.1109\/JSTARS.2019.2918242","volume":"12","author":"P Helber","year":"2019","unstructured":"Helber P, Bischke B, Dengel A, Borth D (2019) Eurosat: a novel dataset and deep learning benchmark for land use and land cover classification. IEEE J Sel Top Appl Earth Obs Remote Sens 12(7):2217\u20132226","journal-title":"IEEE J Sel Top Appl Earth Obs Remote Sens"},{"key":"6919_CR25","doi-asserted-by":"crossref","unstructured":"Cimpoi M, Maji S, Kokkinos I, Mohamed S, Vedaldi A (2014) Describing textures in the wild. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3606\u20133613","DOI":"10.1109\/CVPR.2014.461"},{"key":"6919_CR26","doi-asserted-by":"crossref","unstructured":"Xiao J, Hays J, Ehinger KA, Oliva A, Torralba A (2010) Sun database: large-scale scene recognition from abbey to zoo. In: 2010 IEEE computer society conference on computer vision and pattern recognition. IEEE, pp 3485\u20133492","DOI":"10.1109\/CVPR.2010.5539970"},{"key":"6919_CR27","doi-asserted-by":"crossref","unstructured":"Wortsman M, Ilharco G, Kim JW, Li M, Kornblith S, Roelofs R, Lopes RG, Hajishirzi H, Farhadi A, Namkoong H et al (2022) Robust fine-tuning of zero-shot models. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 7959\u20137971","DOI":"10.1109\/CVPR52688.2022.00780"},{"key":"6919_CR28","doi-asserted-by":"crossref","unstructured":"Anderson P, He X, Buehler C, Teney D, Johnson M, Gould S, Zhang L (2018) Bottom-up and top-down attention for image captioning and visual question answering. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 6077\u20136086","DOI":"10.1109\/CVPR.2018.00636"},{"key":"6919_CR29","unstructured":"Dong L, Yang N, Wang W, Wei F, Liu X, Wang Y, Gao J, Zhou M, Hon H-W (2019) Unified language model pre-training for natural language understanding and generation. Adv Neural Inf Process Syst 32"},{"key":"6919_CR30","doi-asserted-by":"crossref","unstructured":"Yu Z, Yu J, Cui Y, Tao D, Tian Q (2019) Deep modular co-attention networks for visual question answering. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 6281\u20136290","DOI":"10.1109\/CVPR.2019.00644"},{"key":"6919_CR31","unstructured":"Devlin J, Chang M-W, Lee K, Toutanova K (2018) Bert: pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805"},{"key":"6919_CR32","unstructured":"Jia C, Yang Y, Xia Y, Chen Y-T, Parekh Z, Pham H, Le Q, Sung Y-H, Li Z, Duerig T (2021) Scaling up visual and vision-language representation learning with noisy text supervision. In: International conference on machine learning. PMLR, pp 4904\u20134916"},{"issue":"9","key":"6919_CR33","doi-asserted-by":"publisher","first-page":"2337","DOI":"10.1007\/S11263-022-01653-1","volume":"130","author":"K Zhou","year":"2022","unstructured":"Zhou K, Yang J, Loy CC, Liu Z (2022) Learning to prompt for vision-language models. Int J Comput Vis 130(9):2337\u20132348. https:\/\/doi.org\/10.1007\/S11263-022-01653-1","journal-title":"Int J Comput Vis"},{"issue":"2","key":"6919_CR34","doi-asserted-by":"publisher","first-page":"581","DOI":"10.1007\/S11263-023-01891-X","volume":"132","author":"P Gao","year":"2024","unstructured":"Gao P, Geng S, Zhang R, Ma T, Fang R, Zhang Y, Li H, Qiao Y (2024) Clip-adapter: better vision-language models with feature adapters. Int J Comput Vis 132(2):581\u2013595. https:\/\/doi.org\/10.1007\/S11263-023-01891-X","journal-title":"Int J Comput Vis"},{"key":"6919_CR35","doi-asserted-by":"crossref","unstructured":"Li XL, Liang P (2021) Prefix-tuning: optimizing continuous prompts for generation. arXiv:2101.00190","DOI":"10.18653\/v1\/2021.acl-long.353"},{"key":"6919_CR36","doi-asserted-by":"crossref","unstructured":"Gu Y, Han X, Liu Z, Huang M (2021) Ppt: pre-trained prompt tuning for few-shot learning. arXiv:2109.04332","DOI":"10.18653\/v1\/2022.acl-long.576"},{"key":"6919_CR37","doi-asserted-by":"crossref","unstructured":"Lester B, Al-Rfou R, Constant N (2021) The power of scale for parameter-efficient prompt tuning. arXiv:2104.08691","DOI":"10.18653\/v1\/2021.emnlp-main.243"},{"key":"6919_CR38","unstructured":"Li J, Li D, Xiong C, Hoi S (2022) Blip: bootstrapping language-image pre-training for unified vision-language understanding and generation. In: International conference on machine learning. PMLR, pp 12888\u201312900"},{"key":"6919_CR39","unstructured":"Li J, Li D, Savarese S, Hoi S (2023) Blip-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In: International conference on machine learning. PMLR, pp 19730\u201319742"},{"key":"6919_CR40","doi-asserted-by":"publisher","first-page":"423","DOI":"10.1162\/tacl_a_00324","volume":"8","author":"Z Jiang","year":"2020","unstructured":"Jiang Z, Xu FF, Araki J, Neubig G (2020) How can we know what language models know? Trans Assoc Comput Linguist 8:423\u2013438","journal-title":"Trans Assoc Comput Linguist"},{"key":"6919_CR41","doi-asserted-by":"crossref","unstructured":"Shin T, Razeghi Y, Logan\u00a0IV RL, Wallace E, Singh S (2020) Autoprompt: eliciting knowledge from language models with automatically generated prompts. arXiv:2010.15980","DOI":"10.18653\/v1\/2020.emnlp-main.346"},{"key":"6919_CR42","unstructured":"Gao T, Fisch A, Chen D (2020) Making pre-trained language models better few-shot learners. arXiv:2012.15723"},{"key":"6919_CR43","unstructured":"Tschannen M, Gritsenko A, Wang X, Naeem MF, Alabdulmohsin I, Parthasarathy N, Evans T, Beyer L, Xia Y, Mustafa B, H\u00e9naff O, Harmsen J, Steiner A, Zhai X (2025) SigLIP 2: multilingual vision-language encoders with improved semantic understanding, localization, and dense features. https:\/\/arxiv.org\/abs\/2502.14786"},{"key":"6919_CR44","unstructured":"Sun Q, Fang Y, Wu L, Wang X, Cao Y (2023) EVA-CLIP: improved training techniques for CLIP at scale. https:\/\/arxiv.org\/abs\/2303.15389"},{"key":"6919_CR45","doi-asserted-by":"crossref","unstructured":"Zhang R, Wei Z, Fang R, Gao P, Li K, Dai J, Qiao Y, Li H (2022) Tip-adapter: training-free adaption of CLIP for few-shot classification. https:\/\/arxiv.org\/abs\/2207.09519","DOI":"10.1007\/978-3-031-19833-5_29"},{"key":"6919_CR46","doi-asserted-by":"crossref","unstructured":"Khattak MU, Wasim ST, Naseer M, Khan S, Yang M-H, Khan FS (2023) Self-regulating prompts: foundational model adaptation without forgetting. https:\/\/arxiv.org\/abs\/2307.06948","DOI":"10.1109\/ICCV51070.2023.01394"},{"key":"6919_CR47","doi-asserted-by":"publisher","DOI":"10.1016\/j.jii.2024.100709","volume":"42","author":"H Ali","year":"2024","unstructured":"Ali H, Safdar R, Rasool MH, Anjum H, Zhou Y, Yao Y, Yao L, Gao F (2024) Advance industrial monitoring of physio-chemical processes using novel integrated machine learning approach. J Ind Inf Integ 42:100709. https:\/\/doi.org\/10.1016\/j.jii.2024.100709","journal-title":"J Ind Inf Integ"},{"key":"6919_CR48","doi-asserted-by":"publisher","first-page":"106361","DOI":"10.1016\/j.conengprac.2025.106361","volume":"162","author":"H Ali","year":"2025","unstructured":"Ali H, Safdar R, Ding W, Zhou Y, Yao Y, Yao L, Gao F (2025) Intelligent machine learning-based multi-model fusion monitoring: application to industrial physio-chemical systems. Control Eng Pract 162:106361. https:\/\/doi.org\/10.1016\/j.conengprac.2025.106361","journal-title":"Control Eng Pract"},{"issue":"1","key":"6919_CR49","doi-asserted-by":"publisher","first-page":"015005","DOI":"10.1088\/2632-2153\/ada088","volume":"6","author":"H Ali","year":"2025","unstructured":"Ali H, Safdar R, Zhou Y, Yao Y, Yao L, Zhang Z, Ding W, Gao F (2025) A novel dynamic machine learning-based explainable fusion monitoring: application to industrial and chemical processes. Mach Learn Sci Technol 6(1):015005. https:\/\/doi.org\/10.1088\/2632-2153\/ada088","journal-title":"Mach Learn Sci Technol"},{"key":"6919_CR50","doi-asserted-by":"publisher","first-page":"100156","DOI":"10.1016\/j.dche.2024.100156","volume":"11","author":"H Ali","year":"2024","unstructured":"Ali H, Zhang Z, Safdar R, Rasool MH, Yao Y, Yao L, Gao F (2024) Fault detection using machine learning based dynamic ica-distributed cca: application to industrial chemical process. Digital Chem Eng 11:100156. https:\/\/doi.org\/10.1016\/j.dche.2024.100156","journal-title":"Digital Chem Eng"}],"container-title":["Applied Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-025-06919-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10489-025-06919-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-025-06919-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,7]],"date-time":"2025-11-07T15:43:26Z","timestamp":1762530206000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10489-025-06919-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10]]},"references-count":50,"journal-issue":{"issue":"15","published-print":{"date-parts":[[2025,10]]}},"alternative-id":["6919"],"URL":"https:\/\/doi.org\/10.1007\/s10489-025-06919-y","relation":{},"ISSN":["0924-669X","1573-7497"],"issn-type":[{"type":"print","value":"0924-669X"},{"type":"electronic","value":"1573-7497"}],"subject":[],"published":{"date-parts":[[2025,10]]},"assertion":[{"value":"28 March 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 September 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 October 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"997"}}