{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,16]],"date-time":"2025-12-16T12:00:22Z","timestamp":1765886422611,"version":"3.32.0"},"reference-count":50,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,11,30]],"date-time":"2024-11-30T00:00:00Z","timestamp":1732924800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2024,11,30]],"date-time":"2024-11-30T00:00:00Z","timestamp":1732924800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"funder":[{"name":"the Collaborative Innovation Key Projects of Zhengzhou","award":["123-32211645"],"award-info":[{"award-number":["123-32211645"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Appl Intell"],"published-print":{"date-parts":[[2025,1]]},"DOI":"10.1007\/s10489-024-06027-3","type":"journal-article","created":{"date-parts":[[2024,11,30]],"date-time":"2024-11-30T08:36:37Z","timestamp":1732955797000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["EDIR: an expert method for describing image regions based on knowledge distillation and triple fusion"],"prefix":"10.1007","volume":"55","author":[{"given":"Kai","family":"Ren","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-7769-8005","authenticated-orcid":false,"given":"Chuanping","family":"Hu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hao","family":"Xi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongqiang","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jinhao","family":"Fan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lihua","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,11,30]]},"reference":[{"key":"6027_CR1","first-page":"23716","volume":"35","author":"J-B Alayrac","year":"2022","unstructured":"Alayrac J-B, Donahue J, Luc P, Miech A, Barr I, Hasson Y, Lenc K, Mensch A, Millican K, Reynolds M et al (2022) Flamingo: a visual language model for few-shot learning. Adv Neural Inf Process Syst 35:23716\u201323736","journal-title":"Adv Neural Inf Process Syst"},{"key":"6027_CR2","unstructured":"Bai J, Bai S, Yang S, Wang S, Tan S, Wang P, Lin J, Zhou C, Zhou J (2023) Qwen-vl: a frontier large vision-language model with versatile abilities. arXiv:2308.12966"},{"key":"6027_CR3","unstructured":"Chen K, Zhang Z, Zeng W, Zhang R, Zhu F, Zhao R (2023) Unleashing multimodal llm\u2019s referential dialogue magic. Shikra"},{"key":"6027_CR4","doi-asserted-by":"crossref","unstructured":"Chen S, Zhu H, Chen X, Lei Y, Yu G, Chen T (2023) End-to-end 3d dense captioning with vote2cap-detr. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 11124\u201311133","DOI":"10.1109\/CVPR52729.2023.01070"},{"key":"6027_CR5","unstructured":"Chen T, Kornblith S, Norouzi M, Hinton G (2020) A simple framework for contrastive learning of visual representations. In: International conference on machine learning. PMLR, pp1597\u20131607"},{"key":"6027_CR6","doi-asserted-by":"crossref","unstructured":"Chen X, Xie S, He K (2021) An empirical study of training self-supervised vision transformers in 2021 IEEE. In: CVF International conference on computer vision (ICCV), pp 9620\u20139629","DOI":"10.1109\/ICCV48922.2021.00950"},{"key":"6027_CR7","doi-asserted-by":"crossref","unstructured":"Chen X, Djolonga J, Padlewski P, Mustafa B, Changpinyo S, Wu J, Ruiz CR, Goodman S, Wang X, Tay Y et\u00a0al (2023) Pali-x: on scaling up a multilingual vision and language model. arXiv:2305.18565","DOI":"10.1109\/CVPR52733.2024.01368"},{"key":"6027_CR8","unstructured":"Chen X, Wang X, Changpinyo S, Piergiovanni AJ, Padlewski P, Salz D, Goodman S, Grycner A, Mustafa B, Beyer L et\u00a0al (2022) Pali: a jointly-scaled multilingual language-image model. arXiv:2209.06794"},{"key":"6027_CR9","doi-asserted-by":"crossref","unstructured":"Chen X, Zhao Z, Zhang Y, Duan M, Qi D, Zhao H (2022) Focalclick: towards practical interactive image segmentation. In: 2022 IEEE\/CVF Conference on computer vision and pattern recognition (CVPR), pp 1290\u20131299","DOI":"10.1109\/CVPR52688.2022.00136"},{"issue":"5","key":"6027_CR10","doi-asserted-by":"publisher","first-page":"1701","DOI":"10.1007\/s11263-023-01949-w","volume":"132","author":"M Cornia","year":"2024","unstructured":"Cornia M, Baraldi L, Fiameni G, Cucchiara R (2024) Generating more pertinent captions by leveraging semantics and style on multi-source datasets. Int J Comput Vis 132(5):1701\u20131720","journal-title":"Int J Comput Vis"},{"key":"6027_CR11","unstructured":"Dai W, Li J, Li D, Tiong A, Zhao J, Wang W, Li B, Fung PN, Hoi S (2023) Instructblip: towards general-purpose vision-language models with instruction tuning. Adv Neural Inf Process Syst 36"},{"key":"6027_CR12","doi-asserted-by":"crossref","unstructured":"Fang Y, Wang W, Xie B, Sun Q, Wu L, Wang X, Huang T, Wang X, Cao Y (2023) Eva: exploring the limits of masked visual representation learning at scale. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 19358\u201319369","DOI":"10.1109\/CVPR52729.2023.01855"},{"key":"6027_CR13","doi-asserted-by":"crossref","unstructured":"Fang Z, Wang J, Hu X, Liang L, Gan Z, Wang L, Yang Y, Liu Z (2022) Injecting semantic concepts into end-to-end image captioning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 18009\u201318019","DOI":"10.1109\/CVPR52688.2022.01748"},{"issue":"3","key":"6027_CR14","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3617592","volume":"56","author":"T Ghandi","year":"2023","unstructured":"Ghandi T, Pourreza H, Mahyar H (2023) Deep learning approaches on image captioning: a review. ACM Comput Surv 56(3):1\u201339","journal-title":"ACM Comput Surv"},{"key":"6027_CR15","doi-asserted-by":"crossref","unstructured":"He K, Chen X, Xie S, Li Y, Doll\u00e1r P, Girshick R (2022) Masked autoencoders are scalable vision learners. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 16000\u201316009","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"6027_CR16","doi-asserted-by":"crossref","unstructured":"Hu X, Gan Z, Wang J, Yang Z, Liu Z, Lu Y, Wang L (2022) Scaling up vision-language pre-training for image captioning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 17980\u201317989","DOI":"10.1109\/CVPR52688.2022.01745"},{"key":"6027_CR17","doi-asserted-by":"crossref","unstructured":"Huang X, Wang J, Tang Y, Zhang Z, Hu H, Lu J, Wang L, Liu Z (2024) Segment and caption anything. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 13405\u201313417","DOI":"10.1109\/CVPR52733.2024.01273"},{"key":"6027_CR18","doi-asserted-by":"crossref","unstructured":"Jain J, Li J, Chiu MT, Hassani A, Orlov N, Shi H (2023) Oneformer: one transformer to rule universal image segmentation. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 2989\u20132998","DOI":"10.1109\/CVPR52729.2023.00292"},{"key":"6027_CR19","unstructured":"Jia C, Yang Y, Xia Y, Chen Y-T, Parekh Z, Pham H, Le Q, Sung Y-H, Li Z, Duerig T (2021) Scaling up visual and vision-language representation learning with noisy text supervision. In: International conference on machine learning. PMLR, pp 4904\u20134916"},{"key":"6027_CR20","doi-asserted-by":"crossref","unstructured":"Jiang H, Misra I, Rohrbach M, Learned-Miller E, Chen X (2020) In defense of grid features for visual question answering. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 10267\u201310276","DOI":"10.1109\/CVPR42600.2020.01028"},{"key":"6027_CR21","doi-asserted-by":"crossref","unstructured":"Kirillov A, Mintun E, Ravi N, Mao H, Rolland C, Gustafson L, Xiao T, Whitehead S, Berg AC, Lo W-Y, Dollar P, Girshick R (2023) Segment anything. In: Proceedings of the IEEE\/CVF international conference on computer vision (ICCV), pp 4015\u20134026","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"6027_CR22","unstructured":"Li J, Li D, Savarese S, Hoi S (2023) Blip-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In: International conference on machine learning. PMLR, pp 19730\u201319742"},{"key":"6027_CR23","unstructured":"Li J, Li D, Xiong C, Hoi S (2022) Blip: bootstrapping language-image pre-training for unified vision-language understanding and generation. In: International conference on machine learning. PMLR, pp 12888\u201312900"},{"key":"6027_CR24","doi-asserted-by":"crossref","unstructured":"Li Y, Fan H, Hu R, Feichtenhofer C, He K (2023) Scaling language-image pre-training via masking. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 23390\u20132340","DOI":"10.1109\/CVPR52729.2023.02240"},{"key":"6027_CR25","doi-asserted-by":"crossref","unstructured":"Li Y, Du Y, Zhou K, Wang J, Zhao X, Wen J-R (2023) Evaluating object hallucination in large vision-language models. In: Bouamor H, Pino J, Bali K (eds) Proceedings of the 2023 conference on empirical methods in natural language processing. Singapore, pp 292\u2013305. Association for Computational Linguistics","DOI":"10.18653\/v1\/2023.emnlp-main.20"},{"key":"6027_CR26","unstructured":"Liu H, Li C, Wu Q, Lee YJ (2024) Visual instruction tuning. Adv Neural Inf Process Syst 36"},{"key":"6027_CR27","doi-asserted-by":"crossref","unstructured":"Liu Q, Xu Z, Bertasius G, Niethammer M (2023) Simpleclick: interactive image segmentation with simple vision transformers. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp 22290\u201322300","DOI":"10.1109\/ICCV51070.2023.02037"},{"key":"6027_CR28","doi-asserted-by":"crossref","unstructured":"Long Y, Wen Y, Han J, Xu H, Ren P, Zhang W, Zhao S, Liang X (2023) Capdet: unifying dense captioning and open-world detection pretraining. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 15233\u201315243","DOI":"10.1109\/CVPR52729.2023.01462"},{"issue":"7","key":"6027_CR29","first-page":"3523","volume":"44","author":"S Minaee","year":"2021","unstructured":"Minaee S, Boykov Y, Porikli F, Plaza A, Kehtarnavaz N, Terzopoulos D (2021) Image segmentation using deep learning: a survey. IEEE Trans Pattern Anal Mach Intell 44(7):3523\u20133542","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"6027_CR30","unstructured":"Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J et al (2021) Learning transferable visual models from natural language supervision. In: International conference on machine learning. PMLR, pp 8748\u20138763"},{"key":"6027_CR31","unstructured":"Ramesh A, Pavlov M, Goh G, Gray S, Voss C, Radford A, Chen M, Sutskever I (2021)Zero-shot text-to-image generation.In: International conference on machine learning. PMLR, pp 8821\u20138831"},{"key":"6027_CR32","doi-asserted-by":"crossref","unstructured":"Ren K, Hu C, Xi H (2024) Rlm-tracking: online multi-pedestrian tracking supported by relative location mapping. Int J Mach Learn Cybern 1\u201317","DOI":"10.1007\/s13042-023-02070-7"},{"key":"6027_CR33","doi-asserted-by":"crossref","unstructured":"Ren S, Wei F, Zhang Z, Hu H (2023) Tinymim: an empirical study of distilling mim pre-trained models. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR), pp 3687\u20133697","DOI":"10.1109\/CVPR52729.2023.00359"},{"key":"6027_CR34","unstructured":"Ridnik T, Ben-Baruch E, Noy A, Zelnik L (2021) Imagenet-21k pretraining for the masses. In: Vanschoren J, Yeung S (eds) Proceedings of the neural information processing systems track on datasets and benchmarks,vol 1"},{"key":"6027_CR35","doi-asserted-by":"crossref","unstructured":"Shao Z, Han J, Debattista K, Pang Y (2023) Textual context-aware dense captioning with diverse words. IEEE Trans Multimedia","DOI":"10.1109\/TMM.2023.3241517"},{"key":"6027_CR36","unstructured":"Sun Q, Fang Y, Wu L, Wang X, Cao Y (2023) Eva-clip: Improved training techniques for clip at scale. arXiv:2303.15389"},{"key":"6027_CR37","doi-asserted-by":"crossref","unstructured":"Sun Z, Fang Y, Wu T, Zhang P, Zang Y, Kong S, Xiong Y, Lin D, Wang J (2024) Alpha-clip: a clip model focusing on wherever you want. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 13019\u201313029","DOI":"10.1109\/CVPR52733.2024.01237"},{"key":"6027_CR38","unstructured":"Wang J, Yang Z, Hu X, Li X, Lin K, Gan Z, Liu Z, Liu C, Wang L (2022) Git: a generative image-to-text transformer for vision and language"},{"key":"6027_CR39","unstructured":"Wang T, Zhang J, Fei J, Ge Y, Zheng H, Tang Y, Li Z, Gao M, Zhao S, Shan Y et al (2023) Caption anything: interactive image description with diverse multimodal controls. arXiv:2305.02677"},{"key":"6027_CR40","unstructured":"Wang W, Lv Q, Yu W, Hong W, Qi J, Wang Y, Ji J, Yang Z, Zhao L, Song X et\u00a0al (2023) Cogvlm: visual expert for pretrained language models. arXiv:2311.03079"},{"key":"6027_CR41","unstructured":"Wang Z, Yu J, Yu AW, Dai Z, Tsvetkov Y, Cao Y (2021) Simvlm: simple visual language model pretraining with weak supervision. arXiv:2108.10904"},{"key":"6027_CR42","doi-asserted-by":"crossref","unstructured":"Wu K, Peng H, Zhou Z, Xiao B, Liu M, Yuan L, Xuan H, Valenzuela M, Chen XS, Wang X, Chao H (2023) Tinyclip: clip distillation via affinity mimicking and weight inheritance","DOI":"10.1109\/ICCV51070.2023.02008"},{"key":"6027_CR43","doi-asserted-by":"crossref","unstructured":"Wu K, Zhang J, Peng H, Liu M, Xiao B, Fu J, Yuan L (2022) Tinyvit: fast pretraining distillation for small vision transformers. In: European conference on computer vision. Springer, pp 68\u201385","DOI":"10.1007\/978-3-031-19803-8_5"},{"key":"6027_CR44","unstructured":"Yu J, Wang Z, Vasudevan V, Yeung L, Seyedhosseini M, Wu Y (2022) Coca: contrastive captioners are image-text foundation models. arXiv:2205.01917"},{"key":"6027_CR45","unstructured":"Zhang A, Yao Y, Ji W, Liu Z, Chua T-S (2023) Next-chat: an lmm for chat, detection and segmentation"},{"key":"6027_CR46","unstructured":"Zhang C, Han D, Qiao Y, Kim JU, Bae S-H, Lee S, Hong CS (2023) Faster segment anything: towards lightweight sam for mobile applications"},{"issue":"2","key":"6027_CR47","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3478642","volume":"18","author":"F Zhang","year":"2022","unstructured":"Zhang F, Xu M, Xu C (2022) Tell, imagine, and search: end-to-end learning for composing text and image to image retrieval. ACM Trans Multimed Comput Commun Appl (TOMM) 18(2):1\u201323","journal-title":"ACM Trans Multimed Comput Commun Appl (TOMM)"},{"key":"6027_CR48","doi-asserted-by":"crossref","unstructured":"Zhou Y, Zhang R, Chen C, Li C, Tensmeyer C, Yu T, Gu J, Xu J, Sun T (2022) Towards language-free training for text-to-image generation. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 17907\u201317917","DOI":"10.1109\/CVPR52688.2022.01738"},{"key":"6027_CR49","unstructured":"Zhu D, Chen J, Shen X, Li X, Elhoseiny M (2023) Minigpt-4: enhancing vision-language understanding with advanced large language models. arXiv:2304.10592"},{"key":"6027_CR50","unstructured":"Zou X, Yang J, Zhang H, Li F, Li L, Wang J, Wang L, Gao J, Lee YJ (2024) Segment everything everywhere all at once. Adv Neural Inf Process Syst 36"}],"container-title":["Applied Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-024-06027-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10489-024-06027-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-024-06027-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,2]],"date-time":"2025-01-02T06:24:56Z","timestamp":1735799096000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10489-024-06027-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,30]]},"references-count":50,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1]]}},"alternative-id":["6027"],"URL":"https:\/\/doi.org\/10.1007\/s10489-024-06027-3","relation":{},"ISSN":["0924-669X","1573-7497"],"issn-type":[{"type":"print","value":"0924-669X"},{"type":"electronic","value":"1573-7497"}],"subject":[],"published":{"date-parts":[[2024,11,30]]},"assertion":[{"value":"18 September 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 November 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant financial interests in the manuscript and no other potential conflicts of interest to disclose.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Informed consent"}}],"article-number":"62"}}