{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,21]],"date-time":"2026-06-21T05:44:42Z","timestamp":1782020682029,"version":"3.54.5"},"reference-count":63,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,6,28]],"date-time":"2024-06-28T00:00:00Z","timestamp":1719532800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,6,28]],"date-time":"2024-06-28T00:00:00Z","timestamp":1719532800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62176169"],"award-info":[{"award-number":["62176169"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["070-63243150"],"award-info":[{"award-number":["070-63243150"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Intell."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The advent of large vision-language models (LVLMs) represents a remarkable advance in the quest for artificial general intelligence. However, the models\u2019 effectiveness in both specialized and general tasks warrants further investigation. This paper endeavors to evaluate the competency of popular LVLMs in specialized and general tasks, respectively, aiming to offer a comprehensive understanding of these novel models. To gauge their effectiveness in specialized tasks, we employ six challenging tasks in three different application scenarios: natural, healthcare, and industrial. These six tasks include salient\/camouflaged\/transparent object detection, as well as polyp detection, skin lesion detection, and industrial anomaly detection. We examine the performance of three recent open-source LVLMs, including MiniGPT-v2, LLaVA-1.5, and Shikra, on both visual recognition and localization in these tasks. Moreover, we conduct empirical investigations utilizing the aforementioned LVLMs together with GPT-4V, assessing their multi-modal understanding capabilities in general tasks including object counting, absurd question answering, affordance reasoning, attribute recognition, and spatial relation reasoning. Our investigations reveal that these LVLMs demonstrate limited proficiency not only in specialized tasks but also in general tasks. We delve deep into this inadequacy and uncover several potential factors, including limited cognition in specialized tasks, object hallucination, text-to-image interference, and decreased robustness in complex problems. We hope that this study can provide useful insights for the future development of LVLMs, helping researchers improve LVLMs for both general and specialized applications.<\/jats:p>","DOI":"10.1007\/s44267-024-00050-1","type":"journal-article","created":{"date-parts":[[2024,6,28]],"date-time":"2024-06-28T10:03:34Z","timestamp":1719569014000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":37,"title":["Effectiveness assessment of recent large vision-language models"],"prefix":"10.1007","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-0812-1036","authenticated-orcid":false,"given":"Yao","family":"Jiang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5401-3304","authenticated-orcid":false,"given":"Xinyu","family":"Yan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7092-2877","authenticated-orcid":false,"given":"Ge-Peng","family":"Ji","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3195-2077","authenticated-orcid":false,"given":"Keren","family":"Fu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8691-8677","authenticated-orcid":false,"given":"Meijun","family":"Sun","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5401-7834","authenticated-orcid":false,"given":"Huan","family":"Xiong","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5245-7518","authenticated-orcid":false,"given":"Deng-Ping","family":"Fan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4263-3143","authenticated-orcid":false,"given":"Fahad Shahbaz","family":"Khan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,6,28]]},"reference":[{"key":"50_CR1","first-page":"1877","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"T. Brown","year":"2020","unstructured":"Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., et al. (2020). Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, et al. (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 1877\u20131901). Red Hook: Curran Associates."},{"key":"50_CR2","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., et\u00a0al. (2023). LLaMA: open and efficient foundation language models. arXiv preprint. arXiv:2302.13971."},{"key":"50_CR3","first-page":"1","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"H. Liu","year":"2023","unstructured":"Liu, H., Li, C., Wu, Q., & Lee, Y. J. (2023). Visual instruction tuning. In A. Oh, T. Naumann, A. Globerson, et al. (Eds.), Proceedings of the 37th international conference on neural information processing systems (pp. 1\u201325). Red Hook: Curran Associates."},{"key":"50_CR4","unstructured":"Chen, J., Zhu, D., Shen, X., Li, X., Liu, Z., Zhang, P., et\u00a0al. (2023). Minigpt-v2: large language model as a unified interface for vision-language multi-task learning. arXiv preprint. arXiv:2310.09478."},{"key":"50_CR5","unstructured":"Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., et\u00a0al. (2023). Gpt-4 technical report. arXiv preprint. arXiv:2303.08774."},{"key":"50_CR6","unstructured":"Liu, H., Li, C., Li, Y., & Lee, Y. J. (2023). Improved baselines with visual instruction tuning. arXiv preprint. arXiv:2310.03744."},{"key":"50_CR7","unstructured":"Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., & Zhao, R. (2023). Shikra: unleashing multimodal LLM\u2019s referential dialogue magic. arXiv preprint. arXiv:2306.15195."},{"key":"50_CR8","unstructured":"Fu, C., Zhang, R., Lin, H., Wang, Z., Gao, T., Luo, Y., et\u00a0al. (2023). A challenger to GPT-4v? early explorations of gemini in visual expertise. arXiv preprint. arXiv:2312.12436."},{"issue":"5","key":"50_CR9","doi-asserted-by":"publisher","first-page":"605","DOI":"10.1007\/s11633-023-1469-x","volume":"20","author":"H. Qin","year":"2023","unstructured":"Qin, H., Ji, G.-P., Khan, S., Fan, D.-P., Khan, F. S., & Gool, L. V. (2023). How good is Google bard\u2019s visual understanding? An empirical study on open challenges. Machine Intelligence Research, 20(5), 605\u2013613.","journal-title":"Machine Intelligence Research"},{"key":"50_CR10","unstructured":"Xie, L., Wei, L., Zhang, X., Bi, K., Gu, X., Chang, J., et\u00a0al. (2023). Towards AGI in computer vision: lessons learned from GPT and large language models. arXiv preprint. arXiv:2306.08641."},{"key":"50_CR11","unstructured":"Zhang, J., Chen, X., Xue, Z., Wang, Y., Wang, C., & Liu, Y. (2023). Exploring grounding potential of VQA-oriented GPT-4v for zero-shot anomaly detection. arXiv preprint. arXiv:2311.02612."},{"key":"50_CR12","unstructured":"Tang, L., Jiang, P.-T., Shen, Z., Zhang, H., Chen, J., & Li, B. (2023). Generalization and hallucination of large vision-language models through a camouflaged lens. arXiv preprint. arXiv:2311.11273."},{"issue":"12","key":"50_CR13","doi-asserted-by":"publisher","first-page":"6074","DOI":"10.1109\/JBHI.2023.3316750","volume":"27","author":"J. Qiu","year":"2023","unstructured":"Qiu, J., Li, L., Sun, J., Peng, J., Shi, P., Zhang, R., et al. (2023). Large AI models in health informatics: applications, challenges, and the future. IEEE Journal of Biomedical and Health Informatics, 27(12), 6074\u20136087.","journal-title":"IEEE Journal of Biomedical and Health Informatics"},{"key":"50_CR14","first-page":"1932","volume-title":"Proceedings of the 38th AAAI conference on artificial intelligence","author":"Z. Gu","year":"2024","unstructured":"Gu, Z., Zhu, B., Zhu, G., Chen, Y., Tang, M., & Wang, J. (2024). AnomalyGPT: detecting industrial anomalies using large vision-language models. In M. J. Wooldridge, J. G. Dy, & S. Natarajan (Eds.), Proceedings of the 38th AAAI conference on artificial intelligence (pp. 1932\u20131940). Palo Alto: AAAI Press."},{"key":"50_CR15","unstructured":"Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., et\u00a0al. (2023). MME: a comprehensive evaluation benchmark for multimodal large language models. arXiv preprint. arXiv:2306.13394."},{"key":"50_CR16","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., et\u00a0al. (2014). Microsoft coco: common objects in context. arXiv preprint. arXiv:1405.0312.","DOI":"10.1007\/978-3-319-10602-1_48"},{"issue":"11","key":"50_CR17","first-page":"13083","volume":"45","author":"R. Song","year":"2023","unstructured":"Song, R., Zhang, W., Zhao, Y., Liu, Y., & Rosin, P. L. (2023). 3D visual saliency: an independent perceptual measure or a derivative of 2D image saliency? IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(11), 13083\u201313099.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"9","key":"50_CR18","first-page":"5541","volume":"44","author":"K. Fu","year":"2021","unstructured":"Fu, K., Fan, D.-P., Ji, G.-P., Zhao, Q., Shen, J., & Zhu, C. (2021). Siamese network for RGB-D salient object detection and beyond. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(9), 5541\u20135559.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"50_CR19","first-page":"3052","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"K. Fu","year":"2020","unstructured":"Fu, K., Fan, D.-P., Ji, G.-P., & Zhao, Q. (2020). JL-DCF: joint learning and densely-cooperative fusion framework for RGB-D salient object detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 3052\u20133062). Piscataway: IEEE."},{"key":"50_CR20","first-page":"696","volume-title":"Proceedings of the 16th European conference on computer vision","author":"E. Xie","year":"2020","unstructured":"Xie, E., Wang, W., Wang, W., Ding, M., Shen, C., & Luo, P. (2020). Segmenting transparent objects in the wild. In A. Vedaldi, H. Bischof, T. Brox, et al. (Eds.), Proceedings of the 16th European conference on computer vision (pp. 696\u2013711). Cham: Springer."},{"issue":"10","key":"50_CR21","doi-asserted-by":"publisher","first-page":"6024","DOI":"10.1109\/TPAMI.2021.3085766","volume":"44","author":"D.-P. Fan","year":"2021","unstructured":"Fan, D.-P., Ji, G.-P., Cheng, M.-M., & Shao, L. (2021). Concealed object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10), 6024\u20136042.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"50_CR22","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.108414","volume":"123","author":"G.-P. Ji","year":"2022","unstructured":"Ji, G.-P., Zhu, L., Zhuge, M., & Fu, K. (2022). Fast camouflaged object detection via edge-based reversible re-calibration network. Pattern Recognition, 123, 108414.","journal-title":"Pattern Recognition"},{"key":"50_CR23","first-page":"168","volume-title":"Proceedings of the IEEE international symposium on biomedical imaging","author":"N. C. Codella","year":"2018","unstructured":"Codella, N. C., Gutman, D., Celebi, M. E., Helba, B., Marchetti, M. A., Dusza, S. W., et al. (2018). Skin lesion analysis toward melanoma detection: a challenge at the 2017 international symposium on biomedical imaging (isbi), hosted by the international skin imaging collaboration (isic). In Proceedings of the IEEE international symposium on biomedical imaging (pp. 168\u2013172). Piscataway: IEEE."},{"issue":"2","key":"50_CR24","doi-asserted-by":"publisher","first-page":"630","DOI":"10.1109\/TMI.2015.2487997","volume":"35","author":"N. Tajbakhsh","year":"2015","unstructured":"Tajbakhsh, N., Gurudu, S. R., & Liang, J. (2015). Automated polyp detection in colonoscopy videos using shape and context information. IEEE Transactions on Medical Imaging, 35(2), 630\u2013644.","journal-title":"IEEE Transactions on Medical Imaging"},{"issue":"4","key":"50_CR25","doi-asserted-by":"publisher","first-page":"1038","DOI":"10.1007\/s11263-020-01400-4","volume":"129","author":"P. Bergmann","year":"2021","unstructured":"Bergmann, P., Batzner, K., Fauser, M., Sattlegger, D., & Steger, C. (2021). The MVTec anomaly detection dataset: a comprehensive real-world dataset for unsupervised anomaly detection. International Journal of Computer Vision, 129(4), 1038\u20131059.","journal-title":"International Journal of Computer Vision"},{"key":"50_CR26","first-page":"30662","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"A. Conti","year":"2023","unstructured":"Conti, A., Fini, E., Mancini, M., Rota, P., Wang, Y., & Ricci, E. (2023). Vocabulary-free image classification. In A. Oh, T. Naumann, A. Globerson, et al. (Eds.), Proceedings of the 37th international conference on neural information processing systems (pp. 30662\u201330680). Red Hook: Curran Associates."},{"key":"50_CR27","first-page":"136","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"L. Wang","year":"2017","unstructured":"Wang, L., Lu, H., Wang, Y., Feng, M., Wang, D., Yin, B., et al. (2017). Learning to detect salient objects with image-level supervision. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 136\u2013145). Piscataway: IEEE."},{"key":"50_CR28","first-page":"186","volume-title":"Proceedings of the 15th European conference on computer vision","author":"D.-P. Fan","year":"2018","unstructured":"Fan, D.-P., Cheng, M.-M., Liu, J.-J., Gao, S.-H., Hou, Q., & Borji, A. (2018). Salient objects in clutter: bringing salient object detection to the foreground. In V. Ferrari, M. Hebert, & C. Sminchisescu (Eds.), Proceedings of the 15th European conference on computer vision (pp. 186\u2013202). Cham: Springer."},{"issue":"2","key":"50_CR29","doi-asserted-by":"publisher","first-page":"283","DOI":"10.1007\/s11548-013-0926-3","volume":"9","author":"J. Silva","year":"2014","unstructured":"Silva, J., Histace, A., Romain, O., Dray, X., & Granado, B. (2014). Toward embedded detection of polyps in WCE images for early diagnosis of colorectal cancer. International Journal of Computer Assisted Radiology and Surgery, 9(2), 283\u2013293.","journal-title":"International Journal of Computer Assisted Radiology and Surgery"},{"key":"50_CR30","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s12880-020-00482-3","volume":"20","author":"W. Wang","year":"2020","unstructured":"Wang, W., Tian, J., Zhang, C., Luo, Y., Wang, X., & Li, J. (2020). An improved deep learning approach and its applications on colonic polyp images detection. BMC Medical Imaging, 20, 1\u201314.","journal-title":"BMC Medical Imaging"},{"key":"50_CR31","first-page":"392","volume-title":"Proceedings of the 17th European conference on computer vision","author":"Y. Zou","year":"2022","unstructured":"Zou, Y., Jeong, J., Pemula, L., Zhang, D., & Dabeer, O. (2022). Spot-the-difference self-supervised pre-training for anomaly detection and segmentation. In S. Avidan, G. Brostow, M. Ciss\u00e9, et al. (Eds.), Proceedings of the 17th European conference on computer vision (pp. 392\u2013408). Cham: Springer."},{"key":"50_CR32","doi-asserted-by":"publisher","first-page":"292","DOI":"10.18653\/v1\/2023.emnlp-main.20","volume-title":"Proceedings of the 2023 conference on empirical methods in natural language processing","author":"Y. Li","year":"2023","unstructured":"Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., & Wen, J.-R. (2023). Evaluating object hallucination in large vision-language models. In H. Bouamor, J. Pino, & K. Bali (Eds.), Proceedings of the 2023 conference on empirical methods in natural language processing (pp. 292\u2013305). Stroudsburg: ACL."},{"key":"50_CR33","unstructured":"Xu, P., Shao, W., Zhang, K., Gao, P., Liu, S., Lei, M., et\u00a0al. (2023). LVLM-eHub: a comprehensive evaluation benchmark for large vision-language models. arXiv preprint. arXiv:2306.09265."},{"key":"50_CR34","unstructured":"Cui, C., Zhou, Y., Yang, X., Wu, S., Zhang, L., Zou, J., et\u00a0al. (2023). Holistic analysis of hallucination in GPT-4V (ision): bias and interference challenges. arXiv preprint. arXiv:2311.03287."},{"key":"50_CR35","first-page":"4015","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"A. Kirillov","year":"2023","unstructured":"Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., et al. (2023). Segment anything. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 4015\u20134026). Piscataway: IEEE."},{"key":"50_CR36","first-page":"1","volume-title":"Proceedings of the 9th international conference on learning representations","author":"A. Dosovitskiy","year":"2021","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., et al. (2021). An image is worth 16x16 words: transformers for image recognition at scale. In Proceedings of the 9th international conference on learning representations (pp. 1\u201321). Retrieved June 4, 2024, from https:\/\/openreview.net\/forum?id=YicbFdNTTy."},{"key":"50_CR37","doi-asserted-by":"publisher","DOI":"10.3390\/electronics10030279","volume":"10","author":"R. Padilla","year":"2021","unstructured":"Padilla, R., Passos, W. L., Dias, T. L., Netto, S. L., & Da Silva, E. A. (2021). A comparative analysis of object detection metrics with a companion open-source toolkit. Electronics, 10, 279.","journal-title":"Electronics"},{"key":"50_CR38","first-page":"733","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"F. Perazzi","year":"2012","unstructured":"Perazzi, F., Kr\u00e4henb\u00fchl, P., Pritch, Y., & Hornung, A. (2012). Saliency filters: contrast based filtering for salient region detection. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 733\u2013740). Piscataway: IEEE."},{"key":"50_CR39","first-page":"4558","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"D.-P. Fan","year":"2017","unstructured":"Fan, D.-P., Cheng, M.-M., Liu, Y., Li, T., & Borji, A. (2017). Structure-measure: a new way to evaluate foreground maps. In Proceedings of the IEEE international conference on computer vision (pp. 4558\u20134567). Piscataway: IEEE."},{"key":"50_CR40","first-page":"1597","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"R. Achanta","year":"2009","unstructured":"Achanta, R., Hemami, S., Estrada, F., & Susstrunk, S. (2009). Frequency-tuned salient region detection. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1597\u20131604). Piscataway: IEEE."},{"key":"50_CR41","unstructured":"Gu, J., Han, Z., Chen, S., Beirami, A., He, B., Zhang, G., et\u00a0al. (2023). A systematic survey of prompt engineering on vision-language foundation models. arXiv preprint. arXiv:2307.12980."},{"key":"50_CR42","first-page":"28541","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"C. Li","year":"2023","unstructured":"Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., et al. (2023). LLaVA-Med: training a large language-and-vision assistant for biomedicine in one day. In A. Oh, T. Naumann, A. Globerson, et al. (Eds.), Proceedings of the 37th international conference on neural information processing systems (pp. 28541\u201328564). Red Hook: Curran Associates."},{"key":"50_CR43","unstructured":"Liu, X., Fu, K., & Zhao, Q. (2023). Promoting segment anything model towards highly accurate dichotomous image segmentation. arXiv preprint. arXiv:2401.00248."},{"key":"50_CR44","unstructured":"Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., et\u00a0al. (2023). Analyzing and mitigating object hallucination in large vision-language models. arXiv preprint. arXiv:2310.00754."},{"key":"50_CR45","unstructured":"Qian, Y., Zhang, H., Yang, Y., & Gan, Z. (2024). How easy is it to fool your multimodal LLMs? An empirical analysis on deceptive prompts. arXiv preprint. arXiv:2402.13220."},{"key":"50_CR46","first-page":"2584","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"J. M. Kim","year":"2023","unstructured":"Kim, J. M., Koepke, A., Schmid, C., & Akata, Z. (2023). Exposing and mitigating spurious correlations for cross-modal retrieval. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 2584\u20132594). Piscataway: IEEE."},{"key":"50_CR47","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-023-3911-2","author":"Y. Wu","year":"2023","unstructured":"Wu, Y., Zhao, Y., Li, Z., Qin, B., & Xiong, K. (2023). Improving cross-task generalization with step-by-step instructions. Science China. Information Sciences. Advance online publication. https:\/\/doi.org\/10.1007\/s11432-023-3911-2.","journal-title":"Science China. Information Sciences"},{"issue":"6","key":"50_CR48","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s11432-023-3740-x","volume":"66","author":"H. Chen","year":"2023","unstructured":"Chen, H., Yuan, K., Huang, Y., Guo, L., Wang, Y., & Chen, J. (2023). Feedback is all you need: from chatgpt to autonomous driving. Science China. Information Sciences, 66(6), 1\u20133.","journal-title":"Science China. Information Sciences"},{"key":"50_CR49","unstructured":"Yan, S., Bai, M., Chen, W., Zhou, X., Huang, Q., & Li, L. E. (2024). ViGoR: improving visual grounding of large vision language models with fine-grained reward modeling. arXiv preprint. arXiv:2402.06118."},{"key":"50_CR50","unstructured":"Jiao, Q., Chen, D., Huang, Y., Li, Y., & Shen, Y. (2024). Enhancing multimodal large language models with vision detection models: an empirical study. arXiv preprint. arXiv:2401.17981."},{"key":"50_CR51","unstructured":"Yao, Z., Wu, X., Li, C., Zhang, M., Qi, H., Ruwase, O., et\u00a0al. (2023). DeepSpeed-VisualChat: multi-round multi-image interleave chat via multi-modal causal attention. arXiv preprint. arXiv:2309.14327."},{"issue":"4","key":"50_CR52","doi-asserted-by":"publisher","first-page":"509","DOI":"10.1007\/s41095-021-0256-2","volume":"8","author":"K. Fu","year":"2022","unstructured":"Fu, K., Jiang, Y., Ji, G.-P., Zhou, T., Zhao, Q., & Fan, D.-P. (2022). Light field salient object detection: a review and benchmark. Computational Visual Media, 8(4), 509\u2013534.","journal-title":"Computational Visual Media"},{"issue":"10","key":"50_CR53","doi-asserted-by":"crossref","first-page":"2860","DOI":"10.11834\/jig.211068","volume":"27","author":"J. He","year":"2022","unstructured":"He, J., & Fu, K. (2022). RGB-D salient object detection of using few-shot learning. International Journal of Image and Graphics, 27(10), 2860\u20132872.","journal-title":"International Journal of Image and Graphics"},{"key":"50_CR54","first-page":"4681","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"T. Zhou","year":"2021","unstructured":"Zhou, T., Fu, H., Chen, G., Zhou, Y., Fan, D.-P., & Shao, L. (2021). Specificity-preserving RGB-D saliency detection. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 4681\u20134691). Piscataway: IEEE."},{"key":"50_CR55","first-page":"1063","volume-title":"Proceedings of the 35th AAAI conference on artificial intelligence","author":"Q. Chen","year":"2021","unstructured":"Chen, Q., Liu, Z., Zhang, Y., Fu, K., Zhao, Q., & Du, H. (2021). RGB-D salient object detection via 3D convolutional neural networks. In Proceedings of the 35th AAAI conference on artificial intelligence (pp. 1063\u20131071). Palo Alto: AAAI Press."},{"key":"50_CR56","doi-asserted-by":"publisher","first-page":"69","DOI":"10.1016\/j.neucom.2019.04.062","volume":"356","author":"K. Fu","year":"2019","unstructured":"Fu, K., Zhao, Q., Gu, I. Y.-H., & Yang, J. (2019). Deepside: a general deep framework for salient object detection. Neurocomputing, 356, 69\u201382.","journal-title":"Neurocomputing"},{"key":"50_CR57","doi-asserted-by":"publisher","first-page":"731","DOI":"10.1145\/3474085.3475240","volume-title":"Proceedings of the 29th ACM international conference on multimedia","author":"W. Zhang","year":"2021","unstructured":"Zhang, W., Ji, G.-P., Wang, Z., Fu, K., & Zhao, Q. (2021). Depth quality-inspired feature manipulation for efficient RGB-D salient object detection. In H. T. Shen, Y. Zhuang, J. Smith, et al. (Eds.), Proceedings of the 29th ACM international conference on multimedia (pp. 731\u2013740). New York: ACM."},{"key":"50_CR58","unstructured":"Zhong, L., Liao, X., Zhang, S., Zhang, X., & Wang, G. (2024). VLM-CPL: consensus pseudo labels from vision-language models for human annotation-free pathological image classification. arXiv preprint. arXiv:2403.15836."},{"key":"50_CR59","first-page":"8483","volume-title":"Proceedings of the 36th international conference on neural information processing systems","author":"Z. Wang","year":"2022","unstructured":"Wang, Z., Li, M., Xu, R., Zhou, L., Lei, J., Lin, X., et al. (2022). Language models with image descriptors are strong few-shot video-language learners. In Proceedings of the 36th international conference on neural information processing systems (pp. 8483\u20138497). Red Hook: Curran Associates."},{"key":"50_CR60","unstructured":"He, S., & Ding, H. (2024). Decoupling static and hierarchical motion perception for referring video segmentation. arXiv preprint. arXiv:2404.03645."},{"key":"50_CR61","first-page":"2694","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"H. Ding","year":"2023","unstructured":"Ding, H., Liu, C., He, S., Jiang, X., & Loy, C. C. (2023). MeViS: a large-scale benchmark for video segmentation with motion expressions. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 2694\u20132703). Piscataway: IEEE."},{"key":"50_CR62","first-page":"20224","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"H. Ding","year":"2023","unstructured":"Ding, H., Liu, C., He, S., Jiang, X., Torr, P. H., & Bai, S. (2023). MOSE: a new dataset for video object segmentation in complex scenes. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 20224\u201320234). Piscataway: IEEE."},{"issue":"12","key":"50_CR63","doi-asserted-by":"publisher","first-page":"3088","DOI":"10.1109\/TPAMI.2019.2920899","volume":"42","author":"W. Zhang","year":"2019","unstructured":"Zhang, W., Wang, B., Ma, L., & Liu, W. (2019). Reconstruct and represent video contents for captioning via reinforcement learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(12), 3088\u20133101.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"}],"container-title":["Visual Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-024-00050-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44267-024-00050-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-024-00050-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,22]],"date-time":"2024-11-22T23:52:02Z","timestamp":1732319522000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44267-024-00050-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,28]]},"references-count":63,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["50"],"URL":"https:\/\/doi.org\/10.1007\/s44267-024-00050-1","relation":{},"ISSN":["2731-9008"],"issn-type":[{"value":"2731-9008","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,28]]},"assertion":[{"value":"15 April 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 June 2024","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 June 2024","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 June 2024","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Deng-Ping Fan is an Associate Editor at Visual Intelligence and was not involved in the editorial review of this article or the decision to publish it. The authors declare that they have no other competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"17"}}