{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T23:35:20Z","timestamp":1784158520248,"version":"3.55.0"},"reference-count":54,"publisher":"Springer Science and Business Media LLC","issue":"32","license":[{"start":{"date-parts":[[2025,9,25]],"date-time":"2025-09-25T00:00:00Z","timestamp":1758758400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,9,25]],"date-time":"2025-09-25T00:00:00Z","timestamp":1758758400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2025,11]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>State-of-the-art (SoTA) image captioning models are often trained on the MicroSoft Common Objects in Context (MS-COCO) dataset, which contains human-annotated captions with an average length of approximately ten tokens. Although effective for general scene understanding, these short captions often fail to capture complex scenes and convey detailed information. Moreover, captioning models tend to exhibit bias toward the \u201caverage\u201d caption, which captures only the more general aspects, thus overlooking finer details. In this paper, we present a novel approach to generate richer and more informative image captions by combining the captions generated from different SoTA captioning models. Our proposed method requires no additional model training: given an image, it leverages pretrained models from the literature to generate the initial captions, and then ranks them using a newly introduced image-text-based metric, which we name BLIPScore. Subsequently, the top two captions are fused using a Large Language Model to produce the final, more detailed description. Experimental results on the MS-COCO and Flickr30k test sets demonstrate the effectiveness of our approach in terms of caption-image alignment and hallucination reduction according to the ALOHa, CAPTURE, and Polos metrics. A subjective study lends additional support to these results, suggesting that the captions produced by our model are generally perceived as more consistent with human judgment. By combining the strengths of diverse SoTA models, our method enhances the quality and appeal of image captions, bridging the gap between automated systems and the rich and informative nature of human-generated descriptions. This advance enables the generation of more suitable captions for the training of both vision-language and captioning models.<\/jats:p>","DOI":"10.1007\/s00521-025-11672-x","type":"journal-article","created":{"date-parts":[[2025,9,25]],"date-time":"2025-09-25T13:13:35Z","timestamp":1758806015000},"page":"27279-27299","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Improving image captioning descriptiveness by ranking and LLM-based fusion"],"prefix":"10.1007","volume":"37","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5925-2646","authenticated-orcid":false,"given":"Luigi","family":"Celona","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Simone","family":"Bianco","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marco","family":"Donzella","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Paolo","family":"Napoletano","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,9,25]]},"reference":[{"key":"11672_CR1","first-page":"23716","volume":"35","author":"JB Alayrac","year":"2022","unstructured":"Alayrac JB, Donahue J, Luc P et al (2022) Flamingo: a visual language model for few-shot learning. Adv Neural Inf Process Syst 35:23716\u201323736","journal-title":"Adv Neural Inf Process Syst"},{"key":"11672_CR2","doi-asserted-by":"crossref","unstructured":"Anderson P, Fernando B, Johnson M et\u00a0al (2016) Spice: Semantic propositional image caption evaluation. In: European Conference on Computer Vision (ECCV), Springer, pp 382\u2013398","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"11672_CR3","doi-asserted-by":"crossref","unstructured":"Anderson P, He X, Buehler C et\u00a0al (2018) Bottom-up and top-down attention for image captioning and visual question answering. In: Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 6077\u20136086","DOI":"10.1109\/CVPR.2018.00636"},{"key":"11672_CR4","doi-asserted-by":"crossref","unstructured":"Aneja J, Agrawal H, Batra D et\u00a0al (2019) Sequential latent spaces for modeling the intention during diverse image captioning. In: International Conference on Computer Vision (ICCV). IEEE\/CVF, pp 4261\u20134270","DOI":"10.1109\/ICCV.2019.00436"},{"key":"11672_CR5","unstructured":"Banerjee S, Lavie A (2005) Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In: Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization, pp 65\u201372"},{"key":"11672_CR6","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown T, Mann B, Ryder N et al (2020) Language models are few-shot learners. Adv Neural Inf Process Syst 33:1877\u20131901","journal-title":"Adv Neural Inf Process Syst"},{"key":"11672_CR7","doi-asserted-by":"crossref","unstructured":"Changpinyo S, Sharma P, Ding N et\u00a0al (2021) Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In: Conference on Computer Vision and Pattern Recognition, IEEE\/CVF, pp 3558\u20133568","DOI":"10.1109\/CVPR46437.2021.00356"},{"key":"11672_CR8","unstructured":"Chen Q, Deng C, Wu Q (2022) Learning distinct and representative modes for image captioning. In: Advances in Neural Information Processing Systems"},{"key":"11672_CR9","unstructured":"Cho J, Lei J, Tan H et\u00a0al (2021) Unifying vision-and-language tasks via text generation. In: International Conference on Machine Learning (ICML). PMLR, pp 1931\u20131942"},{"key":"11672_CR10","doi-asserted-by":"crossref","unstructured":"Donahue J, Anne\u00a0Hendricks L, Guadarrama S et\u00a0al (2015) Long-term recurrent convolutional networks for visual recognition and description. In: Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 2625\u20132634","DOI":"10.1109\/CVPR.2015.7298878"},{"key":"11672_CR11","unstructured":"Dong H, Li J, Wu B et\u00a0al (2024) Benchmarking and improving detail image caption. arXiv preprint arXiv:2405.19092"},{"key":"11672_CR12","first-page":"35544","volume":"36","author":"L Fan","year":"2024","unstructured":"Fan L, Krishnan D, Isola P et al (2024) Improving clip training with language rewrites. Adv Neural Inf Process Syst 36:35544\u201335575","journal-title":"Adv Neural Inf Process Syst"},{"key":"11672_CR13","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2024.125847","volume":"264","author":"N Gao","year":"2025","unstructured":"Gao N, Yao R, Chen P et al (2025) Multi-granularity semantic relational mapping for image caption. Expert Syst Appl 264:125847","journal-title":"Expert Syst Appl"},{"key":"11672_CR14","doi-asserted-by":"crossref","unstructured":"Ge Y, Zeng X, Huffman JS et\u00a0al (2024) Visual fact checker: Enabling high-fidelity detailed caption generation. In: Conference on Computer Vision and Pattern Recognition (CVPR). IEEE\/CVF, pp 14033\u201314042","DOI":"10.1109\/CVPR52733.2024.01331"},{"key":"11672_CR15","doi-asserted-by":"crossref","unstructured":"Hu JC, Cavicchioli R, Capotondi A (2022a) Expansionnet v2: Block static expansion in fast end to end training for image captioning. arXiv preprint arXiv:2208.06551","DOI":"10.1109\/BigData59044.2023.10386812"},{"key":"11672_CR16","doi-asserted-by":"crossref","unstructured":"Hu X, Gan Z, Wang J et\u00a0al (2022b) Scaling up vision-language pre-training for image captioning. In: Computer Vision and Pattern Recognition (CVPR). IEEE\/CVF, pp 17980\u201317989","DOI":"10.1109\/CVPR52688.2022.01745"},{"key":"11672_CR17","doi-asserted-by":"crossref","unstructured":"Huang L, Wang W, Chen J et\u00a0al (2019) Attention on attention for image captioning. In: International Conference on Computer Vision (ICCV). IEEE\/CVF, pp 4634\u20134643","DOI":"10.1109\/ICCV.2019.00473"},{"key":"11672_CR18","doi-asserted-by":"crossref","unstructured":"Karpathy A, Fei-Fei L (2015) Deep visual-semantic alignments for generating image descriptions. In: Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 3128\u20133137","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"11672_CR19","doi-asserted-by":"crossref","unstructured":"Kim T, Lee S, Kim SW et\u00a0al (2025) Vipcap: Retrieval text-based visual prompts for lightweight image captioning. In: AAAI Conference on Artificial Intelligence, pp 4320\u20134328","DOI":"10.1609\/aaai.v39i4.32454"},{"key":"11672_CR20","doi-asserted-by":"crossref","unstructured":"Lai Z, Zhang H, Zhang B et\u00a0al (2025) Veclip: Improving clip training via visual-enriched captions. In: European Conference on Computer Vision. Springer, pp 111\u2013127","DOI":"10.1007\/978-3-031-72946-1_7"},{"key":"11672_CR21","first-page":"9694","volume":"34","author":"J Li","year":"2021","unstructured":"Li J, Selvaraju R, Gotmare A et al (2021) Align before fuse: vision and language representation learning with momentum distillation. Adv Neural Inf Process Syst 34:9694\u20139705","journal-title":"Adv Neural Inf Process Syst"},{"key":"11672_CR22","unstructured":"Li J, Li D, Xiong C et\u00a0al (2022) BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In: International Conference on Machine Learning (ICML). PMLR"},{"key":"11672_CR23","unstructured":"Li J, Li D, Savarese S et\u00a0al (2023) BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In: International Conference on Machine Learning (ICML). PMLR, pp 19730\u201319742"},{"key":"11672_CR24","doi-asserted-by":"crossref","unstructured":"Li J, Vo DM, Sugimoto A et\u00a0al (2024) Evcap: Retrieval-augmented image captioning with external visual-name memory for open-world comprehension. In: Conference on Computer Vision and Pattern Recognition, IEEE\/CVF, pp 13733\u201313742","DOI":"10.1109\/CVPR52733.2024.01303"},{"key":"11672_CR25","doi-asserted-by":"crossref","unstructured":"Li W, Gao C, Niu G et\u00a0al (2021b) Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp 2592\u20132607","DOI":"10.18653\/v1\/2021.acl-long.202"},{"key":"11672_CR26","doi-asserted-by":"crossref","unstructured":"Li X, Yin X, Li C et\u00a0al (2020) Oscar: Object-semantics aligned pre-training for vision-language tasks. In: European Conference on Computer Vision (ECCV). Springer, pp 121\u2013137","DOI":"10.1007\/978-3-030-58577-8_8"},{"key":"11672_CR27","doi-asserted-by":"crossref","unstructured":"Lin TY, Maire M, Belongie S et\u00a0al (2014) Microsoft coco: Common objects in context. In: European Conference on Computer Vision (ECCV). Springer, pp 740\u2013755","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"11672_CR28","doi-asserted-by":"crossref","unstructured":"Luo Z, Hu Z, Xi Y et al (2023) I-tuning: tuning frozen language models with image for lightweight image captioning. International Conference on Acoustics. IEEE, Speech and Signal Processing (ICASSP), pp 1\u20135","DOI":"10.1109\/ICASSP49357.2023.10096424"},{"key":"11672_CR29","doi-asserted-by":"crossref","unstructured":"Luu DT, Le VT, Vo DM (2024) Questioning, answering, and captioning for zero-shot detailed image caption. In: Asian Conference on Computer Vision (ACCV), pp 242\u2013259","DOI":"10.1007\/978-981-96-2641-0_17"},{"key":"11672_CR30","doi-asserted-by":"publisher","unstructured":"NLP Connect (2022) vit-gpt2-image-captioning (revision 0e334c7). https:\/\/doi.org\/10.57967\/hf\/0222, https:\/\/huggingface.co\/nlpconnect\/vit-gpt2-image-captioning","DOI":"10.57967\/hf\/0222"},{"key":"11672_CR31","unstructured":"Ordonez V, Kulkarni G, Berg T (2011) Im2text: Describing images using 1 million captioned photographs. Advances in neural information processing systems 24"},{"key":"11672_CR32","doi-asserted-by":"crossref","unstructured":"Papineni K, Roukos S, Ward T et\u00a0al (2002) Bleu: a method for automatic evaluation of machine translation. In: Annual meeting of the Association for Computational Linguistics, pp 311\u2013318","DOI":"10.3115\/1073083.1073135"},{"key":"11672_CR33","doi-asserted-by":"crossref","unstructured":"Petryk S, Chan DM, Kachinthaya A et\u00a0al (2024) Aloha: A new measure for hallucination in captioning models. In: Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics","DOI":"10.18653\/v1\/2024.naacl-short.30"},{"key":"11672_CR34","unstructured":"Radford A, Kim JW, Hallacy C et\u00a0al (2021) Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning (ICML). PMLR, pp 8748\u20138763"},{"key":"11672_CR35","doi-asserted-by":"crossref","unstructured":"Rohrbach A, Hendricks LA, Burns K et\u00a0al (2018) Object hallucination in image captioning. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp 4035\u20134045","DOI":"10.18653\/v1\/D18-1437"},{"key":"11672_CR36","doi-asserted-by":"crossref","unstructured":"Rotstein N, Bensa\u00efd D, Brody S et\u00a0al (2024) Fusecap: Leveraging large language models for enriched fused image captions. In: Winter Conference on Applications of Computer Vision (WACV). IEEE\/CVF, pp 5677\u20135688","DOI":"10.1109\/WACV57701.2024.00559"},{"key":"11672_CR37","doi-asserted-by":"crossref","unstructured":"Sharma P, Ding N, Goodman S et\u00a0al (2018) Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In: Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp 2556\u20132565","DOI":"10.18653\/v1\/P18-1238"},{"key":"11672_CR38","unstructured":"Shen S, Li LH, Tan H et\u00a0al (2022) How much can clip benefit vision-and-language tasks? In: International Conference on Learning Representations (ICLR)"},{"key":"11672_CR39","doi-asserted-by":"crossref","unstructured":"Shi Z, Liu H, Zhu X (2021) Enhancing descriptive image captioning with natural language inference. In: Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pp 269\u2013277","DOI":"10.18653\/v1\/2021.acl-short.36"},{"key":"11672_CR40","doi-asserted-by":"crossref","unstructured":"Smith NA, Eisner J (2005) Contrastive estimation: Training log-linear models on unlabeled data. In: Annual Meeting of the Association for Computational Linguistics (ACL), pp 354\u2013362","DOI":"10.3115\/1219840.1219884"},{"key":"11672_CR41","doi-asserted-by":"crossref","unstructured":"Sun C, Shrivastava A, Singh S et\u00a0al (2017) Revisiting unreasonable effectiveness of data in deep learning era. In: International Conference on Computer Vision. IEEE, pp 843\u2013852","DOI":"10.1109\/ICCV.2017.97"},{"key":"11672_CR42","doi-asserted-by":"crossref","unstructured":"Vedantam R, Lawrence\u00a0Zitnick C, Parikh D (2015) Cider: Consensus-based image description evaluation. In: Computer Vision and Pattern Recognition (CVPR). IEEE, pp 4566\u20134575","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"11672_CR43","doi-asserted-by":"crossref","unstructured":"Wada Y, Kaneda K, Saito D et\u00a0al (2024) Polos: Multimodal metric learning from human feedback for image captioning. In: Conference on Computer Vision and Pattern Recognition, IEEE\/CVF, pp 13559\u201313568","DOI":"10.1109\/CVPR52733.2024.01287"},{"key":"11672_CR44","unstructured":"Wang J, Yang Z, Hu X et\u00a0al (2022a) GIT: A generative image-to-text transformer for vision and language. Transactions on Machine Learning Research"},{"key":"11672_CR45","unstructured":"Wang P, Yang A, Men R et\u00a0al (2022b) OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In: International Conference on Machine Learning (ICML). PMLR, pp 23318\u201323340"},{"key":"11672_CR46","doi-asserted-by":"crossref","unstructured":"Wang S, Yao Z, Wang R et\u00a0al (2021) Faier: Fidelity and adequacy ensured image caption evaluation. In: Conference on Computer Vision and Pattern Recognition (CVPR). IEEE\/CVF, pp 14050\u201314059","DOI":"10.1109\/CVPR46437.2021.01383"},{"key":"11672_CR47","unstructured":"Wang Z, Yu J, Yu AW et\u00a0al (2022c) SimVLM: Simple visual language model pretraining with weak supervision. In: International Conference on Learning Representations (ICLR)"},{"key":"11672_CR48","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2024.123847","volume":"250","author":"X Yang","year":"2024","unstructured":"Yang X, Yang Y, Wu J et al (2024) Ca-captioner: a novel concentrated attention for image captioning. Expert Syst Appl 250:123847","journal-title":"Expert Syst Appl"},{"key":"11672_CR49","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1162\/tacl_a_00166","volume":"2","author":"P Young","year":"2014","unstructured":"Young P, Lai A, Hodosh M et al (2014) From image descriptions to visual denotations: new similarity metrics for semantic inference over event descriptions. Trans Assoc Comput linguist 2:67\u201378","journal-title":"Trans Assoc Comput linguist"},{"key":"11672_CR50","unstructured":"Yu J, Wang Z, Vasudevan V et\u00a0al (2022) Coca: Contrastive captioners are image-text foundation models. Transactions on Machine Learning Research"},{"key":"11672_CR51","doi-asserted-by":"crossref","unstructured":"Yu Q, Sun Q, Zhang X et\u00a0al (2024) Capsfusion: Rethinking image-text data at scale. In: Conference on Computer Vision and Pattern Recognition (CVPR). IEEE\/CVF, pp 14022\u201314032","DOI":"10.1109\/CVPR52733.2024.01330"},{"key":"11672_CR52","doi-asserted-by":"crossref","unstructured":"Zha D, Bhat ZP, Lai KH et\u00a0al (2023) Data-centric ai: Perspectives and challenges. In: International Conference on Data Mining (SDM). SIAM, pp 945\u2013948","DOI":"10.1137\/1.9781611977653.ch106"},{"key":"11672_CR53","unstructured":"Zhou C, Gu J, Neubig G (2020) Understanding knowledge distillation in non-autoregressive machine translation. In: International Conference on Learning Representations (ICLR)"},{"key":"11672_CR54","unstructured":"Zhu D, Chen J, Haydarov K et\u00a0al (2024) Chatgpt asks, blip-2 answers: Automatic questioning towards enriched visual descriptions. Transactions on Machine Learning Research"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-025-11672-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-025-11672-x","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-025-11672-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,6]],"date-time":"2026-01-06T08:32:03Z","timestamp":1767688323000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-025-11672-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,25]]},"references-count":54,"journal-issue":{"issue":"32","published-print":{"date-parts":[[2025,11]]}},"alternative-id":["11672"],"URL":"https:\/\/doi.org\/10.1007\/s00521-025-11672-x","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,25]]},"assertion":[{"value":"16 April 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 August 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 September 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}