{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T17:19:05Z","timestamp":1782407945787,"version":"3.54.5"},"reference-count":88,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2024,10,8]],"date-time":"2024-10-08T00:00:00Z","timestamp":1728345600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,10,8]],"date-time":"2024-10-08T00:00:00Z","timestamp":1728345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Research Grants Council of the Hong Kong Special Administrative Region","award":["Project No. CityU 11215820"],"award-info":[{"award-number":["Project No. CityU 11215820"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["Grant 62301063"],"award-info":[{"award-number":["Grant 62301063"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["No.070323006"],"award-info":[{"award-number":["No.070323006"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2025,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Recent advances in image captioning have focused on enhancing accuracy by substantially increasing the dataset and model size. While conventional captioning models exhibit high performance on established metrics such as BLEU, CIDEr, and SPICE, the capability of captions to distinguish the target image from other similar images is under-explored. To generate distinctive captions, a few pioneers employed contrastive learning or re-weighted the ground-truth captions. However, these approaches often overlook the relationships among objects in a similar image group\u00a0(e.g., items or properties within the same album or fine-grained events). In this paper, we introduce a novel approach to enhance the distinctiveness of image captions, namely Group-based Differential Distinctive Captioning Method, which visually compares each image with other images in one similar group and highlights the uniqueness of each image. In particular, we introduce a Group-based Differential Memory Attention\u00a0(GDMA) module, designed to identify and emphasize object features in an image that are uniquely distinguishable within its image group, i.e., those exhibiting low similarity with objects in other images. This mechanism ensures that such unique object features are prioritized during caption generation for the image, thereby enhancing the distinctiveness of the resulting captions. To further refine this process, we select distinctive words from the ground-truth captions to guide both the language decoder and the GDMA module. Additionally, we propose a new evaluation metric, the Distinctive Word Rate (DisWordRate), to quantitatively assess caption distinctiveness. Quantitative results indicate that the proposed method significantly improves the distinctiveness of several baseline models, and achieves state-of-the-art performance on distinctiveness while not excessively sacrificing accuracy. Moreover, the results of our user study are consistent with the quantitative evaluation and demonstrate the rationality of the new metric DisWordRate.<\/jats:p>","DOI":"10.1007\/s11263-024-02220-6","type":"journal-article","created":{"date-parts":[[2024,10,8]],"date-time":"2024-10-08T07:02:54Z","timestamp":1728370974000},"page":"1435-1455","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Group-Based Distinctive Image Captioning with Memory Difference Encoding and Attention"],"prefix":"10.1007","volume":"133","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6113-0066","authenticated-orcid":false,"given":"Jiuniu","family":"Wang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenjia","family":"Xu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qingzhong","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Antoni B.","family":"Chan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,10,8]]},"reference":[{"key":"2220_CR1","doi-asserted-by":"crossref","unstructured":"Anderson, P., Fernando, B., Johnson, M., & Gould, S. (2016). SPICE: Semantic propositional image caption evaluation. In ECCV.","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"2220_CR2","doi-asserted-by":"crossref","unstructured":"Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., & Zhang, L. (2018). Bottom-up and top-down attention for image captioning and visual question answering. In CVPR.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"2220_CR3","doi-asserted-by":"crossref","unstructured":"Aneja, J., Agrawal, H., Batra, D., & Schwing, A. (2019). Sequential latent spaces for modeling the intention during diverse image captioning. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 4261\u20134270).","DOI":"10.1109\/ICCV.2019.00436"},{"key":"2220_CR4","doi-asserted-by":"crossref","unstructured":"Aneja, J., Deshpande, A., & Schwing, A. G. (2018). Convolutional image captioning. In CVPR.","DOI":"10.1109\/CVPR.2018.00583"},{"key":"2220_CR5","doi-asserted-by":"crossref","unstructured":"Cai, D., Zhao, L., Zhang, J., Sheng, L., & Xu, D. (2022). 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 16464\u201316473).","DOI":"10.1109\/CVPR52688.2022.01597"},{"key":"2220_CR6","doi-asserted-by":"crossref","unstructured":"Chen, J., Guo, H., Yi, K., Li, B., & Elhoseiny, M. (2022). VisualGPT: Data-efficient adaptation of pretrained language models for image captioning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 18030\u201318040).","DOI":"10.1109\/CVPR52688.2022.01750"},{"key":"2220_CR7","doi-asserted-by":"crossref","unstructured":"Chen, F., Ji, R., Su, J., Wu, Y., & Wu, Y. (2017a). Structcap: Structured semantic embedding for image captioning. In ACM MM.","DOI":"10.1145\/3123266.3123275"},{"key":"2220_CR8","doi-asserted-by":"crossref","unstructured":"Chen, F., Ji, R., Sun, X., Wu, Y., & Su, J. (2018). Groupcap: Group-based image captioning with structured relevance and diversity constraints. In CVPR.","DOI":"10.1109\/CVPR.2018.00146"},{"key":"2220_CR9","doi-asserted-by":"crossref","unstructured":"Chen, L., Zhang, H., Xiao, J., Nie, L., Shao, J., Liu, W., & Chua, T. S. (2017b). SCA-CNN: Spatial and channel-wise attention in convolutional networks for image captioning. In CVPR.","DOI":"10.1109\/CVPR.2017.667"},{"key":"2220_CR10","doi-asserted-by":"crossref","unstructured":"Cornia, M., Stefanini, M., Baraldi, L., & Cucchiara, R. (2020). Meshed-memory transformer for image captioning. In CVPR.","DOI":"10.1109\/CVPR42600.2020.01059"},{"key":"2220_CR11","unstructured":"Dai, B., & Lin, D. (2017). Contrastive learning for image captioning. In NeurIPS."},{"key":"2220_CR12","doi-asserted-by":"crossref","unstructured":"Dai, B., Fidler, S., Urtasun, R., & Lin, D. (2017). Towards diverse and natural image descriptions via a conditional GAN. In ICCV.","DOI":"10.1109\/ICCV.2017.323"},{"key":"2220_CR13","doi-asserted-by":"crossref","DOI":"10.1016\/j.imavis.2022.104515","volume":"124","author":"B Demirel","year":"2022","unstructured":"Demirel, B., & Cinbis, R. G. (2022). Caption generation on scenes with seen and unseen object categories. Image and Vision Computing, 124, 104515.","journal-title":"Image and Vision Computing"},{"key":"2220_CR14","doi-asserted-by":"crossref","unstructured":"Denkowski, M., & Lavie, A. (2014). Meteor universal: Language specific translation evaluation for any target language. In EACL workshop.","DOI":"10.3115\/v1\/W14-3348"},{"key":"2220_CR15","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., & Uszkoreit, J. (2020). An image is worth 16 $$\\times $$ 16 words: Transformers for image recognition at scale. In ICLR."},{"key":"2220_CR16","unstructured":"Faghri, F., Fleet, D. J., Kiros, J. R., & Fidler, S. (2018). VSE++: Improving visual-semantic embeddings with hard negatives. In BMVC."},{"key":"2220_CR17","doi-asserted-by":"crossref","unstructured":"Fei, J., Wang, T., Zhang, J., He, Z., Wang, C., & Zheng, F. (2023). Transferable decoding with visual entities for zero-shot image captioning. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 3136\u20133146).","DOI":"10.1109\/ICCV51070.2023.00291"},{"key":"2220_CR18","doi-asserted-by":"crossref","unstructured":"Guo, L., Liu, J., Zhu, X., Yao, P., Lu, S., & Lu, H. (2020). Normalized and geometry-aware self-attention network for image captioning. In CVPR.","DOI":"10.1109\/CVPR42600.2020.01034"},{"key":"2220_CR19","doi-asserted-by":"crossref","unstructured":"Hessel, J., Holtzman, A., Forbes, M., Bras, R. L., & Choi, Y. (2021). Clipscore: A reference-free evaluation metric for image captioning. In EMNLP.","DOI":"10.18653\/v1\/2021.emnlp-main.595"},{"issue":"9","key":"2220_CR20","doi-asserted-by":"crossref","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","volume":"37","author":"K He","year":"2015","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2015). Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 37(9), 1904\u20131916.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2220_CR21","doi-asserted-by":"crossref","unstructured":"Hirota, Y., Nakashima, Y., & Garcia, N. (2022). Quantifying societal bias amplification in image captioning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 13450\u201313459).","DOI":"10.1109\/CVPR52688.2022.01309"},{"key":"2220_CR22","doi-asserted-by":"crossref","unstructured":"Hu, X., Gan, Z., Wang, J., Yang, Z., Liu, Z., Lu, Y., & Wang, L. (2022). Scaling up vision-language pre-training for image captioning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 17980\u201317989).","DOI":"10.1109\/CVPR52688.2022.01745"},{"key":"2220_CR23","doi-asserted-by":"crossref","unstructured":"Huang, L., Wang, W., Chen, J., & Wei, X. Y. (2019). Attention on attention for image captioning. In ICCV.","DOI":"10.1109\/ICCV.2019.00473"},{"key":"2220_CR24","doi-asserted-by":"crossref","first-page":"4013","DOI":"10.1109\/TIP.2020.2969330","volume":"29","author":"Y Huang","year":"2020","unstructured":"Huang, Y., Chen, J., Ouyang, W., Wan, W., & Xue, Y. (2020). Image captioning with end-to-end attribute detection and subsequent attributes prediction. IEEE Transactions on Image processing, 29, 4013\u20134026.","journal-title":"IEEE Transactions on Image processing"},{"key":"2220_CR25","doi-asserted-by":"crossref","unstructured":"Jain, U., Zhang, Z., & Schwing, A. G. (2017). Creativity: Generating diverse questions using variational autoencoders. In CVPR.","DOI":"10.1109\/CVPR.2017.575"},{"key":"2220_CR26","unstructured":"Jia, C., Yang, Y., Xia, Y., Chen, Y. T., Parekh, Z., Pham, H., Le, Q., Sung, Y. H., Li, Z., & Duerig, T. (2021). Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning, PMLR (pp. 4904\u20134916)."},{"key":"2220_CR27","doi-asserted-by":"crossref","first-page":"3920","DOI":"10.1109\/TIP.2022.3177318","volume":"31","author":"W Jiang","year":"2022","unstructured":"Jiang, W., Zhu, M., Fang, Y., Shi, G., Zhao, X., & Liu, Y. (2022). Visual cluster grounding for image captioning. IEEE Transactions on Image Processing, 31, 3920.","journal-title":"IEEE Transactions on Image Processing"},{"key":"2220_CR28","doi-asserted-by":"crossref","unstructured":"Karpathy, A., & Fei-Fei, L. (2015). Deep visual-semantic alignments for generating image descriptions. In CVPR.","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"2220_CR29","doi-asserted-by":"crossref","unstructured":"Kim, T., Ahn, P., Kim, S., Lee, S., Marsden, M., Sala, A., Kim, S. H., Han, B., Lee, K. M., Lee, H., Bae, K., Wu, X., Gao, Y., Zhang, H., Yang, Y., Guo, W., Lu, J., Oh, Y., Cho, J. W., ... & Sun, M. (2023). NICE: CVPR 2023 challenge on zero-shot image captioning. 2309.01961","DOI":"10.1109\/CVPRW63382.2024.00731"},{"key":"2220_CR30","doi-asserted-by":"crossref","unstructured":"Kuo, C. W., & Kira, Z. (2022). Beyond a pre-trained object detector: Cross-modal textual and visual context for image captioning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 17969\u201317979).","DOI":"10.1109\/CVPR52688.2022.01744"},{"key":"2220_CR31","doi-asserted-by":"crossref","unstructured":"Kuo, C. W., & Kira, Z. (2023). HAAV: Hierarchical aggregation of augmented views for image captioning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 11039\u201311049).","DOI":"10.1109\/CVPR52729.2023.01062"},{"key":"2220_CR32","doi-asserted-by":"crossref","unstructured":"Lee, H., Yoon, S., Dernoncourt, F., Bui, T., & Jung, K. (2021). UMIC: An unreferenced metric for image captioning via contrastive learning. In ACL.","DOI":"10.18653\/v1\/2021.acl-short.29"},{"key":"2220_CR33","doi-asserted-by":"crossref","unstructured":"Lee, H., Yoon, S., Dernoncourt, F., Kim, D. S., Bui, T., & Jung, K. (2020). Vilbertscore: Evaluating image caption using vision-and-language bert. In Proceedings of the first workshop on evaluation and comparison of NLP systems (pp. 34\u201339).","DOI":"10.18653\/v1\/2020.eval4nlp-1.4"},{"key":"2220_CR34","unstructured":"Li, J., Li, D., Savarese, S., & Hoi, S. (2023). BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, PMLR (pp. 19730\u201319742)."},{"key":"2220_CR35","doi-asserted-by":"crossref","unstructured":"Li, Z., Tran, Q., Mai, L., Lin, Z., & Yuille, A. L. (2020). Context-aware group captioning via self-attention and contrastive features. In CVPR.","DOI":"10.1109\/CVPR42600.2020.00350"},{"key":"2220_CR36","doi-asserted-by":"crossref","unstructured":"Li, G., Zhu, L., Liu, P., & Yang, Y. (2019). Entangled transformer for image captioning. In ICCV.","DOI":"10.1109\/ICCV.2019.00902"},{"key":"2220_CR37","doi-asserted-by":"crossref","unstructured":"Liu, X., Li, H., Shao, J., Chen, D., & Wang, X. (2018). Show, tell and discriminate: Image captioning by self-retrieval with partially labeled data. In ECCV.","DOI":"10.1007\/978-3-030-01267-0_21"},{"key":"2220_CR38","unstructured":"Lu, J., Batra, D., Parikh, D., Lee, S. (2019). ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In NeurIPS (Vol.\u00a032)."},{"key":"2220_CR39","unstructured":"Luo, R., & Shakhnarovich, G. (2019). Analysis of diversity-accuracy tradeoff in image captioning. In ICCV Workshop."},{"key":"2220_CR40","doi-asserted-by":"crossref","unstructured":"Luo, R., Price, B., Cohen, S., & Shakhnarovich, G. (2018). Discriminability objective for training descriptive captions. In CVPR.","DOI":"10.1109\/CVPR.2018.00728"},{"key":"2220_CR41","first-page":"3613","volume":"33","author":"S Mahajan","year":"2020","unstructured":"Mahajan, S., & Roth, S. (2020). Diverse image captioning with context-object split latent spaces. Advances in Neural Information Processing Systems, 33, 3613\u20133624.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2220_CR42","unstructured":"Mao, J., Xu, W., Yang, Y., Wang, J., Huang, Z., & Yuille, A. (2015). Deep captioning with multimodal recurrent neural networks (m-RNN). In ICLR."},{"key":"2220_CR43","doi-asserted-by":"crossref","unstructured":"Mohamed, Y., Khan, F. F., Haydarov, K., & Elhoseiny, M. (2022). It is okay to not be okay: Overcoming emotional bias in affective image captioning by contrastive data collection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 21263\u201321272).","DOI":"10.1109\/CVPR52688.2022.02058"},{"key":"2220_CR44","doi-asserted-by":"crossref","unstructured":"Pan, Y., Yao, T., Li, Y., & Mei, T. (2020). X-linear attention networks for image captioning. In CVPR.","DOI":"10.1109\/CVPR42600.2020.01098"},{"key":"2220_CR45","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002). BLEU: A method for automatic evaluation of machine translation. In ACL.","DOI":"10.3115\/1073083.1073135"},{"key":"2220_CR46","unstructured":"Radford, A., Kim, JW., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., & Krueger, G. (2021). Learning transferable visual models from natural language supervision. In International conference on machine learning, PMLR (pp. 8748\u20138763)."},{"key":"2220_CR47","unstructured":"Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training."},{"issue":"8","key":"2220_CR48","first-page":"9","volume":"1","author":"A Radford","year":"2019","unstructured":"Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI blog, 1(8), 9.","journal-title":"OpenAI blog"},{"key":"2220_CR49","unstructured":"Ramachandran, P., Parmar, N., Vaswani, A., Bello, I., Levskaya, A., & Shlens, J. (2019). Stand-alone self-attention in vision models. In NeurIPS."},{"key":"2220_CR50","doi-asserted-by":"crossref","unstructured":"Rasiwasia, N., Costa\u00a0Pereira, J., Coviello, E., Doyle, G., Lanckriet, G. R., Levy, R., & Vasconcelos, N. (2010). A new approach to cross-modal multimedia retrieval. In Proceedings of the 18th ACM international conference on multimedia (pp. 251\u2013260).","DOI":"10.1145\/1873951.1873987"},{"key":"2220_CR51","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 779\u2013788).","DOI":"10.1109\/CVPR.2016.91"},{"key":"2220_CR52","unstructured":"Ren, S., He, K., Girshick, R., & Sun, J. (2015) Faster R-CNN: Towards real-time object detection with region proposal networks. In NeurIPS."},{"key":"2220_CR53","doi-asserted-by":"crossref","unstructured":"Rennie, S. J., Marcheret, E., Mroueh, Y., Ross, J., & Goel, V. (2017). Self-critical sequence training for image captioning. In CVPR.","DOI":"10.1109\/CVPR.2017.131"},{"key":"2220_CR54","doi-asserted-by":"crossref","unstructured":"Sarto, S., Barraco, M., Cornia, M., Baraldi, L., & Cucchiara, R. (2023). Positive-augmented contrastive learning for image and video captioning evaluation. In CVPR (pp. 6914\u20136924).","DOI":"10.1109\/CVPR52729.2023.00668"},{"key":"2220_CR55","doi-asserted-by":"crossref","unstructured":"Sharma, P., Ding, N., Goodman, S., & Soricut, R. (2018). Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th annual meeting of the association for computational linguistics (Volume 1: Long Papers) (pp. 2556\u20132565).","DOI":"10.18653\/v1\/P18-1238"},{"key":"2220_CR56","doi-asserted-by":"crossref","unstructured":"Shetty, R., Rohrbach, M., & Hendricks, L. A. (2017). Speaking the same language: Matching machine to human captions by adversarial training. In ICCV.","DOI":"10.1109\/ICCV.2017.445"},{"issue":"1","key":"2220_CR57","doi-asserted-by":"crossref","first-page":"539","DOI":"10.1109\/TPAMI.2022.3148210","volume":"45","author":"M Stefanini","year":"2022","unstructured":"Stefanini, M., Cornia, M., Baraldi, L., Cascianelli, S., Fiameni, G., & Cucchiara, R. (2022). From show to tell: A survey on deep learning-based image captioning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1), 539.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2220_CR58","unstructured":"Su, W., Zhu, X., Cao, Y., Li, B., Lu, L., Wei, F., & Dai, J. (2020). VL-BERT: Pre-training of generic visual-linguistic representations. In ICLR."},{"key":"2220_CR59","doi-asserted-by":"crossref","unstructured":"Tewel, Y., Shalev, Y., Schwartz, I., & Wolf, L. (2022). Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 17918\u201317928).","DOI":"10.1109\/CVPR52688.2022.01739"},{"key":"2220_CR60","unstructured":"Van\u00a0Miltenburg, E., Elliott, D., & Vossen, P. (2018). Measuring the diversity of automatic image descriptions. In Proceedings of the 27th international conference on computational linguistics (pp. 1730\u20131741)."},{"key":"2220_CR61","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, \u0141., & Polosukhin, I. (2017). Attention is all you need. In NeurIPS."},{"key":"2220_CR62","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Bengio, S., Murphy, K., Parikh, D., & Chechik, G. (2017). Context-aware captions from context-agnostic supervision. In CVPR.","DOI":"10.1109\/CVPR.2017.120"},{"key":"2220_CR63","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Lawrence\u00a0Zitnick, C., & Parikh, D. (2015). CIDEr: Consensus-based image description evaluation. In CVPR (pp. 4566\u20134575).","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"2220_CR64","doi-asserted-by":"crossref","unstructured":"Vered, G., Oren, G., Atzmon, Y., & Chechik, G. (2019). Joint optimization for cooperative image captioning. In CVPR.","DOI":"10.1109\/ICCV.2019.00899"},{"key":"2220_CR65","doi-asserted-by":"crossref","unstructured":"Vinyals, O., Toshev, A., Bengio, S., & Erhan, D. (2015). Show and tell: A neural image caption generator. In CVPR.","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"2220_CR66","unstructured":"Wang, Q., & Chan, A. B. (2018a). CNN\u00a0+\u00a0CNN: Convolutional decoders for image captioning. In CVPR Workshop."},{"key":"2220_CR67","doi-asserted-by":"crossref","unstructured":"Wang, Q., & Chan, A. B. (2018b). Gated hierarchical attention for image captioning. In ACCV.","DOI":"10.1007\/978-3-030-20870-7_2"},{"key":"2220_CR68","doi-asserted-by":"crossref","unstructured":"Wang, Q., & Chan, A. B. (2019). Describing like humans: On diversity in image captioning. In CVPR.","DOI":"10.1109\/CVPR.2019.00432"},{"key":"2220_CR69","unstructured":"Wang, Q., & Chan, A. B. (2020). Towards diverse and accurate image captions via reinforcing determinantal point process. In TPAMI."},{"key":"2220_CR70","unstructured":"Wang, L., Schwing, A. G., & Lazebnik, S. (2017). Diverse and accurate image description using a variational auto-encoder with an additive Gaussian encoding space. In NeurIPS."},{"key":"2220_CR71","unstructured":"Wang, Q., Wan, J., & Chan, A. B. (2020b). On diversity in image captioning: Metrics and methods. In IEEE TPAMI."},{"key":"2220_CR72","unstructured":"Wang, Q., Wang, J., Chan, A. B., Huang, S., Xiong, H., Li, X., & Dou, D. (2020c). Neighbours matter: Image captioning with similar images. In 31st British machine vision virtual conference (BMVC 2020)."},{"key":"2220_CR73","doi-asserted-by":"crossref","unstructured":"Wang, J., Xu, W., Wang, Q., & Chan, A. B. (2020a). Compare and reweight: Distinctive image captioning using similar images sets. In ECCV.","DOI":"10.1007\/978-3-030-58452-8_22"},{"key":"2220_CR74","doi-asserted-by":"crossref","unstructured":"Wang, J., Xu, W., Wang, Q., & Chan, A. B. (2021). Group-based distinctive image captioning with memory attention. In Proceedings of the 29th ACM international conference on multimedia (pp. 5020\u20135028).","DOI":"10.1145\/3474085.3475215"},{"key":"2220_CR75","doi-asserted-by":"crossref","first-page":"2088","DOI":"10.1109\/TPAMI.2022.3159811","volume":"45","author":"J Wang","year":"2022","unstructured":"Wang, J., Xu, W., Wang, Q., & Chan, A. B. (2022). On distinctive image captioning via comparing and reweighting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45, 2088.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2220_CR76","unstructured":"Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., & Bengio, Y. (2015). Show, attend and tell: Neural image caption generation with visual attention. In ICML."},{"key":"2220_CR77","unstructured":"Xu, H., Ye, Q., Yan, M., Shi, Y., Ye, J., Xu, Y., Li, C., Bi, B., Qian, Q., Wang, W., Xu, G., Zhang, J., Huang, S., Huang, F., & Zhou, J. (2023). mPLUG-2: A modularized multi-modal foundation model across text, image and video."},{"key":"2220_CR78","doi-asserted-by":"crossref","unstructured":"Yan, F., & Mikolajczyk, K. (2015). Deep correlation for matching images and text. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3441\u20133450).","DOI":"10.1109\/CVPR.2015.7298966"},{"key":"2220_CR79","doi-asserted-by":"crossref","unstructured":"Yang, Z., Garcia, N., Chu, C., Otani, M., Nakashima, Y., & Takemura, H. (2020). BERT representations for video question answering. In WACV.","DOI":"10.1109\/WACV45572.2020.9093596"},{"key":"2220_CR80","doi-asserted-by":"crossref","unstructured":"Ye, L., Rochan, M., Liu, Z., & Wang, Y. (2019). Cross-modal self-attention network for referring image segmentation. In CVPR.","DOI":"10.1109\/CVPR.2019.01075"},{"key":"2220_CR81","doi-asserted-by":"crossref","unstructured":"You, Q., Jin, H., Wang, Z., Fang, C., & Luo, J. (2016). Image captioning with semantic attention. In CVPR.","DOI":"10.1109\/CVPR.2016.503"},{"key":"2220_CR82","doi-asserted-by":"crossref","first-page":"1723","DOI":"10.1109\/TIP.2022.3145158","volume":"31","author":"J Yuan","year":"2022","unstructured":"Yuan, J., Zhu, S., Huang, S., Zhang, H., Xiao, Y., Li, Z., & Wang, M. (2022). Discriminative style learning for cross-domain image captioning. IEEE Transactions on Image Processing, 31, 1723\u20131736.","journal-title":"IEEE Transactions on Image Processing"},{"key":"2220_CR83","doi-asserted-by":"crossref","unstructured":"Zeng, Z., Zhang, H., Lu, R., Wang, D., Chen, B., & Wang, Z. (2023) Conzic: Controllable zero-shot image captioning by sampling-based polishing. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 23465\u201323476).","DOI":"10.1109\/CVPR52729.2023.02247"},{"key":"2220_CR84","unstructured":"Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2019). BERTScore: Evaluating text generation with bert. In ICLR."},{"key":"2220_CR85","doi-asserted-by":"crossref","unstructured":"Zhao, D., Wang, A., & Russakovsky, O. (2021). Understanding and evaluating racial biases in image captioning. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 14830\u201314840).","DOI":"10.1109\/ICCV48922.2021.01456"},{"key":"2220_CR86","doi-asserted-by":"crossref","first-page":"1180","DOI":"10.1109\/TIP.2020.3042086","volume":"30","author":"W Zhao","year":"2020","unstructured":"Zhao, W., Wu, X., & Luo, J. (2020). Cross-domain image captioning via cross-modal retrieval and model adaptation. IEEE Transactions on Image Processing, 30, 1180\u20131192.","journal-title":"IEEE Transactions on Image Processing"},{"key":"2220_CR87","doi-asserted-by":"crossref","unstructured":"Zhong, Y., Yang, J., Zhang, P., Li, C., Codella, N., Li, L. H., Zhou, L., Dai, X., Yuan, L., Li, Y., & Gao, J. (2022). Regionclip: Region-based language-image pretraining. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 16793\u201316803).","DOI":"10.1109\/CVPR52688.2022.01629"},{"key":"2220_CR88","doi-asserted-by":"crossref","unstructured":"Zhou, Y., Wang, M., Liu, D., Hu, Z., & Zhang, H. (2020). More grounded image captioning by distilling image-text matching model. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 4777\u20134786).","DOI":"10.1109\/CVPR42600.2020.00483"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-024-02220-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-024-02220-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-024-02220-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,3,30]],"date-time":"2025-03-30T22:11:59Z","timestamp":1743372719000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-024-02220-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,8]]},"references-count":88,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,4]]}},"alternative-id":["2220"],"URL":"https:\/\/doi.org\/10.1007\/s11263-024-02220-6","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,8]]},"assertion":[{"value":"3 April 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 July 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 October 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}