{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T15:40:29Z","timestamp":1783438829062,"version":"3.54.6"},"reference-count":49,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2024,4,1]],"date-time":"2024-04-01T00:00:00Z","timestamp":1711929600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,4,1]],"date-time":"2024-04-01T00:00:00Z","timestamp":1711929600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Process Lett"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Image captioning, which involves automatically generating textual descriptions based on the content of images, has garnered increasing attention from researchers. Recently, Transformers have emerged as the preferred choice for the language model in image captioning models. Transformers leverage self-attention mechanisms to address gradient accumulation issues and eliminate the risk of gradient explosion commonly associated with RNN networks. However, a challenge arises when the input features of the self-attention mechanism belong to different categories, as it may result in ineffective highlighting of important features. To address this issue, our paper proposes a novel attention mechanism called Self-Enhanced Attention (SEA), which replaces the self-attention mechanism in the decoder part of the Transformer model. In our proposed SEA, after generating the attention weight matrix, it further adjusts the matrix based on its own distribution to effectively highlight important features. To evaluate the effectiveness of SEA, we conducted experiments on the COCO dataset, comparing the results with different visual models and training strategies. The experimental results demonstrate that when using SEA, the CIDEr score is significantly higher compared to the scores obtained without using SEA. This indicates the successful addressing of the challenge of effectively highlighting important features with our proposed mechanism.<\/jats:p>","DOI":"10.1007\/s11063-024-11527-x","type":"journal-article","created":{"date-parts":[[2024,4,1]],"date-time":"2024-04-01T09:13:14Z","timestamp":1711962794000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["Self-Enhanced Attention for Image Captioning"],"prefix":"10.1007","volume":"56","author":[{"given":"Qingyu","family":"Sun","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Juan","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhijun","family":"Fang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yongbin","family":"Gao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,4,1]]},"reference":[{"key":"11527_CR1","doi-asserted-by":"crossref","unstructured":"Allaouzi I, Ben Ahmed M, Benamrou B, Ouardouz M (2018) Automatic caption generation for medical images. In: Proceedings of the 3rd International Conference on Smart City Applications, pp 1\u20136","DOI":"10.1145\/3286606.3286863"},{"key":"11527_CR2","doi-asserted-by":"crossref","unstructured":"Rennie S J, Marcheret E, Mroueh Y, Ross J, Goel V (2017) Self-critical sequence training for image captioning. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 7008\u20137024","DOI":"10.1109\/CVPR.2017.131"},{"key":"11527_CR3","doi-asserted-by":"publisher","unstructured":"Xiong Y, Du B, Yan P (2019) Reinforced transformer for medical image captioning. In Machine Learning in Medical Imaging: 10th International Workshop, MLMI 2019, Held in Conjunction with MICCAI 2019, Shenzhen, China, October 13, 2019, Proceedings 10, pp 673-680. Springer International Publishing. https:\/\/doi.org\/10.1007\/978-3-030-32692-0_77","DOI":"10.1007\/978-3-030-32692-0_77"},{"key":"11527_CR4","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.107856","volume":"114","author":"H Ayesha","year":"2021","unstructured":"Ayesha H, Iqbal S, Tariq M, Abrar M, Sanaullah M, Abbas I, Hussain S et al (2021) Automatic medical image interpretation: State of the art and future directions. Pattern Recogn 114:107856","journal-title":"Pattern Recogn"},{"issue":"13","key":"11527_CR5","doi-asserted-by":"publisher","first-page":"16747","DOI":"10.1007\/s10489-022-04313-6","volume":"53","author":"J Yu","year":"2023","unstructured":"Yu J, Zhang J, Gao Y (2023) MACFNet: multi-attention complementary fusion network for image denoising. Appl Intell 53(13):16747\u201316761","journal-title":"Appl Intell"},{"key":"11527_CR6","doi-asserted-by":"crossref","unstructured":"Huang L, Wang W, Chen J, Wei XY (2019) Attention on attention for image captioning. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 4634\u20134643","DOI":"10.1109\/ICCV.2019.00473"},{"key":"11527_CR7","doi-asserted-by":"crossref","unstructured":"Anderson P, He X, Buehler C, Teney D, Johnson M, Gould S, Zhang L (2018) Bottom-up and top-down attention for image captioning and visual question answering. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 6077\u20136086","DOI":"10.1109\/CVPR.2018.00636"},{"key":"11527_CR8","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2021.106832","volume":"218","author":"Y Zhang","year":"2021","unstructured":"Zhang Y, Zhang J, Huang B, Fang Z (2021) Single-image deraining via a recurrent memory unit network. Knowl-Based Syst 218:106832","journal-title":"Knowl-Based Syst"},{"key":"11527_CR9","doi-asserted-by":"crossref","unstructured":"Yang X, Tang K, Zhang H, Cai J (2019) Auto-encoding scene graphs for image captioning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 10685\u201310694","DOI":"10.1109\/CVPR.2019.01094"},{"key":"11527_CR10","doi-asserted-by":"crossref","unstructured":"Aneja J, Deshpande A, Schwing AG (2018) Convolutional image captioning. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 5561\u20135570","DOI":"10.1109\/CVPR.2018.00583"},{"key":"11527_CR11","doi-asserted-by":"crossref","unstructured":"Papineni K, Roukos S, Ward T, Zhu WJ (2002) Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp 311\u2013318","DOI":"10.3115\/1073083.1073135"},{"key":"11527_CR12","unstructured":"ROUGE LC (2004). A package for automatic evaluation of summaries. In: Proceedings of Workshop on Text Summarization of ACL, Spain"},{"key":"11527_CR13","doi-asserted-by":"crossref","unstructured":"Gu J, Wang G, Cai J, Chen T (2017) An empirical study of language cnn for image captioning. In: Proceedings of the IEEE international conference on computer vision, pp 1222\u20131231","DOI":"10.1109\/ICCV.2017.138"},{"key":"11527_CR14","doi-asserted-by":"crossref","unstructured":"Denkowski M, Lavie A (2014) Meteor universal: Language specific translation evaluation for any target language. In: Proceedings of the ninth workshop on statistical machine translation, pp 376\u2013380","DOI":"10.3115\/v1\/W14-3348"},{"key":"11527_CR15","doi-asserted-by":"crossref","unstructured":"Vedantam R, Lawrence Zitnick C, Parikh D (2015) Cider: Consensus-based image description evaluation. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4566\u20134575","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"11527_CR16","first-page":"517","volume":"2022","author":"J Cho","year":"2022","unstructured":"Cho J, Yoon S, Kale A, Dernoncourt F, Bui T, Bansal M (2022) Fine-grained Image Captioning with CLIP Reward. Find Assoc Comput Linguistics: NAACL 2022:517\u2013527","journal-title":"Find Assoc Comput Linguistics: NAACL"},{"key":"11527_CR17","doi-asserted-by":"crossref","unstructured":"Anderson P, Fernando B, Johnson M, Gould S (2016) Spice: Semantic propositional image caption evaluation. In: Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11\u201314, 2016, Proceedings, Part V 14, pp 382\u2013398. Springer","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"11527_CR18","doi-asserted-by":"crossref","unstructured":"Barraco M., Stefanini M., Cornia M., Cascianelli S, Baraldi L, Cucchiara R (2022, August) CaMEL: mean teacher learning for image captioning. In 2022 26th International Conference on Pattern Recognition (ICPR), pp 4087- 4094.","DOI":"10.1109\/ICPR56361.2022.9955644"},{"key":"11527_CR19","doi-asserted-by":"crossref","unstructured":"He S, Liao W, Tavakoli HR, Yang M, Rosenhahn B, Pugeault N (2020) Image captioning through image transformer. In: Proceedings of the Asian conference on computer vision","DOI":"10.1007\/978-3-030-69538-5_10"},{"key":"11527_CR20","doi-asserted-by":"crossref","unstructured":"Huang X, Belongie S (2017) Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE international conference on computer vision, pp 1501\u20131510","DOI":"10.1109\/ICCV.2017.167"},{"key":"11527_CR21","doi-asserted-by":"publisher","first-page":"249","DOI":"10.1016\/j.neucom.2020.03.087","volume":"401","author":"H Wang","year":"2020","unstructured":"Wang H, Wang H, Xu K (2020) Evolutionary recurrent neural network for image captioning. Neurocomputing 401:249\u2013256. https:\/\/doi.org\/10.1016\/j.neucom.2020.03.087","journal-title":"Neurocomputing"},{"key":"11527_CR22","doi-asserted-by":"crossref","unstructured":"Cornia M, Stefanini M, Baraldi L, Cucchiara R (2020) Meshed-memory transformer for image captioning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 10578\u201310587","DOI":"10.1109\/CVPR42600.2020.01059"},{"key":"11527_CR23","doi-asserted-by":"crossref","unstructured":"Kim Y, Soh JW, Park GY, Cho, NI (2020) Transfer learning from synthetic to real-noise denoising with adaptive instance normalization. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 3482\u20133492","DOI":"10.1109\/CVPR42600.2020.00354"},{"key":"11527_CR24","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, Polosukhin I et al (2017) Attention is all you need. Advances in neural information processing systems, 30."},{"key":"11527_CR25","doi-asserted-by":"crossref","unstructured":"Ling J, Xue H, Song L, Xie R., Gu X (2021) Region-aware adaptive instance normalization for ima-ge harmonization. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 9361\u20139370","DOI":"10.1109\/CVPR46437.2021.00924"},{"issue":"4","key":"11527_CR26","doi-asserted-by":"publisher","first-page":"652","DOI":"10.1109\/TPAMI.2016.2587640","volume":"39","author":"O Vinyals","year":"2016","unstructured":"Vinyals O, Toshev A, Bengio S, Erhan D (2016) Show and tell: Lessons learned from the 2015 mscoco image captioning challenge. IEEE Trans Pattern Anal Mach Intell 39(4):652\u2013663","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11527_CR27","doi-asserted-by":"publisher","DOI":"10.1007\/s11063-022-11106-y","author":"H Sharma","year":"2022","unstructured":"Sharma H, Srivastava S (2022) A Framework for Image Captioning Based on Relation Network and Multilevel Attention Mechanism. Neural Process Letters. https:\/\/doi.org\/10.1007\/s11063-022-11106-y","journal-title":"Neural Process Letters"},{"key":"11527_CR28","doi-asserted-by":"crossref","unstructured":"Pan Y, Yao T, Li Y, Mei T (2020) X-linear attention networks for image captioning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 10971\u201310980","DOI":"10.1109\/CVPR42600.2020.01098"},{"key":"11527_CR29","doi-asserted-by":"crossref","unstructured":"Lu J, Xiong C, Parikh D, Socher R (2017) Knowing when to look: Adaptive attention via a visual sentinel for image captioning. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 375\u2013383","DOI":"10.1109\/CVPR.2017.345"},{"key":"11527_CR30","doi-asserted-by":"crossref","unstructured":"Li G, Zhu L, Liu P, Yang Y (2019) Entangled transformer for image captioning. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 8928\u20138937","DOI":"10.1109\/ICCV.2019.00902"},{"key":"11527_CR31","doi-asserted-by":"crossref","unstructured":"Yao T, Pan Y, Li Y, Mei T (2019) Hierarchy parsing for image captioning. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 2621\u20132629","DOI":"10.1109\/ICCV.2019.00271"},{"key":"11527_CR32","unstructured":"Ribeiro AH, Tiels K, Aguirre LA, Sch\u00f6n T (2020) Beyond exploding and vanishing gradients: analysing RNN training using attractors and smoothness. In: International Conference on Artificial Intelligence and Statistics, pp 2370\u20132380. PMLR"},{"key":"11527_CR33","doi-asserted-by":"crossref","unstructured":"Zhao H, Jia J, Koltun V (2020) Exploring self-attention for image recognition. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 10076\u201310085","DOI":"10.1109\/CVPR42600.2020.01009"},{"issue":"11","key":"11527_CR34","doi-asserted-by":"publisher","first-page":"5514","DOI":"10.1109\/TIP.2018.2855406","volume":"27","author":"S Ye","year":"2018","unstructured":"Ye S, Han J, Liu N (2018) Attentive linear transformation for image captioning. IEEE Trans Image Process 27(11):5514\u20135524","journal-title":"IEEE Trans Image Process"},{"key":"11527_CR35","doi-asserted-by":"publisher","unstructured":"Sarto S, Cornia M, Baraldi L, Cucchiara R (2022) Retrieval-augmented transformer for image captioning. In: Proceedings of the 19th International Conference on Content-based Multimedia Indexing, pp 1\u20137. https:\/\/doi.org\/10.1145\/3549555.3549585","DOI":"10.1145\/3549555.3549585"},{"key":"11527_CR36","doi-asserted-by":"crossref","unstructured":"Gao J, Wang S, Wang S, Ma S, Gao W (2019) Self-critical n-step training for image captioning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 6300\u20136308","DOI":"10.1109\/CVPR.2019.00646"},{"key":"11527_CR37","doi-asserted-by":"crossref","unstructured":"Mishra SK, Dhir R, Saha S, Bhattacharyya P, Singh AK (2021) Image captioning in Hindi language using transformer networks. Computers & Electrical Engineering, 92, 107114. 10.1016 \/j.compeleceng.2021.107114","DOI":"10.1016\/j.compeleceng.2021.107114"},{"key":"11527_CR38","doi-asserted-by":"publisher","first-page":"1101","DOI":"10.1007\/s11063-021-10431-y","volume":"53","author":"H Zhu","year":"2021","unstructured":"Zhu H, Wang R, Zhang X (2021) Image Captioning with Dense Fusion Connection and Improved Stacked Attention Module. Neural Process Lett 53:1101\u20131118. https:\/\/doi.org\/10.1007\/s11063-021-10431-y","journal-title":"Neural Process Lett"},{"issue":"2","key":"11527_CR39","doi-asserted-by":"publisher","first-page":"1655","DOI":"10.1609\/aaai.v35i2.16258","volume":"35","author":"J Ji","year":"2021","unstructured":"Ji J, Luo Y, Sun X, Chen F, Luo G, Wu Y, Gao Y, Ji R (2021) Improving Image Captioning by Leveraging Intra-and Inter-layer Global Representation in Transformer Network. Proc AAAI Conf Artif Intell 35(2):1655\u20131663. https:\/\/doi.org\/10.1609\/aaai.v35i2.16258","journal-title":"Proc AAAI Conf Artif Intell"},{"issue":"3","key":"11527_CR40","doi-asserted-by":"publisher","first-page":"3801","DOI":"10.1007\/s11042-022-13443-5","volume":"82","author":"T Tiwary","year":"2023","unstructured":"Tiwary T, Mahapatra RP (2023) An accurate generation of image captions for blind people using extended convolutional atom neural network. Multimed Tools Appl 82(3):3801\u20133830","journal-title":"Multimed Tools Appl"},{"key":"11527_CR41","doi-asserted-by":"crossref","unstructured":"Jiang W, Ma L, Jiang YG, Liu W, Zhang T (2018) Recurrent fusion network for image captioning. In: Proceedings of the European conference on computer vision (ECCV), pp 499\u2013515.","DOI":"10.1007\/978-3-030-01216-8_31"},{"key":"11527_CR42","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2023.107234","volume":"127","author":"PD Chen","year":"2024","unstructured":"Chen PD, Zhang J, Gao YB, Fang ZJ, Hwang JN (2024) A lightweight RGB superposition effect adjustment network for low-light image enhancement and denoising. Eng Appl Artif Intell 127:107234","journal-title":"Eng Appl Artif Intell"},{"key":"11527_CR43","unstructured":"Chen X, Fang H, Lin TY, Vedantam R, Gupta S, Doll\u00e1r P, Zitnick CL (2015) Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325"},{"key":"11527_CR44","doi-asserted-by":"publisher","unstructured":"Lin TY, Maire M, Belongie SJ, Hays J, Perona P, Ramanan D et al (2014) Microsoft COCO: Common Objects in Context. In: Fleet D, Pajdla T, Schiele B, Tuytelaars T (eds) Computer Vision \u2013 ECCV 2014. ECCV 2014. Lecture Notes in Computer Science, vol 8693. Springer, Cham. https:\/\/doi.org\/10.1007\/978-3-319-10602-1_48","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"11527_CR45","doi-asserted-by":"crossref","unstructured":"Karpathy A, Fei-Fei L (2015) Deep visual-semantic alignments for generating image descriptions. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3128\u20133137","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"11527_CR46","doi-asserted-by":"crossref","unstructured":"Sennrich R, Haddow B, Birch A (2015) Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909","DOI":"10.18653\/v1\/P16-1162"},{"key":"11527_CR47","doi-asserted-by":"crossref","unstructured":"Vinyals O, Toshev A, Bengio S, Erhan D (2015) Show and tell: A neural image caption generator. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3156\u20133164.","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"11527_CR48","doi-asserted-by":"crossref","unstructured":"Yao T, Pan Y, Li Y, Qiu Z, Mei T (2017) Boosting image captioning with attributes. In Proceedings of the IEEE international conference on computer vision, pp 4894\u20134902","DOI":"10.1109\/ICCV.2017.524"},{"key":"11527_CR49","doi-asserted-by":"crossref","unstructured":"Zhang X, Sun X, Luo Y, Ji J, Zhou Y, Wu Y, Ji R (2021) Rstnet: Captioning with adaptive attention on visual and non-visual words. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 15465\u201315474.","DOI":"10.1109\/CVPR46437.2021.01521"}],"container-title":["Neural Processing Letters"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-024-11527-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11063-024-11527-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-024-11527-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,15]],"date-time":"2024-11-15T11:27:23Z","timestamp":1731670043000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11063-024-11527-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,1]]},"references-count":49,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2024,4]]}},"alternative-id":["11527"],"URL":"https:\/\/doi.org\/10.1007\/s11063-024-11527-x","relation":{},"ISSN":["1573-773X"],"issn-type":[{"value":"1573-773X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,1]]},"assertion":[{"value":"8 January 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 April 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Qingyu Sun, Juan Zhang, Zhijun Fang, and Yongbin Gao declare no conflict of interest. This paper also did not plagiarize the research results of other authors.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interests"}}],"article-number":"131"}}