{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T06:54:27Z","timestamp":1780988067858,"version":"3.54.1"},"reference-count":58,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2023,7,20]],"date-time":"2023-07-20T00:00:00Z","timestamp":1689811200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China","award":["62273142"],"award-info":[{"award-number":["62273142"]}]},{"name":"National Natural Science Foundation of China","award":["22A0490"],"award-info":[{"award-number":["22A0490"]}]},{"name":"National Natural Science Foundation of China","award":["2021GK5074"],"award-info":[{"award-number":["2021GK5074"]}]},{"name":"National Natural Science Foundation of China","award":["2021GK2010"],"award-info":[{"award-number":["2021GK2010"]}]},{"name":"Research Foundation of the Education Bureau of Hunan Province, China","award":["62273142"],"award-info":[{"award-number":["62273142"]}]},{"name":"Research Foundation of the Education Bureau of Hunan Province, China","award":["22A0490"],"award-info":[{"award-number":["22A0490"]}]},{"name":"Research Foundation of the Education Bureau of Hunan Province, China","award":["2021GK5074"],"award-info":[{"award-number":["2021GK5074"]}]},{"name":"Research Foundation of the Education Bureau of Hunan Province, China","award":["2021GK2010"],"award-info":[{"award-number":["2021GK2010"]}]},{"name":"Hunan Enterprise Science and Technology Commissioner program","award":["62273142"],"award-info":[{"award-number":["62273142"]}]},{"name":"Hunan Enterprise Science and Technology Commissioner program","award":["22A0490"],"award-info":[{"award-number":["22A0490"]}]},{"name":"Hunan Enterprise Science and Technology Commissioner program","award":["2021GK5074"],"award-info":[{"award-number":["2021GK5074"]}]},{"name":"Hunan Enterprise Science and Technology Commissioner program","award":["2021GK2010"],"award-info":[{"award-number":["2021GK2010"]}]},{"name":"science and technology innovation program of Hunan Province","award":["62273142"],"award-info":[{"award-number":["62273142"]}]},{"name":"science and technology innovation program of Hunan Province","award":["22A0490"],"award-info":[{"award-number":["22A0490"]}]},{"name":"science and technology innovation program of Hunan Province","award":["2021GK5074"],"award-info":[{"award-number":["2021GK5074"]}]},{"name":"science and technology innovation program of Hunan Province","award":["2021GK2010"],"award-info":[{"award-number":["2021GK2010"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Image captioning is a challenging task, which generates a sentence for a given image. The earlier captioning methods mainly decode the visual features to generate caption sentences for the image. However, the visual features lack the context semantic information which is vital for generating an accurate caption sentence. To address this problem, this paper first proposes the Attention-Aware (AA) mechanism which can filter out erroneous or irrelevant context semantic information. And then, AA is utilized to constitute a Context Semantic Auxiliary Network (CSAN), which can capture the effective context semantic information to regenerate or polish the image caption. Moreover, AA can capture the visual feature information needed to generate a caption. Experimental results show that our proposed CSAN outperforms the compared image captioning methods on MS COCO \u201cKarpathy\u201d offline test split and the official online testing server.<\/jats:p>","DOI":"10.3390\/info14070419","type":"journal-article","created":{"date-parts":[[2023,7,21]],"date-time":"2023-07-21T01:58:38Z","timestamp":1689904718000},"page":"419","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["A Context Semantic Auxiliary Network for Image Captioning"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4104-8401","authenticated-orcid":false,"given":"Jianying","family":"Li","sequence":"first","affiliation":[{"name":"School of Computer and Electrical Engineering, Hunan University of Arts and Science, Changde 415000, China"},{"name":"Key Laboratory of Hunan Province for Control Technology of Distributed Electric Propulsion Air Vehicle, Changde 415000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiangjun","family":"Shao","sequence":"additional","affiliation":[{"name":"School of Computer and Electrical Engineering, Hunan University of Arts and Science, Changde 415000, China"},{"name":"School of Computer Science, Wuhan University, Wuhan 430072, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,7,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2015, January 7\u201312). Show and tell: A neural image caption generator. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Donahue, J., Hendricks, L.A., Guadarrama, S., Rohrbach, M., Venugopalan, S., Darrell, T., and Saenko, K. (2015, January 7\u201312). Long-term recurrent convolutional networks for visual recognition and description. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298878"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Jia, X., Gavves, E., Fernando, B., and Tuytelaars, T. (2015, January 7\u201313). Guiding the long-short term mem-ory model for image caption generation. Proceedings of the 2015 IEEE International Conference on Computer Vision, ICCV, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.277"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wang, C., Yang, H., Bartz, C., and Meinel, C. (2016, January 15\u201319). Image Captioning with Deep Bidirectional LSTMs. Proceedings of the 2016 ACM Conference on Multimedia Conference, MM, Amsterdam, The Netherlands.","DOI":"10.1145\/2964284.2964299"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A.L., and Murphy, K. (2016, January 27\u201330). Generation and Comprehension of Unambiguous Object Descriptions. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.9"},{"key":"ref_6","first-page":"294","article-title":"Rethinking the Form of Latent States in Image Captioning","volume":"Volume 11209","author":"Dai","year":"2018","journal-title":"Proceedings of the Computer Vision\u2014ECCV"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"You, Q., Jin, H., Wang, Z., Fang, C., and Luo, J. (2016, January 27\u201330). Image Captioning with Semantic Attention. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.503"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Wang, Y., Lin, Z., Shen, X., Cohen, S., and Cottrell, G.W. (2017, January 21\u201326). Skeleton Key: Image Captioning by Skeleton-Attribute Decomposition. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.780"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Yao, T., Pan, Y., Li, Y., Qiu, Z., and Mei, T. (2017, January 22\u201329). Boosting image captioning with attributes. Proceedings of the 2017 IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.524"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Li, N., and Chen, Z. (2018, January 13\u201319). Image Cationing with Visual-Semantic LSTM. Proceedings of the 2018 Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI, Stockholm, Sweden.","DOI":"10.24963\/ijcai.2018\/110"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Huang, F., Li, Z., Chen, S., Zhang, C., and Ma, H. (2020, January 19\u201323). Image Captioning with Internal and External Knowledge. Proceedings of the CIKM \u201920: The 29th ACM International Conference on Information and Knowledge Management, CIKM, Virtual Event, Ireland.","DOI":"10.1145\/3340531.3411948"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Yao, T., Pan, Y., Li, Y., and Mei, T. (2018, January 8\u201314). Exploring Visual Relationship for Image Captioning. Proceedings of the Computer Vision\u2014ECCV, Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_42"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Chen, F., Ji, R., Sun, X., Wu, Y., and Su, J. (2018, January 18\u201322). GroupCap: Group-Based Image Captioning With Structured Relevance and Diversity Constraints. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00146"},{"key":"ref_14","unstructured":"Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y. (2015, January 6\u201311). Show, attend and tell: Neural image caption generation with visual attention. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_15","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2015, January 7\u20139). Neural Machine Translation by Jointly Learning to Align and Translate. Proceedings of the ICLR, San Diego, CA, USA."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Luo, Y., Ji, J., Sun, X., Cao, L., Wu, Y., Huang, F., Lin, C., and Ji, R. (2021, January 2\u20139). Dual-level Collaborative Transformer for Image Captioning. Proceedings of the Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI, Virtual Event.","DOI":"10.1609\/aaai.v35i3.16328"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"69700","DOI":"10.1109\/ACCESS.2021.3067607","article-title":"Multi-Gate Attention Network for Image Captioning","volume":"9","author":"Jiang","year":"2021","journal-title":"IEEE Access"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"129","DOI":"10.1016\/j.neunet.2022.01.011","article-title":"Dual Global Enhanced Transformer for image captioning","volume":"148","author":"Xian","year":"2022","journal-title":"Neural Netw."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Chen, L., Zhang, H., Xiao, J., Nie, L., Shao, J., Liu, W., and Chua, T. (2017, January 21\u201326). SCA-CNN: Spatial and Channel-Wise Attention in Convolutional Networks for Image Captioning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.667"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"2321","DOI":"10.1109\/TPAMI.2016.2642953","article-title":"Aligning Where to See and What to Tell: Image Captioning with Region-Based Attention and Scene-Specific Contexts","volume":"39","author":"Fu","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","unstructured":"Singh, S.P., and Markovitch, S. (2017, January 4\u20139). Attention Correctness in Neural Image Captioning. Proceedings of the Thirty-First AAAI, San Francisco, CA, USA."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Pedersoli, M., Lucas, T., Schmid, C., and Verbeek, J. (2017, January 22\u201329). Areas of Attention for Image Captioning. Proceedings of the IEEE International Conference on Computer Vision, ICCV, Venice, Italy.","DOI":"10.1109\/ICCV.2017.140"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"2149","DOI":"10.1109\/TMM.2019.2951226","article-title":"Show, Tell, and Polish: Ruminant Decoding for Image Captioning","volume":"22","author":"Guo","year":"2020","journal-title":"IEEE Trans. Multim."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Song, Z., Zhou, X., Mao, Z., and Tan, J. (2021, January 2\u20139). Image Captioning with Context-Aware Auxiliary Guidance. Proceedings of the Thirty-Fifth AAAI Conference on Artificial Intelligence, Virtual Event.","DOI":"10.1609\/aaai.v35i3.16361"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Liang, C., Wang, W., Zhou, T., Miao, J., Luo, Y., and Yang, Y. (2022). Local-Global Context Aware Transformer for Language-Guided Video Segmentation. arXiv.","DOI":"10.1109\/TPAMI.2023.3262578"},{"key":"ref_26","unstructured":"Precup, D., and Teh, Y.W. (2017, January 6\u201311). Language Modeling with Gated Convolutional Networks. Proceedings of the 2017 34th International Conference on Machine Learning, ICML, Sydney, Australia. Proceedings of Machine Learning Research."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L. (2018, January 18\u201322). Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Ke, L., Pei, W., Li, R., Shen, X., and Tai, Y. (November, January 27). Reflective Decoding Network for Image Captioning. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision, ICCV, Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00898"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Huang, L., Wang, W., Chen, J., and Wei, X. (November, January 27). Attention on Attention for Image Captioning. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision, ICCV, Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00473"},{"key":"ref_30","first-page":"15","article-title":"Every Picture Tells a Story: Generating Sentences from Images","volume":"Volume 6314","author":"Farhadi","year":"2010","journal-title":"Proceedings of the Computer Vision\u2014ECCV"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"2891","DOI":"10.1109\/TPAMI.2012.162","article-title":"BabyTalk: Understanding and Generating Simple Image Descriptions","volume":"35","author":"Kulkarni","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_32","unstructured":"Mitchell, M., Dodge, J., Goyal, A., Yamaguchi, K., Stratos, K., Han, X., Mensch, A.C., Berg, A.C., Berg, T.L., and Daume, H. (2012, January 23\u201327). Midge: Generating Image Descriptions From Computer Vision Detections. Proceedings of the EACL, Avignon, France."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Cho, K., van Merrienboer, B., G\u00fcl\u00e7ehre, \u00c7., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014, January 25\u201329). Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. Proceedings of the 2014 the Conference on Empirical Methods in Natural Language Processing, EMNLP, Doha, Qatar.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Liang, C., Wang, W., Zhou, T., and Yang, Y. (2022, January 19\u201320). Visual Abductive Reasoning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01512"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Lu, J., Xiong, C., Parikh, D., and Socher, R. (2017, January 21\u201326). Knowing when to look: Adaptive attention via a visual sentinel for image captioning. Proceedings of the 2017 The IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.345"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Yu, D., Fu, J., Mei, T., and Rui, Y. (2017, January 21\u201326). Multi-level Attention Networks for Visual Question Answering. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.446"},{"key":"ref_37","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017, January 4\u20139). Attention is All you Need. Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems, NIPS, Long Beach, CA, USA."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the CVPR, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Lin, T., Maire, M., Belongie, S.J., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft COCO: Common Objects in Context. Proceedings of the Computer Vision\u2014ECCV, Zurich, Switzerland. Lecture Notes in Computer Science.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W. (2002, January 6\u201312). Bleu: A Method for Automatic Evaluation of Machine Translation. Proceedings of the 2002 40th Annual Meeting of the Association for Computational Linguistics, ACL, Philadephia, PA, USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_42","unstructured":"Banerjee, S., and Lavie, A. (, January June). METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. Proceedings of the 2005 ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization, Ann Arbor, MI, USA."},{"key":"ref_43","unstructured":"Lin, C.Y. (2004, January 25\u201326). Rouge: A package for automatic evaluation of summaries. Proceedings of the Text Summarization Branches Out, Barcelona, Spain."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Zitnick, C.L., and Parikh, D. (2015, January 7\u201312). CIDEr: Consensus-based image description evaluation. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition, CVPR, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Leibe, B., Matas, J., Sebe, N., and Welling, M. (2016, January 11\u201314). SPICE: Semantic Propositional Image Caption Evaluation. Proceedings of the Computer Vision\u2014ECCV, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46478-7"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Li, Z., Li, Y., and Lu, H. (2019, January 12\u201315). Improve Image Captioning by Self-attention. Proceedings of the Neural Information Processing\u201426th International Conference, ICONIP, Sydney, NSW, Australia.","DOI":"10.1007\/978-3-030-36802-9_11"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"103138","DOI":"10.1016\/j.jvcir.2021.103138","article-title":"Attention-guided image captioning with adaptive global and local feature fusion","volume":"78","author":"Zhong","year":"2021","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"710","DOI":"10.1109\/TPAMI.2019.2909864","article-title":"Context-Aware Visual Policy Network for Fine-Grained Image Captioning","volume":"44","author":"Zha","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_49","unstructured":"Fei, Z. (March, January 22). Attention-Aligned Transformer for Image Captioning. Proceedings of the Thirty-Sixth AAAI Conference on Artificial Intelligence, AAAI, Virtual."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"103068","DOI":"10.1016\/j.cviu.2020.103068","article-title":"The synergy of double attention: Combine sentence-level and word-level attention for image captioning","volume":"201","author":"Wei","year":"2020","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"107075","DOI":"10.1016\/j.patcog.2019.107075","article-title":"Learning visual relationship and context-aware attention for image captioning","volume":"98","author":"Wang","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"75:1","DOI":"10.1145\/3386725","article-title":"Constrained LSTM and Residual Attention for Image Captioning","volume":"16","author":"Yang","year":"2020","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1016\/j.patrec.2020.12.020","article-title":"Image captioning with transformer and knowledge graph","volume":"143","author":"Zhang","year":"2021","journal-title":"Pattern Recognit. Lett."},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"103707","DOI":"10.1016\/j.dsp.2022.103707","article-title":"Local-global visual interaction attention for image captioning","volume":"130","author":"Wang","year":"2022","journal-title":"Digit. Signal Process."},{"key":"ref_55","first-page":"2361","article-title":"Review networks for caption generation","volume":"29","author":"Yang","year":"2016","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Jiang, W., Ma, L., Jiang, Y., Liu, W., and Zhang, T. (2018, January 8\u201314). Recurrent Fusion Network for Image Captioning. Proceedings of the Computer Vision\u2014ECCV, Munich, Germany.","DOI":"10.1007\/978-3-030-01216-8_31"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Chen, F., Xie, S., Li, X., Tang, J., Pang, K., Li, S., and Wang, T. (2021, January 5\u20139). Show, Rethink, And Tell: Image Caption Generation With Hierarchical Topic Cues. Proceedings of the IEEE International Conference on Multimedia and Expo, ICME, Shenzhen, China.","DOI":"10.1109\/ICME51207.2021.9428353"},{"key":"ref_58","doi-asserted-by":"crossref","first-page":"117174","DOI":"10.1016\/j.eswa.2022.117174","article-title":"Geometry Attention Transformer with position-aware LSTMs for image captioning","volume":"201","author":"Wang","year":"2022","journal-title":"Expert Syst. Appl."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/14\/7\/419\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:15:56Z","timestamp":1760127356000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/14\/7\/419"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,20]]},"references-count":58,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2023,7]]}},"alternative-id":["info14070419"],"URL":"https:\/\/doi.org\/10.3390\/info14070419","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,20]]}}}