{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T16:24:30Z","timestamp":1781367870536,"version":"3.54.1"},"reference-count":37,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2020,3,13]],"date-time":"2020-03-13T00:00:00Z","timestamp":1584057600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"The National Natural Science Foundation of China under Grants","award":["61772399, U1701267, 61773304, 61672405, and 61772400"],"award-info":[{"award-number":["61772399, U1701267, 61773304, 61672405, and 61772400"]}]},{"name":"The Key Research and Development Program in Shaanxi Province of China","award":["2019ZDLGY09-05"],"award-info":[{"award-number":["2019ZDLGY09-05"]}]},{"name":"The Program for Cheung Kong Scholars and Innovative Research Team in University","award":["IRT_15R53"],"award-info":[{"award-number":["IRT_15R53"]}]},{"name":"The Technology Foundation for Selected Overseas Chinese Scholar in Shaanxi","award":["2017021 and 2018021"],"award-info":[{"award-number":["2017021 and 2018021"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>The task of image captioning involves the generation of a sentence that can describe an image appropriately, which is the intersection of computer vision and natural language. Although the research on remote sensing image captions has just started, it has great significance. The attention mechanism is inspired by the way humans think, which is widely used in remote sensing image caption tasks. However, the attention mechanism currently used in this task is mainly aimed at images, which is too simple to express such a complex task well. Therefore, in this paper, we propose a multi-level attention model, which is a closer imitation of attention mechanisms of human beings. This model contains three attention structures, which represent the attention to different areas of the image, the attention to different words, and the attention to vision and semantics. Experiments show that our model has achieved better results than before, which is currently state-of-the-art. In addition, the existing datasets for remote sensing image captioning contain a large number of errors. Therefore, in this paper, a lot of work has been done to modify the existing datasets in order to promote the research of remote sensing image captioning.<\/jats:p>","DOI":"10.3390\/rs12060939","type":"journal-article","created":{"date-parts":[[2020,3,18]],"date-time":"2020-03-18T08:13:27Z","timestamp":1584519207000},"page":"939","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":64,"title":["A Multi-Level Attention Model for Remote Sensing Image Captions"],"prefix":"10.3390","volume":"12","author":[{"given":"Yangyang","family":"Li","sequence":"first","affiliation":[{"name":"Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education, International Research Center for Intelligent Perception and Computation, Joint International Research Laboratory of Intelligent Perception and Computation, School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuangkang","family":"Fang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education, International Research Center for Intelligent Perception and Computation, Joint International Research Laboratory of Intelligent Perception and Computation, School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Licheng","family":"Jiao","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education, International Research Center for Intelligent Perception and Computation, Joint International Research Laboratory of Intelligent Perception and Computation, School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ruijiao","family":"Liu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education, International Research Center for Intelligent Perception and Computation, Joint International Research Laboratory of Intelligent Perception and Computation, School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ronghua","family":"Shang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education, International Research Center for Intelligent Perception and Computation, Joint International Research Laboratory of Intelligent Perception and Computation, School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2020,3,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Qu, B., Li, X., and Tao, D. (2016, January 6\u20138). Deep semantic understanding of high resolution remote sensing image. Proceedings of the International Conference on Computer, Information and Telecommunication Systems (CITS), Kunming, China.","DOI":"10.1109\/CITS.2016.7546397"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"2183","DOI":"10.1109\/TGRS.2017.2776321","article-title":"Exploring Models and Data for Remote Sensing Image Caption Generation","volume":"56","author":"Lu","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1349","DOI":"10.1109\/TGRS.2015.2478379","article-title":"Unsupervised Deep Feature Extraction for Remote Sensing Image Classification","volume":"54","author":"Romero","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"5148","DOI":"10.1109\/TGRS.2017.2702596","article-title":"Remote Sensing Scene Classification by Unsupervised Representation Learning","volume":"55","author":"Lu","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2811","DOI":"10.1109\/TGRS.2017.2748120","article-title":"When Deep Learning Meets Metric Learning: Remote Sensing Image Scene Classification via Learning Discriminative CNNs","volume":"56","author":"Gong","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1720","DOI":"10.1109\/LGRS.2015.2421736","article-title":"A New Approach to Segmentation of Multispectral Remote Sensing Images Based on MRF","volume":"12","author":"Baumgartner","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"760","DOI":"10.1109\/36.917889","article-title":"Real-time processing algorithms for target detection and classification in hyperspectral imagery","volume":"39","author":"Chang","year":"2001","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"4955","DOI":"10.1109\/TGRS.2013.2286195","article-title":"Hyperspectral Remote Sensing Image Subpixel Target Detection Based on Supervised Metric Learning","volume":"52","author":"Zhang","year":"2014","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Vinyals, O., Toshev, A., and Bengio, S. (2015, January 7\u201312). Show and tell: A neural image caption generator. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"ref_10","unstructured":"Xu, K., Ba, J., and Kiros, R. (2015, January 6\u201311). Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Lu, J., Xiong, C., and Parikh, D. (2017, January 21\u201326). Knowing When to Look: Adaptive Attention via a Visual Sentinel for Image Captioning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.345"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"3623","DOI":"10.1109\/TGRS.2017.2677464","article-title":"Can a Machine Generate Humanlike Language Descriptions for a Remote Sensing Image?","volume":"55","author":"Shi","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"664","DOI":"10.1109\/TPAMI.2016.2598339","article-title":"Deep Visual-Semantic Alignments for Generating Image Descriptions","volume":"39","author":"Karpathy","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","unstructured":"Wu, Z., and Cohen, R. (2020, March 09). Encode, Review, and Decode: Reviewer Module for Caption Generation. Available online: https:\/\/www.researchgate.net\/publication\/303521432_Encode_Review_and_Decode_Reviewer_Module_for_Caption_Generation."},{"key":"ref_15","unstructured":"Sadeghi, M.A., Sadeghi, M.A., and Sadeghi, M.A. (2010, January 5\u201311). Every picture tells a story: Generating sentences from images. Proceedings of the European Conference on Computer Vision, Crete, Greece."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Gupta, A., and Mannem, P. (2012, January 12\u201315). From Image Annotation to Image Description. Proceedings of the International Conference on Neural Information Processing, Doha, Qatar.","DOI":"10.1007\/978-3-642-34500-5_24"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Kulkarni, G., Premraj, V., and Dhar, S. (2011, January 20\u201325). Baby Talk: Understanding and Generating Simple Image Descriptions. Proceedings of the Computer Vision and Pattern Recognition (CVPR), Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995466"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Donahue, J., Anne Hendricks, L., and Guadarrama, S. (2015, January 7\u201312). Long-term recurrent convolutional networks for visual recognition and description. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298878"},{"key":"ref_19","unstructured":"Mao, J., Xu, W., and Yang, Y. (2020, March 09). Deep Captioning with Multimodal Recurrent Neural Networks (M-Rnn). Available online: https:\/\/arxiv.org\/abs\/1412.6632."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"416","DOI":"10.1016\/j.neucom.2017.07.014","article-title":"A region-based image caption generator with refined descriptions","volume":"272","author":"Kinghorn","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Anderson, P., He, X., and Buehler, C. (2018, January 18\u201323). Bottom-up and top-down attention for image captioning and visual question answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Lu, J., Yang, J., and Batra, D. (2018, January 18\u201323). Neural baby talk. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00754"},{"key":"ref_23","unstructured":"Lillesand, T., Kiefer, R.W., and Chipman, J. (2015). Remote Sensing and Image Interpretation, John Wiley & Sons."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"652","DOI":"10.1109\/TPAMI.2016.2587640","article-title":"Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge","volume":"39","author":"Vinyals","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_25","unstructured":"Simonyan, K., and Zisserman, A. (2020, March 09). Very Deep Convolutional Networks for Large-Scale Image Recognition. Available online: https:\/\/arxiv.org\/abs\/1409.1556."},{"key":"ref_26","unstructured":"He, K., Zhang, X., and Ren, S. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_27","unstructured":"Cho, K., Van Merri\u00ebnboer, B., and Gulcehre, C. (2020, March 09). Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation. Available online: https:\/\/arxiv.org\/abs\/1406.1078."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"291","DOI":"10.1016\/j.neucom.2018.05.080","article-title":"A survey on automatic image caption generation","volume":"311","author":"Bai","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3295748","article-title":"A comprehensive survey of deep learning for image captioning","volume":"51","author":"Hossain","year":"2019","journal-title":"ACM Comput. Surv."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"77","DOI":"10.1016\/j.patrec.2017.09.013","article-title":"A hierarchical and regional deep learning architecture for image description generation","volume":"119","author":"Kinghorn","year":"2019","journal-title":"Pattern Recognit. Lett."},{"key":"ref_32","unstructured":"Rumelhart, D.E., Hinton, G.E., and Williams, R.J. (2020, March 09). Learning Internal Representations by Error Propagation. Available online: https:\/\/web.stanford.edu\/class\/psych209a\/ReadingsByDate\/02_06\/PDPVolIChapter8.pdf."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., and Ward, T. (2002, January 7\u201312). BLEU: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, Philadelphia, PA, USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Denkowski, M., and Lavie, A. (2014, January 26\u201327). Meteor universal: Language specific translation evaluation for any target language. Proceedings of the Ninth Workshop on Statistical Machine Translation, Baltimore, MD, USA.","DOI":"10.3115\/v1\/W14-3348"},{"key":"ref_35","unstructured":"Lin C, Y. (2004, January 25\u201326). ROUGE: A Package for Automatic Evaluation of summaries. Proceedings of the Workshop on Text Summarization Branches Out (WAS 2004), Barcelona, Spain."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Lawrence Zitnick, C., and Parikh, D. (2015, January 7\u201312). Cider: Consensus-based image description evaluation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Anderson, P., Fernando, B., and Johnson, M. (2016, January 8\u201316). Spice: Semantic propositional image caption evaluation. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46454-1_24"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/6\/939\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:06:49Z","timestamp":1760173609000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/6\/939"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,3,13]]},"references-count":37,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2020,3]]}},"alternative-id":["rs12060939"],"URL":"https:\/\/doi.org\/10.3390\/rs12060939","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,3,13]]}}}