{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,28]],"date-time":"2026-07-28T04:24:53Z","timestamp":1785212693431,"version":"3.55.0"},"reference-count":44,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2019,3,13]],"date-time":"2019-03-13T00:00:00Z","timestamp":1552435200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Image captioning generates a semantic description of an image. It deals with image understanding and text mining, which has made great progress in recent years. However, it is still a great challenge to bridge the \u201csemantic gap\u201d between low-level features and high-level semantics in remote sensing images, in spite of the improvement of image resolutions. In this paper, we present a new model with an attribute attention mechanism for the description generation of remote sensing images. Therefore, we have explored the impact of the attributes extracted from remote sensing images on the attention mechanism. The results of our experiments demonstrate the validity of our proposed model. The proposed method obtains six higher scores and one slightly lower, compared against several state of the art techniques, on the Sydney Dataset and Remote Sensing Image Caption Dataset (RSICD), and receives all seven higher scores on the UCM Dataset for remote sensing image captioning, indicating that the proposed framework achieves robust performance for semantic description in high-resolution remote sensing images.<\/jats:p>","DOI":"10.3390\/rs11060612","type":"journal-article","created":{"date-parts":[[2019,3,14]],"date-time":"2019-03-14T04:15:29Z","timestamp":1552536929000},"page":"612","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":140,"title":["Description Generation for Remote Sensing Images Using Attribute Attention Mechanism"],"prefix":"10.3390","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0379-2042","authenticated-orcid":false,"given":"Xiangrong","family":"Zhang","sequence":"first","affiliation":[{"name":"Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education, International Research Center for Intelligent Perception and Computation, Joint International Research Laboratory of Intelligent Perception and Computation, School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xin","family":"Wang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education, International Research Center for Intelligent Perception and Computation, Joint International Research Laboratory of Intelligent Perception and Computation, School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1375-0778","authenticated-orcid":false,"given":"Xu","family":"Tang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education, International Research Center for Intelligent Perception and Computation, Joint International Research Laboratory of Intelligent Perception and Computation, School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huiyu","family":"Zhou","sequence":"additional","affiliation":[{"name":"Department of Informatics, University of Leicester, Leicester LE1 7RH, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chen","family":"Li","sequence":"additional","affiliation":[{"name":"Computer Science Department, Xi\u2019an Jiaotong University, Xi\u2019an 710049, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2019,3,13]]},"reference":[{"key":"ref_1","first-page":"22","article-title":"A survey, journal of photogrammetry and remote sensing","volume":"115","author":"Toth","year":"2016","journal-title":"Remote Sens. Platforms Sens."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Rahaman, K.R., and Hassan, Q.K. (2016, January 13\u201314). Application of remote sensing to quantify local warming trends: A review. Proceedings of the 5th International Conference on Informatics, Electronics and Vision (ICIEV), Dhaka, Bangladesh.","DOI":"10.1109\/ICIEV.2016.7760006"},{"key":"ref_3","unstructured":"Aroma, R.J., and Raimond, K. (2015, January 10\u201312). A review on availability of remote sensing data. Proceedings of the IEEE Technological Innovation in ICT for Agriculture and Rural Development (TIAR), Chennai, India."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"645","DOI":"10.1109\/TGRS.2016.2612821","article-title":"Convolutional neural networks for large-scale remote-sensing image classification","volume":"55","author":"Maggiori","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2811","DOI":"10.1109\/TGRS.2017.2783902","article-title":"When deep learning meets metric learning: Remote sensing image scene classification via learning discriminative CNNs","volume":"56","author":"Cheng","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_6","first-page":"1","article-title":"Remote sensing scene classification by unsupervised representation learning","volume":"99","author":"Lu","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Han, X., Zhong, Y., and Zhang, L. (2017). An efficient and robust integrated geospatial object detection framework for high spatial resolution remote sensing imagery. Remote Sens., 9.","DOI":"10.3390\/rs9070666"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1511","DOI":"10.1109\/JSTARS.2016.2620900","article-title":"Airport detection and aircraft recognition based on two-layer saliency model in high spatial resolution remote-sensing images","volume":"10","author":"Zhang","year":"2017","journal-title":"IEEE J. Sel. Top. Appl. Earth Observ. Remote Sens."},{"key":"ref_9","first-page":"684","article-title":"A remote sensing image retrieval model based on semantic mining","volume":"34","author":"Liu","year":"2009","journal-title":"Geomatics Inf. Sci. Wuhan Univ."},{"key":"ref_10","unstructured":"Zhu, Q., Zhong, Y., and Zhang, L. (2014, January 13\u201318). Multi-feature probability topic scene classifier for high spatial resolution remote sensing imagery. Proceedings of the 2014 IEEE Geoscience and Remote Sensing Symposium, Quebec, QC, Canada."},{"key":"ref_11","first-page":"3069","article-title":"Remote sensing image semantic labeling based on conditional random field. Acta Aeronaut","volume":"36","author":"Yang","year":"2015","journal-title":"Et Astronaut. Sin."},{"key":"ref_12","first-page":"48","article-title":"Research on key technologies of remote sensing image data retrieval based on semantics","volume":"40","author":"Wang","year":"2012","journal-title":"Comput. Digit. Eng."},{"key":"ref_13","unstructured":"Chen, K.M., Zhou, Z.X., Guo, J.E., Zhang, D.B., and Sun, X. (2013, January 1). Semantic scene understanding oriented high resolution remote sensing image change information analysis. Proceedings of the Annual Conference on High Resolution Earth Observation, Beijing, China."},{"key":"ref_14","unstructured":"Li, Y. (2012). Target Detection Method of High Resolution Remote Sensing Image Based on Semantic Model, Graduate University of Chinese Academy of Sciences."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"652","DOI":"10.1109\/TPAMI.2016.2587640","article-title":"Show and tell: Lessons learned from the mscoco image captioning challenge","volume":"39","author":"Vinyals","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_16","unstructured":"Mao, J., Xu, W., Yang, Y., Wang, J., Huang, Z., and Yuille, A. (arXiv, 2014). Deep caption with multimodal recurrent neural networks (m-rnn), arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Chen, X., and Lawrence Zitnick, C. (2015, January 7\u201312). Mind\u2019s eye: A recurrent visual representation for image captioning generation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298856"},{"key":"ref_18","unstructured":"Karpathy, A., Joulin, A., and Fei-Fei, L.F. (2014, January 8\u201313). Deep fragment embeddings for bidirectional image sentence mapping. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Qu, B., Li, X., Tao, D., and Lu, X. (2016, January 6\u20138). Deep semantic understanding of high resolution remote sensing image. Proceedings of the 2016 International Conference on Computer, Information and Telecommunication Systems (CITS), Kunming, China.","DOI":"10.1109\/CITS.2016.7546397"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"3623","DOI":"10.1109\/TGRS.2017.2677464","article-title":"Can a machine generate humanlike language descriptions for a remote sensing image?","volume":"55","author":"Shi","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2183","DOI":"10.1109\/TGRS.2017.2776321","article-title":"Exploring models and data for remote sensing image caption generation","volume":"56","author":"Lu","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"219","DOI":"10.1162\/089892904322984526","article-title":"A feedback model of visual attention","volume":"16","author":"Spratling","year":"2004","journal-title":"J. Cogn. Neurosci."},{"key":"ref_23","unstructured":"Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y. (2015, January 6\u201311). Show, attend and tell: Neural image captioning generation with visual attention. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Lu, J., Xiong, C., Parikh, D., and Socher, R. (2017, January 21\u201326). Knowing when to look: Adaptive attention via a visual sentinel for image captioning. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.345"},{"key":"ref_25","unstructured":"You, Q., Jin, H., Wang, Z., Fang, C., and Luo, J. (July, January 26). Image captioning with semantic attention. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Yao, T., Pan, Y., Li, Y., Qiu, Z., and Mei, T. (2017, January 22\u201329). Boosting image captioning with attributes. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.524"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1367","DOI":"10.1109\/TPAMI.2017.2708709","article-title":"Image captioning and visual question answering based on attributes and external knowledge","volume":"40","author":"Wu","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_28","unstructured":"Hossain, M.D., Sohel, F., Shiratuddin, M.F., and Laga, H. (2019, March 13). A Comprehensive Survey of Deep Learning for Image Captioning. Available online: https:\/\/arxiv.org\/pdf\/1810.04020.pdf."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Farhadi, A., Hejrati, M., Sadeghi, M.A., Young, P., Rashtchian, C., Hockenmaier, J., and Forsyth, D. (2010, January 5\u201311). Every picture tells a story: Generating sentences from images. Proceedings of the 11th European Conference on Computer Vision: Part IV, ECCV\u201910, Crete, Greece.","DOI":"10.1007\/978-3-642-15561-1_2"},{"key":"ref_30","unstructured":"Li, S., Kulkarni, G., Berg, T.L., Berg, A.C., and Choi, Y. (2011, January 23\u201324). Composing simple image descriptions using web-scale n-grams. Proceedings of the Fifteenth Conference on Computational Natural Language Learning, Portland, OR, USA."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Kulkarni, G., Premraj, V., Dhar, S., Li, S., Choi, Y., Berg, A.C., and Berg, T.L. (2011, January 20\u201325). Baby talk: Understanding and generating image descriptions. Proceedings of the 24th CVPR, Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995466"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Gong, Y., Wang, L., Hodosh, M., Hockenmaier, J., and Lazebnik, S. (2014). Improving image-sentence embeddings using large weakly annotated photo collections. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10593-2_35"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Sun, C., Gan, C., and Nevatia, R. (2015, January 13\u201316). Automatic concept discovery from parallel text and visual corpora. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.298"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"853","DOI":"10.1613\/jair.3994","article-title":"Framing image description as a ranking task: Data, models and evaluation metrics","volume":"47","author":"Hodosh","year":"2013","journal-title":"J. Artif. Intell. Res."},{"key":"ref_35","unstructured":"Wu, Q., Shen, C., Liu, L., Dick, A., and Van Den Hengel, A. (July, January 26). What value do explicit high level concepts have in vision to language problems?. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, America."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Yang, Y., and Newsam, S. (2010, January 2\u20135). Bag-of-visual-words and spatial extensions for land-use classification. Proceedings of the 18th SIGSPATIAL International Conference on Advances in Geographic Information Systems, San Jose, CA, USA.","DOI":"10.1145\/1869790.1869829"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"2175","DOI":"10.1109\/TGRS.2014.2357078","article-title":"Saliency-guided unsupervised feature learning for scene classification","volume":"53","author":"Zhang","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Wang, B., Lu, X., Zheng, X., and Li, X. Semantic Descriptions of High-Resolution Remote Sensing Images. IEEE Geosci. Remote Sens. Lett., 2019.","DOI":"10.1109\/LGRS.2019.2893772"},{"key":"ref_39","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (arXiv, 2014). Neural machine translation by jointly learning to align and translate, arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2002, January 7\u201312). BLEU: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, Philadelphia, PA, USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_41","unstructured":"Lin, C.Y. (2004, January 25\u201326). Rouge: A package for automatic evaluation of summaries. Proceedings of the Text Summarization Branches Out, Barcelona, Spain."},{"key":"ref_42","unstructured":"Banerjee, S., and Lavie, A. (2005, January 11). METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization, Ann Arbor, MI, USA."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Lawrence Zitnick, C., and Parikh, D. (2015, January 7\u201312). Cider: Consensus-based image description evaluation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"3965","DOI":"10.1109\/TGRS.2017.2685945","article-title":"AID: A benchmark data set for performance evaluation of aerial scene classification","volume":"55","author":"Xia","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/11\/6\/612\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:38:28Z","timestamp":1760186308000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/11\/6\/612"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,3,13]]},"references-count":44,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2019,3]]}},"alternative-id":["rs11060612"],"URL":"https:\/\/doi.org\/10.3390\/rs11060612","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,3,13]]}}}