{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,16]],"date-time":"2026-05-16T16:11:04Z","timestamp":1778947864616,"version":"3.51.4"},"reference-count":36,"publisher":"MDPI AG","issue":"20","license":[{"start":{"date-parts":[[2019,10,10]],"date-time":"2019-10-10T00:00:00Z","timestamp":1570665600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["41701508"],"award-info":[{"award-number":["41701508"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Significant progress has been made in remote sensing image captioning by encoder-decoder frameworks. The conventional attention mechanism is prevalent in this task but still has some drawbacks. The conventional attention mechanism only uses visual information about the remote sensing images without considering using the label information to guide the calculation of attention masks. To this end, a novel attention mechanism, namely Label-Attention Mechanism (LAM), is proposed in this paper. LAM additionally utilizes the label information of high-resolution remote sensing images to generate natural sentences to describe the given images. It is worth noting that, instead of high-level image features, the predicted categories\u2019 word embedding vectors are adopted to guide the calculation of attention masks. Representing the content of images in the form of word embedding vectors can filter out redundant image features. In addition, it can also preserve pure and useful information for generating complete sentences. The experimental results from UCM-Captions, Sydney-Captions and RSICD demonstrate that LAM can improve the model\u2019s performance for describing high-resolution remote sensing images and obtain better     S m     scores compared with other methods.     S m     score is a hybrid scoring method derived from the AI Challenge 2017 scoring method. In addition, the validity of LAM is verified by the experiment of using true labels.<\/jats:p>","DOI":"10.3390\/rs11202349","type":"journal-article","created":{"date-parts":[[2019,10,11]],"date-time":"2019-10-11T03:07:11Z","timestamp":1570763231000},"page":"2349","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":56,"title":["LAM: Remote Sensing Image Captioning with Label-Attention Mechanism"],"prefix":"10.3390","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2383-7738","authenticated-orcid":false,"given":"Zhengyuan","family":"Zhang","sequence":"first","affiliation":[{"name":"Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"Key Laboratory of Network Information System Technology (NIST), Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenhui","family":"Diao","sequence":"additional","affiliation":[{"name":"Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"Key Laboratory of Network Information System Technology (NIST), Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenkai","family":"Zhang","sequence":"additional","affiliation":[{"name":"Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"Key Laboratory of Network Information System Technology (NIST), Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Menglong","family":"Yan","sequence":"additional","affiliation":[{"name":"Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"Key Laboratory of Network Information System Technology (NIST), Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xin","family":"Gao","sequence":"additional","affiliation":[{"name":"Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"Key Laboratory of Network Information System Technology (NIST), Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xian","family":"Sun","sequence":"additional","affiliation":[{"name":"Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"Key Laboratory of Network Information System Technology (NIST), Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,10,10]]},"reference":[{"key":"ref_1","first-page":"45","article-title":"A Fast Target Detection Algorithm for High Resolution SAR Imagery","volume":"9","author":"Zhang","year":"2005","journal-title":"J. Remote Sens."},{"key":"ref_2","first-page":"195","article-title":"An Aircraft Detection Method Based on Convolutional Neural Networks in High-Resolution SAR Images","volume":"6","author":"Wang","year":"2017","journal-title":"J. Radars"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1600","DOI":"10.1109\/LGRS.2018.2846802","article-title":"Cloud and cloud shadow detection using multilevel feature fused segmentation network","volume":"15","author":"Yan","year":"2018","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"39401","DOI":"10.1109\/ACCESS.2018.2856088","article-title":"An End-to-End Neural Network for Road Extraction From Remote Sensing Imagery by Multiple Feature Pyramid Network","volume":"6","author":"Gao","year":"2018","journal-title":"IEEE Access"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2183","DOI":"10.1109\/TGRS.2017.2776321","article-title":"Exploring Models and Data for Remote Sensing Image Caption Generation","volume":"56","author":"Lu","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_6","unstructured":"Ordonez, V., Kulkarni, G., and Berg, T.L. (2011, January 12\u201315). Im2Text: Describing Images Using 1 Million Captioned Photographs. Proceedings of the 24th International Conference on Neural Information Processing Systems, Granada, Spain."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"853","DOI":"10.1613\/jair.3994","article-title":"Framing image description as a ranking task: Data, models and evaluation metrics","volume":"47","author":"Hodosh","year":"2013","journal-title":"J. Artif. Intell. Res."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Sun, C., Gan, C., and Nevatia, R. (2015, January 7\u201313). Automatic Concept Discovery from Parallel Text and Visual Corpora. Proceedings of the 2015 IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.298"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Gong, Y., Wang, L., Hodosh, M., Hockenmaier, J., and Lazebnik, S. (2014, January 6\u201312). Improving Image-Sentence Embeddings Using Large Weakly Annotated Photo Collections. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10593-2_35"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"46","DOI":"10.1007\/s11263-015-0840-y","article-title":"Large Scale Retrieval and Generation of Image Descriptions","volume":"119","author":"Ordonez","year":"2016","journal-title":"Int. J. Comput. Vis."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Farhadi, A., Hejrati, M., Sadeghi, M.A., Young, P., Rashtchian, C., Hockenmaier, J., and Forsyth, D.A. (2010, January 5\u201311). Every picture tells a story: Generating sentences from images. Proceedings of the 11th European Conference on Computer Vision, Heraklion, Greece.","DOI":"10.1007\/978-3-642-15561-1_2"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Kulkarni, G., Premraj, V., Dhar, S., Li, S., Choi, Y., Berg, A.C., and Berg, T.L. (2011, January 20\u201325). Baby talk: Understanding and generating simple image descriptions. Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition, Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995466"},{"key":"ref_13","unstructured":"Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A.C., Salakhudinov, R., Zemel, R., and Bengio, Y. (2015, January 6\u201311). Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. Proceedings of the 32nd International Conference on Machine Learning, Lille, France."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Lu, J., Xiong, C., Parikh, D., and Socher, R. (2017, January 21\u201326). Knowing When to Look: Adaptive Attention via a Visual Sentinel for Image Captioning. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.345"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Chen, S., and Zhao, Q. (2018, January 8\u201314). Boosted Attention: Leveraging Human Attention for Image Captioning. Proceedings of the 15th European Conference on Computer Vision\u2014ECCV2018, Munich, Germany.","DOI":"10.1007\/978-3-030-01252-6_5"},{"key":"ref_16","unstructured":"Mao, J., Xu, W., Yang, Y., Wang, J., and Yuille, A. (2014). Explain Images with Multimodal Recurrent Neural Networks. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Aneja, J., Deshpande, A., and Schwing, A.G. (2017). Convolutional Image Captioning. arXiv.","DOI":"10.1109\/CVPR.2018.00583"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zhang, X., Li, X., An, J., Gao, L., Hou, B., and Li, C. (2017, January 23\u201328). Natural language description of remote sensing images based on deep learning. Proceedings of the 2017 IEEE International Geoscience and Remote Sensing Symposium, IGARSS 2017, Fort Worth, TX, USA.","DOI":"10.1109\/IGARSS.2017.8128075"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhang, X., Wang, X., Tang, X., Zhou, H., and Li, C. (2019). Description Generation for Remote Sensing Images Using Attribute Attention Mechanism. Remote Sens., 11.","DOI":"10.3390\/rs11060612"},{"key":"ref_20","unstructured":"Mao, J., Xu, W., Yang, Y., Wang, J., Huang, Z., and Yuille, A. (2015). Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN). arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Karpathy, A., and Feifei, L. (2015, January 7\u201312). Deep visual-semantic alignments for generating image descriptions. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Chen, X., and Zitnick, C. (2015, January 7\u201312). Learning a Recurrent Visual Representation for Image Caption Generation. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298856"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Fang, H., Gupta, S., Iandola, F.N., Srivastava, R.K., Deng, L., Dollar, P., Gao, J., He, X., Mitchell, M., and Platt, J. (2015, January 7\u201312). From captions to visual concepts and back. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298754"},{"key":"ref_24","unstructured":"Karpathy, A., Joulin, A., and Li, F. (2014, January 8\u201313). Deep Fragment Embeddings for Bidirectional Image Sentence Mapping. Proceedings of the 27th International Conference on Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_25","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G. (2012, January 3\u20136). ImageNet Classification with Deep Convolutional Neural Networks. Proceedings of the 25th International Conference on Neural Information Processing Systems, Lake Tahoe, NV, USA."},{"key":"ref_26","unstructured":"Simonyan, K., and Zisserman, A. (2015). Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S.E., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Madison, WI, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2015, January 7\u201312). Show and tell: A neural image caption generator. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Qu, B., Li, X., Tao, D., and Lu, X. (2016, January 6\u20138). Deep semantic understanding of high resolution remote sensing image. Proceedings of the 2016 International Conference on Computer, Information and Telecommunication Systems, Kunming, China.","DOI":"10.1109\/CITS.2016.7546397"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Yang, Y., and Newsam, S. (2010, January 2\u20135). Bag-of-visual-words and spatial extensions for land-use classification. Proceedings of the 18th SIGSPATIAL International Conference on Advances in Geographic Information Systems, San Jose, CA, USA.","DOI":"10.1145\/1869790.1869829"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"2175","DOI":"10.1109\/TGRS.2014.2357078","article-title":"Saliency-Guided Unsupervised Feature Learning for Scene Classification","volume":"53","author":"Zhang","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Li, L., Tang, S., Deng, L., Zhang, Y., and Tian, Q. (2017, January 4\u20139). Image Caption with Global-Local Attention. Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.11236"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Yao, T., Pan, Y., Li, Y., and Mei, T. (2017, January 21\u201326). Incorporating Copying Mechanism in Image Captioning for Learning Novel Objects. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2017.559"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wang, Y., Lin, Z., Shen, X., Cohen, S., and Cottrell, G.W. (2017, January 21\u201326). Skeleton Key: Image Captioning by Skeleton-Attribute Decomposition. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2017.780"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"ImageNet Large Scale Visual Recognition Challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/11\/20\/2349\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:29:07Z","timestamp":1760189347000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/11\/20\/2349"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,10,10]]},"references-count":36,"journal-issue":{"issue":"20","published-online":{"date-parts":[[2019,10]]}},"alternative-id":["rs11202349"],"URL":"https:\/\/doi.org\/10.3390\/rs11202349","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,10,10]]}}}