{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T23:58:26Z","timestamp":1780444706162,"version":"3.54.1"},"reference-count":53,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2022,6,20]],"date-time":"2022-06-20T00:00:00Z","timestamp":1655683200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62102296"],"award-info":[{"award-number":["62102296"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2022JQ-661"],"award-info":[{"award-number":["2022JQ-661"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["XJS222215"],"award-info":[{"award-number":["XJS222215"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["XJS222221"],"award-info":[{"award-number":["XJS222221"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"the Natural Science Foundation of Shaanxi Province","award":["62102296"],"award-info":[{"award-number":["62102296"]}]},{"name":"the Natural Science Foundation of Shaanxi Province","award":["2022JQ-661"],"award-info":[{"award-number":["2022JQ-661"]}]},{"name":"the Natural Science Foundation of Shaanxi Province","award":["XJS222215"],"award-info":[{"award-number":["XJS222215"]}]},{"name":"the Natural Science Foundation of Shaanxi Province","award":["XJS222221"],"award-info":[{"award-number":["XJS222221"]}]},{"name":"the Fundamental Research Funds for the Central Universities","award":["62102296"],"award-info":[{"award-number":["62102296"]}]},{"name":"the Fundamental Research Funds for the Central Universities","award":["2022JQ-661"],"award-info":[{"award-number":["2022JQ-661"]}]},{"name":"the Fundamental Research Funds for the Central Universities","award":["XJS222215"],"award-info":[{"award-number":["XJS222215"]}]},{"name":"the Fundamental Research Funds for the Central Universities","award":["XJS222221"],"award-info":[{"award-number":["XJS222221"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Remote sensing image captioning aims to describe the content of images using natural language. In contrast with natural images, the scale, distribution, and number of objects generally vary in remote sensing images, making it hard to capture global semantic information and the relationships between objects at different scales. In this paper, in order to improve the accuracy and diversity of captioning, a mask-guided Transformer network with a topic token is proposed. Multi-head attention is introduced to extract features and capture the relationships between objects. On this basis, a topic token is added into the encoder, which represents the scene topic and serves as a prior in the decoder to help us focus better on global semantic information. Moreover, a new Mask-Cross-Entropy strategy is designed in order to improve the diversity of the generated captions, which randomly replaces some input words with a special word (named [Mask]) in the training stage, with the aim of enhancing the model\u2019s learning ability and forcing exploration of uncommon word relations. Experiments on three data sets show that the proposed method can generate captions with high accuracy and diversity, and the experimental results illustrate that the proposed method can outperform state-of-the-art models. Furthermore, the CIDEr score on the RSICD data set increased from 275.49 to 298.39.<\/jats:p>","DOI":"10.3390\/rs14122939","type":"journal-article","created":{"date-parts":[[2022,6,21]],"date-time":"2022-06-21T04:39:55Z","timestamp":1655786395000},"page":"2939","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":34,"title":["A Mask-Guided Transformer Network with Topic Token for Remote Sensing Image Captioning"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1520-0737","authenticated-orcid":false,"given":"Zihao","family":"Ren","sequence":"first","affiliation":[{"name":"Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education, School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2619-6481","authenticated-orcid":false,"given":"Shuiping","family":"Gou","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education, School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7586-4020","authenticated-orcid":false,"given":"Zhang","family":"Guo","sequence":"additional","affiliation":[{"name":"Academy of Advanced Interdisciplinary Research, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3308-1794","authenticated-orcid":false,"given":"Shasha","family":"Mao","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education, School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2393-2225","authenticated-orcid":false,"given":"Ruimin","family":"Li","sequence":"additional","affiliation":[{"name":"Academy of Advanced Interdisciplinary Research, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,6,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"4955","DOI":"10.1109\/TGRS.2013.2286195","article-title":"Hyperspectral remote sensing image subpixel target detection based on supervised metric learning","volume":"52","author":"Zhang","year":"2013","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Xu, Y., Du, B., and Zhang, L. (2018, January 22\u201327). Multi-source remote sensing data classification via fully convolutional networks and post-classification processing. Proceedings of the IGARSS 2018-2018 IEEE International Geoscience and Remote Sensing Symposium, Valencia, Spain.","DOI":"10.1109\/IGARSS.2018.8518295"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"2811","DOI":"10.1109\/TGRS.2017.2783902","article-title":"When deep learning meets metric learning: Remote sensing image scene classification via learning discriminative CNNs","volume":"56","author":"Cheng","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1720","DOI":"10.1109\/LGRS.2015.2421736","article-title":"A new approach to segmentation of multispectral remote sensing images based on mrf","volume":"12","author":"Baumgartner","year":"2015","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"4376","DOI":"10.1109\/TIP.2019.2910667","article-title":"Weakly supervised adversarial domain adaptation for semantic segmentation in urban scenes","volume":"28","author":"Wang","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Qu, B., Li, X., Tao, D., and Lu, X. (2016, January 16\u201318). Deep semantic understanding of high resolution remote sensing image. Proceedings of the 2016 International Conference on Computer, Information and Telecommunication Systems (Cits), Istanbul, Turkey.","DOI":"10.1109\/CITS.2016.7546397"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"3623","DOI":"10.1109\/TGRS.2017.2677464","article-title":"Can a machine generate humanlike language descriptions for a remote sensing image?","volume":"55","author":"Shi","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"355","DOI":"10.1109\/TIP.2016.2627801","article-title":"Latent semantic minimal hashing for image retrieval","volume":"26","author":"Lu","year":"2016","journal-title":"IEEE Trans. Image Process."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"566","DOI":"10.1016\/j.procs.2016.07.144","article-title":"Geological disaster recognition on optical remote sensing images using deep learning","volume":"91","author":"Liu","year":"2016","journal-title":"Procedia Comput. Sci."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"46","DOI":"10.1007\/s11263-015-0840-y","article-title":"Large scale retrieval and generation of image descriptions","volume":"119","author":"Ordonez","year":"2016","journal-title":"Int. J. Comput. Vis."},{"key":"ref_11","unstructured":"Kuznetsova, P., Ordonez, V., Berg, A., Berg, T., and Choi, Y. (2012, January 8\u201314). Collective generation of natural image descriptions. Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics, Jeju Island, Korea."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"2891","DOI":"10.1109\/TPAMI.2012.162","article-title":"Babytalk: Understanding and generating simple image descriptions","volume":"35","author":"Kulkarni","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Gupta, A., and Mannem, P. (2012, January 12\u201315). From image annotation to image description. Proceedings of the International Conference on Neural Information Processing, Doha, Qatar.","DOI":"10.1007\/978-3-642-34500-5_24"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2015, January 7\u201312). Show and tell: A neural image caption generator. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"ref_15","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015, January 7\u201312). Faster r-cnn: Towards real-time object detection with region proposal networks. Proceedings of the Advance in Neural Information Processing Systems (NIPS), Montreal, QC, USA."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_17","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Rennie, S.J., Marcheret, E., Mroueh, Y., Ross, J., and Goel, V. (2017, January 21\u201326). Self-critical sequence training for image captioning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.131"},{"key":"ref_19","unstructured":"Li, Y. (2017). Deep reinforcement learning: An overview. arXiv."},{"key":"ref_20","unstructured":"Ranzato, M., Chopra, S., Auli, M., and Zaremba, W. (2015). Sequence level training with recurrent neural networks. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L. (2018, January 18\u201323). Bottom-up and top-down attention for image captioning and visual question answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"ref_23","unstructured":"Huang, L., Wang, W., Chen, J., and Wei, X.Y. (November, January 27). Attention on attention for image captioning. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_24","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. In Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Cornia, M., Stefanini, M., Baraldi, L., and Cucchiara, R. (2020, January 13\u201319). Meshed-memory transformer for image captioning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01059"},{"key":"ref_26","unstructured":"Herdade, S., Kappeler, A., Boakye, K., and Soares, J. (2019, January 8\u201314). Image captioning: Transforming objects into words. Proceedings of the Conference on Neural Information Processing Systems (NIPS 2019), Vancouver, BC, Canada."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"2183","DOI":"10.1109\/TGRS.2017.2776321","article-title":"Exploring models and data for remote sensing image caption generation","volume":"56","author":"Lu","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2020.3042202","article-title":"High-resolution remote sensing image captioning based on structured attention","volume":"60","author":"Zhao","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1109\/LGRS.2020.2980933","article-title":"Denoising-based multiscale feature fusion for remote sensing image captioning","volume":"18","author":"Huang","year":"2020","journal-title":"IEEE Geosci. Remote. Sens. Lett."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Li, Y., Fang, S., Jiao, L., Liu, R., and Shang, R. (2020). A multi-level attention model for remote sensing image captions. Remote Sens., 12.","DOI":"10.3390\/rs12060939"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"105920","DOI":"10.1016\/j.knosys.2020.105920","article-title":"Remote sensing image captioning via Variational Autoencoder and Reinforcement Learning","volume":"203","author":"Shen","year":"2020","journal-title":"Knowl.-Based Syst."},{"key":"ref_32","first-page":"1","article-title":"Recurrent Attention and Semantic Gate for Remote Sensing Image Captioning","volume":"60","author":"Li","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2002, January 7\u201312). Bleu: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Philadelphia, PA, USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Lawrence Zitnick, C., and Parikh, D. (2015, January 7\u201312). Cider: Consensus-based image description evaluation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Dai, B., Fidler, S., Urtasun, R., and Lin, D. (2017, January 22\u201329). Towards diverse and natural image descriptions via a conditional gan. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.323"},{"key":"ref_36","unstructured":"Wang, L., Schwing, A., and Lazebnik, S. (2017, January 4\u20139). Diverse and accurate image description using a variational auto-encoder with an additive gaussian encoding space. Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_37","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014, January 8\u201313). Generative adversarial nets. Proceedings of the Conference on Neural Information Processing Systems (NIPS 2014), Montreal, QC, Canada."},{"key":"ref_38","unstructured":"Kingma, D.P., and Welling, M. (2013). Auto-encoding variational bayes. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Chen, J., and Jin, Q. (2020, January 13\u201319). Better captioning with sequence-level exploration. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01090"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Wang, Q., and Chan, A.B. (2019, January 15\u201320). Describing like humans: On diversity in image captioning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00432"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Shi, J., Li, Y., and Wang, S. (2021, January 10\u201317). Partial Off-Policy Learning: Balance Accuracy and Diversity for Human-Oriented Image Captioning. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00219"},{"key":"ref_42","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv."},{"key":"ref_43","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"26","DOI":"10.9781\/ijimai.2016.415","article-title":"Multilayer perceptron: Architecture optimization and training","volume":"4","author":"Ramchoun","year":"2016","journal-title":"IJIMAI"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"He, K., Chen, X., Xie, S., Li, Y., Doll\u00e1r, P., and Girshick, R. (2021). Masked autoencoders are scalable vision learners. arXiv.","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"ref_46","unstructured":"Bao, H., Dong, L., and Wei, F. (2021). Beit: Bert pre-training of image transformers. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Yang, Y., and Newsam, S. (2010, January 2\u20135). Bag-of-visual-words and spatial extensions for land-use classification. Proceedings of the 18th SIGSPATIAL International Conference on Advances in Geographic Information Systems, San Jose, CA, USA.","DOI":"10.1145\/1869790.1869829"},{"key":"ref_48","unstructured":"Lin, C.Y. (2004, January 25\u201326). Rouge: A package for automatic evaluation of summaries. Proceedings of the Text Summarization Branches Out, Barcelona, Spain."},{"key":"ref_49","unstructured":"Banerjee, S., and Lavie, A. (2005, January 29). METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization, Ann Arbor, MI, USA."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Graves, A. (2012). Sequence transduction with recurrent neural networks. arXiv.","DOI":"10.1007\/978-3-642-24797-2"},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"154","DOI":"10.1007\/s11263-013-0620-5","article-title":"Selective search for object recognition","volume":"104","author":"Uijlings","year":"2013","journal-title":"Int. J. Comput. Vis."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Zhang, X., Wang, X., Tang, X., Zhou, H., and Li, C. (2019). Description generation for remote sensing images using attribute attention mechanism. Remote Sens., 11.","DOI":"10.3390\/rs11060612"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/14\/12\/2939\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:35:24Z","timestamp":1760139324000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/14\/12\/2939"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,20]]},"references-count":53,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2022,6]]}},"alternative-id":["rs14122939"],"URL":"https:\/\/doi.org\/10.3390\/rs14122939","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,20]]}}}