{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T19:05:53Z","timestamp":1782587153929,"version":"3.54.5"},"reference-count":63,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2019,5,2]],"date-time":"2019-05-02T00:00:00Z","timestamp":1556755200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key R&amp;D Program of China","award":["2018YFC0810600c"],"award-info":[{"award-number":["2018YFC0810600c"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>A comprehensive interpretation of remote sensing images involves not only remote sensing object recognition but also the recognition of spatial relations between objects. Especially in the case of different objects with the same spectrum, the spatial relationship can help interpret remote sensing objects more accurately. Compared with traditional remote sensing object recognition methods, deep learning has the advantages of high accuracy and strong generalizability regarding scene classification and semantic segmentation. However, it is difficult to simultaneously recognize remote sensing objects and their spatial relationship from end-to-end only relying on present deep learning networks. To address this problem, we propose a multi-scale remote sensing image interpretation network, called the MSRIN. The architecture of the MSRIN is a parallel deep neural network based on a fully convolutional network (FCN), a U-Net, and a long short-term memory network (LSTM). The MSRIN recognizes remote sensing objects and their spatial relationship through three processes. First, the MSRIN defines a multi-scale remote sensing image caption strategy and simultaneously segments the same image using the FCN and U-Net on different spatial scales so that a two-scale hierarchy is formed. The output of the FCN and U-Net are masked to obtain the location and boundaries of remote sensing objects. Second, using an attention-based LSTM, the remote sensing image captions include the remote sensing objects (nouns) and their spatial relationships described with natural language. Finally, we designed a remote sensing object recognition and correction mechanism to build the relationship between nouns in captions and object mask graphs using an attention weight matrix to transfer the spatial relationship from captions to objects mask graphs. In other words, the MSRIN simultaneously realizes the semantic segmentation of the remote sensing objects and their spatial relationship identification end-to-end. Experimental results demonstrated that the matching rate between samples and the mask graph increased by 67.37 percentage points, and the matching rate between nouns and the mask graph increased by 41.78 percentage points compared to before correction. The proposed MSRIN has achieved remarkable results.<\/jats:p>","DOI":"10.3390\/rs11091044","type":"journal-article","created":{"date-parts":[[2019,5,7]],"date-time":"2019-05-07T03:15:46Z","timestamp":1557198946000},"page":"1044","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":61,"title":["Multi-Scale Semantic Segmentation and Spatial Relationship Recognition of Remote Sensing Images Based on an Attention Model"],"prefix":"10.3390","volume":"11","author":[{"given":"Wei","family":"Cui","sequence":"first","affiliation":[{"name":"School of Resources and Environmental Engineering, Wuhan University of Technology, Wuhan 430070, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fei","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Resources and Environmental Engineering, Wuhan University of Technology, Wuhan 430070, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xin","family":"He","sequence":"additional","affiliation":[{"name":"School of Resources and Environmental Engineering, Wuhan University of Technology, Wuhan 430070, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dongyou","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Resources and Environmental Engineering, Wuhan University of Technology, Wuhan 430070, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xuxiang","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Resources and Environmental Engineering, Wuhan University of Technology, Wuhan 430070, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Meng","family":"Yao","sequence":"additional","affiliation":[{"name":"School of Resources and Environmental Engineering, Wuhan University of Technology, Wuhan 430070, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ziwei","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Resources and Environmental Engineering, Wuhan University of Technology, Wuhan 430070, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2595-4946","authenticated-orcid":false,"given":"Jiejun","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Resources and Environmental Engineering, Wuhan University of Technology, Wuhan 430070, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2019,5,2]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"Lecun","year":"2015","journal-title":"Nature"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1016\/j.neunet.2014.09.003","article-title":"Deep learning in neural networks: An overview","volume":"61","author":"Schmidhuber","year":"2015","journal-title":"Neural Netw."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"Lecun","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"4238","DOI":"10.1109\/TGRS.2015.2393857","article-title":"Effective and efficient midlevel visual elements-oriented land-use classification using vhr remote sensing images","volume":"53","author":"Cheng","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"7405","DOI":"10.1109\/TGRS.2016.2601622","article-title":"Learning rotation-invariant convolutional neural networks for object detection in vhr optical remote sensing images","volume":"54","author":"Cheng","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"3325","DOI":"10.1109\/TGRS.2014.2374218","article-title":"Object detection in optical remote sensing images based on weakly supervised learning and high-level feature learning","volume":"53","author":"Han","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"514","DOI":"10.1080\/01431161.2016.1266059","article-title":"Scene classification based on a hierarchical convolutional sparse auto-encoder for high spatial resolution imagery","volume":"38","author":"Han","year":"2017","journal-title":"Int. J. Remote Sens."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"14680","DOI":"10.3390\/rs71114680","article-title":"Transferring deep convolutional neural networks for the scene classification of high-resolution remote sensing imagery","volume":"7","author":"Hu","year":"2015","journal-title":"Remote Sens."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1155\/2015\/258619","article-title":"Deep convolutional neural networks for hyperspectral image classification","volume":"2015","author":"Hu","year":"2015","journal-title":"J. Sens."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"025006","DOI":"10.1117\/1.JRS.10.025006","article-title":"Large patch convolutional neural networks for the scene classification of high spatial resolution imagery","volume":"10","author":"Zhong","year":"2016","journal-title":"J. Appl. Remote Sens."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"334","DOI":"10.1080\/2150704X.2017.1420265","article-title":"Application of a parallel spectral\u2013spatial convolution neural network in object-oriented remote sensing land use classification","volume":"9","author":"Cui","year":"2018","journal-title":"Remote Sens. Lett."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Cui, W., Zhou, Q., and Zheng, Z. (2018). Application of a hybrid model based on a convolutional auto-encoder and convolutional neural network in object-oriented remote sensing classification. Algorithms, 11.","DOI":"10.3390\/a11010009"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"400","DOI":"10.1109\/TGRS.1986.289598","article-title":"Segmentation of a thematic mapper image using the fuzzy c-means clusterng algorthm","volume":"GE-24","author":"Cannon","year":"1986","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"663","DOI":"10.1109\/36.158859","article-title":"Classification with spatio-temporal interpixel class dependency contexts","volume":"30","author":"Jeon","year":"1992","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_15","unstructured":"Baatz, M., Schape, A., and Multiresolution segmentation: An optimization approach for high quality multi-scale image segmentation (2019, April 02). Angew. Available online: https:\/\/pdfs.semanticscholar.org\/364c\/c1ff514a2e11d21a101dc072575e5487d17e.pdf?_ga=2.55340014.416308819.1554177081-320853791.1554177081."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2015, January 7\u201312). Show and tell: A neural image caption generator. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"ref_17","unstructured":"Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhutdinov, R., Zemel, R., and Bengio, Y. (2015, January 6\u201311). Show, attend and tell: Neural image caption generation with visual attention. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Chen, F., Ji, R., Sun, X., Wu, Y., and Su, J. (2018, January 18\u201322). GroupCap: Group-based image captioning with structured relevance and diversity constraints. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00146"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1016\/j.neucom.2018.10.059","article-title":"3G structure for image caption generation","volume":"330","author":"Yuan","year":"2019","journal-title":"Neurocomputing"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Chen, H., Ding, G., Lin, Z., Zhao, S., and Han, J. (2018, January 13\u201319). Show, observe and tell: Attribute-driven attention model for image captioning. Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, Stockholm, Sweden.","DOI":"10.24963\/ijcai.2018\/84"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Khademi, M., and Schulte, O. (2018, January 18\u201322). Image caption generation with hierarchical contextual visual spatial attention. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00260"},{"key":"ref_22","unstructured":"Ranzato, M., Chopra, S., Auli, M., and Zaremba, W. (2019, April 02). Sequence Level Training with Recurrent Neural Networks. Available online: https:\/\/arxiv.org\/abs\/1511.06732."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Shi, H., Li, P., Wang, B., and Wang, Z. (2018, January 17\u201319). Image captioning based on deep reinforcement learning. Proceedings of the 10th International Conference on Internet Multimedia Computing and Service(ICIMCS), Nanjing, China. Available online: https:\/\/arxiv.org\/10.1145\/3240876.3240900.","DOI":"10.1145\/3240876.3240900"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Karpathy, A., and Li, F.-F. (2015, January 7\u201312). Deep visual-semantic alignments for generating image descriptions. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Lu, J., Xiong, C., Parikh, D., and Socher, R. (2017, January 21\u201326). Knowing when to look: adaptive attention via a visual sentinel for image captioning. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.345"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Qu, B., Li, X., Tao, D., and Lu, X. (2016, January 6\u20138). Deep semantic understanding of high resolution remote sensing image. Proceedings of the 2016 International Conference on Computer, Information and Telecommunication Systems (CITS), Kunming, China.","DOI":"10.1109\/CITS.2016.7546397"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"3623","DOI":"10.1109\/TGRS.2017.2677464","article-title":"Can a machine generate humanlike language descriptions for a remote sensing image?","volume":"55","author":"Shi","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"2183","DOI":"10.1109\/TGRS.2017.2776321","article-title":"Exploring models and data for remote sensing image caption generation","volume":"56","author":"Lu","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_29","first-page":"1","article-title":"Semantic descriptions of high-resolution remote sensing images","volume":"99","author":"Wang","year":"2019","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhang, X., Wang, X., Tang, X., Zhou, H., and Li, C. (2019). Description generation for remote sensing images using attribute attention mechanism. Remote Sens., 11.","DOI":"10.3390\/rs11060612"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Wang, Y., Lin, Z., Shen, X., Cohen, S., and Cottrell, G.W. (2017, January 21\u201326). Skeleton key: Image captioning by skeleton-attribute decomposition. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.780"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"234","DOI":"10.2307\/143141","article-title":"A computer movie simulating urban growth in the Detroit region","volume":"46","author":"Tobler","year":"1970","journal-title":"Econ. Geogr."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"539","DOI":"10.1016\/j.patcog.2016.07.001","article-title":"Towards better exploiting convolutional neural networks for remote sensing scene classification","volume":"61","author":"Nogueira","year":"2017","journal-title":"Pattern Recognit."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"1793","DOI":"10.1109\/TGRS.2015.2488681","article-title":"Scene classification via a gradient boosting random convolutional network framework","volume":"54","author":"Zhang","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"5653","DOI":"10.1109\/TGRS.2017.2711275","article-title":"Integrating multilayer features of convolutional neural networks for remote sensing scene classification","volume":"55","author":"Li","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1002","DOI":"10.1080\/2150704X.2018.1504334","article-title":"Region-based cascade pooling of convolutional features for HRRS image retrieval","volume":"9","author":"Ge","year":"2018","journal-title":"Remote Sens. Lett."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"640","DOI":"10.1109\/TPAMI.2016.2572683","article-title":"Fully convolutional networks for semantic segmentation","volume":"39","author":"Shelhamer","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Huang, Z., Cheng, G., Wang, H., Li, H., Shi, L., and Pan, C. (2016, January 10\u201315). Building extraction from multi-source remote sensing images via deep deconvolution neural networks. Proceedings of the 2016 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Beijing, China.","DOI":"10.1109\/IGARSS.2016.7729471"},{"key":"ref_39","first-page":"463","article-title":"Study on the optimal segmentation scale based on fractual dimension of remote sensing images","volume":"33","author":"Cui","year":"2011","journal-title":"J. Wuhan Univ. Technol."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"1138","DOI":"10.1080\/2150704X.2018.1513662","article-title":"A multi-depth convolutional neural network for SAR image classification","volume":"9","author":"Xia","year":"2018","journal-title":"Remote Sens. Lett."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Li, L., Liang, J., Weng, M., and Zhu, H. (2018). A multiple-feature reuse network to extract buildings from remote sensing imagery. Remote Sens., 10.","DOI":"10.3390\/rs10091350"},{"key":"ref_42","first-page":"234","article-title":"U-Net: Convolutional networks for biomedical image segmentation","volume":"Volume 9351","author":"Navab","year":"2015","journal-title":"Medical Image Computing and Computer-Assisted Intervention\u2014MICCAI 2015"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., van der Maaten, L., and Weinberger, K.Q. (2017, January 21\u201326). Densely connected convolutional networks. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Tao, Y., Xu, M., Lu, Z., and Zhong, Y. (2018). DenseNet-based depth-width double reinforced deep learning neural network for high-resolution remote sensing image per-pixel classification. Remote Sens., 10.","DOI":"10.3390\/rs10050779"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Dr\u00f6nner, J., Korfhage, N., Egli, S., M\u00fchling, M., Thies, B., Bendix, J., Freisleben, B., and Seeger, B. (2018). Fast cloud segmentation using convolutional neural networks. Remote Sens., 10.","DOI":"10.3390\/rs10111782"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Yang, H., Wu, P., Yao, X., Wu, Y., Wang, B., and Xu, Y. (2018). Building extraction in very high resolution imagery by dense-attention networks. Remote Sens., 10.","DOI":"10.3390\/rs10111768"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Panboonyuen, T., Jitkajornwanich, K., Lawawirojwong, S., Srestasathiern, P., and Vateekul, P. (2019). Semantic segmentation on remotely sensed images using an enhanced global convolutional network with channel attention and domain specific transfer learning. Remote Sens., 11.","DOI":"10.20944\/preprints201812.0090.v3"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Zhang, T., and Tang, H. (2019). A comprehensive evaluation of approaches for built-up area extraction from landsat oli images using massive samples. Remote Sens., 11.","DOI":"10.20944\/preprints201812.0067.v1"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Sun, G., Huang, H., Zhang, A., Li, F., Zhao, H., and Fu, H. (2019). Fusion of multiscale convolutional neural networks for building extraction in very high-resolution images. Remote Sens., 11.","DOI":"10.3390\/rs11030227"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Fu, Y., Liu, K., Shen, Z., Deng, J., Gan, M., Liu, X., Lu, D., and Wang, K. (2019). Mapping impervious surfaces in town\u2013rural transition belts using China\u2019s GF-2 imagery and object-based deep CNNs. Remote Sens., 11.","DOI":"10.3390\/rs11030280"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Li, W., Dong, R., Fu, H., and Yu, L. (2019). Large-scale oil palm tree detection from high-resolution satellite images using two-stage convolutional neural networks. Remote Sens., 11.","DOI":"10.3390\/rs11010011"},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"3639","DOI":"10.1109\/TGRS.2016.2636241","article-title":"Deep recurrent neural networks for hyperspectral image classification","volume":"55","author":"Mou","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Wu, H., and Prasad, S. (2017). Convolutional recurrent neural networks forhyperspectral data classification. Remote Sens., 9.","DOI":"10.3390\/rs9030298"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Ndikumana, E., Minh, D.H.T., Baghdadi, N., Courault, D., and Hossard, L. (2018). Deep recurrent neural network for agricultural classification using multitemporal sar sentinel-1 for Camargue, France. Remote Sens., 10.","DOI":"10.3390\/rs10081217"},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"1118","DOI":"10.1080\/2150704X.2018.1511933","article-title":"Spectral-spatial classification of hyperspectral imagery based on recurrent neural networks","volume":"9","author":"Liu","year":"2018","journal-title":"Remote Sens. Lett."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Liu, Q., Zhou, F., Hang, R., and Yuan, X. (2017). Bidirectional-convolutional LSTM based spectral-spatial feature learning for hyperspectral image classification. Remote Sens., 9.","DOI":"10.3390\/rs9121330"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Ma, A., Filippi, A.M., Wang, Z., and Yin, Z. (2019). Hyperspectral image classification using similarity measurements-based deep recurrent neural networks. Remote Sens., 11.","DOI":"10.3390\/rs11020194"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002, January 7\u201312). BLEU: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting on Association for Computational Linguistics\u2014ACL\u201902, Philadelphia, PA, USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Lavie, A., and Agarwal, A. (2007, January 23). Meteor: An automatic metric for MT evaluation with high levels of correlation with human judgments. Proceedings of the Second Workshop on Statistical Machine Translation\u2014StatMT\u201907, Prague, Czech Republic.","DOI":"10.3115\/1626355.1626389"},{"key":"ref_61","unstructured":"Lin, C. (2004, January 25\u201326). ROUGE: A package for automatic evaluation of summaries. Proceedings of the Workshop on Text Summarization Branches Out\u2014ACL\u201905, Barcelona, Spain. Available online: https:\/\/www.aclweb.org\/anthology\/W04-1013."},{"key":"ref_62","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Zitnick, C.L., and Parikh, D. (2015, January 7\u201312). CIDEr: Consensus-based image description evaluation. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"You, Q., Jin, H., Wang, Z., Fang, C., and Luo, J. (July, January 26). Image captioning with semantic attention. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.503"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/11\/9\/1044\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:48:49Z","timestamp":1760186929000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/11\/9\/1044"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,5,2]]},"references-count":63,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2019,5]]}},"alternative-id":["rs11091044"],"URL":"https:\/\/doi.org\/10.3390\/rs11091044","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,5,2]]}}}