{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,16]],"date-time":"2025-10-16T10:10:07Z","timestamp":1760609407830,"version":"build-2065373602"},"reference-count":72,"publisher":"MDPI AG","issue":"22","license":[{"start":{"date-parts":[[2021,11,10]],"date-time":"2021-11-10T00:00:00Z","timestamp":1636502400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key R&amp;D Program of China","award":["2018YFC1407201"],"award-info":[{"award-number":["2018YFC1407201"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>The semantic segmentation of remote sensing images requires distinguishing local regions of different classes and exploiting a uniform global representation of the same-class instances. Such requirements make it necessary for the segmentation methods to extract discriminative local features between different classes and to explore representative features for all instances of a given class. While common deep convolutional neural networks (DCNNs) can effectively focus on local features, they are limited by their receptive field to obtain consistent global information. In this paper, we propose a memory-augmented transformer (MAT) to effectively model both the local and global information. The feature extraction pipeline of the MAT is split into a memory-based global relationship guidance module and a local feature extraction module. The local feature extraction module mainly consists of a transformer, which is used to extract features from the input images. The global relationship guidance module maintains a memory bank for the consistent encoding of the global information. Global guidance is performed by memory interaction. Bidirectional information flow between the global and local branches is conducted by a memory-query module, as well as a memory-update module, respectively. Experiment results on the ISPRS Potsdam and ISPRS Vaihingen datasets demonstrated that our method can perform competitively with state-of-the-art methods.<\/jats:p>","DOI":"10.3390\/rs13224518","type":"journal-article","created":{"date-parts":[[2021,11,11]],"date-time":"2021-11-11T23:04:46Z","timestamp":1636671886000},"page":"4518","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Memory-Augmented Transformer for Remote Sensing Image Semantic Segmentation"],"prefix":"10.3390","volume":"13","author":[{"given":"Xin","family":"Zhao","sequence":"first","affiliation":[{"name":"Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"Key Laboratory of Technology in Geo-Spatial Information Processing and Application System, Beijing 100190, China"},{"name":"School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing 101408, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiayi","family":"Guo","sequence":"additional","affiliation":[{"name":"Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"Key Laboratory of Technology in Geo-Spatial Information Processing and Application System, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yueting","family":"Zhang","sequence":"additional","affiliation":[{"name":"Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"Key Laboratory of Technology in Geo-Spatial Information Processing and Application System, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yirong","family":"Wu","sequence":"additional","affiliation":[{"name":"Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"Key Laboratory of Technology in Geo-Spatial Information Processing and Application System, Beijing 100190, China"},{"name":"School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing 101408, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,11,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Neupane, B., Horanont, T., and Aryal, J. (2021). Deep Learning-Based Semantic Segmentation of Urban Features in Satellite Images: A Review and Meta-Analysis. Remote Sens., 13.","DOI":"10.3390\/rs13040808"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"114417","DOI":"10.1016\/j.eswa.2020.114417","article-title":"A review of deep learning methods for semantic segmentation of remote sensing imagery","volume":"169","author":"Yuan","year":"2021","journal-title":"Expert Syst. Appl."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1016\/j.neucom.2019.02.003","article-title":"Survey on semantic segmentation using deep learning techniques","volume":"338","author":"Lateef","year":"2019","journal-title":"Neurocomputing"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1016\/j.isprsjprs.2016.10.010","article-title":"MRF-based segmentation and unsupervised classification for building and road detection in peri-urban areas of high-resolution satellite images","volume":"122","author":"Grinias","year":"2016","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"69","DOI":"10.1080\/01431160903439882","article-title":"Information fusion of aerial images and LIDAR data in urban areas: Vector-stacking, re-classification and post-processing approaches","volume":"32","author":"Huang","year":"2011","journal-title":"Int. J. Remote Sens."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1731","DOI":"10.1109\/TPAMI.2011.208","article-title":"Layered object models for image segmentation","volume":"34","author":"Yang","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"205","DOI":"10.1016\/j.isprsjprs.2020.10.015","article-title":"Mapping forest tree species in high resolution UAV-based RGB-imagery by means of convolutional neural networks","volume":"170","author":"Schiefer","year":"2020","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Nezami, S., Khoramshahi, E., Nevalainen, O., P\u00f6l\u00f6nen, I., and Honkavaara, E. (2020). Tree species classification of drone hyperspectral and rgb imagery with deep learning convolutional neural networks. Remote Sens., 12.","DOI":"10.20944\/preprints202002.0334.v1"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Mou, L., Hua, Y., and Zhu, X.X. (2019, January 15\u201320). A relation-augmented fully convolutional network for semantic segmentation in aerial scenes. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01270"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Peng, C., Zhang, K., Ma, Y., and Ma, J. (2021). Cross Fusion Net: A Fast Semantic Segmentation Network for Small-Scale Semantic Information Capturing in Aerial Scenes. IEEE Trans. Geosci. Remote. Sens.","DOI":"10.1109\/TGRS.2021.3053062"},{"key":"ref_11","unstructured":"Chen, L.C., Papandreou, G., Schroff, F., and Adam, H. (2017). Rethinking atrous convolution for semantic image segmentation. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_13","unstructured":"Yuan, Y., Huang, L., Guo, J., Zhang, C., Chen, X., and Wang, J. (2018). Ocnet: Object context network for scene parsing. arXiv."},{"key":"ref_14","unstructured":"Tao, A., Sapra, K., and Catanzaro, B. (2020). Hierarchical multi-scale attention for semantic segmentation. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., and Wei, Y. (2017, January 22\u201329). Deformable convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.89"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zhu, X., Hu, H., Lin, S., and Dai, J. (2019, January 15\u201320). Deformable convnets v2: More deformable, better results. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00953"},{"key":"ref_17","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"180","DOI":"10.1016\/j.isprsjprs.2013.09.014","article-title":"Geographic object-based image analysis\u2013towards a new paradigm","volume":"87","author":"Blaschke","year":"2014","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_19","unstructured":"Derivaux, S., Lefevre, S., Wemmert, C., and Korczak, J. (August, January 31). Watershed segmentation of remotely sensed images based on a supervised fuzzy pixel classification. Proceedings of the IEEE International Geosciences And Remote Sensing Symposium (IGARSS), Denver, CO, USA."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"319","DOI":"10.1016\/j.isprsjprs.2018.12.003","article-title":"Scale-variable region-merging for high resolution remote sensing image segmentation","volume":"147","author":"Su","year":"2019","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1109\/36.905239","article-title":"A new approach for the morphological segmentation of high-resolution satellite imagery","volume":"39","author":"Pesaresi","year":"2001","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"4758","DOI":"10.1080\/01431161.2014.930199","article-title":"Object-based change detection in wind storm-damaged forest using high-resolution multispectral images","volume":"35","author":"Chehata","year":"2014","journal-title":"Int. J. Remote Sens."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"152","DOI":"10.1016\/j.isprsjprs.2020.01.028","article-title":"A framework for large-scale mapping of human settlement extent from Sentinel-2 images via fully convolutional neural networks","volume":"163","author":"Qiu","year":"2020","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"025010","DOI":"10.1117\/1.JRS.12.025010","article-title":"Using convolutional neural network to identify irregular segmentation objects from very high-resolution remote sensing imagery","volume":"12","author":"Fu","year":"2018","journal-title":"J. Appl. Remote Sens."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"5367","DOI":"10.1109\/TGRS.2020.2964675","article-title":"Semantic segmentation of large-size VHR remote sensing images using a two-stage multiscale training architecture","volume":"58","author":"Ding","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"905","DOI":"10.1109\/LGRS.2020.2988294","article-title":"SCAttNet: Semantic segmentation network with spatial and channel attention mechanism for high-resolution remote sensing images","volume":"18","author":"Li","year":"2020","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_29","unstructured":"Burtsev, M.S., Kuratov, Y., Peganov, A., and Sapunov, G.V. (2020). Memory transformer. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020, January 23\u201328). End-to-end object detection with transformers. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_31","unstructured":"Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J. (2020). Deformable detr: Deformable transformers for end-to-end object detection. arXiv."},{"key":"ref_32","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. arXiv.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L. (2021). Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv.","DOI":"10.1109\/ICCV48922.2021.00061"},{"key":"ref_35","unstructured":"Chu, X., Tian, Z., Wang, Y., Zhang, B., Ren, H., Wei, X., Xia, H., and Shen, C. (2021). Twins: Revisiting the design of spatial attention in vision transformers. arXiv."},{"key":"ref_36","unstructured":"Sun, P., Jiang, Y., Zhang, R., Xie, E., Cao, J., Hu, X., Kong, T., Yuan, Z., Wang, C., and Luo, P. (2020). Transtrack: Multiple-object tracking with transformer. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Yan, B., Peng, H., Fu, J., Wang, D., and Lu, H. (2021). Learning spatio-temporal transformer for visual tracking. arXiv.","DOI":"10.1109\/ICCV48922.2021.01028"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Hirose, S., Wada, N., Katto, J., and Sun, H. (2021, January 25\u201327). ViT-GAN: Using Vision Transformer as Discriminator with Adaptive Data Augmentation. Proceedings of the 2021 3rd International Conference on Computer Communication and the Internet (ICCCI), Nagoya, Japan.","DOI":"10.1109\/ICCCI51764.2021.9486805"},{"key":"ref_39","unstructured":"Lee, K., Chang, H., Jiang, L., Zhang, H., Tu, Z., and Liu, C. (2021). ViTGAN: Training GANs with Vision Transformers. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Esser, P., Rombach, R., and Ommer, B. (2021, January 19\u201325). Taming transformers for high-resolution image synthesis. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01268"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Engel, N., Belagiannis, V., and Dietmayer, K. (2020). Point transformer. arXiv.","DOI":"10.1109\/ACCESS.2021.3116304"},{"key":"ref_42","unstructured":"Guo, M.H., Cai, J.X., Liu, Z.N., Mu, T.J., Martin, R.R., and Hu, S.M. (2020). PCT: Point cloud transformer. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Qing, Y., Liu, W., Feng, L., and Gao, W. (2021). Improved Transformer Net for Hyperspectral Image Classification. Remote Sens., 13.","DOI":"10.3390\/rs13112216"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Bazi, Y., Bashmal, L., Rahhal, M.M.A., Dayil, R.A., and Ajlan, N.A. (2021). Vision Transformers for Remote Sensing Image Classification. Remote Sens., 13.","DOI":"10.3390\/rs13030516"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"He, X., Chen, Y., and Lin, Z. (2021). Spatial-Spectral Transformer for Hyperspectral Image Classification. Remote Sens., 13.","DOI":"10.3390\/rs13030498"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Li, W., Cao, D., Peng, Y., and Yang, C. (2021). MSNet: A Multi-Stream Fusion Network for Remote Sensing Spatiotemporal Fusion Based on Transformer and Convolution. Remote Sens., 13.","DOI":"10.3390\/rs13183724"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Yu, Y., Zhao, J., Gong, Q., Huang, C., Zheng, G., and Ma, J. (2021). Real-Time Underwater Maritime Object Detection in Side-Scan Sonar Images Based on Transformer-YOLOv5. Remote Sens., 13.","DOI":"10.3390\/rs13183555"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Xu, Z., Zhang, W., Zhang, T., Yang, Z., and Li, J. (2021). Efficient Transformer for Remote Sensing Image Segmentation. Remote Sens., 13.","DOI":"10.3390\/rs13183585"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Wang, L., Li, R., Wang, D., Duan, C., Wang, T., and Meng, X. (2021). Transformer Meets Convolution: A Bilateral Awareness Network for Semantic Segmentation of Very Fine Resolution Urban Scene Images. Remote Sens., 13.","DOI":"10.3390\/rs13163065"},{"key":"ref_50","unstructured":"Oord, A.V.D., Vinyals, O., and Kavukcuoglu, K. (2017). Neural discrete representation learning. arXiv."},{"key":"ref_51","unstructured":"Razavi, A., van den Oord, A., and Vinyals, O. (2019, January 8\u201314). Generating diverse high-fidelity images with vq-vae-2. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, QC, Canada."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Han, T., Xie, W., and Zisserman, A. (2020, January 23\u201328). Memory-augmented dense predictive coding for video representation learning. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK.","DOI":"10.1007\/978-3-030-58580-8_19"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Oh, S.W., Lee, J.Y., Xu, N., and Kim, S.J. (2019, January 16\u201317). Video object segmentation using space-time memory networks. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Long Beach, CA, USA.","DOI":"10.1109\/ICCV.2019.00932"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Gong, D., Liu, L., Le, V., Saha, B., Mansour, M.R., Venkatesh, S., and Hengel, A.V.D. (2019, January 16\u201317). Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Long Beach, CA, USA.","DOI":"10.1109\/ICCV.2019.00179"},{"key":"ref_55","unstructured":"Kim, Y., Kim, M., and Kim, G. (2018). Memorization precedes generation: Learning unsupervised gans with memory networks. arXiv."},{"key":"ref_56","unstructured":"Santoro, A., Bartunov, S., Botvinick, M., Wierstra, D., and Lillicrap, T. (2016, January 20\u201322). Meta-learning with memory-augmented neural networks. Proceedings of the International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Guo, M.H., Liu, Z.N., Mu, T.J., and Hu, S.M. (2021). Beyond self-attention: External attention using two linear layers for visual tasks. arXiv.","DOI":"10.1109\/TPAMI.2022.3211006"},{"key":"ref_58","unstructured":"Hendrycks, D., and Gimpel, K. (2020). Gaussian Error Linear Units (GELUs). arXiv."},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Vaswani, A., Ramachandran, P., Srinivas, A., Parmar, N., Hechtman, B., and Shlens, J. (2021, January 19\u201325). Scaling local self-attention for parameter efficient visual backbones. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01270"},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Wu, H., Xiao, B., Codella, N., Liu, M., Dai, X., Yuan, L., and Zhang, L. (2021). Cvt: Introducing convolutions to vision transformers. arXiv.","DOI":"10.1109\/ICCV48922.2021.00009"},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Niu, R., Sun, X., Tian, Y., Diao, W., Chen, K., and Fu, K. (2021). Hybrid multiple attention network for semantic segmentation in aerial images. IEEE Trans. Geosci. Remote Sens.","DOI":"10.1109\/TGRS.2021.3065112"},{"key":"ref_62","doi-asserted-by":"crossref","unstructured":"Yun, S., Han, D., Oh, S.J., Chun, S., Choe, J., and Yoo, Y. (2019, January 16\u201317). Cutmix: Regularization strategy to train strong classifiers with localizable features. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Long Beach, CA, USA.","DOI":"10.1109\/ICCV.2019.00612"},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Zhang, H., Cisse, M., Dauphin, Y.N., and Lopez-Paz, D. (2017). mixup: Beyond empirical risk minimization. arXiv.","DOI":"10.1007\/978-1-4899-7687-1_79"},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Pan, X., Shi, J., Luo, P., Wang, X., and Tang, X. (2018, January 2\u20137). Spatial as deep: Spatial cnn for traffic scene understanding. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12301"},{"key":"ref_66","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1016\/j.isprsjprs.2018.06.005","article-title":"Developing a multi-filter convolutional neural network for semantic segmentation using high-resolution aerial imagery and LiDAR data","volume":"143","author":"Sun","year":"2018","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_67","doi-asserted-by":"crossref","first-page":"881","DOI":"10.1109\/TGRS.2016.2616585","article-title":"Dense semantic labeling of subdecimeter resolution images with convolutional neural networks","volume":"55","author":"Volpi","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_68","doi-asserted-by":"crossref","first-page":"7503","DOI":"10.1109\/TGRS.2019.2913861","article-title":"Dynamic multicontext segmentation of remote sensing images based on convolutional networks","volume":"57","author":"Nogueira","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Shi, H., Fan, J., Wang, Y., and Chen, L. (2021). Dual Attention Feature Fusion and Adaptive Context for Accurate Segmentation of Very High-Resolution Remote Sensing Images. Remote Sens., 13.","DOI":"10.3390\/rs13183715"},{"key":"ref_70","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1016\/j.isprsjprs.2018.01.021","article-title":"Land cover mapping at very high resolution with rotation equivariant CNNs: Towards small yet accurate models","volume":"145","author":"Marcos","year":"2018","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_71","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1016\/j.isprsjprs.2020.01.023","article-title":"Aerial image semantic segmentation using DCNN predicted distance maps","volume":"161","author":"Chai","year":"2020","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_72","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Hinton","year":"2008","journal-title":"J. Mach. Learn. Res."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/22\/4518\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:28:09Z","timestamp":1760167689000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/22\/4518"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,11,10]]},"references-count":72,"journal-issue":{"issue":"22","published-online":{"date-parts":[[2021,11]]}},"alternative-id":["rs13224518"],"URL":"https:\/\/doi.org\/10.3390\/rs13224518","relation":{},"ISSN":["2072-4292"],"issn-type":[{"type":"electronic","value":"2072-4292"}],"subject":[],"published":{"date-parts":[[2021,11,10]]}}}