{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T16:51:56Z","timestamp":1784652716741,"version":"3.55.0"},"reference-count":167,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2023,3,30]],"date-time":"2023-03-30T00:00:00Z","timestamp":1680134400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Deep learning-based algorithms have seen a massive popularity in different areas of remote sensing image analysis over the past decade. Recently, transformer-based architectures, originally introduced in natural language processing, have pervaded computer vision field where the self-attention mechanism has been utilized as a replacement to the popular convolution operator for capturing long-range dependencies. Inspired by recent advances in computer vision, the remote sensing community has also witnessed an increased exploration of vision transformers for a diverse set of tasks. Although a number of surveys have focused on transformers in computer vision in general, to the best of our knowledge we are the first to present a systematic review of recent advances based on transformers in remote sensing. Our survey covers more than 60 recent transformer-based methods for different remote sensing problems in sub-areas of remote sensing: very high-resolution (VHR), hyperspectral (HSI) and synthetic aperture radar (SAR) imagery. We conclude the survey by discussing different challenges and open issues of transformers in remote sensing.<\/jats:p>","DOI":"10.3390\/rs15071860","type":"journal-article","created":{"date-parts":[[2023,3,31]],"date-time":"2023-03-31T01:37:02Z","timestamp":1680226622000},"page":"1860","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":381,"title":["Transformers in Remote Sensing: A Survey"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6114-0494","authenticated-orcid":false,"given":"Abdulaziz Amer","family":"Aleissaee","sequence":"first","affiliation":[{"name":"Computer Vision Faculty, Mohamed bin Zayed University of Artificial Intelligence, Building 1B, Masdar City, Abu Dhabi P.O. Box 5224, United Arab Emirates"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Amandeep","family":"Kumar","sequence":"additional","affiliation":[{"name":"Computer Vision Faculty, Mohamed bin Zayed University of Artificial Intelligence, Building 1B, Masdar City, Abu Dhabi P.O. Box 5224, United Arab Emirates"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rao Muhammad","family":"Anwer","sequence":"additional","affiliation":[{"name":"Computer Vision Faculty, Mohamed bin Zayed University of Artificial Intelligence, Building 1B, Masdar City, Abu Dhabi P.O. Box 5224, United Arab Emirates"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Salman","family":"Khan","sequence":"additional","affiliation":[{"name":"Computer Vision Faculty, Mohamed bin Zayed University of Artificial Intelligence, Building 1B, Masdar City, Abu Dhabi P.O. Box 5224, United Arab Emirates"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hisham","family":"Cholakkal","sequence":"additional","affiliation":[{"name":"Computer Vision Faculty, Mohamed bin Zayed University of Artificial Intelligence, Building 1B, Masdar City, Abu Dhabi P.O. Box 5224, United Arab Emirates"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7660-6090","authenticated-orcid":false,"given":"Gui-Song","family":"Xia","sequence":"additional","affiliation":[{"name":"School of Computer Science, Wuhan University, Wuchang District, Wuhan 430072, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fahad Shahbaz","family":"Khan","sequence":"additional","affiliation":[{"name":"Computer Vision Faculty, Mohamed bin Zayed University of Artificial Intelligence, Building 1B, Masdar City, Abu Dhabi P.O. Box 5224, United Arab Emirates"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,30]]},"reference":[{"key":"ref_1","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3\u20137). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. Proceedings of the ICLR, Virtual-Only."},{"key":"ref_2","unstructured":"Naseer, M., Ranasinghe, K., Khan, S., Hayat, M., Khan, F.S., and Yang, M.H. (2021, January 7\u201310). Intriguing Properties of Vision Transformers. Proceedings of the NeurIPS, Virtual-Only."},{"key":"ref_3","unstructured":"Park, N., and Kim, S. (2022, January 25). How Do Vision Transformers Work?. Proceedings of the ICLR, Virtual-Only."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Bazi, Y., Bashmal, L., Rahhal, M.M.A., Dayil, R.A., and Ajlan, N.A. (2021). Vision transformers for remote sensing image classification. Remote Sens., 13.","DOI":"10.3390\/rs13030516"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Hao, S., Wu, B., Zhao, K., Ye, Y., and Wang, W. (2022). Two-Stream Swin Transformer with Differentiable Sobel Operator for Remote Sensing Image Classification. Remote Sens., 14.","DOI":"10.3390\/rs14061507"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"2223","DOI":"10.1109\/JSTARS.2022.3155665","article-title":"Homo\u2013Heterogenous Transformer Learning Framework for RS Scene Classification","volume":"15","author":"Ma","year":"2022","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Wang, D., Zhang, J., Du, B., Xia, G.S., and Tao, D. (2022). An Empirical Study of Remote Sensing Pretraining. IEEE Trans. Geosci. Remote Sens.","DOI":"10.1109\/TGRS.2022.3176603"},{"key":"ref_8","first-page":"5518615","article-title":"SpectralFormer: Rethinking hyperspectral image classification with transformers","volume":"60","author":"Hong","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1080\/22797254.2021.2023910","article-title":"DSS-TRM: Deep spatial\u2013spectral transformer for hyperspectral image classification","volume":"55","author":"Liu","year":"2022","journal-title":"Eur. J. Remote Sens."},{"key":"ref_10","first-page":"1","article-title":"Convolutional Transformer Network for Hyperspectral Image Classification","volume":"19","author":"Zhao","year":"2022","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_11","first-page":"5528715","article-title":"Hyperspectral Image Transformer Classification Networks","volume":"60","author":"Yang","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_12","unstructured":"Jia, S., and Wang, Y. (2022). Multiscale Convolutional Transformer with Center Mask Pretraining for Hyperspectral Image Classification. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"606","DOI":"10.1109\/JSTSP.2011.2139193","article-title":"A survey of active learning algorithms for supervised remote sensing image classification","volume":"5","author":"Tuia","year":"2011","journal-title":"IEEE J. Sel. Top. Signal Process."},{"key":"ref_14","first-page":"45","article-title":"Advances in hyperspectral image classification: Earth monitoring with statistical learning methods","volume":"31","author":"Tuia","year":"2013","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1109\/MGRS.2017.2762307","article-title":"Deep learning in remote sensing: A comprehensive review and list of resources","volume":"5","author":"Zhu","year":"2017","journal-title":"IEEE Geosci. Remote Sens. Mag."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"166","DOI":"10.1016\/j.isprsjprs.2019.04.015","article-title":"Deep learning in remote sensing applications: A meta-analysis and review","volume":"152","author":"Ma","year":"2019","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_17","first-page":"600","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"NeurIPS"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3505244","article-title":"Transformers in Vision: A Survey","volume":"54","author":"Khan","year":"2021","journal-title":"ACM Comput. Surv."},{"key":"ref_19","unstructured":"Shamshad, F., Khan, S., Zamir, S.W., Khan, M.H., Hayat, M., Khan, F.S., and Fu, H. (2022). Transformers in medical imaging: A survey. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Selva, J., Johansen, A., Escalera, S., Nasrollahi, K., Moeslund, T., and Clapes, A. (2022). Video Transformers: A Survey. arXiv.","DOI":"10.1109\/TPAMI.2023.3243465"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Teng, M.Y., Mehrubeoglu, R., King, S.A., Cammarata, K., and Simons, J. (2013, January 26\u201328). Investigation of epifauna coverage on seagrass blades using spatial and spectral analysis of hyperspectral images. Proceedings of the 2013 5th Workshop on Hyperspectral Image and Signal Processing: Evolution in Remote Sensing (WHISPERS), Gainesville, FL, USA.","DOI":"10.1109\/WHISPERS.2013.8080658"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Notesco, G., Dor, E.B., and Brook, A. (2014, January 24\u201327). Mineral mapping of makhtesh ramon in israel using hyperspectral remote sensing day and night LWIR images. Proceedings of the 2014 6th Workshop on Hyperspectral Image and Signal Processing: Evolution in Remote Sensing (WHISPERS), Lausanne, Switzerland.","DOI":"10.1109\/WHISPERS.2014.8077538"},{"key":"ref_23","first-page":"84","article-title":"Imagenet classification with deep convolutional neural networks","volume":"60","author":"Krizhevsky","year":"2012","journal-title":"NeurIPS"},{"key":"ref_24","first-page":"1137","article-title":"Faster r-cnn: Towards real-time object detection with region proposal networks","volume":"28","author":"Ren","year":"2015","journal-title":"NeurIPS"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei, L. (2009, January 20\u201325). ImageNet: A large-scale hierarchical image database. Proceedings of the CVPR, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_26","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the CVPR, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_28","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the CVPR, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L. (2021, January 10\u201317). Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions. Proceedings of the ICCV, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00061"},{"key":"ref_31","first-page":"1","article-title":"When CNNs meet vision transformer: A joint framework for remote sensing scene classification","volume":"19","author":"Deng","year":"2021","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zhang, J., Zhao, H., and Li, J. (2021). TRS: Transformers for Remote Sensing Scene Classification. Remote Sens., 13.","DOI":"10.3390\/rs13204143"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"4205","DOI":"10.1109\/JSTARS.2021.3070368","article-title":"On Creating Benchmark Dataset for Aerial Image Interpretation: Reviews, Guidances and Million-AID","volume":"14","author":"Long","year":"2021","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Chattopadhay, A., Sarkar, A., Howlader, P., and Balasubramanian, V.N. (2018, January 12\u201315). Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks. Proceedings of the 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe, NV, USA.","DOI":"10.1109\/WACV.2018.00097"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"3965","DOI":"10.1109\/TGRS.2017.2685945","article-title":"AID: A benchmark data set for performance evaluation of aerial scene classification","volume":"55","author":"Xia","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020, January 23\u201328). End-to-end object detection with transformers. Proceedings of the ECCV, Glasgow, UK.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Xu, X., Feng, Z., Cao, C., Li, M., Wu, J., Wu, Z., Shang, Y., and Ye, S. (2021). An Improved Swin Transformer-Based Model for Remote Sensing Object Detection and Instance Segmentation. Remote Sens., 13.","DOI":"10.3390\/rs13234779"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask r-cnn. Proceedings of the ICCV, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Li, Q., Chen, Y., and Zeng, Y. (2022). Transformer with Transfer CNN for Remote-Sensing-Image Object Detection. Remote Sens., 14.","DOI":"10.3390\/rs14040984"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Liu, X., Wa, S., Chen, S., and Ma, Q. (2022). GANsformer: A Detection Network for Aerial Images with High Performance Combining Convolutional Network and Transformer. Remote Sens., 14.","DOI":"10.3390\/rs14040923"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Zheng, Y., Sun, P., Zhou, Z., Xu, W., and Ren, Q. (2021). ADT-Det: Adaptive Dynamic Refined Single-Stage Transformer Detector for Arbitrary-Oriented Object Detection in Satellite Optical Imagery. Remote Sens., 13.","DOI":"10.3390\/rs13132623"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Tang, J., Zhang, W., Liu, H., Yang, M., Jiang, B., Hu, G., and Bai, X. (2022, January 19\u201324). Few Could Be Better Than All: Feature Sampling and Grouping for Scene Text Detection. Proceedings of the CVPR, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00452"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Dai, Y., Yu, J., Zhang, D., Hu, T., and Zheng, X. (2022). RODFormer: High-Precision Design for Rotating Object Detection with Transformers. Sensors, 22.","DOI":"10.3390\/s22072633"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Zhou, Q., and Yu, C. (2022). Point RCNN: An Angle-Free Framework for Rotated Object Detection. Remote Sens., 14.","DOI":"10.3390\/rs14112605"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Liu, X., Ma, S., He, L., Wang, C., and Chen, Z. (2022). Hybrid Network Model: TransConvNet for Oriented Object Detection in Remote Sensing Images. Remote Sens., 14.","DOI":"10.3390\/rs14092090"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Li, W., Chen, Y., Hu, K., and Zhu, J. (2021, January 20\u201325). Oriented RepPoints for Aerial Object Detection. Proceedings of the IEEE\/CVF, Nashville, TN, USA.","DOI":"10.1109\/CVPR52688.2022.00187"},{"key":"ref_47","unstructured":"Ma, T., Mao, M., Zheng, H., Gao, P., Wang, X., Han, S., Ding, E., Zhang, B., and Doermann, D. (2021). Oriented Object Detection with Transformer. arXiv."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Dai, L., Liu, H., Tang, H., Wu, Z., and Song, P. (2022). AO2-DETR: Arbitrary-Oriented Object Detection Transformer. arXiv.","DOI":"10.1109\/TCSVT.2022.3222906"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Xia, G.S., Bai, X., Ding, J., Zhu, Z., Belongie, S., Luo, J., Datcu, M., Pelillo, M., and Zhang, L. (2018, January 18\u201322). DOTA: A large-scale dataset for object detection in aerial images. Proceedings of the CVPR, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00418"},{"key":"ref_50","unstructured":"Muzein, B.S. (2006). Remote Sensing & GIS for Land Cover, Land Use Change Detection and Analysis in the Semi-Natural Ecosystems and Agriculture Landscapes of the Central Ethiopian Rift Valley. [Ph.D. Thesis, Institute of Photogrammetry and Remote Sensing, Technology University of Dresden]."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"65","DOI":"10.1080\/10106049809354643","article-title":"Remote sensing change detection of irrigated agriculture in Afghanistan","volume":"13","author":"Haack","year":"1998","journal-title":"Geocarto Int."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Bolorinos, J., Ajami, N.K., and Rajagopal, R. (2020). Consumption change detection for urban planning: Monitoring and segmenting water customers during drought. Water Resour. Res., 56.","DOI":"10.1029\/2019WR025812"},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1016\/S0924-2716(99)00023-4","article-title":"Change detection assessment using fuzzy sets and remotely sensed data: An application of topographic map revision","volume":"54","author":"Metternicht","year":"1999","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_54","first-page":"5607514","article-title":"Remote Sensing Image Change Detection with Transformers","volume":"60","author":"Chen","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_55","first-page":"3131993","article-title":"Deep multiscale Siamese network with parallel convolutional structure and self-attention for change detection","volume":"60","author":"Guo","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"5224713","DOI":"10.1109\/TGRS.2022.3221492","article-title":"SwinSUNet: Pure Transformer Network for Remote Sensing Image Change Detection","volume":"60","author":"Zhang","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Wang, G., Li, B., Zhang, T., and Zhang, S. (2022). A Network Combining a Transformer and a Convolutional Neural Network for Remote Sensing Image Change Detection. Remote Sens., 14.","DOI":"10.3390\/rs14092228"},{"key":"ref_58","first-page":"5622519","article-title":"TransUNetCD: A Hybrid Transformer Network for Change Detection in Optical Remote-Sensing Images","volume":"60","author":"Li","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Ke, Q., and Zhang, P. (2022). Hybrid-TransCD: A Hybrid Transformer Remote Sensing Image Change Detection Network via Token Aggregation. Int. J. Geo-Inform., 11.","DOI":"10.3390\/ijgi11040263"},{"key":"ref_60","doi-asserted-by":"crossref","first-page":"574","DOI":"10.1109\/TGRS.2018.2858817","article-title":"Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set","volume":"57","author":"Ji","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Chen, H., and Shi, Z. (2020). A spatial-temporal attention-based method and a new dataset for remote sensing image change detection. Remote Sens., 12.","DOI":"10.3390\/rs12101662"},{"key":"ref_62","unstructured":"Daudt, R.C., Le Saux, B., and Boulch, A. (2018, January 7). Fully convolutional siamese networks for change detection. Proceedings of the ICIP, Athens, Greece."},{"key":"ref_63","doi-asserted-by":"crossref","first-page":"1301","DOI":"10.1007\/s10514-018-9734-5","article-title":"Street-view change detection with deconvolutional networks","volume":"42","author":"Alcantarilla","year":"2018","journal-title":"Auton. Robot."},{"key":"ref_64","doi-asserted-by":"crossref","first-page":"1194","DOI":"10.1109\/JSTARS.2020.3037893","article-title":"DASNet: Dual attentive fully convolutional Siamese networks for change detection in high-resolution satellite images","volume":"14","author":"Chen","year":"2020","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Xu, Z., Zhang, W., Zhang, T., Yang, Z., and Li, J. (2021). Efficient transformer for remote sensing image segmentation. Remote Sens., 13.","DOI":"10.3390\/rs13183585"},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"Wang, H., Chen, X., Zhang, T., Xu, Z., and Li, J. (2022). CCTNet: Coupled CNN and Transformer Network for Crop Segmentation of Remote Sensing Images. Remote Sens., 14.","DOI":"10.3390\/rs14091956"},{"key":"ref_67","doi-asserted-by":"crossref","first-page":"10990","DOI":"10.1109\/JSTARS.2021.3119654","article-title":"STransFuse: Fusing Swin Transformer and Convolutional Neural Network for Remote Sensing Image Semantic Segmentation","volume":"14","author":"Gao","year":"2021","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_68","first-page":"1","article-title":"Transformer and CNN Hybrid Deep Neural Network for Semantic Segmentation of Very-High-Resolution Remote Sensing Imagery","volume":"60","author":"Zhang","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Panboonyuen, T., Jitkajornwanich, K., Lawawirojwong, S., Srestasathiern, P., and Vateekul, P. (2021). Transformer-Based Decoder Designs for Semantic Segmentation on Remotely Sensed Images. Remote Sens., 13.","DOI":"10.3390\/rs13245100"},{"key":"ref_70","unstructured":"(2022, August 27). Available online: https:\/\/www.isprs.org\/education\/benchmarks\/UrbanSemLab\/2d-sem-label-potsdam.aspx."},{"key":"ref_71","unstructured":"(2022, August 27). Available online: https:\/\/www.isprs.org\/education\/benchmarks\/UrbanSemLab\/2d-sem-label-vaihingen.aspx."},{"key":"ref_72","doi-asserted-by":"crossref","unstructured":"Chen, K., Zou, Z., and Shi, Z. (2021). Building extraction from remote sensing images with sparse token transformers. Remote Sens., 13.","DOI":"10.3390\/rs13214441"},{"key":"ref_73","doi-asserted-by":"crossref","unstructured":"Xiao, X., Guo, W., Chen, R., Hui, Y., Wang, J., and Zhao, H. (2022). A Swin Transformer-Based Encoding Booster Integrated in U-Shaped Network for Building Extraction. Remote Sens., 14.","DOI":"10.3390\/rs14112611"},{"key":"ref_74","first-page":"2611","article-title":"Building extraction with vision transformer","volume":"14","author":"Wang","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_75","doi-asserted-by":"crossref","first-page":"4104","DOI":"10.1109\/JSTARS.2022.3175200","article-title":"Transferring transformer-based models for cross-area building extraction from remote sensing images","volume":"15","author":"Qiu","year":"2022","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_76","doi-asserted-by":"crossref","unstructured":"Yang, Y., and Newsam, S. (2010, January 2\u20135). Bag-of-visual-words and spatial extensions for land-use classification. Proceedings of the SIGSPATIAL, San Jose, CA, USA.","DOI":"10.1145\/1869790.1869829"},{"key":"ref_77","doi-asserted-by":"crossref","first-page":"1155","DOI":"10.1109\/TGRS.2018.2864987","article-title":"Scene classification with recurrent attention of VHR remote sensing images","volume":"57","author":"Wang","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_78","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1109\/LGRS.2017.2731997","article-title":"Remote sensing image scene classification using bag of convolutional features","volume":"14","author":"Cheng","year":"2017","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_79","doi-asserted-by":"crossref","first-page":"10590","DOI":"10.1109\/TGRS.2020.3047447","article-title":"Learning deep cross-modal embedding networks for zero-shot remote sensing image scene classification","volume":"59","author":"Li","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_80","unstructured":"Waqas Zamir, S., Arora, A., Gupta, A., Khan, S., Sun, G., Shahbaz Khan, F., Zhu, F., Shao, L., Xia, G.S., and Bai, X. (2019, January 16\u201320). Isaid: A large-scale dataset for instance segmentation in aerial images. Proceedings of the CVPR Workshops, Long Beach, CA, USA."},{"key":"ref_81","doi-asserted-by":"crossref","unstructured":"Liu, Z., Yuan, L., Weng, L., and Yang, Y. (2017, January 24\u201326). A high resolution optical satellite image dataset for ship recognition and some new baselines. Proceedings of the ICPRAM, Porto, Portugal.","DOI":"10.5220\/0006120603240331"},{"key":"ref_82","first-page":"324","article-title":"Change Detection in remote sensing images using conditional adversarial networks","volume":"42","author":"Lebedev","year":"2018","journal-title":"Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci."},{"key":"ref_83","doi-asserted-by":"crossref","first-page":"296","DOI":"10.1016\/j.isprsjprs.2019.11.023","article-title":"Object detection in optical remote sensing images: A survey and a new benchmark","volume":"159","author":"Li","year":"2020","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_84","doi-asserted-by":"crossref","first-page":"5535","DOI":"10.1109\/TGRS.2019.2900302","article-title":"Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection","volume":"57","author":"Zhang","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_85","doi-asserted-by":"crossref","first-page":"7405","DOI":"10.1109\/TGRS.2016.2601622","article-title":"Learning rotation-invariant convolutional neural networks for object detection in VHR optical remote sensing images","volume":"54","author":"Cheng","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_86","doi-asserted-by":"crossref","unstructured":"Zhu, H., Chen, X., Dai, W., Fu, K., Ye, Q., and Jiao, J. (2015, January 27\u201330). Orientation robust object detection in aerial images using deep convolutional neural network. Proceedings of the ICIP, Quebec City, QC, Canada.","DOI":"10.1109\/ICIP.2015.7351502"},{"key":"ref_87","doi-asserted-by":"crossref","first-page":"187","DOI":"10.1016\/j.jvcir.2015.11.002","article-title":"Vehicle detection in aerial imagery: A small target detection benchmark","volume":"34","author":"Razakarivony","year":"2016","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_88","doi-asserted-by":"crossref","unstructured":"Pan, X., Ren, Y., Sheng, K., Dong, W., Yuan, H., Guo, X., Ma, C., and Xu, C. (2020, January 13\u201319). Dynamic refinement network for oriented and densely packed object detection. Proceedings of the CVPR, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01122"},{"key":"ref_89","doi-asserted-by":"crossref","unstructured":"Gupta, A., Vedaldi, A., and Zisserman, A. (2016, January 27\u201330). Synthetic data for text localisation in natural images. Proceedings of the CVPR, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.254"},{"key":"ref_90","doi-asserted-by":"crossref","unstructured":"Karatzas, D., Gomez-Bigorda, L., Nicolaou, A., Ghosh, S., Bagdanov, A., Iwamura, M., Matas, J., Neumann, L., Chandrasekhar, V.R., and Lu, S. (2015, January 23\u201326). ICDAR 2015 competition on robust reading. Proceedings of the ICDAR, Tunis, Tunisia.","DOI":"10.1109\/ICDAR.2015.7333942"},{"key":"ref_91","doi-asserted-by":"crossref","unstructured":"Nayef, N., Yin, F., Bizid, I., Choi, H., Feng, Y., Karatzas, D., Luo, Z., Pal, U., Rigaud, C., and Chazalon, J. (2017, January 9\u201315). Icdar2017 robust reading challenge on multi-lingual scene text detection and script identification-rrc-mlt. Proceedings of the ICDAR, Kyoto, Japan.","DOI":"10.1109\/ICDAR.2017.237"},{"key":"ref_92","unstructured":"Yao, C., Bai, X., Liu, W., Ma, Y., and Tu, Z. (2012, January 16\u201321). Detecting texts of arbitrary orientations in natural images. Proceedings of the CVPR, Providence, RI, USA."},{"key":"ref_93","doi-asserted-by":"crossref","unstructured":"He, M., Liu, Y., Yang, Z., Zhang, S., Luo, C., Gao, F., Zheng, Q., Wang, Y., Zhang, X., and Jin, L. (2018, January 20\u201324). ICPR2018 contest on robust reading for multi-type web images. Proceedings of the ICPR, Beijing, China.","DOI":"10.1109\/ICPR.2018.8546143"},{"key":"ref_94","doi-asserted-by":"crossref","unstructured":"Ch\u2019ng, C.K., and Chan, C.S. (2017, January 9\u201315). Total-text: A comprehensive dataset for scene text detection and recognition. Proceedings of the ICDAR, Kyoto, Japan.","DOI":"10.1109\/ICDAR.2017.157"},{"key":"ref_95","unstructured":"Yuliang, L., Lianwen, J., Shuaitao, Z., and Sheng, Z. (2017). Detecting curve text in the wild: New dataset and new solution. arXiv."},{"key":"ref_96","doi-asserted-by":"crossref","first-page":"183","DOI":"10.1016\/j.isprsjprs.2020.06.003","article-title":"A deeply supervised image fusion network for change detection in high resolution bi-temporal remote sensing images","volume":"166","author":"Zhang","year":"2020","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_97","doi-asserted-by":"crossref","first-page":"26661","DOI":"10.1007\/s11042-020-09294-7","article-title":"Remote sensing image caption generation via transformer and reinforcement learning","volume":"79","author":"Shen","year":"2020","journal-title":"Multi. Tools Appl."},{"key":"ref_98","first-page":"6506605","article-title":"Remote-Sensing Image Captioning Based on Multilayer Aggregated Transformer","volume":"19","author":"Liu","year":"2022","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_99","doi-asserted-by":"crossref","unstructured":"Ren, Z., Gou, S., Guo, Z., Mao, S., and Li, R. (2022). A Mask-Guided Transformer Network with Topic Token for Remote Sensing Image Captioning. Remote Sens., 14.","DOI":"10.3390\/rs14122939"},{"key":"ref_100","first-page":"5615611","article-title":"Transformer-Based Multistage Enhancement for Remote Sensing Image Super-Resolution","volume":"60","author":"Lei","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_101","first-page":"905","article-title":"A Super-resolution Method of Remote Sensing Image Using Transformers","volume":"2","author":"Ye","year":"2021","journal-title":"IDAACS"},{"key":"ref_102","doi-asserted-by":"crossref","first-page":"1373","DOI":"10.1109\/JSTARS.2022.3143532","article-title":"TR-MISR: Multiimage Super-Resolution Based on Feature Fusion with Transformers","volume":"15","author":"An","year":"2022","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_103","first-page":"5604816","article-title":"A deeply supervised attention metric-based network and an open aerial image dataset for remote sensing change detection","volume":"60","author":"Shi","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_104","doi-asserted-by":"crossref","unstructured":"Daudt, R.C., Le Saux, B., Boulch, A., and Gousseau, Y. (2018, January 22\u201327). Urban change detection for multispectral earth observation using convolutional neural networks. Proceedings of the IEEE International Geoscience and Remote Sensing Symposium, Valencia, Spain.","DOI":"10.1109\/IGARSS.2018.8518015"},{"key":"ref_105","doi-asserted-by":"crossref","first-page":"102783","DOI":"10.1016\/j.cviu.2019.07.003","article-title":"Multitask learning for large-scale semantic change detection","volume":"187","author":"Daudt","year":"2019","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_106","doi-asserted-by":"crossref","unstructured":"Shen, L., Lu, Y., Chen, H., Wei, H., Xie, D., Yue, J., Chen, R., Lv, S., and Jiang, B. (2021). S2Looking: A satellite side-looking dataset for building change detection. Remote Sens., 13.","DOI":"10.3390\/rs13245094"},{"key":"ref_107","unstructured":"(2022, August 27). Barley Remote Sensing Dataset. Available online: https:\/\/tianchi.aliyun.com\/dataset\/dataDetail?dataId=74952."},{"key":"ref_108","doi-asserted-by":"crossref","unstructured":"Maggiori, E., Tarabalka, Y., Charpiat, G., and Alliez, P. (2017, January 23\u201328). Can semantic labeling methods generalize to any city? The inria aerial image labeling benchmark. Proceedings of the IGARSS, Fort Worth, TX, USA.","DOI":"10.1109\/IGARSS.2017.8127684"},{"key":"ref_109","doi-asserted-by":"crossref","first-page":"2183","DOI":"10.1109\/TGRS.2017.2776321","article-title":"Exploring models and data for remote sensing image caption generation","volume":"56","author":"Lu","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_110","unstructured":"(2022, August 27). MEGA. Available online: https:\/\/mega.nz\/folder\/wCpSzSoS#RXzIlrv\u2013TDt3ENZdKN8JA."},{"key":"ref_111","unstructured":"(2022, August 27). MEGA. Available online: https:\/\/mega.nz\/folder\/pG4yTYYA#4c4buNFLibryZnlujsrwEQ."},{"key":"ref_112","doi-asserted-by":"crossref","first-page":"387","DOI":"10.1007\/s42064-019-0059-8","article-title":"Super-resolution of PROBA-V images using convolutional neural networks","volume":"3","author":"Izzo","year":"2019","journal-title":"Astrodynamics"},{"key":"ref_113","unstructured":"(2022, August 27). Available online: http:\/\/weegee.vision.ucmerced.edu\/datasets\/landuse.html."},{"key":"ref_114","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1109\/TGRS.2019.2934760","article-title":"HSI-BERT: Hyperspectral image classification using the bidirectional encoder representation from transformers","volume":"58","author":"He","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_115","first-page":"5514715","article-title":"Spectral-spatial transformer network for hyperspectral image classification: A factorized architecture search framework","volume":"60","author":"Zhong","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_116","doi-asserted-by":"crossref","first-page":"5522214","DOI":"10.1109\/TGRS.2022.3221534","article-title":"Spectral\u2013Spatial Feature Tokenization Transformer for Hyperspectral Image Classification","volume":"60","author":"Sun","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_117","unstructured":"Roy, S.K., Deria, A., Hong, D., Rasti, B., Plaza, A., and Chanussot, J. (2022). Multimodal fusion transformer for remote sensing image classification. arXiv."},{"key":"ref_118","doi-asserted-by":"crossref","first-page":"3095","DOI":"10.1109\/TIP.2022.3162964","article-title":"Deep Hierarchical Vision Transformer for Hyperspectral and LiDAR Data Classification","volume":"31","author":"Xue","year":"2022","journal-title":"IEEE Trans. Image Process."},{"key":"ref_119","first-page":"258619","article-title":"Deep Convolutional Neural Networks for Hyperspectral Image Classification","volume":"2015","author":"Hu","year":"2015","journal-title":"Sensors"},{"key":"ref_120","doi-asserted-by":"crossref","first-page":"844","DOI":"10.1109\/TGRS.2016.2616355","article-title":"Hyperspectral Image Classification Using Deep Pixel-Pair Features","volume":"2","author":"Li","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_121","doi-asserted-by":"crossref","unstructured":"Zhang, F., Zhang, K., and Sun, J. (2022). Multiscale Spatial\u2013Spectral Interaction Transformer for Pan-Sharpening. Remote Sens., 14.","DOI":"10.3390\/rs14071736"},{"key":"ref_122","doi-asserted-by":"crossref","unstructured":"Li, S., Guo, Q., and Li, A. (2022). Pan-Sharpening Based on CNN+ Pyramid Transformer by Using No-Reference Loss. Remote Sens., 14.","DOI":"10.3390\/rs14030624"},{"key":"ref_123","doi-asserted-by":"crossref","first-page":"5512805","DOI":"10.1109\/LGRS.2022.3170904","article-title":"PMACNet: Parallel Multiscale Attention Constraint Network for Pan-Sharpening","volume":"19","author":"Liang","year":"2022","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_124","first-page":"5407423","article-title":"Transformer-Based Regression Network for Pansharpening Remote Sensing Images","volume":"60","author":"Su","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_125","doi-asserted-by":"crossref","first-page":"3553","DOI":"10.1609\/aaai.v36i3.20267","article-title":"Pan-Sharpening with Customized Transformer and Invertible Neural Network","volume":"36","author":"Zhou","year":"2022","journal-title":"AAAI"},{"key":"ref_126","doi-asserted-by":"crossref","unstructured":"Bandara, W., and Patel, V. (2022, January 19\u201324). HyperTransformer: A Textural and Spectral Feature Fusion Transformer for Pansharpening. Proceedings of the CVPR, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00181"},{"key":"ref_127","unstructured":"(2022, August 27). 220 Band AVIRIS Hyperspectral Image Data Set: June 12, 1992 Indian Pine Test Site 3. Available online: https:\/\/purr.purdue.edu\/publications\/1947\/1."},{"key":"ref_128","unstructured":"(2022, August 27). Available online: https:\/\/www.ehu.eus\/ccwintco\/index.php\/Hyperspectral_Remote_Sensing_Scenes#Pavia_Centre_and_University."},{"key":"ref_129","unstructured":"(2022, August 27). Available online: https:\/\/hyperspectral.ee.uh.edu\/?page_id=459."},{"key":"ref_130","unstructured":"(2022, August 27). Available online: https:\/\/www.ehu.eus\/ccwintco\/index.php\/Hyperspectral_Remote_Sensing_Scenes#Salinas."},{"key":"ref_131","unstructured":"Gader, P., Zare, A., Close, R., Aitken, J., and Tuell, G. (2013). Muufl Gulfport Hyperspectral and Lidar Airborne Data Set, University of Florida. Technical Report REP-2013-570."},{"key":"ref_132","unstructured":"(2022, August 27). Hyperspectral Image Analysis Lab. Available online: https:\/\/hyperspectral.ee.uh.edu\/?page_id=1075."},{"key":"ref_133","unstructured":"(2022, August 27). Pavia Centre Scene. Available online: https:\/\/www.ehu.eus\/ccwintco\/index.php\/Hyperspectral_Remote_Sensing_Scenes#Pavia_Centre_scene."},{"key":"ref_134","doi-asserted-by":"crossref","unstructured":"Zhou, H., Liu, Q., and Wang, Y. (2022). PanFormer: A Transformer Based Model for Pan-sharpening. arXiv.","DOI":"10.1109\/ICME52920.2022.9859770"},{"key":"ref_135","unstructured":"(2022, August 27). WorldView-2 Full Archive and Tasking. Available online: https:\/\/earth.esa.int\/eogateway\/catalog\/worldview-2-full-archive-and-tasking."},{"key":"ref_136","unstructured":"(2022, August 27). WorldView-3 Full Archive and Tasking. Available online: https:\/\/earth.esa.int\/eogateway\/catalog\/worldview-3-full-archive-and-tasking."},{"key":"ref_137","unstructured":"(2022, August 27). Botswana. Available online: https:\/\/www.ehu.eus\/ccwintco\/index.php\/Hyperspectral_Remote_Sensing_Scenes#Botswana."},{"key":"ref_138","unstructured":"Yokoya, N., and Iwasaki, A. (2016). Airborne Hyperspectral Data over Chikusei, Space Application Laboratory, University of Tokyo. Technical Report."},{"key":"ref_139","unstructured":"(2022, August 27). Pleiades. Available online: https:\/\/pleiades.stoa.org\/downloads."},{"key":"ref_140","unstructured":"(2022, August 27). QuickBird Full Archive. Available online: https:\/\/earth.esa.int\/eogateway\/catalog\/quickbird-full-archive."},{"key":"ref_141","doi-asserted-by":"crossref","first-page":"5219715","DOI":"10.1109\/TGRS.2021.3137383","article-title":"Exploring Vision Transformers for Polarimetric SAR Image Classification","volume":"60","author":"Dong","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_142","first-page":"4505405","article-title":"High Resolution SAR Image Classification Using Global-Local Network Structure Based on Vision Transformer and CNN","volume":"19","author":"Liu","year":"2022","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_143","doi-asserted-by":"crossref","unstructured":"Cai, J., Zhang, Y., Guo, J., Zhao, X., Lv, J., and Hu, Y. (2022). ST-PN: A Spatial Transformed Prototypical Network for Few-Shot SAR Image Classification. Remote Sens., 14.","DOI":"10.3390\/rs14092019"},{"key":"ref_144","doi-asserted-by":"crossref","unstructured":"Ke, X., Zhang, X., and Zhang, T. (2022). GCBANet: A Global Context Boundary-Aware Network for SAR Ship Instance Segmentation. Remote Sens., 14.","DOI":"10.3390\/rs14092165"},{"key":"ref_145","doi-asserted-by":"crossref","unstructured":"Xia, R., Chen, J., Huang, Z., Wan, H., Wu, B., Sun, L., Yao, B., Xiang, H., and Xing, M. (2022). CRTransSar: A Visual Transformer Based on Contextual Joint Representation Learning for SAR Ship Detection. Remote Sens., 14.","DOI":"10.3390\/rs14061488"},{"key":"ref_146","first-page":"1","article-title":"Geospatial transformer is what you need for aircraft detection in SAR Imagery","volume":"60","author":"Chen","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_147","doi-asserted-by":"crossref","unstructured":"Zhang, P., Xu, H., Tian, T., Gao, P., and Tian, J. (2022). SFRE-Net: Scattering Feature Relation Enhancement Network for Aircraft Detection in SAR Images. Remote Sens., 14.","DOI":"10.3390\/rs14092076"},{"key":"ref_148","first-page":"5217619","article-title":"End-to-End Method with Transformer for 3D Detection of Oil Tank from Single SAR Image","volume":"60","author":"Ma","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_149","doi-asserted-by":"crossref","unstructured":"Perera, M., Bandara, W., Valanarasu, J., and Patel, V. (2022). Transformer-based SAR Image Despeckling. arXiv.","DOI":"10.1109\/IGARSS46834.2022.9884596"},{"key":"ref_150","doi-asserted-by":"crossref","unstructured":"Dong, H., Ma, W., Jiao, L., Liu, F., Shang, R., Li, Y., and Bai, J. (2022). A Contrastive Learning Transformer for Change Detection in High-Resolution SAR Images, SSRN. SSRN 4169439.","DOI":"10.2139\/ssrn.4169439"},{"key":"ref_151","doi-asserted-by":"crossref","unstructured":"Fan, Y., Wang, F., and Wang, H. (2022). A Transformer-Based Coarse-to-Fine Wide-Swath SAR Image Registration Method under Weak Texture Conditions. Remote Sens., 14.","DOI":"10.3390\/rs14051175"},{"key":"ref_152","unstructured":"Norikane, L., Broek, B., and Freeman, A. (1992, January 1\u20135). Application of modified VICAR\/IBIS GIS to analysis of July 1991 Flevoland AIRSAR data. Proceedings of the AIRSAR Workshop, Pasadena, CA, USA."},{"key":"ref_153","unstructured":"(2022, August 27). E-SAR\u2014The Airborne SAR System of DLR. Available online: https:\/\/www.dlr.de\/hr\/en\/desktopdefault.aspx\/tabid-2326\/3776_read-5679\/."},{"key":"ref_154","unstructured":"(2022, August 27). Available online: https:\/\/ietr-lab.univ-rennes1.fr\/polsarpro-bio\/san-francisco\/dataset\/SAN_FRANCISCO_AIRSAR.zip."},{"key":"ref_155","unstructured":"(2022, August 27). Use Data. Available online: https:\/\/www.eorc.jaxa.jp\/ALOS\/en\/alos-2\/a2_data_e.htm."},{"key":"ref_156","unstructured":"(2022, August 27). GF-3 (Gaofen-3). Available online: https:\/\/directory.eoportal.org\/web\/eoportal\/satellite-missions\/g\/gaofen-3."},{"key":"ref_157","unstructured":"(2022, August 27). F-SAR\u2014The New Airborne SAR System. Available online: https:\/\/www.dlr.de\/hr\/en\/desktopdefault.aspx\/tabid-2326\/3776_read-5691\/."},{"key":"ref_158","unstructured":"(2022, August 27). MSTAR Overview. Available online: https:\/\/www.sdms.afrl.af.mil\/index.php?collection=mstar."},{"key":"ref_159","doi-asserted-by":"crossref","unstructured":"Li, J., Qu, C., and Shao, J. (2017, January 3\u201314). Ship detection in SAR images based on an improved faster R-CNN. Proceedings of the BIGSARDATA, Beijing, China.","DOI":"10.1109\/BIGSARDATA.2017.8124934"},{"key":"ref_160","doi-asserted-by":"crossref","first-page":"120234","DOI":"10.1109\/ACCESS.2020.3005861","article-title":"HRSID: A High-Resolution SAR Images Dataset for Ship Detection and Instance Segmentation","volume":"8","author":"Wei","year":"2020","journal-title":"IEEE Access"},{"key":"ref_161","unstructured":"(2022, August 27). CryoSat Products. Available online: https:\/\/earth.esa.int\/eogateway\/catalog\/cryosat-products."},{"key":"ref_162","unstructured":"Martin, D., Fowlkes, C., Tal, D., and Malik, J. (2001, January 7\u201314). A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. Proceedings of the ICCV, Vancouver, BC, Canada."},{"key":"ref_163","unstructured":"(2022, August 27). TerraSAR-X ESA Archive. Available online: https:\/\/earth.esa.int\/eogateway\/catalog\/terrasar-x-esa-archive."},{"key":"ref_164","doi-asserted-by":"crossref","unstructured":"Li, Z., and Snavely, N. (2018, January 18\u201323). MegaDepth: Learning Single-View Depth Prediction from Internet Photos. Proceedings of the CVPR, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00218"},{"key":"ref_165","doi-asserted-by":"crossref","unstructured":"Dong, X., Bao, J., Chen, D., Zhang, W., Yu, N., Yuan, L., Chen, D., and Guo, B. (2022, January 19\u201324). CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows. Proceedings of the CVPR, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01181"},{"key":"ref_166","unstructured":"Mehta, S., and Rastegari, M. (2022, January 25). MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer. Proceedings of the ICLR, Virtual-Only."},{"key":"ref_167","unstructured":"Yanghao, L., Wu, C.Y., Fan, H., Mangalam, K., Xiong, B., Malik, J., and Feichtenhofer, C. (2022, January 19\u201324). MViTv2: Improved Multiscale Vision Transformers for Classification and Detection. Proceedings of the CVPR, New Orleans, LA, USA."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/7\/1860\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:07:37Z","timestamp":1760123257000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/7\/1860"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,30]]},"references-count":167,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2023,4]]}},"alternative-id":["rs15071860"],"URL":"https:\/\/doi.org\/10.3390\/rs15071860","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,30]]}}}