{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T00:16:35Z","timestamp":1783642595479,"version":"3.55.0"},"reference-count":69,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2023,5,29]],"date-time":"2023-05-29T00:00:00Z","timestamp":1685318400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2022YFF0711602"],"award-info":[{"award-number":["2022YFF0711602"]}]},{"name":"National Key Research and Development Program of China","award":["CAS-WX2021SF-0106"],"award-info":[{"award-number":["CAS-WX2021SF-0106"]}]},{"name":"14th Five-year Informatization Plan of Chinese Academy of Sciences","award":["2022YFF0711602"],"award-info":[{"award-number":["2022YFF0711602"]}]},{"name":"14th Five-year Informatization Plan of Chinese Academy of Sciences","award":["CAS-WX2021SF-0106"],"award-info":[{"award-number":["CAS-WX2021SF-0106"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Semantic segmentation with deep learning networks has become an important approach to the extraction of objects from very high-resolution remote sensing images. Vision Transformer networks have shown significant improvements in performance compared to traditional convolutional neural networks (CNNs) in semantic segmentation. Vision Transformer networks have different architectures to CNNs. Image patches, linear embedding, and multi-head self-attention (MHSA) are several of the main hyperparameters. How we should configure them for the extraction of objects in VHR images and how they affect the accuracy of networks are topics that have not been sufficiently investigated. This article explores the role of vision Transformer networks in the extraction of building footprints from very-high-resolution (VHR) images. Transformer-based models with different hyperparameter values were designed and compared, and their impact on accuracy was analyzed. The results show that smaller image patches and higher-dimension embeddings result in better accuracy. In addition, the Transformer-based network is shown to be scalable and can be trained with general-scale graphics processing units (GPUs) with comparable model sizes and training times to convolutional neural networks while achieving higher accuracy. The study provides valuable insights into the potential of vision Transformer networks in object extraction using VHR images.<\/jats:p>","DOI":"10.3390\/s23115166","type":"journal-article","created":{"date-parts":[[2023,5,29]],"date-time":"2023-05-29T07:49:23Z","timestamp":1685346563000},"page":"5166","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["Transformer-Based Semantic Segmentation for Extraction of Building Footprints from Very-High-Resolution Images"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9051-1925","authenticated-orcid":false,"given":"Jia","family":"Song","sequence":"first","affiliation":[{"name":"State Key Laboratory of Resources and Environmental Information System, Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences, Beijing 100101, China"},{"name":"Jiangsu Center for Collaborative Innovation in Geographical Information Resource Development and Application, Nanjing 210023, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"A-Xing","family":"Zhu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Resources and Environmental Information System, Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences, Beijing 100101, China"},{"name":"Department of Geography, University of Wisconsin, Madison, WI 53706, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yunqiang","family":"Zhu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Resources and Environmental Information System, Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences, Beijing 100101, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,5,29]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"114417","DOI":"10.1016\/j.eswa.2020.114417","article-title":"A review of deep learning methods for semantic segmentation of remote sensing imagery","volume":"169","author":"Yuan","year":"2021","journal-title":"Expert Syst. Appl."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1007\/s13735-017-0141-z","article-title":"A review of semantic segmentation using deep neural networks","volume":"7","author":"Guo","year":"2018","journal-title":"Int. J. Multimedia Inf. Retr."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Blaschke, T., Lang, S., and Hay, G.J. (2008). Object-Based Image Analysis: Spatial Concepts for Knowledge-Driven Remote Sensing Applications, Springer.","DOI":"10.1007\/978-3-540-77058-9"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Blaschke, T., and Lang, S.H.G.J. (2008). Object-Based Image Analysis: Spatial Concepts for Knowledge-Driven Remote Sensing Applications, Springer.","DOI":"10.1007\/978-3-540-77058-9"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"113","DOI":"10.1109\/JSTARS.2019.2953234","article-title":"Very High Resolution Remote Sensing Imagery Classification Using a Fusion of Random Forest and Deep Learning Technique\u2014Subtropical Area for Example","volume":"13","author":"Dong","year":"2020","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"112589","DOI":"10.1016\/j.rse.2021.112589","article-title":"Deep building footprint update network: A semi-supervised method for updating existing building footprint from bi-temporal remote sensing images","volume":"264","author":"Guo","year":"2021","journal-title":"Remote Sens. Environ."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"353","DOI":"10.1016\/j.isprsjprs.2021.03.016","article-title":"A Global Context-aware and Batch-independent Network for road extraction from VHR satellite imagery","volume":"175","author":"Zhu","year":"2021","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"240","DOI":"10.1016\/j.isprsjprs.2021.11.005","article-title":"A coarse-to-fine boundary refinement network for building footprint extraction from remote sensing imagery","volume":"183","author":"Guo","year":"2022","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1016\/j.isprsjprs.2021.12.007","article-title":"CMGFNet: A deep cross-modal gated fusion network for building extraction from very high-resolution remote sensing images","volume":"184","author":"Hosseinpour","year":"2022","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"200","DOI":"10.1007\/s11036-020-01703-3","article-title":"Convolutional Neural Network for the Semantic Segmentation of Remote Sensing Images","volume":"26","author":"Alam","year":"2021","journal-title":"Mob. Netw. Appl."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"4101","DOI":"10.1109\/JSTARS.2021.3068864","article-title":"A Pixel Cluster CNN and Spectral-Spatial Fusion Algorithm for Hyperspectral Image Classification with Small-Size Training Samples","volume":"14","author":"Dong","year":"2021","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Pan, X., and Zhao, J. (2018). High-Resolution Remote Sensing Image Classification Method Based on Convolutional Neural Network and Restricted Conditional Random Field. Remote Sens., 10.","DOI":"10.3390\/rs10060920"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1786","DOI":"10.1109\/LGRS.2020.3008051","article-title":"An End-to-End Hyperspectral Image Classification Method Using Deep Convolutional Neural Network with Spatial Constraint","volume":"18","author":"Jia","year":"2021","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"111322","DOI":"10.1016\/j.rse.2019.111322","article-title":"Land-cover classification with high-resolution remote sensing images using transferable deep models","volume":"237","author":"Tong","year":"2020","journal-title":"Remote Sens. Environ."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"L\u00e4ngkvist, M., Kiselev, A., Alirezaie, M., and Loutfi, A. (2016). Classification and Segmentation of Satellite Orthoimagery Using Convolutional Neural Networks. Remote Sens., 8.","DOI":"10.3390\/rs8040329"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"5573","DOI":"10.1080\/01431161.2020.1734251","article-title":"A deep residual learning serial segmentation network for extracting buildings from remote sensing imagery","volume":"41","author":"Liu","year":"2020","journal-title":"Int. J. Remote Sens."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1016\/j.isprsjprs.2022.01.022","article-title":"Estimating building height in China from ALOS AW3D30","volume":"185","author":"Huang","year":"2022","journal-title":"ISPRS-J. Photogramm. Remote Sens."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"2784","DOI":"10.1080\/01431161.2018.1433343","article-title":"Implementation of machine-learning classification in remote sensing: An applied review","volume":"39","author":"Maxwell","year":"2018","journal-title":"Int. J. Remote Sens."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"012010","DOI":"10.1088\/1755-1315\/620\/1\/012010","article-title":"Urban building detection using object-based image analysis (OBIA) and machine learning (ML) algorithms","volume":"620","author":"Norman","year":"2021","journal-title":"IOP Conf. Ser. Earth Environ. Sci."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"153","DOI":"10.3390\/rs70100153","article-title":"Comparing Machine Learning Classifiers for Object-Based Land Cover Classification Using Very High Resolution Imagery","volume":"7","author":"Qian","year":"2015","journal-title":"Remote Sens."},{"key":"ref_21","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is All You Need. Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_22","unstructured":"Conneau, A., and Lample, G. (2019, January 8\u201314). Cross-Lingual Language Model Pretraining. Proceedings of the 33rd International Conference on Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_23","unstructured":"Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019, January 2\u20137). BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MI, USA."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Yu, P., Fei, H., and Li, P. (2021, January 12\u201316). Cross-lingual Language Model Pretraining for Retrieval. Proceedings of the Web Conference, Ljubljana, Slovenia.","DOI":"10.1145\/3442381.3449830"},{"key":"ref_25","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3\u20137). An Image is Worth 16 \u00d7 16 Words: Transformers for Image Recognition at Scale. Proceedings of the International Conference on Learning Representations, Virtual Event."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"48","DOI":"10.1016\/j.neucom.2021.03.091","article-title":"A review on the attention mechanism of deep learning","volume":"452","author":"Niu","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Ghaffarian, S., Valente, J., van der Voort, M., and Tekinerdogan, B. (2021). Effect of Attention Mechanism in Deep Learning-Based Remote Sensing Image Processing: A Systematic Literature Review. Remote Sens., 13.","DOI":"10.3390\/rs13152965"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"29","DOI":"10.3389\/fncom.2020.00029","article-title":"Attention in Psychology, Neuroscience, and Machine Learning","volume":"14","author":"Lindsay","year":"2020","journal-title":"Front. Comput. Neurosci."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Chen, K., Zou, Z., and Shi, Z. (2021). Building Extraction from Remote Sensing Images with Sparse Token Transformers. Remote Sens., 13.","DOI":"10.3390\/rs13214441"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Chen, C.F.R., Fan, Q., and Panda, R. (2021, January 10\u201317). Crossvit: Cross-Attention Multi-Scale Vision Transformer for Image Classification. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00041"},{"key":"ref_31","first-page":"1","article-title":"A survey: Object detection methods from CNN to transformer","volume":"27","author":"Arkin","year":"2022","journal-title":"Multimedia Tools Appl."},{"key":"ref_32","unstructured":"Cao, F., and Lu, X. (2021, January 19\u201321). Self-Attention Technology in Image Segmentation. Proceedings of the International Conference on Intelligent Traffic Systems and Smart City, Zhengzhou, China."},{"key":"ref_33","first-page":"200","article-title":"Transformers in Vision: A Survey","volume":"54","author":"Khan","year":"2021","journal-title":"ACM Comput. Surv."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1109\/TPAMI.2022.3152247","article-title":"A Survey on Vision Transformer","volume":"45","author":"Han","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L. (2021, January 10\u201317). Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00061"},{"key":"ref_36","unstructured":"Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J.M., and Luo, P. (2021, January 6\u201314). SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers. Proceedings of the 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Montreal, QC, Canada."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2022, July 15). Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. Available online: https:\/\/arxiv.org\/abs\/2103.14030.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Dong, X., Bao, J., Chen, D., Zhang, W., Yu, N., Yuan, L., Chen, D., and Guo, B. (2022). CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows. arXiv.","DOI":"10.1109\/CVPR52688.2022.01181"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Bazi, Y., Bashmal, L., Al Rahhal, M.M., Al Dayil, R., and Al Ajlan, N. (2021). Vision Transformers for Remote Sensing Image Classification. Remote Sens., 13.","DOI":"10.3390\/rs13030516"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Reedha, R., Dericquebourg, E., Canals, R., and Hafiane, A. (2022). Transformer Neural Network for Weed and Crop Classification of High Resolution UAV Images. Remote Sens., 14.","DOI":"10.3390\/rs14030592"},{"key":"ref_41","unstructured":"Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jegou, H. (2021, January 18\u201324). Training Data-Efficient Image Transformers & Distillation through Attention. Proceedings of the 38th International Conference on Machine Learning, Virtual."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Avidan, S., Brostow, G., Ciss\u00e9, M., Farinella, G.M., and Hassner, T. (2022). Computer Vision\u2013ECCV 2022, Springer. Lecture Notes in Computer Science.","DOI":"10.1007\/978-3-031-19827-4"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1007\/s10462-010-9200-z","article-title":"Locally linear embedding: A survey","volume":"36","author":"Chen","year":"2011","journal-title":"Artif. Intell. Rev."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"6615","DOI":"10.1109\/TCSVT.2022.3176055","article-title":"Spatial-Temporal Based Multihead Self-Attention for Remote Sensing Image Change Detection","volume":"32","author":"Zhou","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z., Tay, F.E.H., Feng, J., and Yan, S. (2021, January 10\u201317). Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00060"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Yuan, W., and Xu, W. (2021). MSST-Net: A Multi-Scale Adaptive Network for Building Extraction from Remote Sensing Images Based on Swin Transformer. Remote Sens., 13.","DOI":"10.3390\/rs13234743"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"1422","DOI":"10.11834\/jrs.20210360","article-title":"Global-Local-Aware conditional random fields based building extraction for high spatial resolution remote sensing images","volume":"25","author":"Zhu","year":"2020","journal-title":"Natl. Remote Sens. Bull."},{"key":"ref_48","first-page":"102768","article-title":"Multi-scale attention integrated hierarchical networks for high-resolution building footprint extraction","volume":"109","author":"Liu","year":"2022","journal-title":"Int. J. Appl. Earth Obs."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"3308","DOI":"10.1080\/01431161.2018.1528024","article-title":"A scale robust convolutional neural network for automatic building extraction from aerial and satellite imagery","volume":"40","author":"Ji","year":"2018","journal-title":"Int. J. Remote Sens."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"16","DOI":"10.1109\/JSTARS.2010.2049478","article-title":"Improved Textural Built-Up Presence Index for Automatic Recognition of Human Settlements in Arid Regions with Scattered Vegetation","volume":"4","author":"Pesaresi","year":"2011","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1109\/LGRS.2009.2028744","article-title":"Urban Area Detection Using Local Feature Points and Spatial Voting","volume":"7","author":"Sirmacek","year":"2010","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"2078","DOI":"10.1109\/JSTARS.2015.2394504","article-title":"Cauchy Graph Embedding Optimization for Built-Up Areas Detection From High-Resolution Remote Sensing Images","volume":"8","author":"Li","year":"2015","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"487","DOI":"10.1109\/LGRS.2014.2347332","article-title":"An Efficient Approach for Automatic Rectangular Building Extraction From Very High Resolution Optical Satellite Imagery","volume":"12","author":"Wang","year":"2015","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"357","DOI":"10.1080\/17538947.2020.1831087","article-title":"Incorporating DeepLabv3+ and object-based image analysis for semantic segmentation of very high resolution remote sensing images","volume":"14","author":"Du","year":"2021","journal-title":"Int. J. Digit. Earth"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Chen, H., and Lu, S. (2019, January 5\u20137). Building Extraction from Remote Sensing Images Using SegNet. Proceedings of the 2019 IEEE 4th International Conference on Image, Vision and Computing (ICIVC), Xiamen, China.","DOI":"10.1109\/ICIVC47709.2019.8981046"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Chen, D.-Y., Peng, L., Li, W.-C., and Wang, Y.-D. (2021). Building Extraction and Number Statistics in WUI Areas Based on UNet Structure and Ensemble Learning. Remote Sens., 13.","DOI":"10.3390\/rs13061172"},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"645","DOI":"10.1109\/TGRS.2016.2612821","article-title":"Convolutional Neural Networks for Large-Scale Remote-Sensing Image Classification","volume":"55","author":"Maggiori","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Tong, Z., Li, Y., Li, Y., Fan, K., Si, Y., and He, L. (October, January 26). New Network Based on Unet++ and Densenet for Building Extraction from High Resolution Satellite Imagery. Proceedings of the IGARSS 2020\u20142020 IEEE International Geoscience and Remote Sensing Symposium, Waikoloa, HI, USA.","DOI":"10.1109\/IGARSS39084.2020.9324166"},{"key":"ref_59","doi-asserted-by":"crossref","first-page":"012145","DOI":"10.1088\/1742-6596\/1651\/1\/012145","article-title":"Building extraction from remote sensing image based on improved segnet neural network and image pyramid","volume":"1651","author":"Yu","year":"2020","journal-title":"J. Phys. Conf. Ser."},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Angelis, G.-E., Domi, A., Zamichos, A., Tsourma, M., Drosou, A., and Tzovaras, D. (2022, January 5). On The Exploration of Vision Transformers in Remote Sensing Building Extraction. Proceedings of the 2022 IEEE International Symposium on Multimedia (ISM), Naples, Italy.","DOI":"10.1109\/ISM55400.2022.00046"},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"369","DOI":"10.1109\/JSTARS.2022.3225150","article-title":"Improved Swin Transformer-Based Semantic Segmentation of Postearthquake Dense Buildings in Urban Areas Using Remote Sensing Images","volume":"16","author":"Cui","year":"2023","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_62","doi-asserted-by":"crossref","unstructured":"Yuan, W., Zhang, X., Shi, J., and Wang, J. (2023). LiteST-Net: A Hybrid Model of Lite Swin Transformer and Convolution for Building Extraction from Remote Sensing Image. Remote Sens., 15.","DOI":"10.3390\/rs15081996"},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Sun, Z., Zhou, W., Ding, C., and Xia, M. (2022). Multi-Resolution Transformer Network for Building and Road Segmentation of Remote Sensing Image. ISPRS Int. J. Geo-Inf., 11.","DOI":"10.3390\/ijgi11030165"},{"key":"ref_64","first-page":"1","article-title":"Building Extraction with Vision Transformer","volume":"60","author":"Wang","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Xiao, X., Guo, W., Chen, R., Hui, Y., Wang, J., and Zhao, H. (2022). A Swin Transformer-Based Encoding Booster Integrated in U-Shaped Network for Building Extraction. Remote Sens., 14.","DOI":"10.3390\/rs14112611"},{"key":"ref_66","unstructured":"Ba, J.L., Kiros, J.R., and Hinton, G.E. (2016). Layer Normalization. arXiv, Available online: http:\/\/arxiv.org\/abs\/1607.06450."},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Xiao, T., Liu, Y., Zhou, B., Jiang, Y., and Sun, J. (2018, January 8\u201314). Unified Perceptual Parsing for Scene Understanding. Proceedings of the Lecture Notes in Computer Science, Computer Vision\u2014ECCV 2018, Munich, Germany.","DOI":"10.1007\/978-3-030-01228-1_26"},{"key":"ref_68","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Dollar, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature Pyramid Networks for Object Detection. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_69","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1177\/001316446002000104","article-title":"A Coefficient of Agreement for Nominal Scales","volume":"20","author":"Cohen","year":"1960","journal-title":"Educ. Psychol. Meas."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/11\/5166\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:44:15Z","timestamp":1760125455000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/11\/5166"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,29]]},"references-count":69,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2023,6]]}},"alternative-id":["s23115166"],"URL":"https:\/\/doi.org\/10.3390\/s23115166","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,29]]}}}