{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T13:57:59Z","timestamp":1784123879240,"version":"3.55.0"},"reference-count":48,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2022,4,7]],"date-time":"2022-04-07T00:00:00Z","timestamp":1649289600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&amp;D Program of China","doi-asserted-by":"publisher","award":["2021YFB3901300"],"award-info":[{"award-number":["2021YFB3901300"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["42071301"],"award-info":[{"award-number":["42071301"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["41671341"],"award-info":[{"award-number":["41671341"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Jiangsu Province Water Conservancy Science and Technology Project","award":["2021064"],"award-info":[{"award-number":["2021064"]}]},{"name":"Chongqing Agricultural Industry Digital Map Project","award":["21C00346"],"award-info":[{"award-number":["21C00346"]}]},{"name":"Postgraduate Research and Practice Innovation Program of Jiangsu Province","award":["KYCX21_1349"],"award-info":[{"award-number":["KYCX21_1349"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Semantic segmentation is a crucial approach for remote sensing interpretation. High-precision semantic segmentation results are obtained at the cost of manually collecting massive pixelwise annotations. Remote sensing imagery contains complex and variable ground objects and obtaining abundant manual annotations is expensive and arduous. The semi-supervised learning (SSL) strategy can enhance the generalization capability of a model with a small number of labeled samples. In this study, a novel semi-supervised adversarial semantic segmentation network is developed for remote sensing information extraction. A multiscale input convolution module (MICM) is designed to extract sufficient local features, while a Transformer module (TM) is applied for long-range dependency modeling. These modules are integrated to construct a segmentation network with a double-branch encoder. Additionally, a double-branch discriminator network with different convolution kernel sizes is proposed. The segmentation network and discriminator network are jointly trained under the semi-supervised adversarial learning (SSAL) framework to improve its segmentation accuracy in cases with small amounts of labeled data. Taking building extraction as a case study, experiments on three datasets with different resolutions are conducted to validate the proposed network. Semi-supervised semantic segmentation models, in which DeepLabv2, the pyramid scene parsing network (PSPNet), UNet and TransUNet are taken as backbone networks, are utilized for performance comparisons. The results suggest that the approach effectively improves the accuracy of semantic segmentation. The F1 and mean intersection over union (mIoU) accuracy measures are improved by 0.82\u201311.83% and 0.74\u20137.5%, respectively, over those of other methods.<\/jats:p>","DOI":"10.3390\/rs14081786","type":"journal-article","created":{"date-parts":[[2022,4,7]],"date-time":"2022-04-07T21:08:22Z","timestamp":1649365702000},"page":"1786","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":24,"title":["Semi-Supervised Adversarial Semantic Segmentation Network Using Transformer and Multiscale Convolution for High-Resolution Remote Sensing Imagery"],"prefix":"10.3390","volume":"14","author":[{"given":"Yalan","family":"Zheng","sequence":"first","affiliation":[{"name":"Key Laboratory of Virtual Geographic Environment (Nanjing Normal University), Ministry of Education, Nanjing 210023, China"},{"name":"School of Geography, Nanjing Normal University, Nanjing 210023, China"},{"name":"Jiangsu Center for Collaborative Innovation in Geographical Information Resource Development and Application, Nanjing 210023, China"},{"name":"State Key Laboratory Cultivation Base of Geographical Environment Evolution (Jiangsu Province), Nanjing 210023, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mengyuan","family":"Yang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Virtual Geographic Environment (Nanjing Normal University), Ministry of Education, Nanjing 210023, China"},{"name":"School of Geography, Nanjing Normal University, Nanjing 210023, China"},{"name":"Jiangsu Center for Collaborative Innovation in Geographical Information Resource Development and Application, Nanjing 210023, China"},{"name":"State Key Laboratory Cultivation Base of Geographical Environment Evolution (Jiangsu Province), Nanjing 210023, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Min","family":"Wang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Virtual Geographic Environment (Nanjing Normal University), Ministry of Education, Nanjing 210023, China"},{"name":"School of Geography, Nanjing Normal University, Nanjing 210023, China"},{"name":"Jiangsu Center for Collaborative Innovation in Geographical Information Resource Development and Application, Nanjing 210023, China"},{"name":"State Key Laboratory Cultivation Base of Geographical Environment Evolution (Jiangsu Province), Nanjing 210023, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaojun","family":"Qian","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, Nanjing Normal University, Nanjing 210097, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rui","family":"Yang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Virtual Geographic Environment (Nanjing Normal University), Ministry of Education, Nanjing 210023, China"},{"name":"School of Geography, Nanjing Normal University, Nanjing 210023, China"},{"name":"Jiangsu Center for Collaborative Innovation in Geographical Information Resource Development and Application, Nanjing 210023, China"},{"name":"State Key Laboratory Cultivation Base of Geographical Environment Evolution (Jiangsu Province), Nanjing 210023, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0394-7972","authenticated-orcid":false,"given":"Xin","family":"Zhang","sequence":"additional","affiliation":[{"name":"Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100101, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wen","family":"Dong","sequence":"additional","affiliation":[{"name":"Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100101, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,4,7]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"10548","DOI":"10.1109\/JSTARS.2021.3119286","article-title":"PiCoCo: Pixelwise Contrast and Consistency Learning for Semisupervised Building Footprint Segmentation","volume":"14","author":"Kang","year":"2021","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Su, Y., Cheng, J., Bai, H., Liu, H., and He, C. (2022). Semantic Segmentation of Very-High-Resolution Remote Sensing Images via Deep Multi-Feature Learning. Remote Sens., 14.","DOI":"10.3390\/rs14030533"},{"key":"ref_3","first-page":"640","article-title":"Fully convolutional networks for semantic segmentation","volume":"39","author":"Long","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1016\/j.isprsjprs.2017.05.002","article-title":"Simultaneous extraction of roads and buildings in remote sensing imagery with convolutional neural networks","volume":"130","author":"Alshehhi","year":"2017","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Li, Y., Lu, H., Liu, Q., Zhang, Y., and Liu, X. (2022). SSDBN: A Single-Side Dual-Branch Network with Encoder\u2013Decoder for Building Extraction. Remote Sens., 14.","DOI":"10.3390\/rs14030768"},{"key":"ref_6","first-page":"102499","article-title":"Multi-scale context extractor network for water-body extraction from high-resolution optical remotely sensed images","volume":"103","author":"Kang","year":"2021","journal-title":"Int. J. Appl. Earth Obs. Geoinf."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 21\u201326). Pyramid scene parsing network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., and Lu, H. (2019, January 15\u201320). Dual Attention Network for Scene Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, NY, USA.","DOI":"10.1109\/CVPR.2019.00326"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Tong, X., Xia, G., Lu, Q., Shen, H., Li, S., You, S., and Zhang, L. (2019). Land-Cover Classification with High-Resolution Remote Sensing Images Using Transferable Deep Models. arXiv, Available online: https:\/\/arxiv.org\/abs\/1807.05713.","DOI":"10.1016\/j.rse.2019.111322"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Zhang, M., Hu, X., Zhao, L., Lv, Y., and Luo, M. (2017). Learning dual multi-scale manifold ranking for semantic segmentation of high-resolution images. Remote Sens., 9.","DOI":"10.20944\/preprints201704.0061.v1"},{"key":"ref_12","unstructured":"Gerke, M., Rottensteiner, F., Wegner, J.D., and Sohn, G. (2014, September 07). ISPRS Semantic Labeling Contest. Available online: https:\/\/www.isprs.org\/education\/benchmarks\/UrbanSemLab\/2d-sem-label-potsdam.aspx."},{"key":"ref_13","first-page":"6214","article-title":"Low-shot learning for the semantic segmentation of remote sensing imagery","volume":"56","author":"Kemker","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_14","first-page":"102603","article-title":"Hyperspectral image classification on insufficient-sample and feature learning using deep neural networks: A review","volume":"105","author":"Wambugu","year":"2021","journal-title":"Int. J. Appl. Earth Obs. Geoinf."},{"key":"ref_15","unstructured":"Lee, D.H. (2013, January 16\u201321). Pseudo-Label: The Simple and Efficient Semi-Supervised Learning Method for Deep Neural Networks. Proceedings of the 30th International Conference on Machine Learning, Atlanta, GA, USA."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Qiao, S., Shen, W., Zhang, Z., Wang, B., and Yuille, A. (2018, January 8\u201314). Deep Co-Training for Semi-Supervised Image Recognition. Proceedings of the 15th European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01267-0_9"},{"key":"ref_17","unstructured":"Laine, S., and Aila, T. (2017). Temporal ensembling for semisupervised learning. arXiv, Available online: https:\/\/arxiv.org\/abs\/1610.02242."},{"key":"ref_18","unstructured":"Tarvainen, A., and Valpola, H. (2017). Mean teachers are better role models: Weight-averaged consistency targets improve semisupervised deep learning results. arXiv, Available online: https:\/\/arxiv.org\/abs\/1703.01780."},{"key":"ref_19","unstructured":"Berthelot, D., Carlini, N., Goodfellow, I., Oliver, A., Papernot, N., and Raffel, C. (2019). MixMatch: A holistic approach to semi-supervised learning. arXiv, Available online: https:\/\/arxiv.org\/abs\/1905.02249."},{"key":"ref_20","unstructured":"Sohn, K., Berthelot, D., Li, C., Zhang, Z., Carlini, N., Cubuk, E.D., Kurakin, A., Zhang, H., and Raffel, C. (2020). FixMatch: Simplifying semi-supervised learning with consistency and confidence. arXiv, Available online: https:\/\/arxiv.org\/abs\/2001.07685v2."},{"key":"ref_21","unstructured":"Odena, A. (2016). Semi-supervised learning with generative adversarial networks. arXiv."},{"key":"ref_22","first-page":"1","article-title":"CCS-GAN: A semi-supervised generative adversarial network for image classification","volume":"4","author":"Wang","year":"2021","journal-title":"Vis. Comput."},{"key":"ref_23","unstructured":"Luc, P., Couprie, C., Chintala, S., and Verbeek, J. (2016). Semantic segmentation using adversarial networks. arXiv, Available online: https:\/\/arxiv.org\/abs\/1611.08408."},{"key":"ref_24","unstructured":"Hung, W.C., Tsai, Y.H., Liou, Y.T., Lin, Y.Y., and Yang, M.H. (2018). Adversarial learning for semi-supervised semantic segmentation. arXiv, Available online: https:\/\/arxiv.org\/abs\/1802.07934."},{"key":"ref_25","unstructured":"Goodfellow, I.J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014). Generative adversarial networks. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zheng, S., Lu, J., Zhao, H., Zhu, X., and Zhang, L. (2020). Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. arXiv, Available online: https:\/\/arxiv.org\/abs\/2012.15840.","DOI":"10.1109\/CVPR46437.2021.00681"},{"key":"ref_27","first-page":"2341","article-title":"Adaboost-like End-to-End multiple lightweight U-nets for road extraction from optical remote sensing images","volume":"100","author":"Chen","year":"2021","journal-title":"Int. J. Appl. Earth Obs. Geoinf."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"2011","DOI":"10.1109\/TPAMI.2019.2913372","article-title":"Squeeze-and-excitation networks","volume":"42","author":"Hu","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021). Swin Transformer: Hierarchical vision transformer using shifted windows. arXiv.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Yang, F., Yang, H., Fu, J., Lu, H., and Guo, B. (2020, January 13\u201319). Learning texture transformer network for image super-resolution. Proceedings of the Conference on Computer Vision and Pattern Recognition, Seattle, DC, USA.","DOI":"10.1109\/CVPR42600.2020.00583"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wang, Z., Zhao, J., Zhang, R., Li, Z., Lin, Q., and Wang, X. (2022). UATNet: U-Shape Attention-Based Transformer Net for Meteorological Satellite Cloud Recognition. Remote Sens., 14.","DOI":"10.3390\/rs14010104"},{"key":"ref_33","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017, January 4\u20139). Attention is All you Need. Proceedings of the 31st Conference on Neural Information Processing Systems, Long Beach, NY, USA."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Liu, H., and Hu, Q. (2021). TransFuse: Fusing transformers and cnns for medical image segmentation. arXiv.","DOI":"10.1007\/978-3-030-87193-2_2"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"574","DOI":"10.1109\/TGRS.2018.2858817","article-title":"Fully convolutional networks for multi-source building extraction from an open aerial and satellite imagery dataset","volume":"57","author":"Ji","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_36","unstructured":"Mnih, V. (2013). Machine Learning for Aerial Image Labeling. [Ph.D. Dissertation, Department Computer Science]."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"1369","DOI":"10.1109\/TPAMI.2019.2960224","article-title":"Semi-supervised semantic segmentation with high- and low-level consistency","volume":"43","author":"Mittal","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"He, Y., Wang, J., Liao, C., Shan, B., and Zhou, X. (2022). ClassHyPer: ClassMix-Based Hybrid Perturbations for Deep Semi-Supervised Semantic Segmentation of Remote Sensing Imagery. Remote Sens., 14.","DOI":"10.3390\/rs14040879"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Souly, N., Spampinato, C., and Shah, M. (2017, January 22\u201329). Semi Supervised Semantic Segmentation Using Generative Adversarial Network. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.606"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Zhang, J., Li, Z., Zhang, C., and Ma, H. (2020, January 25\u201328). Robust Adversarial Learning for Semi-Supervised Semantic Segmentation. Proceedings of the IEEE International Conference on Image Processing, Abu Dhabi, United Arab Emirates.","DOI":"10.1109\/ICIP40778.2020.9190911"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"5398","DOI":"10.1109\/JSTARS.2020.3021098","article-title":"BAS4Net: Boundary-aware semi-supervised semantic segmentation network for very high resolution remote sensing images","volume":"13","author":"Sun","year":"2020","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"3492","DOI":"10.1109\/JSTARS.2019.2930724","article-title":"High-resolution aerial images semantic segmentation using deep fully convolutional network with channel attention mechanism","volume":"12","author":"Luo","year":"2019","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"4490","DOI":"10.1109\/JSTARS.2021.3073935","article-title":"Attention-guided label refinement network for semantic segmentation of very high resolution aerial orthoimages","volume":"14","author":"Huang","year":"2021","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_44","unstructured":"Chen, J., Lu, Y., Yu, Q., Luo, X., Adeli, E., Wang, Y., Lu, L., Yuille, L.A., and Zhou, Y. (2021). TransUNet: Transformers make strong encoders for medical image segmentation. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Hou, Q., Zhang, L., Cheng, M., and Feng, J. (2020, January 13\u201319). Strip Pooling: Rethinking Spatial Pooling for Scene Parsing. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00406"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-Net: Convolutional Networks for Biomedical Image Segmentation. Proceedings of the Medical Image Computing and Computer Assisted Intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_47","unstructured":"Kingma, D., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv, Available online: https:\/\/arxiv.org\/abs\/1412.6980."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Li, X., Wang, W., Hu, X., and Yang, J. (2019, January 15\u201320). Selective Kernel Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, NY, USA.","DOI":"10.1109\/CVPR.2019.00060"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/14\/8\/1786\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:50:05Z","timestamp":1760136605000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/14\/8\/1786"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,7]]},"references-count":48,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2022,4]]}},"alternative-id":["rs14081786"],"URL":"https:\/\/doi.org\/10.3390\/rs14081786","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,4,7]]}}}