{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T17:01:21Z","timestamp":1779382881926,"version":"3.53.1"},"reference-count":43,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2022,12,31]],"date-time":"2022-12-31T00:00:00Z","timestamp":1672444800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"China National Postdoctoral Program for Innovative Talents","award":["BX2021223"],"award-info":[{"award-number":["BX2021223"]}]},{"name":"China National Postdoctoral Program for Innovative Talents","award":["2021M702510"],"award-info":[{"award-number":["2021M702510"]}]},{"name":"China Postdoctoral Science Foundation","award":["BX2021223"],"award-info":[{"award-number":["BX2021223"]}]},{"name":"China Postdoctoral Science Foundation","award":["2021M702510"],"award-info":[{"award-number":["2021M702510"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Recently, deep learning has been widely used in the segmentation tasks of remote sensing images. However, the existing deep learning method most focus on local contextual information and has limited field of perception, which makes it difficult to capture the long-range contextual feature of objects at large scales form very-high-resolution (VHR) images. In this paper, we present a novel Local\u2013global Framework consisting of the dual-source fusion network and local\u2013global transformer modules, which efficiently utilize features extracted from multiple sources and fully capture features of local and global regions. The dual-source fusion network is an encoder designed to extract features from multiple sources such as spectra, synthetic aperture radar, and elevations, which selective fuse features from multiple sources and reduce the interference of redundant features. The local\u2013global transformer module is proposed to capture fine-grained local features and coarse-grained global features, which enables the framework to focus on recognizing multiple-scale objects from the local and global regions. Moreover, we propose a pixelwise contrastive loss, which could encourage that the prediction is pulled closer to the ground truth. The Local\u2013global Framework achieves state-of-the-art performance with 90.45% mean f1 score on the ISPRS Vaihingen dataset and 93.20% mean f1 score on the ISPRS Potsdam dataset.<\/jats:p>","DOI":"10.3390\/rs15010231","type":"journal-article","created":{"date-parts":[[2023,1,2]],"date-time":"2023-01-02T02:44:03Z","timestamp":1672627443000},"page":"231","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":20,"title":["A Local\u2013Global Framework for Semantic Segmentation of Multisource Remote Sensing Images"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8158-4604","authenticated-orcid":false,"given":"Luyi","family":"Qiu","sequence":"first","affiliation":[{"name":"School of Information and Software Engineering, University of Electronic Science and Technology of China, Chengdu 610054, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1720-8302","authenticated-orcid":false,"given":"Dayu","family":"Yu","sequence":"additional","affiliation":[{"name":"School of Remote Sensing and Information Engineering, Wuhan University, Wuhan 430070, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chenxiao","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Remote Sensing and Information Engineering, Wuhan University, Wuhan 430070, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaofeng","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, Shanghai 200240, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,12,31]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"He, J., Jia, X., Chen, S., and Liu, J. (2021, January 19\u201325). Multi-source domain adaptation with collaborative learning for semantic segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01086"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1029","DOI":"10.1109\/TGRS.2020.2999558","article-title":"Integrating multiresolution and multitemporal Sentinel-2 imagery for land-cover mapping in the Xiongan New Area, China","volume":"59","author":"Luo","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1016\/j.neucom.2019.10.118","article-title":"A comprehensive survey on support vector machine classification: Applications, challenges and trends","volume":"408","author":"Cervantes","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"191","DOI":"10.1016\/j.neucom.2019.01.108","article-title":"Learning longitudinal classification-regression model for infant hippocampus segmentation","volume":"391","author":"Guo","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"105596","DOI":"10.1016\/j.knosys.2020.105596","article-title":"A review of deep learning with special emphasis on architectures, applications and recent trends","volume":"194","author":"Sengupta","year":"2020","journal-title":"Knowl.-Based Syst."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"e5074","DOI":"10.1002\/cpe.5074","article-title":"Pansharpening multispectral remote-sensing images with guided filter for monitoring impact of human behavior on environment","volume":"33","author":"Li","year":"2021","journal-title":"Concurr. Comput. Pract. Exp."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2300","DOI":"10.1109\/TGRS.2002.803623","article-title":"Context-driven fusion of high spatial and spectral resolution images based on oversampled multiresolution analysis","volume":"40","author":"Aiazzi","year":"2002","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_8","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention Is All You Need. Proceedings of the Annual Conference on Neural Information Processing Systems, Long Beach City, CA, USA."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zhang, H., Dana, K., Shi, J., Zhang, Z., Wang, X., Tyagi, A., and Agrawal, A. (2018, January 18\u201323). Context encoding for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00747"},{"key":"ref_10","first-page":"1","article-title":"Multilevel Adaptive-Scale Context Aggregating Network for Semantic Segmentation in High-Resolution Remote Sensing Images","volume":"19","author":"Li","year":"2021","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Diao, Q., Dai, Y., Zhang, C., Wu, Y., Feng, X., and Pan, F. (2022). Superpixel-based attention graph neural network for semantic segmentation in aerial images. Remote Sens., 14.","DOI":"10.3390\/rs14020305"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Luo, A., Li, X., Yang, F., Jiao, Z., Cheng, H., and Lyu, S. (2020, January 23\u201328). Cascade graph neural networks for RGB-D salient object detection. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58610-2_21"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"476","DOI":"10.1016\/j.patcog.2017.11.024","article-title":"Material based salient object detection from hyperspectral images","volume":"76","author":"Liang","year":"2018","journal-title":"Pattern Recognit."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"108873","DOI":"10.1016\/j.patcog.2022.108873","article-title":"HCFNN: High-order coverage function neural network for image classification","volume":"131","author":"Ning","year":"2022","journal-title":"Pattern Recognit."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"108498","DOI":"10.1016\/j.patcog.2021.108498","article-title":"Uncertainty estimation for stereo matching based on evidential deep learning","volume":"124","author":"Wang","year":"2022","journal-title":"Pattern Recognit."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Xiong, D., He, C., Liu, X., and Liao, M. (2020). An end-to-end Bayesian segmentation network based on a generative adversarial network for remote sensing images. Remote Sens., 12.","DOI":"10.3390\/rs12020216"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Ren, Y., Yu, Y., and Guan, H. (2020). DA-CapsUNet: A dual-attention capsule U-Net for road extraction from remote sensing imagery. Remote Sens., 12.","DOI":"10.3390\/rs12182866"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"78","DOI":"10.1016\/j.isprsjprs.2017.12.007","article-title":"Semantic labeling in very high resolution images via a self-cascaded convolutional neural network","volume":"145","author":"Liu","year":"2018","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"7557","DOI":"10.1109\/TGRS.2020.2979552","article-title":"Relation matters: Relational context-aware fully convolutional network for semantic segmentation of high-resolution aerial images","volume":"58","author":"Mou","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"6106","DOI":"10.1109\/TGRS.2020.3022410","article-title":"Multiscale U-shaped CNN building instance extraction framework with edge constraint for high-spatial-resolution remote sensing imagery","volume":"59","author":"Liu","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_21","first-page":"1","article-title":"Hybrid multiple attention network for semantic segmentation in aerial images","volume":"60","author":"Niu","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"10015","DOI":"10.1109\/TGRS.2019.2930982","article-title":"CAD-Net: A context-aware detection network for objects in remote sensing imagery","volume":"57","author":"Zhang","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"310","DOI":"10.1109\/LGRS.2018.2872355","article-title":"Multiscale visual attention networks for object detection in VHR remote sensing images","volume":"16","author":"Wang","year":"2018","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Xing, Q., Xu, M., Li, T., and Guan, Z. (2020, January 23\u201328). Early exit or not: Resource-efficient blind quality enhancement for compressed images. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58517-4_17"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3484440","article-title":"Multi-stage fusion and multi-source attention network for multi-modal remote sensing image segmentation","volume":"12","author":"Zhao","year":"2021","journal-title":"ACM Trans. Intell. Syst. Technol. (TIST)"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"101842","DOI":"10.1016\/j.media.2020.101842","article-title":"Efficient and robust instrument segmentation in 3D ultrasound using patch-of-interest-FuseNet with hybrid loss","volume":"67","author":"Yang","year":"2021","journal-title":"Med. Image Anal."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_28","unstructured":"Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, L., Shazeer, N., Ku, A., and Tran, D. (2018, January 10\u201315). Image transformer. Proceedings of the International Conference on Machine Learning, PMLR, Stockholm, Sweden."},{"key":"ref_29","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_30","unstructured":"Chen, J., Lu, Y., Yu, Q., Luo, X., Adeli, E., Wang, Y., Lu, L., Yuille, A.L., and Zhou, Y. (2021). Transunet: Transformers make strong encoders for medical image segmentation. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020, January 13\u201319). Momentum contrast for unsupervised visual representation learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., Levine, S., and Brain, G. (2018, January 21\u201325). Time-contrastive networks: Self-supervised learning from video. Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia.","DOI":"10.1109\/ICRA.2018.8462891"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wu, H., Qu, Y., Lin, S., Zhou, J., Qiao, R., Zhang, Z., Xie, Y., and Ma, L. (2021, January 20\u201325). Contrastive learning for compact single image dehazing. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01041"},{"key":"ref_34","first-page":"21271","article-title":"Bootstrap your own latent-a new approach to self-supervised learning","volume":"33","author":"Grill","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_35","unstructured":"Henaff, O. (2020, January 3\u201318). Data-efficient image recognition with contrastive predictive coding. Proceedings of the International Conference on Machine Learning, PMLR, Virtual Event."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"158","DOI":"10.1016\/j.isprsjprs.2017.11.009","article-title":"Classification with an edge: Improving semantic image segmentation with boundary detection","volume":"135","author":"Marmanis","year":"2018","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"7503","DOI":"10.1109\/TGRS.2019.2913861","article-title":"Dynamic multicontext segmentation of remote sensing images based on convolutional networks","volume":"57","author":"Nogueira","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"297","DOI":"10.1016\/j.neucom.2018.11.051","article-title":"Problems of encoder-decoder frameworks for high-resolution remote sensing image segmentation: Structural stereotype and insufficient learning","volume":"330","author":"Sun","year":"2019","journal-title":"Neurocomputing"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"426","DOI":"10.1109\/TGRS.2020.2994150","article-title":"LANet: Local attention embedding to improve the semantic segmentation of remote sensing images","volume":"59","author":"Ding","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"2245","DOI":"10.1109\/TGRS.2020.3006872","article-title":"Class-guided feature decoupling network for airborne image segmentation","volume":"59","author":"Zhou","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"5260","DOI":"10.1109\/JSTARS.2021.3076035","article-title":"Boundary-Aware Dual-Stream Network for VHR Remote Sensing Images Semantic Segmentation","volume":"14","author":"Nong","year":"2021","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_42","first-page":"1","article-title":"Semantic Segmentation Network Using Local Relationship Upsampling for Remote Sensing Images","volume":"19","author":"Lin","year":"2021","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_43","first-page":"1","article-title":"Semantic segmentation for high-resolution remote-sensing images via dynamic graph context reasoning","volume":"19","author":"Su","year":"2022","journal-title":"IEEE Geosci. Remote Sens. Lett."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/1\/231\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:49:16Z","timestamp":1760147356000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/1\/231"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,31]]},"references-count":43,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,1]]}},"alternative-id":["rs15010231"],"URL":"https:\/\/doi.org\/10.3390\/rs15010231","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,31]]}}}