{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T17:07:08Z","timestamp":1783444028656,"version":"3.54.6"},"reference-count":68,"publisher":"MDPI AG","issue":"17","license":[{"start":{"date-parts":[[2023,8,28]],"date-time":"2023-08-28T00:00:00Z","timestamp":1693180800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National NSF of China","doi-asserted-by":"publisher","award":["U19A2058"],"award-info":[{"award-number":["U19A2058"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National NSF of China","doi-asserted-by":"publisher","award":["41971362"],"award-info":[{"award-number":["41971362"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National NSF of China","doi-asserted-by":"publisher","award":["41871248"],"award-info":[{"award-number":["41871248"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National NSF of China","doi-asserted-by":"publisher","award":["62106276"],"award-info":[{"award-number":["62106276"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Road extraction from a remote sensing image is a research hotspot due to its broad range of applications. Despite recent advancements, achieving precise road extraction remains challenging. Since a road is thin and long, roadside objects and shadows cause occlusions, thus influencing the distinguishment of the road. Masked image modeling reconstructs masked areas from unmasked areas, which is similar to the process of inferring occluded roads from nonoccluded areas. Therefore, we believe that mask image modeling is beneficial for indicating occluded areas from other areas, thus alleviating the occlusion issue in remote sensing image road extraction. In this paper, we propose a remote sensing image road extraction network named RemainNet, which is based on mask image modeling. RemainNet consists of a backbone, image prediction module, and semantic prediction module. An image prediction module reconstructs a masked area RGB value from unmasked areas. Apart from reconstructing original remote sensing images, a semantic prediction module of RemainNet also extracts roads from masked images. Extensive experiments are carried out on the Massachusetts Roads dataset and DeepGlobe Road Extraction dataset; the proposed RemainNet improves 0.82\u20131.70% IoU compared with other state-of-the-art road extraction methods.<\/jats:p>","DOI":"10.3390\/rs15174215","type":"journal-article","created":{"date-parts":[[2023,8,28]],"date-time":"2023-08-28T05:46:47Z","timestamp":1693201607000},"page":"4215","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["RemainNet: Explore Road Extraction from Remote Sensing Image Using Mask Image Modeling"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8111-6803","authenticated-orcid":false,"given":"Zhenghong","family":"Li","sequence":"first","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7880-3394","authenticated-orcid":false,"given":"Hao","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China"},{"name":"Key Laboratory of Natural Resources Monitoring and Supervision in Southern Hilly Region, Ministry of Natural Resources, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ning","family":"Jing","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China"},{"name":"Key Laboratory of Natural Resources Monitoring and Supervision in Southern Hilly Region, Ministry of Natural Resources, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jun","family":"Li","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,8,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Abdollahi, A., Pradhan, B., Shukla, N., Chakraborty, S., and Alamri, A. (2020). Deep learning approaches applied to remote sensing datasets for road extraction: A state-of-the-art review. Remote Sens., 12.","DOI":"10.3390\/rs12091444"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zi, W., Xiong, W., Chen, H., Li, J., and Jing, N. (2021). SGA-Net: Self-constructing graph attention neural network for semantic segmentation of remote sensing images. Remote Sens., 13.","DOI":"10.3390\/rs13214201"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Song, J., Chen, H., Du, C., and Li, J. (2023). Semi-MapGen: Translation of Remote Sensing Image into Map via Semi-supervised Adversarial Learning. IEEE Trans. Geosci. Remote. Sens., 61.","DOI":"10.1109\/TGRS.2023.3263897"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"274","DOI":"10.1016\/j.ins.2021.01.065","article-title":"TAGCN: Station-level demand prediction for bike-sharing system via a temporal attention graph convolution network","volume":"561","author":"Zi","year":"2021","journal-title":"Inf. Sci."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"5489","DOI":"10.1109\/JSTARS.2020.3023549","article-title":"Road extraction methods in high-resolution remote sensing images: A comprehensive review","volume":"13","author":"Lian","year":"2020","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Feng, S., Ji, K., Wang, F., Zhang, L., Ma, X., and Kuang, G. (2023). PAN: Part Attention Network Integrating Electromagnetic Characteristics for Interpretable SAR Vehicle Target Recognition. IEEE Trans. Geosci. Remote Sens., 61.","DOI":"10.1109\/TGRS.2023.3256399"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Wu, S., Du, C., Chen, H., Xu, Y., Guo, N., and Jing, N. (2019). Road extraction from very high resolution images using weakly labeled OpenStreetMap centerline. ISPRS Int. J. Geo-Inf., 8.","DOI":"10.3390\/ijgi8110478"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Chen, H., Peng, S., Du, C., Li, J., and Wu, S. (2022). SW-GAN: Road Extraction from Remote Sensing Imagery Using Semi-Weakly Supervised Adversarial Learning. Remote Sens., 14.","DOI":"10.3390\/rs14174145"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"8540","DOI":"10.1109\/TIP.2021.3117076","article-title":"CoANet: Connectivity attention network for road extraction from satellite imagery","volume":"30","author":"Mei","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Wang, Y., Seo, J., and Jeon, T. (2021). NL-LinkNet: Toward lighter but more accurate road extraction with nonlocal operations. IEEE Geosci. Remote Sens. Lett., 19.","DOI":"10.1109\/LGRS.2021.3050477"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Chen, S.B., Ji, Y.X., Tang, J., Luo, B., Wang, W.Q., and Lv, K. (2021). DBRANet: Road extraction by dual-branch encoder and regional attention decoder. IEEE Geosci. Remote Sens. Lett., 19.","DOI":"10.1109\/LGRS.2021.3074524"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"786","DOI":"10.1109\/LGRS.2020.2985774","article-title":"Gated auxiliary edge detection task for road extraction with weight-balanced loss","volume":"18","author":"Li","year":"2020","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"169","DOI":"10.1016\/j.isprsjprs.2023.03.012","article-title":"SemiRoadExNet: A semi-supervised network for road extraction from remote sensing imagery via adversarial learning","volume":"198","author":"Chen","year":"2023","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"8919","DOI":"10.1109\/TGRS.2020.2991733","article-title":"Simultaneous road surface and centerline extraction from large-scale remote sensing images using CNN-based segmentation and tracing","volume":"58","author":"Wei","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Xu, Y., Chen, H., Du, C., and Li, J. (2021). MSACon: Mining spatial attention-based contextual information for road extraction. IEEE Trans. Geosci. Remote Sens., 60.","DOI":"10.1109\/TGRS.2021.3073923"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"10243","DOI":"10.1109\/TGRS.2020.3034011","article-title":"DiResNet: Direction-aware residual network for road extraction in VHR remote sensing images","volume":"59","author":"Ding","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Yang, Z., Zhou, D., Yang, Y., Zhang, J., and Chen, Z. (2022). Road Extraction From Satellite Imagery by Road Context and Full-Stage Feature. IEEE Geosci. Remote. Sens. Lett., 20.","DOI":"10.1109\/LGRS.2022.3228967"},{"key":"ref_18","unstructured":"Li, S., Wu, D., Wu, F., Zang, Z., Sun, B., Li, H., Xie, X., and Li, S. (2022). Architecture-Agnostic Masked Image Modeling\u2013From ViT back to CNN. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_20","unstructured":"Khan, S.H., Bennamoun, M., Sohel, F., and Togneri, R. (2014). European Conference on Computer Vision, Springer."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1007\/978-981-15-6014-9_27","article-title":"A Review on Image Segmentation","volume":"2020","author":"Jaiswal","year":"2021","journal-title":"Rising Threat. Expert Appl. Solut."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Yuan, X., Shi, J., and Gu, L. (2021). A review of deep learning methods for semantic segmentation of remote sensing imagery. Expert Syst. Appl., 169.","DOI":"10.1016\/j.eswa.2020.114417"},{"key":"ref_23","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2023, January 17\u201324). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada."},{"key":"ref_24","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the Medical Image Computing and Computer-Assisted Intervention\u2013MICCAI 2015: 18th International Conference, Munich, Germany. Proceedings, Part III 18."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. (2018, January 8\u201314). Encoder-decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Ulku, I., and Akag\u00fcnd\u00fcz, E. (2022). A survey on deep learning-based architectures for semantic segmentation on 2d images. Appl. Artif. Intell., 36.","DOI":"10.1080\/08839514.2022.2032924"},{"key":"ref_29","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S.A. (2021, January 3\u20137). An Image is Worth 16 \u00d7 16 Words: Transformers for Image Recognition at Scale. Proceedings of the International Conference on Learning Representations, Virtual Event."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Lv, P., Wu, W., Zhong, Y., and Zhang, L. (2022, January 17\u201322). Review of Vision Transformer Models for Remote Sensing Image Scene Classification. Proceedings of the IGARSS 2022\u20142022 IEEE International Geoscience and Remote Sensing Symposium, Kuala Lumpur, Malaysia.","DOI":"10.1109\/IGARSS46834.2022.9883054"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1109\/TPAMI.2022.3152247","article-title":"A Survey on Vision Transformer","volume":"45","author":"Han","year":"2023","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_32","unstructured":"Meila, M., and Zhang, T. (2021, January 18\u201324). Training data-efficient image transformers & distillation through attention. Proceedings of the 38th International Conference on Machine Learning, Virtual Event."},{"key":"ref_33","unstructured":"Chu, X., Tian, Z., Zhang, B., Wang, X., Wei, X., Xia, H., and Shen, C. (2021). Conditional Positional Encodings for Vision Transformers. arXiv."},{"key":"ref_34","unstructured":"Li, Y., Zhang, K., Cao, J., Timofte, R., and Gool, L.V. (2021). LocalViT: Bringing Locality to Vision Transformers. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Strudel, R., Garcia, R., Laptev, I., and Schmid, C. (2021, January 2\u20136). Segmenter: Transformer for Semantic Segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV48922.2021.00717"},{"key":"ref_36","unstructured":"Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J.W. (2021). Advances in Neural Information Processing Systems, IEEE."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1016\/j.patrec.2021.04.024","article-title":"TrSeg: Transformer for semantic segmentation","volume":"148","author":"Jin","year":"2021","journal-title":"Pattern Recognit. Lett."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Li, R., Duan, C., Zheng, S., Zhang, C., and Atkinson, P.M. (2021). MACU-Net for Semantic Segmentation of Fine-Resolution Remotely Sensed Images. IEEE Geosci. Remote. Sens. Lett., 19.","DOI":"10.1109\/LGRS.2021.3052886"},{"key":"ref_39","unstructured":"Wan, Q., Huang, Z., Lu, J., Yu, G., and Zhang, L. (2023). SeaFormer: Squeeze-enhanced Axial Transformer for Mobile Semantic Segmentation. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Yuan, F., Zhang, Z., and Fang, Z. (2023). An effective CNN and Transformer complementary network for medical image segmentation. Pattern Recognit., 136.","DOI":"10.1016\/j.patcog.2022.109228"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Chen, Z., Deng, L., Luo, Y., Li, D., Marcato Junior, J., Nunes Gon\u00e7alves, W., Awal Md Nurunnabi, A., Li, J., Wang, C., and Li, D. (2022). Road extraction in remote sensing data: A survey. Int. J. Appl. Earth Obs. Geoinf., 112.","DOI":"10.1016\/j.jag.2022.102833"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"1946","DOI":"10.1109\/JSTARS.2015.2449296","article-title":"Road extraction from very high resolution remote sensing optical images based on texture analysis and beamlet transform","volume":"9","author":"Sghaier","year":"2015","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_43","unstructured":"Wang, J., Qin, Q., Yang, X., Wang, J., Ye, X., and Qin, X. (2014, January 13\u201318). Automated road extraction from multi-resolution images using spectral information and texture. Proceedings of the 2014 IEEE Geoscience and Remote Sensing Symposium, Quebec City, QC, Canada."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"1373","DOI":"10.1109\/JSTARS.2012.2219614","article-title":"Road extraction from SAR imagery based on multiscale geometric analysis of detector responses","volume":"5","author":"He","year":"2012","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"709","DOI":"10.1109\/LGRS.2017.2672734","article-title":"Road structure refined CNN for road extraction in aerial image","volume":"14","author":"Wei","year":"2017","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"391","DOI":"10.1080\/12265934.2019.1596040","article-title":"Extraction of road features from UAV images using a novel level set segmentation approach","volume":"23","author":"Abdollahi","year":"2019","journal-title":"Int. J. Urban Sci."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Xin, J., Zhang, X., Zhang, Z., and Fang, W. (2019). Road extraction of high-resolution remote sensing images derived from DenseUNet. Remote Sens., 11.","DOI":"10.3390\/rs11212499"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"64381","DOI":"10.1109\/ACCESS.2021.3075951","article-title":"Improving road semantic segmentation using generative adversarial network","volume":"9","author":"Abdollahi","year":"2021","journal-title":"IEEE Access"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Zhou, L., Zhang, C., and Wu, M. (2018, January 8\u201323). D-LinkNet: LinkNet with pretrained encoder and dilated convolution for high resolution satellite imagery road extraction. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00034"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"423","DOI":"10.1007\/s12524-017-0702-x","article-title":"Investigation of SVM and level set interactive methods for road extraction from google earth images","volume":"46","author":"Abdollahi","year":"2018","journal-title":"J. Indian Soc. Remote Sens."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"155","DOI":"10.1016\/j.isprsjprs.2019.10.001","article-title":"Spatial information inference net: Road extraction using road-specific contextual information","volume":"158","author":"Tao","year":"2019","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Zhou, Q., Yu, C., Luo, H., Wang, Z., and Li, H. (2022, January 10\u201314). MimCo: Masked Image Modeling Pre-training with Contrastive Teacher. Proceedings of the 30th ACM International Conference on Multimedia, Lisboa, Portugal.","DOI":"10.1145\/3503161.3548173"},{"key":"ref_53","unstructured":"Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (2023, June 20). Improving Language Understanding by Generative Pre-Training; 2018; p. 12. Available online: https:\/\/s3-us-west-2.amazonaws.com\/openai-assets\/research-covers\/language-unsupervised\/language_understanding_paper.pdf."},{"key":"ref_54","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_55","unstructured":"Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I. (2020, January 13\u201318). Generative pretraining from pixels. Proceedings of the International Conference on Machine Learning, Virtual Event."},{"key":"ref_56","unstructured":"Bao, H., Dong, L., and Wei, F. (2021). BEiT: BERT Pre-Training of Image Transformers. arXiv."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Caron, M., Touvron, H., Misra, I., J\u00e9gou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021, January 10\u201317). Emerging properties in self-supervised vision transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"He, K., Chen, X., Xie, S., Li, Y., Doll\u00e1r, P., and Girshick, R. (2022, January 19\u201320). Masked autoencoders are scalable vision learners. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Zhang, C., Zhang, C., Song, J., Yi, J.S.K., Zhang, K., and Kweon, I.S. (2022). A Survey on Masked Autoencoder for Self-supervised Learning in Vision and Beyond. arXiv.","DOI":"10.24963\/ijcai.2023\/762"},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., and Hu, H. (2022, January 18\u201324). Simmim: A simple framework for masked image modeling. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00943"},{"key":"ref_61","unstructured":"Li, G., Zheng, H., Liu, D., Su, B., and Zheng, C. (2022). SemMAE: Semantic-Guided Masking for Learning Masked Autoencoders. arXiv."},{"key":"ref_62","doi-asserted-by":"crossref","unstructured":"Xue, H., Gao, P., Li, H., Qiao, Y., Sun, H., Li, H., and Luo, J. (2023, January 17\u201324). Stare at What You See: Masked Image Modeling Without Reconstruction. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02177"},{"key":"ref_63","unstructured":"Qi, G.J., and Shah, M. (2022). Adversarial Pretraining of Self-Supervised Deep Networks: Past, Present and Future. arXiv."},{"key":"ref_64","unstructured":"Chen, X., Ding, M., Wang, X., Xin, Y., Mo, S., Wang, Y., Han, S., Luo, P., Zeng, G., and Wang, J. (2022). Context Autoencoder for Self-Supervised Representation Learning. arXiv."},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Wei, C., Fan, H., Xie, S., Wu, C.Y., Yuille, A., and Feichtenhofer, C. (2022, January 18\u201324). Masked Feature Prediction for Self-Supervised Visual Pre-Training. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01426"},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"Chen, X., Liu, W., Liu, X., Zhang, Y., Han, J., and Mei, T. (2022, January 10\u201314). MAPLE: Masked Pseudo-Labeling AutoEncoder for Semi-Supervised Point Cloud Action Recognition. Proceedings of the 30th ACM International Conference on Multimedia, New York, NY, USA.","DOI":"10.1145\/3503161.3547892"},{"key":"ref_67","unstructured":"Mnih, V. (2013). Machine Learning for Aerial Image Labeling, University of Toronto."},{"key":"ref_68","doi-asserted-by":"crossref","unstructured":"Demir, I., Koperski, K., Lindenbaum, D., Pang, G., Huang, J., Basu, S., Hughes, F., Tuia, D., and Raskar, R. (2018, January 17\u201324). DeepGlobe 2018: A Challenge to Parse the Earth Through Satellite Images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, Vancouver, BC, Canada.","DOI":"10.1109\/CVPRW.2018.00031"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/17\/4215\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:40:36Z","timestamp":1760128836000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/17\/4215"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,28]]},"references-count":68,"journal-issue":{"issue":"17","published-online":{"date-parts":[[2023,9]]}},"alternative-id":["rs15174215"],"URL":"https:\/\/doi.org\/10.3390\/rs15174215","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,28]]}}}