{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,19]],"date-time":"2026-04-19T08:22:08Z","timestamp":1776586928557,"version":"3.51.2"},"reference-count":56,"publisher":"MDPI AG","issue":"18","license":[{"start":{"date-parts":[[2020,9,10]],"date-time":"2020-09-10T00:00:00Z","timestamp":1599696000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the Strategic Priority Research Program of Chinese Academy of Sciences","award":["XDA19040501"],"award-info":[{"award-number":["XDA19040501"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["41471376"],"award-info":[{"award-number":["41471376"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>It is challenging for semantic segmentation of buildings based on high-resolution remote sensing images, given high variability of appearance and complicated backgrounds of the buildings and their images. In this communication, we proposed an ensemble multi-scale residual deep learning method with the regularizer of shape representation for semantic segmentation of buildings. Based on the U-Net architecture using residual connections and multi-scale ASPP (atrous spatial pyramid pooling) modules, our method introduced the regularizer of shape representation and ensemble learning of multi-scale models to enhance model training and reduce over-fitting. In our method, the shape representation was coded in an antoencoder that was used to encode and reconstruct the shape characteristics of the buildings. In prediction, we consider multi-scale trained models for different resolution inputs and side effects to obtain an optimal semantic segmentation. With the high-resolution image of the Changshan, an island county in China, we used two-thirds of the study region image to train the model and the remaining one-third for the independent test. We obtained the accuracy of 0.98\u20130.99, mean intersection over union (MIoU) of 0.91\u20130.93 and Jaccard coefficient of 0.89\u20130.92 in validation. In the independent test, our method achieved state-of-the-art performance (MIoU: 0.83; Jaccard index: 0.81). By comparing with the existing representative methods on four different data sets, the proposed method consistently improved the learning process and generalization. The study shows important contributions of ensemble learning of multi-scale residual models and regularizer of shape representation to semantic segmentation of buildings.<\/jats:p>","DOI":"10.3390\/rs12182932","type":"journal-article","created":{"date-parts":[[2020,9,10]],"date-time":"2020-09-10T09:10:09Z","timestamp":1599729009000},"page":"2932","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["Multi-Scale Residual Deep Network for Semantic Segmentation of Buildings with Regularizer of Shape Representation"],"prefix":"10.3390","volume":"12","author":[{"given":"Chengyi","family":"Wang","sequence":"first","affiliation":[{"name":"National Engineering Research Center for Geomatics, Aerospace Information Research Institute, Chinese Academy of Sciences, Datun Road, Beijing 100101, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9382-8637","authenticated-orcid":false,"given":"Lianfa","family":"Li","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Resources and Environmental Information Systems, Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences, Datun Road, Beijing 100101, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,9,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Bischke, B., Helber, P., Folz, J., Borth, D., and Dengel, A. (2019, January 22\u201325). Multi-task learning for segmentation of building footprints with deep neural networks. Proceedings of the 2019 IEEE International Conference on Image Processing (ICIP), Taipei, Taiwan.","DOI":"10.1109\/ICIP.2019.8803050"},{"key":"ref_2","first-page":"724","article-title":"Object-based morphological building index for building extractionfrom high resolution remote sensing imagery","volume":"46","author":"Lin","year":"2017","journal-title":"Acta Geod. Cartogr. Sin."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Yi, Y.N., Zhang, Z.J., Zhang, W.C., Zhang, C.R., Li, W.D., and Zhao, T. (2019). Semantic Segmentation of Urban Buildings from VHR Remote Sensing Imagery Using a Deep Convolutional Neural Network. Remote Sens., 11.","DOI":"10.3390\/rs11151774"},{"key":"ref_4","first-page":"653","article-title":"A Survey of Building Extraction Methods from Optical High Resolution Remote Sensing Imagery","volume":"31","author":"Wang","year":"2016","journal-title":"Remote Sens. Technol. Appl."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2097","DOI":"10.1109\/TGRS.2008.916644","article-title":"Automatic detection of geospatial objects using multiple hierarchical segmentations","volume":"46","author":"Aksoy","year":"2008","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"180","DOI":"10.1016\/j.isprsjprs.2013.09.014","article-title":"Geographic object-based image analysis\u2014towards a new paradigm","volume":"87","author":"Blaschke","year":"2014","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1502","DOI":"10.3724\/SP.J.1004.2010.01502","article-title":"Towards Automatic Building Extraction: Variational Level Set Model Using Prior Shape Knowledge","volume":"36","author":"Tian","year":"2010","journal-title":"Acta Autom. Sin."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"721","DOI":"10.14358\/PERS.77.7.721","article-title":"A multidirectional and multiscale morphological index for automatic building extraction from multispectral GeoEye-1 imagery","volume":"77","author":"Huang","year":"2011","journal-title":"Photogramm. Eng. Remote Sens."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"180","DOI":"10.1109\/JSTARS.2008.2002869","article-title":"A robust built-up area presence index by anisotropic rotation-invariant textural measure","volume":"1","author":"Pesaresi","year":"2008","journal-title":"IEEE J. Select. Top. Appl. Earth Obser. Remote Sens."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"161","DOI":"10.1109\/JSTARS.2011.2168195","article-title":"Morphological building\/shadow index for building extraction from high-resolution imagery over urban areas","volume":"5","author":"Huang","year":"2011","journal-title":"IEEE J. Select. Top. Appl. Earth Obser. Remote Sens."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"641","DOI":"10.1109\/34.295913","article-title":"Seeded region growing","volume":"16","author":"Adams","year":"1994","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"603","DOI":"10.1109\/34.1000236","article-title":"Mean shift: A robust approach toward feature space analysis","volume":"24","author":"Comaniciu","year":"2002","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1145\/1015706.1015720","article-title":"Interactive foreground extraction using iterated graph cuts","volume":"23","author":"Rother","year":"2004","journal-title":"ACM Trans. Gr."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"2152","DOI":"10.1080\/01431161.2011.606852","article-title":"Unsupervised building detection in complex urban environments from multispectral satellite imagery","volume":"33","author":"Erener","year":"2012","journal-title":"Int. J. Remote Sens."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"3906","DOI":"10.1109\/TGRS.2011.2136381","article-title":"Use of salient features for the design of a multistage framework to extract roads from high-resolution multispectral satellite images","volume":"49","author":"Das","year":"2011","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1365","DOI":"10.14358\/PERS.70.12.1365","article-title":"Road extraction using SVM and image segmentation","volume":"70","author":"Song","year":"2004","journal-title":"Photogramm. Eng. Remote Sens."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Wang, Y., Song, H., and Zhang, Y. (2016). Spectral-spatial classification of hyperspectral images using joint bilateral filter and graph cut based model. Remote Sens., 8.","DOI":"10.20944\/preprints201608.0022.v1"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Tian, S., Zhang, X., Tian, J., and Sun, Q. (2016). Random forest classification of wetland landcovers from multi-sensor data in the arid region of Xinjiang, China. Remote Sens., 8.","DOI":"10.3390\/rs8110954"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Li, L.F. (2019). Deep Residual Autoencoder with Multiscaling for Semantic Segmentation of Land-Use Images. Remote Sens., 11.","DOI":"10.3390\/rs11182142"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1109\/MGRS.2016.2540798","article-title":"Deep Learning for Remote Sensing Data A technical tutorial on the state of the art","volume":"4","author":"Zhang","year":"2016","journal-title":"IEEE Geosci. Remote Sens. Mag."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1109\/MGRS.2017.2762307","article-title":"Deep Learning in Remote Sensing","volume":"5","author":"Zhu","year":"2017","journal-title":"IEEE Geosci. Remote Sens. Mag."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVRP), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1016\/j.neucom.2018.03.037","article-title":"Methods and datasets on semantic segmentation: A review","volume":"304","author":"Yu","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015). U-Net: Convolutional Networks for Biomedical Image Segmentation. arXiv.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_27","unstructured":"Chen, L.-C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A.L. (2014). Semantic image segmentation with deep convolutional nets and fully connected CRFs. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Lin, G., Milan, A., Shen, C., and Reid, I. (2017, January 21\u201326). Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.549"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Peng, C., Zhang, X., Yu, G., Luo, G., and Sun, J. (2017, January 21\u201326). Large Kernel Matters\u2014Improve Semantic Segmentation by Global Convolutional Network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.189"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. (2018, January 8\u201314). Encoder-decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"ref_32","unstructured":"Zuo, T. (2017). Research of Building Extraction Technology for High-Resolution Remote Sensing Images, University of Science and Technology of China. (In Chinese)."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"645","DOI":"10.1109\/TGRS.2016.2612821","article-title":"Convolutional neural networks for large-scale remote-sensing image classification","volume":"55","author":"Maggiori","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_34","first-page":"188","article-title":"Application of Convolutional Neural Netowrk using region information to remote sensing image classification","volume":"54","author":"Yang","year":"2018","journal-title":"Comput. Eng. Appl."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Qin, Y., Wu, Y., Li, B., Gao, S., Liu, M., and Zhan, Y. (2019). Semantic Segmentation of Building Roof in Dense Urban Environment with Deep Convolutional Neural Network: A Case Study Using GF2 VHR Imagery in China. Sensors, 19.","DOI":"10.3390\/s19051164"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"184","DOI":"10.1016\/j.isprsjprs.2019.11.004","article-title":"Building segmentation through a gated graph convolutional neural network with deep structured feature embedding","volume":"159","author":"Shi","year":"2020","journal-title":"ISPRS J. Photogramm."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"102897","DOI":"10.1016\/j.earscirev.2019.102897","article-title":"Principles and methods of scaling geospatial Earth science data","volume":"197","author":"Ge","year":"2019","journal-title":"Earth Sci. Rev."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Li, L., Fang, Y., Wu, J., Wang, C., and Ge, Y. (2020). Encoder-Decoder Full Residual Deep Networks for Robust Regression and Spatiotemporal Estimation. IEEE Trans. Nerual Netw. Learn. Syst.","DOI":"10.1109\/TNNLS.2020.3017200"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Garcia-Garcia, A., Orts-Escolano, S., Oprea, S., Villena-Martinez, V., and Garcia-Rodriguez, J. (2017). A review on deep learning techniques applied to semantic segmentation. arXiv.","DOI":"10.1016\/j.asoc.2018.05.018"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"He, K.M., Zhang, X.Y., Ren, S.Q., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"630","DOI":"10.1007\/978-3-319-46493-0_38","article-title":"Identity Mappings in Deep Residual Networks","volume":"Volume 9908","author":"He","year":"2016","journal-title":"Lecture Notes in Computer Science"},{"key":"ref_42","unstructured":"(2020, April 01). Wiki, Residual Neural Network. Available online: https:\/\/en.wikipedia.org\/wiki\/Residual_neural_network."},{"key":"ref_43","unstructured":"Chen, L.-C., Papandreou, G., Schroff, F., and Adam, H. (2017). Rethinking atrous convolution for semantic image segmentation. arXiv."},{"key":"ref_44","unstructured":"Sethi, A. (2020, February 01). One-Hot Encoding vs. Label Encoding using Scikit-Learn. Available online: https:\/\/www.analyticsvidhya.com\/blog\/2020\/03\/one-hot-encoding-vs-label-encoding-using-scikit-learn."},{"key":"ref_45","first-page":"281","article-title":"Random Search for Hyper-Parameter Optimization","volume":"13","author":"Bergstra","year":"2012","journal-title":"Mach. Learn. Res."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Cui, W., Wang, F., He, X., Zhang, D., Xu, X., Yao, M., Wang, Z., and Huang, J. (2019). Multi-Scale Semantic Segmentation and Spatial Relationship Recognition of Remote Sensing Images Based on an Attention Model. Remote Sens., 11.","DOI":"10.3390\/rs11091044"},{"key":"ref_47","unstructured":"Iglovikov, V., Mushinskiy, S., and Osin, V. (2017). Satellite imagery feature detection using deep convolutional neural network: A kaggle competition. arXiv."},{"key":"ref_48","unstructured":"(2020, January 10). Dstl Satellite Imagery Feature Detection. Available online: https:\/\/www.kaggle.com\/c\/dstl-satellite-imagery-feature-detection."},{"key":"ref_49","unstructured":"Padwick, C., Deskevich, M., Pacifici, F., and Smallwood, S. (2010, January 26\u201330). WorldView-2 pan-sharpening. Proceedings of the American Society for Photogrammetry and Remote Sensing Annual Conference, San Diego, CA, USA."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Volpi, M., and Ferrari, V. (2015, January 7\u201312). Semantic segmentation of urban scenes by learning local class interactions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Workshops. Looking from Above: When Earth Observation Meets Vision (EARTHVISION), Boston, MA, USA.","DOI":"10.1109\/CVPRW.2015.7301377"},{"key":"ref_51","first-page":"380","article-title":"The simulation and prediction of spatio-temporal urban growth trends using cellular automata models: A review","volume":"52","author":"Aburas","year":"2016","journal-title":"Int. J. Appl. Earth Obs. Geoinf."},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"4558","DOI":"10.1002\/mp.13147","article-title":"Fully automatic multi-organ segmentation for head and neck cancer radiotherapy using shape representation model constrained fully convolutional neural networks","volume":"45","author":"Tong","year":"2018","journal-title":"Med. Phys."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Zhang, J., Lin, S., Ding, L., and Bruzzone, L. (2020). Multi-Scale Context Aggregation for Semantic Segmentation of Remote Sensing Images. Remote Sens., 12.","DOI":"10.3390\/rs12040701"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Lin, D., Ji, Y., Lischinski, D., Cohen-Or, D., and Huang, H. (2018, January 8\u201314). Multi-scale context intertwining for semantic segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01219-9_37"},{"key":"ref_55","unstructured":"Kaggle Team (2019, November 10). Dstl Satellite Imagery Competition, 1st Place Winner\u2019s Interview: Kyle Lee. Available online: https:\/\/medium.com\/kaggle-blog\/dstl-satellite-imagery-competition-1st-place-winners-interview-kyle-lee-6571ce640253."},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"659","DOI":"10.1109\/TGRS.2014.2326886","article-title":"Conditional random fields for multitemporal and multiscale classification of optical satellite imagery","volume":"53","author":"Hoberg","year":"2014","journal-title":"IEEE Trans. Geosci. Remote Sens."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/18\/2932\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:08:37Z","timestamp":1760177317000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/18\/2932"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,10]]},"references-count":56,"journal-issue":{"issue":"18","published-online":{"date-parts":[[2020,9]]}},"alternative-id":["rs12182932"],"URL":"https:\/\/doi.org\/10.3390\/rs12182932","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,9,10]]}}}