{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,3]],"date-time":"2026-03-03T16:19:20Z","timestamp":1772554760300,"version":"3.50.1"},"reference-count":38,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2020,4,8]],"date-time":"2020-04-08T00:00:00Z","timestamp":1586304000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the Key Projects of Science and Technology Agency of Guangxi province, China","award":["Guike AA 17129002"],"award-info":[{"award-number":["Guike AA 17129002"]}]},{"name":"National Science and Technology Key Program of China","award":["2013GS500303"],"award-info":[{"award-number":["2013GS500303"]}]},{"name":"the Municipal Science and Technology Project of CQMMC, China","award":["2017030502"],"award-info":[{"award-number":["2017030502"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>With the development of deep learning technology, an enormous number of convolutional neural network (CNN) models have been proposed to address the challenging building extraction task from very high-resolution (VHR) remote sensing images. However, searching for better CNN architectures is time-consuming, and the robustness of a new CNN model cannot be guaranteed. In this paper, an improved boundary-aware perceptual (BP) loss is proposed to enhance the building extraction ability of CNN models. The proposed BP loss consists of a loss network and transfer loss functions. The usage of the boundary-aware perceptual loss has two stages. In the training stage, the loss network learns the structural information from circularly transferring between the building mask and the corresponding building boundary. In the refining stage, the learned structural information is embedded into the building extraction models via the transfer loss functions without additional parameters or postprocessing. We verify the effectiveness and efficiency of the proposed BP loss both on the challenging WHU aerial dataset and the INRIA dataset. Substantial performance improvements are observed within two representative CNN architectures: PSPNet and UNet, which are widely used on pixel-wise labelling tasks. With BP loss, UNet with ResNet101 achieves 90.78% and 76.62% on IoU (intersection over union) scores on the WHU aerial dataset and the INRIA dataset, respectively, which are 1.47% and 1.04% higher than those simply trained with the cross-entropy loss function. Additionally, similar improvements (0.64% on the WHU aerial dataset and 1.69% on the INRIA dataset) are also observed on PSPNet, which strongly supports the robustness of the proposed BP loss.<\/jats:p>","DOI":"10.3390\/rs12071195","type":"journal-article","created":{"date-parts":[[2020,4,9]],"date-time":"2020-04-09T03:40:19Z","timestamp":1586403619000},"page":"1195","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":24,"title":["An Improved Boundary-Aware Perceptual Loss for Building Extraction from VHR Images"],"prefix":"10.3390","volume":"12","author":[{"given":"Yan","family":"Zhang","sequence":"first","affiliation":[{"name":"Key Lab of Optoelectronic Technology &amp; Systems of Education Ministry, Chongqing University, Chongqing 400044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weihong","family":"Li","sequence":"additional","affiliation":[{"name":"Key Lab of Optoelectronic Technology &amp; Systems of Education Ministry, Chongqing University, Chongqing 400044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weiguo","family":"Gong","sequence":"additional","affiliation":[{"name":"Key Lab of Optoelectronic Technology &amp; Systems of Education Ministry, Chongqing University, Chongqing 400044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zixu","family":"Wang","sequence":"additional","affiliation":[{"name":"Key Lab of Optoelectronic Technology &amp; Systems of Education Ministry, Chongqing University, Chongqing 400044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingxi","family":"Sun","sequence":"additional","affiliation":[{"name":"Key Lab of Optoelectronic Technology &amp; Systems of Education Ministry, Chongqing University, Chongqing 400044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,4,8]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Huang, H., and Xu, K. (2019). Combing Triple-Part Features of Convolutional Neural Networks for Scene Classification in Remote Sensing. Remote. Sens., 11.","DOI":"10.3390\/rs11141687"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zhu, R., Yan, L., Mo, N., and Liu, Y. (2020). AttentionBased Deep Feature Fusion for the Scene Classification of HighResolution Remote Sensing Images. Remote. Sens., 12.","DOI":"10.3390\/rs12040742"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Cui, B., Zhang, Y., Yan, L., Wei, J., and Wu, H. (2019). An Unsupervised SAR Change Detection Method Based on Stochastic Subspace Ensemble Learning. Remote. Sens., 11.","DOI":"10.3390\/rs11111314"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Li, L., Wang, C., Zhang, H., Zhang, B., and Wu, F. (2019). Urban Building Change Detection in SAR Images Using Combined Differential Image and Residual U-Net Network. Remote. Sens., 11.","DOI":"10.3390\/rs11091091"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Mahdavi, S., Salehi, B., Huang, W., Amani, M., and Brisco, B. (2019). A PolSAR Change Detection Index Based on Neighborhood Information for Flood Mapping. Remote. Sens., 11.","DOI":"10.3390\/rs11161854"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Chen, C., Gong, W., Chen, Y., and Li, W. (2019). Object Detection in Remote Sensing Images Based on a Scene-Contextual Feature Pyramid Network. Remote. Sens., 11.","DOI":"10.3390\/rs11030339"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Pan, X., Yang, F., Gao, L., Chen, Z., Zhang, B., Fan, H., and Ren, J. (2019). Building Extraction from High-Resolution Aerial Imagery Using a Generative Adversarial Network with Spatial and Channel Attention Mechanisms. Remote Sens., 11.","DOI":"10.3390\/rs11080917"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Gong, W., Sun, J., and Li, W. (2019). Web-Net: A Novel Nest Networks with Ultra-Hierarchical Sampling for Building Extraction from Aerial Imageries. Remote. Sens., 11.","DOI":"10.3390\/rs11161897"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Neuville, R., Pouliot, J., Poux, F., and Billen, R. (2019). 3D Viewpoint Management and Navigation in Urban Planning: Application to the Exploratory Phase. Remote. Sens., 11.","DOI":"10.3390\/rs11030236"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Khanal, N., Uddin, K., Matin, M., and Tenneson, K. (2019). Automatic Detection of Spatiotemporal Urban Expansion Patterns by Fusing OSM and Landsat Data in Kathmandu. Remote. Sens., 11.","DOI":"10.3390\/rs11192296"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully Convolutional Networks for Semantic Segmentation. Proceedings of the 2015 Ieee Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"74","DOI":"10.1016\/j.neunet.2019.08.025","article-title":"MultiResUNet : Rethinking the U-Net architecture for multimodal biomedical image segmentation","volume":"121","author":"Ibtehaz","year":"2020","journal-title":"Neural Netw."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zhao, J., He, X., Li, J., Feng, T., Ye, C., and Xiong, L. (2019). Automatic Vector-Based Road Structure Mapping Using Multibeam LiDAR. Remote. Sens., 11.","DOI":"10.3390\/rs11141726"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1016\/j.isprsjprs.2019.02.019","article-title":"Automatic building extraction from high-resolution aerial images and LiDAR data using gated residual refinement network","volume":"151","author":"Huang","year":"2019","journal-title":"ISPRS J. Photogramm. Remote. Sens."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Sun, G., Huang, H., Zhang, A., Li, F., Zhao, H., and Fu, H. (2019). Fusion of Multiscale Convolutional Neural Networks for Building Extraction in Very High-Resolution Images. Remote. Sens., 11.","DOI":"10.3390\/rs11030227"},{"key":"ref_16","first-page":"234","article-title":"U-Net: Convolutional Networks for Biomedical Image Segmentation","volume":"Volume 9351","author":"Navab","year":"2015","journal-title":"Medical Image Computing and Computer-Assisted Intervention, Pt Iii"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Peng, D., Zhang, Y., and Guan, H. (2019). Guan End-to-End Change Detection for High Resolution Satellite Images Using Improved UNet++. Remote. Sens., 11.","DOI":"10.3390\/rs11111382"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.isprsjprs.2019.07.007","article-title":"TreeUNet: Adaptive Tree convolutional neural networks for subdecimeter aerial image segmentation","volume":"156","author":"Yue","year":"2019","journal-title":"ISPRS J. Photogramm. Remote. Sens."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/978-3-030-00889-5_1","article-title":"UNet plus plus : A Nested U-Net Architecture for Medical Image Segmentation","volume":"Volume 11045","author":"Stoyanov","year":"2018","journal-title":"Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support, Dlmia 2018"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Wu, G., Shao, X., Guo, Z., Chen, Q., Yuan, W., Shi, X., Xu, Y., and Shibasaki, R. (2018). Automatic Building Segmentation of Aerial Imagery Using Multi-Constraint Fully Convolutional Networks. Remote. Sens., 10.","DOI":"10.3390\/rs10030407"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 21\u201326). Pyramid Scene Parsing Network. Proceedings of the 30th Ieee Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. (2018, January 8\u201314). Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"ref_24","unstructured":"Kr\u00e4henb\u00fchl, P., and Koltun, V. (2020, April 08). Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials. Available online: http:\/\/papers.nips.cc\/paper\/4296-efficient-inference-in-fully-connected-crfs-with-gaussian-edge-potentials.pdf."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zheng, S., Jayasumana, S., Romera-Paredes, B., Vineet, V., Su, Z., Du, D., Huang, C., and Torr, P.H.S. (2015, January 7\u201313). Conditional Random Fields as Recurrent Neural Networks. Proceedings of the 2015 Ieee International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.179"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Bertels, J., Eelbode, T., Berman, M., Vandermeulen, D., Maes, F., Bisschops, R., and Blaschko, M.B. (2019). Optimizing the Dice score and Jaccard index for medical image segmentation: Theory and practice. International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer.","DOI":"10.1007\/978-3-030-32245-8_11"},{"key":"ref_27","unstructured":"Iglovikov, V., and Shvets, A. (2018). TernausNet: U-Net with VGG11 Encoder Pre-Trained on ImageNet for Image Segmentation. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"694","DOI":"10.1007\/978-3-319-46475-6_43","article-title":"Perceptual Losses for Real-Time Style Transfer and Super-Resolution","volume":"Volume 9906","author":"Leibe","year":"2016","journal-title":"Computer Vision-Eccv 2016, Pt Ii"},{"key":"ref_29","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Chen, Y., Dapogny, A., and Cord, M. (2019). SEMEDA: Enhancing Segmentation Precision with Semantic Edge Aware Loss. arXiv.","DOI":"10.1016\/j.patcog.2020.107557"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"574","DOI":"10.1109\/TGRS.2018.2858817","article-title":"Fully Convolutional Networks for Multisource Building Extraction From an Open Aerial and Satellite Imagery Data Set","volume":"57","author":"Ji","year":"2018","journal-title":"IEEE Trans. Geosci. Remote. Sens."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Maggiori, E., Tarabalka, Y., Charpiat, G., and Alliez, P. (2017, January 23\u201328). Can semantic labeling methods generalize to any city? The inria aerial image labeling benchmark. Proceedings of the 2017 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Worth, TX, USA.","DOI":"10.1109\/IGARSS.2017.8127684"},{"key":"ref_33","unstructured":"Sobel, I. (2020, April 08). History and Definition of the Sobel Operator. Available online: https:\/\/www.researchgate.net\/publication\/239398674_An_Isotropic_3x3_Image_Gradient_Operator."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 Ieee Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Liu, P., Liu, X., Liu, M., Shi, Q., Yang, J., Xu, X., and Zhang, Y. (2019). Building Footprint Extraction from High-Resolution Images via Spatial Residual Inception Convolutional Neural Network. Remote. Sens., 11.","DOI":"10.3390\/rs11070830"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Liu, H., Luo, J., Huang, B., Hu, X., Sun, Y., Yang, Y., Xu, N., and Zhou, N. (2019). DE-Net: Deep Encoding Network for Building Extraction from High-Resolution Remote Sensing Imagery. Remote. Sens., 11.","DOI":"10.3390\/rs11202380"},{"key":"ref_37","unstructured":"Bischke, B., Helber, P., Folz, J., Borth, D., and Dengel, A. (2017). Multi-task learning for segmentation of building footprints with deep neural networks. arXiv."},{"key":"ref_38","unstructured":"Mou, L., and Zhu, X.X. (2018). RiFCN: Recurrent network in fully convolutional network for semantic segmentation of high resolution remote sensing images. arXiv."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/7\/1195\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:16:31Z","timestamp":1760174191000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/7\/1195"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,4,8]]},"references-count":38,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2020,4]]}},"alternative-id":["rs12071195"],"URL":"https:\/\/doi.org\/10.3390\/rs12071195","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,4,8]]}}}