{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T16:35:21Z","timestamp":1777653321926,"version":"3.51.4"},"reference-count":55,"publisher":"MDPI AG","issue":"16","license":[{"start":{"date-parts":[[2019,8,14]],"date-time":"2019-08-14T00:00:00Z","timestamp":1565740800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Key Projects of Science and Technology Agency of Guangxi province, China","award":["Guike AA 17129002"],"award-info":[{"award-number":["Guike AA 17129002"]}]},{"name":"National Science and Technology Key Program of China","award":["2013GS500303"],"award-info":[{"award-number":["2013GS500303"]}]},{"name":"Municipal Science and Technology Project of CQMMC, China","award":["2017030502"],"award-info":[{"award-number":["2017030502"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>How to efficiently utilize vast amounts of easily accessed aerial imageries is a critical challenge for researchers with the proliferation of high-resolution remote sensing sensors and platforms. Recently, the rapid development of deep neural networks (DNN) has been a focus in remote sensing, and the networks have achieved remarkable progress in image classification and segmentation tasks. However, the current DNN models inevitably lose the local cues during the downsampling operation. Additionally, even with skip connections, the upsampling methods cannot properly recover the structural information, such as the edge intersections, parallelism, and symmetry. In this paper, we propose the Web-Net, which is a nested network architecture with hierarchical dense connections, to handle these issues. We design the Ultra-Hierarchical Sampling (UHS) block to absorb and fuse the inter-level feature maps to propagate the feature maps among different levels. The position-wise downsampling\/upsampling methods in the UHS iteratively change the shape of the inputs while preserving the number of their parameters, so that the low-level local cues and high-level semantic cues are properly preserved. We verify the effectiveness of the proposed Web-Net in the Inria Aerial Dataset and WHU Dataset. The results of the proposed Web-Net achieve an overall accuracy of 96.97% and an IoU (Intersection over Union) of 80.10% on the Inria Aerial Dataset, which surpasses the state-of-the-art SegNet 1.8% and 9.96%, respectively; the results on the WHU Dataset also support the effectiveness of the proposed Web-Net. Additionally, benefitting from the nested network architecture and the UHS block, the extracted buildings on the prediction maps are obviously sharper and more accurately identified, and even the building areas that are covered by shadows can also be correctly extracted. The verified results indicate that the proposed Web-Net is both effective and efficient for building extraction from high-resolution remote sensing images.<\/jats:p>","DOI":"10.3390\/rs11161897","type":"journal-article","created":{"date-parts":[[2019,8,14]],"date-time":"2019-08-14T03:59:26Z","timestamp":1565755166000},"page":"1897","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":40,"title":["Web-Net: A Novel Nest Networks with Ultra-Hierarchical Sampling for Building Extraction from Aerial Imageries"],"prefix":"10.3390","volume":"11","author":[{"given":"Yan","family":"Zhang","sequence":"first","affiliation":[{"name":"Key Lab of Optoelectronic Technology &amp; Systems of Education Ministry, Chongqing University, Chongqing 400044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weiguo","family":"Gong","sequence":"additional","affiliation":[{"name":"Key Lab of Optoelectronic Technology &amp; Systems of Education Ministry, Chongqing University, Chongqing 400044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingxi","family":"Sun","sequence":"additional","affiliation":[{"name":"Key Lab of Optoelectronic Technology &amp; Systems of Education Ministry, Chongqing University, Chongqing 400044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weihong","family":"Li","sequence":"additional","affiliation":[{"name":"Key Lab of Optoelectronic Technology &amp; Systems of Education Ministry, Chongqing University, Chongqing 400044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,8,14]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Xu, Y., Wu, L., Xie, Z., and Chen, Z. (2018). Building Extraction in Very High Resolution Remote Sensing Imagery Using Deep Learning and Guided Filters. Remote Sens., 10.","DOI":"10.3390\/rs10010144"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Gao, L., Shi, W., Miao, Z., and Lv, Z. (2018). Method based on edge constraint and fast marching for road centerline extraction from very high-resolution remote sensing images. Remote Sens., 10.","DOI":"10.3390\/rs10060900"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Audebert, N., Le Saux, B., and Lef\u00e8vre, S. (2017). Segment-before-detect: Vehicle detection and classification through semantic segmentation of aerial images. Remote Sens., 9.","DOI":"10.3390\/rs9040368"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"2327","DOI":"10.1109\/JSTARS.2013.2242846","article-title":"Airborne vehicle detection in dense urban areas using HoG features and disparity maps","volume":"6","author":"Tuermer","year":"2013","journal-title":"IEEE J. Sel. Top. Appl. Earth Observ. Remote Sens."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"313","DOI":"10.1109\/TGRS.2012.2200689","article-title":"Automatic rooftop extraction in nadir aerial imagery of suburban regions using corners and variational level set evolution","volume":"51","author":"Cote","year":"2013","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Yang, Y., and Newsam, S. (2008, January 12\u201315). Comparing SIFT descriptors and Gabor texture features for classification of remote sensed imagery. Proceedings of the 2008 15th IEEE International Conference on Image Processing, San Diego, CA, USA.","DOI":"10.1109\/ICIP.2008.4712139"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Lowe, D.G. (1999, January 20\u201327). Object recognition from local scale-invariant features. Proceedings of the 7th IEEE International Conference on Computer Vision, Kerkyra, Greece.","DOI":"10.1109\/ICCV.1999.790410"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"247","DOI":"10.1016\/j.isprsjprs.2010.11.001","article-title":"Support vector machines in remote sensing: A review","volume":"66","author":"Mountrakis","year":"2011","journal-title":"ISPRS J. Photogramm."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1023\/A:1025623527461","article-title":"Improved rooftop detection in aerial images with machine learning","volume":"53","author":"Maloof","year":"2003","journal-title":"Mach. Learn."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1295","DOI":"10.1109\/JSTARS.2013.2249498","article-title":"Building detection with decision fusion","volume":"6","author":"Senaras","year":"2013","journal-title":"IEEE J. Sel. Top. Appl. Earth Observ. Remote Sens."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 8\u201310). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_12","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A.A. (2017, January 4\u20139). Inception-v4, inception-resnet and the impact of residual connections on learning. Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Xie, S., Girshick, R., Doll\u00e1r, P., Tu, Z., and He, K. (2017, January 22\u201325). Aggregated residual transformations for deep neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.634"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Chollet, F. (2017, January 22\u201325). Xception: Deep learning with depthwise separable convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.195"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K.Q. (2017, January 22\u201325). Densely connected convolutional networks. Proceedings of the IEEE conference on computer vision and pattern recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 22\u201325). Pyramid scene parsing network. Proceedings of the IEEE conference on computer vision and pattern recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. (2018, January 8\u201314). Encoder-decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Huang, Z., Cheng, G., Wang, H., Li, H., Shi, L., and Pan, C. (2016, January 10\u201315). Building extraction from multi-source remote sensing images via deep deconvolution neural networks. Proceedings of the 2016 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Beijing, China.","DOI":"10.1109\/IGARSS.2016.7729471"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhong, Z., Li, J., Cui, W., and Jiang, H. (2016, January 10\u201315). Fully convolutional networks for building and road extraction: Preliminary results. Proceedings of the 2016 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Beijing, China.","DOI":"10.1109\/IGARSS.2016.7729406"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Bittner, K., Cui, S., and Reinartz, P. (2017, January 6\u20139). Building Extraction from Remote Sensing Data Using fully Convolutional networks. Proceedings of the International Archives of the Photogrammetry, Remote Sensing & Spatial Information Sciences, Hannover, Germany.","DOI":"10.5194\/isprs-archives-XLII-1-W1-481-2017"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1016\/j.isprsjprs.2017.05.002","article-title":"Simultaneous extraction of roads and buildings in remote sensing imagery with convolutional neural networks","volume":"130","author":"Alshehhi","year":"2017","journal-title":"ISPRS J. Photogramm."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Wu, G., Shao, X., Guo, Z., Chen, Q., Yuan, W., Shi, X., Xu, Y., and Shibasaki, R. (2018). Automatic building segmentation of aerial imagery using multi-constraint fully convolutional networks. Remote Sens., 10.","DOI":"10.3390\/rs10030407"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"173","DOI":"10.1109\/LGRS.2017.2778181","article-title":"Semantic segmentation of aerial images with shuffling convolutional neural networks","volume":"15","author":"Chen","year":"2018","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Yang, H., Wu, P., Yao, X., Wu, Y., Wang, B., and Xu, Y. (2018). Building Extraction in Very High Resolution Imagery by Dense-Attention Networks. Remote Sens., 10.","DOI":"10.3390\/rs10111768"},{"key":"ref_30","unstructured":"Mou, L., and Zhu, X.X. (2018). RiFCN: Recurrent network in fully convolutional network for semantic segmentation of high resolution remote sensing images. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"3639","DOI":"10.1109\/TGRS.2016.2636241","article-title":"Deep recurrent neural networks for hyperspectral image classification","volume":"55","author":"Mou","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1016\/j.isprsjprs.2017.11.011","article-title":"Beyond RGB: Very high resolution urban remote sensing with multimodal deep networks","volume":"140","author":"Audebert","year":"2018","journal-title":"ISPRS J. Photogramm."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"108","DOI":"10.1016\/j.isprsjprs.2017.11.003","article-title":"MugNet: Deep learning for hyperspectral image classification using limited samples","volume":"145","author":"Pan","year":"2018","journal-title":"ISPRS J. Photogramm."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"2615","DOI":"10.1109\/JSTARS.2018.2849363","article-title":"Building Footprint Extraction From VHR Remote Sensing Images Combined with Normalized DSMs Using Fused Fully Convolutional Networks","volume":"11","author":"Bittner","year":"2018","journal-title":"IEEE J. Sel. Top. Appl. Earth Observ. Remote Sens."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Shrestha, S., and Vanneschi, L. (2018). Improved Fully Convolutional Network with Conditional Random Fields for Building Extraction. Remote Sens., 10.","DOI":"10.3390\/rs10071135"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Wang, Y., Liang, B., Ding, M., and Li, J. (2019). Dense Semantic Labelling with Atrous Spatial Pyramid Pooling and Decoder for High-Resolution Remote Sensing Imagery. Remote Sens., 11.","DOI":"10.3390\/rs11010020"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Maggiori, E., Tarabalka, Y., Charpiat, G., and Alliez, P. (2017, January 23\u201328). Can semantic labelling methods generalize to any city? the inria aerial image labelling benchmark. Proceedings of the 2017 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Fort Worth, TX, USA.","DOI":"10.1109\/IGARSS.2017.8127684"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Fourure, D., Emonet, R., Fromont, E., Muselet, D., Tremeau, A., and Wolf, C. (2017). Residual conv-deconv grid network for semantic segmentation. arXiv.","DOI":"10.5244\/C.31.181"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Pohlen, T., Hermans, A., Mathias, M., and Leibe, B. (2017, January 22\u201325). Full-resolution residual networks for semantic segmentation in street scenes. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.353"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Zhou, Z., Siddiquee, M.M.R., Tajbakhsh, N., and Liang, J. (2018). Unet++: A nested u-net architecture for medical image segmentation. Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support, Springer.","DOI":"10.1007\/978-3-030-00889-5_1"},{"key":"ref_41","unstructured":"Wang, L., Lee, C.-Y., Tu, Z., and Lazebnik, S. (2015). Training deeper convolutional networks with deep supervision. arXiv."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201322). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Wang, P., Chen, P., Yuan, Y., Liu, D., Huang, Z., Hou, X., and Cottrell, G. (2018, January 12\u201315). Understanding convolution for semantic segmentation. Proceedings of the 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe, CA, USA.","DOI":"10.1109\/WACV.2018.00163"},{"key":"ref_44","first-page":"574","article-title":"Fully Convolutional Networks for Multisource Building Extraction From an Open Aerial and Satellite Imagery Data Set","volume":"574","author":"Ji","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_45","unstructured":"Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017, January 9). Automatic differentiation in pytorch. Proceedings of the NIPS 2017 Autodiff Workshop: The Future of Gradient-basedMachine Learning Software and Techniques, Long Beach, CA, USA."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Li, F.-F. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_47","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_48","unstructured":"Pleiss, G., Chen, D., Huang, G., Li, T., van der Maaten, L., and Weinberger, K.Q. (2017). Memory-efficient implementation of densenets. arXiv."},{"key":"ref_49","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014, January 8\u201313). Generative adversarial nets. Proceedings of the Advances in Neural Information Processing Systems, Montr\u00e9al, Canada."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1016\/j.isprsjprs.2018.11.011","article-title":"Aerial imagery for roof segmentation: A large-scale dataset towards automatic mapping of buildings","volume":"147","author":"Chen","year":"2019","journal-title":"ISPRS J. Photogramm."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201325). Mask R-CNN. Proceedings of the IEEE International Conference on Computer Vision, Honolulu, HI, USA.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_52","unstructured":"Bischke, B., Helber, P., Folz, J., Borth, D., and Dengel, A. (2017). Multi-task learning for segmentation of building footprints with deep neural networks. arXiv."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Lu, K., Sun, Y., and Ong, S.-H. (2018, January 20\u201324). Dual-Resolution U-Net: Building Extraction from Aerial Images. Proceedings of the 2018 24th International Conference on Pattern Recognition (ICPR), Beijing, China.","DOI":"10.1109\/ICPR.2018.8545190"},{"key":"ref_54","unstructured":"Khalel, A., and El-Saban, M. (2018). Automatic pixelwise object labelling for aerial imagery using stacked u-nets. arXiv."},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"3680","DOI":"10.1109\/JSTARS.2018.2865187","article-title":"Building-A-Nets: Robust Building Extraction From High-Resolution Remote Sensing Images With Adversarial Networks","volume":"11","author":"Li","year":"2018","journal-title":"IEEE J. Sel. Top. Appl. Earth Observ. Remote Sens."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/11\/16\/1897\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:10:56Z","timestamp":1760188256000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/11\/16\/1897"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,8,14]]},"references-count":55,"journal-issue":{"issue":"16","published-online":{"date-parts":[[2019,8]]}},"alternative-id":["rs11161897"],"URL":"https:\/\/doi.org\/10.3390\/rs11161897","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,8,14]]}}}