{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T05:21:43Z","timestamp":1781760103389,"version":"3.54.5"},"reference-count":33,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2018,11,5]],"date-time":"2018-11-05T00:00:00Z","timestamp":1541376000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Semantic segmentation of high-resolution aerial images is of great importance in certain fields, but the increasing spatial resolution brings large intra-class variance and small inter-class differences that can lead to classification ambiguities. Based on high-level contextual features, the deep convolutional neural network (DCNN) is an effective method to deal with semantic segmentation of high-resolution aerial imagery. In this work, a novel dense pyramid network (DPN) is proposed for semantic segmentation. The network starts with group convolutions to deal with multi-sensor data in channel wise to extract feature maps of each channel separately; by doing so, more information from each channel can be preserved. This process is followed by the channel shuffle operation to enhance the representation ability of the network. Then, four densely connected convolutional blocks are utilized to both extract and take full advantage of features. The pyramid pooling module combined with two convolutional layers are set to fuse multi-resolution and multi-sensor features through an effective global scenery prior manner, producing the probability graph for each class. Moreover, the median frequency balanced focal loss is proposed to replace the standard cross entropy loss in the training phase to deal with the class imbalance problem. We evaluate the dense pyramid network on the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam 2D semantic labeling dataset, and the results demonstrate that the proposed framework exhibits better performances, compared to the state of the art baseline.<\/jats:p>","DOI":"10.3390\/s18113774","type":"journal-article","created":{"date-parts":[[2018,11,5]],"date-time":"2018-11-05T10:43:45Z","timestamp":1541414625000},"page":"3774","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":32,"title":["High-Resolution Aerial Imagery Semantic Labeling with Dense Pyramid Network"],"prefix":"10.3390","volume":"18","author":[{"given":"Xuran","family":"Pan","sequence":"first","affiliation":[{"name":"Key Laboratory of Digital Earth Science, Institute of Remote Sensing and Digital Earth, Chinese Academy of Sciences, Beijing 100094, China"},{"name":"School of Electronics and Information Engineering, Hebei University of Technology, Tianjin 300401, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3888-8124","authenticated-orcid":false,"given":"Lianru","family":"Gao","sequence":"additional","affiliation":[{"name":"Key Laboratory of Digital Earth Science, Institute of Remote Sensing and Digital Earth, Chinese Academy of Sciences, Beijing 100094, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0319-7753","authenticated-orcid":false,"given":"Bing","family":"Zhang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Digital Earth Science, Institute of Remote Sensing and Digital Earth, Chinese Academy of Sciences, Beijing 100094, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fan","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Engineering, Hebei University of Technology, Tianjin 300401, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenzhi","family":"Liao","sequence":"additional","affiliation":[{"name":"Department of Telecommunications and Information Processing, Ghent University, 9000 Ghent, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2018,11,5]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1016\/j.neucom.2018.01.076","article-title":"A deep learning based feature hybrid framework for spatiotemporal saliency detection inside videos","volume":"287","author":"Wang","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"3325","DOI":"10.1109\/TGRS.2014.2374218","article-title":"Object detection in optical remote sensing images based on weakly supervised learning and high-level feature learning","volume":"53","author":"Han","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1309","DOI":"10.1109\/TCSVT.2014.2381471","article-title":"Background prior-based salient object detection via deep reconstruction residual","volume":"25","author":"Han","year":"2015","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"573","DOI":"10.1016\/j.mcm.2011.10.063","article-title":"Automatic remotely sensed image classification in a grid environment based on the maximum likelihood method","volume":"58","author":"Sun","year":"2013","journal-title":"Math. Comput. Model."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"65","DOI":"10.1016\/j.patcog.2018.02.004","article-title":"Unsupervised image saliency detection with Gestalt-laws guided optimization and visual attention based refinement","volume":"79","author":"Yan","year":"2018","journal-title":"Pattern Recognit."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"144","DOI":"10.1016\/j.knosys.2011.07.016","article-title":"ANN vs. SVM: Which one performs better in classification of MCCs in mammogram imaging","volume":"26","author":"Ren","year":"2012","journal-title":"Knowl. Based Syst."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"4238","DOI":"10.1109\/TGRS.2015.2393857","article-title":"Effective and efficient midlevel visual elements-oriented land-use classification using VHR remote sensing images","volume":"53","author":"Cheng","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"343","DOI":"10.14358\/PERS.80.4.343","article-title":"Mapping Impervious Surfaces Using Object-Oriented Classification in a Semiarid Urban Region","volume":"80","author":"Sugg","year":"2015","journal-title":"Photogramm. Eng. Remote Sens."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1613","DOI":"10.1109\/JSTARS.2015.2508285","article-title":"One-Class Classification of Remote Sensing Images Using Kernel Sparse Representation","volume":"9","author":"Song","year":"2016","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Paisitkriangkrai, S., Sherrah, J., Janney, P., and Hengel, A. (2015, January 7\u201312). Effective semantic pixel labelling with convolutional networks and conditional random fields. Proceedings of the Conference on Computer Vision and Pattern Recognition Workshops (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPRW.2015.7301381"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"881","DOI":"10.1109\/TGRS.2016.2616585","article-title":"Dense semantic labeling of subdecimeter resolution images with convolutional neural networks","volume":"55","author":"Volpi","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Liu, Y., Piramanayagam, S., Monteiro, S., and Saber, E. (2017, January 21\u201326). Dense semantic labeling of very-high-resolution aerial imagery and lidar with fully-convolutional neural networks and higher-order CRFs. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.200"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_14","unstructured":"Sherrah, J. (arXiv, 2016). Fully convolutional networks for dense semantic labelling of high-resolution aerial imagery, arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Li, R., Liu, W., Yang, L., Sun, S., Hu, W., Zhang, F., and Li, W. (arXiv, 2017). DeepUNet: A Deep Fully Convolutional Network for Pixel-Level Sea-Land Segmentation, arXiv.","DOI":"10.1109\/JSTARS.2018.2833382"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Audebert, N., Le Saux, B., and Lef\u00e8vre, S. (2016, January 21\u201323). Semantic segmentation of earth observation data using multimodal and multi-scale deep networks. Proceedings of the Asian Conference on Computer Vision (ACCV), Taipei, Taiwan.","DOI":"10.1007\/978-3-319-54181-5_12"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"7092","DOI":"10.1109\/TGRS.2017.2740362","article-title":"High-resolution aerial image labeling with convolutional neural networks","volume":"55","author":"Maggiori","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Liu, Y., Minh Nguyen, D., Deligiannis, N., Ding, W., and Munteanu, A. (2017). Hourglass-shapenetwork based semantic segmentation for high resolution aerial imagery. Remote Sens., 9.","DOI":"10.3390\/rs9060522"},{"key":"ref_21","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20136). Imagenet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems, Nevada, NV, USA."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhou, X., Lin, M., and Sun, J. (arXiv, 2017). Shufflenet: An extremely efficient convolutional neural network for mobile devices, arXiv.","DOI":"10.1109\/CVPR.2018.00716"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., and Weinberger, K.Q. (2017, January 21\u201326). Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Pan, X., Gao, L., Marinoni, A., Zhang, B., Yang, F., and Gamba, P. (2018). Semantic Labeling of High Resolution Aerial Imagery and LiDAR Data with Fine Segmentation Network. Remote Sens., 10.","DOI":"10.3390\/rs10050743"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 21\u201326). Pyramid scene parsing network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Lin, T., Goyal, P., Girshick, R., He, K., and Dollar, P. (2017). Focal Loss for Dense Object Detection. IEEE Trans. Pattern Anal. Mach. Intell., 2999\u20133007.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Eigen, D., and Fergus, R. (2015, January 7\u201313). Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Los Alamitos, CA, USA.","DOI":"10.1109\/ICCV.2015.304"},{"key":"ref_28","unstructured":"Kingma, D., and Ba, J. (2014, January 14\u201316). Adam: A method for stochastic optimization. Proceedings of the International Conference on Learning Representations (ICLR), Banff, AB, Canada."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"200","DOI":"10.1016\/j.knosys.2017.10.018","article-title":"A stability constrained adaptive alpha for gravitational search algorithm","volume":"139","author":"Sun","year":"2018","journal-title":"Knowl. Based Syst."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1109\/TCYB.2016.2641986","article-title":"A dynamic neighborhood learning-based gravitational search algorithm","volume":"48","author":"Zhang","year":"2018","journal-title":"IEEE Trans. Cybern."},{"key":"ref_31","unstructured":"(2017, April 01). ISPRS Vaihingen 2D Semantic Labeling Dataset. Available online: http:\/\/www2.isprs.org\/commissions\/comm3\/wg4\/tests.html."},{"key":"ref_32","unstructured":"Gerke, M. (2015). Use of the Stair Vision Library within the ISPRS 2D Semantic Labeling Benchmark (Vaihingen), University of Twente."},{"key":"ref_33","unstructured":"(2017, April 01). ISPRS Test Project on Urban Classification, 3D Building Reconstruction and Semantic Labeling. Available online: http:\/\/www2.isprs.org\/commissions\/comm2\/wg4\/potsdam-2d-semantic-labeling.html."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/11\/3774\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T23:45:18Z","timestamp":1775259918000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/11\/3774"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,11,5]]},"references-count":33,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2018,11]]}},"alternative-id":["s18113774"],"URL":"https:\/\/doi.org\/10.3390\/s18113774","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,11,5]]}}}