{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T15:44:59Z","timestamp":1783439099796,"version":"3.54.6"},"reference-count":44,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2017,5,25]],"date-time":"2017-05-25T00:00:00Z","timestamp":1495670400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100003130","name":"Fonds Wetenschappelijk Onderzoek","doi-asserted-by":"publisher","award":["G084117"],"award-info":[{"award-number":["G084117"]}],"id":[{"id":"10.13039\/501100003130","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Brussels Institute for Research and Innovation (Innoviris)","award":["3DLicornea"],"award-info":[{"award-number":["3DLicornea"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>A new convolution neural network (CNN) architecture for semantic segmentation of high resolution aerial imagery is proposed in this paper. The proposed architecture follows an hourglass-shaped network (HSN) design being structured into encoding and decoding stages. By taking advantage of recent advances in CNN designs, we use the composed inception module to replace common convolutional layers, providing the network with multi-scale receptive areas with rich context. Additionally, in order to reduce spatial ambiguities in the up-sampling stage, skip connections with residual units are also employed to feed forward encoding-stage information directly to the decoder. Moreover, overlap inference is employed to alleviate boundary effects occurring when high resolution images are inferred from small-sized patches. Finally, we also propose a post-processing method based on weighted belief propagation to visually enhance the classification results. Extensive experiments based on the Vaihingen and Potsdam datasets demonstrate that the proposed architectures outperform three reference state-of-the-art network designs both numerically and visually.<\/jats:p>","DOI":"10.3390\/rs9060522","type":"journal-article","created":{"date-parts":[[2017,5,30]],"date-time":"2017-05-30T04:35:42Z","timestamp":1496118942000},"page":"522","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":129,"title":["Hourglass-ShapeNetwork Based Semantic Segmentation for High Resolution Aerial Imagery"],"prefix":"10.3390","volume":"9","author":[{"given":"Yu","family":"Liu","sequence":"first","affiliation":[{"name":"ETRO Department, Vrije Universiteit Brussel, Pleinlaan 2, 1050 Brussels, Belgium"},{"name":"School of Electronic and Information Engineering, Beihang University, 37 Xueyuan Rd., Haidian District, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Duc","family":"Minh Nguyen","sequence":"additional","affiliation":[{"name":"ETRO Department, Vrije Universiteit Brussel, Pleinlaan 2, 1050 Brussels, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9300-5860","authenticated-orcid":false,"given":"Nikos","family":"Deligiannis","sequence":"additional","affiliation":[{"name":"ETRO Department, Vrije Universiteit Brussel, Pleinlaan 2, 1050 Brussels, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenrui","family":"Ding","sequence":"additional","affiliation":[{"name":"School of Electronic and Information Engineering, Beihang University, 37 Xueyuan Rd., Haidian District, Beijing 100191, China"},{"name":"Collaborative Innovation Center of Geospatial Technology, Wuhan 430079, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Adrian","family":"Munteanu","sequence":"additional","affiliation":[{"name":"ETRO Department, Vrije Universiteit Brussel, Pleinlaan 2, 1050 Brussels, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2017,5,25]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Rees, W.G. (2013). Physical Principles of Remote Sensing, Cambridge University Press.","DOI":"10.1017\/CBO9781139017411"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1279","DOI":"10.1109\/TNNLS.2015.2477537","article-title":"Salient band selection for hyperspectral image classification via manifold ranking","volume":"27","author":"Wang","year":"2016","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"3123","DOI":"10.1109\/TCYB.2015.2497711","article-title":"Hyperspectral anomaly detection by graph pixel selection","volume":"46","author":"Yuan","year":"2016","journal-title":"IEEE Trans. Cybern."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1023\/A:1011139631724","article-title":"Modeling the shape of the scene: A holistic representation of the spatial envelope","volume":"42","author":"Oliva","year":"2001","journal-title":"Int. J. Comput. Vis."},{"key":"ref_5","unstructured":"Huang, J., Kumar, S.R., Mitra, M., Zhu, W.J., and Zabih, R. (1997, January 17\u201319). Image indexing using color correlograms. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Stehling, R.O., Nascimento, M.A., and Falc\u00e3o, A.X. (2002, January 4\u20139). A Compact and Efficient Image Retrieval Approach Based on Border\/Interior Pixel Classification. Proceedings of the International Conference on Information and Knowledge Management (CIKM), McLean, VA, USA.","DOI":"10.1145\/584792.584812"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"453","DOI":"10.1016\/j.cviu.2012.09.007","article-title":"Pooling in image representation: The visual codeword point of view","volume":"117","author":"Avila","year":"2013","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_8","unstructured":"Lazebnik, S., Schmid, C., and Ponce, J. (2006, January 17\u201322). Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), New York, NY, USA."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1007\/s10618-005-1396-1","article-title":"Automatic subspace clustering of high dimensional data","volume":"11","author":"Agrawal","year":"2005","journal-title":"Data Min. Knowl. Discov."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1540","DOI":"10.1016\/j.patcog.2011.01.004","article-title":"A survey of multilinear subspace learning for tensor data","volume":"44","author":"Lu","year":"2011","journal-title":"Pattern Recognit."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1053","DOI":"10.1109\/TCYB.2016.2536752","article-title":"Constructing the L2-graph for robust subspace learning and subspace clustering","volume":"47","author":"Peng","year":"2017","journal-title":"IEEE Trans. Cybern."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"280","DOI":"10.1109\/TGRS.2014.2321423","article-title":"Features, Color Spaces, and Boosting: New Insights on Semantic Classification of Remote Sensing Images","volume":"53","author":"Tokarczyk","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"439","DOI":"10.1109\/TGRS.2013.2241444","article-title":"Unsupervised feature learning for aerial scene classification","volume":"52","author":"Cheriyadat","year":"2014","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1109\/LGRS.2011.2161569","article-title":"Automatic Target Detection in High-Resolution Remote Sensing Images Using Spatial Sparse Coding Bag-of-Words Model","volume":"9","author":"Sun","year":"2012","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_15","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20138). ImageNet Classification with Deep Convolutional Neural Networks. Proceedings of the Advances in Neural Information Processing Systems (NIPS), Stateline, NV, USA."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Razavian, A.S., Azizpour, H., Sullivan, J., and Carlsson, S. (2014, January 23\u201328). CNN Features Off-the-Shelf: An Astounding Baseline for Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Columbus, OH, USA.","DOI":"10.1109\/CVPRW.2014.131"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_19","unstructured":"Badrinarayanan, V., Kendall, A., and Cipolla, R. (2015). Segnet: A deep convolutional encoder-decoder architecture for image segmentation. Comput. Vis. Pattern Recognit."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Paisitkriangkrai, S., Sherrah, J., Janney, P., and Hengel, V.D. (2015, January 7\u201312). Effective semantic pixel labeling with convolutional networks and conditional random fields. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Boston, MA, USA.","DOI":"10.1109\/CVPRW.2015.7301381"},{"key":"ref_21","unstructured":"Kampffmeyer, M., Salberg, A.B., and Jenssen, R. (July, January 26). Semantic Segmentation of Small Objects and Modeling of Uncertainty in Urban Remote Sensing Images Using Deep Convolutional Neural Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Las Vegas, NV, USA."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2016.2616585","article-title":"Dense Semantic Labeling of Subdecimeter Resolution Images With Convolutional Neural Networks","volume":"55","author":"Volpi","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zeiler, M.D., Taylor, G.W., and Fergus, R. (2011, January 6\u201313). Adaptive deconvolutional networks for mid and high level feature learning. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126474"},{"key":"ref_24","unstructured":"Glorot, X., Bordes, A., and Bengio, Y. (2011, January 11\u201313). Deep Sparse Rectifier Neural Networks. Proceedings of the International Conference on Artificial Intelligence and Statistics (AISTATS), Fort Lauderdale, FL, USA."},{"key":"ref_25","unstructured":"Maas, A.L., Hannun, A.Y., and Ng, A.Y. (2013, January 16\u201321). Rectifier nonlinearities improve neural network acoustic models. Proceedings of the International Conference on Machine Learning (ICML), Atlanta, GA, USA."},{"key":"ref_26","unstructured":"Saxe, A., Koh, P.W., Chen, Z., Bhand, M., Suresh, B., and Ng, A.Y. (July, January 28). On random weights and unsupervised feature learning. Proceedings of the International conference on machine learning (ICML), Bellevue, WA, USA."},{"key":"ref_27","unstructured":"Springenberg, J.T., Dosovitskiy, A., Brox, T., and Riedmiller, M. (2015, January 7\u20139). Striving for simplicity: The all convolutional net. Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA."},{"key":"ref_28","unstructured":"Lee, C.Y., Gallagher, P.W., and Tu, Z. (2016, January 9\u201311). Generalizing pooling functions in convolutional neural networks: Mixed, gated, and tree. Proceedings of the International Conference on Artificial Intelligence and Statistics (AISTATS), Cadiz, Spain."},{"key":"ref_29","unstructured":"Sermanet, P., Eigen, D., Zhang, X., Mathieu, M., Fergus, R., and LeCun, Y. (2014, January 14\u201316). Overfeat: Integrated recognition, localization and detection using convolutional networks. Proceedings of the International Conference on Learning Representations (ICLR), Banff, AB, Canada."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1915","DOI":"10.1109\/TPAMI.2012.231","article-title":"Learning hierarchical features for scene labeling","volume":"35","author":"Farabet","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_31","unstructured":"Pinheiro, P., and Collobert, R. (2014, January 21\u201326). Recurrent convolutional neural networks for scene parsing. Proceedings of the International Conference on Machine Learning (ICML), Beijing, China."},{"key":"ref_32","unstructured":"Simonyan, K., and Zisserman, A. (2015, January 7\u20139). Very deep convolutional networks for large-scale image recognition. Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA."},{"key":"ref_33","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 7\u20139). Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. Proceedings of the International Conference on Machine Learning (ICML), San Diego, CA, USA."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_35","unstructured":"Chen, W., Fu, Z., Yang, D., and Deng, J. (2016, January 5\u201310). Single-image depth perception in the wild. Proceedings of the Advances in Neural Information Processing Systems (NIPS), Barcelona, Spain."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Newell, A., Yang, K., and Deng, J. (2016, January 8\u201316). Stacked Hourglass Networks for Human Pose Estimation. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Eigen, D., and Fergus, R. (2015, January 7\u201313). Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Los Alamitos, CA, USA.","DOI":"10.1109\/ICCV.2015.304"},{"key":"ref_38","first-page":"26","article-title":"Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude","volume":"4","author":"Tieleman","year":"2012","journal-title":"Neural Netw. Mach. Learn."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2015, January 7\u201313). Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Los Alamitos, CA, USA.","DOI":"10.1109\/ICCV.2015.123"},{"key":"ref_40","unstructured":"Murphy, K.P., Weiss, Y., and Jordan, M.I. (August, January 30). Loopy belief propagation for approximate inference: An empirical study. Proceedings of the Conference on Uncertainty in artificial intelligence (UAI), Stockholm, Sweden."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"498","DOI":"10.1109\/18.910572","article-title":"Factor graphs and the sum-product algorithm","volume":"47","author":"Kschischang","year":"2001","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_42","unstructured":"(2017, May 24). ISPRS Vaihingen 2D Semantic Labeling Dataset. Available online: http:\/\/www2.isprs.org\/commissions\/comm3\/wg4\/2d-sem-label-vaihingen.html."},{"key":"ref_43","unstructured":"(2017, May 24). ISPRS Potsdam 2D Semantic Labeling Dataset. Available online: http:\/\/www2.isprs.org\/commissions\/comm3\/wg4\/2d-sem-label-potsdam.html."},{"key":"ref_44","unstructured":"Gerke, M. (2015). Use of the Stair Vision Library within the ISPRS 2D Semantic Labeling Benchmark (Vaihingen), University of Twente. Technical Report."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/9\/6\/522\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T18:36:55Z","timestamp":1760207815000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/9\/6\/522"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,5,25]]},"references-count":44,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2017,6]]}},"alternative-id":["rs9060522"],"URL":"https:\/\/doi.org\/10.3390\/rs9060522","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,5,25]]}}}