{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,20]],"date-time":"2026-01-20T08:10:28Z","timestamp":1768896628044,"version":"3.49.0"},"reference-count":34,"publisher":"MDPI AG","issue":"21","license":[{"start":{"date-parts":[[2021,10,24]],"date-time":"2021-10-24T00:00:00Z","timestamp":1635033600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["41971280"],"award-info":[{"award-number":["41971280"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Semantic and instance segmentation methods are commonly used to build extraction from high-resolution images. The semantic segmentation method involves assigning a class label to each pixel in the image, thus ignoring the geometry of the building rooftop, which results in irregular shapes of the rooftop edges. As for instance segmentation, there is a strong assumption within this method that there exists only one outline polygon along the rooftop boundary. In this paper, we present a novel method to sequentially delineate exterior and interior contours of rooftops with holes from VHR aerial images, where most of the buildings have holes, by integrating semantic segmentation and polygon delineation. Specifically, semantic segmentation from the Mask R-CNN is used as a prior for hole detection. Then, the holes are used as objects for generating the internal contours of the rooftop. The external and internal contours of the rooftop are inferred separately using a convolutional recurrent neural network. Experimental results showed that the proposed method can effectively delineate the rooftops with both one and multiple polygons and outperform state-of-the-art methods in terms of the visual results and six statistical indicators, including IoU, OA, F1, BoundF, RE and Hd.<\/jats:p>","DOI":"10.3390\/rs13214271","type":"journal-article","created":{"date-parts":[[2021,10,24]],"date-time":"2021-10-24T22:07:11Z","timestamp":1635113231000},"page":"4271","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":12,"title":["Sequentially Delineation of Rooftops with Holes from VHR Aerial Images Using a Convolutional Recurrent Neural Network"],"prefix":"10.3390","volume":"13","author":[{"given":"Wei","family":"Huang","sequence":"first","affiliation":[{"name":"State Key Laboratory of Remote Sensing Science, Faculty of Geographical Science, Beijing Normal University, Beijing 100875, China"},{"name":"Beijing Key Laboratory of Environmental Remote Sensing and Digital Cities, Faculty of Geographical Science, Beijing Normal University, Beijing 100875, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2898-0023","authenticated-orcid":false,"given":"Zeping","family":"Liu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Remote Sensing Science, Faculty of Geographical Science, Beijing Normal University, Beijing 100875, China"},{"name":"Beijing Key Laboratory of Environmental Remote Sensing and Digital Cities, Faculty of Geographical Science, Beijing Normal University, Beijing 100875, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4091-0175","authenticated-orcid":false,"given":"Hong","family":"Tang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Remote Sensing Science, Faculty of Geographical Science, Beijing Normal University, Beijing 100875, China"},{"name":"Beijing Key Laboratory of Environmental Remote Sensing and Digital Cities, Faculty of Geographical Science, Beijing Normal University, Beijing 100875, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiayi","family":"Ge","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Remote Sensing Science, Faculty of Geographical Science, Beijing Normal University, Beijing 100875, China"},{"name":"Beijing Key Laboratory of Environmental Remote Sensing and Digital Cities, Faculty of Geographical Science, Beijing Normal University, Beijing 100875, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,10,24]]},"reference":[{"key":"ref_1","first-page":"1097","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"2793","DOI":"10.1109\/TPAMI.2017.2750680","article-title":"Learning building extraction in aerial scenes with convolutional networks","volume":"40","author":"Yuan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"7502","DOI":"10.1109\/TGRS.2020.2973720","article-title":"Building footprint generation by integrating convolution neural network with feature pairwise conditional random field (FPCRF)","volume":"58","author":"Li","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Yang, N., and Tang, H. (2021). Semantic Segmentation of Satellite Images: A Deep Learning Approach Integrated with Geospatial Hash Codes. Remote Sens., 13.","DOI":"10.3390\/rs13142723"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2600","DOI":"10.1109\/JSTARS.2018.2835377","article-title":"Building extraction at scale using convolutional neural network: Mapping of the united states","volume":"11","author":"Yang","year":"2018","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"574","DOI":"10.1109\/TGRS.2018.2858817","article-title":"Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set","volume":"57","author":"Ji","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2627","DOI":"10.1109\/JSTARS.2019.2924582","article-title":"Transferable object-based framework based on deep convolutional neural networks for building extraction","volume":"12","author":"Majd","year":"2019","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1842","DOI":"10.1109\/JSTARS.2020.2991391","article-title":"Refined extraction of building outlines from high-resolution remote sensing imagery based on a multifeature convolutional neural network and morphological filtering","volume":"13","author":"Xie","year":"2020","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_9","first-page":"91","article-title":"Faster r-cnn: Towards real-time object detection with region proposal networks","volume":"28","author":"Ren","year":"2015","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 8\u201310). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1998","DOI":"10.1109\/LGRS.2017.2745900","article-title":"Rural building detection in high-resolution imagery based on a two-stage CNN model","volume":"14","author":"Sun","year":"2017","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Yang, N., and Tang, H. (2020). GeoBoost: An incremental deep learning approach toward global mapping of buildings from VHR remote sensing images. Remote Sens., 12.","DOI":"10.3390\/rs12111794"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Cheng, D., Liao, R., Fidler, S., and Urtasun, R. (2019, January 16\u201320). Darnet: Deep active ray network for building segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00761"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Xu, Y., Wu, L., Zhong, X., and Chen, Z. (2018). Building Extraction in Very High Resolution Remote Sensing Imagery Using Deep Learning and Guided Filters. Remote Sens., 10.","DOI":"10.3390\/rs10010144"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Huang, Z., Cheng, G., Wang, H., Li, H., Shi, L., and Pan, C. (2016, January 10\u201315). Building extraction from multi-source remote sensing images via deep deconvolution neural networks. Proceedings of the IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Beijing, China.","DOI":"10.1109\/IGARSS.2016.7729471"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zhao, K., Kang, J., Jung, J., and Sohn, G. (2018, January 18\u201322). Building extraction from satellite images using mask R-CNN with building boundary regularization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00045"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"2178","DOI":"10.1109\/TGRS.2019.2954461","article-title":"Toward automatic building footprint delineation from aerial images using cnn and regularization","volume":"58","author":"Wei","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Girard, N., Smirnov, D., Solomon, J., and Tarabalka, Y. (2020, January 26). Regularized Building Segmentation by Frame Field Learning. Proceedings of the IGARSS 2020\u20132020 IEEE International Geoscience and Remote Sensing Symposium, Brussels, Belgium.","DOI":"10.1109\/IGARSS39084.2020.9324080"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_22","unstructured":"Gur, S., Shaharabany, T., and Wolf, L. (2019). End to end trainable active contours via differentiable rendering. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1016\/j.isprsjprs.2021.02.014","article-title":"Building outline delineation: From aerial images to polygons with an improved end-to-end learning framework","volume":"175","author":"Zhao","year":"2021","journal-title":"ISPRS J. Photogramm."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Li, Z., Wegner, J., and Lucchi, A. (2019, January 27\u201331). Topological map extraction from overhead images. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00180"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Huang, W., Tang, H., and Xu, P. (2021). OEC-RNN: Object-oriented delineation of rooftops with edges and corners using the recurrent neural network from the aerial images. IEEE Trans. Geosci. Remote Sens., Online Early Access.","DOI":"10.1109\/TGRS.2021.3076098"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Castrejon, L., Kundu, K., Urtasun, R., and Fidler, S. (2017). Annotating object instances with a polygon-RNN. arXiv.","DOI":"10.1109\/CVPR.2017.477"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Acuna, D., Ling, H., Kar, A., and Fidler, S. (2018, January 18\u201322). Efficient interactive annotation of segmentation datasets with polygon-rnn++. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00096"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Liu, Y., Cheng, M.-M., Hu, X., Wang, K., and Bai, X. (2017, January 21\u201326). Richer convolutional features for edge detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.622"},{"key":"ref_29","unstructured":"Xingjian, S.H.I., Chen, Z., Wang, H., Yeung, D.-Y., Wong, W.-K., and Woo, W. (2015, January 7\u201312). Convolutional LSTM network: A machine learning approach for precipitation nowcasting. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Perazzi, F., Point-Tuset, J., McWilliams, B., Van Gool, L., Gross, M., and Sorkine-Hornung, A. (2016, January 27\u201330). A benchmark dataset and evaluation methodology for video object segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.85"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Maggiori, E., Tarabalka, Y., Charpiat, G., and Alliez, P. (2017, January 23\u201328). Can semantic labeling methods generalize to any city? The inria aerial image labeling benchmark. Proceedings of the 2017 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Fort Worth, TX, USA.","DOI":"10.1109\/IGARSS.2017.8127684"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1590\/S0104-65002004000100006","article-title":"The Douglas-peucker algorithm: Sufficiency conditions for non-self-intersections","volume":"9","author":"Wu","year":"2004","journal-title":"J. Braz. Comput. Soc."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/21\/4271\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:22:21Z","timestamp":1760167341000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/21\/4271"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,24]]},"references-count":34,"journal-issue":{"issue":"21","published-online":{"date-parts":[[2021,11]]}},"alternative-id":["rs13214271"],"URL":"https:\/\/doi.org\/10.3390\/rs13214271","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,10,24]]}}}