{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T21:14:17Z","timestamp":1784236457802,"version":"3.55.0"},"reference-count":49,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2019,12,10]],"date-time":"2019-12-10T00:00:00Z","timestamp":1575936000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the key research and development task of Sichuan science and technology planning project","award":["2019YFS0067"],"award-info":[{"award-number":["2019YFS0067"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IJGI"],"abstract":"<jats:p>Road extraction is a unique and difficult problem in the field of semantic segmentation because roads have attributes such as slenderness, long span, complexity, and topological connectivity, etc. Therefore, we propose a novel road extraction network, abbreviated HsgNet, based on high-order spatial information global perception network using bilinear pooling. HsgNet, taking the efficient LinkNet as its basic architecture, embeds a Middle Block between the Encoder and Decoder. The Middle Block learns to preserve global-context semantic information, long-distance spatial information and relationships, and different feature channels\u2019 information and dependencies. It is different from other road segmentation methods which lose spatial information, such as those using dilated convolution and multiscale feature fusion to record local-context semantic information. The Middle Block consists of three important steps: (1) forming a feature resource pool to gather high-order global spatial information; (2) selecting a feature weight distribution, enabling each pixel position to obtain complementary features according to its own needs; and (3) inversely mapping the intermediate output feature encoding to the size of the input image by expanding the number of channels of the intermediate output feature. We compared multiple road extraction methods on two open datasets, SpaceNet and DeepGlobe. The results show that compared to the efficient road extraction model D-LinkNet, our model has fewer parameters and better performance: we achieved higher mean intersection over union (71.1%), and the model parameters were reduced in number by about 1\/4.<\/jats:p>","DOI":"10.3390\/ijgi8120571","type":"journal-article","created":{"date-parts":[[2019,12,10]],"date-time":"2019-12-10T10:52:41Z","timestamp":1575975161000},"page":"571","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":55,"title":["HsgNet: A Road Extraction Network Based on Global Perception of High-Order Spatial Information"],"prefix":"10.3390","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4888-3880","authenticated-orcid":false,"given":"Yan","family":"Xie","sequence":"first","affiliation":[{"name":"College of Geophysics, Chengdu University of Technology, Chengdu 610059, China"},{"name":"Geological Team 103, Guizhou Bureau of Geology Mineral Exploration Development, Tongren 554300, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fang","family":"Miao","sequence":"additional","affiliation":[{"name":"Big Data Research Institute, Chengdu University, Chengdu 610106, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kai","family":"Zhou","sequence":"additional","affiliation":[{"name":"College of Computer science, Sichuan University, Chengdu 610065, China"},{"name":"Science and Technology Information Department, Sichuan Provincial Department of Public Security, Chengdu 610041, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jing","family":"Peng","sequence":"additional","affiliation":[{"name":"Science and Technology Information Department, Sichuan Provincial Department of Public Security, Chengdu 610041, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2019,12,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"245","DOI":"10.1016\/j.isprsjprs.2017.02.008","article-title":"Hierarchical graph-based segmentation for extracting road networks from high-resolution satellite images","volume":"126","author":"Alshehhi","year":"2017","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"749","DOI":"10.1109\/LGRS.2018.2802944","article-title":"Road Extraction by Deep Residual U-Net","volume":"15","author":"Zhang","year":"2018","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1186\/s13640-015-0062-9","article-title":"Connected component-based technique for automatic extraction of road centerline in high resolution satellite images","volume":"2015","author":"Sujatha","year":"2015","journal-title":"J. Image Video Proc."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1007\/s001380050121","article-title":"Automatic extraction of roads from aerial images based on scale space and snakes","volume":"12","author":"Laptev","year":"2000","journal-title":"Mach. Vis. Appl."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Zhang, X., Sun, Y., and Zhang, P. (2018). Road Centerline Extraction from Very-High-Resolution Aerial Image and LiDAR Data Based on Road Connectivity. Remote Sens., 10.","DOI":"10.3390\/rs10081284"},{"key":"ref_6","unstructured":"Shelhamer, E., Long, J., and Darrell, T. (2016). Fully Convolutional Networks for Semantic Segmentation. arXiv."},{"key":"ref_7","unstructured":"Chen, L.-C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A.L. (2016). DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2016). Pyramid Scene Parsing Network. arXiv.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015). U-Net: Convolutional Networks for Biomedical Image Segmentation. arXiv.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Chaurasia, A., and Culurciello, E. (2017, January 10\u201313). LinkNet: Exploiting Encoder Representations for Efficient Semantic Segmentation. Proceedings of the 2017 IEEE Visual Communications and Image Processing (VCIP), Saint Petersburg, FL, USA.","DOI":"10.1109\/VCIP.2017.8305148"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Zhou, L., Zhang, C., and Wu, M. (2018, January 18\u201322). D-LinkNet: LinkNet with Pretrained Encoder and Dilated Convolution for High Resolution Satellite Imagery Road Extraction. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00034"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Demir, I., Koperski, K., Lindenbaum, D., Pang, G., Huang, J., Basu, S., Hughes, F., Tuia, D., and Raskar, R. (2018, January 18\u201322). DeepGlobe 2018: A Challenge to Parse the Earth through Satellite Images. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00031"},{"key":"ref_13","unstructured":"Van Etten, A., Lindenbaum, D., and Bacastow, T.M. (2018). SpaceNet: A Remote Sens. Dataset and Challenge Series. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Wegner, J.D., Montoya-Zegarra, J.A., and Schindler, K. (2013, January 23\u201328). A Higher-Order CRF Model for Road Network Extraction. Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.222"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Chai, D., Forstner, W., and Lafarge, F. (2013, January 23\u201328). Recovering Line-Networks in Images by Junction-Point Processes. Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.247"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Liu, J., Qin, Q., Li, J., and Li, Y. (2017). Rural Road Extraction from High-Resolution Remote Sens. Images Based on Geometric Feature Inference. IJGI, 6.","DOI":"10.3390\/ijgi6100314"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1365","DOI":"10.14358\/PERS.70.12.1365","article-title":"Road Extraction Using SVM and Image Segmentation","volume":"70","author":"Song","year":"2004","journal-title":"Photogramm. Eng. Remote Sens."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"3906","DOI":"10.1109\/TGRS.2011.2136381","article-title":"Use of Salient Features for the Design of a Multistage Framework to Extract Roads From High-Resolution Multispectral Satellite Images","volume":"49","author":"Das","year":"2011","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_19","unstructured":"Mnih, V. (2013). Machine Learning for Aerial Image Labeling. [Ph.D. Thesis, Department of Computer Science, University of Toronto]."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"104021","DOI":"10.2352\/J.ImagingSci.Technol.2016.60.1.010402","article-title":"Multiple Object Extraction from Aerial Imagery with Convolutional Neural Networks","volume":"60","author":"Saito","year":"2016","journal-title":"J. Imaging Sci. Technol."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Bastani, F., He, S., Abbar, S., Alizadeh, M., Balakrishnan, H., Chawla, S., Madden, S., and DeWitt, D. (2018, January 18\u201322). RoadTracer: Automatic Extraction of Road Networks from Aerial Images. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00496"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Xia, W., Zhang, Y.-Z., Liu, J., Luo, L., and Yang, K. (2018). Road Extraction from High Resolution Image with Deep Convolution Network\u2014A Case Study of GF-2 Image. Proceedings, 2.","DOI":"10.3390\/ecrs-2-05138"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Batra, A., Singh, S., Pang, G., Basu, S., Jawahar, C.V., and Paluri, M. (2019, January 16\u201320). Improved Road Connectivity by Joint Learning of Orientation and Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01063"},{"key":"ref_24","unstructured":"(2018). Qiqi Zhu; Yanfei Zhong; Yanfei Liu; Liangpei Zhang; Deren Li A Deep-Local-Global Feature Fusion Framework for High Spatial Resolution Imagery Scene Classification. Remote Sens., 10."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Xu, Y., Xie, Z., Feng, Y., and Chen, Z. (2018). Road Extraction from High-Resolution Remote Sens. Imagery Using Deep Learning. Remote Sens., 10.","DOI":"10.3390\/rs10091461"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., RoyChowdhury, A., and Maji, S. (2015, January 7\u201313). Bilinear CNN Models for Fine-Grained Visual Recognition. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.170"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Gao, Y., Beijbom, O., Zhang, N., and Darrell, T. (July, January 26). Compact Bilinear Pooling. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.41"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Fukui, A., Park, D.H., Yang, D., Rohrbach, A., Darrell, T., and Rohrbach, M. (2016, January 2\u20136). Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Austin, TX, USA.","DOI":"10.18653\/v1\/D16-1044"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Kong, S., and Fowlkes, C. (2017, January 21\u201326). Low-Rank Bilinear Pooling for Fine-Grained Classification. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.743"},{"key":"ref_30","unstructured":"Kim, J.-H., and On, K.-W. (2017). Hadamard Product for Low-Rank Bilinear Pooling. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"365","DOI":"10.1007\/978-3-030-01219-9_22","article-title":"Grassmann Pooling as Compact Homogeneous Bilinear Pooling for Fine-Grained Visual Classification","volume":"Volume 11207","author":"Ferrari","year":"2018","journal-title":"Computer Vision\u2014ECCV 2018"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"5947","DOI":"10.1109\/TNNLS.2018.2817340","article-title":"Beyond Bilinear: Generalized Multimodal Factorized High-order Pooling for Visual Question Answering","volume":"29","author":"Yu","year":"2018","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"595","DOI":"10.1007\/978-3-030-01270-0_35","article-title":"Hierarchical Bilinear Pooling for Fine-Grained Visual Recognition","volume":"Volume 11220","author":"Ferrari","year":"2018","journal-title":"Computer Vision\u2014ECCV 2018"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Li, P., Xie, J., Wang, Q., and Zuo, W. (2017, January 22\u201329). Is Second-Order Information Helpful for Large-Scale Visual Recognition?. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.228"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"430","DOI":"10.1007\/978-3-642-33786-4_32","article-title":"Semantic Segmentation with Second-Order Pooling","volume":"Volume 7578","author":"Fitzgibbon","year":"2012","journal-title":"Computer Vision\u2014ECCV 2012"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2015). Deep Residual Learning for Image Recognition. arXiv.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). ImageNet: A Large-Scale Hierarchical Image Database. Proceedings of the Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_38","unstructured":"Yosinski, J., Clune, J., Bengio, Y., and Lipson, H. (2014). How transferable are features in deep neural networks?. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Zeiler, M.D., Taylor, G.W., and Fergus, R. (2011, January 6\u201313). Adaptive deconvolutional networks for mid and high level feature learning. Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126474"},{"key":"ref_40","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is All you Need. Proceedings of the Neural Information Processing Systems (NIPS), Long Beach, CA, USA."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Chen, L.-C., Yang, Y., Wang, J., Xu, W., and Yuille, A.L. (2016). Attention to Scale: Scale-aware Semantic Image Segmentation. arXiv.","DOI":"10.1109\/CVPR.2016.396"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Liu, M., and Yin, H. (2019). Cross Attention Network for Semantic Segmentation. arXiv.","DOI":"10.1109\/ICIP.2019.8803320"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., and Lu, H. (2018). Dual Attention Network for Scene Segmentation. arXiv.","DOI":"10.1109\/CVPR.2019.00326"},{"key":"ref_44","first-page":"1","article-title":"Squeeze-and-Excitation Networks","volume":"10","author":"Hu","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., Gupta, A., and He, K. (2018, January 18\u201322). Non-local Neural Networks. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"ref_46","first-page":"352","article-title":"A^2-Nets: Double Attention Networks","volume":"10","author":"Chen","year":"2018","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_47","unstructured":"Kingma, D.P., and Lei, J. (2015). Adam: A Method for Stochastic Optimization. arXiv."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Huang, Z., Wang, X., Huang, L., Huang, C., Wei, Y., and Liu, W. (2018). CCNet: Criss-Cross Attention for Semantic Segmentation. arXiv.","DOI":"10.1109\/ICCV.2019.00069"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"270","DOI":"10.1007\/978-3-030-01240-3_17","article-title":"PSANet: Point-wise Spatial Attention Network for Scene Parsing","volume":"Volume 11213","author":"Ferrari","year":"2018","journal-title":"Computer Vision\u2014ECCV 2018"}],"container-title":["ISPRS International Journal of Geo-Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2220-9964\/8\/12\/571\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:41:07Z","timestamp":1760190067000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2220-9964\/8\/12\/571"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,12,10]]},"references-count":49,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2019,12]]}},"alternative-id":["ijgi8120571"],"URL":"https:\/\/doi.org\/10.3390\/ijgi8120571","relation":{},"ISSN":["2220-9964"],"issn-type":[{"value":"2220-9964","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,12,10]]}}}