{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T04:22:36Z","timestamp":1783570956127,"version":"3.55.0"},"reference-count":43,"publisher":"MDPI AG","issue":"20","license":[{"start":{"date-parts":[[2021,10,16]],"date-time":"2021-10-16T00:00:00Z","timestamp":1634342400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key R &amp; D Program of China","award":["2018YFC0810600, 2018YFC0810605"],"award-info":[{"award-number":["2018YFC0810600, 2018YFC0810605"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Traditional pixel-based semantic segmentation methods for road extraction take each pixel as the recognition unit. Therefore, they are constrained by the restricted receptive field, in which pixels do not receive global road information. These phenomena greatly affect the accuracy of road extraction. To improve the limited receptive field, a non-local neural network is generated to let each pixel receive global information. However, its spatial complexity is enormous, and this method will lead to considerable information redundancy in road extraction. To optimize the spatial complexity, the Crisscross Network (CCNet), with a crisscross shaped attention area, is applied. The key aspect of CCNet is the Crisscross Attention (CCA) module. Compared with non-local neural networks, CCNet can let each pixel only perceive the correlation information from horizontal and vertical directions. However, when using CCNet in road extraction of remote sensing (RS) images, the directionality of its attention area is insufficient, which is restricted to the horizontal and vertical direction. Due to the recurrent mechanism, the similarity of some pixel pairs in oblique directions cannot be calculated correctly and will be intensely dilated. To address the above problems, we propose a special attention module called the Dual Crisscross Attention (DCCA) module for road extraction, which consists of the CCA module, Rotated Crisscross Attention (RCCA) module and Self-adaptive Attention Fusion (SAF) module. The DCCA module is embedded into the Dual Crisscross Network (DCNet). In the CCA module and RCCA module, the similarities of pixel pairs are represented by an energy map. In order to remove the influence from the heterogeneous part, a heterogeneous filter function (HFF) is used to filter the energy map. Then the SAF module can distribute the weights of the CCA module and RCCA module according to the actual road shape. The DCCA module output is the fusion of the CCA module and RCCA module with the help of the SAF module, which can let pixels perceive local information and eight-direction non-local information. The geometric information of roads improves the accuracy of road extraction. The experimental results show that DCNet with the DCCA module improves the road IOU by 4.66% compared to CCNet with a single CCA module and 3.47% compared to CCNet with a single RCCA module.<\/jats:p>","DOI":"10.3390\/s21206873","type":"journal-article","created":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T23:25:15Z","timestamp":1634513115000},"page":"6873","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Dual Crisscross Attention Module for Road Extraction from Remote Sensing Images"],"prefix":"10.3390","volume":"21","author":[{"given":"Chuan","family":"Chen","sequence":"first","affiliation":[{"name":"TUM Department of Aerospace and Geodesy, Technical University of Munich, 80333 Munich, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huilin","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Resources and Environmental Engineering, Wuhan University of Technology, Wuhan 430070, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wei","family":"Cui","sequence":"additional","affiliation":[{"name":"School of Resources and Environmental Engineering, Wuhan University of Technology, Wuhan 430070, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xin","family":"He","sequence":"additional","affiliation":[{"name":"School of Resources and Environmental Engineering, Wuhan University of Technology, Wuhan 430070, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2021,10,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Geng, K., Sun, X., Yan, Z., Diao, W., and Gao, X. (2020). Topological Space Knowledge Distillation for Compact Road Extraction in Optical Remote Sensing Images. Remote Sens., 12.","DOI":"10.3390\/rs12193175"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., Gupta, A., and He, K. (2018, January 18\u201323). Non-local neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"ref_3","first-page":"330","article-title":"An road extraction method for remote sensing image based on Encoder-Decoder network","volume":"48","author":"He","year":"2019","journal-title":"Acta Geod. Cartogr. Sin."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_5","unstructured":"Luo, W., Li, Y., Urtasun, R., and Zemel, R. (2016). Understanding the effective receptive field in deep convolutional neural networks. arXiv."},{"key":"ref_6","unstructured":"Liu, W., Rabinovich, A., and Berg, A.C. (2015). Parsenet: Looking wider to see better. arXiv."},{"key":"ref_7","unstructured":"Yu, F., and Koltun, V. (2015). Multi-scale context aggregation by dilated convolutions. arXiv."},{"key":"ref_8","unstructured":"Huang, Z., Wang, X., Huang, L., Huang, C., Wei, Y., and Liu, W. (November, January 27). Ccnet: Criss-cross attention for semantic segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European conference on computer vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_11","first-page":"1097","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015). U-Net: Convolutional Networks for Biomedical Image Segmentation. Medical Image Computing and Computer-Assisted Intervention\u2014MICCAI 2015, Springer.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_15","unstructured":"Chen, L.C., Papandreou, G., Schroff, F., and Adam, H. (2017). Rethinking atrous convolution for semantic image segmentation. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. (2018, January 8\u201314). Encoder-decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 21\u201326). Pyramid scene parsing network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zhao, H., Zhang, Y., Liu, S., Shi, J., Loy, C.C., Lin, D., and Jia, J. (2018, January 8\u201314). Psanet: Point-wise spatial attention network for scene parsing. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01240-3_17"},{"key":"ref_19","unstructured":"Mnih, V., Heess, N., and Graves, A. (2014). Recurrent models of visual attention. arXiv."},{"key":"ref_20","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv."},{"key":"ref_21","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017). Attention is all you need. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., and Lu, H. (2019, January 15\u201320). Dual attention network for scene segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00326"},{"key":"ref_24","unstructured":"Ulku, I., and Akagunduz, E. (2019). A Survey on Deep Learning-based Architectures for Semantic Segmentation on 2D images. arXiv."},{"key":"ref_25","unstructured":"Cordonnier, J.B., Loukas, A., and Jaggi, M. (2019). On the Relationship between Self-Attention and Convolutional Layers. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1977","DOI":"10.1080\/01431160802546837","article-title":"Road centreline extraction from high-resolution imagery based on multiscale structural features and support vector machines","volume":"30","author":"Huang","year":"2009","journal-title":"Int. J. Remote Sens."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Mnih, V., and Hinton, G.E. (2010). Learning to detect roads in high-resolution aerial images. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-642-15567-3_16"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"4441","DOI":"10.1109\/TGRS.2012.2190078","article-title":"Road network detection using probabilistic and graph theoretical methods","volume":"50","author":"Unsalan","year":"2012","journal-title":"IEEE Trans. Geosci. Remote. Sens."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Cheng, G., Wang, Y., Gong, Y., Zhu, F., and Pan, C. (2014, January 27\u201330). Urban road extraction via graph cuts based probability propagation. Proceedings of the 2014 IEEE International Conference on Image Processing (ICIP), Paris, France.","DOI":"10.1109\/ICIP.2014.7026027"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1","DOI":"10.2352\/ISSN.2470-1173.2016.10.ROBVIS-392","article-title":"Multiple object extraction from aerial imagery with convolutional neural networks","volume":"2016","author":"Saito","year":"2016","journal-title":"Electron. Imaging"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"245","DOI":"10.1016\/j.isprsjprs.2017.02.008","article-title":"Hierarchical graph-based segmentation for extracting road networks from high-resolution satellite images","volume":"126","author":"Alshehhi","year":"2017","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Liu, B., Wu, H., Wang, Y., and Liu, W. (2015). Main road extraction from zy-3 grayscale imagery based on directional mathematical morphology and vgi prior knowledge in urban areas. PLoS ONE, 10.","DOI":"10.1371\/journal.pone.0138071"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s13640-015-0062-9","article-title":"Connected component-based technique for automatic extraction of road centerline in high resolution satellite images","volume":"2015","author":"Sujatha","year":"2015","journal-title":"EURASIP J. Image Video Process."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"3322","DOI":"10.1109\/TGRS.2017.2669341","article-title":"Automatic road detection and centerline extraction via cascaded end-to-end convolutional neural network","volume":"55","author":"Cheng","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"545","DOI":"10.1109\/LGRS.2016.2524025","article-title":"Road centerline extraction via semisupervised segmentation and multidirection nonmaximum suppression","volume":"13","author":"Cheng","year":"2016","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"89","DOI":"10.1007\/BF02826642","article-title":"Automatic road change detection and GIS updating from high spatial remotely-sensed imagery","volume":"7","author":"Qiaoping","year":"2004","journal-title":"Geo-Spat. Inf. Sci."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"1365","DOI":"10.14358\/PERS.70.12.1365","article-title":"Road extraction using SVM and image segmentation","volume":"70","author":"Song","year":"2004","journal-title":"Photogramm. Eng. Remote Sens."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"3906","DOI":"10.1109\/TGRS.2011.2136381","article-title":"Use of salient features for the design of a multistage framework to extract roads from high-resolution multispectral satellite images","volume":"49","author":"Das","year":"2011","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Alvarez, J.M., Gevers, T., LeCun, Y., and Lopez, A.M. (2012). Road scene segmentation from a single image. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-642-33786-4_28"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Zhong, Z., Li, J., Cui, W., and Jiang, H. (2016, January 10\u201315). Fully convolutional networks for building and road extraction: Preliminary results. Proceedings of the 2016 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Beijing, China.","DOI":"10.1109\/IGARSS.2016.7729406"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"749","DOI":"10.1109\/LGRS.2018.2802944","article-title":"Road extraction by deep residual u-net","volume":"15","author":"Zhang","year":"2018","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Peng, B., Li, Y., Fan, K., Yuan, L., Tong, L., and He, L. (August, January 28). New Network Based on D-Linknet and Densenet for High Resolution Satellite Imagery Road Extraction. Proceedings of the IGARSS 2019-2019 IEEE International Geoscience and Remote Sensing Symposium, Yokohama, Japan.","DOI":"10.1109\/IGARSS.2019.8898640"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Demir, I., Koperski, K., Lindenbaum, D., Pang, G., Huang, J., Basu, S., Hughes, F., Tuia, D., and Raskar, R. (2018, January 18\u201322). DeepGlobe 2018: A Challenge to Parse the Earth through Satellite Images. Proceedings of the2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00031"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/20\/6873\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:16:16Z","timestamp":1760166976000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/20\/6873"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,16]]},"references-count":43,"journal-issue":{"issue":"20","published-online":{"date-parts":[[2021,10]]}},"alternative-id":["s21206873"],"URL":"https:\/\/doi.org\/10.3390\/s21206873","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,10,16]]}}}