{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T17:06:56Z","timestamp":1783444016162,"version":"3.54.6"},"reference-count":44,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2023,2,15]],"date-time":"2023-02-15T00:00:00Z","timestamp":1676419200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the National Key Research and Development Program of China","award":["2018YFC0807504"],"award-info":[{"award-number":["2018YFC0807504"]}]},{"name":"the National Key Research and Development Program of China","award":["21-1-2-18-xx"],"award-info":[{"award-number":["21-1-2-18-xx"]}]},{"name":"The Key R&amp;D Projects of Qingdao Science and Technology Plan","award":["2018YFC0807504"],"award-info":[{"award-number":["2018YFC0807504"]}]},{"name":"The Key R&amp;D Projects of Qingdao Science and Technology Plan","award":["21-1-2-18-xx"],"award-info":[{"award-number":["21-1-2-18-xx"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>The accurate detection and extraction of roads using remote sensing technology are crucial to the development of the transportation industry and intelligent perception tasks. Recently, in view of the advantages of CNNs in feature extraction, its related road extraction methods have been proposed successively. However, due to the limitation of kernel size, they perform less effectively at capturing long-range information and global context, which are crucial for road targets distributed over long distances and highly structured. To deal with this problem, a novel model named RoadFormer with a Swin Transformer as the backbone is developed in this paper. Firstly, to extract long-range information effectively, a Swin Transformer multi-scale encoder is adopted in our model. Secondly, to enhance the feature representation capability of the model, we design an innovative bottleneck module, in which the spatial and channel separable convolution is employed to obtain fine-grained and globe features, and then a dilated block is connected after the spatial convolution module to capture more integrated road structures. Finally, a lightweight decoder consisting of transposed convolution and skip connection generates the final extraction results. Extensive experimental results confirm the advantages of RoadFormer on the Deepglobe and Massachusetts datasets. The comparative results of visualization and quantification demonstrate that our model outperforms comparable methods.<\/jats:p>","DOI":"10.3390\/rs15041049","type":"journal-article","created":{"date-parts":[[2023,2,15]],"date-time":"2023-02-15T02:31:23Z","timestamp":1676428283000},"page":"1049","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":45,"title":["RoadFormer: Road Extraction Using a Swin Transformer Combined with a Spatial and Channel Separable Convolution"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2751-6096","authenticated-orcid":false,"given":"Xiangzeng","family":"Liu","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ziyao","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinting","family":"Wan","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Juli","family":"Zhang","sequence":"additional","affiliation":[{"name":"Academy of Advanced Interdisciplinary Research, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yue","family":"Xi","sequence":"additional","affiliation":[{"name":"Guangzhou Institute of Technology, Xidian University, Guangzhou 510555, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ruyi","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2872-388X","authenticated-orcid":false,"given":"Qiguang","family":"Miao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,2,15]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"8919","DOI":"10.1109\/TGRS.2020.2991733","article-title":"Simultaneous Road Surface and Centerline Extraction from Large-Scale Remote Sensing Images Using CNN-Based Segmentation and Tracing","volume":"58","author":"Wei","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"107141","DOI":"10.1016\/j.patcog.2019.107141","article-title":"A Fusion Network for Road Detection via Spatial Propagation and Spatial Transformation","volume":"100","author":"Yang","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1120","DOI":"10.1016\/j.patrec.2009.12.018","article-title":"Advanced Directional Mathematical Morphology for the Detection of the Road Network in Very High Resolution Remote Sensing Images","volume":"31","author":"Valero","year":"2010","journal-title":"Pattern Recognit. Lett."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1538","DOI":"10.1109\/JSTARS.2012.2199085","article-title":"Semi-Automated Road Detection From High Resolution Satellite Images by Directional Morphological Enhancement and Segmentation Techniques","volume":"5","author":"Chaudhuri","year":"2012","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1858","DOI":"10.1109\/LGRS.2015.2431268","article-title":"Automatic Road Extraction From Remote Sensing Images Based on a Normalized Second Derivative Map","volume":"12","author":"Bae","year":"2015","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"749","DOI":"10.1109\/LGRS.2018.2802944","article-title":"Road Extraction by Deep Residual U-Net","volume":"15","author":"Zhang","year":"2018","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Raj, J.S., Iliyasu, A.M., Bestak, R., and Baig, Z.A. (2021). Innovative Data Communication Technologies and Application, Springer.","DOI":"10.1007\/978-981-15-9651-3"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Abdollahi, A., Pradhan, B., Shukla, N., Chakraborty, S., and Alamri, A. (2020). Deep Learning Approaches Applied to Remote Sensing Datasets for Road Extraction: A State-Of-The-Art Review. Remote Sens., 12.","DOI":"10.3390\/rs12091444"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Mendes, C.C.T., Fr\u00e9mont, V., and Wolf, D.F. (2016, January 16\u201320). Exploiting Fully Convolutional Neural Networks for Fast Road Detection. Proceedings of the 2016 IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden.","DOI":"10.1109\/ICRA.2016.7487486"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"245","DOI":"10.1016\/j.isprsjprs.2017.02.008","article-title":"Hierarchical Graph-Based Segmentation for Extracting Road Networks from High-Resolution Satellite Images","volume":"126","author":"Alshehhi","year":"2017","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Costea, D., and Leordeanu, M. (2016). Aerial Image Geolocalization from Recognition and Matching of Roads and Intersections. arXiv.","DOI":"10.5244\/C.30.118"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Bastani, F., He, S., Abbar, S., Alizadeh, M., Balakrishnan, H., Chawla, S., Madden, S., and DeWitt, D. (2018, January 18\u201322). RoadTracer: Automatic Extraction of Road Networks from Aerial Images. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00496"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"2043","DOI":"10.1109\/TGRS.2018.2870871","article-title":"RoadNet: Learning to Comprehensively Analyze Road Networks in Complex Urban Scenes from High-Resolution Remotely Sensed Images","volume":"57","author":"Liu","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_14","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An Image Is Worth 16 \u00d7 16 Words: Transformers for Image Recognition at Scale. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Liu, X., Gao, H., Miao, Q., Xi, Y., Ai, Y., and Gao, D. (2022). MFST: Multi-Modal Feature Self-Adaptive Transformer for Infrared and Visible Image Fusion. Remote Sens., 14.","DOI":"10.3390\/rs14133233"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zhong, Z., Li, J., Cui, W., and Jiang, H. (2016, January 10\u201315). Fully Convolutional Networks for Building and Road Extraction: Preliminary Results. Proceedings of the 2016 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Beijing, China.","DOI":"10.1109\/IGARSS.2016.7729406"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-Net: Convolutional Networks for Biomedical Image Segmentation. Proceedings of the Medical Image Computing and Computer-Assisted Intervention\u2014MICCAI 2015, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhou, Z., Siddiquee, M.M.R., Tajbakhsh, N., and Liang, J. (2018). UNet++: A Nested U-Net Architecture for Medical Image Segmentation. CoRR, Available online: https:\/\/link.springer.com\/chapter\/10.1007\/978-3-030-00889-5_1.","DOI":"10.1007\/978-3-030-00889-5_1"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 21\u201326). Pyramid Scene Parsing Network. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. (2018, January 8\u201314). Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. Proceedings of the Computer Vision\u2014ECCV 2018, Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Chaurasia, A., and Culurciello, E. (2017, January 10\u201313). LinkNet: Exploiting Encoder Representations for Efficient Semantic Segmentation. Proceedings of the 2017 IEEE Visual Communications and Image Processing (VCIP), St. Petersburg, FL, USA.","DOI":"10.1109\/VCIP.2017.8305148"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhou, L., Zhang, C., and Wu, M. (2018, January 18\u201322). D-LinkNet: LinkNet with Pretrained Encoder and Dilated Convolution for High Resolution Satellite Imagery Road Extraction. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00034"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Li, J., Liu, Y., Zhang, Y., and Zhang, Y. (2021). Cascaded Attention DenseUNet (CADUNet) for Road Extraction from Very-High-Resolution Images. ISPRS Int. J. Geo-Inf., 10.","DOI":"10.3390\/ijgi10050329"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Newell, A., Yang, K., and Deng, J. (2016, January 11\u201314). Stacked Hourglass Networks for Human Pose Estimation. Proceedings of the Computer Vision\u2014ECCV 2016, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Batra, A., Singh, S., Pang, G., Basu, S., Jawahar, C., and Paluri, M. (2019, January 15\u201320). Improved Road Connectivity by Joint Learning of Orientation and Segmentation. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01063"},{"key":"ref_28","first-page":"5614115","article-title":"Split Depth-Wise Separable Graph-Convolution Network for Road Extraction in Complex Environments From High-Resolution Remote-Sensing Images","volume":"60","author":"Zhou","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_29","unstructured":"Cao, H., Wang, Y., Chen, J., Jiang, D., Zhang, X., Tian, Q., and Wang, M. (2021). Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"135","DOI":"10.1007\/s41064-022-00194-z","article-title":"Survey of Road Extraction Methods in Remote Sensing Images Based on Deep Learning","volume":"90","author":"Liu","year":"2022","journal-title":"PFG"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"626","DOI":"10.1109\/JSTARS.2010.2094181","article-title":"Application of a Fast Linear Feature Detector to Road Extraction From Remotely Sensed Imagery","volume":"4","author":"Shao","year":"2011","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1016\/j.isprsjprs.2017.05.002","article-title":"Simultaneous Extraction of Roads and Buildings in Remote Sensing Imagery with Convolutional Neural Networks","volume":"130","author":"Alshehhi","year":"2017","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Cui, F., Feng, R., Wang, L., and Wei, L. (2021, January 11\u201316). Joint Superpixel Segmentation and Graph Convolutional Network Road Extration for High-Resolution Remote Sensing Imagery. Proceedings of the 2021 IEEE International Geoscience and Remote Sensing Symposium IGARSS, Brussels, Belgium.","DOI":"10.1109\/IGARSS47720.2021.9554635"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"016020","DOI":"10.1117\/1.JRS.12.016020","article-title":"UFCN: A Fully Convolutional Neural Network for Road Extraction in RGB Imagery Acquired by Remote Sensing from an Unmanned Aerial Vehicle","volume":"12","author":"Kestur","year":"2018","journal-title":"J. Appl. Remote Sens."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Varia, N., Dokania, A., and Senthilnath, J. (2018, January 18\u201321). DeepExt: A Convolution Neural Network for Road Extraction Using RGB Images Captured by UAV. Proceedings of the 2018 IEEE Symposium Series on Computational Intelligence (SSCI), Bengaluru, India.","DOI":"10.1109\/SSCI.2018.8628717"},{"key":"ref_36","unstructured":"Oktay, O., Schlemper, J., Folgoc, L.L., Lee, M., Heinrich, M., Misawa, K., Mori, K., McDonagh, S., Hammerla, N.Y., and Kainz, B. (2018). Attention U-Net: Learning Where to Look for the Pancreas. arXiv."},{"key":"ref_37","unstructured":"Chen, L.-C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A.L. (2014). Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected Crfs. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_39","first-page":"5998","article-title":"Attention Is All You Need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_40","unstructured":"Park, N., and Kim, S. (2022). How Do Vision Transformers Work?. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"155","DOI":"10.1016\/j.isprsjprs.2019.10.001","article-title":"Spatial Information Inference Net: Road Extraction Using Road-Specific Contextual Information","volume":"158","author":"Tao","year":"2019","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_42","unstructured":"Sifre, L., and Mallat, S. (2014). Rigid-Motion Scattering for Texture Classification. arXiv."},{"key":"ref_43","unstructured":"Yu, F., and Koltun, V. (2015). Multi-Scale Context Aggregation by Dilated Convolutions. arXiv."},{"key":"ref_44","first-page":"102833","article-title":"Road Extraction in Remote Sensing Data: A Survey","volume":"112","author":"Chen","year":"2022","journal-title":"Int. J. Appl. Earth Obs. Geoinf."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/4\/1049\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:36:02Z","timestamp":1760121362000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/4\/1049"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,15]]},"references-count":44,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2023,2]]}},"alternative-id":["rs15041049"],"URL":"https:\/\/doi.org\/10.3390\/rs15041049","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,15]]}}}