{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T15:11:07Z","timestamp":1783955467332,"version":"3.55.0"},"reference-count":44,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2022,4,29]],"date-time":"2022-04-29T00:00:00Z","timestamp":1651190400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China","award":["61971006"],"award-info":[{"award-number":["61971006"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>In recent years, with the development of deep learning, semantic segmentation for remote sensing images has gradually become a hot issue in computer vision. However, segmentation for multicategory targets is still a difficult problem. To address the issues regarding poor precision and multiple scales in different categories, we propose a UNet, based on multi-attention (MA-UNet). Specifically, we propose a residual encoder, based on a simple attention module, to improve the extraction capability of the backbone for fine-grained features. By using multi-head self-attention for the lowest level feature, the semantic representation of the given feature map is reconstructed, further implementing fine-grained segmentation for different categories of pixels. Then, to address the problem of multiple scales in different categories, we increase the number of down-sampling to subdivide the feature sizes of the target at different scales, and use channel attention and spatial attention in different feature fusion stages, to better fuse the feature information of the target at different scales. We conducted experiments on the WHDLD datasets and DLRSD datasets. The results show that, with multiple visual attention feature enhancements, our method achieves 63.94% mean intersection over union (IOU) on the WHDLD datasets; this result is 4.27% higher than that of UNet, and on the DLRSD datasets, the mean IOU of our methods improves UNet\u2019s 56.17% to 61.90%, while exceeding those of other advanced methods.<\/jats:p>","DOI":"10.3390\/sym14050906","type":"journal-article","created":{"date-parts":[[2022,4,28]],"date-time":"2022-04-28T23:52:30Z","timestamp":1651189950000},"page":"906","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":82,"title":["A Multi-Attention UNet for Semantic Segmentation in Remote Sensing Images"],"prefix":"10.3390","volume":"14","author":[{"given":"Yu","family":"Sun","sequence":"first","affiliation":[{"name":"School of Electronics and Communications Engineering, North China University of Technology, Beijing 100144, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fukun","family":"Bi","sequence":"additional","affiliation":[{"name":"School of Electronics and Communications Engineering, North China University of Technology, Beijing 100144, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yangte","family":"Gao","sequence":"additional","affiliation":[{"name":"School of Information and Electronics, Beijing Institute of Technology, Beijing 100081, China"},{"name":"Qian Xuesen Laboratory of Space Technology, China Academy of Space Technology, Beijing 100094, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Liang","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Information and Electronics, Beijing Institute of Technology, Beijing 100081, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Suting","family":"Feng","sequence":"additional","affiliation":[{"name":"School of Electronics and Communications Engineering, North China University of Technology, Beijing 100144, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,4,29]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Zhan, J., Hu, Y., Cai, W., Zhou, G., and Li, L. (2021). PDAM\u2013STPNNet: A Small Target Detection Approach for Wildland Fire Smoke through Remote Sensing Images. Symmetry, 13.","DOI":"10.3390\/sym13122260"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Wang, S., Sun, X., Liu, P., Xu, K., Wu, C., and Wu, C. (2021). Research on Remote Sensing Image Matching with Special Texture Background. Symmetry, 13.","DOI":"10.3390\/sym13081380"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Kai, Y.K., and Rajendran, P. (2021). A Descriptor-Based Advanced Feature Detector for Improved Visual Tracking. Symmetry, 13.","DOI":"10.3390\/sym13081337"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Ren, Y., Yu, Y., and Guan, H. (2020). DA-CapsUNet: A Dual-Attention Capsule U-Net for Road Extraction from Remote Sensing Imagery. Remote Sens., 12.","DOI":"10.3390\/rs12182866"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1016\/j.infrared.2018.03.012","article-title":"Sea-Land Segmentation for Infrared Remote Sensing Images based on Superpixels and Multi-scale Features","volume":"91","author":"Lei","year":"2018","journal-title":"Infrared Phys. Technol."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Xi, C., Yulong, G., and He, R. (2022). The Use of Remote Sensing to Quantitatively Assess the Visual Effect of Urban Landscape-A Case Study of Zhengzhou. China Remote Sens., 14.","DOI":"10.3390\/rs14010203"},{"key":"ref_7","first-page":"1","article-title":"Multilevel Mapping from Remote Sensing Images: A Case Study of Urban Buildings","volume":"99","author":"Shen","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Abdollahi, A., Pradhan, B., and Shukla, N. (2020). Deep Learning Approaches Applied to Remote Sensing Datasets for Road Extraction: A State-Of-The-Art Review. Remote Sens., 12.","DOI":"10.3390\/rs12091444"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1074","DOI":"10.3390\/rs70101074","article-title":"UAV Remote Sensing for Urban Vegetation Mapping Using Random Forest and Texture Analysis","volume":"7","author":"Feng","year":"2015","journal-title":"Remote Sens."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"2589","DOI":"10.1109\/TGRS.2011.2109389","article-title":"Automatic Image Registration Through Image Segmentation and SIFT","volume":"49","author":"Goncalves","year":"2011","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_11","unstructured":"Wang, Y., Chen, D.R., and Shen, M.L. (2008). Watershed segmentation based on morphological gradient reconstruction. J. Optoelectron. Laser."},{"key":"ref_12","unstructured":"Blake, A., Criminisi, A., and Cross, G. (2010). Image Segmentation of Foreground from Background Layers. (US20100119147 A1), US Patent."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"60","DOI":"10.1016\/j.dsp.2017.02.003","article-title":"Automated segmentation of iris images acquired in an unconstrained environment using HOG-SVM and GrowCut","volume":"64","author":"Radman","year":"2017","journal-title":"Digit. Signal Processing"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Dong, C., Liu, J., and Xu, F. (2019). Ship detection from optical remote sensing images using multi-scale analysis and Fourier HOG descriptor. Remote Sens., 11.","DOI":"10.3390\/rs11131529"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1451","DOI":"10.1109\/LGRS.2015.2408355","article-title":"Unsupervised ship detection based on saliency and S-HOG descriptor from optical satellite images","volume":"12","author":"Qi","year":"2015","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_16","unstructured":"Simonyan, K., Vedaldi, A., and Zisserman, A. (2013). Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps. Comput. Sci., preprint."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015). U-Net: Convolutional Networks for Biomedical Image Segmentation, Springer International Publishing.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., and Qi, X. (2017, January 21\u201326). Pyramid Scene Parsing Network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"4257","DOI":"10.1007\/s11063-021-10592-w","article-title":"CT-UNet: Context-Transfer-UNet for Building Segmentation in Remote Sensing Images","volume":"53","author":"Liu","year":"2021","journal-title":"Neural Processing Lett."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Abdollahi, A., Pradhan, B., and Shukla, N. (2021). Multi-Object Segmentation in Complex Urban Scenes from High-Resolution Remote Sensing Data. Remote Sens., 13.","DOI":"10.3390\/rs13183710"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wang, S., Chen, W., and Xie, S.M. (2020). Weakly supervised deep learning for segmentation of remote sensing imagery. Remote Sens., 12.","DOI":"10.3390\/rs12020207"},{"key":"ref_24","first-page":"1","article-title":"Road Extraction by Deep Residual U-Net","volume":"99","author":"Zhang","year":"2017","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhang, J., Lin, S., and Ding, L. (2020). Multi-scale context aggregation for semantic segmentation of remote sensing images. Remote Sens., 12.","DOI":"10.3390\/rs12040701"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zhou, Z., Siddiquee, M.M., and Tajbakhsh, N. (2018). Unet++: A Nested U-Net Architecture for Medical Image Segmentation. Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support, Springer.","DOI":"10.1007\/978-3-030-00889-5_1"},{"key":"ref_27","unstructured":"Oktay, O., Schlemper, J., and Folgoc, L.L. (2018). Attention u-net: Learning where to look for the pancreas. arXiv."},{"key":"ref_28","unstructured":"Chen, L.Y., and Yu, Q. (2021). TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation. arXiv."},{"key":"ref_29","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., and Lee, J.Y. (2018, January 23\u201328). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"783","DOI":"10.1007\/s11263-019-01283-0","article-title":"A Simple and Light-Weight Attention Module for Convolutional Neural Networks","volume":"128","author":"Park","year":"2020","journal-title":"Int. J. Comput. Vis."},{"key":"ref_32","unstructured":"Roy, A.G., Navab, N., and Wachinger, C. (October, January 27). Concurrent spatial and channel \u201csqueeze & excitation\u201d in fully convolutional networks. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Strasbourg, France."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., and Ren, S. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Xie, S., Girshick, R., and Doll\u00e1r, P. (2017, January 21\u201326). Aggregated residual transformations for deep neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.634"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"318","DOI":"10.1109\/JSTARS.2019.2961634","article-title":"Multilabel Remote Sensing Image Retrieval Based on Fully Convolutional Network","volume":"13","author":"Shao","year":"2020","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Shao, Z., Yang, K., and Zhou, W. (2018). Performance Evaluation of Single-Label and Multi-Label Remote Sensing Image Retrieval Using a Dense Labeling Dataset. Remote Sens., 10.","DOI":"10.3390\/rs10060964"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"7448","DOI":"10.1109\/TGRS.2014.2312793","article-title":"Road Centerline Extraction in Complex Urban Scenes from LiDAR Data Based on Multiple Features","volume":"52","author":"Hu","year":"2014","journal-title":"Geosci. Remote Sens."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Cubuk, E.D., Zoph, B., and Mane, D. (2018). Autoaugment: Learning augmentation policies from data. arXiv.","DOI":"10.1109\/CVPR.2019.00020"},{"key":"ref_40","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Zhou, L., Zhang, C., and Ming, W. (2018, January 18\u201323). D-LinkNet: LinkNet with Pretrained Encoder and Dilated Convolution for High Resolution Satellite Imagery Road Extraction. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00034"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"2532","DOI":"10.3390\/rs13132532","article-title":"SAFFNet: Self-Attention-Based Feature Fusion Network for Remote Sensing Few-Shot Scene Classification","volume":"13","author":"Chi","year":"2021","journal-title":"Remote Sens."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"4597","DOI":"10.3390\/rs13224597","article-title":"Attention-Guided Siamese Fusion Network for Change Detection of Remote Sensing Images","volume":"13","author":"Jiao","year":"2021","journal-title":"Remote Sens."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Bai, T., Li, D., Sun, K., Chen, Y., and Li, W. (2016). Cloud Detection for High-Resolution Satellite Imagery Using Machine Learning and Multi-Feature Fusion. Remote Sensing, 8.","DOI":"10.3390\/rs8090715"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/14\/5\/906\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:03:45Z","timestamp":1760137425000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/14\/5\/906"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,29]]},"references-count":44,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2022,5]]}},"alternative-id":["sym14050906"],"URL":"https:\/\/doi.org\/10.3390\/sym14050906","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,4,29]]}}}