{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T23:22:52Z","timestamp":1781738572656,"version":"3.54.5"},"reference-count":39,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2023,3,31]],"date-time":"2023-03-31T00:00:00Z","timestamp":1680220800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62005307"],"award-info":[{"award-number":["62005307"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61975228"],"award-info":[{"award-number":["61975228"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2019GXRC042"],"award-info":[{"award-number":["2019GXRC042"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["YJKYYQ20210031"],"award-info":[{"award-number":["YJKYYQ20210031"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["YJKYYQ20180032"],"award-info":[{"award-number":["YJKYYQ20180032"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2021135"],"award-info":[{"award-number":["2021135"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Jinan Innovation Team","award":["62005307"],"award-info":[{"award-number":["62005307"]}]},{"name":"Jinan Innovation Team","award":["61975228"],"award-info":[{"award-number":["61975228"]}]},{"name":"Jinan Innovation Team","award":["2019GXRC042"],"award-info":[{"award-number":["2019GXRC042"]}]},{"name":"Jinan Innovation Team","award":["YJKYYQ20210031"],"award-info":[{"award-number":["YJKYYQ20210031"]}]},{"name":"Jinan Innovation Team","award":["YJKYYQ20180032"],"award-info":[{"award-number":["YJKYYQ20180032"]}]},{"name":"Jinan Innovation Team","award":["2021135"],"award-info":[{"award-number":["2021135"]}]},{"name":"Scientisc Research and Equipment Development Project of Chinese Academy of Sciences","award":["62005307"],"award-info":[{"award-number":["62005307"]}]},{"name":"Scientisc Research and Equipment Development Project of Chinese Academy of Sciences","award":["61975228"],"award-info":[{"award-number":["61975228"]}]},{"name":"Scientisc Research and Equipment Development Project of Chinese Academy of Sciences","award":["2019GXRC042"],"award-info":[{"award-number":["2019GXRC042"]}]},{"name":"Scientisc Research and Equipment Development Project of Chinese Academy of Sciences","award":["YJKYYQ20210031"],"award-info":[{"award-number":["YJKYYQ20210031"]}]},{"name":"Scientisc Research and Equipment Development Project of Chinese Academy of Sciences","award":["YJKYYQ20180032"],"award-info":[{"award-number":["YJKYYQ20180032"]}]},{"name":"Scientisc Research and Equipment Development Project of Chinese Academy of Sciences","award":["2021135"],"award-info":[{"award-number":["2021135"]}]},{"name":"Jiangsu Key Disciplines of the Fourteenth Five-Year Plan","award":["62005307"],"award-info":[{"award-number":["62005307"]}]},{"name":"Jiangsu Key Disciplines of the Fourteenth Five-Year Plan","award":["61975228"],"award-info":[{"award-number":["61975228"]}]},{"name":"Jiangsu Key Disciplines of the Fourteenth Five-Year Plan","award":["2019GXRC042"],"award-info":[{"award-number":["2019GXRC042"]}]},{"name":"Jiangsu Key Disciplines of the Fourteenth Five-Year Plan","award":["YJKYYQ20210031"],"award-info":[{"award-number":["YJKYYQ20210031"]}]},{"name":"Jiangsu Key Disciplines of the Fourteenth Five-Year Plan","award":["YJKYYQ20180032"],"award-info":[{"award-number":["YJKYYQ20180032"]}]},{"name":"Jiangsu Key Disciplines of the Fourteenth Five-Year Plan","award":["2021135"],"award-info":[{"award-number":["2021135"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>This study aimed to address the problems of low detection accuracy and inaccurate positioning of small-object detection in remote sensing images. An improved architecture based on the Swin Transformer and YOLOv5 is proposed. First, Complete-IOU (CIOU) was introduced to improve the K-means clustering algorithm, and then an anchor of appropriate size for the dataset was generated. Second, a modified CSPDarknet53 structure combined with Swin Transformer was proposed to retain sufficient global context information and extract more differentiated features through multi-head self-attention. Regarding the path-aggregation neck, a simple and efficient weighted bidirectional feature pyramid network was proposed for effective cross-scale feature fusion. In addition, extra prediction head and new feature fusion layers were added for small objects. Finally, Coordinate Attention (CA) was introduced to the YOLOv5 network to improve the accuracy of small-object features in remote sensing images. Moreover, the effectiveness of the proposed method was demonstrated by several kinds of experiments on the DOTA (Dataset for Object detection in Aerial images). The mean average precision on the DOTA dataset reached 74.7%. Compared with YOLOv5, the proposed method improved the mean average precision (mAP) by 8.9%, which can achieve a higher accuracy of small-object detection in remote sensing images.<\/jats:p>","DOI":"10.3390\/s23073634","type":"journal-article","created":{"date-parts":[[2023,3,31]],"date-time":"2023-03-31T03:02:01Z","timestamp":1680231721000},"page":"3634","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":60,"title":["Swin-Transformer-Based YOLOv5 for Small-Object Detection in Remote Sensing Images"],"prefix":"10.3390","volume":"23","author":[{"given":"Xuan","family":"Cao","sequence":"first","affiliation":[{"name":"School of Physical Science and Technology, Suzhou University of Science and Technology, Suzhou 215009, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanwei","family":"Zhang","sequence":"additional","affiliation":[{"name":"Suzhou Institute of Biomedical Engineering and Technology, Chinese Academy of Sciences, Suzhou 215613, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Song","family":"Lang","sequence":"additional","affiliation":[{"name":"Suzhou Institute of Biomedical Engineering and Technology, Chinese Academy of Sciences, Suzhou 215613, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yan","family":"Gong","sequence":"additional","affiliation":[{"name":"Suzhou Institute of Biomedical Engineering and Technology, Chinese Academy of Sciences, Suzhou 215613, China"},{"name":"Jinan Guoke Medical Technology Development Co., Ltd., Jinan 250104, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,31]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Zhu, X., Lyu, S., Wang, X., and Zhao, Q. (2021, January 11\u201317). In TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00312"},{"key":"ref_2","unstructured":"Ding, Y. (2020). Research and Implementation of Small Target Detection Network in Complex Background. [Master\u2019s Thesis, Beijing University of Posts and Telecommunications]."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"8509","DOI":"10.1007\/s13369-021-05471-4","article-title":"An improved faster-RCNN model for handwritten character recognition","volume":"46","author":"Albahli","year":"2021","journal-title":"Arab. J. Sci. Eng."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., and Berg, A.C. (2016). Ssd: Single Shot Multibox Detector, European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Gong, H., Mu, T., Li, Q., Dai, H., Li, C., He, Z., Wang, W., Han, F., Tuniyazi, A., and Li, H. (2022). Swin-Transformer-Enabled YOLOv5 with Attention Mechanism for Small Object Detection on Satellite Images. Remote Sens., 14.","DOI":"10.3390\/rs14122861"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Cheng, G., Lang, C., Wu, M., Xie, X., Yao, X., and Han, J. (2021). Feature enhancement network for object detection in optical remote sensing images. J. Remote Sens.","DOI":"10.34133\/2021\/9805389"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2486","DOI":"10.1109\/TGRS.2016.2645610","article-title":"Accurate object localization in remote sensing images based on convolutional neural networks","volume":"55","author":"Long","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Lu, X., Cao, G., Yang, Y., Jiao, L., and Liu, F. (2021, January 11\u201317). ViT-YOLO: Transformer-Basd YOLO for Object Detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00314"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"165","DOI":"10.9734\/jerr\/2022\/v23i12774","article-title":"Review of Typical Vehicle Detection Algorithms Based on Deep Learning","volume":"23","author":"Dong","year":"2022","journal-title":"J. Eng. Res. Rep."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). In Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"52","DOI":"10.1016\/j.jmsy.2019.03.002","article-title":"Machine vision intelligence for product defect inspection based on deep learning and Hough transform","volume":"51","author":"Wang","year":"2019","journal-title":"J. Manuf. Syst."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Girdhar, R., Gkioxari, G., Torresani, L., Paluri, M., and Tran, D. (2018, January 18\u201322). Detect-and-track: Efficient pose estimation in videos. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00044"},{"key":"ref_13","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_14","unstructured":"Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J. (2020). Deformable detr: Deformable transformers for end-to-end object detection. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 11\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1109\/TPAMI.2022.3152247","article-title":"A survey on vision transformer","volume":"45","author":"Han","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Wu, H., Xiao, B., Codella, N., Liu, M., Dai, X., Yuan, L., and Zhang, L. (2021, January 11\u201317). Cvt: Introducing convolutions to vision transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00009"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"124963","DOI":"10.1109\/ACCESS.2021.3109798","article-title":"An Improved Light-Weight Traffic Sign Recognition Algorithm Based on YOLOv4-Tiny","volume":"9","author":"Wang","year":"2021","journal-title":"IEEE Access"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Saleem, M.H., Potgieter, J., and Arif, K.M. (2022). Weed detection by faster RCNN model: An enhanced anchor box approach. Agronomy, 12.","DOI":"10.3390\/agronomy12071580"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Li, X., Wang, C., Ju, H., and Li, Z. (2022). Surface defect detection model for aero-engine components based on improved YOLOv5. Appl. Sci., 12.","DOI":"10.3390\/app12147235"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zheng, Z., Wang, P., Liu, W., Li, J., Ye, R., and Ren, D. (2020, January 7\u201312). Distance-IoU loss: Faster and better learning for bounding box regression. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6999"},{"key":"ref_22","unstructured":"Ba, J.L., Kiros, J.R., and Hinton, G.E. (2016). Layer normalization. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1007\/s11554-023-01268-w","article-title":"Real-time detection algorithm of helmet and reflective vest based on improved YOLOv5","volume":"20","author":"Chen","year":"2023","journal-title":"J. Real-Time Image Process."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Hou, Q., Zhou, D., and Feng, J. (2021, January 20\u201325). Coordinate attention for efficient mobile network design. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01350"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Ren, Z., Yu, Z., Yang, X., Liu, M.-Y., Lee, Y.J., Schwing, A.G., and Kautz, J. (2020, January 13\u201319). Instance-aware, context-focused, and memory-efficient weakly supervised object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01061"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"2301","DOI":"10.1109\/TIP.2020.3038483","article-title":"MSB-FCN: Multi-scale bidirectional fcn for object skeleton extraction","volume":"30","author":"Yang","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Tan, M., Pang, R., and Le, Q.V. (2020, January 13\u201319). Efficientdet: Scalable and efficient object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"ref_28","first-page":"264","article-title":"Improved Surface Defect Detection of YOLOV5 Aluminum Profiles based on CBAM and BiFPN","volume":"8","author":"Hua","year":"2022","journal-title":"Int. Core J. Eng."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Xia, G.S., Bai, X., Ding, J., Zhu, Z., Belongie, S., Luo, J., Datcu, M., Pelillo, M., and Zhang, L. (2018, January 18\u201323). DOTA: A Large-scale Dataset for Object Detection in Aerial Images. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00418"},{"key":"ref_30","unstructured":"Li, C., Li, L., Jiang, H., Weng, K., Geng, Y., Li, L., Ke, Z., Li, Q., Cheng, M., and Nie, W. (2022). YOLOv6: A single-stage object detection framework for industrial applications. arXiv."},{"key":"ref_31","unstructured":"Wang, C.-Y., Bochkovskiy, A., and Liao, H.-Y.M. (2022). YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv."},{"key":"ref_32","unstructured":"Bochkovskiy, A., Wang, C.-Y., and Liao, H.-Y.M. (2020). Yolov4: Optimal speed and accuracy of object detection. arXiv."},{"key":"ref_33","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_34","first-page":"427","article-title":"Remote sensing image target detection based on multi-scale feature fusion network","volume":"59","author":"Tian","year":"2022","journal-title":"Laser Optoelectron. Prog."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"3377","DOI":"10.1109\/TGRS.2019.2954328","article-title":"FMSSD: Feature-merged single-shot detection for multiscale objects in large-scale remote sensing imagery","volume":"58","author":"Wang","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Ding, J., Xue, N., Long, Y., Xia, G.-S., and Lu, Q. (2019, January 15\u201320). Learning roi transformer for oriented object detection in aerial images. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00296"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Azimi, S.M., Vig, E., Bahmanyar, R., K\u00f6rner, M., and Reinartz, P. (2018). Towards Multi-Class Object Detection in Unconstrained Remote Sensing Imagery, Asian Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-20893-6_10"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Acatay, O., Sommer, L., Schumann, A., and Beyerer, J. (2018, January 27\u201330). Comprehensive evaluation of deep learning based detection methods for vehicle detection in aerial imagery. Proceedings of the 2018 15th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), Auckland, New Zealand.","DOI":"10.1109\/AVSS.2018.8639127"},{"key":"ref_39","first-page":"1","article-title":"RetinaNet with difference channel attention and adaptively spatial feature fusion for steel surface defect detection","volume":"70","author":"Cheng","year":"2020","journal-title":"IEEE Trans. Instrum. Meas."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/7\/3634\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:07:57Z","timestamp":1760123277000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/7\/3634"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,31]]},"references-count":39,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2023,4]]}},"alternative-id":["s23073634"],"URL":"https:\/\/doi.org\/10.3390\/s23073634","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,31]]}}}