{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T16:36:14Z","timestamp":1781714174415,"version":"3.54.5"},"reference-count":33,"publisher":"MDPI AG","issue":"15","license":[{"start":{"date-parts":[[2022,7,30]],"date-time":"2022-07-30T00:00:00Z","timestamp":1659139200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62162059"],"award-info":[{"award-number":["62162059"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"publisher","award":["12061072"],"award-info":[{"award-number":["12061072"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2021xjkk1404"],"award-info":[{"award-number":["2021xjkk1404"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Third Xinjiang Scientific Expedition Program","award":["62162059"],"award-info":[{"award-number":["62162059"]}]},{"name":"Third Xinjiang Scientific Expedition Program","award":["12061072"],"award-info":[{"award-number":["12061072"]}]},{"name":"Third Xinjiang Scientific Expedition Program","award":["2021xjkk1404"],"award-info":[{"award-number":["2021xjkk1404"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In view of the existence of remote sensing images with large variations in spatial resolution, small and dense objects, and the inability to determine the direction of motion, all these components make object detection from remote sensing images very challenging. In this paper, we propose a single-stage detection network based on YOLOv5. This method introduces the MS Transformer module at the end of the feature extraction network of the original network to enhance the feature extraction capability of the network model and integrates the Convolutional Block Attention Model (CBAM) to find the attention area in dense scenes. In addition, the YOLOv5 target detection network is improved by incorporating a rotation angle approach from the a priori frame design and the bounding box regression formulation to make it suitable for rotating frame-based detection scenarios. Finally, the weighted combination of the two difficult sample mining methods is used to improve the focal loss function, so as to improve the detection accuracy. The average accuracy of the test results of the improved algorithm on the DOTA data set is 77.01%, which is higher than the previous detection algorithm. Compared with the average detection accuracy of YOLOv5, the average detection accuracy is improved by 8.83%. The experimental results show that the algorithm has higher detection accuracy than other algorithms in remote sensing scenes.<\/jats:p>","DOI":"10.3390\/s22155716","type":"journal-article","created":{"date-parts":[[2022,8,1]],"date-time":"2022-08-01T23:49:27Z","timestamp":1659397767000},"page":"5716","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":42,"title":["R-YOLO: A YOLO-Based Method for Arbitrary-Oriented Target Detection in High-Resolution Remote Sensing Images"],"prefix":"10.3390","volume":"22","author":[{"given":"Yongjie","family":"Hou","sequence":"first","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gang","family":"Shi","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yingxiang","family":"Zhao","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4392-7821","authenticated-orcid":false,"given":"Fan","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xian","family":"Jiang","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rujun","family":"Zhuang","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yunfei","family":"Mei","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinjiang","family":"Ma","sequence":"additional","affiliation":[{"name":"Geomatics School of Earth Sciences and Engineering, Hohai University, Nanjing 211100, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,7,30]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, faster, stronger. Proceedings of the IEEE conference on computer vision and pattern recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_3","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_4","unstructured":"Bochkovskiy, A., Wang, C.Y., and Liao, H.Y.M. (2020). Yolov4: Optimal speed and accuracy of object detection. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016). SSD: Single Shot MultiBox Detector. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Neubeck, A., and Van Gool, L. (2006, January 20\u201324). Efficient non-maximum suppression. Proceedings of the IEEE Conference on Pattern Recognition (ICPR), Hong Kong, China.","DOI":"10.1109\/ICPR.2006.479"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast-RCNN. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster r-cnn: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"310","DOI":"10.1109\/LGRS.2018.2872355","article-title":"Multiscale visual attention networks for object detection in VHR remote sensing images","volume":"16","author":"Wang","year":"2018","journal-title":"IEEE Geosci Remote Sens. Lett."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"5535","DOI":"10.1109\/TGRS.2019.2900302","article-title":"Hierarchical and Robust Convolutional Neural Network for Very High-Resolution Remote Sensing Object Detection","volume":"57","author":"Zhang","year":"2019","journal-title":"IEEE Trans Geosci Remote Sens."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"2337","DOI":"10.1109\/TGRS.2017.2778300","article-title":"Rotation-Insensitive and Context-Augmented Object Detection in Remote Sensing Images","volume":"56","author":"Li","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_13","first-page":"1","article-title":"R\u00b2-CNN: Fast Tiny Object Detection in Large-Scale Remote Sensing Images","volume":"56","author":"Pang","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_14","first-page":"9","article-title":"High-resolution remote sensing image object detection algorithm combining RPN network and SSD algorithm","volume":"46","author":"Cheng","year":"2021","journal-title":"Sci. Surv. Mapp."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Liu, Y., Yang, J., and Cui, W. (2020, January 16\u201326). Simple, Fast, Accurate Object Detection based on Anchor-Free Method for High Resolution Remote Sensing Images. Proceedings of the IGARSS 2020-2020 IEEE International Geoscience and Remote Sensing Symposium, Waikoloa Village, HI, USA.","DOI":"10.1109\/IGARSS39084.2020.9324301"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Qu, Z., Zhu, F., and Qi, C. (2021). Remote Sensing Image Target Detection: Improvement of the YOLOv3 Model with Auxiliary Networks. Remote Sens., 13.","DOI":"10.3390\/rs13193908"},{"key":"ref_17","unstructured":"Van Etten, A. (2018). You only look twice: Rapid multi-scale object detection in satellite imagery. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., and Guo, B. (2021, January 10\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yu, H., Zhang, Z., Qin, Z., Wu, H., Li, D., Zhao, J., and Lu, X. (2018, January 8\u201313). Loss rank mining: A general hard example mining method for real-time detectors. Proceedings of the IEEE International Joint Conference on Neural Networks (IJCNN), Rio de Janeiro, Brazil.","DOI":"10.1109\/IJCNN.2018.8489071"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201323). Path aggregation network for instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Yang, X., and Yan, J. (2020). Arbitrary-oriented object detection with circular smooth label. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-58598-3_40"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Xia, G.S., Bai, X., Ding, J., Zhu, Z., Belongie, S., Luo, J., and Zhang, L. (2018, January 18\u201323). DOTA: A large-scale dataset for object detection in aerial images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake, UT, USA.","DOI":"10.1109\/CVPR.2018.00418"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Ding, J., Xue, N., Long, Y., Xia, G.S., and Lu, Q. (2018). Learning roi transformer for detecting oriented objects in aerial images. arXiv.","DOI":"10.1109\/CVPR.2019.00296"},{"key":"ref_27","unstructured":"Lin, Y., Feng, P., Guan, J., Wang, W., and Chambers, J. (2019). IENet: Interacting embranchment one stage anchor free detector for orientation aerial object detection. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Jiang, Y., Zhu, X., Wang, X., Yang, S., Li, W., Wang, H., and Luo, Z. (2017). R2cnn: Rotational region cnn for orientation robust scene text detection. arXiv, 1\u20138.","DOI":"10.1109\/ICPR.2018.8545598"},{"key":"ref_29","unstructured":"Yang, X., Liu, Q., Yan, J., Li, A., Zhang, Z., and Yu, G. (2019). R3det: Refined single-stage detector with feature refinement for rotating object. arXiv."},{"key":"ref_30","first-page":"2866","article-title":"Telemetry Based on Rotation Center Point Estimation Accurate detection algorithm of sensory target","volume":"38","author":"Jiang","year":"2021","journal-title":"Comput. Appl. Res."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Yang, X., Yang, J., Yan, J., Zhang, Y., Zhang, T., Guo, Z., and Fu, K. (2019, January 20\u201326). Scrdet: Towards more robust detection for small, cluttered and rotated objects. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00832"},{"key":"ref_32","unstructured":"Qian, W., Yang, X., Peng, S., Guo, Y., and Yan, J. (2019). Learning modulated loss for rotated object detection. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wang, J., Ding, J., Guo, H., Cheng, W., Pan, T., and Yang, W. (2019). Mask OBB: A semantic attention-based mask oriented bounding box representation for multi-category object detection in aerial images. Remote Sens., 11.","DOI":"10.3390\/rs11242930"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/15\/5716\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:59:53Z","timestamp":1760140793000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/15\/5716"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,30]]},"references-count":33,"journal-issue":{"issue":"15","published-online":{"date-parts":[[2022,8]]}},"alternative-id":["s22155716"],"URL":"https:\/\/doi.org\/10.3390\/s22155716","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,7,30]]}}}