{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T11:05:16Z","timestamp":1785409516967,"version":"3.56.0"},"reference-count":54,"publisher":"MDPI AG","issue":"16","license":[{"start":{"date-parts":[[2023,8,15]],"date-time":"2023-08-15T00:00:00Z","timestamp":1692057600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100008869","name":"Graduate Innovative Fund of Wuhan Institute of Technology","doi-asserted-by":"publisher","award":["CX2022148"],"award-info":[{"award-number":["CX2022148"]}],"id":[{"id":"10.13039\/501100008869","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Unmanned aerial vehicle (UAV) object detection plays a crucial role in civil, commercial, and military domains. However, the high proportion of small objects in UAV images and the limited platform resources lead to the low accuracy of most of the existing detection models embedded in UAVs, and it is difficult to strike a good balance between detection performance and resource consumption. To alleviate the above problems, we optimize YOLOv8 and propose an object detection model based on UAV aerial photography scenarios, called UAV-YOLOv8. Firstly, Wise-IoU (WIoU) v3 is used as a bounding box regression loss, and a wise gradient allocation strategy makes the model focus more on common-quality samples, thus improving the localization ability of the model. Secondly, an attention mechanism called BiFormer is introduced to optimize the backbone network, which improves the model\u2019s attention to critical information. Finally, we design a feature processing module named Focal FasterNet block (FFNB) and propose two new detection scales based on this module, which makes the shallow features and deep features fully integrated. The proposed multiscale feature fusion network substantially increased the detection performance of the model and reduces the missed detection rate of small objects. The experimental results show that our model has fewer parameters compared to the baseline model and has a mean detection accuracy higher than the baseline model by 7.7%. Compared with other mainstream models, the overall performance of our model is much better. The proposed method effectively improves the ability to detect small objects. There is room to optimize the detection effectiveness of our model for small and feature-less objects (such as bicycle-type vehicles), as we will address in subsequent research.<\/jats:p>","DOI":"10.3390\/s23167190","type":"journal-article","created":{"date-parts":[[2023,8,15]],"date-time":"2023-08-15T11:15:08Z","timestamp":1692098108000},"page":"7190","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":878,"title":["UAV-YOLOv8: A Small-Object-Detection Model Based on Improved YOLOv8 for UAV Aerial Photography Scenarios"],"prefix":"10.3390","volume":"23","author":[{"given":"Gang","family":"Wang","sequence":"first","affiliation":[{"name":"Hubei Key Laboratory of Optical Information and Pattern Recognition, School of Electrical and Information Engineering, Wuhan Institute of Technology, Wuhan 430205, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanfei","family":"Chen","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Optical Information and Pattern Recognition, School of Electrical and Information Engineering, Wuhan Institute of Technology, Wuhan 430205, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pei","family":"An","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Optical Information and Pattern Recognition, School of Electrical and Information Engineering, Wuhan Institute of Technology, Wuhan 430205, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hanyu","family":"Hong","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Optical Information and Pattern Recognition, School of Electrical and Information Engineering, Wuhan Institute of Technology, Wuhan 430205, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinghu","family":"Hu","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Optical Information and Pattern Recognition, School of Electrical and Information Engineering, Wuhan Institute of Technology, Wuhan 430205, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tiange","family":"Huang","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Optical Information and Pattern Recognition, School of Electrical and Information Engineering, Wuhan Institute of Technology, Wuhan 430205, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,8,15]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Li, Z., Zhang, Y., Wu, H., Suzuki, S., Namiki, A., and Wang, W. (2023). Design and Application of a UAV Autonomous Inspection System for High-Voltage Power Transmission Lines. Remote Sens., 15.","DOI":"10.3390\/rs15030865"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Byun, S., Shin, I.-K., Moon, J., Kang, J., and Choi, S.-I. (2021). Road Traffic Monitoring from UAV Images Using Deep Learning Networks. Remote Sens., 13.","DOI":"10.3390\/rs13204027"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1297","DOI":"10.1007\/s10586-022-03627-x","article-title":"A survey on deep learning-based identification of plant and crop diseases from UAV-based aerial images","volume":"26","author":"Bouguettaya","year":"2022","journal-title":"Cluster. Comput."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","article-title":"Object Detection with Discriminatively Trained Part-Based Models","volume":"32","author":"Felzenszwalb","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, Faster, Stronger. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_10","unstructured":"Redmon, J., and Farhadi, A. (2018). YOLOv3: An Incremental Improvement. arXiv."},{"key":"ref_11","unstructured":"Bochkovskiy, A., Wang, C.Y., and Liao, H.Y.M. (2020). YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv."},{"key":"ref_12","unstructured":"Li, C., Li, L., Jiang, H., Weng, K., Geng, Y., Li, L., Ke, Z., Li, Q., Cheng, M., and Nie, W. (2022). YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, C.Y., Bochkovskiy, A., and Liao, H.Y.M. (2022). YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv.","DOI":"10.1109\/UV56588.2022.10185474"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016). SSD: Single Shot MultiBox Detector. arXiv.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_15","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Advances in Neural Information Processing Systems, MIT Press."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020). End-to-end object detection with transformers. arXiv.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Luo, X., Wu, Y., and Wang, F. (2022). Target Detection Method of UAV Aerial Imagery Based on Improved YOLOv5. Remote Sens., 14.","DOI":"10.3390\/rs14195063"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zhou, H., Ma, A., Niu, Y., and Ma, Z. (2022). Small-Object Detection for UAV-Based Images Using a Distance Metric Method. Drones, 6.","DOI":"10.3390\/drones6100308"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Du, B., Huang, Y., Chen, J., and Huang, D. (2023, January 18\u201322). Adaptive Sparse Convolutional Networks with Global Context Enhancement for Faster Object Detection on Drone Images. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01291"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"108054","DOI":"10.1016\/j.ijepes.2022.108054","article-title":"Research on edge intelligent recognition method oriented to transmission line insulator fault detection","volume":"139","author":"Deng","year":"2022","journal-title":"Int. J. Electr. Power Energy Syst."},{"key":"ref_21","unstructured":"Howard, A., Pang, R., Adam, H., Le, Q., Sandler, M., Chen, B., Wang, W., Chen, L.C., Tan, M., and Chu, G. (November, January 27). Searching for MobileNetV3. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"95","DOI":"10.1016\/j.isprsjprs.2021.01.008","article-title":"Growing status observation for oil palm trees using Unmanned Aerial Vehicle (UAV) images","volume":"173","author":"Zheng","year":"2021","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Liu, M., Wang, X., Zhou, A., Fu, X., Ma, Y., and Piao, C. (2020). UAV-YOLO: Small Object Detection on Unmanned Aerial Vehicle Perspective. Sensors, 20.","DOI":"10.3390\/s20082238"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"He, K.M., Zhang, X.Y., Ren, S.Q., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Liu, B., and Luo, H. (2022). An Improved Yolov5 for Multi-Rotor UAV Detection. Electronics, 11.","DOI":"10.3390\/electronics11152330"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wang, J., Zhang, F., Zhang, Y., Liu, Y., and Cheng, T. (2023). Lightweight Object Detection Algorithm for UAV Aerial Imagery. Sensors, 23.","DOI":"10.3390\/s23135786"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.-Y., and Kweon, I.S. (2018, January 8\u201314). CBAM: Convolutional Block Attention Module. Proceedings of the Computer Vision\u2014ECCV 2018, Cham, Switzerland.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"145740","DOI":"10.1109\/ACCESS.2020.3014910","article-title":"Small-object detection in UAV-captured images via multi-branch parallel feature pyramid networks","volume":"8","author":"Liu","year":"2020","journal-title":"IEEE Access"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Chen, J., Kao, S.-H., He, H., Zhuo, W., Wen, S., Lee, C.-H., and Chan, S.-H.G. (2023, January 18\u201322). Run, Don\u2019t walk: Chasing higher FLOPS for faster neural networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01157"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhu, L., Wang, X., Ke, Z., Zhang, W., and Lau, R. (2023). BiFormer: Vision Transformer with Bi-Level Routing Attention. arXiv.","DOI":"10.1109\/CVPR52729.2023.00995"},{"key":"ref_31","unstructured":"Tong, Z., Chen, Y., Xu, Z., and Yu, R. (2023). Wise-IoU: Bounding Box Regression Loss with Dynamic Focusing Mechanism. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","article-title":"Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition","volume":"37","author":"He","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201323). Path Aggregation Network for Instance Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Li, X., Wang, W., Wu, L., Chen, S., Hu, X., Li, J., Tang, J., and Yang, J. (2020). Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection. arXiv.","DOI":"10.1109\/CVPR46437.2021.01146"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Zheng, Z., Wang, P., Liu, W., Li, J., Ye, R., and Ren, D. (2020, January 7\u201312). Distance-IoU loss: Faster and better learning for bounding box regression. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6999"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Feng, C., Zhong, Y., Gao, Y., Scott, M.R., and Huang, W. (2021, January 10\u201317). TOOD: Task-Aligned One-Stage Object Detection. Proceedings of the 2021 IEEE International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00349"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1016\/j.neucom.2022.07.042","article-title":"Focal and efficient IOU loss for accurate bounding box regression","volume":"506","author":"Zhang","year":"2022","journal-title":"Neurocomputing"},{"key":"ref_38","unstructured":"Gevorgyan, Z. (2022). SIoU Loss: More Powerful Learning for Bounding Box Regression. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Cao, X., Zhang, Y., Lang, S., and Gong, Y. (2023). Swin-Transformer-Based YOLOv5 for Small-Object Detection in Remote Sensing Images. Sensors, 23.","DOI":"10.3390\/s23073634"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Lu, S., Lu, H., Dong, J., and Wu, S. (2023). Object Detection for UAV Aerial Scenarios Based on Vectorized IOU. Sensors, 23.","DOI":"10.3390\/s23063061"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Zhang, T., Zhang, Y., Xin, M., Liao, J., and Xie, Q. (2023). A Light-Weight Network for Small Insulator and Defect Detection Using UAV Imaging Based on Improved YOLOv5. Sensors, 23.","DOI":"10.20944\/preprints202305.0796.v1"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Jiang, X., Cui, Q., Wang, C., Wang, F., Zhao, Y., Hou, Y., Zhuang, R., Mei, Y., and Shi, G. (2023). A Model for Infrastructure Detection along Highways Based on Remote Sensing Images from UAVs. Sensors, 23.","DOI":"10.3390\/s23083847"},{"key":"ref_43","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhou, X., Lin, M., and Sun, J. (2017). ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices. arXiv.","DOI":"10.1109\/CVPR.2018.00716"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Han, K., Wang, Y.H., Tian, Q., Guo, J.Y., Xu, C.J., and Xu, C. (2020, January 14\u201319). GhostNet: More Features from Cheap Operations. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00165"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"7380","DOI":"10.1109\/TPAMI.2021.3119563","article-title":"Detection and Tracking Meet Drones Challenge","volume":"44","author":"Zhu","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., and Savarese, S. (2019, January 15\u201320). Generalized intersection over union: A metric and a loss for bounding box regression. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00075"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Wang, C.Y., Liao, H.Y.M., Wu, Y.H., Chen, P.Y., Hsieh, J.W., and Yeh, I.H. (2020, January 13\u201319). CSPNet: A new backbone that can enhance learning capability of CNN. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00203"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Cai, Z., and Vasconcelos, N. (2018, January 18\u201322). Cascade R-CNN: Delving into High Quality Object Detection. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00644"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Lin, T., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal Loss for Dense Object Detection. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_51","unstructured":"Duan, K., Bai, S., Xie, L., Qi, H., Huang, Q., and Tian, Q. (November, January 27). Centernet: Keypoint triplets for object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Zhu, C., He, Y., and Savvides, M. (2019, January 16\u201320). Feature Selective Anchor-Free Module for Single-Shot Object Detection. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00093"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Zhang, S., Chi, C., Yao, Y., Lei, Z., and Li, S.Z. (2020, January 14\u201319). Bridging the Gap Between Anchor-Based and Anchor-Free Detection via Adaptive Training Sample Selection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00978"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. (2017, January 22\u201329). Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.74"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/16\/7190\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:34:17Z","timestamp":1760128457000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/16\/7190"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,15]]},"references-count":54,"journal-issue":{"issue":"16","published-online":{"date-parts":[[2023,8]]}},"alternative-id":["s23167190"],"URL":"https:\/\/doi.org\/10.3390\/s23167190","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,15]]}}}