{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T01:37:00Z","timestamp":1783647420288,"version":"3.55.0"},"reference-count":53,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2023,3,13]],"date-time":"2023-03-13T00:00:00Z","timestamp":1678665600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Central Government\u2019s Local Science and Technology Development Special Foundation Projects of China","award":["No. S202107d08050071"],"award-info":[{"award-number":["No. S202107d08050071"]}]},{"name":"Central Government\u2019s Local Science and Technology Development Special Foundation Projects of China","award":["No. 202107d08050031"],"award-info":[{"award-number":["No. 202107d08050031"]}]},{"name":"Central Government\u2019s Local Science and Technology Development Special Foundation Projects of China","award":["No. (2020)4001"],"award-info":[{"award-number":["No. (2020)4001"]}]},{"name":"Central Government\u2019s Local Science and Technology Development Special Foundation Projects of China","award":["(2020)1Y155"],"award-info":[{"award-number":["(2020)1Y155"]}]},{"name":"Science and Technology Foundation of Guizhou Province","award":["No. S202107d08050071"],"award-info":[{"award-number":["No. S202107d08050071"]}]},{"name":"Science and Technology Foundation of Guizhou Province","award":["No. 202107d08050031"],"award-info":[{"award-number":["No. 202107d08050031"]}]},{"name":"Science and Technology Foundation of Guizhou Province","award":["No. (2020)4001"],"award-info":[{"award-number":["No. (2020)4001"]}]},{"name":"Science and Technology Foundation of Guizhou Province","award":["(2020)1Y155"],"award-info":[{"award-number":["(2020)1Y155"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Object detection in unmanned aerial vehicle (UAV) images is an extremely challenging task and involves problems such as multi-scale objects, a high proportion of small objects, and high overlap between objects. To address these issues, first, we design a Vectorized Intersection Over Union (VIOU) loss based on YOLOv5s. This loss uses the width and height of the bounding box as a vector to construct a cosine function that corresponds to the size of the box and the aspect ratio and directly compares the center point value of the box to improve the accuracy of the bounding box regression. Second, we propose a Progressive Feature Fusion Network (PFFN) that addresses the issue of insufficient semantic extraction of shallow features by Panet. This allows each node of the network to fuse semantic information from deep layers with features from the current layer, thus significantly improving the detection ability of small objects in multi-scale scenes. Finally, we propose an Asymmetric Decoupled (AD) head, which separates the classification network from the regression network and improves the classification and regression capabilities of the network. Our proposed method results in significant improvements on two benchmark datasets compared to YOLOv5s. On the VisDrone 2019 dataset, the performance increased by 9.7% from 34.9% to 44.6%, and on the DOTA dataset, the performance increased by 2.1%.<\/jats:p>","DOI":"10.3390\/s23063061","type":"journal-article","created":{"date-parts":[[2023,3,13]],"date-time":"2023-03-13T03:28:33Z","timestamp":1678678113000},"page":"3061","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Object Detection for UAV Aerial Scenarios Based on Vectorized IOU"],"prefix":"10.3390","volume":"23","author":[{"given":"Shun","family":"Lu","sequence":"first","affiliation":[{"name":"College of Big Data and Information Engineering, Guizhou University, Guiyang 550025, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hanyu","family":"Lu","sequence":"additional","affiliation":[{"name":"College of Big Data and Information Engineering, Guizhou University, Guiyang 550025, China"},{"name":"Bijie 5G Innovation and Application Research Institute, Guizhou University of Engineering Science, Bijie 551700, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9039-4685","authenticated-orcid":false,"given":"Jun","family":"Dong","sequence":"additional","affiliation":[{"name":"Hefei Institutes of Physical Science, Chinese Academy of Sciences, Hefei 230031, China"},{"name":"Anhui Zhongke Deji Intelligence Technology Co., Ltd., Hefei 230045, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuang","family":"Wu","sequence":"additional","affiliation":[{"name":"Hefei Institutes of Physical Science, Chinese Academy of Sciences, Hefei 230031, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2013). Rich feature hierarchies for accurate object detection and semantic segmentation. arXiv.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_3","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015). Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. arXiv."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_5","unstructured":"Redmon, J., and Farhadi, A. (2018). YOLOv3: An Incremental Improvement. arXiv."},{"key":"ref_6","unstructured":"Bochkovskiy, A., Wang, C.Y., and Liao, H.Y.M. (2020). YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv."},{"key":"ref_7","unstructured":"Jocher, G. (2021, October 12). Stoken Yolo v5. Available online: https:\/\/github.com\/ultralytics\/yolov5\/releases\/tag\/v6.0."},{"key":"ref_8","unstructured":"Li, C., Li, L., Jiang, H., Weng, K., Geng, Y., Li, L., Ke, Z., Li, Q., Cheng, M., and Nie, W. (2022). YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, Faster, Stronger. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_10","unstructured":"Wang, C.Y., Bochkovskiy, A., and Liao, H.Y.M. (2022). YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv."},{"key":"ref_11","unstructured":"Ultralytics, G.J. (2023, January 09). Yolo v8. Available online: https:\/\/github.com\/ultralytics\/ultralytics.git."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016, January 11\u201314). Ssd: Single shot multibox detector. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Lin, T., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal Loss for Dense Object Detection. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_14","unstructured":"Duan, K., Bai, S., Xie, L., Qi, H., Huang, Q., and Tian, Q. (November, January 27). Centernet: Keypoint triplets for object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Ding, J., Xue, N., Xia, G.S., Bai, X., Yang, W., Yang, M.Y., Belongie, S., Luo, J., Datcu, M., and Pelillo, M. (2021). Object Detection in Aerial Images: A Large-Scale Benchmark and Challenges. arXiv.","DOI":"10.1109\/TPAMI.2021.3117983"},{"key":"ref_16","unstructured":"Shadab Malik, H., Sobirov, I., and Mohamed, A. (2022). Object Detection in Aerial Images: What Improves the Accuracy?. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"116675","DOI":"10.1016\/j.image.2022.116675","article-title":"Focus-and-Detect: A small object detection framework for aerial images","volume":"104","author":"Koyun","year":"2022","journal-title":"Signal Process. Image Commun."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Li, C., Yang, T., Zhu, S., Chen, C., and Guan, S. (2020, January 14\u201319). Density map guided object detection in aerial images. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00103"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Duan, C., Wei, Z., Zhang, C., Qu, S., and Wang, H. (2021, January 11\u201317). Coarse-grained Density Map Guided Object Detection in Aerial Images. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00313"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zhu, X., Lyu, S., Wang, X., and Zhao, Q. (2021). TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios. arXiv.","DOI":"10.1109\/ICCVW54120.2021.00312"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Luo, X., Wu, Y., and Zhao, L. (2022). YOLOD: A Target Detection Method for UAV Aerial Imagery. Remote Sens., 14.","DOI":"10.3390\/rs14143240"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Liu, H., Mu, C., Yang, R., He, Y., and Wu, N. (2021, January 17\u201319). Research on Object Detection Algorithm Based on UVA Aerial Image. Proceedings of the 2021 7th IEEE International Conference on Network Intelligence and Digital Content (IC-NIDC), Beijing, China.","DOI":"10.1109\/IC-NIDC54101.2021.9660571"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Li, Z., Sun, S., Li, Y., Sun, B., Tian, K., Qiao, L., and Lu, X. (2021, January 13\u201316). Aerial Image Object Detection Method Based on Adaptive ClusDet Network. Proceedings of the 2021 IEEE 21st International Conference on Communication Technology (ICCT), Tianjin, China.","DOI":"10.1109\/ICCT52962.2021.9657834"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Cao, C., Wu, J., Zeng, X., Feng, Z., Wang, T., Yan, X., Wu, Z., Wu, Q., and Huang, Z. (2020). Research on Airplane and Ship Detection of Aerial Remote Sensing Images Based on Convolutional Neural Network. Sensors, 20.","DOI":"10.3390\/s20174696"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zheng, Z., Wang, P., Liu, W., Li, J., Ye, R., and Ren, D. (2020, January 7\u201312). Distance-IoU loss: Faster and better learning for bounding box regression. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6999"},{"key":"ref_26","unstructured":"Ge, Z., Liu, S., Wang, F., Li, Z., and Sun, J. (2021). YOLOX: Exceeding YOLO Series in 2021. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"7380","DOI":"10.1109\/TPAMI.2021.3119563","article-title":"Detection and Tracking Meet Drones Challenge","volume":"44","author":"Zhu","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Xia, G.S., Bai, X., Ding, J., Zhu, Z., Belongie, S., Luo, J., Datcu, M., Pelillo, M., and Zhang, L. (2018, January 18\u201322). DOTA: A Large-Scale Dataset for Object Detection in Aerial Images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00418"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Ding, J., Xue, N., Long, Y., Xia, G.S., and Lu, Q. (2019, January 16\u201320). Learning RoI Transformer for Detecting Oriented Objects in Aerial Images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00296"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhou, D., Fang, J., Song, X., Guan, C., Yin, J., Dai, Y., and Yang, R. (2019, January 16\u201319). IoU Loss for 2D\/3D Object Detection. Proceedings of the 2019 International Conference on 3D Vision (3DV), Quebec City, QC, Canada.","DOI":"10.1109\/3DV.2019.00019"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., and Savarese, S. (2019, January 15\u201320). Generalized intersection over union: A metric and a loss for bounding box regression. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00075"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Xu, C., Wang, J., Yang, W., and Yu, L. (2021, January 19\u201325). Dot Distance for Tiny Object Detection in Aerial Images. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Virtual.","DOI":"10.1109\/CVPRW53098.2021.00130"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"012001","DOI":"10.1088\/1742-6596\/1924\/1\/012001","article-title":"EIoU: An Improved Vehicle Detection Algorithm Based on VehicleNet Neural Network","volume":"1924","author":"Yang","year":"2021","journal-title":"J. Phys. Conf. Ser."},{"key":"ref_34","unstructured":"Gevorgyan, Z. (2022). SIoU Loss: More Powerful Learning for Bounding Box Regression. arXiv."},{"key":"ref_35","unstructured":"He, J., Erfani, S., Ma, X., Bailey, J., Chi, Y., and Hua, X.S. (2021). Alpha-IoU: A Family of Power Intersection over Union Losses for Bounding Box Regression. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Dollar, P., and Girshick, R. (2017, January 22\u201329). Mask R-CNN. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Cai, Z., and Vasconcelos, N. (2018, January 18\u201322). Cascade R-CNN: Delving Into High Quality Object Detection. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00644"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lin, T., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature Pyramid Networks for Object Detection. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_39","unstructured":"Wang, K., Liew, J.H., Zou, Y., Zhou, D., and Feng, J. (November, January 27). PANet: Few-Shot Image Semantic Segmentation With Prototype Alignment. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Huang, W., Li, G., Chen, Q., Ju, M., and Qu, J. (2021). CF2PN: A Cross-Scale Feature Fusion Pyramid Network Based Remote Sensing Target Detection. Remote. Sens., 13.","DOI":"10.3390\/rs13050847"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Zhou, L., Rao, X., Li, Y., Zuo, X., Qiao, B., and Lin, Y. (2022). A Lightweight Object Detection Method in Aerial Images Based on Dense Feature Fusion Path Aggregation Network. Isprs Int. J. Geo-Inf., 11.","DOI":"10.3390\/ijgi11030189"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Tayara, H., and Chong, K.T. (2018). Object Detection in Very High-Resolution Aerial Images Using One-Stage Densely Connected Feature Pyramid Network. Sensors, 18.","DOI":"10.3390\/s18103341"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Tian, H., Zheng, Y., and Jin, Z. (2020, January 18\u201320). Improved RetinaNet model for the application of small target detection in the aerial images. Proceedings of the IOP Conference Series: Earth and Environmental Science, Changsha, China.","DOI":"10.1088\/1755-1315\/585\/1\/012142"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"1147","DOI":"10.1016\/0043-1354(89)90158-9","article-title":"Kinetic analysis of aerated submerged fixed-film (ASFF) bioreactors","volume":"23","author":"Hamoda","year":"1989","journal-title":"Water Res."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Ghiasi, G., Lin, T.Y., and Le, Q.V. (2019, January 15\u201320). NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object Detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00720"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Tan, M., Pang, R., and Le, Q.V. (2020, January 13\u201319). Efficientdet: Scalable and efficient object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Qiao, S., Chen, L.C., and Yuille, A. (2021, January 20\u201325). Detectors: Detecting objects with recursive feature pyramid and switchable atrous convolution. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01008"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Jiang, B., Luo, R., Mao, J., Xiao, T., and Jiang, Y. (2018, January 8\u201314). Acquisition of Localization Confidence for Accurate Object Detection. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_48"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Song, G., Liu, Y., and Wang, X. (2020, January 13\u201319). Revisiting the Sibling Head in Object Detector. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01158"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Wang, C.Y., Bochkovskiy, A., and Liao, H.Y.M. (2021, January 20\u201325). Scaled-YOLOv4: Scaling Cross Stage Partial Network. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01283"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Wu, Y., Chen, Y., Yuan, L., Liu, Z., Wang, L., Li, H., and Fu, Y. (2020, January 13\u201319). Rethinking Classification and Localization for Object Detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01020"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Wang, J., Song, L., Li, Z., Sun, H., Sun, J., and Zheng, N. (2021, January 20\u201325). End-to-End Object Detection With Fully Convolutional Network. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01559"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Chen, Q., Wang, Y., Yang, T., Zhang, X., Cheng, J., and Sun, J. (2021, January 20\u201325). You Only Look One-Level Feature. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01284"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/6\/3061\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:53:28Z","timestamp":1760122408000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/6\/3061"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,13]]},"references-count":53,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["s23063061"],"URL":"https:\/\/doi.org\/10.3390\/s23063061","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,13]]}}}