{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,24]],"date-time":"2026-03-24T03:11:26Z","timestamp":1774321886970,"version":"3.50.1"},"reference-count":42,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2024,6,24]],"date-time":"2024-06-24T00:00:00Z","timestamp":1719187200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Object detection in unmanned aerial vehicle (UAV) scenes is a challenging task due to the varying scales and complexities of targets. To address this, we propose a novel object detection model, AFE-YOLOv8, which integrates three innovative modules: the Multi-scale Nonlinear Fusion Module (MNFM), the Adaptive Feature Enhancement Module (AFEM), and the Receptive Field Expansion Module (RFEM). The MNFM introduces nonlinear mapping by exploiting the property that deformable convolution can dynamically adjust the shape of the convolution kernel according to the shape of the target, and it effectively enhances the feature extraction capability of the backbone network by integrating multi-scale feature maps from different mapping branches. Meanwhile, the AFEM introduces an adaptive fusion factor, and through the fusion factor, it adaptively integrates the small-target features contained in the feature maps of different detection branches into the small-target detection branch, thus enhancing the expression of the small-target features contained in the feature maps of the small-target detection branch. Furthermore, the RFEM expands the receptive field of the feature maps of the large- and medium-scale target detection branches through stacked convolution, so as to make the model\u2019s receptive field cover the whole target, and thereby learn more rich and comprehensive features of the target. The experimental results demonstrate the superior performance of the proposed model compared to the baseline in detecting objects of various scales. On the VisDrone dataset, the proposed model achieves a 4.5% enhancement in mean average precision (mAP) and a 5.45% improvement in average precision at an IOU threshold of 0.5 (AP50). Additionally, ablation experiments conducted on the challenging DOTA dataset showcase the model\u2019s robustness and generalization capabilities.<\/jats:p>","DOI":"10.3390\/a17070276","type":"journal-article","created":{"date-parts":[[2024,6,24]],"date-time":"2024-06-24T05:16:18Z","timestamp":1719206178000},"page":"276","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["AFE-YOLOv8: A Novel Object Detection Model for Unmanned Aerial Vehicle Scenes with Adaptive Feature Enhancement"],"prefix":"10.3390","volume":"17","author":[{"given":"Shijie","family":"Wang","sequence":"first","affiliation":[{"name":"College of Electronic Information, Qingdao University, Qingdao 266071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zekun","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Electronic Information, Qingdao University, Qingdao 266071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qingqing","family":"Chao","sequence":"additional","affiliation":[{"name":"College of Electronic Information, Qingdao University, Qingdao 266071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Teng","family":"Yu","sequence":"additional","affiliation":[{"name":"College of Electronic Information, Qingdao University, Qingdao 266071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,6,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Kim, B., Lee, J., Kang, J., Kim, E.S., and Kim, H.J. (2021, January 20\u201325). Hotr: End-to-end human-object interaction detection with transformers. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00014"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Vielzeuf, V., Lechervy, A., Pateux, S., and Jurie, F. (2018, January 8\u201314). Centralnet: A multilayer approach for multimodal fusion. Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Munich, Germany.","DOI":"10.1007\/978-3-030-11024-6_44"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Zhu, C., He, Y., and Savvides, M. (2019, January 15\u201320). Feature selective anchor-free module for single-shot object detection. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00093"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wang, X., Zhang, S., Yu, Z., Feng, L., and Zhang, W. (2020, January 13\u201319). Scale-equalizing pyramid convolution for object detection. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01337"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"106094","DOI":"10.1016\/j.engappai.2023.106094","article-title":"Texture-aware gray-scale image colorization using a bistream generative adversarial network with multi scale attention structure","volume":"122","author":"Zang","year":"2023","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Xie, S., Girshick, R., Doll\u00e1r, P., Tu, Z., and He, K. (2017, January 21\u201326). Aggregated residual transformations for deep neural networks. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.634"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"261","DOI":"10.1007\/s11263-019-01247-4","article-title":"Deep learning for generic object detection: A survey","volume":"128","author":"Liu","year":"2020","journal-title":"Int. J. Comput. Vis."},{"key":"ref_9","unstructured":"Dai, J., Li, Y., He, K., and Sun, J. (2016). R-fcn: Object Detection via Region-Based Fully Convolutional Networks, Curran Associates Inc."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast r-cnn. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster r-cnn: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Cai, Z., and Vasconcelos, N. (2018, January 18\u201323). Cascade r-cnn: Delving into high quality object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00644"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). Yolo9000: Better, faster, stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_15","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_16","unstructured":"Bochkovskiy, A., Wang, C.Y., and Liao, H.Y.M. (2020). Yolov4: Optimal speed and accuracy of object detection. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Wang, C.Y., Bochkovskiy, A., and Liao, H. (2022). Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv.","DOI":"10.1109\/CVPR52729.2023.00721"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"318","DOI":"10.1109\/TPAMI.2018.2858826","article-title":"Focal loss for dense object detection","volume":"42","author":"Lin","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Dollar, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Yang, C., Huang, Z., and Wang, N. (2022, January 18\u201324). Querydet: Cascaded sparse query for accelerating high-resolution small object detection. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01330"},{"key":"ref_21","unstructured":"Zhuang, J., Qin, Z., Yu, H., and Chen, X. (2023). Task-specific context decoupling for object detection. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhu, X., Hu, H., Lin, S., and Dai, J. (2019, January 15\u201320). Deformable convnets v2: More deformable, better results. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00953"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wang, W., Dai, J., Chen, Z., Huang, Z., Li, Z., Zhu, X., Hu, X., Lu, T., Lu, L., and Li, H. (2022). Internimage: Exploring large-scale vision foundation models with deformable convolutions. arXiv.","DOI":"10.1109\/CVPR52729.2023.01385"},{"key":"ref_24","unstructured":"Berg, A.C., Fu, C.Y., Szegedy, C., Anguelov, D., Erhan, D., Reed, S., and Liu, W. (2016). Ssd: Single shot multibox detector. Computer Vision\u2014ECCV 2016, Springer."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Varghese, R., and Sambath, M. (2024, January 18\u201319). Yolov8: A novel object detection algorithm with enhanced performance and robustness. Proceedings of the 2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS), Chennai, India.","DOI":"10.1109\/ADICS58448.2024.10533619"},{"key":"ref_26","unstructured":"Wang, C.Y., Yeh, I.H., and Liao, H.Y.M. (2024). Yolov9: Learning what you want to learn using programmable gradient information. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhao, Y., Lv, W., Xu, S., Wei, J., Wang, G., Dang, Q., Liu, Y., and Chen, J. (2024, January 17\u201321). Detrs beat yolos on real-time object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01605"},{"key":"ref_28","unstructured":"Wang, A., Chen, H., Liu, L., Chen, K., Lin, Z., Han, J., and Ding, G. (2024). Yolov10: Real-time end-to-end object detection. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Akyon, F.C., Onur Altinuc, S., and Temizel, A. (2022, January 16\u201319). Slicing aided hyper inference and fine-tuning for small object detection. Proceedings of the 2022 IEEE International Conference on Image Processing (ICIP), Bordeaux, France.","DOI":"10.1109\/ICIP46576.2022.9897990"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhang, S., Zhu, X., Lei, Z., Shi, H., Wang, X., and Li, S.Z. (2017, January 1\u20134). Faceboxes: A cpu real-time face detector with high accuracy. Proceedings of the 2017 IEEE International Joint Conference on Biometrics (IJCB), Denver, CO, USA.","DOI":"10.1109\/BTAS.2017.8272675"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"4909","DOI":"10.1109\/TITS.2021.3054625","article-title":"Deep reinforcement learning for autonomous driving: A survey","volume":"23","author":"Kiran","year":"2022","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zhang, S., Wen, L., Bian, X., Lei, Z., and Li, S.Z. (2018, January 18\u201323). Single-shot refinement neural network for object detection. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00442"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Li, Y., Chen, Y., Wang, N., and Zhang, Z. (November, January 27). Scale-aware trident networks for object detection. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00615"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zhu, X., Lyu, S., Wang, X., and Zhao, Q. (2021, January 11\u201317). Tph-yolov5: Improved yolov5 based on transformer prediction head for object detection on drone-captured scenarios. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision Workshops (ICCVW), Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00312"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wang, P., Xue, D., Zhu, Y., Sun, J., Yan, Q., Yoon, S.E., and Zhang, Y. (2023). Take a prior from other tasks for severe blur removal. arXiv.","DOI":"10.2139\/ssrn.4435376"},{"key":"ref_36","unstructured":"Du, D., Zhu, P., Wen, L., Bian, X., and Liu, Z.M. (2019, January 27\u201328). Visdrone-det2019: The vision meets drone object detection in image challenge results. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision Workshop (ICCVW), Seoul, Republic of Korea."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Xia, G.S., Bai, X., Ding, J., Zhu, Z., Belongie, S., Luo, J., Datcu, M., Pelillo, M., and Zhang, L. (2018, January 18\u201323). Dota: A large-scale dataset for object detection in aerial images. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00418"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Zitnick, C.L., and Doll\u00e1r, P. (2014). Microsoft coco: Common objects in context. Computer Vision\u2014ECCV 2014, Springer.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Chen, C., Zhang, Y., Lv, Q., Wei, S., Wang, X., Sun, X., and Dong, J. (2019, January 27\u201328). Rrnet: A hybrid detector for object detection in drone-captured images. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision Workshop (ICCVW), Seoul, Republic of Korea.","DOI":"10.1109\/ICCVW.2019.00018"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Cao, Y., He, Z., Wang, L., Wang, W., Yuan, Y., Zhang, D., Zhang, J., Zhu, P., Van Gool, L., and Han, J. (2021, January 11\u201317). Visdrone-det2021: The vision meets drone object detection challenge results. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision Workshop (ICCVW), Virtual.","DOI":"10.1109\/ICCVW54120.2021.00319"},{"key":"ref_41","unstructured":"Xu, S., Wang, X., Lv, W., Chang, Q., Cui, C., Deng, K., Wang, G., Dang, Q., Wei, S., and Du, Y. (2022). Pp-yoloe: An evolved version of yolo. arXiv."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Meethal, A., Granger, E., and Pedersoli, M. (2023, January 17\u201324). Cascaded zoom-in detector for high resolution aerial images. Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vancouver, BC, Canada.","DOI":"10.1109\/CVPRW59228.2023.00198"}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/17\/7\/276\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:03:23Z","timestamp":1760108603000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/17\/7\/276"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,24]]},"references-count":42,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2024,7]]}},"alternative-id":["a17070276"],"URL":"https:\/\/doi.org\/10.3390\/a17070276","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,24]]}}}