{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T05:04:15Z","timestamp":1787029455676,"version":"3.56.0"},"reference-count":39,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2023,5,24]],"date-time":"2023-05-24T00:00:00Z","timestamp":1684886400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Xi\u2019an Key Laboratory of Advanced Control and Intelligent Process","award":["2019220714SYS022CG04"],"award-info":[{"award-number":["2019220714SYS022CG04"]}]},{"name":"Xi\u2019an Key Laboratory of Advanced Control and Intelligent Process","award":["2021ZDLGY04-04"],"award-info":[{"award-number":["2021ZDLGY04-04"]}]},{"name":"Key R&amp;D plan of Shaanxi Province","award":["2019220714SYS022CG04"],"award-info":[{"award-number":["2019220714SYS022CG04"]}]},{"name":"Key R&amp;D plan of Shaanxi Province","award":["2021ZDLGY04-04"],"award-info":[{"award-number":["2021ZDLGY04-04"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>In remote sensing images, small objects have too few discriminative features, are easily confused with background information, and are difficult to locate, leading to a degradation in detection accuracy when using general object detection networks for aerial images. To solve the above problems, we propose a remote sensing small object detection network based on the attention mechanism and multi-scale feature fusion, and name it AMMFN. Firstly, a detection head enhancement module (DHEM) was designed to strengthen the characterization of small object features through a combination of multi-scale feature fusion and attention mechanisms. Secondly, an attention mechanism based channel cascade (AMCC) module was designed to reduce the redundant information in the feature layer and protect small objects from information loss during feature fusion. Then, the Normalized Wasserstein Distance (NWD) was introduced and combined with Generalized Intersection over Union (GIoU) as the location regression loss function to improve the optimization weight of the model for small objects and the accuracy of the regression boxes. Finally, an object detection layer was added to improve the object feature extraction ability at different scales. Experimental results from the Unmanned Aerial Vehicles (UAV) dataset VisDrone2021 and the homemade dataset show that the AMMFN improves the APs values by 2.4% and 3.2%, respectively, compared with YOLOv5s, which represents an effective improvement in the detection accuracy of small objects.<\/jats:p>","DOI":"10.3390\/rs15112728","type":"journal-article","created":{"date-parts":[[2023,5,25]],"date-time":"2023-05-25T02:00:55Z","timestamp":1684980055000},"page":"2728","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":55,"title":["Remote Sensing Small Object Detection Network Based on Attention Mechanism and Multi-Scale Feature Fusion"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4781-260X","authenticated-orcid":false,"given":"Junsuo","family":"Qu","sequence":"first","affiliation":[{"name":"Xi\u2019an Key Laboratory of Advanced Control and Intelligent Process, School of Automation, Xi\u2019an Robertic Intelligent Systems International Science and Technology Cooperation Base, Xi\u2019an University of Posts and Telecommunications, Xi\u2019an 710121, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8048-5671","authenticated-orcid":false,"given":"Zongbing","family":"Tang","sequence":"additional","affiliation":[{"name":"Xi\u2019an Key Laboratory of Advanced Control and Intelligent Process, School of Automation, Xi\u2019an Robertic Intelligent Systems International Science and Technology Cooperation Base, Xi\u2019an University of Posts and Telecommunications, Xi\u2019an 710121, China"},{"name":"School of Communication and Information Engineering, Xi\u2019an University of Posts & Telecommunications, Xi\u2019an 710121, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Le","family":"Zhang","sequence":"additional","affiliation":[{"name":"Xi\u2019an Key Laboratory of Advanced Control and Intelligent Process, School of Automation, Xi\u2019an Robertic Intelligent Systems International Science and Technology Cooperation Base, Xi\u2019an University of Posts and Telecommunications, Xi\u2019an 710121, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanghai","family":"Zhang","sequence":"additional","affiliation":[{"name":"Xi\u2019an Key Laboratory of Advanced Control and Intelligent Process, School of Automation, Xi\u2019an Robertic Intelligent Systems International Science and Technology Cooperation Base, Xi\u2019an University of Posts and Telecommunications, Xi\u2019an 710121, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhenguo","family":"Zhang","sequence":"additional","affiliation":[{"name":"Xi\u2019an Key Laboratory of Advanced Control and Intelligent Process, School of Automation, Xi\u2019an Robertic Intelligent Systems International Science and Technology Cooperation Base, Xi\u2019an University of Posts and Telecommunications, Xi\u2019an 710121, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,5,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1016\/j.rse.2018.06.028","article-title":"Detecting Mammals in UAV Images: Best Practices to address a substantially Imbalanced Dataset with Deep Learning","volume":"216","author":"Kellenberger","year":"2018","journal-title":"Remote Sens. Environ."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Kellenberger, B., Volpi, M., and Tuia, D. (2017, January 23\u201328). Fast animal detection in UAV images using convolutional neural networks. Proceedings of the 2017 IEEE International Geoscience and Remote Sensing Symposium, Fort Worth, TX, USA.","DOI":"10.1109\/IGARSS.2017.8127090"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature Pyramid Networks for Object Detection. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) IEEE Computer Society, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201322). Path aggregation network for instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Tan, M., Pang, R., and Le, Q.V. (2020, January 14\u201319). Efficientdet: Scalable and efficient object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"ref_6","unstructured":"Liu, S., Huang, D., and Wang, Y. (2018). Learning spatial fusion for single-shot object detection. arXiv."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Ghiasi, G., Lin, T.Y., and Le, Q.V. (2019, January 15\u201320). NAS-FPN: Learning scalable feature pyramid architecture for object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00720"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Yu, J., Jiang, Y., Wang, Z., Cao, Z., and Huang, T. (2016, January 15\u201319). Unitbox: An advanced object detection network. Proceedings of the 24th ACM International Conference on Multimedia, Amsterdam, The Netherlands.","DOI":"10.1145\/2964284.2967274"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., and Savarese, S. (2019, January 15\u201320). Generalized intersection over union: A metric and a loss for bounding box regression. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00075"},{"key":"ref_11","unstructured":"Zheng, Z., Wang, P., Liu, W., Li, J., Ye, R., and Ren, D. (2019, January 29\u201331). Distance-iou loss: Faster and better learning for bounding box regression. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1016\/j.neucom.2022.07.042","article-title":"Focal and efficient iou loss for accurate bounding box regression","volume":"506","author":"Zhang","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_13","unstructured":"Jocher, G., Stoken, A., Borovec, J., Chaurasia, A., Changyu, L., Hogan, A., Hajek, J., Diaconu, L., Kwon, Y., and Defretin, Y. (Zenodo, 2021). Ultralytics\/Yolov5: v5.0\u2013YOLOv5-P6 1280 Models, AWS, Supervise.ly and YouTube integrations, Zenodo."},{"key":"ref_14","unstructured":"Wang, J., Xu, C., Yang, W., and Yu, L. (2021). A normalized gaussian wasserstein distance for tiny object detection. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"104","DOI":"10.1016\/j.inffus.2022.06.009","article-title":"Radar Sensor Network Resource Allocation for Fused Target Tracking: A Brief Review","volume":"86\u201387","author":"Yan","year":"2022","journal-title":"Inf. Fusion"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_17","first-page":"1137","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_18","first-page":"379","article-title":"R-FCN: Object detection via region-based fully convolutional networks","volume":"29","author":"Dai","year":"2016","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016, January 11\u201314). Ssd: Single shot multibox detector. Proceedings of the Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). Yolo9000: Better, faster, stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_22","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv preprint."},{"key":"ref_23","unstructured":"Bochkovskiy, A., Wang, C.Y., and Liao, H.Y.M. (2020). Yolov4: Optimal speed and accuracy of object detection. arXiv."},{"key":"ref_24","unstructured":"Tian, Z., Shen, C., Chen, H., and He, T. (November, January 27). Fcos: Fully convolutional one-stage object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Dollar, P. (2017, January 22\u201329). Focal Loss for Dense Object Detection. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zhu, Z., Liang, D., Zhang, S., Huang, X., Li, B., and Hu, S. (2016, January 27\u201330). Traffic-Sign Detection and Classification in the Wild. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.232"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"82832","DOI":"10.1109\/ACCESS.2020.2991439","article-title":"Dilated convolution and feature fusion SSD network for small object detection in remote sensing images","volume":"8","author":"Qu","year":"2020","journal-title":"IEEE Access"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1968","DOI":"10.1109\/TMM.2021.3074273","article-title":"Extended Feature Pyramid Network for Small Object Detection","volume":"24","author":"Deng","year":"2021","journal-title":"IEEE Trans. Multimed."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Deng, T., Liu, X., and Mao, G. (2022). Improved YOLOv5 Based on Hybrid Domain Attention for Small Object Detection in Optical Remote Sensing Images. Electronics, 11.","DOI":"10.3390\/electronics11172657"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhu, X., Lyu, S., Wang, X., and Zhao, Q. (2021, January 10\u201317). TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00312"},{"key":"ref_31","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 17\u201324). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_33","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS\u201917), Long Beach, CA, USA."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Shi, T., Gong, J., Hu, J., Zhi, X., Zhang, W., Zhang, Y., Zhang, P., and Bao, G. (2022). Feature-Enhanced CenterNet for Small Object Detection in Remote Sensing Images. Remote Sens., 14.","DOI":"10.3390\/rs14215488"},{"key":"ref_35","first-page":"927","article-title":"Deep-level Small Target Detection Algorithm Based on Attention Mechanism","volume":"16","author":"Zhao","year":"2022","journal-title":"J. Comput. Sci. Explor."},{"key":"ref_36","unstructured":"Zhang, F., Jiao, L., Li, L., Liu, F., and Liu, X. (2020). MultiResolution Attention Extractor for Small Object Detection. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Dai, Y., Gieseke, F., Oehmcke, S., Wu, Y., and Barnard, K. (2020, January 1\u20135). Attentional Feature Fusion. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Village, CO, USA.","DOI":"10.1109\/WACV48630.2021.00360"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Cao, Y., He, Z., Wang, L., Wang, W., Yuan, Y., Zhang, D., Zhang, J., Zhu, P., Van Gool, L., and Han, J. (2021, January 11\u201317). VisDrone-DET2021: The vision meets drone object detection challenge results. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision Workshops (ICCVW), Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00319"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Zitnick, C.L., and Doll\u00e1r, P. (2015). Microsoft COCO: Common Objects in Context. arXiv.","DOI":"10.1007\/978-3-319-10602-1_48"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/11\/2728\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:41:07Z","timestamp":1760125267000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/11\/2728"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,24]]},"references-count":39,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2023,6]]}},"alternative-id":["rs15112728"],"URL":"https:\/\/doi.org\/10.3390\/rs15112728","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,24]]}}}