{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,4]],"date-time":"2026-08-04T22:26:39Z","timestamp":1785882399351,"version":"3.56.0"},"reference-count":48,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2024,2,29]],"date-time":"2024-02-29T00:00:00Z","timestamp":1709164800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"national defense foundation","award":["E21D41B1"],"award-info":[{"award-number":["E21D41B1"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Due to the issues of remote sensing object detection algorithms based on deep learning, such as a high number of network parameters, large model size, and high computational requirements, it is challenging to deploy them on small mobile devices. This paper proposes an extremely lightweight remote sensing aircraft object detection network based on the improved YOLOv5n. This network combines Shufflenet v2 and YOLOv5n, significantly reducing the network size while ensuring high detection accuracy. It substitutes the original CIoU and convolution with EIoU and deformable convolution, optimizing for the small-scale characteristics of aircraft objects and further accelerating convergence and improving regression accuracy. Additionally, a coordinate attention (CA) mechanism is introduced at the end of the backbone to focus on orientation perception and positional information. We conducted a series of experiments, comparing our method with networks like GhostNet, PP-LCNet, MobileNetV3, and MobileNetV3s, and performed detailed ablation studies. The experimental results on the Mar20 public dataset indicate that, compared to the original YOLOv5n network, our lightweight network has only about one-fifth of its parameter count, with only a slight decrease of 2.7% in mAP@0.5. At the same time, compared with other lightweight networks of the same magnitude, our network achieves an effective balance between detection accuracy and resource consumption such as memory and computing power, providing a novel solution for the implementation and hardware deployment of lightweight remote sensing object detection networks.<\/jats:p>","DOI":"10.3390\/rs16050857","type":"journal-article","created":{"date-parts":[[2024,2,29]],"date-time":"2024-02-29T08:13:44Z","timestamp":1709194424000},"page":"857","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":20,"title":["A Lightweight Remote Sensing Aircraft Object Detection Network Based on Improved YOLOv5n"],"prefix":"10.3390","volume":"16","author":[{"given":"Jiale","family":"Wang","sequence":"first","affiliation":[{"name":"Xi\u2019an Institute of Optics and Precision Mechanics of CAS, Xi\u2019an 710119, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhe","family":"Bai","sequence":"additional","affiliation":[{"name":"Xi\u2019an Institute of Optics and Precision Mechanics of CAS, Xi\u2019an 710119, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ximing","family":"Zhang","sequence":"additional","affiliation":[{"name":"Xi\u2019an Institute of Optics and Precision Mechanics of CAS, Xi\u2019an 710119, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuehong","family":"Qiu","sequence":"additional","affiliation":[{"name":"Xi\u2019an Institute of Optics and Precision Mechanics of CAS, Xi\u2019an 710119, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,2,29]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_3","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015). Faster r-cnn: Towards real-time object detection with region proposal networks. arXiv."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, faster, stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_6","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_7","unstructured":"Bochkovskiy, A., Wang, C.-Y., and Liao, H.-Y.M. (2020). Yolov4: Optimal speed and accuracy of object detection. arXiv."},{"key":"ref_8","unstructured":"Jocher, G. (2023, August 01). YOLOv5 by Ultralytics. Available online: https:\/\/github.com\/ultralytics\/yolov5."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020, January 23\u201328). End-to-end object detection with transformers. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_10","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. (2018, January 18\u201323). Mobilenetv2: Inverted residuals and linear bottlenecks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_12","unstructured":"Howard, A., Sandler, M., Chu, G., Chen, L.-C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., and Vasudevan, V. (November, January 27). Searching for mobilenetv3. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Han, K., Wang, Y., Tian, Q., Guo, J., Xu, C., and Xu, C. (2020, January 13\u201319). Ghostnet: More features from cheap operations. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00165"},{"key":"ref_14","unstructured":"Cui, C., Gao, T., Wei, S., Du, Y., Guo, R., Dong, S., Lu, B., Zhou, Y., Lv, X., and Liu, Q. (2021). PP-LCNet: A lightweight CPU convolutional neural network. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhou, X., Lin, M., and Sun, J. (2018, January 18\u201323). Shufflenet: An extremely efficient convolutional neural network for mobile devices. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00716"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Ma, N., Zhang, X., Zheng, H.-T., and Sun, J. (2018, January 8\u201314). Shufflenet v2: Practical guidelines for efficient cnn architecture design. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Jiang, Y., Tang, Y., and Ying, C.J.E. (2023). Finding a Needle in a Haystack: Faint and Small Space Object Detection in 16-Bit Astronomical Images Using a Deep Learning-Based Approach. Electronics, 12.","DOI":"10.3390\/electronics12234820"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1145\/3065386","article-title":"Imagenet classification with deep convolutional neural networks","volume":"60","author":"Krizhevsky","year":"2012","journal-title":"ACM"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., and Berg, A.C. (2016, January 11\u201314). Ssd: Single shot multibox detector. Proceedings of the Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands. Part I 14.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1016\/j.isprsjprs.2018.04.003","article-title":"Multi-scale object detection in remote sensing imagery with convolutional neural networks","volume":"145","author":"Deng","year":"2018","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Cheng, S., Cheng, H., Yang, R., Zhou, J., Li, Z., Shi, B., Lee, M., and Ma, Q.J.P. (2023). A High Performance Wheat Disease Detection Based on Position Information. Plants, 12.","DOI":"10.3390\/plants12051191"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"5184","DOI":"10.1109\/ACCESS.2022.3140876","article-title":"Aircraft target detection in remote sensing images based on improved YOLOv5","volume":"10","author":"Luo","year":"2022","journal-title":"IEEE Access"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1742","DOI":"10.1109\/ACCESS.2023.3233964","article-title":"YOLO-extract: Improved YOLOv5 for aircraft object detection in remote sensing images","volume":"11","author":"Liu","year":"2023","journal-title":"IEEE Access"},{"key":"ref_24","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention is all you need. arXiv."},{"key":"ref_25","unstructured":"Liu, L., Pan, Z., and Lei, B. (2017). Learning a rotation invariant detector with rotatable bounding box. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Yang, X., Yan, J., Feng, Z., and He, T. (2021, January 2\u20139). R3det: Refined single-stage detector with feature refinement for rotating object. Proceedings of the AAAI Conference on Artificial Intelligence, Virtually.","DOI":"10.1609\/aaai.v35i4.16426"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Ding, J., Xue, N., Long, Y., Xia, G.-S., and Lu, Q. (2019, January 15\u201320). Learning RoI transformer for oriented object detection in aerial images. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00296"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"223373","DOI":"10.1109\/ACCESS.2020.3041025","article-title":"Arbitrary-oriented object detection in remote sensing images based on polar coordinates","volume":"8","author":"Zhou","year":"2020","journal-title":"IEEE Access"},{"key":"ref_29","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_30","unstructured":"Iandola, F.N., Han, S., Moskewicz, M.W., Ashraf, K., Dally, W.J., and Keutzer, K. (2016). SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and< 0.5 MB model size. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Gholami, A., Kwon, K., Wu, B., Tai, Z., Yue, X., Jin, P., Zhao, S., and Keutzer, K. (2018, January 18\u201323). Squeezenext: Hardware-aware neural network design. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00215"},{"key":"ref_32","unstructured":"Sifre, L., and Mallat, S. (2014). Rigid-motion scattering for texture classification. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_34","unstructured":"Park, J., Woo, S., Lee, J.-Y., and Kweon, I.S. (2018). Bam: Bottleneck attention module. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.-Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Wang, Q., Wu, B., Zhu, P., Li, P., Zuo, W., and Hu, Q. (2020, January 13\u201319). ECA-Net: Efficient channel attention for deep convolutional neural networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01155"},{"key":"ref_37","unstructured":"Liu, Y., Shao, Z., Teng, Y., and Hoffmann, N. (2021). NAM: Normalization-based attention module. arXiv."},{"key":"ref_38","unstructured":"Yang, L., Zhang, R.-Y., Li, L., and Xie, X. (2021, January 18\u201324). Simam: A simple, parameter-free attention module for convolutional neural networks. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Chen, Q., and Wang, W. (2019). Sequential attention-based network for noetic end-to-end response selection. arXiv.","DOI":"10.1016\/j.csl.2020.101072"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Feng, G., Hu, Z., Zhang, L., and Lu, H. (2021, January 20\u201325). Encoder fusion network with co-attention embedding for referring image segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01525"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"71800","DOI":"10.1109\/ACCESS.2023.3293864","article-title":"An improved YOLOv5 algorithm for wood defect detection based on attention","volume":"11","author":"Han","year":"2023","journal-title":"IEEE Access"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Qiu, S., Li, Y., Zhao, H., Li, X., and Yuan, X. (2022). Foxtail Millet Ear Detection Method Based on Attention Mechanism and Improved YOLOv5. Sensors, 22.","DOI":"10.3390\/s22218206"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"131361","DOI":"10.1109\/ACCESS.2022.3226886","article-title":"Real-Time Detection Algorithm of Marine Organisms Based on Improved YOLOv4-Tiny","volume":"10","author":"Shi","year":"2022","journal-title":"IEEE Access"},{"key":"ref_44","first-page":"2050793","article-title":"X-Ray Small Target Security Inspection Based on TB-YOLOv5","volume":"2022","author":"Wang","year":"2022","journal-title":"Secur. Commun. Netw."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., and Savarese, S. (2019, January 15\u201320). Generalized intersection over union: A metric and a loss for bounding box regression. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00075"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"2688","DOI":"10.11834\/jrs.20222139","article-title":"MAR20: A Benchmark for Military Aircraft Recognition in Remote Sensing Images","volume":"27","author":"Yu","year":"2022","journal-title":"Natl. Remote Sens. Bull."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"296","DOI":"10.1016\/j.isprsjprs.2019.11.023","article-title":"Object detection in optical remote sensing images: A survey and a new benchmark","volume":"159","author":"Li","year":"2020","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_48","unstructured":"Akson, N. (2024, January 31). teknofest2 Dataset. Roboflow Universe. Available online: https:\/\/universe.roboflow.com\/neslihan-akson-rvo7y\/teknofest2-ftdcf."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/16\/5\/857\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:07:11Z","timestamp":1760105231000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/16\/5\/857"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,29]]},"references-count":48,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2024,3]]}},"alternative-id":["rs16050857"],"URL":"https:\/\/doi.org\/10.3390\/rs16050857","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,2,29]]}}}