{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T10:13:38Z","timestamp":1784024018368,"version":"3.55.0"},"reference-count":61,"publisher":"MDPI AG","issue":"22","license":[{"start":{"date-parts":[[2020,11,14]],"date-time":"2020-11-14T00:00:00Z","timestamp":1605312000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Municipal Science and Technology Project of CQMMC, China","award":["2017030502"],"award-info":[{"award-number":["2017030502"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Object detection is one of the core technologies in aerial image processing and analysis. Although existing aerial image object detection methods based on deep learning have made some progress, there are still some problems remained: (1) Most existing methods fail to simultaneously consider multi-scale and multi-shape object characteristics in aerial images, which may lead to some missing or false detections; (2) high precision detection generally requires a large and complex network structure, which usually makes it difficult to achieve the high detection efficiency and deploy the network on resource-constrained devices for practical applications. To solve these problems, we propose a slimmer network for more efficient object detection in aerial images. Firstly, we design a polymorphic module (PM) for simultaneously learning the multi-scale and multi-shape object features, so as to better detect the hugely different objects in aerial images. Then, we design a group attention module (GAM) for better utilizing the diversiform concatenation features in the network. By designing multiple detection headers with adaptive anchors and the above-mentioned two modules, we propose a one-stage network called PG-YOLO for realizing the higher detection accuracy. Based on the proposed network, we further propose a more efficient channel pruning method, which can slim the network parameters from 63.7 million (M) to 3.3M that decreases the parameter size by 94.8%, so it can significantly improve the detection efficiency for real-time detection. Finally, we execute the comparative experiments on three public aerial datasets, and the experimental results show that the proposed method outperforms the state-of-the-art methods.<\/jats:p>","DOI":"10.3390\/rs12223750","type":"journal-article","created":{"date-parts":[[2020,11,16]],"date-time":"2020-11-16T21:48:52Z","timestamp":1605563332000},"page":"3750","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":21,"title":["A Slimmer Network with Polymorphic and Group Attention Modules for More Efficient Object Detection in Aerial Images"],"prefix":"10.3390","volume":"12","author":[{"given":"Wei","family":"Guo","sequence":"first","affiliation":[{"name":"Key Lab of Optoelectronic Technology and Systems Ministry of Education, College of Optoelectronic Engineering, Chongqing University, Chongqing 400044, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6462-1848","authenticated-orcid":false,"given":"Weihong","family":"Li","sequence":"additional","affiliation":[{"name":"Key Lab of Optoelectronic Technology and Systems Ministry of Education, College of Optoelectronic Engineering, Chongqing University, Chongqing 400044, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9898-8974","authenticated-orcid":false,"given":"Zhenghao","family":"Li","sequence":"additional","affiliation":[{"name":"Chongqing Institute of Green and Intelligent Technology, Chinese Academy of Sciences, Chongqing 400714, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weiguo","family":"Gong","sequence":"additional","affiliation":[{"name":"Key Lab of Optoelectronic Technology and Systems Ministry of Education, College of Optoelectronic Engineering, Chongqing University, Chongqing 400044, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinkai","family":"Cui","sequence":"additional","affiliation":[{"name":"Key Lab of Optoelectronic Technology and Systems Ministry of Education, College of Optoelectronic Engineering, Chongqing University, Chongqing 400044, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinran","family":"Wang","sequence":"additional","affiliation":[{"name":"Key Lab of Optoelectronic Technology and Systems Ministry of Education, College of Optoelectronic Engineering, Chongqing University, Chongqing 400044, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2020,11,14]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"296","DOI":"10.1016\/j.isprsjprs.2019.11.023","article-title":"Object detection in optical remote sensing images: A survey and a new benchmark","volume":"159","author":"Li","year":"2020","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_3","unstructured":"Redmon, J., and Farhadi, A.J. (2018). Yolov3: An incremental improvement. arXiv, e-prints."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2020.2980023","article-title":"Adaptive Saliency Biased Loss for Object Detection in Aerial Images","volume":"58","author":"Sun","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_5","unstructured":"Zhu, P., Wen, L., Bian, X., Ling, H., and Hu, Q.J. (2018). Vision meets drones: A challenge. arXiv."},{"key":"ref_6","first-page":"33","article-title":"Pyramid methods in image processing","volume":"29","author":"Adelson","year":"1984","journal-title":"RCA Eng."},{"key":"ref_7","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20138). ImageNet classification with deep convolutional neural networks. Proceedings of the International Conference on Neural Information Processing Systems (NIPS), Lake Tahoe, NV, USA."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017). Feature pyramid networks for object detection. arXiv.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Fu, C., and Berg, A.C. (2016, January 8\u201316). SSD: Single Shot Multibox Detector. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Li, Y., Chen, Y., Wang, N., and Zhang, Z. (2020, January 22\u201323). Scale-Aware Trident Networks for Object Detection. Proceedings of the International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2019.00615"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Qiu, H., Li, H., and Wu, Q. (2019). A2RMNet: Adaptively aspect ratio multi-scale network for object detection in remote sensing images. Remote Sens., 11.","DOI":"10.3390\/rs11131594"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Xu, Z., Xu, X., Wang, L., and Yang, R. (2017). Deformable ConvNet with Aspect Ratio Constrained NMS for Object Detection in Remote Sensing Imagery. Remote Sens., 9.","DOI":"10.3390\/rs9121312"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Yan, J., Wang, H., and Yan, M. (2019). IoU-adaptive deformable R-CNN: Make full use of IoU for multi-class object detection in remote sensing imagery. Remote Sens., 11.","DOI":"10.3390\/rs11030286"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Dai, J., Qi, H., and Xiong, Y. (2017, January 22\u201329). Deformable convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.89"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., and Shlens, J. (2016, January 27\u201330). Rethinking the inception architecture for computer vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Xie, S., and Girshick, R. (2017, January 22\u201325). Aggregated Residual Transformations for Deep Neural Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.634"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Liu, S., and Huang, D. (2018, January 8\u201314). Receptive field block net for accurate and fast object detection. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01252-6_24"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Liu, Z., Li, J., Shen, Z., and Huang, G. (2017, January 22\u201329). Learning efficient convolutional networks through network slimming. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.298"},{"key":"ref_20","unstructured":"Ye, J., Lu, X., and Lin, Z. (2018). Rethinking the smaller-norm-less-informative assumption in channel pruning of convolution layers. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhang, P., Zhong, Y., and Li, X. (2019, January 16\u201320). SlimYOLOv3: Narrower, Faster and Better for Real-Time UAV Applications. Proceedings of the IEEE International Conference on Computer Vision Workshops, Long Beach, CA, USA.","DOI":"10.1109\/ICCVW.2019.00011"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"He, Y., Zhang, X., and Sun, J. (2017, January 22\u201329). Channel pruning for accelerating very deep neural networks. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.155"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Ye, Y., You, G., Zhu, X., and Yang, Q. (2020). Channel Pruning via Optimal Thresholding. arXiv.","DOI":"10.1007\/978-3-030-63823-8_58"},{"key":"ref_24","first-page":"1135","article-title":"Learning both weights and connections for efficient neural network","volume":"1","author":"Han","year":"2015","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_25","unstructured":"Ioffe, S., and Szegedy, C.J. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 11\u201318). Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"142","DOI":"10.1109\/TPAMI.2015.2437384","article-title":"Region-Based Convolutional Networks for Accurate Object Detection and Segmentation","volume":"38","author":"Girshick","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_28","unstructured":"Dai, J.F., Li, Y., He, K.M., and Sun, J. (2016). R-FCN: Object Detection via Region-based Fully Convolutional Networks. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Cai, Z., Fan, Q., and Feris, R.S. (2016, January 8\u201316). A unified multi-scale deep convolutional neural network for fast object detection. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46493-0_22"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Song, G., Liu, Y., and Wang, X.J. (2020, January 13\u201319). Revisiting the Sibling Head in Object Detector. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2020, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01158"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 22\u201325). YOLO9000: Better, faster, stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"241","DOI":"10.1016\/j.ins.2020.02.067","article-title":"DC-SPP-YOLO: Dense connection and spatial pyramid pooling based YOLO for object detection","volume":"522","author":"Huang","year":"2020","journal-title":"Inf. Sci."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., and Girshick, R. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zhang, S., Wen, L., and Bian, X. (2018, January 18\u201322). Single-shot refinement neural network for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00442"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K.Q. (2017, January 21\u201326). Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Xia, G.S., Bai, X., and Ding, J. (2017). DOTA: A Large-scale Dataset for Object Detection in Aerial Images. arXiv.","DOI":"10.1109\/CVPR.2018.00418"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Guo, W., Li, W., and Gong, W. (2020). Extended Feature Pyramid Network with Adaptive Scale Training Strategy and Anchors for Object Detection in Aerial Images. Remote Sens., 12.","DOI":"10.3390\/rs12050784"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Ren, Y., Zhu, C., and Xiao, S. (2018). Deformable Faster R-CNN with Aggregating Multi-Layer Features for Partially Occluded Object Detection in Optical Remote Sensing Images. Remote Sens., 10.","DOI":"10.3390\/rs10091470"},{"key":"ref_39","first-page":"1","article-title":"Toward Fast and Accurate Vehicle Detection in Aerial Images Using Coupled Region-Based Convolutional Neural Networks","volume":"10","author":"Deng","year":"2017","journal-title":"Remote Sens."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Ren, Y., Zhu, C., and Xiao, S.J. (2018). Small Object Detection in Optical Remote Sensing Images via Modified Faster R-CNN. Appl. Sci., 8.","DOI":"10.3390\/app8050813"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Cheng, G., Si, Y., and Hong, H. (2020). Cross-Scale Feature Fusion for Object Detection in Optical Remote Sensing Images. IEEE Geoence Remote Sens. Lett., 1\u20135.","DOI":"10.34133\/2021\/9805389"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Yang, X., Yang, J., Yan, J., Zhang, Y., Zhang, T., and Guo, Z. (November, January 27). SCRDet: Towards More Robust Detection for Small, Cluttered and Rotated Objects. Proceedings of the IEEE International Conference on Computer Vision 2019, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00832"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Azimi, S.M., Vig, E., Bahmanyar, R., and Korner, M. (2018, January 18\u201322). Towards Multi-class Object Detection in Unconstrained Remote Sensing Imagery. Proceedings of the Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1007\/978-3-030-20893-6_10"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Tayara, H., and Chong, K.T. (2018). Object detection in very high-resolution aerial images using one-stage densely connected feature pyramid network. Sensors, 18.","DOI":"10.3390\/s18103341"},{"key":"ref_45","unstructured":"Van Etten, A. (2018). You Only Look Twice: Rapid Multi-Scale Object Detection in Satellite Imagery. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Ma, H., Liu, Y., Ren, Y., and Yu, J. (2019). Detection of Collapsed Buildings in Post-Earthquake Remote Sensing Images Based on the Improved YOLOv3. Remote Sens., 12.","DOI":"10.3390\/rs12010044"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Zhang, X., and Zhu, X. (2019). An Efficient and Scene-Adaptive Algorithm for Vehicle Detection in Aerial Images Using an Improved YOLOv3 Framework. Int. J. Geo-Inf., 8.","DOI":"10.3390\/ijgi8110483"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"2889","DOI":"10.1109\/TPAMI.2018.2873305","article-title":"Holistic CNN Compression via Low-rank Decomposition with Knowledge Transfer","volume":"41","author":"Lin","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Yang, Y., Wu, S., and Deng, L. (2019). Training High-Performance and Large-Scale Deep Neural Networks with Full 8-bit Integers. arXiv.","DOI":"10.1016\/j.neunet.2019.12.027"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"2051","DOI":"10.1109\/TIP.2018.2883743","article-title":"Low-Resolution Face Recognition in the Wild via Selective Knowledge Distillation","volume":"28","author":"Ge","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Ma, N., Zhang, X., Zheng, H., and Sun, J. (2018, January 8\u201314). ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Chollet, F. (2017, January 22\u201325). Xception: Deep Learning with Depthwise Separable Convolutions. Proceedings of the Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.195"},{"key":"ref_53","unstructured":"Xu, B., Wang, N., and Chen, T. (2015). Empirical Evaluation of Rectified Activations in Convolutional Network. Computer ence. arXiv."},{"key":"ref_54","unstructured":"Hu, J., Shen, L., and Albanie, S. (2017). Squeeze-and-Excitation Networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, IEEE Computer Society."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., and Lee, J. (2018, January 8\u201314). CBAM: Convolutional Block Attention Module. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"2337","DOI":"10.1109\/TGRS.2017.2778300","article-title":"Rotation-Insensitive and Context-Augmented Object Detection in Remote Sensing Images","volume":"56","author":"Li","year":"2017","journal-title":"IEEE Trans. Geoence Remote Sens."},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1109\/TIP.2018.2867198","article-title":"Learning Rotation-Invariant and Fisher Discriminative Convolutional Neural Networks for Object Detection","volume":"28","author":"Cheng","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"He, K., Georgia, G., and Piotr, D. (2017). Mask R-CNN. IEEE Transactions on Pattern Analysis and Machine Intelligence, IEEE Computer Society.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201322). Path Aggregation Network for Instance Segmentation. Proceedings of the Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Law, H., and Deng, J. (2018). CornerNet: Detecting Objects as Paired Keypoints. Int. J. Comput. Vis.","DOI":"10.1007\/978-3-030-01264-9_45"},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.isprsjprs.2020.04.019","article-title":"HyNet: Hyper-scale object detection network framework for multiple spatial resolution remote sensing imagery","volume":"166","author":"Zheng","year":"2020","journal-title":"ISPRS J. Photogramm. Remote Sens."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/22\/3750\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:33:26Z","timestamp":1760178806000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/22\/3750"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,11,14]]},"references-count":61,"journal-issue":{"issue":"22","published-online":{"date-parts":[[2020,11]]}},"alternative-id":["rs12223750"],"URL":"https:\/\/doi.org\/10.3390\/rs12223750","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,11,14]]}}}