{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T15:11:34Z","timestamp":1778080294620,"version":"3.51.4"},"reference-count":40,"publisher":"MDPI AG","issue":"22","license":[{"start":{"date-parts":[[2020,11,15]],"date-time":"2020-11-15T00:00:00Z","timestamp":1605398400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the Key Laboratory Foundation of Science and Technology on Reliability and Environmental Engineering Laboratory","award":["6142004190401"],"award-info":[{"award-number":["6142004190401"]}]},{"name":"the Aviation Science Foundation of China","award":["20170751006"],"award-info":[{"award-number":["20170751006"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>There are a large number of studies on geospatial object detection. However, many existing methods only focus on either accuracy or speed. Methods with both fast speed and high accuracy are of great importance in some scenes, like search and rescue, and military information acquisition. In remote sensing images, there are some targets that are small and have few textures and low contrast compared with the background, which impose challenges on object detection. In this paper, we propose an accurate and fast single shot detector (AF-SSD) for high spatial remote sensing imagery to solve these problems. Firstly, we design a lightweight backbone to reduce the number of trainable parameters of the network. In this lightweight backbone, we also use some wide and deep convolutional blocks to extract more semantic information and keep the high detection precision. Secondly, a novel encoding\u2013decoding module is employed to detect small targets accurately. With up-sampling and summation operations, the encoding\u2013decoding module can add strong high-level semantic information to low-level features. Thirdly, we design a cascade structure with spatial and channel attention modules for targets with low contrast (named low-contrast targets) and few textures (named few-texture targets). The spatial attention module can extract long-range features for few-texture targets. By weighting each channel of a feature map, the channel attention module can guide the network to concentrate on easily identifiable features for low-contrast and few-texture targets. The experimental results on the NWPU VHR-10 dataset show that our proposed AF-SSD achieves superior detection performance: parameters 5.7 M, mAP 88.7%, and 0.035 s per image on average on an NVIDIA GTX-1080Ti GPU.<\/jats:p>","DOI":"10.3390\/s20226530","type":"journal-article","created":{"date-parts":[[2020,11,16]],"date-time":"2020-11-16T21:48:52Z","timestamp":1605563332000},"page":"6530","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["AF-SSD: An Accurate and Fast Single Shot Detector for High Spatial Remote Sensing Imagery"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5989-8976","authenticated-orcid":false,"given":"Ruihong","family":"Yin","sequence":"first","affiliation":[{"name":"School of Electronic and Information Engineering, Beihang University, Beijing 100191, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6060-1022","authenticated-orcid":false,"given":"Wei","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Electronic and Information Engineering, Beihang University, Beijing 100191, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xudong","family":"Fan","sequence":"additional","affiliation":[{"name":"School of Electronic and Information Engineering, Beihang University, Beijing 100191, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongfeng","family":"Yin","sequence":"additional","affiliation":[{"name":"School of Reliability and Systems Engineering, Beihang University, Beijing 100191, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,11,15]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1074","DOI":"10.1109\/LGRS.2016.2565705","article-title":"Ship rotated bounding box space for ship extraction from high-resolution optical satellite images with complex backgrounds","volume":"13","author":"Liu","year":"2017","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1016\/j.isprsjprs.2014.10.002","article-title":"Multi-class geospatial object detection and geographic image classification based on collection of part detectors","volume":"98","author":"Cheng","year":"2014","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Xia, G.-S., Bai, X., Ding, J., Zhu, Z., Belongie, S., Luo, J., Datcu, M., Pelillo, M., and Zhang, L. (2018, January 18\u201323). Dota: A large-scale dataset for object detection in aerial images. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00418"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., and Berg, A.C. (2016, January 8\u201316). SSD: Single shot multibox detector. Proceedings of the 2016 European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the 2016 IEEE Conference on Computer vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Liu, S., and Huang, D. (2018, January 8\u201314). Receptive field block net for accurate and fast object detection. Proceedings of the 2018 European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01252-6_24"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Li, Z., Peng, C., Yu, G., Zhang, X., Deng, Y., and Sun, J. (2018). Detnet: A backbone network for object detection. arXiv.","DOI":"10.1007\/978-3-030-01240-3_21"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Cai, Z., Fan, Q., Feris, R.S., and Vasconcelos, N. (2016, January 8\u201316). A unified multi-scale deep convolutional neural network for fast object detection. Proceedings of the 2016 European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46493-0_22"},{"key":"ref_10","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015, January 7\u201312). Faster R-Cnn: Towards real-time object detection with region proposal networks. Proceedings of the 2015 Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"142","DOI":"10.1109\/TPAMI.2015.2437384","article-title":"Region-Based Convolutional Networks for Accurate Object Detection and Segmentation","volume":"38","author":"Girshick","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the 2015 IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_13","unstructured":"Dai, J., Li, Y., He, K., and Sun, J. (2016, January 5\u201310). R-Fcn: Object detection via region-based fully convolutional networks. Proceedings of the 2016 Advances in Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask R-Cnn. Proceedings of the 2017 IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_15","unstructured":"Fu, C.-Y., Liu, W., Ranga, A., Tyagi, A., and Berg, A.C. (2017). DSSD: Deconvolutional Single Shot Detector. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zheng, L., Fu, C., and Zhao, Y. (2018). Extend the shallow part of single shot multibox detector via convolutional neural network. arXiv.","DOI":"10.1117\/12.2503001"},{"key":"ref_17","unstructured":"Li, Z., and Zhou, F. (2017). FSSD: Feature Fusion Single Shot Multibox Detector. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the 2017 IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"154","DOI":"10.1007\/s11263-013-0620-5","article-title":"Selective search for object recognition","volume":"104","author":"Uijlings","year":"2013","journal-title":"Int. J. Comput. Vis."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","article-title":"The Pascal Visual Object Classes (Voc) challenge","volume":"88","author":"Everingham","year":"2010","journal-title":"Int. J. Comput. Vis."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the 2014 European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhu, K., Chen, G., Tan, X., Zhang, L., Dai, F., Liao, P., and Gong, Y. (2019). Geospatial object detection on high resolution remote sensing imagery based on double multi-scale feature pyramid network. Remote Sens., 11.","DOI":"10.3390\/rs11070755"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Mo, N., Yan, L., Zhu, R., and Xie, H. (2019). Class-specific anchor based and context-guided multi-class object detection in high resolution remote sensing imagery with a convolutional neural network. Remote Sens., 11.","DOI":"10.3390\/rs11030272"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhuang, S., Wang, P., Jiang, B., Wang, G., and Wang, C. (2019). A Single shot framework with multi-scale feature fusion for geospatial object detection. Remote Sens., 11.","DOI":"10.3390\/rs11050594"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Xie, W., Qin, H., Li, Y., Wang, Z., and Lei, J. (2019). A novel effectively optimized one-stage network for object detection in remote sensing imagery. Remote Sens., 11.","DOI":"10.3390\/rs11111376"},{"key":"ref_26","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_27","unstructured":"Ioffe, S., and Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv."},{"key":"ref_28","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Ma, N., Zhang, X., Zheng, H.-T., and Sun, J. (2018, January 8\u201314). Shufflenet V2: Practical guidelines for efficient cnn architecture design. Proceedings of the 2018 European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"ref_30","unstructured":"Nair, V., and Hinton, G.E. (2010, January 21\u201324). Rectified linear units improve restricted boltzmann machines. Proceedings of the 27th International Conference on Machine Learning, Haifa, Israel."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Li, Y., Chen, Y., Wang, N., and Zhang, Z. (2019, January 27\u201328). Scale-aware trident networks for object detection. Proceedings of the 2019 IEEE International Conference on Computer Vision, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00615"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Hu, P., and Ramanan, D. (2017, January 21\u201326). Finding tiny faces. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.166"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhang, S., Yang, J., and Schiele, B. (2018, January 18\u201323). Occluded pedestrian detection through guided attention in CNNs. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00731"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Han, X., Zhong, Y., and Zhang, L. (2017). An efficient and robust integrated geospatial object detection framework for high spatial resolution remote sensing imagery. Remote Sens., 9.","DOI":"10.3390\/rs9070666"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2015, January 7\u201313). Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. Proceedings of the 2015 IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.123"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Neubeck, A., and Van Gool, L. (2006, January 20\u201324). Efficient non-maximum suppression. Proceedings of the 18th International Conference on Pattern Recognition (ICPR\u201906), Hong Kong, China.","DOI":"10.1109\/ICPR.2006.479"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). Yolo9000: Better, faster, stronger. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"7405","DOI":"10.1109\/TGRS.2016.2601622","article-title":"Learning Rotation-invariant convolutional neural networks for object detection in VHR optical remote sensing images","volume":"54","author":"Cheng","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_40","unstructured":"Maas, A.L., Hannun, A.Y., and Ng, A.Y. (2013, January 16\u201321). Rectifier nonlinearities improve neural network acoustic models. Proceedings of the 30th International Conference on Machine Learning, Atlanta, GA, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/22\/6530\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:33:39Z","timestamp":1760178819000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/22\/6530"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,11,15]]},"references-count":40,"journal-issue":{"issue":"22","published-online":{"date-parts":[[2020,11]]}},"alternative-id":["s20226530"],"URL":"https:\/\/doi.org\/10.3390\/s20226530","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,11,15]]}}}