{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T04:32:05Z","timestamp":1781497925471,"version":"3.54.1"},"reference-count":45,"publisher":"MDPI AG","issue":"15","license":[{"start":{"date-parts":[[2020,7,28]],"date-time":"2020-07-28T00:00:00Z","timestamp":1595894400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61471370"],"award-info":[{"award-number":["61471370"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61901500"],"award-info":[{"award-number":["61901500"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Nowadays, object detection methods based on deep learning are applied more and more to the interpretation of optical remote sensing images. However, the complex background and the wide range of object sizes in remote sensing images increase the difficulty of object detection. In this paper, we improve the detection performance by combining the attention information, and generate adaptive anchor boxes based on the attention map. Specifically, the attention mechanism is introduced into the proposed method to enhance the features of the object regions while reducing the influence of the background. The generated attention map is then used to obtain diverse and adaptable anchor boxes using the guided anchoring method. The generated anchor boxes can match better with the scene and the objects, compared with the traditional proposal boxes. Finally, the modulated feature adaptation module is applied to transform the feature maps to adapt to the diverse anchor boxes. Comprehensive evaluations on the DIOR dataset demonstrate the superiority of the proposed method over the state-of-the-art methods, such as RetinaNet, FCOS and CornerNet. The mean average precision of the proposed method is 4.5% higher than the feature pyramid network. In addition, the ablation experiments are also implemented to further analyze the respective influence of different blocks on the performance improvement.<\/jats:p>","DOI":"10.3390\/rs12152416","type":"journal-article","created":{"date-parts":[[2020,7,28]],"date-time":"2020-07-28T10:16:49Z","timestamp":1595931409000},"page":"2416","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":21,"title":["Generating Anchor Boxes Based on Attention Mechanism for Object Detection in Remote Sensing Images"],"prefix":"10.3390","volume":"12","author":[{"given":"Zhuangzhuang","family":"Tian","sequence":"first","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6799-620X","authenticated-orcid":false,"given":"Ronghui","family":"Zhan","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiemin","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Informatics Science and Technology, Zhejiang Sci-Tech University, Hangzhou 310018, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wei","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhiqiang","family":"He","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhaowen","family":"Zhuang","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2020,7,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"4511","DOI":"10.1109\/TGRS.2013.2282355","article-title":"Ship Detection in High-Resolution Optical Imagery Based on Anomaly Detector and Local Shape Feature","volume":"52","author":"Shi","year":"2014","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"30","DOI":"10.1016\/j.isprsjprs.2014.10.007","article-title":"A generic discriminative part-based model for geospatial object detection in optical remote sensing images","volume":"99","author":"Zhang","year":"2015","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"6508","DOI":"10.1109\/TGRS.2013.2296782","article-title":"VHR Object Detection Based on Structural Feature Extraction and Query Expansion","volume":"52","author":"Bai","year":"2014","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_4","unstructured":"Cheng, G., Han, J., Zhou, P., and Guo, L. (2014, January 13\u201318). Scalable multi-class geospatial object detection in high-spatial-resolution remote sensing images. Proceedings of the 2014 IEEE Geoscience and Remote Sensing Symposium, Quebec City, QC, Canada."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"7405","DOI":"10.1109\/TGRS.2016.2601622","article-title":"Learning Rotation-Invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing Images","volume":"54","author":"Cheng","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"7147","DOI":"10.1109\/TGRS.2018.2848901","article-title":"HSF-Net: Multiscale Deep Feature Embedding for Ship Detection in Optical Remote Sensing Imagery","volume":"56","author":"Li","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Ding, J., Xue, N., Long, Y., Xia, G.S., and Lu, Q. (2019, January 15\u201320). Learning RoI Transformer for Oriented Object Detection in Aerial Images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00296"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1640","DOI":"10.1109\/LGRS.2019.2904076","article-title":"Remote Sensing Airport Detection Based on End-to-End Deep Transferable Convolutional Neural Networks","volume":"16","author":"Li","year":"2019","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, Faster, Stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_11","unstructured":"Joseph, R., and Ali, F. (2018). YOLOv3: An Incremental Improvement. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Leibe, B., Matas, J., Sebe, N., and Welling, M. (2016). SSD: Single Shot MultiBox Detector. Computer Vision\u2014ECCV 2016, Springer International Publishing.","DOI":"10.1007\/978-3-319-46454-1"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Lin, T., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal Loss for Dense Object Detection. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"154","DOI":"10.1007\/s11263-013-0620-5","article-title":"Selective Search for Object Recognition","volume":"104","author":"Uijlings","year":"2013","journal-title":"Int. J. Comput. Vis."},{"key":"ref_16","unstructured":"Cortes, C., Lawrence, N.D., Lee, D.D., Sugiyama, M., and Garnett, R. (2015). Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. Advances in Neural Information Processing Systems 28, Curran Associates, Inc."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zhu, X., Hu, H., Lin, S., and Dai, J. (2018). Deformable ConvNets v2: More Deformable, Better Results. arXiv.","DOI":"10.1109\/CVPR.2019.00953"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1016\/j.isprsjprs.2016.03.014","article-title":"A survey on object detection in optical remote sensing images","volume":"117","author":"Cheng","year":"2016","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"382","DOI":"10.1016\/j.isprsjprs.2007.10.005","article-title":"On-line boosting-based car detection from aerial images","volume":"63","author":"Grabner","year":"2008","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1279","DOI":"10.1109\/TGRS.2009.2031812","article-title":"A Nonparametric Feature Extraction and Its Application to Nearest Neighbor Classification for Hyperspectral Image Data","volume":"48","author":"Yang","year":"2010","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Lang, F., Yang, J., Yan, S., and Qin, F. (2018). Superpixel Segmentation of Polarimetric Synthetic Aperture Radar (SAR) Images Based on Generalized Mean Shift. Remote Sens., 10.","DOI":"10.3390\/rs10101592"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"196","DOI":"10.1016\/j.eswa.2017.04.018","article-title":"River Channel Segmentation in Polarimetric SAR Images","volume":"82","author":"Ciecholewski","year":"2017","journal-title":"Expert Syst. Appl."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1171","DOI":"10.1109\/LGRS.2017.2702062","article-title":"A Median Regularized Level Set for Hierarchical Segmentation of SAR Images","volume":"14","author":"Braga","year":"2017","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"4565","DOI":"10.1109\/JSTARS.2017.2716620","article-title":"Level Set Segmentation Algorithm for High-Resolution Polarimetric SAR Images Based on a Heterogeneous Clutter Model","volume":"10","author":"Jin","year":"2017","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"5512","DOI":"10.1109\/TGRS.2019.2899955","article-title":"R2-CNN: Fast Tiny Object Detection in Large-Scale Remote Sensing Images","volume":"57","author":"Pang","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., Gupta, A., and He, K. (2018, January 18\u201323). Non-local Neural Networks. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Huang, Z., Wang, X., Huang, L., Huang, C., Wei, Y., and Liu, W. (2018). CCNet: Criss-Cross Attention for Semantic Segmentation. arXiv.","DOI":"10.1109\/ICCV.2019.00069"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Cao, Y., Xu, J., Lin, S., Wei, F., and Hu, H. (2019). GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond. arXiv.","DOI":"10.1109\/ICCVW.2019.00246"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-Excitation Networks. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_30","unstructured":"Sermanet, P., Eigen, D., Zhang, X., Mathieu, M., Fergus, R., and LeCun, Y. (2013). OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"167","DOI":"10.1023\/B:VISI.0000022288.19776.77","article-title":"Efficient Graph-Based Image Segmentation","volume":"59","author":"Felzenszwalb","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_32","unstructured":"Ferrari, V., Hebert, M., Sminchisescu, C., and Weiss, Y. (2018). CornerNet: Detecting Objects as Paired Keypoints. Computer Vision\u2014ECCV 2018, Springer International Publishing."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Duan, K., Bai, S., Xie, L., Qi, H., Huang, Q., and Tian, Q. (2019). CenterNet: Keypoint Triplets for Object Detection. arXiv.","DOI":"10.1109\/ICCV.2019.00667"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zhou, X., Zhuo, J., and Krahenbuhl, P. (2019, January 16\u201320). Bottom-Up Object Detection by Grouping Extreme and Center Points. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00094"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Zhu, X., Cheng, D., Zhang, Z., Lin, S., and Dai, J. (2019). An Empirical Study of Spatial Attention Mechanisms in Deep Networks. arXiv.","DOI":"10.1109\/ICCV.2019.00679"},{"key":"ref_36","unstructured":"Ba, J.L., Kiros, J.R., and Hinton, G.E. (2016). Layer Normalization. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Wang, J., Chen, K., Yang, S., Loy, C.C., and Lin, D. (2019, January 16\u201320). Region Proposal by Guided Anchoring. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00308"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"296","DOI":"10.1016\/j.isprsjprs.2019.11.023","article-title":"Object detection in optical remote sensing images: A survey and a new benchmark","volume":"159","author":"Li","year":"2020","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_39","unstructured":"Chen, K., Wang, J., Pang, J., Cao, Y., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., and Xu, J. (2019). MMDetection: Open MMLab Detection Toolbox and Benchmark. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Lin, T., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature Pyramid Networks for Object Detection. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Dollar, P., and Girshick, R. (2017, January 22\u201329). Mask R-CNN. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Tian, Z., Shen, C., Chen, H., and He, T. (2019, April 02). FCOS: Fully Convolutional One-Stage Object Detection. Available online: https:\/\/arxiv.org\/abs\/1904.01355v5.","DOI":"10.1109\/ICCV.2019.00972"},{"key":"ref_44","unstructured":"Kong, T., Sun, F., Liu, H., Jiang, Y., and Shi, J. (2019, April 08). FoveaBox: Beyond Anchor-based Object Detector. Available online: https:\/\/arxiv.org\/abs\/1904.03797v1."},{"key":"ref_45","unstructured":"Simonyan, K., and Zisserman, A. (2014, September 04). Very Deep Convolutional Networks for Large-Scale Image Recognition 2014. Available online: https:\/\/arxiv.org\/abs\/1409.1556."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/15\/2416\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:52:15Z","timestamp":1760176335000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/15\/2416"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,7,28]]},"references-count":45,"journal-issue":{"issue":"15","published-online":{"date-parts":[[2020,8]]}},"alternative-id":["rs12152416"],"URL":"https:\/\/doi.org\/10.3390\/rs12152416","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,7,28]]}}}