{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,24]],"date-time":"2026-04-24T18:49:04Z","timestamp":1777056544049,"version":"3.51.4"},"reference-count":50,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2018,10,9]],"date-time":"2018-10-09T00:00:00Z","timestamp":1539043200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["41622107, 41771385, 41271454 and 41371344"],"award-info":[{"award-number":["41622107, 41771385, 41271454 and 41371344"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>A deep neural network is suitable for remote sensing image pixel-wise classification because it effectively extracts features from the raw data. However, remote sensing images with higher spatial resolution exhibit smaller inter-class differences and greater intra-class differences; thus, feature extraction becomes more difficult. The attention mechanism, as a method that simulates the manner in which humans comprehend and perceive images, is useful for the quick and accurate acquisition of key features. In this study, we propose a novel neural network that incorporates two kinds of attention mechanisms in its mask and trunk branches; i.e., control gate (soft) and feedback attention mechanisms, respectively, based on the branches\u2019 primary roles. Thus, a deep neural network can be equipped with an attention mechanism to perform pixel-wise classification for very high-resolution remote sensing (VHRRS) images. The control gate attention mechanism in the mask branch is utilized to build pixel-wise masks for feature maps, to assign different priorities to different locations on different channels for feature extraction recalibration, to apply stress to the effective features, and to weaken the influence of other profitless features. The feedback attention mechanism in the trunk branch allows for the retrieval of high-level semantic features. Hence, additional aids are provided for lower layers to re-weight the focus and to re-update higher-level feature extraction in a target-oriented manner. These two attention mechanisms are fused to form a neural network module. By stacking various modules with different-scale mask branches, the network utilizes different attention-aware features under different local spatial structures. The proposed method is tested on the VHRRS images from the BJ-02, GF-02, Geoeye, and Quickbird satellites, and the influence of the network structure and the rationality of the network design are discussed. Compared with other state-of-the-art methods, our proposed method achieves competitive accuracy, thereby proving its effectiveness.<\/jats:p>","DOI":"10.3390\/rs10101602","type":"journal-article","created":{"date-parts":[[2018,10,9]],"date-time":"2018-10-09T11:10:44Z","timestamp":1539083444000},"page":"1602","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":72,"title":["Attention-Mechanism-Containing Neural Networks for High-Resolution Remote Sensing Image Classification"],"prefix":"10.3390","volume":"10","author":[{"given":"Rudong","family":"Xu","sequence":"first","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping, and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8757-0174","authenticated-orcid":false,"given":"Yiting","family":"Tao","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping, and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhongyuan","family":"Lu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping, and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9446-5850","authenticated-orcid":false,"given":"Yanfei","family":"Zhong","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping, and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2018,10,9]]},"reference":[{"key":"ref_1","unstructured":"Hwang, J.J., and Liu, T.L. (arXiv, 2015). Pixel-wise deep learning for contour detection, arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"98","DOI":"10.1007\/s11263-014-0733-5","article-title":"The pascal visual object classes challenge: A retrospective","volume":"111","author":"Everingham","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Huang, Z., Cheng, G., Wang, H., Li, H., Shi, L., and Pan, C. (2016, January 10\u201315). Building extraction from multi-source remote sensing images via deep deconvolution neural networks. Proceedings of the 2016 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Beijing, China.","DOI":"10.1109\/IGARSS.2016.7729471"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"709","DOI":"10.1109\/LGRS.2017.2672734","article-title":"Road structure refined cnn for road extraction in aerial image","volume":"14","author":"Wei","year":"2017","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"504","DOI":"10.1126\/science.1127647","article-title":"Reducing the dimensionality of data with neural networks","volume":"313","author":"Hinton","year":"2006","journal-title":"Science"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","article-title":"Spatial pyramid pooling in deep convolutional networks for visual recognition","volume":"37","author":"He","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"120","DOI":"10.1016\/j.isprsjprs.2017.11.021","article-title":"A new deep convolutional neural network for fast hyperspectral image classification","volume":"145","author":"Paoletti","year":"2017","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"2940","DOI":"10.1109\/TGRS.2007.902824","article-title":"An innovative neural-net method to detect temporal changes in high-resolution optical satellite imagery","volume":"45","author":"Pacifici","year":"2007","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"6232","DOI":"10.1109\/TGRS.2016.2584107","article-title":"Deep feature extraction and classification of hyperspectral images based on convolutional neural networks","volume":"54","author":"Chen","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1349","DOI":"10.1109\/TGRS.2015.2478379","article-title":"Unsupervised deep feature extraction for remote sensing image classification","volume":"54","author":"Romero","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_11","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20138). Imagenet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems (NIPS), Harrahs and Harveys, NV, USA."},{"key":"ref_12","unstructured":"Simonyan, K., and Zisserman, A. (arXiv, 2014). Very deep convolutional networks for large-scale image recognition, arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 8\u201310). Going deeper with convolutions. Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_14","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K.Q. (2016, January 11\u201314). Deep networks with stochastic depth. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46493-0_39"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"4544","DOI":"10.1109\/TGRS.2016.2543748","article-title":"Spectral\u2013spatial feature extraction for hyperspectral image classification: A dimension reduction and deep learning approach","volume":"54","author":"Zhao","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"762","DOI":"10.3390\/a6040762","article-title":"Very high resolution satellite image classification using fuzzy rule-based systems","volume":"6","author":"Jabari","year":"2013","journal-title":"Algorithms"},{"key":"ref_18","unstructured":"Larochelle, H., and Hinton, G.E. (2010, January 6\u201311). Learning to combine foveal glimpses with a third-order boltzmann machine. Proceedings of the Advances in Neural Information Processing Systems (NIPS), Vancouver, BC, Canada."},{"key":"ref_19","unstructured":"Mnih, V., Heess, N., and Graves, A. (2014, January 8\u201313). Recurrent models of visual attention. Proceedings of the Advances in Neural Information Processing Systems (NIPS), Montreal, QC, Canada."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wang, F., Jiang, M., Qian, C., Yang, S., Li, C., Zhang, H., Wang, X., and Tang, X. (arXiv, 2017). Residual attention network for image classification, arXiv.","DOI":"10.1109\/CVPR.2017.683"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1487","DOI":"10.1109\/TIP.2017.2774041","article-title":"Object-part attention model for fine-grained image classification","volume":"27","author":"Peng","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"113","DOI":"10.1109\/TIP.2018.2865280","article-title":"Attention couplenet: Fully convolutional attention coupling network for object detection","volume":"28","author":"Zhu","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Cao, C., Liu, X., Yang, Y., Yu, Y., Wang, J., Wang, Z., Huang, Y., Wang, L., Huang, C., and Xu, W. (2015, January 13\u201316). Look and think twice: Capturing top-down visual attention with feedback convolutional neural networks. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.338"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Nam, H., Ha, J.-W., and Kim, J. (arXiv, 2016). Dual attention networks for multimodal reasoning and matching, arXiv.","DOI":"10.1109\/CVPR.2017.232"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"2175","DOI":"10.1109\/TGRS.2014.2357078","article-title":"Saliency-guided unsupervised feature learning for scene classification","volume":"53","author":"Zhang","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Hu, J., Xia, G.-S., Hu, F., Sun, H., and Zhang, L. (2015, January 26\u201331). A comparative study of sampling analysis in scene classification of high-resolution remote sensing imagery. Proceedings of the 2015 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Milan, Italy.","DOI":"10.1109\/IGARSS.2015.7326290"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Chen, J., Wang, C., Ma, Z., Chen, J., He, D., and Ackland, S. (2018). Remote sensing scene classification based on convolutional neural networks pre-trained using attention-guided sparse filters. Remote Sens., 10.","DOI":"10.3390\/rs10020290"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (arXiv, 2017). Squeeze-and-excitation networks, arXiv.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Yang, Y., Zhong, Z., Shen, T., and Lin, Z. (2018, January 19\u201321). Convolutional neural networks with alternately updated clique. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00256"},{"key":"ref_30","unstructured":"Kim, J.-H., Lee, S.-W., Kwak, D., Heo, M.-O., Kim, J., Ha, J.-W., and Zhang, B.-T. (2016, January 5\u201310). Multimodal residual learning for visual qa. Proceedings of the Advances in Neural Information Processing Systems (NIPS), Barcelona, Spain."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Chen, L.-C., Yang, Y., Wang, J., Xu, W., and Yuille, A.L. (2016, January 27\u201330). Attention to scale: Scale-aware semantic image segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.396"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Kong, S., and Fowlkes, C. (arXiv, 2018). Pixel-wise attentional gating for parsimonious pixel labeling, arXiv.","DOI":"10.1109\/WACV.2019.00114"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Fu, J., Liu, J., Tian, H., Fang, Z., and Lu, H. (arXiv, 2018). Dual attention network for scene segmentation, arXiv.","DOI":"10.1109\/CVPR.2019.00326"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"284","DOI":"10.1038\/72999","article-title":"The neural mechanisms of top-down attentional control","volume":"3","author":"Hopfinger","year":"2000","journal-title":"Nat. Neurosci."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K.Q. (2017, January 21\u201326). Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"5653","DOI":"10.1109\/TGRS.2017.2711275","article-title":"Integrating multilayer features of convolutional neural networks for remote sensing scene classification","volume":"55","author":"Li","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_37","first-page":"23","article-title":"An unsupervised convolutional feature fusion network for deep representation of remote sensing images","volume":"15","author":"Yu","year":"2018","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"3173","DOI":"10.1109\/TGRS.2018.2794326","article-title":"Hyperspectral image classification with deep feature fusion network","volume":"56","author":"Song","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"4843","DOI":"10.1109\/TIP.2017.2725580","article-title":"Going deeper with contextual cnn for hyperspectral image classification","volume":"26","author":"Lee","year":"2017","journal-title":"IEEE Trans. Image Process."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Tao, Y., Xu, M., Lu, Z., and Zhong, Y. (2018). Densenet-based depth-width double reinforced deep learning neural network for high-resolution remote sensing image pixel-wise classification. Remote Sens., 10.","DOI":"10.3390\/rs10050779"},{"key":"ref_41","unstructured":"Bansal, A., Chen, X., Russell, B., Gupta, A., and Ramanan, D. (arXiv, 2017). Pixelnet: Representation of the pixels, by the pixels, and for the pixels, arXiv."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"6805","DOI":"10.1109\/TGRS.2017.2734697","article-title":"Unsupervised-restricted deconvolutional neural network for very high resolution remote-sensing image classification","volume":"55","author":"Tao","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_43","unstructured":"Glorot, X., and Bengio, Y. (2010, January 13\u201315). Understanding the difficulty of training deep feedforward neural networks. Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, Sardinia, Italy."},{"key":"ref_44","unstructured":"Yu, F., and Koltun, V. (arXiv, 2015). Multi-scale context aggregation by dilated convolutions, arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Pinheiro, P.O., and Collobert, R. (2015, January 8\u201310). From image-level to pixel-level labeling with convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298780"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H. (2007, January 3\u20136). Greedy layer-wise training of deep networks. Proceedings of the Advances in Neural Information Processing Systems (NIPS), Vancouver, BC, Canada.","DOI":"10.7551\/mitpress\/7503.003.0024"},{"key":"ref_47","unstructured":"Lee, C.Y., Xie, S., Gallagher, P., Zhang, Z., and Tu, Z. (2015, January 9\u201312). Deeply-supervised nets. Proceedings of the Artificial Intelligence and Statistics, San Diego, CA, USA."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"5677","DOI":"10.1109\/TGRS.2015.2427791","article-title":"Domain adaptation for remote sensing image classification: A low-rank reconstruction and instance weighting label propagation inspired algorithm","volume":"53","author":"Shi","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_49","unstructured":"Coates, A., Ng, A., and Lee, H. (2011, January 11\u201313). In An analysis of single-layer networks in unsupervised feature learning. Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, Lauderdale, FL, USA."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"881","DOI":"10.1109\/TGRS.2016.2616585","article-title":"Dense semantic labeling of subdecimeter resolution images with convolutional neural networks","volume":"55","author":"Volpi","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/10\/10\/1602\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T15:24:34Z","timestamp":1760196274000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/10\/10\/1602"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,10,9]]},"references-count":50,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2018,10]]}},"alternative-id":["rs10101602"],"URL":"https:\/\/doi.org\/10.3390\/rs10101602","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,10,9]]}}}