{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T09:06:56Z","timestamp":1776416816637,"version":"3.51.2"},"reference-count":42,"publisher":"MDPI AG","issue":"13","license":[{"start":{"date-parts":[[2021,6,23]],"date-time":"2021-06-23T00:00:00Z","timestamp":1624406400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the National Key R &amp; D Program of China","award":["2016YFA0602302"],"award-info":[{"award-number":["2016YFA0602302"]}]},{"name":"the Key R &amp; D and Transformation Program of Qinghai Province","award":["2020-SF-C37"],"award-info":[{"award-number":["2020-SF-C37"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Convolutional neural network (CNN) is capable of automatically extracting image features and has been widely used in remote sensing image classifications. Feature extraction is an important and difficult problem in current research. In this paper, data augmentation for avoiding over fitting was attempted to enrich features of samples to improve the performance of a newly proposed convolutional neural network with UC-Merced and RSI-CB datasets for remotely sensed scene classifications. A multiple grouped convolutional neural network (MGCNN) for self-learning that is capable of promoting the efficiency of CNN was proposed, and the method of grouping multiple convolutional layers capable of being applied elsewhere as a plug-in model was developed. Meanwhile, a hyper-parameter C in MGCNN is introduced to probe into the influence of different grouping strategies for feature extraction. Experiments on the two selected datasets, the RSI-CB dataset and UC-Merced dataset, were carried out to verify the effectiveness of this newly proposed convolutional neural network, the accuracy obtained by MGCNN was 2% higher than the ResNet-50. An algorithm of attention mechanism was thus adopted and incorporated into grouping processes and a multiple grouped attention convolutional neural network (MGCNN-A) was therefore constructed to enhance the generalization capability of MGCNN. The additional experiments indicate that the incorporation of the attention mechanism to MGCNN slightly improved the accuracy of scene classification, but the robustness of the proposed network was enhanced considerably in remote sensing image classifications.<\/jats:p>","DOI":"10.3390\/rs13132457","type":"journal-article","created":{"date-parts":[[2021,6,23]],"date-time":"2021-06-23T11:28:41Z","timestamp":1624447721000},"page":"2457","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["A Convolutional Neural Network Based on Grouping Structure for Scene Classification"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8031-4154","authenticated-orcid":false,"given":"Xuan","family":"Wu","sequence":"first","affiliation":[{"name":"Laboratory 5, Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China"},{"name":"Aerospace Information Research Institute, University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7276-5649","authenticated-orcid":false,"given":"Zhijie","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Geography, University of Connecticut, Storrs, CT 06269, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2607-4628","authenticated-orcid":false,"given":"Wanchang","family":"Zhang","sequence":"additional","affiliation":[{"name":"Laboratory 5, Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2653-8920","authenticated-orcid":false,"given":"Yaning","family":"Yi","sequence":"additional","affiliation":[{"name":"Laboratory 5, Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China"},{"name":"Aerospace Information Research Institute, University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9165-5584","authenticated-orcid":false,"given":"Chuanrong","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Geography, University of Connecticut, Storrs, CT 06269, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qiang","family":"Xu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Geohazard Prevention and Geo-Environment Protection, Chengdu University of Technology, Chengdu 610059, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,6,23]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Zhao, H., Zhang, Y., Liu, S., and Shi, J. (2018, January 8\u201314). PSANet: Point-wise Spatial Attention Network for Scene Parsing. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01240-3_17"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"4033","DOI":"10.1109\/TIP.2016.2577886","article-title":"High-resolution image classification integrating spectral-spatial-location cues by conditional random fields","volume":"25","author":"Zhao","year":"2016","journal-title":"IEEE Trans. Image Process."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Yi, Y., Zhang, Z., Zhang, W., Zhang, C., Li, W., and Zhao, T. (2019). Semantic segmentation of urban buildings from VHR remote sensing imagery using a deep convolutional neural network. Remote Sens., 11.","DOI":"10.3390\/rs11151774"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"165356","DOI":"10.1016\/j.ijleo.2020.165356","article-title":"Remote sensing image scene classification using CNN-MLP with data augmentation","volume":"221","author":"Shawky","year":"2020","journal-title":"Optik"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zhang, R., Chen, Z., Zhang, S., Song, F., Zhang, G., Zhou, Q., and Lei, T. (2020). Remote sensing image scene classification with noisy label distillation. Remote Sens., 12.","DOI":"10.3390\/rs12152376"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"250","DOI":"10.1016\/j.ins.2020.06.011","article-title":"Two-stream feature aggregation deep neural network for scene classification of remote sensing images","volume":"539","author":"Xu","year":"2020","journal-title":"Inf. Sci."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1016\/j.isprsjprs.2020.11.025","article-title":"SceneNet: Remote sensing scene classification deep learning network using multi-objective neural evolution architecture search","volume":"172","author":"Ma","year":"2021","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"139208","DOI":"10.1016\/j.scitotenv.2020.139208","article-title":"Planning for the wetland restoration potential based on the viability of the seed bank and the land-use change trajectory in the Sanjiang Plain of China","volume":"733","author":"Shi","year":"2020","journal-title":"Sci. Total Environ."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"104851","DOI":"10.1016\/j.catena.2020.104851","article-title":"Landslide susceptibility mapping using multiscale sampling strategy and convolutional neural network: A case study in Jiuzhaigou region","volume":"195","author":"Yi","year":"2020","journal-title":"Catena"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"143179","DOI":"10.1016\/j.scitotenv.2020.143179","article-title":"Planning a Green Infrastructure Network to Integrate Potential Evacuation Routes and the Urban Green Space in a Coastal City: The Case Study of Haeundae District, Busan, South Korea","volume":"761","author":"Jeong","year":"2021","journal-title":"Sci. Total Environ."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"111912","DOI":"10.1016\/j.rse.2020.111912","article-title":"A generalized approach based on convolutional neural networks for large area cropland mapping at very high resolution","volume":"247","author":"Zhang","year":"2020","journal-title":"Remote Sens. Environ."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"73","DOI":"10.1016\/j.rse.2018.04.050","article-title":"Urban land-use mapping using a deep convolutional neural network with high spatial resolution multispectral remote sensing imagery","volume":"214","author":"Huang","year":"2018","journal-title":"Remote Sens. Environ."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"54","DOI":"10.1016\/j.isprsjprs.2020.06.020","article-title":"Multi-level monitoring of three-dimensional building changes for megacities: Trajectory, morphology, and landscape","volume":"167","author":"Cao","year":"2020","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1386","DOI":"10.1016\/j.asr.2020.05.041","article-title":"An object based framework for building change analysis using 2D and 3D information of high resolution satellite images","volume":"66","author":"Mohammadi","year":"2020","journal-title":"Adv. Space Res."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"and Kwon, S. (2020). CLSTM: Deep feature-based speech emotion recognition using the hierarchical convlstm network. Mathematics, 8.","DOI":"10.3390\/math8122133"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive image features from scale-invariant keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Philbin, J., Chum, O., Isard, M., Sivic, J., and Zisserman, A. (2007, January 17\u201322). Object retrieval with large vocabularies and fast spatial matching. Proceedings of the 2007 IEEE Conference on Computer Vision and Pattern Recognition, Minneapolis, MN, USA.","DOI":"10.1109\/CVPR.2007.383172"},{"key":"ref_18","first-page":"1097","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_19","unstructured":"Simonyan, K., and Zisserman, A. (2015, January 7\u20139). Very deep convolutional networks for large-scale image recognition. Proceedings of the 3rd International Conference on Learning Representations, San Diego, CA, USA."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K.Q. (2017, January 21\u201326). Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., Albanie, S., Sun, G., and Wu, E. (2018, January 18\u201323). Squeeze-and-Excitation Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Li, X., Wang, W., Hu, X., and Yang, J. (2019, January 15\u201320). Selective kernel networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00060"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Xie, S., Girshick, R., Doll\u00e1r, P., Tu, Z., and He, K. (2017, January 21\u201326). Aggregated residual transformations for deep neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.634"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"294","DOI":"10.1016\/j.isprsjprs.2020.01.025","article-title":"Rotation-aware and multi-scale convolutional neural network for object detection in remote sensing images","volume":"161","author":"Fu","year":"2020","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"112308","DOI":"10.1016\/j.rse.2021.112308","article-title":"Change detection using deep learning approach with object-based image analysis","volume":"256","author":"Liu","year":"2021","journal-title":"Remote Sens. Environ."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1666","DOI":"10.1109\/LGRS.2016.2601930","article-title":"Feature-Level Change Detection Using Deep Representation and Feature Change Analysis for Multispectral Imagery","volume":"13","author":"Zhang","year":"2016","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"2232","DOI":"10.1016\/j.rse.2011.04.022","article-title":"Using active learning to adapt remote sensing image classifiers","volume":"115","author":"Tuia","year":"2011","journal-title":"Remote Sens. Environ."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1063","DOI":"10.1016\/S0167-8655(02)00053-3","article-title":"A partially unsupervised cascade classifier for the analysis of multitemporal remote-sensing images","volume":"23","author":"Bruzzone","year":"2002","journal-title":"Pattern Recognit. Lett."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Han, X., Zhong, Y., Cao, L., and Zhang, L. (2017). Pre-trained alexnet architecture with pyramid pooling and supervision for high spatial resolution remote sensing image scene classification. Remote Sens., 9.","DOI":"10.3390\/rs9080848"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Gong, X., Xie, Z., Liu, Y., Shi, X., and Zheng, Z. (2018). Deep salient feature based anti-noise transfer network for scene classification of remote sensing imagery. Remote Sens., 10.","DOI":"10.3390\/rs10030410"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Li, L., Liang, P., Ma, J., Jiao, L., Guo, X., Liu, F., and Sun, C. (2020). A multiscale self-adaptive attention network for remote sensing scene classification. Remote Sens., 12.","DOI":"10.3390\/rs12142209"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"1155","DOI":"10.1109\/TGRS.2018.2864987","article-title":"Scene Classification With Recurrent Attention of VHR Remote Sensing Images","volume":"57","author":"Wang","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"114177","DOI":"10.1016\/j.eswa.2020.114177","article-title":"MLT-DNet: Speech emotion recognition using 1D dilated CNN based on multi-learning trick approach","volume":"167","author":"Mustaqeem","year":"2021","journal-title":"Expert Syst. Appl."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Zhao, X., Zhang, J., Tian, J., Zhuo, L., and Zhang, J. (2020). Residual dense network based on channel-spatial attention for the scene classification of a high-resolution remote sensing image. Remote Sens., 12.","DOI":"10.3390\/rs12111887"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"6344","DOI":"10.1109\/ACCESS.2019.2963769","article-title":"Scene Classification of Remote Sensing Images Based on Saliency Dual Attention Residual Network","volume":"8","author":"Guo","year":"2020","journal-title":"IEEE Access"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Chen, L., Zhang, H., Xiao, J., Nie, L., Shao, J., Liu, W., and Chua, T.S. (2017, January 21\u201326). SCA-CNN: Spatial and channel-wise attention in convolutional networks for image captioning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.667"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"107101","DOI":"10.1016\/j.asoc.2021.107101","article-title":"Att-Net: Enhanced emotion recognition system using lightweight self-attention module","volume":"102","author":"Mustaqeem","year":"2021","journal-title":"Appl. Soft Comput."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Li, H., Dou, X., Tao, C., Wu, Z., Chen, J., Peng, J., Deng, M., and Zhao, L. (2020). Rsi-cb: A large-scale remote sensing image classification benchmark using crowdsourced data. Sensors, 20.","DOI":"10.3390\/s20061594"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Yang, Y., and Newsam, S. (2010, January 2\u20135). Bag-of-visual-words and spatial extensions for land-use classification. Proceedings of the 18th SIGSPATIAL International Symposium on Advances in Geographic Information Systems, San Jose, CA, USA.","DOI":"10.1145\/1869790.1869829"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1016\/j.neucom.2021.02.087","article-title":"Fast hierarchical tucker decomposition with single-mode preservation and tensor subspace analysis for feature extraction from augmented multimodal data","volume":"445","author":"Zdunek","year":"2021","journal-title":"Neurocomputing"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/13\/2457\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:22:15Z","timestamp":1760163735000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/13\/2457"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,6,23]]},"references-count":42,"journal-issue":{"issue":"13","published-online":{"date-parts":[[2021,7]]}},"alternative-id":["rs13132457"],"URL":"https:\/\/doi.org\/10.3390\/rs13132457","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,6,23]]}}}