{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T15:48:34Z","timestamp":1778082514975,"version":"3.51.4"},"reference-count":56,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2018,3,12]],"date-time":"2018-03-12T00:00:00Z","timestamp":1520812800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>With the large number of high-resolution images now being acquired, high spatial resolution (HSR) remote sensing imagery scene classification has drawn great attention but is still a challenging task due to the complex arrangements of the ground objects in HSR imagery, which leads to the semantic gap between low-level features and high-level semantic concepts. As a feature representation method for automatically learning essential features from image data, convolutional neural networks (CNNs) have been introduced for HSR remote sensing image scene classification due to their excellent performance in natural image classification. However, some scene classes of remote sensing images are object-centered, i.e., the scene class of an image is decided by the objects it contains. Although previous methods based on CNNs have achieved comparatively high classification accuracies compared with the traditional methods with handcrafted features, they do not consider the scale variation of the objects in the scenes. This makes it difficult to directly utilize CNNs on those remote sensing images belonging to object-centered classes to extract features that are robust to scale variation, leading to wrongly classified scene images. To solve this problem, scene classification based on a deep random-scale stretched convolutional neural network (SRSCNN) for HSR remote sensing imagery is proposed in this paper. In the proposed method, patches with a random scale are cropped from the image and stretched to the specified scale as the input to train the CNN. This forces the CNN to extract features that are robust to the scale variation. Furthermore, to further improve the performance of the CNN, a robust scene classification strategy is adopted, i.e., multi-perspective fusion. The experimental results obtained using three datasets\u2014the UC Merced dataset, the Google dataset of SIRI-WHU, and the Wuhan IKONOS dataset\u2014confirm that the proposed method performs better than the traditional scene classification methods.<\/jats:p>","DOI":"10.3390\/rs10030444","type":"journal-article","created":{"date-parts":[[2018,3,12]],"date-time":"2018-03-12T13:13:48Z","timestamp":1520860428000},"page":"444","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":80,"title":["Scene Classification Based on a Deep Random-Scale Stretched Convolutional Neural Network"],"prefix":"10.3390","volume":"10","author":[{"given":"Yanfei","family":"Liu","sequence":"first","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9446-5850","authenticated-orcid":false,"given":"Yanfei","family":"Zhong","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Feng","family":"Fei","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qiqi","family":"Zhu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qianqing","family":"Qin","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2018,3,12]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"778","DOI":"10.1016\/j.imavis.2006.07.015","article-title":"Which is the best way to organize\/classify images by content?","volume":"25","author":"Bosch","year":"2007","journal-title":"Image Vis. Comput."},{"key":"ref_2","unstructured":"Yang, Y., and Newsam, S. (2011, January 6\u201313). Spatial pyramid co-occurrence for image classification. Proceedings of the IEEE International Conference on Computer Vision, Barcelona, Spain."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"2296","DOI":"10.1080\/01431161.2014.890762","article-title":"A 2-D wavelet decomposition-based bag-of-visual-words model for land-use scene classification","volume":"35","author":"Zhao","year":"2014","journal-title":"Int. J. Remote Sens."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"4620","DOI":"10.1109\/JSTARS.2014.2339842","article-title":"Land-use scene classification using a concentric circle-structured multiscale bag-of-visual-words model","volume":"7","author":"Zhao","year":"2014","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1947","DOI":"10.1109\/TGRS.2014.2351395","article-title":"Pyramid of spatial relations for scene-level land use classification","volume":"53","author":"Chen","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Hu, F., Xia, G., Wang, Z., Huang, X., Zhang, L., and Sun, H. (2015). Unsupervised feature learning via spectral clustering of multidimensional patches for remotely sensed scene classification. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., 8.","DOI":"10.1109\/JSTARS.2015.2444405"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"73","DOI":"10.1016\/j.isprsjprs.2016.03.004","article-title":"A spectral-structural bag-of-features scene classifier for very high spatial resolution remote sensing imagery","volume":"116","author":"Zhao","year":"2016","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zhong, Y., Wu, S., and Zhao, B. (2017). Scene Semantic Understanding Based on the Spatial Context Relations of Multiple Objects. Remote Sens., 9.","DOI":"10.3390\/rs9101030"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1204","DOI":"10.1080\/2150704X.2013.858843","article-title":"Scene classification via latent Dirichlet allocation using a hybrid generative\/discriminative strategy for high spatial resolution remote sensing imagery","volume":"4","author":"Zhao","year":"2013","journal-title":"Remote Sens. Lett."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"6207","DOI":"10.1109\/TGRS.2015.2435801","article-title":"Scene classification based on the multifeature fusion probabilistic topic model for high spatial resolution remote sensing imagery","volume":"53","author":"Zhong","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_11","first-page":"993","article-title":"Latent Dirichlet allocation","volume":"3","author":"Blei","year":"2003","journal-title":"J. Mach. Learn. Res."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"28","DOI":"10.1109\/LGRS.2009.2023536","article-title":"Semantic annotation of satellite images using latent Dirichlet allocation","volume":"7","author":"Lienou","year":"2010","journal-title":"IEEE Geosci. Remote. Sens. Lett."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1356","DOI":"10.1109\/TGRS.2013.2250978","article-title":"Semantic annotation of satellite images using author\u2013genre\u2013topic model","volume":"52","author":"Luo","year":"2014","journal-title":"IEEE Trans. Geosci. Remote. Sens."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"2108","DOI":"10.1109\/TGRS.2015.2496185","article-title":"Dirichlet-derived multiple topic scene classification model for high spatial resolution remote sensing imagery","volume":"54","author":"Zhao","year":"2016","journal-title":"IEEE Trans. Geosci. Remote. Sens."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zhu, Q., Zhong, Y., Zhang, L., and Li, D. (2018). Scene Classification Based on the Sparse Homogeneous-Heterogeneous Topic Feature Model. IEEE Trans. Geosci. Remote. Sens.","DOI":"10.1109\/TGRS.2017.2781712"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"5525","DOI":"10.1109\/TGRS.2017.2709802","article-title":"Scene Classification Based on the Fully Sparse Semantic Topic Model","volume":"55","author":"Zhu","year":"2017","journal-title":"IEEE Trans. Geosci. Remote. Sens."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Taigman, Y., Yang, M., Ranzato, M., and Wolf, L. (2014, January 24\u201327). Deepface: Closing the gap to human-level performance in face verification. Proceedings of the 27th IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.220"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Schroff, F., Kalenichenko, D., and Philbin, J. (2015, January 7\u201312). Facenet: A unified embedding for face recognition and clustering. Proceedings of the 28th IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Sun, Y., Wang, X., and Tang, X. (2014, January 24\u201327). Deep learning face representation from predicting 10,000 classes. Proceedings of the 27th IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.244"},{"key":"ref_20","unstructured":"Sun, Y., Liang, D., Wang, X., and Tang, X. (2015). Deepid3: Face recognition with very deep neural networks. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Kontschieder, P., Fiterau, M., Criminisi, A., and Rota Bulo, S. (2015, January 7\u201313). Deep neural decision forests. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.172"},{"key":"ref_22","unstructured":"Zhou, B., Lapedriza, A., Xiao, J., Torralba, A., and Oliva, A. (2014, January 8\u201313). Learning deep features for scene recognition using places database. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"295","DOI":"10.1109\/TPAMI.2015.2439281","article-title":"Image super-resolution using deep convolutional networks","volume":"38","author":"Dong","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1109\/TPAMI.2012.59","article-title":"3D convolutional neural networks for human action recognition","volume":"35","author":"Ji","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_25","unstructured":"Wallach, I., Dzamba, M., and Heifets, A. (2015). AtomNet: A Deep Convolutional Neural Network for Bioactivity Prediction in Structure-based Drug Discovery. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"75","DOI":"10.1016\/j.asoc.2017.11.045","article-title":"Computational Intelligence in Optical Remote Sensing Image Processing","volume":"64","author":"Zhong","year":"2018","journal-title":"Appl. Soft Comput."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"2094","DOI":"10.1109\/JSTARS.2014.2329330","article-title":"Deep learning-based classification of hyperspectral data","volume":"7","author":"Chen","year":"2014","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"4073","DOI":"10.1109\/JSTARS.2016.2517204","article-title":"Spectral\u2013Spatial Classification of Hyperspectral Image Based on Deep Auto-Encoder","volume":"9","author":"Ma","year":"2016","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Slavkovikj, V., Verstockt, S., De Neve, W., Van Hoecke, S., and Van de Walle, R. (2015, January 26\u201330). Hyperspectral image classification with convolutional neural networks. Proceedings of the 23rd ACM International Conference on Multimedia, Brisbane, Australia.","DOI":"10.1145\/2733373.2806306"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"3325","DOI":"10.1109\/TGRS.2014.2374218","article-title":"Object detection in optical remote sensing images based on weakly supervised learning and high-level feature learning","volume":"53","author":"Han","year":"2015","journal-title":"IEEE Trans. Geosci. Remote. Sens."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Salberg, A.B. (2015, January 26\u201331). Detection of seals in remote sensing images using features extracted from deep convolutional neural networks. Proceedings of the IEEE International Geoscience and Remote Sensing Symposium, Milan, Italy.","DOI":"10.1109\/IGARSS.2015.7326163"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Song, X., Rui, T., Zha, Z., Wang, X., and Fang, H. (2015, January 19\u201321). The AdaBoost algorithm for vehicle detection based on CNN features. Proceedings of the 7th International Conference on Internet Multimedia Computing and Service, Zhangjiajie, China.","DOI":"10.1145\/2808492.2808497"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"136","DOI":"10.1080\/2150704X.2016.1235299","article-title":"SatCNN: Satellite Image Dataset Classification Using Agile Convolutional Neural Networks","volume":"8","author":"Zhong","year":"2017","journal-title":"Remote Sens. Lett."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"14680","DOI":"10.3390\/rs71114680","article-title":"Transferring deep convolutional neural networks for the scene classification of high-resolution remote sensing imagery","volume":"7","author":"Hu","year":"2015","journal-title":"Remote Sens."},{"key":"ref_35","unstructured":"Castelluccio, M., Poggi, G., Sansone, C., and Verdoliva, L. (2015). Land use classification in remote sensing images by convolutional neural networks. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"2448","DOI":"10.1109\/LGRS.2015.2483680","article-title":"Multiview deep learning for land-use classification","volume":"12","author":"Luus","year":"2015","journal-title":"IEEE Geosci. Remote. Sens. Lett."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"2175","DOI":"10.1109\/TGRS.2014.2357078","article-title":"Saliency-guided unsupervised feature learning for scene classification","volume":"53","author":"Zhang","year":"2015","journal-title":"IEEE Trans. Geosci. Remote. Sens."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"1349","DOI":"10.1109\/TGRS.2015.2478379","article-title":"Unsupervised deep feature extraction for remote sensing image classification","volume":"54","author":"Romero","year":"2016","journal-title":"IEEE Trans. Geosci. Remote. Sens."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"2321","DOI":"10.1109\/LGRS.2015.2475299","article-title":"Deep learning based feature selection for remote sensing scene classification","volume":"12","author":"Zou","year":"2015","journal-title":"IEEE Geosci. Remote. Sens. Lett."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"1793","DOI":"10.1109\/TGRS.2015.2488681","article-title":"Scene classification via a gradient boosting random convolutional network framework","volume":"54","author":"Zhang","year":"2016","journal-title":"IEEE Trans. Geosci. Remote. Sens."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"025006","DOI":"10.1117\/1.JRS.10.025006","article-title":"Large patch convolutional neural networks for the scene classification of high spatial resolution imagery","volume":"10","author":"Zhong","year":"2016","journal-title":"J. Appl. Remote Sens."},{"key":"ref_42","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_44","unstructured":"Lin, M., Chen, Q., and Yan, S. (2013). Network in network. arXiv."},{"key":"ref_45","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G. (2012, January 3\u20138). Imagenet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (arXiv, 2015). Deep residual learning for image recognition, arXiv.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (arXiv, 2016). Identity Mappings in Deep Residual Networks, arXiv.","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"ref_48","unstructured":"Yosinski, J., Clune, J., Nguyen, A., Fuchs, T., and Lipson, H. (arXiv, 2015). Understanding neural networks through deep visualization, arXiv."},{"key":"ref_49","unstructured":"Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Torralba, A. (arXiv, 2014). Object detectors emerge in deep scene CNNs, arXiv."},{"key":"ref_50","unstructured":"Glorot, X., Bordes, A., and Bengio, Y. (2011, January 11\u201313). Deep sparse rectifier neural networks. Proceedings of the International Conference on Artificial Intelligence and Statistics, Ft. Lauderdale, FL, USA."},{"key":"ref_51","unstructured":"Maas, L., Hannun, A.Y., and Ng, A.Y. (2013, January 16\u201321). Rectifier nonlinearities improve neural network acoustic models. Proceedings of the 30th International Conference on Machine Learning, Atlanta, GA, USA."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Zeiler, M.D., Ranzato, M.A., Monga, R., Mao, M., Yang, K., Le, Q.V., Nguyen, P., Senior, A., Vanhoucke, V., and Dean, J. (2013, January 26\u201331). On rectified linear units for speech processing. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6638312"},{"key":"ref_53","unstructured":"Boureau, Y.L., Ponce, J., and LeCun, Y. (2010, January 21\u201324). A theoretical analysis of feature pooling in visual recognition. Proceedings of the 27th International Conference on Machine Learning, Haifa, Israel."},{"key":"ref_54","unstructured":"Hinton, G.E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R.R. (2012). Improving neural networks by preventing co-adaptation of feature detectors. arXiv."},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"439","DOI":"10.1109\/TGRS.2013.2241444","article-title":"Unsupervised feature learning for aerial scene classification","volume":"52","author":"Cheriyadat","year":"2014","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Zhao, B., Zhong, Y., Zhang, L., and Huang, B. (2016). The Fisher Kernel coding framework for high spatial resolution scene classification. Remote. Sens., 8.","DOI":"10.3390\/rs8020157"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/10\/3\/444\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T14:56:44Z","timestamp":1760194604000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/10\/3\/444"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,3,12]]},"references-count":56,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2018,3]]}},"alternative-id":["rs10030444"],"URL":"https:\/\/doi.org\/10.3390\/rs10030444","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,3,12]]}}}