{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T15:14:37Z","timestamp":1781018077915,"version":"3.54.1"},"reference-count":47,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2018,4,6]],"date-time":"2018-04-06T00:00:00Z","timestamp":1522972800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>High spatial resolution (HSR) imagery scene classification has recently attracted increased attention. The bag-of-visual-words (BoVW) model is an effective method for scene classification. However, it can only extract handcrafted features, and it disregards the spatial layout information, whereas deep learning can automatically mine the intrinsic features as well as preserve the spatial location, but it may lose the characteristic information of the HSR images. Although previous methods based on the combination of BoVW and deep learning have achieved comparatively high classification accuracies, they have not explored the combination of handcrafted and deep features, and they just used the BoVW model as a feature coding method to encode the deep features. This means that the intrinsic characteristics of these models were not combined in the previous works. In this paper, to discover more discriminative semantics for HSR imagery, the deep-local-global feature fusion (DLGFF) framework is proposed for HSR imagery scene classification. Differing from the conventional scene classification methods, which utilize only handcrafted features or deep features, DLGFF establishes a framework integrating multi-level semantics from the global texture feature\u2013based method, the BoVW model, and a pre-trained convolutional neural network (CNN). In DLGFF, two different approaches are proposed, i.e., the local and global features fused with the pooling-stretched convolutional features (LGCF) and the local and global features fused with the fully connected features (LGFF), to exploit the multi-level semantics for complex scenes. The experimental results obtained with three HSR image classification datasets confirm the effectiveness of the proposed DLGFF framework. Compared with the published results of the previous scene classification methods, the classification accuracies of the DLGFF framework on the 21-class UC Merced dataset and 12-class Google dataset of SIRI-WHU can reach 99.76%, which is superior to the current state-of-the-art methods. The classification accuracy of the DLGFF framework on the 45-class NWPU-RESISC45 dataset, 96.37 \u00b1 0.05%, is an increase of about 6% when compared with the current state-of-the-art methods. This indicates that the fusion of the global low-level feature, the local mid-level feature, and the deep high-level feature can provide a representative description for HSR imagery.<\/jats:p>","DOI":"10.3390\/rs10040568","type":"journal-article","created":{"date-parts":[[2018,4,10]],"date-time":"2018-04-10T13:06:08Z","timestamp":1523365568000},"page":"568","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":61,"title":["A Deep-Local-Global Feature Fusion Framework for High Spatial Resolution Imagery Scene Classification"],"prefix":"10.3390","volume":"10","author":[{"given":"Qiqi","family":"Zhu","sequence":"first","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9446-5850","authenticated-orcid":false,"given":"Yanfei","family":"Zhong","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanfei","family":"Liu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Liangpei","family":"Zhang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Deren","family":"Li","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2018,4,6]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"180","DOI":"10.1016\/j.isprsjprs.2013.09.014","article-title":"Geographic object-based image analysis\u2014Towards a new paradigm","volume":"87","author":"Blaschke","year":"2014","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"327","DOI":"10.1016\/S0924-2716(02)00162-4","article-title":"A comparison of three image-object methods for the multiscale analysis of landscape structure","volume":"57","author":"Hay","year":"2003","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"4454","DOI":"10.1109\/TGRS.2012.2190079","article-title":"Best merge region-growing segmentation with integrated nonadjacent region object aggregation","volume":"50","author":"Tilton","year":"2012","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"193","DOI":"10.1109\/JSTARS.2010.2081349","article-title":"Bridging the semantic gap for satellite image annotation and automatic mapping applications","volume":"4","author":"Bratasanu","year":"2011","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"439","DOI":"10.1109\/TGRS.2013.2241444","article-title":"Unsupervised feature learning for aerial scene classification","volume":"52","author":"Cheriyadat","year":"2014","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"4157","DOI":"10.1109\/TGRS.2017.2689071","article-title":"Zero-shot scene classification for high spatial resolution remote sensing images","volume":"55","author":"Li","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1023\/A:1011139631724","article-title":"Modeling the shape of the scene: A holistic representation of the spatial envelope","volume":"42","author":"Oliva","year":"2001","journal-title":"Int. J. Comput. Vis."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"971","DOI":"10.1109\/TPAMI.2002.1017623","article-title":"Multiresolution gray-scale and rotation invariant texture classification with local binary patterns","volume":"24","author":"Ojala","year":"2002","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1947","DOI":"10.1109\/TGRS.2014.2351395","article-title":"Pyramid of spatial relatons for scene-level land use classification","volume":"53","author":"Chen","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Yang, Y., and Newsam, S. (2010, January 2\u20135). Bag-of-visual-words and spatial extensions for land-use classification. Proceedings of the 18th SIGSPATIAL International Conference on Advances in Geographic Information Systems, San Jose, CA, USA.","DOI":"10.1145\/1869790.1869829"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"4620","DOI":"10.1109\/JSTARS.2014.2339842","article-title":"Land-use scene classification using a concentric circle-structured multiscale bag-of-visual-words model","volume":"7","author":"Zhao","year":"2014","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"747","DOI":"10.1109\/LGRS.2015.2513443","article-title":"Bag-of-visual-words scene classifier with local and global features for high spatial resolution remote sensing imagery","volume":"13","author":"Zhu","year":"2016","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1204","DOI":"10.1080\/2150704X.2013.858843","article-title":"Scene classification via latent Dirichlet allocation using a hybrid generative\/discriminative strategy for high spatial resolution remote sensing imagery","volume":"4","author":"Zhao","year":"2013","journal-title":"Remote Sens. Lett."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"2250","DOI":"10.1109\/TGRS.2016.2640186","article-title":"Unsupervised feature learning for land-use scene recognition","volume":"55","author":"Fan","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"4427","DOI":"10.1109\/TGRS.2017.2692280","article-title":"Learning a discriminative distance metric with label consistency for scene classification","volume":"55","author":"Wang","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"0950640","DOI":"10.1117\/1.JRS.9.095064","article-title":"Scene classification based on multifeature probabilistic latent semantic analysis for high spatial resolution remote sensing images","volume":"9","author":"Zhong","year":"2015","journal-title":"J. Appl. Remote Sens."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"6207","DOI":"10.1109\/TGRS.2015.2435801","article-title":"Scene classification based on the multifeature fusion probabilistic topic model for high spatial resolution remote sensing imagery","volume":"53","author":"Zhong","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"5525","DOI":"10.1109\/TGRS.2017.2709802","article-title":"Scene classification based on the fully sparse semantic topic model","volume":"55","author":"Zhu","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_19","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20138). Imagenet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, CA, USA."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1109\/TPAMI.2012.59","article-title":"3D convolutional neural networks for human action recognition","volume":"35","author":"Ji","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Taigman, Y., Yang, M., Ranzato, M.A., and Wolf, L. (2014, January 23\u201328). Deepface: Closing the gap to human-level performance in face verification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.220"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Schroff, F., Kalenichenko, D., and Philbin, J. (2015, January 7\u201312). Facenet: A unified embedding for face recognition and clustering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"ref_23","unstructured":"Wallach, I., Dzamba, M., and Heifets, A. (arXiv, 2015). Atomnet: A deep convolutional neural network for bioactivity prediction in structure-based drug discovery, arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"3325","DOI":"10.1109\/TGRS.2014.2374218","article-title":"Object detection in optical remote sensing images based on weakly supervised learning and high-level feature learning","volume":"53","author":"Han","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"4073","DOI":"10.1109\/JSTARS.2016.2517204","article-title":"Spectral\u2013spatial classification of hyperspectral image based on deep auto-encoder","volume":"9","author":"Ma","year":"2016","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1865","DOI":"10.1109\/JPROC.2017.2675998","article-title":"Remote sensing image scene classification: Benchmark and state of the art","volume":"105","author":"Cheng","year":"2017","journal-title":"Proc. IEEE"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database, Computer Vision and Pattern Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Penatti, O.A., Nogueira, K., and dos Santos, J.A. (2015, January 7\u201313). Do deep features generalize from everyday objects to remote sensing and aerial scenes domains?. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Boston, MA, USA.","DOI":"10.1109\/CVPRW.2015.7301382"},{"key":"ref_29","unstructured":"Liu, Q., Hang, R., Song, H., Zhu, F., Plaza, J., and Plaza, A. (arXiv, 2016). Adaptive deep pyramid matching for remote sensing scene classification, arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Wang, J., Luo, C., Huang, H., Zhao, H., and Wang, S. (2017). Transferring pre-trained deep CNNs for remote scene classification with general features learned from linear PCA network. Remote Sens., 9.","DOI":"10.3390\/rs9030225"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Oquab, M., Bottou, L., Laptev, I., and Sivic, J. (2014, January 23\u201328). Learning and transferring mid-level image representations using convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.222"},{"key":"ref_32","unstructured":"Yosinski, J., Clune, J., Bengio, Y., and Lipson, H. (2014, January 8\u201313). How transferable are features in deep neural networks?. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_33","unstructured":"Castelluccio, M., Poggi, G., Sansone, C., and Verdoliva, L. (arXiv, 2015). Land use classification in remote sensing images by convolutional neural networks, arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"220","DOI":"10.1109\/JSTARS.2017.2761800","article-title":"Scene classification via triplet networks","volume":"11","author":"Liu","year":"2017","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"14680","DOI":"10.3390\/rs71114680","article-title":"Transferring deep convolutional neural networks for the scene classification of high-resolution remote sensing imagery","volume":"7","author":"Hu","year":"2015","journal-title":"Remote Sens."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1109\/LGRS.2017.2731997","article-title":"Remote sensing image scene classification using bag of convolutional features","volume":"14","author":"Cheng","year":"2017","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"5653","DOI":"10.1109\/TGRS.2017.2711275","article-title":"Integrating multilayer features of convolutional neural networks for remote sensing scene classification","volume":"55","author":"Li","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"382","DOI":"10.1007\/s11263-009-0312-3","article-title":"Shape-based invariant texture indexing","volume":"88","author":"Xia","year":"2010","journal-title":"Int. J. Comput. Vis."},{"key":"ref_39","unstructured":"Boureau, Y.-L., Ponce, J., and LeCun, Y. (2010, January 21\u201324). A theoretical analysis of feature pooling in visual recognition. Proceedings of the 27th International Conference on Machine Learning (ICML-10), Haifa, Israel."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"539","DOI":"10.1016\/j.patcog.2016.07.001","article-title":"Towards better exploiting convolutional neural networks for remote sensing scene classification","volume":"61","author":"Nogueira","year":"2017","journal-title":"Pattern Recognit."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive image features from scale-invariant keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_42","unstructured":"Fei-Fei, L., and Perona, P. (2005, January 20\u201325). A Bayesian hierarchical model for learning natural scene categories. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Diego, CA, USA."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., and Darrell, T. (2014, January 3\u20137). Caffe: Convolutional architecture for fast feature embedding. Proceedings of the 22nd ACM International Conference on Multimedia, Orlando, FL, USA.","DOI":"10.1145\/2647868.2654889"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"Imagenet large scale visual recognition challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_45","unstructured":"Barla, A., Odone, F., and Verri, A. (2003, January 14\u201317). Histogram intersection kernel for image classification. Proceedings of the International Conference on Image Processing, Barcelona, Spain."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"2108","DOI":"10.1109\/TGRS.2015.2496185","article-title":"Dirichlet-derived multiple topic scene classification model for high spatial resolution remote sensing imagery","volume":"54","author":"Zhao","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"73","DOI":"10.1016\/j.isprsjprs.2016.03.004","article-title":"A spectral\u2013structural bag-of-features scene classifier for very high spatial resolution remote sensing imagery","volume":"116","author":"Zhao","year":"2016","journal-title":"ISPRS J. Photogramm. Remote Sens."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/10\/4\/568\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T14:59:51Z","timestamp":1760194791000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/10\/4\/568"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,4,6]]},"references-count":47,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2018,4]]}},"alternative-id":["rs10040568"],"URL":"https:\/\/doi.org\/10.3390\/rs10040568","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,4,6]]}}}