{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T17:42:21Z","timestamp":1783100541510,"version":"3.54.6"},"reference-count":51,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2024,2,14]],"date-time":"2024-02-14T00:00:00Z","timestamp":1707868800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National NSF of China","award":["U19A2058"],"award-info":[{"award-number":["U19A2058"]}]},{"name":"National NSF of China","award":["41971362"],"award-info":[{"award-number":["41971362"]}]},{"name":"National NSF of China","award":["41871248"],"award-info":[{"award-number":["41871248"]}]},{"name":"National NSF of China","award":["62106276"],"award-info":[{"award-number":["62106276"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Cross-view geo-localization aims to locate street-view images by matching them with a collection of GPS-tagged remote sensing (RS) images. Due to the significant viewpoint and appearance differences between street-view images and RS images, this task is highly challenging. While deep learning-based methods have shown their dominance in the cross-view geo-localization task, existing models have difficulties in extracting comprehensive meaningful features from both domains of images. This limitation results in not establishing accurate and robust dependencies between street-view images and the corresponding RS images. To address the aforementioned issues, this paper proposes a novel and lightweight neural network for cross-view geo-localization. Firstly, in order to capture more diverse information, we propose a module for extracting multi-scale features from images. Secondly, we introduce contrastive learning and design a contrastive loss to further enhance the robustness in extracting and aligning meaningful multi-scale features. Finally, we conduct comprehensive experiments on two open benchmarks. The experimental results have demonstrated the superiority of the proposed method over the state-of-the-art methods.<\/jats:p>","DOI":"10.3390\/rs16040678","type":"journal-article","created":{"date-parts":[[2024,2,14]],"date-time":"2024-02-14T06:59:26Z","timestamp":1707893966000},"page":"678","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["GeoViewMatch: A Multi-Scale Feature-Matching Network for Cross-View Geo-Localization Using Swin-Transformer and Contrastive Learning"],"prefix":"10.3390","volume":"16","author":[{"given":"Wenhui","family":"Zhang","sequence":"first","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhinong","family":"Zhong","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7880-3394","authenticated-orcid":false,"given":"Hao","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China"},{"name":"Key Laboratory of Natural Resources Monitoring and Supervision in Southern Hilly Region, Ministry of Natural Resources, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ning","family":"Jing","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China"},{"name":"Key Laboratory of Natural Resources Monitoring and Supervision in Southern Hilly Region, Ministry of Natural Resources, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,2,14]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"3681","DOI":"10.1109\/TKDE.2020.3025580","article-title":"Deep learning for spatio-temporal data mining: A survey","volume":"34","author":"Wang","year":"2020","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"169","DOI":"10.1016\/j.isprsjprs.2023.03.012","article-title":"SemiRoadExNet: A semi-supervised network for road extraction from remote sensing imagery via adversarial learning","volume":"198","author":"Chen","year":"2023","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Sun, C., Wu, J., Chen, H., and Du, C. (2022). SemiSANet: A semi-supervised high-resolution remote sensing image change detection model using Siamese networks with graph attention. Remote Sens., 14.","DOI":"10.3390\/rs14122801"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Brosh, E., Friedmann, M., Kadar, I., Yitzhak Lavy, L., Levi, E., Rippa, S., Lempert, Y., Fernandez-Ruiz, B., Herzig, R., and Darrell, T. (2019, January 15\u201320). Accurate visual localization for automotive applications. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Long Beach, CA, USA.","DOI":"10.1109\/CVPRW.2019.00170"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"867","DOI":"10.1109\/TCSVT.2021.3061265","article-title":"Each part matters: Local patterns facilitate cross-view geo-localization","volume":"32","author":"Wang","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Zhu, S., Yang, T., and Chen, C. (2021, January 20\u201325). Vigor: Cross-view image geo-localization beyond one-to-one retrieval. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00364"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Weyand, T., Kostrikov, I., and Philbin, J. (2016, January 11\u201314). Planet-photo geolocation with convolutional neural networks. Proceedings of the Computer Vision\u2014ECCV 2016: 14th European Conference, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_3"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Sun, B., Chen, C., Zhu, Y., and Jiang, J. (2019, January 8\u201312). Geocapsnet: Ground to aerial view image geo-localization using capsule network. Proceedings of the 2019 IEEE International Conference on Multimedia and Expo (ICME), Shanghai, China.","DOI":"10.1109\/ICME.2019.00133"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Hu, S., Feng, M., Nguyen, R.M., and Lee, G.H. (2018, January 18\u201323). Cvm-net: Cross-view matching network for image-based ground-to-aerial geo-localization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00758"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Liu, L., and Li, H. (2019, January 15\u201320). Lending orientation to neural networks for cross-view geo-localization. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00577"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Xu, Y., Wu, S., Du, C., Li, J., and Jing, N. (2022, January 15\u201318). UAV Image Geo-Localization by Point-Line-Patch Feature Matching and ICLK Optimization. Proceedings of the 2022 29th International Conference on Geoinformatics, Beijing, China.","DOI":"10.1109\/Geoinformatics57846.2022.9963796"},{"key":"ref_12","unstructured":"Shi, Y., Liu, L., Yu, X., and Li, H. (2019). Spatial-aware feature aggregation for image based cross-view geo-localization. Adv. Neural Inf. Process. Syst., 32."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Tian, Y., Chen, C., and Shah, M. (2017, January 21\u201326). Cross-view image matching for geo-localization in urban environments. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.216"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zhu, S., Yang, T., and Chen, C. (2021, January 3\u20138). Revisiting street-to-aerial view image geo-localization and orientation estimation. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV48630.2021.00080"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Shi, Y., Yu, X., Campbell, D., and Li, H. (2020, January 13\u201319). Where am i looking at?. joint location and orientation estimation by cross-view matching. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00412"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Vo, N.N., and Hays, J. (2016, January 11\u201314). Localizing and orienting street views using overhead imagery. Proceedings of the Computer Vision\u2014ECCV 2016: 14th European Conference, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_30"},{"key":"ref_17","unstructured":"Mirza, M., and Osindero, S. (2014). Conditional generative adversarial nets. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Lu, X., Li, Z., Cui, Z., Oswald, M.R., Pollefeys, M., and Qin, R. (2020, January 13\u201319). Geometry-aware satellite-to-ground image synthesis for urban areas. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00094"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Regmi, K., and Borji, A. (2018, January 18\u201323). Cross-view image synthesis using conditional gans. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00369"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Toker, A., Zhou, Q., Maximov, M., and Leal-Taix\u00e9, L. (2021, January 20\u201325). Coming down to earth: Satellite-to-street view synthesis for geo-localization. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00642"},{"key":"ref_21","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"4376","DOI":"10.1109\/TCSVT.2021.3135013","article-title":"A transformer-based feature segmentation and region alignment method for UAV-view geo-localization","volume":"32","author":"Dai","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zhu, S., Shah, M., and Chen, C. (2022, January 18\u201324). Transgeo: Transformer is all you need for cross-view image geo-localization. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00123"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Pramanick, S., Nowara, E.M., Gleason, J., Castillo, C.D., and Chellappa, R. (2022, January 23\u201327). Where in the world is this image?. transformer-based geo-localization in the wild. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19839-7_12"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 11\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wu, Z., Xiong, Y., Yu, S.X., and Lin, D. (2018, January 18\u201323). Unsupervised feature learning via non-parametric instance discrimination. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00393"},{"key":"ref_27","unstructured":"Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. (2020, January 13\u201318). A simple framework for contrastive learning of visual representations. Proceedings of the International Conference on Machine Learning, PMLR, Online."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"105","DOI":"10.1145\/2001269.2001293","article-title":"Building rome in a day","volume":"54","author":"Agarwal","year":"2011","journal-title":"Commun. ACM"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zamir, A.R., and Shah, M. (2010, January 5\u201311). Accurate image localization based on google maps street view. Proceedings of the Computer Vision\u2014ECCV 2010: 11th European Conference on Computer Vision, Heraklion, Crete, Greece.","DOI":"10.1007\/978-3-642-15561-1_19"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Baatz, G., Saurer, O., K\u00f6ser, K., and Pollefeys, M. (2012, January 7\u201313). Large scale visual geo-localization of images in mountainous terrain. Proceedings of the Computer Vision\u2014ECCV 2012: 12th European Conference on Computer Vision, Florence, Italy.","DOI":"10.1007\/978-3-642-33709-3_37"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Belongie, S., and Hays, J. (2013, January 23\u201328). Cross-view image geolocalization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.120"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Workman, S., Souvenir, R., and Jacobs, N. (2015, January 7\u201313). Wide-area image geolocalization with aerial reference imagery. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.451"},{"key":"ref_33","unstructured":"Regmi, K., and Shah, M. (November, January 27). Bridging the domain gap for ground-to-aerial image matching. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Shi, Y., Yu, X., Liu, L., Zhang, T., and Li, H. (2020, January 7\u201312). Optimal feature transport for cross-view image geo-localization. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6875"},{"key":"ref_35","first-page":"29009","article-title":"Cross-view geo-localization with layer-to-layer transformer","volume":"34","author":"Yang","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_36","unstructured":"Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and J\u00e9gou, H. (2021, January 18\u201324). Training data-efficient image transformers & distillation through attention. Proceedings of the International Conference on Machine Learning, PMLR, Online."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"4486","DOI":"10.1109\/TCSVT.2021.3127149","article-title":"SwinNet: Swin transformer drives edge-aware RGB-D and RGB-T salient object detection","volume":"32","author":"Liu","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Li, Z., Chen, H., Jing, N., and Li, J. (2023). RemainNet: Explore Road Extraction from Remote Sensing Image Using Mask Image Modeling. Remote Sens., 15.","DOI":"10.3390\/rs15174215"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Liang, J., Cao, J., Sun, G., Zhang, K., Van Gool, L., and Timofte, R. (2021, January 11\u201317). Swinir: Image restoration using swin transformer. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00210"},{"key":"ref_40","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser,, L., and Polosukhin,, I. (2017). Attention is all you need. Adv. Neural Inf. Process. Syst., 30."},{"key":"ref_41","first-page":"22243","article-title":"Big self-supervised models are strong semi-supervised learners","volume":"33","author":"Chen","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020, January 13\u201319). Momentum contrast for unsupervised visual representation learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"ref_43","unstructured":"Zhang, H., Li, F., Liu, S., Zhang, L., Su, H., Zhu, J., Ni, L.M., and Shum, H.Y. (2022). Dino: Detr with improved denoising anchor boxes for end-to-end object detection. arXiv."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Chen, X., Xie, S., and He, K. (2021, January 11\u201317). An empirical study of training self-supervised vision transformers. Proceedings of the CVF International Conference on Computer Vision (ICCV), Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00950"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Zhai, M., Bessinger, Z., Workman, S., and Jacobs, N. (2017, January 21\u201326). Predicting ground-level scene layout from aerial imagery. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.440"},{"key":"ref_46","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., and Antiga, L. (2019). Pytorch: An imperative style, high-performance deep learning library. Adv. Neural Inf. Process. Syst., 32."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., and Dong, L. (2022, January 18\u201324). Swin transformer v2: Scaling up capacity and resolution. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01170"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_49","unstructured":"Loshchilov, I., and Hutter, F. (2017). Decoupled weight decay regularization. arXiv."},{"key":"ref_50","unstructured":"Loshchilov, I., and Hutter, F. (2016). Sgdr: Stochastic gradient descent with warm restarts. arXiv."},{"key":"ref_51","unstructured":"Kwon, J., Kim, J., Park, H., and Choi, I.K. (2021, January 18\u201324). Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks. Proceedings of the International Conference on Machine Learning, PMLR, Online."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/16\/4\/678\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T13:59:35Z","timestamp":1760104775000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/16\/4\/678"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,14]]},"references-count":51,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2024,2]]}},"alternative-id":["rs16040678"],"URL":"https:\/\/doi.org\/10.3390\/rs16040678","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,2,14]]}}}