{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T23:21:28Z","timestamp":1784071288487,"version":"3.55.0"},"reference-count":50,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2023,1,8]],"date-time":"2023-01-08T00:00:00Z","timestamp":1673136000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"The Science and Technology Development Center of the Ministry of Education of China","award":["2021ZYA08008"],"award-info":[{"award-number":["2021ZYA08008"]}]},{"name":"The Science and Technology Development Center of the Ministry of Education of China","award":["22DZ1100803"],"award-info":[{"award-number":["22DZ1100803"]}]},{"DOI":"10.13039\/501100003399","name":"Science and Technology Commission of Shanghai Municipality","doi-asserted-by":"publisher","award":["2021ZYA08008"],"award-info":[{"award-number":["2021ZYA08008"]}],"id":[{"id":"10.13039\/501100003399","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003399","name":"Science and Technology Commission of Shanghai Municipality","doi-asserted-by":"publisher","award":["22DZ1100803"],"award-info":[{"award-number":["22DZ1100803"]}],"id":[{"id":"10.13039\/501100003399","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Visual geo-localization plays a crucial role in positioning and navigation for unmanned aerial vehicles, whose goal is to match the same geographic target from different views. This is a challenging task due to the drastic variations in different viewpoints and appearances. Previous methods have been focused on mining features inside the images. However, they underestimated the influence of external elements and the interaction of various representations. Inspired by multimodal and bilinear pooling, we proposed a pioneering feature fusion network (MBF) to address these inherent differences between drone and satellite views. We observe that UAV\u2019s status, such as flight height, leads to changes in the size of image field of view. In addition, local parts of the target scene act a role of importance in extracting discriminative features. Therefore, we present two approaches to exploit those priors. The first module is to add status information to network by transforming them into word embeddings. Note that they concatenate with image embeddings in Transformer block to learn status-aware features. Then, global and local part feature maps from the same viewpoint are correlated and reinforced by hierarchical bilinear pooling (HBP) to improve the robustness of feature representation. By the above approaches, we achieve more discriminative deep representations facilitating the geo-localization more effectively. Our experiments on existing benchmark datasets show significant performance boosting, reaching the new state-of-the-art result. Remarkably, the recall@1 accuracy achieves 89.05% in drone localization task and 93.15% in drone navigation task in University-1652, and shows strong robustness at different flight heights in the SUES-200 dataset.<\/jats:p>","DOI":"10.3390\/s23020720","type":"journal-article","created":{"date-parts":[[2023,1,9]],"date-time":"2023-01-09T07:05:09Z","timestamp":1673247909000},"page":"720","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":59,"title":["UAV\u2019s Status Is Worth Considering: A Fusion Representations Matching Method for Geo-Localization"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8105-2336","authenticated-orcid":false,"given":"Runzhe","family":"Zhu","sequence":"first","affiliation":[{"name":"School of Electronic and Electrical Engineering, Shanghai University of Engineering Science, Shanghai 201602, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3645-7550","authenticated-orcid":false,"given":"Mingze","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Electronic and Electrical Engineering, Shanghai University of Engineering Science, Shanghai 201602, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ling","family":"Yin","sequence":"additional","affiliation":[{"name":"School of Electronic and Electrical Engineering, Shanghai University of Engineering Science, Shanghai 201602, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fei","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Electronic and Electrical Engineering, Shanghai University of Engineering Science, Shanghai 201602, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuncheng","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Electronic and Electrical Engineering, Shanghai University of Engineering Science, Shanghai 201602, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,1,8]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Wang, Y., Li, S., Lin, Y., and Wang, M. (2021). Lightweight Deep Neural Network Method for Water Body Extraction from High-Resolution Remote Sensing Images with Multisensors. Sensors, 21.","DOI":"10.3390\/s21217397"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Suo, C., Zhao, J., Zhang, W., Li, P., Huang, R., Zhu, J., and Tan, X. (2021). Research on UAV Three-Phase Transmission Line Tracking and Localization Method Based on Electric Field Sensor Array. Sensors, 21.","DOI":"10.3390\/s21248400"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Zhu, C., Zhu, J., Bu, T., and Gao, X. (2022). Monitoring and Identification of Road Construction Safety Factors via UAV. Sensors, 22.","DOI":"10.3390\/s22228797"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Chen, C.L., He, R., and Peng, C.C. (2022). Development of an Online Adaptive Parameter Tuning vSLAM Algorithm for UAVs in GPS-Denied Environments. Sensors, 22.","DOI":"10.3390\/s22208067"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Hassan, S.I., Alam, M.M., Zia, M.Y.I., Rashid, M., Illahi, U., and Su\u2019ud, M.M. (2022). Rice Crop Counting Using Aerial Imagery and GIS for the Assessment of Soil Health to Increase Crop Yield. Sensors, 22.","DOI":"10.3390\/s22218567"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Oh, D., and Han, J. (2021). Smart Search System of Autonomous Flight UAVs for Disaster Rescue. Sensors, 21.","DOI":"10.3390\/s21206810"},{"key":"ref_7","unstructured":"Bansal, M., Sawhney, H.S., Cheng, H., and Daniilidis, K. (December, January 28). Geo-localization of street views with aerial image databases. Proceedings of the 19th ACM international conference on Multimedia, Scottsdale, AZ, USA."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Senlet, T., and Elgammal, A. (2011, January 6\u201313). A framework for global vehicle localization using stereo images and satellite and road maps. Proceedings of the 2011 IEEE International Conference on Computer Vision Workshops (ICCV Workshops), Barcelona, Spain.","DOI":"10.1109\/ICCVW.2011.6130498"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Belongie, S., and Hays, J. (2013, January 23\u201328). Cross-view image geolocalization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.120"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Castaldo, F., Zamir, A., Angst, R., Palmieri, F., and Savarese, S. (2015, January 7\u201313). Semantic cross-view matching. Proceedings of the IEEE International Conference on Computer Vision Workshops, Santiago, Chile.","DOI":"10.1109\/ICCVW.2015.137"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Gao, J., and Sun, Z. (2022). An Improved ASIFT Image Feature Matching Algorithm Based on POS Information. Sensors, 22.","DOI":"10.3390\/s22207749"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Cui, Y., Belongie, S., and Hays, J. (2015, January 7\u201312). Learning deep representations for ground-to-aerial geolocalization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299135"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Tian, Y., Chen, C., and Shah, M. (2017, January 21\u201326). Cross-view image matching for geo-localization in urban environments. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.216"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Workman, S., Souvenir, R., and Jacobs, N. (2015, January 7\u201313). Wide-area image geolocalization with aerial reference imagery. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.451"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Liu, L., and Li, H. (2019, January 15\u201320). Lending orientation to neural networks for cross-view geo-localization. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00577"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zheng, Z., Wei, Y., and Yang, Y. (2020, January 12\u201316). University-1652: A multi-view multi-source benchmark for drone-based geo-localization. Proceedings of the 28th ACM international conference on Multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3413896"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Ding, L., Zhou, J., Meng, L., and Long, Z. (2020). A practical cross-view image matching method between UAV and satellite for UAV-based geo-localization. Remote Sens., 13.","DOI":"10.3390\/rs13010047"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"867","DOI":"10.1109\/TCSVT.2021.3061265","article-title":"Each part matters: Local patterns facilitate cross-view geo-localization","volume":"32","author":"Wang","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zhuang, J., Dai, M., Chen, X., and Zheng, E. (2021). A Faster and More Effective Cross-View Matching Method of UAV and Satellite Images for UAV Geolocalization. Remote Sens., 13.","DOI":"10.3390\/rs13193979"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"4804","DOI":"10.1109\/TCSVT.2021.3121987","article-title":"UAV-Satellite View Synthesis for Cross-view Geo-Localization","volume":"32","author":"Tian","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"4376","DOI":"10.1109\/TCSVT.2021.3135013","article-title":"A Transformer-Based Feature Segmentation and Region Alignment Method For UAV-View Geo-Localization","volume":"32","author":"Dai","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_23","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_24","first-page":"29009","article-title":"Cross-view geo-localization with layer-to-layer transformer","volume":"34","author":"Yang","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhu, S., Yang, T., and Chen, C. (2021, January 20\u201325). Vigor: Cross-view image geo-localization beyond one-to-one retrieval. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00364"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Chen, Y.C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J. (2020, January 23\u201328). Uniter: Universal image-text representation learning. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58577-8_7"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Huang, Z., Zeng, Z., Huang, Y., Liu, B., Fu, D., and Fu, J. (2021, January 20\u201325). Seeing out of the box: End-to-end pre-training for vision-language representation learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01278"},{"key":"ref_28","unstructured":"Kim, W., Son, B., and Kim, I. (2021, January 18\u201324). Vilt: Vision-and-language transformer without convolution or region supervision. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Tan, H., and Bansal, M. (2019). Lxmert: Learning cross-modality encoder representations from transformers. arXiv.","DOI":"10.18653\/v1\/D19-1514"},{"key":"ref_30","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"458","DOI":"10.1109\/TIP.2020.3037470","article-title":"Data-level recombination and lightweight fusion scheme for RGB-D salient object detection","volume":"30","author":"Wang","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"George, A., and Marcel, S. (2021, January 20\u201325). Cross modal focal loss for rgbd face anti-spoofing. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00779"},{"key":"ref_33","unstructured":"Zheng, A., Wang, Z., Chen, Z., Li, C., and Tang, J. (2021, January 2\u20139). Robust Multi-Modality Person Re-identification. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual."},{"key":"ref_34","first-page":"251","article-title":"Visual instance retrieval with deep convolutional networks","volume":"4","author":"Razavian","year":"2016","journal-title":"ITE Trans. Media Technol. Appl."},{"key":"ref_35","unstructured":"Babenko, A., and Lempitsky, V. (2015). Aggregating deep convolutional features for image retrieval. arXiv."},{"key":"ref_36","unstructured":"Mousavian, A., and Kosecka, J. (2015). Deep convolutional features for image based retrieval and scene categorization. arXiv."},{"key":"ref_37","first-page":"1655","article-title":"Fine-tuning CNN image retrieval with no human annotation","volume":"41","author":"Tolias","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., RoyChowdhury, A., and Maji, S. (2015, January 7\u201313). Bilinear CNN models for fine-grained visual recognition. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.170"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Gao, Y., Beijbom, O., Zhang, N., and Darrell, T. (2016, January 27\u201330). Compact bilinear pooling. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.41"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Fukui, A., Park, D.H., Yang, D., Rohrbach, A., Darrell, T., and Rohrbach, M. (2016, January 1\u20135). Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding. Proceedings of the EMNLP, Austin, TX, USA.","DOI":"10.18653\/v1\/D16-1044"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Yu, C., Zhao, X., Zheng, Q., Zhang, P., and You, X. (2018, January 8\u201314). Hierarchical bilinear pooling for fine-grained visual recognition. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01270-0_35"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 11\u201314). Identity mappings in deep residual networks. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Zhu, R. (2022). SUES-200: A Multi-height Multi-scene Cross-view Image Benchmark Across Drone and Satellite. arXiv.","DOI":"10.1109\/TCSVT.2023.3249204"},{"key":"ref_44","unstructured":"Chopra, S., Hadsell, R., and LeCun, Y. (2005, January 20\u201326). Learning a similarity metric discriminatively, with application to face verification. Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905), San Diego, CA, USA."},{"key":"ref_45","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_46","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Suh, Y., Wang, J., Tang, S., Mei, T., and Lee, K.M. (2018, January 8\u201314). Part-aligned bilinear representations for person re-identification. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_25"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"108335","DOI":"10.1016\/j.compeleceng.2022.108335","article-title":"Learning discriminative representations via variational self-distillation for cross-view geo-localization","volume":"103","author":"Hu","year":"2022","journal-title":"Comput. Electr. Eng."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"34277","DOI":"10.1109\/ACCESS.2022.3162693","article-title":"A Semantic Guidance and Transformer-Based Matching Method for UAVs and Satellite Images for UAV Geo-Localization","volume":"10","author":"Zhuang","year":"2022","journal-title":"IEEE Access"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1016\/j.inffus.2021.02.012","article-title":"A review of multimodal image matching: Methods and applications","volume":"73","author":"Jiang","year":"2021","journal-title":"Inf. Fusion"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/2\/720\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:03:36Z","timestamp":1760119416000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/2\/720"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,8]]},"references-count":50,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2023,1]]}},"alternative-id":["s23020720"],"URL":"https:\/\/doi.org\/10.3390\/s23020720","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,8]]}}}