{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,16]],"date-time":"2026-05-16T04:05:23Z","timestamp":1778904323158,"version":"3.51.4"},"reference-count":33,"publisher":"MDPI AG","issue":"24","license":[{"start":{"date-parts":[[2022,12,15]],"date-time":"2022-12-15T00:00:00Z","timestamp":1671062400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Local feature matching is a part of many large vision tasks. Local feature matching usually consists of three parts: feature detection, description, and matching. The matching task usually serves a downstream task, such as camera pose estimation, so geometric information is crucial for the matching task. We propose the geometric feature embedding matching method (GFM) for local feature matching. We propose the adaptive keypoint geometric embedding module dynamic adjust keypoint position information and the orientation geometric embedding displayed modeling of geometric information about rotation. Subsequently, we interleave the use of self-attention and cross-attention for local feature enhancement. The predicted correspondences are multiplied by the local features. The correspondences are solved by computing dual-softmax. An intuitive human extraction and matching scheme is implemented. In order to verify the effectiveness of our proposed method, we performed validation on three datasets (MegaDepth, Hpatches, Aachen Day-Night v1.1) according to their respective metrics, and the results showed that our method achieved satisfactory results in all scenes.<\/jats:p>","DOI":"10.3390\/s22249882","type":"journal-article","created":{"date-parts":[[2022,12,16]],"date-time":"2022-12-16T03:55:39Z","timestamp":1671162939000},"page":"9882","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Learning Geometric Feature Embedding with Transformers for Image Matching"],"prefix":"10.3390","volume":"22","author":[{"given":"Xiaohu","family":"Nan","sequence":"first","affiliation":[{"name":"Shanghai Institute of Technical Physics, Chinese Academy of Sciences, Shanghai 200083, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"},{"name":"Key Laboratory of Infrared System Detection and Imaging Technology, Chinese Academy of Sciences, Shanghai 200083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lei","family":"Ding","sequence":"additional","affiliation":[{"name":"Shanghai Institute of Technical Physics, Chinese Academy of Sciences, Shanghai 200083, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"},{"name":"Key Laboratory of Infrared System Detection and Imaging Technology, Chinese Academy of Sciences, Shanghai 200083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,12,15]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"611","DOI":"10.1109\/TPAMI.2017.2658577","article-title":"Direct Sparse Odometry","volume":"40","author":"Engel","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Lindenberger, P., Sarlin, P.E., Larsson, V., and Pollefeys, M. (2021, January 10\u201317). Pixel-Perfect Structure-from-Motion with Featuremetric Refinement. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV) 2021, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00593"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Sch\u00f6nberger, J.L., and Frahm, J.M. (2016, January 27\u201330). Structure-from-Motion Revisited. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.445"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Sarlin, P.E., Cadena, C., Siegwart, R., and Dymczyk, M. (2019, January 15\u201320). From Coarse to Fine: Robust Hierarchical Localization at Large Scale. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01300"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1293","DOI":"10.1109\/TPAMI.2019.2952114","article-title":"InLoc: Indoor Visual Localization with Dense Matching and View Synthesis","volume":"43","author":"Taira","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"710","DOI":"10.1109\/TPAMI.2019.2909864","article-title":"Context-Aware Visual Policy Network for Fine-Grained Image Captioning","volume":"44","author":"Zha","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Sarlin, P.E., DeTone, D., Malisiewicz, T., and Rabinovich, A. (2020, January 14\u201319). SuperGlue: Learning Feature Matching with Graph Neural Networks. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seoul, Republic of Korea.","DOI":"10.1109\/CVPR42600.2020.00499"},{"key":"ref_8","unstructured":"Battaglia, P.W., Hamrick, J.B., Bapst, V., Sanchez-Gonzalez, A., Zambaldi, V.F., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., and Faulkner, R. (2018). Relational inductive biases, deep learning, and graph networks. arXiv."},{"key":"ref_9","unstructured":"Li, X., Han, K., Li, S., and Prisacariu, V.A. (2020). Dual-Resolution Correspondence Networks. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Li, Z., and Snavely, N. (2018, January 18\u201323). MegaDepth: Learning Single-View Depth Prediction from Internet Photos. Proceedings of the Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00218"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Balntas, V., Lenc, K., Vedaldi, A., and Mikolajczyk, K. (2017, January 21\u201326). HPatches: A Benchmark and Evaluation of Handcrafted and Learned Local Descriptors. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.410"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Sattler, T., Maddern, W., Toft, C., Torii, A., Hammarstrand, L., Stenborg, E., Safari, D., Okutomi, M., Pollefeys, M., and Sivic, J. (2018, January 18\u201323). Benchmarking 6DOF Outdoor Visual Localization in Changing Conditions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00897"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Leonardis, A., Bischof, H., and Pinz, A. (2006, January 7\u201313). Machine Learning for High-Speed Corner Detection. Proceedings of the Computer Vision\u2014ECCV 2006, Graz, Austria.","DOI":"10.1007\/11744047"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Rublee, E., Rabaud, V., Konolige, K., and Bradski, G. (2011, January 6\u201313). ORB: An efficient alternative to SIFT or SURF. Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126544"},{"key":"ref_15","unstructured":"DeTone, D., Malisiewicz, T., and Rabinovich, A. (2017). Toward Geometric Deep SLAM. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"DeTone, D., Malisiewicz, T., and Rabinovich, A. (2018, January 18\u201322). SuperPoint: Self-Supervised Interest Point Detection and Description. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00060"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Dusmanu, M., Rocco, I., Pajdla, T., Pollefeys, M., Sivic, J., Torii, A., and Sattler, T. (2019, January 15\u201320). D2-Net: A Trainable CNN for Joint Description and Detection of Local Features. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00828"},{"key":"ref_18","unstructured":"Revaud, J., Weinzaepfel, P., De Souza, C., Pion, N., Csurka, G., Cabon, Y., and Humenberger, M. (2019). R2d2: Repeatable and reliable detector and descriptor. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"355","DOI":"10.1561\/2200000073","article-title":"Computational Optimal Transport","volume":"11","author":"Cuturi","year":"2019","journal-title":"Found. Trends Mach. Learn."},{"key":"ref_20","unstructured":"Cuturi, M. (2013, January 5\u20138). Sinkhorn Distances: Lightspeed Computation of Optimal Transport. Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1020","DOI":"10.1109\/TPAMI.2020.3016711","article-title":"NCNet: Neighbourhood Consensus Networks for Estimating Image Correspondences","volume":"44","author":"Rocco","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Rocco, I., Arandjelovi\u2019c, R., and Sivic, J. (2020, January 23\u201328). Efficient Neighbourhood Consensus Networks via Submanifold Sparse Convolutions. Proceedings of the ECCV, Glasgow, UK.","DOI":"10.1007\/978-3-030-58545-7_35"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2022.3231215","article-title":"A Weakly Supervised Graph Deep Learning Framework for Point Cloud Registration","volume":"60","author":"Sun","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Sitzmann, V., Thies, J., Heide, F., Nie\u00dfner, M., Wetzstein, G., and Zollh\u00f6fer, M. (2019, January 15\u201320). DeepVoxels: Learning Persistent 3D Feature Embeddings. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00254"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Qin, Z., Yu, H., Wang, C., Guo, Y., Peng, Y., and Xu, K. (2022, January 19\u201324). Geometric Transformer for Fast and Robust Point Cloud Registration. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01086"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Jiang, W., Trulls, E., Hosang, J., Tagliasacchi, A., and Yi, K.M. (2021, January 10\u201317). COTR: Correspondence Transformer for Matching Across Images. Proceedings of the ICCV, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00615"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Sun, J., Shen, Z., Wang, Y., Bao, H., and Zhou, X. (2021, January 19\u201321). LoFTR: Detector-Free Local Feature Matching with Transformers. Proceedings of the CVPR, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00881"},{"key":"ref_28","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L.u., and Polosukhin, I. (2017, January 4\u20139). Attention is All you Need. Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_29","first-page":"14254","article-title":"DISK: Learning local features with policy gradient","volume":"33","author":"Tyszkiewicz","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Agarwal, P.K., Aronov, B., Har-Peled, S., Phillips, J.M., Yi, K., and Zhang, W. (2016). Nearest-Neighbor Searching Under Uncertainty II. arXiv.","DOI":"10.1145\/2955098"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Barath, D., and Valasek, G. (2022). Space-Partitioning RANSAC. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-031-19824-3_42"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zhou, Y., Fan, H., Gao, S., Yang, Y., Zhang, X., Li, J., and Guo, Y. (June, January 30). Retrieval and Localization with Observation Constraints. Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi\u2019an, China.","DOI":"10.1109\/ICRA48506.2021.9560987"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhou, Q., Sattler, T., and Leal-Taixe, L. (2021, January 19\u201321). Patch2Pix: Epipolar-Guided Pixel-Level Correspondences. Proceedings of the CVPR, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00464"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/24\/9882\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:42:19Z","timestamp":1760146939000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/24\/9882"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,15]]},"references-count":33,"journal-issue":{"issue":"24","published-online":{"date-parts":[[2022,12]]}},"alternative-id":["s22249882"],"URL":"https:\/\/doi.org\/10.3390\/s22249882","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,15]]}}}