{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T05:02:12Z","timestamp":1781154132267,"version":"3.54.1"},"reference-count":51,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2026,6,7]],"date-time":"2026-06-07T00:00:00Z","timestamp":1780790400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the Key R&amp;D Program of Jiangsu Province","award":["BE2023010-3"],"award-info":[{"award-number":["BE2023010-3"]}]},{"name":"Basic Research Program of Jiangsu","award":["BK20253037"],"award-info":[{"award-number":["BK20253037"]}]},{"name":"Advanced Technology Research and Development Program of Jiangsu","award":["BF2025018"],"award-info":[{"award-number":["BF2025018"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Local feature matching plays a critical role in robotic SLAM and visual localization. However, in weakly textured indoor industrial environments, lightweight appearance-based methods often struggle to learn discriminative and stable local features. To address this challenge, this paper proposes GAEFeat, short for Geometry-Aware Efficient Feature, a lightweight vision\u2013geometric feature learning network. To address the scarcity of specialized training data, we integrated robotic arm pose priors with depth information to automatically generate cross-view supervision signals and surface-normal labels. Based on this strategy, we constructed two complementary datasets, including a simulated dataset and a real-world dataset, to support feature learning and evaluation in weakly textured indoor industrial environments. For feature extraction, we design a dual enhancement mechanism consisting of a geometric auxiliary branch and a geometry-aware enhancement (GAE) module. The former guides the network to perceive local surface structures through surface normal supervision, while the latter utilizes a gating mechanism to achieve deep fusion between geometric priors and 2D texture descriptors. Experimental results demonstrate that GAEFeat achieves strong robustness and high inference efficiency in relative pose estimation, homography estimation, and visual localization tasks, with particularly notable advantages in near-field, weakly textured industrial scenes. The framework achieves an inference latency of only 3.9 ms on the NVIDIA Jetson AGX Orin edge platform, demonstrating its real-time capability and practical potential for deployment in edge computing environments.<\/jats:p>","DOI":"10.3390\/jimaging12060253","type":"journal-article","created":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T02:32:46Z","timestamp":1780885966000},"page":"253","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["3D Geometry-Aware Efficient Feature Matching for Weakly Textured Scenes"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7838-9410","authenticated-orcid":false,"given":"Libo","family":"Sun","sequence":"first","affiliation":[{"name":"School of Instrument Science and Engineering, Southeast University, Nanjing 210096, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-8746-5937","authenticated-orcid":false,"given":"Yidong","family":"Yan","sequence":"additional","affiliation":[{"name":"School of Instrument Science and Engineering, Southeast University, Nanjing 210096, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenqi","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Instrument Science and Engineering, Southeast University, Nanjing 210096, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenhu","family":"Qin","sequence":"additional","affiliation":[{"name":"School of Instrument Science and Engineering, Southeast University, Nanjing 210096, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,6,7]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Liu, L., Wang, C., Feng, C., Gong, W., Zhang, L., Liao, L., and Feng, C. (2024). Incremental SFM 3D reconstruction based on deep learning. Electronics, 13.","DOI":"10.3390\/electronics13142850"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"230","DOI":"10.1016\/j.isprsjprs.2020.04.016","article-title":"Efficient structure from motion for large-scale UAV images: A review and a comparison of SfM tools","volume":"167","author":"Jiang","year":"2020","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1293","DOI":"10.1109\/TPAMI.2019.2952114","article-title":"InLoc: Indoor visual localization with dense matching and view synthesis","volume":"43","author":"Taira","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Sarlin, P.E., Cadena, C., Siegwart, R., and Dymczyk, M. (2019). From coarse to fine: Robust hierarchical localization at large scale. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15\u201320 June 2019, IEEE.","DOI":"10.1109\/CVPR.2019.01300"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Chen, W., Liao, X., Sun, Y., and Wang, Q. (2020). Improved orb-slam based 3d dense reconstruction for monocular endoscopic image. Proceedings of the 2020 International Conference on Virtual Reality and Visualization (ICVRV), IEEE.","DOI":"10.1109\/ICVRV51359.2020.00030"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"5191","DOI":"10.1109\/LRA.2021.3068640","article-title":"DynaSLAM II: Tightly-coupled multi-object tracking and SLAM","volume":"6","author":"Bescos","year":"2021","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Sun, J., Wang, Z., Zhang, S., He, X., Zhao, H., Zhang, G., and Zhou, X. (2022). Onepose: One-shot object pose estimation without cad models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18\u201324 June 2022, IEEE.","DOI":"10.1109\/CVPR52688.2022.00670"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Nguyen, V.N., Groueix, T., Salzmann, M., and Lepetit, V. (2024). Gigapose: Fast and robust novel object pose estimation via one correspondence. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16\u201322 June 2024, IEEE.","DOI":"10.1109\/CVPR52733.2024.00945"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Rublee, E., Rabaud, V., Konolige, K., and Bradski, G. (2011). ORB: An efficient alternative to SIFT or SURF. Proceedings of the 2011 International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV.2011.6126544"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Bay, H., Tuytelaars, T., and Van Gool, L. (2006). Surf: Speeded up robust features. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/11744023_32"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"DeTone, D., Malisiewicz, T., and Rabinovich, A. (2018). Superpoint: Self-supervised interest point detection and description. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA, 18\u201323 June 2018, IEEE.","DOI":"10.1109\/CVPRW.2018.00060"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Dusmanu, M., Rocco, I., Pajdla, T., Pollefeys, M., Sivic, J., Torii, A., and Sattler, T. (2019). D2-net: A trainable cnn for joint description and detection of local features. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15\u201320 June 2019, IEEE.","DOI":"10.1109\/CVPR.2019.00828"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Sun, J., Shen, Z., Wang, Y., Bao, H., and Zhou, X. (2021). LoFTR: Detector-free local feature matching with transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20\u201325 June 2021, IEEE.","DOI":"10.1109\/CVPR46437.2021.00881"},{"key":"ref_14","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems 30 (NIPS 2017), NIPS."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Potje, G., Cadar, F., Araujo, A., Martins, R., and Nascimento, E.R. (2024). Xfeat: Accelerated features for lightweight image matching. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16\u201322 June 2024, IEEE.","DOI":"10.1109\/CVPR52733.2024.00259"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Li, Z., and Snavely, N. (2018). Megadepth: Learning single-view depth prediction from internet photos. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18\u201323 June 2018, IEEE.","DOI":"10.1109\/CVPR.2018.00218"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Wang, M., Lian, Z., N\u00fa\u00f1ez-Andr\u00e9s, M.A., Wang, P., Tian, Y., Yue, Z., and Gu, L. (2024). Robot localization method based on multi-sensor fusion in low-light environment. Electronics, 13.","DOI":"10.3390\/electronics13224346"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Erdei, T.I., Kapusi, T.P., Hajdu, A., and Husi, G. (2024). Image-to-image translation-based deep learning application for object identification in industrial robot systems. Robotics, 13.","DOI":"10.3390\/robotics13060088"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"He, Z., Yang, W., Liu, Y., Zheng, A., Liu, J., Lou, T., and Zhang, J. (2024). Insulator defect detection based on YOLOv8s-SwinT. Information, 15.","DOI":"10.3390\/info15040206"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Nasim, M., Mumtaz, R., Ahmad, M., and Ali, A. (2024). Fabric defect detection in real world manufacturing using deep learning. Information, 15.","DOI":"10.3390\/info15080476"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Huang, X., Wei, D., Chang, Q., Liu, J., and Jing, Q. (2024). Copper Nodule Defect Detection in Industrial Processes Using Deep Learning. Information, 15.","DOI":"10.3390\/info15120802"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Kostusiak, A., and Skrzypczy\u0144ski, P. (2024). Enhancing Visual Odometry with Estimated Scene Depth: Leveraging RGB-D Data with Deep Learning. Electronics, 13.","DOI":"10.3390\/electronics13142755"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Song, J., Jo, H., Jin, Y., and Lee, S.J. (2024). Uncertainty-aware depth network for visual inertial odometry of mobile robots. Sensors, 24.","DOI":"10.3390\/s24206665"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Wang, S., Kannala, J., Pollefeys, M., and Barath, D. (2023). Guiding local feature matching with surface curvature. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France, 1\u20136 October 2023, IEEE.","DOI":"10.1109\/ICCV51070.2023.01648"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Liu, Y., Lai, W., Zhao, Z., Xiong, Y., Zhu, J., Cheng, J., and Xu, Y. (2025). LiftFeat: 3D geometry-aware local feature matching. Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA), IEEE.","DOI":"10.1109\/ICRA55743.2025.11127853"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Ku, J., Harakeh, A., and Waslander, S.L. (2018). In defense of classical image processing: Fast depth completion on the cpu. Proceedings of the 2018 15th Conference on Computer and Robot Vision (CRV), IEEE.","DOI":"10.1109\/CRV.2018.00013"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Dai, A., Chang, A.X., Savva, M., Halber, M., Funkhouser, T., and Nie\u00dfner, M. (2017). Scannet: Richly-annotated 3d reconstructions of indoor scenes. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21\u201326 July 2017, IEEE.","DOI":"10.1109\/CVPR.2017.261"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Balntas, V., Lenc, K., Vedaldi, A., and Mikolajczyk, K. (2017). HPatches: A benchmark and evaluation of handcrafted and learned local descriptors. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21\u201326 July 2017, IEEE.","DOI":"10.1109\/CVPR.2017.410"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Sattler, T., Maddern, W., Toft, C., Torii, A., Hammarstrand, L., Stenborg, E., Safari, D., Okutomi, M., Pollefeys, M., and Sivic, J. (2018). Benchmarking 6dof outdoor visual localization in changing conditions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18\u201323 June 2018, IEEE.","DOI":"10.1109\/CVPR.2018.00897"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Sarlin, P.E., DeTone, D., Malisiewicz, T., and Rabinovich, A. (2020). Superglue: Learning feature matching with graph neural networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13\u201319 June 2020, IEEE.","DOI":"10.1109\/CVPR42600.2020.00499"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Chen, H., Luo, Z., Zhou, L., Tian, Y., Zhen, M., Fang, T., Mckinnon, D., Tsin, Y., and Quan, L. (2022). Aspanformer: Detector-free image matching with adaptive span transformer. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-031-19824-3_2"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Edstedt, J., Athanasiadis, I., Wadenb\u00e4ck, M., and Felsberg, M. (2023). DKM: Dense kernelized feature matching for geometry estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17\u201324 June 2023, IEEE.","DOI":"10.1109\/CVPR52729.2023.01704"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Lindenberger, P., Sarlin, P.E., and Pollefeys, M. (2023). Lightglue: Local feature matching at light speed. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France, 1\u20136 October 2023, IEEE.","DOI":"10.1109\/ICCV51070.2023.01616"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Opanin Gyamfi, E., Qin, Z., Mantebea Danso, J., and Adu-Gyamfi, D. (2024). Hierarchical graph neural network: A lightweight image matching model with enhanced message passing of local and global information in hierarchical graph neural networks. Information, 15.","DOI":"10.3390\/info15100602"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Ying, S., Zhao, J., Li, G., and Dai, J. (2025). LIM: Lightweight image local feature matching. J. Imaging, 11.","DOI":"10.3390\/jimaging11050164"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Chao, Q., Wang, S., and Yu, T. (2024). A lightweight face detector via bi-stream convolutional neural network and vision transformer. Information, 15.","DOI":"10.3390\/info15050290"},{"key":"ref_38","first-page":"4826","article-title":"Working Hard to Know Your Neighbor\u2019s Margins: Local Descriptor Learning Loss","volume":"Volume 30","author":"Mishchuk","year":"2017","journal-title":"Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"ref_39","unstructured":"Tang, J., Kim, H., Guizilini, V., Pillai, S., and Ambrus, R. (2019). Neural outlier rejection for self-supervised keypoint learning. arXiv."},{"key":"ref_40","unstructured":"Zhang, Y., and Zhao, X. (2023). Searching from area to point: A hierarchical framework for semantic-geometric combined feature matching. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., and Lo, W.Y. (2023). Segment anything. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France, 1\u20136 October 2023, IEEE.","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"ref_42","unstructured":"He, X., Yu, H., Peng, S., Tan, D., Shen, Z., Bao, H., and Zhou, X. (2025). Matchanything: Universal cross-modality image matching with large-scale pre-training. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Gleize, P., Wang, W., and Feiszli, M. (2023). Silk: Simple learned keypoints. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France, 1\u20136 October 2023, IEEE.","DOI":"10.1109\/ICCV51070.2023.02056"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Yang, L., Kang, B., Huang, Z., Xu, X., Feng, J., and Zhao, H. (2024). Depth anything: Unleashing the power of large-scale unlabeled data. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16\u201322 June 2024, IEEE.","DOI":"10.1109\/CVPR52733.2024.00987"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Keselman, L., Iselin Woodfill, J., Grunnet-Jepsen, A., and Bhowmik, A. (2017). Intel realsense stereoscopic depth cameras. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Honolulu, HI, USA, 21\u201326 July 2017, IEEE.","DOI":"10.1109\/CVPRW.2017.167"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"184","DOI":"10.1109\/TRO.2018.2875382","article-title":"Canny-vo: Visual odometry with rgb-d cameras based on geometric 3-d\u20132-d edge alignment","volume":"35","author":"Zhou","year":"2018","journal-title":"IEEE Trans. Robot."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Porzi, L., Penate-Sanchez, A., Ricci, E., and Moreno-Noguer, F. (2017). Depth-aware convolutional neural networks for accurate 3D pose estimation in RGB-D images. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE.","DOI":"10.1109\/IROS.2017.8206469"},{"key":"ref_48","unstructured":"Jiang, J., Zheng, L., Luo, F., and Zhang, Z. (2018). Rednet: Residual encoder-decoder network for indoor rgb-d semantic segmentation. arXiv."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"3101","DOI":"10.1109\/TMM.2022.3155927","article-title":"Alike: Accurate and lightweight keypoint detection and descriptor extraction","volume":"25","author":"Zhao","year":"2022","journal-title":"IEEE Trans. Multimed."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Barath, D., Noskova, J., Ivashechkin, M., and Matas, J. (2020). MAGSAC++, a fast, reliable and accurate robust estimator. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13\u201319 June 2020, IEEE.","DOI":"10.1109\/CVPR42600.2020.00138"},{"key":"ref_51","unstructured":"Wang, Z., Chen, S., Yang, L., Wang, J., Zhang, Z., Zhao, H., and Zhao, Z. (2026, January 23\u201327). Depth Anything with Any Prior. Proceedings of the International Conference on Learning Representations, Rio de Janeiro, Brazil."}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/6\/253\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T04:45:08Z","timestamp":1781153108000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/6\/253"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,7]]},"references-count":51,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2026,6]]}},"alternative-id":["jimaging12060253"],"URL":"https:\/\/doi.org\/10.3390\/jimaging12060253","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,7]]}}}