{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T03:24:17Z","timestamp":1784690657286,"version":"3.55.0"},"reference-count":44,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2025,9,24]],"date-time":"2025-09-24T00:00:00Z","timestamp":1758672000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"NSF","award":["2142428"],"award-info":[{"award-number":["2142428"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Image matching plays a critical role in a wide range of computer vision applications, including object recognition, 3D reconstruction, aiming-point and six-degree-of-freedom detection for aiming devices, and video surveillance. Over the past three decades, image-matching algorithms and techniques have evolved significantly, from handcrafted feature extraction algorithms to modern approaches powered by deep learning neural networks and attention mechanisms. This paper provides a comprehensive review of image-matching techniques, aiming to offer researchers valuable insights into the evolving landscape of this field. It traces the historical development of feature-based methods and examines the transition to neural network-based approaches that leverage large-scale data and learned representations. Additionally, this paper discusses the current state of the field, highlighting key algorithms, benchmarks, and real-world applications. Furthermore, this study introduces some recent contributions to this area and outlines promising directions for future research, including H-matrix optimization, LoFTR model speedup, and performance improvements. It also identifies persistent challenges such as robustness to viewpoint and illumination changes, scalability, and matching under extreme conditions. Finally, this paper summarizes future trends for research and development in this field.<\/jats:p>","DOI":"10.3390\/jimaging11100329","type":"journal-article","created":{"date-parts":[[2025,9,24]],"date-time":"2025-09-24T10:39:42Z","timestamp":1758710382000},"page":"329","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Image Matching: Foundations, State of the Art, and Future Directions"],"prefix":"10.3390","volume":"11","author":[{"given":"Ming","family":"Yang","sequence":"first","affiliation":[{"name":"Department of Information Technology, Kennesaw State University, Kennesaw, GA 30060, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6134-5393","authenticated-orcid":false,"given":"Rui","family":"Wu","sequence":"additional","affiliation":[{"name":"Department of Information Technology, Kennesaw State University, Kennesaw, GA 30060, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yunxuan","family":"Yang","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, Columbia University, New York, NY 10027, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Liang","family":"Tao","sequence":"additional","affiliation":[{"name":"Department of Information Technology, Kennesaw State University, Kennesaw, GA 30060, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5832-2622","authenticated-orcid":false,"given":"Yifan","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Missouri State University, Springfield, MO 65897, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4436-3474","authenticated-orcid":false,"given":"Yixin","family":"Xie","sequence":"additional","affiliation":[{"name":"Department of Information Technology, Kennesaw State University, Kennesaw, GA 30060, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gnana Prakash Reddy Donthi","family":"Reddy","sequence":"additional","affiliation":[{"name":"Department of Information Technology, Kennesaw State University, Kennesaw, GA 30060, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,9,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1109\/MRA.2006.1678144","article-title":"Simultaneous localization and mapping: Part I","volume":"13","author":"Bailey","year":"2006","journal-title":"IEEE Robot. Autom. Mag."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"371","DOI":"10.1243\/PIME_PROC_1965_180_029_02","article-title":"A platform with six degrees of freedom","volume":"180","author":"Stewart","year":"1965","journal-title":"Proc. Inst. Mech. Eng."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"751","DOI":"10.1007\/s11760-023-02802-w","article-title":"Self-adaptive SURF for image-to-video matching","volume":"18","author":"Yang","year":"2024","journal-title":"Signal Image Video Process."},{"key":"ref_4","unstructured":"Schmidhuber, J. (2012, January 16\u201321). Multi-column deep neural networks for image classification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Providence, RI, USA."},{"key":"ref_5","first-page":"1097","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_6","first-page":"128","article-title":"Deep learning vs. traditional computer vision","volume":"Volume 943","author":"Campbell","year":"2020","journal-title":"Advances in Computer Vision. CVC 2019"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"7068349","DOI":"10.1155\/2018\/7068349","article-title":"Deep learning for computer vision: A brief review","volume":"2018","author":"Voulodimos","year":"2018","journal-title":"Comput. Intell. Neurosci."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Hassaballah, M., and Awad, A.I. (2020). Deep Learning in Computer Vision: Principles and Applications, CRC Press.","DOI":"10.1201\/9781351003827"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Lowe, D.G. (1999, January 20\u201325). Object recognition from local scale-invariant features. Proceedings of the Seventh IEEE International Conference on Computer Vision, Kerkyra, Greece.","DOI":"10.1109\/ICCV.1999.790410"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"346","DOI":"10.1016\/j.cviu.2007.09.014","article-title":"Speeded-up robust features (SURF)","volume":"110","author":"Bay","year":"2008","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Rublee, E., Rabaud, V., Konolige, K., and Bradski, G. (2011, January 6\u201313). ORB: An efficient alternative to SIFT or SURF. Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126544"},{"key":"ref_12","unstructured":"Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (2017, January 4\u20139). Attention is All you Need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"102344","DOI":"10.1016\/j.inffus.2024.102344","article-title":"Local feature matching using deep learning: A survey","volume":"107","author":"Xu","year":"2024","journal-title":"Inf. Fusion"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1007\/s11263-020-01359-2","article-title":"Image matching from handcrafted to deep features: A survey","volume":"129","author":"Ma","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1385","DOI":"10.1049\/ipr2.13032","article-title":"A survey of feature matching methods","volume":"18","author":"Huang","year":"2024","journal-title":"IET Image Process."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Brunelli, R. (2009). Template Matching Techniques in Computer Vision: Theory and Practice, John Wiley & Sons.","DOI":"10.1002\/9780470744055"},{"key":"ref_17","unstructured":"Cantzler, H. (1981). Random Sample Consensus (Ransac), Institute for Perception, Action and Behaviour, Division of Informatics, University of Edinburgh."},{"key":"ref_18","unstructured":"Zhang, Z. (1999, January 20\u201325). Flexible camera calibration by viewing a plane from unknown orientations. Proceedings of the Seventh IEEE International Conference on Computer Vision, Kerkyra, Greece."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"DeTone, D., Malisiewicz, T., and Rabinovich, A. (2018, January 18\u201322). Superpoint: Self-supervised interest point detection and description. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00060"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Sarlin, P.E., DeTone, D., Malisiewicz, T., and Rabinovich, A. (2020, January 13\u201319). Superglue: Learning feature matching with graph neural networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00499"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Sun, J., Shen, Z., Wang, Y., Bao, H., and Zhou, X. (2021, January 20\u201325). LoFTR: Detector-Free Local Feature Matching with Transformers. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00881"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"876","DOI":"10.1214\/aoms\/1177703591","article-title":"A relationship between arbitrary positive matrices and doubly stochastic matrices","volume":"35","author":"Sinkhorn","year":"1964","journal-title":"Ann. Math. Stat."},{"key":"ref_23","first-page":"17346","article-title":"Dual-resolution correspondence networks","volume":"33","author":"Li","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wang, Y., He, X., Peng, S., Tan, D., and Zhou, X. (2024, January 16\u201322). Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like Speed. Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02047"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Wu, R., Dascalu, S.M., and Harris, F.C. (2024). Sparse transformer with local and seasonal adaptation for multivariate time series forecasting. Sci. Rep., 14.","DOI":"10.1038\/s41598-024-66886-1"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_29","unstructured":"Nam, J., Lee, G., Kim, S., Kim, H., Cho, H., Kim, S., and Kim, S. (2024, January 7\u201311). Diffusion Model for Dense Matching. Proceedings of the Twelfth International Conference on Learning Representations, Vienna, Austria."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Youssef, A., and Vasconcelos, F. (2025). NeRF-Supervised Feature Point Detection and Description. Computer Vision\u2014ECCV 2024 Workshops: Milan, Italy, September 29\u2013October 4, 2024, Proceedings, Part XXIII, Springer.","DOI":"10.1007\/978-3-031-91989-3_7"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1145\/3503250","article-title":"NeRF: Representing scenes as neural radiance fields for view synthesis","volume":"65","author":"Mildenhall","year":"2021","journal-title":"Commun. ACM"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Xue, F., Budvytis, I., and Cipolla, R. (2023, January 17\u201324). SFD2: Semantic-Guided Feature Detection and Description. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00504"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Strecha, C., Lindner, A., Ali, K., and Fua, P. (2009, January 9\u201311). Training for task specific keypoint detection. Proceedings of the Pattern Recognition: 31st DAGM Symposium, Jena, Germany.","DOI":"10.1007\/978-3-642-03798-6_16"},{"key":"ref_34","first-page":"12414","article-title":"R2d2: Reliable and repeatable detector and descriptor","volume":"32","author":"Revaud","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Edstedt, J., Sun, Q., B\u00f6kman, G., Wadenb\u00e4ck, M., and Felsberg, M. (2024, January 16\u201324). RoMa: Robust Dense Feature Matching. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01871"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Barroso-Laguna, A., Munukutla, S., Prisacariu, V.A., and Brachmann, E. (2024, January 16\u201324). Matching 2D Images in 3D: Metric Relative Pose from Metric Correspondences. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00464"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Arsalan Soltani, A., Huang, H., Wu, J., Kulkarni, T.D., and Tenenbaum, J.B. (2017, January 21\u201326). Synthesizing 3d shapes via modeling multi-view depth maps and silhouettes with deep generative networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.269"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lin, C.H., Ma, W.C., Torralba, A., and Lucey, S. (2021, January 10\u201317). Barf: Bundle-adjusting neural radiance fields. Proceedings of the IEEE\/CVF international Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00569"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Lindenberger, P., Sarlin, P.E., and Pollefeys, M. (2023, January 1\u20136). Lightglue: Local feature matching at light speed. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.01616"},{"key":"ref_40","unstructured":"Mehta, S., and Rastegari, M. (2021). Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Wu, K., Zhang, J., Peng, H., Liu, M., Xiao, B., Fu, J., and Yuan, L. (2022, January 23\u201327). Tinyvit: Fast pretraining distillation for small vision transformers. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19803-8_5"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Potje, G., Cadar, F., Araujo, A., Martins, R., and Nascimento, E.R. (2024, January 16\u201322). XFeat: Accelerated Features for Lightweight Image Matching. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00259"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Teed, Z., and Deng, J. (2020). Raft: Recurrent all-pairs field transforms for optical flow. Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part II, Springer.","DOI":"10.1007\/978-3-030-58536-5_24"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Jiang, W., Trulls, E., Hosang, J., Tagliasacchi, A., and Yi, K.M. (2021, January 10\u201317). Cotr: Correspondence transformer for matching across images. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00615"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/11\/10\/329\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:48:49Z","timestamp":1760035729000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/11\/10\/329"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,24]]},"references-count":44,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2025,10]]}},"alternative-id":["jimaging11100329"],"URL":"https:\/\/doi.org\/10.3390\/jimaging11100329","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,24]]}}}