{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,29]],"date-time":"2025-10-29T06:21:39Z","timestamp":1761718899376,"version":"build-2065373602"},"reference-count":56,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2020,2,20]],"date-time":"2020-02-20T00:00:00Z","timestamp":1582156800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Detecting similarities between image patches and measuring their mutual displacement are important parts in the registration of multimodal remote sensing (RS) images. Deep learning approaches advance the discriminative power of learned similarity measures (SM). However, their ability to find the best spatial alignment of the compared patches is often ignored. We propose to unify the patch discrimination and localization problems by assuming that the more accurately two patches can be aligned, the more similar they are. The uncertainty or confidence in the localization of a patch pair serves as a similarity measure of these patches. We train a two-channel patch matching convolutional neural network (CNN), called DLSM, to solve a regression problem with uncertainty. This CNN inputs two multimodal patches, and outputs a prediction of the translation vector between the input patches as well as the uncertainty of this prediction in the form of an error covariance matrix of the translation vector. The proposed patch matching CNN predicts a normal two-dimensional distribution of the translation vector rather than a simple value of it. The determinant of the covariance matrix is used as a measure of uncertainty in the matching of patches and also as a measure of similarity between patches. For training, we used the Siamese architecture with three towers. During training, the input of two towers is the same pair of multimodal patches but shifted by a random translation; the last tower is fed by a pair of dissimilar patches. Experiments performed on a large base of real RS images show that the proposed DLSM has both a higher discriminative power and a more precise localization compared to existing hand-crafted SMs and SMs trained with conventional losses. Unlike existing SMs, DLSM correctly predicts translation error distribution ellipse for different modalities, noise level, isotropic, and anisotropic structures.<\/jats:p>","DOI":"10.3390\/rs12040703","type":"journal-article","created":{"date-parts":[[2020,2,21]],"date-time":"2020-02-21T08:59:47Z","timestamp":1582275587000},"page":"703","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["Efficient Discrimination and Localization of Multimodal Remote Sensing Images Using CNN-Based Prediction of Localization Uncertainty"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4485-9207","authenticated-orcid":false,"given":"Mykhail","family":"Uss","sequence":"first","affiliation":[{"name":"Department of Information-Communication Technologies, National Aerospace University, 61070 Kharkiv, Ukraine"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1920-2847","authenticated-orcid":false,"given":"Benoit","family":"Vozel","sequence":"additional","affiliation":[{"name":"IETR UMR CNRS 6164, University of Rennes 1, Enssat, 22305 Lannion, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1443-9685","authenticated-orcid":false,"given":"Vladimir","family":"Lukin","sequence":"additional","affiliation":[{"name":"Department of Information-Communication Technologies, National Aerospace University, 61070 Kharkiv, Ukraine"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kacem","family":"Chehdi","sequence":"additional","affiliation":[{"name":"IETR UMR CNRS 6164, University of Rennes 1, Enssat, 22305 Lannion, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,2,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"6587","DOI":"10.1109\/TGRS.2016.2587321","article-title":"Multimodal remote sensing images registration with accuracy estimation at local and global scales","volume":"54","author":"Uss","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1706","DOI":"10.1109\/TIP.2014.2307478","article-title":"Robust Point Matching via Vector Field Consensus","volume":"23","author":"Ma","year":"2014","journal-title":"IEEE Trans. Image Process."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Le Moigne, J., Netanyahu, N.S., and Eastman, R.D. (2011). Image Registration for Remote Sensing, Cambridge University Press.","DOI":"10.1017\/CBO9780511777684"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"En, S., Lechervy, A., and Jurie, F. (2018, January 7\u201310). TS-NET: Combining Modality Specific and Common Features for Multimodal Patch Matching. Proceedings of the 2018 25th IEEE International Conference on Image Processing (ICIP), Athens, Greece.","DOI":"10.1109\/ICIP.2018.8451804"},{"key":"ref_5","unstructured":"Aguilera, C.A., Aguilera, F.J., Sappa, A.D., Aguilera, C., and Toledo, R. (July, January 26). Learning cross-spectral similarity measures with deep convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Las Vegas, NV, USA."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Aguilera, C.A., Sappa, A.D., Aguilera, C., and Toledo, R. (2017). Cross-Spectral Local Descriptors via Quadruplet Network. Sensors, 17.","DOI":"10.20944\/preprints201703.0061.v1"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Goshtasby, A., and Le Moign, J. (2012). Image Registration: Principles, Tools and Methods, Springer.","DOI":"10.1007\/978-1-4471-2458-0_11"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"977","DOI":"10.1016\/S0262-8856(03)00137-9","article-title":"Image registration methods: A survey","volume":"21","author":"Flusser","year":"2003","journal-title":"Image Vis. Comput."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Roche, A., Malandain, G., Pennec, X., and Ayache, N. (1998). The correlation ratio as a new similarity measure for multimodal image registration. Medical Image Computing and Computer-Assisted Interventation\u2014MICCAI\u201998, Springer.","DOI":"10.1007\/BFb0056301"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"188","DOI":"10.1109\/83.988953","article-title":"Extension of phase correlation to subpixel registration","volume":"11","author":"Foroosh","year":"2002","journal-title":"IEEE Trans. Image Process."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"939","DOI":"10.1109\/TGRS.2009.2034842","article-title":"Mutual-Information-Based Registration of TerraSAR-X and Ikonos Imagery in Urban Areas","volume":"48","author":"Suri","year":"2010","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Uss, M., Vozel, B., Lukin, V., and Chehdi, K. (2016). Statistical power of intensity- and feature-based similarity measures for registration of multimodal remote sensing images. Proc. SPIE, 10004.","DOI":"10.1117\/12.2240895"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"2941","DOI":"10.1109\/TGRS.2017.2656380","article-title":"Robust Registration of Multimodal Remote Sensing Images Based on Structural Similarity","volume":"55","author":"Ye","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive Image Features from Scale-Invariant Keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"243","DOI":"10.1080\/19479832.2010.495322","article-title":"Modifications in the SIFT operator for effective SAR image matching","volume":"1","author":"Suri","year":"2010","journal-title":"Int. J. Image Data Fusion"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1423","DOI":"10.1016\/j.media.2012.05.008","article-title":"MIND: Modality independent neighbourhood descriptor for multi-modal deformable registration","volume":"16","author":"Heinrich","year":"2012","journal-title":"Med. Image Anal."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Yi, K.M., Trulls, E., Lepetit, V., and Fua, P. (2016). Lift: Learned invariant feature transform. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46466-4_28"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zagoruyko, S., and Komodakis, N. (2015, January 7\u201312). Learning to compare image patches via convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299064"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zeng, A., Song, S., Nie\u00dfner, M., Fisher, M., Xiao, J., and Funkhouser, T. (2017, January 21\u201326). 3dmatch: Learning local geometric descriptors from rgb-d reconstructions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.29"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Schonberger, J.L., Hardmeier, H., Sattler, T., and Pollefeys, M. (2017, January 21\u201326). Comparative evaluation of hand-crafted and learned local features. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.736"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"378","DOI":"10.1016\/j.neuroimage.2017.07.008","article-title":"Quicksilver: Fast predictive image registration\u2014A deep learning approach","volume":"158","author":"Yang","year":"2017","journal-title":"NeuroImage"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Balakrishnan, G., Zhao, A., Sabuncu, M.R., Guttag, J., and Dalca, A.V. (2018, January 18\u201322). An unsupervised learning model for deformable medical image registration. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00964"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Altwaijry, H., Trulls, E., Hays, J., Fua, P., and Belongie, S. (2016, January 27\u201330). Learning to Match Aerial Images with Deep Attentive Architectures. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.385"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Merkle, N., Luo, W., Auer, S., M\u00fcller, R., and Urtasun, R. (2017). Exploiting Deep Matching and SAR Data for the Geo-Localization Accuracy Improvement of Optical Satellite Images. Remote Sens., 9.","DOI":"10.3390\/rs9060586"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"3333","DOI":"10.1109\/TGRS.2013.2272559","article-title":"A precise lower bound on image subpixel registration accuracy","volume":"52","author":"Uss","year":"2013","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"138","DOI":"10.1006\/cviu.1999.0832","article-title":"MLESAC: A New Robust Estimator with Application to Estimating Image Geometry","volume":"78","author":"Torr","year":"2000","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Tian, Y., Fan, B., and Wu, F. (2017, January 21\u201326). L2-net: Deep learning of discriminative patch descriptor in euclidean space. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.649"},{"key":"ref_28","unstructured":"Han, X., Leung, T., Jia, Y., Sukthankar, R., and Berg, A.C. (2015, January 7\u201312). Matchnet: Unifying feature and metric learning for patch-based matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_29","first-page":"2287","article-title":"Stereo matching by training a convolutional neural network to compare image patches","volume":"17","author":"LeCun","year":"2016","journal-title":"J. Mach. Learn. Res."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Simo-Serra, E., Trulls, E., Ferraz, L., Kokkinos, I., Fua, P., and Moreno-Noguer, F. (2015, January 7\u201312). Discriminative learning of deep convolutional feature point descriptors. Proceedings of the IEEE International Conference on Computer Vision, Boston, MA, USA.","DOI":"10.1109\/ICCV.2015.22"},{"key":"ref_31","unstructured":"Hadsell, R., Chopra, S., and LeCun, Y. (2006, January 17\u201322). Dimensionality reduction by learning an invariant mapping. Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201906), New York, NY, USA."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Georgakis, G., Karanam, S., Wu, Z., Ernst, J., and Ko\u0161eck\u00e1, J. (2018, January 18\u201322). End-to-end learning of keypoint detector and descriptor for pose invariant 3D matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00210"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Mobahi, H., Collobert, R., and Weston, J. (2009). Deep learning from temporal coherence in video. Proceedings of the 26th Annual International Conference on Machine Learning, ACM.","DOI":"10.1145\/1553374.1553469"},{"key":"ref_34","unstructured":"Balntas, V., Johns, E., Tang, L., and Mikolajczyk, K. (2016). PN-Net: Conjoined triple deep network for learning local image descriptors. arXiv."},{"key":"ref_35","unstructured":"Choy, C.B., Gwak, J., Savarese, S., and Chandraker, M. (2016, January 5\u201310). Universal correspondence network. Proceedings of the Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Deng, H., Birdal, T., and Ilic, S. (2018, January 18\u201322). Ppfnet: Global context aware local features for robust 3d point matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00028"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Hoffer, E., and Ailon, N. (2015). Deep Metric Learning Using Triplet Network. International Workshop on Similarity-Based Pattern Recognition, Springer International Publishing.","DOI":"10.1007\/978-3-319-24261-3_7"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Khoury, M., Zhou, Q.Y., and Koltun, V. (2017, January 22\u201329). Learning compact geometric features. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.26"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Masci, J., Migliore, D., Bronstein, M.M., and Schmidhuber, J. (2014). Descriptor learning for omnidirectional image matching. Registration and Recognition in Images and Videos, Springer.","DOI":"10.1007\/978-3-642-44907-9_3"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Wang, J., Song, Y., Leung, T., Rosenberg, C., Wang, J., Philbin, J., Chen, B., and Wu, Y. (2014, January 23\u201328). Learning fine-grained image similarity with deep ranking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.180"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Su\u00e1rez, P.L., Sappa, A.D., and Vintimilla, B.X. (2017, January 24\u201326). Cross-spectral image patch similarity using convolutional neural network. Proceedings of the 2017 IEEE International Workshop of Electronics, Control, Measurement, Signals and their Application to Mechatronics (ECMSM), Donostia-San Sebastian, Spain.","DOI":"10.1109\/ECMSM.2017.7945888"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"He, H., Chen, M., Chen, T., and Li, D. (2018). Matching of Remote Sensing Images with Complex Background Variations via Siamese Convolutional Neural Network. Remote Sens., 10.","DOI":"10.3390\/rs10020355"},{"key":"ref_43","unstructured":"Kumar, B., Carneiro, G., and Reid, I. (2016, January 27\u201330). Learning local image descriptors with deep siamese and triplet convolutional networks by minimising global loss functions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"38544","DOI":"10.1109\/ACCESS.2018.2853100","article-title":"Multi-Temporal Remote Sensing Image Registration Using Deep Convolutional Features","volume":"6","author":"Yang","year":"2018","journal-title":"IEEE Access"},{"key":"ref_45","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L., Kai, L., and Li, F.F. (2009, January 20\u201325). ImageNet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_47","unstructured":"Luo, W., Schwing, A.G., and Urtasun, R. (1, January 26). Efficient deep learning for stereo matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_48","unstructured":"Dosovitskiy, A., Springenberg, J.T., Riedmiller, M., and Brox, T. (2014, January 8\u201313). Discriminative unsupervised feature learning with convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"9059","DOI":"10.1109\/TGRS.2019.2924684","article-title":"Fast and Robust Matching for Multimodal Remote Sensing Image Registration","volume":"57","author":"Ye","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"2589","DOI":"10.1109\/TGRS.2011.2109389","article-title":"Automatic Image Registration Through Image Segmentation and SIFT","volume":"49","author":"Goncalves","year":"2011","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"861","DOI":"10.1016\/j.patrec.2005.10.010","article-title":"An introduction to ROC analysis","volume":"27","author":"Fawcett","year":"2006","journal-title":"Pattern Recognit. Lett."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Huber, P.J. (2011). Robust Statistics, Springer.","DOI":"10.1007\/978-3-642-04898-2_594"},{"key":"ref_53","unstructured":"Gurevich, P., and Stuke, H. (2017). Learning uncertainty in regression tasks by deep neural networks. arXiv."},{"key":"ref_54","unstructured":"Kendall, A., and Gal, Y. (2017, January 4\u20139). What uncertainties do we need in bayesian deep learning for computer vision?. Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"809","DOI":"10.1109\/42.876307","article-title":"Image registration by maximization of combined mutual information and gradient information","volume":"19","author":"Pluim","year":"2000","journal-title":"IEEE Trans. Med. Imag."},{"key":"ref_56","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/4\/703\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T08:59:33Z","timestamp":1760173173000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/4\/703"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,2,20]]},"references-count":56,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2020,2]]}},"alternative-id":["rs12040703"],"URL":"https:\/\/doi.org\/10.3390\/rs12040703","relation":{},"ISSN":["2072-4292"],"issn-type":[{"type":"electronic","value":"2072-4292"}],"subject":[],"published":{"date-parts":[[2020,2,20]]}}}