{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,23]],"date-time":"2026-03-23T10:32:16Z","timestamp":1774261936136,"version":"3.50.1"},"reference-count":48,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2021,1,5]],"date-time":"2021-01-05T00:00:00Z","timestamp":1609804800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The application of deep learning is blooming in the field of visual place recognition, which plays a critical role in visual Simultaneous Localization and Mapping (vSLAM) applications. The use of convolutional neural networks (CNNs) achieve better performance than handcrafted feature descriptors. However, visual place recognition is still a challenging task due to two major problems, i.e., perceptual aliasing and perceptual variability. Therefore, designing a customized distance learning method to express the intrinsic distance constraints in the large-scale vSLAM scenarios is of great importance. Traditional deep distance learning methods usually use the triplet loss which requires the mining of anchor images. This may, however, result in very tedious inefficient training and anomalous distance relationships. In this paper, a novel deep distance learning framework for visual place recognition is proposed. Through in-depth analysis of the multiple constraints of the distance relationship in the visual place recognition problem, the multi-constraint loss function is proposed to optimize the distance constraint relationships in the Euclidean space. The new framework can support any kind of CNN such as AlexNet, VGGNet and other user-defined networks to extract more distinguishing features. We have compared the results with the traditional deep distance learning method, and the results show that the proposed method can improve the performance by 19\u201328%. Additionally, compared to some contemporary visual place recognition techniques, the proposed method can improve the performance by 40%\/36% and 27%\/24% in average on VGGNet\/AlexNet using the New College and the TUM datasets, respectively. It\u2019s verified the method is capable to handle appearance changes in complex environments.<\/jats:p>","DOI":"10.3390\/s21010310","type":"journal-article","created":{"date-parts":[[2021,1,5]],"date-time":"2021-01-05T10:35:12Z","timestamp":1609842912000},"page":"310","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":12,"title":["Towards a Robust Visual Place Recognition in Large-Scale vSLAM Scenarios Based on a Deep Distance Learning"],"prefix":"10.3390","volume":"21","author":[{"given":"Liang","family":"Chen","sequence":"first","affiliation":[{"name":"School of Mechanical and Electric Engineering, Soochow University, Suzhou 215131, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sheng","family":"Jin","sequence":"additional","affiliation":[{"name":"School of Mechanical and Electric Engineering, Soochow University, Suzhou 215131, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhoujun","family":"Xia","sequence":"additional","affiliation":[{"name":"School of Mechanical and Electric Engineering, Soochow University, Suzhou 215131, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,1,5]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TRO.2015.2496823","article-title":"Visual place recognition: A survey","volume":"32","author":"Lowry","year":"2016","journal-title":"IEEE Trans. Rob."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1255","DOI":"10.1109\/TRO.2017.2705103","article-title":"ORB-SLAM2: An Open-Source SLAM System for Monocular, Stereo, and RGB-D Cameras","volume":"33","year":"2017","journal-title":"IEEE Trans. Robot."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1309","DOI":"10.1109\/TRO.2016.2624754","article-title":"Past, present, and future of simultaneous localization and mapping: Towards the robust-perception age","volume":"32","author":"Cadena","year":"2016","journal-title":"IEEE Trans. Rob."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"495","DOI":"10.1007\/s10846-017-0718-z","article-title":"Fast and Effective Loop Closure Detection to Improve SLAM Performance","volume":"93","author":"Guclu","year":"2019","journal-title":"J. Intell Robot. Syst."},{"key":"ref_5","unstructured":"Zaffar, M. (2020). Visual Place Recognition for Autonomous Robots. [Master\u2019s Thesis, University of Essex]."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive Image Features from Scale-Invariant Keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vision."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"346","DOI":"10.1016\/j.cviu.2007.09.014","article-title":"Speeded-up robust features (SURF)","volume":"110","author":"Bay","year":"2008","journal-title":"Comput. Vision Image Underst."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Rublee, E., Rabaud, V., Konolige, K., and Bradski, G. (2011, January 6\u201313). ORB: An efficient alternative to SIFT or SURF. Proceedings of the IEEE International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126544"},{"key":"ref_9","unstructured":"Dalal, N., and Triggs, B. (2005, January 20\u201326). Histograms of Oriented Gradients for Human Detection. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Diego, CA, USA."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Sivic, J., and Zisserman, A. (2003, January 13\u201316). Video Google: A Text Retrieval Approach to Object Matching in Videos. Proceedings of the IEEE International Conference on Computer Vision, Beijing, China.","DOI":"10.1109\/ICCV.2003.1238663"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1188","DOI":"10.1109\/TRO.2012.2197158","article-title":"Bags of binary words for fast place recognition in image sequences","volume":"28","author":"Tardos","year":"2012","journal-title":"IEEE Trans. Rob."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Liu, H., Wang, R., Shan, S., and Chen, X. (2016, January 27\u201330). Deep Supervised Hashing for Fast Image Retrieval. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.227"},{"key":"ref_13","unstructured":"Krizhevsky, A., and Hinton, G.E. (2012, January 3\u20136). ImageNet Classification with Deep Convolutional Neural Networks. Proceedings of the Neural Information Processing Systems, Lake Tahoe, CA, USA."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Babenko, A., Slesarev, A., Chigorin, A., and Lempitsky, V. (2014, January 6\u201312). Neural Codes for Image Retrieval. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10590-1_38"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Chatfield, K., Simonyan, K., Vedaldi, A., and Zisserman, A. (2014, January 1\u20135). Return of the Devil in the Details: Delving Deep into Convolutional Nets. Proceedings of the British Machine Vision Conference, Nottingham, UK.","DOI":"10.5244\/C.28.6"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Wan, J., Wang, D., Chu Hong Hoi, S., Wu, P., Zhu, J., Zhang, Y., and Li, J. (2014, January 3\u20137). Deep Learning for Content-Based Image Retrieval: A Comprehensive Study. Proceedings of the 22nd ACM International Conference on Multimedia, Orlando, FL, USA.","DOI":"10.1145\/2647868.2654948"},{"key":"ref_17","unstructured":"S\u00fcnderhauf, N., Shirazi, S., Dayoub, F., and Upcroft, B. (October, January 28). On the performance of ConvNet features for place recognition. Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems, Hamburg, Germany."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Li, F.-F. (2009, January 20\u201325). ImageNet: A Large-Scale Hierarchical Image Database. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Xia, Y., Li, J., Qi, L., and Fan, H. (2016, January 24\u201329). Loop closure detection for visual SLAM using PCANet features. Proceedings of the International Joint Conference on Neural Networks, Vancouver, BC, Canada.","DOI":"10.1109\/IJCNN.2016.7727481"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Chen, Z., Jacobson, A., S\u00fcnderhauf, N., Upcroft, B., Liu, L., Shen, C., Reid, I., and Milford, M. (June, January 29). Deep learning features at scale for visual place recognition. Proceedings of the 2017 IEEE International Conference on Robotics and Automation, Singapore.","DOI":"10.1109\/ICRA.2017.7989366"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"12175","DOI":"10.1109\/JSEN.2019.2937740","article-title":"Point-cloud-based place recognition using CNN feature extraction","volume":"19","author":"Sun","year":"2019","journal-title":"IEEE Sens. J."},{"key":"ref_22","unstructured":"Camara, L.G., G\u00e4bert, C., and P\u0159eu\u010dil, L. (August, January 31). Highly Robust Visual Place Recognition through Spatial Matching of CNN Features. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Paris, France."},{"key":"ref_23","unstructured":"Karen, S., and Andrew, Z. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_24","unstructured":"Tolias, G., Sicre, R., and Jegou, H. (2016, January 2\u20134). Particular object retrieval with integral max-pooling of CNN activations. Proceedings of the International Conference on Learning Representations, San Juan, Puerto Rico."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"561","DOI":"10.1109\/TRO.2019.2956352","article-title":"A holistic visual place recognition approach using lightweight cnns for significant viewpoint and appearance changes","volume":"36","author":"Khaliq","year":"2019","journal-title":"IEEE Trans. Robot."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Jegou, H., Douze, M., Schmid, C., and Perez, P. (2010, January 13\u201318). Aggregating local descriptors into a compact image representation. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5540039"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zitnick, C., and Dollar, P. (2014, January 6\u201312). Edge boxes: Locating object proposals from edges. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_26"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Sunderhauf, N., Shirazi, S., Jacobson, A., Dayoub, F., Pepperell, E., Upcroft, B., and Milford, M. (2015, January 13\u201317). Place recognition with ConvNet landmarks: Viewpoint-robust, condition-robust, training-free. Proceedings of the Robotics: Science and Systems, Rome, Italy.","DOI":"10.15607\/RSS.2015.XI.022"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Milford, M.J., and Wyeth, G.F. (2012, January 14\u201318). SeqSLAM: Visual route-based navigation for sunny summer days and stormy winter nights. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Saint Paul, MN, USA.","DOI":"10.1109\/ICRA.2012.6224623"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1016\/j.robot.2018.10.014","article-title":"SeqSLAM++: View-based robot localization and navigation","volume":"112","author":"Oishi","year":"2019","journal-title":"Robot. Auton. Syst."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Johns, E., and Yang, G.Z. (2013, January 6\u201310). Feature co-occurrence maps: Appearance-based localisation throughout the day. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Karlsruhe, Germany.","DOI":"10.1109\/ICRA.2013.6631024"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"261","DOI":"10.1007\/s11263-006-0020-1","article-title":"Detecting loop closure with scene sequences","volume":"74","author":"Ho","year":"2007","journal-title":"Int. J. Comput. Vision."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1016\/j.patcog.2017.10.013","article-title":"Recent advances in convolutional neural networks","volume":"77","author":"Gu","year":"2018","journal-title":"Pattern Recognit."},{"key":"ref_34","unstructured":"Hermans, A., Beyer, L., and Leibe, B. (2017). In defense of the triplet loss for person re-identification. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"3492","DOI":"10.1109\/TIP.2017.2700762","article-title":"End-to-end comparative attention networks for person re-identification","volume":"26","author":"Liu","year":"2017","journal-title":"IEEE Trans. Image Process."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Xie, S., Pan, C., Peng, Y., Liu, K., and Ying, S. (2020). Large-Scale Place Recognition Based on Camera-LiDAR Fused Descriptor. Sensors, 20.","DOI":"10.3390\/s20102870"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Martini, D., Gadd, M., and Newman, P. (2020). kRadar++: Coarse-to-Fine FMCW Scanning Radar Localisation. Sensors, 20.","DOI":"10.3390\/s20216002"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"S\u0103ftescu, \u015e., Gadd, M., Martini, D., Barnes, D., and Newman, P. (2020). Kidnapped Radar: Topological Radar Localisation using Rotationally-Invariant Metric Learning. arXiv.","DOI":"10.1109\/ICRA40945.2020.9196682"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Gadd, M., Martini, D.D., and Newman, P. (2020, January 20\u201323). Look around You: Sequence-based Radar Place Recognition with Learned Rotational Invariance. Proceedings of the IEEE\/ION Position, Location and Navigation Symposium (PLANS), Portland, OR, USA.","DOI":"10.1109\/PLANS46316.2020.9109951"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"1224","DOI":"10.1109\/TPAMI.2017.2709749","article-title":"SIFT meets CNN: A decade survey of instance retrieval","volume":"40","author":"Zheng","year":"2018","journal-title":"IEEE Trans. Pattern. Anal. Mach. Intell."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"1790","DOI":"10.1109\/TPAMI.2015.2500224","article-title":"Factors of Transferability for a Generic ConvNet Representation","volume":"38","author":"Azizpour","year":"2016","journal-title":"IEEE Trans. Pattern. Anal. Mach. Intell."},{"key":"ref_42","unstructured":"Zheng, L., Zhao, Y., Wang, S., Wang, J., and Tian, Q. (2016). Good practice in CNN feature transfer. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Jin, S., Gao, Y., and Chen, L. (2020). Improved Deep Distance Learning for Visual Loop Closure Detection in Smart City. Peer-to-Peer Netw. Appl.","DOI":"10.1007\/s12083-019-00861-w"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Cieslewski, T., Choudhary, S., and Scaramuzza, D. (2018, January 21\u201325). Data-efficient decentralized visual SLAM. Proceedings of the 2018 IEEE International Conference on Robotics and Automation, Brisbane, Australia.","DOI":"10.1109\/ICRA.2018.8461155"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"1437","DOI":"10.1109\/TPAMI.2017.2711011","article-title":"NetVLAD: CNN Architecture for Weakly Supervised Place Recognition","volume":"40","author":"Arandjelovi","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"J\u00e9gou, H., and Chum, O. (2012, January 7\u201313). Negative Evidences and Co-occurences in Image Retrieval: The Benefit of PCA and Whitening. Proceedings of the European Conference on Computer Vision, Firenze, Italy.","DOI":"10.1007\/978-3-642-33709-3_55"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"647","DOI":"10.1177\/0278364908090961","article-title":"FAB-MAP: Probabilistic localization and mapping in the space of appearance","volume":"27","author":"Cummins","year":"2008","journal-title":"Int. J. Rob. Res."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Sturm, J., Engelhard, N., Endres, F., Burgard, W., and Cremers, D. (2012, January 7\u201312). A benchmark for the evaluation of RGB-D SLAM systems. Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems, Vilamoura, Portugal.","DOI":"10.1109\/IROS.2012.6385773"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/1\/310\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:07:09Z","timestamp":1760159229000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/1\/310"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,1,5]]},"references-count":48,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2021,1]]}},"alternative-id":["s21010310"],"URL":"https:\/\/doi.org\/10.3390\/s21010310","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,1,5]]}}}