{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,28]],"date-time":"2026-01-28T06:43:01Z","timestamp":1769582581063,"version":"3.49.0"},"reference-count":52,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2019,1,10]],"date-time":"2019-01-10T00:00:00Z","timestamp":1547078400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61633002"],"award-info":[{"award-number":["61633002"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>This paper proposes a novel multi-sensor-based indoor global localization system integrating visual localization aided by CNN-based image retrieval with a probabilistic localization approach. The global localization system consists of three parts: coarse place recognition, fine localization and re-localization from kidnapping. Coarse place recognition exploits a monocular camera to realize the initial localization based on image retrieval, in which off-the-shelf features extracted from a pre-trained Convolutional Neural Network (CNN) are adopted to determine the candidate locations of the robot. In the fine localization, a laser range finder is equipped to estimate the accurate pose of a mobile robot by means of an adaptive Monte Carlo localization, in which the candidate locations obtained by image retrieval are considered as seeds for initial random sampling. Additionally, to address the problem of robot kidnapping, we present a closed-loop localization mechanism to monitor the state of the robot in real time and make adaptive adjustments when the robot is kidnapped. The closed-loop mechanism effectively exploits the correlation of image sequences to realize the re-localization based on Long-Short Term Memory (LSTM) network. Extensive experiments were conducted and the results indicate that the proposed method not only exhibits great improvement on accuracy and speed, but also can recover from localization failures compared to two conventional localization methods.<\/jats:p>","DOI":"10.3390\/s19020249","type":"journal-article","created":{"date-parts":[[2019,1,11]],"date-time":"2019-01-11T04:10:16Z","timestamp":1547179816000},"page":"249","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":56,"title":["A Robust Indoor Localization System Integrating Visual Localization Aided by CNN-Based Image Retrieval with Monte Carlo Localization"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4263-5632","authenticated-orcid":false,"given":"Song","family":"Xu","sequence":"first","affiliation":[{"name":"School of Mechanical Engineering and Automation, Beihang University, Beijing 100191, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wusheng","family":"Chou","sequence":"additional","affiliation":[{"name":"School of Mechanical Engineering and Automation, Beihang University, Beijing 100191, China"},{"name":"State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing 100191, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongyi","family":"Dong","sequence":"additional","affiliation":[{"name":"School of Mechanical Engineering and Automation, Beihang University, Beijing 100191, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,1,10]]},"reference":[{"key":"ref_1","unstructured":"Thrun, S., Burgard, W., and Fox, D. (2005). Probabilistic Robotics, MIT Press."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1016\/S0004-3702(01)00069-8","article-title":"Robust Monte Carlo localization for mobile robots","volume":"128","author":"Thrun","year":"2001","journal-title":"Artif. Intell."},{"key":"ref_3","unstructured":"Thrun, S., Fox, D., and Burgard, W. (August, January 30). Monte Carlo Localization with Mixture Proposal Distribution. Proceedings of the National Conference on Artificial Intelligence (AAAI), Austin, TX, USA."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1067","DOI":"10.1109\/TSMCC.2007.905750","article-title":"Survey of Wireless Indoor Positioning Techniques and Systems","volume":"37","author":"Liu","year":"2007","journal-title":"IEEE Trans. Syst. Man Cybern. C."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Jiang, Y., Pan, X., Li, K., Lv, Q., Dick, RP., Hannigan, M., and Shang, L. (2012, January 5\u20138). ARIEL: Automatic wi-fi based room fingerprinting for indoor localization. Proceedings of the ACM International Conference on Ubiquitous Computing (UbiComp), Pittsburgh, PA, USA.","DOI":"10.1145\/2370216.2370282"},{"key":"ref_6","unstructured":"Dellaert, F., Fox, D., Burgard, W., and Thrun, S. (1999, January 10\u201315). Monte Carlo localization for mobile robots. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Detroit, MI, USA."},{"key":"ref_7","first-page":"327","article-title":"Markov localization for mobile robots in dynamic environments","volume":"2","author":"Fox","year":"1999","journal-title":"J. Artif. Intell. Res."},{"key":"ref_8","unstructured":"Roumeliotis, S.I., Bekey, G.A., Burgard, W., and Thrun, S. (2000, January 24\u201328). Bayesian estimation and Kalman filtering: A unified framework for mobile robot localization. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), San Francisco, CA, USA."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"528","DOI":"10.1109\/TRO.2016.2544301","article-title":"Coarse-to-Fine Localization for a Mobile Robot Based on Place Learning With a 2-D Range Scan","volume":"32","author":"Park","year":"2016","journal-title":"IEEE Trans. Robot."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"413","DOI":"10.1109\/TSMCB.2005.859085","article-title":"Coarse-to-fine vision-based localization by indexing scale-Invariant features","volume":"36","author":"Wang","year":"2006","journal-title":"IEEE Trans. Syst. Man Cybern. B."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Sattler, T., Leibe, B., and Kobbelt, L. (2012, January 3\u20137). Image retrieval for image-based localization revisited. Proceedings of the British Machine Vision Conference (BMVC), Guildford, Surrey, UK.","DOI":"10.5244\/C.26.76"},{"key":"ref_12","unstructured":"Sattler, T., Havlena, M., Scjindler, K., and Pollefeys, M. (July, January 26). Large-Scale Location Recognition and the Geometric Burstiness Problem. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Sunderhaulf, N., Dayoub, F., McMahon, S., Talbot, B., Schulz, R., Corke, P., Wyeth, G., Upcroft, B., and Milford, M. (2016, January 16\u201321). Place categorization and semantic mapping on a mobile robot. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden.","DOI":"10.1109\/ICRA.2016.7487796"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zeisl, B., Sattler, T., and Pollefeys, M. (2015, January 13\u201316). Camera Pose Voting for Large-Scale Image-Based Localization. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.310"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1546","DOI":"10.1109\/TPAMI.2014.2299799","article-title":"Image Geo-Localization Based on MultipleNearest Neighbor Feature Matching UsingGeneralized Graphs","volume":"36","author":"Zamir","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Philbin, J., Chum, O., Isard, M., Sivic, J., and Zisserman, A. (2007, January 18\u201323). Object retrieval with large vocabularies and fast spatial matching. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Minneapolis, MN, USA.","DOI":"10.1109\/CVPR.2007.383172"},{"key":"ref_17","unstructured":"Biswas, J., and Veloso, M. (July, January 24). Multi-sensor Mobile Robot Localization for Diverse Environments. Proceedings of the Robot Soccer World Cup (RobotCup), Eindhoven, The Netherlands."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Srinivasan, K., and Gu, J. (2007, January 22\u201326). Multiple Sensor Fusion in Mobile Robot Localization. Proceedings of the Canadian Conference on Electrical and Computer Engineering (CCECE), Vancouver, BC, Canada.","DOI":"10.1109\/CCECE.2007.308"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Duan, P., Tian, G., and Wu, H. (2015, January 6\u20139). A multi-sensor-based mobile robot localization framework. Proceedings of the IEEE International Conference on Robotics and Biomimetics (ROBIO), Zhuhai, China.","DOI":"10.1109\/ROBIO.2014.7090403"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"208","DOI":"10.1109\/TRO.2004.835453","article-title":"Robust vision-based localization by combining an image-retrieval system with Monte Carlo localization","volume":"21","author":"Wolf","year":"2005","journal-title":"IEEE Trans. Robot"},{"key":"ref_21","unstructured":"Irschara, A., Zach, C., Frahm, J.M., and Bischof, H. (October, January 27). From structure-from-motion point clouds to fast location recognition. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Kyoto, Japan."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1744","DOI":"10.1109\/TPAMI.2016.2611662","article-title":"Efficient & Effective Prioritized Matching for Large-Scale Image-Based Localization","volume":"39","author":"Sattler","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zamir, A.R., and Shah, M. (2010, January 5\u201311). Accurate Image Localization Based on Google Maps Street View. Proceedings of the IEEE International Conference on Computer Vision (ECCV), Heraklion, Crete, Greece.","DOI":"10.1007\/978-3-642-15561-1_19"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"21636","DOI":"10.3390\/s150921636","article-title":"A Probabilistic Feature Map-Based Localization System Using a Monocular Camera","volume":"15","author":"Kim","year":"2015","journal-title":"Sensors"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1349","DOI":"10.1109\/34.895972","article-title":"Content-Based Image Retrieval at the End of the Early Years","volume":"22","author":"Smeulders","year":"2000","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"979","DOI":"10.1109\/TIP.2005.847289","article-title":"A unified framework for image retrieval using keyword and visual features","volume":"14","author":"Jing","year":"2005","journal-title":"IEEE Trans. Image Process."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"536","DOI":"10.1007\/s00530-002-0070-3","article-title":"Relevance feedback in image retrieval: A comprehensive review","volume":"8","author":"Zhou","year":"2003","journal-title":"Multimedia Syst."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Smith, J.R., and Chang, S.F. (1996, January 18\u201322). VisualSEEk: A fully automated content-based image query system. Proceedings of the Acm International Conference on Multimedia (ACMMM), Boston, MA, USA.","DOI":"10.1145\/244130.244151"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1224","DOI":"10.1109\/34.809116","article-title":"TextFinder: An Automatic System to Detect and Recognize Text In Images","volume":"21","author":"Wu","year":"1999","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"977","DOI":"10.1016\/j.patcog.2003.10.012","article-title":"Text information extraction in images and video: A survey","volume":"37","author":"Jung","year":"2004","journal-title":"Pattern Recognit."},{"key":"ref_31","unstructured":"Er, N., and Stewennius, H. (2006, January 17\u201322). Scalable recognition with a vocabulary tree. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), New York, NY, USA."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zhu, C.Z., and Satoh, S. (2012, January 5\u20138). Large vocabulary quantization for searching instances from videos. Proceedings of the ACM International Conference on Multimedia Retrieval (ICMR), Hong Kong, China.","DOI":"10.1145\/2324796.2324856"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Jegou, H., Douze, M., and Schmid, C. (2008, January 12\u201318). Hamming Embedding and Weak Geometry Consistency for Large Scale Image Search. Proceedings of the IEEE International Conference on Computer Vision (ECCV), Marseille, France.","DOI":"10.1007\/978-3-540-88682-2_24"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Jegou, H., Douze, M., Schmid, C., and Perez, P. (2010, January 13\u201318). Aggregating local descriptors into a compact image representation. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5540039"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Gong, Y., Wang, L., Guo, R., and Lazebnik, S. (2014, January 6\u201312). Multi-scale Orderless Pooling of Deep Convolutional Activation Features. Proceedings of the IEEE International Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10584-0_26"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Arandjelovic, R., Gronat, P., Torii, A., Pajdla, T., and Sivic, J. (2016, January 27\u201330). NetVLAD: CNN Architecture for Weakly Supervised Place Recognition. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.572"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Radenovi\u0107, F., Tolias, G., and Chum, O. (2016, January 11\u201314). CNN Image Retrieval Learns from BoW: Unsupervised Fine-Tuning with Hard Examples. Proceedings of the IEEE International Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_1"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Liu, H., Tian, Y., Wang, Y., Pang, L., and Huang, T. (2016, January 27\u201330). Deep Relative Distance Learning: Tell the Difference between Similar Vehicles. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.238"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Babenko, A., Slesarev, A., Chigorin, A., and Lempitsky, V. (2014, January 6\u201312). Neural Codes for Image Retrieval. Proceedings of the IEEE International Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10590-1_38"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Ng, J.Y.H., Yang, F., and Davis, L.S. (2015, January 7\u201312). Exploiting local features from deep networks for image retrieval. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Boston, MA, USA.","DOI":"10.1109\/CVPRW.2015.7301272"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Gordo, A., Almaz\u00e1n, J., Revaud, J., and Larlus, D. (2016, January 11\u201314). Deep Image Retrieval: Learning Global Representations for Image Search. Proceedings of the IEEE International Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46466-4_15"},{"key":"ref_42","unstructured":"Salvador, A., Giroinieto, X., Marques, F., and Satoh, S. (July, January 26). Faster R-CNN Features for Instance Search. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)."},{"key":"ref_43","unstructured":"Merwe, R.V.D., Doucet, A., and Freitas, N.D. (2000, January 1\u20132). The unscented particle filter. Proceedings of the 13th International Conference on Neural Information Processing Systems (NIPS), Denver, CO, USA."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"128","DOI":"10.1016\/j.cviu.2014.04.002","article-title":"Rao-Blackwellized particle filtering with Gaussian mixture models for robust visual tracking","volume":"125","author":"Kim","year":"2014","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"229","DOI":"10.1017\/S0263574711000567","article-title":"Self-adaptive Monte Carlo localization for mobile robots using range finders","volume":"30","author":"Zhang","year":"2012","journal-title":"Robot"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Sattler, T., Torii, A., Sivic, J., Pollefeys, M., Taira, H., Okutomi, M., and Pajdla, T. (2017, January 21\u201326). Are Large-Scale 3D Models Really Necessary for Accurate Visual Localization?. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.654"},{"key":"ref_47","unstructured":"Hao, J.D., Dong, J., Wang, W., and Tan, T.N. (2017, January 24\u201326). What Is the Best Practice for CNNs Applied to Visual Instance Retrieval?. Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"1155","DOI":"10.1109\/ACCESS.2017.2778011","article-title":"Action Recognition in Video Sequences using Deep Bi-Directional LSTM With CNN Features","volume":"6","author":"Ullah","year":"2018","journal-title":"IEEE Access"},{"key":"ref_49","unstructured":"Zaremba, W., Sutskever, I., and Vinyals, O. (arXiv, 2014). Recurrent Neural Network Regularization, arXiv."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Luong, M.T., Sutskever, I., Le, Q.V., Vinyals, O., and Zaremba, W. (arXiv, 2014). Addressing the Rare Word Problem in Neural Machine Translation, arXiv.","DOI":"10.3115\/v1\/P15-1002"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Varior, R.R., Shuai, B., Lu, J., Xu, D., and Wang, G. (2016, January 11\u201314). A Siamese Long Short-Term Memory Architecture for Human Re-identification. Proceedings of the IEEE International Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46478-7_9"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Donahue, J., Hendricks, L.A., Marcus, R., Venugopalan, S., Guadarrama, S., Saenko, K., and Darrell, T. (2015, January 7\u201312). Long-term Recurrent Convolutional Networks for Visual Recognition and Description. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298878"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/2\/249\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:24:58Z","timestamp":1760185498000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/2\/249"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,1,10]]},"references-count":52,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2019,1]]}},"alternative-id":["s19020249"],"URL":"https:\/\/doi.org\/10.3390\/s19020249","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,1,10]]}}}