{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T10:13:47Z","timestamp":1781086427187,"version":"3.54.1"},"reference-count":56,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2021,5,13]],"date-time":"2021-05-13T00:00:00Z","timestamp":1620864000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Precise localization and pose estimation in indoor environments are commonly employed in a wide range of applications, including robotics, augmented reality, and navigation and positioning services. Such applications can be solved via visual-based localization using a pre-built 3D model. The increase in searching space associated with large scenes can be overcome by retrieving images in advance and subsequently estimating the pose. The majority of current deep learning-based image retrieval methods require labeled data, which increase data annotation costs and complicate the acquisition of data. In this paper, we propose an unsupervised hierarchical indoor localization framework that integrates an unsupervised network variational autoencoder (VAE) with a visual-based Structure-from-Motion (SfM) approach in order to extract global and local features. During the localization process, global features are applied for the image retrieval at the level of the scene map in order to obtain candidate images, and are subsequently used to estimate the pose from 2D-3D matches between query and candidate images. RGB images only are used as the input of the proposed localization system, which is both convenient and challenging. Experimental results reveal that the proposed method can localize images within 0.16 m and 4\u00b0 in the 7-Scenes data sets and 32.8% within 5 m and 20\u00b0 in the Baidu data set. Furthermore, our proposed method achieves a higher precision compared to advanced methods.<\/jats:p>","DOI":"10.3390\/s21103406","type":"journal-article","created":{"date-parts":[[2021,5,14]],"date-time":"2021-05-14T03:28:36Z","timestamp":1620962916000},"page":"3406","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["A Visual and VAE Based Hierarchical Indoor Localization Method"],"prefix":"10.3390","volume":"21","author":[{"given":"Jie","family":"Jiang","sequence":"first","affiliation":[{"name":"College of Systems Engineering, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5717-7187","authenticated-orcid":false,"given":"Yin","family":"Zou","sequence":"additional","affiliation":[{"name":"College of Systems Engineering, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lidong","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Systems Engineering, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yujie","family":"Fang","sequence":"additional","affiliation":[{"name":"College of Systems Engineering, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2021,5,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1461","DOI":"10.1109\/TMC.2018.2857772","article-title":"ViNav: A vision-based indoor navigation system for smartphones","volume":"18","author":"Dong","year":"2018","journal-title":"Ieee Trans. Mob. Comput."},{"key":"ref_2","first-page":"217","article-title":"Emergency rescue localization (ERL) using GPS, wireless LAN and camera","volume":"9","author":"Bejuri","year":"2015","journal-title":"Int. J. Softw Eng. Appl."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Dickinson, P., Cielniak, G., Szymanezyk, O., and Mannion, M. (2016, January 4\u20137). Indoor positioning of shoppers using a network of Bluetooth Low Energy beacons. Proceedings of the 2016 International Conference on Indoor Positioning and Indoor Navigation (IPIN), Alcala de Henares, Spain.","DOI":"10.1109\/IPIN.2016.7743684"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Xia, S., Liu, Y., Yuan, G., Zhu, M., and Wang, Z. (2017). Indoor fingerprint positioning based on Wi-Fi: An overview. Isprs Int. J. Geo-Inf., 6.","DOI":"10.3390\/ijgi6050135"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2933232","article-title":"A survey on wireless indoor localization from the device perspective","volume":"49","author":"Xiao","year":"2016","journal-title":"Acm Comput. Surv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Xu, H., Ding, Y., Li, P., Wang, R., and Li, Y. (2017). An RFID indoor positioning algorithm based on Bayesian probability and K-nearest neighbor. Sensors, 17.","DOI":"10.3390\/s17081806"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"358","DOI":"10.1007\/s10489-013-0461-5","article-title":"A machine learning based intelligent vision system for autonomous object detection and recognition","volume":"40","author":"Sabourin","year":"2014","journal-title":"Appl. Intell."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Laskar, Z., Melekhov, I., Kalia, S., and Kannala, J. (2017, January 22\u201329). Camera Relocalization by Computing Pairwise Relative Poses Using Convolutional Neural Network. Proceedings of the 2017 IEEE International Conference on Computer Vision Workshops (ICCVW), Venice, Italy.","DOI":"10.1109\/ICCVW.2017.113"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Engel, J., Sch\u00f6ps, T., and Cremers, D. (2014). LSD-SLAM: Large-scale direct monocular SLAM. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10605-2_54"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1147","DOI":"10.1109\/TRO.2015.2463671","article-title":"ORB-SLAM: A versatile and accurate monocular SLAM system","volume":"31","author":"Montiel","year":"2015","journal-title":"Ieee Trans. Robot."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Schonberger, J.L., and Frahm, J.-M. (2016, January 27\u201330). Structure-from-motion revisited. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.445"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Taira, H., Okutomi, M., Sattler, T., Cimpoi, M., Pollefeys, M., Sivic, J., Pajdla, T., and Torii, A. (2018, January 18\u201323). InLoc: Indoor visual localization with dense matching and view synthesis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00752"},{"key":"ref_13","unstructured":"Sarlin, P.-E., Debraine, F., Dymczyk, M., Siegwart, R., and Cadena, C. (2018, January 29\u201331). Leveraging deep visual descriptors for hierarchical efficient localization. Proceedings of the Conference on Robot Learning, Z\u00fcrich, Switzerland."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Noh, H., Araujo, A., Sim, J., Weyand, T., and Han, B. (2017, January 22\u201329). Large-scale image retrieval with attentive deep local features. Proceedings of the IEEE international conference on computer vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.374"},{"key":"ref_15","unstructured":"Revaud, J., Almaz\u00e1n, J., Rezende, R.S., and Souza, C.R.d. (November, January 27). Learning with average precision: Training image retrieval with a listwise loss. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Sarlin, P.-E., Cadena, C., Siegwart, R., and Dymczyk, M. (2019, January 15\u201320). From coarse to fine: Robust hierarchical localization at large scale. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01300"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Kendall, A., Grimes, M., and Cipolla, R. (2015, January 7\u201313). Posenet: A convolutional network for real-time 6-dof camera relocalization. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.336"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Balntas, V., Li, S., and Prisacariu, V. (2018, January 8\u201314). RelocNet: Continuous Metric Learning Relocalisation Using Neural Nets. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_46"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Cimarelli, C., Cazzato, D., Olivares-Mendez, M.A., and Voos, H. (2019). Faster Visual-Based Localization with Mobile-PoseNet, Interdisciplinary Center for Security Reliability and Trust (SnT) University of Luxembourg.","DOI":"10.1007\/978-3-030-29891-3_20"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Sattler, T., Zhou, Q., Pollefeys, M., and Leal-Taixe, L. (2019, January 15\u201320). Understanding the limitations of cnn-based absolute camera pose regression. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00342"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Seifi, S., and Tuytelaars, T. (2019, January 27\u201328). How to Improve CNN-Based 6-DoF Camera Pose Estimation. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision Workshop (ICCVW), Seoul, Korea.","DOI":"10.1109\/ICCVW.2019.00471"},{"key":"ref_22","first-page":"1245","article-title":"Image-similarity-based Convolutional Neural Network for Robot Visual Relocalization","volume":"32","author":"Wang","year":"2020","journal-title":"Sens. Mater."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Sch\u00f6nberger, J.L., Zheng, E., Frahm, J.-M., and Pollefeys, M. (2016, January 8\u201316). Pixelwise view selection for unstructured multi-view stereo. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46487-9_31"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1744","DOI":"10.1109\/TPAMI.2016.2611662","article-title":"Efficient & effective prioritized matching for large-scale image-based localization","volume":"39","author":"Sattler","year":"2016","journal-title":"Ieee Trans. Pattern Anal. Mach. Intell."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Brachmann, E., Krull, A., Nowozin, S., Shotton, J., Michel, F., Gumhold, S., and Rother, C. (2017, January 21\u201326). Dsac-differentiable ransac for camera localization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.267"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Brachmann, E., and Rother, C. (2018, January 18\u201322). Learning Less is More\u20146D Camera Localization via 3D Surface Regression. Proceedings of the Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00489"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Meng, L., Chen, J., Tung, F., Little, J.J., Valentin, J., and de Silva, C.W. (2017, January 24\u201328). Backtracking regression forests for accurate camera relocalization. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8206611"},{"key":"ref_28","unstructured":"Nister, D., and Stewenius, H. (2006, January 17\u201322). Scalable recognition with a vocabulary tree. Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201906), New York, NY, USA."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Li, Y., Snavely, N., and Huttenlocher, D.P. (2010, January 5\u201311). Location recognition using prioritized feature matching. Proceedings of the European Conference on Computer Vision, Heraklion, Crete.","DOI":"10.1007\/978-3-642-15552-9_57"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Middelberg, S., Sattler, T., Untzelmann, O., and Kobbelt, L. (2014, January 6\u201312). Scalable 6-dof localization on mobile devices. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10605-2_18"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive image features from scale-invariant keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"1188","DOI":"10.1109\/TRO.2012.2197158","article-title":"Bags of binary words for fast place recognition in image sequences","volume":"28","author":"Tardos","year":"2012","journal-title":"Ieee Trans. Robot."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"J\u00e9gou, H., Douze, M., Schmid, C., and P\u00e9rez, P. (2010, January 13\u201318). Aggregating local descriptors into a compact image representation. Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5540039"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Torii, A., Arandjelovic, R., Sivic, J., Okutomi, M., and Pajdla, T. (2015, January 7\u201312). 24\/7 place recognition by view synthesis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298790"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"1224","DOI":"10.1109\/TPAMI.2017.2709749","article-title":"SIFT meets CNN: A decade survey of instance retrieval","volume":"40","author":"Zheng","year":"2017","journal-title":"Ieee Trans. Pattern Anal. Mach. Intell."},{"key":"ref_36","unstructured":"Chen, W., Liu, Y., Wang, W., Bakker, E., Georgiou, T., Fieguth, P., Liu, L., and Lew, M.S. (2021). Deep Image Retrieval: A Survey. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Arandjelovic, R., Gronat, P., Torii, A., Pajdla, T., and Sivic, J. (2016, January 27\u201330). NetVLAD: CNN architecture for weakly supervised place recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.572"},{"key":"ref_38","unstructured":"Torii, A., Taira, H., Sivic, J., Pollefeys, M., Okutomi, M., Pajdla, T., and Sattler, T. (2017, January 21\u201326). Are large-scale 3d models really necessary for accurate visual localization?. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Camposeco, F., Cohen, A., Pollefeys, M., and Sattler, T. (2018, January 18\u201323). Hybrid camera pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00022"},{"key":"ref_40","unstructured":"Sarlin, P.-E., Debraine, F., Dymczyk, M., Siegwart, R., and Cadena, C. (2018). Leveraging deep visual descriptors for hierarchical efficient localization. arXiv."},{"key":"ref_41","unstructured":"Hinton, G., Vinyals, O., and Dean, J. (2015). Distilling the knowledge in a neural network. arXiv."},{"key":"ref_42","unstructured":"Kingma, D.P., and Welling, M. (2013). Auto-encoding variational bayes. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"307","DOI":"10.1561\/2200000056","article-title":"An Introduction to Variational Autoencoders","volume":"12","author":"Kingma","year":"2019","journal-title":"Found. Trends\u00ae Mach. Learn."},{"key":"ref_44","unstructured":"Lucas, J., Tucker, G., Grosse, R., and Norouzi, M. (2019, January 6). Understanding posterior collapse in generative latent variable models. Proceedings of the 2019 Deep Generative Models for Highly Structured Data, New Orleans, LA, USA."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Wu, H., and Fieri, M. (2019, January 11\u201314). Learning product codebooks using vector-quantized autoencoders for image retrieval. Proceedings of the 7th IEEE Global Conference on Signal and Information Processing, GlobalSIP 2019, Ottawa, ON, Canada.","DOI":"10.1109\/GlobalSIP45357.2019.8969272"},{"key":"ref_46","unstructured":"Li, X., Chen, Z., Poon, L.K., and Zhang, N.L. (2018). Learning latent superstructures in variational autoencoders for deep multidimensional clustering. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Jiang, Z., Zheng, Y., Tan, H., Tang, B., and Zhou, H. (2016). Variational deep embedding: An unsupervised and generative approach to clustering. arXiv.","DOI":"10.24963\/ijcai.2017\/273"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Sch\u00f6nberger, J.L., Pollefeys, M., Geiger, A., and Sattler, T. (2018, January 18\u201323). Semantic visual localization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00721"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"199440","DOI":"10.1109\/ACCESS.2020.3034828","article-title":"Balancing reconstruction error and Kullback-Leibler divergence in Variational Autoencoders","volume":"8","author":"Asperti","year":"2020","journal-title":"IEEE Access"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Bowman, S.R., Vilnis, L., Vinyals, O., Dai, A.M., Jozefowicz, R., and Bengio, S. (2015). Generating sentences from a continuous space. arXiv.","DOI":"10.18653\/v1\/K16-1002"},{"key":"ref_51","unstructured":"Kingma, D.P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I., and Welling, M. (2016). Improving variational inference with inverse autoregressive flow.(nips). arXiv."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Shotton, J., Glocker, B., Zach, C., Izadi, S., Criminisi, A., and Fitzgibbon, A. (2013, January 23\u201328). Scene coordinate regression forests for camera relocalization in RGB-D images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.377"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Sun, X., Xie, Y., Luo, P., and Wang, L. (2017, January 21\u201326). A dataset for benchmarking image-based localization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.598"},{"key":"ref_54","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_55","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Hinton","year":"2008","journal-title":"J. Mach. Learn. Res."},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Weinzaepfel, P., Csurka, G., Cabon, Y., and Humenberger, M. (2019, January 16\u201320). Visual Localization by Learning Objects-Of-Interest Dense Match Regression. Proceedings of the Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00578"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/10\/3406\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:00:28Z","timestamp":1760162428000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/10\/3406"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,13]]},"references-count":56,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2021,5]]}},"alternative-id":["s21103406"],"URL":"https:\/\/doi.org\/10.3390\/s21103406","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,5,13]]}}}