{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,21]],"date-time":"2026-02-21T01:09:37Z","timestamp":1771636177972,"version":"3.50.1"},"reference-count":38,"publisher":"MDPI AG","issue":"19","license":[{"start":{"date-parts":[[2020,9,25]],"date-time":"2020-09-25T00:00:00Z","timestamp":1600992000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Recently, deep convolutional neural networks (CNN) have become popular for indoor visual localisation, where the networks learn to regress the camera pose from images directly. However, these approaches perform a 3D image-based reconstruction of the indoor spaces beforehand to determine camera poses, which is a challenge for large indoor spaces. Synthetic images derived from 3D indoor models have been used to eliminate the requirement of 3D reconstruction. A limitation of the approach is the low accuracy that occurs as a result of estimating the pose of each image frame independently. In this article, a visual localisation approach is proposed that exploits the spatio-temporal information from synthetic image sequences to improve localisation accuracy. A deep Bayesian recurrent CNN is fine-tuned using synthetic image sequences obtained from a building information model (BIM) to regress the pose of real image sequences. The results of the experiments indicate that the proposed approach estimates a smoother trajectory with smaller inter-frame error as compared to existing methods. The achievable accuracy with the proposed approach is 1.6 m, which is an improvement of approximately thirty per cent compared to the existing approaches. A Keras implementation can be found in our Github repository.<\/jats:p>","DOI":"10.3390\/s20195492","type":"journal-article","created":{"date-parts":[[2020,9,25]],"date-time":"2020-09-25T08:57:32Z","timestamp":1601024252000},"page":"5492","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":32,"title":["A Recurrent Deep Network for Estimating the Pose of Real Indoor Images from Synthetic Image Sequences"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6704-5778","authenticated-orcid":false,"given":"Debaditya","family":"Acharya","sequence":"first","affiliation":[{"name":"Department of Infrastructure Engineering, The University of Melbourne, Parkville, Victoria 3010, Australia"},{"name":"Department of Manufacturing, Materials and Mechatronics, RMIT University, Carlton, Victoria 3053, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2440-8484","authenticated-orcid":false,"given":"Sesa","family":"Singha Roy","sequence":"additional","affiliation":[{"name":"Institute for Sustainable Industries and Livable Cities, Victoria University, Werribee, Victoria 3030, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6639-1727","authenticated-orcid":false,"given":"Kourosh","family":"Khoshelham","sequence":"additional","affiliation":[{"name":"Department of Infrastructure Engineering, The University of Melbourne, Parkville, Victoria 3010, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3403-6939","authenticated-orcid":false,"given":"Stephan","family":"Winter","sequence":"additional","affiliation":[{"name":"Department of Infrastructure Engineering, The University of Melbourne, Parkville, Victoria 3010, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,9,25]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Kendall, A., Grimes, M., and Cipolla, R. (2015, January 7\u201313). PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.336"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Kendall, A., and Cipolla, R. (2016, January 16\u201321). Modelling uncertainty in deep learning for camera relocalization. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden.","DOI":"10.1109\/ICRA.2016.7487679"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Kendall, A., and Cipolla, R. (2017, January 21\u201326). Geometric loss functions for camera pose regression with deep learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.694"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Walch, F., Hazirbas, C., Leal-Taixe, L., Sattler, T., Hilsenbeck, S., and Cremers, D. (2017, January 22\u201329). Image-based localization using lstms for structured feature correlation. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.75"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1362","DOI":"10.1109\/TPAMI.2009.161","article-title":"Accurate, dense, and robust multiview stereopsis","volume":"32","author":"Furukawa","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3322241","article-title":"Indoor Localization Improved by Spatial Context\u2014A Survey","volume":"52","author":"Gu","year":"2019","journal-title":"ACM Comput. Surv."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"245","DOI":"10.1016\/j.isprsjprs.2019.02.020","article-title":"BIM-PoseNet: Indoor camera localisation using a 3D indoor model and deep learning from synthetic images","volume":"150","author":"Acharya","year":"2019","journal-title":"ISPRS J. Photogramm. Remote. Sens."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"247","DOI":"10.5194\/isprs-annals-IV-2-W5-247-2019","article-title":"Modelling uncertainty of single image indoor localisation using a 3D model and deep learning","volume":"IV-2\/W5","author":"Acharya","year":"2019","journal-title":"Isprs Ann. Photogramm. Remote. Sens. Spat. Inf. Sci."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Davison, A.J. (2003, January 14\u201317). Real-time simultaneous localisation and mapping with a single camera. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Nice, France.","DOI":"10.1109\/ICCV.2003.1238654"},{"key":"ref_10","unstructured":"Nister, D., Naroditsky, O., and Bergen, J. (July, January 27). Visual odometry. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Washington, DC, USA."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1016\/j.isprsjprs.2019.02.014","article-title":"BIM-Tracker: A model-based visual tracking approach for indoor localisation using a 3D building model","volume":"150","author":"Acharya","year":"2019","journal-title":"ISPRS J. Photogramm. Remote. Sens."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"90","DOI":"10.1016\/j.patcog.2017.09.013","article-title":"A survey on Visual-Based Localization: On the benefit of heterogeneous data","volume":"74","author":"Piasco","year":"2018","journal-title":"Pattern Recognit."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Irschara, A., Zach, C., Frahm, J.M., and Bischof, H. (2009, January 20\u201325). From structure-from-motion point clouds to fast location recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA.","DOI":"10.1109\/CVPRW.2009.5206587"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Kneip, L., Scaramuzza, D., and Siegwart, R. (2011, January 20\u201325). A novel parametrization of the perspective-three-point problem for a direct computation of absolute camera position and orientation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995464"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Shotton, J., Glocker, B., Zach, C., Izadi, S., Criminisi, A., and Fitzgibbon, A. (2013, January 23\u201328). Scene Coordinate Regression Forests for Camera Relocalization in RGB-D Images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.377"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Brachmann, E., and Rother, C. (2018, January 18\u201322). Learning less is more-6d camera localization via 3d surface regression. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00489"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Cavallari, T., Golodetz, S., Lord, N.A., Valentin, J., Di Stefano, L., and Torr, P.H. (2017, January 21\u201326). On-the-fly adaptation of regression forests for online camera relocalisation. Proceedings of the IEEE conference on computer vision and pattern recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.31"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Brachmann, E., Krull, A., Nowozin, S., Shotton, J., Michel, F., Gumhold, S., and Rother, C. (2017, January 21\u201326). Dsac-differentiable ransac for camera localization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.267"},{"key":"ref_19","unstructured":"Wu, J., Ma, L., and Hu, X. (June, January 29). Delving deeper into convolutional neural networks for camera relocalization. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Singapore."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Clark, R., Wang, S., Markham, A., Trigoni, N., and Wen, H. (2017, January 21\u201326). VidLoc: A deep spatio-temporal model for 6-dof video-clip relocalization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.284"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1016\/j.buildenv.2018.05.026","article-title":"Image retrieval using BIM and features from pretrained VGG network for indoor localization","volume":"140","author":"Ha","year":"2018","journal-title":"Build. Environ."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going Deeper With Convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016, January 27\u201330). Rethinking the inception architecture for computer vision. Proceedings of the IEEE conference on computer vision and pattern recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_25","unstructured":"Zhou, B., Lapedriza, A., Xiao, J., Torralba, A., and Oliva, A. (2014, January 8\u201313). Learning deep features for scene recognition using places database. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zhu, L., and Laptev, N. (2017, January 28\u201321). Deep and confident prediction for time series at uber. Proceedings of the 2017 IEEE International Conference on Data Mining Workshops (ICDMW), New Orleans, LA, USA.","DOI":"10.1109\/ICDMW.2017.19"},{"key":"ref_27","unstructured":"Gal, Y., and Ghahramani, Z. (2015). Bayesian convolutional neural networks with Bernoulli approximate variational inference. arXiv."},{"key":"ref_28","unstructured":"Chollet, F. (2020, August 15). Keras. Available online: https:\/\/keras.io."},{"key":"ref_29","unstructured":"Abadi, M., Agarwal, A., Barham, P., Brevdo, E., and Chen, Z. (2020, August 15). TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. Available online: tensorflow.org."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"367","DOI":"10.5194\/isprs-archives-XLII-2-W7-367-2017","article-title":"The isprs benchmark on indoor modelling","volume":"42","author":"Khoshelham","year":"2017","journal-title":"Int. Arch. Photogramm. Remote. Sens. Spat. Inf. Sci."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"371","DOI":"10.5194\/isprs-annals-IV-2-W4-371-2017","article-title":"Indoor positioning by visual-inertial odometry","volume":"4","author":"Ramezani","year":"2017","journal-title":"ISPRS Ann. Photogramm. Remote. Sens. Spat. Inf. Sci."},{"key":"ref_32","first-page":"297","article-title":"An evaluation framework for benchmarking indoor modelling methods","volume":"XLII-4","author":"Khoshelham","year":"2018","journal-title":"ISPRS Int. Arch. Photogramm. Remote. Sens. Spat. Inf. Sci."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1016\/j.isprsjprs.2019.01.012","article-title":"Geometric comparison and quality evaluation of 3D models of indoor environments","volume":"149","author":"Tran","year":"2019","journal-title":"ISPRS J. Photogramm. Remote. Sens."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1016\/j.autcon.2013.10.023","article-title":"Building Information Modeling (BIM) for existing buildings\u2014Literature review and future needs","volume":"38","author":"Volk","year":"2014","journal-title":"Autom. Constr."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"204","DOI":"10.1016\/j.autcon.2012.12.004","article-title":"Building Information Modeling (BIM) partnering framework for public construction projects","volume":"31","author":"Porwal","year":"2013","journal-title":"Autom. Constr."},{"key":"ref_36","unstructured":"Gal, Y., and Ghahramani, Z. (2016, January 5\u201310). A theoretically grounded application of dropout in recurrent neural networks. Proceedings of the Advances in Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"450","DOI":"10.3390\/make1010027","article-title":"Guidelines and Benchmarks for Deployment of Deep Learning Models on Smartphones as Real-Time Apps","volume":"1","author":"Sehgal","year":"2019","journal-title":"Mach. Learn. Knowl. Extr."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/19\/5492\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:13:31Z","timestamp":1760177611000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/19\/5492"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,25]]},"references-count":38,"journal-issue":{"issue":"19","published-online":{"date-parts":[[2020,10]]}},"alternative-id":["s20195492"],"URL":"https:\/\/doi.org\/10.3390\/s20195492","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,9,25]]}}}