{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:19:17Z","timestamp":1760235557806,"version":"build-2065373602"},"reference-count":41,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2021,8,29]],"date-time":"2021-08-29T00:00:00Z","timestamp":1630195200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Provincia Autonoma di Trento LP6\/99","award":["S067\/12.2-2018-93\/MA"],"award-info":[{"award-number":["S067\/12.2-2018-93\/MA"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Indoor environment modeling has become a relevant topic in several application fields, including augmented, virtual, and extended reality. With the digital transformation, many industries have investigated two possibilities: generating detailed models of indoor environments, allowing viewers to navigate through them; and mapping surfaces so as to insert virtual elements into real scenes. The scope of the paper is twofold. We first review the existing state-of-the-art (SoA) of learning-based methods for 3D scene reconstruction based on structure from motion (SFM) that predict depth maps and camera poses from video streams. We then present an extensive evaluation using a recent SoA network, with particular attention on the capability of generalizing on new unseen data of indoor environments. The evaluation was conducted by using the absolute relative (AbsRel) measure of the depth map prediction as the baseline metric.<\/jats:p>","DOI":"10.3390\/jimaging7090167","type":"journal-article","created":{"date-parts":[[2021,8,30]],"date-time":"2021-08-30T11:01:37Z","timestamp":1630321297000},"page":"167","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Mobile-Based 3D Modeling: An In-Depth Evaluation for the Application in Indoor Scenarios"],"prefix":"10.3390","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2658-3086","authenticated-orcid":false,"given":"Martin","family":"De Pellegrini","sequence":"first","affiliation":[{"name":"ARCODA s.r.l., 38121 Trento, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8376-043X","authenticated-orcid":false,"given":"Lorenzo","family":"Orlandi","sequence":"additional","affiliation":[{"name":"ARCODA s.r.l., 38121 Trento, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Daniele","family":"Sevegnani","sequence":"additional","affiliation":[{"name":"ARCODA s.r.l., 38121 Trento, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7858-0928","authenticated-orcid":false,"given":"Nicola","family":"Conci","sequence":"additional","affiliation":[{"name":"Department of Information Engineering and Computer Science, University of Trento, 38123 Trento, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,8,29]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Fazakas, T., and Fekete, R.T. (2010, January 18\u201320). 3D reconstruction system for autonomous robot navigation. Proceedings of the 2010 11th International Symposium on Computational Intelligence and Informatics (CINTI), Budapest, Hungary.","DOI":"10.1109\/CINTI.2010.5672236"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"903","DOI":"10.1007\/s12524-017-0727-1","article-title":"Application of drone for landslide mapping, dimension estimation and its 3D reconstruction","volume":"46","author":"Gupta","year":"2018","journal-title":"J. Indian Soc. Remote Sens."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"339","DOI":"10.1109\/TMM.2012.2229264","article-title":"Real-time, realistic, full 3-D reconstruction of moving humans from multiple Kinect streams","volume":"15","author":"Alexiadis","year":"2013","journal-title":"IEEE Trans. Multimed."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wu, C. (July, January 29). Towards linear-time incremental structure from motion. Proceedings of the 2013 International Conference on 3D Vision-3DV 2013, Seattle, WA, USA.","DOI":"10.1109\/3DV.2013.25"},{"key":"ref_5","unstructured":"Wu, C. (2021, May 03). VisualSFM: A Visual Structure from Motion System, Available online: http:\/\/ccwu.me\/vsfm\/."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Xie, J., Ross, G., and Ali, F. (2016). Deep3d: Fully automatic 2d-to-3d video conversion with deep convolutional neural networks. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46493-0_51"},{"key":"ref_7","unstructured":"Vijayanarasimhan, S., Ricco, S., Schmid, C., Sukthankar, R., and Fragkiadaki, K. (2017). Sfm-net: Learning of structure and motion from video. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Khilar, R., Chitrakala, S., and SelvamParvathy, S. (2013, January 2\u20133). 3D image reconstruction: Techniques, applications and challenges. Proceedings of the 2013 International Conference on Optical Imaging Sensor and Security (ICOSS), Coimbatore, India.","DOI":"10.1109\/ICOISS.2013.6678395"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"319","DOI":"10.5194\/isprsarchives-XL-1-W5-319-2015","article-title":"3D Reconstruction from Multi-View Medical X-ray images\u2013review and evaluation of existing methods","volume":"XL-1-W5","author":"Hosseinian","year":"2015","journal-title":"Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"181","DOI":"10.1007\/BF00992782","article-title":"The lexical integrity principle: Evidence from Bantu","volume":"13","author":"Bresnan","year":"1995","journal-title":"Nat. Lang. Linguist. Theory"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Szeliski, R. (2010). Computer Vision: Algorithms and Applications, Springer Science & Business Media.","DOI":"10.1007\/978-1-84882-935-0"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Lowe, D.G. (1999, January 20\u201325). Object recognition from local scale-invariant features. Proceedings of the Seventh IEEE International Conference on Computer Vision, Corfu, Greece.","DOI":"10.1109\/ICCV.1999.790410"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Mayer, N., Ilg, E., Hausser, P., Fischer, P., Cremers, D., Dosovitskiy, A., and Brox, T. (2016, January 27\u201330). A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.438"},{"key":"ref_14","unstructured":"Flynn, J., Snavely, K., Neulander, I., and Philbin, J. (2018). Deepstereo: Learning to Predict New Views from Real World Imagery. (No. 9,916,679), U.S. Patent."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Garg, R., Vijay Kumar, B.G., Gustavo, C., and Ian, R. (2016). Unsupervised cnn for single view depth estimation: Geometry to the rescue. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46484-8_45"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Agrawal, P., Joao, C., and Jitendra, M. (2015, January 7\u201312). Learning to see by moving. Proceedings of the IEEE International Conference on Computer Vision, Boston, MA, USA.","DOI":"10.1109\/ICCV.2015.13"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Jayaraman, D., and Kristen, G. (2015, January 13\u201316). Learning image representations equivariant to ego-motion. Proceedings of the 2015 International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.166"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Goroshin, R., Bruna, J., and Tompson, J. (2015, January 7\u201313). Unsupervised learning of spatiotemporally coherent metrics. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.465"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Misra, I., Zitnick, C.L., and Hebert, M. (2016). Shuffle and learn: Unsupervised learning using temporal order verification. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46448-0_32"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Pathak, D., Girshick, R., Dollar, P., Darrell, T., and Hariharan, B. (2017, January 21\u201326). Learning features by watching objects move. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.638"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Wang, X., and Gupta, A. (2015, January 7\u201313). Unsupervised learning of visual representations using videos. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.320"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhou, T., Brown, M., Snavely, N., and Lowe, D.G. (2017, January 21\u201326). Unsupervised learning of depth and ego-motion from video. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.700"},{"key":"ref_23","unstructured":"Bian, J.-W., Li, Z., Wang, N., Zhan, H., Shen, C., Cheng, M., and Reid, I. (2019). Unsupervised scale-consistent depth and ego-motion learning from monocular video. arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The kitti dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_25","unstructured":"Cordts, M., Omran, M., and Ramos, S. (2015, January 11). The cityscapes dataset. Proceedings of the CVPR Workshop on the Future of Datasets in Vision, Boston, MA, USA."},{"key":"ref_26","unstructured":"Bian, J.-W., Zhan, H., Wang, N., Chin, T.-J., Shen, C., and Reid, I. (2020). Unsupervised depth learning in challenging indoor video: Weak rectification to rescue. arXiv."},{"key":"ref_27","unstructured":"(2021, May 05). NYU Depth Datadet Version 2. Available online: https:\/\/cs.nyu.edu\/~silberman\/datasets\/nyu_depth_v2.html."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Silberman, N., Hoiem, D., Kohli, P., and Fergus, R. (2012). Indoor segmentation and support inference from rgbd images. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-642-33715-4_54"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Glocker, B., Izadi, S., Shotton, J., and Criminisi, A. (2013, January 1\u20134). Real-time RGB-D camera relocalization. Proceedings of the 2013 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), Adelaide, SA, Australia.","DOI":"10.1109\/ISMAR.2013.6671777"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Li, F.F. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 21\u201326). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_32","unstructured":"Bian, J. (2021, May 03). Unsupervised-Indoor-Depth. Available online: https:\/\/github.com\/JiawangBian\/Unsupervised-Indoor-Depth."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Sturm, J., Engelhard, N., Endres, F., Burgard, W., and Cremers, D. (2012, January 7\u201312). A benchmark for the evaluation of RGB-D SLAM systems. Proceedings of the 2012 IEEE\/RSJ International Conference on Intelligent Robots and Systems, Vilamoura-Algarve, Portugal.","DOI":"10.1109\/IROS.2012.6385773"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Lai, K., Bo, L., Ren, X., and Fox, D. (2011, January 9\u201313). A large-scale hierarchical multi-view rgb-d object dataset. Proceedings of the 2011 IEEE International Conference on Robotics and Automation, Shanghai, China.","DOI":"10.1109\/ICRA.2011.5980382"},{"key":"ref_35","unstructured":"Shuran, S., Lichtenberg, S.P., and Xiao, J. (2015, January 7\u201312). Sun rgb-d: A rgb-d scene understanding benchmark suite. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_36","unstructured":"(2021, May 05). Computer Vision Group TUM Department of Informatics Technical University of Munich, RGB-D SLAM Dataset. Available online: https:\/\/vision.in.tum.de\/data\/datasets\/rgbd-dataset\/download."},{"key":"ref_37","unstructured":"(2021, April 30). RGB-D Dataset 7-Scene. Available online: https:\/\/www.microsoft.com\/en-us\/research\/project\/rgb-d-dataset-7-scenes\/."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Janoch, A., Karayev, S., Jia, Y., Barron, J.T., Fritz, M., Saenko, K., and Darrell, T. (2013). A category-level 3d object dataset: Putting the kinect to work. Consumer Depth Cameras for Computer Vision, Springer.","DOI":"10.1007\/978-1-4471-4640-7_8"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Xiao, J., Andrew, O., and Antonio, T. (2013, January 1\u20138). Sun3d: A database of big spaces reconstructed using sfm and object labels. Proceedings of the IEEE International Conference on Computer Vision, Sydney, Australia.","DOI":"10.1109\/ICCV.2013.458"},{"key":"ref_40","unstructured":"Wasenm\u00fcller, O., and Didier, S. (2016). Comparison of kinect v1 and v2 depth images in terms of accuracy and precision. Asian Conference on Computer Vision, Springer."},{"key":"ref_41","unstructured":"Godard, C., Aodha, O.M., Firman, M., and Brostow, G.J. (November, January 27). Digging into self-supervised monocular depth estimation. Proceedings of the International Conference on Computer Vision (ICCV), Seoul, Korea."}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/7\/9\/167\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:54:57Z","timestamp":1760165697000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/7\/9\/167"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,29]]},"references-count":41,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2021,9]]}},"alternative-id":["jimaging7090167"],"URL":"https:\/\/doi.org\/10.3390\/jimaging7090167","relation":{},"ISSN":["2313-433X"],"issn-type":[{"type":"electronic","value":"2313-433X"}],"subject":[],"published":{"date-parts":[[2021,8,29]]}}}