{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T16:27:58Z","timestamp":1783009678487,"version":"3.54.5"},"reference-count":32,"publisher":"MDPI AG","issue":"24","license":[{"start":{"date-parts":[[2022,12,13]],"date-time":"2022-12-13T00:00:00Z","timestamp":1670889600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Flemish Government"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>3D reconstruction is the computer vision task of reconstructing the 3D shape of an object from multiple 2D images. Most existing algorithms for this task are designed for offline settings, producing a single reconstruction from a batch of images taken from diverse viewpoints. Alongside reconstruction accuracy, additional considerations arise when 3D reconstructions are used in real-time processing pipelines for applications such as robot navigation or manipulation. In these cases, an accurate 3D reconstruction is already required while the data gathering is still in progress. In this paper, we demonstrate how existing batch-based reconstruction algorithms lead to suboptimal reconstruction quality when used for online, iterative 3D reconstruction and propose appropriate modifications to the existing Pix2Vox++ architecture. When additional viewpoints become available at a high rate, e.g., from a camera mounted on a drone, selecting the most informative viewpoints is important in order to mitigate long term memory loss and to reduce the computational footprint. We present qualitative and quantitative results on the optimal selection of viewpoints and show that state-of-the-art reconstruction quality is already obtained with elementary selection algorithms.<\/jats:p>","DOI":"10.3390\/s22249782","type":"journal-article","created":{"date-parts":[[2022,12,14]],"date-time":"2022-12-14T03:21:52Z","timestamp":1670988112000},"page":"9782","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Iterative Online 3D Reconstruction from RGB Images"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1991-2478","authenticated-orcid":false,"given":"Thorsten","family":"Cardoen","sequence":"first","affiliation":[{"name":"IDLab, Department of Information and Technology, Ghent University-imec, 9052 Ghent, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3792-5026","authenticated-orcid":false,"given":"Sam","family":"Leroux","sequence":"additional","affiliation":[{"name":"IDLab, Department of Information and Technology, Ghent University-imec, 9052 Ghent, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9569-9373","authenticated-orcid":false,"given":"Pieter","family":"Simoens","sequence":"additional","affiliation":[{"name":"IDLab, Department of Information and Technology, Ghent University-imec, 9052 Ghent, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,12,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"15249","DOI":"10.1038\/s41598-021-94634-2","article-title":"2D\u20133D reconstruction of distal forearm bone from actual X-ray images of the wrist using convolutional neural networks","volume":"11","author":"Shiode","year":"2021","journal-title":"Sci. Rep."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"102053","DOI":"10.1016\/j.displa.2021.102053","article-title":"Review of multi-view 3D object recognition methods based on deep learning","volume":"69","author":"Qi","year":"2021","journal-title":"Displays"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Ren, R., Fu, H., Xue, H., Sun, Z., Ding, K., and Wang, P. (2021). Towards a Fully Automated 3D Reconstruction System Based on LiDAR and GNSS in Challenging Scenarios. Remote Sens., 13.","DOI":"10.3390\/rs13101981"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wang, S., Guo, J., Zhang, Y., Hu, Y., Ding, C., and Wu, Y. (2021). Single Target SAR 3D Reconstruction Based on Deep Learning. Sensors, 21.","DOI":"10.3390\/s21030964"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"625","DOI":"10.1111\/cgf.13386","article-title":"State of the Art on 3D Reconstruction with RGB-D Cameras","volume":"37","author":"Stotko","year":"2018","journal-title":"Comput. Graph. Forum"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"628","DOI":"10.1007\/978-3-319-46484-8_38","article-title":"3D-R2N2: A Unified Approach for Single and Multi-view 3D Object Reconstruction","volume":"Volume 9912","author":"Leibe","year":"2016","journal-title":"Proceedings of the Computer Vision\u2014ECCV 2016\u201414th European Conference"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1007\/s11263-019-01217-w","article-title":"Robust Attentional Aggregation of Deep Feature Sets for Multi-view 3D Reconstruction","volume":"128","author":"Yang","year":"2020","journal-title":"Int. J. Comput. Vis."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Xie, H., Yao, H., Sun, X., Zhou, S., and Zhang, S. (November, January 27). Pix2Vox: Context-aware 3 D Reconstruction from Single and Multiview Images. Proceedings of the IEEE\/CVF International Conference on Computer Vision 2019, Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00278"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"2919","DOI":"10.1007\/s11263-020-01347-6","article-title":"Pix2Vox++: Multi-scale Context-aware 3D Object Reconstruction from Single and Multiple Images","volume":"128","author":"Xie","year":"2020","journal-title":"Int. J. Comput. Vis."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Wang, D., Cui, X., Chen, X., Zou, Z., Shi, T., Salcudean, S., Wang, Z.J., and Ward, R. (2021, January 10\u201317). Multi-view 3D Reconstruction with Transformers. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00567"},{"key":"ref_11","unstructured":"Yagubbayli, F., Tonioni, A., and Tombari, F. (2021). LegoFormer: Transformers for Block-by-Block Multi-view 3D Reconstruction. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Peng, K., Islam, R., Quarles, J., and Desai, K. (2022, January 18\u201324). TMVNet: Using Transformers for Multi-View Voxel-Based 3D Reconstruction. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPRW56347.2022.00036"},{"key":"ref_13","first-page":"405","article-title":"NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis","volume":"Volume 12346","author":"Vedaldi","year":"2020","journal-title":"Proceedings of the Computer Vision\u2014ECCV 2020\u201416th European Conference"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Xu, C., Liu, Z., and Li, Z. (2021). Robust Visual-Inertial Navigation System for Low Precision Sensors under Indoor and Outdoor Environments. Remote Sens., 13.","DOI":"10.3390\/rs13040772"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1002\/rob.21732","article-title":"Autonomous aerial navigation using monocular visual-inertial fusion","volume":"35","author":"Lin","year":"2018","journal-title":"J. Field Robot."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1578","DOI":"10.1109\/TPAMI.2019.2954885","article-title":"Image-Based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Era","volume":"43","author":"Han","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_17","unstructured":"Wu, Y., Kirillov, A., Massa, F., Lo, W.Y., and Girshick, R. (2022, October 08). Detectron2. Available online: https:\/\/github.com\/facebookresearch\/detectron2."},{"key":"ref_18","unstructured":"Kar, A., H\u00e4ne, C., and Malik, J. (2017). Learning a Multi-View Stereo Machine. Advances in Neural Information Processing Systems, Curran Associates, Inc."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Sch\u00f6nberger, J.L., and Frahm, J.M. (2016, January 27\u201330). Structure-from-Motion Revisited. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.445"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1007\/s10462-012-9365-8","article-title":"Visual simultaneous localization and mapping: A survey","volume":"43","author":"Ascencio","year":"2015","journal-title":"Artif. Intell. Rev."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Ji, M., Gall, J., Zheng, H., Liu, Y., and Fang, L. (2017, January 22\u201329). SurfaceNet: An End-to-End 3D Neural Network for Multiview Stereopsis. Proceedings of the IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy.","DOI":"10.1109\/ICCV.2017.253"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"4078","DOI":"10.1109\/TPAMI.2020.2996798","article-title":"SurfaceNet+: An End-to-end 3D Neural Network for Very Sparse Multi-View Stereopsis","volume":"43","author":"Ji","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"785","DOI":"10.1007\/978-3-030-01237-3_47","article-title":"MVSNet: Depth Inference for Unstructured Multi-view Stereo","volume":"Volume 11212","author":"Ferrari","year":"2018","journal-title":"Proceedings of the Computer Vision\u2014ECCV 2018\u201415th European Conference"},{"key":"ref_24","first-page":"665","article-title":"RC-MVSNet: Unsupervised Multi-View Stereo with Neural Rendering","volume":"Volume 13691","author":"Avidan","year":"2022","journal-title":"Proceedings of the Computer Vision\u2014ECCV 2022\u201417th European Conference"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Yao, Y., Luo, Z., Li, S., Shen, T., Fang, T., and Quan, L. (2019, January 16\u201320). Recurrent MVSNet for High-Resolution Multi-View Stereo Depth Inference. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00567"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zhao, H., Li, T., Xiao, Y., and Wang, Y. (2020). Improving Multi-Agent Generative Adversarial Nets with Variational Latent Representation. Entropy, 22.","DOI":"10.3390\/e22091055"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"104105","DOI":"10.1016\/j.autcon.2021.104105","article-title":"Near-real-time gradually expanding 3D land surface reconstruction in disaster areas by sequential drone imagery","volume":"135","author":"Cheng","year":"2022","journal-title":"Autom. Constr."},{"key":"ref_28","unstructured":"Ravi, N., Reizenstein, J., Novotny, D., Gordon, T., Lo, W.Y., Johnson, J., and Gkioxari, G. (2020). Accelerating 3D Deep Learning with PyTorch3D. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Collins, J., Goel, S., Deng, K., Luthra, A., Xu, L., Gundogdu, E., Zhang, X., Yago Vicente, T.F., Dideriksen, T., and Arora, H. (2022, January 19\u201320). ABO: Dataset and Benchmarks for Real-World 3D Object Understanding. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.02045"},{"key":"ref_30","unstructured":"Min, P. (2022, October 08). Binvox. 2004\u20132019. Available online: http:\/\/www.patrickmin.com\/binvox."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"191","DOI":"10.1109\/TVCG.2003.1196006","article-title":"Simplification and Repair of Polygonal Models Using Volumetric Techniques","volume":"9","author":"Nooruddin","year":"2003","journal-title":"IEEE Trans. Vis. Comput. Graph."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1162\/neco_a_01246","article-title":"Toward Training Recurrent Neural Networks for Lifelong Learning","volume":"32","author":"Sodhani","year":"2020","journal-title":"Neural Comput."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/24\/9782\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:40:28Z","timestamp":1760146828000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/24\/9782"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,13]]},"references-count":32,"journal-issue":{"issue":"24","published-online":{"date-parts":[[2022,12]]}},"alternative-id":["s22249782"],"URL":"https:\/\/doi.org\/10.3390\/s22249782","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,13]]}}}