{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:27:34Z","timestamp":1760146054586,"version":"build-2065373602"},"reference-count":37,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2024,9,16]],"date-time":"2024-09-16T00:00:00Z","timestamp":1726444800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Reconstructing 3D indoor scenes from 2D images has always been an important task in computer vision and graphics applications. For indoor scenes, traditional 3D reconstruction methods have problems such as missing surface details, poor reconstruction of large plane textures and uneven illumination areas, and many wrongly reconstructed floating debris noises in the reconstructed models. This paper proposes a 3D reconstruction method for indoor scenes that combines neural radiation field (NeRFs) and signed distance function (SDF) implicit expressions. The volume density of the NeRF is used to provide geometric information for the SDF field, and the learning of geometric shapes and surfaces is strengthened by adding an adaptive normal prior optimization learning process. It not only preserves the high-quality geometric information of the NeRF, but also uses the SDF to generate an explicit mesh with a smooth surface, significantly improving the reconstruction quality of large plane textures and uneven illumination areas in indoor scenes. At the same time, a new regularization term is designed to constrain the weight distribution, making it an ideal unimodal compact distribution, thereby alleviating the problem of uneven density distribution and achieving the effect of floating debris removal in the final model. Experiments show that the 3D reconstruction effect of this paper on ScanNet, Hypersim, and Replica datasets outperforms the state-of-the-art methods.<\/jats:p>","DOI":"10.3390\/jimaging10090231","type":"journal-article","created":{"date-parts":[[2024,9,16]],"date-time":"2024-09-16T07:36:00Z","timestamp":1726472160000},"page":"231","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Three-Dimensional Reconstruction of Indoor Scenes Based on Implicit Neural Representation"],"prefix":"10.3390","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3192-6602","authenticated-orcid":false,"given":"Zhaoji","family":"Lin","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, Sanjiang University, Nanjing 210012, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yutao","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Southeast University, Nanjing 211189, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2930-8407","authenticated-orcid":false,"given":"Li","family":"Yao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Southeast University, Nanjing 211189, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,9,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Kang, Z., Yang, J., Yang, Z., and Cheng, S. (2020). A review of techniques for 3d reconstruction of indoor environments. ISPRS Int. J. Geo-Inf., 9.","DOI":"10.3390\/ijgi9050330"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"369","DOI":"10.1007\/s41095-021-0250-8","article-title":"High-quality indoor scene 3d reconstruction with rgb-d cameras: A brief review","volume":"8","author":"Li","year":"2022","journal-title":"Comput. Vis. Media"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Hess, W., Kohler, D., Rapp, H., and Andor, D. (2016, January 16\u201321). Real-Time Loop Closure in 2d Lidar Slam. Proceedings of the2016 IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden.","DOI":"10.1109\/ICRA.2016.7487258"},{"key":"ref_4","first-page":"1","article-title":"Loam: Lidar odometry and mapping in real-time","volume":"Volume 2","author":"Zhang","year":"2014","journal-title":"Robotics: Science and Systems"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Shan, T., and Englot, B. (2018, January 1\u20135). Lego-loam: Lightweight and Ground-Optimized Lidar Odometry and Mapping on Variable Terrain. Proceedings of the 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain.","DOI":"10.1109\/IROS.2018.8594299"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Curless, B., and Levoy, M. (1996, January 4\u20139). A Volumetric Method for Building Complex Models from Range Images. Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques, New Orleans, LA, USA.","DOI":"10.1145\/237170.237269"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Newcombe, R.A., Izadi, S., Hilliges, O., Molyneaux, D., Kim, D., Davison, A.J., Kohi, P., Shotton, J., Hodges, S., and Fitzgibbon, A. (2011, January 26\u201329). Kinectfusion: Real-Time Dense Surface Mapping and Tracking. Proceedings of the 2011 10th IEEE International Symposium on Mixed and Augmented Reality, Basel, Switzerland.","DOI":"10.1109\/ISMAR.2011.6092378"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Murez, Z., Van As, T., Bartolozzi, J., Sinha, A., Badrinarayanan, V., and Rabinovich, A. (2020). Atlas: End-to-end 3d scene reconstruction from posed images. Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part VII 16, Springer.","DOI":"10.1007\/978-3-030-58571-6_25"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Sun, J., Xie, Y., Chen, L., Zhou, X., and Bao, H. (2021, January 20\u201325). Neuralrecon: Real-Time Coherent 3d Reconstruction from Monocular Video. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01534"},{"key":"ref_10","unstructured":"Schonberger, J.L., and Frahm, J.M. (July, January 26). Structure-from-Motion Revisited. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_11","unstructured":"Sch\u00f6nberger, J.L., Zheng, E., Frahm, J.M., and Pollefeys, M. (2016). Pixelwise View Selection for Unstructured Multi-View Stereo. Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11\u201314, 2016, Proceedings, Part III 14, Springer."},{"key":"ref_12","first-page":"12516","article-title":"Planar prior assisted patchmatch multi-view stereo","volume":"34","author":"Xu","year":"2020","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_13","unstructured":"Im, S., Jeon, H.G., Lin, S., and Kweon, I.S. (2019). Dpsnet: End-to-end deep plane sweep stereo. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Wang, F., Galliani, S., Vogel, C., Speciale, P., and Pollefeys, M. (2021, January 20\u201325). Patchmatchnet: Learned Multi-View Patchmatch Stereo. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01397"},{"key":"ref_15","unstructured":"Xu, Q., and Tao, W. (2020). Pvsnet: Pixelwise visibility-aware multi-view stereo network. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Yao, Y., Luo, Z., Li, S., Fang, T., and Quan, L. (2018, January 8\u201314). Mvsnet: Depth Inference for Unstructured Multi-View Stereo. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01237-3_47"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Yu, Z., and Gao, S. (2020, January 13\u201319). Fast-Mvsnet: Sparse-to-Dense Multi-View Stereo with Learned Propagation and Gauss-Newton Refinement. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00202"},{"key":"ref_18","unstructured":"Teed, Z., and Deng, J. (2018). Deepv2d: Video to depth with differentiable structure from motion. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Huang, P.H., Matzen, K., Kopf, J., Ahuja, N., and Huang, J.B. (2018, January 18\u201322). Deepmvs: Learning Multi-View Stereopsis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00298"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Cheng, S., Xu, Z., Zhu, S., Li, Z., Li, L.E., Ramamoorthi, R., and Su, H. (2020, January 13\u201319). Deep Stereo using Adaptive thin Volume Representation with Uncertainty Awareness. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00260"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2396","DOI":"10.11834\/jig.220376","article-title":"The growth of image-related three dimensional reconstruction techniquesin deep learning-driven era: A critical summary","volume":"28","author":"Yang","year":"2023","journal-title":"J. Image Graph."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Liu, S., Zhang, Y., Peng, S., Shi, B., Pollefeys, M., and Cui, Z. (2020, January 13\u201319). Dist: Rendering deep implicit signed distance function with differentiable sphere tracing. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00209"},{"key":"ref_23","first-page":"2492","article-title":"Multiview neural surface reconstruction by disentangling geometry and appearance","volume":"33","author":"Yariv","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_24","unstructured":"Wang, P., Liu, L., Liu, Y., Theobalt, C., Komura, T., and Wang, W. (2021). Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. arXiv."},{"key":"ref_25","first-page":"4805","article-title":"Volume rendering of neural implicit surfaces","volume":"34","author":"Yariv","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Barron, J.T., Mildenhall, B., Verbin, D., Srinivasan, P.P., and Hedman, P. (2022, January 19\u201324). Mip-nerf 360: Unbounded Anti-Aliased Neural Radiance Fields. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00539"},{"key":"ref_27","first-page":"15651","article-title":"Neural sparse voxel fields","volume":"33","author":"Liu","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Rebain, D., Matthews, M., Yi, K.M., Lagun, D., and Tagliasacchi, A. (2022, January 19\u201324). Lolnerf: Learn from one look. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00161"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Gafni, G., Thies, J., Zollhofer, M., and Nie\u00dfner, M. (2021, January 20\u201325). Dynamic Neural Radiance Fields for Monocular 4d Facial Avatar Reconstruction. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00854"},{"key":"ref_30","unstructured":"Do, T., Vuong, K., Roumeliotis, S.I., and Park, H.S. (2020). Surface Normal Estimation of Tilted Images Via Spatial Rectifier. Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part IV 16, Springer International Publishing."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Dai, A., Chang, A.X., Savva, M., Halber, M., Funkhouser, T., and Nie\u00dfner, M. (2017, January 21\u201326). Scannet: Richly-Annotated 3d Reconstructions of Indoor Scenes. Proceedings of the IEEE Conference on Computer vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.261"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Roberts, M., Ramapuram, J., Ranjan, A., Kumar, A., Bautista, M.A., Paczan, N., Webb, R., and Susskind, J.M. (2021, January 10\u201317). Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01073"},{"key":"ref_33","unstructured":"Straub, J., Whelan, T., Ma, L., Chen, Y., Wijmans, E., Green, S., Engel, J.J., Mur-Artal, R., Ren, C., and Verma, S. (2019). The replica dataset: A digital replica of indoor spaces. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1145\/3503250","article-title":"Nerf: Representing scenes as neural radiance fields for view synthesis","volume":"65","author":"Mildenhall","year":"2021","journal-title":"Commun. ACM"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Guo, H., Peng, S., Lin, H., Wang, Q., Zhang, G., Bao, H., and Zhou, X. (2022, January 19\u201324). Neural 3d scene reconstruction with the Manhattan-world assumption. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00543"},{"key":"ref_36","first-page":"25018","article-title":"Monosdf: Exploring monocular geometric cues for neural implicit surface reconstruction","volume":"35","author":"Yu","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Zhu, J., Huo, Y., Ye, Q., Luan, F., Li, J., Xi, D., Wang, L., Tang, R., Hua, W., and Bao, H. (2023, January 18\u201322). I2-SDF: Intrinsic Indoor Scene Reconstruction and Editing via Raytracing in Neural SDFs. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01202"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/10\/9\/231\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:57:24Z","timestamp":1760111844000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/10\/9\/231"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,16]]},"references-count":37,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2024,9]]}},"alternative-id":["jimaging10090231"],"URL":"https:\/\/doi.org\/10.3390\/jimaging10090231","relation":{},"ISSN":["2313-433X"],"issn-type":[{"type":"electronic","value":"2313-433X"}],"subject":[],"published":{"date-parts":[[2024,9,16]]}}}