{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T03:26:42Z","timestamp":1784690802402,"version":"3.55.0"},"reference-count":47,"publisher":"MDPI AG","issue":"17","license":[{"start":{"date-parts":[[2022,9,1]],"date-time":"2022-09-01T00:00:00Z","timestamp":1661990400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62125102"],"award-info":[{"award-number":["62125102"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>The remote sensing 3D reconstruction of mountain areas has a wide range of applications in surveying, visualization, and game modeling. Different from indoor objects, outdoor mountain reconstruction faces additional challenges, including illumination changes, diversity of textures, and highly irregular surface geometry. Traditional neural network-based methods that lack discriminative features struggle to handle the above challenges, and thus tend to generate incomplete and inaccurate reconstructions. Truncated signed distance function (TSDF) is a commonly used parameterized representation of 3D structures, which is naturally convenient for neural network computation and computer storage. In this paper, we propose a novel deep learning method with TSDF-based representations for robust 3D reconstruction from images containing mountain terrains. The proposed method takes in a set of images captured around an outdoor mountain and produces high-quality TSDF representations of the mountain areas. To address the aforementioned challenges, such as lighting variations and texture diversity, we propose a view fusion strategy based on reweighted mechanisms (VRM) to better integrate multi-view 2D features of the same voxel. A feature enhancement (FE) module is designed for providing better discriminative geometry prior in the feature decoding process. We also propose a spatial\u2013temporal aggregation (STA) module to reduce the ambiguity between temporal features and improve the accuracy of the reconstruction surfaces. A synthetic dataset for reconstructing images containing mountain terrains is built. Our method outperforms the previous state-of-the-art TSDF-based and depth-based reconstruction methods in terms of both 2D and 3D metrics. Furthermore, we collect real-world multi-view terrain images from Google Map. Qualitative results demonstrate the good generalization ability of the proposed method.<\/jats:p>","DOI":"10.3390\/rs14174333","type":"journal-article","created":{"date-parts":[[2022,9,2]],"date-time":"2022-09-02T00:19:01Z","timestamp":1662077941000},"page":"4333","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["3D Reconstruction of Remote Sensing Mountain Areas with TSDF-Based Neural Networks"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8571-6942","authenticated-orcid":false,"given":"Zipeng","family":"Qi","sequence":"first","affiliation":[{"name":"Image Processing Center, School of Astronautics, Beihang University, Beijing 100191, China"},{"name":"Beijing Key Laboratory of Digital Media, Beihang University, Beijing 100191, China"},{"name":"State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhengxia","family":"Zou","sequence":"additional","affiliation":[{"name":"Department of Guidance, Navigation and Control, School of Astronautics, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6418-3761","authenticated-orcid":false,"given":"Hao","family":"Chen","sequence":"additional","affiliation":[{"name":"Image Processing Center, School of Astronautics, Beihang University, Beijing 100191, China"},{"name":"Beijing Key Laboratory of Digital Media, Beihang University, Beijing 100191, China"},{"name":"State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4772-3172","authenticated-orcid":false,"given":"Zhenwei","family":"Shi","sequence":"additional","affiliation":[{"name":"Image Processing Center, School of Astronautics, Beihang University, Beijing 100191, China"},{"name":"Beijing Key Laboratory of Digital Media, Beihang University, Beijing 100191, China"},{"name":"State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,9,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Li, H., Chen, S., Wang, Z., and Li, W. (2010, January 25\u201330). Fusion of LiDAR data and orthoimage for automatic building reconstruction. Proceedings of the IEEE International Geoscience and Remote Sensing Symposium, IGARSS 2010, Honolulu, HI, USA.","DOI":"10.1109\/IGARSS.2010.5654163"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Xiong, B., Jiang, W., Li, D., and Qi, M. (2021). Voxel Grid-Based Fast Registration of Terrestrial Point Cloud. Remote Sens., 13.","DOI":"10.3390\/rs13101905"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"4700","DOI":"10.1109\/JSTARS.2016.2543301","article-title":"Reconstruction of the Radar Image From Actual DDMs Collected by TechDemoSat-1 GNSS-R Mission","volume":"9","author":"Schiavulli","year":"2016","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1063","DOI":"10.1109\/JSTARS.2019.2903398","article-title":"Corrections to \u201cRegularization of SAR Tomography for 3-D Height Reconstruction in Urban Areas\u201d [Feb 19 648\u2013659]","volume":"12","author":"Aghababaee","year":"2019","journal-title":"Sel. Top. Appl. Earth Obs. Remote Sens. IEEE J."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"4476","DOI":"10.1109\/JSTARS.2020.3014696","article-title":"CSR-Net: A Novel Complex-valued Network for Fast and Precise 3-D Microwave Sparse Reconstruction","volume":"13","author":"Wang","year":"2020","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Yang, Y.C., Lu, C.Y., Huang, S.J., Yang, T.Z., Chang, Y.C., and Ho, C.R. (2022). On the Reconstruction of Missing Sea Surface Temperature Data from Himawari-8 in Adjacent Waters of Taiwan Using DINEOF Conducted with 25-h Data. Remote Sens., 14.","DOI":"10.3390\/rs14122818"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zhang, E., Fu, Y., Wang, J., Liu, L., Yu, K., and Peng, J. (2022). MSAC-Net: 3D Multi-Scale Attention Convolutional Network for Multi-Spectral Imagery Pansharpening. Remote Sens., 14.","DOI":"10.3390\/rs14122761"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1109\/MRA.2011.2181769","article-title":"Terrain Reconstruction of Glacial Surfaces: Robotic Surveying Techniques","volume":"19","author":"Williams","year":"2012","journal-title":"Robot. Autom. Mag. IEEE"},{"key":"ref_9","unstructured":"Kazhdan, M., Bolitho, M., and Hoppe, H. (2006, January 26\u201328). Poisson surface reconstruction. Proceedings of the Fourth Eurographics Symposium on Geometry Processing, Cagliari Sardinia, Italy."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"163","DOI":"10.1145\/37402.37422","article-title":"Marching cubes: A high resolution 3D surface construction algorithm","volume":"21","author":"Lorensen","year":"1987","journal-title":"ACM Siggraph Comput. Graph."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Murez, Z., van As, T., Bartolozzi, J., Sinha, A., Badrinarayanan, V., and Rabinovich, A. (2020, January 23\u201328). Atlas: End-to-end 3d scene reconstruction from posed images. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Part VII 16.","DOI":"10.1007\/978-3-030-58571-6_25"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Choe, J., Im, S., Rameau, F., Kang, M., and Kweon, I.S. (2021, January 11). Volumefusion: Deep depth fusion for 3d scene reconstruction. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01578"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Sun, J., Xie, Y., Chen, L., Zhou, X., and Bao, H. (2021, January 20\u201325). NeuralRecon: Real-Time Coherent 3D Reconstruction from Monocular Video. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01534"},{"key":"ref_14","first-page":"1403","article-title":"Transformerfusion: Monocular rgb scene reconstruction using transformers","volume":"34","author":"Bozic","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_15","unstructured":"Chang, A.X., Funkhouser, T., Guibas, L., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., and Su, H. (2015). ShapeNet: An Information-Rich 3D Model Repository. arXiv."},{"key":"ref_16","unstructured":"Li, W., Saeedi, S., McCormac, J., Clark, R., Tzoumanikas, D., Ye, Q., Huang, Y., Tang, R., and Leutenegger, S. (2018). Interiornet: Mega-scale multi-sensor photo-realistic indoor scenes dataset. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Grinvald, M., Tombari, F., Siegwart, R., and Nieto, J. (June, January 30). TSDF++: A Multi-Object Formulation for Dynamic Object Tracking and Reconstruction. Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi\u2019an, China.","DOI":"10.1109\/ICRA48506.2021.9560923"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Kim, H., and Lee, B. (June, January 30). Probabilistic TSDF Fusion Using Bayesian Deep Learning for Dense 3D Reconstruction with a Single RGB Camera. Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Xi\u2019an, China.","DOI":"10.1109\/ICRA40945.2020.9196663"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"15501","DOI":"10.1016\/j.ifacol.2020.12.2376","article-title":"Surface-driven Next-Best-View planning for exploration of large-scale 3D environments","volume":"53","author":"Hardouin","year":"2020","journal-title":"IFAC-PapersOnLine"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Yao, Y., Luo, Z., Li, S., Fang, T., and Quan, L. (2018, January 8\u201314). Mvsnet: Depth inference for unstructured multi-view stereo. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01237-3_47"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Wang, F., Galliani, S., Vogel, C., Speciale, P., and Pollefeys, M. (2021, January 20\u201325). PatchmatchNet: Learned Multi-View Patchmatch Stereo. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01397"},{"key":"ref_22","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Newcombe, R.A., Izadi, S., Hilliges, O., Molyneaux, D., Kim, D., Davison, A.J., Kohi, P., Shotton, J., Hodges, S., and Fitzgibbon, A. (2011, January 26\u201329). Kinectfusion: Real-time dense surface mapping and tracking. Proceedings of the 2011 10th IEEE International Symposium on Mixed and Augmented Reality, Basel, Switzerland.","DOI":"10.1109\/ISMAR.2011.6092378"},{"key":"ref_24","unstructured":"Im, S., Jeon, H.G., Lin, S., and Kweon, I.S. (2019). Dpsnet: End-to-end deep plane sweep stereo. arXiv."},{"key":"ref_25","unstructured":"Hou, Y., Kannala, J., and Solin, A. (November, January 27). Multi-view stereo by temporal nonparametric fusion. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wei, Y., Liu, S., Rao, Y., Zhao, W., Lu, J., and Zhou, J. (2021, January 11). Nerfingmvs: Guided optimization of neural radiance fields for indoor multi-view stereo. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00556"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2487228.2487237","article-title":"Screened poisson surface reconstruction","volume":"32","author":"Kazhdan","year":"2013","journal-title":"ACM Trans. Graph. (ToG)"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"2275","DOI":"10.1111\/j.1467-8659.2009.01530.x","article-title":"Robust and efficient surface reconstruction from range data","volume":"Volume 28","author":"Labatut","year":"2009","journal-title":"Computer Graphics Forum"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Weder, S., Schonberger, J.L., Pollefeys, M., and Oswald, M.R. (2021, January 20\u201325). NeuralFusion: Online Depth Fusion in Latent Space. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00318"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Mildenhall, B., Srinivasan, P.P., Tancik, M., Barron, J.T., Ramamoorthi, R., and Ng, R. (2020, January 23\u201328). Nerf: Representing scenes as neural radiance fields for view synthesis. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58452-8_24"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Martin-Brualla, R., Radwan, N., Sajjadi, M.S., Barron, J.T., Dosovitskiy, A., and Duckworth, D. (2021, January 20\u201325). Nerf in the wild: Neural radiance fields for unconstrained photo collections. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00713"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Chen, Y., Liu, S., and Wang, X. (2021, January 20\u201325). Learning continuous image representation with local implicit image function. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00852"},{"key":"ref_33","unstructured":"Xu, X., Wang, Z., and Shi, H. (2021). UltraSR: Spatial Encoding is a Missing Key for Implicit Image Function-based Arbitrary-Scale Super-Resolution. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Skorokhodov, I., Ignatyev, S., and Elhoseiny, M. (2021, January 20\u201325). Adversarial generation of continuous images. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01061"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Anokhin, I., Demochkin, K., Khakhulin, T., Sterkin, G., and Korzhenkov, D. (2021, January 20\u201325). Image Generators with Conditionally-Independent Pixel Synthesis. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01405"},{"key":"ref_36","unstructured":"Dupont, E., Teh, Y.W., and Doucet, A. (2021). Generative Models as Distributions of Functions. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Chen, Z., and Zhang, H. (2019, January 15\u201320). Learning Implicit Fields for Generative Shape Modeling. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00609"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Mescheder, L., Oechsle, M., Niemeyer, M., Nowozin, S., and Geiger, A. (2019, January 15\u201320). Occupancy Networks: Learning 3D Reconstruction in Function Space. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00459"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Park, J., Florence, P., Straub, J., Newcombe, R., and Lovegrove, S. (2019, January 15\u201320). DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00025"},{"key":"ref_40","unstructured":"Tancik, M., Srinivasan, P.P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Singhal, U., Ramamoorthi, R., Barron, J.T., and Ng, R. (2020). Fourier features let networks learn high frequency functions in low dimensional domains. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_42","unstructured":"Wen, Z., Lin, W., Wang, T., and Xu, G. (2021). Distract Your Attention: Multi-head Cross Attention Network for Facial Expression Recognition. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Xu, X., and Hao, J. (2022). U-Former: Improving Monaural Speech Enhancement with Multi-head Self and Cross Attention. arXiv.","DOI":"10.1109\/ICPR56361.2022.9956638"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Tan, M., Chen, B., Pang, R., Vasudevan, V., Sandler, M., Howard, A., and Le, Q.V. (2019, January 15\u201320). Mnasnet: Platform-aware neural architecture search for mobile. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00293"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Tang, H., Liu, Z., Zhao, S., Lin, Y., Lin, J., Wang, H., and Han, S. (2020, January 23\u201328). Searching efficient 3d architectures with sparse point-voxel convolution. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58604-1_41"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Schonberger, J.L., and Frahm, J.M. (2016, January 27\u201330). Structure-from-motion revisited. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.445"},{"key":"ref_47","unstructured":"Eigen, D., Puhrsch, C., and Fergus, R. (2014). Depth map prediction from a single image using a multi-scale deep network. arXiv."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/14\/17\/4333\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T00:21:55Z","timestamp":1760142115000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/14\/17\/4333"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,1]]},"references-count":47,"journal-issue":{"issue":"17","published-online":{"date-parts":[[2022,9]]}},"alternative-id":["rs14174333"],"URL":"https:\/\/doi.org\/10.3390\/rs14174333","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,1]]}}}