{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T18:04:50Z","timestamp":1779300290571,"version":"3.51.4"},"reference-count":40,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2021,4,26]],"date-time":"2021-04-26T00:00:00Z","timestamp":1619395200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Ling Zou","award":["61902201"],"award-info":[{"award-number":["61902201"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>For the task of monocular depth estimation, self-supervised learning supervises training by calculating the pixel difference between the target image and the warped reference image, obtaining results comparable to those with full supervision. However, the problematic pixels in low-texture regions are ignored, since most researchers think that no pixels violate the assumption of camera motion, taking stereo pairs as the input in self-supervised learning, which leads to the optimization problem in these regions. To tackle this problem, we perform photometric loss using the lowest-level feature maps instead and implement first- and second-order smoothing to the depth, ensuring consistent gradients ring optimization. Given the shortcomings of ResNet as the backbone, we propose a new depth estimation network architecture to improve edge location accuracy and obtain clear outline information even in smoothed low-texture boundaries. To acquire more stable and reliable quantitative evaluation results, we introce a virtual data set in the self-supervised task because these have dense depth maps corresponding to pixel by pixel. We achieve performance that exceeds that of the prior methods on both the Eigen Splits of the KITTI and VKITTI2 data sets taking stereo pairs as the input.<\/jats:p>","DOI":"10.3390\/rs13091673","type":"journal-article","created":{"date-parts":[[2021,4,27]],"date-time":"2021-04-27T06:19:11Z","timestamp":1619504351000},"page":"1673","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Self-Supervised Monocular Depth Learning in Low-Texture Areas"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0966-6207","authenticated-orcid":false,"given":"Wanpeng","family":"Xu","sequence":"first","affiliation":[{"name":"Science and Technology on Complex Electronic System Simulation Laboratory, Space Engineering University, Beijing 101416, China"},{"name":"Peng Cheng Laboratory, Shenzhen 518055, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ling","family":"Zou","sequence":"additional","affiliation":[{"name":"Digital Media School, Beijing Film Academy, Beijing 100088, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lingda","family":"Wu","sequence":"additional","affiliation":[{"name":"Science and Technology on Complex Electronic System Simulation Laboratory, Space Engineering University, Beijing 101416, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhipeng","family":"Fu","sequence":"additional","affiliation":[{"name":"Peng Cheng Laboratory, Shenzhen 518055, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,4,26]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 16\u201321). Are we ready for autonomous driving? The kitti vision benchmark suite. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_2","unstructured":"Casser, V., Pirk, S., Mahjourian, R., and Angelova, A. (February, January 27). Depth prediction without the sensors: Leveraging structure for unsupervised learning from monocular videos. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Yang, Z., Wang, P., Wang, Y., Xu, W., and Nevatia, R. (2018, January 18\u201323). Lego: Learning edge with geometry all at once by watching videos. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00031"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Resial Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Yin, Z., and Shi, J. (2018, January 18\u201323). Geonet: Unsupervised learning of dense depth, optical flow and camera pose. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00212"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016, January 27\u201330). Rethinking the inception architecture for computer vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_7","unstructured":"Eigen, D., Puhrsch, C., and Fergus, R. (2014). Depth map prediction from a single image using a multi-scale deep network. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Eigen, D., and Fergus, R. (2015, January 7\u201313). Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.304"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Laina, I., Rupprecht, C., Belagiannis, V., Tombari, F., and Navab, N. (2016, January 25\u201328). Deeper depth prediction with fully convolutional resial networks. Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV), Stanford, CA, USA.","DOI":"10.1109\/3DV.2016.32"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Ramamonjisoa, M.Y., and Lepetit, V. (2020, January 13\u201319). Predicting sharp and accurate occlusion boundaries in monocular depth estimation using displacement fields. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01466"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Zhan, H., Garg, R., Weerasekera, C.S., Li, K., Agarwal, H., and Reid, I. (2018, January 18\u201323). Unsupervised learning of monocular depth estimation and visual odometry with deep feature reconstruction. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00043"},{"key":"ref_12","unstructured":"Chen, W., Fu, Z., Yang, D., and Deng, J. (2016). Single-image depth perception in the wild. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zoran, D., Isola, P., Krishnan, D., and Freeman, W.T. (2015, January 7\u201313). Learning ordinal relationships for mid-level vision. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.52"},{"key":"ref_14","unstructured":"Wu, Y., Ying, S., and Zheng, L. (2018). Size-to-depth: A new perspective for single image depth estimation. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Atapour-Abarghouei, A., and Breckon, T.P. (2018, January 18\u201323). Real-time monocular depth estimation using synthetic data with domain adaptation via image style transfer. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00296"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Garg, R., Bg, V.K., Carneiro, G., and Reid, I. (2016, January 8\u201316). Unsupervised cnn for single view depth estimation: Geometry to the rescue. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_45"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Godard, C., Mac Aodha, O., and Brostow, G.J. (2017, January 21\u201326). Unsupervised monocular depth estimation with left-right consistency. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.699"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Kuznietsov, Y., Stuckler, J., and Leibe, B. (2017, January 21\u201326). Semi-supervised deep learning for monocular depth map prediction. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.238"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Pilzer, A., Xu, D., Puscas, M., Ricci, E., and Sebe, N. (2018, January 5\u20138). Unsupervised adversarial depth estimation using cycled generative networks. Proceedings of the 2018 International Conference on 3D Vision (3DV), Verona, Italy.","DOI":"10.1109\/3DV.2018.00073"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Poggi, M., Aleotti, F., Tosi, F., and Mattoccia, S. (2018, January 1\u20135). Towards real-time unsupervised monocular depth estimation on cpu. Proceedings of the 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain.","DOI":"10.1109\/IROS.2018.8593814"},{"key":"ref_21","unstructured":"Watson, J., Firman, M., Brostow, G.J., and Turmukhambetov, D. (November, January 27). Self-supervised monocular depth hints. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhou, T., Brown, M., Snavely, N., and Lowe, D.G. (2017, January 21\u201326). Unsupervised learning of depth and ego-motion from video. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.700"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zou, Y., Luo, Z., and Huang, J.B. (2018, January 8\u201314). Df-net: Unsupervised joint learning of depth and flow using cross-task consistency. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01228-1_3"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Mahjourian, R., Wicke, M., and Angelova, A. (2018, January 18\u201323). Unsupervised learning of depth and ego-motion from monocular video using 3d geometric constraints. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00594"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"2624","DOI":"10.1109\/TPAMI.2019.2930258","article-title":"Every pixel counts++: Joint learning of geometry and motion with 3d holistic understanding","volume":"42","author":"Luo","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_26","unstructured":"Gordon, A., Li, H., Jonschkowski, R., and Angelova, A. (November, January 27). Depth from videos in the wild: Unsupervised monocular depth learning from unknown cameras. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_27","unstructured":"Godard, C., Mac Aodha, O., Firman, M., and Brostow, G.J. (November, January 27). Digging into self-supervised monocular depth estimation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Shu, C., Yu, K., an, Z., and Yang, K. (2020, January 23\u201328). Feature-metric loss for self-supervised learning of depth and egomotion. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58529-7_34"},{"key":"ref_29","unstructured":"Lopez-Rodriguez, A., and Mikolajczyk, K. (2020). DESC: Domain Adaptation for Depth Estimation via Semantic Consistency. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Guizilini, V., Ambrus, R., Pillai, S., Raventos, A., and Gaidon, A. (2020, January 13\u201319). 3d packing for self-supervised monocular depth estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00256"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Lyu, X., Liu, L., Wang, M., Kong, X., Liu, L., Liu, Y., Chen, X., and Yuan, Y. (2020). HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation. arXiv.","DOI":"10.1609\/aaai.v35i3.16329"},{"key":"ref_32","unstructured":"Duta, I.C., Li, L., Zhu, F., and Ling, S. (2020). Improved residual networks for image and video recognition. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 8\u201316). Identity mappings in deep resial networks. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"ref_34","unstructured":"Zhou, X.Y., Sun, J., Ye, N., Lan, X., Luo, Q., Lai, B.L., Esperanca, P., Yang, G.Z., and Li, Z. (2020). Batch Group Normalization. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Zhou, Z., Siddiquee, M.M.R., Tajbakhsh, N., and Liang, J. (2018). Unet++: A nested u-net architecture for medical image segmentation. Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support, Springer.","DOI":"10.1007\/978-3-030-00889-5_1"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","article-title":"Image quality assessment: From error visibility to structural similarity","volume":"13","author":"Wang","year":"2004","journal-title":"IEEE Trans. Image Process."},{"key":"ref_37","unstructured":"Geiger, A., Lenz, P., Stiller, C., and Urtasun, R. (2020, November 22). The KITTI Vision Benchmark Suite, 2015, 2. Available online: http:\/\/www.Cvlibs.Net\/datasets\/kitti."},{"key":"ref_38","unstructured":"Cabon, Y., Murray, N., and Humenberger, M. (2020). Virtual kitti 2. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. (2016, January 27\u201330). The Cityscapes Dataset for Semantic Urban Scene Understanding. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.350"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"824","DOI":"10.1109\/TPAMI.2008.132","article-title":"Make3D: Learning 3D Scene Structure from a Single Still Image","volume":"31","author":"Saxena","year":"2009","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/9\/1673\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:52:49Z","timestamp":1760161969000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/9\/1673"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,26]]},"references-count":40,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2021,5]]}},"alternative-id":["rs13091673"],"URL":"https:\/\/doi.org\/10.3390\/rs13091673","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,4,26]]}}}