{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,17]],"date-time":"2026-01-17T04:28:10Z","timestamp":1768624090105,"version":"3.49.0"},"reference-count":34,"publisher":"MDPI AG","issue":"21","license":[{"start":{"date-parts":[[2023,10,30]],"date-time":"2023-10-30T00:00:00Z","timestamp":1698624000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China","award":["62073004"],"award-info":[{"award-number":["62073004"]}]},{"name":"National Natural Science Foundation of China","award":["GXWD20201231165807007- 20200807164903001"],"award-info":[{"award-number":["GXWD20201231165807007- 20200807164903001"]}]},{"name":"National Natural Science Foundation of China","award":["JCYJ20200109140410340"],"award-info":[{"award-number":["JCYJ20200109140410340"]}]},{"name":"Shenzhen Fundamental Research Program","award":["62073004"],"award-info":[{"award-number":["62073004"]}]},{"name":"Shenzhen Fundamental Research Program","award":["GXWD20201231165807007- 20200807164903001"],"award-info":[{"award-number":["GXWD20201231165807007- 20200807164903001"]}]},{"name":"Shenzhen Fundamental Research Program","award":["JCYJ20200109140410340"],"award-info":[{"award-number":["JCYJ20200109140410340"]}]},{"name":"Science and Technology Plan of Shenzhen","award":["62073004"],"award-info":[{"award-number":["62073004"]}]},{"name":"Science and Technology Plan of Shenzhen","award":["GXWD20201231165807007- 20200807164903001"],"award-info":[{"award-number":["GXWD20201231165807007- 20200807164903001"]}]},{"name":"Science and Technology Plan of Shenzhen","award":["JCYJ20200109140410340"],"award-info":[{"award-number":["JCYJ20200109140410340"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Self-supervised monocular depth estimation, which has attained remarkable progress for outdoor scenes in recent years, often faces greater challenges for indoor scenes. These challenges comprise: (i) non-textured regions: indoor scenes often contain large areas of non-textured regions, such as ceilings, walls, floors, etc., which render the widely adopted photometric loss as ambiguous for self-supervised learning; (ii) camera pose: the sensor is mounted on a moving vehicle in outdoor scenes, whereas it is handheld and moves freely in indoor scenes, which results in complex motions that pose challenges for indoor depth estimation. In this paper, we propose a novel self-supervised indoor depth estimation framework-PMIndoor that addresses these two challenges. We use multiple loss functions to constrain the depth estimation for non-textured regions. We introduce a pose rectified network that only estimates the rotation transformation between two adjacent frames of images for the camera pose problem, and improves the pose estimation results with the pose rectified network loss. We also incorporate a multi-head self-attention module in the depth estimation network to enhance the model\u2019s accuracy. Extensive experiments are conducted on the benchmark indoor dataset NYU Depth V2, demonstrating that our method achieves excellent performance and is better than previous state-of-the-art methods.<\/jats:p>","DOI":"10.3390\/s23218821","type":"journal-article","created":{"date-parts":[[2023,10,30]],"date-time":"2023-10-30T13:26:55Z","timestamp":1698672415000},"page":"8821","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["PMIndoor: Pose Rectified Network and Multiple Loss Functions for Self-Supervised Monocular Indoor Depth Estimation"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-1754-5108","authenticated-orcid":false,"given":"Siyu","family":"Chen","sequence":"first","affiliation":[{"name":"Institute of Artificial Intelligence, University of Science and Technology Beijing, Beijing 100083, China"},{"name":"Key Laboratory of Machine Perception, Shenzhen Graduate School, Peking University, Shenzhen 518055, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ying","family":"Zhu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Machine Perception, Shenzhen Graduate School, Peking University, Shenzhen 518055, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hong","family":"Liu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Machine Perception, Shenzhen Graduate School, Peking University, Shenzhen 518055, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,10,30]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"16940","DOI":"10.1109\/TITS.2022.3160741","article-title":"Towards real-time monocular depth estimation for robotics: A survey","volume":"23","author":"Dong","year":"2022","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Walz, S., Gruber, T., Ritter, W., and Dietmayer, K. (2020, January 20\u201323). Uncertainty depth estimation with gated images for 3D reconstruction. Proceedings of the 2020 IEEE 23rd International Conference on Intelligent Transportation Systems (ITSC), Rhodes, Greece.","DOI":"10.1109\/ITSC45102.2020.9294571"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"978564","DOI":"10.3389\/fpls.2022.978564","article-title":"LANet: Stereo matching network based on linear-attention mechanism for depth estimation optimization in 3D reconstruction of inter-forest scene","volume":"13","author":"Liu","year":"2022","journal-title":"Front. Plant Sci."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Xue, F., Zhuo, G., Huang, Z., Fu, W., Wu, Z., and Ang, M.H. (2020, January 25\u201329). Toward hierarchical self-supervised monocular absolute depth estimation for autonomous driving applications. Proceedings of the 2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA.","DOI":"10.1109\/IROS45743.2020.9340802"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Kalia, M., Navab, N., and Salcudean, T. (2019, January 20\u201324). A real-time interactive augmented reality depth estimation technique for surgical robotics. Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8793610"},{"key":"ref_6","unstructured":"Eigen, D., Puhrsch, C., and Fergus, R. (2014). Depth map prediction from a single image using a multi-scale deep network. Adv. Neural Inf. Process. Syst., 27."},{"key":"ref_7","unstructured":"Bhat, S.F., Alhashim, I., and Wonka, P. (2021, January 20\u201325). Adabins: Depth estimation using adaptive bins. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Li, J., Klein, R., and Yao, A. (2017, January 22\u201329). A two-streamed network for estimating fine-scaled depth maps from single rgb images. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.365"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"3174","DOI":"10.1109\/TCSVT.2017.2740321","article-title":"Estimating depth from monocular images as classification using deep fully convolutional residual networks","volume":"28","author":"Cao","year":"2017","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"2674","DOI":"10.1109\/TCSVT.2019.2929202","article-title":"Monocular depth estimation with augmented ordinal depth relationships","volume":"30","author":"Cao","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"4381","DOI":"10.1109\/TCSVT.2021.3049869","article-title":"Monocular depth estimation using laplacian pyramid-based depth residuals","volume":"31","author":"Song","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Xu, D., Wang, W., Tang, H., Liu, H., Sebe, N., and Ricci, E. (2018, January 18\u201323). Structured attention guided convolutional neural fields for monocular depth estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00412"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Garg, R., Bg, V.K., Carneiro, G., and Reid, I. (2016, January 11\u201314). Unsupervised cnn for single view depth estimation: Geometry to the rescue. Proceedings of the Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands. Proceedings, Part VIII 14.","DOI":"10.1007\/978-3-319-46484-8_45"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Godard, C., Mac Aodha, O., and Brostow, G.J. (2017, January 21\u201326). Unsupervised monocular depth estimation with left-right consistency. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.699"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zhou, T., Brown, M., Snavely, N., and Lowe, D.G. (2017, January 21\u201326). Unsupervised learning of depth and ego-motion from video. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.700"},{"key":"ref_16","unstructured":"Bian, J., Li, Z., Wang, N., Zhan, H., Shen, C., Cheng, M.M., and Reid, I. (2019). Unsupervised scale-consistent depth and ego-motion learning from monocular video. Adv. Neural Inf. Process. Syst., 32."},{"key":"ref_17","unstructured":"Godard, C., Mac Aodha, O., Firman, M., and Brostow, G.J. (November, January 27). Digging into self-supervised monocular depth estimation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_18","unstructured":"Zhou, J., Wang, Y., Qin, K., and Zeng, W. (November, January 27). Moving indoor: Unsupervised video depth learning in challenging environments. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yu, Z., Jin, L., and Gao, S. (2020, January 23\u201328). P2net: Patch-match and plane-regularization for unsupervised indoor depth estimation. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58586-0_13"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Li, B., Huang, Y., Liu, Z., Zou, D., and Yu, W. (2021, January 11\u201317). StructDepth: Leveraging the structural regularities for self-supervised indoor depth estimation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.01243"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"830","DOI":"10.1109\/TCSVT.2022.3207105","article-title":"Monoindoor++: Towards better practice of self-supervised monocular depth estimation for indoor environments","volume":"33","author":"Li","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"9802","DOI":"10.1109\/TPAMI.2021.3136220","article-title":"Auto-rectify network for unsupervised indoor depth estimation","volume":"44","author":"Bian","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Silberman, N., Hoiem, D., Kohli, P., and Fergus, R. (2012, January 7\u201313). Indoor segmentation and support inference from rgbd images. Proceedings of the Computer Vision\u2013ECCV 2012: 12th European Conference on Computer Vision, Florence, Italy. Proceedings, Part V 12.","DOI":"10.1007\/978-3-642-33715-4_54"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Wang, C., Buenaposada, J.M., Zhu, R., and Lucey, S. (2018, January 18\u201323). Learning depth from monocular videos using direct methods. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00216"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"167","DOI":"10.1023\/B:VISI.0000022288.19776.77","article-title":"Efficient graph-based image segmentation","volume":"59","author":"Felzenszwalb","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_26","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhao, W., Liu, S., Shu, Y., and Liu, Y.J. (2020, January 14\u201319). Towards better generalization: Joint depth-pose learning without posenet. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00917"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"824","DOI":"10.1109\/TPAMI.2008.132","article-title":"Make3d: Learning 3d scene structure from a single still image","volume":"31","author":"Saxena","year":"2008","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Liu, M., Salzmann, M., and He, X. (2014, January 23\u201328). Discrete-continuous depth estimation from a single image. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.97"},{"key":"ref_30","unstructured":"Wang, P., Shen, X., Lin, Z., Cohen, S., Price, B., and Yuille, A.L. (2015, January 7\u201312). Towards unified depth and semantic prediction from a single image. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Eigen, D., and Fergus, R. (2015, January 7\u201313). Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.304"},{"key":"ref_32","unstructured":"Chakrabarti, A., Shao, J., and Shakhnarovich, G. (2016). Depth from a single image by harmonizing overcomplete local network predictions. Adv. Neural Inf. Process. Syst., 29."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Laina, I., Rupprecht, C., Belagiannis, V., Tombari, F., and Navab, N. (2016, January 25\u201328). Deeper depth prediction with fully convolutional residual networks. Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV), Stanford, CA, USA.","DOI":"10.1109\/3DV.2016.32"},{"key":"ref_34","unstructured":"Yin, W., Liu, Y., Shen, C., and Yan, Y. (November, January 27). Enforcing geometric constraints of virtual normal for depth prediction. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/21\/8821\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:14:13Z","timestamp":1760130853000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/21\/8821"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,30]]},"references-count":34,"journal-issue":{"issue":"21","published-online":{"date-parts":[[2023,11]]}},"alternative-id":["s23218821"],"URL":"https:\/\/doi.org\/10.3390\/s23218821","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,30]]}}}