{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,12]],"date-time":"2025-11-12T03:29:57Z","timestamp":1762918197980,"version":"build-2065373602"},"reference-count":42,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2021,1,30]],"date-time":"2021-01-30T00:00:00Z","timestamp":1611964800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61671153"],"award-info":[{"award-number":["61671153"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Estimating the depth of image and egomotion of agent are important for autonomous and robot in understanding the surrounding environment and avoiding collision. Most existing unsupervised methods estimate depth and camera egomotion by minimizing photometric error between adjacent frames. However, the photometric consistency sometimes does not meet the real situation, such as brightness change, moving objects and occlusion. To reduce the influence of brightness change, we propose a feature pyramid matching loss (FPML) which captures the trainable feature error between a current and the adjacent frames and therefore it is more robust than photometric error. In addition, we propose the occlusion-aware mask (OAM) network which can indicate occlusion according to change of masks to improve estimation accuracy of depth and camera pose. The experimental results verify that the proposed unsupervised approach is highly competitive against the state-of-the-art methods, both qualitatively and quantitatively. Specifically, our method reduces absolute relative error (Abs Rel) by 0.017\u20130.088.<\/jats:p>","DOI":"10.3390\/s21030923","type":"journal-article","created":{"date-parts":[[2021,1,30]],"date-time":"2021-01-30T06:22:20Z","timestamp":1611987740000},"page":"923","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Unsupervised Learning of Depth and Camera Pose with Feature Map Warping"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6772-3954","authenticated-orcid":false,"given":"Ente","family":"Guo","sequence":"first","affiliation":[{"name":"College of Physics and Information Engineering, Fuzhou University, Fuzhou 350108, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhifeng","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Physics and Information Engineering, Fuzhou University, Fuzhou 350108, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yanlin","family":"Zhou","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, University of Florida, Gainesville, FL 32611, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dapeng Oliver","family":"Wu","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, University of Florida, Gainesville, FL 32611, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,1,30]]},"reference":[{"key":"ref_1","unstructured":"Pierce, J.S., Agrawala, M., and Klemmer, S.R. (2011, January 16\u201319). KinectFusion: Real-time 3D reconstruction and interaction using a moving depth camera. Proceedings of the 24th Annual ACM Symposium on User Interface Software and Technology, Santa Barbara, CA, USA."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Lin, J., and Zhang, F. (August, January 31). Loam livox: A fast, robust, high-precision LiDAR odometry and mapping package for LiDARs of small FoV. Proceedings of the 2020 IEEE International Conference on Robotics and Automation, ICRA 2020, Paris, France.","DOI":"10.1109\/ICRA40945.2020.9197440"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Rosin, P.L., Lai, Y.K., Shao, L., and Liu, Y. (2019). RGB-D Image Analysis and Processing, Springer.","DOI":"10.1007\/978-3-030-28603-3"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Zhang, J., and Singh, S. (2015, January 26\u201330). Visual-lidar odometry and mapping: Low-drift, robust, and fast. Proceedings of the IEEE International Conference on Robotics and Automation, ICRA 2015, Seattle, WA, USA.","DOI":"10.1109\/ICRA.2015.7139486"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Bloesch, M., Czarnowski, J., Clark, R., Leutenegger, S., and Davison, A.J. (2018, January 18\u201322). CodeSLAM\u2014Learning a Compact, Optimisable Representation for Dense Visual SLAM. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00271"},{"key":"ref_6","unstructured":"Bian, J., Li, Z., Wang, N., Zhan, H., Shen, C., Cheng, M., and Reid, I.D. (2019;, January 8\u201314). Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video. Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, Vancouver, BC, Canada."},{"key":"ref_7","unstructured":"Eigen, D., Puhrsch, C., and Fergus, R. (2014;, January 8\u201313). Depth Map Prediction from a Single Image using a Multi-Scale Deep Network. Proceedings of the Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, Montreal, QC, Canada."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Eigen, D., and Fergus, R. (2015, January 7\u201313). Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-scale Convolutional Architecture. Proceedings of the 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.304"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Godard, C., Aodha, O.M., and Brostow, G.J. (2017, January 21\u201326). Unsupervised Monocular Depth Estimation with Left-Right Consistency. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.699"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhou, T., Brown, M., Snavely, N., and Lowe, D.G. (2017, January 21\u201326). Unsupervised Learning of Depth and Ego-Motion from Video. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.700"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Yin, Z., and Shi, J. (2018, January 18\u201322). GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00212"},{"key":"ref_12","unstructured":"Vijayanarasimhan, S., Ricco, S., Schmid, C., Sukthankar, R., and Fragkiadaki, K. (2017). Sfm-net: Learning of structure and motion from video. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Casser, V., Pirk, S., Mahjourian, R., and Angelova, A. (February, January 27). Depth Prediction without the Sensors: Leveraging Structure for Unsupervised Learning from Monocular Videos. Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, Honolulu, Hawaii, USA.","DOI":"10.1609\/aaai.v33i01.33018001"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Yang, Z., Wang, P., Wang, Y., Xu, W., and Nevatia, R. (2018, January 18\u201322). LEGO: Learning Edge With Geometry All at Once by Watching Videos. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00031"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Yang, Z., Wang, P., Xu, W., Zhao, L., and Nevatia, R. (2018, January 2\u20137). Unsupervised Learning of Geometry From Videos With Edge-Aware Depth-Normal Consistency. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12257"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Mahjourian, R., Wicke, M., and Angelova, A. (2018, January 18\u201322). Unsupervised Learning of Depth and Ego-Motion From Monocular Video Using 3D Geometric Constraints. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00594"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Wang, C., Buenaposada, J.M., Zhu, R., and Lucey, S. (2018, January 18\u201322). Learning Depth From Monocular Videos Using Direct Methods. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00216"},{"key":"ref_18","first-page":"38","article-title":"DF-Net: Unsupervised Joint Learning of Depth and Flow Using Cross-Task Consistency","volume":"Volume 11209","author":"Ferrari","year":"2018","journal-title":"Computer Vision\u2013ECCV 2018, Proceedings of the 15th European Conference, Munich, Germany, 8\u201314 September 2018"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yang, N., von Stumberg, L., Wang, R., and Cremers, D. (2020, January 13\u201319). D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00136"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","article-title":"Image quality assessment: From error visibility to structural similarity","volume":"13","author":"Wang","year":"2004","journal-title":"IEEE Trans. Image Process."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Li, J., Klein, R., and Yao, A. (2016). Learning fine-scaled depth maps from single rgb images. arXiv.","DOI":"10.1109\/ICCV.2017.365"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Laina, I., Rupprecht, C., Belagiannis, V., Tombari, F., and Navab, N. (2016, January 25\u201328). Deeper Depth Prediction with Fully Convolutional Residual Networks. Proceedings of the Fourth International Conference on 3D Vision, 3DV 2016, Stanford, CA, USA.","DOI":"10.1109\/3DV.2016.32"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Xu, D., Ricci, E., Ouyang, W., Wang, X., and Sebe, N. (2017, January 21\u201326). Multi-scale Continuous CRFs as Sequential Deep Networks for Monocular Depth Estimation. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.25"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"942","DOI":"10.1007\/s11263-018-1082-6","article-title":"What Makes Good Synthetic Training Data for Learning Disparity and Optical Flow Estimation?","volume":"126","author":"Mayer","year":"2018","journal-title":"Int. J. Comput. Vis."},{"key":"ref_25","first-page":"740","article-title":"Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue","volume":"Volume 9912","author":"Leibe","year":"2016","journal-title":"Computer Vision\u2013ECCV 2016, Proceedings of the 14th European Conference, Amsterdam, The Netherlands, 11\u201314 October 2016"},{"key":"ref_26","unstructured":"Harltey, A., and Zisserman, A. (2006). Multiple View Geometry in Computer Vision, Cambridge University Press. [2nd ed.]."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R.B. (2017, January 22\u201329). Mask R-CNN. Proceedings of the IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_28","first-page":"691","article-title":"Every Pixel Counts: Unsupervised Geometry Learning with Holistic 3D Motion Understanding","volume":"Volume 11133","author":"Roth","year":"2018","journal-title":"Computer Vision\u2014ECCV 2018 Workshops, Munich, Germany, 8\u201314 September 2018"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Godard, C., Aodha, O.M., Firman, M., and Brostow, G.J. (November, January 27). Digging Into Self-Supervised Monocular Depth Estimation. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00393"},{"key":"ref_30","first-page":"713","article-title":"Unsupervised Learning of Multi-Frame Optical Flow with Occlusions","volume":"Volume 11220","author":"Ferrari","year":"2018","journal-title":"Computer Vision\u2014ECCV 2018, Proceedings of the 15th European Conference, Munich, Germany, 8\u201314 September 2018"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Sun, D., Yang, X., Liu, M., and Kautz, J. (2018, January 18\u201322). PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00931"},{"key":"ref_32","unstructured":"Jaderberg, M., Simonyan, K., Zisserman, A., and Kavukcuoglu, K. (2015, January 7\u201312). Spatial Transformer Networks. Proceedings of the Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, Montreal, QC, Canada."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_34","unstructured":"Clevert, D., Unterthiner, T., and Hochreiter, S. (2016, January 2\u20134). Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs). Proceedings of the 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico."},{"key":"ref_35","unstructured":"Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017, January 4\u20139). Automatic Differentiation in Pytorch. Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_36","unstructured":"Kingma, D.P., and Ba, J. (2015, January 7\u20139). Adam: A Method for Stochastic Optimization. Proceedings of the 3rd International Conference on Learning Representations (ICLR 2015), San Diego, CA, USA."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 16\u201321). Are we ready for autonomous driving? The KITTI vision benchmark suite. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Sturm, J., Engelhard, N., Endres, F., Burgard, W., and Cremers, D. (2012, January 7\u201312). A benchmark for the evaluation of RGB-D SLAM systems. Proceedings of the 2012 IEEE\/RSJ International Conference on Intelligent Robots and Systems\u2014IROS 2012, Vilamoura, Algarve, Portugal.","DOI":"10.1109\/IROS.2012.6385773"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"2024","DOI":"10.1109\/TPAMI.2015.2505283","article-title":"Learning Depth from Single Monocular Images Using Deep Convolutional Neural Fields","volume":"38","author":"Liu","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Ranjan, A., Jampani, V., Balles, L., Kim, K., Sun, D., Wulff, J., and Black, M.J. (2019, January 16\u201320). Competitive Collaboration: Joint Unsupervised Learning of Depth, Camera Motion, Optical Flow and Motion Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01252"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"2624","DOI":"10.1109\/TPAMI.2019.2930258","article-title":"Every Pixel Counts ++: Joint Learning of Geometry and Motion with 3D Holistic Understanding","volume":"42","author":"Luo","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"1147","DOI":"10.1109\/TRO.2015.2463671","article-title":"ORB-SLAM: A Versatile and Accurate Monocular SLAM System","volume":"31","author":"Montiel","year":"2015","journal-title":"IEEE Trans. Robot."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/3\/923\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:17:28Z","timestamp":1760159848000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/3\/923"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,1,30]]},"references-count":42,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2021,2]]}},"alternative-id":["s21030923"],"URL":"https:\/\/doi.org\/10.3390\/s21030923","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2021,1,30]]}}}