{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T16:09:37Z","timestamp":1784131777930,"version":"3.55.0"},"reference-count":37,"publisher":"MDPI AG","issue":"14","license":[{"start":{"date-parts":[[2024,7,13]],"date-time":"2024-07-13T00:00:00Z","timestamp":1720828800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U2033201"],"award-info":[{"award-number":["U2033201"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["20220058052001"],"award-info":[{"award-number":["20220058052001"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["KYCX19_0194"],"award-info":[{"award-number":["KYCX19_0194"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012130","name":"Aeronautical Science Foundation of China","doi-asserted-by":"publisher","award":["U2033201"],"award-info":[{"award-number":["U2033201"]}],"id":[{"id":"10.13039\/501100012130","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012130","name":"Aeronautical Science Foundation of China","doi-asserted-by":"publisher","award":["20220058052001"],"award-info":[{"award-number":["20220058052001"]}],"id":[{"id":"10.13039\/501100012130","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012130","name":"Aeronautical Science Foundation of China","doi-asserted-by":"publisher","award":["KYCX19_0194"],"award-info":[{"award-number":["KYCX19_0194"]}],"id":[{"id":"10.13039\/501100012130","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Postgraduate Research &amp; Practice Innovation Program of Jiangsu Province","award":["U2033201"],"award-info":[{"award-number":["U2033201"]}]},{"name":"Postgraduate Research &amp; Practice Innovation Program of Jiangsu Province","award":["20220058052001"],"award-info":[{"award-number":["20220058052001"]}]},{"name":"Postgraduate Research &amp; Practice Innovation Program of Jiangsu Province","award":["KYCX19_0194"],"award-info":[{"award-number":["KYCX19_0194"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In recent years, there has been extensive research and application of unsupervised monocular depth estimation methods for intelligent vehicles. However, a major limitation of most existing approaches is their inability to predict absolute depth values in physical units, as they generally suffer from the scale problem. Furthermore, most research efforts have focused on ground vehicles, neglecting the potential application of these methods to unmanned aerial vehicles (UAVs). To address these gaps, this paper proposes a novel absolute depth estimation method specifically designed for flight scenes using a monocular vision sensor, in which a geometry-based scale recovery algorithm serves as a post-processing stage of relative depth estimation results with scale consistency. By exploiting the feature correspondence between successive images and using the pose data provided by equipped navigation sensors, the scale factor between relative and absolute scales is calculated according to a multi-view geometry model, and then absolute depth maps are generated by pixel-wise multiplication of relative depth maps with the scale factor. As a result, the unsupervised monocular depth estimation technology is extended from relative depth estimation in semi-structured scenes to absolute depth estimation in unstructured scenes. Experiments on the publicly available Mid-Air dataset and customized data demonstrate the effectiveness of our method in different cases and settings, as well as its robustness to navigation sensor noise. The proposed method only requires UAVs to be equipped with monocular camera and common navigation sensors, and the obtained absolute depth information can be directly used for downstream tasks, which is significant for this kind of vehicle that has rarely been explored in previous depth estimation studies.<\/jats:p>","DOI":"10.3390\/s24144541","type":"journal-article","created":{"date-parts":[[2024,7,15]],"date-time":"2024-07-15T14:15:49Z","timestamp":1721052949000},"page":"4541","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Monocular Absolute Depth Estimation from Motion for Small Unmanned Aerial Vehicles by Geometry-Based Scale Recovery"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5460-8682","authenticated-orcid":false,"given":"Chuanqi","family":"Zhang","sequence":"first","affiliation":[{"name":"College of Astronautics, Nanjing University of Aeronautics and Astronautics, Nanjing 211106, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiangrui","family":"Weng","sequence":"additional","affiliation":[{"name":"College of Astronautics, Nanjing University of Aeronautics and Astronautics, Nanjing 211106, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yunfeng","family":"Cao","sequence":"additional","affiliation":[{"name":"College of Astronautics, Nanjing University of Aeronautics and Astronautics, Nanjing 211106, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6664-9350","authenticated-orcid":false,"given":"Meng","family":"Ding","sequence":"additional","affiliation":[{"name":"College of Civil Aviation, Nanjing University of Aeronautics and Astronautics, Nanjing 211106, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,7,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Khan, F., Salahuddin, S., and Javidnia, H. (2020). Deep learning-based monocular depth estimation methods\u2014A state-of-the-art review. Sensors, 20.","DOI":"10.3390\/s20082272"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Masoumian, A., Rashwan, H.A., Cristiano, J., Asif, M.S., and Puig, D. (2022). Monocular depth estimation using deep learning: A review. Sensors, 22.","DOI":"10.3390\/s22145353"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"16940","DOI":"10.1109\/TITS.2022.3160741","article-title":"Towards real-time monocular depth estimation for robotics: A survey","volume":"23","author":"Dong","year":"2022","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Xue, F., Zhuo, G., Huang, Z., Fu, W., Wu, Z., and Ang, M.H. (2020\u201324, January 24). Toward hierarchical self-supervised monocular absolute depth estimation for autonomous driving applications. Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA.","DOI":"10.1109\/IROS45743.2020.9340802"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Florea, H., and Nedevschi, S. (2022, January 22\u201324). Survey on monocular depth estimation for unmanned aerial vehicles using deep learning. Proceedings of the IEEE International Conference on Intelligent Computer Communication and Processing (ICCP), Cluj-Napoca, Romania.","DOI":"10.1109\/ICCP56966.2022.10053950"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Romero-Lugo, A., Magadan-Salazar, A., Fuentes-Pacheco, J., and Pinto-El\u00edas, R. (2022). A comparison of deep neural networks for monocular depth map estimation in natural environments flying at low altitude. Sensors, 22.","DOI":"10.3390\/s22249830"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"105523","DOI":"10.1016\/j.compag.2020.105523","article-title":"UAV environmental perception and autonomous obstacle avoidance: A deep learning and depth camera combined solution","volume":"175","author":"Wang","year":"2020","journal-title":"Comput. Electron. Agric."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Xiong, M., and Xiong, H. (2019, January 6\u20137). Monocular depth estimation for UAV obstacle avoidance. Proceedings of the International Conference on Cloud Computing and Internet of Things (CCIOT), Changchun, China.","DOI":"10.1109\/CCIOT48581.2019.8980350"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1098\/rspb.1979.0006","article-title":"The interpretation of structure from motion","volume":"203","author":"Ullman","year":"1979","journal-title":"Proc. R. Soc. Lond. B"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1612","DOI":"10.1007\/s11431-020-1582-8","article-title":"Monocular depth estimation based on deep learning: An overview","volume":"63","author":"Zhao","year":"2020","journal-title":"Sci. China-Technol. Sci."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1016\/j.neucom.2020.12.089","article-title":"Deep learning for monocular depth estimation: A review","volume":"438","author":"Ming","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"811","DOI":"10.1360\/N112015-00285","article-title":"The inherent ambiguity in scene depth learning from single images","volume":"46","author":"He","year":"2016","journal-title":"Sci. Sin. Inform."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Godard, C., Mac Aodha, O., and Brostow, G.J. (2017, January 21\u201326). Unsupervised monocular depth estimation with left-right consistency. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.699"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zhou, T., Brown, M., Snavely, N., and Lowe, D.G. (2017, January 21\u201326). Unsupervised learning of depth and ego-motion from video. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.700"},{"key":"ref_15","unstructured":"Godard, C., Mac Aodha, O., Firman, M., and Brostow, G.J. (November, January 27). Digging into self-supervised monocular depth estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seoul, Republic of Korea."},{"key":"ref_16","unstructured":"Fonder, M., and Van Droogenbroeck, M. (November, January 27). Mid-Air: A multi-modal dataset for extremely low altitude drone flights. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, Seoul, Republic of Korea."},{"key":"ref_17","first-page":"2366","article-title":"Depth map prediction from a single image using a multi-scale deep network","volume":"27","author":"Eigen","year":"2014","journal-title":"Adv. Neural Inf. Process. Syst. (NIPS)"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Laina, I., Rupprecht, C., Belagiannis, V., Tombari, F., and Navab, N. (2016, January 25\u201328). Deeper depth prediction with fully convolutional residual networks. Proceedings of the International Conference on 3D Vision (3DV), Stanford, CA, USA.","DOI":"10.1109\/3DV.2016.32"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Fu, H., Gong, M., Wang, C., Batmanghelich, K., and Tao, D. (2018, January 18\u201323). Deep ordinal regression network for monocular depth estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00214"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Garg, R., Bg, V.K., Carneiro, G., and Reid, I. (2016, January 8\u201316). Unsupervised CNN for single view depth estimation: Geometry to the rescue. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_45"},{"key":"ref_21","first-page":"2017","article-title":"Spatial transformer networks","volume":"28","author":"Jaderberg","year":"2015","journal-title":"Adv. Neural Inf. Process. Syst. (NIPS)"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"He, L., Yang, J., Kong, B., and Wang, C. (2017). An automatic measurement method for absolute depth of objects in two monocular images based on SIFT feature. Appl. Sci., 7.","DOI":"10.20944\/preprints201705.0028.v1"},{"key":"ref_23","first-page":"214","article-title":"Object depth measurement and filtering from monocular images for unmanned aerial vehicles","volume":"19","author":"Zhang","year":"2022","journal-title":"J. Aerosp. Inf. Syst."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Petrovai, A., and Nedevschi, S. (2022, January 18\u201324). Exploiting pseudo labels in a self-supervised learning framework for improved monocular depth estimation. Proceedings of the Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00163"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhang, S., Zhang, J., and Tao, D. (2022, January 23\u201327). Towards scale-aware, robust, and generalizable unsupervised monocular depth estimation by integrating IMU motion dynamics. Proceedings of the European Conference on Computer Vision (ECCV), Tel-Aviv, Israel.","DOI":"10.1007\/978-3-031-19839-7_9"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Pinard, C., Chevalley, L., Manzanera, A., and Filliat, D. (2018, January 8\u201314). Learning structure-from-motion from motion. Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Munich, Germany.","DOI":"10.1007\/978-3-030-11015-4_27"},{"key":"ref_27","first-page":"35","article-title":"Unsupervised scale-consistent depth and ego-motion learning from monocular video","volume":"32","author":"Bian","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst. (NeurIPS)"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"2548","DOI":"10.1007\/s11263-021-01484-6","article-title":"Unsupervised scale-consistent depth learning from video","volume":"129","author":"Bian","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Mayer, N., Ilg, E., Hausser, P., Fischer, P., Cremers, D., Dosovitskiy, A., and Brox, T. (2016, January 27\u201330). A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.438"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). ImageNet: A large-scale hierarchical image database. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","article-title":"Image quality assessment: From error visibility to structural similarity","volume":"13","author":"Wang","year":"2004","journal-title":"IEEE Trans. Image Process."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive image features from scale-invariant keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Shah, S., Dey, D., Lovett, C., and Kapoor, A. (2018, January 12\u201315). AirSim: High-fidelity visual and physical simulation for autonomous vehicles. Proceedings of the International Conference on Field and Service Robotics (FSR), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-67361-5_40"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The KITTI dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_36","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_37","first-page":"8026","article-title":"PyTorch: An imperative style, high-performance deep learning library","volume":"32","author":"Paszke","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst. (NeurIPS)"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/14\/4541\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:16:18Z","timestamp":1760109378000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/14\/4541"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,13]]},"references-count":37,"journal-issue":{"issue":"14","published-online":{"date-parts":[[2024,7]]}},"alternative-id":["s24144541"],"URL":"https:\/\/doi.org\/10.3390\/s24144541","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,13]]}}}