{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:47:38Z","timestamp":1760240858068,"version":"build-2065373602"},"reference-count":39,"publisher":"MDPI AG","issue":"20","license":[{"start":{"date-parts":[[2019,10,13]],"date-time":"2019-10-13T00:00:00Z","timestamp":1570924800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Depth estimation is a crucial and fundamental problem in the computer vision field. Conventional methods re-construct scenes using feature points extracted from multiple images; however, these approaches require multiple images and thus are not easily implemented in various real-time applications. Moreover, the special equipment required by hardware-based approaches using 3D sensors is expensive. Therefore, software-based methods for estimating depth from a single image using machine learning or deep learning are emerging as new alternatives. In this paper, we propose an algorithm that generates a depth map in real time using a single image and an optimized lightweight efficient neural network (L-ENet) algorithm instead of physical equipment, such as an infrared sensor or multi-view camera. Because depth values have a continuous nature and can produce locally ambiguous results, pixel-wise prediction with ordinal depth range classification was applied in this study. In addition, in our method various convolution techniques are applied to extract a dense feature map, and the number of parameters is greatly reduced by reducing the network layer. By using the proposed L-ENet algorithm, an accurate depth map can be generated from a single image quickly and, in a comparison with the ground truth, we can produce depth values closer to those of the ground truth with small errors. Experiments confirmed that the proposed L-ENet can achieve a significantly improved estimation performance over the state-of-the-art algorithms in depth estimation based on a single image.<\/jats:p>","DOI":"10.3390\/s19204434","type":"journal-article","created":{"date-parts":[[2019,10,14]],"date-time":"2019-10-14T03:54:13Z","timestamp":1571025253000},"page":"4434","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Fast Depth Estimation in a Single Image Using Lightweight Efficient Neural Network"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7452-3897","authenticated-orcid":false,"given":"Sangwon","family":"Kim","sequence":"first","affiliation":[{"name":"Department of Computer Engineering, Keimyung University, Daegu 42601, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jaeyeal","family":"Nam","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Keimyung University, Daegu 42601, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7284-0768","authenticated-orcid":false,"given":"Byoungchul","family":"Ko","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Keimyung University, Daegu 42601, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,10,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1007\/s11554-012-0313-2","article-title":"Review of stereo vision algorithms and their suitability for resource-limited systems","volume":"11","author":"Tippetts","year":"2016","journal-title":"J. Real-Time Image Process."},{"key":"ref_2","unstructured":"Ha, H., Im, S., Park, J., Jeon, H.G., and Kwoen, I.S. (July, January 26). High quality depth from uncalibrated small motion clip. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1521","DOI":"10.1109\/TPAMI.2004.102","article-title":"Depth estimation and image restoration using defocused stereo pairs","volume":"26","author":"Rajagopalan","year":"2004","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1632","DOI":"10.1016\/j.patcog.2005.01.006","article-title":"Towards a real-time 3D shape reconstruction using a structured light system","volume":"38","author":"Dipanda","year":"2005","journal-title":"Pattern Recognit."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Paragios, N., Chen, Y., and Faugeras, O.D. (2006). Handbook of Mathematical Models in Computer Vision, Springer.","DOI":"10.1007\/0-387-28831-7"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Konolige, K. (2010, January 3\u20138). Projected texture stereo. Proceedings of the IEEE International Conference on Robotics and Automation, Anchorage, AK, USA.","DOI":"10.1109\/ROBOT.2010.5509796"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2","DOI":"10.1177\/1729881418760623","article-title":"Advances in sensing and processing methods for three-dimensional robot vision","volume":"15","author":"He","year":"2018","journal-title":"Int. J. Adv. Robot. Syst."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Gandhi, V., \u010cech, J., and Horaud, R. (2012, January 14\u201318). High-resolution depth maps based on TOF-stereo fusion. Proceedings of the IEEE International Conference on Robotics and Automation, Saint Paul, MN, USA.","DOI":"10.1109\/ICRA.2012.6224771"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"2024","DOI":"10.1109\/TPAMI.2015.2505283","article-title":"Learning depth from single monocular images using deep convolutional neural fields","volume":"38","author":"Liu","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_10","unstructured":"Wang, P., Shen, X., Lin, Z., Cohen, S., Price, B., and Yuille, A. (2015, January 7\u201312). Towards unified depth and semantic prediction from a single image. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Eigen, D., and Fergus, R. (2015, January 13\u201316). Predicting depth, surface normal and semantic labels with a common multi-scale convolutional architecture. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.304"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Kim, S., Park, K., Sohn, K., and Lin, S. (2016, January 8\u201316). Unified depth prediction and intrinsic image decomposition from a single image via joint convolutional neural fields. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_9"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Kuznietsov, Y., Stuckler, J., and Leib, B. (2017, January 21\u201326). Semi-supervised deep learning for monocular depth map prediction. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.238"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Fu, H., Gong, M., Wang, C., Batmanghelich, K., and Tao, D. (2018, January 18\u201322). Deep ordinal regression network for monocular depth estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00214"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"271","DOI":"10.1007\/BF02028349","article-title":"Depth from defocus: A spatial domain approach","volume":"13","author":"Subbarao","year":"1994","journal-title":"Int. J. Comput. Vis."},{"key":"ref_16","unstructured":"Hiura, S., and Matsuyama, T. (1998, January 23\u201325). Depth measurement by the multi-focus camera. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Santa Barbara, CA, USA."},{"key":"ref_17","unstructured":"Saxena, A., Chung, S.H., and Ng, A.Y. (2006, January 4\u20137). Learning depth from single monocular images. Proceedings of the Advances in Neural Information Processing Systems (NIPS), Vancouver, BC, Canada."},{"key":"ref_18","unstructured":"Tompson, J.J., Jain, A., LeCun, Y., and Bregler, C. (2014, January 8\u201313). Joint training of a convolutional network and a graphical model for human pose estimation. Proceedings of the Advances in Neural Information Systems (NIPS), Montr\u00e9al, QC, Canada."},{"key":"ref_19","unstructured":"Li, B., Shen, C., Dai, Y., Hengel, A.V.D., and He, M. (2015, January 7\u201312). Depth and surface normal estimation from monocular images using regression on deep features and hierarchical CRFs. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA."},{"key":"ref_20","unstructured":"Luo, W., Schwing, A.G., and Urtasun, R. (July, January 26). Efficient deep learning for stereo matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Godard, C., Aodha, O.M., and Brostow, G.J. (2017, January 21\u201326). Unsupervised monocular depth estimation with left-right consistency. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.699"},{"key":"ref_22","unstructured":"Roy, A., and Todorovic, S. (July, January 26). Monocular depth estimation using neural regression forest. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Laina, I., Rupprecht, C., Belagiannis, V., Tombari, F., and Navab, N. (2016, January 25\u201328). Deeper depth prediction with fully convolutional residual networks. Proceedings of the International Conference on 3D Vision (3DV), Stanford, CA, USA.","DOI":"10.1109\/3DV.2016.32"},{"key":"ref_24","unstructured":"Chakrabarti, A., Shao, J., and Shakhnarovich, G. (2016, January 5\u201310). Depth from a single image by harmonizing overcomplete local network predictions. Proceedings of the Advances in Neural Information Processing Systems (NIPS), Barcelona, Spain."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Lee, J.H., Heo, M., Kim, K., and Kim, C.S. (2018, January 18\u201322). Single-image depth estimation based on fourier domain analysis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00042"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Goldman, M., Hassner, T., and Avidan, S. (2019, January 16\u201317). Learn stereo, infer mono: Siamese networks for self-supervised, monocular, depth estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, Long Beach, CA, USA.","DOI":"10.1109\/CVPRW.2019.00348"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Diaz, R., and Marathe, A. (2019, January 18\u201320). Soft labels for ordinal regression. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00487"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"4676","DOI":"10.1109\/TIP.2018.2832296","article-title":"Learning depth from single images with deep neural network embedding focal length","volume":"27","author":"He","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Garg, R., BG, V.K., Carneiro, G., and Reid, I. (2016, January 8\u201316). Unsupervised CNN for single view depth estimation: Geometry to the rescue. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_45"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Xie, J., Girshick, R., and Farhadi, A. (2016, January 8\u201316). Deep3d: Fully automatic 2d-to-3d video conversion with deep convolutional neural networks. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46493-0_51"},{"key":"ref_31","unstructured":"Wen, W., Wu, C., Wang, Y., Chen, Y., and Li, H. (2016, January 5\u201310). Learning structured sparsity in deep neural networks. Proceedings of the Advances in Neural Information Processing Systems (NIPS), Barcelona, Spain."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Huang, Z., and Wang, N. (2018, January 8\u201314). Data-driven sparse structure selection for deep neural networks. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01270-0_19"},{"key":"ref_33","unstructured":"Paszke, A., Chaurasia, A., Kim, S., and Culurciello, E. (2016). ENet: A deep neural network architecture for real-time semantic segmentation. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2015, January 13\u201316). Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.123"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Chollet, F. (2017, January 21\u201326). Xception: Deep learning with depthwise separable convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.195"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Silberman, P.K.N., Hoiem, D., and Fergu, R. (2012, January 7\u201313). Indoor segmentation and support inference from rgbd images. Proceedings of the European Conference on Computer Vision (ECCV), Firenze, Italy.","DOI":"10.1007\/978-3-642-33715-4_54"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The KITTI dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"824","DOI":"10.1109\/TPAMI.2008.132","article-title":"Make3d: Learning 3d scene structure from a single still image","volume":"31","author":"Saxena","year":"2009","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Jiang, H., Larsson, G., Marie, M., Shakhnarovich, G., and Miller, E.L. (2018, January 8\u201314). Self-supervised relative depth learning for urban scene understanding. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01252-6_2"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/20\/4434\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:25:53Z","timestamp":1760189153000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/20\/4434"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,10,13]]},"references-count":39,"journal-issue":{"issue":"20","published-online":{"date-parts":[[2019,10]]}},"alternative-id":["s19204434"],"URL":"https:\/\/doi.org\/10.3390\/s19204434","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2019,10,13]]}}}