{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T18:04:47Z","timestamp":1779300287872,"version":"3.51.4"},"reference-count":65,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2019,4,10]],"date-time":"2019-04-10T00:00:00Z","timestamp":1554854400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100004663","name":"Ministry of Science and Technology, Taiwan","doi-asserted-by":"publisher","award":["107-2218-E-011-014"],"award-info":[{"award-number":["107-2218-E-011-014"]}],"id":[{"id":"10.13039\/501100004663","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004663","name":"Ministry of Science and Technology, Taiwan","doi-asserted-by":"publisher","award":["106-2221-E-011-154 -MY2"],"award-info":[{"award-number":["106-2221-E-011-154 -MY2"]}],"id":[{"id":"10.13039\/501100004663","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Ministry of Education (MOE)","award":["The Featured Areas Research Center Program within the framework of the Higher Education Sprout Project"],"award-info":[{"award-number":["The Featured Areas Research Center Program within the framework of the Higher Education Sprout Project"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Depth has been a valuable piece of information for perception tasks such as robot grasping, obstacle avoidance, and navigation, which are essential tasks for developing smart homes and smart cities. However, not all applications have the luxury of using depth sensors or multiple cameras to obtain depth information. In this paper, we tackle the problem of estimating the per-pixel depths from a single image. Inspired by the recent works on generative neural network models, we formulate the task of depth estimation as a generative task where we synthesize an image of the depth map from a single Red, Green, and Blue (RGB) input image. We propose a novel generative adversarial network that has an encoder-decoder type generator with residual transposed convolution blocks trained with an adversarial loss. Quantitative and qualitative experimental results demonstrate the effectiveness of our approach over several depth estimation works.<\/jats:p>","DOI":"10.3390\/s19071708","type":"journal-article","created":{"date-parts":[[2019,4,10]],"date-time":"2019-04-10T11:25:08Z","timestamp":1554895508000},"page":"1708","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Single-Image Depth Inference Using Generative Adversarial Networks"],"prefix":"10.3390","volume":"19","author":[{"given":"Daniel Stanley","family":"Tan","sequence":"first","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Taiwan University of Science and Technology, Taipei 10607, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chih-Yuan","family":"Yao","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Taiwan University of Science and Technology, Taipei 10607, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0956-7201","authenticated-orcid":false,"suffix":"Jr.","given":"Conrado","family":"Ruiz","sequence":"additional","affiliation":[{"name":"Software Technology Department, De La Salle University, Manila 1004, Philippines"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7735-243X","authenticated-orcid":false,"given":"Kai-Lung","family":"Hua","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Taiwan University of Science and Technology, Taipei 10607, Taiwan"},{"name":"Center for Cyber-Physical System Innovation, National Taiwan University of Science and Technology, Taipei 10607, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,4,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"705","DOI":"10.1177\/0278364914549607","article-title":"Deep learning for detecting robotic grasps","volume":"34","author":"Lenz","year":"2015","journal-title":"Int. J. Robot. Res."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Michels, J., Saxena, A., and Ng, A.Y. (2005, January 7\u201311). High speed obstacle avoidance using monocular vision and reinforcement learning. Proceedings of the ACM 22nd International Conference on Machine Learning, Bonn, Germany.","DOI":"10.1145\/1102351.1102426"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"120","DOI":"10.1002\/rob.20276","article-title":"Learning long-range vision for autonomous off-road driving","volume":"26","author":"Hadsell","year":"2009","journal-title":"J. Field Robot."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"2625","DOI":"10.1109\/TMM.2017.2694218","article-title":"Improved depth-assisted error concealment algorithm for 3D video transmission","volume":"19","author":"Huang","year":"2017","journal-title":"IEEE Trans. Multimed."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"822","DOI":"10.1109\/TMM.2016.2626969","article-title":"Sleep apnea detection via depth video and audio feature learning","volume":"19","author":"Yang","year":"2017","journal-title":"IEEE Trans. Multimed."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1625","DOI":"10.1109\/TMM.2017.2672198","article-title":"Estimating heart rate and rhythm via 3D motion tracking in depth video","volume":"19","author":"Yang","year":"2017","journal-title":"IEEE Trans. Multimed."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"116","DOI":"10.1145\/2398356.2398381","article-title":"Real-time human pose recognition in parts from single depth images","volume":"56","author":"Shotton","year":"2013","journal-title":"Commun. ACM"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1053","DOI":"10.1109\/TCYB.2013.2279071","article-title":"Exemplar-based human action pose correction","volume":"44","author":"Shen","year":"2014","journal-title":"IEEE Trans. Cybern."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"2379","DOI":"10.1109\/TCYB.2014.2307121","article-title":"Spherical blurred shape model for 3-D object and pose recognition: Quantitative analysis and HCI applications in smart environments","volume":"44","author":"Lopes","year":"2014","journal-title":"IEEE Trans. Cybern."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"742","DOI":"10.1109\/TCYB.2014.2335540","article-title":"Real-time human movement retrieval and assessment with kinect sensor","volume":"45","author":"Hu","year":"2015","journal-title":"IEEE Trans. Cybern."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1340","DOI":"10.1109\/TCYB.2014.2350774","article-title":"3-d human action recognition by shape analysis of motion trajectories on riemannian manifold","volume":"45","author":"Devanne","year":"2015","journal-title":"IEEE Trans. Cybern."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1194","DOI":"10.1109\/TCYB.2014.2347057","article-title":"Multiple\/single-view human action recognition via part-induced multitask structural learning","volume":"45","author":"Liu","year":"2015","journal-title":"IEEE Trans. Cybern."},{"key":"ref_13","unstructured":"Wang, P., Shen, X., Lin, Z., Cohen, S., Price, B., and Yuille, A.L. (2015, January 7\u201312). Towards unified depth and semantic prediction from a single image. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"927","DOI":"10.1109\/TCYB.2014.2340032","article-title":"Graph-based segmentation for RGB-D data using 3-D geometry enhanced superpixels","volume":"45","author":"Yang","year":"2015","journal-title":"IEEE Trans. Cybern."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"266","DOI":"10.1109\/TCYB.2014.2324815","article-title":"Consistent depth video segmentation using adaptive surface models","volume":"45","author":"Husain","year":"2015","journal-title":"IEEE Trans. Cybern."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Scharstein, D., Szeliski, R., and Zabih, R. (2001). A taxonomy and evaluation of dense two-frame stereo correspondence algorithms. Proceedings of the IEEE Workshop on Stereo and Multi-Baseline Vision, (SMBV 2001), IEEE Computer Society.","DOI":"10.1109\/SMBV.2001.988771"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1983","DOI":"10.1109\/TIP.2015.2409551","article-title":"Depth reconstruction from sparse samples: Representation, algorithm, and sampling","volume":"24","author":"Liu","year":"2015","journal-title":"IEEE Trans. Image Process."},{"key":"ref_18","unstructured":"Memisevic, R., and Conrad, C. (2011, January 16). Stereopsis via deep learning. Proceedings of the NIPS Workshop on Deep Learning and Unsupervised Feature Learning, Granada, Spain."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Sinz, F.H., Candela, J.Q., Bak\u0131r, G.H., Rasmussen, C.E., and Franz, M.O. (2004). Learning depth from stereo. Joint Pattern Recognition Symposium, Springer.","DOI":"10.1007\/978-3-540-28649-3_30"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"685","DOI":"10.1109\/TMM.2016.2646179","article-title":"Accurate depth extraction method for multiple light-coding-based depth cameras","volume":"19","author":"Pan","year":"2017","journal-title":"IEEE Trans. Multimed."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"726","DOI":"10.1109\/TIP.2015.2507984","article-title":"Depth estimation using a sliding camera","volume":"25","author":"Ge","year":"2016","journal-title":"IEEE Trans. Image Process."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"3257","DOI":"10.1109\/TIP.2015.2440760","article-title":"Continuous depth map reconstruction from light fields","volume":"24","author":"Li","year":"2015","journal-title":"IEEE Trans. Image Process."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"97","DOI":"10.1109\/TIP.2013.2286901","article-title":"Real-time scalable depth sensing with hybrid structured light illumination","volume":"23","author":"Zhang","year":"2014","journal-title":"IEEE Trans. Image Process."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1235","DOI":"10.1109\/TIP.2015.2397591","article-title":"Depth from water reflection","volume":"24","author":"Yang","year":"2015","journal-title":"IEEE Trans. Image Process."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Howard, I.P. (2012). Perceiving in Depth, Basic Mechanisms, Oxford University Press.","DOI":"10.1093\/acprof:oso\/9780199764143.001.0001"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"76900B","DOI":"10.1117\/12.850094","article-title":"Depth cues in human visual perception and their realization in 3D displays","volume":"7690","author":"Reichelt","year":"2010","journal-title":"Proc. SPIE"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1111\/1467-8721.ep10772783","article-title":"Visual perception of location and distance","volume":"5","author":"Loomis","year":"1996","journal-title":"Curr. Dir. Psychol. Sci."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"2","DOI":"10.1167\/5.2.2","article-title":"Reaching for visual cues to depth: The brain combines depth cues differently for motor control and perception","volume":"5","author":"Knill","year":"2005","journal-title":"J. Vis."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"122","DOI":"10.1109\/TMM.2013.2283451","article-title":"Robust semi-automatic depth map generation in unconstrained images and video sequences for 2D to stereoscopic 3D conversion","volume":"16","author":"Phan","year":"2014","journal-title":"IEEE Trans. Multimed."},{"key":"ref_30","unstructured":"Eigen, D., Puhrsch, C., and Fergus, R. (2014, January 8\u201313). Depth map prediction from a single image using a multi-scale deep network. Proceedings of the Neural Information Processing Systems 2014 (NIPS 2014), Montreal, QC, Canada."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Eigen, D., and Fergus, R. (2015, January 7\u201312). Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. Proceedings of the IEEE International Conference on Computer Vision, Boston, MA, USA.","DOI":"10.1109\/ICCV.2015.304"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Isola, P., Zhu, J.Y., Zhou, T., and Efros, A.A. (2017, January 21\u201326). Image-to-image translation with conditional adversarial networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.632"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Tran, L., Yin, X., and Liu, X. (2017, January 21\u201326). Disentangled representation learning gan for pose-invariant face recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.141"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zhu, J.Y., Park, T., Isola, P., and Efros, A.A. (2017, January 22\u201329). Unpaired image-to-image translation using cycle-consistent adversarial networks. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.244"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Hoiem, D., Efros, A.A., and Hebert, M. (2005, January 17\u201320). Geometric context from a single image. Proceedings of the Tenth IEEE International Conference on Computer Vision (ICCV 2005), Beijing, China.","DOI":"10.1109\/ICCV.2005.107"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Ladicky, L., Shi, J., and Pollefeys, M. (2014, January 23\u201328). Pulling things out of perspective. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.19"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Karsch, K., Liu, C., and Kang, S.B. (2012). Depth extraction from video using non-parametric sampling. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-642-33715-4_56"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Konrad, J., Wang, M., and Ishwar, P. (2012, January 16\u201321). 2d-to-3d image conversion by learning depth from examples. Proceedings of the 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Providence, RI, USA.","DOI":"10.1109\/CVPRW.2012.6238903"},{"key":"ref_39","unstructured":"Saxena, A., Chung, S.H., and Ng, A.Y. (2019, April 01). Learning depth from single monocular images. Available online: http:\/\/59.80.44.45\/papers.nips.cc\/paper\/2921-learning-depth-from-single-monocular-images.pdf."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Liu, B., Gould, S., and Koller, D. (2010, January 13\u201318). Single image depth estimation from predicted semantic labels. Proceedings of the 2010 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539823"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Liu, M., Salzmann, M., and He, X. (2014, January 23\u201328). Discrete-continuous depth estimation from a single image. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.97"},{"key":"ref_42","unstructured":"Roy, A., and Todorovic, S. (July, January 26). Monocular depth estimation using neural regression forest. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegan, NV, USA."},{"key":"ref_43","unstructured":"Chakrabarti, A., Shao, J., and Shakhnarovich, G. (2016, January 5\u201310). Depth from a single image by harmonizing overcomplete local network predictions. Proceedings of the Neural Information Processing Systems Conference (NIPS 2016), Barcelona, Spain."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"2024","DOI":"10.1109\/TPAMI.2015.2505283","article-title":"Learning depth from single monocular images using deep convolutional neural fields","volume":"38","author":"Liu","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_45","unstructured":"Bousmalis, K., Trigeorgis, G., Silberman, N., Krishnan, D., and Erhan, D. (2016, January 5\u201310). Domain separation networks. Proceedings of the Neural Information Processing Systems Conference (NIPS 2016), Barcelona, Spain."},{"key":"ref_46","unstructured":"Hoffman, J., Wang, D., Yu, F., and Darrell, T. (arXiv, 2016). Fcns in the wild: Pixel-level adversarial and constraint-based adaptation, arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Wulfmeier, M., Bewley, A., and Posner, I. (arXiv, 2017). Addressing appearance change in outdoor robotics with adversarial domain adaptation, arXiv.","DOI":"10.1109\/IROS.2017.8205961"},{"key":"ref_48","unstructured":"Radford, A., Metz, L., and Chintala, S. (arXiv, 2015). Unsupervised representation learning with deep convolutional generative adversarial networks, arXiv."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Zhang, H., Xu, T., Li, H., Zhang, S., Huang, X., Wang, X., and Metaxas, D. (arXiv, 2016). Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks, arXiv.","DOI":"10.1109\/ICCV.2017.629"},{"key":"ref_50","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegan, NV, USA."},{"key":"ref_51","unstructured":"Simonyan, K., and Zisserman, A. (arXiv, 2014). Very deep convolutional networks for large-scale image recognition, arXiv."},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"e3","DOI":"10.23915\/distill.00003","article-title":"Deconvolution and checkerboard artifacts","volume":"1","author":"Odena","year":"2016","journal-title":"Distill"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Sun, J., and Ren, S. (2016). Identity mappings in deep residual networks. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"ref_54","unstructured":"Srivastava, R.K., Greff, K., and Schmidhuber, J. (2015, January 7\u201312). Training very deep networks. Proceedings of the Neural Information Processing Systems 2015, Montreal, QC, Canada."},{"key":"ref_55","unstructured":"Mao, X., Shen, C., and Yang, Y.B. (2016, January 5\u201310). Image restoration using very deep convolutional encoder-decoder networks with symmetric skip connections. Proceedings of the Neural Information Processing Systems 2016 (NIPS 2016), Barcelona, Spain."},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015). U-net: Convolutional networks for biomedical image segmentation. International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Silberman, N., Hoiem, D., Kohli, P., and Fergus, R. (2012). Indoor Segmentation and Support Inference from rgbd Images, Springer.","DOI":"10.1007\/978-3-642-33715-4_54"},{"key":"ref_58","doi-asserted-by":"crossref","first-page":"689","DOI":"10.1145\/1015706.1015780","article-title":"Colorization using optimization","volume":"Volume 23","author":"Levin","year":"2004","journal-title":"ACM Transactions on Graphics"},{"key":"ref_59","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014, January 8\u201313). Generative adversarial nets. Proceedings of the Neural Information Processing Systems (NIPS 2014), Montreal, QC, Canada."},{"key":"ref_60","unstructured":"Zhang, J., Mitliagkas, I., and R\u00e9, C. (arXiv, 2017). YellowFin and the art of momentum tuning, arXiv."},{"key":"ref_61","unstructured":"Torralba, A., Murphy, K., and Freeman, W. (July, January 27). Sharing features: Efficient boosting procedures for multiclass object detection. Proceedings of the 2004 IEEE Conference on Computer Vision and Pattern Recognition, Washington, DC, USA."},{"key":"ref_62","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1023\/A:1011174803800","article-title":"Contour and texture analysis for image segmentation","volume":"43","author":"Malik","year":"2001","journal-title":"Int. J. Comput. Vis."},{"key":"ref_63","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive image features from scale-invariant keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"Ul Hussain, S., and Triggs, B. (2012). Visual recognition using local quantized patterns. European Conference on Computer Vision, Springer.","DOI":"10.5244\/C.26.99"},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Shechtman, E., and Irani, M. (2007, January 17\u201322). Matching local self-similarities across images and videos. Proceedings of the 2007 IEEE Conference on Computer Vision and Pattern Recognition, Minneapolis, MN, USA.","DOI":"10.1109\/CVPR.2007.383198"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/7\/1708\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:44:12Z","timestamp":1760186652000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/7\/1708"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,4,10]]},"references-count":65,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2019,4]]}},"alternative-id":["s19071708"],"URL":"https:\/\/doi.org\/10.3390\/s19071708","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,4,10]]}}}