{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:36:24Z","timestamp":1760240184178,"version":"build-2065373602"},"reference-count":34,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2019,4,2]],"date-time":"2019-04-02T00:00:00Z","timestamp":1554163200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100004663","name":"Ministry of Science and Technology, Taiwan","doi-asserted-by":"publisher","award":["106-2221-E-011-154 -MY2","107-2218-E-011-014 -","107-2218-E-027-020"],"award-info":[{"award-number":["106-2221-E-011-154 -MY2","107-2218-E-011-014 -","107-2218-E-027-020"]}],"id":[{"id":"10.13039\/501100004663","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Ministry of Education, Taiwan","award":["The Featured Areas Research Center Program within the framework of the Higher Education 198 Sprout Project"],"award-info":[{"award-number":["The Featured Areas Research Center Program within the framework of the Higher Education 198 Sprout Project"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Autonomous robots for smart homes and smart cities mostly require depth perception in order to interact with their environments. However, depth maps are usually captured in a lower resolution as compared to RGB color images due to the inherent limitations of the sensors. Naively increasing its resolution often leads to loss of sharpness and incorrect estimates, especially in the regions with depth discontinuities or depth boundaries. In this paper, we propose a novel Generative Adversarial Network (GAN)-based framework for depth map super-resolution that is able to preserve the smooth areas, as well as the sharp edges at the boundaries of the depth map. Our proposed model is trained on two different modalities, namely color images and depth maps. However, at test time, our model only requires the depth map in order to produce a higher resolution version. We evaluated our model both quantitatively and qualitatively, and our experiments show that our method performs better than existing state-of-the-art models.<\/jats:p>","DOI":"10.3390\/s19071587","type":"journal-article","created":{"date-parts":[[2019,4,3]],"date-time":"2019-04-03T03:39:28Z","timestamp":1554262768000},"page":"1587","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Depth Map Upsampling via Multi-Modal Generative Adversarial Network"],"prefix":"10.3390","volume":"19","author":[{"given":"Daniel Stanley","family":"Tan","sequence":"first","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Taiwan University of Science and Technology, Taipei 10607, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun-Ming","family":"Lin","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Taiwan University of Science and Technology, Taipei 10607, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8578-3101","authenticated-orcid":false,"given":"Yu-Chi","family":"Lai","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Taiwan University of Science and Technology, Taipei 10607, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7601-6605","authenticated-orcid":false,"given":"Joel","family":"Ilao","sequence":"additional","affiliation":[{"name":"Center for Automation Research, College of Computer Studies, De La Salle University, Manila 1004, Philippines"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7735-243X","authenticated-orcid":false,"given":"Kai-Lung","family":"Hua","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Taiwan University of Science and Technology, Taipei 10607, Taiwan"},{"name":"Center for Cyber-Physical System Innovation, National Taiwan University of Science and Technology, Taipei 10607, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,4,2]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Schuon, S., Theobalt, C., Davis, J., and Thrun, S. (2008, January 23\u201328). High-quality scanning using time-of-flight depth superresolution. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops (CVPRW\u201908), Anchorage, AK, USA.","DOI":"10.1109\/CVPRW.2008.4563171"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Guomundsson, S.A., Aanaes, H., and Larsen, R. (2007, January 12\u201313). Environmental effects on measurement uncertainties of Time-of-Flight cameras. Proceedings of the International Symposium on Signals, Circuits and Systems (ISSCS), Iasi, Romania.","DOI":"10.1109\/ISSCS.2007.4292664"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Lo, K.H., Hua, K.L., and Wang, Y.C.F. (2013, January 26\u201331). Depth map super-resolution via Markov random fields without texture-copying artifacts. Proceedings of the 2013 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6637884"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Lo, K.H., Wang, Y.C.F., and Hua, K.L. (2013, January 17\u201320). Joint trilateral filtering for depth map super-resolution. Proceedings of the Visual Communications and Image Processing (VCIP), Kuching, Malaysia.","DOI":"10.1109\/VCIP.2013.6706444"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Isola, P., Zhu, J.Y., Zhou, T., and Efros, A.A. (2017, January 21\u201326). Image-to-Image Translation with Conditional Adversarial Networks. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.632"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Zhu, J.Y., Park, T., Isola, P., and Efros, A.A. (2017, January 22\u201329). Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.244"},{"key":"ref_7","unstructured":"Aodha, O.M., Campbell, N.D., Nair, A., and Brostow, G.J. (2012, January 7\u201313). Patch based synthesis for single depth image super-resolution. Proceedings of the European Conference on Computer Vision (ECCV), Florence, Italy."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Hornacek, M., Rhemann, C., Gelautz, M., and Rother, C. (2013, January 23\u201328). Depth super-resolution by rigid body self-similarity in 3d. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.149"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Li, J., Lu, Z., Zeng, G., Gan, R., and Zha, H. (2014, January 24\u201327). Similarity-aware patchwork assembly for depth image super-resolution. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.431"},{"key":"ref_10","unstructured":"Tomasi, C., and Manduchi, R. (1998, January 4\u20137). Bilateral filtering for gray and color images. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Bombay, India."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1145\/1276377.1276497","article-title":"Joint bilateral upsampling","volume":"26","author":"Kopf","year":"2007","journal-title":"ACM Trans. Graph."},{"key":"ref_12","unstructured":"Chan, D., Buisman, H., Theobalt, C., and Thrun, S. (2008, January 18). A noise-aware filter for real-time depth upsampling. Proceedings of the ECCV Workshop on Multi-camera and Multi-modal Sensor Fusion Algorithms and Applications, Marseille, France."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Li, Y., Zhang, L., Zhang, Y., Xuan, H., and Dai, Q. (2014, January 7\u201310). Depth map super-resolution via iterative joint-trilateral-upsampling. Proceedings of the IEEE Visual Communications and Image Processing (VCIP), Valletta, Malta.","DOI":"10.1109\/VCIP.2014.7051587"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Yang, Q., Yang, R., Davis, J., and Nister, D. (2007, January 18\u201323). Spatial-depth super-resolution for range images. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Minneapolis, MN, USA.","DOI":"10.1109\/CVPR.2007.383211"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Kim, J., Lee, J., Han, S., Kim, D., Min, J., and Kim, C. (2013, January 10\u201312). Trilateral filter construction for depth map upsampling. Proceedings of the IEEE IVMSP Workshop, Seoul, Korea.","DOI":"10.1109\/IVMSPW.2013.6611911"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Liu, M.Y., Tuzel, O., and Taguchi, Y. (2013, January 23\u201328). Joint geodesic upsampling of depth images. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.29"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"258","DOI":"10.1109\/TCSVT.2012.2203734","article-title":"Enhancement of Image and Depth Map Using Adaptive Joint Trilateral Filter","volume":"23","author":"Jung","year":"2013","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"3321","DOI":"10.1109\/TIP.2014.2329766","article-title":"A Consensus-Driven Approach for Structure and Texture Aware Depth Map Upsampling","volume":"23","author":"Choi","year":"2014","journal-title":"IEEE Trans. Image Process."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1176","DOI":"10.1109\/TIP.2011.2163164","article-title":"Depth Video Enhancement Based on Weighted Mode Filtering","volume":"21","author":"Min","year":"2012","journal-title":"IEEE Trans. Image Process."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1109\/MMUL.2015.52","article-title":"Extended Guided Filtering for Depth Map Upsampling","volume":"23","author":"Hua","year":"2016","journal-title":"IEEE MultiMedia Mag."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"371","DOI":"10.1109\/TCYB.2016.2637661","article-title":"Edge-preserving depth map upsampling by joint trilateral filter","volume":"48","author":"Lo","year":"2018","journal-title":"IEEE Trans. Cybern."},{"key":"ref_22","unstructured":"Diebel, J., and Thrun, S. An application of Markov random fields to range sensing. Proceedings of the MIT Press Conference on Neural Information Processing Systems (NIPS), Vancouver, BC, Canada, 5\u20138 December 2005."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Lu, J., Min, D., Pahwa, R.S., and Do, M.N. (2011, January 22\u201327). A revisit to MRF-based depth map super-resolution and enhancement. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Prague, Czech Republic.","DOI":"10.1109\/ICASSP.2011.5946571"},{"key":"ref_24","unstructured":"Kim, D., and Yoon, K. (October, January 30). High quality depth map up-sampling robust to edge noise of range sensors. Proceedings of the IEEE International Conference on Image Processing (ICIP), Orlando, FL, USA."},{"key":"ref_25","unstructured":"Lu, J., and Forsyth, D. (, January 7\u201312). Sparse Depth super-resolution. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Park, J., Kim, H., Tai, Y.W., Brown, M., and Kweon, I. (2011, January 6\u201313). High quality depth map upsampling for 3d-tof cameras. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126423"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Ferstl, D., Reinbacher, C., Ranftl, R., Ruether, M., and Bischof, H. (2013, January 1\u20138). Image guided depth upsampling using anisotropic total generalized variation. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Sydney, Australia.","DOI":"10.1109\/ICCV.2013.127"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Ledig, C., Theis, L., Huszar, F., Caballero, J., Aitken, A., Tejani, A., Totz, J., Wang, Z., and Shi, W. (2017, January 21\u201326). Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.19"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Lim, B., Son, S., Kim, H., Nah, S., and Lee, K.M. (2017, January 21\u201326). Enhanced Deep Residual Networks for Single Image Super-Resolution. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.151"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Shi, W., Caballero, J., Husz\u00e1r, F., Totz, J., Aitken, A.P., Bishop, R., Rueckert, D., and Wang, Z. (2016, January 27\u201330). Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.207"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Silberman, N., Hoiem, D., Kohli, P., and Fergus, R. (2012, January 7\u201313). Indoor segmentation and support inference from RGBD images. Proceedings of the European Conference on Computer Vision (ECCV), Florence, Italy.","DOI":"10.1007\/978-3-642-33715-4_54"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1023\/A:1014573219977","article-title":"A taxonomy and evaluation of dense two-frame stereo correspondence algorithms","volume":"47","author":"Scharstein","year":"2001","journal-title":"Int. J. Comput. Vis."},{"key":"ref_33","unstructured":"Kingma, D., and Ba, J. (2015, January 7\u20139). Adam: A method for stochastic optimization. Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA."},{"key":"ref_34","unstructured":"(2019, January 04). Middlebury Stereo. Available online: http:\/\/vision.middlebury.edu\/stereo\/."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/7\/1587\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:42:16Z","timestamp":1760186536000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/7\/1587"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,4,2]]},"references-count":34,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2019,4]]}},"alternative-id":["s19071587"],"URL":"https:\/\/doi.org\/10.3390\/s19071587","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2019,4,2]]}}}