{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T16:08:05Z","timestamp":1784131685216,"version":"3.55.0"},"reference-count":213,"publisher":"Association for Computing Machinery (ACM)","issue":"12","license":[{"start":{"date-parts":[[2024,10,3]],"date-time":"2024-10-03T00:00:00Z","timestamp":1727913600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2024,12,31]]},"abstract":"<jats:p>Estimating depth from single RGB images and videos is of widespread interest due to its applications in many areas, including autonomous driving, 3D reconstruction, digital entertainment, and robotics. More than 500 deep learning-based papers have been published in the past 10 years, which indicates the growing interest in the task. This paper presents a comprehensive survey of the existing deep learning-based methods, the challenges they address, and how they have evolved in their architecture and supervision methods. It provides a taxonomy for classifying the current work based on their input and output modalities, network architectures, and learning methods. It also discusses the major milestones in the history of monocular depth estimation, and different pipelines, datasets, and evaluation metrics used in existing methods.<\/jats:p>","DOI":"10.1145\/3677327","type":"journal-article","created":{"date-parts":[[2024,7,15]],"date-time":"2024-07-15T11:03:48Z","timestamp":1721041428000},"page":"1-51","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":60,"title":["Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey"],"prefix":"10.1145","volume":"56","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1174-8263","authenticated-orcid":false,"given":"Uchitha","family":"Rajapaksha","sequence":"first","affiliation":[{"name":"School of Information Technology, Murdoch University, Murdoch, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1557-4907","authenticated-orcid":false,"given":"Ferdous","family":"Sohel","sequence":"additional","affiliation":[{"name":"School of Information Technology, Murdoch University, Murdoch, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4758-7510","authenticated-orcid":false,"given":"Hamid","family":"Laga","sequence":"additional","affiliation":[{"name":"School of Information Technology, Murdoch University, Murdoch, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1535-8019","authenticated-orcid":false,"given":"Dean","family":"Diepeveen","sequence":"additional","affiliation":[{"name":"Murdoch University, Murdoch, Australia and Western Australia Department of Primary Industries and Regional Development, South Perth, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6603-3257","authenticated-orcid":false,"given":"Mohammed","family":"Bennamoun","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Software Engineering, The University of Western Australia, Perth Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,10,3]]},"reference":[{"issue":"2","key":"e_1_3_3_2_2","doi-asserted-by":"crossref","first-page":"153","DOI":"10.1007\/s11263-016-0902-9","article-title":"Large-scale data for multiple-view stereopsis","volume":"120","author":"Aan\u00e6s Henrik","year":"2016","unstructured":"Henrik Aan\u00e6s, Rasmus Ramsb\u00f8l Jensen, George Vogiatzis, Engin Tola, and Anders Bjorholm Dahl. 2016. Large-scale data for multiple-view stereopsis. International Journal of Computer Vision 120, 2 (2016), 153\u2013168.","journal-title":"International Journal of Computer Vision"},{"key":"e_1_3_3_3_2","first-page":"0","volume-title":"Proceedings of the European Conference on Computer Vision Workshops","author":"Aleotti Filippo","year":"2018","unstructured":"Filippo Aleotti, Fabio Tosi, Matteo Poggi, and Stefano Mattoccia. 2018. Generative adversarial networks for unsupervised monocular depth prediction. In Proceedings of the European Conference on Computer Vision Workshops. 0\u20130."},{"key":"e_1_3_3_4_2","doi-asserted-by":"crossref","first-page":"602","DOI":"10.1109\/ROBIO49542.2019.8961504","volume-title":"2019 IEEE International Conference on Robotics and Biomimetics (ROBIO\u201919)","author":"Amiri Ali Jahani","year":"2019","unstructured":"Ali Jahani Amiri, Shing Yan Loo, and Hong Zhang. 2019. Semi-supervised monocular depth estimation with left-right consistency using deep neural network. In 2019 IEEE International Conference on Robotics and Biomimetics (ROBIO\u201919). IEEE, 602\u2013607."},{"key":"e_1_3_3_5_2","first-page":"2800","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Atapour-Abarghouei Amir","year":"2018","unstructured":"Amir Atapour-Abarghouei and Toby P. Breckon. 2018. Real-time monocular depth estimation using synthetic data with domain adaptation via image style transfer. In IEEE Conference on Computer Vision and Pattern Recognition. 2800\u20132810."},{"key":"e_1_3_3_6_2","first-page":"2039","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Auty Dylan","year":"2023","unstructured":"Dylan Auty and Krystian Mikolajczyk. 2023. Learning to prompt CLIP for monocular depth estimation: Exploring the limits of human language. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 2039\u20132047."},{"key":"e_1_3_3_7_2","article-title":"A study on the generality of neural network structures for monocular depth estimation","author":"Bae Jinwoo","year":"2023","unstructured":"Jinwoo Bae, Kyumin Hwang, and Sunghoon Im. 2023. A study on the generality of neural network structures for monocular depth estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence (2023).","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_8_2","article-title":"Deep digging into the generalization of self-supervised monocular depth estimation","author":"Bae Jinwoo","year":"2022","unstructured":"Jinwoo Bae, Sungho Moon, and Sunghoon Im. 2022. Deep digging into the generalization of self-supervised monocular depth estimation. arXiv preprint arXiv:2205.11083 (2022).","journal-title":"arXiv preprint arXiv:2205.11083"},{"key":"e_1_3_3_9_2","article-title":"MonoFormer: Towards generalization of self-supervised monocular depth estimation with transformers","author":"Bae Jinwoo","year":"2022","unstructured":"Jinwoo Bae, Sungho Moon, and Sunghoon Im. 2022. MonoFormer: Towards generalization of self-supervised monocular depth estimation with transformers. arXiv preprint arXiv:2205.11083 (2022).","journal-title":"arXiv preprint arXiv:2205.11083"},{"key":"e_1_3_3_10_2","first-page":"187","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"37","author":"Bae Jinwoo","year":"2023","unstructured":"Jinwoo Bae, Sungho Moon, and Sunghoon Im. 2023. Deep digging into the generalization of self-supervised monocular depth estimation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 187\u2013196."},{"key":"e_1_3_3_11_2","article-title":"Semi-supervised learning with mutual distillation for monocular depth estimation","author":"Baek Jongbeom","year":"2022","unstructured":"Jongbeom Baek, Gyeongnyeon Kim, and Seungryong Kim. 2022. Semi-supervised learning with mutual distillation for monocular depth estimation. arXiv preprint arXiv:2203.09737 (2022).","journal-title":"arXiv preprint arXiv:2203.09737"},{"key":"e_1_3_3_12_2","article-title":"Self-supervised deep monocular depth estimation with ambiguity boosting","author":"Bello Juan Luis Gonzalez","year":"2021","unstructured":"Juan Luis Gonzalez Bello and Munchurl Kim. 2021. Self-supervised deep monocular depth estimation with ambiguity boosting. IEEE Transactions on Pattern Analysis and Machine Intelligence (2021).","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_13_2","article-title":"Self-supervised monocular depth estimation with positional shift depth variance and adaptive disparity quantization","author":"Bello Juan Luis Gonzalez","year":"2024","unstructured":"Juan Luis Gonzalez Bello, Jaeho Moon, and Munchurl Kim. 2024. Self-supervised monocular depth estimation with positional shift depth variance and adaptive disparity quantization. IEEE Transactions on Image Processing (2024).","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_3_14_2","first-page":"4009","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Bhat Shariq Farooq","year":"2021","unstructured":"Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka. 2021. AdaBins: Depth estimation using adaptive bins. In IEEE Conference on Computer Vision and Pattern Recognition. 4009\u20134018."},{"key":"e_1_3_3_15_2","article-title":"Monocular depth estimation: A survey","author":"Bhoi Amlaan","year":"2019","unstructured":"Amlaan Bhoi. 2019. Monocular depth estimation: A survey. arXiv preprint arXiv:1901.09402 (2019).","journal-title":"arXiv preprint arXiv:1901.09402"},{"key":"e_1_3_3_16_2","article-title":"Unsupervised scale-consistent depth and ego-motion learning from monocular video","volume":"32","author":"Bian Jiawang","year":"2019","unstructured":"Jiawang Bian, Zhichao Li, Naiyan Wang, Huangying Zhan, Chunhua Shen, Ming-Ming Cheng, and Ian Reid. 2019. Unsupervised scale-consistent depth and ego-motion learning from monocular video. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_17_2","article-title":"TransformerFusion: Monocular RGB scene reconstruction using transformers","volume":"34","author":"Bozic Aljaz","year":"2021","unstructured":"Aljaz Bozic, Pablo Palafox, Justus Thies, Angela Dai, and Matthias Nie\u00dfner. 2021. TransformerFusion: Monocular RGB scene reconstruction using transformers. Advances in Neural Information Processing Systems 34 (2021).","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"10","key":"e_1_3_3_18_2","doi-asserted-by":"crossref","first-page":"1157","DOI":"10.1177\/0278364915620033","article-title":"The EuRoC micro aerial vehicle datasets","volume":"35","author":"Burri Michael","year":"2016","unstructured":"Michael Burri, Janosch Nikolic, Pascal Gohl, Thomas Schneider, Joern Rehder, Sammy Omari, Markus W. Achtelik, and Roland Siegwart. 2016. The EuRoC micro aerial vehicle datasets. The International Journal of Robotics Research 35, 10 (2016), 1157\u20131163.","journal-title":"The International Journal of Robotics Research"},{"key":"e_1_3_3_19_2","first-page":"611","volume-title":"European Conference on Computer Vision","author":"Butler Daniel J.","year":"2012","unstructured":"Daniel J. Butler, Jonas Wulff, Garrett B. Stanley, and Michael J. Black. 2012. A naturalistic open source movie for optical flow evaluation. In European Conference on Computer Vision. Springer, 611\u2013625."},{"issue":"11","key":"e_1_3_3_20_2","first-page":"3174","article-title":"Estimating depth from monocular images as classification using deep fully convolutional residual networks","volume":"28","author":"Cao Yuanzhouhan","year":"2017","unstructured":"Yuanzhouhan Cao, Zifeng Wu, and Chunhua Shen. 2017. Estimating depth from monocular images as classification using deep fully convolutional residual networks. IEEE Transactions on Circuits and Systems for Video Technology 28, 11 (2017), 3174\u20133182.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_3_21_2","first-page":"0","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"Casser Vincent","year":"2019","unstructured":"Vincent Casser, Soeren Pirk, Reza Mahjourian, and Anelia Angelova. 2019. Unsupervised monocular depth and ego-motion learning with structure and semantics. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops. 0\u20130."},{"key":"e_1_3_3_22_2","first-page":"2624","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Chen Po-Yi","year":"2019","unstructured":"Po-Yi Chen, Alexander H. Liu, Yen-Cheng Liu, and Yu-Chiang Frank Wang. 2019. Towards scene understanding: Unsupervised monocular depth estimation with semantic-aware representation. In IEEE Conference on Computer Vision and Pattern Recognition. 2624\u20132632."},{"key":"e_1_3_3_23_2","article-title":"Rethinking monocular depth estimation with adversarial training","author":"Chen Richard","year":"2018","unstructured":"Richard Chen, Faisal Mahmood, Alan Yuille, and Nicholas J. Durr. 2018. Rethinking monocular depth estimation with adversarial training. arXiv preprint arXiv:1808.07528 (2018).","journal-title":"arXiv preprint arXiv:1808.07528"},{"key":"e_1_3_3_24_2","first-page":"90","volume-title":"European Conference on Computer Vision","author":"Chen Tian","year":"2020","unstructured":"Tian Chen, Shijie An, Yuan Zhang, Chongyang Ma, Huayan Wang, Xiaoyan Guo, and Wen Zheng. 2020. Improving monocular depth estimation by leveraging structural awareness and complementary datasets. In European Conference on Computer Vision. Springer, 90\u2013108."},{"key":"e_1_3_3_25_2","article-title":"Single-image depth perception in the wild","volume":"29","author":"Chen Weifeng","year":"2016","unstructured":"Weifeng Chen, Zhao Fu, Dawei Yang, and Jia Deng. 2016. Single-image depth perception in the wild. Advances in Neural Information Processing Systems 29 (2016).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_26_2","first-page":"3034","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Xiaotian","year":"2021","unstructured":"Xiaotian Chen, Yuwang Wang, Xuejin Chen, and Wenjun Zeng. 2021. S2R-DepthNet: Learning a generalizable depth-specific structural representation. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3034\u20133043."},{"issue":"6","key":"e_1_3_3_27_2","doi-asserted-by":"crossref","first-page":"1583","DOI":"10.1007\/s13042-020-01251-y","article-title":"Attention-based context aggregation network for monocular depth estimation","volume":"12","author":"Chen Yuru","year":"2021","unstructured":"Yuru Chen, Haitao Zhao, Zhengwei Hu, and Jingchao Peng. 2021. Attention-based context aggregation network for monocular depth estimation. International Journal of Machine Learning and Cybernetics 12, 6 (2021), 1583\u20131596.","journal-title":"International Journal of Machine Learning and Cybernetics"},{"key":"e_1_3_3_28_2","first-page":"15529","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Chen Zhi","year":"2021","unstructured":"Zhi Chen, Xiaoqing Ye, Wei Yang, Zhenbo Xu, Xiao Tan, Zhikang Zou, Errui Ding, Xinming Zhang, and Liusheng Huang. 2021. Revealing the reciprocal relations between self-supervised stereo and monocular depth estimation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 15529\u201315538."},{"issue":"23","key":"e_1_3_3_29_2","doi-asserted-by":"crossref","first-page":"26912","DOI":"10.1109\/JSEN.2021.3120753","article-title":"Swin-Depth: Using transformers and multi-scale fusion for monocular-based depth estimation","volume":"21","author":"Cheng Zeyu","year":"2021","unstructured":"Zeyu Cheng, Yi Zhang, and Chengkai Tang. 2021. Swin-Depth: Using transformers and multi-scale fusion for monocular-based depth estimation. IEEE Sensors Journal 21, 23 (2021), 26912\u201326920.","journal-title":"IEEE Sensors Journal"},{"key":"e_1_3_3_30_2","doi-asserted-by":"crossref","first-page":"114877","DOI":"10.1016\/j.eswa.2021.114877","article-title":"Deep monocular depth estimation leveraging a large-scale outdoor stereo dataset","volume":"178","author":"Cho Jaehoon","year":"2021","unstructured":"Jaehoon Cho, Dongbo Min, Youngjung Kim, and Kwanghoon Sohn. 2021. Deep monocular depth estimation leveraging a large-scale outdoor stereo dataset. Expert Systems with Applications 178 (2021), 114877.","journal-title":"Expert Systems with Applications"},{"key":"e_1_3_3_31_2","first-page":"3213","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Cordts Marius","year":"2016","unstructured":"Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. 2016. The Cityscapes dataset for semantic urban scene understanding. In IEEE Conference on Computer Vision and Pattern Recognition. 3213\u20133223."},{"key":"e_1_3_3_32_2","first-page":"283","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition Workshops","author":"Kumar Arun C. S.","year":"2018","unstructured":"Arun C. S. Kumar, Suchendra M. Bhandarkar, and Mukta Prasad. 2018. DepthNet: A recurrent neural network architecture for monocular depth prediction. In IEEE Conference on Computer Vision and Pattern Recognition Workshops. 283\u2013291."},{"key":"e_1_3_3_33_2","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition Workshops","author":"Kumar Arun C. S.","year":"2018","unstructured":"Arun C. S. Kumar, Suchendra M. Bhandarkar, and Mukta Prasad. 2018. Monocular depth prediction using generative adversarial networks. In IEEE Conference on Computer Vision and Pattern Recognition Workshops."},{"key":"e_1_3_3_34_2","first-page":"4738","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Diaz Raul","year":"2019","unstructured":"Raul Diaz and Amit Marathe. 2019. Soft labels for ordinal regression. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 4738\u20134747."},{"key":"e_1_3_3_35_2","first-page":"2183","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Dijk Tom van","year":"2019","unstructured":"Tom van Dijk and Guido de Croon. 2019. How do neural networks see depth in single images?. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 2183\u20132191."},{"key":"e_1_3_3_36_2","first-page":"2650","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Eigen David","year":"2015","unstructured":"David Eigen and Rob Fergus. 2015. Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. In Proceedings of the IEEE International Conference on Computer Vision. 2650\u20132658."},{"key":"e_1_3_3_37_2","article-title":"Depth map prediction from a single image using a multi-scale deep network","volume":"27","author":"Eigen David","year":"2014","unstructured":"David Eigen, Christian Puhrsch, and Rob Fergus. 2014. Depth map prediction from a single image using a multi-scale deep network. Advances in Neural Information Processing Systems 27 (2014).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_38_2","doi-asserted-by":"crossref","first-page":"4290","DOI":"10.1109\/ICIP.2019.8803544","volume-title":"2019 IEEE International Conference on Image Processing (ICIP\u201919)","author":"Elkerdawy Sara","year":"2019","unstructured":"Sara Elkerdawy, Hong Zhang, and Nilanjan Ray. 2019. Lightweight monocular depth estimation model by joint end-to-end filter pruning. In 2019 IEEE International Conference on Image Processing (ICIP\u201919). IEEE, 4290\u20134294."},{"key":"e_1_3_3_39_2","first-page":"2002","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Fu Huan","year":"2018","unstructured":"Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao. 2018. Deep ordinal regression network for monocular depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition. 2002\u20132011."},{"key":"e_1_3_3_40_2","first-page":"740","volume-title":"European Conference on Computer Vision","author":"Garg Ravi","year":"2016","unstructured":"Ravi Garg, Vijay Kumar Bg, Gustavo Carneiro, and Ian Reid. 2016. Unsupervised CNN for single view depth estimation: Geometry to the rescue. In European Conference on Computer Vision. Springer, 740\u2013756."},{"key":"e_1_3_3_41_2","first-page":"8177","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Gasperini Stefano","year":"2023","unstructured":"Stefano Gasperini, Nils Morbitzer, HyunJun Jung, Nassir Navab, and Federico Tombari. 2023. Robust monocular depth estimation under challenging conditions. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 8177\u20138186."},{"issue":"2","key":"e_1_3_3_42_2","doi-asserted-by":"crossref","first-page":"2822","DOI":"10.1109\/LRA.2021.3060707","article-title":"Combining events and frames using recurrent asynchronous multimodal networks for monocular depth prediction","volume":"6","author":"Gehrig Daniel","year":"2021","unstructured":"Daniel Gehrig, Michelle R\u00fcegg, Mathias Gehrig, Javier Hidalgo-Carri\u00f3, and Davide Scaramuzza. 2021. Combining events and frames using recurrent asynchronous multimodal networks for monocular depth prediction. IEEE Robotics and Automation Letters 6, 2 (2021), 2822\u20132829.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"e_1_3_3_43_2","article-title":"A2D2: Audi autonomous driving dataset","author":"Geyer Jakob","year":"2020","unstructured":"Jakob Geyer, Yohannes Kassahun, Mentar Mahmudi, Xavier Ricou, Rupesh Durgesh, Andrew S. Chung, Lorenz Hauswald, Viet Hoang Pham, Maximilian M\u00fchlegg, Sebastian Dorn, Tiffany Fernandez, Martin J\u00e4nicke, Sudesh Mirashi, Chiragkumar Savani, Martin Sturm, Oleksandr Vorobiov, Martin Oelker, Sebastian Garreis, and Peter Schuberth. 2020. A2D2: Audi autonomous driving dataset. arXiv preprint arXiv:2004.06320 (2020).","journal-title":"arXiv preprint arXiv:2004.06320"},{"key":"e_1_3_3_44_2","first-page":"270","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Godard Cl\u00e9ment","year":"2017","unstructured":"Cl\u00e9ment Godard, Oisin Mac Aodha, and Gabriel J. Brostow. 2017. Unsupervised monocular depth estimation with left-right consistency. In IEEE Conference on Computer Vision and Pattern Recognition. 270\u2013279."},{"key":"e_1_3_3_45_2","first-page":"3828","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Godard Cl\u00e9ment","year":"2019","unstructured":"Cl\u00e9ment Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J. Brostow. 2019. Digging into self-supervised monocular depth estimation. In Proceedings of the IEEE International Conference on Computer Vision. 3828\u20133838."},{"key":"e_1_3_3_46_2","first-page":"2485","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Guizilini Vitor","year":"2020","unstructured":"Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Allan Raventos, and Adrien Gaidon. 2020. 3D packing for self-supervised monocular depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition. 2485\u20132494."},{"key":"e_1_3_3_47_2","first-page":"503","volume-title":"Conference on Robot Learning","author":"Guizilini Vitor","year":"2020","unstructured":"Vitor Guizilini, Jie Li, Rares Ambrus, Sudeep Pillai, and Adrien Gaidon. 2020. Robust semi-supervised monocular depth estimation with reprojected distances. In Conference on Robot Learning. PMLR, 503\u2013512."},{"key":"e_1_3_3_48_2","first-page":"9233","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Guizilini Vitor","year":"2023","unstructured":"Vitor Guizilini, Igor Vasiljevic, Dian Chen, Rare\u015f Ambru\u015f, and Adrien Gaidon. 2023. Towards zero-shot scale-aware monocular depth estimation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 9233\u20139243."},{"key":"e_1_3_3_49_2","first-page":"1432","volume-title":"2019 IEEE Intelligent Transportation Systems Conference (ITSC\u201919)","author":"Guo Rui","year":"2019","unstructured":"Rui Guo, Babajide Ayinde, Hao Sun, Haritha Muralidharan, and Kentaro Oguchi. 2019. Monocular depth estimation using synthetic images with shadow removal. In 2019 IEEE Intelligent Transportation Systems Conference (ITSC\u201919). IEEE, 1432\u20131439."},{"issue":"5","key":"e_1_3_3_50_2","first-page":"1578","article-title":"Image-based 3D object reconstruction: State-of-the-art and trends in the deep learning era","volume":"43","author":"Han Xian-Feng","year":"2019","unstructured":"Xian-Feng Han, Hamid Laga, and Mohammed Bennamoun. 2019. Image-based 3D object reconstruction: State-of-the-art and trends in the deep learning era. IEEE Transactions on Pattern Analysis and Machine Intelligence 43, 5 (2019), 1578\u20131604.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_51_2","first-page":"304","volume-title":"2018 International Conference on 3D Vision (3DV\u201918)","author":"Hao Zhixiang","year":"2018","unstructured":"Zhixiang Hao, Yu Li, Shaodi You, and Feng Lu. 2018. Detail preserving depth estimation from a single image using attention guided networks. In 2018 International Conference on 3D Vision (3DV\u201918). IEEE, 304\u2013313."},{"key":"e_1_3_3_52_2","doi-asserted-by":"crossref","first-page":"251","DOI":"10.1016\/j.neucom.2021.01.126","article-title":"SOSD-Net: Joint semantic object segmentation and depth estimation from monocular images","volume":"440","author":"He Lei","year":"2021","unstructured":"Lei He, Jiwen Lu, Guanghui Wang, Shiyu Song, and Jie Zhou. 2021. SOSD-Net: Joint semantic object segmentation and depth estimation from monocular images. Neurocomputing 440 (2021), 251\u2013263.","journal-title":"Neurocomputing"},{"key":"e_1_3_3_53_2","first-page":"565","volume-title":"European Conference on Computer Vision","author":"He Mu","year":"2022","unstructured":"Mu He, Le Hui, Yikai Bian, Jian Ren, Jin Xie, and Jian Yang. 2022. RA-Depth: Resolution adaptive self-supervised monocular depth estimation. In European Conference on Computer Vision. Springer, 565\u2013581."},{"key":"e_1_3_3_54_2","first-page":"36","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201918)","author":"Heo Minhyeok","year":"2018","unstructured":"Minhyeok Heo, Jaehan Lee, Kyung-Rae Kim, Han-Ul Kim, and Chang-Su Kim. 2018. Monocular depth estimation using whole strip masking and reliability-based refinement. In Proceedings of the European Conference on Computer Vision (ECCV\u201918). 36\u201351."},{"key":"e_1_3_3_55_2","first-page":"3869","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Hu Junjie","year":"2019","unstructured":"Junjie Hu, Yan Zhang, and Takayuki Okatani. 2019. Visualization of convolutional neural networks for monocular depth estimation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 3869\u20133878."},{"key":"e_1_3_3_56_2","first-page":"5594","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Hu Xueting","year":"2024","unstructured":"Xueting Hu, Ce Zhang, Yi Zhang, Bowen Hai, Ke Yu, and Zhihai He. 2024. Learning to adapt CLIP for few-shot monocular depth estimation. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 5594\u20135603."},{"issue":"10","key":"e_1_3_3_57_2","doi-asserted-by":"crossref","first-page":"2702","DOI":"10.1109\/TPAMI.2019.2926463","article-title":"The ApolloScape open dataset for autonomous driving and its application","volume":"42","author":"Huang Xinyu","year":"2019","unstructured":"Xinyu Huang, Peng Wang, Xinjing Cheng, Dingfu Zhou, Qichuan Geng, and Ruigang Yang. 2019. The ApolloScape open dataset for autonomous driving and its application. IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 10 (2019), 2702\u20132719.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_58_2","doi-asserted-by":"crossref","first-page":"3429","DOI":"10.1109\/TIP.2019.2960589","article-title":"HMS-Net: Hierarchical multi-scale sparsity-invariant network for sparse depth completion","volume":"29","author":"Huang Zixuan","year":"2019","unstructured":"Zixuan Huang, Junming Fan, Shenggan Cheng, Shuai Yi, Xiaogang Wang, and Hongsheng Li. 2019. HMS-Net: Hierarchical multi-scale sparsity-invariant network for sparse depth completion. IEEE Transactions on Image Processing 29 (2019), 3429\u20133441.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_3_59_2","first-page":"1675","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Hui Tak-Wai","year":"2022","unstructured":"Tak-Wai Hui. 2022. RM-Depth: Unsupervised learning of recurrent monocular depth in dynamic scenes. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 1675\u20131684."},{"key":"e_1_3_3_60_2","first-page":"581","volume-title":"European Conference on Computer Vision","author":"Huynh Lam","year":"2020","unstructured":"Lam Huynh, Phong Nguyen-Ha, Jiri Matas, Esa Rahtu, and Janne Heikkil\u00e4. 2020. Guiding monocular depth estimation using depth-attention volume. In European Conference on Computer Vision. Springer, 581\u2013597."},{"key":"e_1_3_3_61_2","first-page":"12753","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Jafarian Yasamin","year":"2021","unstructured":"Yasamin Jafarian and Hyun Soo Park. 2021. Learning high fidelity depths of dressed humans by watching social media dance videos. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 12753\u201312762."},{"key":"e_1_3_3_62_2","doi-asserted-by":"crossref","DOI":"10.1109\/TPAMI.2022.3231558","article-title":"Self-supervised 3D representation learning of dressed humans from social media videos","author":"Jafarian Yasamin","year":"2022","unstructured":"Yasamin Jafarian and Hyun Soo Park. 2022. Self-supervised 3D representation learning of dressed humans from social media videos. IEEE Transactions on Pattern Analysis and Machine Intelligence (2022).","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_63_2","first-page":"12787","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Ji Pan","year":"2021","unstructured":"Pan Ji, Runze Li, Bir Bhanu, and Yi Xu. 2021. MonoIndoor: Towards good practice of self-supervised monocular depth estimation for indoor environments. In IEEE Conference on Computer Vision and Pattern Recognition. 12787\u201312796."},{"issue":"10","key":"e_1_3_3_64_2","first-page":"2410","article-title":"Semi-supervised adversarial monocular depth estimation","volume":"42","author":"Ji Rongrong","year":"2019","unstructured":"Rongrong Ji, Ke Li, Yan Wang, Xiaoshuai Sun, Feng Guo, Xiaowei Guo, Yongjian Wu, Feiyue Huang, and Jiebo Luo. 2019. Semi-supervised adversarial monocular depth estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 10 (2019), 2410\u20132422.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_65_2","first-page":"53","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201918)","author":"Jiao Jianbo","year":"2018","unstructured":"Jianbo Jiao, Ying Cao, Yibing Song, and Rynson Lau. 2018. Look deeper into depth: Monocular depth estimation with semantic booster and attention-driven loss. In Proceedings of the European Conference on Computer Vision (ECCV\u201918). 53\u201369."},{"key":"e_1_3_3_66_2","first-page":"4756","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Johnston Adrian","year":"2020","unstructured":"Adrian Johnston and Gustavo Carneiro. 2020. Self-supervised monocular trained depth estimation using self-attention and discrete disparity volume. In IEEE Conference on Computer Vision and Pattern Recognition. 4756\u20134765."},{"key":"e_1_3_3_67_2","doi-asserted-by":"crossref","first-page":"1717","DOI":"10.1109\/ICIP.2017.8296575","volume-title":"2017 IEEE International Conference on Image Processing (ICIP\u201917)","author":"Jung Hyungjoo","year":"2017","unstructured":"Hyungjoo Jung, Youngjung Kim, Dongbo Min, Changjae Oh, and Kwanghoon Sohn. 2017. Depth prediction from a single image with conditional adversarial networks. In 2017 IEEE International Conference on Image Processing (ICIP\u201917). IEEE, 1717\u20131721."},{"issue":"11","key":"e_1_3_3_68_2","doi-asserted-by":"crossref","first-page":"2144","DOI":"10.1109\/TPAMI.2014.2316835","article-title":"Depth transfer: Depth extraction from video using non-parametric sampling","volume":"36","author":"Karsch Kevin","year":"2014","unstructured":"Kevin Karsch, Ce Liu, and Sing Bing Kang. 2014. Depth transfer: Depth extraction from video using non-parametric sampling. IEEE Transactions on Pattern Analysis and Machine Intelligence 36, 11 (2014), 2144\u20132158.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_69_2","article-title":"Transformers in vision: A survey","author":"Khan Salman","year":"2021","unstructured":"Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. 2021. Transformers in vision: A survey. ACM Computing Surveys (CSUR) (2021).","journal-title":"ACM Computing Surveys (CSUR)"},{"key":"e_1_3_3_70_2","first-page":"3235","volume-title":"2020 IEEE Conference on Computer Vision and Pattern Recognition","author":"Kim Sungyeon","year":"2020","unstructured":"Sungyeon Kim, Dongwon Kim, Minsu Cho, and Suha Kwak. 2020. Proxy anchor loss for deep metric learning. In 2020 IEEE Conference on Computer Vision and Pattern Recognition. 3235\u20133244. DOI:10.1109\/CVPR42600.2020.00330"},{"issue":"8","key":"e_1_3_3_71_2","doi-asserted-by":"crossref","first-page":"4131","DOI":"10.1109\/TIP.2018.2836318","article-title":"Deep monocular depth estimation via integration of global and local predictions","volume":"27","author":"Kim Youngjung","year":"2018","unstructured":"Youngjung Kim, Hyungjoo Jung, Dongbo Min, and Kwanghoon Sohn. 2018. Deep monocular depth estimation via integration of global and local predictions. IEEE Transactions on Image Processing 27, 8 (2018), 4131\u20134144.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_3_72_2","first-page":"4015","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Kirillov Alexander","year":"2023","unstructured":"Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Doll\u00e1r, and Ross Girshick. 2023. Segment anything. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 4015\u20134026."},{"key":"e_1_3_3_73_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00166"},{"key":"e_1_3_3_74_2","first-page":"596","volume-title":"2012 19th International Conference on Systems, Signals and Image Processing (IWSSIP\u201912)","author":"Kostadinov Dimce","year":"2012","unstructured":"Dimce Kostadinov and Zoran Ivanovski. 2012. Single image depth estimation using local gradient-based features. In 2012 19th International Conference on Systems, Signals and Image Processing (IWSSIP\u201912). IEEE, 596\u2013599."},{"key":"e_1_3_3_75_2","first-page":"2656","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Kundu Jogendra Nath","year":"2018","unstructured":"Jogendra Nath Kundu, Phani Krishna Uppala, Anuj Pahuja, and R. Venkatesh Babu. 2018. AdaDepth: Unsupervised content congruent adaptation for depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition. 2656\u20132665."},{"key":"e_1_3_3_76_2","first-page":"2907","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Kuznietsov Yevhen","year":"2021","unstructured":"Yevhen Kuznietsov, Marc Proesmans, and Luc Van Gool. 2021. CoMoDA: Continuous monocular depth adaptation using past experiences. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 2907\u20132917."},{"key":"e_1_3_3_77_2","first-page":"6647","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Kuznietsov Yevhen","year":"2017","unstructured":"Yevhen Kuznietsov, Jorg Stuckler, and Bastian Leibe. 2017. Semi-supervised deep learning for monocular depth map prediction. In IEEE Conference on Computer Vision and Pattern Recognition. 6647\u20136655."},{"issue":"4","key":"e_1_3_3_78_2","doi-asserted-by":"crossref","first-page":"1738","DOI":"10.1109\/TPAMI.2020.3032602","article-title":"A survey on deep learning techniques for stereo-based depth estimation","volume":"44","author":"Laga Hamid","year":"2020","unstructured":"Hamid Laga, Laurent Valentin Jospin, Farid Boussaid, and Mohammed Bennamoun. 2020. A survey on deep learning techniques for stereo-based depth estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 4 (2020), 1738\u20131764.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_79_2","doi-asserted-by":"crossref","first-page":"239","DOI":"10.1109\/3DV.2016.32","volume-title":"2016 Fourth International Conference on 3D Vision (3DV\u201916)","author":"Laina Iro","year":"2016","unstructured":"Iro Laina, Christian Rupprecht, Vasileios Belagiannis, Federico Tombari, and Nassir Navab. 2016. Deeper depth prediction with fully convolutional residual networks. In 2016 Fourth International Conference on 3D Vision (3DV\u201916). IEEE, 239\u2013248."},{"key":"e_1_3_3_80_2","article-title":"From big to small: Multi-scale local planar guidance for monocular depth estimation","author":"Lee Jin Han","year":"2019","unstructured":"Jin Han Lee, Myung-Kyu Han, Dong Wook Ko, and Il Hong Suh. 2019. From big to small: Multi-scale local planar guidance for monocular depth estimation. arXiv preprint arXiv:1907.10326 (2019).","journal-title":"arXiv preprint arXiv:1907.10326"},{"key":"e_1_3_3_81_2","first-page":"330","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Lee Jae-Han","year":"2018","unstructured":"Jae-Han Lee, Minhyeok Heo, Kyung-Rae Kim, and Chang-Su Kim. 2018. Single-image depth estimation based on Fourier domain analysis. In IEEE Conference on Computer Vision and Pattern Recognition. 330\u2013339."},{"key":"e_1_3_3_82_2","first-page":"2858","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Lee Minhyeok","year":"2022","unstructured":"Minhyeok Lee, Sangwon Hwang, Chaewon Park, and Sangyoun Lee. 2022. EdgeConv with attention module for monocular depth estimation. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 2858\u20132867."},{"key":"e_1_3_3_83_2","doi-asserted-by":"crossref","first-page":"343","DOI":"10.1016\/j.neucom.2020.11.002","article-title":"Attention based multilayer feature fusion convolutional neural network for unsupervised monocular depth estimation","volume":"423","author":"Lei Zeyu","year":"2021","unstructured":"Zeyu Lei, Yan Wang, Zijian Li, and Junyao Yang. 2021. Attention based multilayer feature fusion convolutional neural network for unsupervised monocular depth estimation. Neurocomputing 423 (2021), 343\u2013352.","journal-title":"Neurocomputing"},{"key":"e_1_3_3_84_2","first-page":"1119","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Li Bo","year":"2015","unstructured":"Bo Li, Chunhua Shen, Yuchao Dai, Anton van den Hengel, and Mingyi He. 2015. Depth and surface normal estimation from monocular images using regression on deep features and hierarchical CRFs. In IEEE Conference on Computer Vision and Pattern Recognition. 1119\u20131127."},{"key":"e_1_3_3_85_2","first-page":"663","volume-title":"Asian Conference on Computer Vision","author":"Li Ruibo","year":"2018","unstructured":"Ruibo Li, Ke Xian, Chunhua Shen, Zhiguo Cao, Hao Lu, and Lingxiao Hang. 2018. Deep attention-based classification network for robust depth prediction. In Asian Conference on Computer Vision. Springer, 663\u2013678."},{"key":"e_1_3_3_86_2","first-page":"2041","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Li Zhengqi","year":"2018","unstructured":"Zhengqi Li and Noah Snavely. 2018. MegaDepth: Learning single-view depth prediction from internet photos. In IEEE Conference on Computer Vision and Pattern Recognition. 2041\u20132050."},{"key":"e_1_3_3_87_2","article-title":"KITTI-360: A novel dataset and benchmarks for urban scene understanding in 2D and 3D","author":"Liao Yiyi","year":"2022","unstructured":"Yiyi Liao, Jun Xie, and Andreas Geiger. 2022. KITTI-360: A novel dataset and benchmarks for urban scene understanding in 2D and 3D. IEEE Transactions on Pattern Analysis and Machine Intelligence (2022).","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_88_2","first-page":"14595","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lienen Julian","year":"2021","unstructured":"Julian Lienen, Eyke Hullermeier, Ralph Ewerth, and Nils Nommensen. 2021. Monocular depth estimation via listwise ranking using the Plackett-Luce model. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 14595\u201314604."},{"issue":"11","key":"e_1_3_3_89_2","doi-asserted-by":"crossref","first-page":"4545","DOI":"10.1109\/TIP.2013.2274389","article-title":"Absolute depth estimation from a single defocused image","volume":"22","author":"Lin Jingyu","year":"2013","unstructured":"Jingyu Lin, Xiangyang Ji, Wenli Xu, and Qionghai Dai. 2013. Absolute depth estimation from a single defocused image. IEEE Transactions on Image Processing 22, 11 (2013), 4545\u20134550.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_3_90_2","article-title":"Unsupervised monocular depth estimation using attention and multi-warp reconstruction","author":"Ling Chuanwu","year":"2021","unstructured":"Chuanwu Ling, Xiaogang Zhang, and Hua Chen. 2021. Unsupervised monocular depth estimation using attention and multi-warp reconstruction. IEEE Transactions on Multimedia (2021).","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_3_91_2","first-page":"5162","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Liu Fayao","year":"2015","unstructured":"Fayao Liu, Chunhua Shen, and Guosheng Lin. 2015. Deep convolutional neural fields for depth estimation from a single image. In IEEE Conference on Computer Vision and Pattern Recognition. 5162\u20135170."},{"issue":"10","key":"e_1_3_3_92_2","first-page":"2024","article-title":"Learning depth from single monocular images using deep convolutional neural fields","volume":"38","author":"Liu Fayao","year":"2015","unstructured":"Fayao Liu, Chunhua Shen, Guosheng Lin, and Ian Reid. 2015. Learning depth from single monocular images using deep convolutional neural fields. IEEE Transactions on Pattern Analysis and Machine Intelligence 38, 10 (2015), 2024\u20132039.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_93_2","first-page":"5137","volume-title":"2020 25th International Conference on Pattern Recognition (ICPR\u201921)","author":"Liu Jing","year":"2021","unstructured":"Jing Liu, Xiaona Zhang, Zhaoxin Li, and Tianlu Mao. 2021. Multi-scale residual pyramid attention network for monocular depth estimation. In 2020 25th International Conference on Pattern Recognition (ICPR\u201921). IEEE, 5137\u20135144."},{"key":"e_1_3_3_94_2","first-page":"12737","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Liu Lina","year":"2021","unstructured":"Lina Liu, Xibin Song, Mengmeng Wang, Yong Liu, and Liangjun Zhang. 2021. Self-supervised monocular depth estimation for all day images using domain separation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 12737\u201312746."},{"key":"e_1_3_3_95_2","first-page":"716","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Liu Miaomiao","year":"2014","unstructured":"Miaomiao Liu, Mathieu Salzmann, and Xuming He. 2014. Discrete-continuous depth estimation from a single image. In IEEE Conference on Computer Vision and Pattern Recognition. 716\u2013723."},{"key":"e_1_3_3_96_2","doi-asserted-by":"crossref","first-page":"184437","DOI":"10.1109\/ACCESS.2020.3030097","article-title":"Joint attention mechanisms for monocular depth estimation with multi-scale convolutions and adaptive weight adjustment","volume":"8","author":"Liu Peng","year":"2020","unstructured":"Peng Liu, Zonghua Zhang, Zhaozong Meng, and Nan Gao. 2020. Joint attention mechanisms for monocular depth estimation with multi-scale convolutions and adaptive weight adjustment. IEEE Access 8 (2020), 184437\u2013184450.","journal-title":"IEEE Access"},{"key":"e_1_3_3_97_2","article-title":"Recent advances of monocular 2D and 3D human pose estimation: A deep learning perspective","author":"Liu Wu","year":"2022","unstructured":"Wu Liu and Tao Mei. 2022. Recent advances of monocular 2D and 3D human pose estimation: A deep learning perspective. ACM Computing Surveys (CSUR) (2022).","journal-title":"ACM Computing Surveys (CSUR)"},{"key":"e_1_3_3_98_2","article-title":"Online mutual adaptation of deep depth prediction and visual SLAM.","author":"Loo Shing Yan","year":"2021","unstructured":"Shing Yan Loo, Moein Shakeri, Sai Hong Tang, Syamsiah Mashohor, and Hong Zhang. 2021. Online mutual adaptation of deep depth prediction and visual SLAM.CoRR (2021).","journal-title":"CoRR"},{"key":"e_1_3_3_99_2","first-page":"2329","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Lopes Ivan","year":"2023","unstructured":"Ivan Lopes, Tuan-Hung Vu, and Raoul de Charette. 2023. Cross-task attention mechanism for dense multi-task learning. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 2329\u20132338."},{"issue":"4","key":"e_1_3_3_100_2","article-title":"Consistent video depth estimation","volume":"39","author":"Luo Xuan","year":"2020","unstructured":"Xuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen, and Johannes Kopf. 2020. Consistent video depth estimation. ACM Transactions on Graphics (ToG) 39, 4 (2020), 71\u20131.","journal-title":"ACM Transactions on Graphics (ToG)"},{"key":"e_1_3_3_101_2","volume-title":"Learning and Understanding Single Image Depth Estimation in the Wild","year":"2020","unstructured":"Matteo Poggi. 2020. Learning and Understanding Single Image Depth Estimation in the Wild. https:\/\/drive.google.com\/file\/d\/17Bzlj_KZTXD_WheehKNup7f9BY5_BdHK\/view"},{"issue":"1","key":"e_1_3_3_102_2","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1177\/0278364916679498","article-title":"1 year, 1000 km: The Oxford RobotCar dataset","volume":"36","author":"Maddern Will","year":"2017","unstructured":"Will Maddern, Geoffrey Pascoe, Chris Linegar, and Paul Newman. 2017. 1 year, 1000 km: The Oxford RobotCar dataset. The International Journal of Robotics Research 36, 1 (2017), 3\u201315.","journal-title":"The International Journal of Robotics Research"},{"issue":"3","key":"e_1_3_3_103_2","doi-asserted-by":"crossref","first-page":"1778","DOI":"10.1109\/LRA.2017.2657002","article-title":"Toward domain independence for learning-based monocular depth estimation","volume":"2","author":"Mancini Michele","year":"2017","unstructured":"Michele Mancini, Gabriele Costante, Paolo Valigi, Thomas A. Ciarfuglia, Jeffrey Delmerico, and Davide Scaramuzza. 2017. Toward domain independence for learning-based monocular depth estimation. IEEE Robotics and Automation Letters 2, 3 (2017), 1778\u20131785.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"e_1_3_3_104_2","doi-asserted-by":"crossref","first-page":"e317","DOI":"10.7717\/peerj-cs.317","article-title":"Online supervised attention-based recurrent depth estimation from monocular video","volume":"6","author":"Maslov Dmitrii","year":"2020","unstructured":"Dmitrii Maslov and Ilya Makarov. 2020. Online supervised attention-based recurrent depth estimation from monocular video. PeerJ Computer Science 6 (2020), e317.","journal-title":"PeerJ Computer Science"},{"key":"e_1_3_3_105_2","first-page":"4040","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Mayer Nikolaus","year":"2016","unstructured":"Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. 2016. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In IEEE Conference on Computer Vision and Pattern Recognition. 4040\u20134048."},{"key":"e_1_3_3_106_2","first-page":"3061","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Menze Moritz","year":"2015","unstructured":"Moritz Menze and Andreas Geiger. 2015. Object scene flow for autonomous vehicles. In IEEE Conference on Computer Vision and Pattern Recognition. 3061\u20133070."},{"key":"e_1_3_3_107_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00956"},{"key":"e_1_3_3_108_2","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1016\/j.neucom.2020.12.089","article-title":"Deep learning for monocular depth estimation: A review","volume":"438","author":"Ming Yue","year":"2021","unstructured":"Yue Ming, Xuyang Meng, Chunxiao Fan, and Hui Yu. 2021. Deep learning for monocular depth estimation: A review. Neurocomputing 438 (2021), 14\u201333.","journal-title":"Neurocomputing"},{"key":"e_1_3_3_109_2","doi-asserted-by":"crossref","first-page":"355","DOI":"10.1109\/ICIP.2013.6738073","volume-title":"2013 IEEE International Conference on Image Processing","author":"Mutimbu Lawrence","year":"2013","unstructured":"Lawrence Mutimbu and Antonio Robles-Kelly. 2013. A relaxed factorial Markov random field for colour and depth estimation from a single foggy image. In 2013 IEEE International Conference on Image Processing. IEEE, 355\u2013359."},{"key":"e_1_3_3_110_2","first-page":"944","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Naderi Taher","year":"2022","unstructured":"Taher Naderi, Amir Sadovnik, Jason Hayward, and Hairong Qi. 2022. Monocular depth estimation with adaptive geometric attention. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 944\u2013954."},{"key":"e_1_3_3_111_2","first-page":"23296","article-title":"Intriguing properties of vision transformers","volume":"34","author":"Naseer Muhammad Muzammal","year":"2021","unstructured":"Muhammad Muzammal Naseer, Kanchana Ranasinghe, Salman H. Khan, Munawar Hayat, Fahad Shahbaz Khan, and Ming-Hsuan Yang. 2021. Intriguing properties of vision transformers. Advances in Neural Information Processing Systems 34 (2021), 23296\u201323308.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_112_2","doi-asserted-by":"crossref","first-page":"48","DOI":"10.1016\/j.neucom.2021.03.091","article-title":"A review on the attention mechanism of deep learning","volume":"452","author":"Niu Zhaoyang","year":"2021","unstructured":"Zhaoyang Niu, Guoqiang Zhong, and Hui Yu. 2021. A review on the attention mechanism of deep learning. Neurocomputing 452 (2021), 48\u201362.","journal-title":"Neurocomputing"},{"issue":"4","key":"e_1_3_3_113_2","doi-asserted-by":"crossref","first-page":"6813","DOI":"10.1109\/LRA.2020.3017478","article-title":"Don\u2019t forget the past: Recurrent depth estimation from monocular video","volume":"5","author":"Patil Vaishakh","year":"2020","unstructured":"Vaishakh Patil, Wouter Van Gansbeke, Dengxin Dai, and Luc Van Gool. 2020. Don\u2019t forget the past: Recurrent depth estimation from monocular video. IEEE Robotics and Automation Letters 5, 4 (2020), 6813\u20136820.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"e_1_3_3_114_2","first-page":"15560","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Peng Rui","year":"2021","unstructured":"Rui Peng, Ronggang Wang, Yawen Lai, Luyang Tang, and Yangang Cai. 2021. Excavating the potential capacity of self-supervised monocular depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition. 15560\u201315569."},{"key":"e_1_3_3_115_2","first-page":"1578","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Petrovai Andra","year":"2022","unstructured":"Andra Petrovai and Sergiu Nedevschi. 2022. Exploiting pseudo labels in a self-supervised learning framework for improved monocular depth estimation. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 1578\u20131588."},{"key":"e_1_3_3_116_2","first-page":"21477","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Piccinelli Luigi","year":"2023","unstructured":"Luigi Piccinelli, Christos Sakaridis, and Fisher Yu. 2023. iDisc: Internal discretization for monocular depth estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 21477\u201321487."},{"key":"e_1_3_3_117_2","doi-asserted-by":"crossref","first-page":"9250","DOI":"10.1109\/ICRA.2019.8793621","volume-title":"2019 International Conference on Robotics and Automation (ICRA\u201919)","author":"Pillai Sudeep","year":"2019","unstructured":"Sudeep Pillai, Rare\u015f Ambru\u015f, and Adrien Gaidon. 2019. SuperDepth: Self-supervised, super-resolved monocular depth estimation. In 2019 International Conference on Robotics and Automation (ICRA\u201919). 9250\u20139256. DOI:10.1109\/ICRA.2019.8793621"},{"key":"e_1_3_3_118_2","doi-asserted-by":"crossref","first-page":"587","DOI":"10.1109\/3DV.2018.00073","volume-title":"2018 International Conference on 3D Vision (3DV\u201918)","author":"Pilzer Andrea","year":"2018","unstructured":"Andrea Pilzer, Dan Xu, Mihai Puscas, Elisa Ricci, and Nicu Sebe. 2018. Unsupervised adversarial depth estimation using cycled generative networks. In 2018 International Conference on 3D Vision (3DV\u201918). IEEE, 587\u2013595."},{"key":"e_1_3_3_119_2","first-page":"11536","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Pintore Giovanni","year":"2021","unstructured":"Giovanni Pintore, Marco Agus, Eva Almansa, Jens Schneider, and Enrico Gobbetti. 2021. SliceNet: Deep dense depth estimation from a single indoor panorama using a slice-based representation. In IEEE Conference on Computer Vision and Pattern Recognition. 11536\u201311545."},{"key":"e_1_3_3_120_2","article-title":"On the synergies between machine learning and binocular stereo for depth estimation from images: A survey","author":"Poggi Matteo","year":"2021","unstructured":"Matteo Poggi, Fabio Tosi, Konstantinos Batsos, Philippos Mordohai, and Stefano Mattoccia. 2021. On the synergies between machine learning and binocular stereo for depth estimation from images: A survey. IEEE Trans. on Pattern Anal. and Machine Intelligence (2021).","journal-title":"IEEE Trans. on Pattern Anal. and Machine Intelligence"},{"key":"e_1_3_3_121_2","doi-asserted-by":"crossref","first-page":"324","DOI":"10.1109\/3DV.2018.00045","volume-title":"2018 International Conference on 3D Vision (3DV\u201918)","author":"Poggi Matteo","year":"2018","unstructured":"Matteo Poggi, Fabio Tosi, and Stefano Mattoccia. 2018. Learning monocular depth estimation with unsupervised trinocular assumptions. In 2018 International Conference on 3D Vision (3DV\u201918). IEEE, 324\u2013333."},{"key":"e_1_3_3_122_2","doi-asserted-by":"crossref","first-page":"150","DOI":"10.1109\/ICECS.2010.5724476","volume-title":"2010 17th IEEE International Conference on Electronics, Circuits and Systems","author":"Pourazad Mahsa T.","year":"2010","unstructured":"Mahsa T. Pourazad, Panos Nasiopoulos, and Ali Bashashati. 2010. Random forests-based 2D-to-3D video conversion. In 2010 17th IEEE International Conference on Electronics, Circuits and Systems. IEEE, 150\u2013153."},{"key":"e_1_3_3_123_2","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1109\/3DV.2019.00012","volume-title":"2019 International Conference on 3D Vision (3DV\u201919)","author":"Puscas Mihai Marian","year":"2019","unstructured":"Mihai Marian Puscas, Dan Xu, Andrea Pilzer, and Niculae Sebe. 2019. Structured coupled generative adversarial networks for unsupervised monocular depth estimation. In 2019 International Conference on 3D Vision (3DV\u201919). IEEE, 18\u201326."},{"key":"e_1_3_3_124_2","first-page":"3793","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Qin Zequn","year":"2022","unstructured":"Zequn Qin and Xi Li. 2022. MonoGround: Detecting monocular 3D objects from the ground. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3793\u20133802."},{"key":"e_1_3_3_125_2","first-page":"0","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops","author":"Ramamonjisoa Michael","year":"2019","unstructured":"Michael Ramamonjisoa and Vincent Lepetit. 2019. SharpNet: Fast and accurate recovery of occluding contours in monocular depth estimation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops. 0\u20130."},{"key":"e_1_3_3_126_2","first-page":"12179","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Ranftl Ren\u00e9","year":"2021","unstructured":"Ren\u00e9 Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. 2021. Vision transformers for dense prediction. In IEEE Conference on Computer Vision and Pattern Recognition. 12179\u201312188."},{"issue":"3","key":"e_1_3_3_127_2","doi-asserted-by":"crossref","first-page":"1623","DOI":"10.1109\/TPAMI.2020.3019967","article-title":"Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer","volume":"44","author":"Ranftl Ren\u00e9","year":"2020","unstructured":"Ren\u00e9 Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. 2020. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 3 (2020), 1623\u20131637.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_128_2","first-page":"4058","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Ranftl Rene","year":"2016","unstructured":"Rene Ranftl, Vibhav Vineet, Qifeng Chen, and Vladlen Koltun. 2016. Dense monocular depth estimation in complex dynamic scenes. In IEEE Conference on Computer Vision and Pattern Recognition. 4058\u20134066."},{"key":"e_1_3_3_129_2","first-page":"750","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition Workshops","author":"Ren Haoyu","year":"2020","unstructured":"Haoyu Ren, Aman Raj, Mostafa El-Khamy, and Jungwon Lee. 2020. SUW-Learn: Joint supervised, unsupervised, weakly supervised deep learning for monocular depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition Workshops. 750\u2013751."},{"key":"e_1_3_3_130_2","first-page":"209","volume-title":"International Conference on Pattern Recognition and Machine Intelligence","author":"Repala Vamshi Krishna","year":"2019","unstructured":"Vamshi Krishna Repala and Shiv Ram Dubey. 2019. Dual CNN models for unsupervised monocular depth estimation. In International Conference on Pattern Recognition and Machine Intelligence. Springer, 209\u2013217."},{"key":"e_1_3_3_131_2","first-page":"3762","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Rey-Area Manuel","year":"2022","unstructured":"Manuel Rey-Area, Mingze Yuan, and Christian Richardt. 2022. 360MonoDepth: High-resolution 360\u00b0 monocular depth estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3762\u20133772."},{"key":"e_1_3_3_132_2","first-page":"1735","volume-title":"2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS\u201919)","author":"Roussel Tom","year":"2019","unstructured":"Tom Roussel, Luc Van Eycken, and Tinne Tuytelaars. 2019. Monocular depth estimation in new environments with absolute scale. In 2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS\u201919). IEEE, 1735\u20131741."},{"key":"e_1_3_3_133_2","first-page":"5506","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Roy Anirban","year":"2016","unstructured":"Anirban Roy and Sinisa Todorovic. 2016. Monocular depth estimation using neural regression forest. In IEEE Conference on Computer Vision and Pattern Recognition. 5506\u20135514."},{"key":"e_1_3_3_134_2","first-page":"656","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Sagar Abhinav","year":"2022","unstructured":"Abhinav Sagar. 2022. Monocular depth estimation using multi scale neural network and feature fusion. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 656\u2013662."},{"key":"e_1_3_3_135_2","volume-title":"Advances in Neural Information Processing Systems","author":"Saxena Ashutosh","year":"2005","unstructured":"Ashutosh Saxena, Sung Chung, and Andrew Ng. 2005. Learning depth from single monocular images. In Advances in Neural Information Processing Systems, Y. Weiss, B. Sch\u00f6lkopf, and J. Platt (Eds.). Vol. 18. MIT Press."},{"issue":"5","key":"e_1_3_3_136_2","doi-asserted-by":"crossref","first-page":"824","DOI":"10.1109\/TPAMI.2008.132","article-title":"Make3D: Learning 3D scene structure from a single still image","volume":"31","author":"Saxena Ashutosh","year":"2008","unstructured":"Ashutosh Saxena, Min Sun, and Andrew Y. Ng. 2008. Make3D: Learning 3D scene structure from a single still image. IEEE Transactions on Pattern Analysis and Machine Intelligence 31, 5 (2008), 824\u2013840.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_137_2","article-title":"The surprising effectiveness of diffusion models for optical flow and monocular depth estimation","volume":"36","author":"Saxena Saurabh","year":"2024","unstructured":"Saurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar, Mohammad Norouzi, Deqing Sun, and David J. Fleet. 2024. The surprising effectiveness of diffusion models for optical flow and monocular depth estimation. Advances in Neural Information Processing Systems 36 (2024).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_138_2","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1007\/978-3-319-11752-2_3","volume-title":"German Conference on Pattern Recognition","author":"Scharstein Daniel","year":"2014","unstructured":"Daniel Scharstein, Heiko Hirschm\u00fcller, York Kitajima, Greg Krathwohl, Nera Ne\u0161i\u0107, Xi Wang, and Porter Westling. 2014. High-resolution stereo datasets with subpixel-accurate ground truth. In German Conference on Pattern Recognition. Springer, 31\u201342."},{"key":"e_1_3_3_139_2","first-page":"3260","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Schops Thomas","year":"2017","unstructured":"Thomas Schops, Johannes L. Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger. 2017. A multi-view stereo benchmark with high-resolution images and multi-camera videos. In IEEE Conference on Computer Vision and Pattern Recognition. 3260\u20133269."},{"key":"e_1_3_3_140_2","first-page":"1680","volume-title":"2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS\u201918)","author":"Schubert David","year":"2018","unstructured":"David Schubert, Thore Goll, Nikolaus Demmel, Vladyslav Usenko, J\u00f6rg St\u00fcckler, and Daniel Cremers. 2018. The TUM VI benchmark for evaluating visual-inertial odometry. In 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS\u201918). IEEE, 1680\u20131687."},{"key":"e_1_3_3_141_2","article-title":"IEBins: Iterative elastic bins for monocular depth estimation","volume":"36","author":"Shao Shuwei","year":"2024","unstructured":"Shuwei Shao, Zhongcai Pei, Xingming Wu, Zhong Liu, Weihai Chen, and Zhengguo Li. 2024. IEBins: Iterative elastic bins for monocular depth estimation. Advances in Neural Information Processing Systems 36 (2024).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_142_2","first-page":"746","volume-title":"European Conference on Computer Vision","author":"Silberman Nathan","year":"2012","unstructured":"Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. 2012. Indoor segmentation and support inference from RGBD images. In European Conference on Computer Vision. Springer, 746\u2013760."},{"key":"e_1_3_3_143_2","first-page":"4557","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Son Eunjin","year":"2024","unstructured":"Eunjin Son and Sang Jun Lee. 2024. CaBins: CLIP-based adaptive bins for monocular depth estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 4557\u20134567."},{"key":"e_1_3_3_144_2","first-page":"1746","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Song Shuran","year":"2017","unstructured":"Shuran Song, Fisher Yu, Andy Zeng, Angel X. Chang, Manolis Savva, and Thomas Funkhouser. 2017. Semantic scene completion from a single depth image. In IEEE Conference on Computer Vision and Pattern Recognition. 1746\u20131754."},{"issue":"5","key":"e_1_3_3_145_2","doi-asserted-by":"crossref","first-page":"1220","DOI":"10.1109\/TMM.2019.2941776","article-title":"Contextualized CNN for scene-aware depth estimation from single RGB image","volume":"22","author":"Song Wenfeng","year":"2019","unstructured":"Wenfeng Song, Shuai Li, Ji Liu, Aimin Hao, Qinping Zhao, and Hong Qin. 2019. Contextualized CNN for scene-aware depth estimation from single RGB image. IEEE Transactions on Multimedia 22, 5 (2019), 1220\u20131233.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_3_146_2","doi-asserted-by":"crossref","first-page":"4691","DOI":"10.1109\/TIP.2021.3074306","article-title":"MLDA-Net: Multi-level dual attention-based network for self-supervised monocular depth estimation","volume":"30","author":"Song Xibin","year":"2021","unstructured":"Xibin Song, Wei Li, Dingfu Zhou, Yuchao Dai, Jin Fang, Hongdong Li, and Liangjun Zhang. 2021. MLDA-Net: Multi-level dual attention-based network for self-supervised monocular depth estimation. IEEE Transactions on Image Processing 30 (2021), 4691\u20134705.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_3_147_2","first-page":"573","volume-title":"IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"Sturm J\u00fcrgen","year":"2012","unstructured":"J\u00fcrgen Sturm, Nikolas Engelhard, Felix Endres, Wolfram Burgard, and Daniel Cremers. 2012. A benchmark for the evaluation of RGB-D SLAM systems. In IEEE\/RSJ International Conference on Intelligent Robots and Systems. IEEE, 573\u2013580."},{"key":"e_1_3_3_148_2","doi-asserted-by":"crossref","first-page":"114930","DOI":"10.1109\/ACCESS.2020.3003466","article-title":"Soft regression of monocular depth using scale-semantic exchange network","volume":"8","author":"Su Wen","year":"2020","unstructured":"Wen Su and Haifeng Zhang. 2020. Soft regression of monocular depth using scale-semantic exchange network. IEEE Access 8 (2020), 114930\u2013114939.","journal-title":"IEEE Access"},{"issue":"6","key":"e_1_3_3_149_2","first-page":"3491","article-title":"Monocular depth estimation using information exchange network","volume":"22","author":"Su Wen","year":"2020","unstructured":"Wen Su, Haifeng Zhang, Quan Zhou, Wenzhen Yang, and Zengfu Wang. 2020. Monocular depth estimation using information exchange network. IEEE Transactions on Intelligent Transportation Systems 22, 6 (2020), 3491\u20133503.","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"e_1_3_3_150_2","article-title":"SC-DepthV3: Robust self-supervised monocular depth estimation for dynamic scenes","author":"Sun Libo","year":"2023","unstructured":"Libo Sun, Jia-Wang Bian, Huangying Zhan, Wei Yin, Ian Reid, and Chunhua Shen. 2023. SC-DepthV3: Robust self-supervised monocular depth estimation for dynamic scenes. IEEE Transactions on Pattern Analysis and Machine Intelligence (2023).","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_151_2","volume-title":"E3S Web of Conferences","volume":"309","author":"Swaraja K.","year":"2021","unstructured":"K. Swaraja, V. Akshitha, K. Pranav, B. Vyshnavi, V. Sai Akhil, K. Meenakshi, Padmavathi Kora, Himabindu Valiveti, and Chaitanya Duggineni. 2021. Monocular depth estimation using transfer learning-an overview. In E3S Web of Conferences, Vol. 309. EDP Sciences."},{"key":"e_1_3_3_152_2","first-page":"233","volume-title":"Applications of Digital Image Processing XXXII","author":"Tian Dong","year":"2009","unstructured":"Dong Tian, Po-Lin Lai, Patrick Lopez, and Cristina Gomila. 2009. View synthesis techniques for 3D video. In Applications of Digital Image Processing XXXII, Vol. 7443. SPIE, 233\u2013243."},{"key":"e_1_3_3_153_2","doi-asserted-by":"crossref","first-page":"8573","DOI":"10.1109\/ICASSP.2019.8683235","volume-title":"ICASSP 2019\u20132019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201919)","author":"Tian Hu","year":"2019","unstructured":"Hu Tian and Fei Li. 2019. Semi-supervised depth estimation from a single image based on confidence learning. In ICASSP 2019\u20132019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201919). IEEE, 8573\u20138577."},{"key":"e_1_3_3_154_2","first-page":"9799","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Tosi Fabio","year":"2019","unstructured":"Fabio Tosi, Filippo Aleotti, Matteo Poggi, and Stefano Mattoccia. 2019. Learning monocular depth estimation infusing traditional stereo knowledge. In IEEE Conference on Computer Vision and Pattern Recognition. 9799\u20139809."},{"key":"e_1_3_3_155_2","first-page":"5038","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Ummenhofer Benjamin","year":"2017","unstructured":"Benjamin Ummenhofer, Huizhong Zhou, Jonas Uhrig, Nikolaus Mayer, Eddy Ilg, Alexey Dosovitskiy, and Thomas Brox. 2017. DeMoN: Depth and motion network for learning monocular stereo. In IEEE Conference on Computer Vision and Pattern Recognition. 5038\u20135047."},{"key":"e_1_3_3_156_2","first-page":"3937","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Verdi\u00e9 Yannick","year":"2022","unstructured":"Yannick Verdi\u00e9, Jifei Song, Barnab\u00e9 Mas, Benjamin Busam, Ales Leonardis, and Steven McDonagh. 2022. CroMo: Cross-modal learning for monocular depth estimation. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3937\u20133947."},{"key":"e_1_3_3_157_2","article-title":"Outdoor monocular depth estimation: A research review","author":"Vyas Pulkit","year":"2022","unstructured":"Pulkit Vyas, Chirag Saxena, Anwesh Badapanda, and Anurag Goswami. 2022. Outdoor monocular depth estimation: A research review. arXiv preprint arXiv:2205.01399 (2022).","journal-title":"arXiv preprint arXiv:2205.01399"},{"key":"e_1_3_3_158_2","first-page":"2620","volume-title":"2021 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS\u201921)","author":"Wagstaff Brandon","year":"2021","unstructured":"Brandon Wagstaff and Jonathan Kelly. 2021. Self-supervised scale recovery for monocular depth and egomotion estimation. In 2021 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS\u201921). IEEE, 2620\u20132627."},{"key":"e_1_3_3_159_2","first-page":"7561","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wallingford Matthew","year":"2022","unstructured":"Matthew Wallingford, Hao Li, Alessandro Achille, Avinash Ravichandran, Charless Fowlkes, Rahul Bhotika, and Stefano Soatto. 2022. Task adaptive parameter sharing for multi-task learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 7561\u20137570."},{"issue":"1","key":"e_1_3_3_160_2","doi-asserted-by":"crossref","first-page":"308","DOI":"10.1109\/TITS.2020.3010418","article-title":"Unsupervised learning of depth, optical flow and pose with occlusion from 3D geometry","volume":"23","author":"Wang Guangming","year":"2020","unstructured":"Guangming Wang, Chi Zhang, Hesheng Wang, Jingchuan Wang, Yong Wang, and Xinlei Wang. 2020. Unsupervised learning of depth, optical flow and pose with occlusion from 3D geometry. IEEE Transactions on Intelligent Transportation Systems 23, 1 (2020), 308\u2013320.","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"e_1_3_3_161_2","first-page":"541","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang Lijun","year":"2020","unstructured":"Lijun Wang, Jianming Zhang, Oliver Wang, Zhe Lin, and Huchuan Lu. 2020. SDC-Depth: Semantic divide-and-conquer network for monocular depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition. 541\u2013550."},{"key":"e_1_3_3_162_2","first-page":"8515","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Wang Qin","year":"2021","unstructured":"Qin Wang, Dengxin Dai, Lukas Hoyer, Luc Van Gool, and Olga Fink. 2021. Domain adaptive semantic segmentation with self-supervised depth estimation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 8515\u20138525."},{"key":"e_1_3_3_163_2","first-page":"5555","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang Rui","year":"2019","unstructured":"Rui Wang, Stephen M. Pizer, and Jan-Michael Frahm. 2019. Recurrent neural network for (un-) supervised learning of monocular video visual odometry and depth. In IEEE Conference on Computer Vision and Pattern Recognition. 5555\u20135564."},{"key":"e_1_3_3_164_2","first-page":"21425","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Ruoyu","year":"2023","unstructured":"Ruoyu Wang, Zehao Yu, and Shenghua Gao. 2023. PlaneDepth: Self-supervised depth estimation via orthogonal planes. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 21425\u201321434."},{"key":"e_1_3_3_165_2","first-page":"2162","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Watson Jamie","year":"2019","unstructured":"Jamie Watson, Michael Firman, Gabriel J. Brostow, and Daniyar Turmukhambetov. 2019. Self-supervised monocular depth hints. In IEEE Conference on Computer Vision and Pattern Recognition. 2162\u20132171."},{"key":"e_1_3_3_166_2","first-page":"1164","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Watson Jamie","year":"2021","unstructured":"Jamie Watson, Oisin Mac Aodha, Victor Prisacariu, Gabriel Brostow, and Michael Firman. 2021. The temporal opportunist: Self-supervised multi-frame monocular depth. In IEEE Conference on Computer Vision and Pattern Recognition. 1164\u20131174."},{"key":"e_1_3_3_167_2","first-page":"8987","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Won Changhee","year":"2019","unstructured":"Changhee Won, Jongbin Ryu, and Jongwoo Lim. 2019. OmniMVS: End-to-end learning for omnidirectional stereo matching. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 8987\u20138996."},{"key":"e_1_3_3_168_2","first-page":"3","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201918)","author":"Woo Sanghyun","year":"2018","unstructured":"Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. 2018. CBAM: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV\u201918). 3\u201319."},{"key":"e_1_3_3_169_2","first-page":"65","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Xia Zhihao","year":"2020","unstructured":"Zhihao Xia, Patrick Sullivan, and Ayan Chakrabarti. 2020. Generating and exploiting probabilistic monocular depth estimates. In IEEE Conference on Computer Vision and Pattern Recognition. 65\u201374."},{"key":"e_1_3_3_170_2","first-page":"1625","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Xiao Jianxiong","year":"2013","unstructured":"Jianxiong Xiao, Andrew Owens, and Antonio Torralba. 2013. SUN3D: A database of big spaces reconstructed using SfM and object labels. In Proceedings of the IEEE International Conference on Computer Vision. 1625\u20131632."},{"key":"e_1_3_3_171_2","doi-asserted-by":"crossref","first-page":"2436","DOI":"10.1109\/CAC51589.2020.9327548","volume-title":"2020 Chinese Automation Congress (CAC\u201920)","author":"Xiaogang Ruan","year":"2020","unstructured":"Ruan Xiaogang, Yan Wenjing, Huang Jing, Guo Peiyuan, and Guo Wei. 2020. Monocular depth estimation based on deep learning: A survey. In 2020 Chinese Automation Congress (CAC\u201920). IEEE, 2436\u20132440."},{"key":"e_1_3_3_172_2","first-page":"842","volume-title":"European Conference on Computer Vision","author":"Xie Junyuan","year":"2016","unstructured":"Junyuan Xie, Ross Girshick, and Ali Farhadi. 2016. Deep3D: Fully automatic 2D-to-3D video conversion with deep convolutional neural networks. In European Conference on Computer Vision. Springer, 842\u2013857."},{"key":"e_1_3_3_173_2","first-page":"5354","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Xu Dan","year":"2017","unstructured":"Dan Xu, Elisa Ricci, Wanli Ouyang, Xiaogang Wang, and Nicu Sebe. 2017. Multi-scale continuous CRFs as sequential deep networks for monocular depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition. 5354\u20135362."},{"key":"e_1_3_3_174_2","article-title":"Monocular depth estimation using multi-scale continuous CRFs as sequential deep networks","volume":"41","author":"Xu Dan","year":"2019","unstructured":"Dan Xu, Elisa Ricci, Wanli Ouyang, Xiaogang Wang, and Nicu Sebe. 2019. Monocular depth estimation using multi-scale continuous CRFs as sequential deep networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 41 (2019). Issue 6.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_3_175_2","first-page":"3917","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Xu Dan","year":"2018","unstructured":"Dan Xu, Wei Wang, Hao Tang, Hong Liu, Nicu Sebe, and Elisa Ricci. 2018. Structured attention guided convolutional neural fields for monocular depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition. 3917\u20133925."},{"key":"e_1_3_3_176_2","article-title":"Weakly-supervised monocular depth estimation with resolution-mismatched data","author":"Xu Jialei","year":"2021","unstructured":"Jialei Xu, Yuanchao Bai, Xianming Liu, Junjun Jiang, and Xiangyang Ji. 2021. Weakly-supervised monocular depth estimation with resolution-mismatched data. arXiv preprint arXiv:2109.11573 (2021).","journal-title":"arXiv preprint arXiv:2109.11573"},{"key":"e_1_3_3_177_2","doi-asserted-by":"crossref","first-page":"678","DOI":"10.1109\/LSP.2021.3067498","article-title":"Monocular depth estimation with multi-scale feature fusion","volume":"28","author":"Xu Xianfa","year":"2021","unstructured":"Xianfa Xu, Zhe Chen, and Fuliang Yin. 2021. Monocular depth estimation with multi-scale feature fusion. IEEE Signal Processing Letters 28 (2021), 678\u2013682.","journal-title":"IEEE Signal Processing Letters"},{"key":"e_1_3_3_178_2","doi-asserted-by":"crossref","first-page":"8811","DOI":"10.1109\/TIP.2021.3120670","article-title":"Multi-scale spatial attention-guided monocular depth estimation with semantic enhancement","volume":"30","author":"Xu Xianfa","year":"2021","unstructured":"Xianfa Xu, Zhe Chen, and Fuliang Yin. 2021. Multi-scale spatial attention-guided monocular depth estimation with semantic enhancement. IEEE Transactions on Image Processing 30 (2021), 8811\u20138822.","journal-title":"IEEE Transactions on Image Processing"},{"issue":"10","key":"e_1_3_3_179_2","doi-asserted-by":"crossref","first-page":"17039","DOI":"10.1109\/TITS.2021.3093592","article-title":"Unsupervised learning of depth estimation and camera pose with multi-scale GANs","volume":"23","author":"Xu Yufan","year":"2022","unstructured":"Yufan Xu, Yan Wang, Rui Huang, Zeyu Lei, Junyao Yang, and Zijian Li. 2022. Unsupervised learning of depth estimation and camera pose with multi-scale GANs. IEEE Transactions on Intelligent Transportation Systems 23, 10 (2022), 17039\u201317047.","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"e_1_3_3_180_2","first-page":"464","volume-title":"2021 International Conference on 3D Vision (3DV\u201921)","author":"Yan Jiaxing","year":"2021","unstructured":"Jiaxing Yan, Hong Zhao, Penghui Bu, and YuSheng Jin. 2021. Channel-wise attention-based network for self-supervised monocular depth estimation. In 2021 International Conference on 3D Vision (3DV\u201921). IEEE, 464\u2013473."},{"key":"e_1_3_3_181_2","first-page":"5515","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yang Gengshan","year":"2019","unstructured":"Gengshan Yang, Joshua Manela, Michael Happold, and Deva Ramanan. 2019. Hierarchical deep stereo matching on high-resolution images. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 5515\u20135524."},{"key":"e_1_3_3_182_2","first-page":"899","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yang Guorun","year":"2019","unstructured":"Guorun Yang, Xiao Song, Chaoqin Huang, Zhidong Deng, Jianping Shi, and Bolei Zhou. 2019. DrivingStereo: A large-scale dataset for stereo matching in autonomous driving scenarios. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 899\u2013908."},{"key":"e_1_3_3_183_2","first-page":"16269","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Yang Guanglei","year":"2021","unstructured":"Guanglei Yang, Hao Tang, Mingli Ding, Nicu Sebe, and Elisa Ricci. 2021. Transformer-based attention networks for continuous pixel-wise prediction. In IEEE Conference on Computer Vision and Pattern Recognition. 16269\u201316279."},{"key":"e_1_3_3_184_2","article-title":"Depth anything: Unleashing the power of large-scale unlabeled data","author":"Yang Lihe","year":"2024","unstructured":"Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. 2024. Depth anything: Unleashing the power of large-scale unlabeled data. arXiv preprint arXiv:2401.10891 (2024).","journal-title":"arXiv preprint arXiv:2401.10891"},{"key":"e_1_3_3_185_2","first-page":"817","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Yang Nan","year":"2018","unstructured":"Nan Yang, Rui Wang, Jorg Stuckler, and Daniel Cremers. 2018. Deep virtual stereo odometry: Leveraging deep depth prediction for monocular direct sparse odometry. In Proceedings of the European Conference on Computer Vision. 817\u2013833."},{"key":"e_1_3_3_186_2","doi-asserted-by":"crossref","first-page":"67696","DOI":"10.1109\/ACCESS.2021.3076346","article-title":"Monocular depth estimation based on multi-scale depth map fusion","volume":"9","author":"Yang Xin","year":"2021","unstructured":"Xin Yang, Qingling Chang, Xinglin Liu, Siyuan He, and Yan Cui. 2021. Monocular depth estimation based on multi-scale depth map fusion. IEEE Access 9 (2021), 67696\u201367705.","journal-title":"IEEE Access"},{"issue":"11","key":"e_1_3_3_187_2","doi-asserted-by":"crossref","first-page":"2701","DOI":"10.1109\/TMM.2019.2912121","article-title":"Bayesian DeNet: Monocular depth prediction and frame-wise fusion with synchronized uncertainty","volume":"21","author":"Yang Xin","year":"2019","unstructured":"Xin Yang, Yang Gao, Hongcheng Luo, Chunyuan Liao, and Kwang-Ting Cheng. 2019. Bayesian DeNet: Monocular depth prediction and frame-wise fusion with synchronized uncertainty. IEEE Transactions on Multimedia 21, 11 (2019), 2701\u20132713.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_3_188_2","doi-asserted-by":"crossref","first-page":"142","DOI":"10.1016\/j.neucom.2018.10.019","article-title":"Reactive obstacle avoidance of monocular quadrotors with online adapted depth prediction network","volume":"325","author":"Yang Xin","year":"2019","unstructured":"Xin Yang, Hongcheng Luo, Yuhao Wu, Yang Gao, Chunyuan Liao, and Kwang-Ting Cheng. 2019. Reactive obstacle avoidance of monocular quadrotors with online adapted depth prediction network. Neurocomputing 325 (2019), 142\u2013158.","journal-title":"Neurocomputing"},{"key":"e_1_3_3_189_2","first-page":"12719","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Yang Xiaodong","year":"2023","unstructured":"Xiaodong Yang, Zhuang Ma, Zhiyu Ji, and Zhe Ren. 2023. GEDepth: Ground embedding for monocular depth estimation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 12719\u201312727."},{"key":"e_1_3_3_190_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3072215"},{"key":"e_1_3_3_191_2","first-page":"169","volume-title":"International Conference on Multimedia and Expo (ICME\u201919)","author":"Ye Xinchen","year":"2019","unstructured":"Xinchen Ye, Mingliang Zhang, Rui Xu, Wei Zhong, Xin Fan, Zhu Liu, and Jiaao Zhang. 2019. Unsupervised monocular depth estimation based on dual attention mechanism and depth-aware loss. In International Conference on Multimedia and Expo (ICME\u201919). IEEE, 169\u2013174."},{"key":"e_1_3_3_192_2","first-page":"469","volume-title":"IEEE Annual International Conference on CYBER Technology in Automation, Control, and Intelligent Systems","author":"Yingcai Wan","year":"2019","unstructured":"Wan Yingcai, Fang Lijing, and Zhao Qiankun. 2019. Multi-scale deep CNN network for unsupervised monocular depth estimation. In IEEE Annual International Conference on CYBER Technology in Automation, Control, and Intelligent Systems. IEEE, 469\u2013473."},{"key":"e_1_3_3_193_2","first-page":"2428","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yucel Mehmet Kerim","year":"2021","unstructured":"Mehmet Kerim Yucel, Valia Dimaridou, Anastasios Drosou, and Albert Saa-Garriga. 2021. Real-time monocular depth estimation with sparse supervision on mobile. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2428\u20132437."},{"key":"e_1_3_3_194_2","first-page":"298","volume-title":"Asian Conference on Computer Vision","author":"Ramirez Pierluigi Zama","year":"2018","unstructured":"Pierluigi Zama Ramirez, Matteo Poggi, Fabio Tosi, Stefano Mattoccia, and Luigi Di Stefano. 2018. Geometry meets semantics for semi-supervised monocular depth estimation. In Asian Conference on Computer Vision. Springer, 298\u2013313."},{"key":"e_1_3_3_195_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00043"},{"key":"e_1_3_3_196_2","first-page":"1725","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhang Haokui","year":"2019","unstructured":"Haokui Zhang, Chunhua Shen, Ying Li, Yuanzhouhan Cao, Yu Liu, and Youliang Yan. 2019. Exploiting temporal consistency for real-time video depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition. 1725\u20131734."},{"key":"e_1_3_3_197_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.neucom.2020.05.015","article-title":"Unsupervised detail-preserving network for high quality monocular depth estimation","volume":"404","author":"Zhang Mingliang","year":"2020","unstructured":"Mingliang Zhang, Xinchen Ye, and Xin Fan. 2020. Unsupervised detail-preserving network for high quality monocular depth estimation. Neurocomputing 404 (2020), 1\u201313.","journal-title":"Neurocomputing"},{"key":"e_1_3_3_198_2","first-page":"18537","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Ning","year":"2023","unstructured":"Ning Zhang, Francesco Nex, George Vosselman, and Norman Kerle. 2023. Lite-Mono: A lightweight CNN and transformer architecture for self-supervised monocular depth estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 18537\u201318546."},{"issue":"4","key":"e_1_3_3_199_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3450626.3459871","article-title":"Consistent depth of moving objects in video","volume":"40","author":"Zhang Zhoutong","year":"2021","unstructured":"Zhoutong Zhang, Forrester Cole, Richard Tucker, William T. Freeman, and Tali Dekel. 2021. Consistent depth of moving objects in video. ACM Transactions on Graphics (TOG) 40, 4 (2021), 1\u201312.","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_3_3_200_2","first-page":"235","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Zhang Zhenyu","year":"2018","unstructured":"Zhenyu Zhang, Zhen Cui, Chunyan Xu, Zequn Jie, Xiang Li, and Jian Yang. 2018. Joint task-recursive learning for semantic segmentation and depth estimation. In Proceedings of the European Conference on Computer Vision. 235\u2013251."},{"key":"e_1_3_3_201_2","first-page":"4494","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Zhenyu","year":"2020","unstructured":"Zhenyu Zhang, Stephane Lathuiliere, Elisa Ricci, Nicu Sebe, Yan Yan, and Jian Yang. 2020. Online depth learning against forgetting in monocular videos. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 4494\u20134503."},{"key":"e_1_3_3_202_2","first-page":"2614","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Zhang Ziyu","year":"2015","unstructured":"Ziyu Zhang, Alexander G. Schwing, Sanja Fidler, and Raquel Urtasun. 2015. Monocular object instance segmentation and depth ordering with CNNs. In Proceedings of the IEEE International Conference on Computer Vision. 2614\u20132622."},{"issue":"9","key":"e_1_3_3_203_2","doi-asserted-by":"crossref","first-page":"1612","DOI":"10.1007\/s11431-020-1582-8","article-title":"Monocular depth estimation based on deep learning: An overview","volume":"63","author":"Zhao Chaoqiang","year":"2020","unstructured":"Chaoqiang Zhao, Qiyu Sun, Chongzhen Zhang, Yang Tang, and Feng Qian. 2020. Monocular depth estimation based on deep learning: An overview. Science China Technological Sciences 63, 9 (2020), 1612\u20131627.","journal-title":"Science China Technological Sciences"},{"issue":"12","key":"e_1_3_3_204_2","doi-asserted-by":"crossref","first-page":"5392","DOI":"10.1109\/TNNLS.2020.3044181","article-title":"Masked GAN for unsupervised depth and pose prediction with scale consistency","volume":"32","author":"Zhao Chaoqiang","year":"2020","unstructured":"Chaoqiang Zhao, Gary G. Yen, Qiyu Sun, Chongzhen Zhang, and Yang Tang. 2020. Masked GAN for unsupervised depth and pose prediction with scale consistency. IEEE Transactions on Neural Networks and Learning Systems 32, 12 (2020), 5392\u20135403.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_3_205_2","first-page":"9788","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhao Shanshan","year":"2019","unstructured":"Shanshan Zhao, Huan Fu, Mingming Gong, and Dacheng Tao. 2019. Geometry-aware symmetric domain adaptation for monocular depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition. 9788\u20139798."},{"key":"e_1_3_3_206_2","doi-asserted-by":"crossref","first-page":"16323","DOI":"10.1109\/ACCESS.2019.2894651","article-title":"Super-resolution for monocular depth estimation with multi-scale sub-pixel convolutions and a smoothness constraint","volume":"7","author":"Zhao Shiyu","year":"2019","unstructured":"Shiyu Zhao, Lin Zhang, Ying Shen, Shengjie Zhao, and Huijuan Zhang. 2019. Super-resolution for monocular depth estimation with multi-scale sub-pixel convolutions and a smoothness constraint. IEEE Access 7 (2019), 16323\u201316335.","journal-title":"IEEE Access"},{"key":"e_1_3_3_207_2","first-page":"3330","volume-title":"Proceedings of the Conference on Computer Vision and Pattern Recognition","author":"Zhao Yunhan","year":"2020","unstructured":"Yunhan Zhao, Shu Kong, Daeyun Shin, and Charless Fowlkes. 2020. Domain decluttering: Simplifying images to mitigate synthetic-real domain shift and improve depth estimation. In Proceedings of the Conference on Computer Vision and Pattern Recognition. 3330\u20133340."},{"key":"e_1_3_3_208_2","first-page":"767","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201918)","author":"Zheng Chuanxia","year":"2018","unstructured":"Chuanxia Zheng, Tat-Jen Cham, and Jianfei Cai. 2018. T2Net: Synthetic-to-realistic translation for solving single-image depth estimation tasks. In Proceedings of the European Conference on Computer Vision (ECCV\u201918). 767\u2013783."},{"key":"e_1_3_3_209_2","first-page":"1851","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhou Tinghui","year":"2017","unstructured":"Tinghui Zhou, Matthew Brown, Noah Snavely, and David G. Lowe. 2017. Unsupervised learning of depth and ego-motion from video. In IEEE Conference on Computer Vision and Pattern Recognition. 1851\u20131858."},{"key":"e_1_3_3_210_2","doi-asserted-by":"crossref","first-page":"21359","DOI":"10.1109\/ACCESS.2022.3151108","article-title":"Learning depth estimation from memory infusing monocular cues: A generalization prediction approach","volume":"10","author":"Zhou Yakun","year":"2022","unstructured":"Yakun Zhou, Jinting Luo, Musen Hu, Tingyong Wu, Jinkuan Zhu, Xingzhong Xiong, and Jienan Chen. 2022. Learning depth estimation from memory infusing monocular cues: A generalization prediction approach. IEEE Access 10 (2022), 21359\u201321369.","journal-title":"IEEE Access"},{"key":"e_1_3_3_211_2","first-page":"12777","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhou Zhongkai","year":"2021","unstructured":"Zhongkai Zhou, Xinnan Fan, Pengfei Shi, and Yuanxue Xin. 2021. R-MSFM: Recurrent multi-scale feature modulation for monocular depth estimating. In IEEE Conference on Computer Vision and Pattern Recognition. 12777\u201312786."},{"key":"e_1_3_3_212_2","first-page":"614","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhuo Wei","year":"2015","unstructured":"Wei Zhuo, Mathieu Salzmann, Xuming He, and Miaomiao Liu. 2015. Indoor scene structure analysis for single image depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition. 614\u2013622."},{"key":"e_1_3_3_213_2","first-page":"388","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Zoran Daniel","year":"2015","unstructured":"Daniel Zoran, Phillip Isola, Dilip Krishnan, and William T. Freeman. 2015. Learning ordinal relationships for mid-level vision. In Proceedings of the IEEE International Conference on Computer Vision. 388\u2013396."},{"key":"e_1_3_3_214_2","first-page":"36","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201918)","author":"Zou Yuliang","year":"2018","unstructured":"Yuliang Zou, Zelun Luo, and Jia-Bin Huang. 2018. DF-Net: Unsupervised joint learning of depth and flow using cross-task consistency. In Proceedings of the European Conference on Computer Vision (ECCV\u201918). 36\u201353."}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3677327","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3677327","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:04:21Z","timestamp":1750291461000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3677327"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,3]]},"references-count":213,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2024,12,31]]}},"alternative-id":["10.1145\/3677327"],"URL":"https:\/\/doi.org\/10.1145\/3677327","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,3]]},"assertion":[{"value":"2023-03-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-06-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}