{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T14:53:13Z","timestamp":1784904793362,"version":"3.55.0"},"reference-count":105,"publisher":"Association for Computing Machinery (ACM)","issue":"12","funder":[{"name":"National Key Research and Development Program of China","award":["2024YFB3909902"],"award-info":[{"award-number":["2024YFB3909902"]}]},{"DOI":"10.13039\/501100001809","name":"National Nature Science Foundation of China","doi-asserted-by":"crossref","award":["62306294"],"award-info":[{"award-number":["62306294"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,12,31]]},"abstract":"<jats:p>\n                    Self-supervised monocular depth estimation holds significant importance in the fields of autonomous driving and robotics. However, existing methods are typically trained and evaluated on clear, sunny datasets, overlooking the impact of various adverse conditions commonly encountered in real-world applications, such as rainy weather, low visibility, and motion blur. As a result, they often struggle in challenging scenarios and produce artifacts. To address this issue, we propose ER-Depth, a novel two-stage self-supervised framework designed for robust depth estimation. In the first stage, we propose perturbation-invariant depth consistency regularization to propagate reliable supervision from standard to challenging scenes. In the second stage, we adopt the Mean Teacher paradigm for self-distillation and present a novel consistency-based pseudo-label filtering strategy to improve the quality of pseudo-labels. Extensive experiments demonstrate that our method exhibits exceptional robustness in challenging scenarios while maintaining high performance in standard scenes, significantly outperforming existing state-of-the-art methods on challenging KITTI-C, DrivingStereo, and NuScenes-Night benchmarks. Project page:\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/ruijiezhu94.github.io\/ERDepth_page\">https:\/\/ruijiezhu94.github.io\/ERDepth_page<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3750050","type":"journal-article","created":{"date-parts":[[2025,7,23]],"date-time":"2025-07-23T16:19:08Z","timestamp":1753287548000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["ER-Depth: Enhancing the Robustness of Self-Supervised Monocular Depth Estimation in Challenging Scenes"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-6348-8713","authenticated-orcid":false,"given":"Ziyang","family":"Song","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6092-0712","authenticated-orcid":false,"given":"Ruijie","family":"Zhu","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5160-5171","authenticated-orcid":false,"given":"Jing","family":"Wang","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Integrated Space-Time Network and Equipment Technology, Shijiazhuang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1431-7677","authenticated-orcid":false,"given":"Chuxin","family":"Wang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7700-788X","authenticated-orcid":false,"given":"Jianfeng","family":"He","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2838-0378","authenticated-orcid":false,"given":"Jiacheng","family":"Deng","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3599-7659","authenticated-orcid":false,"given":"Wenfei","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0764-6106","authenticated-orcid":false,"given":"Tianzhu","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,11,21]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8793512"},{"key":"e_1_3_1_3_2","first-page":"4009","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Bhat Shariq Farooq","year":"2021","unstructured":"Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka. 2021. Adabins: Depth estimation using adaptive bins. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 4009\u20134018."},{"key":"e_1_3_1_4_2","unstructured":"Shariq Farooq Bhat Reiner Birkl Diana Wofk Peter Wonka and Matthias M\u00fcller. 2023. Zoedepth: Zero-shot transfer by combining relative and metric depth. arXiv:2302.12288. Retrieved from https:\/\/arxiv.org\/abs\/2302.12288"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33783-3_44"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"e_1_3_1_7_2","article-title":"HASSOD: Hierarchical adaptive self-supervised object detection","volume":"36","author":"Cao Shengcao","year":"2024","unstructured":"Shengcao Cao, Dhiraj Joshi, Liangyan Gui, and Yu-Xiong Wang. 2024. HASSOD: Hierarchical adaptive self-supervised object detection. In Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2017.2740321"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2929202"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48891.2023.10161373"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00273"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01061"},{"key":"e_1_3_1_13_2","article-title":"Depth map prediction from a single image using a multi-scale deep network","volume":"27","author":"Eigen David","year":"2014","unstructured":"David Eigen, Christian Puhrsch, and Rob, Fergus. 2014. Depth map prediction from a single image using a multi-scale deep network. In Advances in Neural Information Processing Systems, Vol. 27.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19824-3_14"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-72670-5_14"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_45"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00751"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913491297"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.699"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00393"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i3.32330"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00256"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00185"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_25_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Hendrycks Dan","unstructured":"Dan Hendrycks and Thomas Dietterich. 2019. Benchmarking neural network robustness to common corruptions and perturbations. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_26_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Hendrycks Dan","year":"2019","unstructured":"Dan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. 2019. AugMix: A simple data processing method to improve robustness and uncertainty. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_27_2","first-page":"22106","article-title":"Semi-supervised semantic segmentation via adaptive equalization learning","volume":"34","author":"Hu Hanzhe","year":"2021","unstructured":"Hanzhe Hu, Fangyun Wei, Han Hu, Qiwei Ye, Jinshi Cui, and Liwei Wang. 2021. Semi-supervised semantic segmentation via adaptive equalization learning. In Advances in Neural Information Processing Systems, Vol. 34, 22106\u201322118.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3051462"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00907"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12297"},{"key":"e_1_3_1_31_2","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv:1412.6980. Retrieved from https:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_1_32_2","first-page":"21298","article-title":"Robodepth: Robust out-of-distribution depth estimation under corruptions","volume":"36","author":"Kong Lingdong","year":"2023","unstructured":"Lingdong Kong, Shaoyuan Xie, Hanjiang Hu, Lai Xing Ng, Benoit Cottereau, and Wei Tsang Ooi. 2023. Robodepth: Robust out-of-distribution depth estimation under corruptions. In Advances in Neural Information Processing Systems, Vol. 36, 21298\u201321342.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_33_2","unstructured":"Jin Han Lee Myung-Kyu Han Dong Wook Ko and Il Hong Suh. 2019. From big to small: Multi-scale local planar guidance for monocular depth estimation. arXiv:1907.10326. Retrieved from https:\/\/arxiv.org\/abs\/1907.10326"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00714"},{"key":"e_1_3_1_35_2","first-page":"1908","volume-title":"Proceedings of the Conference on Robot Learning","author":"Li Hanhan","year":"2021","unstructured":"Hanhan Li, Ariel Gordon, Hang Zhao, Vincent Casser, and Anelia Angelova. 2021. Unsupervised monocular depth learning in dynamic scenes. In Proceedings of the Conference on Robot Learning. PMLR, 1908\u20131917."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.108116"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2024.3416065"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3638559"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01250"},{"key":"e_1_3_1_40_2","first-page":"1136","article-title":"Plane2Depth: Hierarchical adaptive plane guidance for monocular depth estimation","volume":"2","author":"Liu Li","year":"2024","unstructured":"Li Liu, Ruijie Zhu, Jiacheng Deng, Ziyang Song, Wenfei Yang, and Tianzhu Zhang. 2024. Plane2Depth: Hierarchical adaptive plane guidance for monocular depth estimation. IEEE Transactions on Circuits and Systems for Video Technology 35, 2 (2024), 1136\u20131149.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_41_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Loshchilov Ilya","unstructured":"Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3674977"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV48630.2021.00388"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392377"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i3.16329"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01225"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0016-0032(96)00063-4"},{"key":"e_1_3_1_48_2","first-page":"2564","article-title":"DS-depth: Dynamic and static depth estimation via a fusion cost volume","volume":"4","author":"Miao Xingyu","year":"2023","unstructured":"Xingyu Miao, Yang Bai, Haoran Duan, Yawen Huang, Fan Wan, Xinxing Xu, Yang Long, and Yefeng Zheng. 2023. DS-depth: Dynamic and static depth estimation via a fusion cost volume. IEEE Transactions on Circuits and Systems for Video Technology 34, 4 (2023), 2564\u20132576.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33715-4_54"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00540"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2907904"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02672"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3663570"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01527"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00163"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8793621"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00329"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00037"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3588571"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01196"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3019967"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01252"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19769-7_6"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00818"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01551"},{"key":"e_1_3_1_66_2","first-page":"3664","article-title":"MonoDiffusion: Self-supervised monocular depth estimation using diffusion model","volume":"4","author":"Shao Shuwei","year":"2024","unstructured":"Shuwei Shao, Zhongcai Pei, Weihai Chen, Dingchi Sun, Peter C. Y. Chen, and Zhengguo Li. 2024. MonoDiffusion: Self-supervised monocular depth estimation using diffusion model. IEEE Transactions on Circuits and Systems for Video Technology 35, 4 (2024), 3664\u20133678.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58529-7_34"},{"key":"e_1_3_1_68_2","first-page":"596","article-title":"Fixmatch: Simplifying semi-supervised learning with consistency and confidence","volume":"33","author":"Sohn Kihyuk","year":"2020","unstructured":"Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A. Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li. 2020. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. In Advances in Neural Information Processing Systems, Vol. 33, 596\u2013608.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3049869"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298655"},{"key":"e_1_3_1_71_2","article-title":"Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results","volume":"30","author":"Tarvainen Antti","year":"2017","unstructured":"Antti Tarvainen and Harri Valpola. 2017. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_72_2","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Tosi Fabio","year":"2024","unstructured":"Fabio Tosi, Pierluigi Zama Ramirez, and Matteo Poggi. 2024. Diffusion models for monocular depth estimation: Overcoming challenging conditions. In Proceedings of the European Conference on Computer Vision."},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-020-01366-3"},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2019.00046"},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA57147.2024.10611100"},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681168"},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.2983686"},{"key":"e_1_3_1_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01575"},{"key":"e_1_3_1_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02052"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00864"},{"key":"e_1_3_1_81_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i6.28383"},{"key":"e_1_3_1_82_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_1_83_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00122"},{"key":"e_1_3_1_84_2","first-page":"4989","article-title":"Self-supervised multi-frame monocular depth estimation for dynamic scenes","volume":"6","author":"Wu Guanghui","year":"2023","unstructured":"Guanghui Wu, Hao Liu, Longguang Wang, Kunhong Li, Yulan Guo, and Zengping Chen. 2023. Self-supervised multi-frame monocular depth estimation for dynamic scenes. IEEE Transactions on Circuits and Systems for Video Technology 34, 6 (2023), 4989\u20135001.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_85_2","first-page":"6256","article-title":"Unsupervised data augmentation for consistency training","volume":"33","author":"Xie Qizhe","year":"2020","unstructured":"Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le. 2020. Unsupervised data augmentation for consistency training. In Advances in Neural Information Processing Systems, Vol. 33, 6256\u20136268.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_86_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00077"},{"key":"e_1_3_1_87_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00305"},{"key":"e_1_3_1_88_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00297"},{"key":"e_1_3_1_89_2","doi-asserted-by":"publisher","DOI":"10.1109\/3DV53792.2021.00056"},{"key":"e_1_3_1_90_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00099"},{"key":"e_1_3_1_91_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00699"},{"key":"e_1_3_1_92_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2020.3023634"},{"key":"e_1_3_1_93_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00212"},{"key":"e_1_3_1_94_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"You Yurong","unstructured":"Yurong You, Yan Wang, Wei-Lun Chao, Divyansh Garg, Geoff Pleiss, Bharath Hariharan, Mark Campbell, and Kilian Q. Weinberger. 2020. Pseudo-LiDAR++: Accurate depth for 3D object detection in autonomous driving. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_95_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00389"},{"key":"e_1_3_1_96_2","first-page":"18408","article-title":"Flexmatch: Boosting semi-supervised learning with curriculum pseudo labeling","volume":"34","author":"Zhang Bowen","year":"2021","unstructured":"Bowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu, Jindong Wang, Manabu Okumura, and Takahiro Shinozaki. 2021. Flexmatch: Boosting semi-supervised learning with curriculum pseudo labeling. In Advances in Neural Information Processing Systems, Vol. 34, 18408\u201318419.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_97_2","doi-asserted-by":"publisher","DOI":"10.1145\/3672397"},{"key":"e_1_3_1_98_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01778"},{"key":"e_1_3_1_99_2","doi-asserted-by":"publisher","DOI":"10.1109\/3DV57658.2022.00077"},{"key":"e_1_3_1_100_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00527"},{"key":"e_1_3_1_101_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i4.16468"},{"key":"e_1_3_1_102_2","article-title":"Self-Supervised monocular depth estimation with internal feature fusion","author":"Zhou Hang","year":"2021","unstructured":"Hang Zhou, David Greenwood, and Sarah Taylor. 2021. Self-Supervised monocular depth estimation with internal feature fusion. In British Machine Vision Conference.","journal-title":"British Machine Vision Conference"},{"key":"e_1_3_1_103_2","doi-asserted-by":"crossref","unstructured":"Hang Zhou Sarah Taylor David Greenwood and Michal Mackiewicz. 2021. Sub-depth: Self-distillation and uncertainty boosting self-supervised monocular depth estimation. arXiv:2111.09692. Retrieved from https:\/\/arxiv.org\/abs\/2111.09692","DOI":"10.5244\/C.35.208"},{"key":"e_1_3_1_104_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.700"},{"key":"e_1_3_1_105_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3335316"},{"key":"e_1_3_1_106_2","unstructured":"Ruijie Zhu Chuxin Wang Ziyang Song Li Liu Tianzhu Zhang and Yongdong Zhang. 2024. Scaledepth: Decomposing metric depth estimation into scale prediction and relative depth estimation. arXiv:2407.08187. Retrieved from https:\/\/arxiv.org\/abs\/2407.08187"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3750050","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,22]],"date-time":"2025-11-22T06:59:30Z","timestamp":1763794770000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3750050"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,21]]},"references-count":105,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2025,12,31]]}},"alternative-id":["10.1145\/3750050"],"URL":"https:\/\/doi.org\/10.1145\/3750050","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,21]]},"assertion":[{"value":"2025-03-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}