{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T09:00:48Z","timestamp":1784797248792,"version":"3.55.0"},"reference-count":64,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T00:00:00Z","timestamp":1781049600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T00:00:00Z","timestamp":1781049600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach. Intell. Res."],"published-print":{"date-parts":[[2026,8]]},"DOI":"10.1007\/s11633-025-1617-6","type":"journal-article","created":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T05:58:37Z","timestamp":1781071117000},"page":"841-854","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Adversarial Patches in 3D: A Stealth Threat to Monocular Depth Estimation"],"prefix":"10.1007","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-3092-1361","authenticated-orcid":false,"given":"Meng","family":"Pan","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yunshu","family":"Dai","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianhuang","family":"Lai","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0310-4679","authenticated-orcid":false,"given":"Xiaohua","family":"Xie","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,6,10]]},"reference":[{"issue":"6","key":"1617_CR1","doi-asserted-by":"publisher","first-page":"999","DOI":"10.1007\/s11633-025-1558-0","volume":"22","author":"Y Hu","year":"2025","unstructured":"Yufan Hu, Longhui Hu, Qingqun Kong, Bin Fan. A Survey on End-to-end Perception and Prediction for Autonomous Driving. Machine Intelligence Research, vol. 22, no. 6, pp. 999\u20131030, 2025. DOI: https:\/\/doi.org\/10.1007\/s11633-025-1558-0.","journal-title":"Machine Intelligence Research"},{"key":"1617_CR2","volume-title":"Proceedings of the 2nd International Conference on Learning Representations","author":"J Bruna","year":"2014","unstructured":"J. Bruna, C. Szegedy, I. Sutskever, I. Goodfellow, W. Zaremba, R. Fergus, D. Erhan. Intriguing properties of neural networks. In Proceedings of the 2nd International Conference on Learning Representations, Banff, Canada, 2014."},{"issue":"12","key":"1617_CR3","doi-asserted-by":"publisher","first-page":"7865","DOI":"10.1109\/TIV.2024.3403667","volume":"9","author":"J Liang","year":"2024","unstructured":"J. Liang, R. Yi, J. Chen, Y. Nie, H. Zhang. Securing autonomous vehicles visual perception: Adversarial patch attack and defense schemes with experimental validations. IEEE Transactions on Intelligent Vehicles, vol. 9, no. 12, pp. 7865\u20137875, 2024. DOI: https:\/\/doi.org\/10.1109\/TIV.2024.3403667.","journal-title":"IEEE Transactions on Intelligent Vehicles"},{"issue":"2","key":"1617_CR4","doi-asserted-by":"publisher","first-page":"832","DOI":"10.1109\/TIV.2024.3418887","volume":"10","author":"X Cai","year":"2025","unstructured":"X. Cai, X. Bai, Z. Cui, P. Hang, H. Yu, Y. Ren. Adversarial stress test for autonomous vehicle via series re-inforcement learning tasks with reward shaping. IEEE Transactions on Intelligent Vehicles, vol. 10, no. 2, pp. 832\u2013847, 2025. DOI: https:\/\/doi.org\/10.1109\/TIV.2024.3418887.","journal-title":"IEEE Transactions on Intelligent Vehicles"},{"key":"1617_CR5","unstructured":"J. Hu, T. Okatani. Analysis of deep networks for monocular depth estimation through adversarial attacks with proposal of a defense method, [Online], Available: https:\/\/arxiv.org\/abs\/1911.08790, 2019."},{"key":"1617_CR6","first-page":"99","volume-title":"Proceedings of the 5th International Conference on Learning Representations","author":"A Kurakin","year":"2017","unstructured":"A. Kurakin, I. J. Goodfellow, S. Bengio. Adversarial examples in the physical world. In Proceedings of the 5th International Conference on Learning Representations, Toulon, France, pp. 99\u2013112, 2017."},{"key":"1617_CR7","unstructured":"Z. Zhang, X. Zhu, Y. Li, X. Chen, Y. Guo. Adversarial attacks on monocular depth estimation, [Online], Available: https:\/\/arxiv.org\/abs\/2003.10315, 2020."},{"key":"1617_CR8","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"A Wong","year":"2020","unstructured":"A. Wong, S. Cicek, S. Soatto. Targeted adversarial perturbations for monocular depth prediction. In Proceedings of the 34th International Conference on Neural Information Processing Systems, Vancouver, Canada, Article number 711, 2020."},{"key":"1617_CR9","doi-asserted-by":"publisher","first-page":"3466","DOI":"10.1109\/SMC52423.2021.9658898","volume-title":"Proceedings of IEEE International Conference on Systems, Man, and Cybernetics","author":"R Daimo","year":"2021","unstructured":"R. Daimo, T. Suzuki, S. Ono. Black-box adversarial attacks on monocular depth estimation using evolutionary multi-objective optimization. In Proceedings of IEEE International Conference on Systems, Man, and Cybernetics, Melbourne, Australia, pp. 3466\u20133471, 2021. DOI: https:\/\/doi.org\/10.1109\/SMC52423.2021.9658898."},{"key":"1617_CR10","unstructured":"T. B. Brown, D. Man\u00e9, A. Roy, M. Abadi, J. Gilmer. Adversarial patch, [Online], Available: https:\/\/arxiv.org\/abs\/1712.09665, 2017."},{"key":"1617_CR11","doi-asserted-by":"publisher","first-page":"1625","DOI":"10.1109\/CVPR.2018.00175","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"K Eykholt","year":"2018","unstructured":"K. Eykholt, I. Evtimov, E. Fernandes, B. Li, A. Rahmati, C. Xiao, A. Prakash, T. Kohno, D. Song. Robust physical-world attacks on deep learning visual classification. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Salt Lake City, USA, pp. 1625\u20131634, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00175."},{"key":"1617_CR12","doi-asserted-by":"publisher","first-page":"7828","DOI":"10.1109\/ICCV48922.2021.00775","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"Y C T Hu","year":"2021","unstructured":"Y. C. T. Hu, J. C. Chen, B. H. Kung, K. L. Hua, D. S. Tan. Naturalistic physical adversarial patch for object detectors. In Proceedings of IEEE\/CVF International Conference on Computer Vision, IEEE, Montreal, Canada, pp. 7828\u20137837, 2021. DOI: https:\/\/doi.org\/10.1109\/ICCV48922.2021.00775."},{"key":"1617_CR13","doi-asserted-by":"publisher","first-page":"14242","DOI":"10.1109\/CV-PR42600.2020.01426","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Z Kong","year":"2020","unstructured":"Z. Kong, J. Guo, A. Li, C. Liu. PhysGAN: Generating physical-world-resilient adversarial examples for autonomous driving. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Seattle, USA, pp. 14242\u201314251, 2020. DOI: https:\/\/doi.org\/10.1109\/CV-PR42600.2020.01426."},{"key":"1617_CR14","doi-asserted-by":"publisher","first-page":"1528","DOI":"10.1145\/2976749.2978392","volume-title":"Proceedings of ACM SIGSAC Conference on Computer and Communications Security","author":"M Sharif","year":"2016","unstructured":"M. Sharif, S. Bhagavatula, L. Bauer, M. K. Reiter. Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition. In Proceedings of ACM SIGSAC Conference on Computer and Communications Security, Vienna, Austria, pp. 1528\u20131540, 2016. DOI: https:\/\/doi.org\/10.1145\/2976749.2978392."},{"key":"1617_CR15","doi-asserted-by":"publisher","first-page":"746","DOI":"10.1007\/978-3-642-33715-4_54","volume-title":"Proceedings of the 12th European Conference on Computer Vision","author":"N Silberman","year":"2012","unstructured":"N. Silberman, D. Hoiem, P. Kohli, R. Fergus. Indoor segmentation and support inference from RGBD images. In Proceedings of the 12th European Conference on Computer Vision, Florence, Italy, vol. 7576, pp. 746\u2013760, 2012. DOI: https:\/\/doi.org\/10.1007\/978-3-642-33715-4_54."},{"key":"1617_CR16","doi-asserted-by":"publisher","first-page":"3354","DOI":"10.1109\/CVPR.2012.6248074","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition","author":"A Geiger","year":"2012","unstructured":"A. Geiger, P. Lenz, R. Urtasun. Are we ready for autonomous driving? The KITTI vision benchmark suite. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Providence, USA, pp. 3354\u20133361, 2012. DOI: https:\/\/doi.org\/10.1109\/CVPR.2012.6248074."},{"key":"1617_CR17","first-page":"2366","volume-title":"Proceedings of the 28th International Conference on Neural Information Processing Systems","author":"D Eigen","year":"2014","unstructured":"D. Eigen, C. Puhrsch, R. Fergus. Depth map prediction from a single image using a multi-scale deep network. In Proceedings of the 28th International Conference on Neural Information Processing Systems, Montreal, Canada, pp. 2366\u20132374, 2014."},{"issue":"25","key":"1617_CR18","doi-asserted-by":"publisher","first-page":"35899","DOI":"10.1007\/s11042-021-11500-z","volume":"81","author":"M Pan","year":"2022","unstructured":"M. Pan, H. Zhang, J. Wu, Z. Jin. Self-distillation framework for indoor and outdoor monocular depth estimation. Multimedia Tools and Applications, vol. 81, no. 25, pp. 35899\u201335913, 2022. DOI: https:\/\/doi.org\/10.1007\/s11042-021-11500-z.","journal-title":"Multimedia Tools and Applications"},{"key":"1617_CR19","doi-asserted-by":"publisher","first-page":"663","DOI":"10.1007\/978-3-030-20870-7_41","volume-title":"Proceedings of the 14th Asian Conference on Computer Vision","author":"R Li","year":"2019","unstructured":"R. Li, K. Xian, C. Shen, Z. Cao, H. Lu, L. Hang. Deep attention-based classification network for robust depth prediction. In Proceedings of the 14th Asian Conference on Computer Vision, Springer, Perth, Australia, vol. 11364, pp. 663\u2013678, 2019. DOI: https:\/\/doi.org\/10.1007\/978-3-030-20870-7_41."},{"key":"1617_CR20","doi-asserted-by":"publisher","first-page":"9721","DOI":"10.1109\/CVPR.2019.00996","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"J H Lee","year":"2019","unstructured":"J. H. Lee, C. S. Kim. Monocular depth estimation using relative depth maps. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Long Beach, USA, pp. 9721\u20139730, 2019. DOI: https:\/\/doi.org\/10.1109\/CVPR.2019.00996."},{"key":"1617_CR21","doi-asserted-by":"publisher","first-page":"388","DOI":"10.1109\/ICCV.2015.52","volume-title":"Proceedings of IEEE International Conference on Computer Vision","author":"D Zoran","year":"2015","unstructured":"D. Zoran, P. Isola, D. Krishnan, W. T. Freeman. Learning ordinal relationships for mid-level vision. In Proceedings of IEEE International Conference on Computer Vision, Santiago, Chile, pp. 388\u2013396, 2015. DOI: https:\/\/doi.org\/10.1109\/ICCV.2015.52."},{"key":"1617_CR22","first-page":"730","volume-title":"Proceedings of the 30th International Conference on Neural Information Processing Systems","author":"W Chen","year":"2016","unstructured":"W. Chen, Z. Fu, D. Yang, J. Deng. Single-image depth perception in the wild. In Proceedings of the 30th International Conference on Neural Information Processing Systems, Barcelona, Spain, pp. 730\u2013738, 2016."},{"key":"1617_CR23","doi-asserted-by":"publisher","first-page":"6101","DOI":"10.1109\/ICRA.2019.8794182","volume-title":"Proceedings of International Conference on Robotics and Automation","author":"D Wofk","year":"2019","unstructured":"D. Wofk, F. Ma, T. J. Yang, S. Karaman, V. Sze. FastDepth: Fast monocular depth estimation on embedded systems. In Proceedings of International Conference on Robotics and Automation, IEEE, Montreal, Canada, pp. 6101\u20136108, 2019. DOI: https:\/\/doi.org\/10.1109\/ICRA.2019.8794182."},{"key":"1617_CR24","doi-asserted-by":"publisher","first-page":"5848","DOI":"10.1109\/IROS.2018.8593814","volume-title":"Proceedings of IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"M Poggi","year":"2018","unstructured":"M. Poggi, F. Aleotti, F. Tosi, S. Mattoccia. Towards real-time unsupervised monocular depth estimation on CPU. In Proceedings of IEEE\/RSJ International Conference on Intelligent Robots and Systems, IEEE, Madrid, Spain, pp. 5848\u20135854, 2018. DOI: https:\/\/doi.org\/10.1109\/IROS.2018.8593814."},{"key":"1617_CR25","doi-asserted-by":"publisher","first-page":"239","DOI":"10.1109\/3DV.2016.32","volume-title":"Proceedings of the 4th International Conference on 3D Vision","author":"I Laina","year":"2016","unstructured":"I. Laina, C. Rupprecht, V. Belagiannis, F. Tombari, N. Navab. Deeper depth prediction with fully convolutional residual networks. In Proceedings of the 4th International Conference on 3D Vision, IEEE, Stanford, USA, pp. 239\u2013248, 2016. DOI: https:\/\/doi.org\/10.1109\/3DV.2016.32."},{"key":"1617_CR26","doi-asserted-by":"publisher","first-page":"12159","DOI":"10.1109\/ICCV48922.2021.01196","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"R Ranftl","year":"2021","unstructured":"R. Ranftl, A. Bochkovskiy, V. Koltun. Vision transformers for dense prediction. In Proceedings of IEEE\/CVF International Conference on Computer Vision, IEEE, Montreal, Canada, pp. 12159\u201312168, 2021. DOI: https:\/\/doi.org\/10.1109\/ICCV48922.2021.01196."},{"key":"1617_CR27","doi-asserted-by":"publisher","first-page":"2183","DOI":"10.1109\/ICCV.2019.00227","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"T Van Dijk","year":"2019","unstructured":"T. Van Dijk, G. De Croon. How do neural networks see depth in single images? In Proceedings of IEEE\/CVF International Conference on Computer Vision, IEEE, Seoul, Korea (South), pp. 2183\u20132191, 2019. DOI: https:\/\/doi.org\/10.1109\/ICCV.2019.00227."},{"key":"1617_CR28","doi-asserted-by":"publisher","first-page":"7101","DOI":"10.1109\/ICRA.2019.8794220","volume-title":"Proceedings of International Conference on Robotics and Automation","author":"V Nekrasov","year":"2019","unstructured":"V. Nekrasov, T. Dharmasiri, A. Spek, T. Drummond, C. Shen, I. Reid. Real-time joint semantic segmentation and depth estimation using asymmetric annotations. In Proceedings of International Conference on Robotics and Automation, IEEE, Montreal, Canada, pp. 7101\u20137107, 2019. DOI: https:\/\/doi.org\/10.1109\/ICRA.2019.8794220."},{"key":"1617_CR29","doi-asserted-by":"publisher","first-page":"538","DOI":"10.1109\/CVPR42600.2020.00062","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"L Wang","year":"2020","unstructured":"L. Wang, J. Zhang, O. Wang, Z. Lin, H. Lu. SDC-Depth: Semantic divide-and-conquer network for monocular depth estimation. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Seattle, USA, pp. 538\u2013547, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00062."},{"key":"1617_CR30","doi-asserted-by":"publisher","first-page":"8811","DOI":"10.1109\/TIP.2021.3120670","volume":"30","author":"X Xu","year":"2021","unstructured":"X. Xu, Z. Chen, F. Yin. Multi-scale spatial attention-guided monocular depth estimation with semantic enhancement. IEEE Transactions on Image Processing, vol. 30, pp. 8811\u20138822, 2021. DOI: https:\/\/doi.org\/10.1109\/TIP.2021.3120670.","journal-title":"IEEE Transactions on Image Processing"},{"issue":"2","key":"1617_CR31","doi-asserted-by":"publisher","first-page":"836","DOI":"10.1109\/TIP.2016.2621673","volume":"26","author":"Y Cao","year":"2017","unstructured":"Y. Cao, C. Shen, H. T. Shen. Exploiting depth from single monocular images for object detection and semantic segmentation. IEEE Transactions on Image Processing, vol. 26, no. 2, pp. 836\u2013846, 2017. DOI: https:\/\/doi.org\/10.1109\/TIP.2016.2621673.","journal-title":"IEEE Transactions on Image Processing"},{"key":"1617_CR32","doi-asserted-by":"publisher","first-page":"6602","DOI":"10.1109\/CVPR.2017.699","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition","author":"C Godard","year":"2017","unstructured":"C. Godard, O. M. Aodha, G. J. Brostow. Unsupervised monocular depth estimation with left-right consistency. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, pp. 6602\u20136611, 2017. DOI: https:\/\/doi.org\/10.1109\/CVPR.2017.699."},{"key":"1617_CR33","doi-asserted-by":"publisher","first-page":"6612","DOI":"10.1109\/CVPR.2017.700","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition","author":"T Zhou","year":"2017","unstructured":"T. Zhou, M. Brown, N. Snavely, D. G. Lowe. Unsupervised learning of depth and ego-motion from video. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, pp. 6612\u20136619, 2017. DOI: https:\/\/doi.org\/10.1109\/CVPR.2017.700."},{"key":"1617_CR34","doi-asserted-by":"publisher","first-page":"2619","DOI":"10.1109\/CVPR.2019.00273","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"P Y Chen","year":"2019","unstructured":"P. Y. Chen, A. H. Liu, Y. C. Liu, Y. C. F. Wang. Towards scene understanding: Unsupervised monocular depth estimation with semantic-aware representation. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Long Beach, USA, pp. 2619\u20132627, 2019. DOI: https:\/\/doi.org\/10.1109\/CVPR.2019.00273."},{"key":"1617_CR35","volume-title":"Proceedings of the 13th International Conference on Learning Representations","author":"Z An","year":"2025","unstructured":"Z. An, G. Sun, Y. Liu, R. Li, M. Wu, M. M. Cheng, E. Konukoglu, S. Belongie. Multimodality helps few-shot 3D point cloud semantic segmentation. In Proceedings of the 13th International Conference on Learning Representations, Singapore, 2025."},{"key":"1617_CR36","doi-asserted-by":"publisher","first-page":"12232","DOI":"10.1109\/CVPR.2019.01252","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"A Ranjan","year":"2019","unstructured":"A. Ranjan, V. Jampani, L. Balles, K. Kim, D. Sun, J. Wulff, M. J. Black. Competitive collaboration: Joint unsupervised learning of depth, camera motion, optical flow and motion segmentation. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Long Beach, USA, pp. 12232\u201312241, 2019. DOI: https:\/\/doi.org\/10.1109\/CVPR.2019.01252."},{"key":"1617_CR37","doi-asserted-by":"publisher","first-page":"1983","DOI":"10.1109\/CVPR.2018.00212","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Z Yin","year":"2018","unstructured":"Z. Yin, J. Shi. GeoNet: Unsupervised learning of dense depth, optical flow and camera pose. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Salt Lake City, USA, pp. 1983\u20131992, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00212."},{"key":"1617_CR38","doi-asserted-by":"publisher","first-page":"1722","DOI":"10.1109\/SP54263.2024.00102","volume-title":"Proceedings of IEEE Symposium on Security and Privacy","author":"H Wang","year":"2024","unstructured":"H. Wang, K. Dong, Z. Zhu, H. Qin, A. Liu, X. Fang, J. Wang, X. Liu. Transferable multimodal attack on vision-language pre-training models. In Proceedings of IEEE Symposium on Security and Privacy, San Francisco, USA, pp. 1722\u20131740, 2024. DOI: https:\/\/doi.org\/10.1109\/SP54263.2024.00102."},{"key":"1617_CR39","doi-asserted-by":"publisher","first-page":"4326","DOI":"10.1109\/CVPRW50498.2020.00510","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"Q Dai","year":"2020","unstructured":"Q. Dai, V. Patii, S. Hecker, D. Dai, L. Van Gool, K. Schindler. Self-supervised object motion and depth estimation from video. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, IEEE, Seattle, USA, pp. 4326\u20134334, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPRW50498.2020.00510."},{"key":"1617_CR40","unstructured":"S. Vijayanarasimhan, S. Ricco, C. Schmid, R. Sukthankar, K. Fragkiadaki. SfM-Net: Learning of structure and motion from video, [Online], Available: https:\/\/arxiv.org\/abs\/1704.07804, 2017."},{"key":"1617_CR41","doi-asserted-by":"publisher","first-page":"283","DOI":"10.1109\/CVPR.2018.00037","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"X Qi","year":"2018","unstructured":"X. Qi, R. Liao, Z. Liu, R. Urtasun, J. Jia. GeoNet: Geometric neural network for joint depth and surface normal estimation. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Salt Lake City, USA, pp. 283\u2013291, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00037."},{"key":"1617_CR42","doi-asserted-by":"publisher","first-page":"7737","DOI":"10.1109\/CVPR52733.2024.00739","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Y Yang","year":"2024","unstructured":"Y. Yang, R. Gao, X. Wang, T. Y. Ho, N. Xu, Q. Xu. MMA-Diffusion: MultiModal attack on diffusion models. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Seattle, USA, pp. 7737\u20137746, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.00739."},{"key":"1617_CR43","doi-asserted-by":"publisher","first-page":"424","DOI":"10.1109\/3DV.2019.00054","volume-title":"Proceedings of International Conference on 3D Vision","author":"L Andraghetti","year":"2019","unstructured":"L. Andraghetti, P. Myriokefalitakis, P. L. Dovesi, B. Luque, M. Poggi, A. Pieropan, S. Mattoccia. Enhancing self-supervised monocular depth estimation with traditional visual odometry. In Proceedings of International Conference on 3D Vision, IEEE, Quebec City, Canada, pp. 424\u2013433, 2019. DOI: https:\/\/doi.org\/10.1109\/3DV.2019.00054."},{"key":"1617_CR44","doi-asserted-by":"publisher","first-page":"2022","DOI":"10.1109\/CVPR.2018.00216","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"C Wang","year":"2018","unstructured":"C. Wang, J. M. Buenaposada, R. Zhu, S. Lucey. Learning depth from monocular videos using direct methods. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Salt Lake City, USA, pp. 2022\u20132030, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00216."},{"key":"1617_CR45","doi-asserted-by":"publisher","first-page":"506","DOI":"10.1007\/978-3-030-01252-6_30","volume-title":"Proceedings of the 15th European Conference on Computer Vision","author":"X Guo","year":"2018","unstructured":"X. Guo, H. Li, S. Yi, J. Ren, X. Wang. Learning monocular depth by distilling cross-domain stereo networks. In Proceedings of the 15th European Conference on Computer Vision, Munich, Germany, vol. 11215, pp. 506\u2013523, 2018. DOI: https:\/\/doi.org\/10.1007\/978-3-030-01252-6_30."},{"key":"1617_CR46","doi-asserted-by":"publisher","first-page":"2162","DOI":"10.1109\/ICCV.2019.00225","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"J Watson","year":"2019","unstructured":"J. Watson, M. Firman, G. Brostow, D. Turmukhambetov. Self-supervised monocular depth hints. In Proceedings of IEEE\/CVF International Conference on Computer Vision, IEEE, Seoul, Korea (South), pp. 2162\u20132171, 2019. DOI: https:\/\/doi.org\/10.1109\/ICCV.2019.00225."},{"key":"1617_CR47","doi-asserted-by":"publisher","first-page":"8976","DOI":"10.1109\/ICCV.2019.00907","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"A Gordon","year":"2019","unstructured":"A. Gordon, H. Li, R. Jonschkowski, A. Angelova. Depth from videos in the wild: Unsupervised monocular depth learning from unknown cameras. In Proceedings of IEEE\/CVF International Conference on Computer Vision, IEEE, Seoul, Republic of Korea, pp. 8976\u20138985, 2019. DOI: https:\/\/doi.org\/10.1109\/ICCV.2019.00907."},{"key":"1617_CR48","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems","author":"J Bian","year":"2019","unstructured":"J. Bian, Z. Li, N. Wang, H. Zhan, C. Shen, M. M. Cheng, I. Reid. Unsupervised scale-consistent depth and ego-motion learning from monocular video. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, Vancouver, Canada, Article number 4, 2019."},{"key":"1617_CR49","doi-asserted-by":"publisher","first-page":"5667","DOI":"10.1109\/CVPR.2018.00594","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"R Mahjourian","year":"2018","unstructured":"R. Mahjourian, M. Wicke, A. Angelova. Unsupervised learning of depth and ego-motion from monocular video using 3D geometric constraints. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Salt Lake City, USA, pp. 5667\u20135675, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00594."},{"key":"1617_CR50","unstructured":"A. Mathew, A. P. Patra, J. Mathew. Monocular depth estimators: Vulnerabilities and attacks, [Online], Available: https:\/\/arxiv.org\/abs\/2005.14302, 2020."},{"key":"1617_CR51","doi-asserted-by":"publisher","first-page":"514","DOI":"10.1007\/978-3-031-19839-7_30","volume-title":"Proceedings of the 17th European Conference on Computer Vision","author":"Z Cheng","year":"2022","unstructured":"Z. Cheng, J. Liang, H. Choi, G. Tao, Z. Cao, D. Liu, X. Zhang. Physical attack on monocular depth estimation with optimal adversarial patches. In Proceedings of the 17th European Conference on Computer Vision, Tel Aviv, Israel, vol. 13698, pp. 514\u2013532, 2022. DOI: https:\/\/doi.org\/10.1007\/978-3-031-19839-7_30."},{"key":"1617_CR52","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1007\/978-3-031-19772-7_31","volume-title":"Proceedings of the 17th European Conference on Computer Vision","author":"Z Chen","year":"2022","unstructured":"Z. Chen, B. Li, S. Wu, J. Xu, S. Ding, W. Zhang. Shape matters: Deformable patch attack. In Proceedings of the 17th European Conference on Computer Vision, Tel Aviv, Israel, vol. 13664, pp. 529\u2013548, 2022. DOI: https:\/\/doi.org\/10.1007\/978-3-031-19772-7_31."},{"issue":"1","key":"1617_CR53","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1587\/transinf.2022MUL0001","volume":"E106.D","author":"R Daimo","year":"2023","unstructured":"R. Daimo, S. Ono. Projection-based physical adversarial attack for monocular depth estimation. IEICE Transactions on Information and Systems, vol. E106.D, no. 1, pp. 31\u201335, 2023. DOI: https:\/\/doi.org\/10.1587\/transinf.2022MUL0001.","journal-title":"IEICE Transactions on Information and Systems"},{"key":"1617_CR54","doi-asserted-by":"publisher","first-page":"13571","DOI":"10.1109\/ACCESS.2024.3353042","volume":"12","author":"A Guesmi","year":"2024","unstructured":"A. Guesmi, M. A. Hanif, B. Ouni, M. Shafique. SAAM: Stealthy adversarial attack on monocular depth estimation. IEEE Access, vol. 12, pp. 13571\u201313585, 2024. DOI: https:\/\/doi.org\/10.1109\/ACCESS.2024.3353042.","journal-title":"IEEE Access"},{"key":"1617_CR55","doi-asserted-by":"publisher","first-page":"24452","DOI":"10.1109\/CVPR52733.2024.02308","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"J Zheng","year":"2024","unstructured":"J. Zheng, C. Lin, J. Sun, Z. Zhao, Q. Li, C. Shen. Physical 3D adversarial attacks against monocular depth estimation in autonomous driving. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Seattle, USA, pp. 24452\u201324461, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.02308."},{"key":"1617_CR56","unstructured":"I. Alhashim, P. Wonka. High quality monocular depth estimation via transfer learning, [Online], Available: https:\/\/arxiv.org\/abs\/1812.11941, 2018."},{"key":"1617_CR57","doi-asserted-by":"publisher","first-page":"3827","DOI":"10.1109\/ICCV.2019.00393","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"C Godard","year":"2019","unstructured":"C. Godard, O. M. Aodha, M. Firman, G. Brostow. Digging into self-supervised monocular depth estimation. In Proceedings of IEEE\/CVF International Conference on Computer Vision, IEEE, Seoul, Republic of Korea, pp. 3827\u20133837, 2019. DOI: https:\/\/doi.org\/10.1109\/ICCV.2019.00393."},{"key":"1617_CR58","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems","author":"A Paszke","year":"2019","unstructured":"A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K\u00f6pf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, S. Chintala. PyTorch: An imperative style, high-performance deep learning library. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, Vancouver, Canada, Article number 721, 2019."},{"key":"1617_CR59","volume-title":"Proceedings of the 38th International Conference on Neural Information Processing Systems","author":"L Yang","year":"2024","unstructured":"L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, H. Zhao. Depth anything V2. In Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, Article number 688, 2024."},{"key":"1617_CR60","doi-asserted-by":"publisher","unstructured":"B. Ke, K. Qu, T. Wang, N. Metzger, S. Huang, B. Li, A. Obukhov, Schindler, K. Marigold: Affordable adaptation of diffusion-based image generators for image analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence, to be published. DOI: https:\/\/doi.org\/10.1109\/TPAMI.2025.3591076.","DOI":"10.1109\/TPAMI.2025.3591076"},{"issue":"22","key":"1617_CR61","doi-asserted-by":"publisher","first-page":"38440","DOI":"10.1109\/JSEN.2024.3472032","volume":"24","author":"C Zhao","year":"2024","unstructured":"C. Zhao, Y. Li, S. Wu, W. Tan, S. Zhou, Q. Pan. Physical adversarial attack on monocular depth estimation via shape-varying patches. IEEE Sensors Journal, vol. 24, no. 22, pp. 38440\u201338452, 2024. DOI: https:\/\/doi.org\/10.1109\/JSEN.2024.3472032.","journal-title":"IEEE Sensors Journal"},{"key":"1617_CR62","volume-title":"Proceedings of the 38th International Conference on Neural Information Processing Systems","author":"H Liu","year":"2024","unstructured":"H. Liu, Z. Wu, H. Wang, X. Han, S. Guo, T. Xiang, T. Zhang. Beware of road markings: A new adversarial patch attack to monocular depth estimation. In Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, Article number 2162, 2024."},{"key":"1617_CR63","doi-asserted-by":"publisher","first-page":"1028","DOI":"10.1609\/aaai.v33i01.33011028","volume-title":"Proceedings of the 33rd AAAI Conference on Artificial Intelligence","author":"A Liu","year":"2019","unstructured":"A. Liu, X. Liu, J. Fan, Y. Ma, A. Zhang, H. Xie, D. Tao. Perceptual-sensitive GAN for generating adversarial patches. In Proceedings of the 33rd AAAI Conference on Artificial Intelligence, Honolulu, USA, pp. 1028\u20131035, 2019. DOI: https:\/\/doi.org\/10.1609\/aaai.v33i01.33011028."},{"key":"1617_CR64","doi-asserted-by":"publisher","first-page":"3381","DOI":"10.1109\/WACV51458.2022.00344","volume-title":"Proceedings of IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"O Jayasinghe","year":"2022","unstructured":"O. Jayasinghe, S. Hemachandra, D. Anhettigama, S. Kariyawasam, R. Rodrigo, P. Jayasekara. CeyMo: See more on roads \u2013 A novel benchmark dataset for road marking detection. In Proceedings of IEEE\/CVF Winter Conference on Applications of Computer Vision, IEEE, Waikoloa, USA, pp. 3381\u20133390, 2022. DOI: https:\/\/doi.org\/10.1109\/WACV51458.2022.00344."}],"container-title":["Machine Intelligence Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11633-025-1617-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11633-025-1617-6","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11633-025-1617-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T08:02:33Z","timestamp":1784793753000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11633-025-1617-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,10]]},"references-count":64,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,8]]}},"alternative-id":["1617"],"URL":"https:\/\/doi.org\/10.1007\/s11633-025-1617-6","relation":{},"ISSN":["2731-538X","2731-5398"],"issn-type":[{"value":"2731-538X","type":"print"},{"value":"2731-5398","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,10]]},"assertion":[{"value":"26 June 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 November 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 June 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}