{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T09:00:53Z","timestamp":1784797253305,"version":"3.55.0"},"reference-count":151,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T00:00:00Z","timestamp":1783900800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T00:00:00Z","timestamp":1783900800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach. Intell. Res."],"published-print":{"date-parts":[[2026,8]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Embodied intelligence (EI), integrating vision-language models (VLMs) with action-oriented capabilities, presents transformative potential for autonomous systems. However, deploying VLMs in safety-critical applications like self-driving cars and collaborative robots introduces significant challenges. Key concerns include adversarial attacks and content manipulation, through which maliciously altered visual or textual inputs could lead to incorrect interpretations and hazardous actions in EI systems. Furthermore, integrating VLMs into embodied agents amplifies these risks, as real-world physical interactions introduce complex safety-critical scenarios where errors in perception or decision-making can have immediate and severe consequences. While VLM integration amplifies safety concerns, these models simultaneously offer a foundation for defensive strategies aimed at enhancing the reliability of EI systems. Additionally, VLM-based EI systems raise critical ethical concerns regarding accountability, embedded biases and societal displacement. This review systematically analyzes current applications, safety risks, protection methods and ethical implications, while proposing research directions to advance trustworthy EI systems.<\/jats:p>","DOI":"10.1007\/s11633-025-1626-x","type":"journal-article","created":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T13:14:17Z","timestamp":1783948457000},"page":"746-769","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Embodied Intelligence Security with Vision-language Models: A Survey"],"prefix":"10.1007","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0218-6924","authenticated-orcid":false,"given":"Junxian","family":"Duan","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haisu","family":"Zhu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2252-1769","authenticated-orcid":false,"given":"Canhui","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shiyi","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiangyun","family":"Tang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yingxue","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,7,13]]},"reference":[{"key":"1626_CR1","doi-asserted-by":"publisher","unstructured":"H. Liu, D. Guo, A. Cangelosi. Embodied intelligence: A synergy of morphology, action, perception and learning. ACM Computing Surveys, vol. 57, no. 7, Article number 186, 2025. DOI: https:\/\/doi.org\/10.1145\/3717059.","DOI":"10.1145\/3717059"},{"issue":"1","key":"1626_CR2","doi-asserted-by":"publisher","first-page":"12","DOI":"10.1038\/s42256-018-0009-9","volume":"1","author":"D Howard","year":"2019","unstructured":"D. Howard, A. E. Eiben, D. F. Kennedy, J. B. Mouret, P. Valencia, D. Winkler. Evolving embodied intelligence from materials to machines. Nature Machine Intelligence, vol. 1, no. 1, pp. 12\u201319, 2019. DOI: https:\/\/doi.org\/10.1038\/s42256-018-0009-9.","journal-title":"Nature Machine Intelligence"},{"key":"1626_CR3","doi-asserted-by":"publisher","unstructured":"C. Bartolozzi, G. Indiveri, E. Donati. Embodied neuromorphic intelligence. Nature Communications, vol. 13, no. 1, Article number 1024, 2022. DOI: https:\/\/doi.org\/10.1038\/s41467-022-28487-2.","DOI":"10.1038\/s41467-022-28487-2"},{"issue":"8","key":"1626_CR4","doi-asserted-by":"publisher","first-page":"5625","DOI":"10.1109\/TPAMI.2024.3369699","volume":"46","author":"J Zhang","year":"2024","unstructured":"J. Zhang, J. Huang, S. Jin, S. Lu. Vision-language models for vision tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 8, pp. 5625\u20135644, 2024. DOI: https:\/\/doi.org\/10.1109\/TPAMI.2024.3369699.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1626_CR5","doi-asserted-by":"publisher","unstructured":"A. Gupta, S. Savarese, S. Ganguli, Li F. F. Embodied intelligence via learning and evolution. Nature Communications, vol. 12, no. 1, Article number 5721, 2021. DOI: https:\/\/doi.org\/10.1038\/s41467-021-25874-z.","DOI":"10.1038\/s41467-021-25874-z"},{"key":"1626_CR6","unstructured":"Y. Ma, Z. Song, Y. Zhuang, J. Hao, I. King. A survey on vision-language-action models for embodied AI, [Online], Available: https:\/\/arxiv.org\/abs\/2405.14093, 2025."},{"key":"1626_CR7","doi-asserted-by":"publisher","unstructured":"S. Zhang, Y. Pan, Q. Liu, Z. Yan, K. K. R. Choo, G. Wang. Backdoor attacks and defenses targeting multidomain AI models: A comprehensive review. ACM Computing Surveys, vol. 57, no. 4, Article number 87, 2025. DOI: https:\/\/doi.org\/10.1145\/3704725.","DOI":"10.1145\/3704725"},{"key":"1626_CR8","doi-asserted-by":"publisher","first-page":"22650","DOI":"10.1609\/aaai.v38i20.30275","volume-title":"Proceedings of the 38th AAAI Conference on Artificial Intelligence","author":"M Vatsa","year":"2024","unstructured":"M. Vatsa, A. Jain, R. Singh. Adventures of trustworthy vision-language models: A survey. In Proceedings of the 38th AAAI Conference on Artificial Intelligence, Vancouver, Canada, pp. 22650\u201322658, 2024. DOI: https:\/\/doi.org\/10.1609\/aaai.v38i20.30275."},{"key":"1626_CR9","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Y Zhao","year":"2023","unstructured":"Y. Zhao, T. Pang, C. Du, X. Yang, C. Li, N. M. Cheung, M. Lin. On evaluating adversarial robustness of large vision-language models. In Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, USA, Article number 2355, 2023."},{"key":"1626_CR10","doi-asserted-by":"publisher","first-page":"135","DOI":"10.1109\/WRCSARA64167.2024.10685688","volume-title":"Proceedings of WRC Symposium on Advanced Robotics and Automation","author":"J Huang","year":"2024","unstructured":"J. Huang, C. Limberg, S. M. N. Arshad, Q. Zhang, Q. Li. Combining VLM and LLM for enhanced semantic object perception in robotic handover tasks. In Proceedings of WRC Symposium on Advanced Robotics and Automation, Beijing, China, pp. 135\u2013140, 2024. DOI: https:\/\/doi.org\/10.1109\/WRCSARA64167.2024.10685688."},{"issue":"4","key":"1626_CR11","doi-asserted-by":"publisher","first-page":"588","DOI":"10.1007\/s11633-025-1542-8","volume":"22","author":"Y Zheng","year":"2025","unstructured":"Y. Zheng, L. Yao, Y. Su, Y. Zhang, Y. Wang, S. Zhao, Y. Zhang, L. P. Chau. A survey of embodied learning for object-centric robotic manipulation. Machine Intelligence Research, vol. 22, no. 4, pp. 588\u2013626, 2025. DOI: https:\/\/doi.org\/10.1007\/s11633-025-1542-8.","journal-title":"Machine Intelligence Research"},{"key":"1626_CR12","doi-asserted-by":"publisher","unstructured":"S. Nordhoff, J. D. Lee, S. C. Calvert, S. Berge, M. Hagenzieker, R. Happee. (Mis-)use of standard autopilot and full self-driving (FSD) Beta: Results from interviews with users of Tesla\u2019s FSD Beta. Frontiers in Psychology, vol. 14, Article number 1101520, 2023. DOI: https:\/\/doi.org\/10.3389\/fpsyg.2023.1101520.","DOI":"10.3389\/fpsyg.2023.1101520"},{"key":"1626_CR13","unstructured":"Tesla, Inc. Autopilot and full self driving capability, [Online], Available: https:\/\/www.tesla.com\/autopilot, 2024."},{"key":"1626_CR14","unstructured":"Boston Dynamics. Spot\u00ae\u2013The Agile Mobile robot, [Online], Available: https:\/\/bostondynamics.com\/products\/spot\/, 2024."},{"key":"1626_CR15","unstructured":"ANYbotics. ANYmal for industrial inspection, [Online], Available: https:\/\/www.anybotics.com\/anymal-industrial-inspection\/, 2024."},{"key":"1626_CR16","unstructured":"C. Cui, P. Ding, W. Song, S. Bai, X. Tong, Z. Ge, R. Suo, W. Zhou, Y. Liu, B. Jia, H. Zhao, S. Huang, D. Wang. OpenHelix: A short survey, empirical analysis, and open-source dual-system VLA model for robotic manipulation, [Online], Available: https:\/\/arxiv.org\/abs\/2505.03912, 2025."},{"key":"1626_CR17","first-page":"8469","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"D Driess","year":"2023","unstructured":"D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, P. Florence. PaLM-E: An embodied multimodal language model. In Proceedings of the 40th International Conference on Machine Learning, Honolulu, USA, pp. 8469\u20138488, 2023."},{"key":"1626_CR18","unstructured":"Amazon, Inc. Amazon has more than 1 million robots that sort, lift, and carry packages\u2013see them in action, [Online], Available: https:\/\/www.aboutamazon.com\/news\/operations\/amazon-robotics-robots-fulfillment-center, 2024."},{"key":"1626_CR19","doi-asserted-by":"publisher","first-page":"16847","DOI":"10.1109\/ICRA55743.2025.11128705","volume-title":"Proceedings of IEEE International Conference on Robotics and Automation","author":"Z Yang","year":"2025","unstructured":"Z. Yang, C. Garrett, D. Fox, T. Lozano-P\u00e9rez, L. P. Kaelbling. Guiding long-horizon task and motion planning with vision language models. In Proceedings of IEEE International Conference on Robotics and Automation, Atlanta, USA, pp. 16847\u201316853, 2025. DOI: https:\/\/doi.org\/10.1109\/ICRA55743.2025.11128705."},{"key":"1626_CR20","unstructured":"X. Zhu, L. Wang, C. Zhou, X. Cao, Y. Gong, L. Chen. A survey on deep learning approaches for data integration in autonomous driving system, [Online], Available: https:\/\/arxiv.org\/abs\/2306.11740, 2023."},{"key":"1626_CR21","unstructured":"Toyota Research Institute. Teaching robots new skills for the home, [Online], Available: https:\/\/www.youtube.com\/watch?v=6IGCIjp2bn4&t=11s, 2020."},{"key":"1626_CR22","unstructured":"SoftBank Robotics. Pepper the humanoid robot, [Online], Available: https:\/\/www.softbank.jp\/en\/robot\/, 2021."},{"key":"1626_CR23","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"H Liu","year":"2023","unstructured":"H. Liu, C. Li, Q. Wu, Y. J. Lee. Visual instruction tuning. In Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, USA, Article number 1516, 2023."},{"key":"1626_CR24","first-page":"16847","volume-title":"Proceedings of the 19th Robotics: Science and Systems","author":"S Belkhale","year":"2024","unstructured":"S. Belkhale, T. Ding, T. Xiao, P. Sermanet, Q. Vuong, J. Tompson, Y. Chebotar, D. Dwibedi, D. Sadigh. RTH: Action hierarchies using language. In Proceedings of the 19th Robotics: Science and Systems, Delft, The Netherlands, pp. 16847\u201316853, 2024."},{"issue":"2","key":"1626_CR25","doi-asserted-by":"publisher","first-page":"230","DOI":"10.1109\/TETCI.2022.3141105","volume":"6","author":"J Duan","year":"2022","unstructured":"J. Duan, S. Yu, H. L. Tan, H. Zhu, C. Tan. A survey of embodied AI: From simulators to research tasks. IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 6, no. 2, pp. 230\u2013244, 2022. DOI: https:\/\/doi.org\/10.1109\/TETCI.2022.3141105.","journal-title":"IEEE Transactions on Emerging Topics in Computational Intelligence"},{"issue":"4","key":"1626_CR26","doi-asserted-by":"publisher","first-page":"713","DOI":"10.1007\/s11633-024-1537-x","volume":"22","author":"H Yuan","year":"2025","unstructured":"H. Yuan, Y. Huang, N. Yu, D. Zhang, Z. Du, Z. Liu, K. Zhang. Multimodal pretrained knowledge for real-world object navigation. Machine Intelligence Research, vol. 22, no. 4, pp. 713\u2013729, 2025. DOI: https:\/\/doi.org\/10.1007\/s11633-024-1537-x.","journal-title":"Machine Intelligence Research"},{"key":"1626_CR27","doi-asserted-by":"publisher","first-page":"7606","DOI":"10.18653\/v1\/2022.acl-long.524","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics","author":"J Gu","year":"2022","unstructured":"J. Gu, E. Stefani, Q. Wu, J. Thomason, X. Wang. Vision-and-language navigation: A survey of tasks, methods, and future directions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, Dublin, Ireland, pp. 7606\u20137623, 2022. DOI: https:\/\/doi.org\/10.18653\/v1\/2022.acl-long.524."},{"key":"1626_CR28","volume-title":"Proceedings of the 38th International Conference on Neural Information Processing Systems","author":"J Zheng","year":"2024","unstructured":"J. Zheng, J. Li, S. Cheng, Y. Zheng, J. Li, J. Liu, Y. Liu, J. Liu, X. Zhan. Instruction-guided visual masking. In Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, Article number 4003, 2024."},{"key":"1626_CR29","unstructured":"Furhat Robotics. Furhat robotics, [Online], Available: https:\/\/furhatrobotics.com\/, 2023."},{"key":"1626_CR30","doi-asserted-by":"publisher","first-page":"14281","DOI":"10.1109\/CVPR52733.2024.01354","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"W Hong","year":"2024","unstructured":"W. Hong, W. Wang, Q. Lv, J. Xu, W. Yu, J. Ji, Y. Wang, Z. Wang, Y. Dong, M. Ding, J. Tang. CogAgent: A visual language model for GUI agents. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 14281\u201314290, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.01354."},{"key":"1626_CR31","doi-asserted-by":"publisher","first-page":"1588","DOI":"10.1109\/HRI61500.2025.10974117","volume-title":"Proceedings of the 20th ACM\/IEEE International Conference on Human-Robot Interaction","author":"O Sautenkov","year":"2025","unstructured":"O. Sautenkov, Y. Yaqoot, A. Lykov, M. A. Mustafa, G. Tadevosyan, A. Akhmetkazy, M. A. Cabrera, M. Martynov, S. Karaf, D. Tsetserukou. UAV-VLA: Vision-language-action system for large scale aerial mission generation. In Proceedings of the 20th ACM\/IEEE International Conference on Human-Robot Interaction, Melbourne, Australia, pp. 1588\u20131592, 2025. DOI: https:\/\/doi.org\/10.1109\/HRI61500.2025.10974117."},{"key":"1626_CR32","doi-asserted-by":"publisher","first-page":"652","DOI":"10.1145\/3678957.3689030","volume-title":"Proceedings of the 26th International Conference on Multimodal Interaction","author":"M Spitale","year":"2024","unstructured":"M. Spitale, M. T. Parreira, M. Stiber, M. Axelsson, N. Kara, G. Kankariya, C. M. Huang, M. Jung, W. Ju, H. Gunes. ERR@HRI 2024 challenge: Multimodal detection of errors and failures in human-robot interactions. In Proceedings of the 26th International Conference on Multimodal Interaction, San Jose, Costa Rica, pp. 652\u2013656, 2024. DOI: https:\/\/doi.org\/10.1145\/3678957.3689030."},{"key":"1626_CR33","doi-asserted-by":"publisher","first-page":"13744","DOI":"10.1109\/CVPR52734.2025.01283","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"K Bae","year":"2025","unstructured":"K. Bae, J. Kim, S. Lee, S. Lee, G. Lee, J. Choi. MASHVLM: Mitigating action-scene hallucination in video-LLMs through disentangled spatial-temporal representations. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, pp. 13744\u201313753, 2025. DOI: https:\/\/doi.org\/10.1109\/CVPR52734.2025.01283."},{"key":"1626_CR34","doi-asserted-by":"publisher","first-page":"12944","DOI":"10.1109\/CVPR52733.2024.01230","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Q Yu","year":"2024","unstructured":"Q. Yu, J. Li, L. Wei, L. Pang, W. Ye, B. Qin, S. Tang, Q. Tian, Y. Zhuang. HalluciDoctor: Mitigating hallucinatory toxicity in visual instruction data. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 12944\u201312953, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.01230."},{"key":"1626_CR35","doi-asserted-by":"publisher","first-page":"7414","DOI":"10.1145\/3664647.3680795","volume-title":"Proceedings of the 32nd ACM International Conference on Multimedia","author":"Z Cai","year":"2024","unstructured":"Z. Cai, S. Ghosh, A. P. Adatia, M. Hayat, A. Dhall, T. Gedeon, K. Stefanov. AV-Deepfake1M: A large-scale LLM-driven audio-visual deepfake dataset. In Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, Australia, pp. 7414\u20137423, 2024. DOI: https:\/\/doi.org\/10.1145\/3664647.3680795."},{"key":"1626_CR36","doi-asserted-by":"publisher","first-page":"19900","DOI":"10.1109\/CVPR52734.2025.01853","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"J Zhang","year":"2025","unstructured":"J. Zhang, J. Ye, X. Ma, Y. Li, Y. Yang, Y. Chen, J. Sang, D. Y. Yeung. Anyattack: Towards large-scale self-supervised adversarial attacks on vision-language models. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, pp. 19900\u201319909, 2025. DOI: https:\/\/doi.org\/10.1109\/CVPR52734.2025.01853."},{"key":"1626_CR37","doi-asserted-by":"publisher","first-page":"3347","DOI":"10.1109\/WACV48630.2021.00339","volume-title":"Proceedings of IEEE Winter Conference on Applications of Computer Vision","author":"S Hussain","year":"2021","unstructured":"S. Hussain, P. Neekhara, M. Jere, F. Koushanfar, J. McAuley. Adversarial deepfakes: Evaluating vulnerability of deepfake detectors to adversarial examples. In Proceedings of IEEE Winter Conference on Applications of Computer Vision, Waikoloa, USA, pp. 3347\u20133356, 2021. DOI: https:\/\/doi.org\/10.1109\/WACV48630.2021.00339."},{"key":"1626_CR38","doi-asserted-by":"publisher","first-page":"24841","DOI":"10.1109\/CVPR52733.2024.02346","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Z Fang","year":"2024","unstructured":"Z. Fang, R. Wang, T. Huang, L. Jing. Strong transferable adversarial attacks via ensembled asymptotically normal distribution learning. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 24841\u201324850, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.02346."},{"key":"1626_CR39","volume-title":"Proceedings of International Conference on Learning Representations","author":"J Li","year":"2024","unstructured":"J. Li, A. Kumar, Y. Zhang, D. Song. Safeguarding data in multimodal AI: A differentially private approach to CLIP training. In Proceedings of International Conference on Learning Representations, Vienna, Austria, 2024."},{"key":"1626_CR40","first-page":"1527","volume-title":"Proceedings of the 27th USENIX Conference on Security Symposium","author":"K Zeng","year":"2018","unstructured":"K. Zeng, S. Liu, Y. Shu, D. Wang, H. Li, Y. Dou, G. Wang, Y. Yang. All your GPS are belong to us: Towards stealthy manipulation of road navigation systems. In Proceedings of the 27th USENIX Conference on Security Symposium, Baltimore, USA, pp. 1527\u20131544, 2018."},{"key":"1626_CR41","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Y Zhou","year":"2024","unstructured":"Y. Zhou, C. Cui, J. Yoon, L. Zhang, Z. Deng, C. Finn, M. Bansal, H. Yao. Analyzing and mitigating object hallucination in large vision-language models. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, 2024."},{"key":"1626_CR42","doi-asserted-by":"publisher","first-page":"10610","DOI":"10.1109\/CVPR52734.2025.00992","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Z Yang","year":"2025","unstructured":"Z. Yang, X. Luo, D. Han, Y. Xu, D. Li. Mitigating hallucinations in large vision-language models via DPO: On-policy data hold the key. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, pp. 10610\u201310620, 2025. DOI: https:\/\/doi.org\/10.1109\/CVPR52734.2025.00992."},{"key":"1626_CR43","doi-asserted-by":"publisher","first-page":"24331","DOI":"10.1109\/CVPR52734.2025.02266","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"C T Lei","year":"2025","unstructured":"C. T. Lei, H. M. Yam, Z. Guo, Y. Qian, C. P. Lau. Instant adversarial purification with adversarial consistency distillation. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, pp. 24331\u201324340, 2025. DOI: https:\/\/doi.org\/10.1109\/CVPR52734.2025.02266."},{"key":"1626_CR44","doi-asserted-by":"publisher","first-page":"13917","DOI":"10.1109\/CVPR52734.2025.01299","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Q Wang","year":"2025","unstructured":"Q. Wang, C. Li, Y. Luo, H. Ling, S. Huang, R. Jia, N. Yu. Detecting adversarial data using perturbation forgery. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, pp. 13917\u201313926, 2025. DOI: https:\/\/doi.org\/10.1109\/CVPR52734.2025.01299."},{"key":"1626_CR45","doi-asserted-by":"publisher","first-page":"24820","DOI":"10.1109\/CVPR52733.2024.02344","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"A M Ishmam","year":"2024","unstructured":"A. M. Ishmam, C. Thomas. Semantic shield: Defending vision-language models against backdooring and poisoning via fine-grained knowledge alignment. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 24820\u201324830, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.02344."},{"key":"1626_CR46","doi-asserted-by":"publisher","first-page":"7766","DOI":"10.1609\/aaai.v38i7.28611","volume-title":"Proceedings of the 38th AAAI Conference on Artificial Intelligence","author":"L Zhu","year":"2024","unstructured":"L. Zhu, R. Ning, J. Li, C. Xin, H. Wu. SEER: Backdoor detection for vision-language models through searching target text and image trigger jointly. In Proceedings of the 38th AAAI Conference on Artificial Intelligence, Vancouver, Canada, pp. 7766\u20137774, 2024. DOI: https:\/\/doi.org\/10.1609\/aaai.v38i7.28611."},{"key":"1626_CR47","volume-title":"Proceedings of the 35th International Conference on Neural Information Processing Systems","author":"B Ghazi","year":"2021","unstructured":"B. Ghazi, N. Golowich, R. Kumar, P. Manurangsi, C. Zhang. Deep learning with label differential privacy. In Proceedings of the 35th International Conference on Neural Information Processing Systems, Article number 2078, 2021."},{"key":"1626_CR48","doi-asserted-by":"publisher","first-page":"9757","DOI":"10.1109\/ICRA48891.2023.10161112","volume-title":"Proceedings of IEEE International Conference on Robotics and Automation","author":"R K Cosner","year":"2023","unstructured":"R. K. Cosner, Y. Chen, K. Leung, M. Pavone. Learning responsibility allocations for safe human-robot interaction with applications to autonomous driving. In Proceedings of IEEE International Conference on Robotics and Automation, London, UK, pp. 9757\u20139763, 2023. DOI: https:\/\/doi.org\/10.1109\/ICRA48891.2023.10161112."},{"key":"1626_CR49","doi-asserted-by":"publisher","DOI":"10.1145\/3544549.3585696","volume-title":"Proceedings of CHI Conference on Human Factors in Computing Systems","author":"E Schneiders","year":"2023","unstructured":"E. Schneiders, E. Papachristos, N. van Berkel, R. M. Jacobsen. \u201cBriefly entertaining but pointless\u201d: Perceived benefits & risks of social robots in the home. In Proceedings of CHI Conference on Human Factors in Computing Systems, Hamburg, Germany, Article number 61, 2023. DOI: https:\/\/doi.org\/10.1145\/3544549.3585696."},{"issue":"1","key":"1626_CR50","doi-asserted-by":"publisher","first-page":"508","DOI":"10.1109\/LRA.2024.3511409","volume":"10","author":"D Song","year":"2025","unstructured":"D. Song, J. Liang, A. Payandeh, A. H. Raj, X. Xiao, D. Manocha. VLM-Social-Nav: Socially aware robot navigation through scoring using vision-language models. IEEE Robotics and Automation Letters, vol. 10, no. 1, pp. 508\u2013515, 2025. DOI: https:\/\/doi.org\/10.1109\/LRA.2024.3511409.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"1626_CR51","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"H You","year":"2024","unstructured":"H. You, H. Zhang, Z. Gan, X. Du, B. Zhang, Z. Wang, L. Cao, S. F. Chang, Y. Yang. Ferret: Refer and ground anything anywhere at any granularity. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, 2024."},{"key":"1626_CR52","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"F Liu","year":"2024","unstructured":"F. Liu, K. Lin, L. Li, J. Wang, Y. Yacoob, L. Wang. Mitigating hallucination in large multi-modal models via robust instruction tuning. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, 2024."},{"key":"1626_CR53","doi-asserted-by":"publisher","first-page":"26753","DOI":"10.1109\/CVPR52733.2024.02527","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Z Li","year":"2024","unstructured":"Z. Li, B. Yang, Q. Liu, Z. Ma, S. Zhang, J. Yang, Y. Sun, Y. Liu, X. Bai. Monkey: Image resolution and text label are important things for large multi-modal models. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 26753\u201326763, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.02527."},{"key":"1626_CR54","doi-asserted-by":"publisher","first-page":"517","DOI":"10.18653\/v1\/2022.findings-naacl.39","volume-title":"Proceedings of Findings of the Association for Computational Linguistics","author":"J Cho","year":"2022","unstructured":"J. Cho, S. Yoon, A. Kale, F. Dernoncourt, T. Bui, M. Bansal. Fine-grained image captioning with CLIP reward. In Proceedings of Findings of the Association for Computational Linguistics, Seattle, USA, pp. 517\u2013527, 2022. DOI: https:\/\/doi.org\/10.18653\/v1\/2022.findings-naacl.39."},{"key":"1626_CR55","doi-asserted-by":"publisher","first-page":"27992","DOI":"10.1109\/CVPR52733.2024.02644","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"J Jain","year":"2024","unstructured":"J. Jain, J. Yang, H. Shi. VCoder: Versatile vision encoders for multimodal large language models. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 27992\u201328002, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.02644."},{"key":"1626_CR56","doi-asserted-by":"publisher","first-page":"26286","DOI":"10.1109\/CVPR52733.2024.02484","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"H Liu","year":"2024","unstructured":"H. Liu, C. Li, Y. Li, Y. J. Lee. Improved baselines with visual instruction tuning. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 26286\u201326296, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.02484."},{"key":"1626_CR57","doi-asserted-by":"publisher","first-page":"59","DOI":"10.18653\/v1\/2024.naacl-long.4","volume-title":"Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Y Li","year":"2024","unstructured":"Y. Li, H. Wang, C. Zhang. Assessing logical puzzle solving in large language models: Insights from a minesweeper case study. In Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Mexico City, Mexico, pp. 59\u201381, 2024. DOI: https:\/\/doi.org\/10.18653\/v1\/2024.naacl-long.4."},{"key":"1626_CR58","doi-asserted-by":"publisher","first-page":"13418","DOI":"10.1109\/CVPR52733.2024.01274","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Q Huang","year":"2024","unstructured":"Q. Huang, X. Dong, P. Zhang, B. Wang, C. He, J. Wang, D. Lin, W. Zhang, N. Yu. OPERA: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 13418\u201313427, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.01274."},{"key":"1626_CR59","doi-asserted-by":"publisher","first-page":"973","DOI":"10.1109\/CVPRW59228.2023.00104","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"R Corvi","year":"2023","unstructured":"R. Corvi, D. Cozzolino, G. Poggi, K. Nagano, L. Verdoliva. Intriguing properties of synthetic images: From generative adversarial networks to diffusion models. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Vancouver, Canada, pp. 973\u2013982, 2023. DOI: https:\/\/doi.org\/10.1109\/CVPRW59228.2023.00104."},{"key":"1626_CR60","doi-asserted-by":"publisher","first-page":"2185","DOI":"10.1109\/CVPR46437.2021.00222","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"H Zhao","year":"2021","unstructured":"H. Zhao, T. Wei, W. Zhou, W. Zhang, D. Chen, N. Yu. Multi-attentional deepfake detection. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, pp. 2185\u20132194, 2021. DOI: https:\/\/doi.org\/10.1109\/CVPR46437.2021.00222."},{"key":"1626_CR61","doi-asserted-by":"publisher","first-page":"3473","DOI":"10.1145\/3474085.3475508","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"Z Gu","year":"2021","unstructured":"Z. Gu, Y. Chen, T. Yao, S. Ding, J. Li, F. Huang, L. Ma. Spatiotemporal inconsistency learning for deepfake video detection. In Proceedings of the 29th ACM International Conference on Multimedia, pp. 3473\u20133481, 2021. DOI: https:\/\/doi.org\/10.1145\/3474085.3475508."},{"key":"1626_CR62","doi-asserted-by":"publisher","DOI":"10.1109\/SPW.2018.00009","volume-title":"Proceedings of IEEE Security and Privacy Workshops","author":"N Carlini","year":"2018","unstructured":"N. Carlini, D. Wagner. Audio adversarial examples: Targeted attacks on speech-to-text. In Proceedings of IEEE Security and Privacy Workshops, San Francisco, USA, 2018. DOI: https:\/\/doi.org\/10.1109\/SPW.2018.00009."},{"key":"1626_CR63","doi-asserted-by":"publisher","first-page":"318","DOI":"10.1007\/978-3-030-98355-0_27","volume-title":"Proceedings of the 28th International Conference on Multimedia Modeling","author":"N H Vo","year":"2022","unstructured":"N. H. Vo, K. D. Phan, A. D. Tran, D. T. D. Nguyen. Adversarial attacks on deepfake detectors: A practical analysis. In Proceedings of the 28th International Conference on Multimedia Modeling, Phu Quoc, Vietnam, pp. 318\u2013330, 2022. DOI: https:\/\/doi.org\/10.1007\/978-3-030-98355-0_27."},{"key":"1626_CR64","doi-asserted-by":"publisher","unstructured":"S. Hussain, P. Neekhara, B. Dolhansky, J. Bitton, C. C. Ferrer, J. McAuley, F. Koushanfar. Exposing vulnerabilities of deepfake detection systems with robust attacks. Digital Threats: Research and Practice, vol. 3, no. 3, Article number 30, 2022. DOI: https:\/\/doi.org\/10.1145\/3464307.","DOI":"10.1145\/3464307"},{"issue":"4","key":"1626_CR65","doi-asserted-by":"publisher","first-page":"2429","DOI":"10.1109\/TPAMI.2024.3519803","volume":"47","author":"X Sun","year":"2025","unstructured":"X. Sun, G. Cheng, H. Li, C. Lang, J. Han. STDatav2: Accessing efficient black-box stealing for adversarial attacks. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 4, pp. 2429\u20132445, 2025. DOI: https:\/\/doi.org\/10.1109\/TPAMI.2024.3519803.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1626_CR66","volume-title":"Proceedings of International Conference on Learning Representations","author":"C Xia","year":"2025","unstructured":"C. Xia, F. Ma, R. Quan, K. Zhan, Y. Yang. Adversarial-guided diffusion for robust and high-fidelity multimodal LLM attacks. In Proceedings of International Conference on Learning Representations, Singapore, 2025."},{"key":"1626_CR67","doi-asserted-by":"publisher","first-page":"130","DOI":"10.1109\/BigDIA60676.2023.10429108","volume-title":"Proceedings of the 9th International Conference on Big Data and Information Analytics","author":"H Sun","year":"2023","unstructured":"H. Sun, Z. Li, L. Liu, B. Li. Real is not true: Backdoor attacks against deepfake detection. In Proceedings of the 9th International Conference on Big Data and Information Analytics, Haikou, China, pp. 130\u2013137, 2023. DOI: https:\/\/doi.org\/10.1109\/BigDIA60676.2023.10429108."},{"key":"1626_CR68","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"J Liang","year":"2024","unstructured":"J. Liang, S. Liang, A. Liu, X. Jia, J. Kuang, X. Cao. Poisoned forgery face: Towards backdoor attacks on face forgery detection. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, 2024."},{"key":"1626_CR69","doi-asserted-by":"publisher","first-page":"1211","DOI":"10.1109\/ICME55011.2023.00211","volume-title":"Proceedings of IEEE International Conference on Multimedia and Expo","author":"Y Zeng","year":"2023","unstructured":"Y. Zeng, J. Tan, Z. You, Z. Qian, X. Zhang. Watermarks for generative adversarial network based on steganographic invisible backdoor. In Proceedings of IEEE International Conference on Multimedia and Expo, Brisbane, Australia, pp. 1211\u20131216, 2023. DOI: https:\/\/doi.org\/10.1109\/ICME55011.2023.00211."},{"key":"1626_CR70","doi-asserted-by":"publisher","first-page":"23951","DOI":"10.1609\/aaai.v39i22.34568","volume-title":"Proceedings of the 39th AAAI Conference on Artificial Intelligence","author":"Y Gong","year":"2025","unstructured":"Y. Gong, D. Ran, J. Liu, C. Wang, T. Cong, A. Wang, S. Duan, X. Wang. FigStep: Jailbreaking large vision-language models via typographic visual prompts. In Proceedings of the 39th AAAI Conference on Artificial Intelligence, Philadelphia, USA, pp. 23951\u201323959, 2025. DOI: https:\/\/doi.org\/10.1609\/aaai.v39i22.34568."},{"key":"1626_CR71","doi-asserted-by":"publisher","first-page":"21527","DOI":"10.1609\/aaai.v38i19.30150","volume-title":"Proceedings of AAAI Conference on Artificial Intelligence","author":"X Qi","year":"2024","unstructured":"X. Qi, K. Huang, A. Panda, P. Henderson, M. Wang, P. Mittal. Visual adversarial examples jailbreak aligned large language models. In Proceedings of AAAI Conference on Artificial Intelligence, Vancouver, Canada, pp. 21527\u201321536, 2024. DOI: https:\/\/doi.org\/10.1609\/aaai.v38i19.30150."},{"issue":"1","key":"1626_CR72","doi-asserted-by":"publisher","first-page":"38","DOI":"10.1007\/s11633-022-1369-5","volume":"20","author":"F L Chen","year":"2023","unstructured":"F. L. Chen, D. Z. Zhang, M. L. Han, X. Y. Chen, J. Shi, S. Xu, B. Xu. VLP: A survey on vision-language pre-training. Machine Intelligence Research, vol. 20, no. 1, pp. 38\u201356, 2023. DOI: https:\/\/doi.org\/10.1007\/s11633-022-1369-5.","journal-title":"Machine Intelligence Research"},{"key":"1626_CR73","doi-asserted-by":"publisher","first-page":"148","DOI":"10.1007\/978-3-031-51630-6_10","volume-title":"Proceedings of the 1st EAI International Conference on Security and Privacy in Cyber-Physical Systems and Smart Vehicles","author":"D Afroze","year":"2024","unstructured":"D. Afroze, Y. Tu, X. Hei. Securing the future: Exploring privacy risks and security questions in robotic systems. In Proceedings of the 1st EAI International Conference on Security and Privacy in Cyber-Physical Systems and Smart Vehicles, Chicago, USA, pp. 148\u2013157, 2024. DOI: https:\/\/doi.org\/10.1007\/978-3-031-51630-6_10."},{"issue":"4","key":"1626_CR74","doi-asserted-by":"publisher","first-page":"3422","DOI":"10.1109\/TDSC.2023.3326299","volume":"21","author":"J Hou","year":"2024","unstructured":"J. Hou, D. Liu, C. Huang, W. Zhuang, X. Shen, R. Sun, B. Ying. Data protection: Privacy-preserving data collection with validation. IEEE Transactions on Dependable and Secure Computing, vol. 21, no. 4, pp. 3422\u20133438, 2024. DOI: https:\/\/doi.org\/10.1109\/TDSC.2023.3326299.","journal-title":"IEEE Transactions on Dependable and Secure Computing"},{"key":"1626_CR75","volume-title":"Proceedings of the 33rd USENIX Conference on Security Symposium","author":"Z Wang","year":"2024","unstructured":"Z. Wang, R. Zhu, D. Zhou, Z. Zhang, J. Mitchell, H. Tang, X. F. Wang. DPAdapter: Improving differentially private deep learning through noise tolerance pretraining. In Proceedings of the 33rd USENIX Conference on Security Symposium, Philadelphia, USA, Article number 56, 2024."},{"issue":"3","key":"1626_CR76","doi-asserted-by":"publisher","first-page":"477","DOI":"10.56553\/popets-2024-0089","volume":"2024","author":"L Ren","year":"2024","unstructured":"L. Ren, Z. Liu, F. Li, K. Liang, Z. Li, B. Luo. PrivDNN: A secure multi-party computation framework for deep learning using partial DNN encryption. Proceedings on Privacy Enhancing Technologies, vol. 2024, no. 3, pp. 477\u2013494, 2024. DOI: https:\/\/doi.org\/10.56553\/popets-2024-0089.","journal-title":"Proceedings on Privacy Enhancing Technologies"},{"issue":"1","key":"1626_CR77","doi-asserted-by":"publisher","first-page":"19","DOI":"10.1007\/s11633-022-1343-2","volume":"20","author":"Q Yang","year":"2023","unstructured":"Q. Yang, A. Huang, L. Fan, C. S. Chan, J. H. Lim, K. W. Ng, D. S. Ong, B. Li. Federated learning with privacy-preserving and model IP-right-protection. Machine Intelligence Research, vol. 20, no. 1, pp. 19\u201337, 2023. DOI: https:\/\/doi.org\/10.1007\/s11633-022-1343-2.","journal-title":"Machine Intelligence Research"},{"key":"1626_CR78","doi-asserted-by":"publisher","unstructured":"K. Rado\u0161, M. Brki\u0107, D. Begu\u0161i\u0107. Recent advances on jamming and spoofing detection in GNSS. Sensors, vol. 24, no. 13, Article number 4210, 2024. DOI: https:\/\/doi.org\/10.3390\/s24134210.","DOI":"10.3390\/s24134210"},{"key":"1626_CR79","doi-asserted-by":"publisher","DOI":"10.1109\/LADC53747.2021.9672561","volume-title":"Proceedings of the 80th Latin-American Symposium on Dependable Computing","author":"G de Carvalho Bertoli","year":"2021","unstructured":"G. de Carvalho Bertoli, L. A. Pereira, O. Saotome. Classification of denial of service attacks on Wi-Fi-based unmanned aerial vehicle. In Proceedings of the 80th Latin-American Symposium on Dependable Computing, Florian\u00f3polis, Brazil, 2021. DOI: https:\/\/doi.org\/10.1109\/LADC53747.2021.9672561."},{"key":"1626_CR80","doi-asserted-by":"publisher","unstructured":"A. Rugo, C. A. Ardagna, N. El Ioini. A security review in the UAVNet era: Threats, countermeasures, and gap analysis. ACM Computing Surveys, vol. 55, no. 1, Article number 21, 2022. DOI: https:\/\/doi.org\/10.1145\/3485272.","DOI":"10.1145\/3485272"},{"key":"1626_CR81","doi-asserted-by":"publisher","first-page":"1032","DOI":"10.1109\/ACSAC63791.2024.00085","volume-title":"Proceedings of Annual Computer Security Applications Conference","author":"B Srimoungchanh","year":"2024","unstructured":"B. Srimoungchanh, J. G. Morris, D. Davidson. Assessing UAV sensor spoofing: More than a GNSS problem. In Proceedings of Annual Computer Security Applications Conference, Honolulu, USA, pp. 1032\u20131046, 2024. DOI: https:\/\/doi.org\/10.1109\/ACSAC63791.2024.00085."},{"key":"1626_CR82","doi-asserted-by":"publisher","unstructured":"A. M. Klein, J. Deutschlander, K. Kolln, M. Rauschenberger, M. J. Escalona. Exploring the context of use for voice user interfaces: Toward context-dependent user experience quality testing. Journal of Software: Evolution and Process, vol. 36, no. 7, Article number e2618, 2024. DOI: https:\/\/doi.org\/10.1002\/smr.2618.","DOI":"10.1002\/smr.2618"},{"key":"1626_CR83","doi-asserted-by":"publisher","first-page":"196","DOI":"10.1007\/978-3-031-73113-6_12","volume-title":"Proceedings of the 18th European Conference on Computer Vision","author":"J Zhang","year":"2025","unstructured":"J. Zhang, T. Wang, H. Zhang, P. Lu, F. Zheng. Reflective instruction tuning: Mitigating hallucinations in large vision-language models. In Proceedings of the 18th European Conference on Computer Vision, Milan, Italy, pp. 196\u2013213, 2025. DOI: https:\/\/doi.org\/10.1007\/978-3-031-73113-6_12."},{"key":"1626_CR84","doi-asserted-by":"publisher","first-page":"18135","DOI":"10.1609\/aaai.v38i16.29771","volume-title":"Proceedings of the 38th AAAI Conference on Artificial Intelligence","author":"A Gunjal","year":"2024","unstructured":"A. Gunjal, J. Yin, E. Bas. Detecting and preventing hallucinations in large vision language models. In Proceedings of the 38th AAAI Conference on Artificial Intelligence, Vancouver, Canada, pp. 18135\u201318143, 2024. DOI: https:\/\/doi.org\/10.1609\/aaai.v38i16.29771."},{"key":"1626_CR85","doi-asserted-by":"publisher","first-page":"13872","DOI":"10.1109\/CVPR52733.2024.01316","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"S Leng","year":"2024","unstructured":"S. Leng, H. Zhang, G. Chen, X. Li, S. Lu, C. Miao, L. Bing. Mitigating object hallucinations in large vision-language models through visual contrastive decoding. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 13872\u201313882, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.01316."},{"key":"1626_CR86","doi-asserted-by":"publisher","first-page":"13258","DOI":"10.18653\/v1\/2024.findings-emnlp.775","volume-title":"Proceedings of Findings of the Association for Computational Linguistics: EMNLP","author":"Y Xie","year":"2024","unstructured":"Y. Xie, G. Li, X. Xu, M. Y. Kan. V-DPO: Mitigating hallucination in large vision language models via vision-guided direct preference optimization. In Proceedings of Findings of the Association for Computational Linguistics: EMNLP, Miami, USA, pp. 13258\u201313273, 2024. DOI: https:\/\/doi.org\/10.18653\/v1\/2024.findings-emnlp.775."},{"key":"1626_CR87","volume-title":"Proceedings of the 13th International Conference on Learning Representations","author":"T Yang","year":"2025","unstructured":"T. Yang, Z. Li, J. Cao, C. Xu. Understanding and mitigating hallucination in large vision-language models via modular attribution and intervention. In Proceedings of the 13th International Conference on Learning Representations, Singapore, 2025."},{"key":"1626_CR88","volume-title":"Proceedings of the 13th International Conference on Learning Representations","author":"F Huo","year":"2025","unstructured":"F. Huo, W. Xu, Z. Zhang, H. Wang, Z. Chen, P. Zhao. Self-introspective decoding: Alleviating hallucinations for large vision-language models. In Proceedings of the 13th International Conference on Learning Representations, Singapore, Singapore, 2025."},{"key":"1626_CR89","doi-asserted-by":"publisher","first-page":"18684","DOI":"10.1609\/aaai.v39i18.34056","volume-title":"Proceedings of the 39th AAAI Conference on Artificial Intelligence","author":"T Liang","year":"2025","unstructured":"T. Liang, Y. Du, J. Huang, M. Kong, L. Chen, Y. Li, S. Chen, Q. Zhu. MoLE: Decoding by mixture of layer experts alleviates hallucination in large vision-language models. In Proceedings of the 39th AAAI Conference on Artificial Intelligence, Philadelphia, USA, pp. 18684\u201318692, 2025. DOI: https:\/\/doi.org\/10.1609\/aaai.v39i18.34056."},{"key":"1626_CR90","doi-asserted-by":"publisher","first-page":"6904","DOI":"10.1109\/CVPR52729.2023.00667","volume-title":"Proceedings of IEEE\/ CVF Conference on Computer Vision and Pattern Recognition","author":"R Shao","year":"2023","unstructured":"R. Shao, T. Wu, Z. Liu. Detecting and grounding multimodal media manipulation. In Proceedings of IEEE\/ CVF Conference on Computer Vision and Pattern Recognition, Vancouver, Canada, pp. 6904\u20136913, 2023. DOI: https:\/\/doi.org\/10.1109\/CVPR52729.2023.00667."},{"issue":"8","key":"1626_CR91","doi-asserted-by":"publisher","first-page":"5556","DOI":"10.1109\/TPAMI.2024.3367749","volume":"46","author":"R Shao","year":"2024","unstructured":"R. Shao, T. Wu, J. Wu, L. Nie, Z. Liu. Detecting and grounding multi-modal media manipulation and beyond. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 8, pp. 5556\u20135574, 2024. DOI: https:\/\/doi.org\/10.1109\/TPAMI.2024.3367749.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"2","key":"1626_CR92","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1145\/3715073.3715079","volume":"26","author":"Y Cui","year":"2024","unstructured":"Y. Cui, J. Ren, H. Xu, P. He, H. Liu, L. Sun, Y. Xing, J. Tang. DiffusionShield: A watermark for data copyright protection against generative diffusion models. ACM SIGKDD Explorations Newsletter, vol. 26, no. 2, pp. 60\u201375, 2024. DOI: https:\/\/doi.org\/10.1145\/3715073.3715079.","journal-title":"ACM SIGKDD Explorations Newsletter"},{"key":"1626_CR93","doi-asserted-by":"publisher","first-page":"4628","DOI":"10.1109\/TIFS.2024.3383648","volume":"19","author":"Y Zhang","year":"2024","unstructured":"Y. Zhang, D. Ye, C. Xie, L. Tang, X. Liao, Z. Liu, C. Chen, J. Deng. Dual defense: Adversarial, traceable, and invisible robust watermarking against face swapping. IEEE Transactions on Information Forensics and Security, vol. 19, pp. 4628\u20134641, 2024. DOI: https:\/\/doi.org\/10.1109\/TIFS.2024.3383648.","journal-title":"IEEE Transactions on Information Forensics and Security"},{"key":"1626_CR94","doi-asserted-by":"publisher","first-page":"11964","DOI":"10.1109\/CVPR52733.2024.01137","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"X Zhang","year":"2024","unstructured":"X. Zhang, R. Li, J. Yu, Y. Xu, W. Li, J. Zhang. Edit-Guard: Versatile image watermarking for tamper localization and copyright protection. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 11964\u201311974, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.01137."},{"key":"1626_CR95","doi-asserted-by":"publisher","first-page":"27092","DOI":"10.1109\/CVPR52733.2024.02559","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"T Oorloff","year":"2024","unstructured":"T. Oorloff, S. Koppisetti, N. Bonettini, D. Solanki, B. Colman, Y. Yacoob, A. Shahriyari, G. Bharaj. AVFF: Audio-visual feature fusion for video deepfake detection. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 27092\u201327102, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.02559."},{"key":"1626_CR96","doi-asserted-by":"publisher","first-page":"5000","DOI":"10.1109\/CVPR42600.2020.00505","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"L Li","year":"2020","unstructured":"L. Li, J. Bao, T. Zhang, H. Yang, D. Chen, F. Wen, B. Guo. Face X-ray for more general face forgery detection. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 5000\u20135009, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00505."},{"key":"1626_CR97","doi-asserted-by":"publisher","first-page":"17395","DOI":"10.1109\/CVPR52733.2024.01647","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"D Nguyen","year":"2024","unstructured":"D. Nguyen, N. Mejri, I. P. Singh, P. Kuleshova, M. Astrid, A. Kacem, E. Ghorbel, D. Aouada. LAA-Net: Localized artifact attention network for quality-agnostic and generalizable deepfake detection. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 17395\u201317405, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.01647."},{"key":"1626_CR98","doi-asserted-by":"publisher","first-page":"180","DOI":"10.1007\/978-3-031-78341-8_12","volume-title":"Proceedings of the 27th International Conference Pattern Recognition","author":"Y Hou","year":"2025","unstructured":"Y. Hou, H. Fu, C. Chen, Z. Li, H. Zhang, J. Zhao. PolyGlotFake: A novel multilingual and multimodal deepfake dataset. In Proceedings of the 27th International Conference Pattern Recognition, Kolkata, India, pp. 180\u2013193, 2025. DOI: https:\/\/doi.org\/10.1007\/978-3-031-78341-8_12."},{"key":"1626_CR99","doi-asserted-by":"publisher","unstructured":"S. Yang, H. Guo, S. Hu, B. Zhu, Y. Fu, S. Lyu, X. Wu, X. Wang. CrossDF: Improving cross-domain deepfake detection with deep information decomposition. Frontiers in Big Data, vol. 8, Article number 1669488, 2023. DOI: https:\/\/doi.org\/10.3389\/fdata.2025.1669488.","DOI":"10.3389\/fdata.2025.1669488"},{"issue":"3","key":"1626_CR100","doi-asserted-by":"publisher","first-page":"209","DOI":"10.1007\/s11633-022-1330-7","volume":"19","author":"M Ren","year":"2022","unstructured":"M. Ren, Y. L. Wang, Z. F. He. Towards interpretable defense against adversarial attacks via causal inference. Machine Intelligence Research, vol. 19, no. 3, pp. 209\u2013226, 2022. DOI: https:\/\/doi.org\/10.1007\/s11633-022-1330-7.","journal-title":"Machine Intelligence Research"},{"key":"1626_CR101","doi-asserted-by":"publisher","first-page":"24408","DOI":"10.1109\/CVPR52733.2024.02304","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"L Li","year":"2024","unstructured":"L. Li, H. Guan, J. Qiu, M. Spratling. One prompt word is enough to boost adversarial robustness for pre-trained vision-language models. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 24408\u201324419, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.02304."},{"key":"1626_CR102","volume-title":"Proceedings of the 38th International Conference on Neural Information Processing Systems","author":"Y Zhou","year":"2024","unstructured":"Y. Zhou, X. Xia, Z. Lin, B. Han, T. Liu. Few-shot adversarial prompt learning on vision-language models. In Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, Article number 102, 2024."},{"key":"1626_CR103","volume-title":"Proceedings of the 13th International Conference on Learning Representations","author":"H Huang","year":"2025","unstructured":"H. Huang, S. M. Erfani, Y. Li, X. Ma, J. Bailey. Detecting backdoor samples in contrastive language image pretraining. In Proceedings of the 13th International Conference on Learning Representations, Singapore, 2025."},{"key":"1626_CR104","doi-asserted-by":"publisher","first-page":"77","DOI":"10.1007\/978-3-031-72661-3_5","volume-title":"Proceedings of the 18th European Conference on Computer Vision","author":"Y Wang","year":"2025","unstructured":"Y. Wang, X. Liu, Y. Li, M. Chen, C. Xiao. AdaShield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting. In Proceedings of the 18th European Conference on Computer Vision, Milan, Italy, pp. 77\u201394, 2025. DOI: https:\/\/doi.org\/10.1007\/978-3-031-72661-3_5."},{"key":"1626_CR105","doi-asserted-by":"publisher","first-page":"25038","DOI":"10.1109\/CVPR52734.2025.02331","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"S S Ghosal","year":"2025","unstructured":"S. S. Ghosal, S. Chakraborty, V. Singh, T. Guan, M. Wang, A. Beirami, F. Huang, A. Velasquez, D. Manocha, A. S. Bedi. Immune: Improving safety against jailbreaks in multi-modal LLMs via inference-time alignment. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, pp. 25038\u201325049, 2025. DOI: https:\/\/doi.org\/10.1109\/CVPR52734.2025.02331."},{"issue":"5","key":"1626_CR106","doi-asserted-by":"publisher","first-page":"456","DOI":"10.1007\/s11633-022-1375-7","volume":"19","author":"K Y Liu","year":"2022","unstructured":"K. Y. Liu, X. Y. Li, Y. R. Lai, H. Su, J. C. Wang, C. X. Guo, H. Xie, J. S. Guan, Y. Zhou. Denoised internal models: A brain-inspired autoencoder against adversarial attacks. Machine Intelligence Research, vol. 19, no. 5, pp. 456\u2013471, 2022. DOI: https:\/\/doi.org\/10.1007\/s11633-022-1375-7.","journal-title":"Machine Intelligence Research"},{"key":"1626_CR107","doi-asserted-by":"publisher","first-page":"265","DOI":"10.1007\/11681878_14","volume-title":"Proceedings of the 3rd Theory of Cryptography Conference","author":"C Dwork","year":"2006","unstructured":"C. Dwork, F. McSherry, K. Nissim, A. Smith. Calibrating noise to sensitivity in private data analysis. In Proceedings of the 3rd Theory of Cryptography Conference, New York, USA, pp. 265\u2013284, 2006. DOI: https:\/\/doi.org\/10.1007\/11681878_14."},{"key":"1626_CR108","doi-asserted-by":"publisher","first-page":"2190","DOI":"10.1109\/SP46215.2023.10179466","volume-title":"Proceedings of the 44th IEEE Symposium on Security and Privacy","author":"W Dong","year":"2023","unstructured":"W. Dong, Q. Luo, K. Yi. Continual observation under user-level differential privacy. In Proceedings of the 44th IEEE Symposium on Security and Privacy, San Francisco, USA, pp. 2190\u20132207, 2023. DOI: https:\/\/doi.org\/10.1109\/SP46215.2023.10179466."},{"key":"1626_CR109","first-page":"26048","volume-title":"Proceedings of the 39th International Conference on Machine Learning","author":"X Zhang","year":"2022","unstructured":"X. Zhang, X. Chen, M. Hong, S. Wu, J. Yi. Understanding clipping for federated learning: Convergence and client-level differential privacy. In Proceedings of the 39th International Conference on Machine Learning, Baltimore, USA, pp. 26048\u201326067, 2022."},{"issue":"1","key":"1626_CR110","doi-asserted-by":"publisher","first-page":"169","DOI":"10.56553\/popets-2025-0010","volume":"2025","author":"S Das","year":"2025","unstructured":"S. Das, S. R. Chowdhury, N. Chandran, D. Gupta, S. Lokam, R. Sharma. Communication efficient secure and private multi-party deep learning. Proceedings on Privacy Enhancing Technologies, vol. 2025, no. 1, pp. 169\u2013183, 2025. DOI: https:\/\/doi.org\/10.56553\/popets-2025-0010.","journal-title":"Proceedings on Privacy Enhancing Technologies"},{"key":"1626_CR111","doi-asserted-by":"publisher","first-page":"542","DOI":"10.1109\/SP54263.2024.00128","volume-title":"Proceedings of IEEE Symposium on Security and Privacy","author":"B Karmakar","year":"2024","unstructured":"B. Karmakar, N. Koti, A. Patra, S. Patranabis, P. Paul, D. Ravi. Asterisk: Super-fast MPC with a friend. In Proceedings of IEEE Symposium on Security and Privacy, San Francisco, USA, pp. 542\u2013560, 2024. DOI: https:\/\/doi.org\/10.1109\/SP54263.2024.00128."},{"key":"1626_CR112","doi-asserted-by":"publisher","DOI":"10.1109\/SP54263.2024.00166","volume-title":"Proceedings of IEEE Symposium on Security and Privacy","author":"Y Lu","year":"2024","unstructured":"Y. Lu, M. Magdon-Ismail, Y. Wei, V. Zikas. Eureka: A general framework for black-box differential privacy estimators. In Proceedings of IEEE Symposium on Security and Privacy, San Francisco, USA, pp. 913931, 2024. DOI: https:\/\/doi.org\/10.1109\/SP54263.2024.00166."},{"key":"1626_CR113","doi-asserted-by":"publisher","first-page":"932","DOI":"10.1109\/SP54263.2024.00088","volume-title":"Proceedings of IEEE Symposium on Security and Privacy","author":"N Ashena","year":"2024","unstructured":"N. Ashena, O. Inel, B. L. Persaud, A. Bernstein. Casual users and rational choices within differential privacy. In Proceedings of IEEE Symposium on Security and Privacy, San Francisco, USA, pp. 932\u2013950, 2024. DOI: https:\/\/doi.org\/10.1109\/SP54263.2024.00088."},{"issue":"9","key":"1626_CR114","doi-asserted-by":"publisher","first-page":"1594","DOI":"10.1109\/TSMC.2017.2681698","volume":"48","author":"H Sedjelmaci","year":"2018","unstructured":"H. Sedjelmaci, S. M. Senouci, N. Ansari. A hierarchical detection and response system to enhance security against lethal cyber-attacks in UAV networks. IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 48, no. 9, pp. 1594\u20131606, 2018. DOI: https:\/\/doi.org\/10.1109\/TSMC.2017.2681698.","journal-title":"IEEE Transactions on Systems, Man, and Cybernetics: Systems"},{"issue":"5","key":"1626_CR115","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1109\/MWC.001.1900028","volume":"26","author":"X Sun","year":"2019","unstructured":"X. Sun, D. W. K. Ng, Z. Ding, Y. Xu, Z. Zhong. Physical layer security in UAV systems: Challenges and opportunities. IEEE Wireless Communications, vol. 26, no. 5, pp. 40\u201347, 2019. DOI: https:\/\/doi.org\/10.1109\/MWC.001.1900028.","journal-title":"IEEE Wireless Communications"},{"issue":"6","key":"1626_CR116","doi-asserted-by":"publisher","first-page":"171","DOI":"10.1109\/MNET.011.2000101","volume":"34","author":"X Chen","year":"2020","unstructured":"X. Chen, D. Li, Z. Yang, Y. Chen, N. Zhao, Z. Ding, F. R. Yu. Securing aerial-ground transmission for NOMAUAV networks. IEEE Network, vol. 34, no. 6, pp. 171\u2013177, 2020. DOI: https:\/\/doi.org\/10.1109\/MNET.011.2000101.","journal-title":"IEEE Network"},{"key":"1626_CR117","doi-asserted-by":"publisher","first-page":"1963","DOI":"10.1109\/TIFS.2023.3318942","volume":"19","author":"Y Wang","year":"2024","unstructured":"Y. Wang, Z. Su, A. Benslimane, Q. Xu, M. Dai, R. Li. Collaborative honeypot defense in UAV networks: A learning-based game approach. IEEE Transactions on Information Forensics and Security, vol. 19, pp. 1963\u20131978, 2024. DOI: https:\/\/doi.org\/10.1109\/TIFS.2023.3318942.","journal-title":"IEEE Transactions on Information Forensics and Security"},{"issue":"4","key":"1626_CR118","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1109\/MWC.01.1900545","volume":"27","author":"A S Abdalla","year":"2020","unstructured":"A. S. Abdalla, K. Powell, V. Marojevic, G. Geraci. UAV-assisted attack prevention, detection, and recovery of 5G networks. IEEE Wireless Communications, vol. 27, no. 4, pp. 40\u201347, 2020. DOI: https:\/\/doi.org\/10.1109\/MWC.01.1900545.","journal-title":"IEEE Wireless Communications"},{"issue":"3","key":"1626_CR119","doi-asserted-by":"publisher","first-page":"3170","DOI":"10.1109\/TNSM.2021.3061486","volume":"18","author":"O Fonseca","year":"2021","unstructured":"O. Fonseca, \u00cd. Cunha, E. Fazzion, W. Meira, B. A. da Silva, R. A. Ferreira, E. Katz-Bassett. Identifying networks vulnerable to IP spoofing. IEEE Transactions on Network and Service Management, vol. 18, no. 3, pp. 3170\u20133183, 2021. DOI: https:\/\/doi.org\/10.1109\/TNSM.2021.3061486.","journal-title":"IEEE Transactions on Network and Service Management"},{"key":"1626_CR120","volume-title":"Proceedings of the 38th International Conference on Neural Information Processing Systems","author":"L Liu","year":"2024","unstructured":"L. Liu, D. Yang, S. Zhong, K. S. S. Tholeti, L. Ding, Y. Zhang, L. H. Gilpin. Right this way: Can VLMs guide us to see more to answer questions? In Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, Article number 4226, 2024."},{"key":"1626_CR121","doi-asserted-by":"publisher","first-page":"625","DOI":"10.1109\/WACV57701.2024.00069","volume-title":"Proceedings of IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"G Zhang","year":"2024","unstructured":"G. Zhang, Y. Zhang, K. Zhang, V. Tresp. Can vision-language models be a good guesser? Exploring VLMs for times and location reasoning. In Proceedings of IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, USA, pp. 625\u2013634, 2024. DOI: https:\/\/doi.org\/10.1109\/WACV57701.2024.00069."},{"key":"1626_CR122","doi-asserted-by":"publisher","first-page":"73","DOI":"10.1007\/978-3-031-72624-8_5","volume-title":"Proceedings of the 18th European Conference on Computer Vision","author":"Y Alaluf","year":"2025","unstructured":"Y. Alaluf, E. Richardson, S. Tulyakov, K. Aberman, D. Cohen-Or. MyVLM: Personalizing VLMs for user-specific queries. In Proceedings of the 18th European Conference on Computer Vision, Milan, Italy, pp. 73\u201391, 2025. DOI: https:\/\/doi.org\/10.1007\/978-3-031-72624-8_5."},{"key":"1626_CR123","doi-asserted-by":"publisher","first-page":"9620","DOI":"10.1109\/IROS58592.2024.10801576","volume-title":"Proceedings of IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"Z Wang","year":"2024","unstructured":"Z. Wang, Q. Liu, J. Qin, M. Li. Ensuring safety in LLM-driven robotics: A cross-layer sequence supervision mechanism. In Proceedings of IEEE\/RSJ International Conference on Intelligent Robots and Systems, Abu Dhabi, UAE, pp. 9620\u20139627, 2024. DOI: https:\/\/doi.org\/10.1109\/IROS58592.2024.10801576."},{"key":"1626_CR124","volume-title":"Proceedings of the 13th International Conference on Learning Representations","author":"W Chow","year":"2025","unstructured":"W. Chow, J. Mao, B. Li, D. Seita, V. C. Guizilini, Y. Wang. Physbench: Benchmarking and enhancing vision-language models for physical world understanding. In Proceedings of the 13th International Conference on Learning Representations, Singapore, 2025."},{"issue":"9","key":"1626_CR125","doi-asserted-by":"publisher","first-page":"389","DOI":"10.1038\/s42256-019-0088-2","volume":"1","author":"A Jobin","year":"2019","unstructured":"A. Jobin, M. Ienca, E. Vayena. The global landscape of AI ethics guidelines. Nature Machine Intelligence, vol. 1, no. 9, pp. 389\u2013399, 2019. DOI: https:\/\/doi.org\/10.1038\/s42256-019-0088-2.","journal-title":"Nature Machine Intelligence"},{"key":"1626_CR126","doi-asserted-by":"publisher","first-page":"197","DOI":"10.1109\/IHTC.2017.8058187","volume-title":"Proceedings of IEEE Canada International Humanitarian Technology Conference","author":"K Shahriari","year":"2017","unstructured":"K. Shahriari, M. Shahriari. IEEE standard review\u2013Ethically aligned design: A vision for prioritizing human wellbeing with artificial intelligence and autonomous systems. In Proceedings of IEEE Canada International Humanitarian Technology Conference, Toronto, Canada, pp. 197\u2013201, 2017. DOI: https:\/\/doi.org\/10.1109\/IHTC.2017.8058187."},{"key":"1626_CR127","volume-title":"High-Level Expert Group on Artificial Intelligence. Ethics Guidelines for Trustworthy AI","author":"European Commission","year":"2019","unstructured":"European Commission. High-Level Expert Group on Artificial Intelligence. Ethics Guidelines for Trustworthy AI, European Commission, Brussels, Belgium, 2019."},{"key":"1626_CR128","unstructured":"Z. Liu, Y. Nie, Y. Tan, X. Yue, Q. Cui, C. J. Wang, X. Zhu, B. Zheng. Enhancing vision-language model safety through progressive concept-bottleneck-driven alignment, [Online], Available: https:\/\/arxiv.org\/abs\/2411.11543v1, 2024."},{"key":"1626_CR129","first-page":"77","volume-title":"Proceedings of Conference on Fairness, Accountability and Transparency","author":"J Buolamwini","year":"2018","unstructured":"J. Buolamwini, T. Gebru. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Proceedings of Conference on Fairness, Accountability and Transparency, New York, USA, pp. 77\u201391, 2018."},{"issue":"12","key":"1626_CR130","doi-asserted-by":"publisher","first-page":"1435","DOI":"10.1038\/s42256-024-00926-3","volume":"6","author":"A A Trotsyuk","year":"2024","unstructured":"A. A. Trotsyuk, Q. Waeiss, R. T. Bhatia, B. J. Aponte, I. M. L. Heffernan, D. Madgavkar, R. M. Felder, L. S. Lehmann, M. J. Palmer, H. Greely, R. Wald, L. Goetz, M. Trengove, R. Vandersluis, H. Lin, M. K. Cho, R. B. Altman, D. Endy, D. A. Relman, M. Levi, D. Satz, D. Magnus. Toward a framework for risk mitigation of potential misuse of artificial intelligence in biomedical research. Nature Machine Intelligence, vol. 6, no. 12, pp. 1435\u20131442, 2024. DOI: https:\/\/doi.org\/10.1038\/s42256-024-00926-3.","journal-title":"Nature Machine Intelligence"},{"key":"1626_CR131","first-page":"513","volume-title":"Proceedings of the 25th USENIX Conference on Security Symposium","author":"N Carlini","year":"2016","unstructured":"N. Carlini, P. Mishra, T. Vaidya, Y. Zhang, M. Sherr, C. Shields, D. Wagner, W. Zhou. Hidden voice commands. In Proceedings of the 25th USENIX Conference on Security Symposium, Austin, USA, pp. 513\u2013530, 2016."},{"key":"1626_CR132","volume-title":"Strategic Planning for Public and Nonprofit Organizations: A Guide to Strengthening and Sustaining Organizational Achievement","author":"J M Bryson","year":"2018","unstructured":"J. M. Bryson. Strategic Planning for Public and Nonprofit Organizations: A Guide to Strengthening and Sustaining Organizational Achievement, 5th ed., Hoboken, USA: Wiley, 2018.","edition":"5th ed."},{"key":"1626_CR133","doi-asserted-by":"publisher","first-page":"59","DOI":"10.1016\/j.nbt.2024.12.003","volume":"85","author":"A Holzinger","year":"2025","unstructured":"A. Holzinger, K. Zatloukal, H. M\u00fcller. Is human oversight to AI systems still possible? New Biotechnology, vol. 85, pp. 59\u201362, 2025. DOI: https:\/\/doi.org\/10.1016\/j.nbt.2024.12.003.","journal-title":"New Biotechnology"},{"key":"1626_CR134","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1109\/SP.2017.41","volume-title":"Proceedings of IEEE Symposium on Security and Privacy","author":"R Shokri","year":"2017","unstructured":"R. Shokri, M. Stronati, C. Song, V. Shmatikov. Membership inference attacks against machine learning models. In Proceedings of IEEE Symposium on Security and Privacy, San Jose, USA, pp. 3\u201318, 2017. DOI: https:\/\/doi.org\/10.1109\/SP.2017.41."},{"key":"1626_CR135","doi-asserted-by":"publisher","DOI":"10.6028\/NIST.CSWP.01162020","volume-title":"Nist Privacy Framework: A Tool for Improving Privacy Through Enterprise Risk Management, Version 1.0","author":"K R Boeckl","year":"2020","unstructured":"K. R. Boeckl, N. B. Lefkovitz. Nist Privacy Framework: A Tool for Improving Privacy Through Enterprise Risk Management, Version 1.0, National Institute of Standards and Technology, Gaithersburg, USA, 2020. DOI: https:\/\/doi.org\/10.6028\/NIST.CSWP.01162020."},{"key":"1626_CR136","volume-title":"Proceedings of the 3rd International Conference on Learning Representations","author":"I J Goodfellow","year":"2015","unstructured":"I. J. Goodfellow, J. Shlens, C. Szegedy. Explaining and harnessing adversarial examples. In Proceedings of the 3rd International Conference on Learning Representations, San Diego, USA, 2015."},{"key":"1626_CR137","doi-asserted-by":"publisher","first-page":"220","DOI":"10.1145\/3287560.3287596","volume-title":"Proceedings of Conference on Fairness, Accountability, and Transparency","author":"M Mitchell","year":"2019","unstructured":"M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, T. Gebru. Model cards for model reporting. In Proceedings of Conference on Fairness, Accountability, and Transparency, Atlanta, USA, pp. 220\u2013229, 2019. DOI: https:\/\/doi.org\/10.1145\/3287560.3287596."},{"key":"1626_CR138","volume-title":"International Scientific Report on the Safety of Advanced AI: Interim Report","author":"Y Bengio","year":"2024","unstructured":"Y. Bengio. International Scientific Report on the Safety of Advanced AI: Interim Report, OGL, 2024."},{"key":"1626_CR139","volume-title":"Recommendation on the Ethics of Artificial Intelligence","author":"UNESCO","year":"2021","unstructured":"UNESCO. Recommendation on the Ethics of Artificial Intelligence, UNESCO, 2021."},{"key":"1626_CR140","doi-asserted-by":"publisher","first-page":"3645","DOI":"10.18653\/v1\/P19-1355","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"E Strubell","year":"2019","unstructured":"E. Strubell, A. Ganesh, A. McCallum. Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, pp. 3645\u20133650, 2019. DOI: https:\/\/doi.org\/10.18653\/v1\/P19-1355."},{"key":"1626_CR141","unstructured":"D. Patterson, J. Gonzalez, Q. Le, C. Liang, L. M. Munguia, D. Rothchild, D. So, M. Texier, J. Dean. Carbon emissions and large neural network training, [Online], Available: https:\/\/arxiv.org\/abs\/2104.10350, 2021."},{"key":"1626_CR142","volume-title":"Energy and AI: Executive Summary","author":"IEA","year":"2025","unstructured":"IEA. Energy and AI: Executive Summary, IEA, Paris, France, 2025."},{"key":"1626_CR143","unstructured":"N. C. Thompson, K. Greenewald, K. Lee, G. F. Manso. The computational limits of deep learning, [Online], Available: https:\/\/arxiv.org\/abs\/2007.05558, 2020."},{"key":"1626_CR144","unstructured":"Is AI closing the door on entry-level job opportunities? World Economic Forum. WEF Stories, 2025."},{"key":"1626_CR145","unstructured":"International Monetary Fund. AI Will Transform the Global Economy. Let\u2019s Make Sure it Benefits Humanity, IMF Blog, 2024."},{"issue":"5","key":"1626_CR146","doi-asserted-by":"publisher","first-page":"1973","DOI":"10.3982\/ECTA19815","volume":"90","author":"D Acemoglu","year":"2022","unstructured":"D. Acemoglu, P. Restrepo. Tasks, automation, and the rise in U.S. wage inequality. Econometrica, vol. 90, no. 5, pp. 1973\u20132016, 2022. DOI: https:\/\/doi.org\/10.3982\/ECTA19815.","journal-title":"Econometrica"},{"issue":"1","key":"1626_CR147","doi-asserted-by":"publisher","first-page":"333","DOI":"10.1257\/mac.20180386","volume":"13","author":"E Brynjolfsson","year":"2021","unstructured":"E. Brynjolfsson, D. Rock, C. Syverson. The productivity J-curve: How intangibles complement general purpose technologies. American Economic Journal: Macroeconomics, vol. 13, no. 1, pp. 333\u2013372, 2021. DOI: https:\/\/doi.org\/10.1257\/mac.20180386.","journal-title":"American Economic Journal: Macroeconomics"},{"key":"1626_CR148","doi-asserted-by":"publisher","first-page":"610","DOI":"10.1145\/3442188.3445922","volume-title":"Proceedings of ACM Conference on Fairness, Accountability, and Transparency","author":"E M Bender","year":"2021","unstructured":"E. M. Bender, T. Gebru, A. McMillan-Major, S. Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of ACM Conference on Fairness, Accountability, and Transparency, Toronto, Canada, pp. 610\u2013623, 2021. DOI: https:\/\/doi.org\/10.1145\/3442188.3445922."},{"key":"1626_CR149","volume-title":"Face Recognition Vendor Test (FRVT) Part 3: Demographic Effects NISTIR 8280","author":"P Grother","year":"2019","unstructured":"P. Grother, M. Ngan, K. Hanaoka. Face Recognition Vendor Test (FRVT) Part 3: Demographic Effects NISTIR 8280, NIST, 2019."},{"issue":"1","key":"1626_CR150","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1007\/s11023-020-09517-8","volume":"30","author":"T Hagendorff","year":"2020","unstructured":"T. Hagendorff. The ethics of AI ethics: An evaluation of guidelines. Minds and Machines, vol. 30, no. 1, pp. 99\u2013120, 2020. DOI: https:\/\/doi.org\/10.1007\/s11023-020-09517-8.","journal-title":"Minds and Machines"},{"key":"1626_CR151","unstructured":"A. Birhane, O. Guest. Towards decolonising computational sciences, [Online], Available: https:\/\/arxiv.org\/abs\/2009.14258, 2020."}],"container-title":["Machine Intelligence Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11633-025-1626-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11633-025-1626-x","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11633-025-1626-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T08:03:04Z","timestamp":1784793784000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11633-025-1626-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,13]]},"references-count":151,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,8]]}},"alternative-id":["1626"],"URL":"https:\/\/doi.org\/10.1007\/s11633-025-1626-x","relation":{},"ISSN":["2731-538X","2731-5398"],"issn-type":[{"value":"2731-538X","type":"print"},{"value":"2731-5398","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,13]]},"assertion":[{"value":"30 June 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 December 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 July 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declared that they have no conflicts of interest to this work.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations of conflict of interest"}}]}}