{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,10]],"date-time":"2025-12-10T04:15:24Z","timestamp":1765340124506,"version":"3.46.0"},"publisher-location":"New York, NY, USA","reference-count":51,"publisher":"ACM","funder":[{"name":"This work is funded by Science and Technology Commission of Shanghai Municipality Program, China","award":["No.24DZ2202100"],"award-info":[{"award-number":["No.24DZ2202100"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,10,27]]},"DOI":"10.1145\/3746027.3754574","type":"proceedings-article","created":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T06:47:18Z","timestamp":1761374838000},"page":"11199-11208","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0452-7593","authenticated-orcid":false,"given":"Wanying","family":"Wang","sequence":"first","affiliation":[{"name":"Shanghai Key Laboratory of Computer Software Testing and Evaluating, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-5154-0392","authenticated-orcid":false,"given":"Zeyu","family":"Ma","sequence":"additional","affiliation":[{"name":"Shanghai Key Laboratory of Computer Software Testing and Evaluating, Shanghai, China and Shanghai Normal University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6668-8145","authenticated-orcid":false,"given":"Han","family":"Zheng","sequence":"additional","affiliation":[{"name":"TrustAI Pte. Ltd., Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9346-1196","authenticated-orcid":false,"given":"Xin","family":"Tan","sequence":"additional","affiliation":[{"name":"East China Normal University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0718-8187","authenticated-orcid":false,"given":"Mingang","family":"Chen","sequence":"additional","affiliation":[{"name":"Shanghai Key Laboratory of Computer Software Testing and Evaluating, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,10,27]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence. 3058-3068","author":"Baechler Gilles","year":"2024","unstructured":"Gilles Baechler, Srinivas Sunkara, Maria Wang, Fedir Zubach, Hassan Mansoor, Vincent Etter, Victor C\u0103rbune, Jason Lin, Jindong Chen, and Abhanshu Sharma. 2024. Screenai: A vision-language model for ui and infographics understanding. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence. 3058-3068."},{"key":"e_1_3_2_1_2_1","unstructured":"Shuai Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge Sibo Song Kai Dang Peng Wang Shijie Wang Jun Tang Humen Zhong Yuanzhi Zhu Mingkun Yang Zhaohai Li Jianqiang Wan Pengfei Wang Wei Ding Zheren Fu Yiheng Xu Jiabo Ye Xi Zhang Tianbao Xie Zesen Cheng Hang Zhang Zhibo Yang Haiyang Xu and Junyang Lin. 2025. Qwen2.5-VL Technical Report. arXiv:2502.13923 [cs.CV] https:\/\/arxiv.org\/abs\/2502.13923"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-0237"},{"key":"e_1_3_2_1_4_1","volume-title":"Language Models are Few-Shot Learners. arXiv: Computation and Language,arXiv: Computation and Language (May","author":"Brown T.B.","year":"2020","unstructured":"T.B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Askell Amanda, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Henighan Tom, Rewon Child, A. Ramesh, DanielM. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, EricJ. Sigler, Mateusz Litwin, Scott Gray, Chess Benjamin, Jack Clark, Christopher Berner, McCandlish Sam, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. arXiv: Computation and Language,arXiv: Computation and Language (May 2020)."},{"key":"e_1_3_2_1_5_1","unstructured":"Trishna Chakraborty Erfan Shayegani Zikui Cai Nael Abu-Ghazaleh M. Salman Asif Yue Dong Amit K. Roy-Chowdhury and Chengyu Song. 2024. Cross-Modal Safety Alignment: Is textual unlearning all you need? arXiv:2406.02575 [cs.CL] https:\/\/arxiv.org\/abs\/2406.02575"},{"key":"e_1_3_2_1_6_1","unstructured":"Jun Chen Deyao Zhu Xiaoqian Shen Xiang Li Zechun Liu Pengchuan Zhang Raghuraman Krishnamoorthi Vikas Chandra Yunyang Xiong and Mohamed Elhoseiny. 2023. MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning. arXiv:2310.09478 [cs.CV] https:\/\/arxiv.org\/abs\/2310.09478"},{"key":"e_1_3_2_1_7_1","volume-title":"ICLR 2024 Workshop on Secure and Trustworthy Large Language Models. https:\/\/openreview.net\/forum?id=WubY1GeLij","author":"Chen Shuo","year":"2024","unstructured":"Shuo Chen, Zhen Han, Bailan He, Zifeng Ding, Wenqian Yu, Philip Torr, Volker Tresp, and Jindong Gu. 2024. Red Teaming GPT-4V: Are GPT-4V Safe Against Uni\/Multi-Modal Jailbreak Attacks?. In ICLR 2024 Workshop on Secure and Trustworthy Large Language Models. https:\/\/openreview.net\/forum?id=WubY1GeLij"},{"key":"e_1_3_2_1_8_1","volume-title":"The Twelfth International Conference on Learning Representations.","author":"Chuang Yung-Sung","year":"2024","unstructured":"Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R. Glass, and Pengcheng He. 2024. DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models. In The Twelfth International Conference on Learning Representations."},{"key":"e_1_3_2_1_9_1","volume-title":"The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=QoDDNkx4fP","author":"Ding Yi","year":"2025","unstructured":"Yi Ding, Bolian Li, and Ruqi Zhang. 2025. ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=QoDDNkx4fP"},{"key":"e_1_3_2_1_10_1","volume-title":"Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment. https:\/\/arxiv.org\/abs\/2411.18688","author":"Ghosal Soumya Suvra","year":"2024","unstructured":"Soumya Suvra Ghosal, Souradip Chakraborty, Vaibhav Singh, Tianrui Guan, Mengdi Wang, Ahmad Beirami, Furong Huang, Alvaro Velasquez, Dinesh Manocha, and Amrit Singh Bedi. 2024. Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment. https:\/\/arxiv.org\/abs\/2411.18688"},{"key":"e_1_3_2_1_11_1","volume-title":"Figstep: Jailbreaking large vision-language models via typographic visual prompts. arXiv preprint arXiv:2311.05608","author":"Gong Yichen","year":"2023","unstructured":"Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, and Xiaoyun Wang. 2023. Figstep: Jailbreaking large vision-language models via typographic visual prompts. arXiv preprint arXiv:2311.05608 (2023)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-72643-9_23"},{"key":"e_1_3_2_1_13_1","volume-title":"Tracing Representation Progression: Analyzing and Enhancing Layer-Wise Similarity. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=vVxeFSR4fU","author":"Jiang Jiachen","year":"2025","unstructured":"Jiachen Jiang, Jinxin Zhou, and Zhihui Zhu. 2025. Tracing Representation Progression: Analyzing and Enhancing Layer-Wise Similarity. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=vVxeFSR4fU"},{"key":"e_1_3_2_1_14_1","volume-title":"Proceedings of the 40th International Conference on Machine Learning. 19730-19742","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning. 19730-19742."},{"key":"e_1_3_2_1_15_1","volume-title":"International conference on machine learning. PMLR, 12888-12900","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning. PMLR, 12888-12900."},{"key":"e_1_3_2_1_16_1","volume-title":"Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update. arXiv preprint arXiv:2501.16378","author":"Li Qing","year":"2025","unstructured":"Qing Li, Jiahui Geng, Zongxiong Chen, Kun Song, Lei Ma, and Fakhri Karray. 2025a. Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update. arXiv preprint arXiv:2501.16378 (2025)."},{"key":"e_1_3_2_1_17_1","volume-title":"The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=kUH1yPMAn7","author":"Li Shen","year":"2025","unstructured":"Shen Li, Liuyi Yao, Lan Zhang, and Yaliang Li. 2025b. Safety Layers in Aligned Large Language Models: The Key to LLM Security. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=kUH1yPMAn7"},{"key":"e_1_3_2_1_18_1","volume-title":"The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=kUH1yPMAn7","author":"Li Shen","year":"2025","unstructured":"Shen Li, Liuyi Yao, Lan Zhang, and Yaliang Li. 2025c. Safety Layers in Aligned Large Language Models: The Key to LLM Security. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=kUH1yPMAn7"},{"key":"e_1_3_2_1_19_1","volume-title":"European Conference on Computer Vision. Springer, 174-189","author":"Li Yifan","year":"2024","unstructured":"Yifan Li, Hangyu Guo, Kun Zhou, Wayne Xin Zhao, and Ji-Rong Wen. 2024. Images are achilles' heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models. In European Conference on Computer Vision. Springer, 174-189."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_1_21_1","unstructured":"Haotian Liu Chunyuan Li Qingyang Wu and Yong Jae Lee. 2023. Visual Instruction Tuning. arXiv:2304.08485 [cs.CV] https:\/\/arxiv.org\/abs\/2304.08485"},{"key":"e_1_3_2_1_22_1","volume-title":"MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models. In European Conference on Computer Vision. 386-403","author":"Liu Xin","year":"2025","unstructured":"Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, and Chao Yang. 2025. MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models. In European Conference on Computer Vision. 386-403."},{"key":"e_1_3_2_1_23_1","unstructured":"Liming Lu Shuchao Pang Siyuan Liang Haotian Zhu Xiyu Zeng Aishan Liu Yunhuai Liu and Yongbin Zhou. 2025. Adversarial Training for Multimodal Large Language Models against Jailbreak Attacks. arXiv:2503.04833 [cs.CV] https:\/\/arxiv.org\/abs\/2503.04833"},{"key":"e_1_3_2_1_24_1","first-page":"2507","article-title":"Learn to explain: Multimodal reasoning via thought chains for science question answering","volume":"35","author":"Lu Pan","year":"2022","unstructured":"Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, Vol. 35 (2022), 2507-2521.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2410.07149"},{"key":"e_1_3_2_1_26_1","volume-title":"Jailbreaking attack against multimodal large language model. arXiv preprint arXiv:2402.02309","author":"Niu Zhenxing","year":"2024","unstructured":"Zhenxing Niu, Haodong Ren, Xinbo Gao, Gang Hua, and Rong Jin. 2024. Jailbreaking attack against multimodal large language model. arXiv preprint arXiv:2402.02309 (2024)."},{"key":"e_1_3_2_1_27_1","unstructured":"Long Ouyang Jeffrey Wu Xu Jiang Diogo Almeida Carroll Wainwright Pamela Mishkin Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems Vol. 35 (2022) 27730-27744."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW60793.2023.00307"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.60"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.895"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i19.30150"},{"key":"e_1_3_2_1_32_1","unstructured":"Qwen: An Yang Baosong Yang Beichen Zhang Binyuan Hui Bo Zheng Bowen Yu Chengyuan Li Dayiheng Liu Fei Huang Haoran Wei Huan Lin Jian Yang Jianhong Tu Jianwei Zhang Jianxin Yang Jiaxi Yang Jingren Zhou Junyang Lin Kai Dang Keming Lu Keqin Bao Kexin Yang Le Yu Mei Li Mingfeng Xue Pei Zhang Qin Zhu Rui Men Runji Lin Tianhao Li Tianyi Tang Tingyu Xia Xingzhang Ren Xuancheng Ren Yang Fan Yang Su Yichang Zhang Yu Wan Yuqiong Liu Zeyu Cui Zhenru Zhang and Zihan Qiu. 2025. Qwen2.5 Technical Report. arXiv:2412.15115 [cs.CL] https:\/\/arxiv.org\/abs\/2412.15115"},{"key":"e_1_3_2_1_33_1","volume-title":"International conference on machine learning. PmLR, 8748-8763","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PmLR, 8748-8763."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW60793.2023.00308"},{"key":"e_1_3_2_1_35_1","volume-title":"The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=plmBsXHxgR","author":"Shayegani Erfan","year":"2024","unstructured":"Erfan Shayegani, Yue Dong, and Nael Abu-Ghazaleh. 2024. Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=plmBsXHxgR"},{"key":"e_1_3_2_1_36_1","volume-title":"The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=s20W12XTF8","author":"Shen Guobin","year":"2025","unstructured":"Guobin Shen, Dongcheng Zhao, Yiting Dong, Xiang He, and Yi Zeng. 2025. Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=s20W12XTF8"},{"key":"e_1_3_2_1_37_1","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale Dan Bikel Lukas Blecher Cristian Canton Ferrer Moya Chen Guillem Cucurull David Esiobu Jude Fernandes Jeremy Fu Wenyin Fu Brian Fuller Cynthia Gao Vedanuj Goswami Naman Goyal Anthony Hartshorn Saghar Hosseini Rui Hou Hakan Inan Marcin Kardas Viktor Kerkez Madian Khabsa Isabel Kloumann Artem Korenev Punit Singh Koura Marie-Anne Lachaux Thibaut Lavril Jenya Lee Diana Liskovich Yinghai Lu Yuning Mao Xavier Martinet Todor Mihaylov Pushkar Mishra Igor Molybog Yixin Nie Andrew Poulton Jeremy Reizenstein Rashi Rungta Kalyan Saladi Alan Schelten Ruan Silva Eric Michael Smith Ranjan Subramanian Xiaoqing Ellen Tan Binh Tang Ross Taylor Adina Williams Jian Xiang Kuan Puxin Xu Zheng Yan Iliyan Zarov Yuchen Zhang Angela Fan Melanie Kambadur Sharan Narang Aurelien Rodriguez Robert Stojnic Sergey Edunov and Thomas Scialom. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv:2307.09288 [cs.CL] https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_2_1_38_1","volume-title":"European Conference on Computer Vision. Springer, 77-94","author":"Wang Yu","year":"2024","unstructured":"Yu Wang, Xiaogeng Liu, Yu Li, Muhao Chen, and Chaowei Xiao. 2024. Adashield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting. In European Conference on Computer Vision. Springer, 77-94."},{"key":"e_1_3_2_1_39_1","volume-title":"Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=45rvZkJbuX","author":"Xu Shicheng","year":"2025","unstructured":"Shicheng Xu, Liang Pang, Yunchang Zhu, Huawei Shen, and Xueqi Cheng. 2025. Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=45rvZkJbuX"},{"key":"e_1_3_2_1_40_1","volume-title":"Defending jailbreak attack in vlms via cross-modality information detector. arXiv e-prints","author":"Xu Yue","year":"2024","unstructured":"Yue Xu, Xiuyuan Qi, Zhan Qin, and Wenjie Wang. 2024. Defending jailbreak attack in vlms via cross-modality information detector. arXiv e-prints (2024), arXiv-2407."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-023-0393-x"},{"key":"e_1_3_2_1_42_1","volume-title":"MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities. In Forty-first International Conference on Machine Learning. https:\/\/openreview.net\/forum?id=KOTutrSR2y","author":"Yu Weihao","year":"2024","unstructured":"Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. 2024. MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities. In Forty-first International Conference on Machine Learning. https:\/\/openreview.net\/forum?id=KOTutrSR2y"},{"key":"e_1_3_2_1_43_1","unstructured":"Zhi Zhang Srishti Yadav Fengze Han and Ekaterina Shutova. 2024. Cross-modal Information Flow in Multimodal Large Language Models. arXiv:2411.18620 [cs.AI] https:\/\/arxiv.org\/abs\/2411.18620"},{"key":"e_1_3_2_1_44_1","volume-title":"Computer Vision - ECCV","author":"Zhao Qinyu","year":"2024","unstructured":"Qinyu Zhao, Ming Xu, Kartik Gupta, Akshay Asthana, Liang Zheng, and Stephen Gould. 2025b. The First to Know: How Token Distributions Reveal Hidden Knowledge in\u00a0Large Vision-Language Models?. In Computer Vision - ECCV 2024, Ale\u0161 Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and G\u00fcl Varol (Eds.). Springer Nature Switzerland, Cham, 127-142."},{"key":"e_1_3_2_1_45_1","volume-title":"Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency. arXiv preprint arXiv:2501.04931","author":"Zhao Shiji","year":"2025","unstructured":"Shiji Zhao, Ranjie Duan, Fengxiang Wang, Chi Chen, Caixin Kang, Jialing Tao, YueFeng Chen, Hui Xue, and Xingxing Wei. 2025a. Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency. arXiv preprint arXiv:2501.04931 (2025)."},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.847"},{"key":"e_1_3_2_1_47_1","volume-title":"Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models. arXiv preprint arXiv:2501.02029","author":"Zheng Ziwei","year":"2025","unstructured":"Ziwei Zheng, Junyao Zhao, Le Yang, Lijun He, and Fan Li. 2025. Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models. arXiv preprint arXiv:2501.02029 (2025)."},{"key":"e_1_3_2_1_48_1","volume-title":"On the Role of Attention Heads in Large Language Model Safety. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=h0Ak8A5yqw","author":"Zhou Zhenhong","year":"2025","unstructured":"Zhenhong Zhou, Haiyang Yu, Xinghua Zhang, Rongwu Xu, Fei Huang, Kun Wang, Yang Liu, Junfeng Fang, and Yongbin Li. 2025. On the Role of Attention Heads in Large Language Model Safety. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=h0Ak8A5yqw"},{"key":"e_1_3_2_1_49_1","volume-title":"The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=1tZbq88f27","author":"Zhu Deyao","year":"2024","unstructured":"Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2024. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=1tZbq88f27"},{"key":"e_1_3_2_1_50_1","volume-title":"Proceedings of the 41st International Conference on Machine Learning. 62867-62891","author":"Zong Yongshuo","year":"2024","unstructured":"Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, and Timothy Hospedales. 2024. Safety fine-tuning at (almost) no cost: a baseline for vision large language models. In Proceedings of the 41st International Conference on Machine Learning. 62867-62891."},{"key":"e_1_3_2_1_51_1","unstructured":"Andy Zou Long Phan Sarah Chen James Campbell Phillip Guo Richard Ren Alexander Pan Xuwang Yin Mantas Mazeika Ann-Kathrin Dombrowski et al. 2023. Representation engineering: A top-down approach to ai transparency. arXiv:2310.01405"}],"event":{"name":"MM '25: The 33rd ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Dublin Ireland","acronym":"MM '25"},"container-title":["Proceedings of the 33rd ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746027.3754574","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,10]],"date-time":"2025-12-10T04:12:21Z","timestamp":1765339941000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746027.3754574"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,27]]},"references-count":51,"alternative-id":["10.1145\/3746027.3754574","10.1145\/3746027"],"URL":"https:\/\/doi.org\/10.1145\/3746027.3754574","relation":{},"subject":[],"published":{"date-parts":[[2025,10,27]]},"assertion":[{"value":"2025-10-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}