{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T19:20:52Z","timestamp":1765308052783,"version":"3.46.0"},"publisher-location":"New York, NY, USA","reference-count":29,"publisher":"ACM","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,10,27]]},"DOI":"10.1145\/3746027.3755003","type":"proceedings-article","created":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T05:47:42Z","timestamp":1761371262000},"page":"3173-3181","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Deciphering Functions of Neurons in Vision-Language Models"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-4081-736X","authenticated-orcid":false,"given":"Jiaqi","family":"Xu","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, Hefei, Anhui, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9145-9957","authenticated-orcid":false,"given":"Cuiling","family":"Lan","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5383-6424","authenticated-orcid":false,"given":"Yan","family":"Lu","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,10,27]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966","author":"Bai Jinze","year":"2023","unstructured":"Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966 (2023)."},{"key":"e_1_3_2_1_2_1","volume-title":"Interpreting Neurons in Deep Vision Networks with Language Models. arXiv preprint arXiv:2403.13771","author":"Bai Nicholas","year":"2025","unstructured":"Nicholas Bai, Rahul A. Iyer, Tuomas Oikarinen, Akshay Kulkarni, and Tsui-Wei Weng. 2025b. Interpreting Neurons in Deep Vision Networks with Language Models. arXiv preprint arXiv:2403.13771 (2025)."},{"key":"e_1_3_2_1_3_1","unstructured":"Shuai Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge Sibo Song Kai Dang Peng Wang Shijie Wang Jun Tang et al. 2025a. Qwen2.5-vl technical report. arXiv preprint arXiv:2502.13923 (2025)."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.354"},{"key":"e_1_3_2_1_5_1","volume-title":"Mechanistic Interpretability for AI Safety-A Review. arXiv preprint arXiv:2404.14082","author":"Bereska Leonard","year":"2024","unstructured":"Leonard Bereska and Efstratios Gavves. 2024. Mechanistic Interpretability for AI Safety-A Review. arXiv preprint arXiv:2404.14082 (2024)."},{"key":"e_1_3_2_1_6_1","volume-title":"Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, et al.","author":"Beyer Lucas","year":"2024","unstructured":"Lucas Beyer, Andreas Steiner, Andr\u00e9 Susano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, et al., 2024. Paligemma: A versatile 3b vlm for transfer. arXiv preprint arXiv:2407.07726 (2024)."},{"key":"e_1_3_2_1_7_1","volume-title":"Language models can explain neurons in language models. URL https:\/\/openaipublic. blob. core. windows. net\/neuron-explainer\/paper\/index. html.(Date accessed: 14.05","author":"Bills Steven","year":"2023","unstructured":"Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. 2023. Language models can explain neurons in language models. URL https:\/\/openaipublic. blob. core. windows. net\/neuron-explainer\/paper\/index. html.(Date accessed: 14.05. 2023), Vol. 2 (2023)."},{"key":"e_1_3_2_1_8_1","first-page":"24804","article-title":"Labeling neural representations with inverse recognition","volume":"36","author":"Bykov Kirill","year":"2023","unstructured":"Kirill Bykov, Laura Kopf, Shinichi Nakajima, Marius Kloft, and Marina H\u00f6hne. 2023. Labeling neural representations with inverse recognition. Advances in Neural Information Processing Systems, Vol. 36 (2023), 24804-24828.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_9_1","unstructured":"Zhe Chen Weiyun Wang Yue Cao Yangzhou Liu Zhangwei Gao Erfei Cui Jinguo Zhu Shenglong Ye Hao Tian Zhaoyang Liu et al. 2024. Expanding performance boundaries of open-source multimodal models with model data and test-time scaling. arXiv preprint arXiv:2412.05271 (2024)."},{"key":"e_1_3_2_1_10_1","first-page":"122867","article-title":"Towards neuron attributions in multi-modal large language models","volume":"37","author":"Fang Junfeng","year":"2024","unstructured":"Junfeng Fang, Zac Bi, Ruipeng Wang, Houcheng Jiang, Yuan Gao, Kun Wang, An Zhang, Jie Shi, Xiang Wang, and Tat-Seng Chua. 2024. Towards neuron attributions in multi-modal large language models. Advances in Neural Information Processing Systems, Vol. 37 (2024), 122867-122890.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_11_1","volume-title":"Anh Tuan Luu, and Cheston Tan","author":"Guertler Leon","year":"2024","unstructured":"Leon Guertler, M Ganesh Kumar, Anh Tuan Luu, and Cheston Tan. 2024. TeLLMe what you see: Using LLMs to Explain Neurons in Vision Models. (2024)."},{"key":"e_1_3_2_1_12_1","volume-title":"International Conference on Learning Representations.","author":"Hernandez Evan","year":"2021","unstructured":"Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas. 2021. Natural language descriptions of deep visual features. In International Conference on Learning Representations."},{"key":"e_1_3_2_1_13_1","volume-title":"International conference on machine learning. PMLR","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning. PMLR, 19730-19742."},{"volume-title":"Microsoft coco: Common objects in context","author":"Lin Tsung-Yi","key":"e_1_3_2_1_14_1","unstructured":"Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In ECCV. Springer, 740-755."},{"key":"e_1_3_2_1_15_1","unstructured":"Haotian Liu Chunyuan Li Yuheng Li and Yong Jae Lee. 2023a. Improved Baselines with Visual Instruction Tuning."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02484"},{"key":"e_1_3_2_1_17_1","unstructured":"Haotian Liu Chunyuan Li Qingyang Wu and Yong Jae Lee. 2023b. Visual Instruction Tuning. In NeurIPS."},{"key":"e_1_3_2_1_18_1","volume-title":"Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian.","author":"Naveed Humza","year":"2023","unstructured":"Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian. 2023. A comprehensive overview of large language models. arXiv preprint arXiv:2307.06435 (2023)."},{"key":"e_1_3_2_1_19_1","volume-title":"Clip-dissect: Automatic description of neuron representations in deep vision networks. arXiv preprint arXiv:2204.10965","author":"Oikarinen Tuomas","year":"2022","unstructured":"Tuomas Oikarinen and Tsui-Wei Weng. 2022. Clip-dissect: Automatic description of neuron representations in deep vision networks. arXiv preprint arXiv:2204.10965 (2022)."},{"key":"e_1_3_2_1_20_1","volume-title":"Linear explanations for individual neurons. arXiv preprint arXiv:2405.06855","author":"Oikarinen Tuomas","year":"2024","unstructured":"Tuomas Oikarinen and Tsui-Wei Weng. 2024. Linear explanations for individual neurons. arXiv preprint arXiv:2405.06855 (2024)."},{"key":"e_1_3_2_1_21_1","volume-title":"Finding and editing multi-modal neurons in pre-trained transformer. arXiv preprint arXiv:2311.07470","author":"Pan Haowen","year":"2023","unstructured":"Haowen Pan, Yixin Cao, Xiaozhi Wang, and Xun Yang. 2023. Finding and editing multi-modal neurons in pre-trained transformer. arXiv preprint arXiv:2311.07470 (2023)."},{"volume-title":"Notes on regression and inheritance in the case of two parents proceedings of the royal society of London","author":"Pearson K","key":"e_1_3_2_1_22_1","unstructured":"K Pearson. 1895. Notes on regression and inheritance in the case of two parents proceedings of the royal society of London, Vol. 58."},{"key":"e_1_3_2_1_23_1","volume-title":"A practical review of mechanistic interpretability for transformer-based language models. arXiv preprint arXiv:2407.02646","author":"Rai Daking","year":"2024","unstructured":"Daking Rai, Yilun Zhou, Shi Feng, Abulhair Saparov, and Ziyu Yao. 2024. A practical review of mechanistic interpretability for transformer-based language models. arXiv preprint arXiv:2407.02646 (2024)."},{"key":"e_1_3_2_1_24_1","unstructured":"Tianhe Ren Shilong Liu Ailing Zeng Jing Lin Kunchang Li He Cao Jiayu Chen Xinyu Huang Yukang Chen Feng Yan et al. 2024. Grounded sam: Assembling open-world models for diverse visual tasks. arXiv preprint arXiv:2401.14159 (2024)."},{"key":"e_1_3_2_1_25_1","volume-title":"Forty-first International Conference on Machine Learning.","author":"Shaham Tamar Rott","year":"2024","unstructured":"Tamar Rott Shaham, Sarah Schwettmann, Franklin Wang, Achyuta Rajaram, Evan Hernandez, Jacob Andreas, and Antonio Torralba. 2024. A multimodal automated interpretability agent. In Forty-first International Conference on Machine Learning."},{"key":"e_1_3_2_1_26_1","volume-title":"Michel Galley, Rich Caruana, and Jianfeng Gao.","author":"Singh Chandan","year":"2024","unstructured":"Chandan Singh, Jeevana Priya Inala, Michel Galley, Rich Caruana, and Jianfeng Gao. 2024. Rethinking interpretability in the era of large language models. arXiv preprint arXiv:2402.01761 (2024)."},{"key":"e_1_3_2_1_27_1","volume-title":"Yaniv Gurwicz, Matthew Lyle Olson, Anahita Bhiwandiwalla, Estelle Aflalo, Chenfei Wu, Nan Duan, Shao-Yen Tseng, and Vasudev Lal.","author":"Melech Stan Gabriela Ben","year":"2024","unstructured":"Gabriela Ben Melech Stan, Raanan Yehezkel Rohekar, Yaniv Gurwicz, Matthew Lyle Olson, Anahita Bhiwandiwalla, Estelle Aflalo, Chenfei Wu, Nan Duan, Shao-Yen Tseng, and Vasudev Lal. 2024. LVLM-Intrepret: An Interpretability Tool for Large Vision-Language Models. arXiv preprint arXiv:2404.03118 (2024)."},{"key":"e_1_3_2_1_28_1","volume-title":"Language-specific neurons: The key to multilingual capabilities in large language models. arXiv preprint arXiv:2402.16438","author":"Tang Tianyi","year":"2024","unstructured":"Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, and Ji-Rong Wen. 2024. Language-specific neurons: The key to multilingual capabilities in large language models. arXiv preprint arXiv:2402.16438 (2024)."},{"key":"e_1_3_2_1_29_1","volume-title":"Explainable AI: A brief survey on history, research areas, approaches and challenges. In Natural language processing and Chinese computing","author":"Xu Feiyu","year":"2019","unstructured":"Feiyu Xu, Hans Uszkoreit, Yangzhou Du, Wei Fan, Dongyan Zhao, and Jun Zhu. 2019. Explainable AI: A brief survey on history, research areas, approaches and challenges. In Natural language processing and Chinese computing. Springer, 563-574."}],"event":{"name":"MM '25: The 33rd ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Dublin Ireland","acronym":"MM '25"},"container-title":["Proceedings of the 33rd ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746027.3755003","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T19:18:04Z","timestamp":1765307884000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746027.3755003"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,27]]},"references-count":29,"alternative-id":["10.1145\/3746027.3755003","10.1145\/3746027"],"URL":"https:\/\/doi.org\/10.1145\/3746027.3755003","relation":{},"subject":[],"published":{"date-parts":[[2025,10,27]]},"assertion":[{"value":"2025-10-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}