{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T14:56:04Z","timestamp":1781535364458,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":28,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T00:00:00Z","timestamp":1781481600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,6,16]]},"DOI":"10.1145\/3805622.3810676","type":"proceedings-article","created":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T14:42:57Z","timestamp":1781534577000},"page":"2714-2722","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Assessing Color Vision Test in Large Vision-language Models"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-5108-0876","authenticated-orcid":false,"given":"Hongfei","family":"Ye","sequence":"first","affiliation":[{"name":"University of Chinese Academy of Sciences, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-6400-2084","authenticated-orcid":false,"given":"Bin","family":"Chen","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2250-8069","authenticated-orcid":false,"given":"Wenxi","family":"Liu","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-1360-8309","authenticated-orcid":false,"given":"Yu","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5056-0351","authenticated-orcid":false,"given":"Zhao","family":"Li","sequence":"additional","affiliation":[{"name":"Zhejiang Lab, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-6781-8072","authenticated-orcid":false,"given":"Dandan","family":"Ni","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7626-0162","authenticated-orcid":false,"given":"Hongyang","family":"Chen","sequence":"additional","affiliation":[{"name":"Zhejiang Lab, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,15]]},"reference":[{"key":"e_1_3_3_1_2_2","unstructured":"Jinze Bai Shuai Bai Shusheng Yang Shijie Wang Sinan Tan Peng Wang Junyang Lin Chang Zhou and Jingren Zhou. 2023. Qwen-VL: A Versatile Vision-Language Model for Understanding Localization Text Reading and Beyond. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2308.12966 (2023)."},{"key":"e_1_3_3_1_3_2","doi-asserted-by":"crossref","unstructured":"Marc\u00a0G Berman Michael\u00a0C Hout Omid Kardan MaryCarol\u00a0R Hunter Grigori Yourganov John\u00a0M Henderson Taylor Hanayik Hossein Karimi and John Jonides. 2014. The perception of naturalness correlates with low-level visual features of environmental scenes. PloS one 9 12 (2014) e114572.","DOI":"10.1371\/journal.pone.0114572"},{"key":"e_1_3_3_1_4_2","doi-asserted-by":"crossref","unstructured":"Manojit Bhattacharya Soumen Pal Srijan Chatterjee Sang-Soo Lee and Chiranjib Chakraborty. 2024. Large language model to multimodal large language model: A journey to shape the biological macromolecules to biological sciences and medicine. Molecular Therapy-Nucleic Acids 35 3 (2024).","DOI":"10.1016\/j.omtn.2024.102255"},{"key":"e_1_3_3_1_5_2","doi-asserted-by":"crossref","unstructured":"Jennifer Birch. 1997. Efficiency of the Ishihara test for identifying red-green colour deficiency. Ophthalmic and Physiological Optics 17 5 (1997) 403\u2013408.","DOI":"10.1111\/j.1475-1313.1997.tb00072.x"},{"key":"e_1_3_3_1_6_2","unstructured":"Lin Chen Jinsong Li Xiaoyi Dong Pan Zhang Yuhang Zang Zehui Chen Haodong Duan Jiaqi Wang Yu Qiao Dahua Lin et\u00a0al. 2024. Are We on the Right Way for Evaluating Large Vision-Language Models? arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2403.20330 (2024)."},{"key":"e_1_3_3_1_7_2","unstructured":"Xiaokang Chen Zhiyu Wu Xingchao Liu Zizheng Pan Wen Liu Zhenda Xie Xingkai Yu and Chong Ruan. 2025. Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2501.17811 (2025)."},{"key":"e_1_3_3_1_8_2","unstructured":"Zhe Chen Weiyun Wang Yue Cao Yangzhou Liu Zhangwei Gao Erfei Cui Jinguo Zhu Shenglong Ye Hao Tian Zhaoyang Liu et\u00a0al. 2024. Expanding performance boundaries of open-source multimodal models with model data and test-time scaling. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2412.05271 (2024)."},{"key":"e_1_3_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACVW60836.2024.00106"},{"key":"e_1_3_3_1_10_2","unstructured":"Team GLM Aohan Zeng Bin Xu Bowen Wang Chenhui Zhang Da Yin Dan Zhang Diego Rojas Guanyu Feng Hanlin Zhao et\u00a0al. 2024. Chatglm: A family of large language models from glm-130b to glm-4 all tools. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2406.12793 (2024)."},{"key":"e_1_3_3_1_11_2","unstructured":"Aaron Hurst Adam Lerer Adam\u00a0P Goucher Adam Perelman Aditya Ramesh Aidan Clark AJ Ostrow Akila Welihinda Alan Hayes Alec Radford et\u00a0al. 2024. Gpt-4o system card. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.21276 (2024)."},{"key":"e_1_3_3_1_12_2","unstructured":"Bohao Li Rui Wang Guangzhi Wang Yuying Ge Yixiao Ge and Ying Shan. 2023. Seed-bench: Benchmarking multimodal llms with generative comprehension. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2307.16125 (2023)."},{"key":"e_1_3_3_1_13_2","unstructured":"Guanzhen Li Yuxi Xie and Min-Yen Kan. 2024. MVP-Bench: Can Large Vision\u2013Language Models Conduct Multi-level Visual Perception Like Humans? arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.04345 (2024)."},{"key":"e_1_3_3_1_14_2","unstructured":"Haotian Liu Chunyuan Li Qingyang Wu and Yong\u00a0Jae Lee. 2024. Visual instruction tuning. Advances in neural information processing systems 36 (2024)."},{"key":"e_1_3_3_1_15_2","first-page":"216","volume-title":"European conference on computer vision","author":"Liu Yuan","year":"2024","unstructured":"Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et\u00a0al. 2024. Mmbench: Is your multi-modal model an all-around player?. In European conference on computer vision. Springer, 216\u2013233."},{"key":"e_1_3_3_1_16_2","doi-asserted-by":"crossref","unstructured":"Alex Melamud Stephanie Hagstrom and Elias Traboulsi. 2004. Color vision testing. Ophthalmic Genetics 25 3 (2004) 159\u2013187.","DOI":"10.1080\/13816810490498341"},{"key":"e_1_3_3_1_17_2","unstructured":"OpenAI. 2023. GPT-4V(ision) System Card. https:\/\/api.semanticscholar.org\/CorpusID:263218031"},{"key":"e_1_3_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-89862-5_374"},{"key":"e_1_3_3_1_19_2","doi-asserted-by":"crossref","unstructured":"Romke Rouw Stephen\u00a0M Kosslyn and Ronald Hamel. 1997. Detecting high-level and low-level properties in visual images and visual percepts. Cognition 63 2 (1997) 209\u2013226.","DOI":"10.1016\/S0010-0277(97)00006-1"},{"key":"e_1_3_3_1_20_2","unstructured":"Ahnaf\u00a0Mozib Samin M\u00a0Firoz Ahmed and Md\u00a0Mushtaq\u00a0Shahriyar Rafee. 2024. ColorFoil: Investigating Color Blindness in Large Vision and Language Models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2405.11685 (2024)."},{"key":"e_1_3_3_1_21_2","doi-asserted-by":"crossref","unstructured":"WS Stiles. 1959. Color vision: the approach through increment-threshold sensitivity.","DOI":"10.1073\/pnas.45.1.100"},{"key":"e_1_3_3_1_22_2","unstructured":"Gemini Team Petko Georgiev Ving\u00a0Ian Lei Ryan Burnell Libin Bai Anmol Gulati Garrett Tanzer Damien Vincent Zhufeng Pan Shibo Wang et\u00a0al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2403.05530 (2024)."},{"key":"e_1_3_3_1_23_2","unstructured":"Qwen Team. 2025. Qwen2.5-VL. https:\/\/qwenlm.github.io\/blog\/qwen2.5-vl\/"},{"key":"e_1_3_3_1_24_2","unstructured":"Qwen Team. 2026. Qwen3. 5: Towards native multimodal agents. URL: https:\/\/qwen. ai\/blog (2026)."},{"key":"e_1_3_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.4324\/9780203082263"},{"key":"e_1_3_3_1_26_2","doi-asserted-by":"crossref","unstructured":"Jinge Wang Qing Ye Li Liu Nancy\u00a0Lan Guo and Gangqing Hu. 2024. Scientific figures interpreted by ChatGPT: strengths in plot recognition and limits in color perception. NPJ Precision Oncology 8 1 (2024) 84.","DOI":"10.1038\/s41698-024-00576-z"},{"key":"e_1_3_3_1_27_2","unstructured":"Zhenhua Xu Yujia Zhang Enze Xie Zhen Zhao Yong Guo Kwan-Yee\u00a0K Wong Zhenguo Li and Hengshuang Zhao. 2024. Drivegpt4: Interpretable end-to-end autonomous driving via large language model. IEEE Robotics and Automation Letters (2024)."},{"key":"e_1_3_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00913"},{"key":"e_1_3_3_1_29_2","first-page":"169","volume-title":"European Conference on Computer Vision","author":"Zhang Renrui","year":"2024","unstructured":"Renrui Zhang, Dongzhi Jiang, Yichi Zhang, Haokun Lin, Ziyu Guo, Pengshuo Qiu, Aojun Zhou, Pan Lu, Kai-Wei Chang, Yu Qiao, et\u00a0al. 2024. Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?. In European Conference on Computer Vision. Springer, 169\u2013186."}],"event":{"name":"ICMR '26: International Conference on Multimedia Retrieval","location":"Amsterdam The Netherlands","acronym":"ICMR '26","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 2026 International Conference on Multimedia Retrieval"],"original-title":[],"deposited":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T14:46:51Z","timestamp":1781534811000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3805622.3810676"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,15]]},"references-count":28,"alternative-id":["10.1145\/3805622.3810676","10.1145\/3805622"],"URL":"https:\/\/doi.org\/10.1145\/3805622.3810676","relation":{},"subject":[],"published":{"date-parts":[[2026,6,15]]},"assertion":[{"value":"2026-06-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}