{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T04:21:43Z","timestamp":1784002903223,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":70,"publisher":"ACM","license":[{"start":{"date-parts":[[2025,4,25]],"date-time":"2025-04-25T00:00:00Z","timestamp":1745539200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,4,26]]},"DOI":"10.1145\/3706598.3713274","type":"proceedings-article","created":{"date-parts":[[2025,4,24]],"date-time":"2025-04-24T04:45:58Z","timestamp":1745469958000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["ArtMentor: AI-Assisted Evaluation of Artworks to Explore Multimodal Large Language Models Capabilities"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1232-0020","authenticated-orcid":false,"given":"Chanjin","family":"Zheng","sequence":"first","affiliation":[{"name":"Shanghai Institute of Artificial Intelligence for Education, East China Normal University, Shanghai, China and Faculty of Education, East China Normal University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-4571-340X","authenticated-orcid":false,"given":"Zengyi","family":"Yu","sequence":"additional","affiliation":[{"name":"Faculty of Education, East China Normal University, Shanghai, China and College of Education, Zhejiang University of Technology, Hangzhou, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-3179-6969","authenticated-orcid":false,"given":"Yilin","family":"Jiang","sequence":"additional","affiliation":[{"name":"College of Education, Zheiiang University of Technology, Hangzhou, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-8539-5559","authenticated-orcid":false,"given":"Mingzi","family":"Zhang","sequence":"additional","affiliation":[{"name":"Faculty of Education, East China Normal University, Shanghai, China and College of Education, Zhejiang Normal University, Jinhua, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-9174-393X","authenticated-orcid":false,"given":"Xunuo","family":"Lu","sequence":"additional","affiliation":[{"name":"School of Economy, Zhejiang University of Technology, Hangzhou, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5986-3363","authenticated-orcid":false,"given":"Jing","family":"Jin","sequence":"additional","affiliation":[{"name":"School of Education, Zhejiang Normal University, Jinhua, Zhejiang, China and Tianchang Guanchao Primary School, Hangzhou, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-4363-2030","authenticated-orcid":false,"given":"Liteng","family":"Gao","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence Science and Technology, University of Shanghai for Science and Technology, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,4,25]]},"reference":[{"key":"e_1_3_3_3_2_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia\u00a0Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et\u00a0al. 2023. Gpt-4 technical report. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2303.08774 (2023)."},{"key":"e_1_3_3_3_3_2","doi-asserted-by":"crossref","unstructured":"John\u00a0R Anderson Albert\u00a0T Corbett Kenneth\u00a0R Koedinger and Ray Pelletier. 1995. Cognitive tutors: Lessons learned. The journal of the learning sciences 4 2 (1995) 167\u2013207.","DOI":"10.1207\/s15327809jls0402_2"},{"key":"e_1_3_3_3_4_2","unstructured":"Anas Awadalla Irena Gao Josh Gardner Jack Hessel Yusuf Hanafy Wanrong Zhu Kalyani Marathe Yonatan Bitton Samir Gadre Shiori Sagawa Jenia Jitsev Simon Kornblith Pang\u00a0Wei Koh Gabriel Ilharco Mitchell Wortsman and Ludwig Schmidt. 2023. OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2308.01390 (2023)."},{"key":"e_1_3_3_3_5_2","unstructured":"Jinze Bai Shuai Bai Shusheng Yang Shijie Wang Sinan Tan Peng Wang Junyang Lin Chang Zhou and Jingren Zhou. 2023. Qwen-vl: A versatile vision-language model for understanding localization text reading and beyond. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2308.12966 3 1 (2023)."},{"key":"e_1_3_3_3_6_2","unstructured":"Shuai Bai Shusheng Yang Jinze Bai Peng Wang Xingxuan Zhang Junyang Lin Xinggang Wang Chang Zhou and Jingren Zhou. 2023. Touchstone: Evaluating vision-language models by language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2308.16890 (2023)."},{"key":"e_1_3_3_3_7_2","doi-asserted-by":"crossref","unstructured":"Siwar Bengamra Olfa Mzoughi Andr\u00e9 Bigand and Ezzeddine Zagrouba. 2024. A comprehensive survey on object detection in Visual Art: taxonomy and challenge. Multimedia Tools and Applications 83 5 (2024) 14637\u201314670.","DOI":"10.1007\/s11042-023-15968-9"},{"key":"e_1_3_3_3_8_2","volume-title":"Essai sur les donn\u00e9es imm\u00e9diates de la conscience","author":"Bergson Henri","year":"1911","unstructured":"Henri Bergson. 1911. Essai sur les donn\u00e9es imm\u00e9diates de la conscience. F. Alcan."},{"key":"e_1_3_3_3_9_2","unstructured":"Anjanava Biswas and Wrick Talukdar. 2024. Robustness of Structured Data Extraction from In-Plane Rotated Documents Using Multi-Modal Large Language Models (LLM). Journal of Artificial Intelligence Research (2024)."},{"key":"e_1_3_3_3_10_2","doi-asserted-by":"crossref","unstructured":"Moinak Biswas. 2021. Realism. BioScope: South Asian Screen Studies 12 1-2 (2021) 158\u2013161.","DOI":"10.1177\/09749276211026084"},{"key":"e_1_3_3_3_11_2","doi-asserted-by":"crossref","unstructured":"Ann\u00a0E Blandford Philip\u00a0J Barnard and Michael\u00a0D Harrison. 1995. Using Interaction Framework to guide the design of interactive systems. International journal of human-computer studies 43 1 (1995) 101\u2013130.","DOI":"10.1006\/ijhc.1995.1037"},{"key":"e_1_3_3_3_12_2","doi-asserted-by":"crossref","unstructured":"Eva Cetinic and James She. 2022. Understanding and creating art with AI: Review and outlook. ACM Transactions on Multimedia Computing Communications and Applications (TOMM) 18 2 (2022) 1\u201322.","DOI":"10.1145\/3475799"},{"key":"e_1_3_3_3_13_2","unstructured":"Jiaxing Chen Yuxuan Liu Dehu Li Xiang An Ziyong Feng Yongle Zhao and Yin Xie. 2024. Plug-and-play grounding of reasoning in multimodal large language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2403.19322 (2024)."},{"key":"e_1_3_3_3_14_2","unstructured":"Pengcheng Chen Jin Ye Guoan Wang Yanjun Li Zhongying Deng Wei Li Tianbin Li Haodong Duan Ziyan Huang Yanzhou Su et\u00a0al. 2024. GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2408.03361 (2024)."},{"key":"e_1_3_3_3_15_2","doi-asserted-by":"crossref","unstructured":"Shih-Yeh Chen Pei-Hsuan Lin and Wei-Che Chien. 2022. Children\u2019s digital art ability training system based on ai-assisted learning: a case study of drawing color perception. Frontiers in psychology 13 (2022) 823078.","DOI":"10.3389\/fpsyg.2022.823078"},{"key":"e_1_3_3_3_16_2","unstructured":"Xinlei Chen Hao Fang Tsung-Yi Lin Ramakrishna Vedantam Saurabh Gupta Piotr Doll\u00e1r and C\u00a0Lawrence Zitnick. 2015. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/1504.00325 (2015)."},{"key":"e_1_3_3_3_17_2","doi-asserted-by":"crossref","unstructured":"David\u00a0H Cropley and Rebecca\u00a0L Marrone. 2022. Automated scoring of figural creativity using a convolutional neural network. Psychology of Aesthetics Creativity and the Arts (2022).","DOI":"10.31234\/osf.io\/8qe7y"},{"key":"e_1_3_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.261"},{"key":"e_1_3_3_3_19_2","unstructured":"Wenliang Dai Junnan Li Dongxu Li Anthony Meng\u00a0Huat Tiong Junqi Zhao Weisheng Wang Boyang Li Pascale Fung and Steven Hoi. 2023. InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. arxiv:https:\/\/arXiv.org\/abs\/2305.06500\u00a0[cs.CV]"},{"key":"e_1_3_3_3_20_2","unstructured":"Keyan Ding Kede Ma Shiqi Wang and Eero\u00a0P Simoncelli. 2020. Image quality assessment: Unifying structure and texture similarity. IEEE transactions on pattern analysis and machine intelligence 44 5 (2020) 2567\u20132581."},{"key":"e_1_3_3_3_21_2","unstructured":"Yingbei Du. 2020. Research on the transformation and innovation of visual art design form based on digital fusion technology. Applied Mathematics and Nonlinear Sciences (2020)."},{"key":"e_1_3_3_3_22_2","doi-asserted-by":"crossref","unstructured":"David\u00a0W Eccles and G\u00fcler Arsal. 2017. The think aloud method: what is it and how do I use it? Qualitative Research in Sport Exercise and Health 9 4 (2017) 514\u2013531.","DOI":"10.1080\/2159676X.2017.1331501"},{"key":"e_1_3_3_3_23_2","doi-asserted-by":"crossref","unstructured":"K\u00a0Anders Ericsson. 2017. Protocol analysis. A companion to cognitive science (2017) 425\u2013432.","DOI":"10.1002\/9781405164535.ch33"},{"key":"e_1_3_3_3_24_2","doi-asserted-by":"crossref","unstructured":"Dejan Grba. 2022. Deep else: A critical framework for ai art. Digital 2 1 (2022) 1\u201332.","DOI":"10.3390\/digital2010001"},{"key":"e_1_3_3_3_25_2","unstructured":"Jiaxing Huang and Jingyi Zhang. 2024. A Survey on Evaluation of Multimodal Large Language Models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2408.15769 (2024)."},{"key":"e_1_3_3_3_26_2","doi-asserted-by":"crossref","unstructured":"Olga\u00a0M Hubard. 2010. Three modes of dialogue about works of art. Art Education 63 3 (2010) 40\u201345.","DOI":"10.1080\/00043125.2010.11519069"},{"key":"e_1_3_3_3_27_2","doi-asserted-by":"crossref","unstructured":"Christian\u00a0P Janssen and Duncan\u00a0P Brumby. 2015. Strategic adaptation to task characteristics incentives and individual differences in dual-tasking. PloS one 10 7 (2015) e0130009.","DOI":"10.1371\/journal.pone.0130009"},{"key":"e_1_3_3_3_28_2","doi-asserted-by":"publisher","unstructured":"Jing Jin and Runzhou Li. 2024. Implications Concerns and Transcendence of AI-Based Art Evaluation: Focusing on \"Elementary School Art\". Curriculum Teaching Material and Method 44 05 (2024) 138\u2013143. 10.19877\/j.cnki.kcjcjf.2024.05.020","DOI":"10.19877\/j.cnki.kcjcjf.2024.05.020"},{"key":"e_1_3_3_3_29_2","doi-asserted-by":"crossref","unstructured":"Bonnie\u00a0E John and David\u00a0E Kieras. 1996. Using GOMS for user interface design and evaluation: Which technique? ACM Transactions on Computer-Human Interaction (TOCHI) 3 4 (1996) 287\u2013319.","DOI":"10.1145\/235833.236050"},{"key":"e_1_3_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN52387.2021.9534264"},{"key":"e_1_3_3_3_31_2","doi-asserted-by":"crossref","unstructured":"Margaret Kelaher Naomi Berman David Dunt Victoria Johnson Steve Curry and Lindy Joubert. 2014. Evaluating community outcomes of participation in community arts: A case for civic dialogue. Journal of Sociology 50 2 (2014) 132\u2013149.","DOI":"10.1177\/1440783312442255"},{"key":"e_1_3_3_3_32_2","first-page":"119","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Kim Chris\u00a0Dongjoo","year":"2019","unstructured":"Chris\u00a0Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. 2019. Audiocaps: Generating captions for audios in the wild. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 119\u2013132."},{"key":"e_1_3_3_3_33_2","doi-asserted-by":"crossref","unstructured":"Maurice Lamb Rachel\u00a0W Kallen Steven\u00a0J Harrison Mario Di\u00a0Bernardo Ali Minai and Michael\u00a0J Richardson. 2017. To pass or not to pass: Modeling the movement and affordance dynamics of a pick and place task. Frontiers in psychology 8 (2017) 1061.","DOI":"10.3389\/fpsyg.2017.01061"},{"key":"e_1_3_3_3_34_2","doi-asserted-by":"crossref","unstructured":"Jason\u00a0J Lau Soumya Gayen Asma Ben\u00a0Abacha and Dina Demner-Fushman. 2018. A dataset of clinically generated visual questions and answers about radiology images. Scientific data 5 1 (2018) 1\u201310.","DOI":"10.1038\/sdata.2018.251"},{"key":"e_1_3_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642697"},{"key":"e_1_3_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3491102.3502030"},{"key":"e_1_3_3_3_37_2","unstructured":"Bohao Li Yuying Ge Yixiao Ge Guangzhi Wang Rui Wang Ruimao Zhang and Ying Shan. 2023. SEED-Bench: Benchmarking Multimodal Large Language Models. (2023). arxiv:https:\/\/arXiv.org\/abs\/2307.16125\u00a0[cs.CV]"},{"key":"e_1_3_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01263"},{"key":"e_1_3_3_3_39_2","unstructured":"Chunyuan Li Cliff Wong Sheng Zhang Naoto Usuyama Haotian Liu Jianwei Yang Tristan Naumann Hoifung Poon and Jianfeng Gao. 2023. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2306.00890 (2023)."},{"key":"e_1_3_3_3_40_2","unstructured":"Weifeng Lin Xinyu Wei Ruichuan An Peng Gao Bocheng Zou Yulin Luo Siyuan Huang Shanghang Zhang and Hongsheng Li. 2024. Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2403.20271 (2024)."},{"key":"e_1_3_3_3_41_2","unstructured":"Haotian Liu Chunyuan Li Qingyang Wu and Yong\u00a0Jae Lee. 2023. Visual instruction tuning. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2304.08485 (2023)."},{"key":"e_1_3_3_3_42_2","unstructured":"Yuan Liu Haodong Duan Yuanhan Zhang Bo Li Songyang Zhang Wangbo Zhao Yike Yuan Jiaqi Wang Conghui He Ziwei Liu et\u00a0al. 2023. MMBench: Is Your Multi-modal Model an All-around Player? arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2307.06281 (2023)."},{"key":"e_1_3_3_3_43_2","unstructured":"Ziqiang Liu Feiteng Fang Xi Feng Xinrun Du Chenhao Zhang Zekun Wang Yuelin Bai Qixuan Zhao Liyang Fan Chengguang Gan et\u00a0al. 2024. II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2406.05862 (2024)."},{"key":"e_1_3_3_3_44_2","doi-asserted-by":"crossref","unstructured":"Paul\u00a0J Locher Pieter\u00a0Jan Stappers and Kees Overbeeke. 1999. An empirical evaluation of the visual rightness theory of pictorial composition. Acta psychologica 103 3 (1999) 261\u2013280.","DOI":"10.1016\/S0001-6918(99)00044-X"},{"key":"e_1_3_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-16811-1_30"},{"key":"e_1_3_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-75018-3_14"},{"key":"e_1_3_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48506.2021.9561649"},{"key":"e_1_3_3_3_48_2","doi-asserted-by":"crossref","unstructured":"John\u00a0D Patterson Baptiste Barbot James Lloyd-Cox and Roger\u00a0E Beaty. 2024. AuDrA: An automated drawing assessment platform for evaluating creativity. Behavior Research Methods 56 4 (2024) 3619\u20133636.","DOI":"10.3758\/s13428-023-02258-3"},{"key":"e_1_3_3_3_49_2","first-page":"279","volume-title":"Educational Data Mining 2009: 2nd International Conference on Educational Data Mining: proceedings [EDM\u201909], Cordoba, Spain. July 1-3, 2009","author":"Pechenizkiy Mykola","year":"2009","unstructured":"Mykola Pechenizkiy, Nikola Trcka, Ekaterina Vasilyeva, Wil\u00a0MP van\u00a0der Aalst, and PME De\u00a0Bra. 2009. Process mining online assessment data. In Educational Data Mining 2009: 2nd International Conference on Educational Data Mining: proceedings [EDM\u201909], Cordoba, Spain. July 1-3, 2009. International Working Group on Educational Data Mining, 279\u2013288."},{"key":"e_1_3_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.2991\/assehr.k.211216.036"},{"key":"e_1_3_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3527927.3532792"},{"key":"e_1_3_3_3_52_2","unstructured":"Machel Reid Nikolay Savinov Denis Teplyashin Dmitry Lepikhin Timothy Lillicrap Jean-baptiste Alayrac Radu Soricut Angeliki Lazaridou Orhan Firat Julian Schrittwieser et\u00a0al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2403.05530 (2024)."},{"key":"e_1_3_3_3_53_2","doi-asserted-by":"crossref","unstructured":"J Rowe L Hughes D Eckstein and AM Owen. 2008. Rule-selection and action-selection have a shared neuroanatomical basis in the human prefrontal and parietal cortex. Cerebral cortex 18 10 (2008) 2275\u20132285.","DOI":"10.1093\/cercor\/bhm249"},{"key":"e_1_3_3_3_54_2","doi-asserted-by":"crossref","unstructured":"Michelle\u00a0J Searle and Lyn\u00a0M Shulha. 2016. Capturing the imagination: Arts-informed inquiry as a method in program evaluation. Canadian Journal of Program Evaluation 31 1 (2016) 34\u201360.","DOI":"10.3138\/cjpe.258"},{"key":"e_1_3_3_3_55_2","volume-title":"ECSCW","author":"Seo Woosuk","year":"2022","unstructured":"Woosuk Seo, Joonyoung Jun, Minki Chun, Hyeonhak Jeong, Sungmin Na, Woohyun Cho, Saeri Kim, and Hyunggu Jung. 2022. Toward an AI-assisted Assessment Tool to Support Online Art Therapy Practices: A Pilot Study.. In ECSCW."},{"key":"e_1_3_3_3_56_2","doi-asserted-by":"crossref","unstructured":"S Sfarra C Ibarra-Castanedo D Ambrosini D Paoletti A Bendada and X Maldague. 2014. Discovering the defects in paintings using non-destructive testing (NDT) techniques and passing through measurements of deformation. Journal of Nondestructive Evaluation 33 (2014) 358\u2013383.","DOI":"10.1007\/s10921-013-0223-7"},{"key":"e_1_3_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.23919\/ACC.1992.4792465"},{"key":"e_1_3_3_3_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/PADS.2007.30"},{"key":"e_1_3_3_3_59_2","unstructured":"Yashar Talebirad and Amirhossein Nadiri. 2023. Multi-agent collaboration: Harnessing the power of intelligent llm agents. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2306.03314 (2023)."},{"key":"e_1_3_3_3_60_2","unstructured":"Gemini Team Rohan Anil Sebastian Borgeaud Yonghui Wu Jean-Baptiste Alayrac Jiahui Yu Radu Soricut Johan Schalkwyk Andrew\u00a0M Dai Anja Hauth et\u00a0al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2312.11805 (2023)."},{"key":"e_1_3_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISDA.2009.159"},{"key":"e_1_3_3_3_62_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24258-3_26"},{"key":"e_1_3_3_3_63_2","unstructured":"Shukang Yin Chaoyou Fu Sirui Zhao Ke Li Xing Sun Tong Xu and Enhong Chen. 2023. A survey on multimodal large language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2306.13549 (2023)."},{"key":"e_1_3_3_3_64_2","unstructured":"Weihao Yu Zhengyuan Yang Linjie Li Jianfeng Wang Kevin Lin Zicheng Liu Xinchao Wang and Lijuan Wang. 2023. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2308.02490 (2023)."},{"key":"e_1_3_3_3_65_2","unstructured":"Yuqian Yuan Wentong Li Jian Liu Dongqi Tang Xinjie Luo Chi Qin Lei Zhang and Jianke Zhu. 2024. Osprey: Pixel Understanding with Visual Instruction Tuning. arxiv:https:\/\/arXiv.org\/abs\/2312.10032\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2312.10032"},{"key":"e_1_3_3_3_66_2","doi-asserted-by":"crossref","unstructured":"Duzhen Zhang Yahan Yu Chenxing Li Jiahua Dong Dan Su Chenhui Chu and Dong Yu. 2024. Mm-llms: Recent advances in multimodal large language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2401.13601 (2024).","DOI":"10.18653\/v1\/2024.findings-acl.738"},{"key":"e_1_3_3_3_67_2","volume-title":"Two Systems of Chinese Life Aesthetics","author":"Zhang Jun","year":"2020","unstructured":"Jun Zhang. 2020. Two Systems of Chinese Life Aesthetics. People\u2019s Publishing House, Beijing. 010\u2013011 pages."},{"key":"e_1_3_3_3_68_2","doi-asserted-by":"crossref","unstructured":"Jiajing Zhang Yongwei Miao and Jinhui Yu. 2021. A comprehensive survey on computational aesthetic evaluation of visual art images: Metrics and challenges. IEEE Access 9 (2021) 77164\u201377187.","DOI":"10.1109\/ACCESS.2021.3083075"},{"key":"e_1_3_3_3_69_2","unstructured":"Wenxuan Zhang Mahani Aljunied Chang Gao Yew\u00a0Ken Chia and Lidong Bing. 2023. M3exam: A multilingual multimodal multilevel benchmark for examining large language models. Advances in Neural Information Processing Systems 36 (2023) 5484\u20135505."},{"key":"e_1_3_3_3_70_2","doi-asserted-by":"crossref","unstructured":"Liang Zhao Eslam Hussam Jin-Taek Seong Assem Elshenawy Mustafa Kamal and Etaf Alshawarbeh. 2024. Revolutionizing art education: Integrating AI and multimedia for enhanced appreciation teaching. Alexandria Engineering Journal 93 (2024) 33\u201343.","DOI":"10.1016\/j.aej.2024.03.011"},{"key":"e_1_3_3_3_71_2","unstructured":"Tiancheng Zhao Tianqi Zhang Mingwei Zhu Haozhan Shen Kyusong Lee Xiaopeng Lu and Jianwei Yin. 2023. VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects Attributes and Relations. arxiv:https:\/\/arXiv.org\/abs\/2207.00221\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2207.00221"}],"event":{"name":"CHI 2025: CHI Conference on Human Factors in Computing Systems","location":"Yokohama Japan","acronym":"CHI '25","sponsor":["SIGCHI ACM Special Interest Group on Computer-Human Interaction"]},"container-title":["Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3706598.3713274","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3706598.3713274","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,4]],"date-time":"2025-07-04T05:55:50Z","timestamp":1751608550000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3706598.3713274"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,25]]},"references-count":70,"alternative-id":["10.1145\/3706598.3713274","10.1145\/3706598"],"URL":"https:\/\/doi.org\/10.1145\/3706598.3713274","relation":{},"subject":[],"published":{"date-parts":[[2025,4,25]]},"assertion":[{"value":"2025-04-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}