{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T08:58:05Z","timestamp":1785488285935,"version":"3.56.0"},"publisher-location":"New York, NY, USA","reference-count":36,"publisher":"ACM","license":[{"start":{"date-parts":[[2025,12,17]],"date-time":"2025-12-17T00:00:00Z","timestamp":1765929600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,12,17]]},"DOI":"10.1145\/3774521.3774566","type":"proceedings-article","created":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T07:34:24Z","timestamp":1785483264000},"page":"1-9","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Towards Structured Multimodal Understanding: Scientific Diagram Captioning with Modern VLMs"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2576-8943","authenticated-orcid":false,"given":"Deepika","family":"Kamboj","sequence":"first","affiliation":[{"name":"Computer Science &amp; Engineering, IIT Jodhpur, Jodhpur, Rajasthan, India and School of Computer Science, UPES, dehradun, Uttarakhand, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7943-0123","authenticated-orcid":false,"given":"Gaurav","family":"Harit","sequence":"additional","affiliation":[{"name":"Computer Science &amp; Engineering, IIT Jodhpur, Jodhpur, Rajasthan, India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,31]]},"reference":[{"key":"e_1_3_3_2_2_2","unstructured":"Marah Abdin Jyoti Aneja Hany Awadalla Ahmed Awadallah Ammar\u00a0Ahmad Awan Nguyen Bach Amit Bahree Arash Bakhtiari Jianmin Bao Harkirat Behl et\u00a0al. 2024. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2404.14219 (2024)."},{"key":"e_1_3_3_2_3_2","unstructured":"Peter Anderson Basura Fernando Mark Johnson and Stephen Gould. 2016. SPICE: Semantic Propositional Image Caption Evaluation. arxiv:https:\/\/arXiv.org\/abs\/1607.08822\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/1607.08822"},{"key":"e_1_3_3_2_4_2","unstructured":"Jinze Bai Shuai Bai Shusheng Yang Shijie Wang Sinan Tan Peng Wang Junyang Lin Chang Zhou and Jingren Zhou. 2023. Qwen-VL: A Versatile Vision-Language Model for Understanding Localization Text Reading and Beyond. arxiv:https:\/\/arXiv.org\/abs\/2308.12966\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2308.12966"},{"key":"e_1_3_3_2_5_2","unstructured":"Shuai Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge Sibo Song Kai Dang Peng Wang Shijie Wang Jun Tang et\u00a0al. 2025. Qwen2. 5-vl technical report. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2502.13923 (2025)."},{"key":"e_1_3_3_2_6_2","first-page":"65","volume-title":"Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization","author":"Banerjee Satanjeev","year":"2005","unstructured":"Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization, Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare Voss (Eds.). Association for Computational Linguistics, Ann Arbor, Michigan, 65\u201372. https:\/\/aclanthology.org\/W05-0909\/"},{"key":"e_1_3_3_2_7_2","unstructured":"Devichand Budagam Ashutosh Kumar Mahsa Khoshnoodi KJ Sankalp Vinija Jain and Aman Chadha. 2025. Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles. (2025)."},{"key":"e_1_3_3_2_8_2","unstructured":"Feilong Chen Xiuyi Chen Fandong Meng Peng Li and Jie Zhou. 2022. GoG: Relation-aware Graph-over-Graph Network for Visual Dialog. arxiv:https:\/\/arXiv.org\/abs\/2109.08475\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2109.08475"},{"key":"e_1_3_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58577-8_7"},{"key":"e_1_3_3_2_10_2","unstructured":"Zhe Chen Weiyun Wang Yue Cao Yangzhou Liu Zhangwei Gao Erfei Cui Jinguo Zhu Shenglong Ye Hao Tian Zhaoyang Liu et\u00a0al. 2024. Expanding performance boundaries of open-source multimodal models with model data and test-time scaling. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2412.05271 (2024)."},{"key":"e_1_3_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Yuren Cong Michael\u00a0Ying Yang and Bodo Rosenhahn. 2023. Reltr: Relation transformer for scene graph generation. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 9 (2023) 11169\u201311183.","DOI":"10.1109\/TPAMI.2023.3268066"},{"key":"e_1_3_3_2_12_2","unstructured":"Wenliang Dai Junnan Li Dongxu Li Anthony Meng\u00a0Huat Tiong Junqi Zhao Weisheng Wang Boyang Li Pascale Fung and Steven Hoi. 2023. InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. arxiv:https:\/\/arXiv.org\/abs\/2305.06500\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2305.06500"},{"key":"e_1_3_3_2_13_2","unstructured":"Yifan Du Zikang Liu Junyi Li and Wayne\u00a0Xin Zhao. 2022. A survey of vision-language pre-trained models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2202.10936 (2022)."},{"key":"e_1_3_3_2_14_2","unstructured":"Jack Hessel Ari Holtzman Maxwell Forbes Ronan\u00a0Le Bras and Yejin Choi. 2022. CLIPScore: A Reference-free Evaluation Metric for Image Captioning. arxiv:https:\/\/arXiv.org\/abs\/2104.08718\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2104.08718"},{"key":"e_1_3_3_2_15_2","unstructured":"Yifan Hou Buse Giledereli Yilei Tu and Mrinmaya Sachan. 2024. Do vision-language models really understand visual language? arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.00193 (2024)."},{"key":"e_1_3_3_2_16_2","unstructured":"Ting-Yao Hsu C\u00a0Lee Giles and Ting-Hao\u2019Kenneth\u2019 Huang. 2021. SciCap: Generating captions for scientific figures. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2110.11624 (2021)."},{"key":"e_1_3_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548112"},{"key":"e_1_3_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_15"},{"key":"e_1_3_3_2_19_2","first-page":"19730","volume-title":"International conference on machine learning","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning. PMLR, 19730\u201319742."},{"key":"e_1_3_3_2_20_2","unstructured":"Liunian\u00a0Harold Li Mark Yatskar Da Yin Cho-Jui Hsieh and Kai-Wei Chang. 2019. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/1908.03557 (2019)."},{"key":"e_1_3_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_3_2_22_2","unstructured":"Jiasen Lu Dhruv Batra Devi Parikh and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems 32 (2019)."},{"key":"e_1_3_3_2_23_2","unstructured":"Pan Lu Liang Qiu Jiaqi Chen Tony Xia Yizhou Zhao Wei Zhang Zhou Yu Xiaodan Liang and Song-Chun Zhu. 2021. Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2110.13214 (2021)."},{"key":"e_1_3_3_2_24_2","doi-asserted-by":"crossref","unstructured":"Ahmed Masry Do\u00a0Xuan Long Jia\u00a0Qing Tan Shafiq Joty and Enamul Hoque. 2022. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2203.10244 (2022).","DOI":"10.18653\/v1\/2022.findings-acl.177"},{"key":"e_1_3_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV48630.2021.00225"},{"key":"e_1_3_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093523"},{"key":"e_1_3_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1238"},{"key":"e_1_3_3_2_29_2","unstructured":"Andreas Steiner Andr\u00e9\u00a0Susano Pinto Michael Tschannen Daniel Keysers Xiao Wang Yonatan Bitton Alexey Gritsenko Matthias Minderer Anthony Sherbondy Shangbang Long et\u00a0al. 2024. Paligemma 2: A family of versatile vlms for transfer. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2412.03555 (2024)."},{"key":"e_1_3_3_2_30_2","unstructured":"Ningyuan Sun Xuefeng Yang and Yunfeng Liu. 2020. Tableqa: a large-scale chinese text-to-sql dataset for table-aware sql generation. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2006.06434 (2020)."},{"key":"e_1_3_3_2_31_2","unstructured":"Ramakrishna Vedantam C.\u00a0Lawrence Zitnick and Devi Parikh. 2015. CIDEr: Consensus-based Image Description Evaluation. arxiv:https:\/\/arXiv.org\/abs\/1411.5726\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/1411.5726"},{"key":"e_1_3_3_2_32_2","unstructured":"Weihan Wang Qingsong Lv Wenmeng Yu Wenyi Hong Ji Qi Yan Wang Junhui Ji Zhuoyi Yang Lei Zhao Xixuan Song Jiazheng Xu Bin Xu Juanzi Li Yuxiao Dong Ming Ding and Jie Tang. 2024. CogVLM: Visual Expert for Pretrained Language Models. arxiv:https:\/\/arXiv.org\/abs\/2311.03079\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2311.03079"},{"key":"e_1_3_3_2_33_2","unstructured":"Ning Xie Farley Lai Derek Doran and Asim Kadav. 2019. Visual entailment: A novel task for fine-grained image understanding. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/1901.06706 (2019)."},{"key":"e_1_3_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403172"},{"key":"e_1_3_3_2_35_2","unstructured":"Junhan Yang Zheng Liu Shitao Xiao Chaozhuo Li Defu Lian Sanjay Agrawal Amit Singh Guangzhong Sun and Xing Xie. 2023. GraphFormers: GNN-nested Transformers for Representation Learning on Textual Graph. arxiv:https:\/\/arXiv.org\/abs\/2105.02605\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2105.02605"},{"key":"e_1_3_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475255"},{"key":"e_1_3_3_2_37_2","unstructured":"Deyao Zhu Jun Chen Xiaoqian Shen Xiang Li and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2304.10592 (2023)."}],"event":{"name":"ICVGIP 2025: Indian Conference on Computer Vision, Graphics, and Image Processing","location":"Mandi Himachal Pradesh India","acronym":"ICVGIP 2025"},"container-title":["Proceedings of the Sixteen Indian Conference on Computer Vision, Graphics and Image Processing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3774521.3774566","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T08:07:36Z","timestamp":1785485256000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3774521.3774566"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,17]]},"references-count":36,"alternative-id":["10.1145\/3774521.3774566","10.1145\/3774521"],"URL":"https:\/\/doi.org\/10.1145\/3774521.3774566","relation":{},"subject":[],"published":{"date-parts":[[2025,12,17]]},"assertion":[{"value":"2026-07-31","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}