{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,22]],"date-time":"2026-06-22T13:58:00Z","timestamp":1782136680315,"version":"3.54.5"},"reference-count":85,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"name":"National Library of Medicine, National Institutes of Health","award":["R00LM014024"],"award-info":[{"award-number":["R00LM014024"]}]},{"name":"National Library of Medicine and the National Eye Institute"},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Comput. Healthcare"],"published-print":{"date-parts":[[2026,7,30]]},"abstract":"<jats:p>The rising prevalence of vision-threatening eye diseases poses a major global health and economic burden, yet timely diagnosis remains limited by workforce shortages, diagnostic delays, and restricted access to specialized care. Artificial intelligence (AI) offers potential solutions. In particular, recent progress in foundation models and large language models\u2014especially multimodal large language models (MLLMs)\u2014has shown promise in medical image interpretation and automated clinical documentation. However, advancing MLLMs for ophthalmology is hindered by the lack of unified, comprehensive benchmark datasets for development and evaluation. Most existing benchmarks were designed for earlier models, which focused on narrow tasks or specific disease conditions. These benchmarks typically provide outputs in the form of disease labels rather than free-text responses. As a result, they are less suitable for assessing emerging generative models.<\/jats:p>\n                  <jats:p>In this work, we present LMOD+, a large-scale multimodal ophthalmology benchmark dataset comprising 32,633 instances with multi-granular annotations across 12 common ophthalmic conditions and 5 imaging modalities. The dataset integrates imaging, anatomical structures, demographics, and free-text annotations. It supports primary ophthalmic applications such as anatomical structure recognition, disease screening, disease staging, and demographic prediction for potential performance bias evaluation. Alongside the dataset, we introduce a systematic and unified data curation pipeline that repurposes existing or new datasets for MLLM development.<\/jats:p>\n                  <jats:p>LMOD+ extends our preliminary LMOD benchmark\u2014the first multimodal ophthalmology benchmark for MLLMs\u2014with three major enhancements. First, we expanded the dataset by nearly 50% (from 21,933 to 32,633 instances). The color fundus photography (CFP) modality, the most accessible imaging modality in ophthalmology, was significantly enlarged to cover a broader range of pathological conditions. Second, we broadened task coverage to include (a) 12 binary disease diagnosis tasks for prevalent conditions such as diabetic retinopathy, age-related macular degeneration, and retinal vein occlusion; (b) multi-class ophthalmic disease diagnosis; (c) disease severity classification, including a diabetic retinopathy staging task, which uses two internationally adopted grading standards: the international clinical diabetic retinopathy classification and the Scottish diabetic retinopathy grading scheme classification; and (d) demographic prediction (age and sex) to assess potential model bias. Third, we systematically evaluated 24 state-of-the-art MLLMs, including recent models from the InternVL, Qwen, and DeepSeek families.<\/jats:p>\n                  <jats:p>Our evaluations highlight both the promise and limitations of current MLLMs in ophthalmology. For example, Qwen-7B and InternVL achieved accuracies of 58.26% and 57.83% in disease screening under a zero-shot setting with a single model\u2014a considerably more challenging paradigm than traditional fine-tuning, where separate models are trained for each specific task. InternVL also demonstrated potential in anatomical recognition. Nonetheless, overall performance remained suboptimal and often close to random baselines for challenging tasks such as disease staging, underscoring the substantial gap between general-domain MLLMs and the specialized requirements of ophthalmology.<\/jats:p>\n                  <jats:p>\n                    We publicly release the dataset, curation pipeline, and leaderboard to encourage community-wide development and evaluation of MLLMs, with the goal of advancing ophthalmic applications and ultimately reducing the global burden of vision-threatening diseases through AI. The dataset website, benchmark leaderboard, and download link are available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/kfzyqin.github.io\/lmod_plus\">https:\/\/kfzyqin.github.io\/lmod_plus<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3801746","type":"journal-article","created":{"date-parts":[[2026,3,16]],"date-time":"2026-03-16T16:15:29Z","timestamp":1773677729000},"page":"1-38","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["LMOD \\(\\boldsymbol{+}\\) : A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology"],"prefix":"10.1145","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3471-6280","authenticated-orcid":false,"given":"Zhenyue","family":"Qin","sequence":"first","affiliation":[{"name":"Department of Biomedical Informatics &amp; Data Science, Yale University, New Haven, Connecticut, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-1734-9789","authenticated-orcid":false,"given":"Yang","family":"Liu","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, Pennsylvania, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0469-6653","authenticated-orcid":false,"given":"Yu","family":"Yin","sequence":"additional","affiliation":[{"name":"School of Engineering, Imperial College London, London, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-4177-2281","authenticated-orcid":false,"given":"Jinyu","family":"Ding","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics &amp; Data Science, Yale University, New Haven, Connecticut, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-1770-958X","authenticated-orcid":false,"given":"Haoran","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics &amp; Data Science, Yale University, New Haven, Connecticut, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3592-4153","authenticated-orcid":false,"given":"Anran","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics &amp; Data Science, Yale University, New Haven, Connecticut, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4717-6850","authenticated-orcid":false,"given":"Dylan","family":"Campbell","sequence":"additional","affiliation":[{"name":"Australian National University, Canberra, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7816-7658","authenticated-orcid":false,"given":"Xuansheng","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Computing, University of Georgia, Athens, Georgia, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2025-9859","authenticated-orcid":false,"given":"Ke","family":"Zou","sequence":"additional","affiliation":[{"name":"Yong Loo Lin School of Medicine, National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2253-1772","authenticated-orcid":false,"given":"Tiarnan D. L.","family":"Keenan","sequence":"additional","affiliation":[{"name":"National Institutes of Health, Bethesda, Maryland, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0999-9802","authenticated-orcid":false,"given":"Emily Y.","family":"Chew","sequence":"additional","affiliation":[{"name":"National Institutes of Health, Bethesda, Maryland, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9998-916X","authenticated-orcid":false,"given":"Zhiyong","family":"Lu","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health, Bethesda, Maryland, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6752-797X","authenticated-orcid":false,"given":"Yih Chung","family":"Tham","sequence":"additional","affiliation":[{"name":"Yong Loo Lin School of Medicine, National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9170-2424","authenticated-orcid":false,"given":"Ninghao","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computing, University of Georgia, Athens, Georgia, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5558-3790","authenticated-orcid":false,"given":"Xiuzhen","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computing Technologies, RMIT University, Melbourne, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6036-1516","authenticated-orcid":false,"given":"Qingyu","family":"Chen","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics &amp; Data Science, Yale University, New Haven, Connecticut, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,22]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et al. 2023. Gpt-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1723"},{"issue":"4","key":"e_1_3_2_4_2","doi-asserted-by":"crossref","first-page":"100324","DOI":"10.1016\/j.xops.2023.100324","article-title":"Evaluating the performance of ChatGPT in ophthalmology: An analysis of its successes and shortcomings","volume":"3","author":"Antaki Fares","year":"2023","unstructured":"Fares Antaki, Samir Touma, Daniel Milad, Jonathan El-Khoury, and Renaud Duval. 2023. Evaluating the performance of ChatGPT in ophthalmology: An analysis of its successes and shortcomings. Ophthalmology Science 3, 4 (2023), 100324.","journal-title":"Ophthalmology Science"},{"key":"e_1_3_2_5_2","unstructured":"Jinze Bai Shuai Bai Shusheng Yang Shijie Wang Sinan Tan Peng Wang Junyang Lin Chang Zhou and Jingren Zhou. 2023. Qwen-VL: A frontier large vision-language model with versatile abilities. arXiv:2308.12966. Retrieved from https:\/\/arxiv.org\/abs\/2308.12966"},{"key":"e_1_3_2_6_2","first-page":"1","volume-title":"Proceedings of the 2020 International Joint Conference on Neural Networks (IJCNN)","author":"Bajwa Muhammad Naseer","year":"2020","unstructured":"Muhammad Naseer Bajwa, Gur Amrit Pal Singh, Wolfgang Neumeier, Muhammad Imran Malik, Andreas Dengel, and Sheraz Ahmed. 2020. G1020: A benchmark retinal fundus image dataset for computer-aided glaucoma detection. In Proceedings of the 2020 International Joint Conference on Neural Networks (IJCNN). IEEE, 1\u20137."},{"issue":"8","key":"e_1_3_2_7_2","doi-asserted-by":"crossref","first-page":"e25165","DOI":"10.2196\/25165","article-title":"Gender prediction for a multiethnic population via deep learning across different retinal fundus photograph fields: Retrospective cross-sectional study","volume":"9","author":"Betzler Bjorn Kaijun","year":"2021","unstructured":"Bjorn Kaijun Betzler, Henrik Hee Seung Yang, Sahil Thakur, Marco Yu, Ten Cheer Quek, Zhi Da Soh, Geunyoung Lee, Yih-Chung Tham, Tien Yin Wong, Tyler Hyungtaek Rim, et al. 2021. Gender prediction for a multiethnic population via deep learning across different retinal fundus photograph fields: Retrospective cross-sectional study. JMIR Medical Informatics 9, 8 (2021), e25165.","journal-title":"JMIR Medical Informatics"},{"key":"e_1_3_2_8_2","volume-title":"Proceedings of the 2023 Annual Meeting of the Midwest Political Science Association (MPSA)","author":"Bosley Mitchell","year":"2023","unstructured":"Mitchell Bosley, Musashi Jacobs-Harukawa, Hauke Licht, and Alexander Hoyle. 2023. Do we still need BERT in the age of GPT? Comparing the benefits of domain-adaptation and in-context-learning approaches to using LLMs for political science research. In Proceedings of the 2023 Annual Meeting of the Midwest Political Science Association (MPSA)."},{"issue":"6","key":"e_1_3_2_9_2","doi-asserted-by":"crossref","first-page":"1150","DOI":"10.1016\/S0161-6420(01)00581-4","article-title":"Long-term follow-up of unoperated macular holes","volume":"108","author":"Casuso Lourdes A.","year":"2001","unstructured":"Lourdes A. Casuso, Ingrid U. Scott, Harry W. Flynn, Jr., J. Donald M. Gass, William E. Smiddy, Mary Lou Lewis, and Joyce Schiffman. 2001. Long-term follow-up of unoperated macular holes. Ophthalmology 108, 6 (2001), 1150\u20131155.","journal-title":"Ophthalmology"},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","first-page":"16","DOI":"10.1016\/j.diabres.2017.03.023","article-title":"The diabetic retinopathy barometer study: Global perspectives on access to and experiences of diabetic retinopathy screening and treatment","volume":"129","author":"Cavan D.","year":"2017","unstructured":"D. Cavan, L. Makaroff, J. da Rocha Fernandes, M. Sylvanowicz, P. Ackland, J. Conlon, D. Chaney, A. Malhi, and J. Barratt. 2017. The diabetic retinopathy barometer study: Global perspectives on access to and experiences of diabetic retinopathy screening and treatment. Diabetes Research and Clinical Practice 129 (2017), 16\u201324.","journal-title":"Diabetes Research and Clinical Practice"},{"issue":"7","key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"e2517204","DOI":"10.1001\/jamanetworkopen.2025.17204","article-title":"AI workflow, external validation, and development in eye disease diagnosis","volume":"8","author":"Chen Qingyu","year":"2025","unstructured":"Qingyu Chen, Tiarnan D. L. Keenan, Elvira Agron, Alexis Allot, Emily Guan, Bryant Duong, Amr Elsawy, Benjamin Hou, Cancan Xue, Sanjeeb Bhandari, et al. 2025. AI workflow, external validation, and development in eye disease diagnosis. JAMA Network Open 8, 7 (2025), e2517204.","journal-title":"JAMA Network Open"},{"key":"e_1_3_2_12_2","first-page":"505","volume-title":"Proceedings of the AMIA Summits on Translational Science Proceedings 2019","author":"Chen Qingyu","year":"2019","unstructured":"Qingyu Chen, Yifan Peng, Tiarnan Keenan, Shazia Dharssi, Elvira Agro, Wai T. Wong, Emily Y. Chew, and Zhiyong Lu. 2019. A multi-task deep learning model for the classification of age-related macular degeneration. In Proceedings of the AMIA Summits on Translational Science Proceedings 2019, 505."},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","unstructured":"Zhe Chen Jiannan Wu Wenhai Wang Weijie Su Guo Chen Sen Xing Muyan Zhong Qinglong Zhang Xizhou Zhu Lewei Lu et al. 2024. InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.","DOI":"10.1109\/CVPR52733.2024.02283"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.3949\/ccjm.91a.24028"},{"key":"e_1_3_2_15_2","first-page":"2020","article-title":"Prediction of sex and age from macular optical coherence tomography images and feature analysis using deep learning","author":"Chueh Kuan-Ming","year":"2020","unstructured":"Kuan-Ming Chueh, Yi-Ting Hsieh, Hao-Hsiang Chen, I-Hsin Ma, and Shih-Len Huang. 2020. Prediction of sex and age from macular optical coherence tomography images and feature analysis using deep learning. American Journal of Ophthalmology 235 (2020), 2020\u20132012.","journal-title":"American Journal of Ophthalmology"},{"key":"e_1_3_2_16_2","first-page":"49250","article-title":"Instructblip: Towards general-purpose vision-language models with instruction tuning","volume":"36","author":"Dai Wenliang","year":"2024","unstructured":"Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N. Fung, and Steven Hoi. 2024. Instructblip: Towards general-purpose vision-language models with instruction tuning. In Advances in Neural Information Processing Systems , Vol. 36, 49250\u201349267.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.3389\/fpubh.2023.1166120"},{"key":"e_1_3_2_18_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","volume":"1","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol. 1, 4171\u20134186."},{"key":"e_1_3_2_19_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1177\/20552076251328120"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ophtha.2012.10.036"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11023-020-09548-1"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1038\/eye.2017.1"},{"issue":"1","key":"e_1_3_2_24_2","doi-asserted-by":"crossref","first-page":"e001903","DOI":"10.1136\/bmjophth-2024-001903","article-title":"Recent advances in the application of artificial intelligence in age-related macular degeneration","volume":"9","author":"Gao Yundi","year":"2024","unstructured":"Yundi Gao, Fen Xiong, Jian Xiong, Zidan Chen, Yucai Lin, Xinjing Xia, Yulan Yang, Guodong Li, and Yunwei Hu. 2024. Recent advances in the application of artificial intelligence in age-related macular degeneration. BMJ Open Ophthalmology 9, 1 (2024), e001903.","journal-title":"BMJ Open Ophthalmology"},{"issue":"1","key":"e_1_3_2_25_2","doi-asserted-by":"crossref","first-page":"373","DOI":"10.1038\/s41597-024-03193-4","article-title":"Cataract-1K dataset for deep-learning-assisted analysis of cataract surgery videos","volume":"11","author":"Ghamsarian Negin","year":"2024","unstructured":"Negin Ghamsarian, Yosuf El-Shabrawi, Sahar Nasirihaghighi, Doris Putzgruber-Adamitsch, Martin Zinkernagel, Sebastian Wolf, Klaus Schoeffmann, and Raphael Sznitman. 2024. Cataract-1K dataset for deep-learning-assisted analysis of cataract surgery videos. Scientific Data 11, 1 (2024), 373.","journal-title":"Scientific Data"},{"key":"e_1_3_2_26_2","unstructured":"Aidan Gilson Xuguang Ai Qianqian Xie Sahana Srinivasan Krithi Pushpanathan Maxwell B. Singer Jimin Huang Hyunjae Kim Erping Long Peixing Wan et al. 2024. Language enhanced model for eye (LEME): An open-source ophthalmology-specific large language model. arXiv:2410.03740. Retrieved from https:\/\/arxiv.org\/abs\/2410.03740"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1001\/jamanetworkopen.2024.40969"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1001\/jama.2016.17216"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3437288"},{"key":"e_1_3_2_30_2","first-page":"15563","article-title":"Using self-supervised learning can improve model robustness and uncertainty","volume":"32","author":"Hendrycks Dan","year":"2019","unstructured":"Dan Hendrycks, Mantas Mazeika, Saurav Kadavath, and Dawn Song. 2019. Using self-supervised learning can improve model robustness and uncertainty. In Advances in Neural Information Processing Systems, Vol. 32, 15563\u201315674.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","unstructured":"Ming Hu Peng Xia Lin Wang Siyuan Yan Feilong Tang Zhongxing Xu Yimin Luo Kaimin Song Jurgen Leitner Xuelian Cheng et al. 2024. OphNet: A large-scale video benchmark for ophthalmic surgical workflow understanding. In Proceedings of the European Conference on Computer Vision 481\u2013500.","DOI":"10.1007\/978-3-031-73235-5_27"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/S2589-7500(20)30240-5"},{"issue":"1","key":"e_1_3_2_33_2","doi-asserted-by":"crossref","first-page":"10286","DOI":"10.1038\/s41598-021-89743-x","article-title":"Predicting sex from retinal fundus photographs using automated deep learning","volume":"11","author":"Korot Edward","year":"2021","unstructured":"Edward Korot, Nikolas Pontikos, Xiaoxuan Liu, Siegfried K. Wagner, Livia Faes, Josef Huemer, Konstantinos Balaskas, Alastair K. Denniston, Anthony Khawaja, and Pearse A. Keane. 2021. Predicting sex from retinal fundus photographs using automated deep learning. Scientific Reports 11, 1 (2021), 10286.","journal-title":"Scientific Reports"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01263"},{"key":"e_1_3_2_35_2","first-page":"28541","article-title":"LLaVA-Med: Training a large language-and-vision assistant for biomedicine in one day","volume":"36","author":"Li Chunyuan","year":"2024","unstructured":"Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. 2024. LLaVA-Med: Training a large language-and-vision assistant for biomedicine in one day. In Advances in Neural Information Processing Systems, Vol. 36, 28541\u201328564.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_36_2","first-page":"19730","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the International Conference on Machine Learning. PMLR, 19730\u201319742."},{"key":"e_1_3_2_37_2","unstructured":"Zongxia Li Xiyang Wu Hongyang Du Fuxiao Liu Huy Nghiem and Guangyao Shi. 2025. A survey of state of the art large vision language models: Alignment benchmark evaluations and challenges. arXiv:2501.02189. Retrieved from https:\/\/arxiv.org\/abs\/2501.02189"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3672758.3672824"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ebiom.2023.104770"},{"key":"e_1_3_2_40_2","first-page":"26689","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Lin Ji","year":"2024","unstructured":"Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han. 2024. VILA: On pre-training for visual language models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 26689\u201326699."},{"key":"e_1_3_2_41_2","first-page":"34892","article-title":"Visual instruction tuning","volume":"36","author":"Liu Haotian","year":"2024","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. In Advances in Neural Information Processing Systems, Vol. 36, 34892\u201334916","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1016\/S2589-7500(19)30123-2"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.metrad.2023.100017"},{"key":"e_1_3_2_44_2","unstructured":"Pan Lu Hritik Bansal Tony Xia Jiacheng Liu Chunyuan Li Hannaneh Hajishirzi Hao Cheng Kai-Wei Chang Michel Galley and Jianfeng Gao. 2023. MathVista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv:2310.02255. Retrieved from https:\/\/arxiv.org\/abs\/2310.02255"},{"issue":"1","key":"e_1_3_2_45_2","first-page":"5278196","article-title":"Applications of artificial intelligence in ophthalmology: General overview","volume":"2018","author":"Lu Wei","year":"2018","unstructured":"Wei Lu, Yan Tong, Yue Yu, Yiqiao Xing, Changzheng Chen, and Yin Shen. 2018. Applications of artificial intelligence in ophthalmology: General overview. Journal of Ophthalmology 2018, 1 (2018), 5278196.","journal-title":"Journal of Ophthalmology"},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","unstructured":"Yan Luo Yu Tian Min Shi Louis R. Pasquale Lucy Q. Shen Nazlee Zebardast Tobias Elze and Mengyu Wang. 2024. Harvard glaucoma fairness: A retinal nerve disease dataset for fairness learning and fair identity normalization. IEEE Transactions on Medical Imaging 43 7 (2024) 2623\u20132633.","DOI":"10.1109\/TMI.2024.3377552"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.173"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-019-1799-6"},{"key":"e_1_3_2_49_2","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1007\/978-3-031-83756-2_8","volume-title":"Artificial Intelligence in Ophthalmology","author":"Mukherjee Souvick","year":"2025","unstructured":"Souvick Mukherjee, Yifan Peng, Qingyu Chen, Tiarnan D. L. Keenan, Emily Y. Chew, and Zhiyong Lu. 2025. Artificial Intelligence in age-related macular degeneration (AMD). In Artificial Intelligence in Ophthalmology. Andrzej Grzybowski (Ed.), Springer, 121\u2013135."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pdig.0000454"},{"issue":"6","key":"e_1_3_2_51_2","doi-asserted-by":"crossref","first-page":"570","DOI":"10.1001\/jamaophthalmol.2017.0830","article-title":"Prevalence of undiagnosed age-related macular degeneration in primary eye care","volume":"135","author":"Neely David C.","year":"2017","unstructured":"David C. Neely, Kevin J. Bray, Carrie E. Huisingh, Mark E. Clark, Gerald McGwin, and Cynthia Owsley. 2017. Prevalence of undiagnosed age-related macular degeneration in primary eye care. JAMA Ophthalmology 135, 6 (2017), 570\u2013575.","journal-title":"JAMA Ophthalmology"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.3390\/medicina61030433"},{"key":"e_1_3_2_53_2","unstructured":"World Health Organization. 2023. Blindness and Vision Impairment. Retrieved from https:\/\/www.who.int\/news-room\/fact-sheets\/detail\/blindness-and-visual-impairmentAccessed [Insert access date here]."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2019.101570"},{"key":"e_1_3_2_55_2","volume-title":"Hypertension Prevalence Among Adults Aged 18 and over: United States, 2017\u20132018","author":"Ostchega Yechiam","year":"2020","unstructured":"Yechiam Ostchega, Cheryl D. Fryar, Tatiana Nwankwo, and Duong T. Nguyen. 2020. Hypertension Prevalence Among Adults Aged 18 and over: United States, 2017\u20132018. CDC Stacks."},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41551-018-0195-0"},{"key":"e_1_3_2_57_2","volume-title":"IEEE Dataport 2","author":"Prasanna Porwal","year":"2018","unstructured":"Porwal Prasanna, Pachade Samiksha, Kamble Ravi, Kokare Manesh, D. Girish, S. Vivek, and Meriaudeau Fabrice. 2018. Indian Diabetic Retinopathy Image Dataset (IDRiD). IEEE Dataport 2."},{"key":"e_1_3_2_58_2","unstructured":"PupiUp. 2023. cau001 Dataset. Roboflow Universe. Retrieved June 3 2024 from https:\/\/universe.roboflow.com\/pupiup-rjvfv\/cau001"},{"key":"e_1_3_2_59_2","unstructured":"Alec Radford Jong Wook Kim Chris Hallacy Aditya Ramesh Gabriel Goh Sandhini Agarwal Girish Sastry Amanda Askell Pamela Mishkin Jack Clark et al. 2021. Learning Transferable Visual Models from Natural Language Supervision. PMLR 8748\u20138763."},{"key":"e_1_3_2_60_2","unstructured":"SRM University Ramapuram. 2023. Cataract Detection 2 Dataset. Roboflow Universe (Visited on June 3 2024)."},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-019-0048-x"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1001\/jamaophthalmol.2020.0273"},{"key":"e_1_3_2_63_2","unstructured":"Sahana Srinivasan Xuguang Ai Thaddaeus Wai Soon Lo Aidan Gilson Minjie Zou Ke Zou Hyunjae Kim Mingjia Yang Krithi Pushpanathan Samantha Yew et al. 2025. BEnchmarking LLMs for ophthalmology (BELO) for ophthalmological knowledge and reasoning. arXiv:2507.15717. Retrieved from https:\/\/arxiv.org\/abs\/2507.15717"},{"key":"e_1_3_2_64_2","unstructured":"Gemini Team Rohan Anil Sebastian Borgeaud Jean-Baptiste Alayrac Jiahui Yu Radu Soricut Johan Schalkwyk Andrew M. Dai Anja Hauth Katie Millican et al. 2023. Gemini: A family of highly capable multimodal models. arXiv:2312.11805. Retrieved from https:\/\/arxiv.org\/abs\/2312.11805"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ophtha.2014.05.013"},{"issue":"1","key":"e_1_3_2_66_2","doi-asserted-by":"crossref","first-page":"bbad493","DOI":"10.1093\/bib\/bbad493","article-title":"Opportunities and challenges for ChatGPT and large language models in biomedicine and health","volume":"25","author":"Tian Shubo","year":"2023","unstructured":"Shubo Tian, Qiao Jin, Lana Yeganova, Po-Ting Lai, Qingqing Zhu, Xiuying Chen, Yifan Yang, Qingyu Chen, Won Kim, Donald C. Comeau, et al. 2023. Opportunities and challenges for ChatGPT and large language models in biomedicine and health. Briefings in Bioinformatics 25, 1 (2023), bbad493.","journal-title":"Briefings in Bioinformatics"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1093\/bib\/bbad493"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1136\/bjophthalmol-2018-313173"},{"issue":"1","key":"e_1_3_2_69_2","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1186\/s40662-020-00183-6","article-title":"Application of machine learning in ophthalmic imaging modalities","volume":"7","author":"Tong Yan","year":"2020","unstructured":"Yan Tong, Wei Lu, Yue Yu, and Yin Shen. 2020. Application of machine learning in ophthalmic imaging modalities. Eye and Vision 7, 1 (2020), 22.","journal-title":"Eye and Vision"},{"key":"e_1_3_2_70_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et al. 2023. LLaMA: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_2_71_2","unstructured":"Weiyun Wang Zhe Chen Wenhai Wang Yue Cao Yangzhou Liu Zhangwei Gao Jinguo Zhu Xizhou Zhu Lewei Lu Yu Qiao et al. 2025. Enhancing the reasoning ability of multimodal large language models via mixed preference optimization. arXiv:2411.10442. Retrieved from https:\/\/arxiv.org\/abs\/2411.10442"},{"issue":"2","key":"e_1_3_2_72_2","first-page":"AIdbp2300092","article-title":"Benchmarking open-source large language models, GPT-4 and Claude 2 on multiple-choice questions in nephrology","volume":"1","author":"Wu Sean","year":"2024","unstructured":"Sean Wu, Michael Koo, Lesley Blum, Andy Black, Liyo Kao, Zhe Fei, Fabien Scalzo, and Ira Kurtz. 2024. Benchmarking open-source large language models, GPT-4 and Claude 2 on multiple-choice questions in nephrology. NEJM AI 1, 2 (2024), AIdbp2300092.","journal-title":"NEJM AI"},{"issue":"11","key":"e_1_3_2_73_2","doi-asserted-by":"crossref","first-page":"714","DOI":"10.21037\/atm-20-976","article-title":"Application of artificial intelligence in anterior segment ophthalmic diseases: Diversity and standardization","volume":"8","author":"Wu Xiaohang","year":"2020","unstructured":"Xiaohang Wu, Lixue Liu, Lanqin Zhao, Chong Guo, Ruiyang Li, Ting Wang, Xiaonan Yang, Peichen Xie, Yizhi Liu, and Haotian Lin. 2020. Application of artificial intelligence in anterior segment ophthalmic diseases: Diversity and standardization. Annals of Translational Medicine 8, 11 (2020), 714.","journal-title":"Annals of Translational Medicine"},{"key":"e_1_3_2_74_2","unstructured":"Zhiyu Wu Xiaokang Chen Zizheng Pan Xingchao Liu Wen Liu Damai Dai Huazuo Gao Yiyang Ma Chengyue Wu Bingxuan Wang et al. 2024. DeepSeek-VL2: Mixture-of-experts vision-language models for advanced multimodal understanding. arXiv:2412.10302. Retrieved from https:\/\/arxiv.org\/abs\/2412.10302"},{"key":"e_1_3_2_75_2","unstructured":"Wenkai Yang Shuming Ma Yankai Lin and Furu Wei. 2025. Towards thinking-optimal scaling of test-time compute for LLM reasoning. arXiv:2502.18080. Retrieved from https:\/\/arxiv.org\/abs\/2502.18080"},{"issue":"1","key":"e_1_3_2_76_2","doi-asserted-by":"crossref","first-page":"769","DOI":"10.1038\/s41597-023-02675-1","article-title":"OIMHS: An optical coherence tomography image dataset based on macular hole manual segmentation","volume":"10","author":"Ye Xin","year":"2023","unstructured":"Xin Ye, Shucheng He, Xiaxing Zhong, Jiafeng Yu, Shangchao Yang, Yingjiao Shen, Yiqi Chen, Yaqi Wang, Xingru Huang, and Lijun Shen. 2023. OIMHS: An optical coherence tomography image dataset based on macular hole manual segmentation. Scientific Data 10, 1 (2023), 769.","journal-title":"Scientific Data"},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.1093\/nsr\/nwae403"},{"key":"e_1_3_2_78_2","unstructured":"Alex Young Bei Chen Chao Li Chengen Huang Ge Zhang Guanwei Zhang Heng Li Jiangcheng Zhu Jianqun Chen Jing Chang et al. 2024. Yi: Open foundation models by 01.AI. arXiv:2403.04652. Retrieved from https:\/\/arxiv.org\/abs\/2403.04652"},{"key":"e_1_3_2_79_2","doi-asserted-by":"crossref","unstructured":"Xiang Yue Yuansheng Ni Kai Zhang Tianyu Zheng Ruoqi Liu Ge Zhang Samuel Stevens Dongfu Jiang Weiming Ren Yuxuan Sun et al. 2024. MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.","DOI":"10.1109\/CVPR52733.2024.00913"},{"key":"e_1_3_2_80_2","doi-asserted-by":"crossref","unstructured":"Duzhen Zhang Yahan Yu Jiahua Dong Li Chenxing Su Dan Chu and Chenhui Dong Yu. 2024. MM-LLMs: Recent advances in multimodal large language models. arXiv:2401.13601. Retrieved from https:\/\/arxiv.org\/abs\/2401.13601","DOI":"10.18653\/v1\/2024.findings-acl.738"},{"key":"e_1_3_2_81_2","unstructured":"Jiawei Zhang Tianyu Pang Chao Du Yi Ren Bo Li and Min Lin. 2024. Benchmarking large multimodal models against common corruptions. arXiv:2401.11943. Retrieved from https:\/\/arxiv.org\/abs\/2401.11943"},{"key":"e_1_3_2_82_2","unstructured":"Yi-Fan Zhang Huanyu Zhang Haochen Tian Chaoyou Fu Shuangqing Zhang Junfei Wu Feng Li Kun Wang Qingsong Wen Zhang Zhang et al. 2024. MME-RealWorld: Could your multimodal LLM challenge high-resolution real-world scenarios that are difficult for humans? arXiv:2408.13257. Retrieved from https:\/\/arxiv.org\/abs\/2408.13257"},{"key":"e_1_3_2_83_2","first-page":"3065","volume-title":"Proceedings of the 2010 Annual International Conference of the IEEE Engineering in Medicine and Biology","author":"Zhang Zhuo","year":"2010","unstructured":"Zhuo Zhang, Feng Shou Yin, Jiang Liu, Wing Kee Wong, Ngan Meng Tan, Beng Hai Lee, Jun Cheng, and Tien Yin Wong. 2010. ORIGA-light: An online retinal fundus image database for glaucoma analysis and research. In Proceedings of the 2010 Annual International Conference of the IEEE Engineering in Medicine and Biology. IEEE, 3065\u20133068."},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-023-06555-x"},{"key":"e_1_3_2_85_2","unstructured":"Deyao Zhu Jun Chen Xiaoqian Shen Xiang Li and Mohamed Elhoseiny. 2023. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. arXiv:2304.10592. Retrieved from https:\/\/arxiv.org\/abs\/2304.10592"},{"key":"e_1_3_2_86_2","unstructured":"Minjie Zou Sahana Srinivasan Thaddaeus Wai Soon Lo Ke Zou Gabriel Dawei Yang Xuguang Ai Hyunjae Kim Maxwell Singer Fares Antaki Kelvin Li et al. 2025. Benchmarking next-generation reasoning-focused large language models in ophthalmology: A head-to-head evaluation on 5 888 items. arXiv:2504.11186. Retrieved from https:\/\/arxiv.org\/abs\/2504.11186"}],"container-title":["ACM Transactions on Computing for Healthcare"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3801746","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,22]],"date-time":"2026-06-22T13:14:05Z","timestamp":1782134045000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3801746"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,22]]},"references-count":85,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,7,30]]}},"alternative-id":["10.1145\/3801746"],"URL":"https:\/\/doi.org\/10.1145\/3801746","relation":{},"ISSN":["2691-1957","2637-8051"],"issn-type":[{"value":"2691-1957","type":"print"},{"value":"2637-8051","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,22]]},"assertion":[{"value":"2025-09-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-03","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-22","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}