{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,2]],"date-time":"2026-04-02T14:23:14Z","timestamp":1775139794950,"version":"3.50.1"},"reference-count":14,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,8,25]],"date-time":"2025-08-25T00:00:00Z","timestamp":1756080000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Digit. Health"],"abstract":"<jats:sec><jats:title>Introduction<\/jats:title><jats:p>Vision language models (VLMs) combine image analysis capabilities with large language models (LLMs). Because of their multimodal capabilities, VLMs offer a clinical advantage over image classification models for the diagnosis of optic disc swelling by allowing a consideration of clinical context. In this study, we compare the performance of non-specialty-trained VLMs with different prompts in the classification of optic disc swelling on fundus photographs.<\/jats:p><\/jats:sec><jats:sec><jats:title>Methods<\/jats:title><jats:p>A diagnostic test accuracy study was conducted utilizing an open-sourced dataset. Five different prompts (increasing in context) were used with each of five different VLMs (Llama 3.2-vision, LLaVA-Med, LLaVA, GPT-4o, and DeepSeek-4V), resulting in 25 prompt-model pairs. The performance of VLMs in classifying photographs with and without optic disc swelling was measured using Youden's index (YI), F1 score, and accuracy rate.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>A total of 779 images of normal optic discs and 295 images of swollen discs were obtained from an open-source image database. Among the 25 prompt-model pairs, valid response rates ranged from 7.8% to 100% (median 93.6%). Diagnostic performance ranged from YI: 0.00 to 0.231 (median 0.042), F1 score: 0.00 to 0.716 (median 0.401), and accuracy rate: 27.5 to 70.5% (median 58.8%). The best-performing prompt-model pair was GPT-4o with role-playing with Chain-of-Thought and few-shot prompting. On average, Llama 3.2-vision performed the best (average YI across prompts 0.181). There was no consistent relationship between the amount of information given in the prompt and the model performance.<\/jats:p><\/jats:sec><jats:sec><jats:title>Conclusions<\/jats:title><jats:p>Non-specialty-trained VLMs could classify photographs of swollen and normal optic discs better than chance, with performance varying by model. Increasing prompt complexity did not consistently improve performance. Specialty-specific VLMs may be necessary to improve ophthalmic image analysis performance.<\/jats:p><\/jats:sec>","DOI":"10.3389\/fdgth.2025.1660887","type":"journal-article","created":{"date-parts":[[2025,8,25]],"date-time":"2025-08-25T05:26:52Z","timestamp":1756099612000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Performance of vision language models for optic disc swelling identification on fundus photographs"],"prefix":"10.3389","volume":"7","author":[{"given":"Kelvin Zhenghao","family":"Li","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tuyet Thao","family":"Nguyen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Heather E.","family":"Moss","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2025,8,25]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"1687","DOI":"10.1056\/NEJMoa1917130","article-title":"Artificial intelligence to detect papilledema from ocular fundus photographs","volume":"382","author":"Milea","year":"2020","journal-title":"N Engl J Med"},{"key":"B2","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1016\/j.ajo.2025.04.006","article-title":"Deep learning approach readily differentiates papilledema, NAION, and healthy eyes","volume":"276","author":"Szanto","year":"2025","journal-title":"Am J Ophthalmol"},{"key":"B3","doi-asserted-by":"publisher","first-page":"100681","DOI":"10.1016\/j.xops.2024.100681","article-title":"Large language models in ophthalmology: a review of publications from top ophthalmology journals","volume":"5","author":"Agnihotri","year":"2025","journal-title":"Ophthalmol Sci"},{"key":"B4","volume-title":"Machine Learning for Pseudopapilledema","author":"Kim","year":"2018"},{"key":"B5","doi-asserted-by":"publisher","first-page":"178","DOI":"10.1186\/s12886-019-1184-0","article-title":"Accuracy of machine learning for differentiation between optic neuropathies and pseudopapilledema","volume":"19","author":"Ahn","year":"2019","journal-title":"BMC Ophthalmol"},{"key":"B6","volume-title":"LLaMA 3.2-Vision: Multimodal Large Language Models","author":"AI","year":"2024"},{"key":"B7","article-title":"Visual instruction tuning","author":"Liu","year":"2023","journal-title":"arXiv"},{"key":"B8","article-title":"LLaVA-Med: training a large language-and-vision assistant for biomedicine in one day","author":"Li","year":"2023","journal-title":"arXiv"},{"key":"B9","article-title":"GPT-4o [large language model]","year":"2024"},{"key":"B10","article-title":"Deepseek-vl2: mixture-of-experts vision-language models for advanced multimodal understanding","author":"Wu","year":""},{"key":"B11","doi-asserted-by":"publisher","first-page":"100556","DOI":"10.1016\/j.xops.2024.100556","article-title":"Interpretation of clinical retinal images using an artificial intelligence Chatbot","volume":"4","author":"Mihalache","year":"2024","journal-title":"Ophthalmol Sci"},{"key":"B12","doi-asserted-by":"publisher","first-page":"e55318","DOI":"10.2196\/55318","article-title":"An empirical evaluation of prompting strategies for large language models in zero-shot clinical natural language processing: algorithm development and validation study","volume":"12","author":"Sivarajkumar","year":"2024","journal-title":"JMIR Med Inform"},{"key":"B13","article-title":"Effects of prompt length on domain-specific tasks for large language models","author":"Liu","year":""},{"key":"B14","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/2024.findings-emnlp.888","article-title":"When \u201ca helpful assistant\u201d is not really helpful: personas in system prompts do not improve performances of large language models","author":"Zheng","year":""}],"container-title":["Frontiers in Digital Health"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdgth.2025.1660887\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,25]],"date-time":"2025-08-25T05:26:53Z","timestamp":1756099613000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdgth.2025.1660887\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,25]]},"references-count":14,"alternative-id":["10.3389\/fdgth.2025.1660887"],"URL":"https:\/\/doi.org\/10.3389\/fdgth.2025.1660887","relation":{},"ISSN":["2673-253X"],"issn-type":[{"value":"2673-253X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,25]]},"article-number":"1660887"}}