{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,26]],"date-time":"2026-02-26T20:09:06Z","timestamp":1772136546252,"version":"3.50.1"},"reference-count":25,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T00:00:00Z","timestamp":1763337600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T00:00:00Z","timestamp":1763337600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["npj Digit. Med."],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>This study evaluates the diagnostic performance of commercial and open-source Vision-Language Models (VLMs) in neuroradiological image interpretation, using a dataset of 100 brain and spine cases from Radiopaedia. Five VLMs (Gemini 2.0, OpenAI o1, Llama 3.2 90b, Qwen 2.5, Grok-2-vision) were compared to expert neuroradiologists in generating differential diagnoses based on brief clinical presentations and imaging. Neuroradiologists achieved a mean accuracy of 86.2%, whereas the best-performing VLM (Gemini 2.0) reached 35%. Evaluation of the top three differentials improved VLM accuracy marginally, but remained inferior to human experts. Clinical harm analysis revealed frequent diagnostic risks, primarily treatment delays, with harmful outputs in up to 45% of cases. Error analysis showed consistent failure modes including incorrect anatomical localization, inaccurate imaging descriptions, and hallucinated findings. These results highlight the current limitations of VLMs and underscore the importance of expert oversight in neuroradiological diagnosis.<\/jats:p>","DOI":"10.1038\/s41746-025-02047-6","type":"journal-article","created":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T21:43:19Z","timestamp":1763415799000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Evaluating the diagnostic accuracy of vision language models for neuroradiological image interpretation"],"prefix":"10.1038","volume":"8","author":[{"given":"Aymen","family":"Meddeb","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ida","family":"Rangus","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Paolo","family":"Pagano","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Insaf","family":"Dkhil","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Soumaya","family":"Jelassi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Keno","family":"Bressem","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael","family":"Scheel","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mike P.","family":"Wattjes","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sonia","family":"Nagi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Laurent","family":"Pierot","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sebastien","family":"Soize","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,11,17]]},"reference":[{"key":"2047_CR1","doi-asserted-by":"publisher","unstructured":"Zhang, H. et al. Evaluating Large Language Models in Extracting Cognitive Exam Dates and Scores. medRxiv 2024, https:\/\/doi.org\/10.1101\/2023.07.10.23292373.","DOI":"10.1101\/2023.07.10.23292373"},{"key":"2047_CR2","doi-asserted-by":"publisher","first-page":"33","DOI":"10.1007\/s10916-023-01925-4","volume":"47","author":"M Cascella","year":"2023","unstructured":"Cascella, M., Montomoli, J., Bellini, V. & Bignami, E. Evaluating the Feasibility of ChatGPT in Healthcare: An Analysis of Multiple Clinical and Research Scenarios. J. Me\u0301d. Syst. 47, 33 (2023).","journal-title":"J. Me\u0301d. Syst."},{"key":"2047_CR3","doi-asserted-by":"publisher","first-page":"6","DOI":"10.1038\/s41746-023-00970-0","volume":"7","author":"M Guevara","year":"2024","unstructured":"Guevara, M. et al. Large Language Models to Identify Social Determinants of Health in Electronic Health Records. npj Digit. Med. 7, 6 (2024).","journal-title":"npj Digit. Med."},{"key":"2047_CR4","unstructured":"Meddeb, A. et al. Evaluating Local Open-Source Large Language Models for Data Extraction from Unstructured Reports on Mechanical Thrombectomy in Patients with Ischemic Stroke. J. NeuroInterventional Surg. 2024, jnis-2024-022078."},{"key":"2047_CR5","doi-asserted-by":"publisher","first-page":"1134","DOI":"10.1038\/s41591-024-02855-5","volume":"30","author":"DV Veen","year":"2024","unstructured":"Veen, D. V. et al. Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization. Nat. Med. 30, 1134\u20131142 (2024).","journal-title":"Nat. Med."},{"key":"2047_CR6","doi-asserted-by":"publisher","DOI":"10.1148\/radiol.241736","volume":"313","author":"A Meddeb","year":"2024","unstructured":"Meddeb, A. et al. Large Language Model Ability to Translate CT and MRI Free-Text Radiology Reports Into Multiple Languages. Radiology 313, e241736 https:\/\/doi.org\/10.1148\/radiol.241736 (2024).","journal-title":"Radiology"},{"key":"2047_CR7","doi-asserted-by":"publisher","unstructured":"Mohamad, F. A. et al. Open-Source Large Language Models Can Generate Labels from Radiology Reports for Training Convolutional Neural Networks. Acad. Radiol. 2025, https:\/\/doi.org\/10.1016\/j.acra.2024.12.028.","DOI":"10.1016\/j.acra.2024.12.028"},{"key":"2047_CR8","doi-asserted-by":"publisher","DOI":"10.1148\/radiol.230725","volume":"307","author":"LC Adams","year":"2023","unstructured":"Adams, L. C. et al. Leveraging GPT-4 for Post Hoc Transformation of Free-Text Radiology Reports into Structured Reporting: A Multilingual Feasibility Study. Radiology 307, e230725 https:\/\/doi.org\/10.1148\/radiol.230725 (2023).","journal-title":"Radiology"},{"key":"2047_CR9","doi-asserted-by":"publisher","first-page":"e2346721","DOI":"10.1001\/jamanetworkopen.2023.46721","volume":"6","author":"MC Schubert","year":"2023","unstructured":"Schubert, M. C., Wick, W. & Venkataramani, V. Performance of Large Language Models on a Neurology Board\u2013Style Examination. JAMA Netw. Open 6, e2346721 (2023).","journal-title":"JAMA Netw. Open"},{"key":"2047_CR10","doi-asserted-by":"publisher","unstructured":"Shu, L. et al. Large Language Model Performance in Neurology Board Questions (S33.001). Neurology 2024, 102, https:\/\/doi.org\/10.1212\/wnl.0000000000204763.","DOI":"10.1212\/wnl.0000000000204763"},{"key":"2047_CR11","doi-asserted-by":"publisher","first-page":"85","DOI":"10.1038\/s41746-025-01486-5","volume":"8","author":"X Yang","year":"2025","unstructured":"Yang, X. et al. Multiple Large Language Models versus Experienced Physicians in Diagnosing Challenging Cases with Gastrointestinal Symptoms. npj Digit. Med. 8, 85 (2025).","journal-title":"npj Digit. Med."},{"key":"2047_CR12","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1016\/j.ejim.2024.09.017","volume":"131","author":"AG Levra","year":"2025","unstructured":"Levra, A. G. et al. A Large Language Model-Based Clinical Decision Support System for Syncope Recognition in the Emergency Department: A Framework for Clinical Workflow Integration. Eur. J. Intern. Med. 131, 113\u2013120 (2025).","journal-title":"Eur. J. Intern. Med."},{"key":"2047_CR13","doi-asserted-by":"publisher","unstructured":"Ong, J. C. L. et al. Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties. arXiv 2024, https:\/\/doi.org\/10.48550\/arxiv.2402.01741.","DOI":"10.48550\/arxiv.2402.01741"},{"key":"2047_CR14","doi-asserted-by":"publisher","unstructured":"Bordes, F. et al. An Introduction to Vision-Language Modeling. arXiv 2024, https:\/\/doi.org\/10.48550\/arxiv.2405.17247.","DOI":"10.48550\/arxiv.2405.17247"},{"key":"2047_CR15","doi-asserted-by":"publisher","unstructured":"Yan, Z. et al. Multimodal ChatGPT for Medical Applications: An Experimental Study of GPT-4V. arXiv 2023, https:\/\/doi.org\/10.48550\/arxiv.2310.19061.","DOI":"10.48550\/arxiv.2310.19061"},{"key":"2047_CR16","doi-asserted-by":"crossref","unstructured":"Brin, D. et al. Assessing GPT-4 Multimodal Performance in Radiological Image Analysis. Eur. Radiol. 2024, 1\u20137.","DOI":"10.1101\/2023.11.15.23298583"},{"key":"2047_CR17","doi-asserted-by":"publisher","DOI":"10.1148\/radiol.230582","volume":"307","author":"R Bhayana","year":"2023","unstructured":"Bhayana, R., Krishna, S. & Bleakney, R. R. Performance of ChatGPT on a Radiology Board-Style Examination: Insights into Current Strengths and Limitations. Radiology 307, e230582 https:\/\/doi.org\/10.1148\/radiol.230582 (2023).","journal-title":"Radiology"},{"key":"2047_CR18","doi-asserted-by":"publisher","first-page":"190","DOI":"10.1038\/s41746-024-01185-7","volume":"7","author":"Q Jin","year":"2024","unstructured":"Jin, Q. et al. Hidden Flaws behind Expert-Level Accuracy of Multimodal GPT-4 Vision in Medicine. npj Digit. Med. 7, 190 (2024).","journal-title":"npj Digit. Med."},{"key":"2047_CR19","doi-asserted-by":"publisher","first-page":"e54948","DOI":"10.2196\/54948","volume":"26","author":"F Busch","year":"2024","unstructured":"Busch, F. et al. Integrating Text and Image Analysis: Exploring GPT-4V\u2019s Capabilities in Advanced Radiological Applications Across Subspecialties. J. Me\u0301d. Internet Res. 26, e54948 (2024).","journal-title":"J. Me\u0301d. Internet Res."},{"key":"2047_CR20","doi-asserted-by":"publisher","DOI":"10.1148\/radiol.240273","volume":"312","author":"PS Suh","year":"2024","unstructured":"Suh, P. S. et al. Comparing Diagnostic Accuracy of Radiologists versus GPT-4V and Gemini Pro Vision Using Image Inputs from Diagnosis Please Cases. Radiology 312, e240273 https:\/\/doi.org\/10.1148\/radiol.240273 (2024).","journal-title":"Radiology"},{"key":"2047_CR21","doi-asserted-by":"publisher","unstructured":"Wu, C. et al. Can GPT-4V(Ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis. arXiv 2023, https:\/\/doi.org\/10.48550\/arxiv.2310.09909.","DOI":"10.48550\/arxiv.2310.09909"},{"key":"2047_CR22","doi-asserted-by":"publisher","DOI":"10.1148\/radiol.240689","volume":"314","author":"S Schramm","year":"2025","unstructured":"Schramm, S. et al. Impact of Multimodal Prompt Elements on Diagnostic Performance of GPT-4V in Challenging Brain MRI Cases. Radiology 314, e240689 https:\/\/doi.org\/10.1148\/radiol.240689 (2025).","journal-title":"Radiology"},{"key":"2047_CR23","doi-asserted-by":"publisher","first-page":"1231","DOI":"10.1007\/s11604-024-01619-y","volume":"42","author":"Y Sonoda","year":"2024","unstructured":"Sonoda, Y. et al. Diagnostic Performances of GPT-4o, Claude 3 Opus, and Gemini 1.5 Pro in \u201cDiagnosis Please\u201d Cases. Jpn. J. Radiol. 42, 1231\u20131235 (2024).","journal-title":"Jpn. J. Radiol."},{"key":"2047_CR24","doi-asserted-by":"publisher","DOI":"10.1148\/radiol.231040","volume":"308","author":"D Ueda","year":"2023","unstructured":"Ueda, D. et al. Diagnostic Performance of ChatGPT from Patient History and Imaging Findings on the Diagnosis Please Quizzes. Radiology 308, e231040 https:\/\/doi.org\/10.1148\/radiol.231040 (2023).","journal-title":"Radiology"},{"key":"2047_CR25","doi-asserted-by":"publisher","first-page":"37","DOI":"10.3174\/ajnr.A2704","volume":"33","author":"LS Babiarz","year":"2012","unstructured":"Babiarz, L. S. & Yousem, D. M. Quality Control in Neuroradiology: Discrepancies in Image Interpretation among Academic Neuroradiologists. Am. J. Neuroradiol. 33, 37\u201342 (2012).","journal-title":"Am. J. Neuroradiol."}],"container-title":["npj Digital Medicine"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.nature.com\/articles\/s41746-025-02047-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41746-025-02047-6","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41746-025-02047-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T21:43:21Z","timestamp":1763415801000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.nature.com\/articles\/s41746-025-02047-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,17]]},"references-count":25,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["2047"],"URL":"https:\/\/doi.org\/10.1038\/s41746-025-02047-6","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-6183659\/v1","asserted-by":"object"}]},"ISSN":["2398-6352"],"issn-type":[{"value":"2398-6352","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,17]]},"assertion":[{"value":"8 March 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 September 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 November 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"666"}}