{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,28]],"date-time":"2026-03-28T09:31:06Z","timestamp":1774690266297,"version":"3.50.1"},"reference-count":29,"publisher":"Oxford University Press (OUP)","issue":"3","license":[{"start":{"date-parts":[[2026,1,11]],"date-time":"2026-01-11T00:00:00Z","timestamp":1768089600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/pages\/standard-publication-reuse-rights"}],"funder":[{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Ministry of Science and ICT of Korea Government","award":["RS-2024-00335644"],"award-info":[{"award-number":["RS-2024-00335644"]}]},{"name":"Ministry of Science and ICT of Korea Government","award":["2025IP0016"],"award-info":[{"award-number":["2025IP0016"]}]},{"DOI":"10.13039\/501100005006","name":"Asan Institute for Life Sciences, Asan Medical Center, Republic of Korea","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100005006","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026,3,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Background<\/jats:title>\n                    <jats:p>Despite rapid integration into clinical decision-making, clinical large language models (LLMs) face substantial translational barriers due to insufficient structural characterization and limited external validation.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Objective<\/jats:title>\n                    <jats:p>We systematically map the clinical LLM research landscape to identify key structural patterns influencing their readiness for real-world clinical deployment.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Methods<\/jats:title>\n                    <jats:p>We identified 73 clinical LLM studies published between January 2020 and March 2025 using a structured evidence-mapping approach. To ensure transparency and reproducibility in study selection, we followed key principles from the PRISMA 2020 framework. Each study was categorized by clinical task, base architecture, alignment strategy, data type, language, study design, validation methods, and evaluation metrics.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>Studies often addressed multiple early stage clinical tasks\u2014question answering (56.2%), knowledge structuring (31.5%), and disease prediction (43.8%)\u2014primarily using text data (52.1%) and English-language resources (80.8%). GPT models favored retrieval-augmented generation (43.8%), and LLaMA models consistently adopted multistage pretraining and fine-tuning strategies. Only 6.9% of studies included external validation, and prospective designs were observed in just 4.1% of cases, reflecting significant gaps in translational reliability. Evaluations were predominantly quantitative only (79.5%), though qualitative and mixed-method approaches are increasingly recognized for assessing clinical usability and trustworthiness.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusion<\/jats:title>\n                    <jats:p>Clinical LLM research remains exploratory, marked by limited generalizability across languages, data types, and clinical environments. To bridge this gap, future studies must prioritize multilingual and multimodal training, prospective study designs with rigorous external validation, and hybrid evaluation frameworks combining quantitative performance with qualitative clinical usability metrics.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1093\/jamia\/ocaf230","type":"journal-article","created":{"date-parts":[[2025,12,16]],"date-time":"2025-12-16T12:42:46Z","timestamp":1765888966000},"page":"732-742","source":"Crossref","is-referenced-by-count":2,"title":["Structural insights into clinical large language models and their barriers to translational readiness"],"prefix":"10.1093","volume":"33","author":[{"given":"Jiwon","family":"You","sequence":"first","affiliation":[{"name":"Department of Medical Informatics and Statistics, Asan Medical Center, Brain Korea 21 Project, University of Ulsan College of Medicine , Seoul 05505,","place":["Republic of Korea"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3353-0310","authenticated-orcid":false,"given":"Hangsik","family":"Shin","sequence":"additional","affiliation":[{"name":"Department of Medical Informatics and Statistics, Asan Medical Center, Brain Korea 21 Project, University of Ulsan College of Medicine , Seoul 05505,","place":["Republic of Korea"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2026,1,11]]},"reference":[{"key":"2026031216473412800_ocaf230-B1","doi-asserted-by":"crossref","first-page":"689","DOI":"10.1186\/s12909-023-04698-z","article-title":"Revolutionizing healthcare: the role of artificial intelligence in clinical practice","volume":"23","author":"Alowais","year":"2023","journal-title":"BMC Med Educ"},{"key":"2026031216473412800_ocaf230-B2","doi-asserted-by":"crossref","first-page":"94","DOI":"10.7861\/futurehosp.6-2-94","article-title":"The potential for artificial intelligence in healthcare","volume":"6","author":"Davenport","year":"2019","journal-title":"Future Healthc J"},{"key":"2026031216473412800_ocaf230-B3","first-page":"22","author":"Suleimenov","year":"2020"},{"key":"2026031216473412800_ocaf230-B4","doi-asserted-by":"crossref","first-page":"128","DOI":"10.1055\/s-0039-1677903","article-title":"Artificial intelligence in clinical decision support: challenges for evaluating AI and practical implications","volume":"28","author":"Magrabi","year":"2019","journal-title":"Yearb Med Inform"},{"key":"2026031216473412800_ocaf230-B5","first-page":"71242","article-title":"Towards foundation models for scientific machine learning: characterizing scaling and transfer behavior","volume":"36","author":"Subramanian","year":"2023","journal-title":"Adv Neural Inf Process Syst"},{"key":"2026031216473412800_ocaf230-B6","first-page":"111301","author":"Chen","year":"2026"},{"key":"2026031216473412800_ocaf230-B7","first-page":"1","author":"Kang"},{"key":"2026031216473412800_ocaf230-B8","doi-asserted-by":"crossref","first-page":"172","DOI":"10.1109\/RBME.2024.3496744","article-title":"Foundation model for advancing healthcare: challenges, opportunities and future directions","volume":"18","author":"He","journal-title":"IEEE Rev Biomed Eng"},{"key":"2026031216473412800_ocaf230-B9","doi-asserted-by":"crossref","first-page":"3046\u2013","DOI":"10.1016\/j.acra.2024.04.006","article-title":"Performance of GPT-4 on the American College of Radiology In-training Examination: evaluating accuracy, model drift, and fine-tuning","volume":"31","author":"Payne","year":"2024","journal-title":"Acad Radiol"},{"key":"2026031216473412800_ocaf230-B10","doi-asserted-by":"crossref","first-page":"e0000198","DOI":"10.1371\/journal.pdig.0000198","article-title":"Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models","volume":"2","author":"Kung","year":"2023","journal-title":"PLoS Digit Health"},{"key":"2026031216473412800_ocaf230-B11","author":"Tan","year":"2024"},{"key":"2026031216473412800_ocaf230-B12","article-title":"Me-LLaMA: foundation large language models for medical applications","author":"Xie","year":"2024","journal-title":"Res Sq"},{"key":"2026031216473412800_ocaf230-B13","doi-asserted-by":"crossref","first-page":"943","DOI":"10.1038\/s41591-024-03423-7","article-title":"Toward expert-level medical question answering with large language models","volume":"31","author":"Singhal","year":"2025","journal-title":"Nat Med"},{"key":"2026031216473412800_ocaf230-B14","author":"Saab","year":"2024"},{"key":"2026031216473412800_ocaf230-B15","doi-asserted-by":"crossref","first-page":"e60695","DOI":"10.2196\/60695","article-title":"Performance of retrieval-augmented large language models to recommend head and neck cancer clinical trials","volume":"26","author":"Hung","year":"2024","journal-title":"J Med Internet Res."},{"key":"2026031216473412800_ocaf230-B16","first-page":"120\u2013","article-title":"VetLLM: large language model for predicting diagnosis from veterinary notes","volume":"29","author":"Jiang","year":"2024","journal-title":"Pac Symp Biocomput"},{"key":"2026031216473412800_ocaf230-B17","doi-asserted-by":"crossref","first-page":"6029\u2013","DOI":"10.1109\/JBHI.2023.3315143","article-title":"TeaBERT: an efficient knowledge infused cross-lingual language model for mapping Chinese medical entities to the unified medical language system","volume":"27","author":"Chen","year":"2023","journal-title":"IEEE J Biomed Health Inform"},{"key":"2026031216473412800_ocaf230-B18","doi-asserted-by":"crossref","first-page":"1929\u2013","DOI":"10.1093\/jamia\/ocae095","article-title":"RT: a Retrieving and Chain-of-Thought framework for few-shot medical named entity recognition","volume":"31","author":"Li","year":"2024","journal-title":"J Am Med Inform Assoc"},{"key":"2026031216473412800_ocaf230-B19","doi-asserted-by":"crossref","first-page":"1833\u2013","DOI":"10.1093\/jamia\/ocae045","article-title":"PMC-LLaMA: toward building open-source language models for medicine","volume":"31","author":"Wu","year":"2024","journal-title":"J Am Med Inform Assoc"},{"key":"2026031216473412800_ocaf230-B20","doi-asserted-by":"crossref","first-page":"104796","DOI":"10.1016\/j.jbi.2025.104796","article-title":"Missing-modality enabled multi-modal fusion architecture for medical data","volume":"164","author":"Wang","year":"2025","journal-title":"J Biomed Inform"},{"key":"2026031216473412800_ocaf230-B21","author":"Blankemeier","year":"2024"},{"key":"2026031216473412800_ocaf230-B22","doi-asserted-by":"crossref","first-page":"btad651","DOI":"10.1093\/bioinformatics\/btad651","article-title":"Medcpt: contrastive pre-trained transformers with large-scale PubMed search logs for zero-shot biomedical information retrieval","volume":"39","author":"Jin","year":"2023","journal-title":"Bioinformatics"},{"key":"2026031216473412800_ocaf230-B23","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1038\/s41746-021-00455-y","article-title":"Med-BERT: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction","volume":"4","author":"Rasmy","year":"2021","journal-title":"NPJ Digit Med"},{"key":"2026031216473412800_ocaf230-B24","doi-asserted-by":"crossref","first-page":"194","DOI":"10.1038\/s41746-022-00742-2","article-title":"A large language model for electronic health records","volume":"5","author":"Yang","year":"2022","journal-title":"NPJ Digit Med"},{"key":"2026031216473412800_ocaf230-B25","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1038\/s41746-024-01101-z","article-title":"FFA-GPT: an automated pipeline for fundus fluorescein angiography interpretation and question-answer","volume":"7","author":"Chen","year":"2024","journal-title":"NPJ Digit Med"},{"key":"2026031216473412800_ocaf230-B26","first-page":"1\u2013","article-title":"COBERT: COVID-19 question answering system using BERT","volume":"48","author":"Alzubi","year":"2021","journal-title":"Arab J Sci Eng"},{"key":"2026031216473412800_ocaf230-B27","first-page":"e40895","article-title":"Chatdoctor: a medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge","volume":"15","author":"Li","year":"2023","journal-title":"Cureus"},{"key":"2026031216473412800_ocaf230-B28","doi-asserted-by":"crossref","first-page":"133","DOI":"10.1038\/s44172-024-00271-8","article-title":"Interactive computer-aided diagnosis on medical image using large language models","volume":"3","author":"Wang","year":"2024","journal-title":"Commun Eng"},{"key":"2026031216473412800_ocaf230-B29","doi-asserted-by":"crossref","first-page":"1208\u2013","DOI":"10.1093\/jamia\/ocac040","article-title":"CancerBERT: a cancer domain-specific language model for extracting breast cancer phenotypes from electronic health records","volume":"29","author":"Zhou","year":"2022","journal-title":"J Am Med Inform Assoc"}],"container-title":["Journal of the American Medical Informatics Association"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jamia\/advance-article-pdf\/doi\/10.1093\/jamia\/ocaf230\/66341939\/ocaf230.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/33\/3\/732\/66341939\/ocaf230.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/33\/3\/732\/66341939\/ocaf230.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,12]],"date-time":"2026-03-12T20:47:41Z","timestamp":1773348461000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jamia\/article\/33\/3\/732\/8419921"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,11]]},"references-count":29,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2026,1,11]]},"published-print":{"date-parts":[[2026,3,1]]}},"URL":"https:\/\/doi.org\/10.1093\/jamia\/ocaf230","relation":{},"ISSN":["1067-5027","1527-974X"],"issn-type":[{"value":"1067-5027","type":"print"},{"value":"1527-974X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026,3]]},"published":{"date-parts":[[2026,1,11]]}}}