{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T14:53:55Z","timestamp":1785336835293,"version":"3.55.0"},"reference-count":62,"publisher":"Oxford University Press (OUP)","issue":"4","license":[{"start":{"date-parts":[[2025,1,15]],"date-time":"2025-01-15T00:00:00Z","timestamp":1736899200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["R00LM014097-02"],"award-info":[{"award-number":["R00LM014097-02"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["R01LM013995-01"],"award-info":[{"award-number":["R01LM013995-01"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Objective<\/jats:title>\n                  <jats:p>The objectives of this study are to synthesize findings from recent research of retrieval-augmented generation (RAG) and large language models (LLMs) in biomedicine and provide clinical development guidelines to improve effectiveness.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Materials and Methods<\/jats:title>\n                  <jats:p>We conducted a systematic literature review and a meta-analysis. The report was created in adherence to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 analysis. Searches were performed in 3 databases (PubMed, Embase, PsycINFO) using terms related to \u201cretrieval augmented generation\u201d and \u201clarge language model,\u201d for articles published in 2023 and 2024. We selected studies that compared baseline LLM performance with RAG performance. We developed a random-effect meta-analysis model, using odds ratio as the effect size.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>Among 335 studies, 20 were included in this literature review. The pooled effect size was 1.35, with a 95% confidence interval of 1.19-1.53, indicating a statistically significant effect (P\u2009=\u2009.001). We reported clinical tasks, baseline LLMs, retrieval sources and strategies, as well as evaluation methods.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Discussion<\/jats:title>\n                  <jats:p>Building on our literature review, we developed Guidelines for Unified Implementation and Development of Enhanced LLM Applications with RAG in Clinical Settings to inform clinical applications using RAG.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Conclusion<\/jats:title>\n                  <jats:p>Overall, RAG implementation showed a 1.35 odds ratio increase in performance compared to baseline LLMs. Future research should focus on (1) system-level enhancement: the combination of RAG and agent, (2) knowledge-level enhancement: deep integration of knowledge into LLM, and (3) integration-level enhancement: integrating RAG systems within electronic health records.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/jamia\/ocaf008","type":"journal-article","created":{"date-parts":[[2025,1,15]],"date-time":"2025-01-15T16:07:57Z","timestamp":1736957277000},"page":"605-615","source":"Crossref","is-referenced-by-count":160,"title":["Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines"],"prefix":"10.1093","volume":"32","author":[{"given":"Siru","family":"Liu","sequence":"first","affiliation":[{"name":"Department of Biomedical Informatics, Vanderbilt University Medical Center , Nashville, TN 37212,","place":["United States"]},{"name":"Department of Computer Science , Vanderbilt University, Nashville, TN 37212,","place":["United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2292-9147","authenticated-orcid":false,"given":"Allison B","family":"McCoy","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Vanderbilt University Medical Center , Nashville, TN 37212,","place":["United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6844-145X","authenticated-orcid":false,"given":"Adam","family":"Wright","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Vanderbilt University Medical Center , Nashville, TN 37212,","place":["United States"]},{"name":"Department of Medicine, Vanderbilt University Medical Center , Nashville, TN 37212,","place":["United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2025,1,15]]},"reference":[{"key":"2025041716422184200_ocaf008-B1","doi-asserted-by":"publisher","first-page":"26839","DOI":"10.1109\/ACCESS.2024.3365742","article-title":"A review on large language models: architectures, applications, taxonomies, open issues and challenges","volume":"12","author":"Raiaan","year":"2024","journal-title":"IEEE Access"},{"key":"2025041716422184200_ocaf008-B2","doi-asserted-by":"publisher","first-page":"1930","DOI":"10.1038\/s41591-023-02448-8","article-title":"Large language models in medicine","volume":"29","author":"Thirunavukarasu","year":"2023","journal-title":"Nat Med"},{"key":"2025041716422184200_ocaf008-B3","doi-asserted-by":"publisher","first-page":"589","DOI":"10.1001\/jamainternmed.2023.1838","article-title":"Comparing physician and artificial intelligence Chatbot responses to patient questions posted to a public social media forum","volume":"183","author":"Ayers","year":"2023","journal-title":"JAMA Intern Med"},{"key":"2025041716422184200_ocaf008-B4","doi-asserted-by":"publisher","first-page":"1237","DOI":"10.1093\/jamia\/ocad072","article-title":"Using AI-generated suggestions from ChatGPT to optimize clinical decision support","volume":"30","author":"Liu","year":"2023","journal-title":"J Am Med Inform Assoc"},{"key":"2025041716422184200_ocaf008-B5","doi-asserted-by":"publisher","first-page":"e240357","DOI":"10.1001\/jamanetworkopen.2024.0357","article-title":"Generative artificial intelligence to transform inpatient discharge summaries to patient-friendly language and format","volume":"7","author":"Zaretsky","year":"2024","journal-title":"JAMA Netw Open"},{"key":"2025041716422184200_ocaf008-B6","article-title":"Retrieval-augmented generation for large language models: a survey","author":"Gao","year":"2023"},{"key":"2025041716422184200_ocaf008-B7","author":"Xu","year":"2024"},{"key":"2025041716422184200_ocaf008-B8","first-page":"3784","author":"Shuster"},{"key":"2025041716422184200_ocaf008-B9","doi-asserted-by":"publisher","first-page":"228","DOI":"10.18653\/v1\/2024.naacl-industry.19","article-title":"Reducing hallucination in structured outputs via Retrieval-Augmented Generation","author":"Ayala","year":"2024"},{"key":"2025041716422184200_ocaf008-B10","doi-asserted-by":"publisher","first-page":"e1313","DOI":"10.1161\/CIR.0000000000001251","article-title":"2024 ACC\/AHA\/AACVPR\/APMA\/ABC\/SCAI\/SVM\/SVN\/SVS\/SIR\/VESS guideline for the management of lower extremity peripheral artery disease: a report of the American College of Cardiology\/American Heart Association Joint Committee on Clinical Practice Guidelines","volume":"149","author":"Gornik","year":"2024","journal-title":"Circulation"},{"key":"2025041716422184200_ocaf008-B11","doi-asserted-by":"publisher","first-page":"e070904","DOI":"10.1136\/bmj-2022-070904","article-title":"Reporting guideline for the early stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI","volume":"377","author":"Vasey","year":"2022","journal-title":"BMJ"},{"key":"2025041716422184200_ocaf008-B12","doi-asserted-by":"publisher","first-page":"e200029","DOI":"10.1148\/ryai.2020200029","article-title":"Checklist for Artificial Intelligence in Medical Imaging (CLAIM): a guide for authors and reviewers","volume":"2","author":"Mongan","year":"2020","journal-title":"Radiol Artif Intell"},{"key":"2025041716422184200_ocaf008-B13","doi-asserted-by":"publisher","first-page":"6376","DOI":"10.1038\/s41467-024-45355-3","article-title":"Concordance of randomised controlled trials for artificial intelligence interventions with the CONSORT-AI reporting guidelines","volume":"15","author":"Martindale","year":"2024","journal-title":"Nat Commun"},{"key":"2025041716422184200_ocaf008-B14","doi-asserted-by":"publisher","first-page":"258","DOI":"10.1038\/s41746-024-01258-7","article-title":"A framework for human evaluation of large language models in healthcare derived from literature review","volume":"7","author":"Tam","year":"2024","journal-title":"NPJ Digit Med"},{"key":"2025041716422184200_ocaf008-B15","doi-asserted-by":"publisher","first-page":"g7647","DOI":"10.1136\/bmj.g7647","article-title":"Preferred reporting items for systematic review and meta-analysis protocols (PRISMA-p) 2015: elaboration and explanation","volume":"350","author":"Shamseer","year":"2015","journal-title":"BMJ"},{"key":"2025041716422184200_ocaf008-B16","author":"Higgins","year":"2024"},{"key":"2025041716422184200_ocaf008-B17","article-title":"Chapter 4: searching for and selecting studies","volume":"6","author":"Lefebvre","journal-title":"Cochrane Handbook for Systematic Reviews of Interventions Version, Vol."},{"key":"2025041716422184200_ocaf008-B18","author":"Chapter 3 Effect Sizes"},{"key":"2025041716422184200_ocaf008-B19","volume-title":"Introduction to Meta-Analysis","author":"Borenstein","year":"2011"},{"key":"2025041716422184200_ocaf008-B20","doi-asserted-by":"publisher","first-page":"1539","DOI":"10.1002\/sim.1186","article-title":"Quantifying heterogeneity in a meta-analysis","volume":"21","author":"Higgins","year":"2002","journal-title":"Stat Med"},{"key":"2025041716422184200_ocaf008-B21","doi-asserted-by":"publisher","first-page":"991","DOI":"10.1016\/j.jclinepi.2007.11.010","article-title":"Contour-enhanced meta-analysis funnel plots help distinguish publication bias from other causes of asymmetry","volume":"61","author":"Peters","year":"2008","journal-title":"J Clin Epidemiol"},{"key":"2025041716422184200_ocaf008-B22","doi-asserted-by":"publisher","first-page":"629","DOI":"10.1136\/bmj.315.7109.629","article-title":"Bias in meta-analysis detected by a simple, graphical test measures of funnel plot asymmetry","volume":"315","author":"Egger","year":"1997","journal-title":"BMJ"},{"key":"2025041716422184200_ocaf008-B23","doi-asserted-by":"publisher","first-page":"983","DOI":"10.3233\/SHTI240575","article-title":"Using retrieval-augmented generation to capture molecularly-driven treatment relationships for precision oncology","volume":"316","author":"Kreimeyer","year":"2024","journal-title":"Stud Health Technol Inform"},{"key":"2025041716422184200_ocaf008-B24","doi-asserted-by":"publisher","first-page":"1356","DOI":"10.1093\/jamia\/ocae039","article-title":"Empowering personalized pharmacogenomics with generative AI solutions","volume":"31","author":"Murugan","year":"2024","journal-title":"J Am Med Inform Assoc"},{"key":"2025041716422184200_ocaf008-B25","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1080\/10903127.2024.2374400","article-title":"Emergency patient triage improvement through a retrieval-augmented generation enhanced large-scale language model","author":"Yazaki","year":"2024","journal-title":"Prehosp Emerg Care"},{"key":"2025041716422184200_ocaf008-B26","doi-asserted-by":"publisher","first-page":"514","DOI":"10.20524\/aog.2024.0907","article-title":"Assessing ChatGPT4 with and without retrieval-augmented generation in anticoagulation management for gastrointestinal procedures","volume":"37","author":"Malik","year":"2024","journal-title":"Ann Gastroenterol"},{"key":"2025041716422184200_ocaf008-B27","doi-asserted-by":"publisher","first-page":"102","DOI":"10.1038\/s41746-024-01091-y","article-title":"Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework","volume":"7","author":"Kresevic","year":"2024","journal-title":"NPJ Digit Med"},{"key":"2025041716422184200_ocaf008-B28","doi-asserted-by":"publisher","DOI":"10.1056\/aioa2300068","article-title":"Almanac\u2013retrieval-augmented language models for clinical medicine","author":"Zakka","year":"2024","journal-title":"NEJM AI"},{"key":"2025041716422184200_ocaf008-B29","doi-asserted-by":"publisher","first-page":"1042","DOI":"10.1002\/ohn.864","article-title":"ChatENT: augmented large language model for expert knowledge retrieval in otolaryngology\u2013head and neck surgery","volume":"171","author":"Long","year":"2024","journal-title":"Otolaryngol Head Neck Surg"},{"key":"2025041716422184200_ocaf008-B30","doi-asserted-by":"publisher","first-page":"e58041","DOI":"10.2196\/58041","article-title":"Enhancement of the performance of large language models in diabetes education through retrieval-augmented generation: comparative study","volume":"26","author":"Wang","year":"2024","journal-title":"J Med Internet Res"},{"key":"2025041716422184200_ocaf008-B31","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1186\/s41747-024-00457-x","article-title":"A retrieval-augmented chatbot based on GPT-4 provides appropriate differential diagnosis in gastrointestinal radiology: a proof of concept study","volume":"8","author":"Rau","year":"2024","journal-title":"Eur Radiol Exp"},{"key":"2025041716422184200_ocaf008-B32","doi-asserted-by":"publisher","DOI":"10.1093\/BIOINFORMATICS\/BTAD080","article-title":"The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information","volume":"39","author":"Morris","year":"2023","journal-title":"Bioinformatics"},{"key":"2025041716422184200_ocaf008-B33","doi-asserted-by":"publisher","first-page":"7","DOI":"10.1145\/3606337","article-title":"Biomedical knowledge graph-optimized prompt generation for large language models","volume":"66","author":"Soman","year":"2023","journal-title":"Commun ACM"},{"key":"2025041716422184200_ocaf008-B34","doi-asserted-by":"publisher","first-page":"i119","DOI":"10.1093\/bioinformatics\/btae238","article-title":"Improving medical reasoning through retrieval and self-reflection with retrieval-augmented large language models","volume":"40","author":"Jeong","year":"2024","journal-title":"Bioinformatics"},{"key":"2025041716422184200_ocaf008-B35","doi-asserted-by":"publisher","first-page":"104662","DOI":"10.1016\/j.jbi.2024.104662","article-title":"Applying generative AI with retrieval augmented generation to summarize and extract key clinical information from electronic health records","volume":"156","author":"Alkhalaf","year":"2024","journal-title":"J Biomed Inform"},{"key":"2025041716422184200_ocaf008-B36","doi-asserted-by":"publisher","first-page":"e0000604","DOI":"10.1371\/journal.pdig.0000604","article-title":"Performance of publicly available large language models on internal medicine board-style questions","volume":"3","author":"Tarabanis","year":"2024","journal-title":"PLOS Digit Health"},{"key":"2025041716422184200_ocaf008-B37","doi-asserted-by":"publisher","first-page":"1921","DOI":"10.1093\/jamia\/ocae103","article-title":"Evaluating the accuracy of a state-of-the-art large language model for prediction of admissions from the emergency room","volume":"31","author":"Glicksberg","year":"2024","journal-title":"J Am Med Inform Assoc"},{"key":"2025041716422184200_ocaf008-B38","doi-asserted-by":"publisher","first-page":"104702","DOI":"10.1016\/j.jbi.2024.104702","article-title":"Rare disease diagnosis using knowledge guided retrieval augmentation for ChatGPT","volume":"157","author":"Zelin","year":"2024","journal-title":"J Biomed Inform"},{"key":"2025041716422184200_ocaf008-B39","doi-asserted-by":"publisher","first-page":"e58158","DOI":"10.2196\/58158","article-title":"Evaluating and enhancing large language models\u2019 performance in domain-specific medicine: development and usability study with DocOA","volume":"26","author":"Chen","year":"2024","journal-title":"J Med Internet Res"},{"key":"2025041716422184200_ocaf008-B40","doi-asserted-by":"publisher","first-page":"105401","DOI":"10.1101\/2024.04.03.24305298","article-title":"Enhancing early detection of cognitive decline in the elderly: a comparative study utilizing large language models in clinical notes","volume":"109","author":"Du","year":"2024","journal-title":"medRxiv"},{"key":"2025041716422184200_ocaf008-B41","author":"Zhang","year":"2023"},{"key":"2025041716422184200_ocaf008-B42","author":"Li","year":"2024"},{"key":"2025041716422184200_ocaf008-B43","author":"Xiong"},{"key":"2025041716422184200_ocaf008-B44","doi-asserted-by":"publisher","first-page":"e70009","DOI":"10.1002\/2056-4538.70009","article-title":"Large language models as a diagnostic support tool in neuropathology","volume":"10","author":"Hewitt","year":"2024","journal-title":"J Pathol Clin Res"},{"key":"2025041716422184200_ocaf008-B45","author":"Allahverdiyev","year":"2024"},{"key":"2025041716422184200_ocaf008-B46","doi-asserted-by":"publisher","first-page":"1019","DOI":"10.1007\/978-3-642-31698-2_144","volume-title":"Advances in Intelligent Systems and Computing, Vol.","author":"Cai","year":"2013"},{"key":"2025041716422184200_ocaf008-B47","author":"Optimizing RAG with Advanced Chunking Techniques","year":""},{"key":"2025041716422184200_ocaf008-B48","doi-asserted-by":"publisher","first-page":"2318","DOI":"10.18653\/V1\/2024.FINDINGS-ACL.137","article-title":"M3-Embedding: multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation","author":"Chen","year":"2024","journal-title":"Findings of the Association for Computational Linguistics ACL 2024"},{"key":"2025041716422184200_ocaf008-B49","doi-asserted-by":"publisher","author":"Sawarkar","DOI":"10.1109\/MIPR62202.2024.00031"},{"key":"2025041716422184200_ocaf008-B50","author":"Edge","year":"2024"},{"key":"2025041716422184200_ocaf008-B51","doi-asserted-by":"publisher","first-page":"353","DOI":"10.18653\/v1\/2024.clinicalnlp-1.33","author":"Wu","year":"2024"},{"key":"2025041716422184200_ocaf008-B52","first-page":"18417","author":"Kwon","year":"2024"},{"key":"2025041716422184200_ocaf008-B53","doi-asserted-by":"publisher","first-page":"105401","DOI":"10.1016\/j.ebiom.2024.105401","article-title":"Enhancing early detection of cognitive decline in the elderly: a comparative study utilizing large language models in clinical notes","volume":"109","author":"Du","year":"2024","journal-title":"EBioMedicine"},{"key":"2025041716422184200_ocaf008-B54","doi-asserted-by":"publisher","first-page":"e214622","DOI":"10.1001\/jamanetworkopen.2021.4622","article-title":"Algorithmovigilance\u2014advancing methods to analyze and monitor artificial intelligence\u2013driven health care for effectiveness and equity","volume":"4","author":"Embi","year":"2021","journal-title":"JAMA Netw Open"},{"key":"2025041716422184200_ocaf008-B55","author":"Xi","year":"2023"},{"key":"2025041716422184200_ocaf008-B56","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/S11704-024-40231-1\/METRICS","article-title":"A survey on large language model based autonomous agents","volume":"18","author":"Wang","year":"2024","journal-title":"Front Comput Sci"},{"key":"2025041716422184200_ocaf008-B57","doi-asserted-by":"publisher","first-page":"9","DOI":"10.1007\/s44336-024-00009-2","article-title":"A survey on LLM-based multi-agent systems: workflow, infrastructure, and challenges","volume":"1","author":"Li","year":"2024","journal-title":"Vicinagearth"},{"key":"2025041716422184200_ocaf008-B58","doi-asserted-by":"publisher","DOI":"10.1109\/JBHI.2024.3464555","article-title":"RDguru: a conversational intelligent agent for rare diseases","author":"Yang","year":"19, 2024.","journal-title":"IEEE J Biomed Health Inform"},{"key":"2025041716422184200_ocaf008-B59","author":"Ren","year":"2023"},{"key":"2025041716422184200_ocaf008-B60","doi-asserted-by":"publisher","first-page":"593","DOI":"10.1056\/NEJM196803142781105\/ASSET\/9EE62BDC-88EB-469C-BCDC-DB379C2CAE47\/ASSETS\/IMAGES\/MEDIUM\/NEJM196803142781105_F2.GIF","article-title":"Medical records that guide and teach","volume":"278","author":"Weed","year":"1968","journal-title":"N Engl J Med"},{"key":"2025041716422184200_ocaf008-B61","doi-asserted-by":"publisher","first-page":"1561","DOI":"10.1056\/NEJMP2405999","article-title":"Large language models and the degradation of the medical record","volume":"391","author":"McCoy","year":"2024","journal-title":"N Engl J Med"},{"key":"2025041716422184200_ocaf008-B62","doi-asserted-by":"publisher","first-page":"899","DOI":"10.1093\/jamia\/ocv189","article-title":"SMART on FHIR: a standards-based, interoperable apps platform for electronic health records","volume":"23","author":"Mandel","year":"2016","journal-title":"J Am Med Inform Assoc"}],"container-title":["Journal of the American Medical Informatics Association"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/32\/4\/605\/61442713\/ocaf008.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/32\/4\/605\/61442713\/ocaf008.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,4,17]],"date-time":"2025-04-17T20:42:37Z","timestamp":1744922557000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jamia\/article\/32\/4\/605\/7954485"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,15]]},"references-count":62,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2025,1,15]]},"published-print":{"date-parts":[[2025,4,1]]}},"URL":"https:\/\/doi.org\/10.1093\/jamia\/ocaf008","relation":{},"ISSN":["1067-5027","1527-974X"],"issn-type":[{"value":"1067-5027","type":"print"},{"value":"1527-974X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2025,4]]},"published":{"date-parts":[[2025,1,15]]}}}