{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T02:37:46Z","timestamp":1785292666115,"version":"3.55.0"},"reference-count":75,"publisher":"Oxford University Press (OUP)","issue":"9","license":[{"start":{"date-parts":[[2024,4,29]],"date-time":"2024-04-29T00:00:00Z","timestamp":1714348800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/pages\/standard-publication-reuse-rights"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2024,9,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Objectives<\/jats:title>\n                  <jats:p>Large Language Models (LLMs) such as ChatGPT and Med-PaLM have excelled in various medical question-answering tasks. However, these English-centric models encounter challenges in non-English clinical settings, primarily due to limited clinical knowledge in respective languages, a consequence of imbalanced training corpora. We systematically evaluate LLMs in the Chinese medical context and develop a novel in-context learning framework to enhance their performance.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Materials and Methods<\/jats:title>\n                  <jats:p>The latest China National Medical Licensing Examination (CNMLE-2022) served as the benchmark. We collected 53 medical books and 381\u00a0149 medical questions to construct the medical knowledge base and question bank. The proposed Knowledge and Few-shot Enhancement In-context Learning (KFE) framework leverages the in-context learning ability of LLMs to integrate diverse external clinical knowledge sources. We evaluated KFE with ChatGPT (GPT-3.5), GPT-4, Baichuan2-7B, Baichuan2-13B, and QWEN-72B in CNMLE-2022 and further investigated the effectiveness of different pathways for incorporating LLMs with medical knowledge from 7 distinct perspectives.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>Directly applying ChatGPT failed to qualify for the CNMLE-2022 at a score of 51. Cooperated with the KFE framework, the LLMs with varying sizes yielded consistent and significant improvements. The ChatGPT\u2019s performance surged to 70.04 and GPT-4 achieved the highest score of 82.59. This surpasses the qualification threshold (60) and exceeds the average human score of 68.70, affirming the effectiveness and robustness of the framework. It also enabled a smaller Baichuan2-13B to pass the examination, showcasing the great potential in low-resource settings.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Discussion and Conclusion<\/jats:title>\n                  <jats:p>This study shed light on the optimal practices to enhance the capabilities of LLMs in non-English medical scenarios. By synergizing medical knowledge through in-context learning, LLMs can extend clinical insight beyond language barriers in healthcare, significantly reducing language-related disparities of LLM applications and ensuring global benefit in this field.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/jamia\/ocae079","type":"journal-article","created":{"date-parts":[[2024,4,30]],"date-time":"2024-04-30T03:30:49Z","timestamp":1714447849000},"page":"2054-2064","source":"Crossref","is-referenced-by-count":36,"title":["Large language models leverage external knowledge to extend clinical insight beyond language boundaries"],"prefix":"10.1093","volume":"31","author":[{"given":"Jiageng","family":"Wu","sequence":"first","affiliation":[{"name":"School of Public Health, Zhejiang University School of Medicine , Hangzhou, 310058, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xian","family":"Wu","sequence":"additional","affiliation":[{"name":"Jarvis Research Center, Tencent YouTu Lab , Beijing, 100101, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhaopeng","family":"Qiu","sequence":"additional","affiliation":[{"name":"Jarvis Research Center, Tencent YouTu Lab , Beijing, 100101, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Minghui","family":"Li","sequence":"additional","affiliation":[{"name":"School of Public Health, Zhejiang University School of Medicine , Hangzhou, 310058, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shixu","family":"Lin","sequence":"additional","affiliation":[{"name":"School of Public Health, Zhejiang University School of Medicine , Hangzhou, 310058, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yingying","family":"Zhang","sequence":"additional","affiliation":[{"name":"Jarvis Research Center, Tencent YouTu Lab , Beijing, 100101, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yefeng","family":"Zheng","sequence":"additional","affiliation":[{"name":"Jarvis Research Center, Tencent YouTu Lab , Beijing, 100101, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Changzheng","family":"Yuan","sequence":"additional","affiliation":[{"name":"School of Public Health, Zhejiang University School of Medicine , Hangzhou, 310058, China"},{"name":"Department of Nutrition, Harvard T.H. Chan School of Public Health , Boston, MA 02115, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jie","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Public Health, Zhejiang University School of Medicine , Hangzhou, 310058, China"},{"name":"Division of Pharmacoepidemiology and Pharmacoeconomics , Department of Medicine, Brigham and Women\u2019s Hospital, , Boston, MA 02115, United States"},{"name":"Harvard Medical School , Department of Medicine, Brigham and Women\u2019s Hospital, , Boston, MA 02115, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2024,4,29]]},"reference":[{"key":"2024082207514264500_ocae079-B1","author":"Zhao"},{"issue":"8","key":"2024082207514264500_ocae079-B2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford","year":"2019","journal-title":"OpenAI Blog"},{"issue":"8","key":"2024082207514264500_ocae079-B3","doi-asserted-by":"crossref","first-page":"1930","DOI":"10.1038\/s41591-023-02448-8","article-title":"Large language models in medicine","volume":"29","author":"Thirunavukarasu","year":"2023","journal-title":"Nat. Med"},{"key":"2024082207514264500_ocae079-B4","author":"Devlin"},{"key":"2024082207514264500_ocae079-B5","author":"Edunov","year":"2019"},{"key":"2024082207514264500_ocae079-B6","first-page":"2463","author":"Petroni"},{"issue":"9","key":"2024082207514264500_ocae079-B7","doi-asserted-by":"crossref","first-page":"1028","DOI":"10.1001\/jamainternmed.2023.2909","article-title":"Chatbot vs medical student performance on free-response clinical reasoning examinations","volume":"183","author":"Strong","year":"2023","journal-title":"JAMA Intern Med"},{"key":"2024082207514264500_ocae079-B8","first-page":"1","author":"Chung"},{"key":"2024082207514264500_ocae079-B9","doi-asserted-by":"crossref","first-page":"357","DOI":"10.1038\/s41586-023-06160-y","article-title":"Health system-scale language models are all-purpose prediction engines","author":"Jiang","year":"2023","journal-title":"Nature"},{"key":"2024082207514264500_ocae079-B10","author":"Wu"},{"key":"2024082207514264500_ocae079-B11","first-page":"13063","article-title":"Unified language model pre-training for natural language understanding and generation","author":"Dong"},{"key":"2024082207514264500_ocae079-B12","first-page":"100905","article-title":"ChatGPT: promise and challenges for deployment in low-and middle-income countries","volume":"41","author":"Wang","year":"2023","journal-title":"Lancet Reg Health West Pac"},{"issue":"13","key":"2024082207514264500_ocae079-B13","doi-asserted-by":"crossref","first-page":"1233","DOI":"10.1056\/NEJMsr2214184","article-title":"Benefits, limits, and risks of GPT-4 as an AI Chatbot for medicine","volume":"388","author":"Lee","year":"2023","journal-title":"N Engl J Med"},{"key":"2024082207514264500_ocae079-B14","author":"Liu"},{"issue":"7","key":"2024082207514264500_ocae079-B15","doi-asserted-by":"crossref","first-page":"1237","DOI":"10.1093\/jamia\/ocad072","article-title":"Using AI-generated suggestions from ChatGPT to optimize clinical decision support","volume":"30","author":"Liu","year":"2023","journal-title":"J Am Med Inform Assoc"},{"issue":"9","key":"2024082207514264500_ocae079-B16","doi-asserted-by":"crossref","first-page":"1026","DOI":"10.1001\/jamainternmed.2023.2561","article-title":"Comparison of history of present illness summaries generated by a Chatbot and senior internal medicine residents","volume":"183","author":"Nayak","year":"2023","journal-title":"JAMA Intern Med"},{"key":"2024082207514264500_ocae079-B17","first-page":"589","author":"Ayers","year":"2023"},{"key":"2024082207514264500_ocae079-B18","first-page":"100906","article-title":"ChatGPT for low-and middle-income countries: a Greek gift?","volume":"41","author":"Lam","year":"2023","journal-title":"Lancet Reg Health West Pac"},{"issue":"10","key":"2024082207514264500_ocae079-B19","doi-asserted-by":"crossref","first-page":"842","DOI":"10.1001\/jama.2023.1044","article-title":"Appropriateness of cardiovascular disease prevention recommendations obtained from a popular online chat-based artificial intelligence model","volume":"329","author":"Sarraju","year":"2023","journal-title":"JAMA"},{"key":"2024082207514264500_ocae079-B20","author":"Nori"},{"issue":"2","key":"2024082207514264500_ocae079-B21","doi-asserted-by":"crossref","first-page":"e0000198","DOI":"10.1371\/journal.pdig.0000198","article-title":"Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models","volume":"2","author":"Kung","year":"2023","journal-title":"PLoS Digit Health"},{"issue":"7972","key":"2024082207514264500_ocae079-B22","doi-asserted-by":"crossref","first-page":"172","DOI":"10.1038\/s41586-023-06291-2","article-title":"Large language models encode clinical knowledge","volume":"620","author":"Singhal","year":"2023","journal-title":"Nature"},{"key":"2024082207514264500_ocae079-B23","author":"Nicholas"},{"key":"2024082207514264500_ocae079-B24","author":"Wang"},{"key":"2024082207514264500_ocae079-B25"},{"key":"2024082207514264500_ocae079-B26","author":"Bang"},{"key":"2024082207514264500_ocae079-B27","author":"Blevins"},{"issue":"1","key":"2024082207514264500_ocae079-B28","first-page":"1","article-title":"Domain-specific language model pretraining for biomedical natural language processing","volume":"3","author":"Gu","year":"2021","journal-title":"ACM Trans Comput Healthc (HEALTH)"},{"key":"2024082207514264500_ocae079-B29","author":"Li\u00e9vin"},{"issue":"9","key":"2024082207514264500_ocae079-B30","doi-asserted-by":"crossref","first-page":"866","DOI":"10.1001\/jama.2023.14217","article-title":"Creation and adoption of large language models in medicine","volume":"330","author":"Shah","year":"2023","journal-title":"JAMA"},{"key":"2024082207514264500_ocae079-B31","author":"Peng"},{"key":"2024082207514264500_ocae079-B32","author":"Rubin"},{"key":"2024082207514264500_ocae079-B33","author":"Gao"},{"issue":"1","key":"2024082207514264500_ocae079-B34","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1007\/s10916-023-01961-0","article-title":"ChatGPT performs on the Chinese national medical licensing examination","volume":"47","author":"Wang","year":"2023","journal-title":"J Med Syst"},{"issue":"1","key":"2024082207514264500_ocae079-B35","doi-asserted-by":"crossref","first-page":"e48002","DOI":"10.2196\/48002","article-title":"Performance of GPT-3.5 and GPT-4 on the Japanese medical licensing examination: comparison study","volume":"9","author":"Takagi","year":"2023","journal-title":"JMIR Med Educ"},{"key":"2024082207514264500_ocae079-B36","author":"Kasai"},{"key":"2024082207514264500_ocae079-B37"},{"key":"2024082207514264500_ocae079-B38"},{"issue":"1","key":"2024082207514264500_ocae079-B39","doi-asserted-by":"crossref","first-page":"4352","DOI":"10.1038\/s41467-018-06799-6","article-title":"Master clinical medical knowledge at certificated-doctor-level with deep learning model","volume":"9","author":"Wu","year":"2018","journal-title":"Nat Commun"},{"key":"2024082207514264500_ocae079-B40","first-page":"1877","article-title":"Language models are few-shot learners","author":"Brown"},{"key":"2024082207514264500_ocae079-B41","first-page":"24824","author":"Wei"},{"issue":"4","key":"2024082207514264500_ocae079-B42","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1561\/1500000019","article-title":"The probabilistic relevance framework: BM25 and beyond","volume":"3","author":"Robertson","year":"2009","journal-title":"Found Trends Inf Retr"},{"key":"2024082207514264500_ocae079-B43","author":"Shiyi"},{"key":"2024082207514264500_ocae079-B44"},{"key":"2024082207514264500_ocae079-B45","author":"Qin"},{"key":"2024082207514264500_ocae079-B46","author":"Yang"},{"key":"2024082207514264500_ocae079-B47","author":"Bai"},{"key":"2024082207514264500_ocae079-B48","first-page":"5706","author":"Zhang"},{"key":"2024082207514264500_ocae079-B49"},{"key":"2024082207514264500_ocae079-B50","author":"Zhang"},{"key":"2024082207514264500_ocae079-B51","author":"Fu"},{"key":"2024082207514264500_ocae079-B52","author":"Shwartz","year":"2020"},{"key":"2024082207514264500_ocae079-B53","author":"Liu","year":"2022"},{"key":"2024082207514264500_ocae079-B54","first-page":"3929","author":"Guu","year":"2020"},{"key":"2024082207514264500_ocae079-B55","author":"Kaplan"},{"key":"2024082207514264500_ocae079-B56","author":"Wei"},{"key":"2024082207514264500_ocae079-B57","doi-asserted-by":"crossref","first-page":"104770","DOI":"10.1016\/j.ebiom.2023.104770","article-title":"Benchmarking large language models\u2019 performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google bard","volume":"95","author":"Lim","year":"2023","journal-title":"EBioMedicine"},{"issue":"10","key":"2024082207514264500_ocae079-B58","doi-asserted-by":"crossref","first-page":"e2338050","DOI":"10.1001\/jamanetworkopen.2023.38050","article-title":"Assessing biases in medical decisions via clinician and AI Chatbot responses to patient vignettes","volume":"6","author":"Kim","year":"2023","journal-title":"JAMA Netw Open"},{"key":"2024082207514264500_ocae079-B59"},{"issue":"4","key":"2024082207514264500_ocae079-B60","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1016\/S1473-3099(23)00113-5","article-title":"ChatGPT and antimicrobial advice: the end of the consulting infection doctor?","volume":"23","author":"Howard","year":"2023","journal-title":"Lancet Infect Dis"},{"issue":"11","key":"2024082207514264500_ocae079-B61","doi-asserted-by":"crossref","first-page":"1220","DOI":"10.1001\/jamasurg.2023.3875","article-title":"Implications of using Chatbots for future surgical education","volume":"158","author":"Grigorian","year":"2023","journal-title":"JAMA Surg"},{"key":"2024082207514264500_ocae079-B62","author":"Zhu"},{"key":"2024082207514264500_ocae079-B63","first-page":"172"},{"key":"2024082207514264500_ocae079-B64","author":"Heim"},{"key":"2024082207514264500_ocae079-B65","author":"Liu"},{"key":"2024082207514264500_ocae079-B66","first-page":"578","author":"Lehman"},{"issue":"9","key":"2024082207514264500_ocae079-B67","doi-asserted-by":"crossref","first-page":"792","DOI":"10.1001\/jama.2023.14311","article-title":"Large language models answer medical questions accurately, but can\u2019t match clinicians\u2019 knowledge","volume":"330","author":"Harris","year":"2023","journal-title":"JAMA"},{"issue":"1","key":"2024082207514264500_ocae079-B68","doi-asserted-by":"crossref","first-page":"93","DOI":"10.1038\/s41746-021-00464-x","article-title":"Considering the possibilities and pitfalls of generative pre-trained transformer 3 (GPT-3) in healthcare delivery","volume":"4","author":"Korngiebel","year":"2021","journal-title":"NPJ Digit Med"},{"key":"2024082207514264500_ocae079-B69","author":"Thompson"},{"issue":"5","key":"2024082207514264500_ocae079-B70","doi-asserted-by":"crossref","first-page":"e230582","DOI":"10.1148\/radiol.230582","article-title":"Performance of ChatGPT on a radiology board-style examination: insights into current strengths and limitations","volume":"307","author":"Bhayana","year":"2023","journal-title":"Radiology"},{"issue":"1","key":"2024082207514264500_ocae079-B71","doi-asserted-by":"crossref","first-page":"5131","DOI":"10.1038\/s41467-020-18918-3","article-title":"Deep transfer learning for reducing health care disparities arising from biomedical data inequality","volume":"11","author":"Gao","year":"2020","journal-title":"Nat Commun"},{"key":"2024082207514264500_ocae079-B72","volume":"3968-3977.","author":"Wu J","year":"2023"},{"issue":"7","key":"2024082207514264500_ocae079-B73","doi-asserted-by":"crossref","first-page":"687","DOI":"10.1038\/s42256-023-00670-0","article-title":"The importance of resource awareness in artificial intelligence for healthcare","volume":"5","author":"Jia","year":"2023","journal-title":"Nat Mach Intell"},{"issue":"5","key":"2024082207514264500_ocae079-B74","doi-asserted-by":"crossref","first-page":"379","DOI":"10.1097\/MLR.0000000000001507","article-title":"Health equity beyond data: health care worker perceptions of race, ethnicity, and language data collection in electronic health records","volume":"59","author":"Cruz","year":"2021","journal-title":"Med Care"},{"issue":"9","key":"2024082207514264500_ocae079-B75","doi-asserted-by":"crossref","first-page":"833","DOI":"10.1056\/NEJMra2214964","article-title":"Considering biased data as informative artifacts in ai-assisted health care","volume":"389","author":"Ferryman","year":"2023","journal-title":"New Engl J Med"}],"container-title":["Journal of the American Medical Informatics Association"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/31\/9\/2054\/58868121\/ocae079.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/31\/9\/2054\/58868121\/ocae079.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,22]],"date-time":"2024-08-22T11:42:13Z","timestamp":1724326933000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jamia\/article\/31\/9\/2054\/7659846"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,29]]},"references-count":75,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2024,4,29]]},"published-print":{"date-parts":[[2024,9,1]]}},"URL":"https:\/\/doi.org\/10.1093\/jamia\/ocae079","relation":{},"ISSN":["1067-5027","1527-974X"],"issn-type":[{"value":"1067-5027","type":"print"},{"value":"1527-974X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024,9]]},"published":{"date-parts":[[2024,4,29]]}}}