{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,28]],"date-time":"2026-08-28T23:54:30Z","timestamp":1787961270164,"version":"build-2784847793"},"reference-count":43,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,2,20]],"date-time":"2024-02-20T00:00:00Z","timestamp":1708387200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,2,20]],"date-time":"2024-02-20T00:00:00Z","timestamp":1708387200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["npj Digit. Med."],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>The use of large language models (LLMs) in clinical medicine is currently thriving. Effectively transferring LLMs\u2019 pertinent theoretical knowledge from computer science to their application in clinical medicine is crucial. Prompt engineering has shown potential as an effective method in this regard. To explore the application of prompt engineering in LLMs and to examine the reliability of LLMs, different styles of prompts were designed and used to ask different LLMs about their agreement with the American Academy of Orthopedic Surgeons (AAOS) osteoarthritis (OA) evidence-based guidelines. Each question was asked 5 times. We compared the consistency of the findings with guidelines across different evidence levels for different prompts and assessed the reliability of different prompts by asking the same question 5 times. gpt-4-Web with ROT prompting had the highest overall consistency (62.9%) and a significant performance for strong recommendations, with a total consistency of 77.5%. The reliability of the different LLMs for different prompts was not stable (Fleiss kappa ranged from \u22120.002 to 0.984). This study revealed that different prompts had variable effects across various models, and the gpt-4-Web with ROT prompt was the most consistent. An appropriate prompt could improve the accuracy of responses to professional medical questions.<\/jats:p>","DOI":"10.1038\/s41746-024-01029-4","type":"journal-article","created":{"date-parts":[[2024,2,20]],"date-time":"2024-02-20T04:02:50Z","timestamp":1708401770000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":386,"title":["Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs"],"prefix":"10.1038","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3260-5998","authenticated-orcid":false,"given":"Li","family":"Wang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xi","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"XiangWen","family":"Deng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hao","family":"Wen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"MingKe","family":"You","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"WeiZhi","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-9181-016X","authenticated-orcid":false,"given":"Qi","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-2666-6219","authenticated-orcid":false,"given":"Jian","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,2,20]]},"reference":[{"key":"1029_CR1","doi-asserted-by":"publisher","first-page":"1233","DOI":"10.1056\/NEJMsr2214184","volume":"388","author":"P Lee","year":"2023","unstructured":"Lee, P., Bubeck, S. & Petro, J. Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. N. Engl. J. Med. 388, 1233\u20131239 (2023).","journal-title":"N. Engl. J. Med."},{"key":"1029_CR2","doi-asserted-by":"publisher","first-page":"3197","DOI":"10.1007\/s11845-023-03377-8","volume":"192","author":"E Waisberg","year":"2023","unstructured":"Waisberg, E. et al. GPT-4: a new era of artificial intelligence in medicine. Ir. J. Med. Sci. 192, 3197\u20133200 (2023).","journal-title":"Ir. J. Med. Sci."},{"key":"1029_CR3","doi-asserted-by":"crossref","unstructured":"Scanlon, M., Breitinger, F., Hargreaves, C., Hilgert, J.-N. & Sheppard, J. ChatGPT for digital forensic investigation: The good, the bad, and the unknown. Forensic Science International: Digital Investigation (2023).","DOI":"10.20944\/preprints202307.0766.v1"},{"key":"1029_CR4","doi-asserted-by":"publisher","first-page":"78","DOI":"10.1001\/jama.2023.8288","volume":"330","author":"Z Kanjee","year":"2023","unstructured":"Kanjee, Z., Crowe, B. & Rodman, A. Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge. JAMA 330, 78\u201380 (2023).","journal-title":"JAMA"},{"key":"1029_CR5","doi-asserted-by":"publisher","first-page":"141","DOI":"10.1016\/j.ajo.2023.05.024","volume":"254","author":"LZ Cai","year":"2023","unstructured":"Cai, L. Z. et al. Performance of Generative Large Language Models on Ophthalmology Board Style Questions. Am. J. Ophthalmol. 254, 141\u2013149 (2023).","journal-title":"Am. J. Ophthalmol."},{"key":"1029_CR6","doi-asserted-by":"publisher","first-page":"e47479","DOI":"10.2196\/47479","volume":"25","author":"HL Walker","year":"2023","unstructured":"Walker, H. L. et al. Reliability of Medical Information Provided by ChatGPT: Assessment Against Clinical Guidelines and Patient Information Quality Instrument. J. Med. Internet Res. 25, e47479 (2023).","journal-title":"J. Med. Internet Res."},{"key":"1029_CR7","doi-asserted-by":"publisher","first-page":"2231","DOI":"10.1002\/alr.23201","volume":"13","author":"Y Yoshiyasu","year":"2023","unstructured":"Yoshiyasu, Y. et al. GPT-4 accuracy and completeness against International Consensus Statement on Allergy and Rhinology: Rhinosinusitis. Int Forum Allergy Rhinol. 13, 2231\u20132234 (2023).","journal-title":"Int Forum Allergy Rhinol."},{"key":"1029_CR8","doi-asserted-by":"publisher","first-page":"172","DOI":"10.1038\/s41586-023-06291-2","volume":"620","author":"K Singhal","year":"2023","unstructured":"Singhal, K. et al. Large language models encode clinical knowledge. Nature 620, 172\u2013180 (2023).","journal-title":"Nature"},{"key":"1029_CR9","unstructured":"Wang, X. et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. Published as a conference paper at ICLR 2023. https:\/\/iclr.cc\/media\/iclr-2023\/Slides\/11718.pdf (2023)."},{"key":"1029_CR10","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/s41746-023-00939-z","volume":"6","author":"JA Omiye","year":"2023","unstructured":"Omiye, J. A., Lester, J. C., Spichak, S., Rotemberg, V. & Daneshjou, R. Large language models propagate race-based medicine. NPJ digital Med. 6, 1\u20134 (2023).","journal-title":"NPJ digital Med."},{"key":"1029_CR11","unstructured":"Strobelt, H. et al. Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language Models. IEEE Trans. Vis. Comput. Graph. 29, 1146\u20131156 (2023)."},{"key":"1029_CR12","unstructured":"Wei, J. et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Preprint at: https:\/\/arxiv.org\/abs\/2201.11903 (2023)."},{"key":"1029_CR13","unstructured":"Yao, S. et al. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. Preprint at: https:\/\/arxiv.org\/abs\/2305.10601 (2023)."},{"key":"1029_CR14","doi-asserted-by":"publisher","first-page":"e50638","DOI":"10.2196\/50638","volume":"25","author":"B Mesk\u00f3","year":"2023","unstructured":"Mesk\u00f3, B. Prompt Engineering as an Important Emerging Skill for Medical Professionals: Tutorial. J. Med. Internet Res. 25, e50638 (2023).","journal-title":"J. Med. Internet Res."},{"key":"1029_CR15","doi-asserted-by":"publisher","first-page":"103024","DOI":"10.1016\/j.media.2023.103024","volume":"91","author":"M Fischer","year":"2023","unstructured":"Fischer, M., Bartler, A. & Yang, B. Prompt tuning for parameter-efficient medical image segmentation. Med. image Anal. 91, 103024 (2023).","journal-title":"Med. image Anal."},{"key":"1029_CR16","doi-asserted-by":"publisher","first-page":"201","DOI":"10.1007\/s11604-023-01491-2","volume":"42","author":"Y Toyama","year":"2023","unstructured":"Toyama, Y. et al. Performance evaluation of ChatGPT, GPT-4, and Bard on the official board examination of the Japan Radiology Society. Jpn. J. Radiol. 42, 201\u2013207 (2023).","journal-title":"Jpn. J. Radiol."},{"key":"1029_CR17","doi-asserted-by":"crossref","unstructured":"Kozachek, D. Investigating the Perception of the Future in GPT-3, -3.5 and GPT-4. C&C \u203223: Creativity and Cognition, 282\u2013287 (2023).","DOI":"10.1145\/3591196.3596827"},{"key":"1029_CR18","unstructured":"2019 Global Burden of Disease (GBD) study, https:\/\/vizhub.healthdata.org\/gbd-results\/ (2019)."},{"key":"1029_CR19","doi-asserted-by":"publisher","first-page":"819","DOI":"10.1136\/annrheumdis-2019-216515","volume":"79","author":"S Safiri","year":"2020","unstructured":"Safiri, S. et al. Global, regional and national burden of osteoarthritis 1990-2017: a systematic analysis of the Global Burden of Disease Study 2017. Ann. Rheum. Dis. 79, 819\u2013828 (2020).","journal-title":"Ann. Rheum. Dis."},{"key":"1029_CR20","first-page":"00990","volume":"S1063-4584","author":"AV Perruccio","year":"2023","unstructured":"Perruccio, A. V. et al. Osteoarthritis Year in Review 2023: Epidemiology & therapy. Osteoarthr. Cartil. S1063-4584, 00990\u201300991 (2023).","journal-title":"Osteoarthr. Cartil."},{"key":"1029_CR21","doi-asserted-by":"publisher","first-page":"353","DOI":"10.1076\/edre.7.4.353.8937","volume":"7","author":"TD Pigott","year":"2001","unstructured":"Pigott, T. D. A Review of Methods for Missing Data. Educ. Res. Eval. 7, 353\u2013383 (2001).","journal-title":"Educ. Res. Eval."},{"key":"1029_CR22","doi-asserted-by":"publisher","unstructured":"Koga, S., Martin, N. B. & Dickson, D. W. Evaluating the performance of large language models: ChatGPT and Google Bard in generating differential diagnoses in clinicopathological conferences of neurodegenerative disorders. Brain Pathol., e13207, https:\/\/doi.org\/10.1111\/bpa.13207 (2023).","DOI":"10.1111\/bpa.13207"},{"key":"1029_CR23","doi-asserted-by":"publisher","first-page":"104770","DOI":"10.1016\/j.ebiom.2023.104770","volume":"95","author":"ZW Lim","year":"2023","unstructured":"Lim, Z. W. et al. Benchmarking large language models\u2019 performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google Bard. EBioMedicine 95, 104770 (2023).","journal-title":"EBioMedicine"},{"key":"1029_CR24","doi-asserted-by":"publisher","first-page":"e49995","DOI":"10.2196\/49995","volume":"11","author":"H Fraser","year":"2023","unstructured":"Fraser, H. et al. Comparison of Diagnostic and Triage Accuracy of Ada Health and WebMD Symptom Checkers, ChatGPT, and Physicians for Patients in an Emergency Department: Clinical Data Analysis Study. JMIR mHealth uHealth 11, e49995 (2023).","journal-title":"JMIR mHealth uHealth"},{"key":"1029_CR25","doi-asserted-by":"publisher","first-page":"1353","DOI":"10.1227\/neu.0000000000002632","volume":"93","author":"R Ali","year":"2023","unstructured":"Ali, R. et al. Performance of ChatGPT and GPT-4 on Neurosurgery Written Board Examinations. Neurosurgery 93, 1353\u20131365 (2023).","journal-title":"Neurosurgery"},{"key":"1029_CR26","doi-asserted-by":"publisher","unstructured":"Fowler, T., Pullen, S. & Birkett, L. Performance of ChatGPT and Bard on the official part 1 FRCOphth practice questions. Br. J. Ophthalmol., bjo-2023-324091, https:\/\/doi.org\/10.1136\/bjo-2023-324091 (2023).","DOI":"10.1136\/bjo-2023-324091"},{"key":"1029_CR27","doi-asserted-by":"publisher","unstructured":"Passby, L., Jenko, N. & Wernham, A. Performance of ChatGPT on dermatology Specialty Certificate Examination multiple choice questions. Clin. Exp. Dermatol., llad197, https:\/\/doi.org\/10.1093\/ced\/llad197 (2023).","DOI":"10.1093\/ced\/llad197"},{"key":"1029_CR28","doi-asserted-by":"publisher","first-page":"876","DOI":"10.1111\/1742-6723.14280","volume":"35","author":"J Smith","year":"2023","unstructured":"Smith, J., Choi, P. M. & Buntine, P. Will code one day run a code? Performance of language models on ACEM primary examinations and implications. Emerg. Med. Australas. 35, 876\u2013878 (2023).","journal-title":"Emerg. Med. Australas."},{"key":"1029_CR29","doi-asserted-by":"publisher","first-page":"108163","DOI":"10.1016\/j.isci.2023.108163","volume":"26","author":"K Pushpanathan","year":"2023","unstructured":"Pushpanathan, K. et al. Popular large language model chatbots\u2019 accuracy, comprehensiveness, and self-awareness in answering ocular symptom queries. iScience 26, 108163 (2023).","journal-title":"iScience"},{"key":"1029_CR30","doi-asserted-by":"publisher","unstructured":"Antaki, F. et al. Capabilities of GPT-4 in ophthalmology: an analysis of model entropy and progress towards human-level medical question answering. Br J Ophthalmol, bjo-2023-324438, https:\/\/doi.org\/10.1136\/bjo-2023-324438 (2023).","DOI":"10.1136\/bjo-2023-324438"},{"key":"1029_CR31","doi-asserted-by":"publisher","DOI":"10.1016\/j.cmi.2023.11.002","author":"WI Wei","year":"2023","unstructured":"Wei, W. I. et al. Extracting symptoms from free-text responses using ChatGPT among COVID-19 cases in Hong Kong. Clin. Microbiol Infect. https:\/\/doi.org\/10.1016\/j.cmi.2023.11.002 (2023).","journal-title":"Clin. Microbiol Infect."},{"key":"1029_CR32","doi-asserted-by":"publisher","DOI":"10.1038\/s41433-023-02772-w","author":"O Kleinig","year":"2023","unstructured":"Kleinig, O. et al. How to use large language models in ophthalmology: from prompt engineering to protecting confidentiality. Eye https:\/\/doi.org\/10.1038\/s41433-023-02772-w (2023).","journal-title":"Eye"},{"key":"1029_CR33","doi-asserted-by":"publisher","DOI":"10.4274\/dir.2023.232417","author":"T Akinci D\u2019Antonoli","year":"2023","unstructured":"Akinci D\u2019Antonoli, T. et al. Large language models in radiology: fundamentals, applications, ethical considerations, risks, and future directions. Diagnostic Interventional Radiol. https:\/\/doi.org\/10.4274\/dir.2023.232417 (2023).","journal-title":"Diagnostic Interventional Radiol."},{"key":"1029_CR34","doi-asserted-by":"publisher","first-page":"100324","DOI":"10.1016\/j.xops.2023.100324","volume":"3","author":"F Antaki","year":"2023","unstructured":"Antaki, F., Touma, S., Milad, D., El-Khoury, J. & Duval, R. Evaluating the Performance of ChatGPT in Ophthalmology: An Analysis of Its Successes and Shortcomings. Ophthalmol. Sci. 3, 100324 (2023).","journal-title":"Ophthalmol. Sci."},{"key":"1029_CR35","unstructured":"Zhu, K. et al. PromptBench: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts. Preprint at https:\/\/arxiv.org\/abs\/2306.04528v4 (2023)."},{"key":"1029_CR36","unstructured":"Newsroom. AAOS Updates Clinical Practice Guideline for Osteoarthritis of the Knee, https:\/\/www.aaos.org\/aaos-home\/newsroom\/press-releases\/aaos-updates-clinical-practice-guideline-for-osteoarthritis-of-the-knee\/ (2021)."},{"key":"1029_CR37","unstructured":"Osteoarthritis of the Knee. Clinical Practice Guideline on Management of Osteoarthritis of the Knee. 3rd ed, https:\/\/www.aaos.org\/quality\/quality-programs\/lower-extremity-programs\/osteoarthritis-of-the-knee\/ (2021)."},{"key":"1029_CR38","unstructured":"The American Academy of Orthopaedic Surgeons Board of Directors. Management of Osteoarthritis of the Knee (Non-Arthroplasty) https:\/\/www.aaos.org\/globalassets\/quality-and-practice-resources\/osteoarthritis-of-the-knee\/oak3cpg.pdf (2019)."},{"key":"1029_CR39","doi-asserted-by":"publisher","first-page":"159","DOI":"10.1080\/03610927808827340","volume":"5","author":"M Goldstein","year":"1976","unstructured":"Goldstein, M., Wolf, E. & Dillon, W. On a test of independence for contingency tables. Commun. Stat. Theory Methods 5, 159\u2013169 (1976).","journal-title":"Commun. Stat. Theory Methods"},{"key":"1029_CR40","doi-asserted-by":"crossref","first-page":"105","DOI":"10.1007\/s40368-018-0397-x","volume":"20","author":"AT Gurcan","year":"2019","unstructured":"Gurcan, A. T. & Seymen, F. Clinical and radiographic evaluation of indirect pulp capping with three different materials: a 2-year follow-up study. Eur. J. Paediatr. Dent. 20, 105\u2013110 (2019).","journal-title":"Eur. J. Paediatr. Dent."},{"key":"1029_CR41","doi-asserted-by":"publisher","first-page":"502","DOI":"10.1111\/opo.12131","volume":"34","author":"RA Armstrong","year":"2014","unstructured":"Armstrong, R. A. When to use the Bonferroni correction. Ophthalmic Physiol. Opt. 34, 502\u2013508 (2014).","journal-title":"Ophthalmic Physiol. Opt."},{"key":"1029_CR42","doi-asserted-by":"publisher","DOI":"10.1186\/s12879-023-08729-4","volume":"23","author":"D Pokutnaya","year":"2023","unstructured":"Pokutnaya, D. et al. Inter-rater reliability of the infectious disease modeling reproducibility checklist (IDMRC) as applied to COVID-19 computational modeling research. BMC Infect. Dis. 23, 733 (2023).","journal-title":"BMC Infect. Dis."},{"key":"1029_CR43","doi-asserted-by":"publisher","DOI":"10.1186\/s12874-016-0200-9","volume":"16","author":"A Zapf","year":"2016","unstructured":"Zapf, A., Castell, S., Morawietz, L. & Karch, A. Measuring inter-rater reliability for nominal data \u2013 which coefficients and confidence intervals are appropriate? BMC Med. Res. Methodol. 16, 93 (2016).","journal-title":"BMC Med. Res. Methodol."}],"container-title":["npj Digital Medicine"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.nature.com\/articles\/s41746-024-01029-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41746-024-01029-4","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41746-024-01029-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,2,20]],"date-time":"2024-02-20T04:19:29Z","timestamp":1708402769000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.nature.com\/articles\/s41746-024-01029-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,20]]},"references-count":43,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["1029"],"URL":"https:\/\/doi.org\/10.1038\/s41746-024-01029-4","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-3336823\/v1","asserted-by":"object"}]},"ISSN":["2398-6352"],"issn-type":[{"value":"2398-6352","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,2,20]]},"assertion":[{"value":"8 September 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 February 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 February 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"41"}}