{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,9]],"date-time":"2026-08-09T09:42:35Z","timestamp":1786268555528,"version":"3.56.0"},"reference-count":93,"publisher":"Springer Science and Business Media LLC","issue":"8067","license":[{"start":{"date-parts":[[2025,4,9]],"date-time":"2025-04-09T00:00:00Z","timestamp":1744156800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,4,9]],"date-time":"2025-04-09T00:00:00Z","timestamp":1744156800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Nature"],"published-print":{"date-parts":[[2025,6,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>At the heart of medicine lies physician\u2013patient dialogue, where skillful history-taking enables effective diagnosis, management and enduring trust<jats:sup>1,2<\/jats:sup>. Artificial intelligence (AI) systems capable of diagnostic dialogue could increase accessibility and quality of care. However, approximating clinicians\u2019 expertise is an outstanding challenge. Here we introduce AMIE (Articulate Medical Intelligence Explorer), a large language model (LLM)-based AI system optimized for diagnostic dialogue. AMIE uses a self-play-based<jats:sup>3<\/jats:sup> simulated environment with automated feedback for scaling learning across disease conditions, specialties and contexts. We designed a framework for evaluating clinically meaningful axes of performance, including history-taking, diagnostic accuracy, management, communication skills and empathy. We compared AMIE\u2019s performance to that of primary care physicians in a randomized, double-blind crossover study of text-based consultations with validated patient-actors similar to objective structured clinical examination<jats:sup>4,5<\/jats:sup>. The study included 159 case scenarios from providers in Canada, the United Kingdom and India, 20 primary care physicians compared to AMIE, and evaluations by specialist physicians and patient-actors. AMIE demonstrated greater diagnostic accuracy and superior performance on 30 out of 32 axes according to the specialist physicians and 25 out of 26 axes according to the patient-actors. Our research has several limitations and should be interpreted with caution. Clinicians used synchronous text chat, which permits large-scale LLM\u2013patient interactions, but this is unfamiliar in clinical practice. While further research is required before AMIE could be translated to real-world settings, the results represent a milestone towards conversational diagnostic AI.<\/jats:p>","DOI":"10.1038\/s41586-025-08866-7","type":"journal-article","created":{"date-parts":[[2025,4,9]],"date-time":"2025-04-09T15:03:37Z","timestamp":1744211017000},"page":"442-450","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":292,"title":["Towards conversational diagnostic artificial intelligence"],"prefix":"10.1038","volume":"642","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9191-7938","authenticated-orcid":false,"given":"Tao","family":"Tu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1735-9680","authenticated-orcid":false,"given":"Mike","family":"Schaekermann","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anil","family":"Palepu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Khaled","family":"Saab","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jan","family":"Freyberg","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8107-6730","authenticated-orcid":false,"given":"Ryutaro","family":"Tanno","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Amy","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Brenna","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohamed","family":"Amin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yong","family":"Cheng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Elahe","family":"Vedadi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1624-0220","authenticated-orcid":false,"given":"Nenad","family":"Tomasev","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7447-6031","authenticated-orcid":false,"given":"Shekoofeh","family":"Azizi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-0286-609X","authenticated-orcid":false,"given":"Karan","family":"Singhal","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Le","family":"Hou","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Albert","family":"Webson","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kavita","family":"Kulkarni","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"S. Sara","family":"Mahdavi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christopher","family":"Semturs","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Juraj","family":"Gottweis","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Joelle","family":"Barral","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Katherine","family":"Chou","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Greg S.","family":"Corrado","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3960-6002","authenticated-orcid":false,"given":"Yossi","family":"Matias","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-4958-5976","authenticated-orcid":false,"given":"Alan","family":"Karthikesalingam","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7849-2074","authenticated-orcid":false,"given":"Vivek","family":"Natarajan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,4,9]]},"reference":[{"key":"8866_CR1","doi-asserted-by":"crossref","unstructured":"Levine, D. History taking is a complex skill. Br. Med. J. 358, j3513 (2017).","DOI":"10.1136\/bmj.j3513"},{"key":"8866_CR2","unstructured":"Engel, G. L. & Morgan, W. L. Interviewing the Patient (W. B. Saunders, 1973)."},{"key":"8866_CR3","unstructured":"Fu, Y., Peng, H., Khot, T. & Lapata, M. Improving language model negotiation with self-play and in-context learning from AI feedback. Preprint at https:\/\/arxiv.org\/abs\/2305.10142 (2023)."},{"key":"8866_CR4","doi-asserted-by":"publisher","first-page":"735","DOI":"10.1097\/00000658-199512000-00007","volume":"222","author":"DA Sloan","year":"1995","unstructured":"Sloan, D. A., Donnelly, M. B., Schwartz, R. W. & Strodel, W. E. The objective structured clinical examination. The new gold standard for evaluating postgraduate clinical performance. Ann. Surg. 222, 735 (1995).","journal-title":"Ann. Surg."},{"key":"8866_CR5","doi-asserted-by":"publisher","first-page":"736","DOI":"10.1001\/archpedi.154.7.736","volume":"154","author":"C Carraccio","year":"2000","unstructured":"Carraccio, C. & Englander, R. The objective structured clinical examination: a step in the direction of competency-based evaluation. Arch. Pediatr. Adolesc. Med. 154, 736\u2013741 (2000).","journal-title":"Arch. Pediatr. Adolesc. Med."},{"key":"8866_CR6","first-page":"163","volume":"156","author":"MC Peterson","year":"1992","unstructured":"Peterson, M. C., Holbrook, J. H., Von Hales, D., Smith, N. & Staker, L. Contributions of the history, physical examination, and laboratory investigation in making medical diagnoses. West. J. Med. 156, 163 (1992).","journal-title":"West. J. Med."},{"key":"8866_CR7","doi-asserted-by":"crossref","unstructured":"Silverman, J., Kurtz, S. & Draper, J. Skills for Communicating with Patients 3rd edn (CRC, 2016).","DOI":"10.1201\/9781910227268"},{"key":"8866_CR8","doi-asserted-by":"publisher","first-page":"2246","DOI":"10.1056\/NEJMc1404326","volume":"370","author":"T Rennie","year":"2014","unstructured":"Rennie, T., Marriott, J. & Brock, T. P. Global supply of health professionals. N. Engl. J. Med. 370, 2246\u20132247 (2014).","journal-title":"N. Engl. J. Med."},{"key":"8866_CR9","unstructured":"OpenAI et al. GPT-4 technical report. Preprint at https:\/\/arxiv.org\/abs\/2303.08774 (2023)."},{"key":"8866_CR10","unstructured":"Anil, R. et al. PaLM 2 technical report. Preprint at https:\/\/arxiv.org\/abs\/2305.10403 (2023)."},{"key":"8866_CR11","unstructured":"Gemini Team Google et al. Gemini: a family of highly capable multimodal models. Preprint at https:\/\/arxiv.org\/abs\/2312.11805 (2023)."},{"key":"8866_CR12","doi-asserted-by":"publisher","first-page":"172","DOI":"10.1038\/s41586-023-06291-2","volume":"620","author":"K Singhal","year":"2023","unstructured":"Singhal, K. et al. Large language models encode clinical knowledge. Nature 620, 172\u2013180 (2023).","journal-title":"Nature"},{"key":"8866_CR13","doi-asserted-by":"crossref","unstructured":"Singhal, K. et al. Toward expert-level medical question answering with large language models. Nat. Med. 31, 943\u2013950 (2025).","DOI":"10.1038\/s41591-024-03423-7"},{"key":"8866_CR14","unstructured":"Nori, H. et al. Can generalist foundation models outcompete special-purpose tuning? Case study in medicine. Preprint at https:\/\/arxiv.org\/abs\/2311.16452 (2023)."},{"key":"8866_CR15","unstructured":"Thoppilan, R. et al. LaMDA: language models for dialog applications. Preprint at https:\/\/arxiv.org\/abs\/2201.08239 (2022)."},{"key":"8866_CR16","unstructured":"Introducing ChatGPT. OpenAI https:\/\/openai.com\/blog\/chatgpt (2022)."},{"key":"8866_CR17","unstructured":"Toma, A. et al. Clinical Camel: an open-source expert-level medical language model with dialogue-based knowledge encoding. Preprint at https:\/\/arxiv.org\/abs\/2305.12031 (2023)."},{"key":"8866_CR18","unstructured":"Chen, Z. et al. MEDITRON-70B: scaling medical pretraining for large language models. Preprint at https:\/\/arxiv.org\/abs\/2311.16079 (2023)."},{"key":"8866_CR19","doi-asserted-by":"publisher","first-page":"385","DOI":"10.4300\/JGME-D-13-00072.1","volume":"5","author":"A King","year":"2013","unstructured":"King, A. & Hoppe, R. B. \u201cBest practice\u201d for patient-centered communication: a narrative review. J. Grad. Med. Educ. 5, 385\u2013393 (2013).","journal-title":"J. Grad. Med. Educ."},{"key":"8866_CR20","doi-asserted-by":"publisher","first-page":"452\u2013459","DOI":"10.7861\/clinmedicine.3-5-452","volume":"3","author":"J Dacre","year":"2003","unstructured":"Dacre, J., Besser, M. & White, P. MRCP(UK) part 2 clinical examination (PACES): a review of the first four examination sessions (June 2001 \u2013 July 2002). Clin. Med. 3, 452\u2013459 (2003).","journal-title":"Clin. Med"},{"key":"8866_CR21","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1186\/s12916-019-1426-2","volume":"17","author":"CJ Kelly","year":"2019","unstructured":"Kelly, C. J., Karthikesalingam, A., Suleyman, M., Corrado, G. & King, D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 17, 195 (2019).","journal-title":"BMC Med."},{"key":"8866_CR22","doi-asserted-by":"publisher","unstructured":"McDuff, D. et al. Towards accurate differential diagnosis with large language models. Nature https:\/\/doi.org\/10.1038\/s41586-025-08869-4 (2025).","DOI":"10.1038\/s41586-025-08869-4"},{"key":"8866_CR23","doi-asserted-by":"crossref","unstructured":"Semigran, H. L., Linder, J. A., Gidengil, C. & Mehrotra, A. Evaluation of symptom checkers for self diagnosis and triage: audit study. Br. Med. J. 351, h3480 (2015).","DOI":"10.1136\/bmj.h3480"},{"key":"8866_CR24","doi-asserted-by":"crossref","unstructured":"Ayers, J. W. et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern. Med. 183, 589\u2013596 (2023).","DOI":"10.1001\/jamainternmed.2023.1838"},{"key":"8866_CR25","unstructured":"Chatgpt. OpenAI https:\/\/chat.openai.com\/chat (2023)."},{"key":"8866_CR26","doi-asserted-by":"publisher","first-page":"168","DOI":"10.1093\/fampra\/cmab077","volume":"39","author":"S Carrillo de Albornoz","year":"2022","unstructured":"Carrillo de Albornoz, S., Sia, K.-L. & Harris, A. The effectiveness of teleconsultations in primary care: systematic review. Fam. Pract. 39, 168\u2013182 (2022).","journal-title":"Fam. Pract."},{"key":"8866_CR27","doi-asserted-by":"crossref","unstructured":"Fuster-Casanovas, A. & Vidal-Alaball, J. Asynchronous remote communication as a tool for care management in primary care: a rapid review of the literature. Int. J. Integr. Care 22, 7 (2022).","DOI":"10.5334\/ijic.6489"},{"key":"8866_CR28","doi-asserted-by":"publisher","first-page":"e595","DOI":"10.3399\/bjgp19X704573","volume":"69","author":"V Hammersley","year":"2019","unstructured":"Hammersley, V. et al. Comparing the content and quality of video, telephone, and face-to-face consultations: a non-randomised, quasi-experimental, exploratory study in UK primary care. Br. J. Gen. Pract. 69, e595\u2013e604 (2019).","journal-title":"Br. J. Gen. Pract."},{"key":"8866_CR29","first-page":"133","volume":"47","author":"DA Gross","year":"1998","unstructured":"Gross, D. A., Zyzanski, S. J., Borawski, E. A., Cebul, R. D. & Stange, K. C. Patient satisfaction with time spent with their physician. J. Fam. Pract. 47, 133\u2013138 (1998).","journal-title":"J. Fam. Pract."},{"key":"8866_CR30","doi-asserted-by":"publisher","first-page":"1814","DOI":"10.1038\/s41591-023-02437-x","volume":"29","author":"K Dvijotham","year":"2023","unstructured":"Dvijotham, K. et al. Enhancing the reliability and accuracy of AI-enabled diagnosis via complementarity-driven deferral to clinicians. Nat. Med. 29, 1814\u20131820 (2023).","journal-title":"Nat. Med."},{"key":"8866_CR31","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver, D. et al. Mastering the game of Go with deep neural networks and tree search. Nature 529, 484\u2013489 (2016).","journal-title":"Nature"},{"key":"8866_CR32","doi-asserted-by":"crossref","unstructured":"Gallegos, I. O. et al. Bias and fairness in large language models: a survey. Comput. Linguist. 50, 1\u201379 (2024).","DOI":"10.1162\/coli_a_00524"},{"key":"8866_CR33","doi-asserted-by":"publisher","first-page":"2084","DOI":"10.2105\/AJPH.94.12.2084","volume":"94","author":"RL Johnson","year":"2004","unstructured":"Johnson, R. L., Roter, D., Powe, N. R. & Cooper, L. A. Patient race\/ethnicity and quality of patient\u2013physician communication during medical visits. Am. J. Public Health 94, 2084\u20132090 (2004).","journal-title":"Am. J. Public Health"},{"key":"8866_CR34","doi-asserted-by":"publisher","first-page":"756","DOI":"10.1001\/jama.288.6.756","volume":"288","author":"DL Roter","year":"2002","unstructured":"Roter, D. L., Hall, J. A. & Aoki, Y. Physician gender effects in medical communication: a meta-analytic review. JAMA 288, 756\u2013764 (2002).","journal-title":"JAMA"},{"key":"8866_CR35","doi-asserted-by":"publisher","first-page":"eabj2836","DOI":"10.1126\/sciadv.abj2836","volume":"7","author":"D Schillinger","year":"2021","unstructured":"Schillinger, D. et al. Precision communication: physicians\u2019 linguistic adaptation to patients\u2019 health literacy. Sci. Adv. 7, eabj2836 (2021).","journal-title":"Sci. Adv."},{"key":"8866_CR36","doi-asserted-by":"crossref","unstructured":"Rahman, U. & Cooling, N. Inter-cultural communication skills training in medical schools: a systematic review. Med. Res. Arch. 11, mra.v11i4.3757 (2023).","DOI":"10.18103\/mra.v11i4.3757"},{"key":"8866_CR37","unstructured":"Ganguli, D. et al. Red teaming language models to reduce harms: methods, scaling behaviors, and lessons learned. Preprint at https:\/\/arxiv.org\/abs\/2209.07858 (2022)."},{"key":"8866_CR38","doi-asserted-by":"crossref","unstructured":"Mitchell, M. et al. Model cards for model reporting. In Proc. Conference on Fairness, Accountability, and Transparency 220\u2013229 (Association for Computing Machinery, 2019).","DOI":"10.1145\/3287560.3287596"},{"key":"8866_CR39","doi-asserted-by":"crossref","unstructured":"Crisan, A., Drouhard, M., Vig, J. & Rajani, N. Interactive model cards: a human-centered approach to model documentation. In Proc. 2022 ACM Conference on Fairness, Accountability, and Transparency 427\u2013439 (Association for Computing Machinery, 2022).","DOI":"10.1145\/3531146.3533108"},{"key":"8866_CR40","doi-asserted-by":"crossref","unstructured":"Pushkarna, M., Zaldivar, A. & Kjartansson, O. Data cards: purposeful and transparent dataset documentation for responsible AI. In Proc. 2022 ACM Conference on Fairness, Accountability, and Transparency 1776\u20131826 (Association for Computing Machinery, 2022).","DOI":"10.1145\/3531146.3533231"},{"key":"8866_CR41","doi-asserted-by":"crossref","unstructured":"Choudhury, M. & Deshpande, A. How linguistically fair are multilingual pre-trained language models? In Proc. AAAI Conference on Artificial Intelligence Vol. 35 12710\u201312718 (Association for the Advancement of Artificial Intelligence, 2021).","DOI":"10.1609\/aaai.v35i14.17505"},{"key":"8866_CR42","doi-asserted-by":"crossref","unstructured":"Nguyen, X.-P., Aljunied, S. M., Joty, S. & Bing, L. Democratizing LLMs for low-resource languages by leveraging their English dominant abilities with linguistically-diverse prompts. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics Vol. 1 (eds Ku, L.-W. et al.) 3501\u20133516 (Association for Computational Linguistics, 2024).","DOI":"10.18653\/v1\/2024.acl-long.192"},{"key":"8866_CR43","doi-asserted-by":"crossref","unstructured":"Naous, T., Ryan, M. J., Ritter, A. & Xu, W. Having beer after prayer? Measuring cultural bias in large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics Vol. 1 (eds Ku, L.-W. et al.) 16366\u201316393 (Association for Computational Linguistics, 2024).","DOI":"10.18653\/v1\/2024.acl-long.862"},{"key":"8866_CR44","doi-asserted-by":"crossref","unstructured":"Ramesh, K., Sitaram, S. & Choudhury, M. Fairness in language models beyond English: gaps and challenges. In Findings of the Association for Computational Linguistics: EACL 2023 (eds Vlachos, A. & Augenstein, I.) 2106\u20132119 (Association for Computational Linguistics, 2023).","DOI":"10.18653\/v1\/2023.findings-eacl.157"},{"key":"8866_CR45","unstructured":"Hada, R. et al. Are large language model-based evaluators the solution to scaling up multilingual evaluation? In Findings of the Association for Computational Linguistics: EACL 2024 (eds Graham, Y. & Purver, M.) 1051\u20131070 (Association for Computational Linguistics, 2024)."},{"key":"8866_CR46","unstructured":"Quach, V. et al. Conformal language modeling. Preprint at https:\/\/arxiv.org\/abs\/2306.10193 (2023)."},{"key":"8866_CR47","first-page":"29348","volume":"34","author":"A Lazaridou","year":"2021","unstructured":"Lazaridou, A. et al. Mind the gap: assessing temporal generalization in neural language models. Adv. Neural Inf. Process. Syst. 34, 29348\u201329363 (2021).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"8866_CR48","doi-asserted-by":"publisher","first-page":"6421","DOI":"10.3390\/app11146421","volume":"11","author":"D Jin","year":"2021","unstructured":"Jin, D. et al. What disease does this patient have? A large-scale open domain question answering dataset from medical exams. Appl. Sci. 11, 6421 (2021).","journal-title":"Appl. Sci."},{"key":"8866_CR49","doi-asserted-by":"publisher","first-page":"160035","DOI":"10.1038\/sdata.2016.35","volume":"3","author":"AE Johnson","year":"2016","unstructured":"Johnson, A. E. et al. MIMIC-III, a freely accessible critical care database. Sci. Data 3, 160035 (2016).","journal-title":"Sci. Data"},{"key":"8866_CR50","doi-asserted-by":"crossref","unstructured":"Chiu, C.-C. et al. Speech recognition for medical conversations. In Proc. Interspeech (ed. Yegnanarayana, B.) 2972\u20132976 (International Speech Communication Association, 2018).","DOI":"10.21437\/Interspeech.2018-40"},{"key":"8866_CR51","doi-asserted-by":"crossref","unstructured":"Sharma, A., Miner, A., Atkins, D. & Althoff, T. A computational approach to understanding empathy expressed in text-based mental health support. In Proc. 2020 Conference on Empirical Methods in Natural Language Processing (eds Webber, B. et al.) 5263\u20135276 (Association for Computational Linguistics, 2020).","DOI":"10.18653\/v1\/2020.emnlp-main.425"},{"key":"8866_CR52","doi-asserted-by":"publisher","unstructured":"Aksitov, R. et al. Rest meets ReAct: self-improvement for multi-step reasoning LLM agent. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2312.10003 (2023).","DOI":"10.48550\/arXiv.2312.10003"},{"key":"8866_CR53","unstructured":"Abacha, A. B., Yim, W.-W., Adams, G., Snider, N. & Yetisgen-Yildiz, M. Overview of the MEDIQA-chat 2023 shared tasks on the summarization & generation of doctor-patient conversations. In Proc. 5th Clinical Natural Language Processing Workshop (eds Naumann, T. et al.) 503\u2013513 (Association for Computational Linguistics, 2023)."},{"key":"8866_CR54","unstructured":"Ionescu, B. et al. in Experimental IR Meets Multilinguality, Multimodality, and Interaction. CLEF 2023 Lecture Notes in Computer Science Vol. 14163 (eds Arampatzis, A. et al.) 370\u2013396 (Springer, 2023)."},{"key":"8866_CR55","unstructured":"He, Z. et al. DIALMED: a dataset for dialogue-based medication recommendation. In Proc. 29th International Conference on Computational Linguistics (eds Calzolari, N. et al.) 721\u2013733 (International Committee on Computational Linguistics, 2022)."},{"key":"8866_CR56","doi-asserted-by":"crossref","unstructured":"Naseem, U., Bandi, A., Raza, S., Rashid, J. & Chakravarthi, B. R. Incorporating medical knowledge to transformer-based language models for medical dialogue generation. In Proc. 21st Workshop on Biomedical Language Processing (eds Demner-Fushman, D. et al.) 110\u2013115 (Association for Computational Linguistics, 2022).","DOI":"10.18653\/v1\/2022.bionlp-1.10"},{"key":"8866_CR57","doi-asserted-by":"crossref","unstructured":"Horowitz, J. L. in Handbook of Econometrics, Vol. 5 (eds Heckman, J. J. & Leamer, E.) 3159\u20133228 (Elsevier, 2001).","DOI":"10.1016\/S1573-4412(01)05005-X"},{"key":"8866_CR58","doi-asserted-by":"publisher","first-page":"289","DOI":"10.1111\/j.2517-6161.1995.tb02031.x","volume":"57","author":"Y Benjamini","year":"1995","unstructured":"Benjamini, Y. & Hochberg, Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J. R. Stat. Soc. Ser. B Methodol. 57, 289\u2013300 (1995).","journal-title":"J. R. Stat. Soc. Ser. B Methodol."},{"key":"8866_CR59","doi-asserted-by":"crossref","unstructured":"Woolson, R. F. in Wiley Encyclopedia of Clinical Trials (eds D\u2019Agostino, R. B. et al.) 1\u20133 (Wiley, 2007).","DOI":"10.1002\/9780471462422.eoct979"},{"key":"8866_CR60","doi-asserted-by":"publisher","first-page":"159","DOI":"10.1186\/s12909-015-0443-x","volume":"15","author":"KE Keifenheim","year":"2015","unstructured":"Keifenheim, K. E. et al. Teaching history taking to medical students: a systematic review. BMC Med. Educ. 15, 159 (2015).","journal-title":"BMC Med. Educ."},{"key":"8866_CR61","doi-asserted-by":"publisher","first-page":"1157","DOI":"10.1001\/jama.290.9.1157","volume":"290","author":"MJ Yedidia","year":"2003","unstructured":"Yedidia, M. J. et al. Effect of communications training on medical student performance. JAMA 290, 1157\u20131165 (2003).","journal-title":"JAMA"},{"key":"8866_CR62","doi-asserted-by":"publisher","first-page":"93","DOI":"10.1001\/jama.289.1.93","volume":"289","author":"G Makoul","year":"2003","unstructured":"Makoul, G. Communication skills education in medical school and beyond. JAMA 289, 93\u201393 (2003).","journal-title":"JAMA"},{"key":"8866_CR63","doi-asserted-by":"publisher","first-page":"483","DOI":"10.1186\/s12909-021-02892-5","volume":"21","author":"XH Tan","year":"2021","unstructured":"Tan, X. H. et al. Teaching and assessing communication skills in the postgraduate medical setting: a systematic scoping review. BMC Med. Educ. 21, 483 (2021).","journal-title":"BMC Med. Educ."},{"key":"8866_CR64","doi-asserted-by":"publisher","first-page":"e202","DOI":"10.1016\/j.jsurg.2015.06.008","volume":"72","author":"SE Raper","year":"2015","unstructured":"Raper, S. E., Gupta, M., Okusanya, O. & Morris, J. B. Improving communication skills: a course for academic medical center surgery residents and faculty. J. Surg. Educ. 72, e202\u2013e211 (2015).","journal-title":"J. Surg. Educ."},{"key":"8866_CR65","doi-asserted-by":"publisher","first-page":"1100","DOI":"10.1111\/j.1365-2923.2008.03137.x","volume":"42","author":"M Von Fragstein","year":"2008","unstructured":"Von Fragstein, M. et al. UK consensus statement on the content of communication curricula in undergraduate medical education. Med. Educ. 42, 1100\u20131107 (2008).","journal-title":"Med. Educ."},{"key":"8866_CR66","doi-asserted-by":"publisher","first-page":"287","DOI":"10.1016\/j.pec.2008.12.006","volume":"74","author":"H De Haes","year":"2009","unstructured":"De Haes, H. & Bensing, J. Endpoints in medical communication research, proposing a framework of functions and outcomes. Patient Educ. Couns. 74, 287\u2013294 (2009).","journal-title":"Patient Educ. Couns."},{"key":"8866_CR67","doi-asserted-by":"crossref","unstructured":"Epstein, R. M. & Street Jr, R. L. Patient-Centered Communication in Cancer Care: Promoting Healing and Reducing Suffering (National Cancer Institute, 2007).","DOI":"10.1037\/e481972008-001"},{"key":"8866_CR68","first-page":"184","volume":"37","author":"JM Schirmer","year":"2005","unstructured":"Schirmer, J. M. et al. Assessing communication competence: a review of current tools. Fam. Med. 37, 184\u201392 (2005).","journal-title":"Fam. Med."},{"key":"8866_CR69","unstructured":"Nichol, J. R., Sundjaja, J. H. & Nelson, G. Medical History (StatPearls, 2018)."},{"key":"8866_CR70","doi-asserted-by":"publisher","first-page":"592","DOI":"10.1177\/1755738013475436","volume":"6","author":"C Denness","year":"2013","unstructured":"Denness, C. What are consultation models for? InnovAiT 6, 592\u2013599 (2013).","journal-title":"InnovAiT"},{"key":"8866_CR71","doi-asserted-by":"publisher","first-page":"226","DOI":"10.1001\/jama.287.2.226","volume":"287","author":"RM Epstein","year":"2002","unstructured":"Epstein, R. M. & Hundert, E. M. Defining and assessing professional competence. JAMA 287, 226\u2013235 (2002).","journal-title":"JAMA"},{"key":"8866_CR72","doi-asserted-by":"crossref","unstructured":"Chan, S. C. C., Choa, G., Kelly, J., Maru, D. & Rashid, M. A. Implementation of virtual OSCE in health professions education: a systematic review. Med. Educ. 57, 833\u2013843 (2023).","DOI":"10.1111\/medu.15089"},{"key":"8866_CR73","doi-asserted-by":"crossref","unstructured":"Budzianowski, P. et al. MultiWOZ\u2013a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling. In Proc. 2018 Conference on Empirical Methods in Natural Language Processing (eds Riloff, E. et al.) 5016\u20135026 (Association for Computational Linguistics, 2018).","DOI":"10.18653\/v1\/D18-1547"},{"key":"8866_CR74","doi-asserted-by":"crossref","unstructured":"Wei, W., Le, Q., Dai, A. & Li, J. AirDialogue: an environment for goal-oriented dialogue research. In Proc. 2018 Conference on Empirical Methods in Natural Language Processing (eds Riloff, E. et al.) 3844\u20133854 (Association for Computational Linguistics, 2018).","DOI":"10.18653\/v1\/D18-1419"},{"key":"8866_CR75","doi-asserted-by":"crossref","unstructured":"Lin, J., Tomlin, N., Andreas, J. & Eisner, J. Decision-oriented dialogue for human-AI collaboration. Trans. Assoc. Comput. Linguist. 12, 892\u2013911 (2023).","DOI":"10.1162\/tacl_a_00679"},{"key":"8866_CR76","unstructured":"Vaswani, A. et al. Attention is all you need. In Proc. 31st Conference on Neural Information Processing Systems (eds Guyon, I. et al.) 6000\u20136010 (Curran Associates, 2017)."},{"key":"8866_CR77","first-page":"27730","volume":"35","author":"L Ouyang","year":"2022","unstructured":"Ouyang, L. et al. Training language models to follow instructions with human feedback. Adv. Neural Inf. Process. Syst. 35, 27730\u201327744 (2022).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"8866_CR78","doi-asserted-by":"crossref","unstructured":"Zhao, J., Khashabi, D., Khot, T., Sabharwal, A. & Chang, K.-W. Ethical-advice taker: do language models understand natural language interventions? In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 (eds Zong, C. et al.) 4158\u20134164 (Association for Computational Linguistics, 2021).","DOI":"10.18653\/v1\/2021.findings-acl.364"},{"key":"8866_CR79","unstructured":"Saunders, W. et al. Self-critiquing models for assisting human evaluators. Preprint at https:\/\/arxiv.org\/abs\/2206.05802 (2022)."},{"key":"8866_CR80","unstructured":"Scheurer, J. et al. Training language models with language feedback at scale. Preprint at https:\/\/arxiv.org\/abs\/2303.16755 (2023)."},{"key":"8866_CR81","unstructured":"Glaese, A. et al. Improving alignment of dialogue agents via targeted human judgements. Preprint at https:\/\/arxiv.org\/abs\/2209.14375 (2022)."},{"key":"8866_CR82","unstructured":"Bai, Y. et al. Constitutional AI: harmlessness from AI feedback. Preprint at https:\/\/arxiv.org\/abs\/2212.08073 (2022)."},{"key":"8866_CR83","unstructured":"Askell, A. et al. A general language assistant as a laboratory for alignment. Preprint at https:\/\/arxiv.org\/abs\/2112.00861 (2021)."},{"key":"8866_CR84","doi-asserted-by":"crossref","unstructured":"Shor, J. et al. Clinical BERTScore: an improved measure of automatic speech recognition performance in clinical settings. In Proc. 5th Clinical Natural Language Processing Workshop (eds Naumann, T. et al.) 1\u20137 (Association for Computational Linguistics, 2023).","DOI":"10.18653\/v1\/2023.clinicalnlp-1.1"},{"key":"8866_CR85","unstructured":"Abacha, A. B., Agichtein, E., Pinter, Y. & Demner-Fushman, D. Overview of the medical question answering task at TREC 2017 LiveQA. In Proc. 26th Text Retrieval Conference, TREC 2017 (eds Voorhees, E. M. & Ellis, A.) 1\u201312 (National Institute of Standards and Technology and the Defense Advanced Research Projects Agency, 2017)."},{"key":"8866_CR86","doi-asserted-by":"publisher","DOI":"10.1038\/s41746-022-00667-w","volume":"5","author":"W Wallace","year":"2022","unstructured":"Wallace, W. et al. The diagnostic and triage accuracy of digital and online symptom checker tools: a systematic review. NPJ Digit. Med. 5, 118 (2022).","journal-title":"NPJ Digit. Med."},{"key":"8866_CR87","doi-asserted-by":"publisher","first-page":"480","DOI":"10.1016\/j.mcpdig.2023.08.002","volume":"1","author":"D Zeltzer","year":"2023","unstructured":"Zeltzer, D. et al. Diagnostic accuracy of artificial intelligence in virtual primary care. Mayo Clin. Proc. Digital Health 1, 480\u2013489 (2023).","journal-title":"Mayo Clin. Proc. Digital Health"},{"key":"8866_CR88","doi-asserted-by":"publisher","unstructured":"Johri, S. et al. Testing the limits of language models: a conversational framework for medical AI assessment. Preprint at medRxiv https:\/\/doi.org\/10.1101\/2023.09.12.23295399 (2023).","DOI":"10.1101\/2023.09.12.23295399"},{"key":"8866_CR89","unstructured":"Wu, C.-K., Chen, W.-L. & Chen, H.-H. Large language models perform diagnostic reasoning. Preprint at https:\/\/arxiv.org\/abs\/2307.08922 (2023)."},{"key":"8866_CR90","doi-asserted-by":"crossref","unstructured":"Zeng, G. et al. MedDialog: large-scale medical dialogue datasets. In Proc. 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (eds Webber, B. et al.) 9241\u20139250 (Association for Computational Linguistics, 2020).","DOI":"10.18653\/v1\/2020.emnlp-main.743"},{"key":"8866_CR91","doi-asserted-by":"crossref","unstructured":"Liu, W. et al. MedDG: an entity-centric medical consultation dataset for entity-aware medical dialogue generation. In Proc. 11th CCF International Conference on Natural Language Processing and Chinese Computing (eds Lu, W. et al.) 447\u2013459 (Springer, 2022).","DOI":"10.1007\/978-3-031-17120-8_35"},{"key":"8866_CR92","doi-asserted-by":"crossref","unstructured":"Varshney, D., Zafar, A., Behera, N. & Ekbal, A. CDialog: a multi-turn COVID-19 conversation dataset for entity-aware dialog generation. In Proc. 2022 Conference on Empirical Methods in Natural Language Processing (eds Goldberg, Y. et al.) 11373\u201311385 (Association for Computational Linguistics, 2022).","DOI":"10.18653\/v1\/2022.emnlp-main.782"},{"key":"8866_CR93","doi-asserted-by":"crossref","unstructured":"Yan, G. et al. ReMeDi: resources for multi-domain, multi-service, medical dialogues. In Proc. 45th International ACM SIGIR Conference on Research and Development in Information Retrieval 3013\u20133024 (Association for Computing Machinery, 2022).","DOI":"10.1145\/3477495.3531809"}],"container-title":["Nature"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.nature.com\/articles\/s41586-025-08866-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41586-025-08866-7","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41586-025-08866-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,11]],"date-time":"2025-06-11T12:39:03Z","timestamp":1749645543000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.nature.com\/articles\/s41586-025-08866-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,9]]},"references-count":93,"journal-issue":{"issue":"8067","published-print":{"date-parts":[[2025,6,12]]}},"alternative-id":["8866"],"URL":"https:\/\/doi.org\/10.1038\/s41586-025-08866-7","relation":{},"ISSN":["0028-0836","1476-4687"],"issn-type":[{"value":"0028-0836","type":"print"},{"value":"1476-4687","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,9]]},"assertion":[{"value":"18 January 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 March 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 April 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"This study was funded by Alphabet Inc. and\/or a subsidiary thereof (\u2018Alphabet\u2019). All authors are employees of Alphabet and may own stock as part of the standard compensation package.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}]}}