{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,1]],"date-time":"2026-08-01T08:54:37Z","timestamp":1785574477192,"version":"3.56.0"},"reference-count":140,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2025,6,27]],"date-time":"2025-06-27T00:00:00Z","timestamp":1750982400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Canadian Institutes of Health Research Planning and Dissemination Grant\u2014Institute Community Support","award":["CIHR PCS\u2013191021"],"award-info":[{"award-number":["CIHR PCS\u2013191021"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Large language models (LLMs) are transforming the capabilities of medical chatbots by enabling more context-aware, human-like interactions. This review presents a comprehensive analysis of their applications, technical foundations, benefits, challenges, and future directions in healthcare. LLMs are increasingly used in patient-facing roles, such as symptom checking, health information delivery, and mental health support, as well as in clinician-facing applications, including documentation, decision support, and education. However, as a study from 2024 warns, there is a need to manage \u201cextreme AI risks amid rapid progress\u201d. We examine transformer-based architectures, fine-tuning strategies, and evaluation benchmarks specific to medical domains to identify their potential to transfer and mitigate AI risks when using LLMs in medical chatbots. While LLMs offer advantages in scalability, personalization, and 24\/7 accessibility, their deployment in healthcare also raises critical concerns. These include hallucinations (the generation of factually incorrect or misleading content by an AI model), algorithmic biases, privacy risks, and a lack of regulatory clarity. Ethical and legal challenges, such as accountability, explainability, and liability, remain unresolved. Importantly, this review integrates broader insights on AI safety, drawing attention to the systemic risks associated with rapid LLM deployment. As highlighted in recent policy research, including work on managing extreme AI risks, there is an urgent need for governance frameworks that extend beyond technical reliability to include societal oversight and long-term alignment. We advocate for responsible innovation and sustained collaboration among clinicians, developers, ethicists, and regulators to ensure that LLM-powered medical chatbots are deployed safely, equitably, and transparently within healthcare systems.<\/jats:p>","DOI":"10.3390\/info16070549","type":"journal-article","created":{"date-parts":[[2025,6,30]],"date-time":"2025-06-30T13:06:17Z","timestamp":1751288777000},"page":"549","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":67,"title":["Large Language Models in Medical Chatbots: Opportunities, Challenges, and the Need to Address AI Risks"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4202-4855","authenticated-orcid":false,"given":"James C. L.","family":"Chow","sequence":"first","affiliation":[{"name":"Radiation Medicine Program, Princess Margaret Cancer Centre, University Health Network, Toronto, ON M5G 1X6, Canada"},{"name":"Department of Radiation Oncology, University of Toronto, Toronto, ON M5T 1P5, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5765-1635","authenticated-orcid":false,"given":"Kay","family":"Li","sequence":"additional","affiliation":[{"name":"Department of English, University of Toronto, Toronto, ON M5R 0A3, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,6,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"e56628","DOI":"10.2196\/56628","article-title":"Transforming health care through chatbots for medical history-taking and future directions: Comprehensive systematic review","volume":"12","author":"Hindelang","year":"2024","journal-title":"JMIR Med. Inform."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"e27850","DOI":"10.2196\/27850","article-title":"Chatbot for health care and oncology applications using artificial intelligence and machine learning: Systematic review","volume":"7","author":"Xu","year":"2021","journal-title":"JMIR Cancer"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"36","DOI":"10.1145\/365153.365168","article-title":"ELIZA\u2014A computer program for the study of natural language communication between man and machine","volume":"9","author":"Weizenbaum","year":"1966","journal-title":"Commun. ACM"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"220","DOI":"10.3390\/encyclopedia1010021","article-title":"Machine learning in healthcare communication","volume":"1","author":"Siddique","year":"2021","journal-title":"Encyclopedia"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1016\/j.tacc.2021.02.007","article-title":"Natural language processing in medicine: A review","volume":"38","author":"Locke","year":"2021","journal-title":"Trends Anaesth. Crit. Care"},{"key":"ref_6","first-page":"100419","article-title":"Bert-based medical chatbot: Enhancing healthcare communication through natural language understanding","volume":"13","author":"Babu","year":"2024","journal-title":"Explor. Res. Clin. Soc. Pharm."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"837","DOI":"10.3390\/biomedinformatics4010047","article-title":"Generative pre-trained transformer-empowered healthcare conversations: Current trends, challenges, and future directions in large language model-enabled medical chatbots","volume":"4","author":"Chow","year":"2024","journal-title":"BioMedInformatics"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"e2457879","DOI":"10.1001\/jamanetworkopen.2024.57879","article-title":"Large language models for chatbot health advice studies: A systematic review","volume":"8","author":"Huo","year":"2025","journal-title":"JAMA Netw. Open"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Chow, J.C. (2021). Artificial intelligence in radiotherapy and patient care. Artificial Intelligence in Medicine, Springer.","DOI":"10.1007\/978-3-030-58080-3_143-1"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Chakraborty, C., Pal, S., Bhattacharya, M., Dash, S., and Lee, S.S. (2023). Overview of chatbots with special emphasis on artificial intelligence-enabled ChatGPT in medical science. Front. Artif. Intell., 6.","DOI":"10.3389\/frai.2023.1237704"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"792","DOI":"10.1001\/jama.2023.14311","article-title":"Large language models answer medical questions accurately, but can\u2019t match clinicians\u2019 knowledge","volume":"330","author":"Harris","year":"2023","journal-title":"JAMA"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Chow, J.C., Sanders, L., and Li, K. (2023). Impact of ChatGPT on medical chatbots as a disruptive technology. Front. Artif. Intell., 6.","DOI":"10.3389\/frai.2023.1166014"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"842","DOI":"10.1126\/science.adn0117","article-title":"Managing extreme AI risks amid rapid progress","volume":"384","author":"Bengio","year":"2024","journal-title":"Science"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"e52399","DOI":"10.2196\/52399","article-title":"Potential of Large Language Models in Health Care: Delphi Study","volume":"26","author":"Denecke","year":"2024","journal-title":"J. Med. Internet Res."},{"key":"ref_15","first-page":"10","article-title":"Healthcare chatbot using SVM & decision tree","volume":"2","author":"Chandel","year":"2025","journal-title":"Trends Health Inform."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"943","DOI":"10.1038\/s41591-024-03423-7","article-title":"Toward expert-level medical question answering with large language models","volume":"31","author":"Singhal","year":"2025","journal-title":"Nat. Med."},{"key":"ref_17","unstructured":"Liu, Z., Hou, Z., Di, Y., Yang, K., Sang, Z., Xie, C., Yang, J., Liu, S., Wang, J., and Li, C. (2025). Infi-Med: Low-resource medical MLLMs with robust reasoning evaluation. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"319","DOI":"10.3390\/ai4010015","article-title":"Design of an educational chatbot using artificial intelligence in radiotherapy","volume":"4","author":"Chow","year":"2023","journal-title":"AI"},{"key":"ref_19","first-page":"2","article-title":"AI-driven healthcare chatbots: Enhancing access to medical information and lowering healthcare costs","volume":"2","author":"Kumar","year":"2023","journal-title":"J. Artif. Intell. Cloud Comput."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zhang, S., and Song, J. (2024). A chatbot-based question and answer system for the auxiliary diagnosis of chronic diseases based on large language model. Sci. Rep., 14.","DOI":"10.1038\/s41598-024-67429-4"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"e39443","DOI":"10.2196\/39443","article-title":"Learning the treatment process in radiotherapy using an artificial intelligence\u2013assisted chatbot: Development study","volume":"6","author":"Rebelo","year":"2022","journal-title":"JMIR Form. Res."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Shiferaw, M.W., Zheng, T., Winter, A., Mike, L.A., and Chan, L.N. (2024). Assessing the accuracy and quality of artificial intelligence (AI) chatbot-generated responses in making patient-specific drug-therapy and healthcare-related decisions. BMC Med. Inform. Decis. Mak., 24.","DOI":"10.1186\/s12911-024-02824-5"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"e59479","DOI":"10.2196\/59479","article-title":"The opportunities and risks of large language models in mental health","volume":"11","author":"Lawrence","year":"2024","journal-title":"JMIR Ment. Health"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"28","DOI":"10.1080\/15265161.2023.2191050","article-title":"Conversational artificial intelligence and distortions of the psychotherapeutic frame: Issues of boundaries, responsibility, and industry interests","volume":"23","author":"Vagwala","year":"2023","journal-title":"Am. J. Bioeth."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"034002","DOI":"10.1088\/2633-1357\/ac1f88","article-title":"An AI-assisted chatbot for radiation safety education in radiotherapy","volume":"2","author":"Kovacek","year":"2021","journal-title":"IOP SciNotes"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"e58329","DOI":"10.2196\/58329","article-title":"Evaluation framework of large language models in medical documentation: Development and usability study","volume":"26","author":"Seo","year":"2024","journal-title":"J. Med. Internet Res."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"2613","DOI":"10.1038\/s41591-024-03097-1","article-title":"Evaluation and mitigation of the limitations of large language models in clinical decision-making","volume":"30","author":"Hager","year":"2024","journal-title":"Nat. Med."},{"key":"ref_28","unstructured":"Li, L., Zhou, J., Gao, Z., Hua, W., Fan, L., Yu, H., Hagen, L., Zhang, Y., Assimes, T.L., and Hemphill, L. (2024). A scoping review of using large language models (LLMs) to investigate electronic health records (EHRs). arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"337","DOI":"10.1055\/a-2491-3872","article-title":"Extracting international classification of diseases codes from clinical documentation using large language models","volume":"16","author":"Simmons","year":"2025","journal-title":"Appl. Clin. Inform."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1166014","DOI":"10.21037\/jmai-23-106","article-title":"Applying large language model artificial intelligence for retina international classification of diseases (ICD) coding","volume":"6","author":"Ong","year":"2023","journal-title":"J. Med. Artif. Intell."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Chow, J.C., Wong, V., Sanders, L., and Li, K. (2023). Developing an AI-assisted educational chatbot for radiotherapy using the IBM Watson assistant platform. Healthcare, 11.","DOI":"10.3390\/healthcare11172417"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Br\u00fcgge, E., Ricchizzi, S., Arenbeck, M., Keller, M.N., Schur, L., Stummer, W., Holling, M., Lu, M.H., and Darici, D. (2024). Large language models improve clinical decision making of medical students through patient simulation and structured feedback: A randomized controlled trial. BMC Med. Educ., 24.","DOI":"10.1186\/s12909-024-06399-7"},{"key":"ref_33","unstructured":"Schumacher, E., Naik, D., and Kannan, A. (2025). Rare Disease Differential Diagnosis with Large Language Models at Scale: From Abdominal Actinomycosis to Wilson\u2019s Disease. arXiv."},{"key":"ref_34","unstructured":"Yuan, S., Bai, Z., Xu, M., Yang, F., Gao, Y., and Yu, H. (2023). ChatGPT-assisted clinical decision-making for rare genetic metabolic disorders: A preliminary case-based study. Front. Genet., 14."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"75","DOI":"10.1038\/s41746-023-00819-6","article-title":"Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers","volume":"6","author":"Gao","year":"2023","journal-title":"NPJ Digit. Med."},{"key":"ref_36","first-page":"5998","article-title":"Attention Is All You Need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_37","unstructured":"Dauphin, Y.N., Fan, A., Auli, M., and Grangier, D. (2017, January 11\u201315). Language modeling with gated convolutional networks. Proceedings of the International Conference on Machine Learning 2017, Sydney, Australia."},{"key":"ref_38","unstructured":"Alto, V. (2023). Modern Generative AI with ChatGPT and OpenAI Models: Leverage the Capabilities of OpenAI\u2019s LLM for Productivity and Innovation with GPT3 and GPT4, Packt Publishing Ltd."},{"key":"ref_39","unstructured":"Koroteev, M.V. (2021). BERT: A review of applications in natural language processing and understanding. arXiv."},{"key":"ref_40","unstructured":"Anil, R., Dai, A.M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., and Chen, Z. (2023). Palm 2 technical report. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Luo, R., Sun, L., Xia, Y., Qin, T., Zhang, S., Poon, H., and Liu, T.Y. (2022). BioGPT: Generative pre-trained transformer for biomedical text generation and mining. Brief. Bioinform., 23.","DOI":"10.1093\/bib\/bbac409"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"AIoa2300138","DOI":"10.1056\/AIoa2300138","article-title":"Towards Generalist Biomedical AI","volume":"1","author":"Tu","year":"2024","journal-title":"NEJM AI"},{"key":"ref_43","unstructured":"Yang, X., Chen, A., PourNejatian, N., Shin, H.C., Smith, K.E., Parisien, C., Compas, C., Martin, C., Flores, M.G., and Zhang, Y. (2022). Gatortron: A large clinical language model to unlock patient information from unstructured electronic health records. arXiv."},{"key":"ref_44","unstructured":"Ross, E., Kansal, Y., Renzella, J., Vassar, A., and Taylor, A. (2025). Supervised fine-tuning LLMs to behave as pedagogical agents in programming education. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Hao, S., and Duan, L. (2025, January 6\u201311). Online learning from strategic human feedback in LLM fine-tuning. Proceedings of the ICASSP 2025 IEEE International Conference on Acoustics, Speech and Signal Processing, Hyderabad, India.","DOI":"10.1109\/ICASSP49660.2025.10887891"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Chaddad, A., Jiang, Y., and He, C. (2023, January 15\u201317). OpenAI ChatGPT: A potential medical application. Proceedings of the 2023 IEEE International Conference on E-Health Networking, Application & Services (Healthcom), Chongqing, China.","DOI":"10.1109\/Healthcom56612.2023.10472387"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"11483","DOI":"10.1007\/s10639-023-12249-8","article-title":"Few-shot is enough: Exploring ChatGPT prompt engineering method for automatic question generation in English education","volume":"29","author":"Lee","year":"2024","journal-title":"Educ. Inf. Technol."},{"key":"ref_48","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"12663","DOI":"10.1007\/s11063-022-10835-4","article-title":"A comprehensive survey on various fully automatic machine translation evaluation metrics","volume":"55","author":"Chauhan","year":"2023","journal-title":"Neural Process. Lett."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"e45312","DOI":"10.2196\/45312","article-title":"How does ChatGPT perform on the United States Medical Licensing Examination (USMLE)?","volume":"9","author":"Gilson","year":"2023","journal-title":"JMIR Med. Educ."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Zhou, Y., Liu, X., Ning, C., and Wu, J. (2024). Multifaceteval: Multifaceted evaluation to probe LLMs in mastering medical knowledge. arXiv.","DOI":"10.24963\/ijcai.2024\/737"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Jin, Q., Dhingra, B., Liu, Z., Cohen, W.W., and Lu, X. (2019). PubmedQA: A dataset for biomedical research question answering. arXiv.","DOI":"10.18653\/v1\/D19-1259"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Cheng, N., Yan, Z., Wang, Z., Li, Z., Yu, J., Zheng, Z., Tu, K., Xu, J., and Han, W. (2024). Potential and limitations of LLMs in capturing structured semantics: A case study on SRL. International Conference on Intelligent Computing, Springer.","DOI":"10.1007\/978-981-97-5663-6_5"},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"e66633","DOI":"10.2196\/66633","article-title":"Developing effective frameworks for large language model\u2013based medical chatbots: Insights from radiotherapy education with ChatGPT","volume":"11","author":"Chow","year":"2025","journal-title":"JMIR Cancer"},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"255","DOI":"10.1002\/hcs2.61","article-title":"Large language models in health care: Development, applications, and challenges","volume":"2","author":"Yang","year":"2023","journal-title":"Health Care Sci."},{"key":"ref_56","first-page":"5484","article-title":"M3Exam: A multilingual, multimodal, multilevel benchmark for examining large language models","volume":"36","author":"Zhang","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_57","first-page":"448","article-title":"Pedagogical alignment of large language models (LLM) for personalized learning: A survey, trends and challenges","volume":"16","author":"Razafinirina","year":"2024","journal-title":"J. Intell. Learn. Syst. Appl."},{"key":"ref_58","doi-asserted-by":"crossref","first-page":"e48291","DOI":"10.2196\/48291","article-title":"Large language models in medical education: Opportunities, challenges, and future directions","volume":"9","author":"AlSaad","year":"2023","journal-title":"JMIR Med. Educ."},{"key":"ref_59","doi-asserted-by":"crossref","first-page":"100085","DOI":"10.1016\/j.apjo.2024.100085","article-title":"Understanding natural language: Potential application of large language models to ophthalmology","volume":"13","author":"Yang","year":"2024","journal-title":"Asia-Pac. J. Ophthalmol."},{"key":"ref_60","doi-asserted-by":"crossref","first-page":"4566","DOI":"10.7150\/thno.108552","article-title":"Artificial intelligence in chronic kidney disease management: A scoping review","volume":"15","author":"Sabanayagam","year":"2025","journal-title":"Theranostics"},{"key":"ref_61","first-page":"23","article-title":"Digital Health Literacy: Empowering Patients in the Era of Electronic Medical Records","volume":"6","author":"Mulukuntla","year":"2020","journal-title":"EPH Int. J. Med. Health Sci."},{"key":"ref_62","doi-asserted-by":"crossref","unstructured":"De Busser, B., Roth, L., and De Loof, H. (2024). The role of large language models in self-care: A study and benchmark on medicines and supplement guidance accuracy. Int. J. Clin. Pharm., 1\u201310.","DOI":"10.1007\/s11096-024-01839-2"},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Kohler, A., Tingstrom, P., Jaarsma, T., and Nilsson, S. (2018). Patient empowerment and general self-efficacy in patients with coronary heart disease: A cross-sectional study. BMC Fam Pract., 19.","DOI":"10.1186\/s12875-018-0749-y"},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"McAllister, M., Dunn, G., Payne, K., Davies, L., and Todd, C. (2012). Patient empowerment: The need to consider it as a measurable patient-reported outcome for chronic conditions. BMC Health Serv. Res., 12.","DOI":"10.1186\/1472-6963-12-157"},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Lin, C., and Kuo, C.F. (2025). Roles and potential of large language models in healthcare: A comprehensive review. Biomed. J.","DOI":"10.1016\/j.bj.2025.100868"},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"Dagli, M.M., Ghenbot, Y., Ahmad, H.S., Chauhan, D., Turlip, R., Wang, P., Welch, W.C., Ozturk, A.K., and Yoon, J.W. (2024). Development and validation of a novel AI framework using NLP with LLM integration for relevant clinical data extraction through automated chart review. Sci. Rep., 14.","DOI":"10.1038\/s41598-024-77535-y"},{"key":"ref_67","first-page":"1234","article-title":"Implementation of an All-Day Artificial Intelligence\u2013Based Triage System in the Emergency Department","volume":"97","author":"Kowalski","year":"2022","journal-title":"Mayo Clin. Proc."},{"key":"ref_68","unstructured":"Nuance Communications (2025, June 16). DAX Copilot to Automate the Creation of Clinical Documentation, Reduce Physician Burnout, and Expand Access to Care Deployed Enterprise-Wide at Stanford Health Care, Nuance Newsroom, Available online: https:\/\/news.nuance.com\/2024-03-11-DAX-Copilot-to-Automate-the-Creation-of-Clinical-Documentation,-Reduce-Physician-Burnout,-and-Expand-Access-to-Care-Deployed-Enterprise-Wide-at-Stanford-Health-Care."},{"key":"ref_69","unstructured":"RTB AI (2025, June 16). LLM Cost Explained: How to Estimate the Price of Using Large Language Models, RTB AI Blog, Available online: https:\/\/www.rtb-ai.com\/post\/how-to-estimate-the-cost-of-using-a-large-language-model."},{"key":"ref_70","unstructured":"Sun, X., Ma, R., Zhao, X., Li, Z., Lindqvist, J., El Ali, A., and Bosch, J.A. (2024). Trusting the Search: Unraveling Human Trust in Health Information from Google and ChatGPT. arXiv."},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"Rezaeian, O., Asan, O., and Bayrak, A.E. (2024). The Impact of AI Explanations on Clinicians\u2019 Trust and Diagnostic Accuracy in Breast Cancer. arXiv.","DOI":"10.1016\/j.apergo.2025.104577"},{"key":"ref_72","first-page":"e23429","article-title":"Exploring the Impact of Chatbot Design on Trust and Engagement in Mental Health Support","volume":"8","author":"Talbot","year":"2021","journal-title":"JMIR Ment. Health"},{"key":"ref_73","doi-asserted-by":"crossref","first-page":"e46901","DOI":"10.2196\/51712","article-title":"The Potential of Chatbots for Emotional Support and Promoting Mental Well-Being in Different Cultures: Mixed Methods Study","volume":"25","author":"Chin","year":"2023","journal-title":"J. Med. Internet Res."},{"key":"ref_74","doi-asserted-by":"crossref","first-page":"e54345","DOI":"10.2196\/54345","article-title":"Reference hallucination score for medical artificial intelligence chatbots: Development and usability study","volume":"12","author":"Aljamaan","year":"2024","journal-title":"JMIR Med. Inform."},{"key":"ref_75","doi-asserted-by":"crossref","first-page":"77","DOI":"10.1038\/s41591-024-03328-5","article-title":"An evaluation framework for clinical use of large language models in patient interaction tasks","volume":"31","author":"Johri","year":"2025","journal-title":"Nat. Med."},{"key":"ref_76","first-page":"127129","article-title":"STaRK: Benchmarking LLM retrieval on textual and relational knowledge bases","volume":"37","author":"Wu","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_77","doi-asserted-by":"crossref","first-page":"AIp2400979","DOI":"10.1056\/AIp2400979","article-title":"The burden of reviewing LLM-generated content","volume":"2","author":"Ohde","year":"2025","journal-title":"NEJM AI"},{"key":"ref_78","doi-asserted-by":"crossref","first-page":"e078538","DOI":"10.1136\/bmj-2023-078538","article-title":"Current safeguards, risk mitigation, and transparency measures of large language models against the generation of health disinformation: Repeated cross-sectional analysis","volume":"384","author":"Menz","year":"2024","journal-title":"BMJ"},{"key":"ref_79","doi-asserted-by":"crossref","first-page":"e53616","DOI":"10.2196\/53616","article-title":"Benefits and risks of AI in health care: Narrative review","volume":"13","author":"Chustecki","year":"2024","journal-title":"Interact. J. Med. Res."},{"key":"ref_80","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1089\/cyber.2024.0199","article-title":"Evaluating for evidence of sociodemographic bias in conversational AI for mental health support","volume":"28","author":"Yeo","year":"2025","journal-title":"Cyberpsychol. Behav. Soc. Netw."},{"key":"ref_81","doi-asserted-by":"crossref","unstructured":"Tan, D.N.H., Tham, Y.-C., Koh, V., Loon, S.C., Aquino, M.C., Lun, K., Cheng, C.-Y., Ngiam, K.Y., and Tan, M. (2024). Evaluating chatbot responses to patient questions in the field of glaucoma. Front. Med., 11.","DOI":"10.3389\/fmed.2024.1359073"},{"key":"ref_82","unstructured":"Chhikara, G., Sharma, A., Ghosh, K., and Chakraborty, A. (2024). Few-shot fairness: Unveiling LLM\u2019s potential for fairness-aware classification. arXiv."},{"key":"ref_83","doi-asserted-by":"crossref","first-page":"e47532","DOI":"10.2196\/47532","article-title":"The Accuracy and Potential Racial and Ethnic Biases of GPT-4 in the Diagnosis and Triage of Health Conditions: Evaluation Study","volume":"9","author":"Ito","year":"2023","journal-title":"JMIR Med. Educ."},{"key":"ref_84","doi-asserted-by":"crossref","first-page":"S64","DOI":"10.1093\/ajcp\/aqad150.143","article-title":"Assessing the Impact of Race, Sexual Orientation, and Gender Identity on USMLE Style Questions: Preliminary Results of a Randomized Controlled Trial","volume":"160","author":"Arnason","year":"2023","journal-title":"Am. J. Clin. Pathol."},{"key":"ref_85","doi-asserted-by":"crossref","unstructured":"Rawat, R., McBride, H., Nirmal, D., Ghosh, R., Moon, J., Alamuri, D., O\u2019Brien, S., and Zhu, K. (2024). DiversityMedQA: Assessing Demographic Biases in Medical Diagnosis using Large Language Models. arXiv.","DOI":"10.18653\/v1\/2024.nlp4pi-1.29"},{"key":"ref_86","doi-asserted-by":"crossref","unstructured":"Rathod, V., Nabavirazavi, S., Zad, S., and Iyengar, S.S. (2025, January 6\u20138). Privacy and security challenges in large language models. Proceedings of the 2025 IEEE 15th Annual Computer Communication Workshop Conference (CCWC), Las Vegas, NV, USA.","DOI":"10.1109\/CCWC62904.2025.10903912"},{"key":"ref_87","doi-asserted-by":"crossref","first-page":"2773","DOI":"10.3390\/ai5040134","article-title":"Trustworthy AI: Securing sensitive data in large language models","volume":"5","author":"Feretzakis","year":"2024","journal-title":"AI"},{"key":"ref_88","first-page":"8","article-title":"Navigating the challenges and opportunities of AI and LLM integration in cloud computing","volume":"2","author":"Srinivasan","year":"2025","journal-title":"Baltic Multidiscip. Res. Lett. J."},{"key":"ref_89","doi-asserted-by":"crossref","unstructured":"Siemens, W., von Elm, E., Binder, H., B\u00f6hringer, D., Eisele-Metzger, A., Gartlehner, G., Hanegraaf, P., Metzendorf, M.-I., Mosselman, J.-J., and Nowak, A. (2025). Opportunities, challenges and risks of using artificial intelligence for evidence synthesis. BMJ Evid.-Based Med.","DOI":"10.1136\/bmjebm-2024-113320"},{"key":"ref_90","doi-asserted-by":"crossref","unstructured":"Khan, Z.Y., Hussain, F.K., and Kurniawan, D. (2024). Communication Efficiency and Non-Independent and Identically Distributed Data Challenge in Federated Learning: A Systematic Mapping Study. Appl. Sci., 14.","DOI":"10.3390\/app14072720"},{"key":"ref_91","unstructured":"Xu, H., Zhang, J., Xu, Y., Liu, Z., Ding, S., and Zhang, Y. (2024). FedVCK: Non-IID Robust and Communication-Efficient Federated Learning via Valuable Condensed Knowledge for Medical Image Analysis. arXiv."},{"key":"ref_92","first-page":"1","article-title":"SCAFFOLD: Stochastic Controlled Averaging for Federated Learning","volume":"21","author":"Karimireddy","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"ref_93","unstructured":"Yao, Z., Zhang, Z., Tang, C., Bian, X., Zhao, Y., Yang, Z., Wang, J., Zhou, H., Jang, W.S., and Ouyang, F. (2024). MedQA-CS: Benchmarking large language models clinical skills using an AI-SCE framework. arXiv."},{"key":"ref_94","unstructured":"Srinivasan, V., Jatav, V., Chandrababu, A., and Sharma, G. (2025). On the performance of an explainable language model on PubMedQA. arXiv."},{"key":"ref_95","first-page":"87621","article-title":"MedJourney: Benchmark and evaluation of large language models over patient clinical journey","volume":"37","author":"Wu","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_96","doi-asserted-by":"crossref","unstructured":"Ch\u2019en, P.Y., Day, W., Pekson, R.C., Barrientos, J., Burton, W.B., Ludwig, A.B., Jariwala, S.P., and Cassese, T. (2025). GPT-4 generated answer rationales to multiple choice assessment questions in undergraduate medical education. BMC Med. Educ., 25.","DOI":"10.1186\/s12909-025-06862-z"},{"key":"ref_97","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2002, January 6\u201312). BLEU: A Method for Automatic Evaluation of Machine Translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL), Philadelphia, PA, USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_98","unstructured":"Van Veen, M., Saavedra, J., and Wang, K. (2025, June 16). Clinical Text Summarization with LLM-Based Evaluation. Stanford University CS224N Final Report, 2024. Available online: https:\/\/web.stanford.edu\/class\/cs224n\/final-reports\/256989380.pdf."},{"key":"ref_99","unstructured":"Schroeder, N.L., Davis Jaldi, C., and Zhang, S. (2025). Large Language Models with Human-In-The-Loop Validation for Systematic Reviews. arXiv."},{"key":"ref_100","first-page":"98","article-title":"Human-in-the-loop AI reviewing: Feasibility, opportunities, and risks","volume":"25","author":"Drori","year":"2024","journal-title":"J. Assoc. Inf. Syst."},{"key":"ref_101","doi-asserted-by":"crossref","first-page":"75735","DOI":"10.1109\/ACCESS.2024.3401547","article-title":"Applications, challenges, and future directions of human-in-the-loop learning","volume":"12","author":"Kumar","year":"2024","journal-title":"IEEE Access"},{"key":"ref_102","doi-asserted-by":"crossref","first-page":"84","DOI":"10.15265\/IY-2017-014","article-title":"Are we there yet? Human factors knowledge and health information technology\u2013the challenges of implementation and impact","volume":"26","author":"Turner","year":"2017","journal-title":"Yearb. Med. Inform."},{"key":"ref_103","doi-asserted-by":"crossref","first-page":"1250","DOI":"10.1007\/s10803-018-3833-1","article-title":"Extending the parent-delivered Early Start Denver Model to young children with fragile X syndrome","volume":"49","author":"Vismara","year":"2019","journal-title":"J. Autism Dev. Disord."},{"key":"ref_104","unstructured":"Mayo Clinic Press (2023). AI in Healthcare: The Future of Patient Care and Health Management, Mayo Clinic Press Healthy Aging. Available online: https:\/\/mcpress.mayoclinic.org\/healthy-aging\/ai-in-healthcare-the-future-of-patient-care-and-health-management\/."},{"key":"ref_105","unstructured":"Stanford HAI (2025, May 16). Large Language Models in Healthcare: Are We There Yet?, Stanford HAI News, Available online: https:\/\/hai.stanford.edu\/news\/large-language-models-healthcare-are-we-there-yet."},{"key":"ref_106","first-page":"3","article-title":"Data sources (LLM) for a clinical decision support model (SSDC) using a healthcare interoperability resources (HL7-FHIR) platform for an ICU ecosystem","volume":"5","author":"Plaza","year":"2024","journal-title":"Prim. Sci. Med. Public Health"},{"key":"ref_107","doi-asserted-by":"crossref","unstructured":"Nazi, Z.A., and Peng, W. (2024). Large language models in healthcare and medical domain: A review. Informatics, 11.","DOI":"10.3390\/informatics11030057"},{"key":"ref_108","doi-asserted-by":"crossref","unstructured":"Markus, A.F., Kors, J.A., and Rijnbeek, P.R. (2020). The Role of Explainability in Creating Trustworthy Artificial Intelligence for Health Care: A Comprehensive Survey of the Terminology, Design Choices, and Evaluation Strategies. arXiv.","DOI":"10.1016\/j.jbi.2020.103655"},{"key":"ref_109","doi-asserted-by":"crossref","unstructured":"Bharati, S., Mondal, M.R.H., and Podder, P. (2023). A Review on Explainable Artificial Intelligence for Healthcare: Why, How, and When?. arXiv.","DOI":"10.1109\/TAI.2023.3266418"},{"key":"ref_110","unstructured":"Ong, J.C.L., Ning, Y., Liu, M., Ma, Y., Liang, Z., Singh, K., Chang, R.T., Vogel, S., Lim, J.C.W., and Tan, I.S.K. (2025). Regulatory Science Innovation for Generative AI and Large Language Models in Health and Medicine: A Global Call for Action. arXiv."},{"key":"ref_111","unstructured":"U.S. Food and Drug Administration (2025, June 16). Artificial Intelligence\/Machine Learning (AI\/ML)-Based Software as a Medical Device (SaMD) Action Plan. FDA 2021, Available online: https:\/\/www.fda.gov\/media\/145022\/download."},{"key":"ref_112","doi-asserted-by":"crossref","first-page":"120","DOI":"10.1038\/s41746-023-00873-0","article-title":"The Imperative for Regulatory Oversight of Large Language Models (or Generative AI) in Healthcare","volume":"6","author":"Topol","year":"2023","journal-title":"NPJ Digit. Med."},{"key":"ref_113","doi-asserted-by":"crossref","unstructured":"Chow, J.C., and Li, K. (2024). Ethical considerations in human-centered AI: Advancing oncology chatbots through large language models. JMIR Bioinform. Biotechnol., 5.","DOI":"10.2196\/64406"},{"key":"ref_114","unstructured":"Ge, Z., Huang, H., Zhou, M., Li, J., Wang, G., Tang, S., and Zhuang, Y. (November, January 28). WorldGPT: Empowering LLM as multimodal world model. Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, VIC, Australia."},{"key":"ref_115","first-page":"28541","article-title":"Llava-med: Training a large language-and-vision assistant for biomedicine in one day","volume":"36","author":"Li","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_116","unstructured":"Koleilat, T., Asgariandehkordi, H., Rivaz, H., and Xiao, Y. (2024). Medclip-samv2: Towards universal text-driven medical image segmentation. arXiv."},{"key":"ref_117","unstructured":"Jeong, C. (2024). Fine-tuning and utilization methods of domain-specific LLMs. arXiv."},{"key":"ref_118","doi-asserted-by":"crossref","first-page":"445","DOI":"10.1038\/s44222-025-00279-5","article-title":"Application of large language models in medicine","volume":"3","author":"Liu","year":"2025","journal-title":"Nat. Rev. Bioeng."},{"key":"ref_119","first-page":"11339","article-title":"CXR-LLaVA: A Multimodal Large Language Model for Interpreting Chest X-Ray Images","volume":"303","author":"Lee","year":"2024","journal-title":"Radiology"},{"key":"ref_120","doi-asserted-by":"crossref","unstructured":"Wang, Z., Wu, Z., Agarwal, D., and Sun, J. (2022, January 7\u201311). MedCLIP: Contrastive Learning from Unpaired Medical Images and Text. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), Abu Dhabi, United Arab Emirates.","DOI":"10.18653\/v1\/2022.emnlp-main.256"},{"key":"ref_121","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.imed.2025.01.001","article-title":"Large language models-powered clinical decision support: Enhancing or replacing human expertise?","volume":"5","author":"Li","year":"2025","journal-title":"Intell. Med."},{"key":"ref_122","doi-asserted-by":"crossref","unstructured":"Rajashekar, N.C., Shin, Y.E., Pu, Y., Chung, S., You, K., Giuffr\u00e8, M., Chan, C.E., Saarinen, T., Hsiao, A., and Sekhon, J.S. (2024, January 11\u201316). Human-algorithmic interaction using a large language model-augmented artificial intelligence clinical decision support system. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA.","DOI":"10.1145\/3613904.3642024"},{"key":"ref_123","doi-asserted-by":"crossref","unstructured":"Chow, J.C. (2025). Quantum computing and machine learning in medical decision-making: A comprehensive review. Algorithms, 18.","DOI":"10.3390\/a18030156"},{"key":"ref_124","doi-asserted-by":"crossref","first-page":"107093","DOI":"10.1016\/j.chb.2021.107093","article-title":"User interactions with chatbot interfaces vs. menu-based interfaces: An empirical study","volume":"128","author":"Nguyen","year":"2022","journal-title":"Comput. Hum. Behav."},{"key":"ref_125","doi-asserted-by":"crossref","first-page":"271","DOI":"10.1080\/10447318.2023.2298534","article-title":"Investigating usability of conversational user interfaces for integrated system-physical interactions: A medical device perspective","volume":"41","author":"Dutta","year":"2025","journal-title":"Int. J. Hum. Comput. Interact."},{"key":"ref_126","doi-asserted-by":"crossref","first-page":"2464746","DOI":"10.1080\/17517575.2025.2464746","article-title":"Development of an intelligent hospital information chatbot and evaluation of its system usability","volume":"19","author":"Chen","year":"2025","journal-title":"Enterp. Inf. Syst."},{"key":"ref_127","doi-asserted-by":"crossref","first-page":"108055","DOI":"10.1016\/j.pec.2023.108055","article-title":"Examining how information presentation methods and a chatbot impact the use and effectiveness of electronic health record patient portals: An exploratory study","volume":"119","author":"Yin","year":"2024","journal-title":"Patient Educ. Couns."},{"key":"ref_128","doi-asserted-by":"crossref","first-page":"4876512","DOI":"10.1155\/2022\/4876512","article-title":"The health chatbots in telemedicine: Intelligent dialog system for remote support","volume":"2022","author":"Vasileiou","year":"2022","journal-title":"J. Healthc. Eng."},{"key":"ref_129","doi-asserted-by":"crossref","first-page":"2302980","DOI":"10.1080\/07853890.2024.2302980","article-title":"A systematic review of artificial intelligence-powered (AI-powered) chatbot intervention for managing chronic illness","volume":"56","author":"Kurniawan","year":"2024","journal-title":"Ann. Med."},{"key":"ref_130","doi-asserted-by":"crossref","first-page":"e47551","DOI":"10.2196\/47551","article-title":"Security implications of AI chatbots in health care","volume":"25","author":"Li","year":"2023","journal-title":"J. Med. Internet Res."},{"key":"ref_131","doi-asserted-by":"crossref","first-page":"10859","DOI":"10.1007\/s10586-024-04515-2","article-title":"Comprehensive framework for implementing blockchain-enabled federated learning and full homomorphic encryption for chatbot security system","volume":"27","author":"Jalali","year":"2024","journal-title":"Clust. Comput."},{"key":"ref_132","doi-asserted-by":"crossref","first-page":"311","DOI":"10.1001\/jama.2023.9618","article-title":"Health care privacy risks of AI chatbots","volume":"330","author":"Kanter","year":"2023","journal-title":"JAMA"},{"key":"ref_133","doi-asserted-by":"crossref","first-page":"527","DOI":"10.1007\/s44230-024-00085-z","article-title":"PharmaLLM: A medicine prescriber chatbot exploiting open-source large language models","volume":"4","author":"Azam","year":"2024","journal-title":"Hum. Cent. Intell. Syst."},{"key":"ref_134","doi-asserted-by":"crossref","first-page":"e27460","DOI":"10.2196\/27460","article-title":"Medical specialty recommendations by an artificial intelligence chatbot on a smartphone: Development and deployment","volume":"23","author":"Lee","year":"2021","journal-title":"J. Med. Internet Res."},{"key":"ref_135","doi-asserted-by":"crossref","first-page":"1597","DOI":"10.1001\/jamaoncol.2024.4324","article-title":"Ensuring safety and consistency in artificial intelligence chatbot responses","volume":"10","author":"Zhu","year":"2024","journal-title":"JAMA Oncol."},{"key":"ref_136","doi-asserted-by":"crossref","first-page":"949","DOI":"10.30574\/wjarr.2025.25.3.0807","article-title":"Regulatory and legal challenges of artificial intelligence in the US healthcare system: Liability, compliance, and patient safety","volume":"25","author":"Osifowokan","year":"2025","journal-title":"World J. Adv. Res. Rev."},{"key":"ref_137","doi-asserted-by":"crossref","first-page":"241","DOI":"10.1001\/jama.2024.21451","article-title":"FDA perspective on the regulation of artificial intelligence in health care and biomedicine","volume":"333","author":"Warraich","year":"2025","journal-title":"JAMA"},{"key":"ref_138","doi-asserted-by":"crossref","first-page":"293","DOI":"10.1016\/j.clinthera.2023.12.006","article-title":"A future European scientific dialogue regulatory framework: Connecting the dots","volume":"46","author":"Hey","year":"2024","journal-title":"Clin. Ther."},{"key":"ref_139","doi-asserted-by":"crossref","unstructured":"Palaniappan, K., Lin, E.Y., and Vogel, S. (2024). Global regulatory frameworks for the use of artificial intelligence (AI) in the healthcare services sector. Healthcare, 12.","DOI":"10.3390\/healthcare12050562"},{"key":"ref_140","doi-asserted-by":"crossref","first-page":"241873","DOI":"10.1098\/rsos.241873","article-title":"Ethical and legal considerations in healthcare AI: Innovation and policy for safe and fair use","volume":"12","author":"Pham","year":"2025","journal-title":"R. Soc. Open Sci."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/7\/549\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:00:18Z","timestamp":1760032818000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/7\/549"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,27]]},"references-count":140,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2025,7]]}},"alternative-id":["info16070549"],"URL":"https:\/\/doi.org\/10.3390\/info16070549","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,27]]}}}