{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T09:31:16Z","timestamp":1780047076845,"version":"3.53.1"},"reference-count":31,"publisher":"Oxford University Press (OUP)","issue":"4","license":[{"start":{"date-parts":[[2024,1,22]],"date-time":"2024-01-22T00:00:00Z","timestamp":1705881600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/pages\/standard-publication-reuse-rights"}],"funder":[{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["R01GM114355"],"award-info":[{"award-number":["R01GM114355"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["R01LM013486"],"award-info":[{"award-number":["R01LM013486"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Woods Foundation"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2024,4,3]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Objective<\/jats:title>\n                  <jats:p>Large language models (LLMs) have shown impressive ability in biomedical question-answering, but have not been adequately investigated for more specific biomedical applications. This study investigates ChatGPT family of models (GPT-3.5, GPT-4) in biomedical tasks beyond question-answering.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Materials and Methods<\/jats:title>\n                  <jats:p>We evaluated model performance with 11 122 samples for two fundamental tasks in the biomedical domain\u2014classification (n\u2009=\u20098676) and reasoning (n\u2009=\u20092446). The first task involves classifying health advice in scientific literature, while the second task is detecting causal relations in biomedical literature. We used 20% of the dataset for prompt development, including zero- and few-shot settings with and without chain-of-thought (CoT). We then evaluated the best prompts from each setting on the remaining dataset, comparing them to models using simple features (BoW with logistic regression) and fine-tuned BioBERT models.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>Fine-tuning BioBERT produced the best classification (F1: 0.800-0.902) and reasoning (F1: 0.851) results. Among LLM approaches, few-shot CoT achieved the best classification (F1: 0.671-0.770) and reasoning (F1: 0.682) results, comparable to the BoW model (F1: 0.602-0.753 and 0.675 for classification and reasoning, respectively). It took 78 h to obtain the best LLM results, compared to 0.078 and 0.008 h for the top-performing BioBERT and BoW models, respectively.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Discussion<\/jats:title>\n                  <jats:p>The simple BoW model performed similarly to the most complex LLM prompting. Prompt engineering required significant investment.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Conclusion<\/jats:title>\n                  <jats:p>Despite the excitement around viral ChatGPT, fine-tuning for two fundamental biomedical natural language processing tasks remained the best strategy.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/jamia\/ocad256","type":"journal-article","created":{"date-parts":[[2024,1,23]],"date-time":"2024-01-23T16:51:16Z","timestamp":1706028676000},"page":"940-948","source":"Crossref","is-referenced-by-count":68,"title":["Evaluating the ChatGPT family of models for biomedical reasoning and classification"],"prefix":"10.1093","volume":"31","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7999-7410","authenticated-orcid":false,"given":"Shan","family":"Chen","sequence":"first","affiliation":[{"name":"Artificial Intelligence in Medicine (AIM) Program, Mass General Brigham, Harvard Medical School , Boston, MA 02115, United States"},{"name":"Department of Radiation Oncology, Brigham and Women\u2019s Hospital\/Dana-Farber Cancer Institute , Boston, MA 02115, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yingya","family":"Li","sequence":"additional","affiliation":[{"name":"Computational Health Informatics Program, Boston Children\u2019s Hospital, and Harvard Medical School , Boston, MA 02115, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sheng","family":"Lu","sequence":"additional","affiliation":[{"name":"Ubiquitous Knowledge Processing Lab (UKP Lab), Technical University of Darmstadt , Darmstadt 64289, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hoang","family":"Van","sequence":"additional","affiliation":[{"name":"Computational Health Informatics Program, Boston Children\u2019s Hospital, and Harvard Medical School , Boston, MA 02115, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hugo J W L","family":"Aerts","sequence":"additional","affiliation":[{"name":"Artificial Intelligence in Medicine (AIM) Program, Mass General Brigham, Harvard Medical School , Boston, MA 02115, United States"},{"name":"Department of Radiation Oncology, Brigham and Women\u2019s Hospital\/Dana-Farber Cancer Institute , Boston, MA 02115, United States"},{"name":"Radiology and Nuclear Medicine, GROW & CARIM, Maastricht University , Maastricht 6211 LK, Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guergana K","family":"Savova","sequence":"additional","affiliation":[{"name":"Computational Health Informatics Program, Boston Children\u2019s Hospital, and Harvard Medical School , Boston, MA 02115, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Danielle S","family":"Bitterman","sequence":"additional","affiliation":[{"name":"Artificial Intelligence in Medicine (AIM) Program, Mass General Brigham, Harvard Medical School , Boston, MA 02115, United States"},{"name":"Department of Radiation Oncology, Brigham and Women\u2019s Hospital\/Dana-Farber Cancer Institute , Boston, MA 02115, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2024,1,22]]},"reference":[{"key":"2024040320031163400_ocad256-B1","author":"Vaswani","year":"2017"},{"key":"2024040320031163400_ocad256-B2","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton","year":"2018","edition":"2nd ed"},{"key":"2024040320031163400_ocad256-B3","author":"Ouyang","year":"2022"},{"key":"2024040320031163400_ocad256-B4","author":"Ouyang","year":"2022"},{"issue":"13","key":"2024040320031163400_ocad256-B5","doi-asserted-by":"crossref","first-page":"1233","DOI":"10.1056\/NEJMsr2214184","article-title":"Benefits, limits, and risks of GPT-4 as an AI Chatbot for medicine","volume":"388","author":"Lee","year":"2023","journal-title":"N Engl J Med"},{"key":"2024040320031163400_ocad256-B6","author":"Reardon"},{"key":"2024040320031163400_ocad256-B7","doi-asserted-by":"crossref","first-page":"e45312","DOI":"10.2196\/45312","article-title":"How does ChatGPT perform on the United States medical licensing examination? The implications of large language models for medical education and knowledge assessment","volume":"9","author":"Gilson","year":"2023","journal-title":"JMIR Med Educ"},{"key":"2024040320031163400_ocad256-B8","author":"Li\u00e9vin","year":"2022"},{"key":"2024040320031163400_ocad256-B9","author":"Zuccon","year":"2023"},{"issue":"10","key":"2024040320031163400_ocad256-B10","doi-asserted-by":"crossref","first-page":"1459","DOI":"10.1001\/jamaoncol.2023.2954","article-title":"Use of artificial intelligence Chatbots for cancer treatment information","volume":"9","author":"Chen","year":"2023","journal-title":"JAMA Oncol"},{"key":"2024040320031163400_ocad256-B11","author":"Lyu","year":"2023"},{"key":"2024040320031163400_ocad256-B12","author":"Singhal","year":"2022"},{"key":"2024040320031163400_ocad256-B13","author":"Lehman","year":"2023"},{"key":"2024040320031163400_ocad256-B14","author":"Wang","year":"2023"},{"key":"2024040320031163400_ocad256-B15","author":"OpenAI API [Internet]"},{"key":"2024040320031163400_ocad256-B16","first-page":"6018","author":"Li","year":"2021"},{"key":"2024040320031163400_ocad256-B17","first-page":"4664","author":"Yu","year":"2019"},{"key":"2024040320031163400_ocad256-B18","author":"Devlin","year":"2018"},{"issue":"4","key":"2024040320031163400_ocad256-B19","doi-asserted-by":"crossref","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","article-title":"BioBERT: a pre-trained biomedical language representation model for biomedical text mining","volume":"36","author":"Lee","year":"2020","journal-title":"Bioinformatics"},{"key":"2024040320031163400_ocad256-B20","author":"Wei","year":"2022"},{"key":"2024040320031163400_ocad256-B21","author":"Taylor","year":"2022"},{"key":"2024040320031163400_ocad256-B22","author":"Brown","year":"2020"},{"key":"2024040320031163400_ocad256-B23","author":"Wei","year":"2022"},{"key":"2024040320031163400_ocad256-B24","author":"Kojima","year":"2022"},{"key":"2024040320031163400_ocad256-B25","author":"Shi","year":"2023"},{"key":"2024040320031163400_ocad256-B26","author":"Wang","year":"2022"},{"issue":"21","key":"2024040320031163400_ocad256-B27","doi-asserted-by":"crossref","first-page":"5463","DOI":"10.1158\/0008-5472.CAN-19-0579","article-title":"Use of natural language processing to extract clinical cancer phenotypes from electronic medical records","volume":"79","author":"Savova","year":"2019","journal-title":"Cancer Res"},{"issue":"9","key":"2024040320031163400_ocad256-B28","doi-asserted-by":"crossref","first-page":"977","DOI":"10.1001\/jamapediatrics.2023.2373","article-title":"Performance of a large language model on practice questions for the neonatal board examination","volume":"177","author":"Beam","year":"2023","journal-title":"JAMA Pediatr"},{"issue":"8","key":"2024040320031163400_ocad256-B29","doi-asserted-by":"crossref","first-page":"e2331205","DOI":"10.1001\/jamanetworkopen.2023.31205","article-title":"Quality of layperson CPR instructions from artificial intelligence voice assistants","volume":"6","author":"Murk","year":"2023","journal-title":"JAMA Netw Open"},{"key":"2024040320031163400_ocad256-B30","author":"Nori","year":"2023"},{"key":"2024040320031163400_ocad256-B31","author":"Guevara","year":"2023"}],"container-title":["Journal of the American Medical Informatics Association"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jamia\/advance-article-pdf\/doi\/10.1093\/jamia\/ocad256\/57148492\/ocad256.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jamia\/advance-article-pdf\/doi\/10.1093\/jamia\/ocad256\/57148492\/ocad256.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,4,3]],"date-time":"2024-04-03T20:03:40Z","timestamp":1712174620000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jamia\/article\/31\/4\/940\/7585396"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,22]]},"references-count":31,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2024,1,22]]},"published-print":{"date-parts":[[2024,4,3]]}},"URL":"https:\/\/doi.org\/10.1093\/jamia\/ocad256","relation":{},"ISSN":["1067-5027","1527-974X"],"issn-type":[{"value":"1067-5027","type":"print"},{"value":"1527-974X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024,4,1]]},"published":{"date-parts":[[2024,1,22]]}}}