{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T15:04:51Z","timestamp":1782313491366,"version":"3.54.5"},"reference-count":36,"publisher":"Oxford University Press (OUP)","issue":"9","license":[{"start":{"date-parts":[[2024,5,24]],"date-time":"2024-05-24T00:00:00Z","timestamp":1716508800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"funder":[{"DOI":"10.13039\/100006108","name":"National Center for Advancing Translational Sciences","doi-asserted-by":"publisher","award":["OT2TR003434"],"award-info":[{"award-number":["OT2TR003434"]}],"id":[{"id":"10.13039\/100006108","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000092","name":"National Library of Medicine","doi-asserted-by":"publisher","award":["1R01LM014344"],"award-info":[{"award-number":["1R01LM014344"]}],"id":[{"id":"10.13039\/100000092","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000092","name":"National Library of Medicine","doi-asserted-by":"publisher","award":["R01LM009886"],"award-info":[{"award-number":["R01LM009886"]}],"id":[{"id":"10.13039\/100000092","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2024,9,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Objectives<\/jats:title>\n                  <jats:p>To automatically construct a drug indication taxonomy from drug labels using generative Artificial Intelligence (AI) represented by the Large Language Model (LLM) GPT-4 and real-world evidence (RWE).<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Materials and Methods<\/jats:title>\n                  <jats:p>We extracted indication terms from 46\u00a0421 free-text drug labels using GPT-4, iteratively and recursively generated indication concepts and inferred indication concept-to-concept and concept-to-term subsumption relations by integrating GPT-4 with RWE, and created a drug indication taxonomy. Quantitative and qualitative evaluations involving domain experts were performed for cardiovascular (CVD), Endocrine, and Genitourinary system diseases.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>2909 drug indication terms were extracted and assigned into 24 high-level indication categories (ie, initially generated concepts), each of which was expanded into a sub-taxonomy. For example, the CVD sub-taxonomy contains 242 concepts, spanning a depth of 11, with 170 being leaf nodes. It collectively covers a total of 234 indication terms associated with 189 distinct drugs. The accuracies of GPT-4 on determining the drug indication hierarchy exceeded 0.7 with \u201cgood to very good\u201d inter-rater reliability. However, the accuracies of the concept-to-term subsumption relation checking varied greatly, with \u201cfair to moderate\u201d reliability.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Discussion and Conclusion<\/jats:title>\n                  <jats:p>We successfully used generative AI and RWE to create a taxonomy, with drug indications adequately consistent with domain expert expectations. We show that LLMs are good at deriving their own concept hierarchies but still fall short in determining the subsumption relations between concepts and terms in unregulated language from free-text drug labels, which is the same hard task for human experts.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/jamia\/ocae105","type":"journal-article","created":{"date-parts":[[2024,5,24]],"date-time":"2024-05-24T18:10:12Z","timestamp":1716574212000},"page":"2065-2075","source":"Crossref","is-referenced-by-count":6,"title":["Knowledge-guided generative artificial intelligence for automated taxonomy learning from drug labels"],"prefix":"10.1093","volume":"31","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2681-1931","authenticated-orcid":false,"given":"Yilu","family":"Fang","sequence":"first","affiliation":[{"name":"Department of Biomedical Informatics, Columbia University , New York, NY 10032, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Patrick","family":"Ryan","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Columbia University , New York, NY 10032, United States"},{"name":"Observational Health Data Analytics, Janssen Research and Development , Titusville, NJ 08560, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chunhua","family":"Weng","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Columbia University , New York, NY 10032, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2024,5,24]]},"reference":[{"key":"2024082207520416200_ocae105-B1","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1016\/j.jbi.2017.11.011","article-title":"Clinical information extraction applications: A literature review","volume":"77","author":"Wang","year":"2018","journal-title":"J Biomed Inform"},{"issue":"1","key":"2024082207520416200_ocae105-B2","doi-asserted-by":"crossref","first-page":"208","DOI":"10.1055\/s-0040-1702001","article-title":"Medical information extraction in the age of deep learning","volume":"29","author":"Hahn","year":"2020","journal-title":"Yearb Med Inform"},{"key":"2024082207520416200_ocae105-B3"},{"key":"2024082207520416200_ocae105-B4","doi-asserted-by":"crossref","first-page":"711467","DOI":"10.3389\/frai.2021.711467","article-title":"DICE: A drug indication classification and encyclopedia for AI-based indication extraction","volume":"4","author":"Bhatt","year":"2021","journal-title":"Front Artif Intell"},{"issue":"3","key":"2024082207520416200_ocae105-B5","doi-asserted-by":"crossref","first-page":"482","DOI":"10.1136\/amiajnl-2012-001291","article-title":"Extracting drug indication information from structured product labels using natural language processing","volume":"20","author":"Fung","year":"2013","journal-title":"J Am Med Inform Assoc"},{"key":"2024082207520416200_ocae105-B6","first-page":"787","author":"Khare","year":"2014"},{"issue":"D1","key":"2024082207520416200_ocae105-B7","doi-asserted-by":"crossref","first-page":"D932","DOI":"10.1093\/nar\/gkw993","article-title":"DrugCentral: online drug compendium","volume":"45","author":"Ursu","year":"2017","journal-title":"Nucleic Acids Res"},{"key":"2024082207520416200_ocae105-B8","doi-asserted-by":"crossref","first-page":"670006","DOI":"10.3389\/frma.2021.670006","article-title":"Information extraction from FDA drug Labeling to enhance product-specific guidance assessment using natural language processing","volume":"6","author":"Shi","year":"2021","journal-title":"Front Res Metr Anal"},{"key":"2024082207520416200_ocae105-B9","first-page":"17","author":"Aronson","year":"2001"},{"key":"2024082207520416200_ocae105-B10"},{"key":"2024082207520416200_ocae105-B11","doi-asserted-by":"crossref","first-page":"295","DOI":"10.1016\/j.jbi.2016.09.002","article-title":"Automated learning of domain taxonomies from text using background knowledge","volume":"63","author":"Hoxha","year":"2016","journal-title":"J Biomed Inform"},{"issue":"7972","key":"2024082207520416200_ocae105-B12","doi-asserted-by":"crossref","first-page":"172","DOI":"10.1038\/s41586-023-06291-2","article-title":"Large language models encode clinical knowledge","volume":"620","author":"Singhal","year":"2023","journal-title":"Nature"},{"key":"2024082207520416200_ocae105-B13","first-page":"1998","author":"Agrawal","year":"2022"},{"issue":"1","key":"2024082207520416200_ocae105-B14","doi-asserted-by":"crossref","first-page":"980","DOI":"10.1002\/pra2.918","article-title":"A generative drug\u2013drug interaction triplets extraction framework based on large language models","volume":"60","author":"Hu","year":"2023","journal-title":"Proc Assoc Inf Sci Technol."},{"key":"2024082207520416200_ocae105-B15","first-page":"396","author":"Kartchner","year":"2023"},{"key":"2024082207520416200_ocae105-B16","author":"Wang"},{"key":"2024082207520416200_ocae105-B17","author":"Cohen","year":". 2023."},{"key":"2024082207520416200_ocae105-B18","author":"Funk","year":"2023"},{"key":"2024082207520416200_ocae105-B19","author":"OpenAI","year":"2023"},{"key":"2024082207520416200_ocae105-B20","author":"Bohn","year":"2023"},{"key":"2024082207520416200_ocae105-B21","volume-title":"Handbook of Inter-Rater Reliability: The Definitive Guide to Measuring the Extent of Agreement among Raters","author":"Gwet","year":"2014"},{"key":"2024082207520416200_ocae105-B22","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1186\/1471-2288-13-61","article-title":"A comparison of Cohen's Kappa and Gwet's AC1 when calculating inter-rater reliability coefficients: a study conducted with personality disorder samples","volume":"13","author":"Wongpakaran","year":"2013","journal-title":"BMC Med Res Methodol"},{"key":"2024082207520416200_ocae105-B23","author":"Zhang","year":". 2023."},{"key":"2024082207520416200_ocae105-B24","author":"Manakul"},{"key":"2024082207520416200_ocae105-B25","author":"Noy"},{"issue":"04\/05","key":"2024082207520416200_ocae105-B26","doi-asserted-by":"crossref","first-page":"394","DOI":"10.1055\/s-0038-1634558","article-title":"Desiderata for controlled medical vocabularies in the twenty-first century","volume":"37","author":"Cimino","year":"1998","journal-title":"Methods Inf Med"},{"issue":"12","key":"2024082207520416200_ocae105-B27","doi-asserted-by":"crossref","first-page":"i292","DOI":"10.1093\/bioinformatics\/bts215","article-title":"Extending ontologies by finding siblings using set expansion techniques","volume":"28","author":"Fabian","year":"2012","journal-title":"Bioinformatics"},{"issue":"1","key":"2024082207520416200_ocae105-B28","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s13326-019-0218-0","article-title":"Combining lexical and context features for automatic ontology extension","volume":"11","author":"Althubaiti","year":"2020","journal-title":"J Biomed Semantics"},{"issue":"5","key":"2024082207520416200_ocae105-B29","doi-asserted-by":"crossref","first-page":"635","DOI":"10.1016\/j.cct.2008.02.004","article-title":"Heterogeneous but \u201cstandard\u201d coding systems for adverse events: Issues in achieving interoperability between apples and oranges","volume":"29","author":"Richesson","year":"2008","journal-title":"Contemporary Clinical Trials"},{"key":"2024082207520416200_ocae105-B30","author":"Touvron","year":"2023"},{"key":"2024082207520416200_ocae105-B31","author":"Anil","year":"2023"},{"key":"2024082207520416200_ocae105-B32","author":"Wu","year":"2023"},{"key":"2024082207520416200_ocae105-B33","author":"Singhal","year":"2023"},{"key":"2024082207520416200_ocae105-B34","doi-asserted-by":"crossref","first-page":"108189","DOI":"10.1016\/j.compbiomed.2024.108189","article-title":"A comprehensive evaluation of large Language models on benchmark biomedical text processing tasks","volume":"171","author":"Jahan","year":"2024","journal-title":"Comput Biol Med"},{"key":"2024082207520416200_ocae105-B35","doi-asserted-by":"crossref","first-page":"e55318","DOI":"10.2196\/55318","article-title":"An empirical evaluation of prompting strategies for large language models in zero-shot clinical natural language processing: algorithm development and validation study","volume":"12","author":"Sivarajkumar","year":"2024","journal-title":"JMIR Med Inform"},{"issue":"Database issue","key":"2024082207520416200_ocae105-B36","doi-asserted-by":"crossref","first-page":"D940","DOI":"10.1093\/nar\/gkr972","article-title":"Disease Ontology: a backbone for disease semantic integration","volume":"40","author":"Schriml","year":"2012","journal-title":"Nucleic Acids Res"}],"container-title":["Journal of the American Medical Informatics Association"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/31\/9\/2065\/58868226\/ocae105.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/31\/9\/2065\/58868226\/ocae105.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,22]],"date-time":"2024-08-22T12:04:36Z","timestamp":1724328276000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jamia\/article\/31\/9\/2065\/7681751"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,24]]},"references-count":36,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2024,5,24]]},"published-print":{"date-parts":[[2024,9,1]]}},"URL":"https:\/\/doi.org\/10.1093\/jamia\/ocae105","relation":{},"ISSN":["1067-5027","1527-974X"],"issn-type":[{"value":"1067-5027","type":"print"},{"value":"1527-974X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024,9]]},"published":{"date-parts":[[2024,5,24]]}}}