{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,3]],"date-time":"2026-08-03T06:13:39Z","timestamp":1785737619242,"version":"3.56.0"},"reference-count":32,"publisher":"Oxford University Press (OUP)","issue":"9","license":[{"start":{"date-parts":[[2024,5,6]],"date-time":"2024-05-06T00:00:00Z","timestamp":1714953600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100008460","name":"National Center for Complementary and Integrative Health","doi-asserted-by":"publisher","award":["R01AT009457"],"award-info":[{"award-number":["R01AT009457"]}],"id":[{"id":"10.13039\/100008460","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000049","name":"National Institute on Aging","doi-asserted-by":"publisher","award":["R01AG078154"],"award-info":[{"award-number":["R01AG078154"]}],"id":[{"id":"10.13039\/100000049","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000054","name":"National Cancer Institute","doi-asserted-by":"publisher","award":["R01CA287413"],"award-info":[{"award-number":["R01CA287413"]}],"id":[{"id":"10.13039\/100000054","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2024,9,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Objectives<\/jats:title>\n                  <jats:p>This article aims to enhance the performance of larger language models (LLMs) on the few-shot biomedical named entity recognition (NER) task by developing a simple and effective method called Retrieving and Chain-of-Thought (RT) framework and to evaluate the improvement after applying RT framework.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Materials and Methods<\/jats:title>\n                  <jats:p>Given the remarkable advancements in retrieval-based language model and Chain-of-Thought across various natural language processing tasks, we propose a pioneering RT framework designed to amalgamate both approaches. The RT approach encompasses dedicated modules for information retrieval and Chain-of-Thought processes. In the retrieval module, RT discerns pertinent examples from demonstrations during instructional tuning for each input sentence. Subsequently, the Chain-of-Thought module employs a systematic reasoning process to identify entities. We conducted a comprehensive comparative analysis of our RT framework against 16 other models for few-shot NER tasks on BC5CDR and NCBI corpora. Additionally, we explored the impacts of negative samples, output formats, and missing data on performance.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>Our proposed RT framework outperforms other LMs for few-shot NER tasks with micro-F1 scores of 93.50 and 91.76 on BC5CDR and NCBI corpora, respectively. We found that using both positive and negative samples, Chain-of-Thought (vs Tree-of-Thought) performed better. Additionally, utilization of a partially annotated dataset has a marginal effect of the model performance.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Discussion<\/jats:title>\n                  <jats:p>This is the first investigation to combine a retrieval-based LLM and Chain-of-Thought methodology to enhance the performance in biomedical few-shot NER. The retrieval-based LLM aids in retrieving the most relevant examples of the input sentence, offering crucial knowledge to predict the entity in the sentence. We also conducted a meticulous examination of our methodology, incorporating an ablation study.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Conclusion<\/jats:title>\n                  <jats:p>The RT framework with LLM has demonstrated state-of-the-art performance on few-shot NER tasks.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/jamia\/ocae095","type":"journal-article","created":{"date-parts":[[2024,5,6]],"date-time":"2024-05-06T11:01:43Z","timestamp":1714993303000},"page":"1929-1938","source":"Crossref","is-referenced-by-count":53,"title":["RT: a Retrieving and Chain-of-Thought framework for few-shot medical named entity recognition"],"prefix":"10.1093","volume":"31","author":[{"given":"Mingchen","family":"Li","sequence":"first","affiliation":[{"name":"Division of Computational Health Sciences, Department of Surgery, University of Minnesota , Minneapolis, MN 55455, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6524-5506","authenticated-orcid":false,"given":"Huixue","family":"Zhou","sequence":"additional","affiliation":[{"name":"Division of Computational Health Sciences, Department of Surgery, University of Minnesota , Minneapolis, MN 55455, United States"},{"name":"Institute for Health Informatics, University of Minnesota , Minneapolis, MN 55455, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-4322-6753","authenticated-orcid":false,"given":"Han","family":"Yang","sequence":"additional","affiliation":[{"name":"Division of Computational Health Sciences, Department of Surgery, University of Minnesota , Minneapolis, MN 55455, United States"},{"name":"Institute for Health Informatics, University of Minnesota , Minneapolis, MN 55455, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8258-3585","authenticated-orcid":false,"given":"Rui","family":"Zhang","sequence":"additional","affiliation":[{"name":"Division of Computational Health Sciences, Department of Surgery, University of Minnesota , Minneapolis, MN 55455, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2024,5,6]]},"reference":[{"issue":"1","key":"2024082207514685200_ocae095-B1","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1186\/s12859-021-04551-4","article-title":"Hierarchical shared transfer learning for biomedical named entity recognition","volume":"23","author":"Chai","year":"2022","journal-title":"BMC Bioinformatics"},{"key":"2024082207514685200_ocae095-B2","author":"Li","year":"2020"},{"issue":"2","key":"2024082207514685200_ocae095-B3","doi-asserted-by":"crossref","first-page":"201","DOI":"10.26599\/BDMA.2022.9020021","article-title":"Medical knowledge graph: data sources, construction, reasoning, and applications","volume":"6","author":"Wu","year":"2023","journal-title":"Big Data Min Anal"},{"key":"2024082207514685200_ocae095-B4","author":"Li","year":"2022"},{"key":"2024082207514685200_ocae095-B5","first-page":"571","author":"Pugachev","year":"2023"},{"key":"2024082207514685200_ocae095-B6","author":"Li","year":"2022"},{"issue":"1","key":"2024082207514685200_ocae095-B7","doi-asserted-by":"crossref","first-page":"bbac498","DOI":"10.1093\/bib\/bbac498","article-title":"Sprda: a link prediction approach based on the structural perturbation to infer disease-associated piwi-interacting RNAs","volume":"24","author":"Zheng","year":"2023","journal-title":"Brief Bioinform"},{"key":"2024082207514685200_ocae095-B8","author":"Li","year":"2023"},{"key":"2024082207514685200_ocae095-B9","author":"Huang","year":"2019"},{"issue":"4","key":"2024082207514685200_ocae095-B10","doi-asserted-by":"crossref","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","article-title":"Biobert: a pre-trained biomedical language representation model for biomedical text mining","volume":"36","author":"Lee","year":"2020","journal-title":"Bioinformatics"},{"issue":"1","key":"2024082207514685200_ocae095-B11","doi-asserted-by":"crossref","first-page":"194","DOI":"10.1038\/s41746-022-00742-2","article-title":"A large language model for electronic health records","volume":"5","author":"Yang","year":"2022","journal-title":"NPJ Digit Med"},{"key":"2024082207514685200_ocae095-B12","first-page":"2515","author":"Huang","year":"2022"},{"key":"2024082207514685200_ocae095-B13","article-title":"Prototypical networks for few-shot learning","volume":"30","author":"Snell","year":"2017","journal-title":"Adv Neural Inf Process Syst"},{"key":"2024082207514685200_ocae095-B14","author":"Wiseman","year":"2019"},{"key":"2024082207514685200_ocae095-B15","author":"Yang","year":"2020"},{"key":"2024082207514685200_ocae095-B16","author":"Das","year":"2022"},{"key":"2024082207514685200_ocae095-B17","author":"Zhang","year":"2022"},{"key":"2024082207514685200_ocae095-B18","author":"Min","year":"2022"},{"key":"2024082207514685200_ocae095-B19","author":"Li","year":"2023"},{"key":"2024082207514685200_ocae095-B20","author":"Ashok","year":"2023"},{"key":"2024082207514685200_ocae095-B21","author":"Wei","year":"2022"},{"key":"2024082207514685200_ocae095-B22","author":"Wang","year":"2023"},{"key":"2024082207514685200_ocae095-B23","doi-asserted-by":"crossref","first-page":"baw068","DOI":"10.1093\/database\/baw068","article-title":"BioCreative V CDR task corpus: a resource for chemical disease relation extraction","volume":"2016","author":"Li","year":"2016","journal-title":"Database (Oxford)"},{"key":"2024082207514685200_ocae095-B24","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.jbi.2013.12.006","article-title":"Ncbi disease corpus: a resource for disease name recognition and concept normalization","volume":"47","author":"Do\u011fan","year":"2014","journal-title":"J Biomed Inform"},{"key":"2024082207514685200_ocae095-B25","doi-asserted-by":"crossref","first-page":"S20","DOI":"10.1016\/j.jbi.2015.07.020","article-title":"Annotating longitudinal clinical narratives for deidentification: the 2014 i2b2\/uthealth corpus","volume":"58(Suppl)","author":"Stubbs","year":"2015","journal-title":"J Biomed Inform"},{"key":"2024082207514685200_ocae095-B26","author":"Devlin","year":"2018"},{"key":"2024082207514685200_ocae095-B27","author":"Chen","year":"2023"},{"key":"2024082207514685200_ocae095-B28","first-page":"993","author":"Fritzler","year":"2019"},{"key":"2024082207514685200_ocae095-B29","author":"Hou","year":"2020"},{"key":"2024082207514685200_ocae095-B30","author":"Ji","year":"2022"},{"issue":"2","key":"2024082207514685200_ocae095-B31","doi-asserted-by":"crossref","first-page":"426","DOI":"10.1093\/jamia\/ocad216","article-title":"Complementary and integrative health information in the literature: its lexicon and named entity recognition","volume":"31","author":"Zhou","year":"2023","journal-title":"J Am Med Inform Assoc"},{"key":"2024082207514685200_ocae095-B32","author":"Yao","year":"2023"}],"container-title":["Journal of the American Medical Informatics Association"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/31\/9\/1929\/58868150\/ocae095.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/31\/9\/1929\/58868150\/ocae095.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,22]],"date-time":"2024-08-22T11:37:54Z","timestamp":1724326674000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jamia\/article\/31\/9\/1929\/7665312"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,6]]},"references-count":32,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2024,5,6]]},"published-print":{"date-parts":[[2024,9,1]]}},"URL":"https:\/\/doi.org\/10.1093\/jamia\/ocae095","relation":{},"ISSN":["1067-5027","1527-974X"],"issn-type":[{"value":"1067-5027","type":"print"},{"value":"1527-974X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024,9]]},"published":{"date-parts":[[2024,5,6]]}}}