{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,17]],"date-time":"2026-08-17T18:45:27Z","timestamp":1786992327713,"version":"build-2736575974"},"reference-count":47,"publisher":"Oxford University Press (OUP)","issue":"5","funder":[{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Central Universities"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2023,5,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>Biomedical named entity recognition (BioNER) seeks to automatically recognize biomedical entities in natural language text, serving as a necessary foundation for downstream text mining tasks and applications such as information extraction and question answering. Manually labeling training data for the BioNER task is costly, however, due to the significant domain expertise required for accurate annotation. The resulting data scarcity causes current BioNER approaches to be prone to overfitting, to suffer from limited generalizability, and to address a single entity type at a time (e.g. gene or disease).<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>We therefore propose a novel all-in-one (AIO) scheme that uses external data from existing annotated resources to enhance the accuracy and stability of BioNER models. We further present AIONER, a general-purpose BioNER tool based on cutting-edge deep learning and our AIO schema. We evaluate AIONER on 14 BioNER benchmark tasks and show that AIONER is effective, robust, and compares favorably to other state-of-the-art approaches such as multi-task learning. We further demonstrate the practical utility of AIONER in three independent tasks to recognize entity types not previously seen in training data, as well as the advantages of AIONER over existing methods for processing biomedical text at a large scale (e.g. the entire PubMed data).<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation<\/jats:title>\n                  <jats:p>The source code, trained models and data for AIONER are freely available at https:\/\/github.com\/ncbi\/AIONER.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btad310","type":"journal-article","created":{"date-parts":[[2023,5,12]],"date-time":"2023-05-12T15:58:45Z","timestamp":1683907125000},"source":"Crossref","is-referenced-by-count":52,"title":["AIONER: all-in-one scheme-based biomedical named entity recognition using deep learning"],"prefix":"10.1093","volume":"39","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5141-0259","authenticated-orcid":false,"given":"Ling","family":"Luo","sequence":"first","affiliation":[{"name":"National Center for Biotechnology Information (NCBI), National Library of Medicine (NLM), National Institutes of Health (NIH) , Bethesda, MD 20894, United States"},{"name":"School of Computer Science and Technology, Dalian University of Technology , Dalian 116024, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5094-7321","authenticated-orcid":false,"given":"Chih-Hsuan","family":"Wei","sequence":"additional","affiliation":[{"name":"National Center for Biotechnology Information (NCBI), National Library of Medicine (NLM), National Institutes of Health (NIH) , Bethesda, MD 20894, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Po-Ting","family":"Lai","sequence":"additional","affiliation":[{"name":"National Center for Biotechnology Information (NCBI), National Library of Medicine (NLM), National Institutes of Health (NIH) , Bethesda, MD 20894, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3296-5766","authenticated-orcid":false,"given":"Robert","family":"Leaman","sequence":"additional","affiliation":[{"name":"National Center for Biotechnology Information (NCBI), National Library of Medicine (NLM), National Institutes of Health (NIH) , Bethesda, MD 20894, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6036-1516","authenticated-orcid":false,"given":"Qingyu","family":"Chen","sequence":"additional","affiliation":[{"name":"National Center for Biotechnology Information (NCBI), National Library of Medicine (NLM), National Institutes of Health (NIH) , Bethesda, MD 20894, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9998-916X","authenticated-orcid":false,"given":"Zhiyong","family":"Lu","sequence":"additional","affiliation":[{"name":"National Center for Biotechnology Information (NCBI), National Library of Medicine (NLM), National Institutes of Health (NIH) , Bethesda, MD 20894, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2023,5,12]]},"reference":[{"key":"2023052518473507000_btad310-B1","first-page":"28","author":"Arighi","year":"2017"},{"key":"2023052518473507000_btad310-B2","first-page":"76","author":"Cariello","year":"2021"},{"key":"2023052518473507000_btad310-B3","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1023\/A:1007379606734","article-title":"Multitask learning","volume":"28","author":"Caruana","year":"1997","journal-title":"Mach Learn"},{"key":"2023052518473507000_btad310-B4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s12859-021-04551-4","article-title":"Hierarchical shared transfer learning for biomedical named entity recognition","volume":"23","author":"Chai","year":"2022","journal-title":"BMC Bioinformatics"},{"key":"2023052518473507000_btad310-B5","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s12859-017-1857-8","article-title":"A method for named entity normalization in biomedical articles: application to diseases and plants","volume":"18","author":"Cho","year":"2017","journal-title":"BMC Bioinformatics"},{"key":"2023052518473507000_btad310-B6","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s12859-017-1776-8","article-title":"A neural network multi-task learning approach to biomedical named entity recognition","volume":"18","author":"Crichton","year":"2017","journal-title":"BMC Bioinformatics"},{"key":"2023052518473507000_btad310-B7","first-page":"4171","author":"Devlin","year":"2019"},{"key":"2023052518473507000_btad310-B8","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.jbi.2013.12.006","article-title":"NCBI disease corpus: a resource for disease name recognition and concept normalization","volume":"47","author":"Do\u011fan","year":"2014","journal-title":"J Biomed Inform"},{"key":"2023052518473507000_btad310-B9","first-page":"272","author":"Fang","year":"2021"},{"key":"2023052518473507000_btad310-B10","doi-asserted-by":"crossref","first-page":"2474","DOI":"10.1093\/bioinformatics\/bty152","article-title":"Exploiting and assessing multi-source data for supervised biomedical named entity recognition","volume":"34","author":"Galea","year":"2018","journal-title":"Bioinformatics"},{"key":"2023052518473507000_btad310-B11","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/1471-2105-11-85","article-title":"LINNAEUS: a species name identification system for biomedical literature","volume":"11","author":"Gerner","year":"2010","journal-title":"BMC Bioinformatics"},{"key":"2023052518473507000_btad310-B12","doi-asserted-by":"crossref","first-page":"280","DOI":"10.1093\/bioinformatics\/btz504","article-title":"Towards reliable named entity recognition in the biomedical domain","volume":"36","author":"Giorgi","year":"2020","journal-title":"Bioinformatics"},{"key":"2023052518473507000_btad310-B13","first-page":"315","author":"Glorot","year":"2011"},{"key":"2023052518473507000_btad310-B14","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3458754","article-title":"Domain-specific language model pretraining for biomedical natural language processing","volume":"3","author":"Gu","year":"2022","journal-title":"ACM Trans Comput Healthcare"},{"key":"2023052518473507000_btad310-B15","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1038\/s41597-021-00875-1","article-title":"NLM-Chem, a new resource for chemical entity recognition in PubMed full text literature","volume":"8","author":"Islamaj","year":"2021","journal-title":"Sci Data"},{"key":"2023052518473507000_btad310-B16","doi-asserted-by":"crossref","first-page":"103779","DOI":"10.1016\/j.jbi.2021.103779","article-title":"NLM-Gene, a richly annotated gold standard dataset for gene entities that addresses ambiguity and multi-species gene recognition","volume":"118","author":"Islamaj","year":"2021","journal-title":"J Biomed Inform"},{"key":"2023052518473507000_btad310-B17","author":"Jeong","year":"2021"},{"key":"2023052518473507000_btad310-B18","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/1758-2946-7-S1-S1","article-title":"The CHEMDNER corpus of chemicals and drugs and its annotation principles","volume":"7","author":"Krallinger","year":"2015","journal-title":"J Cheminform"},{"key":"2023052518473507000_btad310-B19","first-page":"282","author":"Lafferty","year":"2001"},{"key":"2023052518473507000_btad310-B20","first-page":"260","author":"Lample","year":"2016"},{"key":"2023052518473507000_btad310-B21","doi-asserted-by":"crossref","first-page":"26","DOI":"10.1186\/s13326-022-00280-6","article-title":"We are not ready yet: limitations of state-of-the-art disease named entity recognizers","volume":"13","author":"K\u00fchnel","year":"2022","journal-title":"J Biomed Semant"},{"key":"2023052518473507000_btad310-B22","doi-asserted-by":"crossref","first-page":"2909","DOI":"10.1093\/bioinformatics\/btt474","article-title":"DNorm: disease name normalization with pairwise learning to rank","volume":"29","author":"Leaman","year":"2013","journal-title":"Bioinformatics"},{"key":"2023052518473507000_btad310-B23","doi-asserted-by":"crossref","first-page":"S3","DOI":"10.1186\/1758-2946-7-S1-S3","article-title":"tmChem: a high performance approach for chemical named entity recognition and normalization","volume":"7","author":"Leaman","year":"2015","journal-title":"J Cheminform"},{"key":"2023052518473507000_btad310-B24","doi-asserted-by":"crossref","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","article-title":"BioBERT: a pre-trained biomedical language representation model for biomedical text mining","volume":"36","author":"Lee","year":"2020","journal-title":"Bioinformatics"},{"key":"2023052518473507000_btad310-B25","doi-asserted-by":"crossref","first-page":"baw068","DOI":"10.1093\/database\/baw068","article-title":"BioCreative V CDR task corpus: a resource for chemical disease relation extraction","volume":"2016","author":"Li","year":"2016","journal-title":"Database"},{"key":"2023052518473507000_btad310-B26","doi-asserted-by":"crossref","first-page":"bbac282","DOI":"10.1093\/bib\/bbac282","article-title":"BioRED: a rich biomedical relation extraction dataset","volume":"23","author":"Luo","year":"2022","journal-title":"Brief Bioinf"},{"key":"2023052518473507000_btad310-B27","doi-asserted-by":"crossref","first-page":"baac090","DOI":"10.1093\/database\/baac090","article-title":"Assigning species information to corresponding genes by a sequence labeling framework","volume":"2022","author":"Luo","year":"2022","journal-title":"Database"},{"key":"2023052518473507000_btad310-B28","doi-asserted-by":"crossref","first-page":"e65390","DOI":"10.1371\/journal.pone.0065390","article-title":"The SPECIES and ORGANISMS resources for fast and accurate identification of taxonomic names in text","volume":"8","author":"Pafilis","year":"2013","journal-title":"PLoS ONE"},{"key":"2023052518473507000_btad310-B29","first-page":"58","author":"Peng","year":"2019"},{"key":"2023052518473507000_btad310-B30","first-page":"2227","author":"Peters","year":"2018"},{"key":"2023052518473507000_btad310-B31","doi-asserted-by":"crossref","first-page":"868","DOI":"10.1093\/bioinformatics\/btt580","article-title":"Anatomical entity mention recognition at literature scale","volume":"30","author":"Pyysalo","year":"2014","journal-title":"Bioinformatics"},{"key":"2023052518473507000_btad310-B32","doi-asserted-by":"crossref","first-page":"104062","DOI":"10.1016\/j.jbi.2022.104062","article-title":"Effects of data and entity ablation on multitask learning models for biomedical entity recognition","volume":"130","author":"Rodriguez","year":"2022","journal-title":"J Biomed Inf"},{"key":"2023052518473507000_btad310-B33","first-page":"142","author":"Sang","year":"2003"},{"key":"2023052518473507000_btad310-B34","doi-asserted-by":"crossref","first-page":"D10","DOI":"10.1093\/nar\/gkaa892","article-title":"Database resources of the national center for biotechnology information","volume":"49","author":"Sayers","year":"2021","journal-title":"Nucleic Acids Res"},{"key":"2023052518473507000_btad310-B35","doi-asserted-by":"crossref","first-page":"e1005017","DOI":"10.1371\/journal.pcbi.1005017","article-title":"Text mining genotype-phenotype relationships from biomedical literature for database curation and precision medicine","volume":"12","author":"Singhal","year":"2016","journal-title":"PLoS Comput Biol"},{"key":"2023052518473507000_btad310-B36","doi-asserted-by":"crossref","first-page":"4837","DOI":"10.1093\/bioinformatics\/btac598","article-title":"BERN2: an advanced neural biomedical named entity recognition and normalization tool","volume":"38","author":"Sung","year":"2022","journal-title":"Bioinformatics"},{"key":"2023052518473507000_btad310-B37","doi-asserted-by":"crossref","first-page":"3976","DOI":"10.1093\/bioinformatics\/btac422","article-title":"Improving biomedical named entity recognition by dynamic caching inter-sentence information","volume":"38","author":"Tong","year":"2022","journal-title":"Bioinformatics"},{"key":"2023052518473507000_btad310-B38","doi-asserted-by":"crossref","first-page":"260","DOI":"10.1109\/TIT.1967.1054010","article-title":"Error bounds for convolutional codes and an asymptotically optimum decoding algorithm","volume":"13","author":"Viterbi","year":"1967","journal-title":"IEEE Trans Inf Theory"},{"key":"2023052518473507000_btad310-B39","doi-asserted-by":"crossref","first-page":"1745","DOI":"10.1093\/bioinformatics\/bty869","article-title":"Cross-type biomedical named entity recognition with deep multi-task learning","volume":"35","author":"Wang","year":"2019","journal-title":"Bioinformatics"},{"key":"2023052518473507000_btad310-B40","doi-asserted-by":"crossref","first-page":"548","DOI":"10.1002\/asi.1104","article-title":"Using concepts in literature-based discovery: simulating Swanson's Raynaud\u2013fish oil and migraine\u2013magnesium discoveries","volume":"52","author":"Weeber","year":"2001","journal-title":"J Am Soc Inf Sci"},{"key":"2023052518473507000_btad310-B41","doi-asserted-by":"crossref","first-page":"W587","DOI":"10.1093\/nar\/gkz389","article-title":"PubTator Central: automated concept annotation for biomedical full text articles","volume":"47","author":"Wei","year":"2019","journal-title":"Nucleic Acids Res"},{"key":"2023052518473507000_btad310-B42","doi-asserted-by":"crossref","first-page":"4449","DOI":"10.1093\/bioinformatics\/btac537","article-title":"tmVar 3.0: an improved variant concept recognition and normalization tool","volume":"38","author":"Wei","year":"2022","journal-title":"Bioinformatics"},{"key":"2023052518473507000_btad310-B43","first-page":"1","article-title":"GNormPlus: an integrative approach for tagging genes, gene families, and protein domains","volume":"2015","author":"Wei","year":"2015","journal-title":"BioMed Res Int"},{"key":"2023052518473507000_btad310-B44","first-page":"4439","author":"W\u00fchrl","year":"2022"},{"key":"2023052518473507000_btad310-B45","doi-asserted-by":"crossref","first-page":"5586","DOI":"10.1109\/TKDE.2021.3070203","article-title":"A survey on multi-task learning","volume":"34","author":"Zhang","year":"2022","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"2023052518473507000_btad310-B46","doi-asserted-by":"crossref","first-page":"1892","DOI":"10.1093\/jamia\/ocab090","article-title":"Biomedical and clinical English model packages for the Stanza Python NLP library","volume":"28","author":"Zhang","year":"2021","journal-title":"J Am Med Inf Assoc"},{"key":"2023052518473507000_btad310-B47","doi-asserted-by":"crossref","first-page":"4331","DOI":"10.1093\/bioinformatics\/btaa515","article-title":"Dataset-aware multi-task learning approaches for biomedical named entity recognition","volume":"36","author":"Zuo","year":"2020","journal-title":"Bioinformatics"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btad310\/50297092\/btad310.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/39\/5\/btad310\/50453589\/btad310.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/39\/5\/btad310\/50453589\/btad310.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,25]],"date-time":"2023-05-25T19:40:54Z","timestamp":1685043654000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/doi\/10.1093\/bioinformatics\/btad310\/7160912"}},"subtitle":[],"editor":[{"given":"Jonathan","family":"Wren","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2023,5,1]]},"references-count":47,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2023,5,4]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btad310","relation":{},"ISSN":["1367-4811"],"issn-type":[{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023,5,1]]},"published":{"date-parts":[[2023,5,1]]},"article-number":"btad310"}}