{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,27]],"date-time":"2026-01-27T23:14:41Z","timestamp":1769555681464,"version":"3.49.0"},"reference-count":26,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,6,2]],"date-time":"2021-06-02T00:00:00Z","timestamp":1622592000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,6,2]],"date-time":"2021-06-02T00:00:00Z","timestamp":1622592000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Background<\/jats:title>\n                <jats:p>Biomedical named entity recognition is one of the most essential tasks in biomedical information extraction. Previous studies suffer from inadequate annotated datasets, especially the limited knowledge contained in them.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Methods<\/jats:title>\n                <jats:p>To remedy the above issue, we propose a novel Biomedical Named Entity Recognition (BioNER) framework with label re-correction and knowledge distillation strategies, which could not only create large and high-quality datasets but also obtain a high-performance  recognition model. Our framework is inspired by two points: (1) named entity recognition should be considered from the perspective of both coverage and accuracy; (2) trustable annotations should be yielded by iterative correction. Firstly, for coverage, we annotate chemical and disease entities in a large-scale unlabeled dataset by PubTator to generate a weakly labeled dataset. For accuracy, we then filter it by utilizing multiple knowledge bases to generate another weakly labeled dataset. Next, the two datasets are revised by a label re-correction strategy to construct two high-quality datasets, which are used to train two recognition models, respectively. Finally, we compress the knowledge in the two models into a single recognition model with knowledge distillation.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>Experiments on the BioCreative V chemical-disease relation corpus and NCBI Disease corpus show that knowledge from large-scale datasets significantly improves the performance of BioNER, especially the recall of it, leading to new state-of-the-art results.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusions<\/jats:title>\n                <jats:p>We propose a framework with label re-correction and knowledge distillation strategies. Comparison results show that the two perspectives of knowledge in the two re-corrected datasets respectively are complementary and both effective for BioNER.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1186\/s12859-021-04200-w","type":"journal-article","created":{"date-parts":[[2021,6,2]],"date-time":"2021-06-02T11:07:12Z","timestamp":1622632032000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Improving the recall of biomedical named entity recognition with label re-correction and knowledge distillation"],"prefix":"10.1186","volume":"22","author":[{"given":"Huiwei","family":"Zhou","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhe","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chengkun","family":"Lang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yibin","family":"Xu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yingyu","family":"Lin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junjie","family":"Hou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,6,2]]},"reference":[{"issue":"8","key":"4200_CR1","doi-asserted-by":"publisher","first-page":"1381","DOI":"10.1093\/bioinformatics\/btx761","volume":"34","author":"L Luo","year":"2017","unstructured":"Luo L, Yang Z, Yang P, Zhang Y, Wang L, Lin H, Wang J. An attention-based BiLSTM-CRF approach to document-level chemidcal named entity recognition. Bioinformatics. 2017;34(8):1381\u20138.","journal-title":"Bioinformatics"},{"key":"4200_CR2","unstructured":"Wei CH, Peng Y, Leaman R, Davis AP, Mattingly CJ, Li J, et al. Overview of the BioCreative V chemical disease relation (CDR) task. In Proceedings of the fifth BioCreative challenge evaluation workshop. 2015; 14."},{"key":"4200_CR3","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.jbi.2013.12.006","volume":"47","author":"RI Do\u011fan","year":"2014","unstructured":"Do\u011fan RI, Leaman R, Lu Z. NCBI disease corpus: a resource for disease name recognition and concept normalization. J Biomed Inform. 2014;47:1\u201310.","journal-title":"J Biomed Inform"},{"key":"4200_CR4","doi-asserted-by":"crossref","unstructured":"Ma X, Hovy E. End-to-end sequence labeling via bi-directional lstm-cnns-crf. ACL. 2016.","DOI":"10.18653\/v1\/P16-1101"},{"issue":"14","key":"4200_CR5","doi-asserted-by":"publisher","first-page":"i37","DOI":"10.1093\/bioinformatics\/btx228","volume":"33","author":"M Habibi","year":"2017","unstructured":"Habibi M, Weber L, Neves M, Wiegandt DL, Leser U. Deep learning with word embeddings improves biomedical named entity recognition. Bioinformatics. 2017;33(14):i37\u201348.","journal-title":"Bioinformatics"},{"key":"4200_CR6","doi-asserted-by":"publisher","first-page":"151","DOI":"10.1016\/j.artint.2012.03.006","volume":"194","author":"J Nothman","year":"2013","unstructured":"Nothman J, Ringland N, Radford W, Murphy T, Curran JR. Learning multilingual named entity recognition from Wikipedia. Artif Intell. 2013;194:151\u201375.","journal-title":"Artif Intell"},{"key":"4200_CR7","first-page":"413","volume":"1","author":"A Ghaddar","year":"2017","unstructured":"Ghaddar A, Winer LP. A wikipedia annotated corpus for named entity recognition. IJCNLP. 2017;1:413\u201322.","journal-title":"IJCNLP"},{"key":"4200_CR8","unstructured":"Zhu M, Deng Z, Xiong W, Yu M, Zhang M, Wang WY. Towards open-domain named entity recognition via neural correction models. AAAI. 2020."},{"key":"4200_CR9","unstructured":"Bagherinezhad H, Horton M, Rastegari M, Farhadi A. Label refinery: Improving imagenet classification through label progression. 2018. arXiv preprint aXiv:1805.02641."},{"issue":"6","key":"4200_CR10","doi-asserted-by":"publisher","first-page":"793","DOI":"10.1289\/ehp.6028","volume":"111","author":"CJ Mattingly","year":"2003","unstructured":"Mattingly CJ, Colby GT, Forrest JN, Boyer JL. The comparative toxicogenomics database (CTD). Environ Health Perspect. 2003;111(6):793\u20135.","journal-title":"Environ Health Perspect"},{"issue":"3","key":"4200_CR11","first-page":"265","volume":"88","author":"CE Lipscomb","year":"2000","unstructured":"Lipscomb CE. Medical subject headings (MeSH). Bull Med Libr Assoc. 2000;88(3):265.","journal-title":"Bull Med Libr Assoc"},{"issue":"18","key":"4200_CR12","doi-asserted-by":"publisher","first-page":"809","DOI":"10.1152\/physiolgenomics.00065.2013","volume":"45","author":"R Nigam","year":"2013","unstructured":"Nigam R, Laulederkind SJF, Hayman GT, Smith JR, Wang SJ, et al. Rat genome database: a unique resource for rat, human, and mouse quantitative trait locus data. Physiol Genomics. 2013;45(18):809\u201316.","journal-title":"Physiol Genomics"},{"key":"4200_CR13","doi-asserted-by":"crossref","unstructured":"Wei CH, Lee K, Leaman R, Lu Z. Biomedical mention disambiguation using a deep learning approach. ACM. 2019; 307\u2013313.","DOI":"10.1145\/3307339.3342162"},{"issue":"W1","key":"4200_CR14","doi-asserted-by":"publisher","first-page":"W518","DOI":"10.1093\/nar\/gkt441","volume":"41","author":"CH Wei","year":"2013","unstructured":"Wei CH, Kao HY, Lu Z. PubTator: a web-based text mining tool for assisting biocuration. Nucleic Acids Res. 2013;41(W1):W518\u201322.","journal-title":"Nucleic Acids Res"},{"key":"4200_CR15","unstructured":"Hinton G, Vinyals O, Dean J. Distilling the knowledge in a neural network. NIPS. 2015."},{"key":"4200_CR16","doi-asserted-by":"crossref","unstructured":"Li Y, Yang J, Song Y, Cao L, Luo J, Li LJ. Learning from noisy labels with distillation. ICCV. 2017; 1910\u20131918.","DOI":"10.1109\/ICCV.2017.211"},{"key":"4200_CR17","doi-asserted-by":"publisher","first-page":"4886","DOI":"10.1609\/aaai.v33i01.33014886","volume":"33","author":"Z Shen","year":"2019","unstructured":"Shen Z, He Z, Xue X. Meal: multi-model ensemble via adversarial learning. AAAI. 2019;33:4886\u201393.","journal-title":"AAAI"},{"issue":"20","key":"4200_CR18","doi-asserted-by":"publisher","first-page":"3539","DOI":"10.1093\/bioinformatics\/bty356","volume":"34","author":"TH Dang","year":"2018","unstructured":"Dang TH, Le HQ, Nguyen TM, Vu ST. D3NER: biomedical named entity recognition using CRF-biLSTM improved with fine-tuned embeddings of various linguistic information. Bioinformatics. 2018;34(20):3539\u201346.","journal-title":"Bioinformatics"},{"key":"4200_CR19","doi-asserted-by":"crossref","unstructured":"Wang J, Xu W, Fu X, Xu G, Wu Y. ASTRAL: adversarial trained LSTM-CNN for named entity recognition. knowledge-based system. 2020; 197.","DOI":"10.1016\/j.knosys.2020.105842"},{"issue":"18","key":"4200_CR20","doi-asserted-by":"publisher","first-page":"2839","DOI":"10.1093\/bioinformatics\/btw343","volume":"32","author":"R Leaman","year":"2016","unstructured":"Leaman R, Lu Z. TaggerOne: joint named entity recognition and normal-ization with semi-Markov Models. Bioinformatics. 2016;32(18):2839\u201346.","journal-title":"Bioinformatics"},{"issue":"10","key":"4200_CR21","doi-asserted-by":"publisher","first-page":"1745","DOI":"10.1093\/bioinformatics\/bty869","volume":"35","author":"X Wang","year":"2019","unstructured":"Wang X, Zhang Y, Ren X, Zhang Y, Zitnik M, Shang J, et al. Cross-type biomedical named entity recognition with deep multi-task learning. Bioinformatics. 2019;35(10):1745\u201352.","journal-title":"Bioinformatics"},{"issue":"10","key":"4200_CR22","doi-asserted-by":"publisher","first-page":"249","DOI":"10.1186\/s12859-019-2813-6","volume":"20","author":"W Yoon","year":"2019","unstructured":"Yoon W, So CH, Lee J, Kang J. CollaboNet: collaboration of deep neural networks for biomedical named entity recognition. BMC Bioinformatics. 2019;20(10):249.","journal-title":"BMC Bioinformatics"},{"key":"4200_CR23","unstructured":"Devlin J, Chang MW, Lee K, Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. NAACL-HLT. 2019."},{"key":"4200_CR24","doi-asserted-by":"crossref","unstructured":"Lee J, Yoon W, Kim S, Kim D, Kim S, So CH, Kang J. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics. 2020; 1\u20137.","DOI":"10.1093\/bioinformatics\/btz682"},{"issue":"2","key":"4200_CR25","doi-asserted-by":"publisher","first-page":"260","DOI":"10.1109\/TIT.1967.1054010","volume":"13","author":"A Viterbi","year":"1967","unstructured":"Viterbi A. Error bounds for convolutional codes and an asymptotically optimum decoding algorithm. IEEE Trans Inf Theory. 1967;13(2):260\u20139.","journal-title":"IEEE Trans Inf Theory"},{"key":"4200_CR26","unstructured":"Mikolov T, Sutskever I, Chen K, Corrado GS, Dean J. Distributed representations of words and phrases and their compositionality. NIPS. 2013."}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-021-04200-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12859-021-04200-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-021-04200-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,6,2]],"date-time":"2021-06-02T11:08:37Z","timestamp":1622632117000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-021-04200-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,6,2]]},"references-count":26,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["4200"],"URL":"https:\/\/doi.org\/10.1186\/s12859-021-04200-w","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,6,2]]},"assertion":[{"value":"10 December 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 May 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 June 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"295"}}