{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,1]],"date-time":"2026-06-01T14:22:02Z","timestamp":1780323722721,"version":"3.54.1"},"reference-count":56,"publisher":"Oxford University Press (OUP)","issue":"2","license":[{"start":{"date-parts":[[2021,9,23]],"date-time":"2021-09-23T00:00:00Z","timestamp":1632355200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"funder":[{"DOI":"10.13039\/100000001","name":"US National Science Foundation","doi-asserted-by":"crossref","award":["IIS-2008208"],"award-info":[{"award-number":["IIS-2008208"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100000001","name":"US National Science Foundation","doi-asserted-by":"crossref","award":["1934600"],"award-info":[{"award-number":["1934600"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100000001","name":"US National Science Foundation","doi-asserted-by":"crossref","award":["1938167"],"award-info":[{"award-number":["1938167"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100000001","name":"US National Science Foundation","doi-asserted-by":"crossref","award":["1955151"],"award-info":[{"award-number":["1955151"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,1,3]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>Biomedical language models produce meaningful concept representations that are useful for a variety of biomedical natural language processing (bioNLP) applications such as named entity recognition, relationship extraction and question answering. Recent research trends have shown that the contextualized language models (e.g. BioBERT, BioELMo) possess tremendous representational power and are able to achieve impressive accuracy gains. However, these models are still unable to learn high-quality representations for concepts with low context information (i.e. rare words). Infusing the complementary information from knowledge-bases (KBs) is likely to be helpful when the corpus-specific information is insufficient to learn robust representations. Moreover, as the biomedical domain contains numerous KBs, it is imperative to develop approaches that can integrate the KBs in a continual fashion.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>We propose a new representation learning approach that progressively fuses the semantic information from multiple KBs into the pretrained biomedical language models. Since most of the KBs in the biomedical domain are expressed as parent-child hierarchies, we choose to model the hierarchical KBs and propose a new knowledge modeling strategy that encodes their topological properties at a granular level. Moreover, the proposed continual learning technique efficiently updates the concepts representations to accommodate the new knowledge while preserving the memory efficiency of contextualized language models. Altogether, the proposed approach generates knowledge-powered embeddings with high fidelity and learning efficiency. Extensive experiments conducted on bioNLP tasks validate the efficacy of the proposed approach and demonstrates its capability in generating robust concept representations.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btab671","type":"journal-article","created":{"date-parts":[[2021,9,20]],"date-time":"2021-09-20T19:21:27Z","timestamp":1632165687000},"page":"494-502","source":"Crossref","is-referenced-by-count":13,"title":["Continual knowledge infusion into pre-trained biomedical language models"],"prefix":"10.1093","volume":"38","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0826-445X","authenticated-orcid":false,"given":"Kishlay","family":"Jha","sequence":"first","affiliation":[{"name":"Department of Computer Science, University of Virginia , Charlottesville, VA 22903, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Aidong","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Virginia , Charlottesville, VA 22903, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2021,9,23]]},"reference":[{"key":"2023020108470591600_btab671-B1","first-page":"3606","author":"Beltagy","year":"2019"},{"key":"2023020108470591600_btab671-B2","first-page":"6523","author":"Biesialska","year":"2020"},{"key":"2023020108470591600_btab671-B3","first-page":"69","author":"Bird","year":"2006"},{"key":"2023020108470591600_btab671-B4","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1186\/s12859-015-0472-9","article-title":"Extraction of relations between genes and diseases from text and large-scale data analysis: implications for translational research","volume":"16","author":"Bravo","year":"2015","journal-title":"BMC Bioinformatics"},{"key":"2023020108470591600_btab671-B5","doi-asserted-by":"crossref","first-page":"e12402","DOI":"10.1111\/lnc3.12402","article-title":"Word embeddings for biomedical natural language processing: a survey","volume":"14","author":"Chiu","year":"2020","journal-title":"Lang. Linguist. Compass"},{"key":"2023020108470591600_btab671-B6","first-page":"166","author":"Chiu","year":"2016"},{"key":"2023020108470591600_btab671-B7","doi-asserted-by":"crossref","first-page":"317","DOI":"10.1136\/jamia.2001.0080317","article-title":"Medical subject headings used to search the biomedical literature","volume":"8","author":"Coletti","year":"2001","journal-title":"J. Am. Med. Inf. Assoc"},{"key":"2023020108470591600_btab671-B8","first-page":"73","author":"Collier","year":"2004"},{"key":"2023020108470591600_btab671-B9","doi-asserted-by":"crossref","first-page":"S2","DOI":"10.1186\/1472-6947-8-S1-S2","article-title":"Forty years of snomed: a literature review","volume":"8","author":"Cornet","year":"2008","journal-title":"BMC Med. Inf. Decision Mak"},{"key":"2023020108470591600_btab671-B10","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2023020108470591600_btab671-B11","author":"Fan","year":"2019"},{"key":"2023020108470591600_btab671-B12","first-page":"231","volume-title":"et al.","author":"Fellbaum","year":"2010"},{"key":"2023020108470591600_btab671-B13","article-title":"Domain-specific language model pretraining for biomedical natural language processing","author":"Gu","year":"2020"},{"key":"2023020108470591600_btab671-B14","first-page":"2281","article-title":"Integrating graph contextualized knowledge into pre-trained language models","author":"He","year":"2020"},{"key":"2023020108470591600_btab671-B15","first-page":"1061","article-title":"Interpretable word embeddings for medical domain","author":"Jha","year":"2018"},{"key":"2023020108470591600_btab671-B16","first-page":"843","article-title":"Hypothesis generation from text based on co-evolution of biomedical concepts","author":"Jha","year":"2019"},{"key":"2023020108470591600_btab671-B17","doi-asserted-by":"crossref","first-page":"2190","DOI":"10.1093\/bioinformatics\/btab067","article-title":"Continual representation learning for evolving biomedical bipartite networks","author":"Jha","year":"2021","journal-title":"Bioinformatics"},{"key":"2023020108470591600_btab671-B18","first-page":"3077","article-title":"Knowledge-guided efficient representation learning for biomedical domain","author":"Jha","year":"2021"},{"key":"2023020108470591600_btab671-B19","first-page":"82","article-title":"Probing biomedical embeddings from language models","author":"Jin","year":"2019"},{"key":"2023020108470591600_btab671-B20","first-page":"61","article-title":"Temporal analysis of language through neural language models","author":"Kim","year":"2014"},{"key":"2023020108470591600_btab671-B21","doi-asserted-by":"crossref","first-page":"3521","DOI":"10.1073\/pnas.1611835114","article-title":"Overcoming catastrophic forgetting in neural networks","volume":"114","author":"Kirkpatrick","year":"2017","journal-title":"Proc. Natl. Acad. Sci"},{"key":"2023020108470591600_btab671-B22","first-page":"141","article-title":"Overview of the biocreative vi chemical-protein interaction track","author":"Krallinger","year":"2017"},{"key":"2023020108470591600_btab671-B23","author":"Lauscher","year":"2019"},{"key":"2023020108470591600_btab671-B24","doi-asserted-by":"crossref","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","article-title":"Biobert: a pre-trained biomedical language representation model for biomedical text mining","volume":"36","author":"Lee","year":"2020","journal-title":"Bioinformatics"},{"key":"2023020108470591600_btab671-B25","first-page":"4656","article-title":"Sensebert: driving some sense into Bert","author":"Levine","year":"2020"},{"key":"2023020108470591600_btab671-B26","doi-asserted-by":"crossref","first-page":"2935","DOI":"10.1109\/TPAMI.2017.2773081","article-title":"Learning without forgetting","volume":"40","author":"Li","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"2023020108470591600_btab671-B27","first-page":"1014","article-title":"Normalising medical concepts in social media texts by learning semantic representation","author":"Limsopatham","year":"2016"},{"key":"2023020108470591600_btab671-B28","first-page":"2901","author":"Liu","year":"2020"},{"key":"2023020108470591600_btab671-B29","first-page":"6470","article-title":"Gradient episodic memory for continual learning","author":"Lopez-Paz","year":"2017"},{"key":"2023020108470591600_btab671-B30","doi-asserted-by":"crossref","first-page":"1604","DOI":"10.1093\/bib\/bbz176","article-title":"Biomedical data and computational models for drug repositioning: a comprehensive review","volume":"22","author":"Luo","year":"2021","journal-title":"Brief. Bioinf"},{"key":"2023020108470591600_btab671-B31","doi-asserted-by":"crossref","first-page":"103132","DOI":"10.1016\/j.jbi.2019.103132","article-title":"MCN: a comprehensive corpus for medical concept normalization","volume":"92","author":"Luo","year":"2019","journal-title":"J. Biomed. Inf"},{"key":"2023020108470591600_btab671-B32","doi-asserted-by":"crossref","first-page":"287","DOI":"10.1007\/s10479-016-2393-z","article-title":"Data mining and predictive analytics applications for the delivery of healthcare services: a systematic literature review","volume":"270","author":"Malik","year":"2018","journal-title":"Ann. Oper. Res"},{"key":"2023020108470591600_btab671-B33","first-page":"393","article-title":"Deep neural models for medical concept normalization in user-generated texts","author":"Miftahutdinov","year":"2019"},{"key":"2023020108470591600_btab671-B34","first-page":"3111","article-title":"Distributed representations of words and phrases and their compositionality","author":"Mikolov","year":"2013"},{"key":"2023020108470591600_btab671-B35","first-page":"158","author":"Muneeb","year":"2015"},{"key":"2023020108470591600_btab671-B36","first-page":"48","article-title":"Results of the fifth edition of the bioasq challenge","author":"Nentidis","year":"2017"},{"key":"2023020108470591600_btab671-B37","first-page":"553","article-title":"Results of the seventh edition of the BioASQ challenge","author":"Nentidis","year":"2019"},{"key":"2023020108470591600_btab671-B38","doi-asserted-by":"crossref","first-page":"1239","DOI":"10.1007\/s11063-018-9873-x","article-title":"Multi-task character-level attentional networks for medical concept normalization","volume":"49","author":"Niu","year":"2019","journal-title":"Neural Process. Lett"},{"key":"2023020108470591600_btab671-B39","doi-asserted-by":"crossref","first-page":"1620","DOI":"10.1111\/j.1475-6773.2005.00444.x","article-title":"Measuring diagnoses: ICD code accuracy","volume":"40","author":"O\u2019Malley","year":"2005","journal-title":"Health Serv. Res"},{"key":"2023020108470591600_btab671-B40","doi-asserted-by":"crossref","first-page":"54","DOI":"10.1016\/j.neunet.2019.01.012","article-title":"Continual lifelong learning with neural networks: a review","volume":"113","author":"Parisi","year":"2019","journal-title":"Neural Netw"},{"key":"2023020108470591600_btab671-B41","first-page":"1532","article-title":"Glove: global vectors for word representation","author":"Pennington","year":"2014"},{"key":"2023020108470591600_btab671-B42","first-page":"43","article-title":"Knowledge enhanced contextual word representations","author":"Peters","year":"2019"},{"key":"2023020108470591600_btab671-B43","first-page":"15","article-title":"Semantic medline: an advanced information management application for biomedicine","volume":"31","author":"Rindflesch","year":"2011","journal-title":"Inf. Serv. Use"},{"key":"2023020108470591600_btab671-B44","article-title":"Distilbert, a distilled version of Bert: smaller, faster, cheaper","author":"Sanh","year":"2019"},{"key":"2023020108470591600_btab671-B45","doi-asserted-by":"crossref","first-page":"1274","DOI":"10.1093\/jamia\/ocy114","article-title":"Data and systems for medication-related text classification and concept normalization from twitter: insights from the social media mining for health (smm4h)-2017 shared task","volume":"25","author":"Sarker","year":"2018","journal-title":"J. Am. Med. Inf. Assoc"},{"key":"2023020108470591600_btab671-B46","doi-asserted-by":"crossref","first-page":"S2","DOI":"10.1186\/gb-2008-9-s2-s2","article-title":"Overview of biocreative II gene mention recognition","volume":"9","author":"Smith","year":"2008","journal-title":"Genome Biol"},{"key":"2023020108470591600_btab671-B47","first-page":"367","article-title":"Biont: deep learning using multiple biomedical ontologies for relation extraction","volume":"12036","author":"Sousa","year":"2020","journal-title":"Adv. Inf. Retrieval"},{"key":"2023020108470591600_btab671-B48","author":"Sun","year":"2019"},{"key":"2023020108470591600_btab671-B49","first-page":"5998","article-title":"Attention is all you need","author":"Vaswani","year":"2017"},{"key":"2023020108470591600_btab671-B50","first-page":"374","article-title":"Large scale incremental learning","author":"Wu","year":"2019"},{"key":"2023020108470591600_btab671-B51","first-page":"8452","article-title":"A generate-and-rank framework with semantic type regularization for biomedical concept normalization","author":"Xu","year":"2020"},{"key":"2023020108470591600_btab671-B52","doi-asserted-by":"crossref","first-page":"3794","DOI":"10.1093\/bioinformatics\/btz142","article-title":"Meshprobenet: a self-attentive probe net for mesh indexing","volume":"35","author":"Xun","year":"2019","journal-title":"Bioinformatics"},{"key":"2023020108470591600_btab671-B53","author":"Yoon","year":"2018"},{"key":"2023020108470591600_btab671-B54","doi-asserted-by":"crossref","first-page":"52","DOI":"10.1038\/s41597-019-0055-0","article-title":"Biowordvec, improving biomedical word embeddings with subword information and mesh","volume":"6","author":"Zhang","year":"2019","journal-title":"Sci. Data"},{"key":"2023020108470591600_btab671-B55","first-page":"1441","article-title":"Ernie: enhanced language representation with informative entities","author":"Zhang","year":"2019"},{"key":"2023020108470591600_btab671-B56","first-page":"1453","article-title":"Online incremental feature learning with denoising autoencoders","author":"Zhou","year":"2012"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btab671\/40500078\/btab671.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/38\/2\/494\/49007514\/btab671.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/38\/2\/494\/49007514\/btab671.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,1]],"date-time":"2023-02-01T20:03:20Z","timestamp":1675281800000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/38\/2\/494\/6374496"}},"subtitle":[],"editor":[{"given":"Jonathan","family":"Wren","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2021,9,23]]},"references-count":56,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2022,1,3]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btab671","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022,1,15]]},"published":{"date-parts":[[2021,9,23]]}}}