{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T04:35:14Z","timestamp":1759206914694,"version":"3.41.2"},"reference-count":33,"publisher":"Oxford University Press (OUP)","license":[{"start":{"date-parts":[[2023,8,8]],"date-time":"2023-08-08T00:00:00Z","timestamp":1691452800000},"content-version":"vor","delay-in-days":219,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2023,7,26]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Biomedical relation extraction (BioRE) is the task of automatically extracting and classifying relations between two biomedical entities in biomedical literature. Recent advances in BioRE research have largely been powered by supervised learning and large language models (LLMs). However, training of LLMs for BioRE with supervised learning requires human-annotated data, and the annotation process often accompanies challenging and expensive work. As a result, the quantity and coverage of annotated data are limiting factors for BioRE systems. In this paper, we present our system for the BioCreative VII challenge\u2014DrugProt track, a BioRE system that leverages a language model structure and weak supervision. Our system is trained on weakly labelled data and then fine-tuned using human-labelled data. To create the weakly labelled dataset, we combined two approaches. First, we trained a model on the original dataset to predict labels on external literature, which will become a model-labelled dataset. Then, we refined the model-labelled dataset using an external knowledge base. Based on our experiment, our approach using refined weak supervision showed significant performance gain over the model trained using standard human-labelled datasets. Our final model showed outstanding performance at the BioCreative VII challenge, achieving 3rd place (this paper focuses on our participating system in the BioCreative VII challenge).<\/jats:p>\n               <jats:p>Database URL: http:\/\/wonjin.info\/biore-yoon-et-al-2022<\/jats:p>","DOI":"10.1093\/database\/baad054","type":"journal-article","created":{"date-parts":[[2023,8,8]],"date-time":"2023-08-08T10:53:34Z","timestamp":1691492014000},"source":"Crossref","is-referenced-by-count":3,"title":["Biomedical relation extraction with knowledge base\u2013refined weak supervision"],"prefix":"10.1093","volume":"2023","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6435-548X","authenticated-orcid":false,"given":"Wonjin","family":"Yoon","sequence":"first","affiliation":[{"name":"Department of Computer Science and Engineering, Korea University, 145 Anam-ro, Seongbuk-gu, Seoul 02841, South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sean","family":"Yi","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Korea University, 145 Anam-ro, Seongbuk-gu, Seoul 02841, South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Richard","family":"Jackson","sequence":"additional","affiliation":[{"name":"AstraZeneca UK, 1 Francis Crick Ave, Trumpington, Cambridge CB2 0AA, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2996-2564","authenticated-orcid":false,"given":"Hyunjae","family":"Kim","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Korea University, 145 Anam-ro, Seongbuk-gu, Seoul 02841, South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0240-6210","authenticated-orcid":false,"given":"Sunkyu","family":"Kim","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Korea University, 145 Anam-ro, Seongbuk-gu, Seoul 02841, South Korea"},{"name":"AIGEN Sciences Inc., 25 Ttukseom-ro 1-gil, Seongdong-gu, Seoul 04778, South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6798-9106","authenticated-orcid":false,"given":"Jaewoo","family":"Kang","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Korea University, 145 Anam-ro, Seongbuk-gu, Seoul 02841, South Korea"},{"name":"AIGEN Sciences Inc., 25 Ttukseom-ro 1-gil, Seongdong-gu, Seoul 04778, South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2023,7,26]]},"reference":[{"key":"2023080810525883100_R1","doi-asserted-by":"crossref","first-page":"317","DOI":"10.1007\/s10115-019-01351-4","article-title":"Constructing biomedical domain-specific knowledge graph with minimum supervision","volume":"62","author":"Yuan","year":"2020","journal-title":"Knowl. Inf. Syst."},{"key":"2023080810525883100_R2","doi-asserted-by":"crossref","first-page":"360","DOI":"10.1093\/bib\/bbz171","article-title":"Deep learning for\u00a0drug response prediction in\u00a0cancer","volume":"22","author":"Baptista","year":"2021","journal-title":"Briefings Bioinf."},{"key":"2023080810525883100_R3","doi-asserted-by":"crossref","DOI":"10.1016\/j.websem.2022.100756","article-title":"Comparison of\u00a0biomedical relationship extraction methods and\u00a0models for\u00a0knowledge graph creation","volume":"75","author":"Milo\u0161evi\u0107","year":"2023","journal-title":"J. Web Semant."},{"key":"2023080810525883100_R4","doi-asserted-by":"crossref","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","article-title":"Biobert: a pre-trained biomedical language representation model for\u00a0biomedical text mining","volume":"36","author":"Lee","year":"2020","journal-title":"Bioinformatics"},{"key":"2023080810525883100_R5","first-page":"1","article-title":"Domain-specific language model pretraining for\u00a0biomedical natural language processing","volume":"3","author":"Gu","year":"2021","journal-title":"ACM Trans. Comput. Healthcare (HEALTH)"},{"key":"2023080810525883100_R6","first-page":"pp. 4700","article-title":"BioMegatron: larger biomedical domain language model","author":"Shin","year":"2020"},{"key":"2023080810525883100_R7","first-page":"pp. 146","article-title":"Pretrained language models for\u00a0biomedical and\u00a0clinical tasks: understanding and\u00a0extending the state-of-the-art","author":"Lewis","year":"2020"},{"key":"2023080810525883100_R8","doi-asserted-by":"crossref","first-page":"pp. 1","DOI":"10.1093\/database\/baac071","article-title":"Challenges and\u00a0opportunities for\u00a0mining adverse drug reactions: perspectives from pharma, regulatory agencies, healthcare providers and\u00a0consumers","volume":"2022","author":"Gonzalez-Hernandez","year":"2022","journal-title":"Database"},{"key":"2023080810525883100_R9","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s13321-016-0165-z","article-title":"Improving chemical disease relation extraction with rich features and\u00a0weakly labeled data","volume":"8","author":"Peng","year":"2016","journal-title":"J. Cheminf."},{"key":"2023080810525883100_R10","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s12859-015-0472-9","article-title":"Extraction of\u00a0relations between genes and\u00a0diseases from text and\u00a0large-scale data analysis: implications for\u00a0translational research","volume":"16","author":"Bravo","year":"2015","journal-title":"BMC Bioinf."},{"key":"2023080810525883100_R11","first-page":"pp. 872","article-title":"Simultaneously self-attending to all mentions for\u00a0full-abstract biological relation extraction","author":"Verga","year":"2018"},{"key":"2023080810525883100_R12","first-page":"pp. 1775","article-title":"Named entity recognition with small strongly labeled and\u00a0large weakly labeled data","author":"Jiang","year":"2021"},{"key":"2023080810525883100_R13","first-page":"pp. 619","article-title":"Biomedical NER for\u00a0the enterprise with distillated BERN2 and\u00a0the kazu framework","author":"Yoon","year":"2022"},{"key":"2023080810525883100_R14","first-page":"pp. 1003","article-title":"Distant supervision for\u00a0relation extraction without labeled data","author":"Mintz","year":"2009"},{"key":"2023080810525883100_R15","first-page":"pp. 11","article-title":"Distantly supervised relation extraction with sentence reconstruction and\u00a0knowledge base priors","author":"Christopoulou","year":"2021"},{"key":"2023080810525883100_R16","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1093\/bib\/bbac282","article-title":"BioRED: a rich biomedical relation extraction dataset","volume":"23","author":"Luo","year":"2022","journal-title":"Briefings Bioinf."},{"key":"2023080810525883100_R17","first-page":"pp. 362","article-title":"Snorkel drybell: a case study in\u00a0deploying weak supervision at industrial scale","author":"Bach","year":"2019"},{"article-title":"Overview of\u00a0DrugProt BioCreative VII track: quality evaluation and\u00a0large scale text mining of\u00a0drug-gene\/protein relations","year":"2021","author":"Miranda","key":"2023080810525883100_R18"},{"key":"2023080810525883100_R19","first-page":"pp. 141","article-title":"Overview of\u00a0the biocreative vi chemical-protein interaction track","author":"Krallinger","year":"2017"},{"key":"2023080810525883100_R20","first-page":"pp. 101","article-title":"Stanza: A python natural language processing toolkit for\u00a0many human languages","author":"Qi","year":"2020"},{"key":"2023080810525883100_R21","doi-asserted-by":"crossref","first-page":"1892","DOI":"10.1093\/jamia\/ocab090","article-title":"Biomedical and\u00a0clinical english model packages for\u00a0the Stanza Python NLP library","volume":"28","author":"Zhang","year":"2021","journal-title":"J. Am. Med. Inf. Assoc."},{"key":"2023080810525883100_R22","doi-asserted-by":"crossref","first-page":"73729","DOI":"10.1109\/ACCESS.2019.2920708","article-title":"A neural named entity recognition and\u00a0multi-type normalization tool for\u00a0biomedical text mining","volume":"7","author":"Kim","year":"2019","journal-title":"IEEE Access"},{"key":"2023080810525883100_R23","doi-asserted-by":"crossref","first-page":"D939","DOI":"10.1093\/nar\/gkaa980","article-title":"Genenames.org: the HGNC and\u00a0VGNC resources in\u00a02021","volume":"49","author":"Tweedie","year":"2020","journal-title":"Nucleic Acids Res."},{"key":"2023080810525883100_R24","first-page":"pp. 4171","article-title":"BERT: pre-training of\u00a0deep bidirectional transformers for\u00a0language understanding","author":"Devlin","year":"2019"},{"key":"2023080810525883100_R25","doi-asserted-by":"crossref","first-page":"D1138","DOI":"10.1093\/nar\/gkaa891","article-title":"Comparative toxicogenomics database (CTD): update 2021","volume":"49","author":"Peter Davis","year":"2021","journal-title":"Nucleic Acids Res."},{"key":"2023080810525883100_R26","doi-asserted-by":"crossref","first-page":"914","DOI":"10.1016\/j.jbi.2013.07.011","article-title":"The DDI corpus: an annotated corpus with pharmacological substances and\u00a0drug\u2013drug interactions","volume":"46","author":"Herrero-Zazo","year":"2013","journal-title":"J. Biomed. Inf."},{"key":"2023080810525883100_R27","first-page":"pp. 341","article-title":"SemEval-2013 task 9: extraction of\u00a0drug-drug interactions from biomedical texts (DDIExtraction 2013)","author":"Segura-Bedmar","year":"2013"},{"key":"2023080810525883100_R28","first-page":"pp. 298","article-title":"BEEDS: Large-scale biomedical event extraction using distant supervision and\u00a0question answering","author":"David Wang","year":"2022"},{"key":"2023080810525883100_R29","doi-asserted-by":"crossref","first-page":"5678","DOI":"10.1093\/bioinformatics\/btaa1087","article-title":"BERT-GT: cross-sentence n-ary relation extraction with BERT and\u00a0Graph Transformer","volume":"36","author":"Lai","year":"2021","journal-title":"Bioinformatics"},{"key":"2023080810525883100_R30","first-page":"pp. 10","article-title":"A sequence-to-sequence approach for\u00a0document-level relation extraction","author":"Giorgi","year":"2022"},{"key":"2023080810525883100_R31","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1093\/database\/baac098","article-title":"Chemical\u2013protein relation extraction with ensembles of\u00a0carefully tuned pretrained language models","volume":"2022","author":"Weber","year":"2022","journal-title":"Database"},{"key":"2023080810525883100_R32","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1093\/database\/baac058","article-title":"A sequence labeling framework for\u00a0extracting drug\u2013protein relations from biomedical literature","volume":"2022","author":"Luo","year":"2022","journal-title":"Database"},{"key":"2023080810525883100_R33","first-page":"pp. 2054","article-title":"Learning named entity tagger using domain-specific dictionary","author":"Shang","year":"2018"}],"container-title":["Database"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baad054\/51052930\/baad054.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baad054\/51052930\/baad054.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,8,8]],"date-time":"2023-08-08T10:54:08Z","timestamp":1691492048000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/database\/article\/doi\/10.1093\/database\/baad054\/7238699"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,1]]},"references-count":33,"URL":"https:\/\/doi.org\/10.1093\/database\/baad054","relation":{},"ISSN":["1758-0463"],"issn-type":[{"type":"electronic","value":"1758-0463"}],"subject":[],"published-other":{"date-parts":[[2023,1,1]]},"published":{"date-parts":[[2023,1,1]]},"article-number":"baad054"}}