{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T11:32:40Z","timestamp":1780054360693,"version":"3.54.0"},"reference-count":125,"publisher":"MIT Press","issue":"3","license":[{"start":{"date-parts":[[2024,5,22]],"date-time":"2024-05-22T00:00:00Z","timestamp":1716336000000},"content-version":"vor","delay-in-days":142,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,9,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>The sparsity of labeled data is an obstacle to the development of Relation Extraction (RE) models and the completion of databases in various biomedical areas. While being of high interest in drug-discovery, the literature on natural products, reporting the identification of potential bioactive compounds from organisms, is a concrete example of such an overlooked topic. To mark the start of this new task, we created the first curated evaluation dataset and extracted literature items from the LOTUS database to build training sets. To this end, we developed a new sampler, inspired by diversity metrics in ecology, named Greedy Maximum Entropy sampler (https:\/\/github.com\/idiap\/gme-sampler). The strategic optimization of both balance and diversity of the selected items in the evaluation set is important given the resource-intensive nature of manual curation. After quantifying the noise in the training set, in the form of discrepancies between the text of input abstracts and the expected output labels, we explored different strategies accordingly. Framing the task as an end-to-end Relation Extraction, we evaluated the performance of standard fine-tuning (BioGPT, GPT-2, and Seq2rel) and few-shot learning with open Large Language Models (LLMs) (LLaMA 7B-65B). In addition to their evaluation in few-shot settings, we explore the potential of open LLMs as synthetic data generators and propose a new workflow for this purpose. All evaluated models exhibited substantial improvements when fine-tuned on synthetic abstracts rather than the original noisy data. We provide our best performing (F1-score = 59.0) BioGPT-Large model for end-to-end RE of natural products relationships along with all the training and evaluation datasets. See more details at https:\/\/github.com\/idiap\/abroad-re.<\/jats:p>","DOI":"10.1162\/coli_a_00520","type":"journal-article","created":{"date-parts":[[2024,5,22]],"date-time":"2024-05-22T14:30:51Z","timestamp":1716388251000},"page":"953-1000","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":6,"title":["Relation Extraction in Underexplored Biomedical Domains: A Diversity-optimized Sampling and Synthetic Data Generation Approach"],"prefix":"10.1162","volume":"50","author":[{"given":"Maxime","family":"Delmas","sequence":"first","affiliation":[{"name":"Idiap Research Institute Switzerland. maxime.delmas@idiap.ch"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Magdalena","family":"Wysocka","sequence":"additional","affiliation":[{"name":"Digital Cancer Research, CRUK National Biomarker Centre United Kingdom. magdalena.wysocka@digitalecmt.org"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andr\u00e9","family":"Freitas","sequence":"additional","affiliation":[{"name":"Idiap Research Institute Switzerland Department of Computer Science, University of Manchester Digital Cancer Research, CRUK National Biomarker Centre United Kingdom. andre.freitas@idiap.ch, andre.freitas@manchester.ac.uk"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,9,1]]},"reference":[{"key":"2024092014244774100_bib1","doi-asserted-by":"publisher","first-page":"5649","DOI":"10.18653\/v1\/2023.findings-acl.349","article-title":"ECG-QALM: Entity-controlled synthetic text generation using contextual Q&A for NER","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Aggarwal","year":"2023"},{"key":"2024092014244774100_bib2","doi-asserted-by":"publisher","first-page":"7319","DOI":"10.18653\/v1\/2021.acl-long.568","article-title":"Intrinsic dimensionality explains the effectiveness of language model fine-tuning","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Aghajanyan","year":"2021"},{"key":"2024092014244774100_bib3","doi-asserted-by":"publisher","first-page":"224","DOI":"10.1109\/ICOSC.2019.8665584","article-title":"Identifying protein-protein interaction using tree LSTM and structured attention","volume-title":"2019 IEEE 13th International Conference on Semantic Computing (ICSC)","author":"Ahmed","year":"2019"},{"key":"2024092014244774100_bib4","doi-asserted-by":"publisher","first-page":"2623","DOI":"10.1145\/3292500.3330701","article-title":"Optuna: A next-generation hyperparameter optimization framework","volume-title":"Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining","author":"Akiba","year":"2019"},{"key":"2024092014244774100_bib5","doi-asserted-by":"publisher","first-page":"7383","DOI":"10.1609\/aaai.v34i05.6233","article-title":"Do not have enough data? Deep learning to the rescue!","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Anaby-Tavor","year":"2020"},{"key":"2024092014244774100_bib6","doi-asserted-by":"publisher","first-page":"546","DOI":"10.18653\/v1\/S17-2091","article-title":"SemEval 2017 Task 10: ScienceIE - Extracting keyphrases and relations from scientific publications","volume-title":"Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017)","author":"Augenstein","year":"2017"},{"key":"2024092014244774100_bib7","first-page":"355","article-title":"Domain adaptation via pseudo in-domain data selection","volume-title":"Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing","author":"Axelrod","year":"2011"},{"key":"2024092014244774100_bib8","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3531863","article-title":"InPars: Data augmentation for information retrieval using large language models","author":"Bonifacio","year":"2022","journal-title":"ArXiv:2202.05144"},{"key":"2024092014244774100_bib9","first-page":"1877","article-title":"Language models are few-shot learners","volume-title":"Advances in Neural Information Processing Systems","author":"Brown","year":"2020"},{"key":"2024092014244774100_bib10","doi-asserted-by":"publisher","first-page":"1860","DOI":"10.18653\/v1\/2021.acl-long.146","article-title":"Knowledgeable or educated guess? Revisiting language models as knowledge bases","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Cao","year":"2021"},{"key":"2024092014244774100_bib11","doi-asserted-by":"publisher","first-page":"191","DOI":"10.1162\/tacl_a_00542","article-title":"An empirical survey of data augmentation for limited data learning in NLP","volume":"11","author":"Chen","year":"2023","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024092014244774100_bib12","article-title":"Weakly supervised data augmentation through prompting for dialogue understanding","volume-title":"NeurIPS 2022 Workshop on Synthetic Data for Empowering ML Research","author":"Chen","year":"2022"},{"key":"2024092014244774100_bib13","doi-asserted-by":"publisher","first-page":"719","DOI":"10.18653\/v1\/2022.acl-long.53","article-title":"Meta-learning via language model in-context tuning","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Chen","year":"2022"},{"key":"2024092014244774100_bib14","first-page":"35","article-title":"InstructEval: Towards holistic evaluation of instruction-tuned large language models","volume-title":"Proceedings of the First Edition of the Workshop on the Scaling Behavior of Large Language Models (SCALE-LLM 2024)","author":"Chia","year":"2024"},{"issue":"3","key":"2024092014244774100_bib15","first-page":"6","article-title":"Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality","volume":"2","author":"Chiang","year":"2023"},{"key":"2024092014244774100_bib16","doi-asserted-by":"publisher","first-page":"575","DOI":"10.18653\/v1\/2023.acl-long.34","article-title":"Increasing diversity while maintaining accuracy: Text data generation with large language models and human interventions","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Chung","year":"2023"},{"key":"2024092014244774100_bib17","article-title":"Promptagator: Few-shot dense retrieval from 8 examples","volume-title":"The Eleventh International Conference on Learning Representations","author":"Dai","year":"2023"},{"key":"2024092014244774100_bib18","article-title":"8-bit optimizers via block-wise quantization","volume-title":"International Conference on Learning Representations","author":"Dettmers","year":"2022"},{"key":"2024092014244774100_bib19","first-page":"10088","article-title":"QLoRA: Efficient finetuning of quantized LLMs","volume-title":"Advances in Neural Information Processing Systems","author":"Dettmers","year":"2023"},{"key":"2024092014244774100_bib20","doi-asserted-by":"publisher","first-page":"3650","DOI":"10.18653\/v1\/2021.eacl-main.319","article-title":"An end-to-end model for entity-level relation extraction using multi-instance learning","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Eberts","year":"2021"},{"key":"2024092014244774100_bib21","article-title":"Learning what data to learn","author":"Fan","year":"2017","journal-title":"ArXiv:1702.08635"},{"issue":"1","key":"2024092014244774100_bib22","doi-asserted-by":"publisher","first-page":"5779","DOI":"10.1609\/aaai.v32i1.12063","article-title":"Reinforcement learning for relation classification from noisy data","volume":"32","author":"Feng","year":"2018","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2024092014244774100_bib23","doi-asserted-by":"publisher","first-page":"968","DOI":"10.18653\/v1\/2021.findings-acl.84","article-title":"A survey of data augmentation approaches for NLP","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Feng","year":"2021"},{"issue":"4","key":"2024092014244774100_bib24","doi-asserted-by":"publisher","first-page":"736","DOI":"10.1016\/j.talanta.2005.03.025","article-title":"A method for calibration and validation subset partitioning","volume":"67","author":"Galvao","year":"2005","journal-title":"Talanta"},{"key":"2024092014244774100_bib25","article-title":"Self-guided noise-free data generation for efficient zero-shot learning","volume-title":"The Eleventh International Conference on Learning Representations","author":"Gao","year":"2023"},{"issue":"1","key":"2024092014244774100_bib26","doi-asserted-by":"publisher","first-page":"85","DOI":"10.1186\/1471-2105-11-85","article-title":"LINNAEUS: A species name identification system for biomedical literature","volume":"11","author":"Gerner","year":"2010","journal-title":"BMC Bioinformatics"},{"key":"2024092014244774100_bib27","doi-asserted-by":"publisher","first-page":"10","DOI":"10.18653\/v1\/2022.bionlp-1.2","article-title":"A sequence-to-sequence approach for document-level relation extraction","volume-title":"Proceedings of the 21st Workshop on Biomedical Language Processing","author":"Giorgi","year":"2022"},{"key":"2024092014244774100_bib28","doi-asserted-by":"publisher","first-page":"64323","DOI":"10.1109\/ACCESS.2019.2917620","article-title":"Diversity in machine learning","volume":"7","author":"Gong","year":"2019","journal-title":"IEEE Access"},{"key":"2024092014244774100_bib29","article-title":"KeyBERT: Minimal keyword extraction with BERT","author":"Grootendorst","year":"2020"},{"key":"2024092014244774100_bib30","doi-asserted-by":"publisher","first-page":"3309","DOI":"10.18653\/v1\/2022.acl-long.234","article-title":"ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Hartvigsen","year":"2022"},{"key":"2024092014244774100_bib31","doi-asserted-by":"publisher","first-page":"826","DOI":"10.1162\/tacl_a_00492","article-title":"Generate, annotate, and learn: NLP with synthetic text","volume":"10","author":"He","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"issue":"2","key":"2024092014244774100_bib32","doi-asserted-by":"publisher","first-page":"427","DOI":"10.2307\/1934352","article-title":"Diversity and evenness: A unifying notation and its consequences","volume":"54","author":"Hill","year":"1973","journal-title":"Ecology"},{"key":"2024092014244774100_bib33","article-title":"The curious case of neural text degeneration","volume-title":"International Conference on Learning Representations","author":"Holtzman","year":"2020"},{"issue":"22","key":"2024092014244774100_bib34","doi-asserted-by":"publisher","first-page":"5100","DOI":"10.1093\/bioinformatics\/btac648","article-title":"Discovering drug\u2013target interaction knowledge from biomedical literature","volume":"38","author":"Hou","year":"2022","journal-title":"Bioinformatics"},{"key":"2024092014244774100_bib35","article-title":"LoRA: Low-rank adaptation of large language models","volume-title":"International Conference on Learning Representations","author":"Hu","year":"2022"},{"key":"2024092014244774100_bib36","doi-asserted-by":"publisher","first-page":"10221","DOI":"10.18653\/v1\/2023.findings-acl.649","article-title":"GDA: Generative data augmentation techniques for relation extraction tasks","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Hu","year":"2023"},{"key":"2024092014244774100_bib37","article-title":"A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions","author":"Huang","year":"2023","journal-title":"CoRR"},{"key":"2024092014244774100_bib38","doi-asserted-by":"publisher","first-page":"2370","DOI":"10.18653\/v1\/2021.findings-emnlp.204","article-title":"REBEL: Relation extraction by end-to-end language generation","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2021","author":"CabotLlu\u00eds","year":"2021"},{"key":"2024092014244774100_bib39","doi-asserted-by":"publisher","first-page":"161","DOI":"10.18653\/v1\/2022.bionlp-1.16","article-title":"Improving supervised drug-protein relation extraction with distantly supervised models","volume-title":"Proceedings of the 21st Workshop on Biomedical Language Processing","author":"Iinuma","year":"2022"},{"key":"2024092014244774100_bib40","doi-asserted-by":"publisher","first-page":"3561","DOI":"10.1145\/3394486.3406477","article-title":"Overview and importance of data quality for machine learning tasks","volume-title":"Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining","author":"Jain","year":"2020"},{"key":"2024092014244774100_bib41","article-title":"Mixtral of experts","author":"Jiang","year":"2024","journal-title":"ArXiv:2401.04088"},{"key":"2024092014244774100_bib42","doi-asserted-by":"publisher","first-page":"4497","DOI":"10.18653\/v1\/2022.findings-emnlp.329","article-title":"Thinking about GPT-3 in-context learning for biomedical IE? Think again","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Jimenez Gutierrez","year":"2022"},{"issue":"2","key":"2024092014244774100_bib43","doi-asserted-by":"publisher","first-page":"166","DOI":"10.1080\/00401706.2021.1921037","article-title":"SPlit: An optimal method for data splitting","volume":"64","author":"Joseph","year":"2022","journal-title":"Technometrics"},{"key":"2024092014244774100_bib44","doi-asserted-by":"publisher","first-page":"1555","DOI":"10.18653\/v1\/2023.emnlp-main.96","article-title":"Exploiting asymmetry for synthetic training data generation: SynthIE and the case of information extraction","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Josifoski","year":"2023"},{"issue":"2","key":"2024092014244774100_bib45","doi-asserted-by":"publisher","first-page":"363","DOI":"10.1111\/j.2006.0030-1299.14714.x","article-title":"Entropy and diversity","volume":"113","author":"Jost","year":"2006","journal-title":"Oikos"},{"key":"2024092014244774100_bib46","doi-asserted-by":"publisher","first-page":"166","DOI":"10.1007\/978-3-031-14054-9_17","article-title":"Chemical-gene relation extraction with graph neural networks and BERT encoder","volume-title":"Proceedings of the ICR\u201922 International Conference on Innovations in Computing Research","author":"Kambar","year":"2022"},{"key":"2024092014244774100_bib47","doi-asserted-by":"publisher","first-page":"218","DOI":"10.1109\/AIIoT54504.2022.9817231","article-title":"A survey on deep learning techniques for joint named entities and relation extraction","volume-title":"2022 IEEE World AI IoT Congress (AIIoT)","author":"Kambar","year":"2022"},{"issue":"1","key":"2024092014244774100_bib48","doi-asserted-by":"publisher","first-page":"137","DOI":"10.1080\/00401706.1969.10490666","article-title":"Computer aided design of experiments","volume":"11","author":"Kennard","year":"1969","journal-title":"Technometrics"},{"key":"2024092014244774100_bib49","article-title":"Self-generated in-context learning: Leveraging auto-regressive language models as a demonstration generator","author":"Kim","year":"2022","journal-title":"CoRR"},{"issue":"11","key":"2024092014244774100_bib50","doi-asserted-by":"publisher","first-page":"2795","DOI":"10.1021\/acs.jnatprod.1c00399","article-title":"NPClassifier: A deep neural network-based structural classification tool for natural products","volume":"84","author":"Kim","year":"2021","journal-title":"Journal of Natural Products"},{"key":"2024092014244774100_bib51","first-page":"22199","article-title":"Large language models are zero-shot reasoners","volume":"35","author":"Kojima","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"1","key":"2024092014244774100_bib52","doi-asserted-by":"publisher","first-page":"S2","DOI":"10.1186\/1758-2946-7-S1-S2","article-title":"The CHEMDNER corpus of chemicals and drugs and its annotation principles","volume":"7","author":"Krallinger","year":"2015","journal-title":"Journal of Cheminformatics"},{"key":"2024092014244774100_bib53","first-page":"18","article-title":"Data augmentation using pre-trained transformer models","volume-title":"Proceedings of the 2nd Workshop on Life-long Learning for Spoken Language Systems","author":"Kumar","year":"2020"},{"key":"2024092014244774100_bib54","article-title":"BioMistral: A collection of open-source pretrained large language models for medical domains","author":"Labrak","year":"2024","journal-title":"ArXiv:2402.10373"},{"key":"2024092014244774100_bib55","doi-asserted-by":"publisher","DOI":"10.1017\/9781108963558","volume-title":"Entropy and Diversity: The Axiomatic Approach","author":"Leinster","year":"2021"},{"issue":"1","key":"2024092014244774100_bib56","doi-asserted-by":"publisher","first-page":"198","DOI":"10.1186\/s12859-017-1609-9","article-title":"A neural joint model for entity and relation extraction from biomedical text","volume":"18","author":"Li","year":"2017","journal-title":"BMC Bioinformatics"},{"key":"2024092014244774100_bib57","doi-asserted-by":"publisher","first-page":"baw068","DOI":"10.1093\/database\/baw068","article-title":"BioCreative V CDR task corpus: A resource for chemical disease relation extraction","volume":"2016","author":"Li","year":"2016","journal-title":"Database: The Journal of Biological Databases and Curation"},{"key":"2024092014244774100_bib58","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2308.12032","article-title":"From quantity to quality: Boosting LLM performance with self-guided data selection for instruction tuning","author":"Li","year":"2023","journal-title":"ArXiv:2308.12032"},{"key":"2024092014244774100_bib59","doi-asserted-by":"publisher","first-page":"106185","DOI":"10.1016\/j.phrs.2022.106185","article-title":"LTM-TCM: A comprehensive database for the linking of traditional Chinese medicine with modern medicine at molecular and phenotypic levels","volume":"178","author":"Li","year":"2022","journal-title":"Pharmacological Research"},{"issue":"8","key":"2024092014244774100_bib60","doi-asserted-by":"publisher","first-page":"669","DOI":"10.1038\/s42256-022-00516-1","article-title":"Advances, challenges and opportunities in creating data for trustworthy AI","volume":"4","author":"Liang","year":"2022","journal-title":"Nature Machine Intelligence"},{"key":"2024092014244774100_bib61","doi-asserted-by":"publisher","first-page":"100","DOI":"10.18653\/v1\/2022.deelio-1.10","article-title":"What makes good in-context examples for GPT-3?","volume-title":"Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures","author":"Liu","year":"2022"},{"issue":"5","key":"2024092014244774100_bib62","doi-asserted-by":"publisher","first-page":"bbac282","DOI":"10.1093\/bib\/bbac282","article-title":"BioRED: A rich biomedical relation extraction dataset","volume":"23","author":"Luo","year":"2022","journal-title":"Briefings in Bioinformatics"},{"issue":"6","key":"2024092014244774100_bib63","doi-asserted-by":"publisher","first-page":"bbac409","DOI":"10.1093\/bib\/bbac409","article-title":"BioGPT: Generative pre-trained transformer for biomedical text generation and mining","volume":"23","author":"Luo","year":"2022","journal-title":"Briefings in Bioinformatics"},{"key":"2024092014244774100_bib64","doi-asserted-by":"publisher","first-page":"9802","DOI":"10.18653\/v1\/2023.acl-long.546","article-title":"When not to trust language models: Investigating effectiveness of parametric and non-parametric memories","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Mallen","year":"2023"},{"key":"2024092014244774100_bib65","article-title":"PEFT: State-of-the-art parameter-efficient fine-tuning methods","author":"Mangrulkar","year":"2022"},{"key":"2024092014244774100_bib66","first-page":"5320","article-title":"DataPerf: Benchmarks for data-centric AI development","volume-title":"Advances in Neural Information Systems","author":"Mazumder","year":"2023"},{"key":"2024092014244774100_bib67","first-page":"462","article-title":"Generating training data with language models: Towards zero-shot language understanding","volume-title":"Advances in Neural Information Processing Systems","author":"Meng","year":"2022"},{"key":"2024092014244774100_bib68","first-page":"24457","article-title":"Tuning language models as training data generators for augmentation-enhanced few-shot learning","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Meng","year":"2023"},{"key":"2024092014244774100_bib69","doi-asserted-by":"publisher","first-page":"1003","DOI":"10.3115\/1690219.1690287","article-title":"Distant supervision for relation extraction without labeled data","volume-title":"Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP","author":"Mintz","year":"2009"},{"issue":"5","key":"2024092014244774100_bib70","doi-asserted-by":"publisher","first-page":"323","DOI":"10.1080\/00107510500052444","article-title":"Power laws, Pareto distributions and Zipf\u2019s law","volume":"46","author":"Newman","year":"2005","journal-title":"Contemporary Physics"},{"key":"2024092014244774100_bib71","doi-asserted-by":"publisher","first-page":"1373","DOI":"10.1613\/jair.1.12125","article-title":"Confident learning: Estimating uncertainty in dataset labels","volume":"70","author":"Northcutt","year":"2021","journal-title":"Journal of Artificial Intelligence Research"},{"key":"2024092014244774100_bib72","first-page":"1373","article-title":"Pervasive label errors in test sets destabilize machine learning benchmarks","volume-title":"Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1)","author":"Northcutt","year":"2021"},{"key":"2024092014244774100_bib73","article-title":"Structured prediction as translation between augmented natural languages","volume-title":"International Conference on Learning Representations","author":"Paolini","year":"2021"},{"key":"2024092014244774100_bib74","article-title":"DARE: Data augmented relation extraction with GPT-2","author":"Papanikolaou","year":"2020","journal-title":"ArXiv:(2004.13845"},{"key":"2024092014244774100_bib75","doi-asserted-by":"publisher","first-page":"311","DOI":"10.3115\/1073083.1073135","article-title":"BLEU: A method for automatic evaluation of machine translation","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics - ACL \u201902","author":"Papineni","year":"2001"},{"key":"2024092014244774100_bib76","doi-asserted-by":"publisher","first-page":"109803","DOI":"10.1016\/j.asoc.2022.109803","article-title":"Data augmentation techniques in natural language processing","volume":"132","author":"Pellicer","year":"2023","journal-title":"Applied Soft Computing"},{"key":"2024092014244774100_bib77","doi-asserted-by":"publisher","first-page":"96","DOI":"10.1109\/ICMLA.2015.22","article-title":"The effect of dataset size on training tweet sentiment classifiers","volume-title":"2015 IEEE 14th International Conference on Machine Learning and Applications (ICMLA)","author":"Prusa","year":"2015"},{"key":"2024092014244774100_bib78","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-srw.1","article-title":"ChatGPT vs human-authored text: Insights into controllable text summarization and sentence style transfer","author":"Pu","year":"2023","journal-title":"ArXiv:2306.07799"},{"key":"2024092014244774100_bib79","article-title":"Language models are unsupervised multitask learners","author":"Radford","year":"2019","journal-title":"OpenAI blog"},{"key":"2024092014244774100_bib80","doi-asserted-by":"publisher","first-page":"e70780","DOI":"10.7554\/eLife.70780","article-title":"The LOTUS initiative for open knowledge management in natural products research","volume":"11","author":"Rutz","year":"2022","journal-title":"eLife"},{"issue":"7","key":"2024092014244774100_bib81","doi-asserted-by":"publisher","first-page":"e0267590","DOI":"10.1371\/journal.pone.0267590","article-title":"Detection of changes in literary writing style using N-grams as style markers and supervised machine learning","volume":"17","author":"R\u00edos-Toledo","year":"2022","journal-title":"PLOS ONE"},{"key":"2024092014244774100_bib82","doi-asserted-by":"publisher","first-page":"83","DOI":"10.18653\/v1\/2022.naacl-srw.11","article-title":"Impact of training instance selection on domain-specific entity extraction using BERT","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Student Research Workshop","author":"Salhofer","year":"2022"},{"key":"2024092014244774100_bib83","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3411764.3445518","article-title":"\u201cEveryone wants to do the model work, not the data work\u201d: Data cascades in high-stakes AI","volume-title":"Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems","author":"Sambasivan","year":"2021"},{"key":"2024092014244774100_bib84","article-title":"BLOOM: A 176B-parameter open-access multilingual language model","author":"Scao","year":"2022","journal-title":"CoRR"},{"key":"2024092014244774100_bib85","doi-asserted-by":"publisher","first-page":"6943","DOI":"10.18653\/v1\/2021.emnlp-main.555","article-title":"Generating datasets with pretrained language models","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Schick","year":"2021"},{"key":"2024092014244774100_bib86","article-title":"A short survey of biomedical relation extraction techniques","author":"Shahab","year":"2017","journal-title":"ArXiv:1707.05850"},{"key":"2024092014244774100_bib87","doi-asserted-by":"publisher","first-page":"2054","DOI":"10.18653\/v1\/D18-1230","article-title":"Learning named entity tagger using domain-specific dictionary","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Shang","year":"2018"},{"key":"2024092014244774100_bib88","doi-asserted-by":"publisher","first-page":"165","DOI":"10.1007\/3-540-29782-0_13","article-title":"KNApSAcK: A comprehensive species-metabolite relationship database","author":"Shinbo","year":"2006","journal-title":"Plant Metabolomics"},{"key":"2024092014244774100_bib89","article-title":"Prompting GPT-3 to be reliable","volume-title":"The Eleventh International Conference on Learning Representations","author":"Si","year":"2023"},{"issue":"3","key":"2024092014244774100_bib90","doi-asserted-by":"publisher","first-page":"853","DOI":"10.1016\/j.eswa.2013.08.015","article-title":"Syntactic N-grams as machine learning features for natural language processing","volume":"41","author":"Sidorov","year":"2014","journal-title":"Expert Systems with Applications"},{"issue":"5","key":"2024092014244774100_bib91","doi-asserted-by":"publisher","first-page":"106:1\u2013106:35","DOI":"10.1145\/3241741","article-title":"Relation extraction using distant supervision: A survey","volume":"51","author":"Smirnova","year":"2018","journal-title":"ACM Computing Surveys"},{"issue":"2","key":"2024092014244774100_bib92","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3617130","article-title":"Language models in the loop: Incorporating prompting into weak supervision","volume":"1","author":"Smith","year":"2024","journal-title":"ACM \/ IMS Journal of Data Science"},{"issue":"1","key":"2024092014244774100_bib93","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s13321-020-00478-9","article-title":"COCONUT online: Collection of open natural products database","volume":"13","author":"Sorokina","year":"2021","journal-title":"Journal of Cheminformatics"},{"key":"2024092014244774100_bib94","doi-asserted-by":"publisher","first-page":"Art. 457","DOI":"10.3389\/fmicb.2017.00457","article-title":"Core microbiota and metabolome of Vitis vinifera L. cv. corvina grapes and musts","volume":"8","author":"Stefanini","year":"2017","journal-title":"Frontiers in Microbiology"},{"issue":"3","key":"2024092014244774100_bib95","doi-asserted-by":"publisher","first-page":"lqab062","DOI":"10.1093\/nargab\/lqab062","article-title":"RENET2: High-performance full-text gene\u2013disease relation extraction with iterative training data expansion","volume":"3","author":"Su","year":"2021","journal-title":"NAR Genomics and Bioinformatics"},{"issue":"7","key":"2024092014244774100_bib96","doi-asserted-by":"publisher","first-page":"e0216913","DOI":"10.1371\/journal.pone.0216913","article-title":"Using distant supervision to augment manually annotated data for relation extraction","volume":"14","author":"Su","year":"2019","journal-title":"PLOS ONE"},{"issue":"7","key":"2024092014244774100_bib97","doi-asserted-by":"publisher","first-page":"109","DOI":"10.1007\/s11306-016-1051-4","article-title":"Recon 2.2: From reconstruction to model of human metabolism","volume":"12","author":"Swainston","year":"2016","journal-title":"Metabolomics"},{"key":"2024092014244774100_bib98","article-title":"Does synthetic data generation of LLMs help clinical text mining?","author":"Tang","year":"2023","journal-title":"ArXiv:2303.04360"},{"issue":"5","key":"2024092014244774100_bib99","doi-asserted-by":"publisher","first-page":"419","DOI":"10.1038\/nbt.2488","article-title":"A community-driven global reconstruction of human metabolism","volume":"31","author":"Thiele","year":"2013","journal-title":"Nature Biotechnology"},{"issue":"2","key":"2024092014244774100_bib100","doi-asserted-by":"publisher","first-page":"184","DOI":"10.1207\/s15327914nc5402_5","article-title":"Phytoestrogen content of foods consumed in Canada, including","volume":"54","author":"Thompson","year":"2006","journal-title":"Nutrition and Cancer"},{"key":"2024092014244774100_bib101","article-title":"Generating faithful synthetic data with large language models: A case study in computational social science","author":"Veselovsky","year":"2023","journal-title":"ArXiv:2305.15041"},{"key":"2024092014244774100_bib102","doi-asserted-by":"publisher","first-page":"3711","DOI":"10.18653\/v1\/2020.emnlp-main.303","article-title":"Global-to-local neural networks for document-level relation extraction","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Wang","year":"2020"},{"key":"2024092014244774100_bib103","article-title":"Towards zero-label language learning","author":"Wang","year":"2021","journal-title":"ArXiv:2109.09193"},{"issue":"W1","key":"2024092014244774100_bib104","doi-asserted-by":"publisher","first-page":"W587\u2013W593","DOI":"10.1093\/nar\/gkz389","article-title":"PubTator central: Automated concept annotation for biomedical full text articles","volume":"47","author":"Wei","year":"2019","journal-title":"Nucleic Acids Research"},{"key":"2024092014244774100_bib105","doi-asserted-by":"publisher","first-page":"ocae045","DOI":"10.1093\/jamia\/ocae045","article-title":"PMC-LLaMA: Toward building open-source language models for medicine","author":"Wu","year":"2024","journal-title":"Journal of the American Medical Informatics Association"},{"key":"2024092014244774100_bib106","doi-asserted-by":"publisher","first-page":"1423","DOI":"10.18653\/v1\/2023.acl-long.79","article-title":"Self-adaptive in-context learning: An information compression perspective for in-context example selection and ordering","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wu","year":"2023"},{"issue":"1","key":"2024092014244774100_bib107","doi-asserted-by":"publisher","first-page":"73","DOI":"10.1162\/coli_a_00462","article-title":"Transformers and the representation of biomedical background knowledge","volume":"49","author":"Wysocki","year":"2023","journal-title":"Computational Linguistics"},{"key":"2024092014244774100_bib108","doi-asserted-by":"publisher","first-page":"8186","DOI":"10.18653\/v1\/2023.acl-long.455","article-title":"S2ynRE: Two-stage self-training with synthetic data for low-resource relation extraction","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Xu","year":"2023"},{"key":"2024092014244774100_bib109","doi-asserted-by":"publisher","first-page":"413","DOI":"10.18653\/v1\/2022.findings-emnlp.29","article-title":"Towards realistic low-resource relation extraction: A benchmark with empirical baseline study","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Xu","year":"2022"},{"issue":"3","key":"2024092014244774100_bib110","doi-asserted-by":"publisher","first-page":"249","DOI":"10.1007\/s41664-018-0068-2","article-title":"On splitting training and validation set: A comparative study of cross-validation, bootstrap and systematic sampling for estimating the generalization performance of supervised learning","volume":"2","author":"Xu","year":"2018","journal-title":"Journal of Analysis and Testing"},{"key":"2024092014244774100_bib111","doi-asserted-by":"publisher","first-page":"1008","DOI":"10.18653\/v1\/2020.findings-emnlp.90","article-title":"Generative data augmentation for commonsense reasoning","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Yang","year":"2020"},{"key":"2024092014244774100_bib112","doi-asserted-by":"publisher","first-page":"11653","DOI":"10.18653\/v1\/2022.emnlp-main.801","article-title":"ZeroGen: Efficient zero-shot learning via dataset generation","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Ye","year":"2022"},{"key":"2024092014244774100_bib113","doi-asserted-by":"publisher","first-page":"2225","DOI":"10.18653\/v1\/2021.findings-emnlp.192","article-title":"GPT3Mix: Leveraging large-scale language models for text augmentation","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2021","author":"Yoo","year":"2021"},{"key":"2024092014244774100_bib114","doi-asserted-by":"publisher","first-page":"baad054","DOI":"10.1093\/database\/baad054","article-title":"Biomedical relation extraction with knowledge base\u2013refined weak supervision","volume":"2023","author":"Yoon","year":"2023","journal-title":"Database"},{"key":"2024092014244774100_bib115","article-title":"Can data diversity enhance learning generalization?","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Yu","year":"2022"},{"key":"2024092014244774100_bib116","doi-asserted-by":"publisher","first-page":"367","DOI":"10.18653\/v1\/D19-1035","article-title":"Learning the extraction order of multiple relational facts in a sentence with reinforcement learning","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Zeng","year":"2019"},{"key":"2024092014244774100_bib117","doi-asserted-by":"publisher","first-page":"506","DOI":"10.18653\/v1\/P18-1047","article-title":"Extracting relational facts by an end-to-end neural model with copy mechanism","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zeng","year":"2018"},{"key":"2024092014244774100_bib118","doi-asserted-by":"publisher","DOI":"10.5772\/intechopen.111542","article-title":"Data-centric artificial intelligence: A survey","author":"Zha","year":"2023","journal-title":"ArXiv:2303.10158"},{"key":"2024092014244774100_bib119","doi-asserted-by":"publisher","first-page":"236","DOI":"10.18653\/v1\/2020.findings-emnlp.23","article-title":"Minimize exposure bias of Seq2Seq models in joint entity and relation extraction","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Zhang","year":"2020"},{"key":"2024092014244774100_bib120","article-title":"Instruction tuning for large language models: A survey","author":"Zhang","year":"2023","journal-title":"ArXiv:2308.10792"},{"issue":"5","key":"2024092014244774100_bib121","doi-asserted-by":"publisher","first-page":"1609","DOI":"10.1093\/bib\/bbz087","article-title":"Deep learning for drug-drug interaction extraction from the literature: A review","volume":"21","author":"Zhang","year":"2020","journal-title":"Briefings in Bioinformatics"},{"issue":"3","key":"2024092014244774100_bib122","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1093\/bib\/bbaa057","article-title":"Recent advances in biomedical literature mining.","volume":"22","author":"Zhao","year":"2020","journal-title":"Briefings in Bioinformatics"},{"key":"2024092014244774100_bib123","article-title":"A comprehensive survey on deep learning for relation extraction: Recent advances and new frontiers","author":"Zhao","year":"2023","journal-title":"CoRR"},{"key":"2024092014244774100_bib124","first-page":"12697","article-title":"Calibrate before use: Improving few-shot performance of language models","volume-title":"Proceedings of the 38th International Conference on Machine Learning","author":"Zhao","year":"2021"},{"key":"2024092014244774100_bib125","first-page":"55006","article-title":"LIMA: Less is more for alignment","volume-title":"Thirty-seventh Conference on Neural Information Processing Systems","author":"Zhou","year":"2023"}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/50\/3\/953\/2471072\/coli_a_00520.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/50\/3\/953\/2471072\/coli_a_00520.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,20]],"date-time":"2024-09-20T14:25:29Z","timestamp":1726842329000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/50\/3\/953\/121178\/Relation-Extraction-in-Underexplored-Biomedical"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":125,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2024,9,1]]},"published-print":{"date-parts":[[2024,9,1]]}},"URL":"https:\/\/doi.org\/10.1162\/coli_a_00520","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"value":"0891-2017","type":"print"},{"value":"1530-9312","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}