{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,3]],"date-time":"2026-05-03T17:49:30Z","timestamp":1777830570148,"version":"3.51.4"},"reference-count":50,"publisher":"Oxford University Press (OUP)","license":[{"start":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T00:00:00Z","timestamp":1759190400000},"content-version":"vor","delay-in-days":272,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000092","name":"National Library of Medicine","doi-asserted-by":"publisher","award":["U24HG010859"],"award-info":[{"award-number":["U24HG010859"]}],"id":[{"id":"10.13039\/100000092","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["U24HG002223"],"award-info":[{"award-number":["U24HG002223"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["U41HG000739"],"award-info":[{"award-number":["U41HG000739"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000051","name":"National Human Genome Research Institute","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000051","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,1,18]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Biological knowledgebases are essential resources for biomedical researchers, providing ready access to gene function and genomic data. Professional, manual curation of knowledgebases, however, is labour-intensive and thus high-performing machine learning (ML) methods that improve biocuration efficiency are needed. Here, we report on sentence-level classification to identify biocuration-relevant sentences in the full text of published references for two gene function data types: gene expression and protein kinase activity. We performed a detailed characterization of sentences from references in the WormBase bibliography and used this characterization to define three tasks for classifying sentences as either (i) fully curatable, (ii) fully and partially curatable, or (iii) all language-related. We evaluated various ML models applied to these tasks and found that GPT and BioBERT achieve the highest average performance, resulting in F1 performance scores ranging from 0.89 to 0.99 depending upon the task. Moreover, our inter-annotator agreement analyses and curator timing exercises demonstrated that curators readily converged on classification of high-quality training sentences that take a relatively short period of time to collect, making expansion of this approach to other data types a realistic addition to existing biocuration workflows. Our findings demonstrate the feasibility of extracting biocuration-relevant sentences from full text. Integrating these models into professional biocuration workflows, such as those used by the Alliance of Genome Resources and the ACKnowledge community curation platform, might well facilitate efficient and accurate annotation of the biomedical literature.<\/jats:p>","DOI":"10.1093\/database\/baaf063","type":"journal-article","created":{"date-parts":[[2025,8,26]],"date-time":"2025-08-26T11:31:33Z","timestamp":1756207893000},"source":"Crossref","is-referenced-by-count":3,"title":["Characterization and automated classification of sentences in the biomedical literature: a case study for biocuration of gene expression and protein kinase activity"],"prefix":"10.1093","volume":"2025","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4945-5837","authenticated-orcid":false,"given":"Daniela","family":"Raciti","sequence":"first","affiliation":[{"name":"Division of Biology and Biological Engineering, California Institute of Technology , 1200 E. California Boulevard, Pasadena, CA 91125 ,","place":["United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1706-4196","authenticated-orcid":false,"given":"Kimberly M","family":"Van\u00a0Auken","sequence":"additional","affiliation":[{"name":"Division of Biology and Biological Engineering, California Institute of Technology , 1200 E. California Boulevard, Pasadena, CA 91125 ,","place":["United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2563-5374","authenticated-orcid":false,"given":"Valerio","family":"Arnaboldi","sequence":"additional","affiliation":[{"name":"Division of Biology and Biological Engineering, California Institute of Technology , 1200 E. California Boulevard, Pasadena, CA 91125 ,","place":["United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8746-0680","authenticated-orcid":false,"given":"Christopher J","family":"Tabone","sequence":"additional","affiliation":[{"name":"The Jackson Laboratory , 600 Main Street, Bar Harbor, ME 04609 ,","place":["United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1036-3051","authenticated-orcid":false,"given":"Hans-Michael","family":"Muller","sequence":"additional","affiliation":[{"name":"Division of Biology and Biological Engineering, California Institute of Technology , 1200 E. California Boulevard, Pasadena, CA 91125 ,","place":["United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7699-0173","authenticated-orcid":false,"given":"Paul W","family":"Sternberg","sequence":"additional","affiliation":[{"name":"Division of Biology and Biological Engineering, California Institute of Technology , 1200 E. California Boulevard, Pasadena, CA 91125 ,","place":["United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2025,9,30]]},"reference":[{"key":"2025093010290058600_bib1","doi-asserted-by":"publisher","first-page":"iyac001","DOI":"10.1093\/genetics\/iyac001","article-title":"Making biological knowledge useful for humans and machines","volume":"220","author":"Wood","year":"2022","journal-title":"Genetics"},{"key":"2025093010290058600_bib2","doi-asserted-by":"publisher","first-page":"160018","DOI":"10.1038\/sdata.2016.18","article-title":"The FAIR Guiding Principles for scientific data management and stewardship","volume":"3","author":"Wilkinson","year":"2016","journal-title":"Sci Data"},{"key":"2025093010290058600_bib3","doi-asserted-by":"publisher","first-page":"e2002846","DOI":"10.1371\/journal.pbio.2002846","article-title":"Biocuration: distilling data into knowledge","volume":"16","author":"International Society for Biocuration","year":"2018","journal-title":"PLoS Biol"},{"key":"2025093010290058600_bib4","doi-asserted-by":"publisher","DOI":"10.1038\/nature.2016.20134","article-title":"Concern over funding cuts for model organism databases","author":"Hayden","year":"2016","journal-title":"Nature"},{"key":"2025093010290058600_bib5","doi-asserted-by":"publisher","first-page":"14","DOI":"10.1126\/science.351.6268.14","article-title":"BIOMEDICAL RESOURCES. Funding for key data resources in jeopardy","volume":"351","author":"Kaiser","year":"2016","journal-title":"Science"},{"key":"2025093010290058600_bib6","doi-asserted-by":"publisher","first-page":"1189","DOI":"10.1534\/genetics.119.302523","article-title":"The Alliance of Genome Resources: building a modern data ecosystem for model organism databases","volume":"213","author":"Alliance of Genome Resources Consortium","year":"2019","journal-title":"Genetics"},{"key":"2025093010290058600_bib7","doi-asserted-by":"publisher","first-page":"228","DOI":"10.1186\/1471-2105-10-228","article-title":"Semi-automated curation of protein subcellular localization: a text mining-based approach to Gene Ontology (GO) Cellular Component curation","volume":"10","author":"Van\u00a0Auken","year":"2009","journal-title":"BMC Bioinf"},{"key":"2025093010290058600_bib8","doi-asserted-by":"publisher","first-page":"16","DOI":"10.1186\/1471-2105-13-16","article-title":"Automatic categorization of diverse experimental information in the bioscience literature","volume":"13","author":"Fang","year":"2012","journal-title":"BMC Bioinf"},{"key":"2025093010290058600_bib9","doi-asserted-by":"publisher","first-page":"bau129","DOI":"10.1093\/database\/bau129","article-title":"OntoMate: a text-mining tool aiding curation at the Rat Genome Database","volume":"2015","author":"Liu","year":"2015","journal-title":"Database"},{"key":"2025093010290058600_bib10","doi-asserted-by":"publisher","first-page":"17","DOI":"10.1109\/TCBB.2014.2372765","article-title":"RLIMS-P 2.0: a generalizable rule-based information extraction system for literature mining of protein phosphorylation information","volume":"12","author":"Torii","year":"2015","journal-title":"IEEE\/ACM Trans Comput Biol Bioinf"},{"key":"2025093010290058600_bib11","doi-asserted-by":"publisher","first-page":"bav020","DOI":"10.1093\/database\/bav020","article-title":"Construction of phosphorylation interaction networks by text mining of full-length articles using the eFIP system","volume":"2015","author":"Tudor","year":"2015","journal-title":"Database"},{"key":"2025093010290058600_bib12","doi-asserted-by":"publisher","first-page":"10148","DOI":"10.1038\/s41598-018-28330-z","article-title":"Using machine learning tools for protein database biocuration assistance","volume":"8","author":"K\u00f6nig","year":"2018","journal-title":"Sci Rep"},{"key":"2025093010290058600_bib13","doi-asserted-by":"publisher","first-page":"154","DOI":"10.3389\/fphys.2019.00154","article-title":"Xenbase: facilitating the use of Xenopus to model human disease","volume":"10","author":"Nenni","year":"2019","journal-title":"Front Physiol"},{"key":"2025093010290058600_bib14","doi-asserted-by":"publisher","first-page":"baaa026","DOI":"10.1093\/database\/baaa026","article-title":"UPCLASS: a deep learning-based classifier for UniProtKB entry publications","volume":"2020","author":"Teodoro","year":"2020","journal-title":"Database"},{"key":"2025093010290058600_bib15","doi-asserted-by":"publisher","first-page":"baab062","DOI":"10.1093\/database\/baab062","article-title":"Classifying domain-specific text documents containing ambiguous keywords","volume":"2021","author":"Karimi","year":"2021","journal-title":"Database"},{"key":"2025093010290058600_bib16","doi-asserted-by":"publisher","first-page":"i468","DOI":"10.1093\/bioinformatics\/btab331","article-title":"Utilizing image and caption information for biomedical document classification","volume":"37","author":"Li","year":"2021","journal-title":"Bioinformatics"},{"key":"2025093010290058600_bib17","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1007\/s00335-021-09921-0","article-title":"Mouse Genome Informatics (MGI): latest news from MGD and GXD","volume":"33","author":"Ringwald","year":"2022","journal-title":"Mamm Genome"},{"key":"2025093010290058600_bib18","doi-asserted-by":"publisher","first-page":"bas030","DOI":"10.1093\/database\/bas030","article-title":"Assessment of community-submitted ontology annotations from a novel database-journal partnership","volume":"2012","author":"Berardini","year":"2012","journal-title":"Database"},{"key":"2025093010290058600_bib19","doi-asserted-by":"publisher","first-page":"bas024","DOI":"10.1093\/database\/bas024","article-title":"Directly e-mailing authors of newly published papers encourages community curation","volume":"2012","author":"Bunt","year":"2012","journal-title":"Database"},{"key":"2025093010290058600_bib20","doi-asserted-by":"publisher","first-page":"baaa006","DOI":"10.1093\/database\/baaa006","article-title":"Text mining meets community curation: a newly designed curation platform to improve author experience and participation at WormBase","volume":"2020","author":"Arnaboldi","year":"2020","journal-title":"Database"},{"key":"2025093010290058600_bib21","doi-asserted-by":"publisher","first-page":"baaa028","DOI":"10.1093\/database\/baaa028","article-title":"Community curation in PomBase: enabling fission yeast experts to provide detailed, standardized, sharable annotation from research publications","volume":"2020","author":"Lock","year":"2020","journal-title":"Database"},{"key":"2025093010290058600_bib22","doi-asserted-by":"publisher","first-page":"D899","DOI":"10.1093\/nar\/gkaa1026","article-title":"FlyBase: updates to the Drosophila melanogaster knowledge base","volume":"49","author":"Larkin","year":"2021","journal-title":"Nucleic Acids Res"},{"key":"2025093010290058600_bib23","doi-asserted-by":"publisher","first-page":"e3001464","DOI":"10.1371\/journal.pbio.3001464","article-title":"A crowdsourcing open platform for literature curation in UniProt","volume":"19","author":"Wang","year":"2021","journal-title":"PLoS Biol"},{"key":"2025093010290058600_bib24","doi-asserted-by":"publisher","first-page":"iyae050","DOI":"10.1093\/genetics\/iyae050","article-title":"WormBase 2024: status and transitioning to Alliance infrastructure","volume":"227","author":"Sternberg","year":"2024","journal-title":"Genetics"},{"key":"2025093010290058600_bib25","doi-asserted-by":"publisher","first-page":"2086","DOI":"10.1093\/bioinformatics\/btn381","article-title":"Multi-dimensional classification of biomedical text: toward automated, practical provision of high-utility text to diverse users","volume":"24","author":"Shatkay","year":"2008","journal-title":"Bioinformatics"},{"key":"2025093010290058600_bib26","doi-asserted-by":"publisher","first-page":"3174","DOI":"10.1093\/bioinformatics\/btp548","article-title":"Automatically classifying sentences in full-text biomedical articles into Introduction, Methods, Results and Discussion","volume":"25","author":"Agarwal","year":"2009","journal-title":"Bioinformatics"},{"key":"2025093010290058600_bib27","doi-asserted-by":"publisher","first-page":"541","DOI":"10.1186\/s12859-018-2496-4","article-title":"Fast and scalable neural embedding models for biomedical sentence classification","volume":"19","author":"Agibetov","year":"2018","journal-title":"BMC Bioinf"},{"key":"2025093010290058600_bib28","doi-asserted-by":"publisher","first-page":"bau074","DOI":"10.1093\/database\/bau074","article-title":"BC4GO: a full-text corpus for the BioCreative IV GO task","volume":"2014","author":"Van\u00a0Auken","year":"2014","journal-title":"Database"},{"key":"2025093010290058600_bib29","doi-asserted-by":"publisher","first-page":"bau086","DOI":"10.1093\/database\/bau086","article-title":"Overview of the gene ontology task at BioCreative IV","volume":"2014","author":"Mao","year":"2014","journal-title":"Database"},{"key":"2025093010290058600_bib30","doi-asserted-by":"publisher","first-page":"baw121","DOI":"10.1093\/database\/baw121","article-title":"BioCreative V BioC track overview: collaborative biocurator assistant task for BioGRID","volume":"2016","author":"Kim","year":"2016","journal-title":"Database"},{"key":"2025093010290058600_bib31","doi-asserted-by":"publisher","first-page":"4171","DOI":"10.18653\/v1\/N19-1423","article-title":"BERT: pre-training of deep bidirectional transformers for language understanding","volume":"1","author":"Devlin","year":"2019","journal-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies."},{"key":"2025093010290058600_bib32","doi-asserted-by":"publisher","first-page":"52","DOI":"10.1038\/s41597-019-0055-0","article-title":"BioWordVec, improving biomedical word embeddings with subword information and MeSH","volume":"6","author":"Zhang","year":"2019","journal-title":"Sci Data"},{"key":"2025093010290058600_bib33","doi-asserted-by":"publisher","DOI":"10.1109\/ICHI.2019.8904728","article-title":"BioSentVec: creating sentence embeddings for biomedical texts","volume-title":"The 7th IEEE International Conference on Healthcare Informatics","author":"Chen","year":"2019"},{"key":"2025093010290058600_bib34","doi-asserted-by":"publisher","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","article-title":"BioBERT: a pre-trained biomedical language representation model for biomedical text mining","volume":"36","author":"Lee","year":"2020","journal-title":"Bioinformatics"},{"key":"2025093010290058600_bib35","doi-asserted-by":"publisher","first-page":"377","DOI":"10.1007\/s00799-023-00392-z","article-title":"Sequential sentence classification in research papers using cross-domain multi-task learning","volume":"25","author":"Brack","year":"2024","journal-title":"Int J Digit Libr"},{"key":"2025093010290058600_bib36","article-title":"Improving language understanding by generative pre-training","author":"Radford","year":"2018","journal-title":"OpenAI"},{"key":"2025093010290058600_bib37","first-page":"1877","article-title":"Language models are few-shot learners","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"Brown","year":"2020"},{"key":"2025093010290058600_bib38","doi-asserted-by":"publisher","first-page":"iyae049","DOI":"10.1093\/genetics\/iyae049","article-title":"Updates to the Alliance of Genome Resources central infrastructure","volume":"227","author":"Alliance of Genome Resources Consortium","year":"2024","journal-title":"Genetics"},{"key":"2025093010290058600_bib39","doi-asserted-by":"publisher","first-page":"861","DOI":"10.21105\/joss.00861","article-title":"UMAP: Uniform Manifold Approximation and Projection","volume":"3","author":"McInnes","year":"2018","journal-title":"J Open Source Software"},{"key":"2025093010290058600_bib40","volume-title":"Foundations of Machine Learning","author":"Mohri","year":"2018"},{"key":"2025093010290058600_bib41","doi-asserted-by":"publisher","first-page":"150","DOI":"10.3390\/info10040150","article-title":"Text classification algorithms: a survey","volume":"10","author":"Kowsari","year":"2019","journal-title":"Information"},{"key":"2025093010290058600_bib42","doi-asserted-by":"publisher","first-page":"785","DOI":"10.1145\/2939672.2939785","article-title":"XGBoost: a scalable tree boosting system","volume-title":"Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Chen","year":"2016"},{"key":"2025093010290058600_bib43","doi-asserted-by":"publisher","first-page":"177","DOI":"10.1007\/978-3-7908-2604-3_16","article-title":"Large-scale machine learning with stochastic gradient descent","volume-title":"Proceedings of COMPSTAT'2010","author":"Bottou","year":"2010"},{"key":"2025093010290058600_bib44","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2303.08774","article-title":"GPT-4 technical report","author":"Open","year":"2024","journal-title":"arXiv"},{"key":"2025093010290058600_bib45","doi-asserted-by":"publisher","first-page":"1429","DOI":"10.1038\/s41588-019-0500-1","article-title":"Gene ontology causal activity modeling (GO-CAM) moves beyond GO annotations to structured descriptions of biological functions and systems","volume":"51","author":"Thomas","year":"2019","journal-title":"Nat Genet"},{"key":"2025093010290058600_bib46","doi-asserted-by":"publisher","first-page":"155","DOI":"10.1186\/1471-2105-15-155","article-title":"A method for increasing expressivity of Gene Ontology annotations using a compositional approach","volume":"15","author":"Huntley","year":"2014","journal-title":"BMC Bioinf"},{"key":"2025093010290058600_bib47","doi-asserted-by":"publisher","first-page":"D1515","DOI":"10.1093\/nar\/gkab1025","article-title":"ECO: the evidence and conclusion ontology, an update for 2022","volume":"50","author":"Nadendla","year":"2022","journal-title":"Nucleic Acids Res"},{"key":"2025093010290058600_bib48","doi-asserted-by":"publisher","first-page":"28167","DOI":"10.1016\/j.jocm.2018.07.002","article-title":"Is your dataset big enough? Sample size requirements when using artificial neural networks for discrete choice analysis","volume":"28","author":"Alwosheel","year":"2018","journal-title":"J Choice Model"},{"key":"2025093010290058600_bib49","doi-asserted-by":"publisher","first-page":"W540","DOI":"10.1093\/nar\/gkae235","article-title":"PubTator 3.0: an AI-powered literature resource for unlocking biomedical knowledge","volume":"52","author":"Wei","year":"2024","journal-title":"Nucleic Acids Res"},{"key":"2025093010290058600_bib50","doi-asserted-by":"publisher","first-page":"bay128","DOI":"10.1093\/database\/bay128","article-title":"iTextMine: integrated text-mining system for large-scale knowledge extraction from the literature","volume":"2018","author":"Ren","year":"2018","journal-title":"Database"}],"container-title":["Database"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baaf063\/64443368\/baaf063.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baaf063\/64443368\/baaf063.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T14:29:07Z","timestamp":1759242547000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/database\/article\/doi\/10.1093\/database\/baaf063\/8268905"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025]]},"references-count":50,"URL":"https:\/\/doi.org\/10.1093\/database\/baaf063","relation":{},"ISSN":["1758-0463"],"issn-type":[{"value":"1758-0463","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2025]]},"published":{"date-parts":[[2025]]},"article-number":"baaf063"}}