{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T16:13:22Z","timestamp":1778861602652,"version":"3.51.4"},"reference-count":12,"publisher":"Springer Science and Business Media LLC","issue":"S1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2005,5]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>The BioCreative text mining evaluation investigated the application of text mining methods to the task of automatically extracting information from text in biomedical research articles. We participated in Task 2 of the evaluation. For this task, we built a system to automatically annotate a given protein with codes from the Gene Ontology (GO) using the text of an article from the biomedical literature as evidence.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Methods<\/jats:title>\n            <jats:p>Our system relies on simple statistical analyses of the full text article provided. We learn <jats:italic>n<\/jats:italic>-gram models for each GO code using statistical methods and use these models to hypothesize annotations. We also learn a set of Na\u00efve Bayes models that identify textual clues of possible connections between the given protein and a hypothesized annotation. These models are used to filter and rank the predictions of the <jats:italic>n<\/jats:italic>-gram models.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>We report experiments evaluating the utility of various components of our system on a set of data held out during development, and experiments evaluating the utility of external data sources that we used to learn our models. Finally, we report our evaluation results from the BioCreative organizers.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion<\/jats:title>\n            <jats:p>We observe that, on the test data, our system performs quite well relative to the other systems submitted to the evaluation. From other experiments on the held-out data, we observe that (i) the Na\u00efve Bayes models were effective in filtering and ranking the initially hypothesized annotations, and (ii) our learned models were significantly more accurate when external data sources were used during learning.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-6-s1-s18","type":"journal-article","created":{"date-parts":[[2005,5,24]],"date-time":"2005-05-24T18:13:44Z","timestamp":1116958424000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":29,"title":["Learning Statistical Models for Annotating Proteins with Function Information using Biomedical Text"],"prefix":"10.1186","volume":"6","author":[{"given":"Soumya","family":"Ray","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mark","family":"Craven","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2005,5,24]]},"reference":[{"key":"653_CR1","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1038\/75556","volume":"25","author":"The Gene Ontology Consortium","year":"2000","unstructured":"The Gene Ontology Consortium: Gene Ontology: tool for the unification of biology. Nature Genetics 2000, 25: 25\u201329. 10.1038\/75556","journal-title":"Nature Genetics"},{"issue":"3","key":"653_CR2","doi-asserted-by":"publisher","first-page":"127","DOI":"10.1108\/eb046814","volume":"14","author":"MF Porter","year":"1980","unstructured":"Porter MF: An Algorithm for Suffix Stripping. Program 1980, 14(3):127\u2013130.","journal-title":"Program"},{"key":"653_CR3","volume-title":"Unified Medical Language System","author":"National Library of Medicine","year":"1999","unstructured":"National Library of Medicine: Unified Medical Language System.1999. [http:\/\/www.nlm.nih.gov\/research\/umls\/umlsmain.html]"},{"key":"653_CR4","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1093\/nar\/25.1.31","volume":"25","author":"A Bairoch","year":"1997","unstructured":"Bairoch A, Apweiler R: The SWISS-PROT Protein Sequence Data Bank and its Supplement TrEMBL. Nucleic Acids Research 1997, 25: 31\u201336. 10.1093\/nar\/25.1.31","journal-title":"Nucleic Acids Research"},{"key":"653_CR5","doi-asserted-by":"publisher","first-page":"464","DOI":"10.1006\/geno.2002.6748","volume":"79","author":"HM Wain","year":"2002","unstructured":"Wain HM, Bruford EA, Lovering RC, Lush MJ, Wright MW, Povey S: Guidelines for Human Gene Nomenclature. Genomics 2002, 79: 464\u2013470. 10.1006\/geno.2002.6748","journal-title":"Genomics"},{"key":"653_CR6","volume-title":"Saccharomyces Genome Database","author":"K Dolinski","year":"2003","unstructured":"Dolinski K, Balakrishnan R, Christie KR, Costanzo MC, Dwight SS, Engel SR, Fisk DG, Hirschman JE, Hong EL, Issel-Tarver L, Sethuraman A, Theesfeld CL, Binkley G, Lane C, Schroeder M, Dong S, Weng S, Andrada R, Botstein D, Cherry JM: Saccharomyces Genome Database.2003. [http:\/\/yeastgenome.org]"},{"key":"653_CR7","doi-asserted-by":"publisher","first-page":"172","DOI":"10.1093\/nar\/gkg094","volume":"31","author":"The FlyBase Consortium","year":"2003","unstructured":"The FlyBase Consortium: The FlyBase database of the Drosophila genome projects and community literature. Nucleic Acids Research 2003, 31: 172\u2013175. [Http:\/\/flybase.org\/] 10.1093\/nar\/gkg094","journal-title":"Nucleic Acids Research"},{"key":"653_CR8","doi-asserted-by":"publisher","first-page":"D411","DOI":"10.1093\/nar\/gkh066","volume":"32","author":"TW Harris","year":"2004","unstructured":"Harris TW, Chen N, Cunningham F, Tello-Ruiz M, Antoshechkin I, Bastiani C, Bieri T, Blasiar D, Bradnam K, Chan J, Chen CK, Chen WJ, Eimear Kenny PD, Kishore R, Lawson D, aymond Lee R, Muller HM, Philip Ozersky CN, Petcherski A, Rogers A, Sabo A, Schwarz EM, Qinghua Wang KVA, Durbin R, Spieth J, Sternberg PW, Stein LD: WormBase: a multi-species resource for nematode biology and genomics. Nucleic Acids Research 2004, 32: D411-D417. 10.1093\/nar\/gkh066","journal-title":"Nucleic Acids Research"},{"key":"653_CR9","doi-asserted-by":"publisher","first-page":"102","DOI":"10.1093\/nar\/29.1.102","volume":"29","author":"E Huala","year":"2001","unstructured":"Huala E, Dickerman A, Garcia-Hernandez M, Weems D, Reiser L, LaFond F, Hanley D, Kiphart D, Zhuang J, Huang W, Mueller L, Bhattacharyya D, Bhaya D, Sobral B, Beavis B, Somerville C, Rhee S: The Arabidopsis Information Resource (TAIR): A comprehensive database and web-based information retrieval, analysis, and visualization system for a model plant. Nucleic Acids Research 2001, 29: 102\u2013105. 10.1093\/nar\/29.1.102","journal-title":"Nucleic Acids Research"},{"key":"653_CR10","doi-asserted-by":"publisher","first-page":"D262","DOI":"10.1093\/nar\/gkh021","volume":"32","author":"E Camon","year":"2004","unstructured":"Camon E, Magrane M, Barrell D, Lee V, Dimmer E, Maslen J, Binns D, Harte N, Lopez R, Apweiler R: The Gene Ontology Annotation (GOA) Database: sharing knowledge in Uniprot with Gene Ontology. Nucleic Acids Research 2004, 32: D262-D266. 10.1093\/nar\/gkh021","journal-title":"Nucleic Acids Research"},{"key":"653_CR11","volume-title":"Machine Learning","author":"TM Mitchell","year":"1997","unstructured":"Mitchell TM: Machine Learning. New York: McGraw-Hill; 1997."},{"issue":"1\u20132","key":"653_CR12","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1016\/S0004-3702(96)00034-3","volume":"89","author":"TG Dietterich","year":"1997","unstructured":"Dietterich TG, Lathrop RH, Lozano-Perez T: Solving the Multiple Instance Problem with Axis-Parallel Rectangles. Artificial Intelligence 1997, 89(1\u20132):31\u201371. 10.1016\/S0004-3702(96)00034-3","journal-title":"Artificial Intelligence"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-6-S1-S18.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T01:32:08Z","timestamp":1630459928000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-6-S1-S18"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,5]]},"references-count":12,"journal-issue":{"issue":"S1","published-print":{"date-parts":[[2005,5]]}},"alternative-id":["653"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-6-s1-s18","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2005,5]]},"assertion":[{"value":"24 May 2005","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"S18"}}