{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,2]],"date-time":"2025-12-02T15:18:47Z","timestamp":1764688727546},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2009,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>The classification of protein domains in the CATH resource is primarily based on structural comparisons, sequence similarity and manual analysis. One of the main bottlenecks in the processing of new entries is the evaluation of 'borderline' cases by human curators with reference to the literature, and better tools for helping both expert and non-expert users quickly identify relevant functional information from text are urgently needed. A text based method for protein classification is presented, which complements the existing sequence and structure-based approaches, especially in cases exhibiting low similarity to existing members and requiring manual intervention. The method is based on the assumption that textual similarity between sets of documents relating to proteins reflects biological function similarities and can be exploited to make classification decisions.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>An optimal strategy for the text comparisons was identified by using an established gold standard enzyme dataset. Filtering of the abstracts using a machine learning approach to discriminate sentences containing functional, structural and classification information that are relevant to the protein classification task improved performance. Testing this classification scheme on a dataset of 'borderline' protein domains that lack significant sequence or structure similarity to classified proteins showed that although, as expected, the structural similarity classifiers perform better on average, there is a significant benefit in incorporating text similarity in logistic regression models, indicating significant orthogonality in this additional information. Coverage was significantly increased especially at low error rates, which is important for routine classification tasks: 15.3% for the combined structure and text classifier compared to 10% for the structural classifier alone, at 10<jats:sup>-3<\/jats:sup> error rate. Finally when only the highest scoring predictions were used to infer classification, an extra 4.2% of correct decisions were made by the combined classifier.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion<\/jats:title>\n            <jats:p>We have described a simple text based method to classify protein domains that demonstrates an improvement over existing methods. The method is unique in incorporating structural and text based classifiers directly and is particularly useful in cases where inconclusive evidence from sequence or structure similarity requires laborious manual classification.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-10-129","type":"journal-article","created":{"date-parts":[[2009,5,5]],"date-time":"2009-05-05T18:15:01Z","timestamp":1241547301000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":12,"title":["Improving classification in protein structure databases using text mining"],"prefix":"10.1186","volume":"10","author":[{"given":"Antonis","family":"Koussounadis","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Oliver C","family":"Redfern","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David T","family":"Jones","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2009,5,5]]},"reference":[{"key":"2859_CR1","doi-asserted-by":"publisher","first-page":"1093","DOI":"10.1016\/S0969-2126(97)00260-8","volume":"5","author":"CA Orengo","year":"1997","unstructured":"Orengo CA, Michie AD, Jones S, Jones DT, Swindells MB, Thornton JM: CATH \u2013 A Hierarchic Classification of Protein Domain Structures. Structure 1997, 5: 1093\u20131108. 10.1016\/S0969-2126(97)00260-8","journal-title":"Structure"},{"key":"2859_CR2","first-page":"536","volume":"247","author":"AG Murzin","year":"1995","unstructured":"Murzin AG, Brenner SE, Hubbard T, Chothia C: SCOP: a structural classification of proteins database for the investigation of sequences and structures. Journal of Molecular Biology 1995, 247: 536\u2013540.","journal-title":"Journal of Molecular Biology"},{"key":"2859_CR3","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4757-2440-0","volume-title":"The Nature of Statistical Learning Theory","author":"VN Vapnik","year":"1995","unstructured":"Vapnik VN: The Nature of Statistical Learning Theory. New York: Springer; 1995."},{"key":"2859_CR4","first-page":"137","volume-title":"Proceedings of 10th European Conference on Machine Learning","author":"T Joachims","year":"1998","unstructured":"Joachims T: Text categorization with support vector machines: learning many relevant features. In Proceedings of 10th European Conference on Machine Learning. Springer-Verlag, Heidelberg; 1998:137\u2013142."},{"key":"2859_CR5","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1186\/1471-2105-4-11","volume":"4","author":"I Donaldson","year":"2003","unstructured":"Donaldson I, Martin J, de Bruijn B, Walting C, Lay V, Tuekam B, et al.: PreBIND and Textomy \u2013 mining the biomedical literature for protein-protein interactions using a support vector machine. BMC Bioinformatics 2003, 4: 11. 10.1186\/1471-2105-4-11","journal-title":"BMC Bioinformatics"},{"key":"2859_CR6","first-page":"374","volume-title":"Pac Symp Biocomput","author":"BJ Stapley","year":"2002","unstructured":"Stapley BJ, Kelley LA, Sternberg MJ: Predicting the sub-cellular location of proteins from text using support vector machines. Pac Symp Biocomput 2002, 374\u2013385."},{"issue":"Suppl 1","key":"2859_CR7","doi-asserted-by":"publisher","first-page":"S22","DOI":"10.1186\/1471-2105-6-S1-S22","volume":"6","author":"SB Rice","year":"2005","unstructured":"Rice SB, Nenadic G, Stapley BI: Mining protein function from text using term-based support vector machine. BMC Bioinformatics 2005, 6(Suppl 1):S22. 10.1186\/1471-2105-6-S1-S22","journal-title":"BMC Bioinformatics"},{"key":"2859_CR8","doi-asserted-by":"publisher","first-page":"370","DOI":"10.1186\/1471-2105-7-370","volume":"7","author":"D Chen","year":"2006","unstructured":"Chen D, Muller H-M, Sternberg PW: Automatic document classification of biological literature. BMC Bioinformatics 2006, 7: 370. 10.1186\/1471-2105-7-370","journal-title":"BMC Bioinformatics"},{"key":"2859_CR9","doi-asserted-by":"publisher","first-page":"445","DOI":"10.1016\/S0092-8674(04)00117-5","volume":"116","author":"M Miaczynska","year":"2004","unstructured":"Miaczynska M, Christoforidis S, Giner A, Shevchenko A, Uttenweiler-Joseph S, Habermann B, Wilm M, Parton RG, Zerial M: APPL proteins link Rab5 to nuclear signal transduction via an endosomal compartment. Cell 2004, 116: 445\u2013456. 10.1016\/S0092-8674(04)00117-5","journal-title":"Cell"},{"key":"2859_CR10","doi-asserted-by":"publisher","first-page":"125","DOI":"10.1093\/bioinformatics\/16.2.125","volume":"16","author":"RM MacCallum","year":"2000","unstructured":"MacCallum RM, Kelley LA, Sternberg MJE: SAWTED: Structure assignment with text description \u2013 Enhanced detection of remote homologues with automated SWISS-PROT annotation comparisons. Bioinformatics 2000, 16: 125\u2013129. 10.1093\/bioinformatics\/16.2.125","journal-title":"Bioinformatics"},{"key":"2859_CR11","doi-asserted-by":"publisher","first-page":"466","DOI":"10.1186\/1471-2105-7-466","volume":"7","author":"CR Bradshaw","year":"2006","unstructured":"Bradshaw CR, Surendranath V, Habermann B: ProFAT: a web-based tool for the functional annotation for protein sequences. BMC Bioinformatics 2006, 7: 466. 10.1186\/1471-2105-7-466","journal-title":"BMC Bioinformatics"},{"issue":"Suppl 1","key":"2859_CR12","doi-asserted-by":"publisher","first-page":"S16","DOI":"10.1186\/1471-2105-6-S1-S16","volume":"6","author":"C Blaschke","year":"2005","unstructured":"Blaschke C, Leon EA, Krallinger M, Valencia A: Evaluation of BioCreAtIvE assessment of task 2. BMC Bioinformatics 2005, 6(Suppl 1):S16. 10.1186\/1471-2105-6-S1-S16","journal-title":"BMC Bioinformatics"},{"issue":"Suppl 1","key":"2859_CR13","doi-asserted-by":"publisher","first-page":"S21","DOI":"10.1186\/1471-2105-6-S1-S21","volume":"6","author":"FM Couto","year":"2005","unstructured":"Couto FM, Silva MJ, Coutinho PM: Finding genomic ontology terms in text using evidence content. BMC Bioinformatics 2005, 6(Suppl 1):S21. 10.1186\/1471-2105-6-S1-S21","journal-title":"BMC Bioinformatics"},{"key":"2859_CR14","doi-asserted-by":"publisher","first-page":"19","DOI":"10.1186\/1747-5333-1-19","volume":"1","author":"FM Couto","year":"2006","unstructured":"Couto FM, Silva MJ, Lee V, Dimmer E, Camon E, Apweiler R, Kirsch H, Rebholz-Schuhmann D: GOAnnotator: linking protein GO annotations to evidence text. Journal of Biomedical Discovery and Collaboration 2006, 1: 19. 10.1186\/1747-5333-1-19","journal-title":"Journal of Biomedical Discovery and Collaboration"},{"key":"2859_CR15","doi-asserted-by":"publisher","first-page":"658","DOI":"10.1093\/bioinformatics\/bti783","volume":"22","author":"P Ruch","year":"2006","unstructured":"Ruch P: Automatic assignment of biomedical categories: toward a generic approach. Bioinformatics 2006, 22: 658\u2013664. 10.1093\/bioinformatics\/bti783","journal-title":"Bioinformatics"},{"key":"2859_CR16","first-page":"342746","volume-title":"EURASIP Journal on Bioinformatics and Systems Biology","author":"S Gaudan","year":"2008","unstructured":"Gaudan S, Jimeno Yepes A, Lee V, Rebholz-Schuhmann D: Combining evidence, specificity, and proximity towards the normalization of Gene Ontology terms in text. EURASIP Journal on Bioinformatics and Systems Biology 2008, 342746."},{"issue":"I","key":"2859_CR17","doi-asserted-by":"publisher","first-page":"R8","DOI":"10.1186\/gb-2006-7-1-r8","volume":"7","author":"SD Brown","year":"2006","unstructured":"Brown SD, Gerlt JA, Seffernick JL, Babbitt PC: A gold standard set of mechanistically diverse enzyme superfamilies. Genome Biology 2006, 7(I):R8. 10.1186\/gb-2006-7-1-r8","journal-title":"Genome Biology"},{"key":"2859_CR18","doi-asserted-by":"publisher","first-page":"C154","DOI":"10.1093\/nar\/gki070","volume":"33","author":"A Bairoch","year":"2005","unstructured":"Bairoch A, Apweiler A, Wu CH, Barker WC, Boeckman B, Ferro S, et al.: The Universal Protein Resource (UniProt). Nucleic Acids Research 2005, 33: C154-D159. 10.1093\/nar\/gki070","journal-title":"Nucleic Acids Research"},{"key":"2859_CR19","doi-asserted-by":"publisher","first-page":"235","DOI":"10.1093\/nar\/28.1.235","volume":"28","author":"HM Berman","year":"2000","unstructured":"Berman HM, Westbrook J, Feng Z, Gilliland G, Bhat TN, Weissig H, Shindyalov IN, Bourne PE: The Protein Data Bank. Nucleic Acids Research 2000, 28: 235\u2013242. 10.1093\/nar\/28.1.235","journal-title":"Nucleic Acids Research"},{"issue":"11","key":"2859_CR20","doi-asserted-by":"publisher","first-page":"e232","DOI":"10.1371\/journal.pcbi.0030232","volume":"3","author":"OC Redfern","year":"2007","unstructured":"Redfern OC, Harrison A, Dallman T, Pearl FM, Orengo CA: CATHEDRAL: a fast and effective algorithm to predict folds and domain boundaries from multidomain protein structures. PLoS Computational Biology 2007, 3(11):e232. 10.1371\/journal.pcbi.0030232","journal-title":"PLoS Computational Biology"},{"key":"2859_CR21","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/0022-2836(89)90084-3","volume":"208","author":"WR Taylor","year":"1989","unstructured":"Taylor WR, Orengo CA: Protein structure alignment. Journal of Molecular Biology 1989, 208: 1\u201322. 10.1016\/0022-2836(89)90084-3","journal-title":"Journal of Molecular Biology"},{"issue":"18","key":"2859_CR22","doi-asserted-by":"publisher","first-page":"2353","DOI":"10.1093\/bioinformatics\/btm355","volume":"23","author":"AJ Reid","year":"2007","unstructured":"Reid AJ, Yeats C, Orengo CA: Methods of remote homology detection can be combined to increase coverage by 10% in the midnight zone. Bioinformatic 2007, 23(18):2353\u201360. 10.1093\/bioinformatics\/btm355","journal-title":"Bioinformatic"},{"key":"2859_CR23","doi-asserted-by":"publisher","first-page":"1201","DOI":"10.1006\/jmbi.1998.2221","volume":"284","author":"J Park","year":"1998","unstructured":"Park J, Karplus K, Barrett C, Hughey R, Haussler D, Hubbard T, Clothia C: Sequence comparisons using multiple sequences detect three times as many remote homologues as pairwise methods. Journal of Molecular Biology 1998, 284: 1201\u20131210. 10.1006\/jmbi.1998.2221","journal-title":"Journal of Molecular Biology"},{"key":"2859_CR24","doi-asserted-by":"publisher","first-page":"739","DOI":"10.1093\/protein\/11.9.739","volume":"11","author":"IN Shindyalov","year":"1998","unstructured":"Shindyalov IN, Bourne PE: Protein structure alignment by incremental combinatorial extension (CE) of the optimal path. Protein Engineering 1998, 11: 739\u2013747. 10.1093\/protein\/11.9.739","journal-title":"Protein Engineering"},{"key":"2859_CR25","doi-asserted-by":"publisher","first-page":"595","DOI":"10.1126\/science.273.5275.595","volume":"273","author":"L Holm","year":"1996","unstructured":"Holm L, Sander C: Mapping the protein universe. Science 1996, 273: 595\u2013603. 10.1126\/science.273.5275.595","journal-title":"Science"},{"key":"2859_CR26","doi-asserted-by":"publisher","first-page":"2256","DOI":"10.1107\/S0907444904026460","volume":"D60","author":"E Krissinel","year":"2004","unstructured":"Krissinel E, Henrick K: Secondary-structure matching (SSM), a new tool for fast protein structure alignment in three dimensions. Acta Crystallogr D Biol Crystallogr 2004, D60: 2256\u20132268. 10.1107\/S0907444904026460","journal-title":"Acta Crystallogr D Biol Crystallogr"},{"key":"2859_CR27","unstructured":"The PSIPRED Protein Structure Prediction Server[http:\/\/bioinf.cs.ucl.ac.uk\/psipred]"},{"key":"2859_CR28","unstructured":"The CATHEDRAL server[http:\/\/www.cathdb.info\/cgi-bin\/CathedralServer.pl]"},{"key":"2859_CR29","doi-asserted-by":"publisher","first-page":"209","DOI":"10.1016\/0010-4825(95)00055-0","volume":"26","author":"WJ Wilbur","year":"1996","unstructured":"Wilbur WJ, Yang YM: An analysis of statistical term strength and its use in the indexing and retrieval of molecular biology texts. Computers in Biology and Medicine 1996, 26: 209\u2013222. 10.1016\/0010-4825(95)00055-0","journal-title":"Computers in Biology and Medicine"},{"key":"2859_CR30","unstructured":"Lucene[http:\/\/lucene.apache.org\/]"},{"key":"2859_CR31","unstructured":"The R Project for Statistical Computing[http:\/\/www.r-project.org]"},{"key":"2859_CR32","doi-asserted-by":"publisher","first-page":"3940","DOI":"10.1093\/bioinformatics\/bti623","volume":"21","author":"T Sing","year":"2005","unstructured":"Sing T, Sander O, Beerenwinkel N, Lengauer T: ROCR: visualizing classifier performance in R. Bioinformatics 2005, 21: 3940\u20133941. 10.1093\/bioinformatics\/bti623","journal-title":"Bioinformatics"},{"key":"2859_CR33","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4757-3462-1","volume-title":"Regression modeling strategies with applications to linear models, logistic regression, and survival analysis","author":"FE Harrell Jr","year":"2001","unstructured":"Harrell FE Jr: Regression modeling strategies with applications to linear models, logistic regression, and survival analysis. New York: Springer; 2001."},{"issue":"1","key":"2859_CR34","doi-asserted-by":"publisher","first-page":"365","DOI":"10.1093\/nar\/gkg095","volume":"31","author":"B Boeckmann","year":"2003","unstructured":"Boeckmann B, Bairoch A, Apweiler R, Blatter MC, Estreicher A, Gasteiger E, Martin MJ, Michoud K, O'Donovan C, Phan I, Pilbout S, Schneider M: The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003. Nucleic Acids Research 2003, 31(1):365\u201370. 10.1093\/nar\/gkg095","journal-title":"Nucleic Acids Research"},{"key":"2859_CR35","first-page":"41","volume-title":"Advances in Kernel Methods \u2013 Support Vector Learning","author":"T Joachims","year":"1999","unstructured":"Joachims T: Making large-Scale SVM Learning Practical. In Advances in Kernel Methods \u2013 Support Vector Learning. Edited by: Sch\u00f6lkopf B, Burges CJC, Smola AJ. Cambridge, MA: MIT Press; 1999:41\u201356."},{"key":"2859_CR36","doi-asserted-by":"publisher","first-page":"423","DOI":"10.1186\/1471-2105-8-423","volume":"8","author":"J Lin","year":"2007","unstructured":"Lin J, Wilbur WJ: PubMed related articles: a probabilistic topic-based model for content similarity. BMC Bioinformatics 2007, 8: 423. 10.1186\/1471-2105-8-423","journal-title":"BMC Bioinformatics"},{"key":"2859_CR37","doi-asserted-by":"publisher","first-page":"130","DOI":"10.1108\/eb046814","volume":"14","author":"MF Porter","year":"1980","unstructured":"Porter MF: An algorithm for suffix stripping. Program 1980, 14: 130\u2013137.","journal-title":"Program"},{"key":"2859_CR38","first-page":"191","volume-title":"Viewing morphology as an inference process","author":"R Krovetz","year":"1993","unstructured":"Krovetz R: Viewing morphology as an inference process. ACM, Pittsburgh; 1993:191\u2013203."}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-10-129.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,8,31]],"date-time":"2021-08-31T21:35:43Z","timestamp":1630445743000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-10-129"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,5,5]]},"references-count":38,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2009,12]]}},"alternative-id":["2859"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-10-129","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2009,5,5]]},"assertion":[{"value":"27 November 2008","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 May 2009","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 May 2009","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"129"}}