{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,5]],"date-time":"2026-02-05T22:38:42Z","timestamp":1770331122641,"version":"3.49.0"},"reference-count":42,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2009,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec><jats:title>Background<\/jats:title><jats:p>We previously developed EFICAz, an enzyme function inference approach that combines predictions from non-completely overlapping component methods. Two of the four components in the original EFICAz are based on the detection of functionally discriminating residues (FDRs). FDRs distinguish between member of an enzyme family that are homofunctional (classified under the EC number of interest) or heterofunctional (annotated with another EC number or lacking enzymatic activity). Each of the two FDR-based components is associated to one of two specific kinds of enzyme families. EFICAz exhibits high precision performance, except when the maximal test to training sequence identity (MTTSI) is lower than 30%. To improve EFICAz's performance in this regime, we: i) increased the number of predictive components and ii) took advantage of consensual information from the different components to make the final EC number assignment.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>We have developed two new EFICAz components, analogs to the two FDR-based components, where the discrimination between homo and heterofunctional members is based on the evaluation, via Support Vector Machine models, of all the aligned positions between the query sequence and the multiple sequence alignments associated to the enzyme families. Benchmark results indicate that: i) the new SVM-based components outperform their FDR-based counterparts, and ii) both SVM-based and FDR-based components generate unique predictions. We developed classification tree models to optimally combine the results from the six EFICAz components into a final EC number prediction. The new implementation of our approach, EFICAz<jats:sup>2<\/jats:sup>, exhibits a highly improved prediction precision at MTTSI &lt; 30% compared to the original EFICAz, with only a slight decrease in prediction recall. A comparative analysis of enzyme function annotation of the human proteome by EFICAz<jats:sup>2<\/jats:sup>and KEGG shows that: i) when both sources make EC number assignments for the same protein sequence, the assignments tend to be consistent and ii) EFICAz<jats:sup>2<\/jats:sup>generates considerably more unique assignments than KEGG.<\/jats:p><\/jats:sec><jats:sec><jats:title>Conclusion<\/jats:title><jats:p>Performance benchmarks and the comparison with KEGG demonstrate that EFICAz<jats:sup>2<\/jats:sup>is a powerful and precise tool for enzyme function annotation, with multiple applications in genome analysis and metabolic pathway reconstruction. The EFICAz<jats:sup>2<\/jats:sup>web service is available at:<jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"http:\/\/cssb.biology.gatech.edu\/skolnick\/webservice\/EFICAz2\/index.html\" ext-link-type=\"uri\">http:\/\/cssb.biology.gatech.edu\/skolnick\/webservice\/EFICAz2\/index.html<\/jats:ext-link><\/jats:p><\/jats:sec>","DOI":"10.1186\/1471-2105-10-107","type":"journal-article","created":{"date-parts":[[2009,4,14]],"date-time":"2009-04-14T06:13:03Z","timestamp":1239689583000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":58,"title":["EFICAz2: enzyme function inference by a combined approach enhanced by machine learning"],"prefix":"10.1186","volume":"10","author":[{"given":"Adrian K","family":"Arakaki","sequence":"first","affiliation":[]},{"given":"Ying","family":"Huang","sequence":"additional","affiliation":[]},{"given":"Jeffrey","family":"Skolnick","sequence":"additional","affiliation":[]}],"member":"297","published-online":{"date-parts":[[2009,4,13]]},"reference":[{"key":"2837_CR1","doi-asserted-by":"publisher","first-page":"315","DOI":"10.1186\/1471-2164-7-315","volume":"7","author":"AK Arakaki","year":"2006","unstructured":"Arakaki AK, Tian W, Skolnick J: High precision multi-genome scale reannotation of enzyme function by EFICAz. BMC Genomics 2006, 7: 315. 10.1186\/1471-2164-7-315","journal-title":"BMC Genomics"},{"issue":"4","key":"2837_CR2","doi-asserted-by":"publisher","first-page":"745","DOI":"10.1016\/j.jmb.2005.04.027","volume":"349","author":"S Freilich","year":"2005","unstructured":"Freilich S, Spriggs RV, George RA, Al-Lazikani B, Swindells M, Thornton JM: The complement of enzymatic sets in different species. J Mol Biol 2005, 349(4):745\u2013763. 10.1016\/j.jmb.2005.04.027","journal-title":"J Mol Biol"},{"key":"2837_CR3","volume-title":"Enzyme nomenclature 1992: recommendations of the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology on the nomenclature and classification of enzymes","author":"EC Webb","year":"1992","unstructured":"Webb EC: Enzyme nomenclature 1992: recommendations of the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology on the nomenclature and classification of enzymes. San Diego: Published for the International Union of Biochemistry and Molecular Biology by Academic Press; 1992."},{"issue":"5","key":"2837_CR4","doi-asserted-by":"publisher","first-page":"492","DOI":"10.1016\/j.cbpa.2006.08.012","volume":"10","author":"ME Glasner","year":"2006","unstructured":"Glasner ME, Gerlt JA, Babbitt PC: Evolution of enzyme superfamilies. Curr Opin Chem Biol 2006, 10(5):492\u2013497. 10.1016\/j.cbpa.2006.08.012","journal-title":"Curr Opin Chem Biol"},{"issue":"1","key":"2837_CR5","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1016\/j.pt.2008.08.012","volume":"25","author":"H Ginsburg","year":"2008","unstructured":"Ginsburg H: Caveat emptor: limitations of the automated reconstruction of metabolic pathways in Plasmodium. Trends Parasitol 2008, 25(1):37\u201343. 10.1016\/j.pt.2008.08.012","journal-title":"Trends Parasitol"},{"key":"2837_CR6","doi-asserted-by":"publisher","first-page":"12","DOI":"10.1186\/1471-2180-5-8","volume":"5","author":"SA Becker","year":"2005","unstructured":"Becker SA, Palsson BO: Genome-scale reconstruction of the metabolic network in Staphylococcus aureus N315: an initial draft to the two-dimensional annotation. BMC Microbiol 2005, 5: 12. 10.1186\/1471-2180-5-8","journal-title":"BMC Microbiol"},{"issue":"13","key":"2837_CR7","doi-asserted-by":"publisher","first-page":"1616","DOI":"10.1093\/bioinformatics\/btm150","volume":"23","author":"R Guimera","year":"2007","unstructured":"Guimera R, Sales-Pardo M, Amaral LAN: A network-based method for target selection in metabolic networks. Bioinformatics 2007, 23(13):1616\u20131622. 10.1093\/bioinformatics\/btm150","journal-title":"Bioinformatics"},{"issue":"11","key":"2837_CR8","doi-asserted-by":"publisher","first-page":"548","DOI":"10.1016\/j.pt.2007.08.013","volume":"23","author":"JW Pinney","year":"2007","unstructured":"Pinney JW, Papp B, Hyland C, Warnbua L, Westhead DR, McConkey GA: Metabolic reconstruction and analysis for parasite genomes. Trends Parasitol 2007, 23(11):548\u2013554. 10.1016\/j.pt.2007.08.013","journal-title":"Trends Parasitol"},{"issue":"1","key":"2837_CR9","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1186\/1476-4598-7-57","volume":"7","author":"A Arakaki","year":"2008","unstructured":"Arakaki A, Mezencev R, Bowen N, Huang Y, McDonald J, Skolnick J: Identification of metabolites with anticancer properties by Computational Metabolomics. Mol Cancer 2008, 7(1):57. 10.1186\/1476-4598-7-57","journal-title":"Mol Cancer"},{"issue":"9\u201310","key":"2837_CR10","doi-asserted-by":"publisher","first-page":"402","DOI":"10.1016\/j.drudis.2008.02.002","volume":"13","author":"H Ma","year":"2008","unstructured":"Ma H, Goryanin I: Human metabolic network reconstruction and its impact on drug discovery and development. Drug Discov Today 2008, 13(9\u201310):402\u2013408. 10.1016\/j.drudis.2008.02.002","journal-title":"Drug Discov Today"},{"key":"2837_CR11","volume-title":"Genome Biol","author":"CA Ouzounis","year":"2002","unstructured":"Ouzounis CA, Karp PD: The past, present and future of genome-wide re-annotation. Genome Biol 2002., 3(2): COMMENT2001."},{"issue":"10","key":"2837_CR12","doi-asserted-by":"publisher","first-page":"e1000160","DOI":"10.1371\/journal.pcbi.1000160","volume":"4","author":"M Punta","year":"2008","unstructured":"Punta M, Ofran Y: The rough guide to in silico function prediction, or how to use sequence and structure information to predict protein function. PLoS Comput Biol 2008, 4(10):e1000160. 10.1371\/journal.pcbi.1000160","journal-title":"PLoS Comput Biol"},{"issue":"5","key":"2837_CR13","doi-asserted-by":"publisher","first-page":"REVIEWS0005","DOI":"10.1186\/gb-2000-1-5-reviews0005","volume":"1","author":"JA Gerlt","year":"2000","unstructured":"Gerlt JA, Babbitt PC: Can sequence determine function? Genome Biol 2000, 1(5):REVIEWS0005. 10.1186\/gb-2000-1-5-reviews0005","journal-title":"Genome Biol"},{"issue":"4","key":"2837_CR14","doi-asserted-by":"publisher","first-page":"863","DOI":"10.1016\/j.jmb.2003.08.057","volume":"333","author":"W Tian","year":"2003","unstructured":"Tian W, Skolnick J: How well is enzyme function conserved as a function of pairwise sequence identity? J Mol Biol 2003, 333(4):863\u2013882. 10.1016\/j.jmb.2003.08.057","journal-title":"J Mol Biol"},{"issue":"4","key":"2837_CR15","doi-asserted-by":"publisher","first-page":"886","DOI":"10.1046\/j.1365-2958.1999.01380.x","volume":"32","author":"NC Kyrpides","year":"1999","unstructured":"Kyrpides NC, Ouzounis CA: Whole-genome sequence annotation: 'Going wrong with confidence'. Mol Microbiol 1999, 32(4):886\u2013887. 10.1046\/j.1365-2958.1999.01380.x","journal-title":"Mol Microbiol"},{"issue":"10","key":"2837_CR16","doi-asserted-by":"publisher","first-page":"1632","DOI":"10.1101\/gr. 183801","volume":"11","author":"H Hegyi","year":"2001","unstructured":"Hegyi H, Gerstein M: Annotation transfer for genomics: measuring functional divergence in multi-domain proteins. Genome Res 2001, 11(10):1632\u20131640. 10.1101\/gr. 183801","journal-title":"Genome Res"},{"issue":"1","key":"2837_CR17","doi-asserted-by":"crossref","first-page":"55","DOI":"10.3233\/ISB-00007","volume":"1","author":"MY Galperin","year":"1998","unstructured":"Galperin MY, Koonin EV: Sources of systematic error in functional annotation of genomes: domain rearrangement, non-orthologous gene displacement and operon disruption. Silico Biol 1998, 1(1):55\u201367.","journal-title":"Silico Biol"},{"issue":"8","key":"2837_CR18","doi-asserted-by":"publisher","first-page":"429","DOI":"10.1016\/S0168-9525(01)02348-4","volume":"17","author":"D Devos","year":"2001","unstructured":"Devos D, Valencia A: Intrinsic errors in genome annotation. Trends Genet 2001, 17(8):429\u2013431. 10.1016\/S0168-9525(01)02348-4","journal-title":"Trends Genet"},{"issue":"4","key":"2837_CR19","doi-asserted-by":"publisher","first-page":"132","DOI":"10.1016\/S0168-9525(99)01706-0","volume":"15","author":"SE Brenner","year":"1999","unstructured":"Brenner SE: Errors in genome annotation. Trends Genet 1999, 15(4):132\u2013133. 10.1016\/S0168-9525(99)01706-0","journal-title":"Trends Genet"},{"issue":"12","key":"2837_CR20","doi-asserted-by":"publisher","first-page":"1641","DOI":"10.1093\/bioinformatics\/18.12.1641","volume":"18","author":"WR Gilks","year":"2002","unstructured":"Gilks WR, Audit B, De Angelis D, Tsoka S, Ouzounis CA: Modeling the percolation of annotation errors in a database of protein sequences. Bioinformatics 2002, 18(12):1641\u20131649. 10.1093\/bioinformatics\/18.12.1641","journal-title":"Bioinformatics"},{"key":"2837_CR21","doi-asserted-by":"publisher","first-page":"9","DOI":"10.1186\/1471-2105-8-170","volume":"8","author":"CE Jones","year":"2007","unstructured":"Jones CE, Brown AL, Baumann U: Estimating the annotation error rate of curated GO database sequence annotations. BMC Bioinformatics 2007, 8: 9. 10.1186\/1471-2105-8-170","journal-title":"BMC Bioinformatics"},{"issue":"7","key":"2837_CR22","doi-asserted-by":"publisher","first-page":"1087","DOI":"10.1093\/bioinformatics\/bth044","volume":"20","author":"AK Arakaki","year":"2004","unstructured":"Arakaki AK, Zhang Y, Skolnick J: Large-scale assessment of the utility of low-resolution protein structures for biochemical function assignment. Bioinformatics 2004, 20(7):1087\u20131096. 10.1093\/bioinformatics\/bth044","journal-title":"Bioinformatics"},{"key":"2837_CR23","doi-asserted-by":"publisher","first-page":"17","DOI":"10.1186\/1471-2105-9-17","volume":"9","author":"DM Kristensen","year":"2008","unstructured":"Kristensen DM, Ward RM, Lisewski AM, Erdin S, Chen BY, Fofanov VY, Kimmel M, Kavraki LE, Lichtarge O: Prediction of enzyme function based on 3D templates of evolutionarily important amino acids. BMC Bioinformatics 2008, 9: 17. 10.1186\/1471-2105-9-17","journal-title":"BMC Bioinformatics"},{"issue":"6","key":"2837_CR24","doi-asserted-by":"publisher","first-page":"723","DOI":"10.1093\/bioinformatics\/btk038","volume":"22","author":"BJ Polacco","year":"2006","unstructured":"Polacco BJ, Babbitt PC: Automated discovery of 3D motifs for protein function annotation. Bioinformatics 2006, 22(6):723\u2013730. 10.1093\/bioinformatics\/btk038","journal-title":"Bioinformatics"},{"key":"2837_CR25","doi-asserted-by":"publisher","first-page":"187","DOI":"10.1007\/978-1-59745-243-4_17","volume-title":"Computational Systems Biology","author":"U Syed","year":"2009","unstructured":"Syed U, Yona G: Enzyme function prediction with interpretable models. In Computational Systems Biology. Volume 541. Edited by: McDermott J, Samudrala R, Bumgarner R, Montgomery K, Ireton R. Totowa, NJ: Humana Press; 2009:187\u2013199."},{"key":"2837_CR26","doi-asserted-by":"publisher","first-page":"177","DOI":"10.1186\/1471-2105-7-177","volume":"7","author":"P Kharchenko","year":"2006","unstructured":"Kharchenko P, Chen L, Freund Y, Vitkup D, Church GM: Identifying metabolic enzymes with multiple types of association evidence. BMC Bioinformatics 2006, 7: 177. 10.1186\/1471-2105-7-177","journal-title":"BMC Bioinformatics"},{"issue":"21","key":"2837_CR27","doi-asserted-by":"publisher","first-page":"6226","DOI":"10.1093\/nar\/gkh956","volume":"32","author":"W Tian","year":"2004","unstructured":"Tian W, Arakaki AK, Skolnick J: EFICAz: a comprehensive approach for accurate genome-scale enzyme function inference. Nucleic Acids Res 2004, 32(21):6226\u20136239. 10.1093\/nar\/gkh956","journal-title":"Nucleic Acids Res"},{"issue":"3","key":"2837_CR28","first-page":"273","volume":"20","author":"C Cortes","year":"1995","unstructured":"Cortes C, Vapnik V: SUPPORT-VECTOR NETWORKS. Mach Learn 1995, 20(3):273\u2013297.","journal-title":"Mach Learn"},{"key":"2837_CR29","volume-title":"Classification and regression trees","author":"L Breiman","year":"1984","unstructured":"Breiman L: Classification and regression trees. Belmont, Calif.: Wadsworth International Group; 1984."},{"key":"2837_CR30","unstructured":"KEGG: Kyoto Encyclopedia of Genes and Genomes[ftp:\/\/ftp.genome.jp\/pub\/kegg\/]"},{"key":"2837_CR31","unstructured":"PROSITE Database[ftp:\/\/us.expasy.org\/databases\/prosite\/]"},{"issue":"3","key":"2837_CR32","doi-asserted-by":"publisher","first-page":"265","DOI":"10.1093\/bib\/3.3.265","volume":"3","author":"CJ Sigrist","year":"2002","unstructured":"Sigrist CJ, Cerutti L, Hulo N, Gattiker A, Falquet L, Pagni M, Bairoch A, Bucher P: PROSITE: a documented database using patterns and profiles as motif descriptors. Brief Bioinform 2002, 3(3):265\u2013274. 10.1093\/bib\/3.3.265","journal-title":"Brief Bioinform"},{"key":"2837_CR33","unstructured":"UniProt Knowledgebase Database[ftp:\/\/us.expasy.org\/databases\/uniprot\/]"},{"issue":"9","key":"2837_CR34","doi-asserted-by":"publisher","first-page":"1011","DOI":"10.1038\/nbt0908-1011","volume":"26","author":"C Kingsford","year":"2008","unstructured":"Kingsford C, Salzberg SL: What are decision trees? Nat Biotechnol 2008, 26(9):1011\u20131013. 10.1038\/nbt0908-1011","journal-title":"Nat Biotechnol"},{"key":"2837_CR35","unstructured":"EFICAz2webservice[http:\/\/cssb.biology.gatech.edu\/skolnick\/webservice\/EFICAz2\/index.html]"},{"key":"2837_CR36","doi-asserted-by":"publisher","first-page":"249","DOI":"10.1186\/1471-2105-9-249","volume":"9","author":"J Espadaler","year":"2008","unstructured":"Espadaler J, Eswar N, Querol E, Avil\u00e9s FX, Sali A, Marti-Renom MA, Oliva B: Prediction of enzyme function by combining sequence similarity and protein interactions. BMC Bioinformatics 2008, 9: 249. 10.1186\/1471-2105-9-249","journal-title":"BMC Bioinformatics"},{"key":"2837_CR37","unstructured":"Pfam Database[ftp:\/\/ftp.sanger.ac.uk\/pub\/databases\/Pfam\/]"},{"issue":"2","key":"2837_CR38","doi-asserted-by":"publisher","first-page":"171","DOI":"10.1038\/nsb0295-171","volume":"2","author":"G Casari","year":"1995","unstructured":"Casari G, Sander C, Valencia A: A method to predict functional residues in proteins. Nat Struct Biol 1995, 2(2):171\u2013178. 10.1038\/nsb0295-171","journal-title":"Nat Struct Biol"},{"issue":"18","key":"2837_CR39","doi-asserted-by":"publisher","first-page":"6395","DOI":"10.1073\/pnas.0408677102","volume":"102","author":"WR Atchley","year":"2005","unstructured":"Atchley WR, Zhao J, Fernandes AD, Dr\u00fcke T: Solving the protein sequence metric problem. Proc Natl Acad Sci USA 2005, 102(18):6395\u20136400. 10.1073\/pnas.0408677102","journal-title":"Proc Natl Acad Sci USA"},{"issue":"15","key":"2837_CR40","doi-asserted-by":"publisher","first-page":"1978","DOI":"10.1093\/bioinformatics\/btg255","volume":"19","author":"Y Zhao","year":"2003","unstructured":"Zhao Y, Pinilla C, Valmori D, Martin R, Simon R: Application of support vector machines for T-cell epitopes prediction. Bioinformatics 2003, 19(15):1978\u20131984. 10.1093\/bioinformatics\/btg255","journal-title":"Bioinformatics"},{"key":"2837_CR41","unstructured":"LIBSVM: a library for support vector machines[http:\/\/www.csie.ntu.edu.tw\/~cjlin\/libsvm]"},{"key":"2837_CR42","volume-title":"R: A Language and Environment for Statistical Computing","author":"R Development Core Team","year":"2008","unstructured":"R Development Core Team: R: A Language and Environment for Statistical Computing. Vienna, Austria: R Foundation for Statistical Computing; 2008."}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-10-107.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,8]],"date-time":"2025-02-08T23:20:10Z","timestamp":1739056810000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-10-107"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,4,13]]},"references-count":42,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2009,12]]}},"alternative-id":["2837"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-10-107","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2009,4,13]]},"assertion":[{"value":"18 November 2008","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 April 2009","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 April 2009","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"107"}}