{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T04:53:47Z","timestamp":1761540827538},"reference-count":52,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2008,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>Automated protein function prediction methods are needed to keep pace with high-throughput sequencing. With the existence of many programs and databases for inferring different protein functions, a pipeline that properly integrates these resources will benefit from the advantages of each method. However, integrated systems usually do not provide mechanisms to generate customized databases to predict particular protein functions. Here, we describe a tool termed PIPA (Pipeline for Protein Annotation) that has these capabilities.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>PIPA annotates protein functions by combining the results of multiple programs and databases, such as InterPro and the Conserved Domains Database, into common Gene Ontology (GO) terms. The major algorithms implemented in PIPA are: (1) a profile database generation algorithm, which generates customized profile databases to predict particular protein functions, (2) an automated ontology mapping generation algorithm, which maps various classification schemes into GO, and (3) a consensus algorithm to reconcile annotations from the integrated programs and databases.<\/jats:p>\n            <jats:p>PIPA's profile generation algorithm is employed to construct the enzyme profile database CatFam, which predicts catalytic functions described by Enzyme Commission (EC) numbers. Validation tests show that CatFam yields average recall and precision larger than 95.0%. CatFam is integrated with PIPA.<\/jats:p>\n            <jats:p>We use an association rule mining algorithm to automatically generate mappings between terms of two ontologies from annotated sample proteins. Incorporating the ontologies' hierarchical topology into the algorithm increases the number of generated mappings. In particular, it generates 40.0% additional mappings from the Clusters of Orthologous Groups (COG) to EC numbers and a six-fold increase in mappings from COG to GO terms. The mappings to EC numbers show a very high precision (99.8%) and recall (96.6%), while the mappings to GO terms show moderate precision (80.0%) and low recall (33.0%).<\/jats:p>\n            <jats:p>Our consensus algorithm for GO annotation is based on the computation and propagation of likelihood scores associated with GO terms. The test results suggest that, for a given recall, the application of the consensus algorithm yields higher precision than when consensus is not used.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion<\/jats:title>\n            <jats:p>The algorithms implemented in PIPA provide automated genome-wide protein function annotation based on reconciled predictions from multiple resources.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-9-52","type":"journal-article","created":{"date-parts":[[2008,1,25]],"date-time":"2008-01-25T19:20:45Z","timestamp":1201288845000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":24,"title":["The development of PIPA: an integrated and automated pipeline for genome-wide protein function annotation"],"prefix":"10.1186","volume":"9","author":[{"given":"Chenggang","family":"Yu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nela","family":"Zavaljevski","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Valmik","family":"Desai","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Seth","family":"Johnson","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fred J","family":"Stevens","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jaques","family":"Reifman","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2008,1,25]]},"reference":[{"issue":"3","key":"2037_CR1","doi-asserted-by":"publisher","first-page":"307","DOI":"10.1017\/S0033583503003901","volume":"36","author":"JC Whisstock","year":"2003","unstructured":"Whisstock JC, Lesk AM: Prediction of protein function from protein sequence and structure.Q Rev Biophys 2004\/03\/20 edition. 2003, 36(3):307\u2013340. 10.1017\/S0033583503003901","journal-title":"Q Rev Biophys"},{"issue":"2","key":"2037_CR2","doi-asserted-by":"publisher","first-page":"170","DOI":"10.1093\/bioinformatics\/bth021","volume":"20","author":"K Sjolander","year":"2004","unstructured":"Sjolander K: Phylogenomic inference of protein molecular function: advances and challenges.Bioinformatics 2004\/01\/22 edition. 2004, 20(2):170\u2013179. 10.1093\/bioinformatics\/bth021","journal-title":"Bioinformatics"},{"issue":"21","key":"2037_CR3","doi-asserted-by":"publisher","first-page":"1475","DOI":"10.1016\/S1359-6446(05)03621-4","volume":"10","author":"Y Ofran","year":"2005","unstructured":"Ofran Y, Punta M, Schneider R, Rost B: Beyond annotation transfer by homology: novel protein-function prediction methods to assist drug discovery.Drug Discov Today 2005\/10\/26 edition. 2005, 10(21):1475\u20131482. 10.1016\/S1359-6446(05)03621-4","journal-title":"Drug Discov Today"},{"issue":"3","key":"2037_CR4","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1093\/bib\/bbl004","volume":"7","author":"I Friedberg","year":"2006","unstructured":"Friedberg I: Automated protein function prediction--the genomic challenge.Brief Bioinform 2006\/06\/15 edition. 2006, 7(3):225\u2013242. 10.1093\/bib\/bbl004","journal-title":"Brief Bioinform"},{"issue":"Database issue","key":"2037_CR5","doi-asserted-by":"publisher","first-page":"D247","DOI":"10.1093\/nar\/gkj149","volume":"34","author":"RD Finn","year":"2006","unstructured":"Finn RD, Mistry J, Schuster-Bockler B, Griffiths-Jones S, Hollich V, Lassmann T, Moxon S, Marshall M, Khanna A, Durbin R, Eddy SR, Sonnhammer EL, Bateman A: Pfam: clans, web tools and services.Nucleic Acids Res 2005\/12\/31 edition. 2006, 34(Database issue):D247\u201351. 10.1093\/nar\/gkj149","journal-title":"Nucleic Acids Res"},{"issue":"Database issue","key":"2037_CR6","doi-asserted-by":"publisher","first-page":"D212","DOI":"10.1093\/nar\/gki034","volume":"33","author":"C Bru","year":"2005","unstructured":"Bru C, Courcelle E, Carrere S, Beausse Y, Dalmar S, Kahn D: The ProDom database of protein domain families: more emphasis on 3D.Nucleic Acids Res 2004\/12\/21 edition. 2005, 33(Database issue):D212\u20135. 10.1093\/nar\/gki034","journal-title":"Nucleic Acids Res"},{"issue":"Database issue","key":"2037_CR7","doi-asserted-by":"publisher","first-page":"D134","DOI":"10.1093\/nar\/gkh044","volume":"32","author":"N Hulo","year":"2004","unstructured":"Hulo N, Sigrist CJ, Le Saux V, Langendijk-Genevaux PS, Bordoli L, Gattiker A, De Castro E, Bucher P, Bairoch A: Recent improvements to the PROSITE database.Nucleic Acids Res 2003\/12\/19 edition. 2004, 32(Database issue):D134\u20137. 10.1093\/nar\/gkh044","journal-title":"Nucleic Acids Res"},{"issue":"1","key":"2037_CR8","doi-asserted-by":"publisher","first-page":"33","DOI":"10.1093\/nar\/28.1.33","volume":"28","author":"RL Tatusov","year":"2000","unstructured":"Tatusov RL, Galperin MY, Natale DA, Koonin EV: The COG database: a tool for genome-scale analysis of protein functions and evolution.Nucleic Acids Res 1999\/12\/11 edition. 2000, 28(1):33\u201336. 10.1093\/nar\/28.1.33","journal-title":"Nucleic Acids Res"},{"key":"2037_CR9","unstructured":"CDD[http:\/\/www.ncbi.nlm.nih.gov\/Structure\/cdd\/cdd.shtml]"},{"issue":"22","key":"2037_CR10","doi-asserted-by":"publisher","first-page":"6633","DOI":"10.1093\/nar\/gkg847","volume":"31","author":"C Claudel-Renard","year":"2003","unstructured":"Claudel-Renard C, Chevalet C, Faraut T, Kahn D: Enzyme-specific profiles for genome annotation: PRIAM.Nucleic Acids Res 2003\/11\/07 edition. 2003, 31(22):6633\u20136639. 10.1093\/nar\/gkg847","journal-title":"Nucleic Acids Res"},{"issue":"21","key":"2037_CR11","doi-asserted-by":"publisher","first-page":"6226","DOI":"10.1093\/nar\/gkh956","volume":"32","author":"W Tian","year":"2004","unstructured":"Tian W, Arakaki AK, Skolnick J: EFICAz: a comprehensive approach for accurate genome-scale enzyme function inference.Nucleic Acids Res 2004\/12\/04 edition. 2004, 32(21):6226\u20136239. 10.1093\/nar\/gkh956","journal-title":"Nucleic Acids Res"},{"key":"2037_CR12","unstructured":"InterPro[http:\/\/www.ebi.ac.uk\/interpro\/]"},{"issue":"Web Server issu","key":"2037_CR13","doi-asserted-by":"publisher","first-page":"W455","DOI":"10.1093\/nar\/gki593","volume":"33","author":"GH Van Domselaar","year":"2005","unstructured":"Van Domselaar GH, Stothard P, Shrivastava S, Cruz JA, Guo A, Dong X, Lu P, Szafron D, Greiner R, Wishart DS: BASys: a web server for automated bacterial genome annotation.Nucleic Acids Res 2005\/06\/28 edition. 2005, 33(Web Server issue):W455\u20139. 10.1093\/nar\/gki593","journal-title":"Nucleic Acids Res"},{"issue":"8","key":"2037_CR14","doi-asserted-by":"publisher","first-page":"2187","DOI":"10.1093\/nar\/gkg312","volume":"31","author":"F Meyer","year":"2003","unstructured":"Meyer F, Goesmann A, McHardy AC, Bartels D, Bekel T, Clausen J, Kalinowski J, Linke B, Rupp O, Giegerich R, Puhler A: GenDB--an open source genome annotation system for prokaryote genomes.Nucleic Acids Res 2003\/04\/12 edition. 2003, 31(8):2187\u20132195. 10.1093\/nar\/gkg312","journal-title":"Nucleic Acids Res"},{"issue":"Database issue","key":"2037_CR15","doi-asserted-by":"publisher","first-page":"D369","DOI":"10.1093\/nar\/gkj095","volume":"34","author":"N Maltsev","year":"2006","unstructured":"Maltsev N, Glass E, Sulakhe D, Rodriguez A, Syed MH, Bompada T, Zhang Y, D'Souza M: PUMA2--grid-based high-throughput analysis of genomes and metabolic pathways.Nucleic Acids Res 2005\/12\/31 edition. 2006, 34(Database issue):D369\u201372. 10.1093\/nar\/gkj095","journal-title":"Nucleic Acids Res"},{"issue":"1","key":"2037_CR16","doi-asserted-by":"publisher","first-page":"53","DOI":"10.1093\/nar\/gkj406","volume":"34","author":"D Vallenet","year":"2006","unstructured":"Vallenet D, Labarre L, Rouy Z, Barbe V, Bocs S, Cruveiller S, Lajus A, Pascal G, Scarpelli C, Medigue C: MaGe: a microbial genome annotation system supported by synteny results.Nucleic Acids Res 2006\/01\/13 edition. 2006, 34(1):53\u201365. 10.1093\/nar\/gkj406","journal-title":"Nucleic Acids Res"},{"issue":"12","key":"2037_CR17","doi-asserted-by":"publisher","first-page":"3533","DOI":"10.1093\/nar\/gkl471","volume":"34","author":"K Bryson","year":"2006","unstructured":"Bryson K, Loux V, Bossy R, Nicolas P, Chaillou S, van de Guchte M, Penaud S, Maguin E, Hoebeke M, Bessieres P, Gibrat JF: AGMIAL: implementing an annotation strategy for prokaryote genomes as a distributed system.Nucleic Acids Res 2006\/07\/21 edition. 2006, 34(12):3533\u20133545. 10.1093\/nar\/gkl471","journal-title":"Nucleic Acids Res"},{"issue":"Database issue","key":"2037_CR18","doi-asserted-by":"publisher","first-page":"D344","DOI":"10.1093\/nar\/gkj024","volume":"34","author":"VM Markowitz","year":"2006","unstructured":"Markowitz VM, Korzeniewski F, Palaniappan K, Szeto E, Werner G, Padki A, Zhao X, Dubchak I, Hugenholtz P, Anderson I, Lykidis A, Mavromatis K, Ivanova N, Kyrpides NC: The integrated microbial genomes (IMG) system.Nucleic Acids Res 2005\/12\/31 edition. 2006, 34(Database issue):D344\u20138. 10.1093\/nar\/gkj024","journal-title":"Nucleic Acids Res"},{"issue":"17","key":"2037_CR19","doi-asserted-by":"publisher","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","volume":"25","author":"SF Altschul","year":"1997","unstructured":"Altschul SF, Madden TL, Schaffer AA, Zhang J, Zhang Z, Miller W, Lipman DJ: Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.Nucleic Acids Res 1997\/09\/01 edition. 1997, 25(17):3389\u20133402. 10.1093\/nar\/25.17.3389","journal-title":"Nucleic Acids Res"},{"key":"2037_CR20","unstructured":"HMMER[http:\/\/hmmer.janelia.org\/]"},{"key":"2037_CR21","unstructured":"Gene Ontology[http:\/\/www.geneontology.org\/GO.indices.shtml]"},{"key":"2037_CR22","volume-title":"VLDB Conference","author":"R Agarwal Srikant R","year":"1999","unstructured":"Agarwal R Srikant R: Fast Algorithm for Mining Association Rules. In VLDB Conference. Santiago, Chile ; 1999."},{"key":"2037_CR23","doi-asserted-by":"publisher","first-page":"304","DOI":"10.1186\/1471-2105-7-304","volume":"7","author":"SH Chiu","year":"2006","unstructured":"Chiu SH, Chen CC, Yuan GF, Lin TH: Association algorithm to mine the rules that govern enzyme definition and to classify protein sequences.BMC Bioinformatics 2006\/06\/17 edition. 2006, 7: 304. 10.1186\/1471-2105-7-304","journal-title":"BMC Bioinformatics"},{"issue":"18","key":"2037_CR24","doi-asserted-by":"publisher","first-page":"2484","DOI":"10.1093\/bioinformatics\/btg338","volume":"19","author":"S Khan","year":"2003","unstructured":"Khan S, Situ G, Decker K, Schmidt CJ: GoFigure: automated Gene Ontology annotation.Bioinformatics 2003\/12\/12 edition. 2003, 19(18):2484\u20132485. 10.1093\/bioinformatics\/btg338","journal-title":"Bioinformatics"},{"key":"2037_CR25","doi-asserted-by":"publisher","first-page":"178","DOI":"10.1186\/1471-2105-5-178","volume":"5","author":"DM Martin","year":"2004","unstructured":"Martin DM, Berriman M, Barton GJ: GOtcha: a new method for prediction of protein function assessed by the annotation of seven genomes.BMC Bioinformatics 2004\/11\/20 edition. 2004, 5: 178. 10.1186\/1471-2105-5-178","journal-title":"BMC Bioinformatics"},{"issue":"5","key":"2037_CR26","doi-asserted-by":"publisher","first-page":"1027","DOI":"10.1016\/j.jmb.2004.03.016","volume":"338","author":"L Kall","year":"2004","unstructured":"Kall L, Krogh A, Sonnhammer EL: A combined transmembrane topology and signal peptide prediction method.J Mol Biol 2004\/04\/28 edition. 2004, 338(5):1027\u20131036. 10.1016\/j.jmb.2004.03.016","journal-title":"J Mol Biol"},{"issue":"5","key":"2037_CR27","doi-asserted-by":"publisher","first-page":"617","DOI":"10.1093\/bioinformatics\/bti057","volume":"21","author":"JL Gardy","year":"2005","unstructured":"Gardy JL, Laird MR, Chen F, Rey S, Walsh CJ, Ester M, Brinkman FS: PSORTb v.2.0: expanded prediction of bacterial protein subcellular localization and insights gained from comparative proteome analysis.Bioinformatics 2004\/10\/27 edition. 2005, 21(5):617\u2013623. 10.1093\/bioinformatics\/bti057","journal-title":"Bioinformatics"},{"key":"2037_CR28","unstructured":"FASTA (Pearson)[http:\/\/www.ebi.ac.uk\/help\/formats_frame.html]"},{"key":"2037_CR29","unstructured":"General Feature Format[http:\/\/www.sanger.ac.uk\/Software\/formats\/GFF\/]"},{"key":"2037_CR30","volume-title":"IEEE Symposium on Computational Intelligence in Bioinformatics and Computational Biology","author":"R Eisner Poulin B,","year":"2005","unstructured":"Eisner R Poulin B, Szafron D, Lu P, Greiner R: Improving protein function prediction using the hierarchical structure of the Gene Ontology. In IEEE Symposium on Computational Intelligence in Bioinformatics and Computational Biology. San Diego, CA ; 2005."},{"issue":"6","key":"2037_CR31","doi-asserted-by":"publisher","first-page":"1544","DOI":"10.1110\/ps.062184006","volume":"15","author":"K Verspoor","year":"2006","unstructured":"Verspoor K, Cohn J, Mniszewski S, Joslyn C: A categorization approach to automated ontological function annotation.Protein Sci 2006\/05\/05 edition. 2006, 15(6):1544\u20131549. 10.1110\/ps.062184006","journal-title":"Protein Sci"},{"key":"2037_CR32","unstructured":"Integrated Microbial Genomes[http:\/\/img.jgi.doe.gov\/pub\/doc\/dataprep.html]"},{"issue":"5","key":"2037_CR33","doi-asserted-by":"publisher","first-page":"635","DOI":"10.1093\/bioinformatics\/btg036","volume":"19","author":"LJ Jensen","year":"2003","unstructured":"Jensen LJ, Gupta R, Staerfeldt HH, Brunak S: Prediction of human protein function according to Gene Ontology categories.Bioinformatics 2003\/03\/26 edition. 2003, 19(5):635\u2013642. 10.1093\/bioinformatics\/btg036","journal-title":"Bioinformatics"},{"issue":"6","key":"2037_CR34","doi-asserted-by":"publisher","first-page":"895","DOI":"10.1093\/bioinformatics\/btg500","volume":"20","author":"M Deng","year":"2004","unstructured":"Deng M, Tu Z, Sun F, Chen T: Mapping Gene Ontology to proteins based on protein-protein interaction data.Bioinformatics 2004\/01\/31 edition. 2004, 20(6):895\u2013902. 10.1093\/bioinformatics\/btg500","journal-title":"Bioinformatics"},{"key":"2037_CR35","doi-asserted-by":"publisher","first-page":"261","DOI":"10.1186\/1471-2105-8-261","volume":"8","author":"I Artamonova","year":"2007","unstructured":"Artamonova I, Frishman G, Frishman D: Applying negative rule mining to improve genome annotation.BMC Bioinformatics 2007\/07\/31 edition. 2007, 8: 261. 10.1186\/1471-2105-8-261","journal-title":"BMC Bioinformatics"},{"key":"2037_CR36","unstructured":"GO Evidence Codes[http:\/\/www.geneontology.org\/GO.evidence.shtml]"},{"issue":"3","key":"2037_CR37","doi-asserted-by":"publisher","first-page":"264","DOI":"10.1145\/331499.331504","volume":"31","author":"AK Jain Murthy MN,","year":"1999","unstructured":"Jain AK Murthy MN, Flynn PJ: Data Clustering: A Review.ACM Computing Surveys 1999, 31(3):264\u2013323. 10.1145\/331499.331504","journal-title":"ACM Computing Surveys"},{"issue":"13","key":"2037_CR38","doi-asserted-by":"publisher","first-page":"3497","DOI":"10.1093\/nar\/gkg500","volume":"31","author":"R Chenna","year":"2003","unstructured":"Chenna R, Sugawara H, Koike T, Lopez R, Gibson TJ, Higgins DG, Thompson JD: Multiple sequence alignment with the Clustal series of programs.Nucleic Acids Res 2003\/06\/26 edition. 2003, 31(13):3497\u20133500. 10.1093\/nar\/gkg500","journal-title":"Nucleic Acids Res"},{"key":"2037_CR39","unstructured":"COG[http:\/\/www.ncbi.nlm.nih.gov\/COG\/grace\/]"},{"key":"2037_CR40","unstructured":"Pfam[http:\/\/pfam.sanger.ac.uk\/]"},{"key":"2037_CR41","unstructured":"TIGRfam[http:\/\/www.tigr.org\/TIGRFAMs\/]"},{"key":"2037_CR42","unstructured":"SMART[http:\/\/smart.embl-heidelberg.de\/]"},{"key":"2037_CR43","unstructured":"Gene3D[http:\/\/cathwww.biochem.ucl.ac.uk:8080\/Gene3D\/]"},{"key":"2037_CR44","unstructured":"FprintScan[http:\/\/www.bioinf.manchester.ac.uk\/dbbrowser\/PRINTS\/]"},{"key":"2037_CR45","unstructured":"PANTHER[http:\/\/www.pantherdb.org\/]"},{"key":"2037_CR46","unstructured":"SUPERFAMILY[http:\/\/supfam.org\/SUPERFAMILY\/index.html]"},{"key":"2037_CR47","unstructured":"ProDom[http:\/\/prodom.prabi.fr\/prodom\/current\/html\/home.php]"},{"key":"2037_CR48","unstructured":"PIR[http:\/\/pir.georgetown.edu\/]"},{"key":"2037_CR49","unstructured":"PROSITE[http:\/\/expasy.org\/prosite\/]"},{"key":"2037_CR50","unstructured":"COILS[http:\/\/www.ch.embnet.org\/software\/COILS_form.html]"},{"key":"2037_CR51","unstructured":"Phobius[http:\/\/phobius.sbc.su.se\/]"},{"key":"2037_CR52","unstructured":"PSORTb[http:\/\/www.psort.org\/psortb\/]"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-9-52.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T10:59:54Z","timestamp":1630493994000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-9-52"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,1,25]]},"references-count":52,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2008,12]]}},"alternative-id":["2037"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-9-52","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2008,1,25]]},"assertion":[{"value":"22 August 2007","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 January 2008","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 January 2008","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"52"}}