{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,5,12]],"date-time":"2023-05-12T09:30:42Z","timestamp":1683883842588},"reference-count":31,"publisher":"Oxford University Press (OUP)","issue":"18","license":[{"start":{"date-parts":[[2016,10,2]],"date-time":"2016-10-02T00:00:00Z","timestamp":1475366400000},"content-version":"vor","delay-in-days":2981,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/2.0\/uk\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2008,9,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: A typical PSI-BLAST search consists of iterative scanning and alignment of a large sequence database during which a scoring profile is progressively built and refined. Such a profile can also be stored and used to search against a different database of sequences. Using it to search against a database of consensus rather than native sequences is a simple add-on that boosts performance surprisingly well. The improvement comes at a price: we hypothesized that random alignment score statistics would differ between native and consensus sequences. Thus PSI-BLAST-based profile searches against consensus sequences might incorrectly estimate statistical significance of alignment scores. In addition, iterative searches against consensus databases may fail. Here, we addressed these challenges in an attempt to harness the full power of the combination of PSI-BLAST and consensus sequences.<\/jats:p>\n               <jats:p>Results: We studied alignment score statistics for various types of consensus sequences. In general, the score distribution parameters of profile-based consensus sequence alignments differed significantly from those derived for the native sequences. PSI-BLAST partially compensated for the parameter variation. We have identified a protocol for building specialized consensus sequences that significantly improved search sensitivity and preserved score distribution parameters. As a result, PSI-BLAST profiles can be used to search specialized consensus sequences without sacrificing estimates of statistical significance. We also provided results indicating that iterative PSI-BLAST searches against consensus sequences could work very well. Overall, we showed how a very popular and effective method could be used to identify significantly more relevant similarities among protein sequences.<\/jats:p>\n               <jats:p>Availability: \u00a0http:\/\/www.rostlab.org\/services\/consensus\/<\/jats:p>\n               <jats:p>Contact: \u00a0dariusz@mit.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/btn384","type":"journal-article","created":{"date-parts":[[2008,8,5]],"date-time":"2008-08-05T00:30:04Z","timestamp":1217896204000},"page":"1987-1993","source":"Crossref","is-referenced-by-count":9,"title":["Powerful fusion: PSI-BLAST and consensus sequences"],"prefix":"10.1093","volume":"24","author":[{"given":"Dariusz","family":"Przybylski","sequence":"first","affiliation":[{"name":"1 Department of Biochemistry and Molecular Biophysics, Columbia University, 630 West 168th Street, New York, NY 10032, 2Broad Institute of MIT and Harvard University, 320 Charles St., Cambridge, MA 02141 and 3Columbia University Center for Computational Biology and Bioinformatics (C2B2), NorthEast Structural Genomics Consortium (NESG), New York Consortium on Membrane Protein Structure (NYCOMPS), 1130 St Nicholas Ave. Rm. 802, New York, NY 10032, USA"},{"name":"1 Department of Biochemistry and Molecular Biophysics, Columbia University, 630 West 168th Street, New York, NY 10032, 2Broad Institute of MIT and Harvard University, 320 Charles St., Cambridge, MA 02141 and 3Columbia University Center for Computational Biology and Bioinformatics (C2B2), NorthEast Structural Genomics Consortium (NESG), New York Consortium on Membrane Protein Structure (NYCOMPS), 1130 St Nicholas Ave. Rm. 802, New York, NY 10032, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Burkhard","family":"Rost","sequence":"additional","affiliation":[{"name":"1 Department of Biochemistry and Molecular Biophysics, Columbia University, 630 West 168th Street, New York, NY 10032, 2Broad Institute of MIT and Harvard University, 320 Charles St., Cambridge, MA 02141 and 3Columbia University Center for Computational Biology and Bioinformatics (C2B2), NorthEast Structural Genomics Consortium (NESG), New York Consortium on Membrane Protein Structure (NYCOMPS), 1130 St Nicholas Ave. Rm. 802, New York, NY 10032, USA"},{"name":"1 Department of Biochemistry and Molecular Biophysics, Columbia University, 630 West 168th Street, New York, NY 10032, 2Broad Institute of MIT and Harvard University, 320 Charles St., Cambridge, MA 02141 and 3Columbia University Center for Computational Biology and Bioinformatics (C2B2), NorthEast Structural Genomics Consortium (NESG), New York Consortium on Membrane Protein Structure (NYCOMPS), 1130 St Nicholas Ave. Rm. 802, New York, NY 10032, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2008,8,4]]},"reference":[{"key":"2023020211100600300_B1","doi-asserted-by":"crossref","first-page":"460","DOI":"10.1016\/S0076-6879(96)66029-7","article-title":"Local alignment statistics","volume":"266","author":"Altschul","year":"1996","journal-title":"Methods Enzymol."},{"key":"2023020211100600300_B2","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1016\/S0022-2836(05)80360-2","article-title":"Basic local alignment search tool","volume":"215","author":"Altschul","year":"1990","journal-title":"J. Mol. Biol."},{"key":"2023020211100600300_B3","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","article-title":"Gapped BLAST and PSI-BLAST: a new generation of protein database search programs","volume":"25","author":"Altschul","year":"1997","journal-title":"Nucleic Acids Res."},{"key":"2023020211100600300_B4","doi-asserted-by":"crossref","first-page":"351","DOI":"10.1093\/nar\/29.2.351","article-title":"The estimation of statistical parameters for local alignment score distributions","volume":"29","author":"Altschul","year":"2001","journal-title":"Nucleic Acids Res."},{"key":"2023020211100600300_B5","doi-asserted-by":"crossref","first-page":"D115","DOI":"10.1093\/nar\/gkh131","article-title":"UniProt: the universal protein knowledgebase","volume":"32","author":"Apweiler","year":"2004","journal-title":"Nucleic Acids Res."},{"key":"2023020211100600300_B6","doi-asserted-by":"crossref","first-page":"352","DOI":"10.1110\/ps.40501","article-title":"LiveBench-1: continuous benchmarking of protein structure prediction servers","volume":"10","author":"Bujnicki","year":"2001","journal-title":"Protein Sci."},{"key":"2023020211100600300_B7","doi-asserted-by":"crossref","first-page":"D247","DOI":"10.1093\/nar\/gkj149","article-title":"Pfam: clans, web tools and services","volume":"34","author":"Finn","year":"2006","journal-title":"Nucleic Acids Res."},{"issue":"(Suppl. 6)","key":"2023020211100600300_B8","doi-asserted-by":"crossref","first-page":"503","DOI":"10.1002\/prot.10538","article-title":"CAFASP3: the third critical assessment of fully automated structure prediction methods","volume":"53","author":"Fischer","year":"2003","journal-title":"Proteins"},{"key":"2023020211100600300_B9","doi-asserted-by":"crossref","first-page":"10915","DOI":"10.1073\/pnas.89.22.10915","article-title":"Amino acid substitution matrices from protein blocks","volume":"89","author":"Henikoff","year":"1992","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020211100600300_B10","doi-asserted-by":"crossref","first-page":"698","DOI":"10.1002\/pro.5560060319","article-title":"Embedding strategies for effective use of information from multiple sequence alignments","volume":"6","author":"Henikoff","year":"1997","journal-title":"Protein Sci."},{"key":"2023020211100600300_B11","doi-asserted-by":"crossref","first-page":"2287","DOI":"10.1093\/bioinformatics\/bti374","article-title":"Quasi-consensus-based comparison of profile hidden Markov models for protein sequences","volume":"21","author":"Kahsay","year":"2005","journal-title":"Bioinformatics"},{"key":"2023020211100600300_B12","doi-asserted-by":"crossref","first-page":"2264","DOI":"10.1073\/pnas.87.6.2264","article-title":"Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes","volume":"87","author":"Karlin","year":"1990","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020211100600300_B13","doi-asserted-by":"crossref","first-page":"D257","DOI":"10.1093\/nar\/gkj079","article-title":"SMART 5: domains in the context of genomes and networks","volume":"34","author":"Letunic","year":"2006","journal-title":"Nucleic Acids Res."},{"key":"2023020211100600300_B14","doi-asserted-by":"crossref","first-page":"282","DOI":"10.1093\/bioinformatics\/17.3.282","article-title":"Clustering of highly homologous sequences to reduce the size of large protein databases","volume":"17","author":"Li","year":"2001","journal-title":"Bioinformatics"},{"key":"2023020211100600300_B15","doi-asserted-by":"crossref","first-page":"281","DOI":"10.1093\/nar\/30.1.281","article-title":"CDD: a database of conserved domain alignments with links to domain three-dimensional structure","volume":"30","author":"Marchler-Bauer","year":"2002","journal-title":"Nucleic Acids Res."},{"key":"2023020211100600300_B16","doi-asserted-by":"crossref","first-page":"51","DOI":"10.1186\/1471-2148-6-51","article-title":"PHOG-BLAST - a new generation tool for fast similarity search of protein families","volume":"6","author":"Merkeev","year":"2006","journal-title":"BMC Evol. Biol."},{"key":"2023020211100600300_B17","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1016\/S0092-8240(05)80176-4","article-title":"Maximum-likelihood estimation of the statistical distribution of Smith-Waterman local sequence similarity scores","volume":"54","author":"Mott","year":"1992","journal-title":"Bull. Math. Biol."},{"key":"2023020211100600300_B18","doi-asserted-by":"crossref","first-page":"536","DOI":"10.1016\/S0022-2836(05)80134-2","article-title":"SCOP: a structural classification of proteins database for the investigation of sequences and structures","volume":"247","author":"Murzin","year":"1995","journal-title":"J. Mol. Biol."},{"key":"2023020211100600300_B19","first-page":"211","article-title":"Rapid assessment of extremal statistics for gapped local alignment","author":"Olsen","year":"1999","journal-title":"Proc. Int. Conf. Intell. Syst. Mol. Biol."},{"key":"2023020211100600300_B20","doi-asserted-by":"crossref","first-page":"567","DOI":"10.1016\/0022-2836(87)90200-2","article-title":"Detecting homology of distantly related proteins with consensus sequences","volume":"198","author":"Patthy","year":"1987","journal-title":"J. Mol. Biol."},{"key":"2023020211100600300_B21","doi-asserted-by":"crossref","first-page":"2238","DOI":"10.1093\/nar\/gkm107","article-title":"Consensus sequences improve PSI-BLAST through mimicking profile-profile alignments","volume":"35","author":"Przybylski","year":"2007","journal-title":"Nucleic Acids Res."},{"key":"2023020211100600300_B22","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1093\/protein\/12.2.85","article-title":"Twilight zone of protein sequence alignments","volume":"12","author":"Rost","year":"1999","journal-title":"Protein Eng."},{"key":"2023020211100600300_B23","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1002\/prot.340090107","article-title":"Database of homology-derived protein structures and the structural meaning of sequence alignment","volume":"9","author":"Sander","year":"1991","journal-title":"Proteins"},{"key":"2023020211100600300_B24","doi-asserted-by":"crossref","first-page":"1000","DOI":"10.1093\/bioinformatics\/15.12.1000","article-title":"IMPALA: matching a protein sequence against a collection of PSI-BLAST-constructed position-specific score matrices","volume":"15","author":"Schaffer","year":"1999","journal-title":"Bioinformatics"},{"key":"2023020211100600300_B25","doi-asserted-by":"crossref","first-page":"2994","DOI":"10.1093\/nar\/29.14.2994","article-title":"Improving the accuracy of PSI-BLAST protein database searches with composition-based statistics and other refinements","volume":"29","author":"Schaffer","year":"2001","journal-title":"Nucleic Acids Res."},{"key":"2023020211100600300_B26","doi-asserted-by":"crossref","first-page":"5857","DOI":"10.1073\/pnas.95.11.5857","article-title":"SMART, a simple modular architecture research tool: identification of signaling domains","volume":"95","author":"Schultz","year":"1998","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020211100600300_B27","doi-asserted-by":"crossref","first-page":"246","DOI":"10.1093\/bib\/3.3.246","article-title":"ProDom: automated clustering of homologous domains","volume":"3","author":"Servant","year":"2002","journal-title":"Brief. Bioinform."},{"key":"2023020211100600300_B28","doi-asserted-by":"crossref","first-page":"482","DOI":"10.1002\/pro.5560030314","article-title":"Modular arrangement of proteins as inferred from analysis of homology","volume":"3","author":"Sonnhammer","year":"1994","journal-title":"Protein Sci."},{"key":"2023020211100600300_B29","doi-asserted-by":"crossref","first-page":"769","DOI":"10.1016\/S0092-8674(00)80587-5","article-title":"A sliding clamp model for the Rad1 family of cell cycle checkpoint proteins","volume":"96","author":"Thelen","year":"1999","journal-title":"Cell"},{"key":"2023020211100600300_B30","doi-asserted-by":"crossref","first-page":"4625","DOI":"10.1073\/pnas.91.11.4625","article-title":"Rapid and accurate estimates of statistical significance for sequence data base searches","volume":"91","author":"Waterman","year":"1994","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020211100600300_B31","doi-asserted-by":"crossref","first-page":"902","DOI":"10.1093\/bioinformatics\/bti070","article-title":"The construction of amino acid substitution matrices for the comparison of proteins with non-standard compositions","volume":"21","author":"Yu","year":"2005","journal-title":"Bioinformatics"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/18\/1987\/49050078\/bioinformatics_24_18_1987.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/18\/1987\/49050078\/bioinformatics_24_18_1987.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,2]],"date-time":"2023-02-02T13:30:51Z","timestamp":1675344651000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/24\/18\/1987\/192491"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,8,4]]},"references-count":31,"journal-issue":{"issue":"18","published-print":{"date-parts":[[2008,9,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btn384","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2008,9,15]]},"published":{"date-parts":[[2008,8,4]]}}}