{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T12:21:12Z","timestamp":1767961272332,"version":"3.49.0"},"reference-count":32,"publisher":"Oxford University Press (OUP)","issue":"11","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2008,6,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: The deluge of biological information from different genomic initiatives and the rapid advancement in biotechnologies have made bioinformatics tools an integral part of modern biology. Among the widely used sequence alignment tools, BLAST and PSI-BLAST are arguably the most popular. PSI-BLAST, which uses an iterative profile position specific score matrix (PSSM)-based search strategy, is more sensitive than BLAST in detecting weak homologies, thus making it suitable for remote homolog detection. Many refinements have been made to improve PSI-BLAST, and its computational efficiency and high specificity have been much touted. Nevertheless, corruption of its profile via the incorporation of false positive sequences remains a major challenge.<\/jats:p>\n               <jats:p>Results: We have developed a simple and elegant approach to resolve the problem of model corruption in PSI-BLAST searches. We hypothesized that combining results from the first (least-corrupted) profile with results from later (most sensitive) iterations of PSI-BLAST provides a better discriminator for true and false hits. Accordingly, we have derived a formula that utilizes the E-values from these two PSI-BLAST iterations to obtain a figure of merit for rank-ordering the hits. Our verification results based on a \u2018gold-standard\u2019 test set indicate that this figure of merit does indeed delineate true positives from false positives better than PSI-BLAST E-values. Perhaps what is most notable about this strategy is that it is simple and straightforward to implement.<\/jats:p>\n               <jats:p>Contact: \u00a0bundschuh@mps.ohio-state.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/btn130","type":"journal-article","created":{"date-parts":[[2008,4,11]],"date-time":"2008-04-11T00:45:47Z","timestamp":1207874747000},"page":"1339-1343","source":"Crossref","is-referenced-by-count":22,"title":["Simple is beautiful: a straightforward approach to improve the delineation of true and false positives in PSI-BLAST searches"],"prefix":"10.1093","volume":"24","author":[{"given":"Marianne M.","family":"Lee","sequence":"first","affiliation":[{"name":"1 The Ohio State Biophysics Program, 2Departments of Biochemistry and Chemistry, Ohio State University, 484 W 12th Av. and 3Department of Physics, Ohio State University, 191 W Woodruff Av., Columbus OH 43210-1117, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael K.","family":"Chan","sequence":"additional","affiliation":[{"name":"1 The Ohio State Biophysics Program, 2Departments of Biochemistry and Chemistry, Ohio State University, 484 W 12th Av. and 3Department of Physics, Ohio State University, 191 W Woodruff Av., Columbus OH 43210-1117, USA"},{"name":"1 The Ohio State Biophysics Program, 2Departments of Biochemistry and Chemistry, Ohio State University, 484 W 12th Av. and 3Department of Physics, Ohio State University, 191 W Woodruff Av., Columbus OH 43210-1117, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ralf","family":"Bundschuh","sequence":"additional","affiliation":[{"name":"1 The Ohio State Biophysics Program, 2Departments of Biochemistry and Chemistry, Ohio State University, 484 W 12th Av. and 3Department of Physics, Ohio State University, 191 W Woodruff Av., Columbus OH 43210-1117, USA"},{"name":"1 The Ohio State Biophysics Program, 2Departments of Biochemistry and Chemistry, Ohio State University, 484 W 12th Av. and 3Department of Physics, Ohio State University, 191 W Woodruff Av., Columbus OH 43210-1117, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2008,4,10]]},"reference":[{"key":"2023020210060575000_B1","doi-asserted-by":"crossref","first-page":"13814","DOI":"10.1073\/pnas.0405612101","article-title":"Comparative homology agreement search: an effective combination of homology-search methods","volume":"101","author":"Alam","year":"2004","journal-title":"Proc. Natl Acad. Sci USA"},{"key":"2023020210060575000_B2","doi-asserted-by":"crossref","first-page":"460","DOI":"10.1016\/S0076-6879(96)66029-7","article-title":"Local alignment statistics","volume":"266","author":"Altschul","year":"1996","journal-title":"Methods Enzymol"},{"key":"2023020210060575000_B3","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1016\/S0022-2836(05)80360-2","article-title":"Basic local alignment search tool","volume":"215","author":"Altschul","year":"1990","journal-title":"J. Mol. Biol"},{"key":"2023020210060575000_B4","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","article-title":"Gapped BLAST and PSI-BLAST: a new generation of protein database search programs","volume":"25","author":"Altschul","year":"1997","journal-title":"Nucleic Acids Res"},{"key":"2023020210060575000_B5","doi-asserted-by":"crossref","first-page":"1023","DOI":"10.1006\/jmbi.1999.2653","article-title":"Gleaning non-trivial structural, functional and evolutionary information about proteins by iterative database searches","volume":"287","author":"Aravind","year":"1999","journal-title":"J. Mol. Biol"},{"key":"2023020210060575000_B6","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1186\/1471-2105-4-61","article-title":"PubMatrix: a tool for multiplex literature mining","volume":"4","author":"Becker","year":"2003","journal-title":"BMC Bioinformatics"},{"key":"2023020210060575000_B7","doi-asserted-by":"crossref","first-page":"164","DOI":"10.1126\/science.1853201","article-title":"A method to identify protein sequences that fold into a known three-dimensional structure","volume":"253","author":"Bowie","year":"1991","journal-title":"Science"},{"key":"2023020210060575000_B8","doi-asserted-by":"crossref","first-page":"598","DOI":"10.1016\/S0076-6879(96)66037-6","article-title":"Three-dimensional profiles for measuring compatibility of amino acid sequence with three-dimensional structure","volume":"266","author":"Bowie","year":"1996","journal-title":"Methods Enzymol"},{"key":"2023020210060575000_B9","first-page":"67","article-title":"The significance of protein sequence similarities","volume":"4","author":"Collins","year":"1988","journal-title":"Comput. Appl. Biosci"},{"key":"2023020210060575000_B10","doi-asserted-by":"crossref","first-page":"361","DOI":"10.1016\/S0959-440X(96)80056-X","article-title":"Hidden Markov models","volume":"6","author":"Eddy","year":"1996","journal-title":"Curr. Opin. Struct. Biol"},{"key":"2023020210060575000_B11","doi-asserted-by":"crossref","first-page":"823","DOI":"10.1006\/jmbi.1996.0679","article-title":"Significant improvement in accuracy of multiple protein sequence alignments by iterative refinement as assessed by reference to structural alignments","volume":"264","author":"Gotoh","year":"1996","journal-title":"J. Mol. Biol"},{"key":"2023020210060575000_B12","doi-asserted-by":"crossref","first-page":"903","DOI":"10.1006\/jmbi.2001.5080","article-title":"Assignment of homology to genome sequences using a library of hidden Markov models that represent all proteins of known structure","volume":"313","author":"Gough","year":"2001","journal-title":"J. Mol. Biol"},{"key":"2023020210060575000_B13","doi-asserted-by":"crossref","first-page":"4355","DOI":"10.1073\/pnas.84.13.4355","article-title":"Profile analysis: detection of distantly related proteins","volume":"84","author":"Gribskov","year":"1987","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020210060575000_B14","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1016\/S0097-8485(96)80004-0","article-title":"Use of receiver operating characteristic (ROC) analysis to evaluate sequence matching","volume":"20","author":"Gribskov","year":"1996","journal-title":"Comput. Chem"},{"key":"2023020210060575000_B15","doi-asserted-by":"crossref","first-page":"479","DOI":"10.1089\/cmb.1998.5.479","article-title":"Homology detection via family pairwise search","volume":"5","author":"Grundy","year":"1998","journal-title":"J. Comput. Biol"},{"key":"2023020210060575000_B16","first-page":"95","article-title":"Hidden Markov models for sequence analysis: extension and analysis of the basic method","volume":"12","author":"Hughey","year":"1996","journal-title":"Comput. Appl. Biosci"},{"key":"2023020210060575000_B17","doi-asserted-by":"crossref","first-page":"1451","DOI":"10.1093\/bioinformatics\/bti233","article-title":"A structure-based method for protein sequence alignment","volume":"21","author":"Kann","year":"2005","journal-title":"Bioinformatics"},{"key":"2023020210060575000_B18","doi-asserted-by":"crossref","first-page":"5617","DOI":"10.1093\/nar\/gkg769","article-title":"PANDORA: keyword-based analysis of protein sets by integration of annotation sources","volume":"31","author":"Kaplan","year":"2003","journal-title":"Nucleic Acids Res"},{"key":"2023020210060575000_B19","doi-asserted-by":"crossref","first-page":"2264","DOI":"10.1073\/pnas.87.6.2264","article-title":"Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes","volume":"87","author":"Karlin","year":"1990","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020210060575000_B20","doi-asserted-by":"crossref","first-page":"5873","DOI":"10.1073\/pnas.90.12.5873","article-title":"Applications and statistics for multiple high-scoring segments in molecular sequences","volume":"90","author":"Karlin","year":"1993","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020210060575000_B21","doi-asserted-by":"crossref","first-page":"113","DOI":"10.2307\/1427732","article-title":"Limit distributions of maximal segmental score among Markov-dependent partial sums","volume":"24","author":"Karlin","year":"1992","journal-title":"Adv. Appl. Prob"},{"key":"2023020210060575000_B22","doi-asserted-by":"crossref","first-page":"846","DOI":"10.1093\/bioinformatics\/14.10.846","article-title":"Hidden Markov models for detecting remote protein homologies","volume":"14","author":"Karplus","year":"1998","journal-title":"Bioinformatics"},{"key":"2023020210060575000_B23","doi-asserted-by":"crossref","first-page":"1409","DOI":"10.1002\/prot.21830","article-title":"Distant homology detection using a LEngth and STructure-based sequence Alignment Tool (LESTAT)","volume":"71","author":"Lee","year":"2007","journal-title":"Proteins"},{"key":"2023020210060575000_B24","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1016\/S0092-8240(05)80176-4","article-title":"Maximum likelihood estimation of the statistical distribution of Smith-Waterman local sequence similarity scores","volume":"54","author":"Mott","year":"1992","journal-title":"Bull. Math. Biol"},{"key":"2023020210060575000_B25","doi-asserted-by":"crossref","first-page":"1201","DOI":"10.1006\/jmbi.1998.2221","article-title":"Sequence comparisons using multiple sequences detect three times as many remote homologues as pairwise methods","volume":"284","author":"Park","year":"1998","journal-title":"J. Mol. Biol"},{"key":"2023020210060575000_B26","doi-asserted-by":"crossref","first-page":"2994","DOI":"10.1093\/nar\/29.14.2994","article-title":"Improving the accuracy of PSI-BLAST protein database searches with composition-based statistics and other refinements","volume":"29","author":"Schaffer","year":"2001","journal-title":"Nucleic Acids Res"},{"key":"2023020210060575000_B27","doi-asserted-by":"crossref","first-page":"1410","DOI":"10.1093\/bioinformatics\/btm115","article-title":"SherLoc: high-accuracy prediction of protein subcellular localization by integrating text and protein sequence data","volume":"23","author":"Shatkay","year":"2007","journal-title":"Bioinformatics"},{"key":"2023020210060575000_B28","doi-asserted-by":"crossref","first-page":"482","DOI":"10.1016\/0196-8858(81)90046-4","article-title":"Comparison of biosequences","volume":"2","author":"Smith","year":"1981","journal-title":"Adv. Appl. Math"},{"key":"2023020210060575000_B29","doi-asserted-by":"crossref","first-page":"645","DOI":"10.1093\/nar\/13.2.645","article-title":"The statistical distribution of nucleic acid similarities","volume":"13","author":"Smith","year":"1985","journal-title":"Nucleic Acids Res"},{"key":"2023020210060575000_B30","doi-asserted-by":"crossref","first-page":"2682","DOI":"10.1093\/nar\/27.13.2682","article-title":"A comprehensive comparison of multiple sequence alignment programs","volume":"27","author":"Thompson","year":"1999","journal-title":"Nucleic Acids Res"},{"key":"2023020210060575000_B31","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/S0022-2836(05)80006-3","article-title":"Sequence alignment and penalty choice. Review of concepts, case studies and implications","volume":"235","author":"Vingron","year":"1994","journal-title":"J. Mol. Biol"},{"key":"2023020210060575000_B32","doi-asserted-by":"crossref","first-page":"4625","DOI":"10.1073\/pnas.91.11.4625","article-title":"Rapid and accurate estimates of statistical significance for sequence data base searches","volume":"91","author":"Waterman","year":"1994","journal-title":"Proc. Natl Acad. Sci. USA"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/11\/1339\/49048511\/bioinformatics_24_11_1339.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/11\/1339\/49048511\/bioinformatics_24_11_1339.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,2]],"date-time":"2023-02-02T11:43:46Z","timestamp":1675338226000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/24\/11\/1339\/191162"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,4,10]]},"references-count":32,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2008,6,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btn130","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2008,6,1]]},"published":{"date-parts":[[2008,4,10]]}}}