{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,30]],"date-time":"2026-08-30T05:33:43Z","timestamp":1788068023055,"version":"build-2784847793"},"reference-count":30,"publisher":"Oxford University Press (OUP)","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2008,2,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: The analyses of the increasing number of genome sequences requires shortcuts for the detection of orthologs, such as Reciprocal Best Hits (RBH), where orthologs are assumed if two genes each in a different genome find each other as the best hit in the other genome. Two BLAST options seem to affect alignment scores the most, and thus the choice of a best hit: the filtering of low information sequence segments and the algorithm used to produce the final alignment. Thus, we decided to test whether such options would help better detect orthologs.<\/jats:p>\n               <jats:p>Results: Using Escherichia coli K12 as an example, we compared the number and quality of orthologs detected as RBH. We tested four different conditions derived from two options: filtering of low-information segments, hard (default) versus soft; and alignment algorithm, default (based on matching words) versus Smith\u2013Waterman. All options resulted in significant differences in the number of orthologs detected, with the highest numbers obtained with the combination of soft filtering with Smith\u2013Waterman alignments. We compared these results with those of Reciprocal Shortest Distances (RSD), supposed to be superior to RBH because it uses an evolutionary measure of distance, rather than BLAST statistics, to rank homologs and thus detect orthologs. RSD barely increased the number of orthologs detected over those found with RBH. Error estimates, based on analyses of conservation of gene order, found small differences in the quality of orthologs detected using RBH. However, RSD showed the highest error rates. Thus, RSD have no advantages over RBH.<\/jats:p>\n               <jats:p>Availability: Orthologs detected as Reciprocal Best Hits using soft masking and Smith\u2013Waterman alignments can be downloaded from http:\/\/popolvuh.wlu.ca\/Orthologs.<\/jats:p>\n               <jats:p>Contact: \u00a0gmoreno@wlu.ca<\/jats:p>","DOI":"10.1093\/bioinformatics\/btm585","type":"journal-article","created":{"date-parts":[[2007,11,28]],"date-time":"2007-11-28T01:13:15Z","timestamp":1196212395000},"page":"319-324","source":"Crossref","is-referenced-by-count":443,"title":["Choosing BLAST options for better detection of orthologs as reciprocal best hits"],"prefix":"10.1093","volume":"24","author":[{"given":"Gabriel","family":"Moreno-Hagelsieb","sequence":"first","affiliation":[{"name":"Department of Biology, Wilfrid Laurier University, 75 University Avenue West, Waterloo, ON, Canada, N2L 3C5"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kristen","family":"Latimer","sequence":"additional","affiliation":[{"name":"Department of Biology, Wilfrid Laurier University, 75 University Avenue West, Waterloo, ON, Canada, N2L 3C5"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2007,11,26]]},"reference":[{"key":"2023061011434056000_B1","doi-asserted-by":"crossref","first-page":"e9","DOI":"10.1093\/bioinformatics\/btl213","article-title":"Automatic clustering of orthologs and inparalogs shared by multiple proteomes","volume":"22","author":"Alexeyenko","year":"2006","journal-title":"Bioinformatics"},{"key":"2023061011434056000_B2","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1016\/S0022-2836(05)80360-2","article-title":"Basic local alignment search tool","volume":"215","author":"Altschul","year":"1990","journal-title":"J. Mol. Biol"},{"key":"2023061011434056000_B3","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","article-title":"Gapped BLAST and PSI-BLAST: a new generation of protein database search programs","volume":"25","author":"Altschul","year":"1997","journal-title":"Nucleic Acids Res"},{"key":"2023061011434056000_B4","doi-asserted-by":"crossref","first-page":"2607","DOI":"10.1093\/nar\/29.12.2607","article-title":"GeneMarkS: a self-training method for prediction of gene starts in microbial genomes implications for finding sequence motifs in regulatory regions","volume":"29","author":"Besemer","year":"2001","journal-title":"Nucleic Acids Res"},{"key":"2023061011434056000_B5","doi-asserted-by":"crossref","first-page":"1453","DOI":"10.1126\/science.277.5331.1453","article-title":"The complete genome sequence ofEscherichia coli K-12","volume":"277","author":"Blattner","year":"1997","journal-title":"Science"},{"key":"2023061011434056000_B6","doi-asserted-by":"crossref","first-page":"365","DOI":"10.1093\/nar\/gkg095","article-title":"The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003","volume":"31","author":"Boeckmann","year":"2003","journal-title":"Nucleic Acids Res"},{"key":"2023061011434056000_B7","doi-asserted-by":"crossref","first-page":"707","DOI":"10.1006\/jmbi.1998.2144","article-title":"Predicting function: from genes to genomes and back","volume":"283","author":"Bork","year":"1998","journal-title":"J. Mol. Biol"},{"key":"2023061011434056000_B8","doi-asserted-by":"crossref","first-page":"6073","DOI":"10.1073\/pnas.95.11.6073","article-title":"Assessing sequence comparison methods with reliable structurally identified distant evolutionary relationships","volume":"95","author":"Brenner","year":"1998","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"2023061011434056000_B9","doi-asserted-by":"crossref","first-page":"2044","DOI":"10.1093\/bioinformatics\/btl286","article-title":"Roundup: a multi-genome repository of orthologs and evolutionary distances","volume":"22","author":"Deluca","year":"2006","journal-title":"Bioinformatics"},{"key":"2023061011434056000_B10","doi-asserted-by":"crossref","first-page":"909","DOI":"10.1038\/nbt0704-909","article-title":"What is dynamic programming?","volume":"22","author":"Eddy","year":"2004","journal-title":"Nat. biotechnol"},{"key":"2023061011434056000_B11","doi-asserted-by":"crossref","first-page":"227","DOI":"10.1016\/S0168-9525(00)02005-9","article-title":"Homology a personal view on some of the problems","volume":"16","author":"Fitch","year":"2000","journal-title":"Trends Genet"},{"key":"2023061011434056000_B12","doi-asserted-by":"crossref","first-page":"573","DOI":"10.1038\/ng1579","article-title":"An adaptive radiation model for the origin of new gene functions","volume":"37","author":"Francino","year":"2005","journal-title":"Nat. Genet"},{"key":"2023061011434056000_B13","doi-asserted-by":"crossref","first-page":"270","DOI":"10.1186\/1471-2105-7-270","article-title":"Improving the specificity of high-throughput ortholog prediction","volume":"7","author":"Fulton","year":"2006","journal-title":"BMC Bioinformatics"},{"key":"2023061011434056000_B14","doi-asserted-by":"crossref","first-page":"49","DOI":"10.1016\/S1476-9271(02)00094-4","article-title":"Automated annotation of microbial proteomes in SWISS-PROT","volume":"27","author":"Gattiker","year":"2003","journal-title":"Comput. Biol. Chem"},{"key":"2023061011434056000_B15","doi-asserted-by":"crossref","first-page":"5392","DOI":"10.1093\/nar\/gkh882","article-title":"Conservation of adjacency as evidence of paralogous operons","volume":"32","author":"Janga","year":"2004","journal-title":"Nucleic Acids Res"},{"key":"2023061011434056000_B16","doi-asserted-by":"crossref","first-page":"540","DOI":"10.1007\/s002390010184","article-title":"The closest BLAST hit is often not the nearest neighbor","volume":"52","author":"Koski","year":"2001","journal-title":"J. Mol. Evol"},{"key":"2023061011434056000_B17","doi-asserted-by":"crossref","first-page":"126","DOI":"10.1093\/nar\/28.1.126","article-title":"NCBI's LocusLink and RefSeq","volume":"28","author":"Maglott","year":"2000","journal-title":"Nucleic Acids Res"},{"key":"2023061011434056000_B18","doi-asserted-by":"crossref","first-page":"S329","DOI":"10.1093\/bioinformatics\/18.suppl_1.S329","article-title":"A powerful non-homology method for the prediction of operons in prokaryotes","volume":"18","author":"Moreno-Hagelsieb","year":"2002","journal-title":"Bioinformatics"},{"key":"2023061011434056000_B19","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-86659-3","volume-title":"Evolution by Gene Duplication.","author":"Ohno","year":"1970"},{"key":"2023061011434056000_B20","doi-asserted-by":"crossref","DOI":"10.1186\/gb-2001-2-10-reviews2002","article-title":"Having a BLAST with bioinformatics (and avoiding BLASTphemy)","volume":"2","author":"Pertsemlidis","year":"2001","journal-title":"Genome Biol"},{"issue":"Database Issue","key":"2023061011434056000_B21","doi-asserted-by":"crossref","first-page":"D501","DOI":"10.1093\/nar\/gki025","article-title":"NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins","volume":"33","author":"Pruitt","year":"2005","journal-title":"Nucleic Acids Res"},{"key":"2023061011434056000_B22","doi-asserted-by":"crossref","first-page":"2994","DOI":"10.1093\/nar\/29.14.2994","article-title":"Improving the accuracy of PSI-BLAST protein database searches with composition-based statistics and other refinements","volume":"29","author":"Schaffer","year":"2001","journal-title":"Nucleic Acids Res"},{"key":"2023061011434056000_B23","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1016\/0022-2836(81)90087-5","article-title":"Identification of common molecular subsequences","volume":"147","author":"Smith","year":"1981","journal-title":"J. Mol. Biol"},{"key":"2023061011434056000_B24","doi-asserted-by":"crossref","first-page":"631","DOI":"10.1126\/science.278.5338.631","article-title":"A genomic perspective on protein families","volume":"278","author":"Tatusov","year":"1997","journal-title":"Science"},{"key":"2023061011434056000_B25","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1186\/1471-2105-4-41","article-title":"The cog database: an updated version includes eukaryotes","volume":"4","author":"Tatusov","year":"2003","journal-title":"BMC Bioinformatics"},{"key":"2023061011434056000_B26","doi-asserted-by":"crossref","first-page":"4673","DOI":"10.1093\/nar\/22.22.4673","article-title":"CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice","volume":"22","author":"Thompson","year":"1994","journal-title":"Nucleic Acids Res"},{"issue":"Database issue","key":"2023061011434056000_B27","doi-asserted-by":"crossref","first-page":"D358","DOI":"10.1093\/nar\/gkl825","article-title":"STRING 7\u2013recent developments in the integration and prediction of protein interactions","volume":"35","author":"von Mering","year":"2007","journal-title":"Nucleic Acids Res"},{"key":"2023061011434056000_B28","doi-asserted-by":"crossref","first-page":"1710","DOI":"10.1093\/bioinformatics\/btg213","article-title":"Detecting putative orthologs","volume":"19","author":"Wall","year":"2003","journal-title":"Bioinformatics"},{"key":"2023061011434056000_B29","first-page":"555","article-title":"PAML: a program package for phylogenetic analysis by maximum likelihood","volume":"13","author":"Yang","year":"1997","journal-title":"Comput. Appl. Biosci"},{"key":"2023061011434056000_B30","doi-asserted-by":"crossref","first-page":"1586","DOI":"10.1093\/molbev\/msm088","article-title":"PAML 4: Phylogenetic analysis by maximum likelihood","volume":"24","author":"Yang","year":"2007","journal-title":"Mol. Biol. Evol"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/3\/319\/50567814\/bioinformatics_24_3_319.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/3\/319\/50567814\/bioinformatics_24_3_319.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,10]],"date-time":"2023-06-10T11:44:24Z","timestamp":1686397464000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/24\/3\/319\/252715"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,11,26]]},"references-count":30,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2008,2,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btm585","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2008,2,1]]},"published":{"date-parts":[[2007,11,26]]}}}