{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T15:54:37Z","timestamp":1768319677269,"version":"3.49.0"},"reference-count":37,"publisher":"Oxford University Press (OUP)","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2008,3,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: As the amount of biological sequence data continues to grow exponentially we face the increasing challenge of assigning function to this enormous molecular \u2018parts list\u2019. The most popular approaches to this challenge make use of the simplifying assumption that similar functional molecules, or proteins, sometimes have similar composition, or sequence. However, these algorithms often fail to identify remote homologs (proteins with similar function but dissimilar sequence) which often are a significant fraction of the total homolog collection for a given sequence. We introduce a Support Vector Machine (SVM)-based tool to detect homology using semi-supervised iterative learning (SVM-HUSTLE) that identifies significantly more remote homologs than current state-of-the-art sequence or cluster-based methods. As opposed to building profiles or position specific scoring matrices, SVM-HUSTLE builds an SVM classifier for a query sequence by training on a collection of representative high-confidence training sets, recruits additional sequences and assigns a statistical measure of homology between a pair of sequences. SVM-HUSTLE combines principles of semi-supervised learning theory with statistical sampling to create many concurrent classifiers to iteratively detect and refine, on-the-fly, patterns indicating homology.<\/jats:p>\n               <jats:p>Results: When compared against existing methods for identifying protein homologs (BLAST, PSI-BLAST, COMPASS, PROF_SIM, RANKPROP and their variants) on two different benchmark datasets SVM-HUSTLE significantly outperforms each of the above methods using the most stringent ROC1 statistic with P-values less than 1e-20. SVM-HUSTLE also yields results comparable to HHSearch but at a substantially reduced computational cost since we do not require the construction of HMMs.<\/jats:p>\n               <jats:p>Availability: The software executable to run SVM-HUSTLE can be downloaded from http:\/\/www.sysbio.org\/sysbio\/networkbio\/svm_hustle<\/jats:p>\n               <jats:p>Contact: \u00a0anuj.shah@pnl.gov<\/jats:p>","DOI":"10.1093\/bioinformatics\/btn028","type":"journal-article","created":{"date-parts":[[2008,2,2]],"date-time":"2008-02-02T01:25:55Z","timestamp":1201915555000},"page":"783-790","source":"Crossref","is-referenced-by-count":34,"title":["SVM-HUSTLE\u2014an iterative semi-supervised machine learning approach for pairwise protein remote homology detection"],"prefix":"10.1093","volume":"24","author":[{"given":"Anuj R.","family":"Shah","sequence":"first","affiliation":[{"name":"1 Scientific Data Management and 2Computational Biology and Bioinformatics, Pacific Northwest National Laboratory, Richland, WA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christopher S.","family":"Oehmen","sequence":"additional","affiliation":[{"name":"1 Scientific Data Management and 2Computational Biology and Bioinformatics, Pacific Northwest National Laboratory, Richland, WA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bobbie-Jo","family":"Webb-Robertson","sequence":"additional","affiliation":[{"name":"1 Scientific Data Management and 2Computational Biology and Bioinformatics, Pacific Northwest National Laboratory, Richland, WA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2008,2,1]]},"reference":[{"key":"2023020209512662900_B1","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1016\/S0022-2836(05)80360-2","article-title":"A basic local alignment search tool","volume":"215","author":"Altschul","year":"1990","journal-title":"J. Mol. Biol"},{"key":"2023020209512662900_B2","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","article-title":"Gapped BLAST and PSI-BLAST: A new generation of protein database search programs","volume":"25","author":"Altschul","year":"1997","journal-title":"Nucl. Acids Res"},{"key":"2023020209512662900_B3","doi-asserted-by":"crossref","first-page":"1429","DOI":"10.1093\/bioinformatics\/bti212","article-title":"Implicit motif distribution based hybrid computational kernel for sequence classification","volume":"21","author":"Atalay","year":"2005","journal-title":"Bioinformatics"},{"key":"2023020209512662900_B4","doi-asserted-by":"crossref","first-page":"1059","DOI":"10.1073\/pnas.91.3.1059","article-title":"Hidden Markov models of biological primary sequence information","volume":"91","author":"Baldi","year":"1994","journal-title":"Proc. Natl Acad. Sci"},{"key":"2023020209512662900_B5","doi-asserted-by":"crossref","first-page":"i26","DOI":"10.1093\/bioinformatics\/btg1002","article-title":"Remote homology detection: a motif based approach","volume":"19","author":"Ben-Hur","year":"2003","journal-title":"Bioinformatics"},{"key":"2023020209512662900_B6","first-page":"191","article-title":"Support vector machines with profile-based kernels for remote protein homology detection","volume":"15","author":"Busuttil","year":"2004","journal-title":"Genome Informatics"},{"key":"2023020209512662900_B7","doi-asserted-by":"crossref","first-page":"6573","DOI":"10.1021\/bi012159+","article-title":"Intrinsic disorder and protein function","volume":"41","author":"Dunker","year":"2002","journal-title":"Biochemistry"},{"key":"2023020209512662900_B8","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1016\/S0097-8485(96)80004-0","article-title":"Use of receiver operating characteristic (ROC) analysis to evaluate sequence matching","volume":"20","author":"Gribskov","year":"1996","journal-title":"Comp Chem"},{"key":"2023020209512662900_B9","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1148\/radiology.143.1.7063747","article-title":"The meaning and use of the area under a receiver operating characteristic (ROC) curve","volume":"143","author":"Hanley","year":"1982","journal-title":"Radiology"},{"key":"2023020209512662900_B10","doi-asserted-by":"crossref","first-page":"2294","DOI":"10.1093\/bioinformatics\/btg317","article-title":"Efficient remote homology detection using local structure","volume":"19","author":"Hou","year":"2003","journal-title":"Bioinformatics"},{"key":"2023020209512662900_B11","doi-asserted-by":"crossref","first-page":"518","DOI":"10.1002\/prot.20221","article-title":"Remote homology detection using local sequence-structure correlations","volume":"57","author":"Hou","year":"2004","journal-title":"Proteins: Structure, Function and Bioinformatics"},{"key":"2023020209512662900_B12","doi-asserted-by":"crossref","first-page":"95","DOI":"10.1089\/10665270050081405","article-title":"A discriminative framework for detecting remote protein homologies","volume":"7","author":"Jaakkola","year":"2000","journal-title":"J. Comput. Biol"},{"key":"2023020209512662900_B13","doi-asserted-by":"crossref","first-page":"527","DOI":"10.1142\/S021972000500120X","article-title":"Profile-based string kernels for remote homology detection and motif extraction","volume":"3","author":"Kuang","year":"2005","journal-title":"J. Bioinform. Computat. Biol"},{"key":"2023020209512662900_B14","doi-asserted-by":"crossref","first-page":"3711","DOI":"10.1093\/bioinformatics\/bti608","article-title":"Motif-based protein ranking by network propagation","volume":"21","author":"Kuang","year":"2005","journal-title":"Bioinformatics"},{"key":"2023020209512662900_B15","first-page":"1","article-title":"Mismatch string kernels for discriminative protein classification","volume":"1","author":"Leslie","year":"2003","journal-title":"Bioinformatics"},{"key":"2023020209512662900_B16","doi-asserted-by":"crossref","first-page":"857","DOI":"10.1089\/106652703322756113","article-title":"Combining pairwise sequence similarity and support vector machines for detecting remote protein evolutionary and structural relationships","volume":"10","author":"Liao","year":"2003","journal-title":"J. Comput. Biol"},{"key":"2023020209512662900_B17","doi-asserted-by":"crossref","first-page":"2224","DOI":"10.1093\/bioinformatics\/btl376","article-title":"Remote homology detection based on oligomer distances","volume":"22","author":"Lingner","year":"2006","journal-title":"Bioinformatics"},{"key":"2023020209512662900_B18","first-page":"1202","article-title":"A discriminative method for remote homology detection based on n-peptide compositions with reduced amino acid alphabets","volume":"284","author":"Ogul","year":"2006","journal-title":"J. Mol. Biol"},{"key":"2023020209512662900_B19","doi-asserted-by":"crossref","first-page":"1202","DOI":"10.1006\/jmbi.1998.2221","article-title":"Sequence comparisons using multiple sequences detect three times as many remote homologues as pairwise methods","volume":"284","author":"Park","year":"1998","journal-title":"J. Mol. Biol"},{"key":"2023020209512662900_B20","doi-asserted-by":"crossref","first-page":"63","DOI":"10.1016\/0076-6879(90)83007-V","article-title":"Rapid and sensitive sequence comparisons with FASTP and FASTA","volume":"183","author":"Pearson","year":"1985","journal-title":"Methods Enzymol"},{"key":"2023020209512662900_B21","doi-asserted-by":"crossref","first-page":"4239","DOI":"10.1093\/bioinformatics\/bti687","article-title":"Profile based direct kernels for remote homology detection and fold recognition","volume":"21","author":"Rangwala","year":"2005","journal-title":"Bioinformatics"},{"key":"2023020209512662900_B22","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1093\/protein\/12.2.85","article-title":"Twilight zone of protein sequence alignments","volume":"12","author":"Rost","year":"1999","journal-title":"Protein Eng"},{"key":"2023020209512662900_B23","doi-asserted-by":"crossref","first-page":"2262","DOI":"10.1110\/ps.03197403","article-title":"Profile-profile comparisons by COMPASS predict intricate homologies between protein families","volume":"12","author":"Sadreyev","year":"2003","journal-title":"Protein Sci"},{"key":"2023020209512662900_B24","doi-asserted-by":"crossref","first-page":"317","DOI":"10.1023\/A:1009752403260","article-title":"On comparing classifiers: pitfalls to avoid and recommended approach","volume":"1","author":"Salzberg","year":"1997","journal-title":"Data Mining Knowledge Discovery"},{"key":"2023020209512662900_B25","doi-asserted-by":"crossref","first-page":"2994","DOI":"10.1093\/nar\/29.14.2994","article-title":"Improving the accuracy of PSI-BLAST protein database searches with composition-based statistics and other refinements","volume":"29","author":"Schaffer","year":"2001","journal-title":"Nucl. Acids Res"},{"key":"2023020209512662900_B26","doi-asserted-by":"crossref","first-page":"138","DOI":"10.1016\/j.compbiolchem.2007.02.012","article-title":"Integrating subcellular location for improving machine learning models of remote homology detection in eukaryotic organisms","volume":"31","author":"Shah","year":"2007","journal-title":"Comput Biol. Chem"},{"key":"2023020209512662900_B27","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1016\/0022-2836(81)90087-5","article-title":"Identification of common molecular subsequences","volume":"147","author":"Smith","year":"1981","journal-title":"J. Mol. Biol"},{"key":"2023020209512662900_B28","doi-asserted-by":"crossref","first-page":"951","DOI":"10.1093\/bioinformatics\/bti125","article-title":"Protein homology detection by HMM-HMM comparison","volume":"21","author":"Soeding","year":"2005","journal-title":"Bioinformatics"},{"key":"2023020209512662900_B29","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4757-2440-0","volume-title":"The nature of Statistical Learning Theory.","author":"Vapnik","year":"1995"},{"key":"2023020209512662900_B30","volume-title":"Statistical Learning Theory. Adaptive and Learning Systems for Signal Processing, Communications, and Control.","author":"Vapnik","year":"1998"},{"key":"2023020209512662900_B31","doi-asserted-by":"crossref","first-page":"440","DOI":"10.1016\/j.compbiolchem.2005.09.006","article-title":"SVM-BALSA: remote homology detection based on Bayesian sequence alignment","volume":"29","author":"Webb-Robertson","year":"2005","journal-title":"Comput. Biol. Chem"},{"key":"2023020209512662900_B32","doi-asserted-by":"crossref","first-page":"6559","DOI":"10.1073\/pnas.0308067101","article-title":"Protein ranking: from local to global structure in the protein similarity network","volume":"101","author":"Weston","year":"2004","journal-title":"Proc. Natl Acad. Sci"},{"key":"2023020209512662900_B33","doi-asserted-by":"crossref","first-page":"S10","DOI":"10.1186\/1471-2105-7-S1-S10","article-title":"Protein ranking by semi-supervised network propagation","volume":"7","author":"Weston","year":"2006","journal-title":"BMC Bioinformatics"},{"key":"2023020209512662900_B34","doi-asserted-by":"crossref","first-page":"3241","DOI":"10.1093\/bioinformatics\/bti497","article-title":"Semi-supervised protein classification using cluster kernels","volume":"21","author":"Weston","year":"2005","journal-title":"Bioinformatics"},{"key":"2023020209512662900_B35","doi-asserted-by":"crossref","first-page":"1257","DOI":"10.1006\/jmbi.2001.5293","article-title":"Within the twilight zone: a sensitive profile-profile comparison tool based on information theory","volume":"315","author":"Yona","year":"2002","journal-title":"J. Mol. Biol"},{"key":"2023020209512662900_B36","article-title":"A comparative analysis of protein homology detection methods","volume":"5","author":"Zaki","year":"2003","journal-title":"J. Theor"},{"key":"2023020209512662900_B37","volume-title":"Semi-supervised Learning Literature Survey.","author":"Zhu","year":"2006"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/6\/783\/49046720\/bioinformatics_24_6_783.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/6\/783\/49046720\/bioinformatics_24_6_783.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,2]],"date-time":"2023-02-02T10:47:07Z","timestamp":1675334827000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/24\/6\/783\/193709"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,2,1]]},"references-count":37,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2008,3,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btn028","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2008,3,15]]},"published":{"date-parts":[[2008,2,1]]}}}