{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T07:41:04Z","timestamp":1781854864710,"version":"3.54.5"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2010,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>Phylogenetic analysis can be used to divide a protein family into subfamilies in the absence of experimental information. Most phylogenetic analysis methods utilize multiple alignment of sequences and are based on an evolutionary model. However, multiple alignment is not an automated procedure and requires human intervention to maintain alignment integrity and to produce phylogenies consistent with the functional splits in underlying sequences. To address this problem, we propose to use the alignment-free Relative Complexity Measure (RCM) combined with reduced amino acid alphabets to cluster protein families into functional subtypes purely on sequence criteria. Comparison with an alignment-based approach was also carried out to test the quality of the clustering.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>We demonstrate the robustness of RCM with reduced alphabets in clustering of protein sequences into families in a simulated dataset and seven well-characterized protein datasets. On protein datasets, crotonases, mandelate racemases, nucleotidyl cyclases and glycoside hydrolase family 2 were clustered into subfamilies with 100% accuracy whereas acyl transferase domains, haloacid dehalogenases, and vicinal oxygen chelates could be assigned to subfamilies with 97.2%, 96.9% and 92.2% accuracies, respectively.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusions<\/jats:title>\n            <jats:p>The overall combination of methods in this paper is useful for clustering protein families into subtypes based on solely protein sequence information. The method is also flexible and computationally fast because it does not require multiple alignment of sequences.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-11-428","type":"journal-article","created":{"date-parts":[[2010,8,18]],"date-time":"2010-08-18T18:18:58Z","timestamp":1282155538000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":20,"title":["Clustering of protein families into functional subtypes using Relative Complexity Measure with reduced amino acid alphabets"],"prefix":"10.1186","volume":"11","author":[{"given":"Aydin","family":"Albayrak","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hasan H","family":"Otu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ugur O","family":"Sezerman","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2010,8,18]]},"reference":[{"key":"3885_CR1","doi-asserted-by":"publisher","first-page":"135","DOI":"10.1186\/1471-2105-8-135","volume":"8","author":"IM Wallace","year":"2007","unstructured":"Wallace IM, Higgins DG: Supervised multivariate analysis of sequence groups to identify specificity determining residues. BMC Bioinformatics 2007, 8: 135. 10.1186\/1471-2105-8-135","journal-title":"BMC Bioinformatics"},{"key":"3885_CR2","doi-asserted-by":"publisher","first-page":"68","DOI":"10.1186\/1472-6807-9-68","volume":"9","author":"B Georgi","year":"2009","unstructured":"Georgi B, Schultz J, Schliep A: Partially-supervised protein subclass discovery with simultaneous annotation of functional residues. BMC Struct Biol 2009, 9: 68. 10.1186\/1472-6807-9-68","journal-title":"BMC Struct Biol"},{"key":"3885_CR3","doi-asserted-by":"publisher","first-page":"286","DOI":"10.1186\/1471-2105-8-286","volume":"8","author":"A Kelil","year":"2007","unstructured":"Kelil A, Wang S, Brzezinski R, Fleury A: CLUSS: clustering of protein sequences based on a new similarity measure. BMC Bioinformatics 2007, 8: 286. 10.1186\/1471-2105-8-286","journal-title":"BMC Bioinformatics"},{"issue":"9","key":"3885_CR4","doi-asserted-by":"publisher","first-page":"1876","DOI":"10.1093\/bioinformatics\/bti244","volume":"21","author":"B Lazareva-Ulitsky","year":"2005","unstructured":"Lazareva-Ulitsky B, Diemer K, Thomas PD: On the quality of tree-based protein classification. Bioinformatics 2005, 21(9):1876\u20131890. 10.1093\/bioinformatics\/bti244","journal-title":"Bioinformatics"},{"issue":"8","key":"3885_CR5","doi-asserted-by":"publisher","first-page":"1435","DOI":"10.1093\/oxfordjournals.molbev.a003929","volume":"18","author":"N Wicker","year":"2001","unstructured":"Wicker N, Perrin GR, Thierry JC, Poch O: Secator: a program for inferring protein subfamilies from phylogenetic trees. Mol Biol Evol 2001, 18(8):1435\u20131441.","journal-title":"Mol Biol Evol"},{"issue":"1","key":"3885_CR6","doi-asserted-by":"publisher","first-page":"27","DOI":"10.1006\/tpbi.2000.1485","volume":"59","author":"L Brocchieri","year":"2001","unstructured":"Brocchieri L: Phylogenetic inferences from molecular sequences: review and critique. Theor Popul Biol 2001, 59(1):27\u201340. 10.1006\/tpbi.2000.1485","journal-title":"Theor Popul Biol"},{"issue":"6","key":"3885_CR7","doi-asserted-by":"publisher","first-page":"345","DOI":"10.1016\/S0168-9525(03)00112-4","volume":"19","author":"SL Baldauf","year":"2003","unstructured":"Baldauf SL: Phylogeny for the faint of heart: a tutorial. Trends Genet 2003, 19(6):345\u2013351. 10.1016\/S0168-9525(03)00112-4","journal-title":"Trends Genet"},{"issue":"16","key":"3885_CR8","doi-asserted-by":"publisher","first-page":"2122","DOI":"10.1093\/bioinformatics\/btg295","volume":"19","author":"HH Otu","year":"2003","unstructured":"Otu HH, Sayood K: A new sequence distance measure for phylogenetic tree construction. Bioinformatics 2003, 19(16):2122\u20132130. 10.1093\/bioinformatics\/btg295","journal-title":"Bioinformatics"},{"issue":"6","key":"3885_CR9","doi-asserted-by":"publisher","first-page":"368","DOI":"10.1007\/BF01734359","volume":"17","author":"J Felsenstein","year":"1981","unstructured":"Felsenstein J: Evolutionary trees from DNA sequences: a maximum likelihood approach. J Mol Evol 1981, 17(6):368\u2013376. 10.1007\/BF01734359","journal-title":"J Mol Evol"},{"key":"3885_CR10","doi-asserted-by":"publisher","first-page":"371","DOI":"10.1146\/annurev.genet.30.1.371","volume":"30","author":"M Nei","year":"1996","unstructured":"Nei M: Phylogenetic analysis in molecular evolutionary genetics. Annu Rev Genet 1996, 30: 371\u2013403. 10.1146\/annurev.genet.30.1.371","journal-title":"Annu Rev Genet"},{"issue":"1","key":"3885_CR11","doi-asserted-by":"publisher","first-page":"61","DOI":"10.1006\/jmbi.2000.4036","volume":"303","author":"SS Hannenhalli","year":"2000","unstructured":"Hannenhalli SS, Russell RB: Analysis and prediction of functional sub-types from protein sequence alignments. J Mol Biol 2000, 303(1):61\u201376. 10.1006\/jmbi.2000.4036","journal-title":"J Mol Biol"},{"issue":"8","key":"3885_CR12","doi-asserted-by":"publisher","first-page":"e160","DOI":"10.1371\/journal.pcbi.0030160","volume":"3","author":"DP Brown","year":"2007","unstructured":"Brown DP, Krishnamurthy N, Sjolander K: Automated protein subfamily identification and classification. PLoS Comput Biol 2007, 3(8):e160. 10.1371\/journal.pcbi.0030160","journal-title":"PLoS Comput Biol"},{"key":"3885_CR13","doi-asserted-by":"publisher","first-page":"337","DOI":"10.1109\/TIT.1977.1055714","volume":"23","author":"J Ziv","year":"1977","unstructured":"Ziv J, Lempel A: A universal algorithm for sequential data compression. IEEE Trans Inf Theory 1977, 23: 337\u2013343. 10.1109\/TIT.1977.1055714","journal-title":"IEEE Trans Inf Theory"},{"issue":"Pt 2","key":"3885_CR14","doi-asserted-by":"publisher","first-page":"117","DOI":"10.1017\/S0953756203009079","volume":"108","author":"DR Bastola","year":"2004","unstructured":"Bastola DR, Otu HH, Doukas SE, Sayood K, Hinrichs SH, Iwen PC: Utilization of the relative complexity measure to construct a phylogenetic tree for fungi. Mycol Res 2004, 108(Pt 2):117\u2013125. 10.1017\/S0953756203009079","journal-title":"Mycol Res"},{"issue":"22","key":"3885_CR15","doi-asserted-by":"publisher","first-page":"5321","DOI":"10.1016\/j.febslet.2006.08.086","volume":"580","author":"N Liu","year":"2006","unstructured":"Liu N, Wang T: Protein-based phylogenetic analysis by using hydropathy profile of amino acids. FEBS Lett 2006, 580(22):5321\u20135327. 10.1016\/j.febslet.2006.08.086","journal-title":"FEBS Lett"},{"key":"3885_CR16","doi-asserted-by":"publisher","first-page":"306","DOI":"10.1186\/1471-2105-9-306","volume":"9","author":"DJ Russell","year":"2008","unstructured":"Russell DJ, Otu HH, Sayood K: Grammar-based distance in progressive multiple sequence alignment. BMC Bioinformatics 2008, 9: 306. 10.1186\/1471-2105-9-306","journal-title":"BMC Bioinformatics"},{"issue":"11","key":"3885_CR17","doi-asserted-by":"publisher","first-page":"1033","DOI":"10.1038\/14918","volume":"6","author":"J Wang","year":"1999","unstructured":"Wang J, Wang W: A computational approach to simplifying the protein folding alphabet. Nat Struct Biol 1999, 6(11):1033\u20131038. 10.1038\/14918","journal-title":"Nat Struct Biol"},{"issue":"8","key":"3885_CR18","doi-asserted-by":"publisher","first-page":"1059","DOI":"10.1007\/s00249-007-0188-5","volume":"36","author":"C Etchebest","year":"2007","unstructured":"Etchebest C, Benros C, Bornot A, Camproux AC, de Brevern AG: A reduced amino acid alphabet for understanding and designing protein adaptation to mutation. Eur Biophys J 2007, 36(8):1059\u20131069. 10.1007\/s00249-007-0188-5","journal-title":"Eur Biophys J"},{"issue":"5","key":"3885_CR19","doi-asserted-by":"publisher","first-page":"323","DOI":"10.1093\/protein\/gzg044","volume":"16","author":"T Li","year":"2003","unstructured":"Li T, Fan K, Wang J, Wang W: Reduction of protein sequence complexity by residue grouping. Protein Eng 2003, 16(5):323\u2013330. 10.1093\/protein\/gzg044","journal-title":"Protein Eng"},{"issue":"8","key":"3885_CR20","doi-asserted-by":"publisher","first-page":"1879","DOI":"10.1093\/molbev\/msp098","volume":"26","author":"W Fletcher","year":"2009","unstructured":"Fletcher W, Yang Z: INDELible: a flexible simulator of biological sequence evolution. Mol Biol Evol 2009, 26(8):1879\u20131888. 10.1093\/molbev\/msp098","journal-title":"Mol Biol Evol"},{"issue":"2","key":"3885_CR21","doi-asserted-by":"publisher","first-page":"193","DOI":"10.1093\/molbev\/msi005","volume":"22","author":"C Kosiol","year":"2005","unstructured":"Kosiol C, Goldman N: Different versions of the Dayhoff rate matrix. Mol Biol Evol 2005, 22(2):193\u2013199. 10.1093\/molbev\/msi005","journal-title":"Mol Biol Evol"},{"issue":"21","key":"3885_CR22","doi-asserted-by":"publisher","first-page":"2947","DOI":"10.1093\/bioinformatics\/btm404","volume":"23","author":"MA Larkin","year":"2007","unstructured":"Larkin MA, Blackshields G, Brown NP, Chenna R, McGettigan PA, McWilliam H, Valentin F, Wallace IM, Wilm A, Lopez R, et al.: Clustal W and Clustal X version 2.0. Bioinformatics 2007, 23(21):2947\u20132948. 10.1093\/bioinformatics\/btm404","journal-title":"Bioinformatics"},{"issue":"9","key":"3885_CR23","doi-asserted-by":"publisher","first-page":"755","DOI":"10.1093\/bioinformatics\/14.9.755","volume":"14","author":"SR Eddy","year":"1998","unstructured":"Eddy SR: Profile hidden Markov models. Bioinformatics 1998, 14(9):755\u2013763. 10.1093\/bioinformatics\/14.9.755","journal-title":"Bioinformatics"},{"issue":"8","key":"3885_CR24","doi-asserted-by":"publisher","first-page":"2545","DOI":"10.1021\/bi052101l","volume":"45","author":"SC Pegg","year":"2006","unstructured":"Pegg SC, Brown SD, Ojha S, Seffernick J, Meng EC, Morris JH, Chang PJ, Huang CC, Ferrin TE, Babbitt PC: Leveraging enzyme structure-function relationships for functional inference and experimental design: the structure-function linkage database. Biochemistry (Mosc) 2006, 45(8):2545\u20132555. 10.1021\/bi052101l","journal-title":"Biochemistry (Mosc)"},{"key":"3885_CR25","doi-asserted-by":"publisher","first-page":"335","DOI":"10.1186\/1471-2105-10-335","volume":"10","author":"P Goldstein","year":"2009","unstructured":"Goldstein P, Zucko J, Vujaklija D, Krisko A, Hranueli D, Long PF, Etchebest C, Basrak B, Cullum J: Clustering of protein domains for functional and evolutionary studies. BMC Bioinformatics 2009, 10: 335. 10.1186\/1471-2105-10-335","journal-title":"BMC Bioinformatics"},{"issue":"6","key":"3885_CR26","doi-asserted-by":"publisher","first-page":"625","DOI":"10.1007\/BF00160408","volume":"39","author":"VB Strelets","year":"1994","unstructured":"Strelets VB, Shindyalov IN, Lim HA: Analysis of peptides from known proteins: clusterization in sequence space. J Mol Evol 1994, 39(6):625\u2013630. 10.1007\/BF00160408","journal-title":"J Mol Evol"},{"issue":"6","key":"3885_CR27","doi-asserted-by":"publisher","first-page":"1501","DOI":"10.1021\/bi00327a032","volume":"24","author":"KA Dill","year":"1985","unstructured":"Dill KA: Theory for the folding and stability of globular proteins. Biochemistry (Mosc) 1985, 24(6):1501\u20131509. 10.1021\/bi00327a032","journal-title":"Biochemistry (Mosc)"},{"issue":"3","key":"3885_CR28","doi-asserted-by":"publisher","first-page":"149","DOI":"10.1093\/protein\/13.3.149","volume":"13","author":"LR Murphy","year":"2000","unstructured":"Murphy LR, Wallqvist A, Levy RM: Simplified amino acid alphabets for protein fold recognition and implications for folding. Protein Eng 2000, 13(3):149\u2013152. 10.1093\/protein\/13.3.149","journal-title":"Protein Eng"},{"issue":"8","key":"3885_CR29","doi-asserted-by":"publisher","first-page":"545","DOI":"10.1093\/protein\/13.8.545","volume":"13","author":"A Prlic","year":"2000","unstructured":"Prlic A, Domingues FS, Sippl MJ: Structure-derived substitution matrices for alignment of distantly related sequences. Protein Eng 2000, 13(8):545\u2013550. 10.1093\/protein\/13.8.545","journal-title":"Protein Eng"},{"issue":"2","key":"3885_CR30","doi-asserted-by":"publisher","first-page":"149","DOI":"10.1002\/(SICI)1097-0134(20000201)38:2<149::AID-PROT4>3.0.CO;2-#","volume":"38","author":"AD Solis","year":"2000","unstructured":"Solis AD, Rackovsky S: Optimized representations and maximal information in proteins. Proteins 2000, 38(2):149\u2013164. 10.1002\/(SICI)1097-0134(20000201)38:2<149::AID-PROT4>3.0.CO;2-#","journal-title":"Proteins"},{"issue":"5","key":"3885_CR31","doi-asserted-by":"publisher","first-page":"311","DOI":"10.1093\/protein\/gzn007","volume":"21","author":"E Munoz","year":"2008","unstructured":"Munoz E, Deem MW: Amino acid alphabet size in protein evolution experiments: better to search a small library thoroughly or a large library sparsely? Protein Eng Des Sel 2008, 21(5):311\u2013317. 10.1093\/protein\/gzn007","journal-title":"Protein Eng Des Sel"},{"issue":"10","key":"3885_CR32","doi-asserted-by":"publisher","first-page":"3986","DOI":"10.1021\/ma00200a030","volume":"22","author":"KF Lau","year":"1989","unstructured":"Lau KF, Dill KA: A lattice statistical mechanics model of the conformational and sequence spaces of proteins. Macromolecules 1989, 22(10):3986\u20133997. 10.1021\/ma00200a030","journal-title":"Macromolecules"},{"issue":"11","key":"3885_CR33","doi-asserted-by":"publisher","first-page":"1356","DOI":"10.1093\/bioinformatics\/btp164","volume":"25","author":"EL Peterson","year":"2009","unstructured":"Peterson EL, Kondev J, Theriot JA, Phillips R: Reduced amino acid alphabets exhibit an improved sensitivity and selectivity in fold assignment. Bioinformatics 2009, 25(11):1356\u20131362. 10.1093\/bioinformatics\/btp164","journal-title":"Bioinformatics"},{"issue":"1","key":"3885_CR34","doi-asserted-by":"publisher","first-page":"75","DOI":"10.1109\/TIT.1976.1055501","volume":"22","author":"A Lempel","year":"1976","unstructured":"Lempel A, Ziv J: On the Complexity of Finite Sequences. IEEE Trans Inf Theory 1976, 22(1):75\u201381. 10.1109\/TIT.1976.1055501","journal-title":"IEEE Trans Inf Theory"},{"issue":"4","key":"3885_CR35","first-page":"406","volume":"4","author":"N Saitou","year":"1987","unstructured":"Saitou N, Nei M: The neighbor-joining method: a new method for reconstructing phylogenetic trees. Mol Biol Evol 1987, 4(4):406\u2013425.","journal-title":"Mol Biol Evol"},{"key":"3885_CR36","first-page":"164","volume":"5","author":"J Felsenstein","year":"1989","unstructured":"Felsenstein J: PHYLIP - Phylogeny Inference Package (Version 3.2). Cladistics 1989, 5: 164\u2013166.","journal-title":"Cladistics"},{"issue":"2","key":"3885_CR37","doi-asserted-by":"publisher","first-page":"241","DOI":"10.1214\/ss\/1063994979","volume":"18","author":"S Holmes","year":"2003","unstructured":"Holmes S: Bootstrapping Phylogenetic Trees: Theory and Methods. Stat Sci 2003, 18(2):241\u2013255. 10.1214\/ss\/1063994979","journal-title":"Stat Sci"},{"key":"3885_CR38","doi-asserted-by":"publisher","first-page":"209","DOI":"10.1146\/annurev.biochem.70.1.209","volume":"70","author":"JA Gerlt","year":"2001","unstructured":"Gerlt JA, Babbitt PC: Divergent evolution of enzymatic function: mechanistically diverse superfamilies and functionally distinct suprafamilies. Annu Rev Biochem 2001, 70: 209\u2013246. 10.1146\/annurev.biochem.70.1.209","journal-title":"Annu Rev Biochem"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-11-428.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T05:19:36Z","timestamp":1630473576000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-11-428"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2010,8,18]]},"references-count":38,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2010,12]]}},"alternative-id":["3885"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-11-428","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2010,8,18]]},"assertion":[{"value":"21 January 2010","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 August 2010","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 August 2010","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"428"}}