{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,3,19]],"date-time":"2025-03-19T14:33:35Z","timestamp":1742394815107},"reference-count":25,"publisher":"Oxford University Press (OUP)","issue":"20","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2006,10,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: The rapid increase in the amount of protein sequence data has created a need for an automated identification of evolutionarily related subgroups from large datasets. The existing methods typically require a priori specification of the number of putative groups, which defines the resolution of the classification solution.<\/jats:p>\n               <jats:p>Results: We introduce a Bayesian model-based approach to simultaneous identification of evolutionary groups and conserved parts of the protein sequences. The model-based approach provides an intuitive and efficient way of determining the number of groups from the sequence data, in contrast to the ad hoc methods often exploited for similar purposes. Our model recognizes the areas in the sequences that are relevant for the clustering and regards other areas as noise. We have implemented the method using a fast stochastic optimization algorithm which yields a clustering associated with the estimated maximum posterior probability. The method has been shown to have high specificity and sensitivity in simulated and real clustering tasks. With real datasets the method also highlights the residues close to the active site.<\/jats:p>\n               <jats:p>Availability: Software \u2018kPax\u2019 is available at<\/jats:p>\n               <jats:p>Contact: \u00a0pekka.marttinen@helsinki.fi<\/jats:p>\n               <jats:p>Supplementary information: \u00a0<\/jats:p>","DOI":"10.1093\/bioinformatics\/btl411","type":"journal-article","created":{"date-parts":[[2006,7,27]],"date-time":"2006-07-27T00:58:50Z","timestamp":1153961930000},"page":"2466-2474","source":"Crossref","is-referenced-by-count":37,"title":["Bayesian search of functionally divergent protein subgroups and their function specific residues"],"prefix":"10.1093","volume":"22","author":[{"given":"Pekka","family":"Marttinen","sequence":"first","affiliation":[{"name":"Department of Mathematics and Statistics, PO Box 68, 00014 University of Helsinki 1 \u00a0 1 \u00a0 \u00a0 Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jukka","family":"Corander","sequence":"additional","affiliation":[{"name":"Department of Mathematics and Statistics, PO Box 68, 00014 University of Helsinki 1 \u00a0 1 \u00a0 \u00a0 Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Petri","family":"T\u00f6r\u00f6nen","sequence":"additional","affiliation":[{"name":"Institute of Biotechnology, PO Box 56, 00014 University of Helsinki 2 \u00a0 2 \u00a0 \u00a0 Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liisa","family":"Holm","sequence":"additional","affiliation":[{"name":"Institute of Biotechnology, PO Box 56, 00014 University of Helsinki 2 \u00a0 2 \u00a0 \u00a0 Finland"},{"name":"Department of Biological and Environmental Sciences, PO Box 56, 00014 University of Helsinki 3 \u00a0 3 \u00a0 \u00a0 Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2006,7,26]]},"reference":[{"key":"2023012409213792800_b1","doi-asserted-by":"crossref","first-page":"276","DOI":"10.1093\/nar\/30.1.276","article-title":"The Pfam protein families database","volume":"30","author":"Bateman","year":"2002","journal-title":"Nucleic Acids Res."},{"key":"2023012409213792800_b2","doi-asserted-by":"crossref","DOI":"10.1002\/9780470316870","volume-title":"Bayesian Theory","author":"Bernardo","year":"1994"},{"key":"2023012409213792800_b3","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1038\/nsb0295-171","article-title":"A method to predict functional residues in proteins","volume":"2","author":"Casari","year":"1995","journal-title":"Nat. Struct. Biol."},{"key":"2023012409213792800_b4","doi-asserted-by":"crossref","first-page":"2363","DOI":"10.1093\/bioinformatics\/bth250","article-title":"BAPS 2: enhanced possibilities for the analysis of genetic population structure","volume":"20","author":"Corander","year":"2004","journal-title":"Bioinformatics"},{"key":"2023012409213792800_b5","article-title":"Bayesian identification of stock mixtures from molecular marker data","volume-title":"Fish. Bull.","author":"Corander","year":"2006"},{"key":"2023012409213792800_b6","article-title":"Bayesian unsupervised classification algorithms based on parallel search strategy","volume-title":"Pattern Recogn.","author":"Corander","year":"2006"},{"key":"2023012409213792800_b7","doi-asserted-by":"crossref","first-page":"2629","DOI":"10.1093\/bioinformatics\/bti396","article-title":"Determining functional specificity from protein sequences","volume":"21","author":"Donald","year":"2005","journal-title":"Bioinformatics"},{"key":"2023012409213792800_b8","doi-asserted-by":"crossref","first-page":"4455","DOI":"10.1093\/nar\/gki755","article-title":"Predicting specificity-determining residues in two large eukaryotic transcription factor families","volume":"33","author":"Donald","year":"2005","journal-title":"Nucleic Acids Res."},{"key":"2023012409213792800_b9","doi-asserted-by":"crossref","first-page":"i130","DOI":"10.1093\/bioinformatics\/btg1017","article-title":"Sensitive pattern discovery with \u2018fuzzy\u2019 alignments of distantly related proteins","volume":"19","author":"Heger","year":"2003","journal-title":"Bioinformatics"},{"key":"2023012409213792800_b10","doi-asserted-by":"crossref","first-page":"848","DOI":"10.1089\/cmb.2004.11.843","article-title":"Accurate detection of very sparse motifs","volume":"11","author":"Heger","year":"2004","journal-title":"J. Comput. Biol."},{"key":"2023012409213792800_b11","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1002\/(SICI)1097-0134(199705)28:1<72::AID-PROT7>3.0.CO;2-L","article-title":"An evolutionary treasure: unification of a broad set of amidohydrolases related to urease","volume":"28","author":"Holm","year":"1998","journal-title":"Proteins"},{"key":"2023012409213792800_b12","doi-asserted-by":"crossref","first-page":"193","DOI":"10.1007\/BF01908075","article-title":"Comparing partitions","author":"Hubert","year":"1985","journal-title":"J. Classif."},{"key":"2023012409213792800_b13","doi-asserted-by":"crossref","first-page":"773","DOI":"10.1080\/01621459.1995.10476572","article-title":"Bayes factors","volume":"90","author":"Kass","year":"1995","journal-title":"J. Am. Stat. Assoc."},{"key":"2023012409213792800_b14","doi-asserted-by":"crossref","first-page":"1154","DOI":"10.1109\/TPAMI.2004.71","article-title":"Simultaneous feature selections and clustering using mixture models","volume":"26","author":"Law","year":"2004","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2023012409213792800_b15","doi-asserted-by":"crossref","first-page":"159","DOI":"10.1023\/A:1026115125950","article-title":"Accurate and scalable identification of functional sites by evolutionary tracing","volume":"4","author":"Lichtarge","year":"2003","journal-title":"J. Struct. Funct. Genomics"},{"key":"2023012409213792800_b16","doi-asserted-by":"crossref","first-page":"1265","DOI":"10.1016\/j.jmb.2003.12.078","article-title":"A family of evolution-entropy hybrid methods for ranking protein residues by importance","volume":"336","author":"Mihalek","year":"2004","journal-title":"J. Mol. Biol."},{"key":"2023012409213792800_b17","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1016\/S0022-2836(02)00587-9","article-title":"Using orthologous and paralogous proteins to identify specificity-determining residues in bacterial transcription factors","volume":"321","author":"Mirny","year":"2002","journal-title":"J. Mol. Biol."},{"key":"2023012409213792800_b18","doi-asserted-by":"crossref","first-page":"90","DOI":"10.1145\/1007730.1007731","article-title":"Subspace clustering for high dimensional data: a review","volume":"6","author":"Parsons","year":"2004","journal-title":"ACM SIGKDD Explor. Newslett."},{"key":"2023012409213792800_b19","first-page":"131","article-title":"Bayesian protein family classifier","volume":"6","author":"Qu","year":"1998","journal-title":"Proc. Int. Conf. Intell. Syst. Mol. Biol."},{"key":"2023012409213792800_b20","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511812651","volume-title":"Pattern Recognition and Neural Networks","author":"Ripley","year":"1996"},{"key":"2023012409213792800_b21","volume-title":"Monte Carlo Statistical Methods","author":"Robert","year":"2005","edition":"2nd edn"},{"key":"2023012409213792800_b22","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1186\/1471-2105-6-82","article-title":"Super paramagnetic clustering of protein sequences","volume":"6","author":"Tetko","year":"2005","journal-title":"BMC Bioinformatics"},{"key":"2023012409213792800_b23","doi-asserted-by":"crossref","first-page":"4673","DOI":"10.1093\/nar\/22.22.4673","article-title":"CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice","volume":"22","author":"Thompson","year":"1994","journal-title":"Nucleic Acids Res."},{"key":"2023012409213792800_b24","first-page":"599","article-title":"Model-based hierarchical clustering","volume-title":"Proceedings of the 16th Annual Conference on Uncertainty in Artificial Intelligence (UAI-00)","author":"Vaithyanathan","year":"2000"},{"key":"2023012409213792800_b25","doi-asserted-by":"crossref","first-page":"379","DOI":"10.1006\/jtbi.2000.2138","article-title":"Information content of protein sequences","volume":"206","author":"Weiss","year":"2000","journal-title":"J. Theoref. Biol."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/20\/2466\/48841661\/bioinformatics_22_20_2466.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/20\/2466\/48841661\/bioinformatics_22_20_2466.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,24]],"date-time":"2023-01-24T10:00:35Z","timestamp":1674554435000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/22\/20\/2466\/218114"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,7,26]]},"references-count":25,"journal-issue":{"issue":"20","published-print":{"date-parts":[[2006,10,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btl411","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2006,10,15]]},"published":{"date-parts":[[2006,7,26]]}}}