{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T07:45:39Z","timestamp":1778053539069,"version":"3.51.4"},"reference-count":33,"publisher":"Oxford University Press (OUP)","issue":"2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2006,1,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Various multiple sequence alignment-based methods have been proposed to detect functional surfaces in proteins, such as active sites or protein interfaces. The effect that the choice of sequences has on the conclusions of such analysis has seldom been discussed. In particular, no method has been discussed in terms of its ability to optimize the sequence selection for the reliable detection of functional surfaces.<\/jats:p>\n               <jats:p>Results: Here we propose, for the case of proteins with known structure, a heuristic Metropolis Monte Carlo strategy to select sequences from a large set of homologues, in order to improve detection of functional surfaces. The quantity guiding the optimization is the clustering of residues which are under increased evolutionary pressure, according to the sample of sequences under consideration. We show that we can either improve the overlap of our prediction with known functional surfaces in comparison with the sequence similarity criteria of selection or match the quality of prediction obtained through more elaborate non-structure based-methods of sequence selection. For the purpose of demonstration we use a set of 50 homodimerizing enzymes which were co-crystallized with their substrates and cofactors.<\/jats:p>\n               <jats:p>Contact: \u00a0imihalek@bcm.tmc.edu<\/jats:p>\n               <jats:p>Supplementary information: Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/bti791","type":"journal-article","created":{"date-parts":[[2005,11,23]],"date-time":"2005-11-23T01:24:56Z","timestamp":1132709096000},"page":"149-156","source":"Crossref","is-referenced-by-count":17,"title":["A structure and evolution-guided Monte Carlo sequence selection strategy for multiple alignment-based analysis of proteins"],"prefix":"10.1093","volume":"22","author":[{"given":"I.","family":"Mihalek","sequence":"first","affiliation":[{"name":"Department of Molecular and Human Genetics, Baylor College of Medicine \u00a0 One Baylor Plaza, Houston, TX 77030, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"I.","family":"Re\u0161","sequence":"additional","affiliation":[{"name":"Department of Molecular and Human Genetics, Baylor College of Medicine \u00a0 One Baylor Plaza, Houston, TX 77030, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"O.","family":"Lichtarge","sequence":"additional","affiliation":[{"name":"Department of Molecular and Human Genetics, Baylor College of Medicine \u00a0 One Baylor Plaza, Houston, TX 77030, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2005,11,22]]},"reference":[{"key":"2023012408305755600_b1","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","article-title":"Gapped blast and psi-blast: a new generation of protein database search programs","volume":"25","author":"Altschul","year":"1997","journal-title":"Nucleic Acids Res."},{"key":"2023012408305755600_b2","doi-asserted-by":"crossref","first-page":"105","DOI":"10.1016\/S0022-2836(02)01036-7","article-title":"Analysis of catalytic residues in enzyme active sites","volume":"324","author":"Bartlett","year":"2002","journal-title":"J. Mol. Biol."},{"key":"2023012408305755600_b3","doi-asserted-by":"crossref","first-page":"235","DOI":"10.1093\/nar\/28.1.235","article-title":"The protein data bank","volume":"28","author":"Berman","year":"2000","journal-title":"Nucleic Acids Res."},{"key":"2023012408305755600_b4","doi-asserted-by":"crossref","first-page":"1487","DOI":"10.1093\/bioinformatics\/bti242","article-title":"Improved prediction of protein\u2013protein binding sites using a support vector machines approach","volume":"21","author":"Bradford","year":"2005","journal-title":"Bioinformatics"},{"key":"2023012408305755600_b5","doi-asserted-by":"crossref","first-page":"190","DOI":"10.1110\/ps.03323604","article-title":"Are protein\u2013protein interfaces more conserved in sequence than the rest of the protein surface?","volume":"13","author":"Caffrey","year":"2004","journal-title":"Protein Sci."},{"key":"2023012408305755600_b6","doi-asserted-by":"crossref","first-page":"2990","DOI":"10.1073\/pnas.061411798","article-title":"Identification of protein oligomerization states by analysis of interface conservation","volume":"98","author":"Elcock","year":"1998","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012408305755600_b7","doi-asserted-by":"crossref","first-page":"1356","DOI":"10.1046\/j.1432-1033.2002.02767.x","article-title":"Prediction of protein\u2013protein interaction sites in heterocomplexes with neural networks","volume":"269","author":"Fariselli","year":"2002","journal-title":"Eur. J. Biochem."},{"key":"2023012408305755600_b8","doi-asserted-by":"crossref","first-page":"2455","DOI":"10.1002\/pro.5560031231","article-title":"The subunit interfaces of oligomeric enzymes are conserved to a similar extent to the overall protein sequence","volume":"3","author":"Grishin","year":"1994","journal-title":"Protein Sci."},{"key":"2023012408305755600_b9","doi-asserted-by":"crossref","first-page":"7189","DOI":"10.1093\/nar\/gkg922","article-title":"Using electrostatic potentials to predict DNA-binding sites on DNA-binding proteins","volume":"31","author":"Jones","year":"2003","journal-title":"Nucleic Acids Res."},{"key":"2023012408305755600_b10","doi-asserted-by":"crossref","first-page":"225","DOI":"10.1006\/jmbi.2001.5344","article-title":"Residues participating in the protein folding nucleus do not exhibit preferentail evolutionary conservation","volume":"316","author":"Larson","year":"2002","journal-title":"J. Mol. Biol."},{"key":"2023012408305755600_b11","doi-asserted-by":"crossref","first-page":"D266","DOI":"10.1093\/nar\/gki001","article-title":"Pdbsum more: new summaries and analyses of the known 3D structures of proteins and nucleic acids","volume":"33","author":"Laskowski","year":"2005","journal-title":"Nucleic Acids Res."},{"key":"2023012408305755600_b12","volume-title":"Molecular Modelling: Principles and Applications","author":"Leach","year":"2001","edition":"2nd edn"},{"key":"2023012408305755600_b13","doi-asserted-by":"crossref","first-page":"1483","DOI":"10.1073\/pnas.93.15.7507","article-title":"Evolutionarily conserved Galphabetagamma binding surfaces support a model of the g protein\u2013receptor complex","volume":"93","author":"Lichtarge","year":"1996","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012408305755600_b14","doi-asserted-by":"crossref","first-page":"342","DOI":"10.1006\/jmbi.1996.0167","article-title":"An evolutionary trace method defines binding surfaces common to protein families","volume":"257","author":"Lichtarge","year":"1996","journal-title":"J. Mol. Biol."},{"key":"2023012408305755600_b15","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1006\/jmbi.2001.5327","article-title":"Structural clusters of evolutionary trace residues are statistically significant and common in proteins","volume":"316","author":"Madabushi","year":"2002","journal-title":"J. Mol. Biol."},{"key":"2023012408305755600_b16","doi-asserted-by":"crossref","first-page":"8126","DOI":"10.1074\/jbc.M312671200","article-title":"Evolutionary trace of G protein-coupled receptors reveals clusters of residues that determine global and class-specific functions","volume":"279","author":"Madabushi","year":"2004","journal-title":"J. Biol. Chem."},{"key":"2023012408305755600_b17","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1016\/S0022-2836(03)00663-6","article-title":"Combining inference from evolution and geometric probability in protein structure evaluation","volume":"331","author":"Mihalek","year":"2003","journal-title":"J. Mol. Biol."},{"key":"2023012408305755600_b18","doi-asserted-by":"crossref","first-page":"1265","DOI":"10.1016\/j.jmb.2003.12.078","article-title":"A family of evolution-entropy hybrid methods for ranking protein residues by importance","volume":"336","author":"Mihalek","year":"2004","journal-title":"J. Mol. Biol."},{"key":"2023012408305755600_b19","doi-asserted-by":"crossref","first-page":"123","DOI":"10.1006\/jmbi.2001.4602","article-title":"Evolutionary conservation of the folding nucleus","volume":"298","author":"Mirny","year":"2001","journal-title":"J. Mol. Biol."},{"key":"2023012408305755600_b20","doi-asserted-by":"crossref","first-page":"205","DOI":"10.1006\/jmbi.2000.4042","article-title":"T-coffee: a novel method for fast and accurate multiple sequence alignment","volume":"302","author":"Notredame","year":"2000","journal-title":"J. Mol. Biol."},{"key":"2023012408305755600_b21","doi-asserted-by":"crossref","first-page":"2176","DOI":"10.1093\/bioinformatics\/btg309","article-title":"Early bioinformatics: the birth of a discipline\u2014a personal view","volume":"19","author":"Ouzounis","year":"2003","journal-title":"Bioinformatics"},{"key":"2023012408305755600_b22","volume-title":"Numerical Recipes in C: The Art of Scientific Computing","author":"Press","year":"1992"},{"key":"2023012408305755600_b23","doi-asserted-by":"crossref","first-page":"402","DOI":"10.1016\/j.jmb.2005.04.054","article-title":"Correlated evolutionary pressure at interacting transcription factors and DNA response elements can guide the rational engineering of dna binding specificity","volume":"350","author":"Raviscioni","year":"2005","journal-title":"J. Mol. Biol."},{"key":"2023012408305755600_b24","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1093\/protein\/12.2.85","article-title":"Twilight zone of protein sequence alignment","volume":"12","author":"Rost","year":"1999","journal-title":"Protein Eng."},{"key":"2023012408305755600_b25","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1002\/prot.340090107","article-title":"Database of homology derived protein structures and the structural meaning of sequence alignment","volume":"9","author":"Sander","year":"1991","journal-title":"Proteins"},{"key":"2023012408305755600_b26","doi-asserted-by":"crossref","first-page":"227","DOI":"10.1016\/j.jmb.2004.03.025","article-title":"Predicting functional sites in proteins: site-specific evolutionary models and their application to neurotransmitter transporters","volume":"339","author":"Soyer","year":"2004","journal-title":"J. Mol. Biol."},{"key":"2023012408305755600_b27","doi-asserted-by":"crossref","first-page":"4673","DOI":"10.1093\/nar\/22.22.4673","article-title":"Clustal W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice","volume":"22","author":"Thompson","year":"1994","journal-title":"Nucleic Acids Res."},{"key":"2023012408305755600_b28","doi-asserted-by":"crossref","first-page":"1113","DOI":"10.1006\/jmbi.2001.4513","article-title":"Evolution of function in protein superfamilies, from a structural perspective","volume":"307","author":"Todd","year":"2001","journal-title":"J. Mol. Biol."},{"key":"2023012408305755600_b29","doi-asserted-by":"crossref","first-page":"227","DOI":"10.1002\/prot.10146","article-title":"Scoring residue conservation","volume":"48","author":"Valdar","year":"2002","journal-title":"Proteins"},{"key":"2023012408305755600_b30","doi-asserted-by":"crossref","first-page":"399","DOI":"10.1006\/jmbi.2001.5034","article-title":"Conservation helps to identify biologically relevant crystal contacts","volume":"313","author":"Valdar","year":"2001","journal-title":"J. Mol. Biol."},{"key":"2023012408305755600_b31","volume-title":"Introduction to Computational Biology","author":"Waterman","year":"2000"},{"key":"2023012408305755600_b32","volume-title":"Enzyme Nomenclature 1992","author":"Webb","year":"1992"},{"key":"2023012408305755600_b33","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1006\/jmbi.2000.3550","article-title":"Assessing annotation transfer for genomics: quantifying the relations between protein sequence, structure and function through traditional and probabilistic scores","volume":"297","author":"Wilson","year":"2000","journal-title":"J. Mol. Biol."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/2\/149\/48838586\/bioinformatics_22_2_149.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/2\/149\/48838586\/bioinformatics_22_2_149.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,24]],"date-time":"2023-01-24T08:38:18Z","timestamp":1674549498000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/22\/2\/149\/425524"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,11,22]]},"references-count":33,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2006,1,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bti791","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2006,1,15]]},"published":{"date-parts":[[2005,11,22]]}}}