{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,29]],"date-time":"2026-01-29T22:14:57Z","timestamp":1769724897339,"version":"3.49.0"},"reference-count":52,"publisher":"Oxford University Press (OUP)","issue":"1","license":[{"start":{"date-parts":[[2018,6,27]],"date-time":"2018-06-27T00:00:00Z","timestamp":1530057600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"funder":[{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["R01 GM118709"],"award-info":[{"award-number":["R01 GM118709"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Extreme Science and Engineering Discovery Environment"},{"DOI":"10.13039\/100000001","name":"NSF","doi-asserted-by":"publisher","award":["ACI-1053575"],"award-info":[{"award-number":["ACI-1053575"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"name":"National Research Service Award","award":["F31GM116570"],"award-info":[{"award-number":["F31GM116570"]}]},{"name":"Medical Scientist Training Program","award":["T32GM007288"],"award-info":[{"award-number":["T32GM007288"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2019,1,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>The analysis of sequence conservation patterns has been widely utilized to identify functionally important (catalytic and ligand-binding) protein residues for over a half-century. Despite decades of development, on average state-of-the-art non-template-based functional residue prediction methods must predict \u223c25% of a protein\u2019s total residues to correctly identify half of the protein\u2019s functional site residues. The overwhelming proportion of false positives results in reported \u2018F-Scores\u2019 of \u223c0.3. We investigated the limits of current approaches, focusing on the so-far neglected impact of the specific choice of homologs included in multiple sequence alignments (MSAs).<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>The limits of conservation-based functional residue prediction were explored by surveying the binding sites of 1023 proteins. A straightforward conservation analysis of MSAs composed of randomly selected homologs sampled from a PSI-BLAST search achieves average F-Scores of \u223c0.3, a performance matching that reported by state-of-the-art methods, which often consider additional features for the prediction in a machine learning setting. Interestingly, we found that a simple combinatorial MSA sampling algorithm will in almost every case produce an MSA with an optimal set of homologs whose conservation analysis reaches average F-Scores of \u223c0.6, doubling state-of-the-art performance. We also show that this is nearly at the theoretical limit of possible performance given the agreement between different binding site definitions. Additionally, we showcase the progress in this direction made by Selection of Alignment by Maximal Mutual Information (SAMMI), an information-theory-based approach to identifying biologically informative MSAs. This work highlights the importance and the unused potential of optimally composed MSAs for conservation analysis.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information<\/jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/bty523","type":"journal-article","created":{"date-parts":[[2018,6,26]],"date-time":"2018-06-26T19:22:35Z","timestamp":1530040955000},"page":"12-19","source":"Crossref","is-referenced-by-count":19,"title":["The choice of sequence homologs included in multiple sequence alignments has a dramatic impact on evolutionary conservation analysis"],"prefix":"10.1093","volume":"35","author":[{"given":"Nelson","family":"Gil","sequence":"first","affiliation":[{"name":"Department of Systems & Computational Biology, Albert Einstein College of Medicine, Bronx, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0085-5335","authenticated-orcid":false,"given":"Andras","family":"Fiser","sequence":"additional","affiliation":[{"name":"Department of Systems & Computational Biology, Albert Einstein College of Medicine, Bronx, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2018,6,27]]},"reference":[{"key":"2023013107222168800_bty523-B1","doi-asserted-by":"crossref","first-page":"5922","DOI":"10.1093\/nar\/gkn573","article-title":"Protein-DNA interactions: structural, thermodynamic and clustering patterns of conserved residues in DNA-binding proteins","volume":"36","author":"Ahmad","year":"2008","journal-title":"Nucleic Acids Res"},{"key":"2023013107222168800_bty523-B2","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1016\/S0022-2836(05)80360-2","article-title":"Basic local alignment search tool","volume":"215","author":"Altschul","year":"1990","journal-title":"J. Mol. Biol"},{"key":"2023013107222168800_bty523-B3","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","article-title":"Gapped BLAST and PSI-BLAST: a new generation of protein database search programs","volume":"25","author":"Altschul","year":"1997","journal-title":"Nucleic Acids Res"},{"key":"2023013107222168800_bty523-B4","doi-asserted-by":"crossref","first-page":"1135","DOI":"10.1016\/j.jmb.2004.10.055","article-title":"Network analysis of protein structures identifies functional residues","volume":"344","author":"Amitai","year":"2004","journal-title":"J. Mol. Biol"},{"key":"2023013107222168800_bty523-B5","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1515\/bchm2.1961.325.1.283","article-title":"[The structure of normal adult human hemoglobins]","volume":"325","author":"Braunitzer","year":"1961","journal-title":"Hoppe Seylers Z Physiol. Chem"},{"key":"2023013107222168800_bty523-B6","doi-asserted-by":"crossref","first-page":"190","DOI":"10.1110\/ps.03323604","article-title":"Are protein-protein interfaces more conserved in sequence than the rest of the protein surface?","volume":"13","author":"Caffrey","year":"2004","journal-title":"Protein Sci"},{"key":"2023013107222168800_bty523-B7","doi-asserted-by":"crossref","first-page":"1875","DOI":"10.1093\/bioinformatics\/btm270","article-title":"Predicting functionally important residues from sequence conservation","volume":"23","author":"Capra","year":"2007","journal-title":"Bioinformatics"},{"key":"2023013107222168800_bty523-B8","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1038\/nsb0295-171","article-title":"A method to predict functional residues in proteins","volume":"2","author":"Casari","year":"1995","journal-title":"Nat. Struct. Biol"},{"key":"2023013107222168800_bty523-B9","doi-asserted-by":"crossref","first-page":"71","DOI":"10.1093\/bib\/bbt092","article-title":"A survey on prediction of specificity-determining sites in proteins","volume":"16","author":"Chakraborty","year":"2015","journal-title":"Brief. Bioinform"},{"key":"2023013107222168800_bty523-B10","doi-asserted-by":"crossref","first-page":"1625","DOI":"10.1093\/molbev\/msu117","article-title":"TCS: a new multiple sequence alignment reliability measure to estimate alignment accuracy and improve phylogenetic tree reconstruction","volume":"31","author":"Chang","year":"2014","journal-title":"Mol. Biol. Evol"},{"key":"2023013107222168800_bty523-B11","doi-asserted-by":"crossref","first-page":"S4.","DOI":"10.1186\/1471-2105-15-S15-S4","article-title":"LigandRFs: random forest ensemble to identify ligand-binding residues from sequence information alone","volume":"15","author":"Chen","year":"2014","journal-title":"BMC Bioinformatics"},{"key":"2023013107222168800_bty523-B12","volume-title":"Elements of Information Theory","author":"Cover","year":"2006"},{"key":"2023013107222168800_bty523-B13","doi-asserted-by":"crossref","first-page":"D667","DOI":"10.1093\/nar\/gkm839","article-title":"LigASite\u2014a database of biologically relevant binding sites in proteins with known apo-structures","volume":"36","author":"Dessailly","year":"2007","journal-title":"Nucleic Acids Res"},{"key":"2023013107222168800_bty523-B14","doi-asserted-by":"crossref","first-page":"63.","DOI":"10.1186\/1471-2105-14-63","article-title":"Protein structure based prediction of catalytic residues","volume":"14","author":"Fajardo","year":"2013","journal-title":"BMC Bioinformatics"},{"key":"2023013107222168800_bty523-B15","doi-asserted-by":"crossref","first-page":"1278","DOI":"10.1093\/bioinformatics\/btx779","article-title":"Identifying functionally informative evolutionary sequence profiles","volume":"34","author":"Gil","year":"2018","journal-title":"Bioinformatics"},{"key":"2023013107222168800_bty523-B16","doi-asserted-by":"crossref","first-page":"2455","DOI":"10.1002\/pro.5560031231","article-title":"The subunit interfaces of oligomeric enzymes are conserved to a similar extent to the overall protein sequences","volume":"3","author":"Grishin","year":"1994","journal-title":"Protein Sci"},{"key":"2023013107222168800_bty523-B17","doi-asserted-by":"crossref","first-page":"15447","DOI":"10.1073\/pnas.0505425102","article-title":"Conservation and relative importance of residues across protein-protein interfaces","volume":"102","author":"Guharoy","year":"2005","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"2023013107222168800_bty523-B18","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1006\/jmbi.2000.4036","article-title":"Analysis and prediction of functional sub-types from protein sequence alignments","volume":"303","author":"Hannenhalli","year":"2000","journal-title":"J. Mol. Biol"},{"key":"2023013107222168800_bty523-B19","doi-asserted-by":"crossref","first-page":"443","DOI":"10.1111\/j.1600-6143.2005.00749.x","article-title":"Rational development of LEA29Y (belatacept), a high-affinity variant of CTLA4-Ig with potent immunosuppressive properties","volume":"5","author":"Larsen","year":"2005","journal-title":"Am. J. Transplant"},{"key":"2023013107222168800_bty523-B20","doi-asserted-by":"crossref","first-page":"342","DOI":"10.1006\/jmbi.1996.0167","article-title":"An evolutionary trace method defines binding surfaces common to protein families","volume":"257","author":"Lichtarge","year":"1996","journal-title":"J. Mol. Biol"},{"key":"2023013107222168800_bty523-B21","doi-asserted-by":"crossref","first-page":"1068","DOI":"10.1016\/j.chembiol.2008.08.007","article-title":"Covalent and noncovalent intermediates of an NAD utilizing enzyme, human CD38","volume":"15","author":"Liu","year":"2008","journal-title":"Chem. Biol"},{"key":"2023013107222168800_bty523-B22","doi-asserted-by":"crossref","first-page":"1885","DOI":"10.1002\/prot.24330","article-title":"DNABind: a hybrid algorithm for structure-based prediction of DNA-binding residues by combining machine learning- and template-based approaches","volume":"81","author":"Liu","year":"2013","journal-title":"Proteins"},{"key":"2023013107222168800_bty523-B23","first-page":"745","article-title":"Protein sequence alignments: a strategy for the hierarchical analysis of residue conservation","volume":"9","author":"Livingstone","year":"1993","journal-title":"Comput. Appl. Biosci"},{"key":"2023013107222168800_bty523-B24","doi-asserted-by":"crossref","first-page":"2592","DOI":"10.1093\/bioinformatics\/btu352","article-title":"SSpro\/ACCpro 5: almost perfect prediction of protein secondary structure and relative solvent accessibility using profiles, machine learning and structural similarity","volume":"30","author":"Magnan","year":"2014","journal-title":"Bioinformatics"},{"key":"2023013107222168800_bty523-B25","doi-asserted-by":"crossref","first-page":"D267","DOI":"10.1093\/nar\/gkt1127","article-title":"FireDB: a compendium of biological and pharmacologically relevant ligands","volume":"42","author":"Maietta","year":"2014","journal-title":"Nucleic Acids Res"},{"key":"2023013107222168800_bty523-B26","doi-asserted-by":"crossref","first-page":"672","DOI":"10.1073\/pnas.50.4.672","article-title":"Primary structure and evolution of cytochrome C","volume":"50","author":"Margoliash","year":"1963","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"2023013107222168800_bty523-B27","doi-asserted-by":"crossref","first-page":"442","DOI":"10.1016\/0005-2795(75)90109-9","article-title":"Comparison of the predicted and observed secondary structure of T4 phage lysozyme","volume":"405","author":"Matthews","year":"1975","journal-title":"Biochim. Biophys. Acta"},{"key":"2023013107222168800_bty523-B28","doi-asserted-by":"crossref","first-page":"D12","DOI":"10.1093\/nar\/gkw1071","article-title":"Database resources of the national center for biotechnology information","volume":"45","author":"NCBI Resource Coordinators","year":"2017","journal-title":"Nucleic Acids Res"},{"key":"2023013107222168800_bty523-B29","doi-asserted-by":"crossref","first-page":"443","DOI":"10.1016\/0022-2836(70)90057-4","article-title":"A general method applicable to the search for similarities in the amino acid sequence of two proteins","volume":"48","author":"Needleman","year":"1970","journal-title":"J. Mol. Biol"},{"key":"2023013107222168800_bty523-B30","doi-asserted-by":"crossref","first-page":"13500","DOI":"10.1093\/nar\/gku1228","article-title":"Prediction of DNA binding motifs from 3D models of transcription factors; identifying TLX3 regulated genes","volume":"42","author":"Pujato","year":"2014","journal-title":"Nucleic Acids Res"},{"key":"2023013107222168800_bty523-B31","doi-asserted-by":"crossref","first-page":"R232.","DOI":"10.1186\/gb-2007-8-11-r232","article-title":"Determinants of protein function revealed by combinatorial entropy optimization","volume":"8","author":"Reva","year":"2007","journal-title":"Genome Biol"},{"key":"2023013107222168800_bty523-B32","first-page":"iii","article-title":"The amino-acid sequence in the glycyl chain of insulin","volume":"52","author":"Sanger","year":"1952","journal-title":"Biochem. J"},{"key":"2023013107222168800_bty523-B33","doi-asserted-by":"crossref","first-page":"617","DOI":"10.1093\/bioinformatics\/btq008","article-title":"Active site prediction using evolutionary and structural information","volume":"26","author":"Sankararaman","year":"2010","journal-title":"Bioinformatics"},{"key":"2023013107222168800_bty523-B34","doi-asserted-by":"crossref","first-page":"2445","DOI":"10.1093\/bioinformatics\/btn474","article-title":"INTREPID\u2014INformation-theoretic TREe traversal for Protein functional site IDentification","volume":"24","author":"Sankararaman","year":"2008","journal-title":"Bioinformatics"},{"key":"2023013107222168800_bty523-B35","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1016\/0022-2836(81)90087-5","article-title":"Identification of common molecular subsequences","volume":"147","author":"Smith","year":"1981","journal-title":"J. Mol. Biol"},{"key":"2023013107222168800_bty523-B36","doi-asserted-by":"crossref","first-page":"327","DOI":"10.1093\/bioinformatics\/15.4.327","article-title":"Automated analysis of interatomic contacts in proteins","volume":"15","author":"Sobolev","year":"1999","journal-title":"Bioinformatics"},{"key":"2023013107222168800_bty523-B37","doi-asserted-by":"crossref","first-page":"951","DOI":"10.1093\/bioinformatics\/bti125","article-title":"Protein homology detection by HMM-HMM comparison","volume":"21","author":"Soding","year":"2005","journal-title":"Bioinformatics"},{"key":"2023013107222168800_bty523-B38","doi-asserted-by":"crossref","first-page":"34044","DOI":"10.1038\/srep34044","article-title":"CRHunter: integrating multifaceted information to predict catalytic residues in enzymes","volume":"6","author":"Sun","year":"2016","journal-title":"Sci. Rep"},{"key":"2023013107222168800_bty523-B39","doi-asserted-by":"crossref","first-page":"1223","DOI":"10.1002\/jcc.24314","article-title":"Sequence-based prediction of protein-peptide binding sites using support vector machine","volume":"37","author":"Taherzadeh","year":"2016","journal-title":"J. Comput. Chem"},{"key":"2023013107222168800_bty523-B40","doi-asserted-by":"crossref","first-page":"477","DOI":"10.1093\/bioinformatics\/btx614","article-title":"Structure-based prediction of protein- peptide binding regions using Random Forest","volume":"34","author":"Taherzadeh","year":"2018","journal-title":"Bioinformatics"},{"key":"2023013107222168800_bty523-B41","doi-asserted-by":"crossref","first-page":"D204","DOI":"10.1093\/nar\/gku989","article-title":"UniProt: a hub for protein information","volume":"43","author":"UniProt","year":"2015","journal-title":"Nucleic Acids Res"},{"key":"2023013107222168800_bty523-B42","doi-asserted-by":"crossref","first-page":"399","DOI":"10.1006\/jmbi.2001.5034","article-title":"Conservation helps to identify biologically relevant crystal contacts","volume":"313","author":"Valdar","year":"2001","journal-title":"J. Mol. Biol"},{"key":"2023013107222168800_bty523-B43","doi-asserted-by":"crossref","first-page":"108","DOI":"10.1002\/1097-0134(20010101)42:1<108::AID-PROT110>3.0.CO;2-O","article-title":"Protein-protein interfaces: analysis of amino acid conservation in homodimers","volume":"42","author":"Valdar","year":"2001","journal-title":"Proteins"},{"key":"2023013107222168800_bty523-B44","doi-asserted-by":"crossref","first-page":"347","DOI":"10.1146\/annurev.med.58.080205.154004","article-title":"T cell costimulation: a rational target in the therapeutic armamentarium for autoimmune diseases and transplantation","volume":"58","author":"Vincenti","year":"2007","journal-title":"Annu. Rev. Med"},{"key":"2023013107222168800_bty523-B45","volume-title":"Data Mining: Practical Machine Learning Tools and Techniques","author":"Witten","year":"2011"},{"key":"2023013107222168800_bty523-B46","doi-asserted-by":"crossref","first-page":"1517","DOI":"10.1109\/TCBB.2013.126","article-title":"Predicting protein-ligand binding site using support vector machine with protein properties","volume":"10","author":"Wong","year":"2013","journal-title":"IEEE\/ACM Trans. Comput. Biol. Bioinform"},{"key":"2023013107222168800_bty523-B47","doi-asserted-by":"crossref","first-page":"88","DOI":"10.1093\/bib\/bbv023","article-title":"A comprehensive comparative review of sequence-based predictors of DNA- and RNA-binding residues","volume":"17","author":"Yan","year":"2016","journal-title":"Brief. Bioinform"},{"key":"2023013107222168800_bty523-B48","doi-asserted-by":"crossref","first-page":"D1096","DOI":"10.1093\/nar\/gks966","article-title":"BioLiP: a semi-manually curated database for biologically relevant ligand-protein interactions","volume":"41","author":"Yang","year":"2012","journal-title":"Nucleic Acids Res"},{"key":"2023013107222168800_bty523-B49","doi-asserted-by":"crossref","first-page":"216","DOI":"10.1110\/ps.062523907","article-title":"Evaluation of features for catalytic residue prediction in novel folds","volume":"16","author":"Youn","year":"2007","journal-title":"Protein Sci"},{"key":"2023013107222168800_bty523-B50","article-title":"Review and comparative assessment of sequence-based predictors of protein-binding residues","author":"Zhang","year":"2017","journal-title":"Brief. Bioinform"},{"key":"2023013107222168800_bty523-B51","doi-asserted-by":"crossref","first-page":"2329","DOI":"10.1093\/bioinformatics\/btn433","article-title":"Accurate sequence-based prediction of catalytic residues","volume":"24","author":"Zhang","year":"2008","journal-title":"Bioinformatics"},{"key":"2023013107222168800_bty523-B52","doi-asserted-by":"crossref","first-page":"957","DOI":"10.1016\/0022-2836(87)90501-8","article-title":"Prediction of protein secondary structure and active sites using the alignment of homologous sequences","volume":"195","author":"Zvelebil","year":"1987","journal-title":"J. Mol. Biol"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/1\/12\/48962886\/bioinformatics_35_1_12.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/1\/12\/48962886\/bioinformatics_35_1_12.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,31]],"date-time":"2023-01-31T10:02:48Z","timestamp":1675159368000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/35\/1\/12\/5045912"}},"subtitle":[],"editor":[{"given":"John","family":"Hancock","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2018,6,27]]},"references-count":52,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2019,1,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bty523","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2019,1,1]]},"published":{"date-parts":[[2018,6,27]]}}}