{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T17:09:45Z","timestamp":1765040985999,"version":"build-2065373602"},"reference-count":35,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2019,11,7]],"date-time":"2019-11-07T00:00:00Z","timestamp":1573084800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100010665","name":"H2020 Marie Sk\u0142odowska-Curie Actions","doi-asserted-by":"publisher","award":["734439 InferNet"],"award-info":[{"award-number":["734439 InferNet"]}],"id":[{"id":"10.13039\/100010665","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Global coevolutionary models of protein families have become increasingly popular due to their capacity to predict residue\u2013residue contacts from sequence information, but also to predict fitness effects of amino acid substitutions or to infer protein\u2013protein interactions. The central idea in these models is to construct a probability distribution, a Potts model, that reproduces single and pairwise frequencies of amino acids found in natural sequences of the protein family. This approach treats sequences from the family as independent samples, completely ignoring phylogenetic relations between them. This simplification is known to lead to potentially biased estimates of the parameters of the model, decreasing their biological relevance. Current workarounds for this problem, such as reweighting sequences, are poorly understood and not principled. Here, we propose an inference scheme that takes the phylogeny of a protein family into account in order to correct biases in estimating the frequencies of amino acids. Using artificial data, we show that a Potts model inferred using these corrected frequencies performs better in predicting contacts and fitness effect of mutations. First, only partially successful tests on real protein data are presented, too.<\/jats:p>","DOI":"10.3390\/e21111090","type":"journal-article","created":{"date-parts":[[2019,11,7]],"date-time":"2019-11-07T11:17:25Z","timestamp":1573125445000},"page":"1090","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":25,"title":["Toward Inferring Potts Models for Phylogenetically Correlated Sequence Data"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6690-629X","authenticated-orcid":false,"given":"Edwin","family":"Rodriguez Horta","sequence":"first","affiliation":[{"name":"Laboratoire de Biologie Computationnelle et Quantitative (LCQB), Institut de Biologie Paris-Seine, Sorbonne Universit\u00e9, Centre national de la recherche scientifique (CNRS), 75005 Paris, France"},{"name":"Group of Complex Systems and Statistical Physics, Department of Theoretical Physics, Physics Faculty, University of Havana, La Habana 10400, Cuba"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pierre","family":"Barrat-Charlaix","sequence":"additional","affiliation":[{"name":"Laboratoire de Biologie Computationnelle et Quantitative (LCQB), Institut de Biologie Paris-Seine, Sorbonne Universit\u00e9, Centre national de la recherche scientifique (CNRS), 75005 Paris, France"},{"name":"Biozentrum, University of Basel, 4056 Basel, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0492-3684","authenticated-orcid":false,"given":"Martin","family":"Weigt","sequence":"additional","affiliation":[{"name":"Laboratoire de Biologie Computationnelle et Quantitative (LCQB), Institut de Biologie Paris-Seine, Sorbonne Universit\u00e9, Centre national de la recherche scientifique (CNRS), 75005 Paris, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,11,7]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"D506","DOI":"10.1093\/nar\/gky1049","article-title":"UniProt: A worldwide hub of protein knowledge","volume":"47","author":"Consortium","year":"2018","journal-title":"Nucleic Acids Res."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"D1099","DOI":"10.1093\/nar\/gku950","article-title":"The Genomes OnLine Database (GOLD) v. 5: A metadata management system based on a four level (meta) genome project classification","volume":"43","author":"Reddy","year":"2014","journal-title":"Nucleic Acids Res."},{"key":"ref_3","first-page":"D427","article-title":"The Pfam protein families database in 2019","volume":"47","author":"Mistry","year":"2018","journal-title":"Nucleic Acids Res."},{"key":"ref_4","first-page":"755","article-title":"Profile hidden Markov models","volume":"14","author":"Eddy","year":"1998","journal-title":"Bioinform. (Oxf. Engl.)"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Durbin, R., Eddy, S.R., Krogh, A., and Mitchison, G. (1998). Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids, Cambridge University Press.","DOI":"10.1017\/CBO9780511790492"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"249","DOI":"10.1038\/nrg3414","article-title":"Emerging methods in protein co-evolution","volume":"14","author":"Pazos","year":"2013","journal-title":"Nat. Rev. Genet."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"032601","DOI":"10.1088\/1361-6633\/aa9965","article-title":"Inverse statistical physics of protein sequences: A key issues review","volume":"81","author":"Cocco","year":"2018","journal-title":"Rep. Prog. Phys."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"E1293","DOI":"10.1073\/pnas.1111471108","article-title":"Direct-coupling analysis of residue coevolution captures native contacts across many protein families","volume":"108","author":"Morcos","year":"2011","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"197","DOI":"10.1080\/00018732.2017.1341604","article-title":"Inverse statistical problems: From the inverse Ising problem to data science","volume":"66","author":"Nguyen","year":"2017","journal-title":"Adv. Phys."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1072","DOI":"10.1038\/nbt.2419","article-title":"Protein structure prediction from sequence variation","volume":"30","author":"Marks","year":"2012","journal-title":"Nat. Biotechnol."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"294","DOI":"10.1126\/science.aah4043","article-title":"Protein structure determination using metagenome sequence data","volume":"355","author":"Ovchinnikov","year":"2017","journal-title":"Science"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1016\/j.sbi.2016.11.004","article-title":"Potts Hamiltonian models of protein co-variation, free energy landscapes, and evolutionary fitness","volume":"43","author":"Levy","year":"2017","journal-title":"Curr. Opin. Struct. Biol."},{"key":"ref_13","unstructured":"Felsenstein, J. (2004). Inferring Phylogenies, Sinauer Associates Sunderland."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"690","DOI":"10.1073\/pnas.1711913115","article-title":"Power Law Tails in Phylogenetic Systems","volume":"115","author":"Qin","year":"2018","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"368","DOI":"10.1007\/BF01734359","article-title":"Evolutionary trees from DNA sequences: A maximum likelihood approach","volume":"17","author":"Felsenstein","year":"1981","journal-title":"J. Mol. Evol."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"van Nimwegen, E. (2007). Finding regulatory elements and regulatory motifs: A general probabilistic framework. BMC Bioinform., 8.","DOI":"10.1186\/1471-2105-8-S6-S4"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1087","DOI":"10.1021\/ci9701042","article-title":"A guided Monte Carlo search algorithm for global optimization of multidimensional functions","volume":"38","author":"Delgoda","year":"1998","journal-title":"J. Chem. Inf. Comput. Sci."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1073\/pnas.0805923106","article-title":"Identification of direct residue contacts in protein\u2013protein interaction by message passing","volume":"106","author":"Weigt","year":"2009","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1061","DOI":"10.1002\/prot.22934","article-title":"Learning generative models for protein fold families","volume":"79","author":"Balakrishnan","year":"2011","journal-title":"Proteins Struct. Funct. Bioinform."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"012707","DOI":"10.1103\/PhysRevE.87.012707","article-title":"Improved contact prediction in proteins: Using pseudolikelihoods to infer Potts models","volume":"87","author":"Ekeberg","year":"2013","journal-title":"Phys. Rev. E"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"512","DOI":"10.1038\/nature03991","article-title":"Evolutionary information for specifying a protein fold","volume":"437","author":"Socolich","year":"2005","journal-title":"Nature"},{"key":"ref_22","first-page":"17","article-title":"On the evolution of random graphs","volume":"5","year":"1960","journal-title":"Publ. Math. Inst. Hung. Acad. Sci."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Mann, J.K., Barton, J.P., Ferguson, A.L., Omarjee, S., Walker, B.D., Chakraborty, A., and Ndung\u2019u, T. (2014). The Fitness Landscape of HIV-1 Gag: Advanced Modeling Approaches and Validation of Model Predictions by In Vitro Testing. PLoS Comput. Biol., 10.","DOI":"10.1371\/journal.pcbi.1003776"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"12408","DOI":"10.1073\/pnas.1413575111","article-title":"Coevolutionary information, protein folding landscapes, and the thermodynamics of natural selection","volume":"111","author":"Morcos","year":"2014","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"268","DOI":"10.1093\/molbev\/msv211","article-title":"Coevolutionary landscape inference and the context-dependence of mutations in beta-lactamase TEM-1","volume":"33","author":"Figliuzzi","year":"2016","journal-title":"Mol. Biol. Evol."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"128","DOI":"10.1038\/nbt.3769","article-title":"Mutation effects predicted from sequence co-variation","volume":"35","author":"Hopf","year":"2017","journal-title":"Nat. Biotechnol."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Feinauer, C., and Weigt, M. (2017). Context-Aware Prediction of Pathogenicity of Missense Mutations Involved in Human Disease. arXiv.","DOI":"10.1101\/103051"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"3812","DOI":"10.1093\/nar\/gkg509","article-title":"SIFT: Predicting amino acid changes that affect protein function","volume":"31","author":"Ng","year":"2003","journal-title":"Nucleic Acids Res."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"248","DOI":"10.1038\/nmeth0410-248","article-title":"A method and server for predicting damaging missense mutations","volume":"7","author":"Adzhubei","year":"2010","journal-title":"Nat. Methods"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1641","DOI":"10.1093\/molbev\/msp077","article-title":"FastTree: Computing large minimum evolution trees with profiles instead of a distance matrix","volume":"26","author":"Price","year":"2009","journal-title":"Mol. Biol. Evol."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Price, M.N., Dehal, P.S., and Arkin, A.P. (2010). FastTree 2\u2013approximately maximum-likelihood trees for large alignments. PLoS ONE, 5.","DOI":"10.1371\/journal.pone.0009490"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Baldassi, C., Zamparo, M., Feinauer, C., Procaccini, A., Zecchina, R., Weigt, M., and Pagnani, A. (2014). Fast and accurate multivariate Gaussian modeling of protein families: Predicting residue contacts and protein-interaction partners. PLoS ONE, 9.","DOI":"10.1371\/journal.pone.0092721"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Cocco, S., Monasson, R., and Weigt, M. (2013). From principal component to direct coupling analysis of coevolution in proteins: Low-eigenvalue modes are needed for structure prediction. PLoS Comput. Biol., 9.","DOI":"10.1371\/journal.pcbi.1003176"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"e39397","DOI":"10.7554\/eLife.39397","article-title":"Learning protein constitutive motifs from sequence data","volume":"8","author":"Tubiana","year":"2019","journal-title":"eLife"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"032128","DOI":"10.1103\/PhysRevE.100.032128","article-title":"Selection of sequence motifs and generative Hopfield-Potts models for protein families","volume":"100","author":"Shimagaki","year":"2019","journal-title":"Phys. Rev. E"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/21\/11\/1090\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:32:36Z","timestamp":1760189556000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/21\/11\/1090"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,11,7]]},"references-count":35,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2019,11]]}},"alternative-id":["e21111090"],"URL":"https:\/\/doi.org\/10.3390\/e21111090","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2019,11,7]]}}}