{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,28]],"date-time":"2026-07-28T20:24:26Z","timestamp":1785270266769,"version":"3.55.0"},"reference-count":51,"publisher":"Oxford University Press (OUP)","issue":"23","license":[{"start":{"date-parts":[[2022,10,13]],"date-time":"2022-10-13T00:00:00Z","timestamp":1665619200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"PRIN 2017","award":["2017483NH8"],"award-info":[{"award-number":["2017483NH8"]}]},{"name":"Italian Ministry of University and Research"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,11,30]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Motivation<\/jats:title>\n                    <jats:p>The advent of massive DNA sequencing technologies is producing a huge number of human single-nucleotide polymorphisms occurring in protein-coding regions and possibly changing their sequences. Discriminating harmful protein variations from neutral ones is one of the crucial challenges in precision medicine. Computational tools based on artificial intelligence provide models for protein sequence encoding, bypassing database searches for evolutionary information. We leverage the new encoding schemes for an efficient annotation of protein variants.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>E-SNPs&amp;GO is a novel method that, given an input protein sequence and a single amino acid variation, can predict whether the variation is related to diseases or not. The proposed method adopts an input encoding completely based on protein language models and embedding techniques, specifically devised to encode protein sequences and GO functional annotations. We trained our model on a newly generated dataset of 101\u00a0146 human protein single amino acid variants in 13\u00a0661 proteins, derived from public resources. When tested on a blind set comprising 10\u00a0266 variants, our method well compares to recent approaches released in literature for the same task, reaching a Matthews Correlation Coefficient score of 0.72. We propose E-SNPs&amp;GO as a suitable, efficient and accurate large-scale annotator of protein variant datasets.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Availability and implementation<\/jats:title>\n                    <jats:p>The method is available as a webserver at https:\/\/esnpsandgo.biocomp.unibo.it. Datasets and predictions are available at https:\/\/esnpsandgo.biocomp.unibo.it\/datasets.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Supplementary information<\/jats:title>\n                    <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btac678","type":"journal-article","created":{"date-parts":[[2022,10,13]],"date-time":"2022-10-13T11:44:50Z","timestamp":1665661490000},"page":"5168-5174","source":"Crossref","is-referenced-by-count":46,"title":["E-SNPs&amp;GO: embedding of protein sequence and function improves the annotation of human pathogenic variants"],"prefix":"10.1093","volume":"38","author":[{"given":"Matteo","family":"Manfredi","sequence":"first","affiliation":[{"name":"Biocomputing Group, Department of Pharmacy and Biotechnology, University of Bologna , Bologna 40126, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7359-0633","authenticated-orcid":false,"given":"Castrense","family":"Savojardo","sequence":"additional","affiliation":[{"name":"Biocomputing Group, Department of Pharmacy and Biotechnology, University of Bologna , Bologna 40126, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0274-5669","authenticated-orcid":false,"given":"Pier Luigi","family":"Martelli","sequence":"additional","affiliation":[{"name":"Biocomputing Group, Department of Pharmacy and Biotechnology, University of Bologna , Bologna 40126, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7462-7039","authenticated-orcid":false,"given":"Rita","family":"Casadio","sequence":"additional","affiliation":[{"name":"Biocomputing Group, Department of Pharmacy and Biotechnology, University of Bologna , Bologna 40126, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2022,10,13]]},"reference":[{"key":"2022113016202726700_btac678-B1","doi-asserted-by":"crossref","first-page":"248","DOI":"10.1038\/nmeth0410-248","article-title":"A method and server for predicting damaging missense mutations","volume":"7","author":"Adzhubei","year":"2010","journal-title":"Nat. Methods"},{"key":"2022113016202726700_btac678-B2","doi-asserted-by":"crossref","first-page":"1315","DOI":"10.1038\/s41592-019-0598-1","article-title":"Unified rational protein engineering with sequence-based deep representation learning","volume":"16","author":"Alley","year":"2019","journal-title":"Nat. Methods"},{"key":"2022113016202726700_btac678-B3","doi-asserted-by":"crossref","first-page":"D1038","DOI":"10.1093\/nar\/gky1151","article-title":"OMIM.org: leveraging knowledge across phenotype\u2013gene relationships","volume":"47","author":"Amberger","year":"2019","journal-title":"Nucleic Acids Res"},{"key":"2022113016202726700_btac678-B4","doi-asserted-by":"crossref","first-page":"e0141287","DOI":"10.1371\/journal.pone.0141287","article-title":"Continuous distributed representation of biological sequences for deep proteomics and genomics","volume":"10","author":"Asgari","year":"2015","journal-title":"PLoS One"},{"key":"2022113016202726700_btac678-B5","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1038\/75556","article-title":"Gene ontology: tool for the unification of biology","volume":"25","author":"Ashburner","year":"2000","journal-title":"Nat. Genet"},{"key":"2022113016202726700_btac678-B6","doi-asserted-by":"crossref","first-page":"5709","DOI":"10.1093\/bioinformatics\/btaa943","article-title":"Calibrating variant-scoring methods for clinical decision making","volume":"36","author":"Benevenuta","year":"2021","journal-title":"Bioinformatics"},{"key":"2022113016202726700_btac678-B7","doi-asserted-by":"crossref","first-page":"654","DOI":"10.1016\/j.cels.2021.05.017","article-title":"Learning the protein language: evolution, structure, and function","volume":"12","author":"Bepler","year":"2021","journal-title":"Cell Syst"},{"key":"2022113016202726700_btac678-B8","doi-asserted-by":"crossref","first-page":"1237","DOI":"10.1002\/humu.21047","article-title":"Functional annotations improve the predictive score of human disease-related mutations in proteins","volume":"30","author":"Calabrese","year":"2009","journal-title":"Hum. Mutat"},{"key":"2022113016202726700_btac678-B9","doi-asserted-by":"crossref","first-page":"S3","DOI":"10.1186\/1471-2164-14-S3-S3","article-title":"Identifying Mendelian disease genes with the variant effect scoring tool","volume":"14 (Suppl. 3)","author":"Carter","year":"2013","journal-title":"BMC Genomics"},{"key":"2022113016202726700_btac678-B10","doi-asserted-by":"crossref","first-page":"1813","DOI":"10.1007\/s10994-021-05997-6","article-title":"OWL2Vec: embedding of OWL ontologies","volume":"110","author":"Chen","year":"2021","journal-title":"Mach. Learn"},{"key":"2022113016202726700_btac678-B11","doi-asserted-by":"crossref","first-page":"e46688","DOI":"10.1371\/journal.pone.0046688","article-title":"Predicting the functional effect of amino acid substitutions and indels","volume":"7","author":"Choi","year":"2012","journal-title":"PLoS One"},{"key":"2022113016202726700_btac678-B12","doi-asserted-by":"crossref","first-page":"e113","DOI":"10.1002\/cpz1.113","article-title":"Learned embeddings from deep learning to visualize and predict protein sets","volume":"1","author":"Dallago","year":"2021","journal-title":"Curr. Protoc"},{"key":"2022113016202726700_btac678-B13","doi-asserted-by":"crossref","first-page":"bbac003","DOI":"10.1093\/bib\/bbac003","article-title":"Anc2vec: embedding gene ontology terms by preserving ancestors relationships","volume":"23","author":"Edera","year":"2022","journal-title":"Brief. Bioinformatics"},{"key":"2022113016202726700_btac678-B14","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TPAMI.2021.3095381","article-title":"ProtTrans: towards cracking the language of life\u2019s code through Self-Supervised deep learning and high performance computing","volume":"14","author":"Elnaggar","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"2022113016202726700_btac678-B15","author":"Grover","year":"2016"},{"key":"2022113016202726700_btac678-B16","doi-asserted-by":"crossref","first-page":"723","DOI":"10.1186\/s12859-019-3220-8","article-title":"Modeling aspects of the language of life through transfer-learning protein sequences","volume":"20","author":"Heinzinger","year":"2019","journal-title":"BMC Bioinformatics"},{"key":"2022113016202726700_btac678-B18","doi-asserted-by":"crossref","first-page":"1581","DOI":"10.1038\/ng.3703","article-title":"M-CAP eliminates a majority of variants of uncertain significance in clinical exomes at high sensitivity","volume":"48","author":"Jagadeesh","year":"2016","journal-title":"Nat. Genet"},{"key":"2022113016202726700_btac678-B19","doi-asserted-by":"crossref","first-page":"e2113348119","DOI":"10.1073\/pnas.2113348119","article-title":"Ultrafast end-to-end protein structure prediction enables high-throughput exploration of uncharacterized proteins","volume":"119","author":"Kandathil","year":"2022","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"2022113016202726700_btac678-B20","doi-asserted-by":"crossref","first-page":"D1062","DOI":"10.1093\/nar\/gkx1153","article-title":"ClinVar: improving access to variant interpretations and supporting evidence","volume":"46","author":"Landrum","year":"2018","journal-title":"Nucleic Acids Res"},{"key":"2022113016202726700_btac678-B21","doi-asserted-by":"crossref","first-page":"1464","DOI":"10.1126\/science.abi8207","article-title":"From variant to function in human disease genetics","volume":"373","author":"Lappalainen","year":"2021","journal-title":"Science"},{"key":"2022113016202726700_btac678-B22","doi-asserted-by":"crossref","first-page":"2744","DOI":"10.1093\/bioinformatics\/btp528","article-title":"Automated inference of molecular mechanisms of disease from amino acid substitutions","volume":"25","author":"Li","year":"2009","journal-title":"Bioinformatics"},{"key":"2022113016202726700_btac678-B23","doi-asserted-by":"crossref","first-page":"1160","DOI":"10.1038\/s41598-020-80786-0","article-title":"Embeddings from deep learning transfer GO annotations beyond homology","volume":"11","author":"Littmann","year":"2021","journal-title":"Sci. Rep"},{"key":"2022113016202726700_btac678-B24","doi-asserted-by":"crossref","first-page":"bbab578","DOI":"10.1093\/bib\/bbab578","article-title":"EGRET: edge aggregated graph attention networks and transfer learning improve protein\u2013protein interaction site prediction","volume":"23","author":"Mahbub","year":"2022","journal-title":"Brief. Bioinformatics"},{"key":"2022113016202726700_btac678-B25","doi-asserted-by":"crossref","first-page":"1629","DOI":"10.1007\/s00439-021-02411-y","article-title":"Embeddings from protein language models predict conservation and variant effects","volume":"141","author":"Marquet","year":"2021","journal-title":"Hum. Genet"},{"key":"2022113016202726700_btac678-B26","first-page":"29287","author":"Meier","year":"2021"},{"key":"2022113016202726700_btac678-B27","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1002\/humu.22204","article-title":"VariBench: a benchmark database for variations","volume":"34","author":"Nair","year":"2013","journal-title":"Hum. Mutat"},{"key":"2022113016202726700_btac678-B28","doi-asserted-by":"crossref","first-page":"863","DOI":"10.1101\/gr.176601","article-title":"Predicting deleterious amino acid substitutions","volume":"11","author":"Ng","year":"2001","journal-title":"Genome Res"},{"key":"2022113016202726700_btac678-B29","first-page":"625","author":"Niculescu-Mizil","year":"2005"},{"key":"2022113016202726700_btac678-B30","doi-asserted-by":"crossref","first-page":"e0117380","DOI":"10.1371\/journal.pone.0117380","article-title":"PON-P2: prediction method for fast and reliable identification of harmful variants","volume":"10","author":"Niroula","year":"2015","journal-title":"PLoS One"},{"key":"2022113016202726700_btac678-B31","doi-asserted-by":"crossref","first-page":"1750","DOI":"10.1016\/j.csbj.2021.03.022","article-title":"The language of proteins: NLP, machine learning & protein sequences","volume":"19","author":"Ofer","year":"2021","journal-title":"Comput. Struct. Biotechnol. J"},{"key":"2022113016202726700_btac678-B32","first-page":"2825","article-title":"Scikit-learn: machine learning in python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"J. Mach. Learn. Res"},{"key":"2022113016202726700_btac678-B33","doi-asserted-by":"crossref","first-page":"5918","DOI":"10.1038\/s41467-020-19669-x","article-title":"Inferring the molecular and phenotypic impact of amino acid variants with MutPred2","volume":"11","author":"Pejaver","year":"2020","journal-title":"Nat. Commun"},{"key":"2022113016202726700_btac678-B34","first-page":"701","author":"Perozzi","year":"2014"},{"key":"2022113016202726700_btac678-B35","doi-asserted-by":"crossref","first-page":"W201","DOI":"10.1093\/nar\/gkx390","article-title":"DEOGEN2: prediction and interactive visualization of single amino acid variant deleteriousness in human proteins","volume":"45","author":"Raimondi","year":"2017","journal-title":"Nucleic Acids Res"},{"key":"2022113016202726700_btac678-B36","doi-asserted-by":"crossref","first-page":"e2016239118","DOI":"10.1073\/pnas.2016239118","article-title":"Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences","volume":"118","author":"Rives","year":"2021","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"2022113016202726700_btac678-B37","doi-asserted-by":"crossref","first-page":"575","DOI":"10.1038\/nmeth0810-575","article-title":"MutationTaster evaluates disease-causing potential of sequence alterations","volume":"7","author":"Schwarz","year":"2010","journal-title":"Nat. Methods"},{"key":"2022113016202726700_btac678-B38","doi-asserted-by":"crossref","first-page":"D704","DOI":"10.1093\/nar\/gkz997","article-title":"The monarch initiative in 2019: an integrative data and analytic platform connecting phenotypes to genotypes across species","volume":"48","author":"Shefchek","year":"2020","journal-title":"Nucleic Acids Res"},{"key":"2022113016202726700_btac678-B39","doi-asserted-by":"crossref","first-page":"1888","DOI":"10.1093\/bioinformatics\/btac053","article-title":"SPOT-Contact-LM: improving single-sequence-based prediction of protein contact map using a transformer language model","volume":"38","author":"Singh","year":"2022","journal-title":"Bioinformatics"},{"key":"2022113016202726700_btac678-B40","doi-asserted-by":"crossref","first-page":"vbab035","DOI":"10.1093\/bioadv\/vbab035","article-title":"Light attention predicts protein location from the language of life","volume":"1","author":"St\u00e4rk","year":"2021","journal-title":"Bioinform. Adv"},{"key":"2022113016202726700_btac678-B41","doi-asserted-by":"crossref","first-page":"1026","DOI":"10.1038\/nbt.3988","article-title":"MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets","volume":"35","author":"Steinegger","year":"2017","journal-title":"Nat. Biotechnol"},{"key":"2022113016202726700_btac678-B42","doi-asserted-by":"crossref","first-page":"2542","DOI":"10.1038\/s41467-018-04964-5","article-title":"Clustering huge protein sequence sets in linear time","volume":"9","author":"Steinegger","year":"2018","journal-title":"Nat. Commun"},{"key":"2022113016202726700_btac678-B43","doi-asserted-by":"crossref","first-page":"603","DOI":"10.1038\/s41592-019-0437-4","article-title":"Protein-level assembly increases protein sequence recovery from metagenomic samples manyfold","volume":"16","author":"Steinegger","year":"2019","journal-title":"Nat. Methods"},{"key":"2022113016202726700_btac678-B44","doi-asserted-by":"crossref","first-page":"2401","DOI":"10.1093\/bioinformatics\/btaa003","article-title":"UDSMProt: universal deep sequence models for protein classification","volume":"36","author":"Strodthoff","year":"2020","journal-title":"Bioinformatics"},{"key":"2022113016202726700_btac678-B45","doi-asserted-by":"crossref","first-page":"926","DOI":"10.1093\/bioinformatics\/btu739","article-title":"UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches","volume":"31","author":"Suzek","year":"2015","journal-title":"Bioinformatics"},{"key":"2022113016202726700_btac678-B46","doi-asserted-by":"crossref","first-page":"1023","DOI":"10.1038\/s41587-021-01156-3","article-title":"SignalP 6.0 predicts all five types of signal peptides using protein language models","volume":"40","author":"Teufel","year":"2022","journal-title":"Nat. Biotechnol"},{"key":"2022113016202726700_btac678-B47","doi-asserted-by":"crossref","first-page":"D480","DOI":"10.1093\/nar\/gkaa1100","article-title":"UniProt: the universal protein knowledgebase in 2021","volume":"49","author":"The UniProt Consortium","year":"2021","journal-title":"Nucleic Acids Res"},{"key":"2022113016202726700_btac678-B48","first-page":"5999","author":"Vaswani","year":"2017"},{"key":"2022113016202726700_btac678-B50","doi-asserted-by":"crossref","first-page":"104","DOI":"10.1016\/j.biochi.2020.10.009","article-title":"Functional effects of protein variants","volume":"180","author":"Vihinen","year":"2021","journal-title":"Biochimie"},{"key":"2022113016202726700_btac678-B51","doi-asserted-by":"crossref","first-page":"1122","DOI":"10.1038\/s41592-021-01205-4","article-title":"DOME: recommendations for supervised machine learning validation in biology","volume":"18","author":"Walsh","year":"2021","journal-title":"Nat. Methods"},{"key":"2022113016202726700_btac678-B52","doi-asserted-by":"crossref","first-page":"867572","DOI":"10.3389\/fmolb.2022.867572","article-title":"PON-All, amino acid substitution tolerance predictor for all organisms","volume":"9","author":"Yang","year":"2022","journal-title":"Front. Mol. Biosci"},{"key":"2022113016202726700_btac678-B53","doi-asserted-by":"crossref","first-page":"918","DOI":"10.1186\/s12864-019-6272-2","article-title":"GO2Vec: transforming GO terms and proteins to vector representations via graph embeddings","volume":"20","author":"Zhong","year":"2019","journal-title":"BMC Genomics"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btac678\/46647989\/btac678.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/38\/23\/5168\/47465983\/btac678.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/38\/23\/5168\/47465983\/btac678.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,11,30]],"date-time":"2022-11-30T12:25:58Z","timestamp":1669811158000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/38\/23\/5168\/6760258"}},"subtitle":[],"editor":[{"given":"Valentina","family":"Boeva","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2022,10,13]]},"references-count":51,"journal-issue":{"issue":"23","published-online":{"date-parts":[[2022,10,13]]},"published-print":{"date-parts":[[2022,11,30]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btac678","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/2022.05.10.491314","asserted-by":"object"}]},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022,12,1]]},"published":{"date-parts":[[2022,10,13]]}}}