{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T23:16:30Z","timestamp":1781651790825,"version":"3.54.5"},"reference-count":40,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,11,14]],"date-time":"2023-11-14T00:00:00Z","timestamp":1699920000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,11,14]],"date-time":"2023-11-14T00:00:00Z","timestamp":1699920000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Janine Bollag Post-Doctoral Fellowship Fund"},{"name":"Teva Pharmaceutical Industries Ltd as part of the Israeli National Forum for BioInnovators"},{"DOI":"10.13039\/501100003977","name":"Israel Science Foundation","doi-asserted-by":"publisher","award":["301\/2021"],"award-info":[{"award-number":["301\/2021"]}],"id":[{"id":"10.13039\/501100003977","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003977","name":"Israel Science Foundation","doi-asserted-by":"publisher","award":["301\/2021"],"award-info":[{"award-number":["301\/2021"]}],"id":[{"id":"10.13039\/501100003977","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003977","name":"Israel Science Foundation","doi-asserted-by":"publisher","award":["301\/2021"],"award-info":[{"award-number":["301\/2021"]}],"id":[{"id":"10.13039\/501100003977","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Background<\/jats:title>\n                    <jats:p>\n                      Determining a protein\u2019s quaternary state,\n                      <jats:italic>i.e.<\/jats:italic>\n                      the number of monomers in a functional unit, is a critical step in protein characterization. Many proteins form multimers for their activity, and over 50% are estimated to naturally form homomultimers. Experimental quaternary state determination can be challenging and require extensive work. To complement these efforts, a number of computational tools have been developed for quaternary state prediction, often utilizing experimentally validated structural information. Recently, dramatic advances have been made in the field of deep learning for predicting protein structure and other characteristics. Protein language models, such as ESM-2, that apply computational natural-language models to proteins successfully capture secondary structure, protein cell localization and other characteristics, from a single sequence. Here we hypothesize that information about the protein quaternary state may be contained within protein sequences as well, allowing us to benefit from these novel approaches in the context of quaternary state prediction.\n                    <\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>We generated ESM-2 embeddings for a large dataset of proteins with quaternary state labels from the curated QSbio dataset. We trained a model for quaternary state classification and assessed it on a non-overlapping set of distinct folds (ECOD family level). Our model, named QUEEN (QUaternary state prediction using dEEp learNing), performs worse than approaches that include information from solved crystal structures. However, it successfully learned to distinguish multimers from monomers, and predicts the specific quaternary state with moderate success, better than simple sequence similarity-based annotation transfer. Our results demonstrate that complex, quaternary state related information is included in such embeddings.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusions<\/jats:title>\n                    <jats:p>\n                      QUEEN is the first to investigate the power of embeddings for the prediction of the quaternary state of proteins. As such, it lays out strengths as well as limitations of a sequence-based protein language model approach, compared to structure-based approaches. Since it does not require any structural information and is fast, we anticipate that it will be of wide use both for in-depth investigation of specific systems, as well as for studies of large sets of protein sequences. A simple colab implementation is available at:\n                      <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/colab.research.google.com\/github\/Furman-Lab\/QUEEN\/blob\/main\/QUEEN_prediction_notebook.ipynb\">https:\/\/colab.research.google.com\/github\/Furman-Lab\/QUEEN\/blob\/main\/QUEEN_prediction_notebook.ipynb<\/jats:ext-link>\n                      .\n                    <\/jats:p>\n                  <\/jats:sec>","DOI":"10.1186\/s12859-023-05549-w","type":"journal-article","created":{"date-parts":[[2023,11,14]],"date-time":"2023-11-14T08:02:58Z","timestamp":1699948978000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":24,"title":["Protein language models can capture protein quaternary state"],"prefix":"10.1186","volume":"24","author":[{"given":"Orly","family":"Avraham","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tomer","family":"Tsaban","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ziv","family":"Ben-Aharon","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Linoy","family":"Tsaban","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ora","family":"Schueler-Furman","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,11,14]]},"reference":[{"key":"5549_CR1","doi-asserted-by":"publisher","first-page":"105","DOI":"10.1146\/annurev.biophys.29.1.105","volume":"29","author":"DS Goodsell","year":"2000","unstructured":"Goodsell DS, Olson AJ. Structural symmetry and protein function. Annu Rev Biophys Biomol Struct. 2000;29:105\u201353.","journal-title":"Annu Rev Biophys Biomol Struct"},{"issue":"39","key":"5549_CR2","doi-asserted-by":"publisher","first-page":"11680","DOI":"10.1039\/D2SC02794A","volume":"13","author":"S Marciano","year":"2022","unstructured":"Marciano S, Dey D, Listov D, Fleishman SJ, Sonn-Segev A, Mertens H, et al. Protein quaternary structures in solution are a mixture of multiple forms. Chem Sci. 2022;13(39):11680\u201395.","journal-title":"Chem Sci"},{"issue":"6483","key":"5549_CR3","doi-asserted-by":"publisher","first-page":"761","DOI":"10.1038\/369761a0","volume":"369","author":"RH Jacobson","year":"1994","unstructured":"Jacobson RH, Zhang XJ, DuBose RF, Matthews BW. Three-dimensional structure of beta-galactosidase from E. coli. Nature. 1994;369(6483):761\u20136.","journal-title":"Nature"},{"issue":"10","key":"5549_CR4","doi-asserted-by":"publisher","first-page":"693","DOI":"10.3389\/fneur.2019.00693","volume":"28","author":"M Uemura","year":"2019","unstructured":"Uemura M, Nozaki H, Koyama A, Sakai N, Ando S, Kanazawa M, et al. HTRA1 mutations identified in symptomatic carriers have the property of interfering the trimer-dependent activation cascade. Front Neurol. 2019;28(10):693.","journal-title":"Front Neurol"},{"issue":"1","key":"5549_CR5","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1038\/75556","volume":"25","author":"M Ashburner","year":"2000","unstructured":"Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, et al. Gene Ontology: tool for the unification of biology. Nat Genet. 2000;25(1):25\u20139.","journal-title":"Nat Genet"},{"issue":"D1","key":"5549_CR6","doi-asserted-by":"publisher","first-page":"D325","DOI":"10.1093\/nar\/gkaa1113","volume":"49","author":"The Gene Ontology Consortium","year":"2021","unstructured":"The Gene Ontology Consortium. The gene ontology resource: enriching a GOld mine. Nucleic Acids Res. 2021;49(D1):D325\u201334.","journal-title":"Nucleic Acids Res"},{"issue":"D1","key":"5549_CR7","doi-asserted-by":"publisher","first-page":"D587","DOI":"10.1093\/nar\/gkac963","volume":"51","author":"M Kanehisa","year":"2023","unstructured":"Kanehisa M, Furumichi M, Sato Y, Kawashima M, Ishiguro-Watanabe M. KEGG for taxonomy-based analysis of pathways and genomes. Nucleic Acids Res. 2023;51(D1):D587\u201392.","journal-title":"Nucleic Acids Res"},{"issue":"12","key":"5549_CR8","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1003926","volume":"10","author":"H Cheng","year":"2014","unstructured":"Cheng H, Schaeffer RD, Liao Y, Kinch LN, Pei J, Shi S, et al. ECOD: an evolutionary classification of protein domains. PLoS Comput Biol. 2014;10(12): e1003926.","journal-title":"PLoS Comput Biol"},{"issue":"D1","key":"5549_CR9","doi-asserted-by":"publisher","first-page":"D412","DOI":"10.1093\/nar\/gkaa913","volume":"49","author":"J Mistry","year":"2021","unstructured":"Mistry J, Chuguransky S, Williams L, Qureshi M, Salazar GA, Sonnhammer ELL, et al. Pfam: the protein families database in 2021. Nucleic Acids Res. 2021;49(D1):D412\u20139.","journal-title":"Nucleic Acids Res"},{"issue":"D1","key":"5549_CR10","doi-asserted-by":"publisher","first-page":"D266","DOI":"10.1093\/nar\/gkaa1079","volume":"49","author":"I Sillitoe","year":"2021","unstructured":"Sillitoe I, Bordin N, Dawson N, Waman VP, Ashford P, Scholes HM, et al. CATH: increased structural coverage of functional space. Nucleic Acids Res. 2021;49(D1):D266\u201373.","journal-title":"Nucleic Acids Res"},{"issue":"2","key":"5549_CR11","doi-asserted-by":"publisher","first-page":"114","DOI":"10.3390\/cryst10020114","volume":"10","author":"K Elez","year":"2020","unstructured":"Elez K, Bonvin AMJJ, Vangone A. Biological vs. crystallographic protein interfaces: an overview of computational approaches for their classification. Crystals. 2020;10(2):114.","journal-title":"Crystals"},{"issue":"3","key":"5549_CR12","doi-asserted-by":"publisher","first-page":"774","DOI":"10.1016\/j.jmb.2007.05.022","volume":"372","author":"E Krissinel","year":"2007","unstructured":"Krissinel E, Henrick K. Inference of macromolecular assemblies from crystalline state. J Mol Biol. 2007;372(3):774\u201397.","journal-title":"J Mol Biol"},{"issue":"13","key":"5549_CR13","doi-asserted-by":"publisher","first-page":"334","DOI":"10.1186\/1471-2105-13-334","volume":"22","author":"JM Duarte","year":"2012","unstructured":"Duarte JM, Srebniak A, Sch\u00e4rer MA, Capitani G. Protein interface classification by evolutionary analysis. BMC Bioinformatics. 2012;22(13):334.","journal-title":"BMC Bioinformatics"},{"issue":"W1","key":"5549_CR14","doi-asserted-by":"publisher","first-page":"W320","DOI":"10.1093\/nar\/gkx246","volume":"45","author":"M Baek","year":"2017","unstructured":"Baek M, Park T, Heo L, Park C, Seok C. GalaxyHomomer: a web server for protein homo-oligomer structure prediction from a monomer sequence or structure. Nucleic Acids Res. 2017;45(W1):W320\u20134.","journal-title":"Nucleic Acids Res"},{"issue":"1","key":"5549_CR15","doi-asserted-by":"publisher","first-page":"711","DOI":"10.1038\/s41467-020-14301-4","volume":"11","author":"Q Xu","year":"2020","unstructured":"Xu Q, Dunbrack RL. ProtCID: a data resource for structural information on protein interactions. Nat Commun. 2020;11(1):711.","journal-title":"Nat Commun"},{"issue":"1","key":"5549_CR16","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1038\/nmeth.4510","volume":"15","author":"S Dey","year":"2018","unstructured":"Dey S, Ritchie DW, Levy ED. PDB-wide identification of biological assemblies from conserved quaternary structure geometry. Nat Methods. 2018;15(1):67\u201372.","journal-title":"Nat Methods"},{"issue":"11","key":"5549_CR17","doi-asserted-by":"publisher","first-page":"1056","DOI":"10.1038\/s41594-022-00849-w","volume":"29","author":"M Akdel","year":"2022","unstructured":"Akdel M, Pires DEV, Pardo EP, J\u00e4nes J, Zalevsky AO, M\u00e9sz\u00e1ros B, et al. A structural biology community assessment of AlphaFold2 applications. Nat Struct Mol Biol. 2022;29(11):1056\u201367.","journal-title":"Nat Struct Mol Biol"},{"key":"5549_CR18","doi-asserted-by":"crossref","unstructured":"Schweke H, Levin T, Pacesa M, Goverde CA, Kumar P, Duhoo Y, et al. An atlas of protein homo-oligomerization across domains of life. BioRxiv. 2023 Jun 11;","DOI":"10.1101\/2023.06.09.544317"},{"key":"5549_CR19","doi-asserted-by":"crossref","unstructured":"Olechnovi\u010d K, Valan\u010dauskas L, Dapk\u016bnas J, Venclovas \u010c. Prediction of protein assemblies by structure sampling followed by interface-focused scoring. Proteins. 2023 Aug 14;","DOI":"10.1101\/2023.03.07.531468"},{"issue":"4","key":"5549_CR20","doi-asserted-by":"publisher","first-page":"1061","DOI":"10.1002\/prot.22934","volume":"79","author":"S Balakrishnan","year":"2011","unstructured":"Balakrishnan S, Kamisetty H, Carbonell JG, Lee S-I, Langmead CJ. Learning generative models for protein fold families. Proteins. 2011;79(4):1061\u201378.","journal-title":"Proteins"},{"issue":"12","key":"5549_CR21","doi-asserted-by":"publisher","first-page":"1315","DOI":"10.1038\/s41592-019-0598-1","volume":"16","author":"EC Alley","year":"2019","unstructured":"Alley EC, Khimulya G, Biswas S, AlQuraishi M, Church GM. Unified rational protein engineering with sequence-based deep representation learning. Nat Methods. 2019;16(12):1315\u201322.","journal-title":"Nat Methods"},{"key":"5549_CR22","unstructured":"Bepler T, Berger B. Learning protein sequence embeddings using information from structure. arXiv. 2019;"},{"key":"5549_CR23","doi-asserted-by":"crossref","unstructured":"Elnaggar A, Heinzinger M, Dallago C, Rihawi G, Wang Y, Jones L, et al. ProtTrans: Towards Cracking the Language of Life\u2019s Code Through Self-Supervised Deep Learning and High Performance Computing. arXiv. 2020;","DOI":"10.1101\/2020.07.12.199554"},{"key":"5549_CR24","doi-asserted-by":"crossref","unstructured":"Verkuil R, Kabeli O, Du Y, Wicky BI, Milles LF, Dauparas J, et al. Language models generalize beyond natural proteins. BioRxiv. 2022 Dec 22;","DOI":"10.1101\/2022.12.21.521521"},{"issue":"6637","key":"5549_CR25","doi-asserted-by":"publisher","first-page":"1123","DOI":"10.1126\/science.ade2574","volume":"379","author":"Z Lin","year":"2023","unstructured":"Lin Z et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379(6637):1123\u20131130.","journal-title":"Science."},{"issue":"1","key":"5549_CR26","doi-asserted-by":"publisher","first-page":"23916","DOI":"10.1038\/s41598-021-03431-4","volume":"11","author":"M Littmann","year":"2021","unstructured":"Littmann M, Heinzinger M, Dallago C, Weissenow K, Rost B. Protein embeddings and deep learning predict binding residues for various ligand classes. Sci Rep. 2021;11(1):23916.","journal-title":"Sci Rep"},{"key":"5549_CR27","doi-asserted-by":"crossref","unstructured":"Bana J, Warwar J, Bayer EA, Livnah O. Self-assembly of a dimeric avidin into unique higher-order oligomers. FEBS J. 2023","DOI":"10.1111\/febs.16764"},{"issue":"1","key":"5549_CR28","doi-asserted-by":"publisher","first-page":"235","DOI":"10.1093\/nar\/28.1.235","volume":"28","author":"HM Berman","year":"2000","unstructured":"Berman HM, Westbrook J, Feng Z, Gilliland G, Bhat TN, Weissig H, et al. The protein data bank. Nucleic Acids Res. 2000;28(1):235\u201342.","journal-title":"Nucleic Acids Res"},{"issue":"1","key":"5549_CR29","doi-asserted-by":"publisher","first-page":"1160","DOI":"10.1038\/s41598-020-80786-0","volume":"11","author":"M Littmann","year":"2021","unstructured":"Littmann M, Heinzinger M, Dallago C, Olenyi T, Rost B. Embeddings from deep learning transfer GO annotations beyond homology. Sci Rep. 2021;11(1):1160.","journal-title":"Sci Rep"},{"key":"5549_CR30","doi-asserted-by":"publisher","first-page":"238","DOI":"10.1016\/j.csbj.2022.11.014","volume":"21","author":"N Ferruz","year":"2023","unstructured":"Ferruz N, Heinzinger M, Akdel M, Goncearenco A, Naef L, Dallago C. From sequence to function through structure: Deep learning for protein design. Comput Struct Biotechnol J. 2023;21:238\u201350.","journal-title":"Comput Struct Biotechnol J"},{"issue":"19","key":"5549_CR31","doi-asserted-by":"publisher","first-page":"1750","DOI":"10.1016\/j.csbj.2021.03.022","volume":"25","author":"D Ofer","year":"2021","unstructured":"Ofer D, Brandes N, Linial M. The language of proteins: NLP, machine learning & protein sequences. Comput Struct Biotechnol J. 2021;25(19):1750\u20138.","journal-title":"Comput Struct Biotechnol J"},{"issue":"6","key":"5549_CR32","doi-asserted-by":"publisher","first-page":"932","DOI":"10.1038\/s41587-021-01179-w","volume":"40","author":"ML Bileschi","year":"2022","unstructured":"Bileschi ML, Belanger D, Bryant DH, Sanderson T, Carter B, Sculley D, et al. Using deep learning to annotate the protein universe. Nat Biotechnol. 2022;40(6):932\u20137.","journal-title":"Nat Biotechnol"},{"key":"5549_CR33","doi-asserted-by":"crossref","unstructured":"UniProt Consortium. The universal protein resource (uniprot) in 2010. Nucleic Acids Res. 2010 (Database issue):D142\u20138.","DOI":"10.1093\/nar\/gkp846"},{"key":"5549_CR34","unstructured":"Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine Learning in Python. arXiv. 2012;"},{"issue":"11","key":"5549_CR35","doi-asserted-by":"publisher","first-page":"1026","DOI":"10.1038\/nbt.3988","volume":"35","author":"M Steinegger","year":"2017","unstructured":"Steinegger M, S\u00f6ding J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat Biotechnol. 2017;35(11):1026\u20138.","journal-title":"Nat Biotechnol"},{"key":"5549_CR36","doi-asserted-by":"crossref","unstructured":"McInnes L, Healy J, Melville J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv. 2018;","DOI":"10.21105\/joss.00861"},{"key":"5549_CR37","unstructured":"Lema\u02c6\u0131tre G, Nogueira F, Aridas CK. Imbalanced-learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning. Journal of Machine Learning Research. 2017"},{"key":"5549_CR38","doi-asserted-by":"crossref","unstructured":"McKinney W. Data structures for statistical computing in python. Proceedings of the 9th Python in Science Conference. SciPy; 2010. p. 56\u201361.","DOI":"10.25080\/Majora-92bf1922-00a"},{"issue":"3","key":"5549_CR39","doi-asserted-by":"publisher","first-page":"90","DOI":"10.1109\/MCSE.2007.55","volume":"9","author":"JD Hunter","year":"2007","unstructured":"Hunter JD. Matplotlib: a 2D graphics environment. Comput Sci Eng. 2007;9(3):90\u20135.","journal-title":"Comput Sci Eng"},{"issue":"60","key":"5549_CR40","doi-asserted-by":"publisher","first-page":"3021","DOI":"10.21105\/joss.03021","volume":"6","author":"M Waskom","year":"2021","unstructured":"Waskom M. Seaborn: statistical data visualization. JOSS. 2021;6(60):3021.","journal-title":"JOSS"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-023-05549-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12859-023-05549-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-023-05549-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,14]],"date-time":"2023-11-14T08:03:22Z","timestamp":1699949002000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-023-05549-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,14]]},"references-count":40,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["5549"],"URL":"https:\/\/doi.org\/10.1186\/s12859-023-05549-w","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/2023.03.30.534955","asserted-by":"object"},{"id-type":"doi","id":"10.21203\/rs.3.rs-2761491\/v1","asserted-by":"object"}]},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,14]]},"assertion":[{"value":"31 March 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 October 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 November 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate."}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"433"}}