{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T23:32:23Z","timestamp":1780702343850,"version":"3.54.1"},"reference-count":77,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2022,11,17]],"date-time":"2022-11-17T00:00:00Z","timestamp":1668643200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Bioinform."],"abstract":"<jats:p>Since 1992, all state-of-the-art methods for fast and sensitive identification of evolutionary, structural, and functional relations between proteins (also referred to as \u201chomology detection\u201d) use sequences and sequence-profiles (PSSMs). Protein Language Models (pLMs) generalize sequences, possibly capturing the same constraints as PSSMs, e.g., through embeddings. Here, we explored how to use such embeddings for nearest neighbor searches to identify relations between protein pairs with diverged sequences (remote homology detection for levels of &amp;lt;20% pairwise sequence identity, PIDE). While this approach excelled for proteins with single domains, we demonstrated the current challenges applying this to multi-domain proteins and presented some ideas how to overcome existing limitations, in principle. We observed that sufficiently challenging data set separations were crucial to provide deeply relevant insights into the behavior of nearest neighbor search when applied to the protein embedding space, and made all our methods readily available for others.<\/jats:p>","DOI":"10.3389\/fbinf.2022.1033775","type":"journal-article","created":{"date-parts":[[2022,11,17]],"date-time":"2022-11-17T03:29:07Z","timestamp":1668655747000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":40,"title":["Nearest neighbor search on embeddings rapidly identifies distant protein relations"],"prefix":"10.3389","volume":"2","author":[{"given":"Konstantin","family":"Sch\u00fctze","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Michael","family":"Heinzinger","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Martin","family":"Steinegger","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Burkhard","family":"Rost","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2022,11,17]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"1315","DOI":"10.1038\/s41592-019-0598-1","article-title":"Unified rational protein engineering with sequence-based deep representation learning","volume":"16","author":"Alley","year":"2019","journal-title":"Nat. Methods"},{"key":"B2","doi-asserted-by":"publisher","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","article-title":"Gapped BLAST and PSI-BLAST: A new generation of protein database search programs","volume":"25","author":"Altschul","year":"1997","journal-title":"Nucleic acids Res."},{"key":"B3","doi-asserted-by":"publisher","first-page":"1247","DOI":"10.1109\/tpami.2014.2361319","article-title":"The inverted multi-index","volume":"37","author":"Babenko","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"B4","first-page":"202","article-title":") Revisiting the inverted indices for billion-scale approximate nearest neighbors","author":"Baranchuk","year":""},{"key":"B5","doi-asserted-by":"publisher","first-page":"263","DOI":"10.1093\/nar\/28.1.263","article-title":"The Pfam protein families database","volume":"28","author":"Bateman","year":"2000","journal-title":"Nucleic Acids Res."},{"key":"B6","first-page":"08661","article-title":"Learning protein sequence embeddings using information from structure","author":"Bepler","year":"2019"},{"key":"B7","unstructured":"Annoy: Approximate nearest neighbors in C++\/Python optimized for memory usage and loading\/saving to disk\n            BernhardssonE.\n          2020"},{"key":"B8","doi-asserted-by":"publisher","first-page":"932","DOI":"10.1038\/s41587-021-01179-w","article-title":"Using deep learning to annotate the protein universe","volume":"40","author":"Bileschi","year":"2022","journal-title":"Nat. Biotechnol."},{"key":"B9","doi-asserted-by":"publisher","first-page":"6073","DOI":"10.1073\/pnas.95.11.6073","article-title":"Assessing sequence comparison methods with reliable structurally identified distant evolutionary relationships","volume":"95","author":"Brenner","year":"1998","journal-title":"Proc. Natl. Acad. Sci. U. S. A."},{"key":"B10","doi-asserted-by":"publisher","first-page":"366","DOI":"10.1038\/s41592-021-01101-x","article-title":"Sensitive protein alignments at tree-of-life scale using DIAMOND","volume":"18","author":"Buchfink","year":"2021","journal-title":"Nat. Methods"},{"key":"B11","doi-asserted-by":"publisher","first-page":"627","DOI":"10.1007\/978-1-4939-7000-1_26","article-title":"Protein Data Bank (PDB): The single global macromolecular structure archive","volume":"1607","author":"Burley","year":"2017","journal-title":"Methods Mol. Biol."},{"key":"B12","doi-asserted-by":"publisher","first-page":"e113","DOI":"10.1002\/cpz1.113","article-title":"Learned embeddings from deep learning to visualize and predict protein sets","volume":"1","author":"Dallago","year":"2021","journal-title":"Curr. Protoc."},{"key":"B13","volume-title":"The origin of species by means of natural selection, or the preservation of favoured races in the struggle for life","author":"Darwin","year":"1859"},{"key":"B14","volume-title":"Of URFs and ORFs: A primer on how to analyze derived amino acid sequences","author":"Doolittle","year":"1986"},{"key":"B15","doi-asserted-by":"crossref","DOI":"10.1101\/2022.05.23.493038","article-title":"High-throughput deep learning variant effect prediction with Sequence UNET","author":"Dunham","year":"2022"},{"key":"B16","doi-asserted-by":"publisher","first-page":"D427","DOI":"10.1093\/nar\/gky995","article-title":"The Pfam protein families database in 2019","volume":"47","author":"El-Gebali","year":"2019","journal-title":"Nucleic acids Res."},{"key":"B17","doi-asserted-by":"crossref","DOI":"10.1109\/TPAMI.2021.3095381","article-title":"ProtTrans: Towards cracking the language of lifes code through self-supervised deep learning and high performance computing","author":"Elnaggar","year":"2021"},{"key":"B18","article-title":"ProtTrans: Towards cracking the language of life\u2019s code through self-supervised deep learning and high performance computing","author":"Elnaggar","year":"2020"},{"key":"B19","doi-asserted-by":"publisher","first-page":"W29","DOI":"10.1093\/nar\/gkr367","article-title":"HMMER web server: Interactive sequence similarity searching","volume":"39","author":"Finn","year":"2011","journal-title":"Nucleic Acids Res."},{"key":"B20","first-page":"162","article-title":"Glimmers in the midnight zone: Characterization of aligned identical residues in sequence-dissimilar proteins sharing a common fold","volume":"8","author":"Friedberg","year":"2000","journal-title":"Proc. Int. Conf. Intell. Syst. Mol. Biol."},{"key":"B21","first-page":"792","article-title":"Protein modeling using hidden Markov models: Analysis of globins","author":"Haussler","year":"1993"},{"key":"B22","doi-asserted-by":"publisher","first-page":"723","DOI":"10.1186\/s12859-019-3220-8","article-title":"Modeling aspects of the language of life through transfer-learning protein sequences","volume":"20","author":"Heinzinger","year":"2019","journal-title":"BMC Bioinforma."},{"key":"B23","doi-asserted-by":"publisher","first-page":"lqac043","DOI":"10.1093\/nargab\/lqac043","article-title":"Contrastive learning on protein embeddings enlightens midnight zone","volume":"4","author":"Heinzinger","year":"2022","journal-title":"Nar. Genom. Bioinform."},{"key":"B24","doi-asserted-by":"publisher","first-page":"151","DOI":"10.1093\/bioinformatics\/5.2.151","article-title":"Fast and sensitive multiple sequence alignments on a microcomputer","volume":"5","author":"Higgins","year":"1989","journal-title":"Bioinformatics"},{"key":"B25","doi-asserted-by":"publisher","first-page":"117","DOI":"10.1109\/tpami.2010.57","article-title":"Product quantization for nearest neighbor search","volume":"33","author":"Jegou","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"B26","article-title":"Billion-scale similarity search with GPUs","author":"Johnson","year":"2019","journal-title":"IEEE Trans. Big Data"},{"key":"B27","doi-asserted-by":"publisher","first-page":"431","DOI":"10.1186\/1471-2105-11-431","article-title":"Hidden Markov model speed heuristic and iterative HMM search procedure","volume":"11","author":"Johnson","year":"2010","journal-title":"BMC Bioinforma."},{"key":"B28","doi-asserted-by":"publisher","first-page":"583","DOI":"10.1038\/s41586-021-03819-2","article-title":"Highly accurate protein structure prediction with AlphaFold","volume":"596","author":"Jumper","year":"2021","journal-title":"Nature"},{"key":"B29","doi-asserted-by":"publisher","first-page":"2264","DOI":"10.1073\/pnas.87.6.2264","article-title":"Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes","volume":"87","author":"Karlin","year":"1990","journal-title":"Proc. Natl. Acad. Sci. U. S. A."},{"key":"B30","article-title":"Generalization through memorization: Nearest neighbor language models","author":"Khandelwal","year":"2019"},{"key":"B31","doi-asserted-by":"publisher","first-page":"1173","DOI":"10.1016\/j.jmb.2004.12.032","article-title":"Comprehensive evaluation of protein structure alignment methods: Scoring by geometric measures","volume":"346","author":"Kolodny","year":"2005","journal-title":"J. Mol. Biol."},{"key":"B32","doi-asserted-by":"publisher","first-page":"1501","DOI":"10.1006\/jmbi.1994.1104","article-title":"Hidden Markov models in computational biology: Applications to protein modeling","volume":"235","author":"Krogh","year":"1994","journal-title":"J. Mol. Biol."},{"key":"B33","first-page":"9","article-title":"The design and implementation of a real time visual search system on JD E-commerce platform","author":"Li","year":""},{"key":"B34","doi-asserted-by":"publisher","first-page":"3449","DOI":"10.1093\/bioinformatics\/btab371","article-title":"Clustering FunFams using sequence embeddings improves EC purity","volume":"37","author":"Littmann","year":"","journal-title":"Bioinformatics"},{"key":"B35","doi-asserted-by":"crossref","DOI":"10.1101\/2020.09.04.282814","article-title":"Embeddings from deep learning transfer GO annotations beyond homology","author":"Littmann","year":"2020"},{"key":"B36","doi-asserted-by":"publisher","first-page":"1160","DOI":"10.1038\/s41598-020-80786-0","article-title":"Embeddings from deep learning transfer GO annotations beyond homology","volume":"11","author":"Littmann","year":"","journal-title":"Sci. Rep."},{"key":"B37","doi-asserted-by":"publisher","first-page":"23916","DOI":"10.1038\/s41598-021-03431-4","article-title":"Protein embeddings and deep learning predict binding residues for various ligand classes","volume":"11","author":"Littmann","year":"","journal-title":"Sci. Rep."},{"key":"B38","doi-asserted-by":"publisher","first-page":"188","DOI":"10.1002\/prot.20012","article-title":"Automatic target selection for structural genomics on eukaryotes","volume":"56","author":"Liu","year":"2004","journal-title":"Proteins."},{"key":"B39","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1016\/s1367-5931(02)00003-0","article-title":"Domains, motifs, and clusters in the protein universe","volume":"7","author":"Liu","year":"2003","journal-title":"Curr. Opin. Chem. Biol."},{"key":"B40","first-page":"28","article-title":"()","author":"Liu","year":"2007","journal-title":"Clustering billions of images with large scale nearest neighbor searchIEEE)"},{"key":"B41","doi-asserted-by":"crossref","DOI":"10.1101\/2020.09.04.283929","article-title":"Self-supervised contrastive learning of protein representations by mutual information maximization","author":"Lu","year":"2020"},{"key":"B42","article-title":"Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs","author":"Malkov","year":"2018"},{"key":"B43","doi-asserted-by":"publisher","first-page":"1629","DOI":"10.1007\/s00439-021-02411-y","article-title":"Embeddings from protein language models predict conservation and variant effects","volume":"141","author":"Marquet","year":"2021","journal-title":"Hum. Genet."},{"key":"B44","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1038\/s41592-021-01359-1","article-title":"Method of the year: Protein structure prediction","volume":"19","author":"Marx","year":"2022","journal-title":"Nat. Methods"},{"key":"B45","doi-asserted-by":"publisher","first-page":"2","DOI":"10.3169\/mta.6.2","article-title":"[Invited paper] A survey of product quantization","volume":"6","author":"Matsui","year":"2018","journal-title":"ITE Trans. Media Technol. Appl."},{"key":"B46","article-title":"ColabFold-Making protein folding accessible to all","author":"Mirdita","year":"2021"},{"key":"B47","doi-asserted-by":"crossref","DOI":"10.1101\/2020.11.03.365932","article-title":"Protein structural alignments from sequence","author":"Morton","year":"2020"},{"key":"B48","doi-asserted-by":"crossref","DOI":"10.1101\/2022.03.10.483805","article-title":"CATHe: Detection of remote homologues for CATH superfamilies using embeddings from protein language models","author":"Nallapareddy","year":"2022"},{"key":"B49","doi-asserted-by":"publisher","first-page":"11703","DOI":"10.1073\/pnas.1707642114","article-title":"Complex evolutionary footprints revealed in an analysis of reused protein segments of diverse lengths","volume":"114","author":"Nepomnyachiy","year":"2017","journal-title":"Proc. Natl. Acad. Sci. U. S. A."},{"key":"B50","doi-asserted-by":"publisher","first-page":"1750","DOI":"10.1016\/j.csbj.2021.03.022","article-title":"the language of proteins: NLP, machine learning & protein sequences","volume":"19","author":"Ofer","year":"2021","journal-title":"Comput. Struct. Biotechnol. J."},{"key":"B51","doi-asserted-by":"publisher","first-page":"139","DOI":"10.1002\/prot.340140203","article-title":"Fast structure alignment for protein databank searching","volume":"14","author":"Orengo","year":"1992","journal-title":"Proteins."},{"key":"B52","doi-asserted-by":"publisher","first-page":"1093","DOI":"10.1016\/s0969-2126(97)00260-8","article-title":"Cath - a hierarchic classification of protein domain structures","volume":"5","author":"Orengo","year":"1997","journal-title":"Structure"},{"key":"B53","doi-asserted-by":"publisher","first-page":"145","DOI":"10.1006\/jsbi.2001.4398","article-title":"Review: What can structural classifications reveal about protein evolution?","volume":"134","author":"Orengo","year":"2001","journal-title":"J. Struct. Biol."},{"key":"B54","doi-asserted-by":"crossref","DOI":"10.5962\/bhl.title.118611","volume-title":"On the archetype and homologies of the vertebrate skeleton","author":"Owen","year":"1848"},{"key":"B55","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/N18-1202","article-title":"Deep contextualized word representations","author":"Peters","year":"2018"},{"key":"B56","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","author":"Raffel","year":"2019"},{"key":"B57","first-page":"9689","article-title":"Evaluating protein transfer learning with TAPE","author":"Rao","year":"2019"},{"key":"B58","doi-asserted-by":"publisher","first-page":"173","DOI":"10.1038\/nmeth.1818","article-title":"HHblits: Lightning-fast iterative protein sequence searching by HMM-HMM alignment","volume":"9","author":"Remmert","year":"2012","journal-title":"Nat. Methods"},{"key":"B59","doi-asserted-by":"publisher","first-page":"e2016239118","DOI":"10.1073\/pnas.2016239118","article-title":"Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences","volume":"118","author":"Rives","year":"2021","journal-title":"Proc. Natl. Acad. Sci. U. S. A."},{"key":"B60","doi-asserted-by":"publisher","first-page":"595","DOI":"10.1016\/s0022-2836(02)00016-5","article-title":"Enzyme function less conserved than anticipated","volume":"318","author":"Rost","year":"2002","journal-title":"J. Mol. Biol."},{"key":"B61","doi-asserted-by":"publisher","first-page":"S19","DOI":"10.1016\/s1359-0278(97)00059-x","article-title":"Protein structures sustain evolutionary drift","volume":"2","author":"Rost","year":"1997","journal-title":"Fold. Des."},{"key":"B62","doi-asserted-by":"publisher","first-page":"85","DOI":"10.1093\/protein\/12.2.85","article-title":"Twilight zone of protein sequence alignments","volume":"12","author":"Rost","year":"1999","journal-title":"Protein Eng. Des. Sel."},{"key":"B63","doi-asserted-by":"publisher","first-page":"D280","DOI":"10.1093\/nar\/gky1097","article-title":"Cath: Expanding the horizons of structure-based functional annotations for genome sequences","volume":"47","author":"Sillitoe","year":"2019","journal-title":"Nucleic acids Res."},{"key":"B64","doi-asserted-by":"publisher","first-page":"128","DOI":"10.1109\/msp.2007.914237","article-title":"Locality-sensitive hashing for finding nearest neighbors [lecture notes]","volume":"25","author":"Slaney","year":"2008","journal-title":"IEEE Signal Process. Mag."},{"key":"B65","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1016\/0022-2836(81)90087-5","article-title":"Identification of common molecular subsequences","volume":"147","author":"Smith","year":"1981","journal-title":"J. Mol. Biol."},{"key":"B66","doi-asserted-by":"publisher","first-page":"405","DOI":"10.1002\/(sici)1097-0134(199707)28:3<405::aid-prot10>3.0.co;2-l","article-title":"Pfam: A comprehensive database of protein domain families based on seed alignments","volume":"28","author":"Sonnhammer","year":"1997","journal-title":"Proteins."},{"key":"B67","doi-asserted-by":"crossref","DOI":"10.1093\/bioadv\/vbab035","article-title":"Light attention predicts protein location from the language of life","author":"Staerk","year":"2021"},{"key":"B68","doi-asserted-by":"publisher","first-page":"603","DOI":"10.1038\/s41592-019-0437-4","article-title":"Protein-level assembly increases protein sequence recovery from metagenomic samples manyfold","volume":"16","author":"Steinegger","year":"2019","journal-title":"Nat. Methods"},{"key":"B69","doi-asserted-by":"publisher","first-page":"2542","DOI":"10.1038\/s41467-018-04964-5","article-title":"Clustering huge protein sequence sets in linear time","volume":"9","author":"Steinegger","year":"2018","journal-title":"Nat. Commun."},{"key":"B70","doi-asserted-by":"publisher","first-page":"1026","DOI":"10.1038\/nbt.3988","article-title":"MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets","volume":"35","author":"Steinegger","year":"2017","journal-title":"Nat. Biotechnol."},{"key":"B71","doi-asserted-by":"publisher","first-page":"926","DOI":"10.1093\/bioinformatics\/btu739","article-title":"UniRef clusters: A comprehensive and scalable alternative for improving sequence similarity searches","volume":"31","author":"Suzek","year":"2015","journal-title":"Bioinformatics"},{"key":"B72","doi-asserted-by":"crossref","DOI":"10.1101\/2021.06.09.447770","article-title":"SignalP 6.0 achieves signal peptide prediction across all types using protein language models","author":"Teufel","year":"2021"},{"key":"B73","doi-asserted-by":"publisher","first-page":"590","DOI":"10.1038\/s41586-021-03828-1","article-title":"Highly accurate protein structure prediction for the human proteome","volume":"596","author":"Tunyasuvunakool","year":"2021","journal-title":"Nature"},{"key":"B74","doi-asserted-by":"publisher","first-page":"D158","DOI":"10.1093\/nar\/gkw1099","article-title":"UniProt: The universal protein knowledgebase","volume":"45","year":"2017","journal-title":"Nucleic Acids Res."},{"key":"B75","first-page":"5998","article-title":"Attention is all you need","author":"Vaswani","year":""},{"key":"B76","doi-asserted-by":"publisher","first-page":"1169","DOI":"10.1016\/j.str.2022.05.001","article-title":"Protein language-model embeddings for fast, accurate, and alignment-free protein structure prediction","volume":"30","author":"Weissenow","year":"2022","journal-title":"Structure"},{"key":"B77","doi-asserted-by":"publisher","first-page":"1257","DOI":"10.1006\/jmbi.2001.5293","article-title":"Within the twilight zone: A sensitive profile-profile comparison tool based on information theory","volume":"315","author":"Yona","year":"2002","journal-title":"J. Mol. Biol."}],"container-title":["Frontiers in Bioinformatics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fbinf.2022.1033775\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,11,17]],"date-time":"2022-11-17T03:29:15Z","timestamp":1668655755000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fbinf.2022.1033775\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,11,17]]},"references-count":77,"alternative-id":["10.3389\/fbinf.2022.1033775"],"URL":"https:\/\/doi.org\/10.3389\/fbinf.2022.1033775","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/2022.09.04.506527","asserted-by":"object"}]},"ISSN":["2673-7647"],"issn-type":[{"value":"2673-7647","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,11,17]]},"article-number":"1033775"}}