{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,8]],"date-time":"2026-08-08T09:05:06Z","timestamp":1786179906489,"version":"3.56.0"},"reference-count":31,"publisher":"Oxford University Press (OUP)","issue":"6","license":[{"start":{"date-parts":[[2016,10,2]],"date-time":"2016-10-02T00:00:00Z","timestamp":1475366400000},"content-version":"vor","delay-in-days":689,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2015,3,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: UniRef databases provide full-scale clustering of UniProtKB sequences and are utilized for a broad range of applications, particularly similarity-based functional annotation. Non-redundancy and intra-cluster homogeneity in UniRef were recently improved by adding a sequence length overlap threshold. Our hypothesis is that these improvements would enhance the speed and sensitivity of similarity searches and improve the consistency of annotation within clusters.<\/jats:p>\n               <jats:p>Results: Intra-cluster molecular function consistency was examined by analysis of Gene Ontology terms. Results show that UniRef clusters bring together proteins of identical molecular function in more than 97% of the clusters, implying that clusters are useful for annotation and can also be used to detect annotation inconsistencies. To examine coverage in similarity results, BLASTP searches against UniRef50 followed by expansion of the hit lists with cluster members demonstrated advantages compared with searches against UniProtKB sequences; the searches are concise (\u223c7 times shorter hit list before expansion), faster (\u223c6 times) and more sensitive in detection of remote similarities (&amp;gt;96% recall at e-value &amp;lt;0.0001). Our results support the use of UniRef clusters as a comprehensive and scalable alternative to native sequence databases for similarity searches and reinforces its reliability for use in functional annotation.<\/jats:p>\n               <jats:p>Availability and implementation: Web access and file download from UniProt website at http:\/\/www.uniprot.org\/uniref and ftp:\/\/ftp.uniprot.org\/pub\/databases\/uniprot\/uniref. BLAST searches against UniRef are available at http:\/\/www.uniprot.org\/blast\/<\/jats:p>\n               <jats:p>Contact: \u00a0huang@dbi.udel.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/btu739","type":"journal-article","created":{"date-parts":[[2014,11,15]],"date-time":"2014-11-15T04:10:48Z","timestamp":1416024648000},"page":"926-932","source":"Crossref","is-referenced-by-count":1931,"title":["UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches"],"prefix":"10.1093","volume":"31","author":[{"given":"Baris E.","family":"Suzek","sequence":"first","affiliation":[{"name":"1 Protein Information Resource, Georgetown University Medical Center, Washington, DC 20007, USA, 2Department of Computer Engineering, Mu\u011fla S\u0131tk\u0131 Ko\u00e7man University, Mu\u011fla 48000, Turkey, 3Center for Bioinformatics and Computational Biology and Protein Information Resource, University of Delaware, Newark, DE 19711, USA, 4European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SD, UK and 5Swiss Institute of Bioinformatics, Centre Medical Universitaire, 1 rue Michel Servet, 1211 Geneva 4, Switzerland"},{"name":"1 Protein Information Resource, Georgetown University Medical Center, Washington, DC 20007, USA, 2Department of Computer Engineering, Mu\u011fla S\u0131tk\u0131 Ko\u00e7man University, Mu\u011fla 48000, Turkey, 3Center for Bioinformatics and Computational Biology and Protein Information Resource, University of Delaware, Newark, DE 19711, USA, 4European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SD, UK and 5Swiss Institute of Bioinformatics, Centre Medical Universitaire, 1 rue Michel Servet, 1211 Geneva 4, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuqi","family":"Wang","sequence":"additional","affiliation":[{"name":"1 Protein Information Resource, Georgetown University Medical Center, Washington, DC 20007, USA, 2Department of Computer Engineering, Mu\u011fla S\u0131tk\u0131 Ko\u00e7man University, Mu\u011fla 48000, Turkey, 3Center for Bioinformatics and Computational Biology and Protein Information Resource, University of Delaware, Newark, DE 19711, USA, 4European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SD, UK and 5Swiss Institute of Bioinformatics, Centre Medical Universitaire, 1 rue Michel Servet, 1211 Geneva 4, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongzhan","family":"Huang","sequence":"additional","affiliation":[{"name":"1 Protein Information Resource, Georgetown University Medical Center, Washington, DC 20007, USA, 2Department of Computer Engineering, Mu\u011fla S\u0131tk\u0131 Ko\u00e7man University, Mu\u011fla 48000, Turkey, 3Center for Bioinformatics and Computational Biology and Protein Information Resource, University of Delaware, Newark, DE 19711, USA, 4European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SD, UK and 5Swiss Institute of Bioinformatics, Centre Medical Universitaire, 1 rue Michel Servet, 1211 Geneva 4, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peter B.","family":"McGarvey","sequence":"additional","affiliation":[{"name":"1 Protein Information Resource, Georgetown University Medical Center, Washington, DC 20007, USA, 2Department of Computer Engineering, Mu\u011fla S\u0131tk\u0131 Ko\u00e7man University, Mu\u011fla 48000, Turkey, 3Center for Bioinformatics and Computational Biology and Protein Information Resource, University of Delaware, Newark, DE 19711, USA, 4European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SD, UK and 5Swiss Institute of Bioinformatics, Centre Medical Universitaire, 1 rue Michel Servet, 1211 Geneva 4, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cathy H.","family":"Wu","sequence":"additional","affiliation":[{"name":"1 Protein Information Resource, Georgetown University Medical Center, Washington, DC 20007, USA, 2Department of Computer Engineering, Mu\u011fla S\u0131tk\u0131 Ko\u00e7man University, Mu\u011fla 48000, Turkey, 3Center for Bioinformatics and Computational Biology and Protein Information Resource, University of Delaware, Newark, DE 19711, USA, 4European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SD, UK and 5Swiss Institute of Bioinformatics, Centre Medical Universitaire, 1 rue Michel Servet, 1211 Geneva 4, Switzerland"},{"name":"1 Protein Information Resource, Georgetown University Medical Center, Washington, DC 20007, USA, 2Department of Computer Engineering, Mu\u011fla S\u0131tk\u0131 Ko\u00e7man University, Mu\u011fla 48000, Turkey, 3Center for Bioinformatics and Computational Biology and Protein Information Resource, University of Delaware, Newark, DE 19711, USA, 4European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SD, UK and 5Swiss Institute of Bioinformatics, Centre Medical Universitaire, 1 rue Michel Servet, 1211 Geneva 4, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"name":"the UniProt Consortium","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2014,11,13]]},"reference":[{"key":"2023020116192713500_btu739-B1","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1016\/S0022-2836(05)80360-2","article-title":"Basic local alignment search tool","volume":"215","author":"Altschul","year":"1990","journal-title":"J. Mol. Biol."},{"key":"2023020116192713500_btu739-B2","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1038\/75556","article-title":"Gene ontology: tool for the unification of biology. The Gene Ontology Consortium","volume":"25","author":"Ashburner","year":"2000","journal-title":"Nat. Genet."},{"key":"2023020116192713500_btu739-B3","doi-asserted-by":"crossref","first-page":"594","DOI":"10.1089\/cmb.2007.R005","article-title":"Clustered sequence representation for fast homology search","volume":"14","author":"Cameron","year":"2007","journal-title":"J. Comput. Biol."},{"key":"2023020116192713500_btu739-B4","doi-asserted-by":"crossref","first-page":"383","DOI":"10.1186\/1471-2105-11-383","article-title":"The oligodeoxynucleotide sequences corresponding to never-expressed peptide motifs are mainly located in the non-coding strand","volume":"11","author":"Capone","year":"2010","journal-title":"BMC Bioinformatics"},{"key":"2023020116192713500_btu739-B5","doi-asserted-by":"crossref","first-page":"S3","DOI":"10.1186\/1471-2105-12-S4-S3","article-title":"Improving the prediction of disease-related variants using protein three-dimensional structure","volume":"12","author":"Capriotti","year":"2011","journal-title":"BMC Bioinformatics"},{"key":"2023020116192713500_btu739-B6","doi-asserted-by":"crossref","first-page":"310","DOI":"10.1016\/j.ygeno.2011.06.010","article-title":"A new disease-specific machine learning approach for the prediction of cancer-causing missense variants","volume":"98","author":"Capriotti","year":"2011","journal-title":"Genomics"},{"key":"2023020116192713500_btu739-B7","doi-asserted-by":"crossref","first-page":"S1","DOI":"10.1186\/1471-2105-13-S4-S1","article-title":"Accurate multiple sequence alignment of transmembrane proteins with PSI-Coffee","volume":"13","author":"Chang","year":"2012","journal-title":"BMC Bioinformatics"},{"key":"2023020116192713500_btu739-B8","doi-asserted-by":"crossref","first-page":"e18910","DOI":"10.1371\/journal.pone.0018910","article-title":"Representative proteomes: a stable, scalable and unbiased proteome set for sequence analysis and functional annotation","volume":"6","author":"Chen","year":"2011","journal-title":"PLoS One"},{"key":"2023020116192713500_btu739-B9","doi-asserted-by":"crossref","first-page":"e3515","DOI":"10.1371\/journal.pone.0003515","article-title":"A computational screen for type I polyketide synthases in metagenomics shotgun data","volume":"3","author":"Foerstner","year":"2008","journal-title":"PLoS One"},{"key":"2023020116192713500_btu739-B10","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1016\/S0097-8485(96)80004-0","article-title":"Use of receiver operating characteristic (ROC) analysis to evaluate sequence matching","volume":"20","author":"Gribskov","year":"1996","journal-title":"Comput. Chem."},{"key":"2023020116192713500_btu739-B11","doi-asserted-by":"crossref","first-page":"D306","DOI":"10.1093\/nar\/gkr948","article-title":"InterPro in 2011: new developments in the family and domain prediction database","volume":"40","author":"Hunter","year":"2012","journal-title":"Nucleic Acids Res."},{"key":"2023020116192713500_btu739-B12","first-page":"93","article-title":"Clustering of database sequences for fast homology search using upper bounds on alignment score","volume":"15","author":"Itoh","year":"2004","journal-title":"Genome Informatics"},{"key":"2023020116192713500_btu739-B13","doi-asserted-by":"crossref","first-page":"2618","DOI":"10.1093\/bioinformatics\/bti386","article-title":"The properties of protein family space depend on experimental design","volume":"21","author":"Kunin","year":"2005","journal-title":"Bioinformatics"},{"key":"2023020116192713500_btu739-B14","doi-asserted-by":"crossref","first-page":"603","DOI":"10.1002\/prot.20409","article-title":"Identification and distribution of protein families in 120 completed genomes using Gene 3D","volume":"59","author":"Lee","year":"2005","journal-title":"Proteins"},{"key":"2023020116192713500_btu739-B15","doi-asserted-by":"crossref","first-page":"1658","DOI":"10.1093\/bioinformatics\/btl158","article-title":"Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences","volume":"22","author":"Li","year":"2006","journal-title":"Bioinformatics"},{"key":"2023020116192713500_btu739-B16","doi-asserted-by":"crossref","first-page":"282","DOI":"10.1093\/bioinformatics\/17.3.282","article-title":"Clustering of highly homologous sequences to reduce the size of large protein databases","volume":"17","author":"Li","year":"2001","journal-title":"Bioinformatics"},{"key":"2023020116192713500_btu739-B17","doi-asserted-by":"crossref","first-page":"643","DOI":"10.1093\/protein\/15.8.643","article-title":"Sequence clustering strategies improve remote homology recognitions while reducing search times","volume":"15","author":"Li","year":"2002","journal-title":"Protein Eng."},{"key":"2023020116192713500_btu739-B18","doi-asserted-by":"crossref","first-page":"i41","DOI":"10.1093\/bioinformatics\/btn174","article-title":"Efficient algorithms for accurate hierarchical clustering of huge datasets: tackling the entire protein space","volume":"24","author":"Loewenstein","year":"2008","journal-title":"Bioinformatics"},{"key":"2023020116192713500_btu739-B19","doi-asserted-by":"crossref","first-page":"238","DOI":"10.4056\/sigs.561626","article-title":"Quantifying protein function specificity in the gene ontology","volume":"2","author":"Louie","year":"2010","journal-title":"Stand. Genomic Sci."},{"key":"2023020116192713500_btu739-B20","doi-asserted-by":"crossref","first-page":"RESEARCH0040","DOI":"10.1186\/gb-2002-3-8-research0040","article-title":"The dominance of the population by a selected few: power-law behaviour applies to a wide variety of genomic properties","volume":"3","author":"Luscombe","year":"2002","journal-title":"Genome Biol."},{"key":"2023020116192713500_btu739-B21","doi-asserted-by":"crossref","first-page":"e54422","DOI":"10.1371\/journal.pone.0054422","article-title":"Increasing sequence search sensitivity with transitive alignments","volume":"8","author":"Malde","year":"2013","journal-title":"PLoS One"},{"key":"2023020116192713500_btu739-B22","doi-asserted-by":"crossref","first-page":"458","DOI":"10.1093\/bioinformatics\/16.5.458","article-title":"RSDB: representative protein sequence databases have high information content","volume":"16","author":"Park","year":"2000","journal-title":"Bioinformatics"},{"key":"2023020116192713500_btu739-B23","doi-asserted-by":"crossref","first-page":"D290","DOI":"10.1093\/nar\/gkr1065","article-title":"The Pfam protein families database","volume":"40","author":"Punta","year":"2012","journal-title":"Nucleic Acids Res."},{"key":"2023020116192713500_btu739-B24","doi-asserted-by":"crossref","first-page":"e1000431","DOI":"10.1371\/journal.pcbi.1000431","article-title":"The Gene Ontology\u2019s Reference Genome Project: a unified framework for functional annotation across species","volume":"5","author":"Reference Genome Group of the Gene Ontology Consortium","year":"2009","journal-title":"PLoS Comput. Biol."},{"key":"2023020116192713500_btu739-B25","doi-asserted-by":"crossref","first-page":"123","DOI":"10.1186\/1471-2148-10-123","article-title":"Gene duplication and the origins of morphological complexity in pancrustacean eyes, a genomic approach","volume":"10","author":"Rivera","year":"2010","journal-title":"BMC Evol. Biol."},{"key":"2023020116192713500_btu739-B26","doi-asserted-by":"crossref","first-page":"539","DOI":"10.1038\/msb.2011.75","article-title":"Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega","volume":"7","author":"Sievers","year":"2011","journal-title":"Mol. Syst. Biol."},{"issue":"Web Server issue","key":"2023020116192713500_btu739-B27","doi-asserted-by":"crossref","first-page":"W452","DOI":"10.1093\/nar\/gks539","article-title":"SIFT web server: predicting effects of amino acid substitutions on proteins","volume":"40","author":"Sim","year":"2012","journal-title":"Nucleic Acids Res."},{"key":"2023020116192713500_btu739-B28","doi-asserted-by":"crossref","first-page":"1282","DOI":"10.1093\/bioinformatics\/btm098","article-title":"UniRef: comprehensive and non-redundant UniProt reference clusters","volume":"23","author":"Suzek","year":"2007","journal-title":"Bioinformatics"},{"key":"2023020116192713500_btu739-B29","first-page":"D43","article-title":"Update on activities at the Universal Protein Resource (UniProt) in 2013","volume":"41","author":"UniProt","year":"2013","journal-title":"Nucleic Acids Res."},{"key":"2023020116192713500_btu739-B30","doi-asserted-by":"crossref","first-page":"427","DOI":"10.4056\/sigs.2945050","article-title":"VIROME: a standard operating procedure for analysis of viral metagenome sequences","volume":"6","author":"Wommack","year":"2012","journal-title":"Stand. Genomic Sci."},{"key":"2023020116192713500_btu739-B31","doi-asserted-by":"crossref","first-page":"D112","DOI":"10.1093\/nar\/gkh097","article-title":"PIRSF: family classification system at the Protein Information Resource","volume":"32","author":"Wu","year":"2004","journal-title":"Nucleic Acids Res."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/31\/6\/926\/49011550\/bioinformatics_31_6_926.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/31\/6\/926\/49011550\/bioinformatics_31_6_926.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,2]],"date-time":"2023-02-02T00:33:23Z","timestamp":1675298003000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/31\/6\/926\/214968"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,11,13]]},"references-count":31,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2015,3,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btu739","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2015,3,15]]},"published":{"date-parts":[[2014,11,13]]}}}