{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T18:55:14Z","timestamp":1781117714465,"version":"3.54.1"},"reference-count":25,"publisher":"Oxford University Press (OUP)","issue":"18","license":[{"start":{"date-parts":[[2020,6,24]],"date-time":"2020-06-24T00:00:00Z","timestamp":1592956800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CCF-15-18897"],"award-info":[{"award-number":["CCF-15-18897"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CNS-15-13263"],"award-info":[{"award-number":["CNS-15-13263"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CCF-19-34884"],"award-info":[{"award-number":["CCF-19-34884"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"name":"VPR office at Iowa State University"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,9,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>As the cost of sequencing decreases, the amount of data being deposited into public repositories is increasing rapidly. Public databases rely on the user to provide metadata for each submission that is prone to user error. Unfortunately, most public databases, such as non-redundant (NR), rely on user input and do not have methods for identifying errors in the provided metadata, leading to the potential for error propagation. Previous research on a small subset of the NR database analyzed misclassification based on sequence similarity. To the best of our knowledge, the amount of misclassification in the entire database has not been quantified. We propose a heuristic method to detect potentially misclassified taxonomic assignments in the NR database. We applied a curation technique and quality control to find the most probable taxonomic assignment. Our method incorporates provenance and frequency of each annotation from manually and computationally created databases and clustering information at 95% similarity.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>We found more than two million potentially taxonomically misclassified proteins in the NR database. Using simulated data, we show a high precision of 97% and a recall of 87% for detecting taxonomically misclassified proteins. The proposed approach and findings could also be applied to other databases.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation<\/jats:title>\n                  <jats:p>Source code, dataset, documentation, Jupyter notebooks and Docker container are available at https:\/\/github.com\/boalang\/nr.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information<\/jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btaa586","type":"journal-article","created":{"date-parts":[[2020,6,16]],"date-time":"2020-06-16T19:12:00Z","timestamp":1592334720000},"page":"4699-4705","source":"Crossref","is-referenced-by-count":41,"title":["Detecting and correcting misclassified sequences in the large-scale public databases"],"prefix":"10.1093","volume":"36","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1540-7797","authenticated-orcid":false,"given":"Hamid","family":"Bagheri","sequence":"first","affiliation":[{"name":"Department of Computer Science , Ames, IA 50011, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrew J","family":"Severin","sequence":"additional","affiliation":[{"name":"Genome Informatics Facility, Iowa State University , Ames, IA 50011, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hridesh","family":"Rajan","sequence":"additional","affiliation":[{"name":"Department of Computer Science , Ames, IA 50011, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2020,6,24]]},"reference":[{"key":"2023062213564374200_btaa586-B1","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1038\/75556","article-title":"Gene ontology: tool for the unification of biology","volume":"25","author":"Ashburner","year":"2000","journal-title":"Nat. Genet"},{"key":"2023062213564374200_btaa586-B2","doi-asserted-by":"crossref","DOI":"10.1186\/s12859-019-2967-2","article-title":"Shared data science infrastructure for genomics data","volume":"20","author":"Bagheri","year":"2019","journal-title":"BMC Bioinformatics"},{"key":"2023062213564374200_btaa586-B3","doi-asserted-by":"crossref","first-page":"D26","DOI":"10.1093\/nar\/gkn723","article-title":"Genbank","volume":"37","author":"Benson","year":"2009","journal-title":"Nucleic Acids Res"},{"key":"2023062213564374200_btaa586-B4","first-page":"394","volume-title":"Protein Structure","author":"Berman","year":"2003"},{"key":"2023062213564374200_btaa586-B5","doi-asserted-by":"crossref","first-page":"365","DOI":"10.1093\/nar\/gkg095","article-title":"The Swiss-Prot protein knowledgebase and its supplement trembl in 2003","volume":"31","author":"Boeckmann","year":"2003","journal-title":"Nucleic Acids Res"},{"key":"2023062213564374200_btaa586-B6","doi-asserted-by":"crossref","first-page":"954","DOI":"10.1101\/gr.245373.118","article-title":"Human contamination in bacterial genomes has created thousands of spurious proteins","volume":"29","author":"Breitwieser","year":"2019","journal-title":"Genome Res"},{"key":"2023062213564374200_btaa586-B7","author":"Chu","year":"2015"},{"key":"2023062213564374200_btaa586-B8","first-page":"D204","article-title":"Uniprot: a hub for protein information","volume":"43","year":"2014","journal-title":"Nucleic Acids Res"},{"key":"2023062213564374200_btaa586-B9","doi-asserted-by":"crossref","first-page":"e5030","DOI":"10.7717\/peerj.5030","article-title":"Taxonomy annotation and guide tree errors in 16s RRNA databases","volume":"6","author":"Edgar","year":"2018","journal-title":"Peer J"},{"key":"2023062213564374200_btaa586-B10","doi-asserted-by":"crossref","first-page":"3150","DOI":"10.1093\/bioinformatics\/bts565","article-title":"Cd-hit: accelerated for clustering the next-generation sequencing data","volume":"28","author":"Fu","year":"2012","journal-title":"Bioinformatics"},{"key":"2023062213564374200_btaa586-B11","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1007\/978-1-4939-3743-1_9","volume-title":"The Gene Ontology Handbook","author":"Holliday","year":"2017"},{"key":"2023062213564374200_btaa586-B12","doi-asserted-by":"crossref","first-page":"1635","DOI":"10.1093\/molbev\/msw046","article-title":"Ete 3: reconstruction, analysis, and visualization of phylogenomic data","volume":"33","author":"Huerta-Cepas","year":"2016","journal-title":"Mol. Biol. Evol"},{"key":"2023062213564374200_btaa586-B13","doi-asserted-by":"crossref","first-page":"5022","DOI":"10.1093\/nar\/gkw396","article-title":"Phylogeny-aware identification and correction of taxonomically mislabeled sequences","volume":"44","author":"Kozlov","year":"2016","journal-title":"Nucleic Acids Res"},{"key":"2023062213564374200_btaa586-B14","doi-asserted-by":"crossref","first-page":"D225","DOI":"10.1093\/nar\/gkq1189","article-title":"Cdd: a conserved domain database for the functional annotation of proteins","volume":"39","author":"Marchler-Bauer","year":"2011","journal-title":"Nucleic Acids Res"},{"key":"2023062213564374200_btaa586-B15","doi-asserted-by":"crossref","first-page":"610","DOI":"10.1038\/ismej.2011.139","article-title":"An improved greengenes taxonomy with explicit ranks for ecological and evolutionary analyses of bacteria and archaea","volume":"6","author":"McDonald","year":"2012","journal-title":"ISME J"},{"key":"2023062213564374200_btaa586-B16","doi-asserted-by":"crossref","first-page":"W479","DOI":"10.1093\/nar\/gky359","article-title":"Aai-profiler: fast proteome-wide exploratory analysis reveals taxonomic identity, misclassification and contamination","volume":"46","author":"Medlar","year":"2018","journal-title":"Nucleic Acids Res"},{"key":"2023062213564374200_btaa586-B17","doi-asserted-by":"crossref","first-page":"2195","DOI":"10.1093\/bioinformatics\/bty099","article-title":"Victree: an automated framework for taxonomic classification from protein sequences","volume":"34","author":"Modha","year":"2018","journal-title":"Bioinformatics"},{"key":"2023062213564374200_btaa586-B18","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1186\/1944-3277-10-18","article-title":"Large-scale contamination of microbial isolate genomes by illumina phix control","volume":"10","author":"Mukherjee","year":"2015","journal-title":"Stand. Genomic Sci"},{"key":"2023062213564374200_btaa586-B19","doi-asserted-by":"crossref","DOI":"10.1093\/database\/bat053","article-title":"Mispred: a resource for identification of erroneous protein sequences in public databases","volume":"2013","author":"Nagy","year":"2013","journal-title":"Database"},{"key":"2023062213564374200_btaa586-B20","doi-asserted-by":"crossref","DOI":"10.1093\/database\/bau032","article-title":"FixPred: a resource for correction of erroneous protein sequences","volume":"2014","author":"Nagy","year":"2014","journal-title":"Database"},{"key":"2023062213564374200_btaa586-B21","doi-asserted-by":"crossref","first-page":"353","DOI":"10.1186\/1471-2105-9-353","article-title":"Identification and correction of abnormal, incomplete and mispredicted proteins in public databases","volume":"9","author":"Nagy","year":"2008","journal-title":"BMC Bioinformatics"},{"key":"2023062213564374200_btaa586-B22","doi-asserted-by":"crossref","first-page":"D61","DOI":"10.1093\/nar\/gkl842","article-title":"NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins","volume":"35","author":"Pruitt","year":"2007","journal-title":"Nucleic Acids Res"},{"key":"2023062213564374200_btaa586-B23","doi-asserted-by":"crossref","first-page":"e1000605","DOI":"10.1371\/journal.pcbi.1000605","article-title":"Annotation error in public databases: misannotation of molecular function in enzyme superfamilies","volume":"5","author":"Schnoes","year":"2009","journal-title":"PLoS Comput. Biol"},{"key":"2023062213564374200_btaa586-B24","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1038\/s41467-018-04964-5","article-title":"Clustering huge protein sequence sets in linear time","volume":"9","author":"Steinegger","year":"2018","journal-title":"Nat. Commun"},{"key":"2023062213564374200_btaa586-B25","doi-asserted-by":"crossref","first-page":"345","DOI":"10.1093\/nar\/gkg040","article-title":"The protein information resource","volume":"31","author":"Wu","year":"2003","journal-title":"Nucleic Acids Res"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btaa586\/33834587\/btaa586.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/36\/18\/4699\/50677613\/btaa586.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/36\/18\/4699\/50677613\/btaa586.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,23]],"date-time":"2023-06-23T19:26:43Z","timestamp":1687548403000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/36\/18\/4699\/5862012"}},"subtitle":[],"editor":[{"given":"Arne","family":"Elofsson","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2020,6,24]]},"references-count":25,"journal-issue":{"issue":"18","published-print":{"date-parts":[[2020,9,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btaa586","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2020,9,15]]},"published":{"date-parts":[[2020,6,24]]}}}