{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T02:01:48Z","timestamp":1787018508850,"version":"3.56.0"},"reference-count":20,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,1,7]],"date-time":"2025-01-07T00:00:00Z","timestamp":1736208000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,1,7]],"date-time":"2025-01-07T00:00:00Z","timestamp":1736208000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"JST NBDC","award":["JPMJND2206"],"award-info":[{"award-number":["JPMJND2206"]}]},{"name":"JSPS KAKENHI","award":["JP22H04925"],"award-info":[{"award-number":["JP22H04925"]}]},{"DOI":"10.13039\/100009619","name":"AMED","doi-asserted-by":"crossref","award":["JP23wm0225029"],"award-info":[{"award-number":["JP23wm0225029"]}],"id":[{"id":"10.13039\/100009619","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100018732","name":"Ohsumi Frontier Science Foundation","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100018732","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Background<\/jats:title>\n                    <jats:p>Accurate taxonomic classification in genome databases is essential for reliable biological research and effective data sharing. Mislabeling or inaccuracies in genome annotations can lead to incorrect scientific conclusions and hinder the reproducibility of research findings. Despite advances in genome analysis techniques, challenges persist in ensuring precise and reliable taxonomic assignments. Existing tools for genome verification often involve extensive computational resources or lengthy processing times, which can limit their accessibility and scalability for large-scale projects. There is a need for more efficient, user-friendly solutions that can handle diverse datasets and provide accurate results with minimal computational demands. This work aimed to address these challenges by introducing a novel tool that enhances taxonomic accuracy, offers a user-friendly interface, and supports large-scale analyses.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>We introduce a novel tool for the quality control and taxonomic classification tool of prokaryotic genomes, called DFAST_QC, which is available as both a command-line tool and a web service. DFAST_QC can quickly identify species based on NCBI and GTDB taxonomies by combining genome-distance calculations using MASH with ANI calculations using Skani. We evaluated DFAST_QC's performance in species identification and found it to be highly consistent with existing taxonomic standards, successfully identifying species across diverse datasets. In several cases, DFAST_QC identified potential mislabeling of species names in public databases and highlighted discrepancies in current classifications, demonstrating its capability to uncover errors and enhance taxonomic accuracy. Additionally, the tool\u2019s efficient design allows it to operate smoothly on local machines with minimal computational requirements, making it a practical choice for large-scale genome projects.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusions<\/jats:title>\n                    <jats:p>\n                      DFAST_QC is a reliable and efficient tool for accurate taxonomic identification and genome quality control, well-suited for large-scale genomic studies. Its compatibility with limited-resource environments, combined with its user-friendly design, ensures seamless integration into existing workflows. DFAST_QC's ability to refine species assignments in public databases highlights its value as a complementary tool for maintaining and enhancing the accuracy of taxonomic data in genomic research. The web version is available at\n                      <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/dfast.ddbj.nig.ac.jp\/dqc\/submit\/\">https:\/\/dfast.ddbj.nig.ac.jp\/dqc\/submit\/<\/jats:ext-link>\n                      , and the source code for local use can be found at\n                      <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/nigyta\/dfast_qc\">https:\/\/github.com\/nigyta\/dfast_qc<\/jats:ext-link>\n                      .\n                    <\/jats:p>\n                  <\/jats:sec>","DOI":"10.1186\/s12859-024-06030-y","type":"journal-article","created":{"date-parts":[[2025,1,7]],"date-time":"2025-01-07T06:16:16Z","timestamp":1736230576000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":31,"title":["DFAST_QC: quality assessment and taxonomic identification tool for prokaryotic Genomes"],"prefix":"10.1186","volume":"26","author":[{"given":"Mohamed","family":"Elmanzalawi","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Takatomo","family":"Fujisawa","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hiroshi","family":"Mori","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yasukazu","family":"Nakamura","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yasuhiro","family":"Tanizawa","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,1,7]]},"reference":[{"key":"6030_CR1","doi-asserted-by":"publisher","first-page":"4699","DOI":"10.1093\/bioinformatics\/btaa586","volume":"36","author":"H Bagheri","year":"2020","unstructured":"Bagheri H, Severin AJ, Rajan H. Detecting and correcting misclassified sequences in the large-scale public databases. Bioinformatics. 2020;36:4699\u2013705.","journal-title":"Bioinformatics"},{"key":"6030_CR2","doi-asserted-by":"publisher","first-page":"bbac416","DOI":"10.1093\/bib\/bbac416","volume":"23","author":"B Goudey","year":"2022","unstructured":"Goudey B, Geard N, Verspoor K, Zobel J. Propagation, detection and correction of errors using the sequence database network. Brief Bioinform. 2022;23:bbac416.","journal-title":"Brief Bioinform"},{"key":"6030_CR3","doi-asserted-by":"publisher","first-page":"2386","DOI":"10.1099\/ijsem.0.002809","volume":"68","author":"S Ciufo","year":"2018","unstructured":"Ciufo S, Kannan S, Sharma S, Badretdin A, Clark K, Turner S, et al. Using average nucleotide identity to improve taxonomic assignments in prokaryotic genomes at the NCBI. Int J Syst Evol Microbiol. 2018;68:2386\u201392.","journal-title":"Int J Syst Evol Microbiol"},{"issue":"Pt 1","key":"6030_CR4","doi-asserted-by":"publisher","first-page":"81","DOI":"10.1099\/ijs.0.64483-0","volume":"57","author":"J Goris","year":"2007","unstructured":"Goris J, Konstantinidis KT, Klappenbach JA, Coenye T, Vandamme P, Tiedje JM. DNA-DNA hybridization values and their relationship to whole-genome sequence similarities. Int J Syst Evol Microbiol. 2007;57(Pt 1):81\u201391.","journal-title":"Int J Syst Evol Microbiol"},{"key":"6030_CR5","first-page":"baaa062","volume":"2020","author":"CL Schoch","year":"2020","unstructured":"Schoch CL, Ciufo S, Domrachev M, Hotton CL, Kannan S, Khovanskaya R, et al. NCBI Taxonomy: a comprehensive update on curation, resources and tools. Database J Biol Databases Curation. 2020;2020:baaa062.","journal-title":"Database J Biol Databases Curation"},{"key":"6030_CR6","doi-asserted-by":"publisher","first-page":"D801","DOI":"10.1093\/nar\/gkab902","volume":"50","author":"JP Meier-Kolthoff","year":"2022","unstructured":"Meier-Kolthoff JP, Carbasse JS, Peinado-Olarte RL, G\u00f6ker M. TYGS and LPSN: a database tandem for fast and reliable genome-based classification and nomenclature of prokaryotes. Nucleic Acids Res. 2022;50:D801\u20137.","journal-title":"Nucleic Acids Res"},{"key":"6030_CR7","doi-asserted-by":"publisher","first-page":"2182","DOI":"10.1038\/s41467-019-10210-3","volume":"10","author":"JP Meier-Kolthoff","year":"2019","unstructured":"Meier-Kolthoff JP, G\u00f6ker M. TYGS is an automated high-throughput platform for state-of-the-art genome-based taxonomy. Nat Commun. 2019;10:2182.","journal-title":"Nat Commun"},{"key":"6030_CR8","doi-asserted-by":"publisher","first-page":"W282","DOI":"10.1093\/nar\/gky467","volume":"46","author":"LM Rodriguez-R","year":"2018","unstructured":"Rodriguez-R LM, Gunturu S, Harvey WT, Rossell\u00f3-Mora R, Tiedje JM, Cole JR, et al. The microbial genomes atlas (MiGA) webserver: taxonomic and gene diversity analysis of Archaea and Bacteria at the whole genome level. Nucleic Acids Res. 2018;46:W282\u20138.","journal-title":"Nucleic Acids Res"},{"key":"6030_CR9","doi-asserted-by":"publisher","first-page":"5315","DOI":"10.1093\/bioinformatics\/btac672","volume":"38","author":"P-A Chaumeil","year":"2022","unstructured":"Chaumeil P-A, Mussig AJ, Hugenholtz P, Parks DH. GTDB-Tk v2: memory friendly classification with the genome taxonomy database. Bioinformatics. 2022;38:5315\u20136.","journal-title":"Bioinformatics"},{"key":"6030_CR10","doi-asserted-by":"publisher","first-page":"1037","DOI":"10.1093\/bioinformatics\/btx713","volume":"34","author":"Y Tanizawa","year":"2018","unstructured":"Tanizawa Y, Fujisawa T, Nakamura Y. DFAST: a flexible prokaryotic genome annotation pipeline for faster genome publication. Bioinformatics. 2018;34:1037\u20139.","journal-title":"Bioinformatics"},{"key":"6030_CR11","doi-asserted-by":"publisher","first-page":"132","DOI":"10.1186\/s13059-016-0997-x","volume":"17","author":"BD Ondov","year":"2016","unstructured":"Ondov BD, Treangen TJ, Melsted P, Mallonee AB, Bergman NH, Koren S, et al. Mash: fast genome and metagenome distance estimation using MinHash. Genome Biol. 2016;17:132.","journal-title":"Genome Biol"},{"key":"6030_CR12","doi-asserted-by":"publisher","first-page":"1661","DOI":"10.1038\/s41592-023-02018-3","volume":"20","author":"J Shaw","year":"2023","unstructured":"Shaw J, Yu YW. Fast and robust metagenomic sequence comparison through sparse chaining with skani. Nat Methods. 2023;20:1661\u20135.","journal-title":"Nat Methods"},{"key":"6030_CR13","doi-asserted-by":"publisher","first-page":"1043","DOI":"10.1101\/gr.186072.114","volume":"25","author":"DH Parks","year":"2015","unstructured":"Parks DH, Imelfort M, Skennerton CT, Hugenholtz P, Tyson GW. CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res. 2015;25:1043\u201355.","journal-title":"Genome Res"},{"key":"6030_CR14","doi-asserted-by":"publisher","first-page":"D785","DOI":"10.1093\/nar\/gkab776","volume":"50","author":"DH Parks","year":"2022","unstructured":"Parks DH, Chuvochina M, Rinke C, Mussig AJ, Chaumeil P-A, Hugenholtz P. GTDB: an ongoing census of bacterial and archaeal diversity through a phylogenetically consistent, rank normalized and complete genome-based taxonomy. Nucleic Acids Res. 2022;50:D785\u201394.","journal-title":"Nucleic Acids Res"},{"key":"6030_CR15","doi-asserted-by":"publisher","first-page":"499","DOI":"10.1038\/s41587-020-0718-6","volume":"39","author":"S Nayfach","year":"2021","unstructured":"Nayfach S, Roux S, Seshadri R, Udwary D, Varghese N, Schulz F, et al. A genomic catalog of Earth\u2019s microbiomes. Nat Biotechnol. 2021;39:499\u2013509.","journal-title":"Nat Biotechnol"},{"key":"6030_CR16","doi-asserted-by":"publisher","DOI":"10.1099\/ijsem.0.005707","volume":"73","author":"S Kannan","year":"2023","unstructured":"Kannan S, Sharma S, Ciufo S, Clark K, Turner S, Kitts PA, et al. Collection and curation of prokaryotic genome assemblies from type strains at NCBI. Int J Syst Evol Microbiol. 2023;73: 005707.","journal-title":"Int J Syst Evol Microbiol"},{"key":"6030_CR17","doi-asserted-by":"publisher","first-page":"895","DOI":"10.1099\/ijsem.0.003276","volume":"69","author":"L Wu","year":"2019","unstructured":"Wu L, Ma J. The Global Catalogue of Microorganisms (GCM) 10K type strain sequencing project: providing services to taxonomists for standard genome sequencing and annotation. Int J Syst Evol Microbiol. 2019;69:895\u20138.","journal-title":"Int J Syst Evol Microbiol"},{"key":"6030_CR18","doi-asserted-by":"publisher","first-page":"D694","DOI":"10.1093\/nar\/gkaa957","volume":"49","author":"W Shi","year":"2021","unstructured":"Shi W, Sun Q, Fan G, Hideaki S, Moriya O, Itoh T, et al. gcType: a high-quality type strain genome database for microbial phylogenetic and functional research. Nucleic Acids Res. 2021;49:D694-705.","journal-title":"Nucleic Acids Res"},{"key":"6030_CR19","doi-asserted-by":"publisher","first-page":"006300","DOI":"10.1099\/ijsem.0.006300","volume":"74","author":"R Riesco","year":"2024","unstructured":"Riesco R, Trujillo ME. Update on the proposed minimal standards for the use of genome data for the taxonomy of prokaryotes. Int J Syst Evol Microbiol. 2024;74:006300.","journal-title":"Int J Syst Evol Microbiol"},{"key":"6030_CR20","doi-asserted-by":"publisher","first-page":"486","DOI":"10.1093\/sysbio\/syad068","volume":"73","author":"SS Renner","year":"2024","unstructured":"Renner SS, Scherz MD, Schoch CL, Gottschling M, Vences M. Improving the gold standard in NCBI GenBank and related databases: DNA sequences from type specimens and type strains. Syst Biol. 2024;73:486\u201394.","journal-title":"Syst Biol"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-024-06030-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12859-024-06030-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-024-06030-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,7]],"date-time":"2025-01-07T18:02:08Z","timestamp":1736272928000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-024-06030-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,7]]},"references-count":20,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["6030"],"URL":"https:\/\/doi.org\/10.1186\/s12859-024-06030-y","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/2024.07.22.604526","asserted-by":"object"}]},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,7]]},"assertion":[{"value":"23 August 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 December 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 January 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"3"}}