{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T13:31:14Z","timestamp":1786627874429,"version":"3.56.0"},"reference-count":13,"publisher":"Oxford University Press (OUP)","issue":"10","license":[{"start":{"date-parts":[[2023,9,27]],"date-time":"2023-09-27T00:00:00Z","timestamp":1695772800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000051","name":"National Human Genome Research Institute","doi-asserted-by":"publisher","award":["R01HG010040"],"award-info":[{"award-number":["R01HG010040"]}],"id":[{"id":"10.13039\/100000051","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100014989","name":"Chan-Zuckerberg Initiative","doi-asserted-by":"crossref","award":["237653"],"award-info":[{"award-number":["237653"]}],"id":[{"id":"10.13039\/100014989","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2023,10,3]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>Evaluating the gene completeness is critical to measuring the quality of a genome assembly. An incomplete assembly can lead to errors in gene predictions, annotation, and other downstream analyses. Benchmarking Universal Single-Copy Orthologs (BUSCO) is a widely used tool for assessing the completeness of genome assembly by testing the presence of a set of single-copy orthologs conserved across a wide range of taxa. However, BUSCO is slow particularly for large genome assemblies. It is cumbersome to apply BUSCO to a large number of assemblies.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>Here, we present compleasm, an efficient tool for assessing the completeness of genome assemblies. Compleasm utilizes the miniprot protein-to-genome aligner and the conserved orthologous genes from BUSCO. It is 14 times faster than BUSCO for human assemblies and reports a more accurate completeness of 99.6% than BUSCO\u2019s 95.7%, which is in close agreement with the annotation completeness of 99.5% for T2T-CHM13.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation<\/jats:title>\n                  <jats:p>https:\/\/github.com\/huangnengCSU\/compleasm.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btad595","type":"journal-article","created":{"date-parts":[[2023,9,28]],"date-time":"2023-09-28T00:03:45Z","timestamp":1695859425000},"source":"Crossref","is-referenced-by-count":367,"title":["compleasm: a faster and more accurate reimplementation of BUSCO"],"prefix":"10.1093","volume":"39","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7187-0749","authenticated-orcid":false,"given":"Neng","family":"Huang","sequence":"first","affiliation":[{"name":"Department of Data Sciences, Dana-Farber Cancer Institute , Boston, MA 02215, United States"},{"name":"Department of Biomedical Informatics, Harvard Medical School , Boston, MA 02115, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Heng","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Data Sciences, Dana-Farber Cancer Institute , Boston, MA 02215, United States"},{"name":"Department of Biomedical Informatics, Harvard Medical School , Boston, MA 02115, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2023,9,27]]},"reference":[{"key":"2023100615595717100_btad595-B1","doi-asserted-by":"crossref","first-page":"1361","DOI":"10.1534\/g3.119.400908","article-title":"Blobtoolkit\u2014interactive quality assessment of genome assemblies","volume":"10","author":"Challis","year":"2020","journal-title":"G3 (Bethesda)"},{"key":"2023100615595717100_btad595-B2","doi-asserted-by":"crossref","first-page":"170","DOI":"10.1038\/s41592-020-01056-5","article-title":"Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm","volume":"18","author":"Cheng","year":"2021","journal-title":"Nat Methods"},{"key":"2023100615595717100_btad595-B3","doi-asserted-by":"crossref","first-page":"1332","DOI":"10.1038\/s41587-022-01261-x","article-title":"Haplotype-resolved assembly of diploid genomes without parental data","volume":"40","author":"Cheng","year":"2022","journal-title":"Nat Biotechnol"},{"key":"2023100615595717100_btad595-B4","doi-asserted-by":"crossref","first-page":"1072","DOI":"10.1093\/bioinformatics\/btt086","article-title":"QUAST: quality assessment tool for genome assemblies","volume":"29","author":"Gurevich","year":"2013","journal-title":"Bioinformatics"},{"key":"2023100615595717100_btad595-B5","doi-asserted-by":"crossref","first-page":"540","DOI":"10.1038\/s41587-019-0072-8","article-title":"Assembly of long, error-prone reads using repeat graphs","volume":"37","author":"Kolmogorov","year":"2019","journal-title":"Nat Biotechnol"},{"key":"2023100615595717100_btad595-B6","doi-asserted-by":"crossref","first-page":"48","DOI":"10.1186\/s40168-020-00808-x","article-title":"Metaeuk-sensitive, high-throughput gene discovery, and annotation for large-scale eukaryotic metagenomics","volume":"8","author":"Levy Karin","year":"2020","journal-title":"Microbiome"},{"key":"2023100615595717100_btad595-B7","doi-asserted-by":"crossref","first-page":"btad014","DOI":"10.1093\/bioinformatics\/btad014","article-title":"Protein-to-genome alignment with miniprot","volume":"39","author":"Li","year":"2023","journal-title":"Bioinformatics"},{"key":"2023100615595717100_btad595-B8","doi-asserted-by":"crossref","first-page":"4647","DOI":"10.1093\/molbev\/msab199","article-title":"Busco update: novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes","volume":"38","author":"Manni","year":"2021","journal-title":"Mol Biol Evol"},{"key":"2023100615595717100_btad595-B9","doi-asserted-by":"crossref","first-page":"e121","DOI":"10.1093\/nar\/gkt263","article-title":"Challenges in homology search: HMMER3 and convergent evolution of coiled\u2013coil regions","volume":"41","author":"Mistry","year":"2013","journal-title":"Nucleic Acids Res"},{"key":"2023100615595717100_btad595-B10","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1126\/science.abj6987","article-title":"The complete sequence of a human genome","volume":"376","author":"Nurk","year":"2022","journal-title":"Science"},{"key":"2023100615595717100_btad595-B11","doi-asserted-by":"crossref","first-page":"3210","DOI":"10.1093\/bioinformatics\/btv351","article-title":"Busco: assessing genome assembly and annotation completeness with single-copy orthologs","volume":"31","author":"Sim\u00e3o","year":"2015","journal-title":"Bioinformatics"},{"key":"2023100615595717100_btad595-B12","doi-asserted-by":"crossref","first-page":"1155","DOI":"10.1038\/s41587-019-0217-9","article-title":"Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome","volume":"37","author":"Wenger","year":"2019","journal-title":"Nat Biotechnol"},{"key":"2023100615595717100_btad595-B13","doi-asserted-by":"crossref","first-page":"D389","DOI":"10.1093\/nar\/gkaa1009","article-title":"OrthoDB in 2020: evolutionary and functional annotations of orthologs","volume":"49","author":"Zdobnov","year":"2021","journal-title":"Nucleic Acids Res"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btad595\/51779686\/btad595.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/39\/10\/btad595\/51913100\/btad595.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/39\/10\/btad595\/51913100\/btad595.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,10,6]],"date-time":"2023-10-06T16:02:10Z","timestamp":1696608130000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/doi\/10.1093\/bioinformatics\/btad595\/7284108"}},"subtitle":[],"editor":[{"given":"Tobias","family":"Marschall","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2023,9,27]]},"references-count":13,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2023,10,3]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btad595","relation":{},"ISSN":["1367-4811"],"issn-type":[{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023,10,1]]},"published":{"date-parts":[[2023,9,27]]},"article-number":"btad595"}}