{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T06:06:51Z","timestamp":1780466811694,"version":"3.54.1"},"reference-count":31,"publisher":"Oxford University Press (OUP)","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2013,2,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: The imperfect sequence data produced by next-generation sequencing technologies have motivated the development of a number of short-read error correctors in recent years. The majority of methods focus on the correction of substitution errors, which are the dominant error source in data produced by Illumina sequencing technology. Existing tools either score high in terms of recall or precision but not consistently high in terms of both measures.<\/jats:p>\n               <jats:p>Results: In this article, we present Musket, an efficient multistage k-mer-based corrector for Illumina short-read data. We use the k-mer spectrum approach and introduce three correction techniques in a multistage workflow: two-sided conservative correction, one-sided aggressive correction and voting-based refinement. Our performance evaluation results, in terms of correction quality and de novo genome assembly measures, reveal that Musket is consistently one of the top performing correctors. In addition, Musket is multi-threaded using a master\u2013slave model and demonstrates superior parallel scalability compared with all other evaluated correctors as well as a highly competitive overall execution time.<\/jats:p>\n               <jats:p>Availability: Musket is available at http:\/\/musket.sourceforge.net.<\/jats:p>\n               <jats:p>Contact: \u00a0liuy@uni-mainz.de or bertil.schmidt@uni-mainz.de<\/jats:p>\n               <jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/bts690","type":"journal-article","created":{"date-parts":[[2012,12,1]],"date-time":"2012-12-01T01:16:58Z","timestamp":1354324618000},"page":"308-315","source":"Crossref","is-referenced-by-count":260,"title":["Musket: a multistage <i>k-<\/i>mer spectrum-based error corrector for Illumina sequence data"],"prefix":"10.1093","volume":"29","author":[{"given":"Yongchao","family":"Liu","sequence":"first","affiliation":[{"name":"1 Institut f\u00fcr Informatik, Johannes Gutenberg Universit\u00e4t Mainz, Mainz 55099, Germany and 2Department of Computing and Information Systems, The University of Melbourne, Parkville 3010, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jan","family":"Schr\u00f6der","sequence":"additional","affiliation":[{"name":"1 Institut f\u00fcr Informatik, Johannes Gutenberg Universit\u00e4t Mainz, Mainz 55099, Germany and 2Department of Computing and Information Systems, The University of Melbourne, Parkville 3010, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bertil","family":"Schmidt","sequence":"additional","affiliation":[{"name":"1 Institut f\u00fcr Informatik, Johannes Gutenberg Universit\u00e4t Mainz, Mainz 55099, Germany and 2Department of Computing and Information Systems, The University of Melbourne, Parkville 3010, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2012,11,29]]},"reference":[{"key":"2023012810214928400_bts690-B1","doi-asserted-by":"crossref","first-page":"422","DOI":"10.1145\/362686.362692","article-title":"Space\/time trade-offs in hash coding with allowable errors","volume":"13","author":"Bloom","year":"1970","journal-title":"Commu. ACM"},{"key":"2023012810214928400_bts690-B2","article-title":"A block-sorting lossless data compression algorithm","author":"Burrows","year":"1994","journal-title":"Technical Report 124 Palo Alto, CA."},{"key":"2023012810214928400_bts690-B3","doi-asserted-by":"crossref","first-page":"810","DOI":"10.1101\/gr.7337908","article-title":"ALLPATHS: de novo assembly of whole-genome shotgun microreads","volume":"18","author":"Butler","year":"2008","journal-title":"Genome Res."},{"key":"2023012810214928400_bts690-B4","doi-asserted-by":"crossref","first-page":"2067","DOI":"10.1093\/bioinformatics\/bth205","article-title":"Fragment assembly with short reads","volume":"20","author":"Chaisson","year":"2004","journal-title":"Bioinformatics"},{"key":"2023012810214928400_bts690-B5","doi-asserted-by":"crossref","first-page":"336","DOI":"10.1101\/gr.079053.108","article-title":"De novo fragment assembly with short mate-paired reads: does the read length matter?","volume":"19","author":"Chaisson","year":"2009","journal-title":"Genome Res."},{"key":"2023012810214928400_bts690-B6","doi-asserted-by":"crossref","first-page":"e105","DOI":"10.1093\/nar\/gkn425","article-title":"Substantial biases in ultra-short read data sets from high-throughput DNA sequencing","volume":"36","author":"Dohm","year":"2008","journal-title":"Nucleic Acids Res."},{"key":"2023012810214928400_bts690-B7","doi-asserted-by":"crossref","first-page":"186","DOI":"10.1101\/gr.8.3.186","article-title":"Base-calling of automated sequencer traces using phred. II. Error probabilities","volume":"8","author":"Ewing","year":"1998","journal-title":"Genome Res."},{"key":"2023012810214928400_bts690-B8","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1145\/1082036.1082039","article-title":"Indexing compressed text","volume":"52","author":"Ferragina","year":"2005","journal-title":"J. ACM"},{"key":"2023012810214928400_bts690-B9","doi-asserted-by":"crossref","first-page":"1513","DOI":"10.1073\/pnas.1017351108","article-title":"High-quality draft assemblies of mammalian genomes from massively parallel sequence data","volume":"108","author":"Gnerre","year":"2010","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012810214928400_bts690-B10","doi-asserted-by":"crossref","first-page":"295","DOI":"10.1093\/bioinformatics\/btq653","article-title":"HiTEC: accurate error correction in high-throughput sequencing data","volume":"27","author":"Ilie","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012810214928400_bts690-B11","doi-asserted-by":"crossref","first-page":"1181","DOI":"10.1101\/gr.111351.110","article-title":"ECHO: a reference-free short-read error correction algorithm","volume":"21","author":"Kao","year":"2011","journal-title":"Genome Res."},{"key":"2023012810214928400_bts690-B12","doi-asserted-by":"crossref","first-page":"R116","DOI":"10.1186\/gb-2010-11-11-r116","article-title":"Quake: quality-aware detection and correction of sequencing errors","volume":"11","author":"Kelley","year":"2010","journal-title":"Genome Biol."},{"key":"2023012810214928400_bts690-B13","doi-asserted-by":"crossref","first-page":"1838","DOI":"10.1093\/bioinformatics\/bts280","article-title":"Exploring single-sample SNP and INDEL calling with whole-genome de novo assembly","volume":"28","author":"Li","year":"2012","journal-title":"Bioinformatics"},{"key":"2023012810214928400_bts690-B14","doi-asserted-by":"crossref","first-page":"311","DOI":"10.1038\/nature08696","article-title":"The sequence and de novo assembly of the giant panda genome","volume":"463","author":"Li","year":"2010","journal-title":"Nature"},{"key":"2023012810214928400_bts690-B15","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1101\/gr.097261.109","article-title":"De novo assembly of human genomes with massively parallel short read sequencing","volume":"20","author":"Li","year":"2010","journal-title":"Genome Res."},{"key":"2023012810214928400_bts690-B16","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1186\/1471-2105-12-85","article-title":"DecGPU: distributed error correction on massively parallel graphics processing units using CUDA and MPI","volume":"12","author":"Liu","year":"2011","journal-title":"BMC Bioinformatics"},{"key":"2023012810214928400_bts690-B17","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1186\/1471-2105-12-354","article-title":"Parallelized short read assembly of large genomes using de Bruijn graphs","volume":"12","author":"Liu","year":"2011","journal-title":"BMC Bioinformatics"},{"key":"2023012810214928400_bts690-B18","doi-asserted-by":"crossref","first-page":"1830","DOI":"10.1093\/bioinformatics\/bts276","article-title":"CUSHAW: a CUDA compatible short read aligner to large genomes based on the Burrows-Wheeler transform","volume":"28","author":"Liu","year":"2012","journal-title":"Bioinformatics"},{"key":"2023012810214928400_bts690-B19","doi-asserted-by":"crossref","first-page":"i137","DOI":"10.1093\/bioinformatics\/btr208","article-title":"Error correction of high-throughput sequencing datasets with non-uniform coverage","volume":"27","author":"Medvedev","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012810214928400_bts690-B20","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1186\/1471-2105-12-333","article-title":"Efficient counting of k-mers in DNA sequences using a bloom filter","volume":"12","author":"Melsted","year":"2011","journal-title":"BMC Bioinformatics"},{"key":"2023012810214928400_bts690-B21","doi-asserted-by":"crossref","first-page":"9748","DOI":"10.1073\/pnas.171285098","article-title":"An Eulerian path approach to DNA fragment assembly","volume":"98","author":"Pevzner","year":"2001","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012810214928400_bts690-B22","doi-asserted-by":"crossref","first-page":"1284","DOI":"10.1093\/bioinformatics\/btq151","article-title":"Correction of sequencing errors in a mixed set of reads","volume":"26","author":"Salmela","year":"2010","journal-title":"Bioinformatics"},{"key":"2023012810214928400_bts690-B23","doi-asserted-by":"crossref","first-page":"1455","DOI":"10.1093\/bioinformatics\/btr170","article-title":"Correcting errors in short reads by multiple alignments","volume":"27","author":"Salmela","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012810214928400_bts690-B24","doi-asserted-by":"crossref","first-page":"557","DOI":"10.1101\/gr.131383.111","article-title":"GAGE: a critical evaluation of genome assemblies and assembly algorithms","volume":"22","author":"Salzberg","year":"2012","journal-title":"Genome Res."},{"key":"2023012810214928400_bts690-B25","doi-asserted-by":"crossref","first-page":"2157","DOI":"10.1093\/bioinformatics\/btp379","article-title":"SHREC: a short-read error correction method","volume":"25","author":"Schr\u00f6der","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012810214928400_bts690-B26","doi-asserted-by":"crossref","first-page":"603","DOI":"10.1089\/cmb.2009.0062","article-title":"A parallel algorithm for error correction in high-throughput short-read data on CUDA-enabled graphics hardware","volume":"17","author":"Shi","year":"2010","journal-title":"J. Comput. Biol."},{"key":"2023012810214928400_bts690-B27","doi-asserted-by":"crossref","first-page":"549","DOI":"10.1101\/gr.126953.111","article-title":"Efficient de novo assembly of large genomes using compressed data structures","volume":"22","author":"Simpson","year":"2012","journal-title":"Genome Res."},{"key":"2023012810214928400_bts690-B28","doi-asserted-by":"crossref","first-page":"1117","DOI":"10.1101\/gr.089532.108","article-title":"ABySS: a parallel assembler for short read sequence data","volume":"19","author":"Simpson","year":"2009","journal-title":"Genome Res."},{"key":"2023012810214928400_bts690-B29","doi-asserted-by":"crossref","first-page":"2526","DOI":"10.1093\/bioinformatics\/btq468","article-title":"Reptile: representative tiling for short read error correction","volume":"26","author":"Yang","year":"2010","journal-title":"Bioinformatics"},{"key":"2023012810214928400_bts690-B30","article-title":"A survey of error-correction methods for next-generation sequencing","author":"Yang","year":"2012","journal-title":"Brief. Bioinform."},{"key":"2023012810214928400_bts690-B31","doi-asserted-by":"crossref","first-page":"821","DOI":"10.1101\/gr.074492.107","article-title":"Velvet: algorithms for de novo short read assembly using de Bruijn graphs","volume":"18","author":"Zerbino","year":"2008","journal-title":"Genome Res."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/29\/3\/308\/48893002\/bioinformatics_29_3_308.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/29\/3\/308\/48893002\/bioinformatics_29_3_308.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,28]],"date-time":"2023-01-28T11:44:59Z","timestamp":1674906299000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/29\/3\/308\/257257"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,11,29]]},"references-count":31,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2013,2,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bts690","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2013,2,1]]},"published":{"date-parts":[[2012,11,29]]}}}