{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T15:31:42Z","timestamp":1776094302834,"version":"3.50.1"},"reference-count":51,"publisher":"Oxford University Press (OUP)","issue":"10","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2014,5,15]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: Rapid advances in next-generation sequencing (NGS) technology have led to exponential increase in the amount of genomic information. However, NGS reads contain far more errors than data from traditional sequencing methods, and downstream genomic analysis results can be improved by correcting the errors. Unfortunately, all the previous error correction methods required a large amount of memory, making it unsuitable to process reads from large genomes with commodity computers.<\/jats:p><jats:p>Results: We present a novel algorithm that produces accurate correction results with much less memory compared with previous solutions. The algorithm, named BLoom-filter-based Error correction Solution for high-throughput Sequencing reads (BLESS), uses a single minimum-sized Bloom filter, and is also able to tolerate a higher false-positive rate, thus allowing us to correct errors with a 40\u00d7 memory usage reduction on average compared with previous methods. Meanwhile, BLESS can extend reads like DNA assemblers to correct errors at the end of reads. Evaluations using real and simulated reads showed that BLESS could generate more accurate results than existing solutions. After errors were corrected using BLESS, 69% of initially unaligned reads could be aligned correctly. Additionally, de novo assembly results became 50% longer with 66% fewer assembly errors.<\/jats:p><jats:p>Availability and implementation: Freely available at http:\/\/sourceforge.net\/p\/bless-ec<\/jats:p><jats:p>Contact: \u00a0dchen@illinois.edu<\/jats:p><jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btu030","type":"journal-article","created":{"date-parts":[[2014,1,23]],"date-time":"2014-01-23T01:47:51Z","timestamp":1390441671000},"page":"1354-1362","source":"Crossref","is-referenced-by-count":108,"title":["BLESS: Bloom filter-based error correction solution for high-throughput sequencing reads"],"prefix":"10.1093","volume":"30","author":[{"given":"Yun","family":"Heo","sequence":"first","affiliation":[{"name":"1 Department of Electrical and Computer Engineering, 2Department of Bioengineering and 3Institute for Genomic Biology, University of Illinois at Urbana-Champaign, Urbana, IL 61801, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiao-Long","family":"Wu","sequence":"additional","affiliation":[{"name":"1 Department of Electrical and Computer Engineering, 2Department of Bioengineering and 3Institute for Genomic Biology, University of Illinois at Urbana-Champaign, Urbana, IL 61801, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Deming","family":"Chen","sequence":"additional","affiliation":[{"name":"1 Department of Electrical and Computer Engineering, 2Department of Bioengineering and 3Institute for Genomic Biology, University of Illinois at Urbana-Champaign, Urbana, IL 61801, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jian","family":"Ma","sequence":"additional","affiliation":[{"name":"1 Department of Electrical and Computer Engineering, 2Department of Bioengineering and 3Institute for Genomic Biology, University of Illinois at Urbana-Champaign, Urbana, IL 61801, USA"},{"name":"1 Department of Electrical and Computer Engineering, 2Department of Bioengineering and 3Institute for Genomic Biology, University of Illinois at Urbana-Champaign, Urbana, IL 61801, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wen-Mei","family":"Hwu","sequence":"additional","affiliation":[{"name":"1 Department of Electrical and Computer Engineering, 2Department of Bioengineering and 3Institute for Genomic Biology, University of Illinois at Urbana-Champaign, Urbana, IL 61801, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2014,1,21]]},"reference":[{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"413","DOI":"10.1016\/j.coviro.2011.07.008","article-title":"Ultra-deep sequencing for the analysis of viral populations","volume":"1","author":"Beerenwinkel","year":"2011","journal-title":"Curr. Opin. Virol."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"422","DOI":"10.1145\/362686.362692","article-title":"Space\/time trade-offs in hash coding with allowable errors","volume":"13","author":"Bloom","year":"1970","journal-title":"Commun. ACM"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"336","DOI":"10.1101\/gr.079053.108","article-title":"De novo fragment assembly with short mate-paired reads: does the read length matter?","volume":"19","author":"Chaisson","year":"2009","journal-title":"Genome Res."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1093\/bioinformatics\/btt310","article-title":"Informed and automated k-mer size selection for genome assembly","volume":"30","author":"Chikhi","year":"2014","journal-title":"Bioinformatics"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"160","DOI":"10.1186\/1471-2105-14-160","article-title":"Disk-based k-mer counting on a PC","volume":"14","author":"Deorowicz","year":"2013","journal-title":"BMC Bioinformatics"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"e105","DOI":"10.1093\/nar\/gkn425","article-title":"Substantial biases in ultra-short read data sets from high-throughput DNA sequencing","volume":"36","author":"Dohm","year":"2008","journal-title":"Nucleic Acids Res."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"1061","DOI":"10.1038\/nature09534","article-title":"A map of human genome variation from population-scale sequencing","volume":"467","author":"Durbin","year":"2010","journal-title":"Nature"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"2224","DOI":"10.1101\/gr.126599.111","article-title":"Assemblathon 1: a competitive assessment of de novo short read assembly methods","volume":"21","author":"Earl","year":"2011","journal-title":"Genome Res."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"281","DOI":"10.1109\/90.851975","article-title":"Summary cache: a scalable wide-area web cache sharing protocol","volume":"8","author":"Fan","year":"2000","journal-title":"IEEE\/ACM Trans. Netw."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"1599","DOI":"10.1101\/gr.146175.112","article-title":"Decoding the human genome","volume":"22","author":"Frazer","year":"2012","journal-title":"Genome Res."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"659","DOI":"10.1093\/jhered\/esp086","article-title":"Genome 10K: a proposal to obtain whole-genome sequence for 10 000 vertebrate species","volume":"100","author":"Haussler","year":"2009","journal-title":"J. Hered."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"295","DOI":"10.1093\/bioinformatics\/btq653","article-title":"HiTEC: accurate error correction in high-throughput sequencing data","volume":"27","author":"Ilie","year":"2011","journal-title":"Bioinformatics"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"187","DOI":"10.1534\/genetics.107.080630","article-title":"Population genetic inference from resequencing data","volume":"181","author":"Jiang","year":"2009","journal-title":"Genetics"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"1181","DOI":"10.1101\/gr.111351.110","article-title":"ECHO: a reference-free short-read error correction algorithm","volume":"21","author":"Kao","year":"2011","journal-title":"Genome Res."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"R116","DOI":"10.1186\/gb-2010-11-11-r116","article-title":"Quake: quality-aware detection and correction of sequencing errors","volume":"11","author":"Kelley","year":"2010","journal-title":"Genome Biol."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"R25","DOI":"10.1186\/gb-2009-10-3-r25","article-title":"Ultrafast and memory-efficient alignment of short DNA sequences to the human genome","volume":"10","author":"Langmead","year":"2009","journal-title":"Genome Biol."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"e109","DOI":"10.1093\/nar\/gkt215","article-title":"Probabilistic error correction for RNA sequencing","volume":"41","author":"Le","year":"2013","journal-title":"Nucleic Acids Res."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"1754","DOI":"10.1093\/bioinformatics\/btp324","article-title":"Fast and accurate short read alignment with Burrows\u2013Wheeler transform","volume":"25","author":"Li","year":"2009","journal-title":"Bioinformatics"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1101\/gr.097261.109","article-title":"De novo assembly of human genomes with massively parallel short read sequencing","volume":"20","author":"Li","year":"2010","journal-title":"Genome Res."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1186\/1471-2105-12-85","article-title":"DecGPU: distributed error correction on massively parallel graphics processing units using CUDA and MPI","volume":"12","author":"Liu","year":"2011","journal-title":"BMC Bioinformatics"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"308","DOI":"10.1093\/bioinformatics\/bts690","article-title":"Musket: a multistage k-mer spectrum-based error corrector for Illumina sequence data","volume":"29","author":"Liu","year":"2013","journal-title":"Bioinformatics"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"434","DOI":"10.1038\/nbt.2198","article-title":"Performance comparison of benchtop high-throughput sequencing platforms","volume":"30","author":"Loman","year":"2012","journal-title":"Nat. Biotechnol."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"764","DOI":"10.1093\/bioinformatics\/btr011","article-title":"A fast, lock-free approach for efficient parallel counting of occurrences of k-mers","volume":"27","author":"Mar\u00e7ais","year":"2011","journal-title":"Bioinformatics"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"i137","DOI":"10.1093\/bioinformatics\/btr208","article-title":"Error correction of high-throughput sequencing datasets with non-uniform coverage","volume":"27","author":"Medvedev","year":"2011","journal-title":"Bioinformatics"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1186\/1471-2105-12-333","article-title":"Efficient counting of k-mers in DNA sequences using a bloom filter","volume":"12","author":"Melsted","year":"2011","journal-title":"BMC Bioinformatics"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1038\/nrg2626","article-title":"Sequencing technologies\u2014the next generation","volume":"11","author":"Metzker","year":"2009","journal-title":"Nat. Rev. Genet."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"9748","DOI":"10.1073\/pnas.171285098","article-title":"An Eulerian path approach to DNA fragment assembly","volume":"98","author":"Pevzner","year":"2001","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","DOI":"10.1038\/srep02837","article-title":"Empirical validation of viral quasispecies assembly algorithms: state-of-the-art and challenges","volume":"3","author":"Prosperi","year":"2013","journal-title":"Sci. Rep."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"1309","DOI":"10.1101\/gr.089151.108","article-title":"Efficient frequency-based de novo short-read clustering for error trimming in next-generation sequencing","volume":"19","author":"Qu","year":"2009","journal-title":"Genome Res."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"652","DOI":"10.1093\/bioinformatics\/btt020","article-title":"DSK: k-mer counting with very low memory usage","volume":"29","author":"Rizk","year":"2013","journal-title":"Bioinformatics"},{"key":"2023041302465875100_","article-title":"Turtle: identifying frequent k-mers with cache-efficient algorithms","author":"Roy","year":"2013"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"1284","DOI":"10.1093\/bioinformatics\/btq151","article-title":"Correction of sequencing errors in a mixed set of reads","volume":"26","author":"Salmela","year":"2010","journal-title":"Bioinformatics"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"1455","DOI":"10.1093\/bioinformatics\/btr170","article-title":"Correcting errors in short reads by multiple alignments","volume":"27","author":"Salmela","year":"2011","journal-title":"Bioinformatics"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"557","DOI":"10.1101\/gr.131383.111","article-title":"GAGE: a critical evaluation of genome assemblies and assembly algorithms","volume":"22","author":"Salzberg","year":"2012","journal-title":"Genome Res."},{"key":"2023041302465875100_","article-title":"Benchmarking of viral haplotype reconstruction programmes: an overview of the capacities and limitations of currently available programmes","author":"Schirmer","year":"2012","journal-title":"Brief. Bioinform"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"2157","DOI":"10.1093\/bioinformatics\/btp379","article-title":"SHREC: a short-read error correction method","volume":"25","author":"Schr\u00f6der","year":"2009","journal-title":"Bioinformatics"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"60","DOI":"10.1109\/IPDPS.2012.16","article-title":"A parallel algorithm for spectrum-based short read error correction","volume-title":"Parallel & Distributed Processing Symposium (IPDPS), 2012 IEEE 26th International","author":"Shah","year":"2012"},{"key":"2023041302465875100_","first-page":"1","article-title":"Accelerating error correction in high-throughput short-read DNA sequencing data with CUDA","volume-title":"Parallel & Distributed Processing, 2009. IPDPS 2009. IEEE International Symposium on","author":"Shi","year":"2009"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"603","DOI":"10.1089\/cmb.2009.0062","article-title":"A parallel algorithm for error correction in high-throughput short-read data on CUDA-enabled graphics hardware","volume":"17","author":"Shi","year":"2010","journal-title":"J. Comput. Biol."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"1129","DOI":"10.1016\/j.procs.2010.04.125","article-title":"Quality-score guided error correction for short-read sequencing data using CUDA","volume":"1","author":"Shi","year":"2010","journal-title":"Procedia Comput. Sci."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"549","DOI":"10.1101\/gr.126953.111","article-title":"Efficient de novo assembly of large genomes using compressed data structures","volume":"22","author":"Simpson","year":"2012","journal-title":"Genome Res."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"185","DOI":"10.1186\/1471-2105-13-185","article-title":"Estimation of sequencing error rates in short reads","volume":"13","author":"Wang","year":"2012","journal-title":"BMC Bioinformatics"},{"key":"2023041302465875100_","first-page":"189","article-title":"Recount: expectation maximization based error correction tool for next generation sequencing data","volume":"23","author":"Wijaya","year":"2009","journal-title":"Genome Inform."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"2101","DOI":"10.1109\/IPDPS.2011.387","article-title":"Error correction and clustering algorithms for next generation sequencing","volume-title":"Parallel and Distributed Processing Workshops and Phd Forum (IPDPSW), 2011 IEEE International Symposium on","author":"Yang","year":"2011"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/1471-2105-12-S1-S52","article-title":"Repeat-aware modeling and correction of short read errors","volume":"12","author":"Yang","year":"2011","journal-title":"BMC Bioinformatics"},{"key":"2023041302465875100_","article-title":"A survey of error-correction methods for next-generation sequencing","author":"Yang","year":"2012","journal-title":"Brief. Bioinform"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"2526","DOI":"10.1093\/bioinformatics\/btq468","article-title":"Reptile: representative tiling for short read error correction","volume":"26","author":"Yang","year":"2010","journal-title":"Bioinformatics"},{"key":"2023041302465875100_","first-page":"1302.0212","article-title":"PREMIER - PRobabilistic Error-correction using Markov Inference in Errored Reads","volume":"2013","author":"Yin","year":"2013","journal-title":"arXiv"},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"821","DOI":"10.1101\/gr.074492.107","article-title":"Velvet: algorithms for de novo short read assembly using de Bruijn graphs","volume":"18","author":"Zerbino","year":"2008","journal-title":"Genome Res."},{"key":"2023041302465875100_","doi-asserted-by":"crossref","first-page":"198","DOI":"10.1007\/978-3-642-22589-5_19","article-title":"An efficient hybrid approach to correcting errors in short reads","volume-title":"Modeling Decision for Artificial Intelligence","author":"Zhao","year":"2011"},{"key":"2023041302465875100_","first-page":"220","article-title":"PSAEC: An Improved Algorithm for Short Read Error Correction Using Partial Suffix Arrays","volume-title":"Proceedings of the 5th Joint International Frontiers in Algorithmics, and 7th International Conference on Algorithmic Aspects in Information and Management","author":"Zhao","year":"2011"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/10\/1354\/49860227\/bioinformatics_30_10_1354.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/10\/1354\/49860227\/bioinformatics_30_10_1354.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,7,10]],"date-time":"2023-07-10T00:01:21Z","timestamp":1688947281000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/30\/10\/1354\/266571"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,1,21]]},"references-count":51,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2014,5,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btu030","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2014,5,15]]},"published":{"date-parts":[[2014,1,21]]}}}