{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T05:05:08Z","timestamp":1780981508751,"version":"3.54.1"},"reference-count":27,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2022,6,13]],"date-time":"2022-06-13T00:00:00Z","timestamp":1655078400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,6,13]],"date-time":"2022-06-13T00:00:00Z","timestamp":1655078400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"DeCoDeML Project by Rhein-Main-University Network"},{"name":"Johannes Gutenberg-Universit\u00e4t Mainz"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2022,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec><jats:title>Background<\/jats:title><jats:p>Next-generation sequencing pipelines often perform error correction as a preprocessing step to obtain cleaned input data. State-of-the-art error correction programs are able to reliably detect and correct the majority of sequencing errors. However, they also introduce new errors by making false-positive corrections. These correction mistakes can have negative impact on downstream analysis, such as<jats:italic>k<\/jats:italic>-mer statistics, de-novo assembly, and variant calling. This motivates the need for more precise error correction tools.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>We present CARE 2.0, a context-aware read error correction tool based on multiple sequence alignment targeting Illumina datasets. In addition to a number of newly introduced optimizations its most significant change is the replacement of CARE 1.0\u2019s hand-crafted correction conditions with a novel classifier based on random decision forests trained on Illumina data. This results in up to two orders-of-magnitude fewer false-positive corrections compared to other state-of-the-art error correction software. At the same time, CARE 2.0 is able to achieve high numbers of true-positive corrections comparable to its competitors. On a simulated full human dataset with 914M reads CARE 2.0 generates only 1.2M false positives (FPs) (and 801.4M true positives (TPs)) at a highly competitive runtime while the best corrections achieved by other state-of-the-art tools contain at least 3.9M FPs and at most 814.5M TPs. Better de-novo assembly and improved<jats:italic>k<\/jats:italic>-mer analysis show the applicability of CARE 2.0 to real-world data.<\/jats:p><\/jats:sec><jats:sec><jats:title>Conclusion<\/jats:title><jats:p>False-positive corrections can negatively influence down-stream analysis. The precision of CARE 2.0 greatly reduces the number of those corrections compared to other state-of-the-art programs including BFC, Karect, Musket, Bcool, SGA, and Lighter. Thus, higher-quality datasets are produced which improve<jats:italic>k<\/jats:italic>-mer analysis and de-novo assembly in real-world datasets which demonstrates the applicability of machine learning techniques in the context of sequencing read error correction. CARE 2.0 is written in C++\/CUDA for Linux systems and can be run on the CPU as well as on CUDA-enabled GPUs. It is available at<jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/fkallen\/CARE\">https:\/\/github.com\/fkallen\/CARE<\/jats:ext-link>.<\/jats:p><\/jats:sec>","DOI":"10.1186\/s12859-022-04754-3","type":"journal-article","created":{"date-parts":[[2022,6,13]],"date-time":"2022-06-13T15:04:05Z","timestamp":1655132645000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":14,"title":["CARE 2.0: reducing false-positive sequencing error corrections using machine learning"],"prefix":"10.1186","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4516-6357","authenticated-orcid":false,"given":"Felix","family":"Kallenborn","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Julian","family":"Cascitti","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2597-8331","authenticated-orcid":false,"given":"Bertil","family":"Schmidt","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,6,13]]},"reference":[{"issue":"1","key":"4754_CR1","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s12859-017-1784-8","volume":"18","author":"M Heydari","year":"2017","unstructured":"Heydari M, Miclotte G, Demeester P, et al. Evaluation of the impact of Illumina error correction tools on de novo genome assembly. BMC Bioinform. 2017;18(1):1\u201313.","journal-title":"BMC Bioinform"},{"issue":"1","key":"4754_CR2","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/s41598-019-51418-z","volume":"9","author":"I Fischer-Hwang","year":"2019","unstructured":"Fischer-Hwang I, Ochoa I, Weissman T, et al. Denoising of aligned genomic data. Sci Rep. 2019;9(1):1\u201311.","journal-title":"Sci Rep"},{"issue":"3","key":"4754_CR3","doi-asserted-by":"publisher","first-page":"549","DOI":"10.1101\/gr.126953.111","volume":"22","author":"JT Simpson","year":"2012","unstructured":"Simpson JT, Durbin R. Efficient de novo assembly of large genomes using compressed data structures. Genome Res. 2012;22(3):549\u201356.","journal-title":"Genome Res"},{"key":"4754_CR4","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/bts690","author":"Y Liu","year":"2013","unstructured":"Liu Y, Schr\u00f6der J, Schmidt B. Musket: a multistage k-mer spectrum-based error corrector for Illumina sequence data. Bioinformatics. 2013. https:\/\/doi.org\/10.1093\/bioinformatics\/bts690.","journal-title":"Bioinformatics"},{"key":"4754_CR5","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btt407","author":"L Ilie","year":"2013","unstructured":"Ilie L, Molnar M. Racer: rapid and accurate correction of errors in reads. Bioinformatics. 2013. https:\/\/doi.org\/10.1093\/bioinformatics\/btt407.","journal-title":"Bioinformatics"},{"key":"4754_CR6","doi-asserted-by":"publisher","DOI":"10.1186\/s13059-014-0509-9","author":"L Song","year":"2014","unstructured":"Song L, Florea L, Langmead B. Lighter: fast and memory-efficient sequencing error correction without counting. Genome Biol. 2014. https:\/\/doi.org\/10.1186\/s13059-014-0509-9.","journal-title":"Genome Biol"},{"issue":"19","key":"4754_CR7","doi-asserted-by":"publisher","first-page":"2723","DOI":"10.1093\/bioinformatics\/btu368","volume":"30","author":"P Greenfield","year":"2014","unstructured":"Greenfield P, Duesing K, Papanicolaou A, et al. Blue: correcting sequencing errors using consensus and context. Bioinformatics. 2014;30(19):2723\u201332.","journal-title":"Bioinformatics"},{"key":"4754_CR8","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btv290","author":"H Li","year":"2015","unstructured":"Li H. BFC: correcting Illumina sequencing errors. Bioinformatics. 2015. https:\/\/doi.org\/10.1093\/bioinformatics\/btv290.","journal-title":"Bioinformatics"},{"issue":"15","key":"4754_CR9","doi-asserted-by":"publisher","first-page":"2369","DOI":"10.1093\/bioinformatics\/btw146","volume":"32","author":"Y Heo","year":"2016","unstructured":"Heo Y, Ramachandran A, Hwu W-M, et al. BLESS 2: accurate, memory-efficient and fast error correction method. Bioinformatics. 2016;32(15):2369\u201371.","journal-title":"Bioinformatics"},{"issue":"7","key":"4754_CR10","doi-asserted-by":"crossref","first-page":"1086","DOI":"10.1093\/bioinformatics\/btw746","volume":"33","author":"M D\u0142ugosz","year":"2017","unstructured":"D\u0142ugosz M, Deorowicz S. RECKONER: read error corrector based on KMC. Bioinformatics. 2017;33(7):1086\u20139.","journal-title":"Bioinformatics"},{"issue":"11","key":"4754_CR11","doi-asserted-by":"publisher","first-page":"1455","DOI":"10.1093\/bioinformatics\/btr170","volume":"27","author":"L Salmela","year":"2011","unstructured":"Salmela L, Schr\u00f6der J. Correcting errors in short reads by multiple alignments. Bioinformatics. 2011;27(11):1455\u201361.","journal-title":"Bioinformatics"},{"issue":"7","key":"4754_CR12","doi-asserted-by":"publisher","first-page":"1181","DOI":"10.1101\/gr.111351.110","volume":"21","author":"W-C Kao","year":"2011","unstructured":"Kao W-C, Chan AH, Song YS. Echo: a reference-free short-read error correction algorithm. Genome Res. 2011;21(7):1181\u201392.","journal-title":"Genome Res"},{"issue":"17","key":"4754_CR13","doi-asserted-by":"publisher","first-page":"i356","DOI":"10.1093\/bioinformatics\/btu440","volume":"30","author":"MH Schulz","year":"2014","unstructured":"Schulz MH, Weese D, Holtgrewe M, et al. Fiona: a parallel and automatic strategy for read error correction. Bioinformatics. 2014;30(17):i356\u201363.","journal-title":"Bioinformatics"},{"issue":"21","key":"4754_CR14","doi-asserted-by":"publisher","first-page":"3421","DOI":"10.1093\/bioinformatics\/btv415","volume":"31","author":"A Allam","year":"2015","unstructured":"Allam A, Kalnis P, Solovyev V. Karect: accurate correction of substitution, insertion and deletion errors for next-generation sequencing data. Bioinformatics. 2015;31(21):3421\u20138.","journal-title":"Bioinformatics"},{"key":"4754_CR15","doi-asserted-by":"publisher","first-page":"1374","DOI":"10.1093\/bioinformatics\/btz102","volume":"36","author":"A Limasset","year":"2019","unstructured":"Limasset A, Flot J, Peterlongo P. Toward perfect reads: self-correction of short reads via mapping on de Bruijn graphs. Bioinformatics. 2019;36:1374\u201381.","journal-title":"Bioinformatics"},{"issue":"1","key":"4754_CR16","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s12859-019-2906-2","volume":"20","author":"M Heydari","year":"2019","unstructured":"Heydari M, Miclotte G, Van de Peer Y, et al. Illumina error correction near highly repetitive DNA regions improves de novo genome assembly. BMC Bioinform. 2019;20(1):1\u201313.","journal-title":"BMC Bioinform"},{"issue":"7","key":"4754_CR17","doi-asserted-by":"publisher","first-page":"889","DOI":"10.1093\/bioinformatics\/btaa738","volume":"37","author":"F Kallenborn","year":"2020","unstructured":"Kallenborn F, Hildebrandt A, Schmidt B. CARE: context-aware sequencing read error correction. Bioinformatics. 2020;37(7):889\u201395. https:\/\/doi.org\/10.1093\/bioinformatics\/btaa738.","journal-title":"Bioinformatics"},{"key":"4754_CR18","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-019-52196-4","author":"M Abdallah","year":"2019","unstructured":"Abdallah M, Mahgoub A, Ahmed H, Chaterji S. Athena: automated tuning of k-mer based genomic error correction algorithms using language models. Sci Rep. 2019. https:\/\/doi.org\/10.1038\/s41598-019-52196-4.","journal-title":"Sci Rep"},{"issue":"1","key":"4754_CR19","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1186\/s12859-021-04547-0","volume":"23","author":"A Sharma","year":"2022","unstructured":"Sharma A, Jain P, Mahgoub A, Zhou Z, Mahadik K, Chaterji S. Lerna: transformer architectures for configuring error correction tools for short- and long-read genome sequencing. BMC Bioinform. 2022;23(1):25. https:\/\/doi.org\/10.1186\/s12859-021-04547-0.","journal-title":"BMC Bioinform"},{"issue":"10","key":"4754_CR20","doi-asserted-by":"publisher","first-page":"1553","DOI":"10.1093\/bioinformatics\/btu856","volume":"31","author":"H Xin","year":"2015","unstructured":"Xin H, Greth J, Emmons J, et al. Shifted hamming distance: a fast and accurate SIMD-friendly filter to accelerate alignment verification in read mapping. Bioinformatics. 2015;31(10):1553\u201360.","journal-title":"Bioinformatics"},{"issue":"4","key":"4754_CR21","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1145\/270563.571472","volume":"28","author":"D Gusfield","year":"1997","unstructured":"Gusfield D. Algorithms on stings, trees, and sequences: computer science and computational biology. Acm Sigact News. 1997;28(4):41\u201360.","journal-title":"Acm Sigact News"},{"key":"4754_CR22","doi-asserted-by":"publisher","first-page":"63","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman L. Random forests. Mach Learn. 2001;45:63\u201379. https:\/\/doi.org\/10.1023\/A:1010933404324.","journal-title":"Mach Learn"},{"issue":"4","key":"4754_CR23","doi-asserted-by":"publisher","first-page":"593","DOI":"10.1093\/bioinformatics\/btr708","volume":"28","author":"W Huang","year":"2012","unstructured":"Huang W, Li L, Myers JR, et al. Art: a next-generation sequencing read simulator. Bioinformatics. 2012;28(4):593\u20134.","journal-title":"Bioinformatics"},{"issue":"5","key":"4754_CR24","doi-asserted-by":"publisher","first-page":"455","DOI":"10.1089\/cmb.2012.0021","volume":"19","author":"A Bankevich","year":"2012","unstructured":"Bankevich A, Nurk S, Antipov D, et al. Spades: a new genome assembly algorithm and its applications to single-cell sequencing. J Comput Biol. 2012;19(5):455\u201377.","journal-title":"J Comput Biol"},{"issue":"8","key":"4754_CR25","doi-asserted-by":"publisher","first-page":"1072","DOI":"10.1093\/bioinformatics\/btt086","volume":"29","author":"A Gurevich","year":"2013","unstructured":"Gurevich A, Saveliev V, Vyahhi N, et al. Quast: quality assessment tool for genome assemblies. Bioinformatics. 2013;29(8):1072\u20135.","journal-title":"Bioinformatics"},{"issue":"6","key":"4754_CR26","doi-asserted-by":"publisher","first-page":"764","DOI":"10.1093\/bioinformatics\/btr011","volume":"27","author":"G Marcais","year":"2011","unstructured":"Marcais G, Kingsford C. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics. 2011;27(6):764\u201370.","journal-title":"Bioinformatics"},{"key":"4754_CR27","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, Vanderplas J, Passos A, Cournapeau D, Brucher M, Perrot M, Duchesnay E. Scikit-learn: machine learning in Python. J Mach Learn Res. 2011;12:2825\u201330.","journal-title":"J Mach Learn Res"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-022-04754-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12859-022-04754-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-022-04754-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,8]],"date-time":"2023-02-08T01:41:07Z","timestamp":1675820467000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-022-04754-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,13]]},"references-count":27,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,12]]}},"alternative-id":["4754"],"URL":"https:\/\/doi.org\/10.1186\/s12859-022-04754-3","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,13]]},"assertion":[{"value":"3 February 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 May 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 June 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}],"article-number":"227"}}