{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,26]],"date-time":"2026-08-26T02:40:54Z","timestamp":1787712054728,"version":"build-2784847793"},"reference-count":26,"publisher":"Oxford University Press (OUP)","issue":"21","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2014,11,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation : Today, the base code of DNA is mostly determined through sequencing by synthesis as provided by the Illumina sequencers. Although highly accurate, resulting reads are short, making their analyses challenging. Recently, a new technology, single molecule real-time (SMRT) sequencing, was developed that could address these challenges, as it generates reads of several thousand bases. But, their broad application has been hampered by a high error rate. Therefore, hybrid approaches that use high-quality short reads to correct erroneous SMRT long reads have been developed. Still, current implementations have great demands on hardware, work only in well-defined computing infrastructures and reject a substantial amount of reads. This limits their usability considerably, especially in the case of large sequencing projects.<\/jats:p>\n               <jats:p>Results : Here we present proovread , a hybrid correction pipeline for SMRT reads, which can be flexibly adapted on existing hardware and infrastructure from a laptop to a high-performance computing cluster. On genomic and transcriptomic test cases covering Escherichia coli , Arabidopsis thaliana and human, proovread achieved accuracies up to 99.9% and outperformed the existing hybrid correction programs. Furthermore, proovread -corrected sequences were longer and the throughput was higher. Thus, proovread combines the most accurate correction results with an excellent adaptability to the available hardware. It will therefore increase the applicability and value of SMRT sequencing.<\/jats:p>\n               <jats:p>Availability and implementation: \u00a0proovread is available at the following URL: http:\/\/proovread.bioapps.biozentrum.uni-wuerzburg.de<\/jats:p>\n               <jats:p>Contact : frank.foerster@biozentrum.uni-wuerzburg.de<\/jats:p>\n               <jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btu392","type":"journal-article","created":{"date-parts":[[2014,7,12]],"date-time":"2014-07-12T03:16:15Z","timestamp":1405134975000},"page":"3004-3011","source":"Crossref","is-referenced-by-count":451,"title":["<i>proovread<\/i>\n            : large-scale high-accuracy PacBio correction through iterative short read consensus"],"prefix":"10.1093","volume":"30","author":[{"given":"Thomas","family":"Hackl","sequence":"first","affiliation":[{"name":"1 Department for Molecular Plant Physiology and Biophysics, University of W\u00fcrzburg, Julius-von-Sachs-Platz 2, 97082 W\u00fcrzburg, Germany and 2 Department of Bioinformatics, University of W\u00fcrzburg, Biocenter, Am Hubland, 97074 W\u00fcrzburg, Germany"},{"name":"1 Department for Molecular Plant Physiology and Biophysics, University of W\u00fcrzburg, Julius-von-Sachs-Platz 2, 97082 W\u00fcrzburg, Germany and 2 Department of Bioinformatics, University of W\u00fcrzburg, Biocenter, Am Hubland, 97074 W\u00fcrzburg, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rainer","family":"Hedrich","sequence":"additional","affiliation":[{"name":"1 Department for Molecular Plant Physiology and Biophysics, University of W\u00fcrzburg, Julius-von-Sachs-Platz 2, 97082 W\u00fcrzburg, Germany and 2 Department of Bioinformatics, University of W\u00fcrzburg, Biocenter, Am Hubland, 97074 W\u00fcrzburg, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"J\u00f6rg","family":"Schultz","sequence":"additional","affiliation":[{"name":"1 Department for Molecular Plant Physiology and Biophysics, University of W\u00fcrzburg, Julius-von-Sachs-Platz 2, 97082 W\u00fcrzburg, Germany and 2 Department of Bioinformatics, University of W\u00fcrzburg, Biocenter, Am Hubland, 97074 W\u00fcrzburg, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Frank","family":"F\u00f6rster","sequence":"additional","affiliation":[{"name":"1 Department for Molecular Plant Physiology and Biophysics, University of W\u00fcrzburg, Julius-von-Sachs-Platz 2, 97082 W\u00fcrzburg, Germany and 2 Department of Bioinformatics, University of W\u00fcrzburg, Biocenter, Am Hubland, 97074 W\u00fcrzburg, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2014,7,10]]},"reference":[{"key":"2023012711573655900_btu392-B1","doi-asserted-by":"crossref","first-page":"R18","DOI":"10.1186\/gb-2011-12-2-r18","article-title":"Analyzing and minimizing pcr amplification bias in illumina sequencing libraries","volume":"12","author":"Aird","year":"2011","journal-title":"Genome Biol."},{"key":"2023012711573655900_btu392-B2","doi-asserted-by":"crossref","first-page":"e46679","DOI":"10.1371\/journal.pone.0046679","article-title":"Improving pacbio long read accuracy by short read alignment","volume":"7","author":"Au","year":"2012","journal-title":"PLoS One"},{"key":"2023012711573655900_btu392-B3","article-title":"A reference-free algorithm for computational normalization of shotgun sequencing data","author":"Brown","year":"2012"},{"key":"2023012711573655900_btu392-B4","doi-asserted-by":"crossref","first-page":"375","DOI":"10.1186\/1471-2164-13-375","article-title":"Pacific biosciences sequencing technology for genotyping and variation discovery in human data","volume":"13","author":"Carneiro","year":"2012","journal-title":"BMC Genomics"},{"key":"2023012711573655900_btu392-B5","doi-asserted-by":"crossref","first-page":"563","DOI":"10.1038\/nmeth.2474","article-title":"Nonhybrid, finished microbial genome assemblies from long-read smrt sequencing data","volume":"10","author":"Chin","year":"2013","journal-title":"Nat. Methods"},{"key":"2023012711573655900_btu392-B6","doi-asserted-by":"crossref","first-page":"1011","DOI":"10.1093\/bioinformatics\/btr046","article-title":"Shrimp2: sensitive yet practical short read mapping","volume":"27","author":"David","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012711573655900_btu392-B7","doi-asserted-by":"crossref","first-page":"e105","DOI":"10.1093\/nar\/gkn425","article-title":"Substantial biases in ultra-short read data sets from high-throughput dna sequencing","volume":"36","author":"Dohm","year":"2008","journal-title":"Nucleic Acids Res."},{"key":"2023012711573655900_btu392-B8","doi-asserted-by":"crossref","first-page":"133","DOI":"10.1126\/science.1162986","article-title":"Real-time dna sequencing from single polymerase molecules","volume":"323","author":"Eid","year":"2009","journal-title":"Science"},{"key":"2023012711573655900_btu392-B9","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1186\/2049-2618-1-10","article-title":"Microbial phylogenetic profiling with the pacific biosciences sequencing platform","volume":"1","author":"Fichot","year":"2013","journal-title":"Microbiome"},{"key":"2023012711573655900_btu392-B10","doi-asserted-by":"crossref","first-page":"1513","DOI":"10.1073\/pnas.1017351108","article-title":"High-quality draft assemblies of mammalian genomes from massively parallel sequence data","volume":"108","author":"Gnerre","year":"2011","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012711573655900_btu392-B11","first-page":"44","article-title":"Slurm: simple linux utility for resource management","volume-title":"In Lecture Notes in Computer Science: Proceedings of Job Scheduling Strategies for Parallel Processing (JSSPP) 2003","author":"Jette","year":"2002"},{"key":"2023012711573655900_btu392-B12","doi-asserted-by":"crossref","first-page":"R116","DOI":"10.1186\/gb-2010-11-11-r116","article-title":"Quake: quality-aware detection and correction of sequencing errors","volume":"11","author":"Kelley","year":"2010","journal-title":"Genome Biol."},{"key":"2023012711573655900_btu392-B13","doi-asserted-by":"crossref","first-page":"693","DOI":"10.1038\/nbt.2280","article-title":"Hybrid error correction and de novo assembly of single-molecule sequencing reads","volume":"30","author":"Koren","year":"2012","journal-title":"Nat. Biotechnol."},{"key":"2023012711573655900_btu392-B14","doi-asserted-by":"crossref","first-page":"357","DOI":"10.1038\/nmeth.1923","article-title":"Fast gapped-read alignment with bowtie 2","volume":"9","author":"Langmead","year":"2012","journal-title":"Nat. Methods"},{"key":"2023012711573655900_btu392-B15","doi-asserted-by":"crossref","first-page":"2078","DOI":"10.1093\/bioinformatics\/btp352","article-title":"The sequence alignment\/map format and samtools","volume":"25","author":"Li","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012711573655900_btu392-B16","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1101\/gr.097261.109","article-title":"De novo assembly of human genomes with massively parallel short read sequencing","volume":"20","author":"Li","year":"2010","journal-title":"Genome Res"},{"key":"2023012711573655900_btu392-B17","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1101\/gr.141705.112","article-title":"Sequencing the unsequenceable: expanded cgg-repeat alleles of the fragile x gene","volume":"23","author":"Loomis","year":"2013","journal-title":"Genome Res."},{"key":"2023012711573655900_btu392-B18","doi-asserted-by":"crossref","first-page":"2818","DOI":"10.1093\/bioinformatics\/btn548","article-title":"Aggressive assembly of pyrosequencing reads with mates","volume":"24","author":"Miller","year":"2008","journal-title":"Bioinformatics"},{"key":"2023012711573655900_btu392-B19","doi-asserted-by":"crossref","first-page":"2196","DOI":"10.1126\/science.287.5461.2196","article-title":"A whole-genome assembly of drosophila","volume":"287","author":"Myers","year":"2000","journal-title":"Science"},{"key":"2023012711573655900_btu392-B20","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1093\/bioinformatics\/bts649","article-title":"Pbsim: Pacbio reads simulator\u2013toward accurate genome assembly","volume":"29","author":"Ono","year":"2013","journal-title":"Bioinformatics"},{"key":"2023012711573655900_btu392-B21","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1186\/gb-2013-14-6-405","article-title":"The advantages of smrt sequencing","volume":"14","author":"Roberts","year":"2013","journal-title":"Genome Biol."},{"key":"2023012711573655900_btu392-B22","doi-asserted-by":"crossref","first-page":"R51","DOI":"10.1186\/gb-2013-14-5-r51","article-title":"Characterizing and measuring bias in sequence data","volume":"14","author":"Ross","year":"2013","journal-title":"Genome Biol."},{"key":"2023012711573655900_btu392-B23","doi-asserted-by":"crossref","first-page":"e68824","DOI":"10.1371\/journal.pone.0068824","article-title":"Advantages of single-molecule real-time sequencing in high-gc content genomes","volume":"8","author":"Shin","year":"2013","journal-title":"PLoS One"},{"key":"2023012711573655900_btu392-B24","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1186\/1471-2105-6-31","article-title":"Automated generation of heuristics for biological sequence comparison","volume":"6","author":"Slater","year":"2005","journal-title":"BMC Bioinformatics"},{"key":"2023012711573655900_btu392-B25","doi-asserted-by":"crossref","first-page":"e159","DOI":"10.1093\/nar\/gkq543","article-title":"A flexible and efficient template format for circular consensus sequencing and snp detection","volume":"38","author":"Travers","year":"2010","journal-title":"Nucleic Acids Res."},{"key":"2023012711573655900_btu392-B26","doi-asserted-by":"crossref","first-page":"1859","DOI":"10.1093\/bioinformatics\/bti310","article-title":"Gmap: a genomic mapping and alignment program for mrna and est sequences","volume":"21","author":"Wu","year":"2005","journal-title":"Bioinformatics"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/21\/3004\/48930910\/bioinformatics_30_21_3004.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/21\/3004\/48930910\/bioinformatics_30_21_3004.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T12:51:52Z","timestamp":1674823912000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/30\/21\/3004\/2422147"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,7,10]]},"references-count":26,"journal-issue":{"issue":"21","published-print":{"date-parts":[[2014,11,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btu392","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2014,11,1]]},"published":{"date-parts":[[2014,7,10]]}}}