{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,3]],"date-time":"2026-03-03T01:14:23Z","timestamp":1772500463735,"version":"3.50.1"},"reference-count":15,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2020,3,4]],"date-time":"2020-03-04T00:00:00Z","timestamp":1583280000000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2020,3,4]],"date-time":"2020-03-04T00:00:00Z","timestamp":1583280000000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2020,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Background<\/jats:title>\n                    <jats:p>Duplex sequencing is the most accurate approach for identification of sequence variants present at very low frequencies. Its power comes from pooling together multiple descendants of both strands of original DNA molecules, which allows distinguishing true nucleotide substitutions from PCR amplification and sequencing artifacts. This strategy comes at a cost\u2014sequencing the same molecule multiple times increases dynamic range but significantly diminishes coverage, making whole genome duplex sequencing prohibitively expensive. Furthermore, every duplex experiment produces a substantial proportion of singleton reads that cannot be used in the analysis and are thrown away.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>In this paper we demonstrate that a significant fraction of these reads contains PCR or sequencing errors within duplex tags. Correction of such errors allows \u201creuniting\u201d these reads with their respective families increasing the output of the method and making it more cost effective.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusions<\/jats:title>\n                    <jats:p>\n                      We combine an error correction strategy with a number of algorithmic improvements in a new version of the duplex analysis software, Du Novo 2.0. It is written in Python, C, AWK, and Bash. It is open source and readily available through Galaxy, Bioconda, and Github:\n                      <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/galaxyproject\/dunovo\">https:\/\/github.com\/galaxyproject\/dunovo<\/jats:ext-link>\n                      .\n                    <\/jats:p>\n                  <\/jats:sec>","DOI":"10.1186\/s12859-020-3419-8","type":"journal-article","created":{"date-parts":[[2020,3,4]],"date-time":"2020-03-04T10:02:39Z","timestamp":1583316159000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":14,"title":["Family reunion via error correction: an efficient analysis of duplex sequencing data"],"prefix":"10.1186","volume":"21","author":[{"given":"Nicholas","family":"Stoler","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Barbara","family":"Arbeithuber","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gundula","family":"Povysil","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Monika","family":"Heinzl","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Renato","family":"Salazar","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kateryna D","family":"Makova","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Irene","family":"Tiemann-Boege","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5987-8032","authenticated-orcid":false,"given":"Anton","family":"Nekrutenko","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2020,3,4]]},"reference":[{"key":"3419_CR1","volume-title":"fgbio fulcrumgenomics","author":"T Fennell","year":"2018","unstructured":"Fennell T, Homer N. fgbio fulcrumgenomics; 2018."},{"key":"3419_CR2","doi-asserted-by":"publisher","first-page":"R25","DOI":"10.1186\/gb-2009-10-3-r25","volume":"10","author":"B Langmead","year":"2009","unstructured":"Langmead B, et al. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol. 2009;10:R25.","journal-title":"Genome Biol"},{"key":"3419_CR3","doi-asserted-by":"publisher","first-page":"62","DOI":"10.1186\/1471-2105-6-62","volume":"6","author":"A Larionov","year":"2005","unstructured":"Larionov A, et al. A standard curve based method for relative real time PCR data processing. BMC Bioinformatics. 2005;6:62.","journal-title":"BMC Bioinformatics"},{"key":"3419_CR4","doi-asserted-by":"publisher","first-page":"858","DOI":"10.1093\/nar\/gkn1006","volume":"37","author":"T Lassmann","year":"2009","unstructured":"Lassmann T, et al. Kalign2: high-performance multiple alignment of protein and nucleotide sequences allowing external features. Nucleic Acids Res. 2009;37:858\u201365.","journal-title":"Nucleic Acids Res"},{"key":"3419_CR5","doi-asserted-by":"publisher","DOI":"10.1101\/429175","volume-title":"A high resolution view of adaptive events","author":"H Mei","year":"2018","unstructured":"Mei H, et al. A high resolution view of adaptive events; 2018."},{"key":"3419_CR6","doi-asserted-by":"publisher","first-page":"15474","DOI":"10.1073\/pnas.1409328111","volume":"111","author":"B Rebolledo Jaramillo","year":"2014","unstructured":"Rebolledo Jaramillo B, et al. Maternal age effect and severe germ-line bottleneck in the inheritance of human mitochondrial DNA. Proc Natl Acad Sci U S A. 2014;111:15474\u20139.","journal-title":"Proc Natl Acad Sci U S A"},{"key":"3419_CR7","doi-asserted-by":"publisher","first-page":"269","DOI":"10.1038\/nrg.2017.117","volume":"19","author":"JJ Salk","year":"2018","unstructured":"Salk JJ, et al. Enhancing the accuracy of next-generation sequencing for detecting rare and subclonal mutations. Nat Rev Genet. 2018;19:269\u201385.","journal-title":"Nat Rev Genet"},{"key":"3419_CR8","doi-asserted-by":"publisher","first-page":"14508","DOI":"10.1073\/pnas.1208715109","volume":"109","author":"MW Schmitt","year":"2012","unstructured":"Schmitt MW, et al. Detection of ultra-rare mutations by next-generation sequencing. Proc Natl Acad Sci U S A. 2012;109:14508\u201313.","journal-title":"Proc Natl Acad Sci U S A"},{"key":"3419_CR9","doi-asserted-by":"publisher","first-page":"423","DOI":"10.1038\/nmeth.3351","volume":"12","author":"MW Schmitt","year":"2015","unstructured":"Schmitt MW, et al. Sequencing small genomic targets with high efficiency and extreme accuracy. Nat Methods. 2015;12:423\u20135.","journal-title":"Nat Methods"},{"key":"3419_CR10","doi-asserted-by":"publisher","first-page":"13","DOI":"10.1371\/journal.pcbi.1005480","volume":"13","author":"M Shugay","year":"2017","unstructured":"Shugay M, et al. MAGERI: computational pipeline for molecular-barcoded targeted resequencing. PLoS Comput Biol. 2017;13:13\u20137.","journal-title":"PLoS Comput Biol"},{"key":"3419_CR11","doi-asserted-by":"publisher","first-page":"653","DOI":"10.1038\/nmeth.2960","volume":"11","author":"M Shugay","year":"2014","unstructured":"Shugay M, et al. Towards error-free profiling of immune repertoires. Nat Methods. 2014;11:653\u20135.","journal-title":"Nat Methods"},{"key":"3419_CR12","doi-asserted-by":"publisher","first-page":"491","DOI":"10.1101\/gr.209601.116","volume":"27","author":"T Smith","year":"2017","unstructured":"Smith T, et al. UMI-tools: modeling sequencing errors in unique molecular identifiers to improve quantification accuracy. Genome Res. 2017;27:491\u20139.","journal-title":"Genome Res"},{"key":"3419_CR13","doi-asserted-by":"publisher","first-page":"180","DOI":"10.1186\/s13059-016-1039-4","volume":"17","author":"N Stoler","year":"2016","unstructured":"Stoler N, et al. Streamlined analysis of duplex sequencing data with Du novo. Genome Biol. 2016;17:180.","journal-title":"Genome Biol"},{"key":"3419_CR14","doi-asserted-by":"publisher","first-page":"83","DOI":"10.1145\/135239.135244","volume":"35","author":"S Wu","year":"1992","unstructured":"Wu S, Manber U. Fast text searching: allowing errors. Commun ACM. 1992;35:83\u201391.","journal-title":"Commun ACM"},{"key":"3419_CR15","doi-asserted-by":"crossref","unstructured":"Xu C, et al. smCounter2: an accurate low-frequency variant caller for targeted sequencing data with unique molecular identifiers. bioRxiv. 2018:281659. https:\/\/www.biorxiv.org\/content\/10.1101\/281659v1.","DOI":"10.1101\/281659"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-020-3419-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/s12859-020-3419-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-020-3419-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,3,3]],"date-time":"2021-03-03T19:07:09Z","timestamp":1614798429000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-020-3419-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,3,4]]},"references-count":15,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,12]]}},"alternative-id":["3419"],"URL":"https:\/\/doi.org\/10.1186\/s12859-020-3419-8","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/469106","asserted-by":"object"}]},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,3,4]]},"assertion":[{"value":"18 March 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 February 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 March 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Not applicable.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"96"}}