{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,4]],"date-time":"2026-03-04T12:39:57Z","timestamp":1772627997737,"version":"3.50.1"},"reference-count":29,"publisher":"Oxford University Press (OUP)","issue":"24","license":[{"start":{"date-parts":[[2016,10,2]],"date-time":"2016-10-02T00:00:00Z","timestamp":1475366400000},"content-version":"vor","delay-in-days":2900,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/2.0\/uk\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2008,12,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: DNA sequence reads from Sanger and pyrosequencing platforms differ in cost, accuracy, typical coverage, average read length and the variety of available paired-end protocols. Both read types can complement one another in a \u2018hybrid\u2019 approach to whole-genome shotgun sequencing projects, but assembly software must be modified to accommodate their different characteristics. This is true even of pyrosequencing mated and unmated read combinations. Without special modifications, assemblers tuned for homogeneous sequence data may perform poorly on hybrid data.<\/jats:p>\n               <jats:p>Results: Celera Assembler was modified for combinations of ABI 3730 and 454 FLX reads. The revised pipeline called CABOG (Celera Assembler with the Best Overlap Graph) is robust to homopolymer run length uncertainty, high read coverage and heterogeneous read lengths. In tests on four genomes, it generated the longest contigs among all assemblers tested. It exploited the mate constraints provided by paired-end reads from either platform to build larger contigs and scaffolds, which were validated by comparison to a finished reference sequence. A low rate of contig mis-assembly was detected in some CABOG assemblies, but this was reduced in the presence of sufficient mate pair data.<\/jats:p>\n               <jats:p>Availability: The software is freely available as open-source from http:\/\/wgs-assembler.sf.net under the GNU Public License.<\/jats:p>\n               <jats:p>Contact: \u00a0jmiller@jcvi.org<\/jats:p>\n               <jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btn548","type":"journal-article","created":{"date-parts":[[2008,10,25]],"date-time":"2008-10-25T00:24:48Z","timestamp":1224894288000},"page":"2818-2824","source":"Crossref","is-referenced-by-count":474,"title":["Aggressive assembly of pyrosequencing reads with mates"],"prefix":"10.1093","volume":"24","author":[{"given":"Jason R.","family":"Miller","sequence":"first","affiliation":[{"name":"1 The J. Craig Venter Institute, 9712 Medical Center Drive, Rockville MD 20850, 2Center for Bioinformatics & Computational Biology, University of Maryland, College Park, MD 20742 and 3White Oak Technologies Inc, 1300 Spring St., Ste 320, Silver Spring, MD 20910, USA"}]},{"given":"Arthur L.","family":"Delcher","sequence":"additional","affiliation":[{"name":"1 The J. Craig Venter Institute, 9712 Medical Center Drive, Rockville MD 20850, 2Center for Bioinformatics & Computational Biology, University of Maryland, College Park, MD 20742 and 3White Oak Technologies Inc, 1300 Spring St., Ste 320, Silver Spring, MD 20910, USA"}]},{"given":"Sergey","family":"Koren","sequence":"additional","affiliation":[{"name":"1 The J. Craig Venter Institute, 9712 Medical Center Drive, Rockville MD 20850, 2Center for Bioinformatics & Computational Biology, University of Maryland, College Park, MD 20742 and 3White Oak Technologies Inc, 1300 Spring St., Ste 320, Silver Spring, MD 20910, USA"}]},{"given":"Eli","family":"Venter","sequence":"additional","affiliation":[{"name":"1 The J. Craig Venter Institute, 9712 Medical Center Drive, Rockville MD 20850, 2Center for Bioinformatics & Computational Biology, University of Maryland, College Park, MD 20742 and 3White Oak Technologies Inc, 1300 Spring St., Ste 320, Silver Spring, MD 20910, USA"}]},{"given":"Brian P.","family":"Walenz","sequence":"additional","affiliation":[{"name":"1 The J. Craig Venter Institute, 9712 Medical Center Drive, Rockville MD 20850, 2Center for Bioinformatics & Computational Biology, University of Maryland, College Park, MD 20742 and 3White Oak Technologies Inc, 1300 Spring St., Ste 320, Silver Spring, MD 20910, USA"}]},{"given":"Anushka","family":"Brownley","sequence":"additional","affiliation":[{"name":"1 The J. Craig Venter Institute, 9712 Medical Center Drive, Rockville MD 20850, 2Center for Bioinformatics & Computational Biology, University of Maryland, College Park, MD 20742 and 3White Oak Technologies Inc, 1300 Spring St., Ste 320, Silver Spring, MD 20910, USA"}]},{"given":"Justin","family":"Johnson","sequence":"additional","affiliation":[{"name":"1 The J. Craig Venter Institute, 9712 Medical Center Drive, Rockville MD 20850, 2Center for Bioinformatics & Computational Biology, University of Maryland, College Park, MD 20742 and 3White Oak Technologies Inc, 1300 Spring St., Ste 320, Silver Spring, MD 20910, USA"}]},{"given":"Kelvin","family":"Li","sequence":"additional","affiliation":[{"name":"1 The J. Craig Venter Institute, 9712 Medical Center Drive, Rockville MD 20850, 2Center for Bioinformatics & Computational Biology, University of Maryland, College Park, MD 20742 and 3White Oak Technologies Inc, 1300 Spring St., Ste 320, Silver Spring, MD 20910, USA"}]},{"given":"Clark","family":"Mobarry","sequence":"additional","affiliation":[{"name":"1 The J. Craig Venter Institute, 9712 Medical Center Drive, Rockville MD 20850, 2Center for Bioinformatics & Computational Biology, University of Maryland, College Park, MD 20742 and 3White Oak Technologies Inc, 1300 Spring St., Ste 320, Silver Spring, MD 20910, USA"}]},{"given":"Granger","family":"Sutton","sequence":"additional","affiliation":[{"name":"1 The J. Craig Venter Institute, 9712 Medical Center Drive, Rockville MD 20850, 2Center for Bioinformatics & Computational Biology, University of Maryland, College Park, MD 20742 and 3White Oak Technologies Inc, 1300 Spring St., Ste 320, Silver Spring, MD 20910, USA"}]}],"member":"286","published-online":{"date-parts":[[2008,10,24]]},"reference":[{"key":"2023020212303141400_B1","doi-asserted-by":"crossref","first-page":"545","DOI":"10.1016\/j.gde.2006.10.009","article-title":"Whole-genome re-sequencing","volume":"16","author":"Bentley","year":"2006","journal-title":"Curr. Opin. Genet. Dev"},{"key":"2023020212303141400_B2","doi-asserted-by":"crossref","first-page":"1453","DOI":"10.1126\/science.277.5331.1453","article-title":"The complete genome sequence of Escherichia coli K-12","volume":"277","author":"Blattner","year":"1997","journal-title":"Science"},{"key":"2023020212303141400_B3","doi-asserted-by":"crossref","first-page":"324","DOI":"10.1101\/gr.7088808","article-title":"Short read fragment assembly of bacterial genomes","volume":"18","author":"Chaisson","year":"2008","journal-title":"Genome Res."},{"key":"2023020212303141400_B4","doi-asserted-by":"crossref","first-page":"1093","DOI":"10.1093\/bioinformatics\/17.12.1093","article-title":"DNA sequence quality trimming and vector removal","volume":"17","author":"Chou","year":"2001","journal-title":"Bioinformatics"},{"key":"2023020212303141400_B5","doi-asserted-by":"crossref","first-page":"1035","DOI":"10.1093\/bioinformatics\/btn074","article-title":"Consensus generation and variant detection by Celera Assembler","volume":"24","author":"Denisov","year":"2008","journal-title":"Bioinformatics"},{"key":"2023020212303141400_B6","doi-asserted-by":"crossref","first-page":"967","DOI":"10.1101\/gr.8.9.967","article-title":"A computer program for aligning a cDNA sequence with a genomic DNA sequence","volume":"8","author":"Florea","year":"1998","journal-title":"Genome Res."},{"key":"2023020212303141400_B7","doi-asserted-by":"crossref","first-page":"11240","DOI":"10.1073\/pnas.0604351103","article-title":"A Sanger\/pyrosequencing hybrid approach for the generation of high-quality draft assemblies of marine microbial genomes","volume":"103","author":"Goldberg","year":"2006","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020212303141400_B8","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511574931","volume-title":"Algorithms on Strings, Trees and Sequences: Computer Science and Computational Biology.","author":"Gusfield","year":"1997"},{"key":"2023020212303141400_B9","doi-asserted-by":"crossref","first-page":"1518","DOI":"10.1242\/jeb.001370","article-title":"Advanced sequencing technologies and their wider impact in microbiology","volume":"210","author":"Hall","year":"2007","journal-title":"J. Exp. Biol."},{"key":"2023020212303141400_B10","doi-asserted-by":"crossref","DOI":"10.1002\/0471250953.bi1103s11","article-title":"Generating a genome assembly with PCAP","author":"Huang","year":"2005","journal-title":"Curr. Protoc. Bioinformatics"},{"key":"2023020212303141400_B11","doi-asserted-by":"crossref","first-page":"1916","DOI":"10.1073\/pnas.0307971100","article-title":"Whole-genome shotgun assembly and comparison of human genome assemblies","volume":"101","author":"Istrail","year":"2004","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020212303141400_B12","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1101\/gr.828403","article-title":"Whole-genome sequence assembly for mammalian genomes: Arachne 2","volume":"13","author":"Jaffe","year":"2003","journal-title":"Genome Res."},{"key":"2023020212303141400_B13","doi-asserted-by":"crossref","first-page":"829","DOI":"10.2144\/000112894","article-title":"De novo assembly and genomic structural variation analysis with genome sequencer FLX 3K long-tag paired end reads","volume":"44","author":"Jarvie","year":"2008","journal-title":"Biotechniques"},{"key":"2023020212303141400_B14","doi-asserted-by":"crossref","first-page":"420","DOI":"10.1126\/science.1149504","article-title":"Paired-end mapping reveals extensive structural variation in the human genome","volume":"318","author":"Korbel","year":"2007","journal-title":"Science"},{"key":"2023020212303141400_B15","doi-asserted-by":"crossref","first-page":"4633","DOI":"10.1093\/nar\/29.22.4633","article-title":"REPuter: the manifold applications of repeat analysis on a genomic scale","volume":"29","author":"Kurtz","year":"2001","journal-title":"Nucleic Acids Res."},{"key":"2023020212303141400_B16","doi-asserted-by":"crossref","first-page":"R12","DOI":"10.1186\/gb-2004-5-2-r12","article-title":"Versatile and open software for comparing large genomes","volume":"5","author":"Kurtz","year":"2004","journal-title":"Genome Biol"},{"key":"2023020212303141400_B17","doi-asserted-by":"crossref","first-page":"e254","DOI":"10.1371\/journal.pbio.0050254","article-title":"The diploid genome sequence of anindividual human","volume":"5","author":"Levy","year":"2007","journal-title":"PLoS Biol"},{"key":"2023020212303141400_B18","doi-asserted-by":"crossref","first-page":"376","DOI":"10.1038\/nature03959","article-title":"Genome sequencing in microfabricated high-density picolitre reactors","volume":"437","author":"Margulies","year":"2005","journal-title":"Nature"},{"key":"2023020212303141400_B19","doi-asserted-by":"crossref","first-page":"2196","DOI":"10.1126\/science.287.5461.2196","article-title":"A whole-genome assembly of Drosophila","volume":"287","author":"Myers","year":"2000","journal-title":"Scienc"},{"key":"2023020212303141400_B20","doi-asserted-by":"crossref","first-page":"5591","DOI":"10.1128\/JB.185.18.5591-5601.2003","article-title":"Complete genome sequence of the oral pathogenic Bacterium Porphyromonas gingivalis strain W83","volume":"185","author":"Nelson","year":"2003","journal-title":"J. Bacteriol."},{"key":"2023020212303141400_B21","doi-asserted-by":"crossref","first-page":"9748","DOI":"10.1073\/pnas.171285098","article-title":"An Eulerian path approach to DNA fragment assembly","volume":"98","author":"Pevzner","year":"2001","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020212303141400_B22","doi-asserted-by":"crossref","first-page":"734","DOI":"10.1089\/cmb.2004.11.734","article-title":"A preprocessor for shotgun assembly of large genomes","volume":"11","author":"Roberts","year":"2004","journal-title":"J. Comput. Biol."},{"key":"2023020212303141400_B23","volume-title":"Genome Sequencer FLX Data Analysis Software Manual.","author":"Roche","year":"2007"},{"key":"2023020212303141400_B24","doi-asserted-by":"crossref","first-page":"927","DOI":"10.1038\/nature03062","article-title":"Shotgun sequence assembly and recent segmental duplications within the human genome","volume":"431","author":"She","year":"2004","journal-title":"Nature"},{"key":"2023020212303141400_B25","doi-asserted-by":"crossref","first-page":"9","DOI":"10.1089\/gst.1995.1.9","article-title":"TIGR Assembler: a new tool for assembling large shotgun sequencing projects","volume":"1","author":"Sutton","year":"1995","journal-title":"Genome Sci. Technol."},{"key":"2023020212303141400_B26","doi-asserted-by":"crossref","first-page":"872","DOI":"10.1038\/nature06884","article-title":"The complete genome of an individual by massively parallel DNA sequencing","volume":"452","author":"Wheeler","year":"2008","journal-title":"Nature"},{"key":"2023020212303141400_B27","doi-asserted-by":"crossref","first-page":"462","DOI":"10.1093\/bioinformatics\/btm632","article-title":"Figaro: a novel statistical method for vector sequence removal","volume":"24","author":"White","year":"2008","journal-title":"Bioinformatics"},{"key":"2023020212303141400_B28","doi-asserted-by":"crossref","first-page":"275","DOI":"10.1186\/1471-2164-7-275","article-title":"454 sequencing put to the test using the complex genome of barley","volume":"7","author":"Wicker","year":"2006","journal-title":"BMC Genomics"},{"key":"2023020212303141400_B29","doi-asserted-by":"crossref","first-page":"821","DOI":"10.1101\/gr.074492.107","article-title":"Velvet: algorithms for de novo short read assembly using de Bruijn graphs","volume":"18","author":"Zerbino","year":"2008","journal-title":"Genome Res."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/24\/2818\/49056211\/bioinformatics_24_24_2818.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/24\/2818\/49056211\/bioinformatics_24_24_2818.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,2]],"date-time":"2023-02-02T15:15:48Z","timestamp":1675350948000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/24\/24\/2818\/197033"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,10,24]]},"references-count":29,"journal-issue":{"issue":"24","published-print":{"date-parts":[[2008,12,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btn548","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2008,12,15]]},"published":{"date-parts":[[2008,10,24]]}}}