{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,10]],"date-time":"2026-02-10T19:51:49Z","timestamp":1770753109915,"version":"3.50.0"},"reference-count":14,"publisher":"Oxford University Press (OUP)","issue":"4","license":[{"start":{"date-parts":[[2016,10,2]],"date-time":"2016-10-02T00:00:00Z","timestamp":1475366400000},"content-version":"vor","delay-in-days":3181,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/2.0\/uk\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2008,2,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Sequences produced by automated Sanger sequencing machines frequently contain fragments of the cloning vector on their ends. Software tools currently available for identifying and removing the vector sequence require knowledge of the vector sequence, specific splice sites and any adapter sequences used in the experiment\u2014information often omitted from public databases. Furthermore, the clipping coordinates themselves are missing or incorrectly reported. As an example, within the \u223c1.24 billion shotgun sequences deposited in the NCBI Trace Archive, as many as \u223c735 million (\u223c60%) lack vector clipping information. Correct clipping information is essential to scientists attempting to validate, improve and even finish the increasingly large number of genomes released at a \u2018draft\u2019 quality level.<\/jats:p>\n               <jats:p>Results: We present here Figaro, a novel software tool for identifying and removing the vector from raw sequence data without prior knowledge of the vector sequence. The vector sequence is automatically inferred by analyzing the frequency of occurrence of short oligo-nucleotides using Poisson statistics. We show that Figaro achieves 99.98% sensitivity when tested on \u223c1.5 million shotgun reads from Drosophila pseudoobscura. We further explore the impact of accurate vector trimming on the quality of whole-genome assemblies by re-assembling two bacterial genomes from shotgun sequences deposited in the Trace Archive. Designed as a module in large computational pipelines, Figaro is fast, lightweight and flexible.<\/jats:p>\n               <jats:p>Availability: Figaro is released under an open-source license through the AMOS package (http:\/\/amos.sourceforge.net\/Figaro).<\/jats:p>\n               <jats:p>Contact: \u00a0mpop@umiacs.umd.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/btm632","type":"journal-article","created":{"date-parts":[[2008,1,18]],"date-time":"2008-01-18T01:14:37Z","timestamp":1200618877000},"page":"462-467","source":"Crossref","is-referenced-by-count":29,"title":["Figaro: a novel statistical method for vector sequence removal"],"prefix":"10.1093","volume":"24","author":[{"given":"James Robert","family":"White","sequence":"first","affiliation":[{"name":"1 Center for Bioinformatics and Computational Biology, University of Maryland \u2013 College Park, MD 20742, 2Applied Mathematics and Scientific Computation Program, University of Maryland \u2013 College Park, MD 20742 and 3Institute for Physical Sciences and Technology, University of Maryland \u2013 College Park, MD 20742, USA"},{"name":"1 Center for Bioinformatics and Computational Biology, University of Maryland \u2013 College Park, MD 20742, 2Applied Mathematics and Scientific Computation Program, University of Maryland \u2013 College Park, MD 20742 and 3Institute for Physical Sciences and Technology, University of Maryland \u2013 College Park, MD 20742, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael","family":"Roberts","sequence":"additional","affiliation":[{"name":"1 Center for Bioinformatics and Computational Biology, University of Maryland \u2013 College Park, MD 20742, 2Applied Mathematics and Scientific Computation Program, University of Maryland \u2013 College Park, MD 20742 and 3Institute for Physical Sciences and Technology, University of Maryland \u2013 College Park, MD 20742, USA"},{"name":"1 Center for Bioinformatics and Computational Biology, University of Maryland \u2013 College Park, MD 20742, 2Applied Mathematics and Scientific Computation Program, University of Maryland \u2013 College Park, MD 20742 and 3Institute for Physical Sciences and Technology, University of Maryland \u2013 College Park, MD 20742, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"James A.","family":"Yorke","sequence":"additional","affiliation":[{"name":"1 Center for Bioinformatics and Computational Biology, University of Maryland \u2013 College Park, MD 20742, 2Applied Mathematics and Scientific Computation Program, University of Maryland \u2013 College Park, MD 20742 and 3Institute for Physical Sciences and Technology, University of Maryland \u2013 College Park, MD 20742, USA"},{"name":"1 Center for Bioinformatics and Computational Biology, University of Maryland \u2013 College Park, MD 20742, 2Applied Mathematics and Scientific Computation Program, University of Maryland \u2013 College Park, MD 20742 and 3Institute for Physical Sciences and Technology, University of Maryland \u2013 College Park, MD 20742, USA"},{"name":"1 Center for Bioinformatics and Computational Biology, University of Maryland \u2013 College Park, MD 20742, 2Applied Mathematics and Scientific Computation Program, University of Maryland \u2013 College Park, MD 20742 and 3Institute for Physical Sciences and Technology, University of Maryland \u2013 College Park, MD 20742, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mihai","family":"Pop","sequence":"additional","affiliation":[{"name":"1 Center for Bioinformatics and Computational Biology, University of Maryland \u2013 College Park, MD 20742, 2Applied Mathematics and Scientific Computation Program, University of Maryland \u2013 College Park, MD 20742 and 3Institute for Physical Sciences and Technology, University of Maryland \u2013 College Park, MD 20742, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2008,1,17]]},"reference":[{"key":"2023020209510448000_B1","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1006\/abio.1996.0138","article-title":"A \u2018double adaptor\u2019 method for improved shotgun library construction","volume":"236","author":"Andersson","year":"1996","journal-title":"Anal. Biochem"},{"key":"2023020209510448000_B2","doi-asserted-by":"crossref","first-page":"1093","DOI":"10.1093\/bioinformatics\/17.12.1093","article-title":"DNA sequence quality trimming and vector removal","volume":"17","author":"Chou","year":"2001","journal-title":"Bioinformatics"},{"key":"2023020209510448000_B3","doi-asserted-by":"crossref","first-page":"2478","DOI":"10.1093\/nar\/30.11.2478","article-title":"Fast algorithms for large-scale genome alignment and comparison","volume":"30","author":"Delcher","year":"2002","journal-title":"Nucleic Acids Res"},{"key":"2023020209510448000_B4","doi-asserted-by":"crossref","first-page":"R12","DOI":"10.1186\/gb-2004-5-2-r12","article-title":"Versatile and open software for comparing large genomes","volume":"5","author":"Kurtz","year":"2004","journal-title":"Genome Biol"},{"key":"2023020209510448000_B5","doi-asserted-by":"crossref","first-page":"376","DOI":"10.1038\/nature03959","article-title":"Genome sequencing in microfabricated high-density picolitre reactors","volume":"437","author":"Margulies","year":"2005","journal-title":"Nature"},{"key":"2023020209510448000_B6","doi-asserted-by":"crossref","first-page":"2196","DOI":"10.1126\/science.287.5461.2196","article-title":"A whole-genome assembly of Drosophila","volume":"287","author":"Myers","year":"2000","journal-title":"Science"},{"key":"2023020209510448000_B7","doi-asserted-by":"crossref","first-page":"9748","DOI":"10.1073\/pnas.171285098","article-title":"An Eulerian path approach to DNA fragment assembly","volume":"98","author":"Pevzner","year":"2001","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020209510448000_B8","doi-asserted-by":"crossref","first-page":"149","DOI":"10.1016\/j.pbi.2006.01.015","article-title":"The maize genome as a model for efficient sequence analysis of large plant genomes","volume":"9","author":"Rabinowicz","year":"2006","journal-title":"Curr. Opin. Plant Biol"},{"key":"2023020209510448000_B9","doi-asserted-by":"crossref","first-page":"2134","DOI":"10.1093\/nar\/gkg321","article-title":"Genome sequence of Chlamydophila caviae (Chlamydia psittaci GPIC): examining the role of niche-specific genes in the evolution of the Chlamydiaceae","volume":"31","author":"Read","year":"2003","journal-title":"Nucleic Acids Res"},{"key":"2023020209510448000_B10","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1101\/gr.3059305","article-title":"Comparative genome sequencing of Drosophila pseudoobscura: chromosomal, gene, and cis-element evolution","volume":"15","author":"Richards","year":"2005","journal-title":"Genome Res"},{"key":"2023020209510448000_B11","doi-asserted-by":"crossref","first-page":"5463","DOI":"10.1073\/pnas.74.12.5463","article-title":"DNA sequencing with chain-terminating inhibitors","volume":"74","author":"Sanger","year":"1977","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020209510448000_B12","doi-asserted-by":"crossref","first-page":"5455","DOI":"10.1073\/pnas.0931379100","article-title":"Complete genome sequence of the Q-fever pathogen Coxiella burnetii","volume":"100","author":"Seshadri","year":"2003","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020209510448000_B13","doi-asserted-by":"crossref","first-page":"1304","DOI":"10.1126\/science.1058040","article-title":"The sequence of the human genome","volume":"291","author":"Venter","year":"2001","journal-title":"Science"},{"key":"2023020209510448000_B14","article-title":"Accession numbers for finished genomes in GenBank: AE015925 and AE016828"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/4\/462\/49045641\/bioinformatics_24_4_462.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/4\/462\/49045641\/bioinformatics_24_4_462.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,2]],"date-time":"2023-02-02T10:40:09Z","timestamp":1675334409000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/24\/4\/462\/207640"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,1,17]]},"references-count":14,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2008,2,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btm632","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2008,2,15]]},"published":{"date-parts":[[2008,1,17]]}}}