{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,12]],"date-time":"2026-03-12T00:22:23Z","timestamp":1773274943096,"version":"3.50.1"},"reference-count":28,"publisher":"Springer Science and Business Media LLC","issue":"S1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2009,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>New short-read sequencing technologies produce enormous volumes of 25\u201330 base paired-end reads. The resulting reads have vastly different characteristics than produced by Sanger sequencing, and require different approaches than the previous generation of sequence assemblers. In this paper, we present a short-read de novo assembler particularly targeted at the new ABI SOLiD sequencing technology.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>This paper presents what we believe to be the first de novo sequence assembly results on <jats:italic>real<\/jats:italic> data from the emerging SOLiD platform, introduced by <jats:italic>Applied Biosystems<\/jats:italic>. Our assembler SHORTY augments short-paired reads using a trivially small number (5 \u2013 10) of <jats:italic>seeds<\/jats:italic> of length 300 \u2013 500 bp. These seeds enable us to produce significant assemblies using short-read coverage no more than 100\u00d7, which can be obtained in a single run of these high-capacity sequencers. SHORTY exploits two ideas which we believe to be of interest to the short-read assembly community: (1) using single seed reads to crystallize assemblies, and (2) estimating intercontig distances accurately from multiple spanning paired-end reads.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion<\/jats:title>\n            <jats:p>We demonstrate effective assemblies (N50 contig sizes ~40 kb) of three different bacterial species using simulated SOLiD data. Sequencing artifacts limit our performance on real data, however our results on this data are substantially better than those achieved by competing assemblers.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-10-s1-s16","type":"journal-article","created":{"date-parts":[[2009,1,30]],"date-time":"2009-01-30T20:04:42Z","timestamp":1233345882000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":32,"title":["Crystallizing short-read assemblies around seeds"],"prefix":"10.1186","volume":"10","author":[{"given":"Mohammad Sajjad","family":"Hossain","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Navid","family":"Azimi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Steven","family":"Skiena","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2009,1,30]]},"reference":[{"issue":"4","key":"3199_CR1","doi-asserted-by":"publisher","first-page":"500","DOI":"10.1093\/bioinformatics\/btl629","volume":"23","author":"RL Warren","year":"2007","unstructured":"Warren RL, Sutton GG, Jones SJM, Holt RA: Assembling millions of short DNA sequences using SSAKE. Bioinformatics 2007, 23(4):500\u2013501. 10.1093\/bioinformatics\/btl629","journal-title":"Bioinformatics"},{"key":"3199_CR2","doi-asserted-by":"publisher","first-page":"1697","DOI":"10.1101\/gr.6435207","volume":"17","author":"JC Dohm","year":"2007","unstructured":"Dohm JC, Lottaz C, Borodina T, Himmelbauer H: SHARCGS, a fast and highly accurate short-read assembly algorithm for de novo genomic sequencing. Genome Research 2007, 17: 1697\u20131706. 10.1101\/gr.6435207","journal-title":"Genome Research"},{"issue":"5","key":"3199_CR3","doi-asserted-by":"publisher","first-page":"603","DOI":"10.1145\/585265.585267","volume":"49","author":"D Huson","year":"2002","unstructured":"Huson D, Reinert K, Myers EW: The greedy path-merging algorithm for contig scaffolding. Journal of the ACM (JACM) 2002, 49(5):603\u2013615. 10.1145\/585265.585267","journal-title":"Journal of the ACM (JACM)"},{"key":"3199_CR4","first-page":"5463","volume-title":"Proc Natl Acad Sci USA","author":"F Sanger","year":"1977","unstructured":"Sanger F, Nicklen S, Coulson A: DNA sequencing with chain-terminating inhibitors. Proc Natl Acad Sci USA 1977, 5463\u20137. 10.1073\/pnas.74.12.5463"},{"key":"3199_CR5","doi-asserted-by":"publisher","first-page":"335","DOI":"10.1038\/nrg1325","volume":"5","author":"J Shendure","year":"2004","unstructured":"Shendure J, Mitra R, Church G: Advanced sequencing technologies: methods and goals. Nature Rev Gen 2004, 5: 335\u2013344. 10.1038\/nrg1325","journal-title":"Nature Rev Gen"},{"key":"3199_CR6","doi-asserted-by":"publisher","first-page":"630","DOI":"10.1038\/76469","volume":"18","author":"S Brenner","year":"2003","unstructured":"Brenner S, Johnson M, Bridgham J, Golda G, Lloyd D, Johnson D, Luo S, McCurdy S, Foy M, Ewan M, et al.: Gene expression analysis by massively parallel signature sequencing (MPSS) on microbead arrays. Nat Biotechnol 2003, 18: 630\u2013634. 10.1038\/76469","journal-title":"Nat Biotechnol"},{"key":"3199_CR7","doi-asserted-by":"publisher","first-page":"1425","DOI":"10.1038\/nbt1203-1425","volume":"21","author":"J Kling","year":"2003","unstructured":"Kling J: Ultrafast DNA sequencing. Nat Biotechol 2003, 21: 1425\u20131427. 10.1038\/nbt1203-1425","journal-title":"Nat Biotechol"},{"key":"3199_CR8","doi-asserted-by":"publisher","first-page":"717","DOI":"10.1101\/gr.886203","volume":"13","author":"R Miller","year":"2003","unstructured":"Miller R, Duan S, Lovins E, Kloss E, Kwok PY: Efficient high-throughput resequencing of genomic DNA. Genome Res 2003, 13: 717\u2013720. 10.1101\/gr.886203","journal-title":"Genome Res"},{"key":"3199_CR9","doi-asserted-by":"publisher","first-page":"363","DOI":"10.1126\/science.281.5375.363","volume":"281","author":"M Ronaghi","year":"1998","unstructured":"Ronaghi M, Uhlen M, Nyren P: DNA sequencing: a sequencing method based on real-time pyrophosphate. Science 1998, 281: 363\u2013365. 10.1126\/science.281.5375.363","journal-title":"Science"},{"key":"3199_CR10","unstructured":"Mitchelson KR, (Ed): . In New High Throughput Technologies For DNA Sequencing And Genomics, of Perspectives. Volume 2. Bioanalysis. Elsevier; 2007."},{"key":"3199_CR11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1093\/nar\/27.24.e34","volume":"27","author":"R Mitra","year":"1999","unstructured":"Mitra R, Church G: In situ localized amplification and contact replication of many individual DNA molecules. Nucleic Acids Research 1999, 27: 1\u20136. 10.1093\/nar\/27.24.e34","journal-title":"Nucleic Acids Research"},{"key":"3199_CR12","volume-title":"Analyt Biochem","author":"R Mitra","year":"2003","unstructured":"Mitra R, Shendure J, Olejnik J, Church G: Fluorescent in situ Sequencing on Polymerase Colonies. Analyt Biochem 2003."},{"key":"3199_CR13","doi-asserted-by":"publisher","first-page":"3960","DOI":"10.1073\/pnas.0230489100","volume":"100","author":"I Braslavsky","year":"2003","unstructured":"Braslavsky I, Hebert B, Kartalov E, Quake S: Sequence Information can be obtained from single DNA molecules. PNAS 2003, 100: 3960\u20133964. 10.1073\/pnas.0230489100","journal-title":"PNAS"},{"key":"3199_CR14","first-page":"106","volume-title":"Science","author":"T Harris","year":"2008","unstructured":"Harris T, Buzby P, et al.: Single-molecule DNA sequencing of a viral genome. Science 2008, 106\u20139. 10.1126\/science.1150427"},{"key":"3199_CR15","doi-asserted-by":"publisher","first-page":"7","DOI":"10.1007\/BF01188580","volume":"13","author":"J Kececioglu","year":"1995","unstructured":"Kececioglu J, Myers E: Combinatorial Algorithms for DNA Sequence Assembly. Algorithmica 1995, 13: 7\u201351. 10.1007\/BF01188580","journal-title":"Algorithmica"},{"key":"3199_CR16","doi-asserted-by":"publisher","first-page":"810","DOI":"10.1101\/gr.7337908","volume":"18","author":"J Butler","year":"2008","unstructured":"Butler J, MacCalluma I, Kleber M, Shlyakhter IA, Belmonte MK, Lander ES, Nusbaum C, Jaffe DB: ALLPATHS: De novo assembly of whole-genome shotgun microreads. Genome Research 2008, 18: 810\u2013820. 10.1101\/gr.7337908","journal-title":"Genome Research"},{"key":"3199_CR17","volume-title":"RECOMB","author":"P Medvedev","year":"2008","unstructured":"Medvedev P, Brudno M: Ab Initio Whole Genome Shotgun Assembly With Mated Short Reads. RECOMB 2008."},{"key":"3199_CR18","volume-title":"Genome Research","author":"D Zerbino","year":"2008","unstructured":"Zerbino D, Birney E: Velvet: Algorithms for De Novo Short Read Assembly Using De Bruijn Graphs. Genome Research 2008."},{"key":"3199_CR19","doi-asserted-by":"crossref","unstructured":"Margulies M, Jarvie TP, Knight JR, Simons JF: New High Throughput Technologies for DNA Sequencing and Genomics, Amsterdam: Elsevier 2007 chap. The 454 Life Sciences Picoliter Sequencing System.151\u2013186.","DOI":"10.1016\/S1871-0069(06)02005-2"},{"issue":"13","key":"3199_CR20","doi-asserted-by":"publisher","first-page":"2067","DOI":"10.1093\/bioinformatics\/bth205","volume":"20","author":"M Chaisson","year":"2004","unstructured":"Chaisson M, Pevzner PA, Tang H: Fragment assembly with short reads. Bioinformatics 2004, 20(13):2067\u20132074. 10.1093\/bioinformatics\/bth205","journal-title":"Bioinformatics"},{"key":"3199_CR21","unstructured":"Chaisson M, Pevzner P: Short Read Fragment Assembly of Bacterial Genomes. , in press."},{"key":"3199_CR22","doi-asserted-by":"crossref","unstructured":"Sundquist A, Ronaghi M, Tang H, Pevzner P, Batzoglou S: Whole-genome sequencing and assembly with high-throughput, short-read technologies. PLoS ONE 2007., 2(5):","DOI":"10.1371\/journal.pone.0000484"},{"key":"3199_CR23","doi-asserted-by":"publisher","first-page":"81","DOI":"10.1101\/gr.731003","volume":"13","author":"J Mullikin","year":"2003","unstructured":"Mullikin J, Ning Z: The Phusion Assembler. Genome Res 2003, 13: 81\u201390. 10.1101\/gr.731003","journal-title":"Genome Res"},{"key":"3199_CR24","volume-title":"15th Annual International Conference on Intelligent Systems for Molecular Biology (ISMB) Poster","author":"T Keane","year":"2007","unstructured":"Keane T, Ning Z: Assessing Assemblability of Reads from New Sequencing Platforms. 15th Annual International Conference on Intelligent Systems for Molecular Biology (ISMB) Poster 2007. [http:\/\/www.iscb.org\/uploaded\/css\/B16Keane.pdf]"},{"key":"3199_CR25","volume-title":"Genome Research","author":"D Hernandez","year":"2008","unstructured":"Hernandez D, Francois P, Farinelli L, Osteras M, Schrenzel J: De novo bacterial genome sequencing: millions of very short reads assembled on a desktop computer. Genome Research 2008."},{"key":"3199_CR26","doi-asserted-by":"publisher","first-page":"175","DOI":"10.1101\/gr.8.3.175","volume":"8","author":"B Ewing","year":"1998","unstructured":"Ewing B, Hillier L, Wendl MC, Green P: Base-Calling of Automated Sequencer Traces Phred. I. Using Accuracy Assessment. Genome Res 1998, 8: 175\u2013185.","journal-title":"Genome Res"},{"key":"3199_CR27","doi-asserted-by":"publisher","first-page":"186","DOI":"10.1101\/gr.8.3.186","volume":"8","author":"B Ewing","year":"1998","unstructured":"Ewing B, Green P: Base-Calling of Automated Sequencer Traces Using Phred. II. Error Probabilities. Genome Res 1998, 8: 186\u2013194.","journal-title":"Genome Res"},{"key":"3199_CR28","volume-title":"Genome Biology","author":"S Kurtz","year":"2004","unstructured":"Kurtz S, Phillippy A, Delcher AL, Smoot M, Shumway M, Antonescu C, Salzberg SL: Versatile and open software for comparing large genomes. Genome Biology 2004., 5:"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-10-S1-S16.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T10:46:22Z","timestamp":1630493182000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-10-S1-S16"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,1]]},"references-count":28,"journal-issue":{"issue":"S1","published-print":{"date-parts":[[2009,1]]}},"alternative-id":["3199"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-10-s1-s16","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2009,1]]},"assertion":[{"value":"30 January 2009","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"S16"}}