{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T19:37:02Z","timestamp":1784921822441,"version":"3.55.0"},"reference-count":20,"publisher":"Oxford University Press (OUP)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2014,1,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Genome assembly tools based on the de Bruijn graph framework rely on a parameter k, which represents a trade-off between several competing effects that are difficult to quantify. There is currently a lack of tools that would automatically estimate the best k to use and\/or quickly generate histograms of k-mer abundances that would allow the user to make an informed decision.<\/jats:p>\n               <jats:p>Results: We develop a fast and accurate sampling method that constructs approximate abundance histograms with several orders of magnitude performance improvement over traditional methods. We then present a fast heuristic that uses the generated abundance histograms for putative k values to estimate the best possible value of k. We test the effectiveness of our tool using diverse sequencing datasets and find that its choice of k leads to some of the best assemblies.<\/jats:p>\n               <jats:p>Availability: Our tool KmerGenie is freely available at: http:\/\/kmergenie.bx.psu.edu\/.<\/jats:p>\n               <jats:p>Contact: \u00a0pashadag@cse.psu.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/btt310","type":"journal-article","created":{"date-parts":[[2013,6,4]],"date-time":"2013-06-04T01:26:49Z","timestamp":1370309209000},"page":"31-37","source":"Crossref","is-referenced-by-count":681,"title":["Informed and automated <i>k<\/i>-mer size selection for genome assembly"],"prefix":"10.1093","volume":"30","author":[{"given":"Rayan","family":"Chikhi","sequence":"first","affiliation":[{"name":"1 Department of Computer Science and Engineering and 2Department of Biochemistry and Molecular Biology, The Pennsylvania State University, University Park, PA 16802, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Paul","family":"Medvedev","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science and Engineering and 2Department of Biochemistry and Molecular Biology, The Pennsylvania State University, University Park, PA 16802, USA"},{"name":"1 Department of Computer Science and Engineering and 2Department of Biochemistry and Molecular Biology, The Pennsylvania State University, University Park, PA 16802, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2013,6,3]]},"reference":[{"key":"2023012710375013900_btt310-B1","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1038\/nmeth.1527","article-title":"Limitations of next-generation genome sequence assembly","volume":"8","author":"Alkan","year":"2011","journal-title":"Nat. Methods"},{"key":"2023012710375013900_btt310-B2","doi-asserted-by":"crossref","first-page":"455","DOI":"10.1089\/cmb.2012.0021","article-title":"SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing","volume":"19","author":"Bankevich","year":"2013","journal-title":"J. Comput. Biol."},{"key":"2023012710375013900_btt310-B3","article-title":"Assemblathon 2: evaluating de novo methods of genome assembly in three vertebrate species. arXiv preprint arXiv:1301.5406","author":"Bradnam","year":"2013"},{"key":"2023012710375013900_btt310-B4","doi-asserted-by":"crossref","first-page":"324","DOI":"10.1101\/gr.7088808","article-title":"Short read fragment assembly of bacterial genomes","volume":"18","author":"Chaisson","year":"2008","journal-title":"Genome Res"},{"key":"2023012710375013900_btt310-B5","doi-asserted-by":"crossref","first-page":"236","DOI":"10.1007\/978-3-642-33122-0_19","article-title":"Space-efficient and exact de Bruijn graph representation based on a bloom filter","volume-title":"Algorithms in Bioinformatics, Lecture Notes in Computer Science","author":"Chikhi","year":"2012"},{"key":"2023012710375013900_btt310-B6","doi-asserted-by":"crossref","first-page":"915","DOI":"10.1038\/nbt.1966","article-title":"Efficient de novo assembly of single-cell bacterial genomes from short-read data sets","volume":"29","author":"Chitsaz","year":"2011","journal-title":"Nat. Biotechnol."},{"key":"2023012710375013900_btt310-B7","first-page":"25","article-title":"Summarizing and mining inverse distributions on data streams via dynamic inverse sampling","volume-title":"Proceedings of the 31st international conference on Very large data bases","author":"Cormode","year":"2005"},{"key":"2023012710375013900_btt310-B8","doi-asserted-by":"crossref","first-page":"2224","DOI":"10.1101\/gr.126599.111","article-title":"Assemblathon 1: a competitive assessment of de novo short read assembly methods","volume":"21","author":"Earl","year":"2011","journal-title":"Genome Res."},{"key":"2023012710375013900_btt310-B9","doi-asserted-by":"crossref","first-page":"1072","DOI":"10.1093\/bioinformatics\/btt086","article-title":"QUAST: quality assessment tool for genome assemblies","volume":"29","author":"Gurevich","year":"2013","journal-title":"Bioinformatics"},{"key":"2023012710375013900_btt310-B10","doi-asserted-by":"crossref","first-page":"R116","DOI":"10.1186\/gb-2010-11-11-r116","article-title":"Quake: quality-aware detection and correction of sequencing errors","volume":"11","author":"Kelley","year":"2010","journal-title":"Genome Biol."},{"key":"2023012710375013900_btt310-B11","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/2047-217X-1-18","article-title":"SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler","volume":"1","author":"Luo","year":"2012","journal-title":"GigaScience"},{"key":"2023012710375013900_btt310-B12","doi-asserted-by":"crossref","first-page":"764","DOI":"10.1093\/bioinformatics\/btr011","article-title":"A fast, lock-free approach for efficient parallel counting of occurrences of k-mers","volume":"27","author":"Mar\u00e7ais","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012710375013900_btt310-B13","doi-asserted-by":"crossref","first-page":"1420","DOI":"10.1093\/bioinformatics\/bts174","article-title":"IDBA-UD: a de novo assembler for single-cell and metagenomic sequencing data with highly uneven depth","volume":"28","author":"Peng","year":"2012","journal-title":"Bioinformatics"},{"key":"2023012710375013900_btt310-B14","doi-asserted-by":"crossref","first-page":"9748","DOI":"10.1073\/pnas.171285098","article-title":"An Eulerian path approach to DNA fragment assembly","volume":"98","author":"Pevzner","year":"2001","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012710375013900_btt310-B15","volume-title":"Numerical Recipes 3rd Edition: The Art of Scientific Computing","author":"Press","year":"2007"},{"key":"2023012710375013900_btt310-B16","doi-asserted-by":"crossref","first-page":"2270","DOI":"10.1101\/gr.141515.112","article-title":"Finished bacterial genomes from shotgun sequence data","volume":"22","author":"Ribeiro","year":"2012","journal-title":"Genome Res."},{"key":"2023012710375013900_btt310-B17","doi-asserted-by":"crossref","first-page":"652","DOI":"10.1093\/bioinformatics\/btt020","article-title":"DSK: k-mer counting with very low memory usage","volume":"29","author":"Rizk","year":"2013","journal-title":"Bioinformatics"},{"key":"2023012710375013900_btt310-B18","doi-asserted-by":"crossref","first-page":"557","DOI":"10.1101\/gr.131383.111","article-title":"GAGE: a critical evaluation of genome assemblies and assembly algorithms","volume":"22","author":"Salzberg","year":"2011","journal-title":"Genome Res."},{"key":"2023012710375013900_btt310-B19","doi-asserted-by":"crossref","first-page":"549","DOI":"10.1101\/gr.126953.111","article-title":"Efficient de novo assembly of large genomes using compressed data structures","volume":"22","author":"Simpson","year":"2011","journal-title":"Genome Res."},{"key":"2023012710375013900_btt310-B20","doi-asserted-by":"crossref","first-page":"821","DOI":"10.1101\/gr.074492.107","article-title":"Velvet: algorithms for de novo short read assembly using de Bruijn graphs","volume":"18","author":"Zerbino","year":"2008","journal-title":"Genome Res."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/1\/31\/48912483\/bioinformatics_30_1_31.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/1\/31\/48912483\/bioinformatics_30_1_31.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T10:38:22Z","timestamp":1674815902000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/30\/1\/31\/235479"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,6,3]]},"references-count":20,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2014,1,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btt310","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2014,1,1]]},"published":{"date-parts":[[2013,6,3]]}}}