{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T09:03:07Z","timestamp":1778058187112,"version":"3.51.4"},"reference-count":28,"publisher":"Oxford University Press (OUP)","issue":"7","license":[{"start":{"date-parts":[[2017,11,29]],"date-time":"2017-11-29T00:00:00Z","timestamp":1511913600000},"content-version":"vor","delay-in-days":1,"URL":"https:\/\/academic.oup.com\/journals\/pages\/about_us\/legal\/notices"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"NSF","doi-asserted-by":"publisher","award":["DBI-ABI 0965596"],"award-info":[{"award-number":["DBI-ABI 0965596"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"NSF","doi-asserted-by":"publisher","award":["DBI-1356529, IIS-1453527, IIS-1421908 and CCF-1439057"],"award-info":[{"award-number":["DBI-1356529, IIS-1453527, IIS-1421908 and CCF-1439057"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000002","name":"NIH","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2018,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Motivation<\/jats:title>\n                    <jats:p>The haploid mammalian Y chromosome is usually under-represented in genome assemblies due to high repeat content and low depth due to its haploid nature. One strategy to ameliorate the low coverage of Y sequences is to experimentally enrich Y-specific material before assembly. As the enrichment process is imperfect, algorithms are needed to identify putative Y-specific reads prior to downstream assembly. A strategy that uses k-mer abundances to identify such reads was used to assemble the gorilla Y. However, the strategy required the manual setting of key parameters, a time-consuming process leading to sub-optimal assemblies.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>We develop a method, RecoverY, that selects Y-specific reads by automatically choosing the abundance level at which a k-mer is deemed to originate from the Y. This algorithm uses prior knowledge about the Y chromosome of a related species or known Y transcript sequences. We evaluate RecoverY on both simulated and real data, for human and gorilla, and investigate its robustness to important parameters. We show that RecoverY leads to a vastly superior assembly compared to alternate strategies of filtering the reads or contigs. Compared to the preliminary strategy used by Tomaszkiewicz et al., we achieve a 33% improvement in assembly size and a 20% improvement in the NG50, demonstrating the power of automatic parameter selection.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Availability and implementation<\/jats:title>\n                    <jats:p>Our tool RecoverY is freely available at https:\/\/github.com\/makovalab-psu\/RecoverY.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Supplementary information<\/jats:title>\n                    <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btx771","type":"journal-article","created":{"date-parts":[[2017,11,27]],"date-time":"2017-11-27T15:22:36Z","timestamp":1511796156000},"page":"1125-1131","source":"Crossref","is-referenced-by-count":15,"title":["RecoverY:\n                    <i>k<\/i>\n                    -mer-based read classification for Y-chromosome-specific sequencing and assembly"],"prefix":"10.1093","volume":"34","author":[{"given":"Samarth","family":"Rangavittal","sequence":"first","affiliation":[{"name":"Department of Biology, Pennsylvania State University, University Park, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Robert S","family":"Harris","sequence":"additional","affiliation":[{"name":"Department of Biology, Pennsylvania State University, University Park, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Monika","family":"Cechova","sequence":"additional","affiliation":[{"name":"Department of Biology, Pennsylvania State University, University Park, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marta","family":"Tomaszkiewicz","sequence":"additional","affiliation":[{"name":"Department of Biology, Pennsylvania State University, University Park, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rayan","family":"Chikhi","sequence":"additional","affiliation":[{"name":"CNRS, CRIStAL, Villeneuve d\u2019Ascq, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kateryna D","family":"Makova","sequence":"additional","affiliation":[{"name":"Department of Biology, Pennsylvania State University, University Park, PA, USA"},{"name":"The Center for Computational Biology and Bioinformatics, Pennsylvania State University, University Park, PA, USA"},{"name":"The Center for Medical Genomics, Pennsylvania State University, University Park, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Paul","family":"Medvedev","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Pennsylvania State University, University Park, PA, USA"},{"name":"The Center for Computational Biology and Bioinformatics, Pennsylvania State University, University Park, PA, USA"},{"name":"Department of Biochemistry and Molecular Biology, Pennsylvania State University, University Park, PA, USA"},{"name":"The Center for Medical Genomics, Pennsylvania State University, University Park, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2017,11,28]]},"reference":[{"key":"2023012712575560400_btx771-B1","doi-asserted-by":"crossref","first-page":"455","DOI":"10.1089\/cmb.2012.0021","article-title":"SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing","volume":"19","author":"Bankevich","year":"2012","journal-title":"J. Comput. Biol"},{"key":"2023012712575560400_btx771-B2","doi-asserted-by":"crossref","first-page":"1894","DOI":"10.1101\/gr.156034.113","article-title":"Efficient identification of Y chromosome sequences in the human and Drosophila genomes","volume":"23","author":"Carvalho","year":"2013","journal-title":"Genome Res"},{"key":"2023012712575560400_btx771-B3","doi-asserted-by":"crossref","first-page":"238.","DOI":"10.1186\/1471-2105-13-238","article-title":"Mapping single molecule sequencing reads using basic local alignment with successive refinement (BLASR): application and theory","volume":"13","author":"Chaisson","year":"2012","journal-title":"BMC Bioinformatics"},{"key":"2023012712575560400_btx771-B4","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1093\/bioinformatics\/btt310","article-title":"Informed and automated k-mer size selection for genome assembly","volume":"30","author":"Chikhi","year":"2014","journal-title":"Bioinformatics"},{"key":"2023012712575560400_btx771-B5","doi-asserted-by":"crossref","first-page":"22.","DOI":"10.1186\/1748-7188-8-22","article-title":"Space-efficient and exact de Bruijn graph representation based on a Bloom filter","volume":"8","author":"Chikhi","year":"2013","journal-title":"Algorithms Mol. Biol"},{"key":"2023012712575560400_btx771-B6","author":"Crusoe","year":"2015"},{"key":"2023012712575560400_btx771-B7","doi-asserted-by":"crossref","first-page":"397","DOI":"10.1007\/s10142-012-0293-0","article-title":"Chromosomes in the flow to simplify genome analysis","volume":"12","author":"Dole\u017eel","year":"2012","journal-title":"Funct. Integr. Genomics"},{"key":"2023012712575560400_btx771-B9","doi-asserted-by":"crossref","first-page":"134","DOI":"10.1007\/s00239-008-9189-y","article-title":"Evolution of X-degenerate Y chromosome genes in greater apes: conservation of gene content in human and gorilla, but not chimpanzee","volume":"68","author":"Goto","year":"2009","journal-title":"J. Mol. Evol"},{"key":"2023012712575560400_btx771-B10","doi-asserted-by":"crossref","first-page":"1072","DOI":"10.1093\/bioinformatics\/btt086","article-title":"QUAST: quality assessment tool for genome assemblies","volume":"29","author":"Gurevich","year":"2013","journal-title":"Bioinformatics"},{"key":"2023012712575560400_btx771-B11","doi-asserted-by":"crossref","first-page":"273.","DOI":"10.1186\/1471-2164-14-273","article-title":"Six novel Y chromosome genes in Anopheles mosquitoes discovered by independently sequencing males and females","volume":"14","author":"Hall","year":"2013","journal-title":"BMC Genomics"},{"key":"2023012712575560400_btx771-B12","doi-asserted-by":"crossref","first-page":"536","DOI":"10.1038\/nature08700","article-title":"Chimpanzee and human Y chromosomes are remarkably divergent in structure and gene content","volume":"463","author":"Hughes","year":"2010","journal-title":"Nature"},{"key":"2023012712575560400_btx771-B13","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1038\/nature10843","article-title":"Strict evolutionary conservation followed rapid gene loss on human and rhesus Y chromosomes","volume":"483","author":"Hughes","year":"2012","journal-title":"Nature"},{"key":"2023012712575560400_btx771-B14","doi-asserted-by":"crossref","first-page":"2759","DOI":"10.1093\/bioinformatics\/btx304","article-title":"KMC 3: counting and manipulating k-mer statistics","volume":"33","author":"Kokot","year":"2017","journal-title":"Bioinformatics"},{"key":"2023012712575560400_btx771-B15","author":"Li"},{"key":"2023012712575560400_btx771-B16","doi-asserted-by":"crossref","first-page":"1754","DOI":"10.1093\/bioinformatics\/btp324","article-title":"Fast and accurate short read alignment with Burrows\u2013Wheeler transform","volume":"25","author":"Li","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012712575560400_btx771-B17","doi-asserted-by":"crossref","first-page":"2078","DOI":"10.1093\/bioinformatics\/btp352","article-title":"The sequence alignment\/map format and SAMtools","volume":"25","author":"Li","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012712575560400_btx771-B19","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1186\/2047-217X-1-18","article-title":"SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler","volume":"1","author":"Luo","year":"2012","journal-title":"GigaScience"},{"key":"2023012712575560400_btx771-B20","doi-asserted-by":"crossref","first-page":"764","DOI":"10.1093\/bioinformatics\/btr011","article-title":"A fast, lock-free approach for efficient parallel counting of occurrences of k-mers","volume":"27","author":"Marcais","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012712575560400_btx771-B21","doi-asserted-by":"crossref","first-page":"333.","DOI":"10.1186\/1471-2105-12-333","article-title":"Efficient counting of k -mers in DNA sequences using a bloom filter","volume":"12","author":"Melsted","year":"2011","journal-title":"BMC Bioinformatics"},{"key":"2023012712575560400_btx771-B22","doi-asserted-by":"crossref","first-page":"652","DOI":"10.1093\/bioinformatics\/btt020","article-title":"DSK: k-mer counting with very low memory usage","volume":"29","author":"Rizk","year":"2013","journal-title":"Bioinformatics"},{"key":"2023012712575560400_btx771-B23","doi-asserted-by":"crossref","first-page":"256","DOI":"10.1006\/geno.2000.6260","article-title":"Four DAZ genes in two clusters found in the AZFc region of the human Y chromosome","volume":"67","author":"Saxena","year":"2000","journal-title":"Genomics"},{"key":"2023012712575560400_btx771-B24","doi-asserted-by":"crossref","first-page":"825","DOI":"10.1038\/nature01722","article-title":"The male-specific region of the human Y chromosome is a mosaic of discrete sequence classes","volume":"423","author":"Skaletsky","year":"2003","journal-title":"Nature"},{"key":"2023012712575560400_btx771-B25","doi-asserted-by":"crossref","first-page":"130","DOI":"10.1101\/gr.188839.114","article-title":"The pig X and Y Chromosomes: structure, sequence, and evolution","volume":"26","author":"Skinner","year":"2016","journal-title":"Genome Res"},{"key":"2023012712575560400_btx771-B26","doi-asserted-by":"crossref","first-page":"800","DOI":"10.1016\/j.cell.2014.09.052","article-title":"Sequencing the Mouse Y chromosome reveals convergent gene acquisition and amplification on both sex chromosomes","volume":"159","author":"Soh","year":"2014","journal-title":"Cell"},{"key":"2023012712575560400_btx771-B27","doi-asserted-by":"crossref","first-page":"530","DOI":"10.1101\/gr.199448.115","article-title":"A time- and cost-effective strategy to sequence mammalian Y Chromosomes: an application to the de novo assembly of gorilla Y","volume":"26","author":"Tomaszkiewicz","year":"2016","journal-title":"Genome Res"},{"key":"2023012712575560400_btx771-B28","author":"Tomaszkiewicz","year":"2017"},{"key":"2023012712575560400_btx771-B29","doi-asserted-by":"crossref","first-page":"1350","DOI":"10.1038\/ng.3121","article-title":"Comprehensive variation discovery in single human genomes","volume":"46","author":"Weisenfeld","year":"2014","journal-title":"Nat. Genet"},{"key":"2023012712575560400_btx771-B30","doi-asserted-by":"crossref","first-page":"67","DOI":"10.2174\/138920207780076929","article-title":"The development of chromosome microdissection and microcloning technique and its applications in genomic research","volume":"8","author":"Zhou","year":"2007","journal-title":"Curr. Genomics"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/34\/7\/1125\/48914854\/bioinformatics_34_7_1125.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/34\/7\/1125\/48914854\/bioinformatics_34_7_1125.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T08:46:24Z","timestamp":1674809184000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/34\/7\/1125\/4670683"}},"subtitle":[],"editor":[{"given":"Inanc","family":"Birol","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2017,11,28]]},"references-count":28,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2018,4,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btx771","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/148114","asserted-by":"object"}]},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2018,4,1]]},"published":{"date-parts":[[2017,11,28]]}}}