{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,26]],"date-time":"2026-02-26T03:53:38Z","timestamp":1772078018482,"version":"3.50.1"},"reference-count":22,"publisher":"Oxford University Press (OUP)","issue":"5","funder":[{"name":"Stanford Graduate Fellowships Program in Science and Engineering"},{"DOI":"10.13039\/501100003086","name":"Basque Government","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100003086","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"NSF","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2015,3,1]]},"abstract":"<jats:p>Motivation: With the release of the latest next-generation sequencing (NGS) machine, the HiSeq X by Illumina, the cost of sequencing a Human has dropped to a mere $4000. Thus we are approaching a milestone in the sequencing history, known as the $1000 genome era, where the sequencing of individuals is affordable, opening the doors to effective personalized medicine. Massive generation of genomic data, including assembled genomes, is expected in the following years. There is crucial need for compression of genomes guaranteed of performing well simultaneously on different species, from simple bacteria to humans, which will ease their transmission, dissemination and analysis. Further, most of the new genomes to be compressed will correspond to individuals of a species from which a reference already exists on the database. Thus, it is natural to propose compression schemes that assume and exploit the availability of such references.<\/jats:p>\n               <jats:p>Results: We propose iDoComp, a compressor of assembled genomes presented in FASTA format that compresses an individual genome using a reference genome for both the compression and the decompression. In terms of compression efficiency, iDoComp outperforms previously proposed algorithms in most of the studied cases, with comparable or better running time. For example, we observe compression gains of up to 60% in several cases, including H.sapiens data, when comparing with the best compression performance among the previously proposed algorithms.<\/jats:p>\n               <jats:p>Availability: iDoComp is written in C and can be downloaded from: http:\/\/www.stanford.edu\/~iochoa\/iDoComp.html (We also provide a full explanation on how to run the program and an example with all the necessary files to run it.).<\/jats:p>\n               <jats:p>Contact: \u00a0iochoa@stanford.edu<\/jats:p>\n               <jats:p>Supplementary information: Supplementary Data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btu698","type":"journal-article","created":{"date-parts":[[2014,10,25]],"date-time":"2014-10-25T03:54:31Z","timestamp":1414209271000},"page":"626-633","source":"Crossref","is-referenced-by-count":41,"title":["iDoComp: a compression scheme for assembled genomes"],"prefix":"10.1093","volume":"31","author":[{"given":"Idoia","family":"Ochoa","sequence":"first","affiliation":[{"name":"1Department of Electrical Engineering, Stanford University, 350 Serra Mall, Stanford, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mikel","family":"Hernaez","sequence":"additional","affiliation":[{"name":"1Department of Electrical Engineering, Stanford University, 350 Serra Mall, Stanford, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tsachy","family":"Weissman","sequence":"additional","affiliation":[{"name":"1Department of Electrical Engineering, Stanford University, 350 Serra Mall, Stanford, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2014,10,24]]},"reference":[{"key":"2023020116162454100_btu698-B1","doi-asserted-by":"crossref","first-page":"1731","DOI":"10.1093\/bioinformatics\/btp319","article-title":"Data structures and compression algorithms for genomic sequence data","volume":"14","author":"Brandon","year":"2009","journal-title":"Bioinformatics"},{"key":"2023020116162454100_btu698-B2","first-page":"Utah","article-title":"A simple statistical algorithm for biological sequence compression","volume-title":"IEEE Data Compression Conference (DCC\u201907)","author":"Cao","year":"2007"},{"key":"2023020116162454100_btu698-B3","first-page":"51","article-title":"DNACompress: fast and effective DNA sequence compression","volume":"10","author":"Chen","year":"2002","journal-title":"Bioinformatics"},{"key":"2023020116162454100_btu698-B4","doi-asserted-by":"crossref","DOI":"10.1109\/ITW.2012.6404708","article-title":"Reference based genome compression","author":"Chern","year":"2012"},{"key":"2023020116162454100_btu698-B5","doi-asserted-by":"crossref","first-page":"274","DOI":"10.1093\/bioinformatics\/btn582","article-title":"Human genomes as email attachments","volume":"2","author":"Christley","year":"2009","journal-title":"Bioinformatics"},{"key":"2023020116162454100_btu698-B6","doi-asserted-by":"crossref","first-page":"2156","DOI":"10.1093\/bioinformatics\/btr330","article-title":"The variant call format and VCFtools","volume":"27","author":"Danecek","year":"2011","journal-title":"Bioinformatics"},{"key":"2023020116162454100_btu698-B7","doi-asserted-by":"crossref","first-page":"2979","DOI":"10.1093\/bioinformatics\/btr505","article-title":"Robust relative compression of genomes with random access","volume":"21","author":"Deorowicz","year":"2011","journal-title":"Bioinformatics"},{"key":"2023020116162454100_btu698-B8","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1186\/1748-7188-8-25","article-title":"Data compression for sequencing data","volume":"8","author":"Deorowicz","year":"2013","journal-title":"Algorithms Mol. Biol."},{"key":"2023020116162454100_btu698-B9","doi-asserted-by":"crossref","first-page":"2572","DOI":"10.1093\/bioinformatics\/btt460","article-title":"Genome compression: a novel approach for large collections","volume":"29","author":"Deorowicz","year":"2013","journal-title":"Bioinformatics"},{"key":"2023020116162454100_btu698-B10","doi-asserted-by":"crossref","first-page":"875","DOI":"10.1016\/0306-4573(94)90014-0","article-title":"A new challenge for compression Algorithms: genetic sequences","volume":"6","author":"Grumbach","year":"1994","journal-title":"Inf. Process Manag."},{"key":"2023020116162454100_btu698-B11","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511574931","author":"Gusfield","year":"1997","journal-title":"Algorithms on strings, trees and sequences: computer science and computational biology"},{"key":"2023020116162454100_btu698-B12","doi-asserted-by":"crossref","first-page":"860","DOI":"10.1038\/35057062","article-title":"Initial sequencing and analysis of the human genome","volume":"409","author":"Int. Human Genome Sequencing Consortium","year":"2001","journal-title":"Nature"},{"key":"2023020116162454100_btu698-B13","first-page":"201","article-title":"Relative lempel-ziv compression of genomes for large-scale storage and retrieval","volume":"6393","author":"Kuruppu","year":"2010","journal-title":"SPIRE 2010. Lecture Notes Comput Sci."},{"key":"2023020116162454100_btu698-B14","first-page":"137","article-title":"Iterative dictionary construction for compression of large DNA data sets","volume":"1","author":"Kuruppu","year":"2010","journal-title":"IEEE\/AMC Trans Comput Biol Bioinform"},{"key":"2023020116162454100_btu698-B15","article-title":"Optimized relative lempel-ziv compression of genomes","volume-title":"34th Australasian Computer Science Conference (ACSC 2011)","author":"Kuruppu","year":"2011"},{"key":"2023020116162454100_btu698-B16","doi-asserted-by":"crossref","first-page":"2199","DOI":"10.1093\/bioinformatics\/btt362","article-title":"The human genome contracts again","volume":"29","author":"Pavlichin","year":"2013","journal-title":"Bioinformatics"},{"key":"2023020116162454100_btu698-B17","doi-asserted-by":"crossref","first-page":"666","DOI":"10.1126\/science.331.6018.666","article-title":"Will computers crash genomics?","volume":"331","author":"Pennisi","year":"2011","journal-title":"Science"},{"key":"2023020116162454100_btu698-B18","doi-asserted-by":"crossref","first-page":"e27","DOI":"10.1093\/nar\/gkr1124","article-title":"GReEn: a tool for efficient compression of genome resequencing data","volume":"40","author":"Pinho","year":"2012","journal-title":"Nucleic Acid Res."},{"key":"2023020116162454100_btu698-B19","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/1748-7188-7-30","article-title":"Adaptive efficient compression of genomes","volume":"7","author":"Wandelt","year":"2012","journal-title":"Algorithms Mol Biol."},{"key":"2023020116162454100_btu698-B20","first-page":"1275","article-title":"FRESCO: referential compression of highly similar sequences","volume-title":"IEEE\/ACM Trans Comput Biol Bioinform (TCBB)","author":"Wandelt","year":"2013"},{"key":"2023020116162454100_btu698-B21","doi-asserted-by":"crossref","first-page":"e45","DOI":"10.1093\/nar\/gkr009","article-title":"A novel compression tool for efficient storage of genome resequencing data","volume":"39","author":"Wang","year":"2011","journal-title":"Nucleic Acid Res."},{"key":"2023020116162454100_btu698-B22","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1093\/bib\/bbt087","article-title":"High-throughput DNA sequence data compression","volume":"16","author":"Zhu","year":"2015","journal-title":"Brief Bioinformatics"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/31\/5\/626\/49011202\/bioinformatics_31_5_626.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/31\/5\/626\/49011202\/bioinformatics_31_5_626.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,2]],"date-time":"2023-02-02T00:28:38Z","timestamp":1675297718000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/31\/5\/626\/2748151"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,10,24]]},"references-count":22,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2015,3,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btu698","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published":{"date-parts":[[2014,10,24]]}}}