{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,12]],"date-time":"2026-03-12T00:10:20Z","timestamp":1773274220077,"version":"3.50.1"},"reference-count":25,"publisher":"Oxford University Press (OUP)","issue":"20","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2016,10,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: The exponential reduction in cost of genome sequencing has resulted in a rapid growth of genomic data. Most of the entropy of short read data lies not in the sequence of read bases themselves but in their Quality Scores\u2014the confidence measurement that each base has been sequenced correctly. Lossless compression methods are now close to their theoretical limits and hence there is a need for lossy methods that further reduce the complexity of these data without impacting downstream analyses.<\/jats:p>\n               <jats:p>Results: We here propose GeneCodeq, a Bayesian method inspired by coding theory for adjusting quality scores to improve the compressibility of quality scores without adversely impacting genotyping accuracy. Our model leverages a corpus of k-mers to reduce the entropy of the quality scores and thereby the compressibility of these data (in FASTQ or SAM\/BAM\/CRAM files), resulting in compression ratios that significantly exceeds those of other methods. Our approach can also be combined with existing lossy compression schemes to further reduce entropy and allows the user to specify a reference panel of expected sequence variations to improve the model accuracy. In addition to extensive empirical evaluation, we also derive novel theoretical insights that explain the empirical performance and pitfalls of corpus-based quality score compression schemes in general. Finally, we show that as a positive side effect of compression, the model can lead to improved genotyping accuracy.<\/jats:p>\n               <jats:p>Availability and implementation: \u00a0GeneCodeq is available at: github.com\/genecodeq\/eval<\/jats:p>\n               <jats:p>Contact: \u00a0dan@petagene.com<\/jats:p>\n               <jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btw385","type":"journal-article","created":{"date-parts":[[2016,6,29]],"date-time":"2016-06-29T14:57:42Z","timestamp":1467212262000},"page":"3124-3132","source":"Crossref","is-referenced-by-count":23,"title":["GeneCodeq: quality score compression and improved genotyping using a Bayesian framework"],"prefix":"10.1093","volume":"32","author":[{"given":"Daniel L.","family":"Greenfield","sequence":"first","affiliation":[{"name":"1 PetaGene, Ideaspace, 3 Charles Babbage Rd, Cambridge CB3 0GT, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Oliver","family":"Stegle","sequence":"additional","affiliation":[{"name":"2 European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SQ, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alban","family":"Rrustemi","sequence":"additional","affiliation":[{"name":"1 PetaGene, Ideaspace, 3 Charles Babbage Rd, Cambridge CB3 0GT, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2016,6,26]]},"reference":[{"key":"2023020113472054500_btw385-B1","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1038\/nature11632","article-title":"An integrated map of genetic variation from 1,092 human genomes","volume":"491","author":"1000 Genomes Project Consortium","year":"2012","journal-title":"Nature"},{"key":"2023020113472054500_btw385-B2","volume-title":"Interscience Tracts in Pure and Applied Mathematics","author":"Ash","year":"1965"},{"key":"2023020113472054500_btw385-B3","doi-asserted-by":"crossref","first-page":"495","DOI":"10.1038\/nmeth0710-495","article-title":"Next-generation sequencing: adjusting to data overload","volume":"7","author":"Baker","year":"2010","journal-title":"Nat. Methods"},{"key":"2023020113472054500_btw385-B4","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1007\/BF03025254","article-title":"At the dawn of the theory of codes","volume":"15","author":"Barg","year":"1993","journal-title":"Math. Intell"},{"key":"2023020113472054500_btw385-B5","doi-asserted-by":"crossref","first-page":"288.","DOI":"10.1186\/s12859-015-0709-7","article-title":"Reference-free compression of high throughput sequencing data with a probabilistic de bruijn graph","volume":"16","author":"Benoit","year":"2015","journal-title":"BMC Bioinformatics"},{"key":"2023020113472054500_btw385-B6","doi-asserted-by":"crossref","first-page":"499","DOI":"10.1097\/GIM.0b013e318220aaba","article-title":"Deploying whole genome sequencing in clinical practice and public health: meeting the challenge one bin at a time","volume":"13","author":"Berg","year":"2011","journal-title":"Genet. Med"},{"key":"2023020113472054500_btw385-B7","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1038\/nrg3433","article-title":"Computational solutions for omics data","volume":"14","author":"Berger","year":"2013","journal-title":"Nat. Rev. Genet"},{"key":"2023020113472054500_btw385-B8","doi-asserted-by":"crossref","first-page":"e59190.","DOI":"10.1371\/journal.pone.0059190","article-title":"Compression of FASTQ and SAM format sequencing data","volume":"8","author":"Bonfield","year":"2013","journal-title":"PLoS One"},{"key":"2023020113472054500_btw385-B9","doi-asserted-by":"crossref","first-page":"2130","DOI":"10.1093\/bioinformatics\/btu183","article-title":"Lossy compression of quality scores in genomic data","volume":"30","author":"C\u00e1novas","year":"2014","journal-title":"Bioinformatics"},{"key":"2023020113472054500_btw385-B10","doi-asserted-by":"crossref","first-page":"1415","DOI":"10.1093\/bioinformatics\/bts173","article-title":"Large-scale compression of genomic sequence databases with the burrows\u2013wheeler transform","volume":"28","author":"Cox","year":"2012","journal-title":"Bioinformatics"},{"key":"2023020113472054500_btw385-B11","doi-asserted-by":"crossref","first-page":"1677","DOI":"10.1093\/bioinformatics\/bts256","article-title":"Onlinecall: fast online parameter estimation and base calling for illumina\u2019s next-generation sequencing","volume":"28","author":"Das","year":"2012","journal-title":"Bioinformatics"},{"key":"2023020113472054500_btw385-B12","doi-asserted-by":"crossref","first-page":"491","DOI":"10.1038\/ng.806","article-title":"A framework for variation discovery and genotyping using next-generation DNA sequencing data","volume":"43","author":"DePristo","year":"2011","journal-title":"Nat. Genet"},{"key":"2023020113472054500_btw385-B13","doi-asserted-by":"crossref","first-page":"186","DOI":"10.1101\/gr.8.3.186","article-title":"Base-calling of automated sequencer traces using phred. II. Error probabilities","volume":"8","author":"Ewing","year":"1998","journal-title":"Genome Res"},{"key":"2023020113472054500_btw385-B14","doi-asserted-by":"crossref","first-page":"1741","DOI":"10.1093\/bioinformatics\/btr295","article-title":"Bioinformatics challenges for personalized medicine","volume":"27","author":"Fernald","year":"2011","journal-title":"Bioinformatics"},{"key":"2023020113472054500_btw385-B15","doi-asserted-by":"crossref","first-page":"734","DOI":"10.1101\/gr.114819.110","article-title":"Efficient storage of high throughput DNA sequencing data using reference-based compression","volume":"21","author":"Fritz","year":"2011","journal-title":"Genome Res"},{"key":"2023020113472054500_btw385-B16","doi-asserted-by":"crossref","first-page":"1389","DOI":"10.1093\/bioinformatics\/btu844","article-title":"Disk-based compression of data from genome sequencing","volume":"31","author":"Grabowski","year":"2015","journal-title":"Bioinformatics"},{"key":"2023020113472054500_btw385-B17","author":"Illumina","year":"2011"},{"key":"2023020113472054500_btw385-B18","author":"Illumina","year":"2014"},{"key":"2023020113472054500_btw385-B19","doi-asserted-by":"crossref","first-page":"2987","DOI":"10.1093\/bioinformatics\/btr509","article-title":"A statistical framework for SNP calling, mutation discovery, association mapping and population genetical parameter estimation from sequencing data","volume":"27","author":"Li","year":"2011","journal-title":"Bioinformatics"},{"key":"2023020113472054500_btw385-B20","doi-asserted-by":"crossref","first-page":"3122","DOI":"10.1093\/bioinformatics\/btv330","article-title":"QVZ: lossy compression of quality values","volume":"31","author":"Malysa","year":"2015","journal-title":"Bioinformatics"},{"key":"2023020113472054500_btw385-B21","doi-asserted-by":"crossref","first-page":"187.","DOI":"10.1186\/1471-2105-14-187","article-title":"QualComp: a new lossy compressor for quality scores based on rate distortion theory","volume":"14","author":"Ochoa","year":"2013","journal-title":"BMC Bioinformatics"},{"key":"2023020113472054500_btw385-B22","doi-asserted-by":"crossref","first-page":"e1002195.","DOI":"10.1371\/journal.pbio.1002195","article-title":"Big data: astronomical or genomical?","volume":"13","author":"Stephens","year":"2015","journal-title":"PLoS Biol"},{"key":"2023020113472054500_btw385-B25","author":"Wetterstrand","year":"2015"},{"key":"2023020113472054500_btw385-B23","doi-asserted-by":"crossref","first-page":"385","DOI":"10.1007\/978-3-319-05269-4_31","volume-title":"Research in Computational Molecular Biology","author":"Yu","year":"2014"},{"key":"2023020113472054500_btw385-B24","doi-asserted-by":"crossref","first-page":"240","DOI":"10.1038\/nbt.3170","article-title":"Quality score compression improves genotyping accuracy","volume":"33","author":"Yu","year":"2015","journal-title":"Nat. Biotechnol"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/32\/20\/3124\/49021777\/bioinformatics_32_20_3124.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/32\/20\/3124\/49021777\/bioinformatics_32_20_3124.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,1]],"date-time":"2023-02-01T23:51:51Z","timestamp":1675295511000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/32\/20\/3124\/2196578"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,6,26]]},"references-count":25,"journal-issue":{"issue":"20","published-print":{"date-parts":[[2016,10,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btw385","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2016,10,15]]},"published":{"date-parts":[[2016,6,26]]}}}