{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,26]],"date-time":"2026-08-26T16:56:09Z","timestamp":1787763369406,"version":"build-2784847793"},"reference-count":16,"publisher":"Oxford University Press (OUP)","issue":"21","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2014,11,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: The increasing interest in rare genetic variants and epistatic genetic effects on complex phenotypic traits is currently pushing genome-wide association study design towards datasets of increasing size, both in the number of studied subjects and in the number of genotyped single nucleotide polymorphisms (SNPs). This, in turn, is leading to a compelling need for new methods for compression and fast retrieval of SNP data.<\/jats:p>\n               <jats:p>Results: We present a novel algorithm and file format for compressing and retrieving SNP data, specifically designed for large-scale association studies. Our algorithm is based on two main ideas: (i) compress linkage disequilibrium blocks in terms of differences with a reference SNP and (ii) compress reference SNPs exploiting information on their call rate and minor allele frequency. Tested on two SNP datasets and compared with several state-of-the-art software tools, our compression algorithm is shown to be competitive in terms of compression rate and to outperform all tools in terms of time to load compressed data.<\/jats:p>\n               <jats:p>Availability and implementation: Our compression and decompression algorithms are implemented in a C++ library, are released under the GNU General Public License and are freely downloadable from http:\/\/www.dei.unipd.it\/~sambofra\/snpack.html .<\/jats:p>\n               <jats:p>Contact: \u00a0sambofra@dei.unipd.it or cobelli@dei.unipd.it .<\/jats:p>","DOI":"10.1093\/bioinformatics\/btu495","type":"journal-article","created":{"date-parts":[[2014,7,27]],"date-time":"2014-07-27T00:29:24Z","timestamp":1406420964000},"page":"3078-3085","source":"Crossref","is-referenced-by-count":11,"title":["Compression and fast retrieval of SNP data"],"prefix":"10.1093","volume":"30","author":[{"given":"Francesco","family":"Sambo","sequence":"first","affiliation":[{"name":"Department of Information Engineering, University of Padova, via Gradenigo 6\/a, 35131 Padova, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Barbara","family":"Di Camillo","sequence":"additional","affiliation":[{"name":"Department of Information Engineering, University of Padova, via Gradenigo 6\/a, 35131 Padova, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gianna","family":"Toffolo","sequence":"additional","affiliation":[{"name":"Department of Information Engineering, University of Padova, via Gradenigo 6\/a, 35131 Padova, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Claudio","family":"Cobelli","sequence":"additional","affiliation":[{"name":"Department of Information Engineering, University of Padova, via Gradenigo 6\/a, 35131 Padova, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2014,7,26]]},"reference":[{"key":"2023012711572066300_btu495-B1","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1038\/nature11632","article-title":"An integrated map of genetic variation from 1,092 human genomes","volume":"491","author":"1000 Genomes Project Consortium","year":"2012","journal-title":"Nature"},{"key":"2023012711572066300_btu495-B2","doi-asserted-by":"crossref","first-page":"1731","DOI":"10.1093\/bioinformatics\/btp319","article-title":"Data structures and compression algorithms for genomic sequence data","volume":"25","author":"Brandon","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012711572066300_btu495-B3","doi-asserted-by":"crossref","first-page":"274","DOI":"10.1093\/bioinformatics\/btn582","article-title":"Human genomes as email attachments","volume":"25","author":"Christley","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012711572066300_btu495-B4","doi-asserted-by":"crossref","first-page":"2156","DOI":"10.1093\/bioinformatics\/btr330","article-title":"The variant call format and VCFtools","volume":"27","author":"Danecek","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012711572066300_btu495-B5","doi-asserted-by":"crossref","first-page":"2572","DOI":"10.1093\/bioinformatics\/btt460","article-title":"Genome compression: a novel approach for large collections","volume":"29","author":"Deorowicz","year":"2013","journal-title":"Bioinformatics"},{"key":"2023012711572066300_btu495-B6","doi-asserted-by":"crossref","first-page":"1266","DOI":"10.1093\/bioinformatics\/btu014","article-title":"Efficient haplotype matching and storage using the positional Burrows-Wheeler transform (PBWT)","volume":"30","author":"Durbin","year":"2014","journal-title":"Bioinformatics"},{"key":"2023012711572066300_btu495-B7","doi-asserted-by":"crossref","first-page":"135","DOI":"10.1038\/nrg3118","article-title":"Rare and common variants: twenty arguments","volume":"13","author":"Gibson","year":"2012","journal-title":"Nat. Rev. Genet."},{"key":"2023012711572066300_btu495-B8","doi-asserted-by":"crossref","first-page":"955","DOI":"10.1038\/ng.2354","article-title":"Fast and accurate genotype imputation in genome-wide association studies through pre-phasing","volume":"44","author":"Howie","year":"2012","journal-title":"Nat. Genet."},{"key":"2023012711572066300_btu495-B9","doi-asserted-by":"crossref","first-page":"832","DOI":"10.1038\/nature09410","article-title":"Hundreds of variants clustered in genomic loci and biological pathways affect human height","volume":"467","author":"Lango Allen","year":"2010","journal-title":"Nature"},{"key":"2023012711572066300_btu495-B10","doi-asserted-by":"crossref","first-page":"549","DOI":"10.1038\/nrg3523","article-title":"Bringing genome-wide association findings into clinical use","volume":"14","author":"Manolio","year":"2013","journal-title":"Nat. Rev. Genet."},{"key":"2023012711572066300_btu495-B11","doi-asserted-by":"crossref","first-page":"747","DOI":"10.1038\/nature08494","article-title":"Finding the missing heritability of complex diseases","volume":"461","author":"Manolio","year":"2009","journal-title":"Nature"},{"key":"2023012711572066300_btu495-B12","doi-asserted-by":"crossref","first-page":"559","DOI":"10.1086\/519795","article-title":"PLINK: a tool set for whole-genome association and population-based linkage analyses","volume":"81","author":"Purcell","year":"2007","journal-title":"Am. J. Hum. Genet."},{"key":"2023012711572066300_btu495-B13","doi-asserted-by":"crossref","first-page":"100","DOI":"10.1186\/1471-2105-13-100","article-title":"Handling the data management needs of high-throughput sequencing data: Speedgene, a compression algorithm for the efficient storage of genetic data","volume":"13","author":"Qiao","year":"2012","journal-title":"BMC Bioinformatics"},{"key":"2023012711572066300_btu495-B14","doi-asserted-by":"crossref","first-page":"587","DOI":"10.1038\/nrg1123","article-title":"Haplotype blocks and linkage disequilibrium in the human genome","volume":"4","author":"Wall","year":"2003","journal-title":"Nat. Rev. Genet."},{"key":"2023012711572066300_btu495-B15","doi-asserted-by":"crossref","first-page":"e45","DOI":"10.1093\/nar\/gkr009","article-title":"A novel compression tool for efficient storage of genome resequencing data","volume":"39","author":"Wang","year":"2011","journal-title":"Nucleic Acids Res."},{"key":"2023012711572066300_btu495-B16","doi-asserted-by":"crossref","first-page":"661","DOI":"10.1038\/nature05911","article-title":"Genome-wide association study of 14,000 cases of seven common diseases and 3,000 shared controls","volume":"447","author":"Wellcome Trust Case Control Consortium","year":"2007","journal-title":"Nature"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/21\/3078\/48930760\/bioinformatics_30_21_3078.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/21\/3078\/48930760\/bioinformatics_30_21_3078.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T12:50:37Z","timestamp":1674823837000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/30\/21\/3078\/2422242"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,7,26]]},"references-count":16,"journal-issue":{"issue":"21","published-print":{"date-parts":[[2014,11,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btu495","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2014,11,1]]},"published":{"date-parts":[[2014,7,26]]}}}