{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,12]],"date-time":"2026-03-12T01:28:39Z","timestamp":1773278919538,"version":"3.50.1"},"reference-count":31,"publisher":"Oxford University Press (OUP)","issue":"12","license":[{"start":{"date-parts":[[2018,11,8]],"date-time":"2018-11-08T00:00:00Z","timestamp":1541635200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"funder":[{"DOI":"10.13039\/501100000923","name":"Australia Research Council","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100000923","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100000163","name":"ARC","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000163","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Discovery Project","award":["DP180100120"],"award-info":[{"award-number":["DP180100120"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["11871061"],"award-info":[{"award-number":["11871061"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Collaborative research project for Overseas Scholars"},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61828203"],"award-info":[{"award-number":["61828203"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2019,6,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>Advanced high-throughput sequencing technologies have produced massive amount of reads data, and algorithms have been specially designed to contract the size of these datasets for efficient storage and transmission. Reordering reads with regard to their positions in de novo assembled contigs or in explicit reference sequences has been proven to be one of the most effective reads compression approach. As there is usually no good prior knowledge about the reference sequence, current focus is on the novel construction of de novo assembled contigs.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>We introduce a new de novo compression algorithm named minicom. This algorithm uses large k-minimizers to index the reads and subgroup those that have the same minimizer. Within each subgroup, a contig is constructed. Then some pairs of the contigs derived from the subgroups are merged into longer contigs according to a (w, k)-minimizer-indexed suffix\u2013prefix overlap similarity between two contigs. This merging process is repeated after the longer contigs are formed until no pair of contigs can be merged. We compare the performance of minicom with two reference-based methods and four de novo methods on 18 datasets (13 RNA-seq datasets and 5 whole genome sequencing datasets). In the compression of single-end reads, minicom obtained the smallest file size for 22 of 34 cases with significant improvement. In the compression of paired-end reads, minicom achieved 20\u201380% compression gain over the best state-of-the-art algorithm. Our method also achieved a 10% size reduction of compressed files in comparison with the best algorithm under the reads-order preserving mode. These excellent performances are mainly attributed to the exploit of the redundancy of the repetitive substrings in the long contigs.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation<\/jats:title>\n                  <jats:p>https:\/\/github.com\/yuansliu\/minicom<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information<\/jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/bty936","type":"journal-article","created":{"date-parts":[[2018,11,7]],"date-time":"2018-11-07T12:09:36Z","timestamp":1541592576000},"page":"2066-2074","source":"Crossref","is-referenced-by-count":33,"title":["Index suffix\u2013prefix overlaps by (<i>w<\/i>, <i>k<\/i>)-minimizer to generate long contigs for reads compression"],"prefix":"10.1093","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7680-3155","authenticated-orcid":false,"given":"Yuansheng","family":"Liu","sequence":"first","affiliation":[{"name":"Advanced Analytics Institute, Faculty of Engineering and IT, University of Technology Sydney, Ultimo, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zuguo","family":"Yu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Computing and Information Processing of Ministry of Education, Hunan Key Laboratory for Computation and Simulation in Science and Engineering, Xiangtan University, Hunan, China"},{"name":"School of Electrical Engineering and Computer Science, Queensland University of Technology, Brisbane, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4423-934X","authenticated-orcid":false,"given":"Marcel E","family":"Dinger","sequence":"additional","affiliation":[{"name":"Kinghorn Centre for Clinical Genomics, Garvan Institute of Medical Research, Sydney, NSW, Australia"},{"name":"St Vincent's Clinical School, University of New South Wales, Sydney, NSW, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1833-7413","authenticated-orcid":false,"given":"Jinyan","family":"Li","sequence":"additional","affiliation":[{"name":"Advanced Analytics Institute, Faculty of Engineering and IT, University of Technology Sydney, Ultimo, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2018,11,8]]},"reference":[{"key":"2023012810014816200_bty936-B1","doi-asserted-by":"crossref","first-page":"288.","DOI":"10.1186\/s12859-015-0709-7","article-title":"Reference-free compression of high throughput sequencing data with a probabilistic de Bruijn graph","volume":"16","author":"Benoit","year":"2015","journal-title":"BMC Bioinformatics"},{"key":"2023012810014816200_bty936-B2","doi-asserted-by":"crossref","first-page":"e59190.","DOI":"10.1371\/journal.pone.0059190","article-title":"Compression of FASTQ and SAM format sequencing data","volume":"8","author":"Bonfield","year":"2013","journal-title":"PLoS One"},{"key":"2023012810014816200_bty936-B3","doi-asserted-by":"crossref","first-page":"2130","DOI":"10.1093\/bioinformatics\/btu183","article-title":"Lossy compression of quality scores in genomic data","volume":"30","author":"C\u00e1novas","year":"2014","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B4","doi-asserted-by":"crossref","first-page":"558","DOI":"10.1093\/bioinformatics\/btx639","article-title":"Compression of genomic sequencing reads via hash-based reordering: algorithm and analysis","volume":"34","author":"Chandak","year":"2018","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B5","doi-asserted-by":"crossref","first-page":"1415","DOI":"10.1093\/bioinformatics\/bts173","article-title":"Large-scale compression of genomic sequence databases with the Burrows\u2013Wheeler transform","volume":"28","author":"Cox","year":"2012","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B6","doi-asserted-by":"crossref","first-page":"2979","DOI":"10.1093\/bioinformatics\/btr505","article-title":"Robust relative compression of genomes with random access","volume":"27","author":"Deorowicz","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B7","doi-asserted-by":"crossref","first-page":"25.","DOI":"10.1186\/1748-7188-8-25","article-title":"Data compression for sequencing data","volume":"8","author":"Deorowicz","year":"2013","journal-title":"Algorithms Mol. Biol"},{"key":"2023012810014816200_bty936-B8","doi-asserted-by":"crossref","first-page":"566","DOI":"10.1038\/s41467-017-02480-6","article-title":"Optimal compressed representation of high throughput sequence data via light assembly","volume":"9","author":"Ginart","year":"2018","journal-title":"Nat. Commun"},{"key":"2023012810014816200_bty936-B9","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1038\/nrg.2016.49","article-title":"Coming of age: ten years of next-generation sequencing technologies","volume":"17","author":"Goodwin","year":"2016","journal-title":"Nat. Rev. Genet"},{"key":"2023012810014816200_bty936-B10","doi-asserted-by":"crossref","first-page":"1389","DOI":"10.1093\/bioinformatics\/btu844","article-title":"Disk-based compression of data from genome sequencing","volume":"31","author":"Grabowski","year":"2015","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B11","doi-asserted-by":"crossref","first-page":"3124","DOI":"10.1093\/bioinformatics\/btw385","article-title":"GeneCodeq: quality score compression and improved genotyping using a Bayesian framework","volume":"32","author":"Greenfield","year":"2016","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B12","doi-asserted-by":"crossref","first-page":"3051","DOI":"10.1093\/bioinformatics\/bts593","article-title":"SCALCE: boosting sequence compression algorithms using locally consistent encoding","volume":"28","author":"Hach","year":"2012","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B13","doi-asserted-by":"crossref","first-page":"1082.","DOI":"10.1038\/nmeth.3133","article-title":"DeeZ: reference-based compression by local assembly","volume":"11","author":"Hach","year":"2014","journal-title":"Nat. Methods"},{"key":"2023012810014816200_bty936-B14","doi-asserted-by":"crossref","first-page":"e171","DOI":"10.1093\/nar\/gks754","article-title":"Compression of next-generation sequencing reads aided by highly efficient de novo assembly","volume":"40","author":"Jones","year":"2012","journal-title":"Nucleic Acids Res"},{"key":"2023012810014816200_bty936-B15","doi-asserted-by":"crossref","first-page":"1920","DOI":"10.1093\/bioinformatics\/btv071","article-title":"Reference-based compression of short-read sequences using path encoding","volume":"31","author":"Kingsford","year":"2015","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B16","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1093\/bib\/bbq016","article-title":"Challenges of sequencing human genomes","volume":"11","author":"Koboldt","year":"2010","journal-title":"Brief. Bioinform"},{"key":"2023012810014816200_bty936-B17","doi-asserted-by":"crossref","first-page":"2103","DOI":"10.1093\/bioinformatics\/btw152","article-title":"Minimap and miniasm: fast mapping and de novo assembly for noisy long sequences","volume":"32","author":"Li","year":"2016","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B18","doi-asserted-by":"crossref","first-page":"3364","DOI":"10.1093\/bioinformatics\/btx412","article-title":"High-speed and high-ratio referential genome compression","volume":"33","author":"Liu","year":"2017","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B19","doi-asserted-by":"crossref","first-page":"3122","DOI":"10.1093\/bioinformatics\/btv330","article-title":"QVZ: lossy compression of quality values","volume":"31","author":"Malysa","year":"2015","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B20","doi-asserted-by":"crossref","first-page":"i110","DOI":"10.1093\/bioinformatics\/btx235","article-title":"Improving the performance of minimizers and winnowing schemes","volume":"33","author":"Mar\u00e7ais","year":"2017","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B21","doi-asserted-by":"crossref","first-page":"1005","DOI":"10.1038\/nmeth.4037","article-title":"Comparison of high-throughput sequencing data compression tools","volume":"13","author":"Numanagi\u0107","year":"2016","journal-title":"Nat. Methods"},{"key":"2023012810014816200_bty936-B22","first-page":"183","article-title":"Effect of lossy compression of quality scores on variant calling","volume":"18","author":"Ochoa","year":"2016","journal-title":"Brief. Bioinform"},{"key":"2023012810014816200_bty936-B23","doi-asserted-by":"crossref","first-page":"2770","DOI":"10.1093\/bioinformatics\/btv248","article-title":"Data-dependent bucketing improves reference-free compression of sequencing reads","volume":"31","author":"Patro","year":"2015","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B24","doi-asserted-by":"crossref","first-page":"3363","DOI":"10.1093\/bioinformatics\/bth408","article-title":"Reducing storage requirements for biological sequence comparison","volume":"20","author":"Roberts","year":"2004","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B25","doi-asserted-by":"crossref","first-page":"2748","DOI":"10.1093\/bioinformatics\/bty205","article-title":"FaStore: a space-saving solution for raw sequencing data","volume":"34","author":"Roguski","year":"2018","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B26","doi-asserted-by":"crossref","first-page":"3380","DOI":"10.1093\/bioinformatics\/btx428","article-title":"Quark enables semi-reference-based compression of RNA-seq data","volume":"33","author":"Sarkar","year":"2017","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B27","doi-asserted-by":"crossref","first-page":"2192","DOI":"10.1093\/bioinformatics\/btq346","article-title":"G-SQZ: compact encoding of genomic sequence and quality data","volume":"26","author":"Tembe","year":"2010","journal-title":"Bioinformatics"},{"key":"2023012810014816200_bty936-B28","doi-asserted-by":"crossref","first-page":"315","DOI":"10.2174\/1574893609666140516010143","article-title":"Trends in genome compression","volume":"9","author":"Wandelt","year":"2014","journal-title":"Curr. Bioinform"},{"key":"2023012810014816200_bty936-B29","doi-asserted-by":"crossref","first-page":"240.","DOI":"10.1038\/nbt.3170","article-title":"Quality score compression improves genotyping accuracy","volume":"33","author":"Yu","year":"2015","journal-title":"Nat. Biotechnol"},{"key":"2023012810014816200_bty936-B30","doi-asserted-by":"crossref","first-page":"188.","DOI":"10.1186\/s12859-015-0628-7","article-title":"Light-weight reference-based compression of FASTQ data","volume":"16","author":"Zhang","year":"2015","journal-title":"BMC Bioinformatics"},{"key":"2023012810014816200_bty936-B31","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1093\/bib\/bbt087","article-title":"High-throughput DNA sequence data compression","volume":"16","author":"Zhu","year":"2015","journal-title":"Brief. Bioinform"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/12\/2066\/48934935\/bioinformatics_35_12_2066.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/12\/2066\/48934935\/bioinformatics_35_12_2066.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,28]],"date-time":"2023-01-28T10:19:16Z","timestamp":1674901156000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/35\/12\/2066\/5165381"}},"subtitle":[],"editor":[{"given":"John","family":"Hancock","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2018,11,8]]},"references-count":31,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2019,6,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bty936","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2019,6]]},"published":{"date-parts":[[2018,11,8]]}}}