{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T23:17:21Z","timestamp":1783811841573,"version":"3.55.0"},"reference-count":21,"publisher":"Oxford University Press (OUP)","issue":"3","license":[{"start":{"date-parts":[[2023,3,2]],"date-time":"2023-03-02T00:00:00Z","timestamp":1677715200000},"content-version":"vor","delay-in-days":1,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100004281","name":"National Science Centre","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100004281","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2023,3,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec><jats:title>Motivation<\/jats:title><jats:p>High-quality sequence assembly is the ultimate representation of complete genetic information of an individual. Several ongoing pangenome projects are producing collections of high-quality assemblies of various species. Each project has already generated assemblies of hundreds of gigabytes on disk, greatly impeding the distribution of and access to such rich datasets.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>Here, we show how to reduce the size of the sequenced genomes by 2\u20133 orders of magnitude. Our tool compresses the genomes significantly better than the existing programs and is much faster. Moreover, its unique feature is the ability to access any contig (or its part) in a fraction of a second and easily append new samples to the compressed collections. Thanks to this, AGC could be useful not only for backup or transfer purposes but also for routine analysis of pangenome sequences in common pipelines. With the rapidly reduced cost and improved accuracy of sequencing technologies, we anticipate more comprehensive pangenome projects with much larger sample sizes. AGC is likely to become a foundation tool to store, distribute and access pangenome data.<\/jats:p><\/jats:sec><jats:sec><jats:title>Availability and implementation<\/jats:title><jats:p>The source code of AGC is available at https:\/\/github.com\/refresh-bio\/agc. The package can be installed via Bioconda at https:\/\/anaconda.org\/bioconda\/agc.<\/jats:p><\/jats:sec><jats:sec><jats:title>Supplementary information<\/jats:title><jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p><\/jats:sec>","DOI":"10.1093\/bioinformatics\/btad097","type":"journal-article","created":{"date-parts":[[2023,3,3]],"date-time":"2023-03-03T05:14:58Z","timestamp":1677820498000},"source":"Crossref","is-referenced-by-count":26,"title":["AGC: compact representation of assembled genomes with fast queries and updates"],"prefix":"10.1093","volume":"39","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9496-733X","authenticated-orcid":false,"given":"Sebastian","family":"Deorowicz","sequence":"first","affiliation":[{"name":"Department of Algorithmics and Software, Faculty of Automatic Control, Electronics and Computer Science, Silesian University of Technology , Akademicka 16 , Gliwice 44-100, Poland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1953-0880","authenticated-orcid":false,"given":"Agnieszka","family":"Danek","sequence":"additional","affiliation":[{"name":"Department of Algorithmics and Software, Faculty of Automatic Control, Electronics and Computer Science, Silesian University of Technology , Akademicka 16 , Gliwice 44-100, Poland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4874-2874","authenticated-orcid":false,"given":"Heng","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Data Sciences, Dana-Farber Cancer Institute , Boston, MA 02215, USA"},{"name":"Department of Biomedical Informatics, Harvard Medical School , Boston, MA 02115, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2023,3,2]]},"reference":[{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"914","DOI":"10.1038\/s41477-020-0733-0","article-title":"Plant pan-genomes are the new reference","volume":"6","author":"Bayer","year":"2020","journal-title":"Nat. Plants"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"e3001421","DOI":"10.1371\/journal.pbio.3001421","article-title":"Exploring bacterial diversity via a curated and searchable snapshot of archived DNA sequences","volume":"19","author":"Blackwell","year":"2021","journal-title":"PLoS Biol"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"170","DOI":"10.1038\/s41592-020-01056-5","article-title":"Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm","volume":"18","author":"Cheng","year":"2021","journal-title":"Nat. Methods"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"2979","DOI":"10.1093\/bioinformatics\/btr505","article-title":"Robust relative compression of genomes with random access","volume":"27","author":"Deorowicz","year":"2011","journal-title":"Bioinformatics"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"11565","DOI":"10.1038\/srep11565","article-title":"GDC 2: compression of large collections of genomes","volume":"5","author":"Deorowicz","year":"2015","journal-title":"Sci. Rep"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"eabf7117","DOI":"10.1126\/science.abf7117","article-title":"Haplotype-resolved diverse human genomes and integrated analysis of structural variation","volume":"372","author":"Ebert","year":"2021","journal-title":"Science"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"giab099","DOI":"10.1093\/gigascience\/giab099","article-title":"MBGC: multiple bacteria genome compressor","volume":"11","author":"Grabowski","year":"2022","journal-title":"Giga Science"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"284","DOI":"10.1038\/s41586-020-2947-8","article-title":"The barley pan-genome reveals the hidden legacy of mutation breeding","volume":"588","author":"Jayakodi","year":"2020","journal-title":"Nature"},{"key":"2023030819152207000_","first-page":"481","volume-title":"Book Man-Machine Interactions 5, Series Advances in Intelligent Systems and Computing","author":"Kokot","year":"2018"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"3826","DOI":"10.1093\/bioinformatics\/btz144","article-title":"Nucleotide archival format (NAF) enables efficient lossless reference-free compression of DNA sequences","volume":"35","author":"Kryukov","year":"2019","journal-title":"Bioinformatics"},{"key":"2023030819152207000_","first-page":"201","volume-title":"Relative Lempel-Ziv Compression of Genomes for Large-Scale Storage and Retrieval","author":"Kuruppu","year":"2010"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"58","DOI":"10.1186\/s13059-022-02602-4","article-title":"Genomic variations and epigenomic landscape of the medaka inbred Kiyosu-Karlsruhe (MIKK) panel","volume":"23","author":"Leger","year":"2022","journal-title":"Genome Biol"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"81","DOI":"10.1146\/annurev-genom-120120-081921","article-title":"The need for a human pangenome reference sequence","volume":"22","author":"Miga","year":"2021","journal-title":"Annu. Rev. Genomics Hum. Genet"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1126\/science.abj6987","article-title":"The complete sequence of a human genome","volume":"376","author":"Nurk","year":"2022","journal-title":"Science"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-84882-903-9","volume-title":"Handbook for Data Compression","author":"Salomon","year":"2010"},{"key":"2023030819152207000_","first-page":"202","author":"Shkarin","year":"2002"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"giaa119","DOI":"10.1093\/gigascience\/giaa119","article-title":"Efficient DNA sequence compression with neural networks","volume":"9","author":"Silva","year":"2020","journal-title":"GigaScience"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"928","DOI":"10.1145\/322344.322346","article-title":"Data compression via textual substitution","volume":"29","author":"Storer","year":"1982","journal-title":"J. ACM"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"3108950","DOI":"10.1155\/2019\/3108950","article-title":"HRCM: an efficient hybrid referential compression method for genomic big data","volume":"2019","author":"Yao","year":"2019","journal-title":"Biomed. Res. Int"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"1275","DOI":"10.1109\/TCBB.2013.122","article-title":"FRESCO: referential compression of highly similar sequences","volume":"10","author":"Wandelt","year":"2013","journal-title":"IEEE\/ACM Trans. Comput. Biol. Bioinform"},{"key":"2023030819152207000_","doi-asserted-by":"crossref","first-page":"437","DOI":"10.1038\/s41586-022-04601-8","article-title":"The human pangenome project: a global resource to map genomic diversity","volume":"604","author":"Wang","year":"2022","journal-title":"Nature"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btad097\/49403972\/btad097.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/39\/3\/btad097\/49457278\/btad097.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/39\/3\/btad097\/49457278\/btad097.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,3,27]],"date-time":"2023-03-27T00:44:02Z","timestamp":1679877842000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/doi\/10.1093\/bioinformatics\/btad097\/7067744"}},"subtitle":[],"editor":[{"given":"Tobias","family":"Marschall","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2023,3,1]]},"references-count":21,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,3,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btad097","relation":{},"ISSN":["1367-4811"],"issn-type":[{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023,3,1]]},"published":{"date-parts":[[2023,3,1]]},"article-number":"btad097"}}