{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,4]],"date-time":"2026-08-04T20:55:46Z","timestamp":1785876946657,"version":"3.56.0"},"reference-count":35,"publisher":"Oxford University Press (OUP)","issue":"1","license":[{"start":{"date-parts":[[2022,11,30]],"date-time":"2022-11-30T00:00:00Z","timestamp":1669766400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000002","name":"NIH","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000026","name":"NIDA","doi-asserted-by":"publisher","award":["U01DA047638"],"award-info":[{"award-number":["U01DA047638"]}],"id":[{"id":"10.13039\/100000026","id-type":"DOI","asserted-by":"publisher"}]},{"name":"NSF PPoSS","award":["#2118709"],"award-info":[{"award-number":["#2118709"]}]},{"name":"Human Technopole in Milan"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2023,1,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Motivation<\/jats:title>\n                    <jats:p>Pangenome variation graphs model the mutual alignment of collections of DNA sequences. A set of pairwise alignments implies a variation graph, but there are no scalable methods to generate such a graph from these alignments. Existing related approaches depend on a single reference, a specific ordering of genomes or a de Bruijn model based on a fixed k-mer length. A scalable, self-contained method to build pangenome graphs without such limitations would be a key step in pangenome construction and manipulation pipelines.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>We design the seqwish algorithm, which builds a variation graph from a set of sequences and alignments between them. We first transform the alignment set into an implicit interval tree. To build up the variation graph, we query this tree-based representation of the alignments to reduce transitive matches into single DNA segments in a sequence graph. By recording the mapping from input sequence to output graph, we can trace the original paths through this graph, yielding a pangenome variation graph. We present an implementation that operates in external memory, using disk-backed data structures and lock-free parallel methods to drive the core graph induction step. We demonstrate that our method scales to very large graph induction problems by applying it to build pangenome graphs for several species.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Availability and implementation<\/jats:title>\n                    <jats:p>seqwish is published as free software under the MIT open source license. Source code and documentation are available at https:\/\/github.com\/ekg\/seqwish. seqwish can be installed via Bioconda https:\/\/bioconda.github.io\/recipes\/seqwish\/README.html or GNU Guix https:\/\/github.com\/ekg\/guix-genomics\/blob\/master\/seqwish.scm.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btac743","type":"journal-article","created":{"date-parts":[[2022,11,30]],"date-time":"2022-11-30T10:00:48Z","timestamp":1669802448000},"source":"Crossref","is-referenced-by-count":82,"title":["Unbiased pangenome graphs"],"prefix":"10.1093","volume":"39","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3821-631X","authenticated-orcid":false,"given":"Erik","family":"Garrison","sequence":"first","affiliation":[{"name":"Department of Genetics, Genomics and Informatics, University of Tennessee Health Science Center , Memphis, TN 38163, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9744-131X","authenticated-orcid":false,"given":"Andrea","family":"Guarracino","sequence":"additional","affiliation":[{"name":"Genomics Research Centre, Human Technopole , Viale Rita Levi-Montalcini 1 , Milan 20157, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2022,11,30]]},"reference":[{"key":"2023010107544180600_btac743-B1","first-page":"370","author":"Anderson","year":"1991"},{"key":"2023010107544180600_btac743-B2","doi-asserted-by":"crossref","first-page":"246","DOI":"10.1038\/s41586-020-2871-y","article-title":"Progressive cactus is a multiple-genome aligner for the thousand-genome era","volume":"587","author":"Armstrong","year":"2020","journal-title":"Nature"},{"key":"2023010107544180600_btac743-B3","first-page":"9:1","author":"Axtmann","year":"2017"},{"key":"2023010107544180600_btac743-B5","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1146\/annurev-genom-120219-080406","article-title":"Pangenome graphs","volume":"21","author":"Eizenga","year":"2020","journal-title":"Annu. Rev. Genomics Hum. Genet"},{"key":"2023010107544180600_btac743-B6","doi-asserted-by":"crossref","first-page":"169","DOI":"10.1007\/978-3-030-80049-9_15","volume-title":"Connecting with Computability","author":"Eizenga","year":"2021"},{"key":"2023010107544180600_btac743-B7","author":"Gao","year":"2020"},{"key":"2023010107544180600_btac743-B8","author":"Garrison","year":"2019"},{"key":"2023010107544180600_btac743-B9","doi-asserted-by":"crossref","first-page":"875","DOI":"10.1038\/nbt.4227","article-title":"Variation graph toolkit improves read mapping by representing genetic variation in the reference","volume":"36","author":"Garrison","year":"2018","journal-title":"Nat. Biotechnol"},{"key":"2023010107544180600_btac743-B10","author":"Garrison","year":"2022"},{"key":"2023010107544180600_btac743-B11","doi-asserted-by":"crossref","first-page":"326","DOI":"10.1007\/978-3-319-07959-2_28","article-title":"From theory to practice: plug and play with succinct data structures","author":"Gog","year":"2014","journal-title":"Experimental Algorithms"},{"key":"2023010107544180600_btac743-B12","doi-asserted-by":"crossref","first-page":"3319","DOI":"10.1093\/bioinformatics\/btac308","article-title":"ODGI: understanding pangenome graphs","volume":"38","author":"Guarracino","year":"2022","journal-title":"Bioinformatics"},{"key":"2023010107544180600_btac743-B13","author":"Harris","year":"2007"},{"key":"2023010107544180600_btac743-B14","first-page":"649","article-title":"A new method that simultaneously aligns and reconstructs ancestral sequences for any number of homologous sequences, when the phylogeny is given","volume":"6","author":"Hein","year":"1989","journal-title":"Mol. Biol. Evol"},{"key":"2023010107544180600_btac743-B15","doi-asserted-by":"crossref","first-page":"35","DOI":"10.1186\/s13059-020-1941-7","article-title":"Genotyping structural variants in pangenome graphs using the vg toolkit","volume":"21","author":"Hickey","year":"2020","journal-title":"Genome Biol"},{"key":"2023010107544180600_btac743-B16","doi-asserted-by":"crossref","first-page":"i748","DOI":"10.1093\/bioinformatics\/bty597","article-title":"A fast adaptive algorithm for computing whole-genome homology maps","volume":"34","author":"Jain","year":"2018","journal-title":"Bioinformatics"},{"key":"2023010107544180600_btac743-B17","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1002\/rsa.3240040303","article-title":"The birth of the giant component","volume":"4","author":"Janson","year":"1993","journal-title":"Random Struct. Algor"},{"key":"2023010107544180600_btac743-B18","doi-asserted-by":"crossref","first-page":"452","DOI":"10.1093\/bioinformatics\/18.3.452","article-title":"Multiple sequence alignment using partial order graphs","volume":"18","author":"Lee","year":"2002","journal-title":"Bioinformatics"},{"key":"2023010107544180600_btac743-B19","doi-asserted-by":"crossref","first-page":"3094","DOI":"10.1093\/bioinformatics\/bty191","article-title":"Minimap2: pairwise alignment for nucleotide sequences","volume":"34","author":"Li","year":"2018","journal-title":"Bioinformatics"},{"key":"2023010107544180600_btac743-B20","doi-asserted-by":"crossref","first-page":"1315","DOI":"10.1093\/bioinformatics\/btaa827","article-title":"Bedtk: finding interval overlap with implicit interval tree","volume":"37","author":"Li","year":"2021","journal-title":"Bioinformatics"},{"key":"2023010107544180600_btac743-B21","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1186\/s13059-020-02168-z","article-title":"The design and construction of reference pangenome graphs with minigraph","volume":"21","author":"Li","year":"2020","journal-title":"Genome Biol"},{"key":"2023010107544180600_btac743-B22","author":"Liao","year":"2022"},{"key":"2023010107544180600_btac743-B23","doi-asserted-by":"crossref","first-page":"456","DOI":"10.1093\/bioinformatics\/btaa777","article-title":"Fast gap-affine pairwise alignment using the wavefront algorithm","volume":"37","author":"Marco-Sola","year":"2021","journal-title":"Bioinformatics"},{"key":"2023010107544180600_btac743-B24","doi-asserted-by":"crossref","first-page":"589","DOI":"10.1016\/j.gde.2005.09.006","article-title":"The microbial pan-genome","volume":"15","author":"Medini","year":"2005","journal-title":"Curr. Opin. Genet. Dev"},{"key":"2023010107544180600_btac743-B25","doi-asserted-by":"crossref","first-page":"btw609","DOI":"10.1093\/bioinformatics\/btw609","article-title":"Twopaco: an efficient algorithm to build the compacted de Bruijn graph from many complete genomes","author":"Minkin","year":"2016","journal-title":"Bioinformatics"},{"key":"2023010107544180600_btac743-B26","doi-asserted-by":"crossref","first-page":"2966","DOI":"10.1093\/bioinformatics\/btz033","article-title":"Improved indel detection in DNA and RNA via realignment with abra2","volume":"35","author":"Mose","year":"2019","journal-title":"Bioinformatics"},{"key":"2023010107544180600_btac743-B27","author":"Nurk","year":"2021"},{"key":"2023010107544180600_btac743-B28","doi-asserted-by":"crossref","first-page":"737","DOI":"10.1038\/s41586-021-03451-0","article-title":"Towards complete and error-free genome assemblies of all vertebrate species","volume":"592","author":"Rhie","year":"2021","journal-title":"Nature"},{"key":"2023010107544180600_btac743-B29","first-page":"410","article-title":"Compressed text databases with efficient query algorithms based on the compressed suffix array","author":"Sadakane","year":"2000"},{"key":"2023010107544180600_btac743-B30","doi-asserted-by":"crossref","first-page":"i487","DOI":"10.1093\/bioinformatics\/btw455","article-title":"PanTools: representation, storage and exploration of pan-genomic data","volume":"32","author":"Sheikhizadeh","year":"2016","journal-title":"Bioinformatics"},{"key":"2023010107544180600_btac743-B31","doi-asserted-by":"crossref","first-page":"243","DOI":"10.1038\/s41576-020-0210-7","article-title":"Pan-genomics in the human genome era","volume":"21","author":"Sherman","year":"2020","journal-title":"Nat. Rev. Genet"},{"key":"2023010107544180600_btac743-B4","first-page":"118","article-title":"Computational pan-genomics: status, promises and challenges","volume":"19","author":"The Computational Pan-Genomics Consortium","year":"2018","journal-title":"Brief Bioinformatics"},{"key":"2023010107544180600_btac743-B32","first-page":"987","author":"Williams","year":"2009"},{"key":"2023010107544180600_btac743-B33","doi-asserted-by":"crossref","first-page":"548","DOI":"10.1186\/s12859-019-3145-2","article-title":"MoMI-G: modular multi-scale integrated genome graph browser","volume":"20","author":"Yokoyama","year":"2019","journal-title":"BMC Bioinformatics"},{"key":"2023010107544180600_btac743-B34","doi-asserted-by":"crossref","first-page":"2471","DOI":"10.1109\/TCBB.2021.3062068","article-title":"Stliter: a novel algorithm to iteratively build the compacted de Bruijn graph from many complete genomes","volume":"19","author":"Yu","year":"2021","journal-title":"IEEE\/ACM Trans. Comput. Biol. Bioinform"},{"key":"2023010107544180600_btac743-B35","doi-asserted-by":"crossref","first-page":"913","DOI":"10.1038\/ng.3847","article-title":"Contrasting evolutionary genome dynamics between domesticated and wild yeasts","volume":"49","author":"Yue","year":"2017","journal-title":"Nat. Genet"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btac743\/47774014\/btac743.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/39\/1\/btac743\/48448986\/btac743.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/39\/1\/btac743\/48448986\/btac743.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,1]],"date-time":"2023-01-01T05:12:54Z","timestamp":1672549974000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/doi\/10.1093\/bioinformatics\/btac743\/6854971"}},"subtitle":[],"editor":[{"given":"Can","family":"Alkan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2022,11,30]]},"references-count":35,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2022,11,30]]},"published-print":{"date-parts":[[2023,1,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btac743","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/2022.02.14.480413","asserted-by":"object"}]},"ISSN":["1367-4811"],"issn-type":[{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023,1,1]]},"published":{"date-parts":[[2022,11,30]]},"article-number":"btac743"}}