{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T19:02:49Z","timestamp":1779130969330,"version":"3.51.4"},"reference-count":22,"publisher":"Oxford University Press (OUP)","issue":"5","license":[{"start":{"date-parts":[[2023,5,12]],"date-time":"2023-05-12T00:00:00Z","timestamp":1683849600000},"content-version":"vor","delay-in-days":11,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["OH 53\/7-2"],"award-info":[{"award-number":["OH 53\/7-2"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2023,5,4]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec><jats:title>Motivation<\/jats:title><jats:p>A pangenome represents many diverse genome sequences of the same species. In order to cope with small variations as well as structural variations, recent research focused on the development of graph-based models of pangenomes. Mapping is the process of finding the original location of a DNA read in a reference sequence, typically a genome. Using a pangenome instead of a (linear) reference genome can, e.g. reduce mapping bias, the tendency to incorrectly map sequences that differ from the reference genome. Mapping reads to a graph, however, is more complex and needs more resources than mapping to a reference genome. Reducing the complexity of the graph by encoding simple variations like SNPs in a simple way can accelerate read mapping and reduce the memory requirements at the same time.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>We introduce graphs based on elastic-degenerate strings (ED strings, EDS) and the linearized form of these EDS graphs as a new representation for pangenomes. In this representation, small variations are encoded directly in the sequence. Structural variations are encoded in a graph structure. This reduces the size of the representation in comparison to sequence graphs. In the linearized form, mapping techniques that are known from ordinary strings can be applied with appropriate adjustments. Since most variations are expressed directly in the sequence, the mapping process rarely has to take edges of the EDS graph into account. We developed a prototypical software tool GED-MAP that uses this representation together with a minimizer index to map short reads to the pangenome. Our experiments show that the new method works on a whole human genome scale, taking structural variants properly into account. The advantage of GED-MAP, compared with other pangenomic short read mappers, is that the new representation allows for a simple indexing method. This makes GED-MAP fast and memory efficient.<\/jats:p><\/jats:sec><jats:sec><jats:title>Availability and implementation<\/jats:title><jats:p>Sources are available at: https:\/\/github.com\/thomas-buechler-ulm\/gedmap.<\/jats:p><\/jats:sec>","DOI":"10.1093\/bioinformatics\/btad320","type":"journal-article","created":{"date-parts":[[2023,5,12]],"date-time":"2023-05-12T15:49:43Z","timestamp":1683906583000},"source":"Crossref","is-referenced-by-count":11,"title":["Efficient short read mapping to a pangenome that is represented by a graph of ED strings"],"prefix":"10.1093","volume":"39","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9273-5439","authenticated-orcid":false,"given":"Thomas","family":"B\u00fcchler","sequence":"first","affiliation":[{"name":"Institute of Theoretical Computer Science, Ulm University , 89075 Ulm, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jannik","family":"Olbrich","sequence":"additional","affiliation":[{"name":"Institute of Theoretical Computer Science, Ulm University , 89075 Ulm, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Enno","family":"Ohlebusch","sequence":"additional","affiliation":[{"name":"Institute of Theoretical Computer Science, Ulm University , 89075 Ulm, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2023,5,12]]},"reference":[{"key":"2023060100193723700_btad320-B1","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1038\/nature15393","article-title":"A global reference for human genetic variation","volume":"526","author":"1000 Genomes Project Consortium","year":"2015","journal-title":"Nature"},{"key":"2023060100193723700_btad320-B2","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","article-title":"Gapped BLAST and PSI-BLAST: a new generation of protein database search programs","volume":"25","author":"Altschul","year":"1997","journal-title":"Nucleic Acids Res"},{"key":"2023060100193723700_btad320-B3","author":"Aoyama","year":"2018"},{"key":"2023060100193723700_btad320-B4","doi-asserted-by":"crossref","first-page":"1413","DOI":"10.1093\/bioinformatics\/btz782","article-title":"An improved encoding of genetic variation in a Burrows\u2013Wheeler transform","volume":"36","author":"B\u00fcchler","year":"2020","journal-title":"Bioinformatics"},{"key":"2023060100193723700_btad320-B5","doi-asserted-by":"crossref","first-page":"4290","DOI":"10.1093\/bioinformatics\/bty506","article-title":"SOPanG: online text searching over a pan-genome","volume":"34","author":"Cis\u0142ak","year":"2018","journal-title":"Bioinformatics"},{"key":"2023060100193723700_btad320-B6","doi-asserted-by":"crossref","first-page":"2369","DOI":"10.1093\/nar\/27.11.2369","article-title":"Alignment of whole genomes","volume":"27","author":"Delcher","year":"1999","journal-title":"Nucleic Acids Res"},{"key":"2023060100193723700_btad320-B7","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1016\/0012-365X(75)90103-X","article-title":"On computing the length of longest increasing subsequences","volume":"11","author":"Fredman","year":"1975","journal-title":"Discret Math"},{"key":"2023060100193723700_btad320-B8","doi-asserted-by":"crossref","first-page":"875","DOI":"10.1038\/nbt.4227","article-title":"Variation graph toolkit improves read mapping by representing genetic variation in the reference","volume":"36","author":"Garrison","year":"2018","journal-title":"Nat Biotechnol"},{"key":"2023060100193723700_btad320-B9","author":"Grossi","year":"2017"},{"key":"2023060100193723700_btad320-B10","first-page":"131","author":"Iliopoulos","year":"2017"},{"key":"2023060100193723700_btad320-B11","first-page":"549","author":"Jacobson","year":"1989"},{"key":"2023060100193723700_btad320-B12","doi-asserted-by":"crossref","first-page":"907","DOI":"10.1038\/s41587-019-0201-4","article-title":"Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype","volume":"37","author":"Kim","year":"2019","journal-title":"Nat Biotechnol"},{"key":"2023060100193723700_btad320-B13","doi-asserted-by":"crossref","first-page":"2103","DOI":"10.1093\/bioinformatics\/btw152","article-title":"Minimap and miniasm: fast mapping and de novo assembly for noisy long sequences","volume":"32","author":"Li","year":"2016","journal-title":"Bioinformatics"},{"key":"2023060100193723700_btad320-B14","doi-asserted-by":"crossref","first-page":"3094","DOI":"10.1093\/bioinformatics\/bty191","article-title":"Minimap2: pairwise alignment for nucleotide sequences","volume":"34","author":"Li","year":"2018","journal-title":"Bioinformatics"},{"key":"2023060100193723700_btad320-B15","first-page":"222","author":"Maciuca","year":"2016"},{"key":"2023060100193723700_btad320-B16","first-page":"50","author":"Proch\u00e1zka","year":"2021"},{"key":"2023060100193723700_btad320-B17","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1186\/s12859-017-1678-9","article-title":"Coordinates and intervals in graph-based reference genomes","volume":"18","author":"Rand","year":"2017","journal-title":"BMC Bioinformatics"},{"key":"2023060100193723700_btad320-B18","doi-asserted-by":"crossref","first-page":"3363","DOI":"10.1093\/bioinformatics\/bth408","article-title":"Reducing storage requirements for biological sequence comparison","volume":"20","author":"Roberts","year":"2004","journal-title":"Bioinformatics"},{"key":"2023060100193723700_btad320-B19","doi-asserted-by":"crossref","first-page":"375","DOI":"10.1109\/TCBB.2013.2297101","article-title":"Indexing graphs for path queries with applications in genome research","volume":"11","author":"Sir\u00e9n","year":"2014","journal-title":"IEEE\/ACM Trans Comput Biol Bioinf"},{"key":"2023060100193723700_btad320-B20","doi-asserted-by":"crossref","first-page":"1461","DOI":"10.1126\/science.abg8871","article-title":"Pangenomics enables genotyping of known structural variants in 5202 diverse genomes","volume":"374","author":"Sir\u00e9n","year":"2021","journal-title":"Science"},{"key":"2023060100193723700_btad320-B21","first-page":"118","article-title":"Computational pan-genomics: status, promises and challenges","volume":"19","author":"The Computational Pan-Genomics Consortium","year":"2016","journal-title":"Brief Bioinf"},{"key":"2023060100193723700_btad320-B22","doi-asserted-by":"crossref","first-page":"160025","DOI":"10.1038\/sdata.2016.25","article-title":"Extensive sequencing of seven human genomes to characterize benchmark reference materials","volume":"3","author":"Zook","year":"2016","journal-title":"Sci Data"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btad320\/50297095\/btad320.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/39\/5\/btad320\/50497654\/btad320.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/39\/5\/btad320\/50497654\/btad320.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,12,12]],"date-time":"2023-12-12T20:41:23Z","timestamp":1702413683000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/doi\/10.1093\/bioinformatics\/btad320\/7160913"}},"subtitle":[],"editor":[{"given":"Yann","family":"Ponty","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2023,5,1]]},"references-count":22,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2023,5,4]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btad320","relation":{},"ISSN":["1367-4811"],"issn-type":[{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023,5,1]]},"published":{"date-parts":[[2023,5,1]]},"article-number":"btad320"}}