{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T10:51:23Z","timestamp":1781779883020,"version":"3.54.5"},"publisher-location":"Cham","reference-count":32,"publisher":"Springer Nature Switzerland","isbn-type":[{"value":"9783031291180","type":"print"},{"value":"9783031291197","type":"electronic"}],"license":[{"start":{"date-parts":[[2023,1,1]],"date-time":"2023-01-01T00:00:00Z","timestamp":1672531200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,4,3]],"date-time":"2023-04-03T00:00:00Z","timestamp":1680480000000},"content-version":"vor","delay-in-days":92,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2023]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The reference indexing problem for <jats:inline-formula><jats:tex-math>$$k$$<\/jats:tex-math><\/jats:inline-formula>-mers is to pre-process a collection of reference genomic sequences <jats:inline-formula><jats:tex-math>$$\\mathcal {R}$$<\/jats:tex-math><\/jats:inline-formula> so that the position of all occurrences of any queried <jats:inline-formula><jats:tex-math>$$k$$<\/jats:tex-math><\/jats:inline-formula>-mer can be rapidly identified. An efficient and scalable solution to this problem is fundamental for many tasks in bioinformatics.<\/jats:p><jats:p>In this work, we introduce the <jats:italic>spectrum preserving tiling<\/jats:italic> (SPT), a general representation of <jats:inline-formula><jats:tex-math>$$\\mathcal {R}$$<\/jats:tex-math><\/jats:inline-formula> that specifies how a set of <jats:italic>tiles<\/jats:italic> repeatedly occur to spell out the constituent reference sequences in <jats:inline-formula><jats:tex-math>$$\\mathcal {R}$$<\/jats:tex-math><\/jats:inline-formula>. By encoding the order and positions where <jats:italic>tiles<\/jats:italic> occur, SPTs enable the implementation and analysis of a general class of modular indexes. An index over an SPT decomposes the reference indexing problem for <jats:inline-formula><jats:tex-math>$$k$$<\/jats:tex-math><\/jats:inline-formula>-mers into: (1) a <jats:inline-formula><jats:tex-math>$$k$$<\/jats:tex-math><\/jats:inline-formula>-mer-to-tile mapping; and (2) a tile-to-occurrence mapping. Recently introduced work to construct and compactly index <jats:inline-formula><jats:tex-math>$$k$$<\/jats:tex-math><\/jats:inline-formula>-mer sets can be used to efficiently implement the <jats:inline-formula><jats:tex-math>$$k$$<\/jats:tex-math><\/jats:inline-formula>-mer-to-tile mapping. However, implementing the tile-to-occurrence mapping remains prohibitively costly in terms of space. As reference collections become large, the space requirements of the tile-to-occurrence mapping dominates that of the <jats:inline-formula><jats:tex-math>$$k$$<\/jats:tex-math><\/jats:inline-formula>-mer-to-tile mapping since the former depends on the amount of total sequence while the latter depends on the number of unique <jats:inline-formula><jats:tex-math>$$k$$<\/jats:tex-math><\/jats:inline-formula>-mers in <jats:inline-formula><jats:tex-math>$$\\mathcal {R}$$<\/jats:tex-math><\/jats:inline-formula>.<\/jats:p><jats:p>To address this, we introduce a class of sampling schemes for SPTs that trade off speed to reduce the size of the tile-to-reference mapping. We implement a practical index with these sampling schemes in the tool . When indexing over 30,000 bacterial genomes,  reduces the size of the tile-to-occurrence mapping from 86.3 GB to 34.6 GB while incurring only a 3.6<jats:inline-formula><jats:tex-math>$$\\times $$<\/jats:tex-math><\/jats:inline-formula> slowdown when querying <jats:inline-formula><jats:tex-math>$$k$$<\/jats:tex-math><\/jats:inline-formula>-mers from a sequenced readset.<\/jats:p><jats:p><jats:bold>Availability:<\/jats:bold> is implemented in Rust and available at <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/COMBINE-lab\/pufferfish2\">https:\/\/github.com\/COMBINE-lab\/pufferfish2<\/jats:ext-link>.<\/jats:p>","DOI":"10.1007\/978-3-031-29119-7_2","type":"book-chapter","created":{"date-parts":[[2023,4,3]],"date-time":"2023-04-03T10:09:17Z","timestamp":1680516557000},"page":"21-40","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":14,"title":["Spectrum Preserving Tilings Enable Sparse and\u00a0Modular Reference Indexing"],"prefix":"10.1007","author":[{"given":"Jason","family":"Fan","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jamshed","family":"Khan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Giulio Ermanno","family":"Pibiri","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rob","family":"Patro","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,4,3]]},"reference":[{"issue":"22","key":"2_CR1","doi-asserted-by":"publisher","first-page":"404","DOI":"10.1093\/bioinformatics\/btab408","volume":"37","author":"F Almodaresi","year":"2021","unstructured":"Almodaresi, F., Zakeri, M., Patro, R.: PuffAligner: a fast, efficient and accurate aligner based on the pufferfish index. Bioinformatics 37(22), 404\u20134055 (2021)","journal-title":"Bioinformatics"},{"issue":"4","key":"2_CR2","doi-asserted-by":"publisher","first-page":"417","DOI":"10.1038\/nmeth.4197","volume":"14","author":"R Patro","year":"2017","unstructured":"Patro, R., Duggal, G., Love, M.I., Irizarry, R.A., Kingsford, C.: Salmon provides fast and bias-aware quantification of transcript expression. Nat. Methods 14(4), 417\u2013419 (2017)","journal-title":"Nat. Methods"},{"issue":"5","key":"2_CR3","doi-asserted-by":"publisher","first-page":"525","DOI":"10.1038\/nbt.3519","volume":"34","author":"NL Bray","year":"2016","unstructured":"Bray, N.L., Pimentel, H., Melsted, P., Pachter, L.: Near-optimal probabilistic RNA-seq quantification. Nat. Biotechnol. 34(5), 525\u2013527 (2016)","journal-title":"Nat. Biotechnol."},{"issue":"5","key":"2_CR4","doi-asserted-by":"publisher","first-page":"462","DOI":"10.1038\/nbt.2862","volume":"32","author":"R Patro","year":"2014","unstructured":"Patro, R., Mount, S.M., Kingsford, C.: Sailfish enables alignment-free isoform quantification from RNA-seq reads using lightweight algorithms. Nat. Biotechnol. 32(5), 462\u2013464 (2014)","journal-title":"Nat. Biotechnol."},{"key":"2_CR5","doi-asserted-by":"crossref","unstructured":"Gagie, T., Navarro, G., Prezza, N.: Optimal-time text indexing in BWT-runs bounded space. In: Proceedings of the Twenty-Ninth Annual ACM-SIAM Symposium on Discrete Algorithms, SODA 2018, USA, pp. 1459\u20131477. Society for Industrial and Applied Mathematics (2018)","DOI":"10.1137\/1.9781611975031.96"},{"issue":"2","key":"2_CR6","doi-asserted-by":"publisher","first-page":"169","DOI":"10.1089\/cmb.2021.0290","volume":"29","author":"M Rossi","year":"2022","unstructured":"Rossi, M., Oliva, M., Langmead, B., Gagie, T., Boucher, C.: MONI: a pangenomic index for finding maximal exact matches. J. Comput. Biol. 29(2), 169\u2013187 (2022). PMID: 35041495","journal-title":"J. Comput. Biol."},{"key":"2_CR7","doi-asserted-by":"crossref","unstructured":"Ahmed, O., Rossi, M., Gagie, T., Boucher, C., Langmead, B.: SPUMONI 2: improved pangenome classification using a compressed index of minimizer digests. BioRxiv (2022)","DOI":"10.1101\/2022.09.08.506805"},{"key":"2_CR8","unstructured":"Burrows, M., Wheeler, D.: A block-sorting lossless data compression algorithm. Digital SRC Research Report, Citeseer (1994)"},{"issue":"13","key":"2_CR9","doi-asserted-by":"publisher","first-page":"i169","DOI":"10.1093\/bioinformatics\/bty292","volume":"34","author":"F Almodaresi","year":"2018","unstructured":"Almodaresi, F., Sarkar, H., Srivastava, A., Patro, R.: A space and time-efficient index for the compacted colored de Bruijn graph. Bioinformatics 34(13), i169\u2013i177 (2018)","journal-title":"Bioinformatics"},{"issue":"8","key":"2_CR10","doi-asserted-by":"publisher","first-page":"907","DOI":"10.1038\/s41587-019-0201-4","volume":"37","author":"D Kim","year":"2019","unstructured":"Kim, D., Paggi, J.M., Park, C., Bennett, C., Salzberg, S.L.: Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat. Biotechnol. 37(8), 907\u2013915 (2019)","journal-title":"Nat. Biotechnol."},{"issue":"9","key":"2_CR11","doi-asserted-by":"publisher","first-page":"875","DOI":"10.1038\/nbt.4227","volume":"36","author":"E Garrison","year":"2018","unstructured":"Garrison, E., et al.: Variation graph toolkit improves read mapping by representing genetic variation in the reference. Nat. Biotechnol. 36(9), 875\u2013879 (2018)","journal-title":"Nat. Biotechnol."},{"issue":"24","key":"2_CR12","doi-asserted-by":"publisher","first-page":"4024","DOI":"10.1093\/bioinformatics\/btw609","volume":"33","author":"I Minkin","year":"2016","unstructured":"Minkin, I., Pham, S., Medvedev, P.: TwoPaCo: an efficient algorithm to build the compacted de Bruijn graph from many complete genomes. Bioinformatics 33(24), 4024\u20134032 (2016)","journal-title":"Bioinformatics"},{"issue":"12","key":"2_CR13","doi-asserted-by":"publisher","first-page":"i201","DOI":"10.1093\/bioinformatics\/btw279","volume":"32","author":"R Chikhi","year":"2016","unstructured":"Chikhi, R., Limasset, A., Medvedev, P.: Compacting de Bruijn graphs from sequencing data quickly and in low memory. Bioinformatics 32(12), i201\u2013i208 (2016)","journal-title":"Bioinformatics"},{"key":"2_CR14","doi-asserted-by":"crossref","unstructured":"Khan, J., Patro, R.: Cuttlefish: fast, parallel and low-memory compaction of de Bruijn graphs from large-scale genome collections. Bioinformatics 37(Supplement_1), i177\u2013i186 (2021)","DOI":"10.1093\/bioinformatics\/btab309"},{"issue":"1","key":"2_CR15","doi-asserted-by":"publisher","first-page":"190","DOI":"10.1186\/s13059-022-02743-6","volume":"23","author":"J Khan","year":"2022","unstructured":"Khan, J., Kokot, M., Deorowicz, S., Patro, R.: Scalable, ultra-fast, and low-memory construction of compacted de Bruijn graphs with Cuttlefish 2. Genome Biol. 23(1), 190 (2022). https:\/\/doi.org\/10.1186\/s13059-022-02743-6","journal-title":"Genome Biol."},{"issue":"10","key":"2_CR16","doi-asserted-by":"publisher","first-page":"958","DOI":"10.1016\/j.cels.2021.08.009","volume":"12","author":"B Ekim","year":"2021","unstructured":"Ekim, B., Berger, B., Chikhi, R.: Minimizer-space de Bruijn graphs: whole-genome assembly of long reads in minutes on a personal computer. Cell Syst. 12(10), 958-968.e6 (2021)","journal-title":"Cell Syst."},{"issue":"9","key":"2_CR17","doi-asserted-by":"publisher","first-page":"1754","DOI":"10.1101\/gr.276607.122","volume":"32","author":"M Karasikov","year":"2022","unstructured":"Karasikov, M., Mustafa, H., R\u00e4tsch, G., Kahles, A.: Lossless indexing with counting de Bruijn graphs. Genome Res. 32(9), 1754\u20131764 (2022)","journal-title":"Genome Res."},{"key":"2_CR18","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"152","DOI":"10.1007\/978-3-030-45257-5_10","volume-title":"Research in Computational Molecular Biology","author":"A Rahman","year":"2020","unstructured":"Rahman, A., Medvedev, P.: Representation of $$k$$-mer sets using spectrum-preserving string sets. In: Schwartz, R. (ed.) RECOMB 2020. LNCS, vol. 12074, pp. 152\u2013168. Springer, Cham (2020). https:\/\/doi.org\/10.1007\/978-3-030-45257-5_10"},{"key":"2_CR19","doi-asserted-by":"crossref","unstructured":"Schmidt, S., Alanko, J.N.: Eulertigs: minimum plain text representation of k-mer sets without repetitions in linear time. BioRxiv (2022)","DOI":"10.1101\/2022.05.17.492399"},{"issue":"1","key":"2_CR20","doi-asserted-by":"publisher","first-page":"96","DOI":"10.1186\/s13059-021-02297-z","volume":"22","author":"K B\u0159inda","year":"2021","unstructured":"B\u0159inda, K., Baym, M., Kucherov, G.: Simplitigs as an efficient and scalable representation of de Bruijn graphs. Genome Biol. 22(1), 96 (2021). https:\/\/doi.org\/10.1186\/s13059-021-02297-z","journal-title":"Genome Biol."},{"key":"2_CR21","doi-asserted-by":"crossref","unstructured":"Pibiri, G.E.: On weighted k-mer dictionaries. In: International Workshop on Algorithms in Bioinformatics (WABI), pp. 9:1\u20139:20 (2022)","DOI":"10.1101\/2022.05.23.493024"},{"key":"2_CR22","doi-asserted-by":"crossref","unstructured":"Pibiri, G.E.: Sparse and skew hashing of k-mers. Bioinformatics 38(Supplement_1), i185\u2013i194 (2022)","DOI":"10.1093\/bioinformatics\/btac245"},{"key":"2_CR23","doi-asserted-by":"crossref","unstructured":"Alanko, J.N., Puglisi, S.J., Vuohtoniemi, J.: Succinct k-mer sets using subset rank queries on the spectral burrows-wheeler transform. BioRxiv (2022)","DOI":"10.1101\/2022.05.19.492613"},{"key":"2_CR24","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"167","DOI":"10.1007\/978-3-642-34109-0_18","volume-title":"String Processing and Information Retrieval","author":"F Claude","year":"2012","unstructured":"Claude, F., Navarro, G.: The wavelet matrix. In: Calder\u00f3n-Benavides, L., Gonz\u00e1lez-Caro, C., Ch\u00e1vez, E., Ziviani, N. (eds.) SPIRE 2012. LNCS, vol. 7608, pp. 167\u2013179. Springer, Heidelberg (2012). https:\/\/doi.org\/10.1007\/978-3-642-34109-0_18"},{"key":"2_CR25","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1007\/978-3-642-33122-0_18","volume-title":"Algorithms in Bioinformatics","author":"A Bowe","year":"2012","unstructured":"Bowe, A., Onodera, T., Sadakane, K., Shibuya, T.: Succinct de Bruijn graphs. In: Raphael, B., Tang, J. (eds.) WABI 2012. LNCS, vol. 7534, pp. 225\u2013235. Springer, Heidelberg (2012). https:\/\/doi.org\/10.1007\/978-3-642-33122-0_18"},{"key":"2_CR26","doi-asserted-by":"crossref","unstructured":"Pibiri, G.E., Venturini, R.: Techniques for inverted index compression. ACM Comput. Surv. 53(6), 125:1\u2013125:36 (2021)","DOI":"10.1145\/3415148"},{"issue":"4","key":"2_CR27","doi-asserted-by":"publisher","first-page":"497","DOI":"10.1093\/bioinformatics\/btv603","volume":"32","author":"U Baier","year":"2015","unstructured":"Baier, U., Beller, T., Ohlebusch, E.: Graphical pan-genome analysis with compressed suffix trees and the Burrows-Wheeler transform. Bioinformatics 32(4), 497\u2013504 (2015)","journal-title":"Bioinformatics"},{"issue":"1","key":"2_CR28","doi-asserted-by":"publisher","first-page":"165","DOI":"10.1186\/s40168-021-01114-w","volume":"9","author":"P Hiseni","year":"2021","unstructured":"Hiseni, P., Rudi, K., Wilson, R.C., Hegge, F.T., Snipen, L.: HumGut: a comprehensive human gut prokaryotic genomes collection filtered by metagenome data. Microbiome 9(1), 165 (2021)","journal-title":"Microbiome"},{"issue":"1","key":"2_CR29","doi-asserted-by":"publisher","DOI":"10.1038\/sdata.2016.25","volume":"3","author":"JM Zook","year":"2016","unstructured":"Zook, J.M., et al.: Extensive sequencing of seven human genomes to characterize benchmark reference materials. Sci. Data 3(1), 160025 (2016)","journal-title":"Sci. Data"},{"issue":"1","key":"2_CR30","doi-asserted-by":"publisher","first-page":"92","DOI":"10.1038\/s41597-020-0427-5","volume":"7","author":"J Mas-Lloret","year":"2020","unstructured":"Mas-Lloret, J., et al.: Gut microbiome diversity detected by high-coverage 16S and shotgun sequencing of paired stool and colon sample. Sci. Data 7(1), 92 (2020)","journal-title":"Sci. Data"},{"issue":"1","key":"2_CR31","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1023\/A:1013002601898","volume":"3","author":"A Moffat","year":"2000","unstructured":"Moffat, A., Stuiver, L.: Binary interpolative coding for effective index compression. Inf. Retrieval 3(1), 25\u201347 (2000). https:\/\/doi.org\/10.1023\/A:1013002601898","journal-title":"Inf. Retrieval"},{"issue":"18","key":"2_CR32","doi-asserted-by":"publisher","first-page":"3094","DOI":"10.1093\/bioinformatics\/bty191","volume":"34","author":"H Li","year":"2018","unstructured":"Li, H.: Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34(18), 3094\u20133100 (2018)","journal-title":"Bioinformatics"}],"container-title":["Lecture Notes in Computer Science","Research in Computational Molecular Biology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-031-29119-7_2","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,4,3]],"date-time":"2023-04-03T10:23:38Z","timestamp":1680517418000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-031-29119-7_2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023]]},"ISBN":["9783031291180","9783031291197"],"references-count":32,"URL":"https:\/\/doi.org\/10.1007\/978-3-031-29119-7_2","relation":{},"ISSN":["0302-9743","1611-3349"],"issn-type":[{"value":"0302-9743","type":"print"},{"value":"1611-3349","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023]]},"assertion":[{"value":"3 April 2023","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"R.P. is a co-founder of Ocean Genomics Inc.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of Interest"}},{"value":"RECOMB","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"International Conference on Research in Computational Molecular Biology","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Istanbul","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"T\u00fcrkiye","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2023","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"16 April 2023","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"19 April 2023","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"27","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"recomb2023","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"http:\/\/recomb2023.bilkent.edu.tr\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Single-blind","order":1,"name":"type","label":"Type","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"Easy Chair","order":2,"name":"conference_management_system","label":"Conference Management System","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"188","order":3,"name":"number_of_submissions_sent_for_review","label":"Number of Submissions Sent for Review","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"11","order":4,"name":"number_of_full_papers_accepted","label":"Number of Full Papers Accepted","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"33","order":5,"name":"number_of_short_papers_accepted","label":"Number of Short Papers Accepted","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"6% - The value is computed by the equation \"Number of Full Papers Accepted \/ Number of Submissions Sent for Review * 100\" and then rounded to a whole number.","order":6,"name":"acceptance_rate_of_full_papers","label":"Acceptance Rate of Full Papers","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"3","order":7,"name":"average_number_of_reviews_per_paper","label":"Average Number of Reviews per Paper","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"6","order":8,"name":"average_number_of_papers_per_reviewer","label":"Average Number of Papers per Reviewer","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"Yes","order":9,"name":"external_reviewers_involved","label":"External Reviewers Involved","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}}]}}