{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,20]],"date-time":"2026-08-20T13:32:13Z","timestamp":1787232733278,"version":"3.56.0"},"reference-count":15,"publisher":"Oxford University Press (OUP)","issue":"5","license":[{"start":{"date-parts":[[2024,5,10]],"date-time":"2024-05-10T00:00:00Z","timestamp":1715299200000},"content-version":"vor","delay-in-days":9,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Library of Medicine Training Program in Biomedical Informatics and Data Science","award":["T15LM007093"],"award-info":[{"award-number":["T15LM007093"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2024,5,2]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Motivation<\/jats:title>\n                    <jats:p>Since 2016, the number of microbial species with available reference genomes in NCBI has more than tripled. Multiple genome alignment, the process of identifying nucleotides across multiple genomes which share a common ancestor, is used as the input to numerous downstream comparative analysis methods. Parsnp is one of the few multiple genome alignment methods able to scale to the current era of genomic data; however, there has been no major release since its initial release in 2014.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>To address this gap, we developed Parsnp v2, which significantly improves on its original release. Parsnp v2 provides users with more control over executions of the program, allowing Parsnp to be better tailored for different use-cases. We introduce a partitioning option to Parsnp, which allows the input to be broken up into multiple parallel alignment processes which are then combined into a final alignment. The partitioning option can reduce memory usage by over 4\u00d7 and reduce runtime by over 2\u00d7, all while maintaining a precise core-genome alignment. The partitioning workflow is also less susceptible to complications caused by assembly artifacts and minor variation, as alignment anchors only need to be conserved within their partition and not across the entire input set. We highlight the performance on datasets involving thousands of bacterial and viral genomes.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Availability and implementation<\/jats:title>\n                    <jats:p>Parsnp v2 is available at https:\/\/github.com\/marbl\/parsnp.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btae311","type":"journal-article","created":{"date-parts":[[2024,5,9]],"date-time":"2024-05-09T21:25:15Z","timestamp":1715289915000},"source":"Crossref","is-referenced-by-count":91,"title":["Parsnp 2.0: scalable core-genome alignment for massive microbial datasets"],"prefix":"10.1093","volume":"40","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2946-6915","authenticated-orcid":false,"given":"Bryce","family":"Kille","sequence":"first","affiliation":[{"name":"Department of Computer Science, Rice University , Houston, TX 77005, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Michael G","family":"Nute","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Rice University , Houston, TX 77005, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Victor","family":"Huang","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Rice University , Houston, TX 77005, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eddie","family":"Kim","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Rice University , Houston, TX 77005, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2983-8934","authenticated-orcid":false,"given":"Adam M","family":"Phillippy","sequence":"additional","affiliation":[{"name":"Genome Informatics Section, Center for Genomics and Data Science Research, National Human Genome Research Institute, National Institutes of Health , Bethesda, MD 20892, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3760-564X","authenticated-orcid":false,"given":"Todd J","family":"Treangen","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Rice University , Houston, TX 77005, United States"},{"name":"Department of Bioengineering, Rice University , Houston, TX 77030, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2024,5,9]]},"reference":[{"key":"2024052610562444000_btae311-B1","doi-asserted-by":"crossref","first-page":"1115","DOI":"10.1093\/molbev\/msr268","article-title":"ALF\u2014a simulation framework for genome evolution","volume":"29","author":"Dalquen","year":"2012","journal-title":"Mol Biol Evol"},{"key":"2024052610562444000_btae311-B2","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1038\/s41587-023-01753-4","article-title":"Inference of phylogenetic trees directly from raw sequencing reads using read2tree","volume":"42","author":"Dylus","year":"2023","journal-title":"Nat Biotechnol"},{"key":"2024052610562444000_btae311-B3","doi-asserted-by":"crossref","first-page":"113","DOI":"10.1186\/1471-2105-5-113","article-title":"Muscle: a multiple sequence alignment method with reduced time and space complexity","volume":"5","author":"Edgar","year":"2004","journal-title":"BMC Bioinformatics"},{"key":"2024052610562444000_btae311-B4","doi-asserted-by":"crossref","first-page":"btad024","DOI":"10.1093\/bioinformatics\/btad024","article-title":"Evaluating impacts of syntenic block detection strategies on rearrangement phylogeny using Mycobacterium tuberculosis isolates","volume":"39","author":"Elghraoui","year":"2023","journal-title":"Bioinformatics"},{"key":"2024052610562444000_btae311-B5","doi-asserted-by":"crossref","first-page":"btad628","DOI":"10.1093\/bioinformatics\/btad628","article-title":"Coredetector: a flexible and efficient program for core-genome alignment of evolutionary diverse genomes","volume":"39","author":"Fruzangohar","year":"2023","journal-title":"Bioinformatics"},{"key":"2024052610562444000_btae311-B6","doi-asserted-by":"crossref","first-page":"1635","DOI":"10.1093\/molbev\/msw046","article-title":"Ete 3: reconstruction, analysis, and visualization of phylogenomic data","volume":"33","author":"Huerta-Cepas","year":"2016","journal-title":"Mol Biol Evol"},{"key":"2024052610562444000_btae311-B7","doi-asserted-by":"crossref","first-page":"5114","DOI":"10.1038\/s41467-018-07641-9","article-title":"High throughput ani analysis of 90k prokaryotic genomes reveals clear species boundaries","volume":"9","author":"Jain","year":"2018","journal-title":"Nat Commun"},{"key":"2024052610562444000_btae311-B8","doi-asserted-by":"crossref","first-page":"182","DOI":"10.1186\/s13059-022-02735-6","article-title":"Multiple genome alignment in the telomere-to-telomere assembly era","volume":"23","author":"Kille","year":"2022","journal-title":"Genome Biol"},{"key":"2024052610562444000_btae311-B9","first-page":"000872","article-title":"A global pangenome for the wheat fungal pathogen pyrenophora tritici-repentis and prediction of effector protein structural homology","volume":"8","author":"Moolhuijzen","year":"2022","journal-title":"Microb Genom"},{"key":"2024052610562444000_btae311-B10","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1126\/science.abj6987","article-title":"The complete sequence of a human genome","volume":"376","author":"Nurk","year":"2022","journal-title":"Science"},{"key":"2024052610562444000_btae311-B11","doi-asserted-by":"crossref","first-page":"3691","DOI":"10.1093\/bioinformatics\/btv421","article-title":"Roary: rapid large-scale prokaryote pan genome analysis","volume":"31","author":"Page","year":"2015","journal-title":"Bioinformatics"},{"key":"2024052610562444000_btae311-B12","doi-asserted-by":"crossref","first-page":"e9490","DOI":"10.1371\/journal.pone.0009490","article-title":"Fasttree 2\u2013approximately maximum-likelihood trees for large alignments","volume":"5","author":"Price","year":"2010","journal-title":"PLoS One"},{"key":"2024052610562444000_btae311-B13","doi-asserted-by":"crossref","first-page":"1312","DOI":"10.1093\/bioinformatics\/btu033","article-title":"Raxml version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies","volume":"30","author":"Stamatakis","year":"2014","journal-title":"Bioinformatics"},{"key":"2024052610562444000_btae311-B14","doi-asserted-by":"crossref","first-page":"524","DOI":"10.1186\/s13059-014-0524-x","article-title":"The harvest suite for rapid core-genome alignment and visualization of thousands of intraspecific microbial genomes","volume":"15","author":"Treangen","year":"2014","journal-title":"Genome Biol"},{"key":"2024052610562444000_btae311-B15","doi-asserted-by":"crossref","first-page":"737","DOI":"10.1101\/gr.214270.116","article-title":"Fast and accurate de novo genome assembly from long uncorrected reads","volume":"27","author":"Vaser","year":"2017","journal-title":"Genome Res"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btae311\/57468753\/btae311.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/40\/5\/btae311\/57909813\/btae311.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/40\/5\/btae311\/57909813\/btae311.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,26]],"date-time":"2024-05-26T06:56:41Z","timestamp":1716706601000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/doi\/10.1093\/bioinformatics\/btae311\/7667868"}},"subtitle":[],"editor":[{"given":"Russell","family":"Schwartz","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2024,5,1]]},"references-count":15,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,5,2]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btae311","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/2024.01.30.577458","asserted-by":"object"}]},"ISSN":["1367-4811"],"issn-type":[{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024,5,1]]},"published":{"date-parts":[[2024,5,1]]},"article-number":"btae311"}}