{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,22]],"date-time":"2025-02-22T00:45:09Z","timestamp":1740185109194,"version":"3.37.3"},"reference-count":28,"publisher":"Oxford University Press (OUP)","issue":"17","license":[{"start":{"date-parts":[[2019,1,14]],"date-time":"2019-01-14T00:00:00Z","timestamp":1547424000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["1447711","1743418"],"award-info":[{"award-number":["1447711","1743418"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2019,9,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>Metagenomics is the study of genetic materials directly sampled from natural habitats. It has the potential to reveal previously hidden diversity of microscopic life largely due to the existence of highly parallel and low-cost next-generation sequencing technology. Conventional approaches align metagenomic reads onto known reference genomes to identify microbes in the sample. Since such a collection of reference genomes is very large, the approach often needs high-end computing machines with large memory which is not often available to researchers. Alternative approaches follow an alignment-free methodology where the presence of a microbe is predicted using the information about the unique k-mers present in the microbial genomes. However, such approaches suffer from high false positives due to trading off the value of k with the computational resources. In this article, we propose a highly efficient metagenomic sequence classification (MSC) algorithm that is a hybrid of both approaches. Instead of aligning reads to the full genomes, MSC aligns reads onto a set of carefully chosen, shorter and highly discriminating model sequences built from the unique k-mers of each of the reference sequences.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>Microbiome researchers are generally interested in two objectives of a taxonomic classifier: (i) to detect prevalence, i.e. the taxa present in a sample, and (ii) to estimate their relative abundances. MSC is primarily designed to detect prevalence and experimental results show that MSC is indeed a more effective and efficient algorithm compared to the other state-of-the-art algorithms in terms of accuracy, memory and runtime. Moreover, MSC outputs an approximate estimate of the abundances.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation<\/jats:title>\n                  <jats:p>The implementations are freely available for non-commercial purposes. They can be downloaded from https:\/\/drive.google.com\/open?id=1XirkAamkQ3ltWvI1W1igYQFusp9DHtVl.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/bty1071","type":"journal-article","created":{"date-parts":[[2019,1,8]],"date-time":"2019-01-08T17:58:56Z","timestamp":1546970336000},"page":"2932-2940","source":"Crossref","is-referenced-by-count":6,"title":["MSC: a metagenomic sequence classification algorithm"],"prefix":"10.1093","volume":"35","author":[{"given":"Subrata","family":"Saha","sequence":"first","affiliation":[{"name":"Healthcare and Life Sciences Division, IBM Thomas J. Watson Research Center , Yorktown Heights, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jethro","family":"Johnson","sequence":"additional","affiliation":[{"name":"The Jackson Laboratory for Genomic Medicine , Farmington, CT, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Soumitra","family":"Pal","sequence":"additional","affiliation":[{"name":"National Center for Biotechnology Information, National Institutes of Health , Bethesda, MD, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"George M","family":"Weinstock","sequence":"additional","affiliation":[{"name":"The Jackson Laboratory for Genomic Medicine , Farmington, CT, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sanguthevar","family":"Rajasekaran","sequence":"additional","affiliation":[{"name":"University of Connecticut Computer Science and Engineering Department, , Storrs, CT, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2019,1,14]]},"reference":[{"key":"2023062803480328100_bty1071-B1","doi-asserted-by":"crossref","first-page":"2253","DOI":"10.1093\/bioinformatics\/btt389","article-title":"Scalable metagenomic taxonomy classification using a reference genome database","volume":"29","author":"Ames","year":"2013","journal-title":"Bioinformatics"},{"key":"2023062803480328100_bty1071-B2","doi-asserted-by":"crossref","first-page":"e94","DOI":"10.1093\/nar\/gks251","article-title":"Grinder: a versatile amplicon and shotgun sequence simulator","volume":"40","author":"Angly","year":"2012","journal-title":"Nucleic Acids Res."},{"key":"2023062803480328100_bty1071-B3","doi-asserted-by":"crossref","first-page":"92","DOI":"10.1186\/1471-2105-13-92","article-title":"A comparative evaluation of sequence classification programs","volume":"13","author":"Bazinet","year":"2012","journal-title":"BMC Bioinformatics"},{"key":"2023062803480328100_bty1071-B4","doi-asserted-by":"crossref","first-page":"D25","DOI":"10.1093\/nar\/gkm929","article-title":"Genbank","volume":"36","author":"Benson","year":"2008","journal-title":"Nucleic Acids Res."},{"key":"2023062803480328100_bty1071-B5","doi-asserted-by":"crossref","first-page":"225","DOI":"10.1089\/10665270252935430","article-title":"Finding motifs using random projections","volume":"9","author":"Buhler","year":"2002","journal-title":"J. Comput. Biol."},{"key":"2023062803480328100_bty1071-B6","doi-asserted-by":"crossref","first-page":"421","DOI":"10.1186\/1471-2105-10-421","article-title":"BLAST: architecture and applications","volume":"10","author":"Camacho","year":"2009","journal-title":"BMC Bioinformatics"},{"key":"2023062803480328100_bty1071-B7","first-page":"265","article-title":"Nonparametric estimation of the number of classes in a population","volume":"11","author":"Chao","year":"1984","journal-title":"Scand. J. Stat."},{"key":"2023062803480328100_bty1071-B8","doi-asserted-by":"crossref","first-page":"182","DOI":"10.1111\/j.2041-1014.2012.00642.x","article-title":"Using high throughput sequencing to explore the biodiversity in oral bacterial communities","volume":"27","author":"Diaz","year":"2012","journal-title":"Mol. Oral Microbiol."},{"key":"2023062803480328100_bty1071-B9","doi-asserted-by":"crossref","first-page":"819","DOI":"10.1007\/s00294-017-0693-8","article-title":"The metagenomics worldwide research","volume":"63","author":"Garrido-Cardenas","year":"2017","journal-title":"Curr. Genet."},{"key":"2023062803480328100_bty1071-B10","doi-asserted-by":"crossref","first-page":"377","DOI":"10.1101\/gr.5969107","article-title":"MEGAN analysis of metagenomic data","volume":"17","author":"Huson","year":"2007","journal-title":"Genome Res."},{"key":"2023062803480328100_bty1071-B11","doi-asserted-by":"crossref","first-page":"853","DOI":"10.1038\/ismej.2016.174","article-title":"Where less may be more: how the rare biosphere pulls ecosystems strings","volume":"11","author":"Jousset","year":"2017","journal-title":"ISME J."},{"key":"2023062803480328100_bty1071-B12","doi-asserted-by":"crossref","DOI":"10.1128\/mSystems.00020-16","article-title":"MetaPalette: a k-mer painting approach for metagenomic taxonomic profiling and quantification of novel strain variation","volume":"1","author":"Koslicki","year":"2016","journal-title":"mSystems"},{"key":"2023062803480328100_bty1071-B13","doi-asserted-by":"crossref","first-page":"e91784","DOI":"10.1371\/journal.pone.0091784","article-title":"WGSQuikr: fast whole-genome shotgun metagenomic classification","volume":"9","author":"Koslicki","year":"2014","journal-title":"PLoS One"},{"key":"2023062803480328100_bty1071-B14","doi-asserted-by":"crossref","first-page":"19233","DOI":"10.1038\/srep19233","article-title":"An evaluation of the accuracy and speed of metagenome analysis tools","volume":"6","author":"Lindgreen","year":"2016","journal-title":"Sci. Rep."},{"key":"2023062803480328100_bty1071-B15","doi-asserted-by":"crossref","DOI":"10.1109\/BIBM.2010.5706544","article-title":"MetaPhyler: taxonomic profiling for metagenomic sequences","volume-title":"IEEE International Conference on Bioinformatics and Biomedicine (BIBM)","author":"Liu","year":"2010"},{"key":"2023062803480328100_bty1071-B16","doi-asserted-by":"crossref","first-page":"e104","DOI":"10.7717\/peerj-cs.104","article-title":"Bracken: estimating species abundance in metagenomics data","volume":"3","author":"Lu","year":"2016","journal-title":"PeerJ Comput. Sci."},{"key":"2023062803480328100_bty1071-B17","doi-asserted-by":"crossref","first-page":"11257","DOI":"10.1038\/ncomms11257","article-title":"Fast and sensitive taxonomic classification for metagenomics with Kaiju","volume":"7","author":"Menzel","year":"2016","journal-title":"Nat. Commun."},{"key":"2023062803480328100_bty1071-B18","doi-asserted-by":"crossref","first-page":"1757","DOI":"10.1093\/bioinformatics\/btn322","article-title":"Database indexing for production MegaBLAST searches","volume":"24","author":"Morgulis","year":"2008","journal-title":"Bioinformatics"},{"key":"2023062803480328100_bty1071-B19","doi-asserted-by":"crossref","first-page":"3740","DOI":"10.1093\/bioinformatics\/btx520","article-title":"MetaCache: context-aware classification of metagenomic reads using minhashing","volume":"33","author":"M\u00fcller","year":"2017","journal-title":"Bioinformatics"},{"key":"2023062803480328100_bty1071-B20","doi-asserted-by":"crossref","first-page":"D733","DOI":"10.1093\/nar\/gkv1189","article-title":"Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation","volume":"44","author":"O\u2019Leary","year":"2015","journal-title":"Nucleic Acids Res."},{"key":"2023062803480328100_bty1071-B21","doi-asserted-by":"crossref","first-page":"3823","DOI":"10.1093\/bioinformatics\/btw542","article-title":"Higher classification sensitivity of short metagenomic reads with CLARK-S","volume":"32","author":"Ounit","year":"2016","journal-title":"Bioinformatics"},{"key":"2023062803480328100_bty1071-B22","doi-asserted-by":"crossref","first-page":"236","DOI":"10.1186\/s12864-015-1419-2","article-title":"CLARK: fast and accurate classification of metagenomic and genomic sequences using discriminative k-mers","volume":"16","author":"Ounit","year":"2015","journal-title":"BMC Genomics"},{"key":"2023062803480328100_bty1071-B23","doi-asserted-by":"crossref","first-page":"2317","DOI":"10.1101\/gr.096651.109","article-title":"The NIH human microbiome project","volume":"19","author":"Peterson","year":"2009","journal-title":"Genome Res."},{"key":"2023062803480328100_bty1071-B24","doi-asserted-by":"crossref","first-page":"2082","DOI":"10.1093\/bioinformatics\/btx106","article-title":"Pseudoalignment for metagenomic read assignment","volume":"33","author":"Schaeffer","year":"2017","journal-title":"Bioinformatics"},{"key":"2023062803480328100_bty1071-B25","doi-asserted-by":"crossref","first-page":"1196","DOI":"10.1038\/nmeth.2693","article-title":"Metagenomic species profiling using universal phylogenetic marker genes","volume":"10","author":"Sunagawa","year":"2013","journal-title":"Nat. Methods"},{"key":"2023062803480328100_bty1071-B26","doi-asserted-by":"crossref","first-page":"902","DOI":"10.1038\/nmeth.3589","article-title":"MetaPhlAn2 for enhanced metagenomic taxonomic profiling","volume":"12","author":"Truong","year":"2015","journal-title":"Nat. Methods"},{"key":"2023062803480328100_bty1071-B27","doi-asserted-by":"crossref","first-page":"R46","DOI":"10.1186\/gb-2014-15-3-r46","article-title":"Kraken: ultrafast metagenomic sequence classification using exact alignments","volume":"15","author":"Wood","year":"2014","journal-title":"Genome Biol."},{"key":"2023062803480328100_bty1071-B28","doi-asserted-by":"crossref","first-page":"138","DOI":"10.1016\/j.gendis.2017.06.001","article-title":"Hypothesis testing and statistical analysis of microbiome","volume":"4","author":"Xia","year":"2017","journal-title":"Genes Dis."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/17\/2932\/50719748\/bioinformatics_35_17_2932.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/17\/2932\/50719748\/bioinformatics_35_17_2932.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,28]],"date-time":"2023-06-28T03:48:34Z","timestamp":1687924114000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/35\/17\/2932\/5288772"}},"subtitle":[],"editor":[{"given":"John","family":"Hancock","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2019,1,14]]},"references-count":28,"journal-issue":{"issue":"17","published-print":{"date-parts":[[2019,9,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bty1071","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"type":"print","value":"1367-4803"},{"type":"electronic","value":"1367-4811"}],"subject":[],"published-other":{"date-parts":[[2019,9,1]]},"published":{"date-parts":[[2019,1,14]]}}}