{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T15:29:39Z","timestamp":1782487779917,"version":"3.54.5"},"reference-count":55,"publisher":"Oxford University Press (OUP)","issue":"1","license":[{"start":{"date-parts":[[2021,9,8]],"date-time":"2021-09-08T00:00:00Z","timestamp":1631059200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000893","name":"Simons Foundation","doi-asserted-by":"publisher","award":["542963"],"award-info":[{"award-number":["542963"]}],"id":[{"id":"10.13039\/100000893","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000879","name":"Sloan Foundation","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100000879","id-type":"DOI","asserted-by":"crossref"}]},{"name":"McKnight Endowment Fund"},{"DOI":"10.13039\/100000001","name":"NSF","doi-asserted-by":"publisher","award":["DBI-1707398"],"award-info":[{"award-number":["DBI-1707398"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100000324","name":"Gatsby Charitable Foundation","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100000324","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2021,12,22]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Motivation<\/jats:title>\n                    <jats:p>The automatic discovery of sparse biomarkers that are associated with an outcome of interest is a central goal of bioinformatics. In the context of high-throughput sequencing (HTS) data, and compositional data (CoDa) more generally, an important class of biomarkers are the log-ratios between the input variables. However, identifying predictive log-ratio biomarkers from HTS data is a combinatorial optimization problem, which is computationally challenging. Existing methods are slow to run and scale poorly with the dimension of the input, which has limited their application to low- and moderate-dimensional metagenomic datasets.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>Building on recent advances from the field of deep learning, we present CoDaCoRe, a novel learning algorithm that identifies sparse, interpretable and predictive log-ratio biomarkers. Our algorithm exploits a continuous relaxation to approximate the underlying combinatorial optimization problem. This relaxation can then be optimized efficiently using the modern ML toolbox, in particular, gradient descent. As a result, CoDaCoRe runs several orders of magnitude faster than competing methods, all while achieving state-of-the-art performance in terms of predictive accuracy and sparsity. We verify the outperformance of CoDaCoRe across a wide range of microbiome, metabolite and microRNA benchmark datasets, as well as a particularly high-dimensional dataset that is outright computationally intractable for existing sparse log-ratio selection methods.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Availability and implementation<\/jats:title>\n                    <jats:p>The CoDaCoRe package is available at https:\/\/github.com\/egr95\/R-codacore. Code and instructions for reproducing our results are available at https:\/\/github.com\/cunningham-lab\/codacore.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Supplementary information<\/jats:title>\n                    <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btab645","type":"journal-article","created":{"date-parts":[[2021,9,8]],"date-time":"2021-09-08T07:46:10Z","timestamp":1631087170000},"page":"157-163","source":"Crossref","is-referenced-by-count":28,"title":["Learning sparse log-ratios for high-throughput sequencing data"],"prefix":"10.1093","volume":"38","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2720-1756","authenticated-orcid":false,"given":"Elliott","family":"Gordon-Rodriguez","sequence":"first","affiliation":[{"name":"Department of Statistics, Columbia University , New York, NY 10025, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0286-6329","authenticated-orcid":false,"given":"Thomas P","family":"Quinn","sequence":"additional","affiliation":[{"name":"Applied Artificial Intelligence Institute, Deakin University , Geelong, VIC 3126, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"John P","family":"Cunningham","sequence":"additional","affiliation":[{"name":"Department of Statistics, Columbia University , New York, NY 10025, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2021,9,8]]},"reference":[{"key":"2023020108403061500_btab645-B1","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1111\/j.2517-6161.1982.tb01195.x","article-title":"The statistical analysis of compositional data","volume":"44","author":"Aitchison","year":"1982","journal-title":"J. R. Stat. Soc. Ser. B (Methodological)"},{"key":"2023020108403061500_btab645-B2","doi-asserted-by":"crossref","first-page":"479","DOI":"10.1158\/2159-8290.CD-15-1483","article-title":"Clinical applications of circulating tumor cells and circulating tumor DNA as liquid biopsy","volume":"6","author":"Alix-Panabi\u00e8res","year":"2016","journal-title":"Cancer Discov"},{"key":"2023020108403061500_btab645-B3","doi-asserted-by":"crossref","first-page":"613","DOI":"10.1111\/biom.12995","article-title":"Log-ratio lasso: scalable, sparse estimation for log-ratio models","volume":"75","author":"Bates","year":"2019","journal-title":"Biometrics"},{"key":"2023020108403061500_btab645-B4","doi-asserted-by":"crossref","first-page":"666","DOI":"10.1016\/j.ccell.2015.09.018","article-title":"RNA-seq of tumor-educated platelets enables blood-based pan-cancer, multiclass, and molecular pathway cancer diagnostics","volume":"28","author":"Best","year":"2015","journal-title":"Cancer Cell"},{"key":"2023020108403061500_btab645-B5","doi-asserted-by":"crossref","first-page":"e6","DOI":"10.5808\/GI.2019.17.1.e6","article-title":"Statistical analysis of metagenomics data","volume":"17","author":"Calle","year":"2019","journal-title":"Genomics Inf"},{"key":"2023020108403061500_btab645-B6","doi-asserted-by":"crossref","first-page":"635","DOI":"10.1038\/s41575-020-0327-3","article-title":"Gut microbiome, big data and machine learning to promote precision medicine for cancer","volume":"17","author":"Cammarota","year":"2020","journal-title":"Nat. Rev. Gastroenterol. Hepatol"},{"key":"2023020108403061500_btab645-B7","doi-asserted-by":"crossref","first-page":"1251","DOI":"10.1038\/s41430-020-0607-6","article-title":"Profile of the gut microbiota of adults with obesity: a systematic review","volume":"74","author":"Crovesy","year":"2020","journal-title":"Eur. J. Clin. Nutr"},{"key":"2023020108403061500_btab645-B8","doi-asserted-by":"crossref","first-page":"671","DOI":"10.1093\/bib\/bbs046","article-title":"A comprehensive evaluation of normalization methods for illumina high-throughput RNA sequencing data analysis","volume":"14","author":"Dillies","year":"2013","journal-title":"Brief. Bioinf"},{"key":"2023020108403061500_btab645-B9","doi-asserted-by":"crossref","first-page":"795","DOI":"10.1007\/s11004-005-7381-9","article-title":"Groups of parts and their balances in compositional data analysis","volume":"37","author":"Egozcue","year":"2005","journal-title":"Math. Geol"},{"key":"2023020108403061500_btab645-B10","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1023\/A:1023818214614","article-title":"Isometric logratio transformations for compositional data analysis","volume":"35","author":"Egozcue","year":"2003","journal-title":"Math. Geol"},{"key":"2023020108403061500_btab645-B11","doi-asserted-by":"crossref","first-page":"e67019","DOI":"10.1371\/journal.pone.0067019","article-title":"Anova-like differential expression (ALDEX) analysis for mixed population RNA-seq","volume":"8","author":"Fernandes","year":"2013","journal-title":"PLoS One"},{"key":"2023020108403061500_btab645-B12","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1186\/2049-2618-2-15","article-title":"Unifying the analysis of high-throughput sequencing datasets: characterizing RNA-seq, 16s RRNA gene sequencing and selective growth experiments by compositional data analysis","volume":"2","author":"Fernandes","year":"2014","journal-title":"Microbiome"},{"key":"2023020108403061500_btab645-B13","doi-asserted-by":"crossref","first-page":"194","DOI":"10.1016\/j.chroma.2014.08.050","article-title":"What can go wrong at the data normalization step for identification of biomarkers?","volume":"1362","author":"Filzmoser","year":"2014","journal-title":"J. Chromatography A"},{"key":"2023020108403061500_btab645-B14","doi-asserted-by":"crossref","first-page":"6100","DOI":"10.1016\/j.scitotenv.2009.08.008","article-title":"Univariate statistical analysis of environmental (compositional) data: problems and possibilities","volume":"407","author":"Filzmoser","year":"2009","journal-title":"Sci. Total Environ"},{"key":"2023020108403061500_btab645-B15","author":"Friedman","year":"2001"},{"key":"2023020108403061500_btab645-B16","doi-asserted-by":"crossref","first-page":"692","DOI":"10.1139\/cjm-2015-0821","article-title":"Compositional analysis: a valid approach to analyze microbiome high-throughput sequencing data","volume":"62","author":"Gloor","year":"2016","journal-title":"Can. J. Microbiol"},{"key":"2023020108403061500_btab645-B17","doi-asserted-by":"crossref","first-page":"322","DOI":"10.1016\/j.annepidem.2016.03.003","article-title":"It\u2019s all relative: analyzing microbiome data as compositions","volume":"26","author":"Gloor","year":"2016","journal-title":"Ann. Epidemiol"},{"key":"2023020108403061500_btab645-B18","doi-asserted-by":"crossref","first-page":"2224","DOI":"10.3389\/fmicb.2017.02224","article-title":"Microbiome datasets are compositional: and this is not optional","volume":"8","author":"Gloor","year":"2017","journal-title":"Front. Microbiol"},{"key":"2023020108403061500_btab645-B19","first-page":"50","article-title":"European union regulations on algorithmic decision-making and a \u201cright to explanation\u201d","volume":"38","author":"Goodman","year":"2017","journal-title":"AI Mag"},{"key":"2023020108403061500_btab645-B20","doi-asserted-by":"crossref","first-page":"644","DOI":"10.1007\/s11749-019-00673-3","article-title":"Comments on: compositional data: the sample space and its structure","volume":"28","author":"Greenacre","year":"2019","journal-title":"TEST"},{"key":"2023020108403061500_btab645-B21","doi-asserted-by":"crossref","first-page":"649","DOI":"10.1007\/s11004-018-9754-x","article-title":"Variable selection in compositional data analysis using pairwise logratios","volume":"51","author":"Greenacre","year":"2019","journal-title":"Math. Geosci"},{"key":"2023020108403061500_btab645-B22","doi-asserted-by":"crossref","first-page":"100017","DOI":"10.1016\/j.acags.2019.100017","article-title":"Amalgamations are valid in compositional data analysis, can be used in agglomerative clustering, and their logratios have an inverse transformation","volume":"5","author":"Greenacre","year":"2020","journal-title":"Appl. Comput. Geosci"},{"key":"2023020108403061500_btab645-B23","first-page":"104621","article-title":"A comparison of isometric and amalgamation logratio balances in compositional data analysis","author":"Greenacre","year":"2020","journal-title":"Computers & Geosciences, 104"},{"key":"2023020108403061500_btab645-B24","author":"He","year":"2013"},{"key":"2023020108403061500_btab645-B25","author":"Jang","year":"2017"},{"key":"2023020108403061500_btab645-B26","doi-asserted-by":"crossref","first-page":"73","DOI":"10.1146\/annurev-statistics-010814-020351","article-title":"Microbiome, metagenomics, and high-dimensional compositional data analysis","volume":"2","author":"Li","year":"2015","journal-title":"Annu. Rev. Stat. Appl"},{"key":"2023020108403061500_btab645-B27","author":"Linderman","year":"2018"},{"key":"2023020108403061500_btab645-B28","doi-asserted-by":"crossref","first-page":"e1004075","DOI":"10.1371\/journal.pcbi.1004075","article-title":"Proportionality: a valid alternative to correlation for relative data","volume":"11","author":"Lovell","year":"2015","journal-title":"PLoS Comput. Biol"},{"key":"2023020108403061500_btab645-B29","doi-asserted-by":"crossref","first-page":"235","DOI":"10.1111\/biom.12956","article-title":"Generalized linear models with linear constraints for microbiome compositional data","volume":"75","author":"Lu","year":"2019","journal-title":"Biometrics"},{"key":"2023020108403061500_btab645-B30","author":"Maddison","year":"2017"},{"key":"2023020108403061500_btab645-B31","doi-asserted-by":"crossref","first-page":"1474","DOI":"10.3390\/nu12051474","article-title":"The firmicutes\/bacteroidetes ratio: a relevant marker of gut dysbiosis in obese patients?","volume":"12","author":"Magne","year":"2020","journal-title":"Nutrients"},{"key":"2023020108403061500_btab645-B32","author":"Mena","year":"2018"},{"key":"2023020108403061500_btab645-B33","doi-asserted-by":"crossref","first-page":"e00162-16","DOI":"10.1128\/mSystems.00162-16","article-title":"Balance trees reveal microbial niche differentiation","volume":"2","author":"Morton","year":"2017","journal-title":"MSystems"},{"key":"2023020108403061500_btab645-B34","doi-asserted-by":"crossref","first-page":"2719","DOI":"10.1038\/s41467-019-10656-5","article-title":"Establishing microbial composition measurement standards with reference frames","volume":"10","author":"Morton","year":"2019","journal-title":"Nat. Commun"},{"key":"2023020108403061500_btab645-B35","doi-asserted-by":"crossref","DOI":"10.1002\/9781119976462","volume-title":"Compositional Data Analysis: Theory and Applications","author":"Pawlowsky-Glahn","year":"2011"},{"key":"2023020108403061500_btab645-B36","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1144\/GSL.SP.2006.264.01.01","article-title":"Compositional data and their analysis: an introduction","volume":"264","author":"Pawlowsky-Glahn","year":"2006","journal-title":"Geol. Soc. Lond. Special Public"},{"key":"2023020108403061500_btab645-B37","doi-asserted-by":"crossref","first-page":"253","DOI":"10.1098\/rsta.1896.0007","article-title":"VII. Mathematical contributions to the theory of evolution. III. Regression, heredity, and panmixia","volume":"187","author":"Pearson","year":"1896","journal-title":"Philos. Trans. R. Soc. Lond. Ser. A"},{"key":"2023020108403061500_btab645-B38","first-page":"33","article-title":"Invertible gaussian reparameterization: revisiting the gumbel-softmax","author":"Potapczynski","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2023020108403061500_btab645-B39","doi-asserted-by":"crossref","first-page":"giaa010","DOI":"10.1093\/gigascience\/giaa010","article-title":"Interpretable and accurate prediction models for metagenomics data","volume":"9","author":"Prifti","year":"2020","journal-title":"GigaScience"},{"key":"2023020108403061500_btab645-B40","author":"Quinn","year":"2020"},{"key":"2023020108403061500_btab645-B41","author":"Quinn","year":"2019"},{"key":"2023020108403061500_btab645-B42","doi-asserted-by":"crossref","first-page":"lqaa076","DOI":"10.1093\/nargab\/lqaa076","article-title":"Amalgams: data-driven amalgamation for the dimensionality reduction of compositional data","volume":"2","author":"Quinn","year":"2020","journal-title":"NAR Genomics Bioinf"},{"key":"2023020108403061500_btab645-B43","doi-asserted-by":"crossref","first-page":"16252","DOI":"10.1038\/s41598-017-16520-0","article-title":"propr: an r-package for identifying proportionally abundant features using compositional data analysis","volume":"7","author":"Quinn","year":"2017","journal-title":"Sci. Rep"},{"key":"2023020108403061500_btab645-B44","doi-asserted-by":"crossref","first-page":"2870","DOI":"10.1093\/bioinformatics\/bty175","article-title":"Understanding sequencing data as compositions: an outlook and review","volume":"34","author":"Quinn","year":"2018","journal-title":"Bioinformatics"},{"key":"2023020108403061500_btab645-B45","doi-asserted-by":"crossref","first-page":"giz107","DOI":"10.1093\/gigascience\/giz107","article-title":"A field guide for the compositional analysis of any-omics data","volume":"8","author":"Quinn","year":"2019","journal-title":"GigaScience"},{"key":"2023020108403061500_btab645-B46","author":"Quinn","year":"2021"},{"key":"2023020108403061500_btab645-B47","doi-asserted-by":"crossref","first-page":"1525","DOI":"10.1038\/ijo.2014.46","article-title":"Evidence for greater production of colonic short-chain fatty acids in overweight than lean humans","volume":"38","author":"Rahat-Rozenbloom","year":"2014","journal-title":"Int. J. Obesity"},{"key":"2023020108403061500_btab645-B48","doi-asserted-by":"crossref","first-page":"e00053-18","DOI":"10.1128\/mSystems.00053-18","article-title":"Balances: a new perspective for microbiome analysis","volume":"3","author":"Rivera-Pinto","year":"2018","journal-title":"MSystems"},{"key":"2023020108403061500_btab645-B49","doi-asserted-by":"crossref","first-page":"8143","DOI":"10.2147\/OTT.S177384","article-title":"Identification of tumor-educated platelet biomarkers of non-small-cell lung cancer","volume":"11","author":"Sheng","year":"2018","journal-title":"OncoTargets Ther"},{"key":"2023020108403061500_btab645-B50","doi-asserted-by":"crossref","first-page":"e21887","DOI":"10.7554\/eLife.21887","article-title":"A phylogenetic transform enhances analysis of compositional microbiota data","volume":"6","author":"Silverman","year":"2017","journal-title":"Elife"},{"key":"2023020108403061500_btab645-B51","doi-asserted-by":"crossref","first-page":"lqaa029","DOI":"10.1093\/nargab\/lqaa029","article-title":"Variable selection in microbiome compositional data analysis","volume":"2","author":"Susin","year":"2020","journal-title":"NAR Genomics and Bioinformatics"},{"key":"2023020108403061500_btab645-B52","doi-asserted-by":"crossref","DOI":"10.1093\/gigascience\/giz042","article-title":"Microbiome Learning Repo (ML Repo): a public repository of microbiome regression and classification tasks","volume":"8","author":"Vangay","year":"2019","journal-title":"GigaScience"},{"key":"2023020108403061500_btab645-B53","doi-asserted-by":"crossref","first-page":"223","DOI":"10.1038\/nrc.2017.7","article-title":"Liquid biopsies come of age: towards implementation of circulating tumour DNA","volume":"17","author":"Wan","year":"2017","journal-title":"Nat. Rev. Cancer"},{"key":"2023020108403061500_btab645-B54","doi-asserted-by":"crossref","first-page":"e2969","DOI":"10.7717\/peerj.2969","article-title":"Phylogenetic factorization of compositional data yields lineage-level associations in microbiome datasets","volume":"5","author":"Washburne","year":"2017","journal-title":"PeerJ"},{"key":"2023020108403061500_btab645-B55","doi-asserted-by":"crossref","first-page":"87494","DOI":"10.18632\/oncotarget.20903","article-title":"Identifying and analyzing different cancer subtypes using RNA-seq data of blood platelets","volume":"8","author":"Zhang","year":"2017","journal-title":"Oncotarget"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btab645\/40416229\/btab645.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/38\/1\/157\/49006577\/btab645.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/38\/1\/157\/49006577\/btab645.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,7]],"date-time":"2024-09-07T21:10:36Z","timestamp":1725743436000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/38\/1\/157\/6366546"}},"subtitle":[],"editor":[{"given":"Pier","family":"Luigi Martelli","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2021,9,8]]},"references-count":55,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12,22]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btab645","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/2021.02.11.430695","asserted-by":"object"}]},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022,1,1]]},"published":{"date-parts":[[2021,9,8]]}}}