{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T13:06:20Z","timestamp":1779195980978,"version":"3.51.4"},"reference-count":40,"publisher":"Oxford University Press (OUP)","issue":"3","license":[{"start":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T00:00:00Z","timestamp":1779148800000},"content-version":"vor","delay-in-days":18,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"name":"French Agence Nationale de la Recherche","award":["ANR-19-P3IA-0001"],"award-info":[{"award-number":["ANR-19-P3IA-0001"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026,5,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Advances in sequencing technologies have enabled the generation of large amounts of data, offering new possibilities to identify relationships between biological units (e.g. genes) and phenotypic traits (e.g. disease outcomes). Yet, identifying these associations using variable selection methods remains challenging due to the high dimension ($p \\gg n$) and the correlation structure of the data. To address these challenges, we study the applicability of the knockoff (KO) procedure. Introduced by Barber and Cand\u00e8s in 2015, the KO variable selection procedure has shown promising results on real biological data, such as Genome-Wide Association Studies. This method seeks to identify the truly important predictors by overcoming the correlation structure between variables while controlling the false discovery rate. Here, we study the applicability of the KO procedure on transcriptomic data in a classification setting. We conduct an extensive simulation study using real transcriptomic data to evaluate the performance of the KO framework in the context of high-dimensional classification. We find that the KO framework outperforms widely used variable selection models, and that using KO aggregation to mitigate the effect of KO stochasticity improves stability while maintaining the same power. Finally, applied to three real transcriptomic datasets, the KO framework made very few discoveries, highlighting its conservative nature and suggesting that other methods may substantially overestimate the number of relevant features.<\/jats:p>","DOI":"10.1093\/bib\/bbag148","type":"journal-article","created":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T12:01:44Z","timestamp":1778673704000},"source":"Crossref","is-referenced-by-count":0,"title":["Statistical knockoffs improve biomarker discovery from transcriptomic data"],"prefix":"10.1093","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-7745-0634","authenticated-orcid":false,"given":"Julie","family":"Cartier","sequence":"first","affiliation":[{"name":"Centre for Computational Biology, Mines Paris, PSL University , 60 bd Saint-Michel, 75272 Paris,","place":["France"]},{"name":"Institut Curie, PSL University, 11 rue Pierre et Marie Curie , 75005 Paris,","place":["France"]},{"name":"U1331, INSERM, 11 rue Pierre et Marie Curie , 75005 Paris,","place":["France"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Johanna","family":"Lagoas","sequence":"additional","affiliation":[{"name":"Centre for Computational Biology, Mines Paris, PSL University , 60 bd Saint-Michel, 75272 Paris,","place":["France"]},{"name":"Institut Curie, PSL University, 11 rue Pierre et Marie Curie , 75005 Paris,","place":["France"]},{"name":"U1331, INSERM, 11 rue Pierre et Marie Curie , 75005 Paris,","place":["France"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Youmna","family":"Ayadi","sequence":"additional","affiliation":[{"name":"Centre for Computational Biology, Mines Paris, PSL University , 60 bd Saint-Michel, 75272 Paris,","place":["France"]},{"name":"Institut Curie, PSL University, 11 rue Pierre et Marie Curie , 75005 Paris,","place":["France"]},{"name":"U1331, INSERM, 11 rue Pierre et Marie Curie , 75005 Paris,","place":["France"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Adeline","family":"Fermanian","sequence":"additional","affiliation":[{"name":"LOPF, LOPF Califrais\u2019Machine Learning Lab , 4 Quai du Val de Loire,\u00a094550 Chevilly Larue,","place":["France"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1003-301X","authenticated-orcid":false,"given":"Chlo\u00e9-Agathe","family":"Azencott","sequence":"additional","affiliation":[{"name":"Centre for Computational Biology, Mines Paris, PSL University , 60 bd Saint-Michel, 75272 Paris,","place":["France"]},{"name":"Institut Curie, PSL University, 11 rue Pierre et Marie Curie , 75005 Paris,","place":["France"]},{"name":"U1331, INSERM, 11 rue Pierre et Marie Curie , 75005 Paris,","place":["France"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5855-0935","authenticated-orcid":false,"given":"Florian","family":"Massip","sequence":"additional","affiliation":[{"name":"Centre for Computational Biology, Mines Paris, PSL University , 60 bd Saint-Michel, 75272 Paris,","place":["France"]},{"name":"Institut Curie, PSL University, 11 rue Pierre et Marie Curie , 75005 Paris,","place":["France"]},{"name":"U1331, INSERM, 11 rue Pierre et Marie Curie , 75005 Paris,","place":["France"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2026,5,19]]},"reference":[{"key":"2026051908091418600_ref1","doi-asserted-by":"publisher","DOI":"10.1126\/science.aan2507","article-title":"A pathology atlas of the human cancer transcriptome","volume":"357","author":"Uhlen","year":"2017","journal-title":"Science"},{"key":"2026051908091418600_ref2","doi-asserted-by":"publisher","DOI":"10.1126\/scitranslmed.aba4448","article-title":"Transcriptomic profiling across the nonalcoholic fatty liver disease spectrum reveals gene signatures for steatohepatitis and fibrosis","volume":"12","author":"Govaere","year":"2020","journal-title":"Sci Transl Med"},{"key":"2026051908091418600_ref3","doi-asserted-by":"publisher","first-page":"97","DOI":"10.1186\/s13059-018-1481-6","article-title":"The healthy ageing gene expression signature for Alzheimer\u2019s disease diagnosis: a random sampling perspective","volume":"19","author":"Jacob","year":"2018","journal-title":"Genome Biol"},{"key":"2026051908091418600_ref4","doi-asserted-by":"publisher","first-page":"171","DOI":"10.1093\/bioinformatics\/bth469","article-title":"Outcome signature genes in breast cancer: is there a unique set?","volume":"21","author":"Ein-Dor","year":"2004","journal-title":"Bioinformatics"},{"key":"2026051908091418600_ref5","doi-asserted-by":"publisher","first-page":"e28210","DOI":"10.1371\/journal.pone.0028210","article-title":"The influence of feature selection methods on accuracy, stability and interpretability of molecular signatures","volume":"6","author":"Haury","year":"2011","journal-title":"PLoS One"},{"key":"2026051908091418600_ref6","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1038\/nrc2294","article-title":"The properties of high-dimensional data spaces: implications for exploring gene and protein expression data","volume":"8","author":"Clarke","year":"2008","journal-title":"Nat Rev Cancer"},{"key":"2026051908091418600_ref7","doi-asserted-by":"publisher","first-page":"4237","DOI":"10.1098\/rsta.2009.0159","article-title":"Statistical challenges of high-dimensional data","volume":"367","author":"Johnstone","year":"2009","journal-title":"Philos Trans A Math Phys Eng Sci"},{"key":"2026051908091418600_ref8","doi-asserted-by":"publisher","first-page":"2055","DOI":"10.1214\/15-AOS1337","article-title":"Controlling the false discovery rate via knockoffs","volume":"43","author":"Barber","year":"2015","journal-title":"Ann Stat"},{"key":"2026051908091418600_ref9","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1093\/biomet\/asy033","article-title":"Gene hunting with knockoffs for hidden markov models","volume":"106","author":"Sesia","year":"2019","journal-title":"Biometrika"},{"key":"2026051908091418600_ref10","doi-asserted-by":"publisher","first-page":"3152","DOI":"10.1038\/s41467-021-22889-4","article-title":"Identification of putative causal loci in whole-genome sequencing data via knockoff statistics","volume":"12","author":"He","year":"2021","journal-title":"Nat Commun"},{"key":"2026051908091418600_ref11","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.2105841118","article-title":"False discovery rate control in genome-wide association studies with population structure","volume":"118","author":"Sesia","year":"2021","journal-title":"Proc Natl Acad Sci"},{"key":"2026051908091418600_ref12","doi-asserted-by":"publisher","first-page":"744","DOI":"10.3390\/cancers11060744","article-title":"False discovery rate control in cancer biomarker selection using knockoffs","volume":"11","author":"Shen","year":"2019","journal-title":"Cancers"},{"key":"2026051908091418600_ref13","doi-asserted-by":"publisher","first-page":"976","DOI":"10.1093\/bioinformatics\/btaa770","article-title":"Knockoff boosted tree for model-free variable selection","volume":"37","author":"Jiang","year":"2021","journal-title":"Bioinformatics"},{"key":"2026051908091418600_ref14","doi-asserted-by":"crossref","DOI":"10.52202\/075280-3417","article-title":"False discovery proportion control for aggregated knockoffs","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Blain","year":"2023;  ,"},{"key":"2026051908091418600_ref15","first-page":"7283","article-title":"Aggregation of multiple knockoffs","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Nguyen","year":"2020"},{"key":"2026051908091418600_ref16","doi-asserted-by":"publisher","first-page":"948","DOI":"10.1080\/01621459.2021.1962720","article-title":"Derandomizing knockoffs","volume":"118","author":"Ren","year":"2021","journal-title":"J Am Stat Assoc"},{"key":"2026051908091418600_ref17","doi-asserted-by":"crossref","first-page":"122","DOI":"10.1093\/jrsssb\/qkad085","article-title":"Derandomised knockoffs: leveraging e-values for false discovery rate control","volume":"86","author":"Ren","year":"2023","journal-title":"J R Stat Soc Series B Stat Methodol"},{"key":"2026051908091418600_ref18","doi-asserted-by":"publisher","first-page":"551","DOI":"10.1111\/rssb.12265","article-title":"Panning for gold: \u2018model-x\u2019 knockoffs for high dimensional controlled variable selection","volume":"80","author":"Cand\u00e8s","year":"2018","journal-title":"J R Stat Soc Series B Stat Methodol"},{"key":"2026051908091418600_ref19","article-title":"Power analysis of knockoff filters for correlated designs","volume":"32","author":"Liu","year":"2019","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026051908091418600_ref20","doi-asserted-by":"publisher","first-page":"252","DOI":"10.1214\/21-AOS2104","article-title":"Powerful knockoffs via minimizing reconstructability","volume":"50","author":"Spector","year":"2022","journal-title":"Ann Stat"},{"key":"2026051908091418600_ref21","first-page":"1","article-title":"Power of knockoff: the impact of ranking algorithm, augmented design, and symmetric statistic","volume":"25","author":"Ke","year":"2024","journal-title":"J Mach Learn Res"},{"key":"2026051908091418600_ref22","first-page":"06892","article-title":"When knockoffs fail: diagnosing and fixing non-exchangeability of knockoffs","volume":"2407","author":"Blain","year":"2024","journal-title":"arXiv"},{"key":"2026051908091418600_ref23","article-title":"KnockoffGAN: generating knockoffs for feature selection using generative adversarial networks","volume-title":"International Conference on Learning Representations (ICLR)","author":"Jordon","year":"2019;  ,"},{"key":"2026051908091418600_ref24","doi-asserted-by":"publisher","first-page":"1861","DOI":"10.1080\/01621459.2019.1660174","article-title":"Deep knockoffs","volume":"115","author":"Romano","year":"2020","journal-title":"J Am Stat Assoc"},{"key":"2026051908091418600_ref25","article-title":"DeepPINK: reproducible feature selection in deep neural networks","volume":"31","author":"Lu","year":"2018","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026051908091418600_ref26","first-page":"08766","article-title":"Asymptotically optimal knockoff statistics via the masked likelihood ratio","volume":"2212","author":"Spector","year":"2022","journal-title":"arXiv"},{"key":"2026051908091418600_ref27","doi-asserted-by":"publisher","first-page":"54","DOI":"10.1186\/s13073-024-01317-4","article-title":"Smoking-associated gene expression alterations in nasal epithelium reveal immune impairment linked to lung cancer risk","volume":"16","author":"De Biase","year":"2024","journal-title":"Genome Med"},{"key":"2026051908091418600_ref28","doi-asserted-by":"publisher","DOI":"10.1093\/jnci\/djw327","article-title":"Shared gene expression alterations in nasal and bronchial epithelium for lung cancer detection","volume":"109","author":"Perez-Rogers","year":"2017","journal-title":"Natl Cancer"},{"key":"2026051908091418600_ref29","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1200\/PO.17.00135","article-title":"Clinical value of rna sequencing\u2013based classifiers for prediction of the five conventional breast cancer biomarkers: a report from the population-based multicenter Sweden cancerome analysis network\u2014Breast initiative","volume":"2","author":"Brueffer","year":"2018","journal-title":"JCO Precis Oncol"},{"key":"2026051908091418600_ref30","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18637\/jss.v033.i01","article-title":"Regularization paths for generalized linear models via coordinate descent","volume":"33","author":"Friedman","year":"2010","journal-title":"J Stat Softw"},{"key":"2026051908091418600_ref31","doi-asserted-by":"publisher","first-page":"79","DOI":"10.1186\/s13059-022-02648-4","article-title":"Exaggerated false positives by popular differential expression methods when analyzing human population samples","volume":"23","author":"Li","year":"2022","journal-title":"Genome Biol"},{"key":"2026051908091418600_ref32","doi-asserted-by":"publisher","first-page":"1166","DOI":"10.1214\/14-AOS1221","article-title":"On asymptotically optimal confidence regions and tests for high-dimensional models","volume":"42","author":"Van De Geer","year":"2014","journal-title":"Ann Stat"},{"key":"2026051908091418600_ref33","doi-asserted-by":"publisher","first-page":"289","DOI":"10.1111\/j.2517-6161.1995.tb02031.x","article-title":"Controlling the false discovery rate: a practical and powerful approach to multiple testing","volume":"57","author":"Benjamini","year":"1995","journal-title":"J R Stat Soc B Methodol"},{"key":"2026051908091418600_ref34","doi-asserted-by":"publisher","first-page":"417","DOI":"10.1111\/j.1467-9868.2010.00740.x","article-title":"Stability selection","volume":"72","author":"Meinshausen","year":"2010","journal-title":"J R Stat Soc Series B Stat Methodol"},{"key":"2026051908091418600_ref35","doi-asserted-by":"publisher","first-page":"55","DOI":"10.1111\/j.1467-9868.2011.01034.x","article-title":"Variable selection with error control: another look at stability selection","volume":"75","author":"Shah","year":"2013","journal-title":"J R Stat Soc Series B Stat Methodol"},{"key":"2026051908091418600_ref36","doi-asserted-by":"publisher","first-page":"281","DOI":"10.1186\/s13059-024-03231-9","article-title":"Neglecting the impact of normalization in semi-synthetic RNA-seq data simulations generates artificial false positives","volume":"25","author":"Hejblum","year":"2024","journal-title":"Genome Biol"},{"key":"2026051908091418600_ref37","doi-asserted-by":"publisher","first-page":"73","DOI":"10.1214\/12-BA703","article-title":"Scalable variational inference for bayesian variable selection in regression, and its accuracy in genetic association studies","volume":"7","author":"Carbonetto","year":"2012","journal-title":"Bayesian Anal"},{"key":"2026051908091418600_ref38","doi-asserted-by":"publisher","first-page":"88","DOI":"10.1093\/bioinformatics\/bti736","article-title":"Gene selection using support vector machines with non-convex penalty","volume":"22","author":"Zhang","year":"2005","journal-title":"Bioinformatics"},{"key":"2026051908091418600_ref39","doi-asserted-by":"publisher","first-page":"334","DOI":"10.1016\/j.cell.2020.11.045","article-title":"A modular master regulator landscape controls cancer transcriptional identity","volume":"184","author":"Paull","year":"2021","journal-title":"Cell"},{"key":"2026051908091418600_ref40","first-page":"1935","article-title":"Feature screening with kernel knockoffs","volume-title":"International Conference on Artificial Intelligence and Statistics","author":"Poignard","year":"2022"}],"container-title":["Briefings in Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bib\/article-pdf\/27\/3\/bbag148\/68334744\/bbag148.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bib\/article-pdf\/27\/3\/bbag148\/68334744\/bbag148.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T12:09:25Z","timestamp":1779192565000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bib\/article\/doi\/10.1093\/bib\/bbag148\/8687371"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5]]},"references-count":40,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,5,4]]}},"URL":"https:\/\/doi.org\/10.1093\/bib\/bbag148","relation":{},"ISSN":["1467-5463","1477-4054"],"issn-type":[{"value":"1467-5463","type":"print"},{"value":"1477-4054","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026,5]]},"published":{"date-parts":[[2026,5]]},"article-number":"bbag148"}}