{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,7,17]],"date-time":"2024-07-17T22:32:03Z","timestamp":1721255523450},"reference-count":38,"publisher":"Oxford University Press (OUP)","issue":"Supplement_2","license":[{"start":{"date-parts":[[2020,12,31]],"date-time":"2020-12-31T00:00:00Z","timestamp":1609372800000},"content-version":"vor","delay-in-days":30,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Strategic Focal Area \u2018Personalized Health and Related Technologies","award":["#2017-110"],"award-info":[{"award-number":["#2017-110"]}]},{"name":"Personalized Swiss Sepsis Study"},{"name":"Alfried Krupp Prize for Young"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,12,30]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>Temporal biomarker discovery in longitudinal data is based on detecting reoccurring trajectories, the so-called shapelets. The search for shapelets requires considering all subsequences in the data. While the accompanying issue of multiple testing has been mitigated in previous work, the redundancy and overlap of the detected shapelets results in an a priori unbounded number of highly similar and structurally meaningless shapelets. As a consequence, current temporal biomarker discovery methods are impractical and underpowered.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>We find that the pre- or post-processing of shapelets does not sufficiently increase the power and practical utility. Consequently, we present a novel method for temporal biomarker discovery: Statistically Significant Submodular Subset Shapelet Mining (S5M) that retrieves short subsequences that are (i) occurring in the data, (ii) are statistically significantly associated with the phenotype and (iii) are of manageable quantity while maximizing structural diversity. Structural diversity is achieved by pruning non-representative shapelets via submodular optimization. This increases the statistical power and utility of S5M compared to state-of-the-art approaches on simulated and real-world datasets. For patients admitted to the intensive care unit (ICU) showing signs of severe organ failure, we find temporal patterns in the sequential organ failure assessment score that are associated with in-ICU mortality.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation<\/jats:title>\n                  <jats:p>S5M is an option in the python package of S3M: github.com\/BorgwardtLab\/S3M.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btaa815","type":"journal-article","created":{"date-parts":[[2020,10,20]],"date-time":"2020-10-20T19:13:55Z","timestamp":1603221235000},"page":"i840-i848","source":"Crossref","is-referenced-by-count":2,"title":["Enhancing statistical power in temporal biomarker discovery through representative shapelet mining"],"prefix":"10.1093","volume":"36","author":[{"given":"Thomas","family":"Gumbsch","sequence":"first","affiliation":[{"name":"Department of Biosystems Science and Engineering, ETH Zurich , Basel 4058, Switzerland"},{"name":"SIB Swiss Institute of Bioinformatics , Lausanne 1015, Switzerland"}]},{"given":"Christian","family":"Bock","sequence":"additional","affiliation":[{"name":"Department of Biosystems Science and Engineering, ETH Zurich , Basel 4058, Switzerland"},{"name":"SIB Swiss Institute of Bioinformatics , Lausanne 1015, Switzerland"}]},{"given":"Michael","family":"Moor","sequence":"additional","affiliation":[{"name":"Department of Biosystems Science and Engineering, ETH Zurich , Basel 4058, Switzerland"},{"name":"SIB Swiss Institute of Bioinformatics , Lausanne 1015, Switzerland"}]},{"given":"Bastian","family":"Rieck","sequence":"additional","affiliation":[{"name":"Department of Biosystems Science and Engineering, ETH Zurich , Basel 4058, Switzerland"},{"name":"SIB Swiss Institute of Bioinformatics , Lausanne 1015, Switzerland"}]},{"given":"Karsten","family":"Borgwardt","sequence":"additional","affiliation":[{"name":"Department of Biosystems Science and Engineering, ETH Zurich , Basel 4058, Switzerland"},{"name":"SIB Swiss Institute of Bioinformatics , Lausanne 1015, Switzerland"}]}],"member":"286","published-online":{"date-parts":[[2020,12,29]]},"reference":[{"key":"2023062409332572900_btaa815-B1","volume-title":"Oh\u2019s Intensive Care Manual E-Book","author":"Bersten","year":"2013"},{"key":"2023062409332572900_btaa815-B2","doi-asserted-by":"crossref","first-page":"i438","DOI":"10.1093\/bioinformatics\/bty246","article-title":"Association mapping in biomedical time series via statistically significant shapelet mining","volume":"34","author":"Bock","year":"2018","journal-title":"Bioinformatics"},{"key":"2023062409332572900_btaa815-B3","first-page":"3","article-title":"Teoria statistica delle classi e calcolo delle probabilit\u00e0","volume":"8","author":"Bonferroni","year":"1936","journal-title":"Pubblicazioni Del R. Istituto Superiore di Scienze Economiche e Commerciali di Firenze"},{"key":"2023062409332572900_btaa815-B4","year":"2019"},{"key":"2023062409332572900_btaa815-B5","first-page":"497","author":"Fang","year":"2018"},{"key":"2023062409332572900_btaa815-B6","doi-asserted-by":"crossref","first-page":"1754","DOI":"10.1001\/jama.286.14.1754","article-title":"Serial evaluation of the sofa score to predict outcome in critically ill patients","volume":"286","author":"Ferreira","year":"2001","journal-title":"JAMA"},{"key":"2023062409332572900_btaa815-B7","volume-title":"Submodular Functions and Optimization","author":"Fujishige","year":"2005"},{"key":"2023062409332572900_btaa815-B8","first-page":"201","author":"Ghalwash","year":"2013"},{"key":"2023062409332572900_btaa815-B9","first-page":"965","author":"Gharghabi","year":"2018"},{"key":"2023062409332572900_btaa815-B10","doi-asserted-by":"crossref","first-page":"409","DOI":"10.1002\/pro.5560010313","article-title":"Selection of representative protein data sets","volume":"1","author":"Hobohm","year":"1992","journal-title":"Protein Sci"},{"key":"2023062409332572900_btaa815-B11","doi-asserted-by":"crossref","first-page":"364","DOI":"10.1038\/s41591-020-0789-4","article-title":"Early prediction of circulatory failure in the intensive care unit using machine learning","volume":"26","author":"Hyland","year":"2020","journal-title":"Nat. Med"},{"key":"2023062409332572900_btaa815-B12","first-page":"382","author":"Imani","year":"2018"},{"key":"2023062409332572900_btaa815-B13","doi-asserted-by":"crossref","first-page":"160035","DOI":"10.1038\/sdata.2016.35","article-title":"MIMIC-III, a freely accessible critical care database","volume":"3","author":"Johnson","year":"2016","journal-title":"Sci. Data"},{"key":"2023062409332572900_btaa815-B14","first-page":"32","article-title":"The MIMIC code repository: enabling reproducibility in critical care research","volume":"25","author":"Johnson","year":"2018","journal-title":"JAMIA"},{"key":"2023062409332572900_btaa815-B15","doi-asserted-by":"crossref","first-page":"1053","DOI":"10.1007\/s10618-016-0473-y","article-title":"Generalized random shapelet forests","volume":"30","author":"Karlsson","year":"2016","journal-title":"Data Min. Knowl. Disc"},{"key":"2023062409332572900_btaa815-B16","first-page":"154","article-title":"Clustering of time-series subsequences is meaningless: implications for previous and future research","volume":"8","author":"Keogh","year":"2005","journal-title":"KAIS"},{"key":"2023062409332572900_btaa815-B17","doi-asserted-by":"crossref","first-page":"454","DOI":"10.1002\/prot.25461","article-title":"Choosing non-redundant representative subsets of protein sequence data sets using submodular optimization","volume":"86","author":"Libbrecht","year":"2018","journal-title":"Proteins Struct. Funct. Bioinf"},{"key":"2023062409332572900_btaa815-B18","author":"Lin","year":"2009"},{"key":"2023062409332572900_btaa815-B19","doi-asserted-by":"crossref","first-page":"2680","DOI":"10.1093\/bioinformatics\/bty1020","article-title":"CASMAP: detection of statistically significant combinations of SNPs in association mapping","volume":"35","author":"Llinares-L\u00f3pez","year":"2019","journal-title":"Bioinformatics"},{"key":"2023062409332572900_btaa815-B20","author":"McGinley","year":"2012"},{"key":"2023062409332572900_btaa815-B21","first-page":"473","author":"Mueen","year":"2009"},{"key":"2023062409332572900_btaa815-B22","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1007\/BF01588971","article-title":"An analysis of approximations for maximizing submodular set functions\u2014I","volume":"14","author":"Nemhauser","year":"1978","journal-title":"Math. Program"},{"key":"2023062409332572900_btaa815-B23","first-page":"2279","author":"Papaxanthos","year":"2016"},{"key":"2023062409332572900_btaa815-B24","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1080\/14786440009463897","article-title":"X. On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can\u00a0be reasonably supposed to have arisen from random sampling","volume":"50","author":"Pearson","year":"1900","journal-title":"Lond. Edinb. Dubl. Phil. Mag"},{"key":"2023062409332572900_btaa815-B25","article-title":"The EICU collaborative research database, a freely available multi-center database for critical care research","volume":"5, 180178","author":"Pollard","year":"2018","journal-title":"Sci. Data"},{"key":"2023062409332572900_btaa815-B26","doi-asserted-by":"crossref","DOI":"10.1017\/9781108377706","volume-title":"Analyzing Network Data in Biology and Medicine: An Interdisciplinary Textbook for Biological, Medical and Computational Scientists","author":"Pr\u017eulj","year":"2019"},{"key":"2023062409332572900_btaa815-B27","first-page":"262","author":"Rakthanmanon","year":"2012"},{"key":"2023062409332572900_btaa815-B28","first-page":"61","author":"Seabold","year":"2010"},{"key":"2023062409332572900_btaa815-B29","doi-asserted-by":"crossref","first-page":"561","DOI":"10.1146\/annurev.ps.46.020195.003021","article-title":"Multiple hypothesis testing","volume":"46","author":"Shaffer","year":"1995","journal-title":"Annu. Rev. Psychol"},{"key":"2023062409332572900_btaa815-B30","doi-asserted-by":"crossref","first-page":"801","DOI":"10.1001\/jama.2016.0287","article-title":"The third international consensus definitions for sepsis and septic shock (sepsis-3)","volume":"315","author":"Singer","year":"2016","journal-title":"JAMA"},{"key":"2023062409332572900_btaa815-B31","doi-asserted-by":"crossref","first-page":"515","DOI":"10.2307\/2531456","article-title":"A modified Bonferroni method for discrete data","volume":"46","author":"Tarone","year":"1990","journal-title":"Biometrics"},{"key":"2023062409332572900_btaa815-B32","doi-asserted-by":"crossref","first-page":"e9654","DOI":"10.1097\/MD.0000000000009654","article-title":"Serial evaluation of the sofa score is reliable for predicting mortality in acute severe pancreatitis","volume":"97","author":"Tee","year":"2018","journal-title":"Medicine"},{"key":"2023062409332572900_btaa815-B33","doi-asserted-by":"crossref","first-page":"12996","DOI":"10.1073\/pnas.1302233110","article-title":"Statistical significance of combinatorial regulations","volume":"110","author":"Terada","year":"2013","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"2023062409332572900_btaa815-B34","author":"Vincent","year":"1996"},{"key":"2023062409332572900_btaa815-B35","first-page":"1954","author":"Wei","year":"2015"},{"key":"2023062409332572900_btaa815-B36","first-page":"28","article-title":"The generalization of student\u2019s\u2019 problem when several different population variances are involved","volume":"34","author":"Welch","year":"1947","journal-title":"Biometrika"},{"key":"2023062409332572900_btaa815-B37","first-page":"947","author":"Ye","year":"2009"},{"key":"2023062409332572900_btaa815-B38","author":"Yilmaz","year":"2019"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/36\/Supplement_2\/i840\/50693587\/btaa815.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/36\/Supplement_2\/i840\/50693587\/btaa815.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,24]],"date-time":"2023-06-24T23:59:13Z","timestamp":1687651153000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/36\/Supplement_2\/i840\/6055899"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,12]]},"references-count":38,"journal-issue":{"issue":"Supplement_2","published-print":{"date-parts":[[2020,12,30]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btaa815","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2020,12]]},"published":{"date-parts":[[2020,12]]}}}