{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,28]],"date-time":"2025-10-28T03:15:13Z","timestamp":1761621313279,"version":"3.37.3"},"reference-count":34,"publisher":"Oxford University Press (OUP)","issue":"18","license":[{"start":{"date-parts":[[2017,4,18]],"date-time":"2017-04-18T00:00:00Z","timestamp":1492473600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/about_us\/legal\/notices"}],"funder":[{"DOI":"10.13039\/100000185","name":"Defense Advanced Research Projects Agency","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000185","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000182","name":"MRMC","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000182","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000183","name":"Army Research Office","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000183","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000183","name":"ARO","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000183","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000005","name":"Department of Defense","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000005","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2017,9,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>This work addresses two common issues in building classification models for biological or medical studies: learning a sparse model, where only a subset of a large number of possible predictors is used, and training in the presence of missing data. This work focuses on supervised generative binary classification models, specifically linear discriminant analysis (LDA). The parameters are determined using an expectation maximization algorithm to both address missing data and introduce priors to promote sparsity. The proposed algorithm, expectation-maximization sparse discriminant analysis (EM-SDA), produces a sparse LDA model for datasets with and without missing data.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>EM-SDA is tested via simulations and case studies. In the simulations, EM-SDA is compared with nearest shrunken centroids (NSCs) and sparse discriminant analysis (SDA) with k-nearest neighbors for imputation for varying mechanism and amount of missing data. In three case studies using published biomedical data, the results are compared with NSC and SDA models with four different types of imputation, all of which are common approaches in the field. EM-SDA is more accurate and sparse than competing methods both with and without missing data in most of the experiments. Furthermore, the EM-SDA results are mostly consistent between the missing and full cases. Biological relevance of the resulting models, as quantified via a literature search, is also presented.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation<\/jats:title>\n                  <jats:p>A Matlab implementation published under GNU GPL v.3 license is available at http:\/\/web.mit.edu\/braatzgroup\/links.html.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information<\/jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btx224","type":"journal-article","created":{"date-parts":[[2017,4,13]],"date-time":"2017-04-13T11:09:27Z","timestamp":1492081767000},"page":"2897-2905","source":"Crossref","is-referenced-by-count":11,"title":["A method for learning a sparse classifier in the presence of missing data for high-dimensional biological datasets"],"prefix":"10.1093","volume":"33","author":[{"given":"Kristen A","family":"Severson","sequence":"first","affiliation":[{"name":"Department of Chemical Engineering, Massachusetts Institute of Technology, Cambridge, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Brinda","family":"Monian","sequence":"additional","affiliation":[{"name":"Department of Chemical Engineering, Massachusetts Institute of Technology, Cambridge, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"J Christopher","family":"Love","sequence":"additional","affiliation":[{"name":"Department of Chemical Engineering, Massachusetts Institute of Technology, Cambridge, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Richard D","family":"Braatz","sequence":"additional","affiliation":[{"name":"Department of Chemical Engineering, Massachusetts Institute of Technology, Cambridge, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2017,4,18]]},"reference":[{"volume-title":"Pattern Recognition and Machine Learning","year":"2007","author":"Bishop","key":"2023020301105216300_btx224-B1"},{"key":"2023020301105216300_btx224-B2","doi-asserted-by":"crossref","first-page":"475","DOI":"10.1089\/cmb.2008.0078","article-title":"A model-based approach to gene clustering with missing observation reconstruction in a Markov random field framework","volume":"16","author":"Blanchet","year":"2009","journal-title":"J. Comput. Biol"},{"key":"2023020301105216300_btx224-B3","doi-asserted-by":"crossref","first-page":"e34","DOI":"10.1093\/nar\/gnh026","article-title":"LSimpute: Accurate estimation of missing values in microarray data with least squares methods","volume":"32","author":"B\u00f8","year":"2004","journal-title":"Nucleic Acids Res"},{"key":"2023020301105216300_btx224-B4","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1186\/1471-2105-9-12","article-title":"Which missing value imputation method to use in expression profiles: A comparative study and two selection schemes","volume":"9","author":"Brock","year":"2008","journal-title":"BMC Bioinformatics"},{"key":"2023020301105216300_btx224-B5","doi-asserted-by":"crossref","first-page":"406","DOI":"10.1198\/TECH.2011.08118","article-title":"Sparse discriminant analysis","volume":"53","author":"Clemmensen","year":"2011","journal-title":"Technometrics"},{"key":"2023020301105216300_btx224-B6","first-page":"1","article-title":"Maximum likelihood from incomplete data via the EM algorithm","volume":"39","author":"Dempster","year":"1977","journal-title":"J. Roy. Stat. Soc. B"},{"key":"2023020301105216300_btx224-B8","doi-asserted-by":"crossref","first-page":"1150","DOI":"10.1109\/TPAMI.2003.1227989","article-title":"Adaptive sparseness for supervised learning","volume":"25","author":"Figueiredo","year":"2003","journal-title":"IEEE T. Pattern Anal"},{"key":"2023020301105216300_btx224-B9","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1007\/s00521-009-0295-6","article-title":"Pattern classification with missing data: a review","volume":"19","author":"Garc\u00eda-Laencina","year":"2010","journal-title":"Neural Comput. Appl"},{"key":"2023020301105216300_btx224-B10","doi-asserted-by":"crossref","first-page":"531","DOI":"10.1126\/science.286.5439.531","article-title":"Molecular classification of cancer: Class discovery and class prediction by gene expression monitoring","volume":"286","author":"Golub","year":"1999","journal-title":"Science"},{"key":"2023020301105216300_btx224-B11","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1371\/journal.pone.0129126","article-title":"Self-organizing feature maps identify proteins critical to learning in a mouse model of Down syndrome","volume":"10","author":"Higuera","year":"2015","journal-title":"PLoS One"},{"key":"2023020301105216300_btx224-B12","doi-asserted-by":"crossref","first-page":"417","DOI":"10.1037\/h0071325","article-title":"Analysis of a complex of statistical variables into principal components","volume":"24","author":"Hotelling","year":"1933","journal-title":"J. Educ. Psychol"},{"key":"2023020301105216300_btx224-B13","first-page":"1957","article-title":"Practical approaches to principal component analysis in the presence of missing values","volume":"11","author":"Ilin","year":"2010","journal-title":"J. Mach. Learn. Res"},{"key":"2023020301105216300_btx224-B14","doi-asserted-by":"crossref","first-page":"187","DOI":"10.1093\/bioinformatics\/bth499","article-title":"Missing value estimation for DNA microarray gene expression data: local least squares imputation","volume":"21","author":"Kim","year":"2004","journal-title":"Bioinformatics"},{"key":"2023020301105216300_btx224-B15","doi-asserted-by":"crossref","first-page":"1495","DOI":"10.1093\/bioinformatics\/btm134","article-title":"Sparse non-negative matrix factorizations via alternating non-negativity-constrained least squares for microarray data analysis","volume":"23","author":"Kim","year":"2007","journal-title":"Bioinformatics"},{"key":"2023020301105216300_btx224-B16","doi-asserted-by":"crossref","DOI":"10.1002\/9781119013563","volume-title":"Statisical Analysis with Missing Data","author":"Little","year":"2002","edition":"2nd edn."},{"year":"2008","author":"Marlin","key":"2023020301105216300_btx224-B17"},{"volume-title":"Machine Learning: A Probabilistic Perspective","year":"2012","author":"Murphy","key":"2023020301105216300_btx224-B18"},{"key":"2023020301105216300_btx224-B19","doi-asserted-by":"crossref","first-page":"2088","DOI":"10.1093\/bioinformatics\/btg287","article-title":"A Bayesian missing value estimation method for gene expression profile data","volume":"19","author":"Oba","year":"2003","journal-title":"Bioinformatics"},{"key":"2023020301105216300_btx224-B20","doi-asserted-by":"crossref","first-page":"917","DOI":"10.1093\/bioinformatics\/bth007","article-title":"Gaussian mixture clustering and imputation of microarray data","volume":"20","author":"Ouyang","year":"2004","journal-title":"Bioinformatics"},{"key":"2023020301105216300_btx224-B21","doi-asserted-by":"crossref","first-page":"681","DOI":"10.1198\/016214508000000337","article-title":"The Bayesian lasso","volume":"103","author":"Park","year":"2008","journal-title":"J. Am. Stat. Assoc"},{"key":"2023020301105216300_btx224-B22","doi-asserted-by":"crossref","first-page":"559","DOI":"10.1080\/14786440109462720","article-title":"On lines and planes of closest fit to systems of points in space","volume":"2","author":"Pearson","year":"1901","journal-title":"Philos. Mag"},{"key":"2023020301105216300_btx224-B23","doi-asserted-by":"crossref","first-page":"2066","DOI":"10.1182\/blood-2006-02-002477","article-title":"Gene expression patterns in blood leukocytes discriminate patients with acute infections","volume":"109","author":"Ramilo","year":"2007","journal-title":"Blood"},{"key":"2023020301105216300_btx224-B24","first-page":"626","article-title":"EM algorithms for PCA and SPCA","author":"Roweis","year":"1998","journal-title":"Adv. Neur. Inf"},{"key":"2023020301105216300_btx224-B25","doi-asserted-by":"crossref","first-page":"581","DOI":"10.1093\/biomet\/63.3.581","article-title":"Inference and missing data","volume":"63","author":"Rubin","year":"1976","journal-title":"Biometrika"},{"year":"2003","author":"Salakhutdinov","key":"2023020301105216300_btx224-B26"},{"key":"2023020301105216300_btx224-B27","doi-asserted-by":"crossref","first-page":"2417","DOI":"10.1093\/bioinformatics\/bti345","article-title":"Collateral missing value imputation: A new robust missing value estimation algorithm for microarray data","volume":"21","author":"Sehgal","year":"2005","journal-title":"Bioinformatics"},{"volume-title":"SpaSM: A Matlab Toolbox for Sparse Statistical Modeling","year":"2012","author":"Sj\u00f6strand","key":"2023020301105216300_btx224-B28"},{"key":"2023020301105216300_btx224-B29","doi-asserted-by":"crossref","first-page":"6567","DOI":"10.1073\/pnas.082099299","article-title":"Diagnosis of multiple cancer types by shrunken centroids of gene expression","volume":"99","author":"Tibshirani","year":"2002","journal-title":"P. Natl. Acad. Sci. USA"},{"key":"2023020301105216300_btx224-B30","doi-asserted-by":"crossref","first-page":"611","DOI":"10.1111\/1467-9868.00196","article-title":"Probabilistic principal component analysis","volume":"61","author":"Tipping","year":"1999","journal-title":"J. Roy. Stat. Soc. B"},{"key":"2023020301105216300_btx224-B31","doi-asserted-by":"crossref","first-page":"520","DOI":"10.1093\/bioinformatics\/17.6.520","article-title":"Missing value estimation methods for DNA microarrays","volume":"17","author":"Troyanskaya","year":"2001","journal-title":"Bioinformatics"},{"key":"2023020301105216300_btx224-B32","doi-asserted-by":"crossref","first-page":"972","DOI":"10.1093\/bioinformatics\/btm046","article-title":"Improved centroids estimation for the nearest strunken centroid classifier","volume":"23","author":"Wang","year":"2007","journal-title":"Bioinformatics"},{"key":"2023020301105216300_btx224-B33","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1186\/1471-2105-7-32","article-title":"Missing value estimation for DNA microarray gene expression data by support vector regression imputation and orthogonal coding scheme","volume":"7","author":"Wang","year":"2006","journal-title":"BMC Bioinformatics"},{"key":"2023020301105216300_btx224-B34","doi-asserted-by":"crossref","first-page":"753","DOI":"10.1111\/j.1467-9868.2011.00783.x","article-title":"Penalized classification using Fisher's linear discriminant","volume":"73","author":"Witten","year":"2011","journal-title":"J. Roy. Stat. Soc. B"},{"key":"2023020301105216300_btx224-B35","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1016\/j.jsb.2010.04.002","article-title":"Probabilistic principal component analysis with expectation maximization (PPCA-EM) facilitates volume classification and estimates the missing data","volume":"171","author":"Yu","year":"2012","journal-title":"J. Struct. Biol"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/33\/18\/2897\/49041187\/bioinformatics_33_18_2897.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/33\/18\/2897\/49041187\/bioinformatics_33_18_2897.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,3]],"date-time":"2023-02-03T01:12:05Z","timestamp":1675386725000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/33\/18\/2897\/3738496"}},"subtitle":[],"editor":[{"given":"Jonathan","family":"Wren","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2017,4,18]]},"references-count":34,"journal-issue":{"issue":"18","published-print":{"date-parts":[[2017,9,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btx224","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"type":"print","value":"1367-4803"},{"type":"electronic","value":"1367-4811"}],"subject":[],"published-other":{"date-parts":[[2017,9,15]]},"published":{"date-parts":[[2017,4,18]]}}}