{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,22]],"date-time":"2025-02-22T00:45:25Z","timestamp":1740185125366,"version":"3.37.3"},"reference-count":18,"publisher":"Oxford University Press (OUP)","issue":"3","license":[{"start":{"date-parts":[[2018,7,13]],"date-time":"2018-07-13T00:00:00Z","timestamp":1531440000000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Centre for Mathematical Sciences"},{"DOI":"10.13039\/501100003252","name":"Lund University","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100003252","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2019,2,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>High throughput biomedical measurements normally capture multiple overlaid biologically relevant signals and often also signals representing different types of technical artefacts like e.g. batch effects. Signal identification and decomposition are accordingly main objectives in statistical biomedical modeling and data analysis. Existing methods, aimed at signal reconstruction and deconvolution, in general, are either supervised, contain parameters that need to be estimated or present other types of ad hoc features. We here introduce SubMatrix Selection Singular Value Decomposition (SMSSVD), a parameter-free unsupervised signal decomposition and dimension reduction method, designed to reduce noise, adaptively for each low-rank-signal in a given data matrix, and represent the signals in the data in a way that enable unbiased exploratory analysis and reconstruction of multiple overlaid signals, including identifying groups of variables that drive different signals.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>The SMSSVD method produces a denoised signal decomposition from a given data matrix. It also guarantees orthogonality between signal components in a straightforward manner and it is designed to make automation possible. We illustrate SMSSVD by applying it to several real and synthetic datasets and compare its performance to golden standard methods like PCA (Principal Component Analysis) and SPC (Sparse Principal Components, using Lasso constraints). The SMSSVD is computationally efficient and despite being a parameter-free method, in general, outperforms existing statistical learning methods.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation<\/jats:title>\n                  <jats:p>A Julia implementation of SMSSVD is openly available on GitHub (https:\/\/github.com\/rasmushenningsson\/SubMatrixSelectionSVD.jl).<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information<\/jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/bty566","type":"journal-article","created":{"date-parts":[[2018,7,13]],"date-time":"2018-07-13T11:21:22Z","timestamp":1531480882000},"page":"478-486","source":"Crossref","is-referenced-by-count":4,"title":["SMSSVD: SubMatrix Selection Singular Value Decomposition"],"prefix":"10.1093","volume":"35","author":[{"given":"Rasmus","family":"Henningsson","sequence":"first","affiliation":[{"name":"The Centre for Mathematical Sciences, Lund University, Lund, Sweden"},{"name":"The International Group for Data Analysis, Institut Pasteur, Paris, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Magnus","family":"Fontes","sequence":"additional","affiliation":[{"name":"The Centre for Mathematical Sciences, Lund University, Lund, Sweden"},{"name":"The International Group for Data Analysis, Institut Pasteur, Paris, France"},{"name":"The Center for Genomic Medicine, Rigshospitalet, Copenhagen, Denmark"},{"name":"Persimune, The Centre of Excellence for Personalized Medicine, Copenhagen, Denmark"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2018,7,13]]},"reference":[{"key":"2023013107242832600_bty566-B1","doi-asserted-by":"crossref","first-page":"R106.","DOI":"10.1186\/gb-2010-11-10-r106","article-title":"Differential expression analysis for sequence count data","volume":"11","author":"Anders","year":"2010","journal-title":"Genome Biol"},{"key":"2023013107242832600_bty566-B2","doi-asserted-by":"crossref","first-page":"e108.","DOI":"10.1371\/journal.pbio.0020108","article-title":"Semi-supervised methods to predict patient survival from gene expression data","volume":"2","author":"Bair","year":"2004","journal-title":"PLoS Biol"},{"key":"2023013107242832600_bty566-B3","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1016\/j.ccr.2006.10.009","article-title":"Genomic and transcriptional aberrations linked to breast cancer pathophysiologies","volume":"10","author":"Chin","year":"2006","journal-title":"Cancer Cell"},{"key":"2023013107242832600_bty566-B4","doi-asserted-by":"crossref","first-page":"i350","DOI":"10.1093\/bioinformatics\/btx265","article-title":"Predicting phenotypes from microarrays using amplified, initially marginal, eigenvector regression","volume":"33","author":"Ding","year":"2017","journal-title":"Bioinformatics"},{"key":"2023013107242832600_bty566-B5","doi-asserted-by":"crossref","first-page":"307.","DOI":"10.1186\/1471-2105-12-307","article-title":"The projection score\u2014an evaluation criterion for variable subset selection in PCA visualization","volume":"12","author":"Fontes","year":"2011","journal-title":"BMC Bioinformatics"},{"key":"2023013107242832600_bty566-B6","first-page":"247346","article-title":"RNA-seq transcript quantification from reduced-representation data in recount2","author":"Fu","year":"2018","journal-title":"bioRxiv"},{"key":"2023013107242832600_bty566-B7","doi-asserted-by":"crossref","first-page":"research0003","DOI":"10.1186\/gb-2000-1-2-research0003","article-title":"Gene shaving\u2019as a method for identifying distinct sets of genes with similar expression patterns","volume":"1","author":"Hastie","year":"2000","journal-title":"Genome Biol"},{"key":"2023013107242832600_bty566-B8","doi-asserted-by":"crossref","first-page":"research0003","DOI":"10.1186\/gb-2001-2-1-research0003","article-title":"Supervised harvesting of expression trees","volume":"2","author":"Hastie","year":"2001","journal-title":"Genome Biol"},{"key":"2023013107242832600_bty566-B9","first-page":"327338","article-title":"DISSEQT-DIStribution based modeling of SEQuence space Time dynamics","author":"Henningsson","year":"2018","journal-title":"bioRxiv"},{"key":"2023013107242832600_bty566-B10","doi-asserted-by":"crossref","first-page":"417.","DOI":"10.1037\/h0071325","article-title":"Analysis of a complex of statistical variables into principal components","volume":"24","author":"Hotelling","year":"1933","journal-title":"J. Educ. Psychol"},{"key":"2023013107242832600_bty566-B11","doi-asserted-by":"crossref","first-page":"R36.","DOI":"10.1186\/gb-2013-14-4-r36","article-title":"Tophat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions","volume":"14","author":"Kim","year":"2013","journal-title":"Genome Biol"},{"key":"2023013107242832600_bty566-B12","doi-asserted-by":"crossref","first-page":"11790","DOI":"10.1038\/ncomms11790","article-title":"Identification of ETV6-RUNX1-like and DUX4-rearranged subtypes in paediatric B-cell precursor acute lymphoblastic leukaemia","volume":"7","author":"Lilljebj\u00f6rn","year":"2016","journal-title":"Nat. Commun"},{"key":"2023013107242832600_bty566-B13","doi-asserted-by":"crossref","first-page":"550.","DOI":"10.1186\/s13059-014-0550-8","article-title":"Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2","volume":"15","author":"Love","year":"2014","journal-title":"Genome Biol"},{"key":"2023013107242832600_bty566-B14","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Maaten","year":"2008","journal-title":"J. Mach. Learn. Res"},{"key":"2023013107242832600_bty566-B15","doi-asserted-by":"crossref","first-page":"R25.","DOI":"10.1186\/gb-2010-11-3-r25","article-title":"A scaling normalization method for differential expression analysis of RNA-seq data","volume":"11","author":"Robinson","year":"2010","journal-title":"Genome Biol"},{"key":"2023013107242832600_bty566-B16","doi-asserted-by":"crossref","first-page":"2951","DOI":"10.1182\/blood-2003-01-0338","article-title":"Classification of pediatric acute lymphoblastic leukemia by gene expression profiling","volume":"102","author":"Ross","year":"2003","journal-title":"Blood"},{"key":"2023013107242832600_bty566-B17","doi-asserted-by":"crossref","first-page":"1113.","DOI":"10.1038\/ng.2764","article-title":"The cancer genome atlas pan-cancer analysis project","volume":"45","author":"Weinstein","year":"2013","journal-title":"Nat. Genet"},{"key":"2023013107242832600_bty566-B18","doi-asserted-by":"crossref","first-page":"515","DOI":"10.1093\/biostatistics\/kxp008","article-title":"A penalized matrix decomposition, with applications to sparse principal components and canonical correlation analysis","volume":"10","author":"Witten","year":"2009","journal-title":"Biostatistics"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/3\/478\/48964918\/bioinformatics_35_3_478.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/3\/478\/48964918\/bioinformatics_35_3_478.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,31]],"date-time":"2023-01-31T10:18:09Z","timestamp":1675160289000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/35\/3\/478\/5053316"}},"subtitle":[],"editor":[{"given":"Inanc","family":"Birol","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2018,7,13]]},"references-count":18,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2019,2,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bty566","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"type":"print","value":"1367-4803"},{"type":"electronic","value":"1367-4811"}],"subject":[],"published-other":{"date-parts":[[2019,2,1]]},"published":{"date-parts":[[2018,7,13]]}}}