{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T19:48:03Z","timestamp":1783972083171,"version":"3.55.0"},"reference-count":19,"publisher":"Oxford University Press (OUP)","issue":"3","license":[{"start":{"date-parts":[[2020,6,10]],"date-time":"2020-06-10T00:00:00Z","timestamp":1591747200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2021,5,20]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Label-free shotgun proteomics is an important tool in biomedical research, where tandem mass spectrometry with data-dependent acquisition (DDA) is frequently used for protein identification and quantification. However, the DDA datasets contain a significant number of missing values (MVs) that severely hinders proper analysis. Existing literature suggests that different imputation methods should be used for the two types of MVs: missing completely at random or missing not at random. However, the simulated or biased datasets utilized by most of such studies offer few clues about the composition and thus proper imputation of MVs in real-life proteomic datasets. Moreover, the impact of imputation methods on downstream differential expression analysis\u2014a critical goal for many biomedical projects\u2014is largely undetermined. In this study, we investigated public DDA datasets of various tissue\/sample types to determine the composition of MVs in them. We then developed simulated datasets that imitate the MV profile of real-life datasets. Using such datasets, we compared the impact of various popular imputation methods on the analysis of differentially expressed proteins. Finally, we make recommendations on which imputation method(s) to use for proteomic data beyond just DDA datasets.<\/jats:p>","DOI":"10.1093\/bib\/bbaa112","type":"journal-article","created":{"date-parts":[[2020,5,13]],"date-time":"2020-05-13T03:26:56Z","timestamp":1589340416000},"source":"Crossref","is-referenced-by-count":79,"title":["Proper imputation of missing values in proteomics datasets for differential expression analysis"],"prefix":"10.1093","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8101-8522","authenticated-orcid":false,"given":"Mingyi","family":"Liu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ashok","family":"Dongre","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2020,6,10]]},"reference":[{"key":"2021052110034268500_ref1","doi-asserted-by":"crossref","first-page":"e30","DOI":"10.1017\/S1462399410001614","article-title":"Mass spectrometry-based proteomics in biomedical research: emerging technologies and future strategies","volume":"12","author":"Walsh","year":"2010","journal-title":"Expert Rev Mol Med"},{"key":"2021052110034268500_ref2","doi-asserted-by":"crossref","first-page":"2343","DOI":"10.1021\/cr3003533","article-title":"Protein analysis by shotgun\/bottom-up proteomics","volume":"113","author":"Zhang","year":"2013","journal-title":"Chem Rev"},{"key":"2021052110034268500_ref3","doi-asserted-by":"crossref","first-page":"2028","DOI":"10.1093\/bioinformatics\/btp362","article-title":"A statistical framework for protein quantitation in bottom-up MS-based proteomics","volume":"25","author":"Karpievitch","year":"2009","journal-title":"Bioinformatics"},{"key":"2021052110034268500_ref4","doi-asserted-by":"crossref","first-page":"1202","DOI":"10.1002\/pmic.200800576","article-title":"Missing values in gel-based proteomics","volume":"10","author":"Albrecht","year":"2010","journal-title":"Proteomics"},{"key":"2021052110034268500_ref5","doi-asserted-by":"crossref","first-page":"S5","DOI":"10.1186\/1471-2105-13-S16-S5","article-title":"Normalization and missing value imputation for label-free LC-MS analysis","volume":"13","author":"Karpievitch","year":"2012","journal-title":"BMC Bioinformatics"},{"key":"2021052110034268500_ref6","doi-asserted-by":"crossref","first-page":"1116","DOI":"10.1021\/acs.jproteome.5b00981","article-title":"Accounting for the multiple natures of missing values in label-free quantitative proteomics data sets to compare imputation strategies","volume":"15","author":"Lazar","year":"2016","journal-title":"J Proteome Res"},{"key":"2021052110034268500_ref7","first-page":"1344","article-title":"A comprehensive evaluation of popular proteomics software workflows for label-free proteome quantification and imputation","volume":"19","author":"V\u00e4likangas","year":"2018","journal-title":"Brief Bioinform"},{"key":"2021052110034268500_ref8","doi-asserted-by":"crossref","first-page":"1993","DOI":"10.1021\/pr501138h","article-title":"Review, evaluation, and discussion of the challenges of missing value imputation for mass spectrometry-based label-free global proteomics","volume":"14","author":"Webb-Robertson","year":"2015","journal-title":"J Proteome Res"},{"key":"2021052110034268500_ref9","doi-asserted-by":"crossref","first-page":"663","DOI":"10.1038\/s41598-017-19120-0","article-title":"Missing value imputation approach for mass spectrometry-based metabolomics data","volume":"8","author":"Wei","year":"2018","journal-title":"Sci Rep"},{"key":"2021052110034268500_ref10","doi-asserted-by":"crossref","first-page":"2075","DOI":"10.1214\/18-AOAS1144","article-title":"The effects of nonignorable missing data on label-free mass spectrometry proteomics experiments","volume":"12","author":"O\u2019Brien","year":"2018","journal-title":"Ann Appl Stat"},{"key":"2021052110034268500_ref11","doi-asserted-by":"crossref","first-page":"3367","DOI":"10.1038\/s41598-017-03650-8","article-title":"In-depth method assessments of differentially expressed protein detection for shotgun proteomics data with missing values","volume":"7","author":"Wang","year":"2017","journal-title":"Sci Rep"},{"key":"2021052110034268500_ref12","doi-asserted-by":"crossref","first-page":"D447","DOI":"10.1093\/nar\/gkv1145","article-title":"2016 update of the PRIDE database and its related tools","volume":"44","author":"Vizca\u00edno","year":"2016","journal-title":"Nucleic Acids Res"},{"key":"2021052110034268500_ref13","doi-asserted-by":"crossref","first-page":"2513","DOI":"10.1074\/mcp.M113.031591","article-title":"Accurate proteome-wide label-free quantification by delayed normalization and maximal peptide ratio extraction, termed MaxLFQ","volume":"13","author":"Cox","year":"2014","journal-title":"Mol Cell Proteomics"},{"key":"2021052110034268500_ref14","doi-asserted-by":"crossref","first-page":"e47","DOI":"10.1093\/nar\/gkv007","article-title":"Limma powers differential expression analyses for RNA-sequencing and microarray studies","volume":"43","author":"Ritchie","year":"2015","journal-title":"Nucleic Acids Res"},{"key":"2021052110034268500_ref15","doi-asserted-by":"crossref","first-page":"971","DOI":"10.1093\/bib\/bbx031","article-title":"Identification of differentially expressed peptides in high-throughput proteomics data","volume":"19","author":"Ooijen","year":"2018","journal-title":"Brief Bioinform"},{"key":"2021052110034268500_ref16","volume-title":"R: A language and environment for statistical computing","author":"R Core Team","year":"2017"},{"key":"2021052110034268500_ref17","volume-title":"RStudio: Integrated Development for R","author":"RStudio Team","year":"2015"},{"key":"2021052110034268500_ref18","doi-asserted-by":"crossref","first-page":"288","DOI":"10.1093\/bioinformatics\/btr645","article-title":"MSnbase-an R\/Bioconductor package for isobaric tagged mass spectrometry data visualization, processing and quantitation","volume":"28","author":"Gatto","year":"2012","journal-title":"Bioinformatics"},{"key":"2021052110034268500_ref19","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-24277-4","volume-title":"ggplot2: Elegant Graphics for Data Analysis","author":"Wickham","year":"2016"}],"container-title":["Briefings in Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/academic.oup.com\/bib\/article-pdf\/22\/3\/bbaa112\/37966016\/bbaa112.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"http:\/\/academic.oup.com\/bib\/article-pdf\/22\/3\/bbaa112\/37966016\/bbaa112.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,5,21]],"date-time":"2021-05-21T10:08:05Z","timestamp":1621591685000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bib\/article\/doi\/10.1093\/bib\/bbaa112\/5855395"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,6,10]]},"references-count":19,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2021,5,20]]}},"URL":"https:\/\/doi.org\/10.1093\/bib\/bbaa112","relation":{},"ISSN":["1467-5463","1477-4054"],"issn-type":[{"value":"1467-5463","type":"print"},{"value":"1477-4054","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2021,5]]},"published":{"date-parts":[[2020,6,10]]},"article-number":"bbaa112"}}