{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,9,6]],"date-time":"2023-09-06T10:10:20Z","timestamp":1693995020082},"reference-count":16,"publisher":"Oxford University Press (OUP)","issue":"5","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2014,3,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Statistical validation of protein identifications is an important issue in shotgun proteomics. The false discovery rate (FDR) is a powerful statistical tool for evaluating the protein identification result. Several research efforts have been made for FDR estimation at the protein level. However, there are still certain drawbacks in the existing FDR estimation methods based on the target-decoy strategy.<\/jats:p>\n               <jats:p>Results: In this article, we propose a decoy-free protein-level FDR estimation method. Under the null hypothesis that each candidate protein matches an identified peptide totally at random, we assign statistical significance to protein identifications in terms of the permutation P-value and use these P-values to calculate the FDR. Our method consists of three key steps: (i) generating random bipartite graphs with the same structure; (ii) calculating the protein scores on these random graphs; and (iii) calculating the permutation P value and final FDR. As it is time-consuming or prohibitive to execute the protein inference algorithms for thousands of times in step ii, we first train a linear regression model using the original bipartite graph and identification scores provided by the target inference algorithm. Then we use the learned regression model as a substitute of original protein inference method to predict protein scores on shuffled graphs. We test our method on six public available datasets. The results show that our method is comparable with those state-of-the-art algorithms in terms of estimation accuracy.<\/jats:p>\n               <jats:p>Availability: The source code of our algorithm is available at: https:\/\/sourceforge.net\/projects\/plfdr\/<\/jats:p>\n               <jats:p>Contact: \u00a0zyhe@dlut.edu.cn<\/jats:p>\n               <jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btt431","type":"journal-article","created":{"date-parts":[[2013,8,8]],"date-time":"2013-08-08T04:52:16Z","timestamp":1375937536000},"page":"675-681","source":"Crossref","is-referenced-by-count":7,"title":["Decoy-free protein-level false discovery rate estimation"],"prefix":"10.1093","volume":"30","author":[{"given":"Ben","family":"Teng","sequence":"first","affiliation":[{"name":"School of Software, Dalian University of Technology, Dalian 116621, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ting","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Software, Dalian University of Technology, Dalian 116621, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zengyou","family":"He","sequence":"additional","affiliation":[{"name":"School of Software, Dalian University of Technology, Dalian 116621, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2013,8,6]]},"reference":[{"key":"2023012710430157800_btt431-B1","doi-asserted-by":"crossref","first-page":"576","DOI":"10.1038\/nbt1300","article-title":"A high-quality catalog of the Drosophila melanogaster proteome","volume":"25","author":"Brunner","year":"2007","journal-title":"Nat. Biotechnol."},{"key":"2023012710430157800_btt431-B2","doi-asserted-by":"crossref","first-page":"1534","DOI":"10.1002\/pmic.200300744","article-title":"Unimod: protein modifications for mass spectrometry","volume":"4","author":"David","year":"2004","journal-title":"Proteomics"},{"key":"2023012710430157800_btt431-B3","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1145\/1297332.1297338","article-title":"Assessing data mining results via swap randomization","volume":"1","author":"Gionis","year":"2007","journal-title":"ACM Trans. Knowl. Discov. Data"},{"key":"2023012710430157800_btt431-B4","doi-asserted-by":"crossref","first-page":"2956","DOI":"10.1093\/bioinformatics\/bts540","article-title":"A linear programming model for protein inference problem in shotgun proteomics","volume":"28","author":"Huang","year":"2012","journal-title":"Bioinformatics"},{"key":"2023012710430157800_btt431-B5","doi-asserted-by":"crossref","first-page":"586","DOI":"10.1093\/bib\/bbs004","article-title":"Protein inference: a review","volume":"13","author":"Huang","year":"2012","journal-title":"Brief. Bioinform."},{"key":"2023012710430157800_btt431-B6","doi-asserted-by":"crossref","first-page":"3354","DOI":"10.1021\/pr8001244","article-title":"Spectral probabilities and generating functions of tandem mass spectra: a strike against decoy databases","volume":"7","author":"Kim","year":"2008","journal-title":"J. Proteome Res."},{"key":"2023012710430157800_btt431-B7","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1021\/pr070244j","article-title":"The Standard Protein Mix Database: a diverse data set to assist in the production of improved peptide and protein identification software tools","volume":"7","author":"Klimek","year":"2008","journal-title":"J. Proteome Res."},{"key":"2023012710430157800_btt431-B8","doi-asserted-by":"crossref","first-page":"4646","DOI":"10.1021\/ac0341261","article-title":"A statistical model for identifying proteins by tandem mass spectrometry","volume":"75","author":"Nesvizhskii","year":"2003","journal-title":"Anal. Chem."},{"key":"2023012710430157800_btt431-B9","doi-asserted-by":"crossref","first-page":"2405","DOI":"10.1038\/nmeth1088","article-title":"Analysis and validation of proteomic data generated by tandem mass spectrometry","volume":"4","author":"Nesvizhskii","year":"2007","journal-title":"Nat. Methods"},{"key":"2023012710430157800_btt431-B10","doi-asserted-by":"crossref","first-page":"2955","DOI":"10.1093\/bioinformatics\/btp461","article-title":"Mining gene functional networks to improve mass-spectrometry based protein identification","volume":"25","author":"Ramakrishnan","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012710430157800_btt431-B11","doi-asserted-by":"crossref","first-page":"1397","DOI":"10.1093\/bioinformatics\/btp168","article-title":"Integrating shotgun proteomics and mRNA expression data to improve protein identification","volume":"25","author":"Ramakrishnan","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012710430157800_btt431-B12","doi-asserted-by":"crossref","first-page":"787","DOI":"10.1074\/mcp.M900317-MCP200","article-title":"Protein identification false discovery rates for very large proteomics data sets generated by tandem mass spectrometry","volume":"8","author":"Reiter","year":"2009","journal-title":"Mol. Cell. Proteomics"},{"key":"2023012710430157800_btt431-B13","doi-asserted-by":"crossref","first-page":"1128","DOI":"10.1093\/bioinformatics\/btr089","article-title":"Assigning spectrum-specific p-values to protein identifications by mass spectrometry","volume":"27","author":"Spirin","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012710430157800_btt431-B14","doi-asserted-by":"crossref","first-page":"479","DOI":"10.1111\/1467-9868.00346","article-title":"A direct approach to false discovery rates","volume":"64","author":"Storey","year":"2002","journal-title":"J. R. Stat. Soc. Ser. B Stat. Methodol."},{"key":"2023012710430157800_btt431-B15","doi-asserted-by":"crossref","first-page":"9440","DOI":"10.1073\/pnas.1530509100","article-title":"Statistical significance for genomewide studies","volume":"100","author":"Storey","year":"2003","journal-title":"Proc. Natl Acad Sci. USA"},{"key":"2023012710430157800_btt431-B16","doi-asserted-by":"crossref","first-page":"654","DOI":"10.1021\/pr0604054","article-title":"Myrimatch: highly accurate tandem mass spectral peptide identification by multivariate hypergeometric analysis","volume":"6","author":"Tabb","year":"2007","journal-title":"J. Proteome Res."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/5\/675\/48917836\/bioinformatics_30_5_675.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/5\/675\/48917836\/bioinformatics_30_5_675.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T11:02:23Z","timestamp":1674817343000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/30\/5\/675\/244620"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,8,6]]},"references-count":16,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2014,3,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btt431","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2014,3,1]]},"published":{"date-parts":[[2013,8,6]]}}}