{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,12]],"date-time":"2026-04-12T07:21:53Z","timestamp":1775978513605,"version":"3.50.1"},"reference-count":18,"publisher":"Oxford University Press (OUP)","issue":"20","license":[{"start":{"date-parts":[[2016,10,2]],"date-time":"2016-10-02T00:00:00Z","timestamp":1475366400000},"content-version":"vor","delay-in-days":476,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2015,10,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Proteogenomics has been well accepted as a tool to discover novel genes. In most conventional proteogenomic studies, a global false discovery rate is used to filter out false positives for identifying credible novel peptides. However, it has been found that the actual level of false positives in novel peptides is often out of control and behaves differently for different genomes.<\/jats:p>\n               <jats:p>Results: To quantitatively model this problem, we theoretically analyze the subgroup false discovery rates of annotated and novel peptides. Our analysis shows that the annotation completeness ratio of a genome is the dominant factor influencing the subgroup FDR of novel peptides. Experimental results on two real datasets of Escherichia coli and Mycobacterium tuberculosis support our conjecture.<\/jats:p>\n               <jats:p>Contact: \u00a0yfu@amss.ac.cn or xupingghy@gmail.com or smhe@ict.ac.cn<\/jats:p>\n               <jats:p>Supplementary information: Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btv340","type":"journal-article","created":{"date-parts":[[2015,6,16]],"date-time":"2015-06-16T00:30:31Z","timestamp":1434414631000},"page":"3249-3253","source":"Crossref","is-referenced-by-count":30,"title":["A note on the false discovery rate of novel peptides in proteogenomics"],"prefix":"10.1093","volume":"31","author":[{"given":"Kun","family":"Zhang","sequence":"first","affiliation":[{"name":"1 Key Lab of Intelligent Information Processing of Chinese Academy of Sciences (CAS), Institute of Computing Technology, CAS, Beijing 100190,"},{"name":"2 University of Chinese Academy of Sciences, Beijing 100049,"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yan","family":"Fu","sequence":"additional","affiliation":[{"name":"3 National Center for Mathematics and Interdisciplinary Sciences, Key Laboratory of Random Complex Structures and Data Science, Academy of Mathematics and Systems Science, Chinese Academy of Sciences, Beijing 100190 and"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wen-Feng","family":"Zeng","sequence":"additional","affiliation":[{"name":"1 Key Lab of Intelligent Information Processing of Chinese Academy of Sciences (CAS), Institute of Computing Technology, CAS, Beijing 100190,"},{"name":"2 University of Chinese Academy of Sciences, Beijing 100049,"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kun","family":"He","sequence":"additional","affiliation":[{"name":"1 Key Lab of Intelligent Information Processing of Chinese Academy of Sciences (CAS), Institute of Computing Technology, CAS, Beijing 100190,"},{"name":"2 University of Chinese Academy of Sciences, Beijing 100049,"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hao","family":"Chi","sequence":"additional","affiliation":[{"name":"1 Key Lab of Intelligent Information Processing of Chinese Academy of Sciences (CAS), Institute of Computing Technology, CAS, Beijing 100190,"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chao","family":"Liu","sequence":"additional","affiliation":[{"name":"1 Key Lab of Intelligent Information Processing of Chinese Academy of Sciences (CAS), Institute of Computing Technology, CAS, Beijing 100190,"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yan-Chang","family":"Li","sequence":"additional","affiliation":[{"name":"4 State Key Laboratory of Proteomics, National Engineering Research Center for Protein Drugs, Beijing Proteome Research Center, National Center for Protein Sciences Beijing, Beijing Institute of Radiation Medicine, Beijing 102206, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuan","family":"Gao","sequence":"additional","affiliation":[{"name":"4 State Key Laboratory of Proteomics, National Engineering Research Center for Protein Drugs, Beijing Proteome Research Center, National Center for Protein Sciences Beijing, Beijing Institute of Radiation Medicine, Beijing 102206, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ping","family":"Xu","sequence":"additional","affiliation":[{"name":"4 State Key Laboratory of Proteomics, National Engineering Research Center for Protein Drugs, Beijing Proteome Research Center, National Center for Protein Sciences Beijing, Beijing Institute of Radiation Medicine, Beijing 102206, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Si-Min","family":"He","sequence":"additional","affiliation":[{"name":"1 Key Lab of Intelligent Information Processing of Chinese Academy of Sciences (CAS), Institute of Computing Technology, CAS, Beijing 100190,"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2015,6,14]]},"reference":[{"key":"2023020202310024500_btv340-B1","doi-asserted-by":"crossref","first-page":"5221","DOI":"10.1021\/pr300411q","article-title":"Addressing statistical biases in nucleotide-derived protein databases for proteogenomic search strategies","volume":"11","author":"Blakeley","year":"2012","journal-title":"J. Proteome Res."},{"key":"2023020202310024500_btv340-B2","doi-asserted-by":"crossref","first-page":"837","DOI":"10.1101\/gr.103119.109","article-title":"Proteogenomics of Pristionchus pacificus reveals distinct proteome structure of nematode models","volume":"20","author":"Borchert","year":"2010","journal-title":"Genome Res."},{"key":"2023020202310024500_btv340-B3","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1038\/nmeth.2732","article-title":"HiRIEF LC-MS enables deep proteome coverage and unbiased proteogenomics","volume":"11","author":"Branca","year":"2014","journal-title":"Nat. Methods"},{"key":"2023020202310024500_btv340-B4","doi-asserted-by":"crossref","first-page":"756","DOI":"10.1101\/gr.114272.110","article-title":"Shotgun proteomics aids discovery of novel protein-coding genes, alternative splicing, and \u201cresurrected\u201d pseudogenes in the mouse genome","volume":"21","author":"Brosch","year":"2011","journal-title":"Genome Res."},{"key":"2023020202310024500_btv340-B5","doi-asserted-by":"crossref","first-page":"1872","DOI":"10.1101\/gr.127951.111","article-title":"A proteogenomic analysis of anopheles gambiae using high-resolution fourier transform mass spectrometry","volume":"21","author":"Chaerkady","year":"2011","journal-title":"Genome Res."},{"key":"2023020202310024500_btv340-B6","doi-asserted-by":"crossref","first-page":"316","DOI":"10.1186\/1471-2164-9-316","article-title":"High accuracy mass spectrometry analysis as a tool to verify and improve gene annotation using Mycobacterium tuberculosis as an example","volume":"9","author":"de Souza","year":"2008","journal-title":"BMC Genomics"},{"key":"2023020202310024500_btv340-B7","doi-asserted-by":"crossref","first-page":"207","DOI":"10.1038\/nmeth1019","article-title":"Target-decoy search strategy for increased confidence in large-scale protein identifications by mass spectrometry","volume":"4","author":"Elias","year":"2007","journal-title":"Nat. Methods"},{"key":"2023020202310024500_btv340-B8","doi-asserted-by":"crossref","first-page":"47","DOI":"10.4310\/SII.2012.v5.n1.a5","article-title":"Bayesian false discovery rates for post-translational modification proteomics","volume":"5","author":"Fu","year":"2012","journal-title":"Stat. Interface"},{"key":"2023020202310024500_btv340-B9","doi-asserted-by":"crossref","first-page":"1359","DOI":"10.1074\/mcp.O113.030189","article-title":"Transferred subgroup false discovery rate for rare post-translational modifications detected by mass spectrometry","volume":"13","author":"Fu","year":"2014","journal-title":"Mol. Cell. Proteomics"},{"key":"2023020202310024500_btv340-B10","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1002\/pmic.200300511","article-title":"Proteogenomic mapping as a complementary method to perform genome annotation","volume":"4","author":"Jaffe","year":"2004","journal-title":"Proteomics"},{"key":"2023020202310024500_btv340-B11","doi-asserted-by":"crossref","first-page":"923","DOI":"10.1038\/nmeth1113","article-title":"Semi-supervised learning for peptide identification from shotgun proteomics datasets","volume":"4","author":"Kall","year":"2007","journal-title":"Nat. Methods"},{"key":"2023020202310024500_btv340-B12","doi-asserted-by":"crossref","first-page":"M111\u2009011627","DOI":"10.1074\/mcp.M111.011627","article-title":"Proteogenomic analysis of Mycobacterium tuberculosis by high resolution mass spectrometry","volume":"10","author":"Kelkar","year":"2011","journal-title":"Mol. Cell. Proteomics"},{"key":"2023020202310024500_btv340-B13","doi-asserted-by":"crossref","first-page":"575","DOI":"10.1038\/nature13302","article-title":"A draft map of the human proteome","volume":"509","author":"Kim","year":"2014","journal-title":"Nature"},{"key":"2023020202310024500_btv340-B14","doi-asserted-by":"crossref","first-page":"3420","DOI":"10.1074\/mcp.M113.029165","article-title":"Deep coverage of the Escherichia coli proteome enables the assessment of false discovery rates in simple proteogenomic experiments","volume":"12","author":"Krug","year":"2013","journal-title":"Mol. Cell. Proteomics"},{"key":"2023020202310024500_btv340-B15","doi-asserted-by":"crossref","first-page":"1660","DOI":"10.1101\/gr.077644.108","article-title":"Use of shotgun proteomics for the identification, confirmation, and correction of C. elegans gene annotations","volume":"18","author":"Merrihew","year":"2008","journal-title":"Genome Res."},{"key":"2023020202310024500_btv340-B16","doi-asserted-by":"crossref","first-page":"1114","DOI":"10.1038\/nmeth.3144","article-title":"Proteogenomics: concepts, applications and computational strategies","volume":"11","author":"Nesvizhskii","year":"2014","journal-title":"Nat. Methods"},{"key":"2023020202310024500_btv340-B17","doi-asserted-by":"crossref","first-page":"620","DOI":"10.1002\/pmic.201000615","article-title":"Proteogenomics","volume":"11","author":"Renuse","year":"2011","journal-title":"Proteomics"},{"key":"2023020202310024500_btv340-B18","doi-asserted-by":"crossref","first-page":"382","DOI":"10.1038\/nature13438","article-title":"Proteogenomic characterization of human colon and rectal cancer","volume":"513","author":"Zhang","year":"2014","journal-title":"Nature"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/31\/20\/3249\/49035577\/bioinformatics_31_20_3249.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/31\/20\/3249\/49035577\/bioinformatics_31_20_3249.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,2]],"date-time":"2023-02-02T03:49:51Z","timestamp":1675309791000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/31\/20\/3249\/195564"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,6,14]]},"references-count":18,"journal-issue":{"issue":"20","published-print":{"date-parts":[[2015,10,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btv340","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2015,10,15]]},"published":{"date-parts":[[2015,6,14]]}}}