{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,11]],"date-time":"2026-02-11T17:25:57Z","timestamp":1770830757785,"version":"3.50.1"},"reference-count":23,"publisher":"Oxford University Press (OUP)","issue":"19","license":[{"start":{"date-parts":[[2021,3,31]],"date-time":"2021-03-31T00:00:00Z","timestamp":1617148800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"funder":[{"name":"Human Genetics Institute of New Jersey"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2021,10,11]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Summary<\/jats:title>\n                  <jats:p>As the next-generation sequencing technology becomes broadly applied, genomics and transcriptomics are becoming more commonly used in both research and clinical settings. However, proteomics is still an obstacle to be conquered. For most peptide search programs in proteomics, a standard reference protein database is used. Because of the thousands of coding DNA variants in each individual, a standard reference database does not provide perfect match for many proteins\/peptides of an individual. A personalized reference database can improve the detection power and accuracy for individual proteomics data. To connect genomics and proteomics, we designed a Python package PrecisionProDB that is specialized for generating a personized protein database for proteomics applications. PrecisionProDB supports multiple popular file formats and reference databases, and can generate a personized database in minutes. To demonstrate the application of PrecisionProDB, we generated human population-specific reference protein databases with PrecisionProDB, which improves the number of identified peptides by 0.34% on average. In addition, by incorporating cell line-specific variants into the protein database, we demonstrated a 0.71% improvement for peptide identification in the Jurkat cell line. With PrecisionProDB and these datasets, researchers and clinicians can improve their peptide search performance by adopting the more representative protein database or adding population and individual-specific proteins to the search database with minimum increase of efforts.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Availabilityand implementation<\/jats:title>\n                  <jats:p>PrecisionProDB and pre-calculated protein databases are freely available at https:\/\/github.com\/ATPs\/PrecisionProDB and https:\/\/github.com\/ATPs\/PrecisionProDB_references.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information<\/jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btab218","type":"journal-article","created":{"date-parts":[[2021,3,30]],"date-time":"2021-03-30T19:13:57Z","timestamp":1617131637000},"page":"3361-3363","source":"Crossref","is-referenced-by-count":7,"title":["PrecisionProDB: improving the proteomics performance for precision medicine"],"prefix":"10.1093","volume":"37","author":[{"given":"Xiaolong","family":"Cao","sequence":"first","affiliation":[{"name":"Department of Genetics, Human Genetic Institute of New Jersey, Rutgers, The State University of New Jersey , Piscataway, NJ 08854, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6469-8733","authenticated-orcid":false,"given":"Jinchuan","family":"Xing","sequence":"additional","affiliation":[{"name":"Department of Genetics, Human Genetic Institute of New Jersey, Rutgers, The State University of New Jersey , Piscataway, NJ 08854, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2021,3,31]]},"reference":[{"key":"2023051608272609200_btab218-B1","doi-asserted-by":"crossref","first-page":"507","DOI":"10.1038\/nrg.2016.86","article-title":"Towards precision medicine","volume":"17","author":"Ashley","year":"2016","journal-title":"Nat. Rev. Genet"},{"key":"2023051608272609200_btab218-B2","doi-asserted-by":"crossref","first-page":"D480","DOI":"10.1093\/nar\/gkaa1100","article-title":"UniProt: the universal protein knowledgebase in 2021","volume":"49","author":"Bateman","year":"2021","journal-title":"Nucleic Acids Res"},{"key":"2023051608272609200_btab218-B3","doi-asserted-by":"crossref","first-page":"1293","DOI":"10.1038\/s41467-020-14968-9","article-title":"Integrated proteogenomic deep sequencing and analytics accurately identify non-canonical peptides in tumor immunopeptidomes","volume":"11","author":"Chong","year":"2020","journal-title":"Nat. Commun"},{"key":"2023051608272609200_btab218-B4","doi-asserted-by":"crossref","first-page":"3681","DOI":"10.1021\/acs.jproteome.8b00295","article-title":"ProteomeGenerator: a framework for comprehensive proteomics based on de novo transcriptome assembly and high-accuracy peptide mass spectral matching","volume":"17","author":"Cifani","year":"2018","journal-title":"J. Proteome Res"},{"key":"2023051608272609200_btab218-B5","doi-asserted-by":"crossref","first-page":"1700259","DOI":"10.1002\/pmic.201700259","article-title":"The role of mass spectrometry and proteogenomics in the advancement of HLA epitope prediction","volume":"18","author":"Creech","year":"2018","journal-title":"Proteomics"},{"key":"2023051608272609200_btab218-B6","doi-asserted-by":"crossref","first-page":"D766","DOI":"10.1093\/nar\/gky955","article-title":"GENCODE reference annotation for the human and mouse genomes","volume":"47","author":"Frankish","year":"2019","journal-title":"Nucleic Acids Res"},{"key":"2023051608272609200_btab218-B7","doi-asserted-by":"crossref","first-page":"434","DOI":"10.1038\/s41586-020-2308-7","article-title":"The mutational constraint spectrum quantified from variation in 141,456 humans","volume":"581","author":"Karczewski","year":"2020","journal-title":"Nature"},{"key":"2023051608272609200_btab218-B8","doi-asserted-by":"crossref","first-page":"2699","DOI":"10.1002\/pmic.201400219","article-title":"Construction and assessment of individualized proteogenomic databases for large-scale analysis of nonsynonymous single nucleotide variants","volume":"14","author":"Krug","year":"2014","journal-title":"Proteomics"},{"key":"2023051608272609200_btab218-B9","doi-asserted-by":"crossref","first-page":"10238","DOI":"10.1038\/ncomms10238","article-title":"Global proteogenomic analysis of human MHC class I-associated peptides derived from non-canonical reading frames","volume":"7","author":"Laumont","year":"2016","journal-title":"Nat. Commun"},{"key":"2023051608272609200_btab218-B10","doi-asserted-by":"crossref","first-page":"eaau5516","DOI":"10.1126\/scitranslmed.aau5516","article-title":"Noncoding regions are the main source of targetable tumor-specific antigens","volume":"10","author":"Laumont","year":"2018","journal-title":"Sci. Transl. Med"},{"key":"2023051608272609200_btab218-B11","doi-asserted-by":"crossref","first-page":"2309","DOI":"10.1021\/acs.jproteome.6b00344","article-title":"JUMPg: an integrative proteogenomics pipeline identifying unannotated proteins in human brain and cancer cells","volume":"15","author":"Li","year":"2016","journal-title":"J. Proteome Res"},{"key":"2023051608272609200_btab218-B12","doi-asserted-by":"crossref","first-page":"1800235","DOI":"10.1002\/pmic.201800235","article-title":"Connecting proteomics to next-generation sequencing: proteogenomics and its current applications in biology","volume":"19","author":"Low","year":"2019","journal-title":"Proteomics"},{"key":"2023051608272609200_btab218-B13","doi-asserted-by":"crossref","first-page":"1114","DOI":"10.1038\/nmeth.3144","article-title":"Proteogenomics: concepts, applications and computational strategies","volume":"11","author":"Nesvizhskii","year":"2014","journal-title":"Nat. Methods"},{"key":"2023051608272609200_btab218-B14","doi-asserted-by":"crossref","first-page":"D733","DOI":"10.1093\/nar\/gkv1189","article-title":"Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation","volume":"44","author":"O'Leary","year":"2016","journal-title":"Nucleic Acids Res"},{"key":"2023051608272609200_btab218-B15","doi-asserted-by":"crossref","first-page":"535","DOI":"10.1016\/j.cell.2018.04.008","article-title":"Revolutionizing precision oncology through collaborative proteogenomics and data sharing","volume":"173","author":"Rodriguez","year":"2018","journal-title":"Cell"},{"key":"2023051608272609200_btab218-B16","doi-asserted-by":"crossref","first-page":"1060","DOI":"10.1074\/mcp.M115.056226","article-title":"An analysis of the sensitivity of proteogenomic mapping of somatic mutations and novel splicing events in cancer","volume":"15","author":"Ruggles","year":"2016","journal-title":"Mol. Cell. Proteomics"},{"key":"2023051608272609200_btab218-B17","doi-asserted-by":"crossref","first-page":"45","DOI":"10.1016\/j.cell.2019.02.003","article-title":"Genomic medicine-progress, pitfalls, and promise","volume":"177","author":"Shendure","year":"2019","journal-title":"Cell"},{"key":"2023051608272609200_btab218-B18","doi-asserted-by":"crossref","first-page":"703","DOI":"10.1186\/1471-2164-15-703","article-title":"Using Galaxy-P to leverage RNA-Seq for the discovery of novel protein variations","volume":"15","author":"Sheynkman","year":"2014","journal-title":"BMC Genomics"},{"key":"2023051608272609200_btab218-B19","doi-asserted-by":"crossref","first-page":"3235","DOI":"10.1093\/bioinformatics\/btt543","article-title":"customProDB: an R package to generate customized protein databases from RNA-Seq data for proteomics search","volume":"29","author":"Wang","year":"2013","journal-title":"Bioinformatics"},{"key":"2023051608272609200_btab218-B20","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1021\/pr400294c","article-title":"Proteogenomic database construction driven from large scale RNA-seq data","volume":"13","author":"Woo","year":"2014","journal-title":"J. Proteome Res"},{"key":"2023051608272609200_btab218-B21","doi-asserted-by":"crossref","first-page":"11778","DOI":"10.1038\/ncomms11778","article-title":"Improving GENCODE reference gene annotation using a high-stringency proteogenomics workflow","volume":"7","author":"Wright","year":"2016","journal-title":"Nat. Commun"},{"key":"2023051608272609200_btab218-B22","first-page":"D682","article-title":"Ensembl 2020","volume":"48","author":"Yates","year":"2020","journal-title":"Nucleic Acids Res"},{"key":"2023051608272609200_btab218-B23","doi-asserted-by":"crossref","first-page":"256","DOI":"10.1038\/s41571-018-0135-7","article-title":"Clinical potential of mass spectrometry-based proteogenomics","volume":"16","author":"Zhang","year":"2019","journal-title":"Nat. Rev. Clin. Oncol"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btab218\/37061800\/btab218.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/37\/19\/3361\/50338257\/btab218.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/37\/19\/3361\/50338257\/btab218.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,16]],"date-time":"2023-05-16T08:41:50Z","timestamp":1684226510000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/37\/19\/3361\/6206360"}},"subtitle":[],"editor":[{"given":"Olga","family":"Vitek","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2021,3,31]]},"references-count":23,"journal-issue":{"issue":"19","published-print":{"date-parts":[[2021,10,11]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btab218","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2021,10,1]]},"published":{"date-parts":[[2021,3,31]]}}}