{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,5]],"date-time":"2025-10-05T04:33:16Z","timestamp":1759638796213},"reference-count":39,"publisher":"Oxford University Press (OUP)","issue":"7","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2013,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: Tandem mass spectrometry (MS\/MS) is a dominant approach for large-scale high-throughput post-translational modification (PTM) profiling. Although current state-of-the-art blind PTM spectral analysis algorithms can predict thousands of modified peptides (PTM predictions) in an MS\/MS experiment, a significant percentage of these predictions have inaccurate modification mass estimates and false modification site assignments. This problem can be addressed by post-processing the PTM predictions with a PTM refinement algorithm. We developed a novel PTM refinement algorithm, iPTMClust, which extends a recently introduced PTM refinement algorithm PTMClust and uses a non-parametric Bayesian model to better account for uncertainties in the quantity and identity of PTMs in the input data. The use of this new modeling approach enables iPTMClust to provide a confidence score per modification site that allows fine-tuning and interpreting resulting PTM predictions.<\/jats:p><jats:p>Results: The primary goal behind iPTMClust is to improve the quality of the PTM predictions. First, to demonstrate that iPTMClust produces sensible and accurate cluster assignments, we compare it with k-means clustering, mixtures of Gaussians (MOG) and PTMClust on a synthetically generated PTM dataset. Second, in two separate benchmark experiments using PTM data taken from a phosphopeptide and a yeast proteome study, we show that iPTMClust outperforms state-of-the-art PTM prediction and refinement algorithms, including PTMClust. Finally, we illustrate the general applicability of our new approach on a set of human chromatin protein complex data, where we are able to identify putative novel modified peptides and modification sites that may be involved in the formation and regulation of protein complexes. Our method facilitates accurate PTM profiling, which is an important step in understanding the mechanisms behind many biological processes and should be an integral part of any proteomic study.<\/jats:p><jats:p>Availability: Our algorithm is implemented in Java and is freely available for academic use from http:\/\/genes.toronto.edu.<\/jats:p><jats:p>Contact: \u00a0frey@psi.utoronto.ca<\/jats:p><jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online<\/jats:p>","DOI":"10.1093\/bioinformatics\/btt056","type":"journal-article","created":{"date-parts":[[2013,2,19]],"date-time":"2013-02-19T12:32:56Z","timestamp":1361277176000},"page":"821-829","source":"Crossref","is-referenced-by-count":8,"title":["Non-parametric Bayesian approach to post-translational modification refinement of predictions from tandem mass spectrometry"],"prefix":"10.1093","volume":"29","author":[{"given":"Clement","family":"Chung","sequence":"first","affiliation":[{"name":"1 Department of Computer Science, University of Toronto, Toronto, M5S 2E4, 2Probabilistic and Statistical Inference Group, University of Toronto, Toronto, M5S 3G4, 3Banting and Best Department of Medical Research, 4Donnelly Centre for Cellular and Biomolecular Research, University of Toronto, Toronto, M5S 3E1 and 5Department of Electrical and Computer Engineering, University of Toronto, Toronto, M5S 3G4, Canada"},{"name":"1 Department of Computer Science, University of Toronto, Toronto, M5S 2E4, 2Probabilistic and Statistical Inference Group, University of Toronto, Toronto, M5S 3G4, 3Banting and Best Department of Medical Research, 4Donnelly Centre for Cellular and Biomolecular Research, University of Toronto, Toronto, M5S 3E1 and 5Department of Electrical and Computer Engineering, University of Toronto, Toronto, M5S 3G4, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrew","family":"Emili","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science, University of Toronto, Toronto, M5S 2E4, 2Probabilistic and Statistical Inference Group, University of Toronto, Toronto, M5S 3G4, 3Banting and Best Department of Medical Research, 4Donnelly Centre for Cellular and Biomolecular Research, University of Toronto, Toronto, M5S 3E1 and 5Department of Electrical and Computer Engineering, University of Toronto, Toronto, M5S 3G4, Canada"},{"name":"1 Department of Computer Science, University of Toronto, Toronto, M5S 2E4, 2Probabilistic and Statistical Inference Group, University of Toronto, Toronto, M5S 3G4, 3Banting and Best Department of Medical Research, 4Donnelly Centre for Cellular and Biomolecular Research, University of Toronto, Toronto, M5S 3E1 and 5Department of Electrical and Computer Engineering, University of Toronto, Toronto, M5S 3G4, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Brendan J.","family":"Frey","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science, University of Toronto, Toronto, M5S 2E4, 2Probabilistic and Statistical Inference Group, University of Toronto, Toronto, M5S 3G4, 3Banting and Best Department of Medical Research, 4Donnelly Centre for Cellular and Biomolecular Research, University of Toronto, Toronto, M5S 3E1 and 5Department of Electrical and Computer Engineering, University of Toronto, Toronto, M5S 3G4, Canada"},{"name":"1 Department of Computer Science, University of Toronto, Toronto, M5S 2E4, 2Probabilistic and Statistical Inference Group, University of Toronto, Toronto, M5S 3G4, 3Banting and Best Department of Medical Research, 4Donnelly Centre for Cellular and Biomolecular Research, University of Toronto, Toronto, M5S 3E1 and 5Department of Electrical and Computer Engineering, University of Toronto, Toronto, M5S 3G4, Canada"},{"name":"1 Department of Computer Science, University of Toronto, Toronto, M5S 2E4, 2Probabilistic and Statistical Inference Group, University of Toronto, Toronto, M5S 3G4, 3Banting and Best Department of Medical Research, 4Donnelly Centre for Cellular and Biomolecular Research, University of Toronto, Toronto, M5S 3E1 and 5Department of Electrical and Computer Engineering, University of Toronto, Toronto, M5S 3G4, Canada"},{"name":"1 Department of Computer Science, University of Toronto, Toronto, M5S 2E4, 2Probabilistic and Statistical Inference Group, University of Toronto, Toronto, M5S 3G4, 3Banting and Best Department of Medical Research, 4Donnelly Centre for Cellular and Biomolecular Research, University of Toronto, Toronto, M5S 3E1 and 5Department of Electrical and Computer Engineering, University of Toronto, Toronto, M5S 3G4, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2013,2,17]]},"reference":[{"key":"2023020303304781700_btt056-B1","article-title":"Spectrum mill","author":"Agilent","year":"2005"},{"key":"2023020303304781700_btt056-B2","doi-asserted-by":"crossref","first-page":"1389","DOI":"10.1074\/mcp.M700468-MCP200","article-title":"A multidimensional chromatography technology for in-depth phosphoproteome analysis","volume":"7","author":"Albuquerque","year":"2008","journal-title":"Mol. Cell Proteomics"},{"key":"2023020303304781700_btt056-B3","doi-asserted-by":"crossref","first-page":"1152","DOI":"10.1214\/aos\/1176342871","article-title":"Mixtures of Dirichlet processes with applications to Bayesian nonparametric problems","volume":"2","author":"Antoniak","year":"1974","journal-title":"Ann. Stat."},{"key":"2023020303304781700_btt056-B4","doi-asserted-by":"crossref","first-page":"1965","DOI":"10.1021\/pr800917p","article-title":"SLoMo: automated site localization of modifications from ETD\/ECD mass spectra","volume":"8","author":"Bailey","year":"2009","journal-title":"J. Proteome Res."},{"key":"2023020303304781700_btt056-B5","doi-asserted-by":"crossref","first-page":"M111","DOI":"10.1074\/mcp.M111.008078","article-title":"Modification site localization scoring integrated into a search engine","volume":"10","author":"Baker","year":"2011","journal-title":"Mol. Cell Proteomics"},{"key":"2023020303304781700_btt056-B6","doi-asserted-by":"crossref","first-page":"12130","DOI":"10.1073\/pnas.0404720101","article-title":"Large-scale characterization of hela cell nuclear phosphoproteins","volume":"101","author":"Beausoleil","year":"2004","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020303304781700_btt056-B7","doi-asserted-by":"crossref","first-page":"1285","DOI":"10.1038\/nbt1240","article-title":"A probability-based approach for high-throughput protein phosphorylation analysis and site localization","volume":"24","author":"Beausoleil","year":"2006","journal-title":"Nat. Biotechnol."},{"key":"2023020303304781700_btt056-B8","volume-title":"Pattern Recognition and Machine Learning","author":"Bishop","year":"2006"},{"key":"2023020303304781700_btt056-B9","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1016\/S0021-9673(04)00971-9","article-title":"Strategies for shotgun identification of post-translational modifications by mass spectrometry","volume":"1053","author":"Cantin","year":"2004","journal-title":"J. Chromatogr."},{"key":"2023020303304781700_btt056-B10","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1074\/mcp.R111.015305","article-title":"Modification site localization scoring: strategies and performance","volume":"11","author":"Chalkley","year":"2012","journal-title":"Mol. Cell Proteomics."},{"key":"2023020303304781700_btt056-B11","doi-asserted-by":"crossref","first-page":"797","DOI":"10.1093\/bioinformatics\/btr017","article-title":"Computational refinement of post-translational modifications predicted from tandem mass spectrometry","volume":"27","author":"Chung","year":"2011","journal-title":"Bioinformatics"},{"key":"2023020303304781700_btt056-B12","doi-asserted-by":"crossref","first-page":"212","DOI":"10.1126\/science.1124619","article-title":"Mass spectrometry and protein analysis","volume":"312","author":"Domon","year":"2006","journal-title":"Science"},{"key":"2023020303304781700_btt056-B13","doi-asserted-by":"crossref","first-page":"577","DOI":"10.1080\/01621459.1995.10476550","article-title":"Bayesian density estimation and inference using mixtures","volume":"90","author":"Escobar","year":"1994","journal-title":"J. Am. Stat. Assoc."},{"key":"2023020303304781700_btt056-B14","doi-asserted-by":"crossref","first-page":"209","DOI":"10.1214\/aos\/1176342360","article-title":"A Bayesian analysis of some nonparametric problems","volume":"1","author":"Ferguson","year":"1973","journal-title":"Ann. Stat."},{"key":"2023020303304781700_btt056-B15","doi-asserted-by":"crossref","first-page":"721","DOI":"10.1109\/TPAMI.1984.4767596","article-title":"Stochastic relaxation, Gibbs distributions, and the Bayesian restoration of images","volume":"PAMI-6","author":"Geman","year":"1984","journal-title":"IEEE Trans. Pattern Anal. Mach. Intel."},{"key":"2023020303304781700_btt056-B16","doi-asserted-by":"crossref","first-page":"697","DOI":"10.1142\/S0219720005001247","article-title":"Spider: software for protein identification from sequence tags with de novo sequencing error","volume":"3","author":"Han","year":"2005","journal-title":"J. Bioinform. Comput. Biol."},{"key":"2023020303304781700_btt056-B17","doi-asserted-by":"crossref","first-page":"158","DOI":"10.1198\/1061860043001","article-title":"A split-merge Markov chain Monte Carlo procedure for the Dirichlet process mixture model","volume":"13","author":"Jain","year":"2000","journal-title":"J. Comput. Graph. Stat."},{"key":"2023020303304781700_btt056-B18","doi-asserted-by":"crossref","first-page":"5383","DOI":"10.1021\/ac025747h","article-title":"Empirical statistical model to estimate the accuracy of peptide identifications made by MS\/MS and database search","volume":"74","author":"Keller","year":"2002","journal-title":"Anal. Chem."},{"key":"2023020303304781700_btt056-B19","doi-asserted-by":"crossref","first-page":"637","DOI":"10.1038\/nature04670","article-title":"Global landscape of protein complexes in the yeast Saccharomyces cerevisiae","volume":"440","author":"Krogan","year":"2006","journal-title":"Nature"},{"key":"2023020303304781700_btt056-B20","volume-title":"Principles of Biochemistry","author":"Lehninger","year":"1993","edition":"2nd"},{"key":"2023020303304781700_btt056-B21","doi-asserted-by":"crossref","first-page":"e307","DOI":"10.1093\/bioinformatics\/btl226","article-title":"Peptide sequence tag-based blind identification of post-translational modifications with point process model","volume":"22","author":"Liu","year":"2006","journal-title":"Bioninformatics"},{"key":"2023020303304781700_btt056-B22","doi-asserted-by":"crossref","first-page":"7846","DOI":"10.1021\/ac8009017","article-title":"Sequential interval motif search: unrestricted database searching of global MS\/MS datasets for unexpected post-translational modifications","volume":"80","author":"Liu","year":"2008","journal-title":"Anal. Chem."},{"key":"2023020303304781700_btt056-B23","doi-asserted-by":"crossref","first-page":"1811","DOI":"10.1016\/j.bbapap.2006.10.003","article-title":"The utility of ETD mass spectrometry in proteomic analysis","volume":"1764","author":"Mikesh","year":"2006","journal-title":"Biochim. Biophys. Acta."},{"key":"2023020303304781700_btt056-B24","doi-asserted-by":"crossref","first-page":"4418","DOI":"10.1021\/pr9001146","article-title":"Prediction of novel modifications by unrestrictive search of tandem mass spectra","volume":"8","author":"Na","year":"2009","journal-title":"J. Proteome Res."},{"key":"2023020303304781700_btt056-B25","first-page":"705","article-title":"Slice sampling","volume":"31","author":"Neal","year":"2000","journal-title":"Ann. Stat."},{"key":"2023020303304781700_btt056-B26","doi-asserted-by":"crossref","first-page":"249","DOI":"10.1080\/10618600.2000.10474879","article-title":"Markov chain sampling methods for Dirichlet process mixture models","volume":"9","author":"Neal","year":"2000","journal-title":"J. Comput. Graph. Stat."},{"key":"2023020303304781700_btt056-B27","doi-asserted-by":"crossref","first-page":"635","DOI":"10.1016\/j.cell.2006.09.026","article-title":"Global, in vivo, and site-specific phosphorylation dynamics in signaling network","volume":"127","author":"Olsen","year":"2006","journal-title":"Cell"},{"key":"2023020303304781700_btt056-B28","first-page":"821","article-title":"Proteomic and phosphoproteomic comparison of human ES and iPS cells","volume":"8","author":"Phanstiel","year":"2011","journal-title":"Mol. Cell Proteomics"},{"key":"2023020303304781700_btt056-B29","doi-asserted-by":"crossref","first-page":"218","DOI":"10.1006\/meth.2001.1183","article-title":"The tandem affinity purification (tap) method: a general procedure of protein complex purification","volume":"24","author":"Puig","year":"2001","journal-title":"Methods"},{"key":"2023020303304781700_btt056-B30","doi-asserted-by":"crossref","first-page":"2955","DOI":"10.1093\/bioinformatics\/btp461","article-title":"Mining gene functional networks to improve mass-spectrometry based protein identification","volume":"25","author":"Ramakrishnan","year":"2009","journal-title":"Bioinformatics"},{"key":"2023020303304781700_btt056-B31","doi-asserted-by":"crossref","first-page":"1397","DOI":"10.1093\/bioinformatics\/btp168","article-title":"Integrating shotgun proteomics and mRNA expression data to improve protein identification","volume":"25","author":"Ramakrishnan","year":"2009","journal-title":"Bioinformatics"},{"key":"2023020303304781700_btt056-B32","first-page":"554","article-title":"The infinite Gaussian mixture model","volume-title":"Advances in Neural Information Processing Systems 12","author":"Rasmussen","year":"2000"},{"key":"2023020303304781700_btt056-B33","doi-asserted-by":"crossref","first-page":"1030","DOI":"10.1038\/13732","article-title":"A generic protein purification method for protein complex characterization and proteome exploration","volume":"17","author":"Rigaut","year":"1999","journal-title":"Nat. Biotechnol."},{"key":"2023020303304781700_btt056-B34","doi-asserted-by":"crossref","first-page":"M110","DOI":"10.1074\/mcp.M110.003830","article-title":"Confident phosphorylation site localization using the mascot delta score","volume":"10","author":"Savitski","year":"2011","journal-title":"Mol. Cell Proteomics"},{"key":"2023020303304781700_btt056-B35","doi-asserted-by":"crossref","first-page":"546","DOI":"10.1021\/pr049781j","article-title":"Identification of protein modifications using MS\/MS de novo sequencing and the opensea alignment algorithm","volume":"4","author":"Searle","year":"2006","journal-title":"J. Proteome Res."},{"key":"2023020303304781700_btt056-B36","doi-asserted-by":"crossref","first-page":"4626","DOI":"10.1021\/ac050102d","article-title":"Inspect: identification of posttranslationally modified peptides from tandem mass spectra","volume":"77","author":"Tanner","year":"2005","journal-title":"Anal. Chem."},{"key":"2023020303304781700_btt056-B37","doi-asserted-by":"crossref","first-page":"170","DOI":"10.1021\/pr070444v","article-title":"Accurate annotation of peptide modifications through unrestrictive database search","volume":"7","author":"Tanner","year":"2008","journal-title":"J. Proteome Res."},{"key":"2023020303304781700_btt056-B38","doi-asserted-by":"crossref","first-page":"5354","DOI":"10.1021\/pr200611n","article-title":"Universal and confident phosphorylation site localization using phosphors","volume":"10","author":"Taus","year":"2011","journal-title":"J. Proteome Res."},{"key":"2023020303304781700_btt056-B39","doi-asserted-by":"crossref","first-page":"1562","DOI":"10.1038\/nbt1168","article-title":"Identification of post-translational modifications by blind search of mass spectra","volume":"23","author":"Tsur","year":"2005","journal-title":"Nat. Biotechnol."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/29\/7\/821\/49060819\/bioinformatics_29_7_821.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/29\/7\/821\/49060819\/bioinformatics_29_7_821.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,6]],"date-time":"2024-05-06T06:07:36Z","timestamp":1714975656000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/29\/7\/821\/253183"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,2,17]]},"references-count":39,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2013,4,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btt056","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2013,4,1]]},"published":{"date-parts":[[2013,2,17]]}}}