{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,8]],"date-time":"2026-03-08T23:42:47Z","timestamp":1773013367485,"version":"3.50.1"},"reference-count":51,"publisher":"Oxford University Press (OUP)","issue":"Supplement_1","license":[{"start":{"date-parts":[[2024,6,28]],"date-time":"2024-06-28T00:00:00Z","timestamp":1719532800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2024,6,28]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>Cross-linking tandem mass spectrometry (XL-MS\/MS) is an established analytical platform used to determine distance constraints between residues within a protein or from physically interacting proteins, thus improving our understanding of protein structure and function. To aid biological discovery with XL-MS\/MS, it is essential that pairs of chemically linked peptides be accurately identified, a process that requires: (i) database search, that creates a ranked list of candidate peptide pairs for each experimental spectrum and (ii) false discovery rate (FDR) estimation, that determines the probability of a false match in a group of top-ranked peptide pairs with scores above a given threshold. Currently, the only available FDR estimation mechanism in XL-MS\/MS is the target-decoy approach (TDA). However, despite its simplicity, TDA has both theoretical and practical limitations that impact the estimation accuracy and increase run time over potential decoy-free approaches (DFAs).<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>We introduce a novel decoy-free framework for FDR estimation in XL-MS\/MS. Our approach relies on multi-sample mixtures of skew normal distributions, where the latent components correspond to the scores of correct peptide pairs (both peptides identified correctly), partially incorrect peptide pairs (one peptide identified correctly, the other incorrectly), and incorrect peptide pairs (both peptides identified incorrectly). To learn these components, we exploit the score distributions of first- and second-ranked peptide-spectrum matches for each experimental spectrum and subsequently estimate FDR using a novel expectation-maximization algorithm with constraints. We evaluate the method on ten datasets and provide evidence that the proposed DFA is theoretically sound and a viable alternative to TDA owing to its good performance in terms of accuracy, variance of estimation, and run time.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation<\/jats:title>\n                  <jats:p>https:\/\/github.com\/shawn-peng\/xlms<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btae233","type":"journal-article","created":{"date-parts":[[2024,6,28]],"date-time":"2024-06-28T09:32:50Z","timestamp":1719567170000},"page":"i428-i436","source":"Crossref","is-referenced-by-count":1,"title":["An algorithm for decoy-free false discovery rate estimation in XL-MS\/MS proteomics"],"prefix":"10.1093","volume":"40","author":[{"given":"Yisu","family":"Peng","sequence":"first","affiliation":[{"name":"Khoury College of Computer Sciences, Northeastern University , Boston, MA 02115, United States"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shantanu","family":"Jain","sequence":"additional","affiliation":[{"name":"Khoury College of Computer Sciences, Northeastern University , Boston, MA 02115, United States"},{"name":"The Institute for Experiential AI, Northeastern University , Boston, MA 02115, United States"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6769-0793","authenticated-orcid":false,"given":"Predrag","family":"Radivojac","sequence":"additional","affiliation":[{"name":"Khoury College of Computer Sciences, Northeastern University , Boston, MA 02115, United States"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2024,6,28]]},"reference":[{"key":"2024071814131602000_btae233-B1","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1007\/978-1-4939-3106-4_7","article-title":"False discovery rate estimation in proteomics","volume":"1362","author":"Aggarwal","year":"2016","journal-title":"Methods Mol Biol"},{"key":"2024071814131602000_btae233-B2","doi-asserted-by":"crossref","first-page":"102","DOI":"10.1093\/bioinformatics\/btm545","article-title":"Fast and accurate identification of semi-tryptic peptides in shotgun proteomics","volume":"24","author":"Alves","year":"2008","journal-title":"Bioinformatics"},{"key":"2024071814131602000_btae233-B3","doi-asserted-by":"crossref","first-page":"581","DOI":"10.1002\/cjs.5550340403","article-title":"A unified view on skewed distributions arising from selections","volume":"34","author":"Arellano-Valle","year":"2006","journal-title":"Can J Statistics"},{"key":"2024071814131602000_btae233-B4","first-page":"171","article-title":"A class of distributions which includes the normal ones","volume":"12","author":"Azzalini","year":"1985","journal-title":"Scand J Stat"},{"key":"2024071814131602000_btae233-B5","doi-asserted-by":"crossref","first-page":"161","DOI":"10.1186\/s13059-018-1547-5","article-title":"SCoPE-MS: mass spectrometry of single mammalian cells quantifies proteome heterogeneity during cell differentiation","volume":"19","author":"Budnik","year":"2018","journal-title":"Genome Biol"},{"key":"2024071814131602000_btae233-B6","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1021\/acs.jproteome.7b00170","article-title":"Gentle introduction to the statistical foundations of false discovery rate in quantitative proteomics","volume":"17","author":"Burger","year":"2018","journal-title":"J Proteome Res"},{"key":"2024071814131602000_btae233-B7","doi-asserted-by":"crossref","first-page":"47","DOI":"10.1021\/pr700747q","article-title":"False discovery rates and related statistical concepts in mass spectrometry-based proteomics","volume":"7","author":"Choi","year":"2008","journal-title":"J Proteome Res"},{"key":"2024071814131602000_btae233-B8","doi-asserted-by":"crossref","first-page":"1432","DOI":"10.1021\/pr101003r","article-title":"The problem with peptide presumption and low mascot scoring","volume":"10","author":"Cooper","year":"2011","journal-title":"J Proteome Res"},{"key":"2024071814131602000_btae233-B9","doi-asserted-by":"crossref","first-page":"9663","DOI":"10.1021\/ac303051s","article-title":"The problem with peptide presumption and the downfall of target-decoy false discovery rates","volume":"84","author":"Cooper","year":"2012","journal-title":"Anal Chem"},{"key":"2024071814131602000_btae233-B10","doi-asserted-by":"crossref","first-page":"327","DOI":"10.1089\/106652799318300","article-title":"De novo peptide sequencing via tandem mass spectrometry","volume":"6","author":"Dancik","year":"1999","journal-title":"J Comput Biol"},{"key":"2024071814131602000_btae233-B11","doi-asserted-by":"crossref","first-page":"2354","DOI":"10.1021\/acs.jproteome.8b00991","article-title":"Bias in false discovery rate estimation in mass-spectrometry-based peptide identification","volume":"18","author":"Danilova","year":"2019","journal-title":"J Proteome Res"},{"key":"2024071814131602000_btae233-B12","first-page":"54","article-title":"Bootstrap methods for standard errors, confidence intervals, and other measures of statistical accuracy","volume":"1","author":"Efron","year":"1986","journal-title":"Stat Sci"},{"key":"2024071814131602000_btae233-B13","doi-asserted-by":"crossref","first-page":"207","DOI":"10.1038\/nmeth1019","article-title":"Target-decoy search strategy for increased confidence in large-scale protein identifications by mass spectrometry","volume":"4","author":"Elias","year":"2007","journal-title":"Nat Methods"},{"key":"2024071814131602000_btae233-B14","doi-asserted-by":"crossref","first-page":"964","DOI":"10.1021\/ac048788h","article-title":"PepNovo: de novo peptide sequencing via probabilistic network modeling","volume":"77","author":"Frank","year":"2005","journal-title":"Anal Chem"},{"key":"2024071814131602000_btae233-B15","doi-asserted-by":"crossref","first-page":"47","DOI":"10.4310\/SII.2012.v5.n1.a5","article-title":"Bayesian false discovery rates for post-translational modification proteomics","volume":"5","author":"Fu","year":"2012","journal-title":"Stat Interface"},{"key":"2024071814131602000_btae233-B16","doi-asserted-by":"crossref","first-page":"1111","DOI":"10.1007\/s13361-011-0139-3","article-title":"Target-decoy approach and false discovery rate: when things may go wrong","volume":"22","author":"Gupta","year":"2011","journal-title":"J Am Soc Mass Spectrom"},{"key":"2024071814131602000_btae233-B17","author":"He","year":"2015"},{"key":"2024071814131602000_btae233-B18","first-page":"271","article-title":"A probabilistic representation of the \u2018skew-normal\u2019 distribution","volume":"13","author":"Henze","year":"1986","journal-title":"Scand J Stat"},{"key":"2024071814131602000_btae233-B19","doi-asserted-by":"crossref","first-page":"24","DOI":"10.1016\/j.jbiotec.2017.06.1201","article-title":"Challenges and perspectives of metaproteomic data analysis","volume":"261","author":"Heyer","year":"2017","journal-title":"J Biotechnol"},{"key":"2024071814131602000_btae233-B20","doi-asserted-by":"crossref","first-page":"2190","DOI":"10.1021\/pr501321h","article-title":"Kojak: efficient analysis of chemically cross-linked protein complexes","volume":"14","author":"Hoopmann","year":"2015","journal-title":"J Proteome Res"},{"key":"2024071814131602000_btae233-B21","doi-asserted-by":"crossref","first-page":"S2","DOI":"10.1186\/1471-2105-13-S16-S2","article-title":"False discovery rates in spectral identification","volume":"13 Suppl 16","author":"Jeong","year":"2012","journal-title":"BMC Bioinformatics"},{"key":"2024071814131602000_btae233-B22","doi-asserted-by":"crossref","first-page":"1830","DOI":"10.1021\/acs.jproteome.6b00004","article-title":"XLSearch: a probabilistic database search algorithm for identifying cross-linked peptides","volume":"15","author":"Ji","year":"2016","journal-title":"J Proteome Res"},{"key":"2024071814131602000_btae233-B23","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1021\/pr700600n","article-title":"Assigning significance to peptides identified by tandem mass spectrometry using decoy databases","volume":"7","author":"K\u00e4ll","year":"2008","journal-title":"J Proteome Res"},{"key":"2024071814131602000_btae233-B24","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1021\/pr700739d","article-title":"Posterior error probabilities and false discovery rates: two sides of the same coin","volume":"7","author":"K\u00e4ll","year":"2008","journal-title":"J Proteome Res"},{"key":"2024071814131602000_btae233-B25","doi-asserted-by":"crossref","first-page":"5383","DOI":"10.1021\/ac025747h","article-title":"Empirical statistical model to estimate the accuracy of peptide identifications made by MS\/MS and database search","volume":"74","author":"Keller","year":"2002","journal-title":"Anal Chem"},{"key":"2024071814131602000_btae233-B26","doi-asserted-by":"crossref","first-page":"5277","DOI":"10.1038\/ncomms6277","article-title":"MS-GF+ makes progress towards a universal database search tool for proteomics","volume":"5","author":"Kim","year":"2014","journal-title":"Nat Commun"},{"key":"2024071814131602000_btae233-B27","doi-asserted-by":"crossref","first-page":"513","DOI":"10.1038\/nmeth.4256","article-title":"MSFragger: ultrafast and comprehensive peptide identification in mass spectrometry-based proteomics","volume":"14","author":"Kong","year":"2017","journal-title":"Nat Methods"},{"key":"2024071814131602000_btae233-B28","author":"Li","year":"2008"},{"key":"2024071814131602000_btae233-B29","doi-asserted-by":"crossref","first-page":"1672","DOI":"10.1074\/mcp.M114.045724","article-title":"An integrated platform for isolation, processing, and mass spectrometry-based proteomic profiling of rare cells in whole blood","volume":"14","author":"Li","year":"2015","journal-title":"Mol Cell Proteomics"},{"key":"2024071814131602000_btae233-B30","doi-asserted-by":"crossref","first-page":"S4","DOI":"10.1186\/1471-2105-13-S16-S4","article-title":"Computational approaches to protein inference in shotgun proteomics","volume":"13(Suppl 16)","author":"Li","year":"2012","journal-title":"BMC Bioinformatics"},{"key":"2024071814131602000_btae233-B31","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1016\/j.jmva.2008.04.010","article-title":"Maximum likelihood estimation for multivariate skew normal mixture models","volume":"100","author":"Lin","year":"2009","journal-title":"J Multivar Anal"},{"key":"2024071814131602000_btae233-B32","first-page":"909","article-title":"Finite mixture modelling using the skew normal distribution","volume":"17","author":"Lin","year":"2007","journal-title":"Stat Sinica"},{"key":"2024071814131602000_btae233-B33","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1093\/biomet\/80.2.267","article-title":"Maximum likelihood estimation via the ECM algorithm: a general framework","volume":"80","author":"Meng","year":"1993","journal-title":"Biometrika"},{"key":"2024071814131602000_btae233-B34","doi-asserted-by":"crossref","first-page":"2092","DOI":"10.1016\/j.jprot.2010.08.009","article-title":"A survey of computational methods and error rate estimation procedures for peptide and protein identification in shotgun proteomics","volume":"73","author":"Nesvizhskii","year":"2010","journal-title":"J Proteomics"},{"key":"2024071814131602000_btae233-B35","doi-asserted-by":"crossref","first-page":"2157","DOI":"10.1074\/mcp.TIR120.002186","article-title":"OpenPepXL: an open-source tool for sensitive identification of cross-linked peptides in XL-MS","volume":"19","author":"Netz","year":"2020","journal-title":"Mol Cell Proteomics"},{"key":"2024071814131602000_btae233-B36","doi-asserted-by":"crossref","first-page":"i745","DOI":"10.1093\/bioinformatics\/btaa807","article-title":"New mixture models for decoy-free false discovery rate estimation in mass spectrometry proteomics","volume":"36","author":"Peng","year":"2020","journal-title":"Bioinformatics"},{"key":"2024071814131602000_btae233-B37","doi-asserted-by":"crossref","first-page":"D543","DOI":"10.1093\/nar\/gkab1038","article-title":"The PRIDE database resources in 2022: a hub for mass spectrometry-based proteomics evidences","volume":"50","author":"Perez-Riverol","year":"2022","journal-title":"Nucleic Acids Res"},{"key":"2024071814131602000_btae233-B38","doi-asserted-by":"crossref","first-page":"3551","DOI":"10.1002\/(SICI)1522-2683(19991201)20:18<3551::AID-ELPS3551>3.0.CO;2-2","article-title":"Probability-based protein identification by searching sequence databases using mass spectrometry data","volume":"20","author":"Perkins","year":"1999","journal-title":"Electrophoresis"},{"key":"2024071814131602000_btae233-B39","doi-asserted-by":"crossref","first-page":"7500","DOI":"10.1021\/acs.chemrev.1c00786","article-title":"Cross-linking mass spectrometry for investigating protein conformations and protein-protein interactions \u2013 a method for all seasons","volume":"122","author":"Piersimoni","year":"2021","journal-title":"Chem Rev"},{"key":"2024071814131602000_btae233-B40","doi-asserted-by":"crossref","first-page":"530","DOI":"10.1016\/j.jsb.2010.10.014","article-title":"The beginning of a beautiful friendship: cross-linking\/mass spectrometry and modelling of proteins and multi-protein complexes","volume":"173","author":"Rappsilber","year":"2011","journal-title":"J Struct Biol"},{"key":"2024071814131602000_btae233-B41","doi-asserted-by":"crossref","first-page":"315","DOI":"10.1038\/nmeth.1192","article-title":"Identification of cross-linked peptides from large sequence databases","volume":"5","author":"Rinner","year":"2008","journal-title":"Nat Methods"},{"key":"2024071814131602000_btae233-B42","doi-asserted-by":"crossref","first-page":"3","DOI":"10.4310\/SII.2012.v5.n1.a2","article-title":"A review of statistical methods for protein identification using tandem mass spectrometry","volume":"5","author":"Serang","year":"2012","journal-title":"Stat Interface"},{"key":"2024071814131602000_btae233-B43","doi-asserted-by":"crossref","first-page":"1225","DOI":"10.1002\/jms.559","article-title":"Chemical cross-linking and mass spectrometry for mapping three-dimensional structures of proteins and protein complexes","volume":"38","author":"Sinz","year":"2003","journal-title":"J Mass Spectrom"},{"key":"2024071814131602000_btae233-B44","doi-asserted-by":"crossref","first-page":"663","DOI":"10.1002\/mas.20082","article-title":"Chemical cross-linking and mass spectrometry to map three-dimensional protein structures and protein-protein interactions","volume":"25","author":"Sinz","year":"2006","journal-title":"Mass Spectrom Rev"},{"key":"2024071814131602000_btae233-B45","doi-asserted-by":"crossref","first-page":"699","DOI":"10.1038\/nrm1468","article-title":"The ABC\u2019s (and XYZ\u2019s) of peptide sequencing","volume":"5","author":"Steen","year":"2004","journal-title":"Nat Rev Mol Cell Biol"},{"key":"2024071814131602000_btae233-B46","doi-asserted-by":"crossref","first-page":"479","DOI":"10.1111\/1467-9868.00346","article-title":"A direct approach to false discovery rate","volume":"64","author":"Storey","year":"2002","journal-title":"J R Statist Soc B"},{"key":"2024071814131602000_btae233-B47","doi-asserted-by":"crossref","first-page":"901","DOI":"10.1038\/nmeth.2103","article-title":"False discovery rate estimation for cross-linked peptides identified by mass spectrometry","volume":"9","author":"Walzthoeni","year":"2012","journal-title":"Nat Methods"},{"key":"2024071814131602000_btae233-B48","doi-asserted-by":"crossref","first-page":"904","DOI":"10.1038\/nmeth.2099","article-title":"Identification of cross-linked peptides from complex samples","volume":"9","author":"Yang","year":"2012","journal-title":"Nat Methods"},{"key":"2024071814131602000_btae233-B49","doi-asserted-by":"crossref","first-page":"1426","DOI":"10.1021\/ac00104a020","article-title":"Method to correlate tandem mass spectra of modified peptides to amino acid sequences in the protein database","volume":"67","author":"Yates","year":"1995","journal-title":"Anal Chem"},{"key":"2024071814131602000_btae233-B50","first-page":"455","article-title":"Algorithm as 76: an integral useful in calculating non-Central t and bivariate normal probabilities","volume":"23","author":"Young","year":"1974","journal-title":"J R Statist Soc C"},{"key":"2024071814131602000_btae233-B51","doi-asserted-by":"crossref","first-page":"144","DOI":"10.1021\/acs.analchem.7b04431","article-title":"Cross-linking mass spectrometry (XL-MS): an emerging technology for interactomics and structural biology","volume":"90","author":"Yu","year":"2018","journal-title":"Anal Chem"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/40\/Supplement_1\/i428\/58585822\/btae233.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/40\/Supplement_1\/i428\/58585822\/btae233.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,18]],"date-time":"2024-07-18T15:36:25Z","timestamp":1721316985000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/40\/Supplement_1\/i428\/7700896"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,28]]},"references-count":51,"journal-issue":{"issue":"Supplement_1","published-print":{"date-parts":[[2024,6,28]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btae233","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024,7]]},"published":{"date-parts":[[2024,6,28]]}}}