{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,20]],"date-time":"2026-02-20T20:43:19Z","timestamp":1771620199822,"version":"3.50.1"},"reference-count":34,"publisher":"Oxford University Press (OUP)","issue":"12","license":[{"start":{"date-parts":[[2016,10,28]],"date-time":"2016-10-28T00:00:00Z","timestamp":1477612800000},"content-version":"vor","delay-in-days":139,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2016,6,15]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: An important problematic of metabolomics is to identify metabolites using tandem mass spectrometry data. Machine learning methods have been proposed recently to solve this problem by predicting molecular fingerprint vectors and matching these fingerprints against existing molecular structure databases. In this work we propose to address the metabolite identification problem using a structured output prediction approach. This type of approach is not limited to vector output space and can handle structured output space such as the molecule space.<\/jats:p><jats:p>Results: We use the Input Output Kernel Regression method to learn the mapping between tandem mass spectra and molecular structures. The principle of this method is to encode the similarities in the input (spectra) space and the similarities in the output (molecule) space using two kernel functions. This method approximates the spectra-molecule mapping in two phases. The first phase corresponds to a regression problem from the input space to the feature space associated to the output kernel. The second phase is a preimage problem, consisting in mapping back the predicted output feature vectors to the molecule space. We show that our approach achieves state-of-the-art accuracy in metabolite identification. Moreover, our method has the advantage of decreasing the running times for the training step and the test step by several orders of magnitude over the preceding methods.<\/jats:p><jats:p>Availability and implementation :<\/jats:p><jats:p>Contact: \u00a0celine.brouard@aalto.fi<\/jats:p><jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btw246","type":"journal-article","created":{"date-parts":[[2016,6,15]],"date-time":"2016-06-15T15:43:52Z","timestamp":1466005432000},"page":"i28-i36","source":"Crossref","is-referenced-by-count":76,"title":["Fast metabolite identification with Input Output Kernel Regression"],"prefix":"10.1093","volume":"32","author":[{"given":"C\u00e9line","family":"Brouard","sequence":"first","affiliation":[{"name":"1 Department of Computer Science, Aalto University, Espoo, Finland"},{"name":"2 Helsinki Institute for Information Technology, Espoo, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huibin","family":"Shen","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science, Aalto University, Espoo, Finland"},{"name":"2 Helsinki Institute for Information Technology, Espoo, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kai","family":"D\u00fchrkop","sequence":"additional","affiliation":[{"name":"3 Chair for Bioinformatics, Friedrich-Schiller University, Jena, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Florence","family":"d'Alch\u00e9-Buc","sequence":"additional","affiliation":[{"name":"4 LTCI, CNRS, T\u00e9l\u00e9com ParisTech, Universit\u00e9 Paris-Saclay, Paris, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sebastian","family":"B\u00f6cker","sequence":"additional","affiliation":[{"name":"3 Chair for Bioinformatics, Friedrich-Schiller University, Jena, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Juho","family":"Rousu","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science, Aalto University, Espoo, Finland"},{"name":"2 Helsinki Institute for Information Technology, Espoo, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2016,6,11]]},"reference":[{"key":"2023020112305102800_btw246-B1","doi-asserted-by":"crossref","first-page":"W94","DOI":"10.1093\/nar\/gku436","article-title":"CFM-ID: a web server for annotation, spectrum prediction and metabolite identification from tandem mass spectra","volume":"42","author":"Allen","year":"2014","journal-title":"Nucleic Acids Res"},{"key":"2023020112305102800_btw246-B2","doi-asserted-by":"crossref","first-page":"98","DOI":"10.1007\/s11306-014-0676-4","article-title":"Competitive fragmentation modeling of ESI-MS\/MS spectra for putative metabolite identification","volume":"11","author":"Allen","year":"2015","journal-title":"Metabolomics"},{"key":"2023020112305102800_btw246-B3","doi-asserted-by":"crossref","first-page":"i49","DOI":"10.1093\/bioinformatics\/btn270","article-title":"Towards de novo identification of metabolites by analyzing tandem mass spectra","volume":"24","author":"B\u00f6cker","year":"2008","journal-title":"Bioinfomatics"},{"key":"2023020112305102800_btw246-B4","first-page":"217","article-title":"PubChem: Integrated platform of small molecules and biological activities","volume":"4","author":"Bolton","year":"2008","journal-title":"Chapter 12 in Annual Reports in Computational Chemistry"},{"key":"2023020112305102800_btw246-B5","author":"Brouard","year":"2011"},{"key":"2023020112305102800_btw246-B6","author":"Brouard","year":"2015"},{"key":"2023020112305102800_btw246-B7","author":"Cortes","year":"2005"},{"key":"2023020112305102800_btw246-B8","first-page":"795","article-title":"Algorithms for learning kernels based on centered alignment","volume":"13","author":"Cortes","year":"2012","journal-title":"J. Mach. Learn. Res"},{"key":"2023020112305102800_btw246-B9","doi-asserted-by":"crossref","first-page":"12549","DOI":"10.1073\/pnas.1516878112","article-title":"Illuminating the dark matter in metabolomics","volume":"112","author":"da Silva","year":"2015","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"2023020112305102800_btw246-B10","doi-asserted-by":"crossref","first-page":"12580","DOI":"10.1073\/pnas.1509788112","article-title":"Searching molecular structure databases with tandem mass spectra using CSI:FingerID","volume":"112","author":"D\u00fchrkop","year":"2015","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"2023020112305102800_btw246-B11","first-page":"615","article-title":"Learning multiple tasks with kernel methods","volume":"6","author":"Evgeniou","year":"2005","journal-title":"J. Mach. Learn. Res"},{"key":"2023020112305102800_btw246-B12","author":"Geurts","year":"2006"},{"key":"2023020112305102800_btw246-B13","doi-asserted-by":"crossref","first-page":"D456","DOI":"10.1093\/nar\/gks1146","article-title":"The ChEBI reference database and ontology for biologically relevant chemistry: enhancements for 2013","volume":"41","author":"Hastings","year":"2013","journal-title":"Nucleic Acids Res"},{"key":"2023020112305102800_btw246-B14","doi-asserted-by":"crossref","first-page":"3043","DOI":"10.1002\/rcm.3701","article-title":"FiD: A software for ab initio structural identification of product ions from tandem mass spectrometric data","volume":"22","author":"Heinonen","year":"2008","journal-title":"Rapid Commun. Mass Spectrom"},{"key":"2023020112305102800_btw246-B15","doi-asserted-by":"crossref","first-page":"2333","DOI":"10.1093\/bioinformatics\/bts437","article-title":"Metabolite identification and molecular fingerprint prediction through machine learning","volume":"28","author":"Heinonen","year":"2012","journal-title":"Bioinformatics"},{"key":"2023020112305102800_btw246-B16","doi-asserted-by":"crossref","first-page":"3111","DOI":"10.1002\/rcm.2177","article-title":"Automated assignment of high-resolution collisionally activated dissociation mass spectra using a systematic bond disconnection approach","volume":"19","author":"Hill","year":"2005","journal-title":"Rapid Commun. Mass Spectrom"},{"key":"2023020112305102800_btw246-B17","doi-asserted-by":"crossref","first-page":"703","DOI":"10.1002\/jms.1777","article-title":"MassBank: A public repository for sharing mass spectral data for life sciences","volume":"45","author":"Horai","year":"2010","journal-title":"J. Mass Spectrom"},{"key":"2023020112305102800_btw246-B18","author":"Kadri","year":"2010"},{"key":"2023020112305102800_btw246-B19","author":"Kadri","year":"2013"},{"key":"2023020112305102800_btw246-B20","doi-asserted-by":"crossref","first-page":"489","DOI":"10.1007\/s10994-014-5479-3","article-title":"Operator-valued kernel-based vector autoregressive models for network inference","volume":"99","author":"Lim","year":"2014","journal-title":"Mach. Learn"},{"key":"2023020112305102800_btw246-B21","volume-title":"Applications of Artificial Intelligence for Organic Chemistry: The DENDRAL Project","author":"Lindsay","year":"1980"},{"key":"2023020112305102800_btw246-B22","first-page":"873","volume-title":"Advances in Neural Information Processing Systems","author":"Marchand","year":"2014"},{"key":"2023020112305102800_btw246-B23","doi-asserted-by":"crossref","first-page":"177","DOI":"10.1162\/0899766052530802","article-title":"On learning vector-valued functions","volume":"17","author":"Micchelli","year":"2005","journal-title":"Neural Comput"},{"key":"2023020112305102800_btw246-B24","doi-asserted-by":"crossref","first-page":"6033","DOI":"10.1021\/ac400861a","article-title":"Automatic chemical structure annotation of an LC\u2013MS n based metabolic profile from green tea","volume":"85","author":"Ridder","year":"2013","journal-title":"Anal. Chem"},{"key":"2023020112305102800_btw246-B25","doi-asserted-by":"crossref","first-page":"105","DOI":"10.7551\/mitpress\/7443.003.0010","volume-title":"Predicting Structured Data","author":"Rousu","year":"2007"},{"key":"2023020112305102800_btw246-B26","doi-asserted-by":"crossref","first-page":"665","DOI":"10.1007\/BF01630739","article-title":"Hilbert spaces of operator-valued functions","volume":"13","author":"Senkene","year":"1973","journal-title":"Lithuanian Math. J"},{"key":"2023020112305102800_btw246-B27","doi-asserted-by":"crossref","first-page":"484","DOI":"10.3390\/metabo3020484","article-title":"Metabolite identification through machine learning\u2013tackling CASMI challenge using FingerID","volume":"3","author":"Shen","year":"2013","journal-title":"Metabolites"},{"key":"2023020112305102800_btw246-B28","doi-asserted-by":"crossref","first-page":"i157","DOI":"10.1093\/bioinformatics\/btu275","article-title":"Metabolite identification through multiple kernel learning on fragmentation trees","volume":"30","author":"Shen","year":"2014","journal-title":"Bioinformatics"},{"key":"2023020112305102800_btw246-B29","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1007\/s10994-014-5465-9","article-title":"Multilabel classification through random graph ensembles","volume":"99","author":"Su","year":"2015","journal-title":"Mach. Learn"},{"key":"2023020112305102800_btw246-B30","first-page":"25","article-title":"Max-margin Markov networks","volume":"16","author":"Taskar","year":"2004","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"2023020112305102800_btw246-B31","author":"Tsochantaridis","year":"2004"},{"key":"2023020112305102800_btw246-B32","doi-asserted-by":"crossref","first-page":"9496","DOI":"10.1021\/ac5014783","article-title":"MIDAS: a database-searching algorithm for metabolite identification in metabolomics","volume":"86","author":"Wang","year":"2014","journal-title":"Anal. Chem"},{"key":"2023020112305102800_btw246-B33","volume-title":"Advances in Neural Information Processing Systems 15","author":"Weston","year":"2003"},{"key":"2023020112305102800_btw246-B34","doi-asserted-by":"crossref","first-page":"148.","DOI":"10.1186\/1471-2105-11-148","article-title":"In silico fragmentation for computer assisted identification of metabolite mass spectra","volume":"11","author":"Wolf","year":"2010","journal-title":"BMC Bioinformatics"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/32\/12\/i28\/49020619\/bioinformatics_32_12_i28.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/32\/12\/i28\/49020619\/bioinformatics_32_12_i28.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,17]],"date-time":"2024-06-17T17:03:21Z","timestamp":1718643801000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/32\/12\/i28\/2288626"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,6,11]]},"references-count":34,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2016,6,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btw246","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2016,6,15]]},"published":{"date-parts":[[2016,6,11]]}}}