{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,21]],"date-time":"2025-11-21T06:03:00Z","timestamp":1763704980920},"reference-count":21,"publisher":"Oxford University Press (OUP)","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2015,2,15]]},"abstract":"<jats:p>Motivation: Tandem mass spectrometry (MS) has become the method of choice for protein identification and quantification. In the era of big data biology, tandem mass spectra are often searched against huge protein databases generated from genomes or RNA-Seq data for peptide identification. However, most existing tools for MS-based peptide identification compare a tandem mass spectrum against all peptides in a database whose molecular masses are similar to the precursor mass of the spectrum, making mass spectral data analysis slow for huge databases. Tag-based methods extract peptide sequence tags from a tandem mass spectrum and use them as a filter to reduce the number of candidate peptides, thus speeding up the database search. Recently, gapped tags have been introduced into mass spectral data analysis because they improve the sensitivity of peptide identification compared with sequence tags. However, the blocked pattern matching (BPM) problem, which is an essential step in gapped tag-based peptide identification, has not been fully solved.<\/jats:p>\n               <jats:p>Results: In this article, we propose a fast and memory-efficient algorithm for the BPM problem. Experiments on both simulated and real datasets showed that the proposed algorithm achieved high speed and high sensitivity for peptide filtration in peptide identification by database search.<\/jats:p>\n               <jats:p>Contact: \u00a0cswangl@cityu.edu.hk or xwliu@iupui.edu<\/jats:p>\n               <jats:p>Supplementary information: \u00a0Supplementary Data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btu678","type":"journal-article","created":{"date-parts":[[2014,10,17]],"date-time":"2014-10-17T05:05:29Z","timestamp":1413522329000},"page":"532-538","source":"Crossref","is-referenced-by-count":12,"title":["An efficient algorithm for the blocked pattern matching problem"],"prefix":"10.1093","volume":"31","author":[{"given":"Fei","family":"Deng","sequence":"first","affiliation":[{"name":"1 \u00a01Department of Computer Science, City University of Hong Kong, Kowloon, Hong Kong, 2Department of BioHealth Informatics, Indiana University\u2014Purdue University Indianapolis and 3Center for Computational Biology and Bioinformatics, Indiana University School of Medicine, Indianapolis, IN 46202, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lusheng","family":"Wang","sequence":"additional","affiliation":[{"name":"1 \u00a01Department of Computer Science, City University of Hong Kong, Kowloon, Hong Kong, 2Department of BioHealth Informatics, Indiana University\u2014Purdue University Indianapolis and 3Center for Computational Biology and Bioinformatics, Indiana University School of Medicine, Indianapolis, IN 46202, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaowen","family":"Liu","sequence":"additional","affiliation":[{"name":"1 \u00a01Department of Computer Science, City University of Hong Kong, Kowloon, Hong Kong, 2Department of BioHealth Informatics, Indiana University\u2014Purdue University Indianapolis and 3Center for Computational Biology and Bioinformatics, Indiana University School of Medicine, Indianapolis, IN 46202, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2014,10,15]]},"reference":[{"key":"2023020108564256600_btu678-B1","doi-asserted-by":"crossref","first-page":"641","DOI":"10.1002\/1615-9861(200104)1:5<641::AID-PROT641>3.0.CO;2-R","article-title":"Mass spectrometry allows direct identification of proteins in large genomes","volume":"1","author":"Andersen","year":"2001","journal-title":"Proteomics"},{"key":"2023020108564256600_btu678-B2","doi-asserted-by":"crossref","first-page":"e8949","DOI":"10.1371\/journal.pone.0008949","article-title":"An integrated mass-spectrometry pipeline identifies novel protein coding-regions in the human genome","volume":"5","author":"Bitton","year":"2010","journal-title":"PLoS One"},{"key":"2023020108564256600_btu678-B3","doi-asserted-by":"crossref","first-page":"2310","DOI":"10.1002\/rcm.1198","article-title":"A method for reducing the time required to match protein sequences with tandem mass spectra","volume":"17","author":"Craig","year":"2003","journal-title":"Rapid Commun. Mass Spectrom."},{"key":"2023020108564256600_btu678-B4","doi-asserted-by":"crossref","first-page":"5002","DOI":"10.1128\/JB.00542-10","article-title":"The human oral microbiome","volume":"192","author":"Dewhirst","year":"2010","journal-title":"J. Bacteriol."},{"key":"2023020108564256600_btu678-B5","doi-asserted-by":"crossref","first-page":"976","DOI":"10.1016\/1044-0305(94)80016-2","article-title":"An approach to correlate tandem mass spectral data of peptides with amino acid sequences in a protein database","volume":"5","author":"Eng","year":"1994","journal-title":"J. Am. Soc. Mass Spectrom."},{"key":"2023020108564256600_btu678-B6","doi-asserted-by":"crossref","first-page":"2377","DOI":"10.1021\/pr1011729","article-title":"Improved peptide identification by targeted fragmentation using CID, HCD and ETD on an LTQ-Orbitrap Velos","volume":"10","author":"Frese","year":"2011","journal-title":"J. Proteome Res."},{"key":"2023020108564256600_btu678-B7","doi-asserted-by":"crossref","first-page":"958","DOI":"10.1021\/pr0499491","article-title":"Open mass spectrometry search algorithm","volume":"3","author":"Geer","year":"2004","journal-title":"J. Proteome Res."},{"key":"2023020108564256600_btu678-B8","doi-asserted-by":"crossref","first-page":"M110.002220","DOI":"10.1074\/mcp.M110.002220","article-title":"Gapped spectral dictionaries and their applications for database searches of tandem mass spectra","volume":"10","author":"Jeong","year":"2011","journal-title":"Mol. Cell. Proteomics"},{"key":"2023020108564256600_btu678-B9","doi-asserted-by":"crossref","first-page":"3354","DOI":"10.1021\/pr8001244","article-title":"Spectral probabilities and generating functions of tandem mass spectra: a strike against decoy databases","volume":"7","author":"Kim","year":"2008","journal-title":"J. Proteome Res."},{"key":"2023020108564256600_btu678-B10","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1074\/mcp.M800103-MCP200","article-title":"Spectral dictionaries: integrating de novo peptide sequencing with database search of tandem mass spectra","volume":"8","author":"Kim","year":"2009","journal-title":"Mol. Cell. Proteomics"},{"key":"2023020108564256600_btu678-B11","doi-asserted-by":"crossref","first-page":"5830","DOI":"10.1021\/pr400849y","article-title":"Identification of ultramodified proteins using top-down tandem mass spectra","volume":"12","author":"Liu","year":"2013","journal-title":"J. Proteome Res."},{"key":"2023020108564256600_btu678-B12","doi-asserted-by":"crossref","first-page":"2337","DOI":"10.1002\/rcm.1196","article-title":"PEAKS: powerful software for peptide de novo sequencing by tandem mass spectrometry","volume":"17","author":"Ma","year":"2003","journal-title":"Rapid Commun. Mass Spectrom."},{"key":"2023020108564256600_btu678-B13","doi-asserted-by":"crossref","first-page":"2896","DOI":"10.1021\/pr200118r","article-title":"ScanRanker: quality assessment of tandem mass spectra via sequence tagging","volume":"10","author":"Ma","year":"2011","journal-title":"J. Proteome Res."},{"key":"2023020108564256600_btu678-B14","doi-asserted-by":"crossref","first-page":"4390","DOI":"10.1021\/ac00096a002","article-title":"Error-tolerant identification of peptides in sequence databases by peptide sequence tags","volume":"66","author":"Mann","year":"1994","journal-title":"Anal. Chem."},{"key":"2023020108564256600_btu678-B15","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-20036-6_27","article-title":"Blocked pattern matching problem and its applications in proteomics","volume-title":"Proceedings of 15th Annual International Conference on Research in Computational Molecular Biology (RECOMB 2011)","author":"Ng","year":"2011"},{"key":"2023020108564256600_btu678-B16","doi-asserted-by":"crossref","first-page":"3551","DOI":"10.1002\/(SICI)1522-2683(19991201)20:18<3551::AID-ELPS3551>3.0.CO;2-2","article-title":"Probability-based protein identification by searching sequence databases using mass spectrometry data","volume":"20","author":"Perkins","year":"1999","journal-title":"Electrophoresis"},{"key":"2023020108564256600_btu678-B17","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1111\/j.2041-1014.2009.00558.x","article-title":"A metaproteomic analysis of the human salivary microbiota by three-dimensional peptide fractionation and tandem mass spectrometry","volume":"25","author":"Rudney","year":"2010","journal-title":"Mol. Oral Microbiol."},{"key":"2023020108564256600_btu678-B18","doi-asserted-by":"crossref","first-page":"6415","DOI":"10.1021\/ac0347462","article-title":"GutenTag: high-throughput sequence tagging via an empirically derived fragmentation model","volume":"75","author":"Tabb","year":"2003","journal-title":"Anal. Chem."},{"key":"2023020108564256600_btu678-B19","doi-asserted-by":"crossref","first-page":"4626","DOI":"10.1021\/ac050102d","article-title":"InsPecT: identification of posttranslationally modified peptides from tandem mass spectra","volume":"77","author":"Tanner","year":"2005","journal-title":"Anal. Chem."},{"key":"2023020108564256600_btu678-B20","doi-asserted-by":"crossref","first-page":"249","DOI":"10.1007\/BF01206331","article-title":"On-line construction of suffix trees","volume":"14","author":"Ukkonen","year":"1995","journal-title":"Algorithmica"},{"key":"2023020108564256600_btu678-B21","doi-asserted-by":"crossref","first-page":"3202","DOI":"10.1021\/ac00114a016","article-title":"Mining genomes: correlating tandem mass spectra of modified and unmodified peptides to sequences in nucleotide databases","volume":"67","author":"Yates","year":"1995","journal-title":"Anal. Chem."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/31\/4\/532\/49011023\/bioinformatics_31_4_532.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/31\/4\/532\/49011023\/bioinformatics_31_4_532.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,1]],"date-time":"2023-02-01T20:24:43Z","timestamp":1675283083000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/31\/4\/532\/2748202"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,10,15]]},"references-count":21,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2015,2,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btu678","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published":{"date-parts":[[2014,10,15]]}}}