{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,5,13]],"date-time":"2023-05-13T09:14:28Z","timestamp":1683969268440},"reference-count":30,"publisher":"Oxford University Press (OUP)","issue":"17","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2013,9,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Protein domain classification is an important step in functional annotation for next-generation sequencing data. For RNA-Seq data of non-model organisms that lack quality or complete reference genomes, existing protein domain analysis pipelines are applied to short reads directly or to contigs that are generated using de novo sequence assembly tools. However, these strategies do not provide satisfactory performance in classifying short reads into their native domain families.<\/jats:p>\n               <jats:p>Results: We introduce SALT, a protein domain classification tool based on profile hidden Markov models and graph algorithms. SALT carefully incorporates the characteristics of reads that are sequenced from the domain regions and assembles them into contigs based on a supervised graph construction algorithm. We applied SALT to two RNA-Seq datasets of different read lengths and quantified its performance using the available protein domain annotations and the reference genomes. Compared with existing strategies, SALT showed better sensitivity and accuracy. In the third experiment, we applied SALT to a non-model organism. The experimental results demonstrated that it identified more transcribed protein domain families than other tested classifiers.<\/jats:p>\n               <jats:p>Availability: The source code and supplementary data are available at https:\/\/sourceforge.net\/projects\/salt1\/<\/jats:p>\n               <jats:p>Contact: \u00a0yannisun@msu.edu<\/jats:p>\n               <jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btt357","type":"journal-article","created":{"date-parts":[[2013,6,20]],"date-time":"2013-06-20T00:15:39Z","timestamp":1371687339000},"page":"2103-2111","source":"Crossref","is-referenced-by-count":10,"title":["A Sensitive and Accurate protein domain cLassification Tool (SALT) for short reads"],"prefix":"10.1093","volume":"29","author":[{"given":"Yuan","family":"Zhang","sequence":"first","affiliation":[{"name":"1 Department of Computer Science and Engineering and 2Center for Microbial Ecology, Michigan State University, East Lansing, MI 48824, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yanni","family":"Sun","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science and Engineering and 2Center for Microbial Ecology, Michigan State University, East Lansing, MI 48824, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"James R.","family":"Cole","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science and Engineering and 2Center for Microbial Ecology, Michigan State University, East Lansing, MI 48824, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2013,6,19]]},"reference":[{"key":"2023012810463756600_btt357-B1","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1016\/S0022-2836(05)80360-2","article-title":"Basisc local alignment search tool","volume":"215","author":"Altschul","year":"1990","journal-title":"J. Mol. Biol."},{"key":"2023012810463756600_btt357-B2","article-title":"A comparative study of k-shortest path algorithms","volume-title":"Proceedings of 11th UK Performance Engineering Workshop for Computer and Telecommunications Systems","author":"Brander","year":"1995"},{"key":"2023012810463756600_btt357-B3","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511790492","volume-title":"Biological Sequence Analysis Probabilistic Models of Proteins and Nucleic Acids","author":"Durbin","year":"1998"},{"key":"2023012810463756600_btt357-B4","first-page":"205","article-title":"A new generation of homology search tools based on probabilistic inference","volume":"23","author":"Eddy","year":"2009","journal-title":"Genome Inform."},{"key":"2023012810463756600_btt357-B5","doi-asserted-by":"crossref","first-page":"154","DOI":"10.1109\/SFCS.1994.365697","article-title":"Finding the k shortest paths","volume-title":"Proceedings of 25th IEEE Annual Symposium on Foundation of Computer Science","author":"Eppstain","year":"1994"},{"key":"2023012810463756600_btt357-B6","doi-asserted-by":"crossref","first-page":"e1000074","DOI":"10.1371\/journal.pcbi.1000074","article-title":"Viral population estimation using pyrosequencing","volume":"4","author":"Eriksson","year":"2008","journal-title":"PLoS Comput. Biol."},{"key":"2023012810463756600_btt357-B7","doi-asserted-by":"crossref","first-page":"317","DOI":"10.1186\/1471-2164-12-317","article-title":"Short read Illumina data for the de novo assembly of a non-model snail species transcriptome (Radix balthica, Basommatophora, Pulmonata), and a comparison of assembler performance","volume":"12","author":"Feldmeyer","year":"2012","journal-title":"BMC Genomics"},{"key":"2023012810463756600_btt357-B8","doi-asserted-by":"crossref","first-page":"D211","DOI":"10.1093\/nar\/gkp985","article-title":"The Pfam protein families database","volume":"38","author":"Finn","year":"2010","journal-title":"Nucleic Acids Res."},{"key":"2023012810463756600_btt357-B9","doi-asserted-by":"crossref","first-page":"371","DOI":"10.1093\/nar\/gkg128","article-title":"The TIGRFAMs database of protein families","volume":"31","author":"Haft","year":"2003","journal-title":"Nucleic Acids Res."},{"key":"2023012810463756600_btt357-B10","doi-asserted-by":"crossref","first-page":"D211","DOI":"10.1093\/nar\/gkn785","article-title":"InterPro: the integrative protein signature database","volume":"37","author":"Hunter","year":"2009","journal-title":"Nucleic Acids Res."},{"key":"2023012810463756600_btt357-B11","doi-asserted-by":"crossref","first-page":"671","DOI":"10.1038\/nrg3068","article-title":"Next-generation transcriptome assembly","volume":"12","author":"Jeffrey","year":"2011","journal-title":"Nature Rev. Genet."},{"key":"2023012810463756600_btt357-B12","doi-asserted-by":"crossref","first-page":"R25","DOI":"10.1186\/gb-2009-10-3-r25","article-title":"Ultrafast and memory-efficient alignment of short dna sequences to the human genome","volume":"10","author":"Langmead","year":"2009","journal-title":"Genome Biol."},{"key":"2023012810463756600_btt357-B13","doi-asserted-by":"crossref","first-page":"540","DOI":"10.1186\/1471-2164-12-540","article-title":"RNA-seq improves annotation of protein-coding genes in the cucumber genome","volume":"12","author":"Li","year":"2011","journal-title":"BMC Genomics"},{"key":"2023012810463756600_btt357-B14","doi-asserted-by":"crossref","first-page":"1184","DOI":"10.1101\/gr.134106.111","article-title":"Transcriptome survey reveals increased complexity of the alternative splicing landscape in Arabidopsis","volume":"22","author":"Marquez","year":"2012","journal-title":"Genome Res."},{"key":"2023012810463756600_btt357-B15","doi-asserted-by":"crossref","first-page":"6643","DOI":"10.1093\/nar\/gkp698","article-title":"FIGfams: yet another set of protein families","volume":"37","author":"Meyer","year":"2009","journal-title":"Nucleic Acids Res"},{"key":"2023012810463756600_btt357-B16","doi-asserted-by":"crossref","first-page":"315","DOI":"10.1016\/j.ygeno.2010.03.001","article-title":"Assembly algorithms for next-generation sequencing data","volume":"95","author":"Miller","year":"2010","journal-title":"Genomics"},{"key":"2023012810463756600_btt357-B17","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1186\/1471-2164-13-99","article-title":"A new RNAseq-based reference transcriptome for sugar beet and its application in transcriptome-scale analysis of vernalization and gibberellin responses","volume":"13","author":"Mutasa-G\u00f6ttgens","year":"2012","journal-title":"BMC Genomics"},{"key":"2023012810463756600_btt357-B18","doi-asserted-by":"crossref","first-page":"e41150","DOI":"10.1371\/journal.pone.0041150","article-title":"RNA-seq analysis of the Sclerotinia homoeocarpa creeping bentgrass pathosystem","volume":"7","author":"Orshinsky","year":"2012","journal-title":"PLoS One"},{"key":"2023012810463756600_btt357-B19","doi-asserted-by":"crossref","first-page":"W116","DOI":"10.1093\/nar\/gki442","article-title":"InterProScan: protein domains identifier","volume":"33","author":"Quevillon","year":"2005","journal-title":"Nucleic Acids Res."},{"key":"2023012810463756600_btt357-B20","doi-asserted-by":"crossref","first-page":"173","DOI":"10.1038\/nmeth.1818","article-title":"HHblits: lightning-fast iterative protein sequence searching by HMM-HMM alignment","volume":"9","author":"Remmert","year":"2012","journal-title":"Nat. Methods"},{"key":"2023012810463756600_btt357-B21","doi-asserted-by":"crossref","first-page":"1165","DOI":"10.1101\/gr.101360.109","article-title":"Assembly of large genomes using second-generation sequencing","volume":"20","author":"Schatz","year":"2010","journal-title":"Genome Res."},{"key":"2023012810463756600_btt357-B22","doi-asserted-by":"crossref","first-page":"e29685","DOI":"10.1371\/journal.pone.0029685","article-title":"A powerful method for transcriptional profiling of specific cell types in eukaryotes: laser-assisted microdissection and RNA sequencing","volume":"7","author":"Schmid","year":"2012","journal-title":"PLoS One"},{"key":"2023012810463756600_btt357-B23","doi-asserted-by":"crossref","first-page":"1086","DOI":"10.1093\/bioinformatics\/bts094","article-title":"Oases: robust de novo RNA-seq assembly across the dynamic range of expression levels","volume":"28","author":"Schulz","year":"2012","journal-title":"Bioinformatics"},{"key":"2023012810463756600_btt357-B24","doi-asserted-by":"crossref","first-page":"500","DOI":"10.1093\/bioinformatics\/btl629","article-title":"Assembling millions of short DNA sequences using SSAKE","volume":"23","author":"Warren","year":"2007","journal-title":"Bioinformatics"},{"key":"2023012810463756600_btt357-B25","doi-asserted-by":"crossref","first-page":"1453","DOI":"10.1128\/AEM.02181-07","article-title":"Metagenomics: read length matters","volume":"74","author":"Wommack","year":"2008","journal-title":"Appl. Environ. Microbiol."},{"key":"2023012810463756600_btt357-B26","doi-asserted-by":"crossref","first-page":"712","DOI":"10.1287\/mnsc.17.11.712","article-title":"Finding the K shortest loopless paths in a network","volume":"17","author":"Yen","year":"1971","journal-title":"Manag. Sci."},{"key":"2023012810463756600_btt357-B27","doi-asserted-by":"crossref","first-page":"3976","DOI":"10.1073\/pnas.0813403106","article-title":"Mapping the burkholderia cenocepacia niche response via high-throughput sequencing","volume":"106","author":"Yoder-Himes","year":"2009","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012810463756600_btt357-B28","doi-asserted-by":"crossref","first-page":"821","DOI":"10.1101\/gr.074492.107","article-title":"Velvet: algorithms for de novo short read assembly using de Bruijn graphs","volume":"18","author":"Zerbino","year":"2008","journal-title":"Genome Res."},{"key":"2023012810463756600_btt357-B29","doi-asserted-by":"crossref","first-page":"198","DOI":"10.1186\/1471-2105-12-198","article-title":"HMM-FRAME: accurate protein domain classification for metagenomic sequences containing frameshift errors","volume":"12","author":"Zhang","year":"2011","journal-title":"BMC Bioinformatics"},{"key":"2023012810463756600_btt357-B30","article-title":"MetaDomain: a profile HMM-based protein domain classification tool for short sequences","volume-title":"Proceedings of Pacific Symposium on Biocomputing (PSB)","author":"Zhang","year":"2012"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/29\/17\/2103\/48892310\/bioinformatics_29_17_2103.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/29\/17\/2103\/48892310\/bioinformatics_29_17_2103.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,28]],"date-time":"2023-01-28T12:35:13Z","timestamp":1674909313000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/29\/17\/2103\/241124"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,6,19]]},"references-count":30,"journal-issue":{"issue":"17","published-print":{"date-parts":[[2013,9,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btt357","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2013,9,1]]},"published":{"date-parts":[[2013,6,19]]}}}