{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,23]],"date-time":"2026-04-23T05:57:11Z","timestamp":1776923831563,"version":"3.51.2"},"reference-count":25,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2004,10,28]],"date-time":"2004-10-28T00:00:00Z","timestamp":1098921600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/2.0\/"},{"start":{"date-parts":[[2004,10,28]],"date-time":"2004-10-28T00:00:00Z","timestamp":1098921600000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/2.0\/"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                        <jats:title>Background<\/jats:title>\n                        <jats:p>Kernel-based learning algorithms are among the most advanced machine learning methods and have been successfully applied to a variety of sequence classification tasks within the field of bioinformatics. Conventional kernels utilized so far do not provide an easy interpretation of the learnt representations in terms of positional and compositional variability of the underlying biological signals.<\/jats:p>\n                     <\/jats:sec><jats:sec>\n                        <jats:title>Results<\/jats:title>\n                        <jats:p>We propose a kernel-based approach to datamining on biological sequences. With our method it is possible to model and analyze positional variability of oligomers of any length in a natural way. On one hand this is achieved by mapping the sequences to an intuitive but high-dimensional feature space, well-suited for interpretation of the learnt models. On the other hand, by means of the kernel trick we can provide a general learning algorithm for that high-dimensional representation because all required statistics can be computed without performing an explicit feature space mapping of the sequences. By introducing a kernel parameter that controls the <jats:italic>degree<\/jats:italic> of position-dependency, our feature space representation can be tailored to the characteristics of the biological problem at hand. A regularized learning scheme enables application even to biological problems for which only small sets of example sequences are available. Our approach includes a visualization method for transparent representation of characteristic sequence features. Thereby importance of features can be measured in terms of discriminative strength with respect to classification of the underlying sequences. To demonstrate and validate our concept on a biochemically well-defined case, we analyze <jats:italic>E. coli<\/jats:italic> translation initiation sites in order to show that we can find biologically relevant signals. For that case, our results clearly show that the Shine-Dalgarno sequence is the most important signal upstream a start codon. The variability in position and composition we found for that signal is in accordance with previous biological knowledge. We also find evidence for signals downstream of the start codon, previously introduced as transcriptional enhancers. These signals are mainly characterized by occurrences of adenine in a region of about 4 nucleotides next to the start codon.<\/jats:p>\n                     <\/jats:sec><jats:sec>\n                        <jats:title>Conclusions<\/jats:title>\n                        <jats:p>We showed that the oligo kernel can provide a valuable tool for the analysis of relevant signals in biological sequences. In the case of translation initiation sites we could clearly deduce the most discriminative motifs and their positional variation from example sequences. Attractive features of our approach are its flexibility with respect to oligomer length and position conservation. By means of these two parameters oligo kernels can easily be adapted to different biological problems.<\/jats:p>\n                     <\/jats:sec>","DOI":"10.1186\/1471-2105-5-169","type":"journal-article","created":{"date-parts":[[2004,11,6]],"date-time":"2004-11-06T07:24:44Z","timestamp":1099725884000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":39,"title":["Oligo kernels for datamining on biological sequences: a case study on prokaryotic translation initiation sites"],"prefix":"10.1186","volume":"5","author":[{"given":"Peter","family":"Meinicke","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Maike","family":"Tech","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Burkhard","family":"Morgenstern","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rainer","family":"Merkl","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2004,10,28]]},"reference":[{"key":"285_CR1","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511790492","volume-title":"Biological Sequence Analysis","author":"R Durbin","year":"1998","unstructured":"Durbin R, Eddy SR, Krogh A: Biological Sequence Analysis. Cambridge University Press; 1998."},{"key":"285_CR2","volume-title":"Bioinformatics \u2013 The machine learning approach","author":"P Baldi","year":"1998","unstructured":"Baldi P, Brunak S: Bioinformatics \u2013 The machine learning approach. Massachusetts Institute of Technology Press; 1998."},{"key":"285_CR3","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511801389","volume-title":"An Introduction to Support Vector Machines and other kernel-based learning methods","author":"N Christiani","year":"2000","unstructured":"Christiani N, Shawe-Taylor J: An Introduction to Support Vector Machines and other kernel-based learning methods. Cambridge University Press; 2000."},{"key":"285_CR4","volume-title":"Solutions of ill-posed problems","author":"AN Tikhonov","year":"1977","unstructured":"Tikhonov AN, Arsenin VY: Solutions of ill-posed problems. Washington, DC: Winston; 1977."},{"issue":"Suppl 2","key":"285_CR5","doi-asserted-by":"publisher","first-page":"75","DOI":"10.1093\/bioinformatics\/18.suppl_2.S75","volume":"18","author":"S Degroeve","year":"2002","unstructured":"Degroeve S, Beats BD, de Peer YV, Rouz\u00e9 P: Feature subset selection for splice site prediction.\n                           Bioinformatics 2002, 18(Suppl 2):75\u201383.","journal-title":"Bioinformatics"},{"key":"285_CR6","volume-title":"Learning with Kernels","author":"B Sch\u00f6lkopf","year":"2002","unstructured":"Sch\u00f6lkopf B, Smola A: Learning with Kernels. MIT Press; 2002."},{"issue":"9","key":"285_CR7","doi-asserted-by":"publisher","first-page":"799","DOI":"10.1093\/bioinformatics\/16.9.799","volume":"16","author":"A Zien","year":"2000","unstructured":"Zien A, R\u00e4tsch G, Mika S, Sch\u00f6lkopf B, Lengauer T, M\u00fcller K: Engineering Support Vector Machine kernels that recognize translation initiation sites.\n                           Bioinformatics 2000, 16(9):799\u2013807. 10.1093\/bioinformatics\/16.9.799","journal-title":"Bioinformatics"},{"key":"285_CR8","first-page":"564","volume-title":"In Proceedings of the Pacific Symposium on Biocomputing, Stanford","author":"C Leslie","year":"2002","unstructured":"Leslie C, Eskin E, Noble W: The Spectrum Kernel: A string kernel for SVM protein classification.\n                           In Proceedings of the Pacific Symposium on Biocomputing, Stanford 2002, 564\u2013575."},{"issue":"3","key":"285_CR9","doi-asserted-by":"publisher","first-page":"377","DOI":"10.1002\/bimj.200390019","volume":"45","author":"F Markowetz","year":"2003","unstructured":"Markowetz F, Edler L, Vingron M: Support Vector Machines for protein fold class prediction.\n                           Biometrical Journal 2003, 45(3):377\u2013389. 10.1002\/bimj.200390019","journal-title":"Biometrical Journal"},{"key":"285_CR10","volume-title":"Bioinformatics","author":"HQ Zhu","year":"2004","unstructured":"Zhu HQ, Hu GQ, Ouyang ZQ, Wang J, She ZS: Accuracy improvement for identifying translation initiation sites in microbial genomes.\n                           Bioinformatics 2004."},{"issue":"6","key":"285_CR11","doi-asserted-by":"publisher","first-page":"1780","DOI":"10.1093\/nar\/gkg254","volume":"31","author":"FB Guo","year":"2000","unstructured":"Guo FB, Hou HY, Zhang CT: ZCURVE: a new system for recognizing protein-coding genes in bacterial and archaeal genomes.\n                           Nucleic Acides Res 2000, 31(6):1780\u20131789. 10.1093\/nar\/gkg254","journal-title":"Nucleic Acides Res"},{"issue":"4","key":"285_CR12","first-page":"441","volume":"3","author":"M Tech","year":"2003","unstructured":"Tech M, Merkl R: YACOP: Enhanced gene prediction obtained by a combination of existing methods.\n                           In Silico Biology 2003, 3(4):441\u201351.","journal-title":"In Silico Biology"},{"key":"285_CR13","volume-title":"Fuzzy logic and its applications","author":"L Zadeh","year":"1965","unstructured":"Zadeh L: Fuzzy logic and its applications. New York: Academic Press; 1965."},{"issue":"12","key":"285_CR14","doi-asserted-by":"publisher","first-page":"2637","DOI":"10.1101\/gr.1679003","volume":"13","author":"XH Zhang","year":"2003","unstructured":"Zhang XH, Heller KA, Hefter I, Leslie CS, Chasin LA: Sequence Information for the Splicing of Human Pre-mRNA Identified by Support Vector Machine Classification.\n                           Genome Res 2003, 13(12):2637\u20132650. 10.1101\/gr.1679003","journal-title":"Genome Res"},{"issue":"3","key":"285_CR15","doi-asserted-by":"publisher","first-page":"273","DOI":"10.1023\/A:1022627411411","volume":"20","author":"C Cortes","year":"1995","unstructured":"Cortes C, Vapnik V: Support-Vector Networks.\n                           Machine Learning 1995, 20(3):273\u2013297. 10.1023\/A:1022627411411","journal-title":"Machine Learning"},{"key":"285_CR16","volume-title":"In Advances in Learning Theory: Methods, Model and Applications NATO Science Series III: Computer and Systems Sciences","author":"R Rifkin","year":"2003","unstructured":"Rifkin R, Yeo G, Poggio T: Regularized Least Squares Classification. In In Advances in Learning Theory: Methods, Model and Applications NATO Science Series III: Computer and Systems Sciences. Volume 190. Amsterdam: IOS Press; 2003."},{"key":"285_CR17","first-page":"169","volume-title":"In Advances in Kernel Methods: Support Vector Machines","author":"T Joachims","year":"1998","unstructured":"Joachims T: Making large-scale support vector machine learning practical. In In Advances in Kernel Methods: Support Vector Machines. MIT Press, Cambridge, MA; 1998:169\u2013184."},{"key":"285_CR18","first-page":"911","volume-title":"In Proc 17th International Conf on Machine Learning","author":"AJ Smola","year":"2000","unstructured":"Smola AJ, Sch\u00f6lkopf B: Sparse Greedy Matrix Approximation for Machine Learning. In In Proc 17th International Conf on Machine Learning. Morgan Kaufmann, San Francisco, CA; 2000:911\u2013918."},{"key":"285_CR19","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1093\/nar\/28.1.60","volume":"28","author":"KE Rudd","year":"2000","unstructured":"Rudd KE: EcoGene: a genome sequence database for Escherichia coli K-12.\n                           Nucleic Acids Res 2000, 28: 60\u201364. [http:\/\/bmb.med.miami.edu\/EcoGene\/EcoWeb\/] 10.1093\/nar\/28.1.60","journal-title":"Nucleic Acids Res"},{"key":"285_CR20","unstructured":"Oligo Plots[http:\/\/gobics.de\/oligo_functions\/oligos.php]"},{"issue":"20","key":"285_CR21","doi-asserted-by":"publisher","first-page":"5733","DOI":"10.1128\/JB.184.20.5733-5745.2002","volume":"184","author":"J Ma","year":"2002","unstructured":"Ma J, Campbell A, Karlin S: Correlation between Shine-Dalgarno sequences and gene features such as predicted expression levels and operon structures.\n                           J Bacteriol 2002, 184(20):5733\u20135745. 10.1128\/JB.184.20.5733-5745.2002","journal-title":"J Bacteriol"},{"key":"285_CR22","doi-asserted-by":"publisher","first-page":"215","DOI":"10.1006\/jmbi.2001.5040","volume":"313","author":"RK Shultzaberger","year":"2001","unstructured":"Shultzaberger RK, Buchheimer RE, Rudd KE, Schneider TD: Anatomy of Escherichia coli ribosome binding sites.\n                           J Mol Biol 2001, 313: 215\u2013228. 10.1006\/jmbi.2001.5040","journal-title":"J Mol Biol"},{"issue":"1\u20132","key":"285_CR23","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/S0378-1119(02)00501-2","volume":"288","author":"CM Stenstrom","year":"2002","unstructured":"Stenstrom CM, Isaksson LA: Influences on translation initiation and early elongation by the messenger RNA region flanking the initiation codon at the 3' side.\n                           Gene 2002, 288(1\u20132):1\u20138. 10.1016\/S0378-1119(02)00501-2","journal-title":"Gene"},{"issue":"1\u20132","key":"285_CR24","doi-asserted-by":"publisher","first-page":"273","DOI":"10.1016\/S0378-1119(00)00550-3","volume":"263","author":"CM Stenstrom","year":"2001","unstructured":"Stenstrom CM, Jin H, Major LL, Tate WP, Isaksson LA: Codon bias at the 3'-side of the initiation codon is correlated with translation initiation efficiency in Escherichia coli.\n                           Gene 2001, 263(1\u20132):273\u2013284. 10.1016\/S0378-1119(00)00550-3","journal-title":"Gene"},{"issue":"6","key":"285_CR25","doi-asserted-by":"publisher","first-page":"851","DOI":"10.1093\/oxfordjournals.jbchem.a002929","volume":"129","author":"T Sato","year":"2001","unstructured":"Sato T, Terabe M, Watanabe H, Gojobori T, Hori-Takemoto C, Miura K: Codon and base biases after the initiation codon of the open reading frames in the Escherichia coli genome and their influence on the translation efficiency.\n                           J Biochem 2001, 129(6):851\u201360.","journal-title":"J Biochem"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-5-169.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/1471-2105-5-169\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-5-169.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,7]],"date-time":"2024-10-07T12:20:04Z","timestamp":1728303604000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-5-169"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2004,10,28]]},"references-count":25,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2004,12]]}},"alternative-id":["285"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-5-169","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2004,10,28]]},"assertion":[{"value":"10 August 2004","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 October 2004","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 October 2004","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"169"}}