{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,20]],"date-time":"2026-01-20T19:43:21Z","timestamp":1768938201498,"version":"3.49.0"},"reference-count":21,"publisher":"Oxford University Press (OUP)","issue":"13","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2007,7,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: The correct identification of translation initiation sites (TIS) remains a challenging problem for computational methods that automatically try to solve this problem. Furthermore, the lion's share of these computational techniques focuses on the identification of TIS in transcript data. However, in the gene prediction context the identification of TIS occurs on the genomic level, which makes things even harder because at the genome level many more pseudo-TIS occur, resulting in models that achieve a higher number of false positive predictions.<\/jats:p>\n               <jats:p>Results: In this article, we evaluate the performance of several \u2018simple\u2019 TIS recognition methods at the genomic level, and compare them to state-of-the-art models for TIS prediction in transcript data. We conclude that the simple methods largely outperform the complex ones at the genomic scale, and we propose a new model for TIS recognition at the genome level that combines the strengths of these simple models. The new model obtains a false positive rate of 0.125 at a sensitivity of 0.80 on a well annotated human chromosome (chromosome 21). Detailed analyses show that the model is useful, both on its own and in a simple gene prediction setting.<\/jats:p>\n               <jats:p>Availability: Datafiles and a web interface for the StartScan program are available at http:\/\/bioinformatics.psb.ugent.be\/supplementary_data\/<\/jats:p>\n               <jats:p>Contact: yvan.saeys@psb.ugent.be<\/jats:p>","DOI":"10.1093\/bioinformatics\/btm177","type":"journal-article","created":{"date-parts":[[2007,7,23]],"date-time":"2007-07-23T16:13:46Z","timestamp":1185207226000},"page":"i418-i423","source":"Crossref","is-referenced-by-count":47,"title":["Translation initiation site prediction on a genomic scale: beauty in simplicity"],"prefix":"10.1093","volume":"23","author":[{"given":"Yvan","family":"Saeys","sequence":"first","affiliation":[{"name":"1 Department of Plant Systems Biology, VIB, Technologiepark 927, B-9052 Ghent, Belgium, 2Department of Molecular Genetics, Ghent University, Ghent, Belgium and 3Pronota, Technologiepark - Zwijnaarde 927, B-9052 Ghent, Belgium"},{"name":"1 Department of Plant Systems Biology, VIB, Technologiepark 927, B-9052 Ghent, Belgium, 2Department of Molecular Genetics, Ghent University, Ghent, Belgium and 3Pronota, Technologiepark - Zwijnaarde 927, B-9052 Ghent, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Thomas","family":"Abeel","sequence":"additional","affiliation":[{"name":"1 Department of Plant Systems Biology, VIB, Technologiepark 927, B-9052 Ghent, Belgium, 2Department of Molecular Genetics, Ghent University, Ghent, Belgium and 3Pronota, Technologiepark - Zwijnaarde 927, B-9052 Ghent, Belgium"},{"name":"1 Department of Plant Systems Biology, VIB, Technologiepark 927, B-9052 Ghent, Belgium, 2Department of Molecular Genetics, Ghent University, Ghent, Belgium and 3Pronota, Technologiepark - Zwijnaarde 927, B-9052 Ghent, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sven","family":"Degroeve","sequence":"additional","affiliation":[{"name":"1 Department of Plant Systems Biology, VIB, Technologiepark 927, B-9052 Ghent, Belgium, 2Department of Molecular Genetics, Ghent University, Ghent, Belgium and 3Pronota, Technologiepark - Zwijnaarde 927, B-9052 Ghent, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yves","family":"Van de Peer","sequence":"additional","affiliation":[{"name":"1 Department of Plant Systems Biology, VIB, Technologiepark 927, B-9052 Ghent, Belgium, 2Department of Molecular Genetics, Ghent University, Ghent, Belgium and 3Pronota, Technologiepark - Zwijnaarde 927, B-9052 Ghent, Belgium"},{"name":"1 Department of Plant Systems Biology, VIB, Technologiepark 927, B-9052 Ghent, Belgium, 2Department of Molecular Genetics, Ghent University, Ghent, Belgium and 3Pronota, Technologiepark - Zwijnaarde 927, B-9052 Ghent, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2007,7,1]]},"reference":[{"key":"2023062708513645900_B1","doi-asserted-by":"crossref","first-page":"123","DOI":"10.1016\/0097-8485(93)85004-V","article-title":"GeneMark: parallel gene recognition for both DNA strands","volume":"17","author":"Borodovsky","year":"1993","journal-title":"Comput. and Chem"},{"key":"2023062708513645900_B2","doi-asserted-by":"crossref","first-page":"4636","DOI":"10.1093\/nar\/27.23.4636","article-title":"Improved microbial gene identification with GLIMMER","volume":"27","author":"Delcher","year":"1999","journal-title":"Nucleic Acids Res"},{"key":"2023062708513645900_B3","doi-asserted-by":"crossref","first-page":"6441","DOI":"10.1093\/nar\/20.24.6441","article-title":"Assessment of protein coding measures","volume":"20","author":"Fickett","year":"1992","journal-title":"Nucleic Acids Res"},{"key":"2023062708513645900_B4","doi-asserted-by":"crossref","first-page":"343","DOI":"10.1093\/bioinformatics\/18.2.343","article-title":"Translation initiation start prediction in human cDNAs with high accuracy","volume":"18","author":"Hatzigeorgiou","year":"2002","journal-title":"Bioinformatics"},{"key":"2023062708513645900_B5","doi-asserted-by":"crossref","first-page":"8125","DOI":"10.1093\/nar\/15.20.8125","article-title":"An analysis of 5'-noncoding sequences from 699 vertebrate messenger RNAs","volume":"15","author":"Kozak","year":"1987","journal-title":"Nucleic Acids Res"},{"key":"2023062708513645900_B6","doi-asserted-by":"crossref","first-page":"229","DOI":"10.1083\/jcb.108.2.229","article-title":"The scanning model for translation: an update","volume":"108","author":"Kozak","year":"1989","journal-title":"J. Cell Biol"},{"key":"2023062708513645900_B7","doi-asserted-by":"crossref","first-page":"187","DOI":"10.1016\/S0378-1119(99)00210-3","article-title":"Initiation of translation in prokaryotes and eukaryotes","volume":"234","author":"Kozak","year":"1999","journal-title":"Gene"},{"key":"2023062708513645900_B8","first-page":"262","article-title":"A class of edit kernels for SVMs to predict translation initiation sites in eukaryotic mRNAs","author":"Li","year":"2004"},{"key":"2023062708513645900_B9","doi-asserted-by":"crossref","first-page":"1152","DOI":"10.1109\/TKDE.2005.133","article-title":"Translation initiation sites prediction with mixture Gaussian models in human cDNA sequences","volume":"8","author":"Li","year":"2005","journal-title":"IEEE Trans. Knowl. Data Eng"},{"key":"2023062708513645900_B10","doi-asserted-by":"crossref","DOI":"10.1142\/9789812562340_0004","article-title":"Techniques for recognition of translation initiation sites","volume-title":"The Practical Bioinformaticion","author":"Li","year":"2004"},{"key":"2023062708513645900_B11","first-page":"255","article-title":"Using amino acid patterns to accurately predict translation initiation sites","volume":"4","author":"Liu","year":"2004","journal-title":"In Silico Biol"},{"key":"2023062708513645900_B12","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1093\/bioinformatics\/16.11.960","article-title":"Prediction whether a human cDNA sequence contains initiation codon by combining statistical information and similarity with protein sequences","volume":"16","author":"Nishikawa","year":"2000","journal-title":"Bioinformatics"},{"key":"2023062708513645900_B13","first-page":"226","article-title":"Neural network prediction of translation initiation sites in eukaryotes: perspectives for EST and genome analysis","author":"Pedersen","year":"1997"},{"key":"2023062708513645900_B14","unstructured":"Saeys\n              Y\n            \n          \n          Feature selection for classification of nucleic acid sequences\n          PhD thesis\n          2004\n          Belgium\n          Ghent University"},{"key":"2023062708513645900_B15","doi-asserted-by":"crossref","first-page":"384","DOI":"10.1093\/bioinformatics\/14.5.384","article-title":"Assessing protein coding region integrity in cDNA sequence projects","volume":"14","author":"Salamov","year":"1998","journal-title":"Bioinformatics"},{"key":"2023062708513645900_B16","first-page":"365","article-title":"A method for identifying splice sites and translational start sites in eukaryotic mRNA","volume":"13","author":"Salzberg","year":"1997","journal-title":"Comput. Appl. Biosci"},{"key":"2023062708513645900_B17","doi-asserted-by":"crossref","first-page":"24","DOI":"10.1006\/geno.1999.5854","article-title":"Interpolated Markov models for eukaryotic gene finding","volume":"59","author":"Salzberg","year":"1999","journal-title":"Genomics"},{"key":"2023062708513645900_B18","first-page":"263","article-title":"Prediction of probable genes by Fourier analysis of genomic sequences","volume":"13","author":"Tiwari","year":"1997","journal-title":"Comput. Appl. Biosci"},{"key":"2023062708513645900_B19","doi-asserted-by":"crossref","first-page":"699","DOI":"10.1089\/106652703322539042","article-title":"Recognition of translation initiation sites of eukaryotic genes based on an EM algorithm","volume":"10","author":"Wang","year":"2003","journal-title":"J. Comput. Biol"},{"key":"2023062708513645900_B20","first-page":"192","article-title":"Using feature generation and feature selection for accurate prediction of translation initiation sites","volume":"13","author":"Zeng","year":"2002","journal-title":"Genome Inform"},{"key":"2023062708513645900_B21","doi-asserted-by":"crossref","first-page":"799","DOI":"10.1093\/bioinformatics\/16.9.799","article-title":"Engineering support vector machine kernels that recognize translation initiation sites","volume":"16","author":"Zien","year":"2000","journal-title":"Bioinformatics"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/13\/i418\/50718225\/bioinformatics_23_13_i418.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/13\/i418\/50718225\/bioinformatics_23_13_i418.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,27]],"date-time":"2023-06-27T08:54:43Z","timestamp":1687856083000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/23\/13\/i418\/227417"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,7,1]]},"references-count":21,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2007,7,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btm177","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2007,7]]},"published":{"date-parts":[[2007,7,1]]}}}