{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,9,1]],"date-time":"2026-09-01T18:36:06Z","timestamp":1788287766252,"version":"build-2803163510"},"reference-count":33,"publisher":"Oxford University Press (OUP)","issue":"5","license":[{"start":{"date-parts":[[2016,10,2]],"date-time":"2016-10-02T00:00:00Z","timestamp":1475366400000},"content-version":"vor","delay-in-days":3174,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/2.0\/uk\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2008,3,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Motivation: Computational annotation of protein coding genes in genomic DNA is a widely used and essential tool for analyzing newly sequenced genomes. However, current methods suffer from inaccuracy and do poorly with certain types of genes. Including additional sources of evidence of the existence and structure of genes can improve the quality of gene predictions. For many eukaryotic genomes, expressed sequence tags (ESTs) are available as evidence for genes. Related genomes that have been sequenced, annotated, and aligned to the target genome provide evidence of existence and structure of genes.<\/jats:p>\n                  <jats:p>Results: We incorporate several different evidence sources into the gene finder AUGUSTUS. The sources of evidence are gene and transcript annotations from related species syntenically mapped to the target genome using TransMap, evolutionary conservation of DNA, mRNA and ESTs of the target species, and retroposed genes. The predictions include alternative splice variants where evidence supports it. Using only ESTs we were able to correctly predict at least one splice form exactly correct in 57% of human genes. Also using evidence from other species and human mRNAs, this number rises to 77%. Syntenic mapping is well-suited to annotate genomes closely related to genomes that are already annotated or for which extensive transcript evidence is available. Native cDNA evidence is most helpful when the alignments are used as compound information rather than independent positionwise information.<\/jats:p>\n                  <jats:p>Availability: AUGUSTUS is open source and available at http:\/\/augustus.gobics.de. The gene predictions for human can be browsed and downloaded at the UCSC Genome Browser (http:\/\/genome.ucsc.edu)<\/jats:p>\n                  <jats:p>Contact: \u00a0mstanke@gwdg.de<\/jats:p>\n                  <jats:p>Supplementary information: Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btn013","type":"journal-article","created":{"date-parts":[[2008,1,24]],"date-time":"2008-01-24T20:24:39Z","timestamp":1201206279000},"page":"637-644","source":"Crossref","is-referenced-by-count":2507,"title":["Using native and syntenically mapped cDNA alignments to improve\n                    <i>de novo<\/i>\n                    gene finding"],"prefix":"10.1093","volume":"24","author":[{"given":"Mario","family":"Stanke","sequence":"first","affiliation":[{"name":"Center for Biomolecular Science and Engineering, University of California Santa Cruz (UCSC), Santa Cruz, CA 95064, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mark","family":"Diekhans","sequence":"additional","affiliation":[{"name":"Center for Biomolecular Science and Engineering, University of California Santa Cruz (UCSC), Santa Cruz, CA 95064, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Robert","family":"Baertsch","sequence":"additional","affiliation":[{"name":"Center for Biomolecular Science and Engineering, University of California Santa Cruz (UCSC), Santa Cruz, CA 95064, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"David","family":"Haussler","sequence":"additional","affiliation":[{"name":"Center for Biomolecular Science and Engineering, University of California Santa Cruz (UCSC), Santa Cruz, CA 95064, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2008,1,24]]},"reference":[{"key":"2023020210104650200_B1","first-page":"14","article-title":"A phylogenetic generalized hidden Markov model for predicting alternatively spliced exons","volume":"1","author":"Allen","year":"2006","journal-title":"AMB"},{"key":"2023020210104650200_B2","article-title":"Evidence combination in hidden Markov models for gene prediction","volume-title":"PhD Thesis.","author":"Brejov\u00e1","year":"2005"},{"issue":"Suppl. 1","key":"2023020210104650200_B3","doi-asserted-by":"crossref","first-page":"i57","DOI":"10.1093\/bioinformatics\/bti1040","article-title":"ExonHunter: a comprehensive approach to gene finding","volume":"21","author":"Brejov\u00e1","year":"2005","journal-title":"Bioinformatics"},{"issue":"Suppl. 2","key":"2023020210104650200_B4","doi-asserted-by":"crossref","first-page":"ii36","DOI":"10.1093\/bioinformatics\/btg1057","article-title":"HMM sampling and applications to gene finding and alternative splicing","volume":"19","author":"Cawley","year":"2003","journal-title":"Bioinformatics"},{"key":"2023020210104650200_B5","doi-asserted-by":"crossref","first-page":"942","DOI":"10.1101\/gr.1858004","article-title":"The Ensembl Automatic Gene Annotation System","volume":"14","author":"Curwen","year":"2004","journal-title":"Genome Res"},{"issue":"Suppl. 1","key":"2023020210104650200_B6","first-page":"S7.1","article-title":"Exogean: a framework for annotating protein-coding genes in eukaryotic genomic DNA","volume":"7","author":"Djebali","year":"2006","journal-title":"BMC Genome Biol"},{"key":"2023020210104650200_B7","doi-asserted-by":"crossref","first-page":"54","DOI":"10.1101\/gr.2889405","article-title":"Gene and alternative splicing annotation with AIR","volume":"15","author":"Florea","year":"2005","journal-title":"Genome Res"},{"key":"2023020210104650200_B8","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1186\/1471-2105-6-25","article-title":"Integrating alternative splicing detection into gene prediction","volume":"6","author":"Foissac","year":"2005","journal-title":"BMC Bioinformatics"},{"key":"2023020210104650200_B9","first-page":"374","article-title":"Using multiple alignments to improve gene prediction","volume-title":"In Proceedings of RECOMB 2005.","author":"Gross","year":"2005"},{"issue":"Suppl. 1","key":"2023020210104650200_B10","first-page":"S2.1","article-title":"EGASP: the human ENCODE Genome Annotation Assessment Project","volume":"7","author":"Guig\u00f3","year":"2006","journal-title":"BMC Genome Biol"},{"key":"2023020210104650200_B11","doi-asserted-by":"crossref","first-page":"5654","DOI":"10.1093\/nar\/gkg770","article-title":"Improving the Arabidopsis genome annotation using maximal transcipt alignment assemblies","volume":"31","author":"Haas","year":"2003","journal-title":"Nucleic Acids Res"},{"issue":"Suppl. 1","key":"2023020210104650200_B12","first-page":"S4.1","article-title":"GENCODE: producing a reference annotation for ENCODE","volume":"7","author":"Harrow","year":"2006","journal-title":"Genome Biol"},{"key":"2023020210104650200_B13","first-page":"656","article-title":"BLAT\u2013The BLAST-Like Alignment Tool","volume":"12","author":"Kent","year":"2002","journal-title":"Genome Res"},{"key":"2023020210104650200_B14","doi-asserted-by":"crossref","first-page":"11484","DOI":"10.1073\/pnas.1932072100","article-title":"Evolution's cauldron: Duplication, deletion, and rearrangement in the mouse and human genomes","volume":"100","author":"Kent","year":"2003","journal-title":"PNAS"},{"key":"2023020210104650200_B15","doi-asserted-by":"crossref","first-page":"S1","DOI":"10.1186\/1471-2105-5-59","article-title":"Gene finding in novel genomes","volume":"5","author":"Korf","year":"2004","journal-title":"BMC Bioinformatics"},{"key":"2023020210104650200_B16","first-page":"179","article-title":"Two methods for improving performance of an HMM and their application for gene finding","volume-title":"AAAI","author":"Krogh","year":"1997"},{"key":"2023020210104650200_B17","doi-asserted-by":"crossref","first-page":"D668","DOI":"10.1093\/nar\/gkl928","article-title":"The UCSC genome browser database: update 2007","volume":"35","author":"Kuhn","year":"2006","journal-title":"Nucl. Acids Res"},{"key":"2023020210104650200_B18","doi-asserted-by":"crossref","first-page":"6494","DOI":"10.1093\/nar\/gki937","article-title":"Gene identification in novel eukaryotic genomes by self-training algorithm","volume":"33","author":"Lomsadze","year":"2005","journal-title":"Nucl. Acids Res"},{"key":"2023020210104650200_B19","doi-asserted-by":"crossref","first-page":"376","DOI":"10.1038\/nature03959","article-title":"Genome sequencing in microfabricated high-density picolitre reactors","volume":"437","author":"Margulies","year":"2005","journal-title":"Nature"},{"key":"2023020210104650200_B20","doi-asserted-by":"crossref","first-page":"1309","DOI":"10.1093\/bioinformatics\/18.10.1309","article-title":"Comparative ab initio prediction of gene structures using pair HMMs","volume":"18","author":"Meyer","year":"2002","journal-title":"Bioinformatics"},{"key":"2023020210104650200_B21","doi-asserted-by":"crossref","first-page":"776","DOI":"10.1093\/nar\/gkh211","article-title":"Gene structure conservation aids similarity based gene prediction","volume":"32","author":"Meyer","year":"2004","journal-title":"Nucl. Acids Res"},{"issue":"Suppl. 1","key":"2023020210104650200_B22","doi-asserted-by":"crossref","first-page":"D61","DOI":"10.1093\/nar\/gkl842","article-title":"NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins","volume":"35","author":"Pruitt","year":"2007","journal-title":"Nucl. Acids Res"},{"key":"2023020210104650200_B23","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1101\/gr.809403","article-title":"Human-Mouse Alignments with BLASTZ","volume":"13","author":"Schwartz","year":"2003","journal-title":"Genome Res"},{"key":"2023020210104650200_B24","first-page":"177","article-title":"Computational identification of evolutionarily conserved exons","author":"Siepel","year":"2004"},{"key":"2023020210104650200_B25","doi-asserted-by":"crossref","first-page":"1034","DOI":"10.1101\/gr.3715005","article-title":"Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes","volume":"15","author":"Siepel","year":"2005","journal-title":"Genome Res"},{"key":"2023020210104650200_B26","doi-asserted-by":"crossref","first-page":"1763","DOI":"10.1101\/gr.7128207","article-title":"Targeted discovery of novel human exons by comparative genomics","volume":"17","author":"Siepel","year":"2007","journal-title":"Genome Res"},{"issue":"Suppl. 2","key":"2023020210104650200_B27","doi-asserted-by":"crossref","first-page":"ii215","DOI":"10.1093\/bioinformatics\/btg1080","article-title":"Gene prediction with a hidden markov model and new intron submodel","volume":"19","author":"Stanke","year":"2003","journal-title":"Bioinformatics"},{"key":"2023020210104650200_B28","doi-asserted-by":"crossref","first-page":"W435","DOI":"10.1093\/nar\/gkl200","article-title":"AUGUSTUS: ab initio prediction of alternative transcripts","volume":"34","author":"Stanke","year":"2006","journal-title":"Nucleic Acids Res"},{"key":"2023020210104650200_B29","doi-asserted-by":"crossref","first-page":"62","DOI":"10.1186\/1471-2105-7-62","article-title":"Gene prediction in eukaryotes with a generalized hidden Markov model that uses hints from external sources","volume":"7","author":"Stanke","year":"2006","journal-title":"BMC Bioinformatics"},{"issue":"Suppl. 1","key":"2023020210104650200_B30","doi-asserted-by":"crossref","first-page":"S12","DOI":"10.1186\/gb-2006-7-s1-s12","article-title":"AceView: a comprehensive cDNA supported gene and transcripts annotation","volume":"7","author":"Thierry-Mieg","year":"2006","journal-title":"BMC Genome Biol"},{"key":"2023020210104650200_B31","doi-asserted-by":"crossref","first-page":"678","DOI":"10.1101\/gr.4766206","article-title":"Iterative gene prediction and pseudogene removal improves genome annotation","volume":"16","author":"van Baren","year":"2006","journal-title":"Genome Res"},{"key":"2023020210104650200_B32","doi-asserted-by":"crossref","first-page":"327","DOI":"10.1186\/1471-2105-7-327","article-title":"Using ESTs to improve the accuracy of de novo gene prediction","volume":"7","author":"Wei","year":"2006","journal-title":"BMC Bioinformatics"},{"key":"2023020210104650200_B33","doi-asserted-by":"publisher","first-page":"e247","DOI":"10.1371\/journal.pcbi.0030247","article-title":"Comparative genomics search for losses of long-established genes on the human lineage","volume":"3","author":"Zhu","year":"2007","journal-title":"PLoS Computational Biol"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/5\/637\/49050288\/bioinformatics_24_5_637.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/5\/637\/49050288\/bioinformatics_24_5_637.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,2]],"date-time":"2023-02-02T06:47:31Z","timestamp":1675320451000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/24\/5\/637\/202844"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,1,24]]},"references-count":33,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2008,3,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btn013","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2008,3,1]]},"published":{"date-parts":[[2008,1,24]]}}}