{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T03:08:50Z","timestamp":1780542530059,"version":"3.54.1"},"reference-count":39,"publisher":"Oxford University Press (OUP)","issue":"4","license":[{"start":{"date-parts":[[2019,9,18]],"date-time":"2019-09-18T00:00:00Z","timestamp":1568764800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"funder":[{"name":"NIAID at NIH","award":["1DP2AI145058-01"],"award-info":[{"award-number":["1DP2AI145058-01"]}]},{"name":"UCSD Graduate Training Program in Cellular and Molecular Pharmacology","award":["T32 GM007752"],"award-info":[{"award-number":["T32 GM007752"]}]},{"DOI":"10.13039\/100000069","name":"NIAMS","doi-asserted-by":"publisher","award":["T32 AR064194"],"award-info":[{"award-number":["T32 AR064194"]}],"id":[{"id":"10.13039\/100000069","id-type":"DOI","asserted-by":"publisher"}]},{"name":"UCSD Microbial Sciences Initiative Graduate Research Fellowship"},{"name":"UCSD Graduate Training Program in Cellular and Molecular Pharmacology"},{"DOI":"10.13039\/100000057","name":"NIGMS","doi-asserted-by":"publisher","award":["T32 GM007752"],"award-info":[{"award-number":["T32 GM007752"]}],"id":[{"id":"10.13039\/100000057","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,2,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>A core task of genomics is to identify the boundaries of protein coding genes, which may cover over 90% of a prokaryote's genome. Several programs are available for gene finding, yet it is currently unclear how well these programs perform and whether any offers superior accuracy. This is in part because there is no universal benchmark for gene finding and, therefore, most developers select their own benchmarking strategy.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>Here, we introduce AssessORF, a new approach for benchmarking prokaryotic gene predictions based on evidence from proteomics data and the evolutionary conservation of start and stop codons. We applied AssessORF to compare gene predictions offered by GenBank, GeneMarkS-2, Glimmer and Prodigal on genomes spanning the prokaryotic tree of life. Gene predictions were 88\u201395% in agreement with the available evidence, with Glimmer performing the worst but no clear winner. All programs were biased towards selecting start codons that were upstream of the actual start. Given these findings, there remains considerable room for improvement, especially in the detection of correct start sites.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation<\/jats:title>\n                  <jats:p>AssessORF is available as an R package via the Bioconductor package repository.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information<\/jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btz714","type":"journal-article","created":{"date-parts":[[2019,9,14]],"date-time":"2019-09-14T11:42:04Z","timestamp":1568461324000},"page":"1022-1029","source":"Crossref","is-referenced-by-count":17,"title":["AssessORF: combining evolutionary conservation and proteomics to assess prokaryotic gene predictions"],"prefix":"10.1093","volume":"36","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6750-5260","authenticated-orcid":false,"given":"Deepank R","family":"Korandla","sequence":"first","affiliation":[{"name":"Department of Biological Sciences , USA"},{"name":"Computational Biology Department, Carnegie Mellon University , Pittsburgh, PA 15213, USA"},{"name":"Department of Biomedical Informatics, University of Pittsburgh School of Medicine , Pittsburgh, PA 15219, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jacob M","family":"Wozniak","sequence":"additional","affiliation":[{"name":"Department of Pharmacology, University of California San Diego , La Jolla, CA 92093, USA"},{"name":"Skaggs School of Pharmacy and Pharmaceutical Sciences, University of California San Diego , La Jolla, CA 92093, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anaamika","family":"Campeau","sequence":"additional","affiliation":[{"name":"Department of Pharmacology, University of California San Diego , La Jolla, CA 92093, USA"},{"name":"Skaggs School of Pharmacy and Pharmaceutical Sciences, University of California San Diego , La Jolla, CA 92093, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"David J","family":"Gonzalez","sequence":"additional","affiliation":[{"name":"Department of Pharmacology, University of California San Diego , La Jolla, CA 92093, USA"},{"name":"Skaggs School of Pharmacy and Pharmaceutical Sciences, University of California San Diego , La Jolla, CA 92093, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1457-4019","authenticated-orcid":false,"given":"Erik S","family":"Wright","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, University of Pittsburgh School of Medicine , Pittsburgh, PA 15219, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2019,9,18]]},"reference":[{"key":"2023013110162188500_btz714-B1","doi-asserted-by":"crossref","first-page":"503","DOI":"10.1016\/j.cbpa.2009.07.026","article-title":"Methods for the proteomic identification of protease substrates","volume":"13","author":"Agard","year":"2009","journal-title":"Curr. Opin. Chem. Biol"},{"key":"2023013110162188500_btz714-B2","doi-asserted-by":"crossref","first-page":"D37","DOI":"10.1093\/nar\/gkw1070","article-title":"GenBank","volume":"45","author":"Benson","year":"2017","journal-title":"Nucleic Acids Res"},{"key":"2023013110162188500_btz714-B3","doi-asserted-by":"crossref","first-page":"2607","DOI":"10.1093\/nar\/29.12.2607","article-title":"GeneMarkS: a self-training method for prediction of gene starts in microbial genomes. Implications for finding sequence motifs in regulatory regions","volume":"29","author":"Besemer","year":"2001","journal-title":"Nucleic Acids Res"},{"key":"2023013110162188500_btz714-B4","doi-asserted-by":"crossref","first-page":"35.","DOI":"10.1186\/1471-2105-12-35","article-title":"VennDiagram: a package for the generation of highly-customizable Venn and Euler diagrams in R","volume":"12","author":"Chen","year":"2011","journal-title":"BMC Bioinformatics"},{"key":"2023013110162188500_btz714-B5","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1016\/j.tube.2012.11.012","article-title":"Reannotation of translational start sites in the genome of Mycobacterium tuberculosis","volume":"93","author":"DeJesus","year":"2013","journal-title":"Tuberculosis (Edinb)"},{"key":"2023013110162188500_btz714-B6","doi-asserted-by":"crossref","first-page":"673","DOI":"10.1093\/bioinformatics\/btm009","article-title":"Identifying bacterial genes and endosymbiont DNA with Glimmer","volume":"23","author":"Delcher","year":"2007","journal-title":"Bioinformatics"},{"key":"2023013110162188500_btz714-B7","doi-asserted-by":"crossref","first-page":"4636","DOI":"10.1093\/nar\/27.23.4636","article-title":"Improved microbial gene identification with GLIMMER","volume":"27","author":"Delcher","year":"1999","journal-title":"Nucleic Acids Res"},{"key":"2023013110162188500_btz714-B8","doi-asserted-by":"crossref","first-page":"125.","DOI":"10.1186\/1471-2164-12-125","article-title":"Consistency of gene starts among Burkholderia genomes","volume":"12","author":"Dunbar","year":"2011","journal-title":"BMC Genomics"},{"key":"2023013110162188500_btz714-B9","doi-asserted-by":"crossref","first-page":"207","DOI":"10.1038\/nmeth1019","article-title":"Target-decoy search strategy for increased confidence in large-scale protein identifications by mass spectrometry","volume":"4","author":"Elias","year":"2007","journal-title":"Nat. Methods"},{"key":"2023013110162188500_btz714-B10","doi-asserted-by":"crossref","first-page":"667","DOI":"10.1038\/nmeth785","article-title":"Comparative evaluation of mass spectrometry platforms used in large-scale proteomics investigations","volume":"2","author":"Elias","year":"2005","journal-title":"Nat. Methods"},{"key":"2023013110162188500_btz714-B11","doi-asserted-by":"crossref","first-page":"976","DOI":"10.1016\/1044-0305(94)80016-2","article-title":"An approach to correlate tandem mass spectral data of peptides with amino acid sequences in a protein database","volume":"5","author":"Eng","year":"1994","journal-title":"J. Am. Soc. Mass Spectrom"},{"key":"2023013110162188500_btz714-B12","first-page":"76","article-title":"Ribosome signatures aid bacterial translation initiation site identification","volume":"15","author":"Giess","year":"2017","journal-title":"BMC Bioinformatics"},{"key":"2023013110162188500_btz714-B13","doi-asserted-by":"crossref","first-page":"1455","DOI":"10.1007\/s00018-004-3466-8","article-title":"Protein N-terminal methionine excision","volume":"61","author":"Giglione","year":"2004","journal-title":"Cell Mol. Life Sci"},{"key":"2023013110162188500_btz714-B14","doi-asserted-by":"crossref","first-page":"3615","DOI":"10.1093\/nar\/gkx070","article-title":"Measurements of translation initiation from all 64 codons in E. coli","volume":"45","author":"Hecht","year":"2017","journal-title":"Nucleic Acids Res"},{"key":"2023013110162188500_btz714-B15","doi-asserted-by":"crossref","first-page":"115","DOI":"10.1038\/nmeth.3252","article-title":"Orchestrating high-throughput genomic analysis with Bioconductor","volume":"12","author":"Huber","year":"2015","journal-title":"Nat. Methods"},{"key":"2023013110162188500_btz714-B16","doi-asserted-by":"crossref","first-page":"e0184119","DOI":"10.1371\/journal.pone.0184119","article-title":"Discovery of numerous novel small genes in the intergenic regions of the Escherichia coli O157:H7 Sakai genome","volume":"12","author":"H\u00fccker","year":"2017","journal-title":"PLoS One"},{"key":"2023013110162188500_btz714-B17","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1186\/1471-2105-11-119","article-title":"Prodigal: prokaryotic gene recognition and translation initiation site identification","volume":"11","author":"Hyatt","year":"2010","journal-title":"BMC Bioinformatics"},{"key":"2023013110162188500_btz714-B18","doi-asserted-by":"crossref","first-page":"e58387.","DOI":"10.1371\/journal.pone.0058387","article-title":"ORFcor: identifying and accommodating ORF prediction inconsistencies for phylogenetic analysis","volume":"8","author":"Klassen","year":"2013","journal-title":"PLoS One"},{"key":"2023013110162188500_btz714-B19","doi-asserted-by":"crossref","first-page":"1079","DOI":"10.1101\/gr.230615.117","article-title":"Modeling leaderless transcription and atypical genes results in more accurate gene prediction in prokaryotes","volume":"28","author":"Lomsadze","year":"2018","journal-title":"Genome Res"},{"key":"2023013110162188500_btz714-B20","doi-asserted-by":"crossref","first-page":"1107","DOI":"10.1093\/nar\/26.4.1107","article-title":"GeneMark.hmm: new solutions for gene finding","volume":"26","author":"Lukashin","year":"1998","journal-title":"Nucleic Acids Res"},{"key":"2023013110162188500_btz714-B21","doi-asserted-by":"crossref","first-page":"551.","DOI":"10.1186\/s12859-018-2550-2","article-title":"Computational discovery and annotation of conserved small open reading frames in fungal genomes","volume":"19","author":"Mat-Sharani","year":"2019","journal-title":"BMC Bioinformatics"},{"key":"2023013110162188500_btz714-B22","doi-asserted-by":"crossref","first-page":"1780","DOI":"10.1074\/mcp.M113.027540","article-title":"Deep proteome coverage based on ribosome profiling aids mass spectrometry-based protein and peptide discovery and provides evidence of alternative translation products and near-cognate translation initiation events","volume":"12","author":"Menschaert","year":"2013","journal-title":"Mol. Cell Proteomics"},{"key":"2023013110162188500_btz714-B23","doi-asserted-by":"crossref","first-page":"481","DOI":"10.1016\/j.molcel.2019.02.017","article-title":"Retapamulin-Assisted Ribosome Profiling Reveals the Alternative Bacterial Proteome","volume":"74","author":"Meydan","year":"2019","journal-title":"Mol. Cell"},{"key":"2023013110162188500_btz714-B24","doi-asserted-by":"crossref","first-page":"e8290","DOI":"10.15252\/msb.20188290","article-title":"Unraveling the hidden universe of small proteins in bacterial genomes","volume":"15","author":"Miravet-Verde","year":"2019","journal-title":"Mol. Syst. Biol"},{"key":"2023013110162188500_btz714-B25","doi-asserted-by":"crossref","first-page":"3922","DOI":"10.1093\/nar\/gkx124","article-title":"Comparative genomic analysis of translation initiation mechanisms for genes lacking the Shine-Dalgarno sequence in prokaryotes","volume":"45","author":"Nakagawa","year":"2017","journal-title":"Nucleic Acids Res"},{"key":"2023013110162188500_btz714-B26","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1021\/pr025556v","article-title":"Evaluation of multidimensional chromatography coupled with tandem mass spectrometry (LC\/LC-MS\/MS) for large-scale protein analysis: the yeast proteome","volume":"2","author":"Peng","year":"2003","journal-title":"J. Proteome Res"},{"key":"2023013110162188500_btz714-B27","year":"2019"},{"key":"2023013110162188500_btz714-B28","doi-asserted-by":"crossref","first-page":"544","DOI":"10.1093\/nar\/26.2.544","article-title":"Microbial gene identification using interpolated Markov models","volume":"26","author":"Salzberg","year":"1998","journal-title":"Nucleic Acids Res"},{"key":"2023013110162188500_btz714-B29","doi-asserted-by":"crossref","first-page":"753","DOI":"10.1146\/annurev-biochem-070611-102400","article-title":"Small proteins can no longer be ignored","volume":"83","author":"Storz","year":"2014","journal-title":"Annu. Rev. Biochem"},{"key":"2023013110162188500_btz714-B30","doi-asserted-by":"crossref","first-page":"1892","DOI":"10.1128\/JB.00202-16","article-title":"Alternative translation initiation of a haloarchaeal serine protease transcript containing two in-frame start codons","volume":"198","author":"Tang","year":"2016","journal-title":"J. Bacteriol"},{"key":"2023013110162188500_btz714-B31","doi-asserted-by":"crossref","first-page":"6614","DOI":"10.1093\/nar\/gkw569","article-title":"NCBI prokaryotic genome annotation pipeline","volume":"44","author":"Tatusova","year":"2016","journal-title":"Nucleic Acids Res"},{"key":"2023013110162188500_btz714-B32","doi-asserted-by":"crossref","first-page":"950","DOI":"10.1038\/nature08080","article-title":"The Listeria transcriptional landscape from saprophytism to virulence","volume":"459","author":"Toledo-Arana","year":"2009","journal-title":"Nature"},{"key":"2023013110162188500_btz714-B33","doi-asserted-by":"crossref","first-page":"e1002284","DOI":"10.1371\/journal.pcbi.1002284","article-title":"Genome majority vote improves gene predictions","volume":"7","author":"Wall","year":"2011","journal-title":"PLoS Comput. Biol"},{"key":"2023013110162188500_btz714-B34","article-title":"Identifying small proteins by ribosome profiling with stalled initiation complexes","volume":"10","author":"Weaver","year":"2019","journal-title":"Mol Biol Physiol"},{"key":"2023013110162188500_btz714-B35","doi-asserted-by":"crossref","first-page":"1064","DOI":"10.1074\/mcp.M116.066662","article-title":"N-terminal proteomics assisted profiling of the unexplored translation initiation landscape in Arabidopsis thaliana","volume":"16","author":"Willems","year":"2017","journal-title":"Mol. Cell Proteomics"},{"key":"2023013110162188500_btz714-B36","doi-asserted-by":"crossref","first-page":"322.","DOI":"10.1186\/s12859-015-0749-z","article-title":"DECIPHER: harnessing local sequence context to improve protein multiple sequence alignment","volume":"16","author":"Wright","year":"2015","journal-title":"BMC Bioinformatics"},{"key":"2023013110162188500_btz714-B37","doi-asserted-by":"crossref","first-page":"352","DOI":"10.32614\/RJ-2016-025","article-title":"Using DECIPHER v2.0 to analyze big biological sequence data in R","volume":"8","author":"Wright","year":"2016","journal-title":"R. J"},{"key":"2023013110162188500_btz714-B38","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1186\/1471-2164-10-61","article-title":"Exploiting proteomic data for genome annotation and gene model validation in Aspergillus niger","volume":"10","author":"Wright","year":"2009","journal-title":"BMC Genomics"},{"key":"2023013110162188500_btz714-B39","doi-asserted-by":"crossref","first-page":"D613","DOI":"10.1093\/nar\/gks1235","article-title":"EcoGene 3.0","volume":"41","author":"Zhou","year":"2013","journal-title":"Nucleic Acids Res"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btz714\/30134403\/btz714.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/36\/4\/1022\/48982520\/btz714.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/36\/4\/1022\/48982520\/btz714.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,31]],"date-time":"2023-01-31T20:20:27Z","timestamp":1675196427000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/36\/4\/1022\/5571369"}},"subtitle":[],"editor":[{"given":"John","family":"Hancock","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2019,9,18]]},"references-count":39,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2020,2,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btz714","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2020,2,15]]},"published":{"date-parts":[[2019,9,18]]}}}