{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,12]],"date-time":"2026-03-12T00:12:30Z","timestamp":1773274350017,"version":"3.50.1"},"reference-count":21,"publisher":"Oxford University Press (OUP)","issue":"7","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2005,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Annotation of operons in a bacterial genome is an important step in determining an organism's transcriptional regulatory program. While extensive studies of operon structure have been carried out in a few species such as Escherichia coli, fewer resources exist to inform operon prediction in newly sequenced genomes. In particular, many extant operon finders require a large body of training examples to learn the properties of operons in the target organism. For newly sequenced genomes, such examples are generally not available; moreover, a model of operons trained on one species may not reflect the properties of other, distantly related organisms. We encountered these issues in the course of predicting operons in the genome of Bacteroides thetaiotaomicron (B.theta), a common anaerobe that is a prominent component of the normal adult human intestinal microbial community.<\/jats:p>\n               <jats:p>Results: We describe an operon predictor designed to work without extensive training data. We rely on a small set of a priori assumptions about the properties of the genome being annotated that permit estimation of the probability that two adjacent genes lie in a common operon. Predictions integrate several sources of information, including intergenic distance, common functional annotation and a novel formulation of conserved gene order. We validate our predictor both on the known operons of E.coli and on the genome of B.theta, using expression data to evaluate our predictions in the latter.<\/jats:p>\n               <jats:p>Availability: The software is available online at http:\/\/www.cse.wustl.edu\/~jbuhler\/research\/operons<\/jats:p>\n               <jats:p>Contact: \u00a0jbuhler@cse.wustl.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/bti123","type":"journal-article","created":{"date-parts":[[2004,11,12]],"date-time":"2004-11-12T01:14:59Z","timestamp":1100222099000},"page":"880-888","source":"Crossref","is-referenced-by-count":67,"title":["Operon prediction without a training set"],"prefix":"10.1093","volume":"21","author":[{"given":"B. P.","family":"Westover","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"J. D.","family":"Buhler","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"J. L.","family":"Sonnenburg","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"J. I.","family":"Gordon","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2004,11,11]]},"reference":[{"key":"2023013107283228700_B1","doi-asserted-by":"crossref","unstructured":"Altschul, S.F., Madden, T.L., Sch\u00e4ffer, A.A., Zhang, J., Zhang, Z., Miller, W., Lipman, D.J. 1997Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res.253389\u20133402","DOI":"10.1093\/nar\/25.17.3389"},{"key":"2023013107283228700_B2","doi-asserted-by":"crossref","unstructured":"Bockhorst, J., Craven, M., Page, D., Shavlik, J., Glasner, J. 2003A Bayesian network approach to operon prediction. Bioinformatics191227\u20131235","DOI":"10.1093\/bioinformatics\/btg147"},{"key":"2023013107283228700_B3","doi-asserted-by":"crossref","unstructured":"Bockhorst, J., Qiu, Y., Glasner, J., Liu, M., Blattner, F., Craven, M. 2003Predicting bacterial transcription units using sequence and expression data. Bioinformatics19i34\u2013i43","DOI":"10.1093\/bioinformatics\/btg1003"},{"key":"2023013107283228700_B4","doi-asserted-by":"crossref","unstructured":"De Hoon, M.J.L., Imoto, S., Kobayashi, K., Ogasawara, N., Miyano, S. 2004Predicting the operon structure of Bacillus subtilis using operon length, intergene distance, and gene expression information. Proceedings of the 2004 Pacific Symposium on BiocomputingSingapore  World Scientific,  pp. 276\u2013287","DOI":"10.1142\/9789812704856_0027"},{"key":"2023013107283228700_B5","doi-asserted-by":"crossref","unstructured":"Durand, D. and Sankoff, D. 2003Tests for gene clustering. J. Comput. Biol.10453\u2013482","DOI":"10.1145\/565196.565214"},{"key":"2023013107283228700_B6","unstructured":"Ermolaeva, M.D., White, O., Salzberg, S.L. 2001Prediction of operons in microbial genomes. Nucleic Acids Res.291216\u20131221"},{"key":"2023013107283228700_B7","unstructured":"Itoh, T. 2004Bacillus subtilis operon predictions, http:\/\/www.cib.nig.ac.jp\/dda\/taitoh\/bsub.operon.html"},{"key":"2023013107283228700_B8","doi-asserted-by":"crossref","unstructured":"Korf, I., Flicek, P., Duan, D., Brent, M.R. 2001Integrating genomic homology into gene structure prediction. Bioinformatics17(Suppl. 1),S140\u2013S148","DOI":"10.1093\/bioinformatics\/17.suppl_1.S140"},{"key":"2023013107283228700_B9","unstructured":"Li, C. and Wong, W. 2001Model-based analysis of oligonucleotide arrays: expression index computation and outlier detection. Proc. Natl Acad. Sci.9831\u201336"},{"key":"2023013107283228700_B10","unstructured":"Mitchell, T. Machine Learning1997, New York  McGraw Hill"},{"key":"2023013107283228700_B11","doi-asserted-by":"crossref","unstructured":"Moreno-Hagelsieb, G. and Collado-Vides, J. 2002A powerful non-homology method for the prediction of operons in eukaryotes. Bioinformatics18,  pp. S329\u2013S336","DOI":"10.1093\/bioinformatics\/18.suppl_1.S329"},{"key":"2023013107283228700_B12","unstructured":"Mushegian, A.R. and Koonin, E.V. 1996Gene order is not conserved in bacterial evolution. Trends Genet.12289\u2013290"},{"key":"2023013107283228700_B13","unstructured":"(Eds.). E.coli Gene Products: Physiological Functions and Common Ancestries1996 2nd edition , Washington, DC  American Society for Microbiology,  pp. 2118\u20132202"},{"key":"2023013107283228700_B14","doi-asserted-by":"crossref","unstructured":"Romero, P.R. and Karp, P.D. 2004Using functional and organizational information to improve genome-wide computational prediction of transcription units on pathway\/genome databases. Bioinformatics20709\u2013717","DOI":"10.1093\/bioinformatics\/btg471"},{"key":"2023013107283228700_B15","unstructured":"Salgado, H., Moreno-Hagelsieb, G., Smith, T.F., Collado-Vides, J. 2000Operons in Escherichia coli: genomic analyses and predictions. Proc. Natl Acad. Sci.976652\u20136657"},{"key":"2023013107283228700_B16","unstructured":"Salgado, H., Gama-Castro, S., Martinez-Antonio, A., Diaz-Peredo, E., Sanchez-Solano, F., Peralta-Gil, M., Garcia-Alonso, D., Jimenez-Jacinto, V., Santos-Zavaleta, A., Bonavides-Martinez, C., et al. 2004RegulonDB (version 4.0): transcriptional regulation, operon organization, and growth conditions in Escherichia coli K-12. Nucleic Acids Res.32D303\u2013D306"},{"key":"2023013107283228700_B17","unstructured":"TIGR.  2004 The Institute for Genome Research. TIGR operon finder"},{"key":"2023013107283228700_B18","doi-asserted-by":"crossref","unstructured":"Tjaden, B., Haynor, D.R., Stolyar, S., Rosenow, C., Kolker, E. 2002Identifying operons and untranslated regions of transcripts using Escherichia coli RNA expression analysis. Bioinformatics18S337\u2013S344","DOI":"10.1093\/bioinformatics\/18.suppl_1.S337"},{"key":"2023013107283228700_B19","unstructured":"Watterson, W.A., Ewens, W.J., Hall, T.E., Morgan, A. 1982The chromosome inversion problem. J. Theoret. Biol.991\u20137"},{"key":"2023013107283228700_B20","unstructured":"Xu, J., Bjursell, M.K., Himrod, J., Deng, S., Carmichael, L.K., Chiang, H.C., Hooper, L.V., Gordon, J.I. 2003A genomic view of the human-Bacteroides thetaiotaomicron symbiosis. Science2992074\u20132076"},{"key":"2023013107283228700_B21","doi-asserted-by":"crossref","unstructured":"Yada, T., Nakao, M., Totoki, Y., Nakai, K. 1999Modeling and predicting transcriptional units of {Escherichia coli} genes using hidden Markov models. Bioinformatics15987\u2013993","DOI":"10.1093\/bioinformatics\/15.12.987"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/7\/880\/48967233\/bioinformatics_21_7_880.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/7\/880\/48967233\/bioinformatics_21_7_880.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,31]],"date-time":"2023-01-31T10:36:58Z","timestamp":1675161418000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/21\/7\/880\/268962"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2004,11,11]]},"references-count":21,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2005,4,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bti123","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2005,4,1]]},"published":{"date-parts":[[2004,11,11]]}}}