{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,1,3]],"date-time":"2025-01-03T22:40:01Z","timestamp":1735944001359,"version":"3.32.0"},"reference-count":29,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec><jats:title>Background<\/jats:title><jats:p>The currently used<jats:italic>k<\/jats:italic><jats:sup><jats:italic>th<\/jats:italic><\/jats:sup>order Markov models estimate the probability of generating a<jats:italic>single<\/jats:italic>nucleotide conditional upon the immediately preceding (<jats:italic>gap<\/jats:italic>= 0)<jats:italic>k<\/jats:italic>units. However, this neither takes into account the joint dependency of<jats:italic>multiple<\/jats:italic>neighboring nucleotides, nor does it consider the long range dependency with<jats:italic>gap<\/jats:italic>&gt;0.<\/jats:p><\/jats:sec><jats:sec><jats:title>Result<\/jats:title><jats:p>We describe a configurable tool to explore generalizations of the standard Markov model. We evaluated whether the sequence classification accuracy can be improved by using an alternative set of model parameters. The evaluation was done on four classes of biological sequences \u2013 CpG-poor promoters, all promoters, exons and nucleosome positioning sequences. Using di- and tri-nucleotide as the model unit significantly improved the sequence classification accuracy relative to the standard single nucleotide model. In the case of nucleosome positioning sequences, optimal accuracy was achieved at a<jats:italic>gap<\/jats:italic>length of 4. Furthermore in the plot of classification accuracy versus the gap, a periodicity of 10\u201311 bps was observed which might indicate structural preferences in the nucleosome positioning sequence. The tool is implemented in Java and is available for download at<jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"ftp:\/\/ftp.pcbi.upenn.edu\/GMM\/\">ftp:\/\/ftp.pcbi.upenn.edu\/GMM\/<\/jats:ext-link>.<\/jats:p><\/jats:sec><jats:sec><jats:title>Conclusion<\/jats:title><jats:p>Markov modeling is an important component of many sequence analysis tools. We have extended the standard Markov model to incorporate joint and long range dependencies between the sequence elements. The proposed generalizations of the Markov model are likely to improve the overall accuracy of sequence analysis tools.<\/jats:p><\/jats:sec>","DOI":"10.1186\/1471-2105-6-219","type":"journal-article","created":{"date-parts":[[2005,9,6]],"date-time":"2005-09-06T18:13:44Z","timestamp":1126030424000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Generalizations of Markov model to characterize biological sequences"],"prefix":"10.1186","volume":"6","author":[{"given":"Junwen","family":"Wang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sridhar","family":"Hannenhalli","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2005,9,6]]},"reference":[{"key":"544_CR1","doi-asserted-by":"publisher","first-page":"799","DOI":"10.1093\/protein\/gzg101","volume":"16","author":"J Wang","year":"2003","unstructured":"Wang J, Feng JA: Exploring the sequence patterns in the alpha-helices of proteins. Protein Eng 2003, 16: 799\u2013807. 10.1093\/protein\/gzg101","journal-title":"Protein Eng"},{"key":"544_CR2","doi-asserted-by":"publisher","first-page":"443","DOI":"10.1016\/0022-2836(70)90057-4","volume":"48","author":"SB Needleman","year":"1970","unstructured":"Needleman SB, Wunsch CD: A general method applicable to the search for similarities in the amino acid sequence of two proteins. J Mol Biol 1970, 48: 443\u2013453. 10.1016\/0022-2836(70)90057-4","journal-title":"J Mol Biol"},{"key":"544_CR3","doi-asserted-by":"publisher","first-page":"628","DOI":"10.1002\/prot.20359","volume":"58","author":"J Wang","year":"2005","unstructured":"Wang J, Feng JA: NdPASA: a novel pairwise protein sequence alignment algorithm that incorporates neighbor-dependent amino acid propensities. Proteins 2005, 58: 628\u2013637. 10.1002\/prot.20359","journal-title":"Proteins"},{"key":"544_CR4","doi-asserted-by":"publisher","first-page":"1255","DOI":"10.1093\/nar\/30.5.1255","volume":"30","author":"ML Bulyk","year":"2002","unstructured":"Bulyk ML, Johnson PL, Church GM: Nucleotides of transcription factor binding sites exert interdependent effects on the binding affinities of transcription factors. Nucleic Acids Res 2002, 30: 1255\u20131261. 10.1093\/nar\/30.5.1255","journal-title":"Nucleic Acids Res"},{"key":"544_CR5","doi-asserted-by":"publisher","first-page":"909","DOI":"10.1093\/bioinformatics\/bth006","volume":"20","author":"Q Zhou","year":"2004","unstructured":"Zhou Q, Liu JS: Modeling within-motif dependence for transcription factor binding site predictions. Bioinformatics 2004, 20: 909\u2013916. 10.1093\/bioinformatics\/bth006","journal-title":"Bioinformatics"},{"key":"544_CR6","volume-title":"Monographs on statistics and applied probability","author":"MHA Davis","year":"1993","unstructured":"Davis MHA: Markov Models & Optimization. In Monographs on statistics and applied probability. Volume 49. , CHAPMAN & HALL; 1993."},{"key":"544_CR7","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511790492","volume-title":"Biological Sequence Analysis","author":"R Durbin","year":"1998","unstructured":"Durbin R, Eddy S, Krogh A, Mitchison G: Biological Sequence Analysis. 1998."},{"key":"544_CR8","doi-asserted-by":"publisher","first-page":"56","DOI":"10.1016\/S0168-9525(00)02174-0","volume":"17","author":"U Ohler","year":"2001","unstructured":"Ohler U, Niemann H: Identification and analysis of eukaryotic promoters: recent computational approaches. Trends Genet 2001, 17: 56\u201360. 10.1016\/S0168-9525(00)02174-0","journal-title":"Trends Genet"},{"key":"544_CR9","doi-asserted-by":"publisher","first-page":"261","DOI":"10.1016\/0022-2836(87)90689-9","volume":"196","author":"M Gardiner-Garden","year":"1987","unstructured":"Gardiner-Garden M, Frommer M: CpG islands in vertebrate genomes. J Mol Biol 1987, 196: 261\u2013282. 10.1016\/0022-2836(87)90689-9","journal-title":"J Mol Biol"},{"key":"544_CR10","doi-asserted-by":"publisher","first-page":"825","DOI":"10.1080\/07391102.1999.10508295","volume":"16","author":"ON Ozoline","year":"1999","unstructured":"Ozoline ON, Deev AA, Trifonov EN: DNA bendability--a novel feature in E. coli promoter recognition. J Biomol Struct Dyn 1999, 16: 825\u2013831.","journal-title":"J Biomol Struct Dyn"},{"key":"544_CR11","doi-asserted-by":"publisher","first-page":"215","DOI":"10.1529\/biophysj.103.020743","volume":"87","author":"RA Dimitrov","year":"2004","unstructured":"Dimitrov RA, Zuker M: Prediction of hybridization and melting for double-stranded nucleic acids. Biophys J 2004, 87: 215\u2013226. 10.1529\/biophysj.103.020743","journal-title":"Biophys J"},{"key":"544_CR12","doi-asserted-by":"publisher","first-page":"891","DOI":"10.1016\/j.jmb.2004.08.068","volume":"343","author":"P Schieg","year":"2004","unstructured":"Schieg P, Herzel H: Periodicities of 10\u201311bp as indicators of the supercoiled state of genomic DNA. J Mol Biol 2004, 343: 891\u2013901. 10.1016\/j.jmb.2004.08.068","journal-title":"J Mol Biol"},{"key":"544_CR13","doi-asserted-by":"crossref","first-page":"528","DOI":"10.1111\/j.2517-6161.1985.tb01383.x","volume":"B47","author":"AE Raftery","year":"1985","unstructured":"Raftery AE: A model for high order Markov chains. J Roy Statst Soc Ser 1985, B47: 528\u2013539.","journal-title":"J Roy Statst Soc Ser"},{"key":"544_CR14","doi-asserted-by":"publisher","first-page":"379","DOI":"10.1111\/1467-9892.00231","volume":"22","author":"A Berchtold","year":"2001","unstructured":"Berchtold A: Estimation in the mixture transition distribution model. J TIme Ser Anal 2001, 22: 379\u2013397. 10.1111\/1467-9892.00231","journal-title":"J TIme Ser Anal"},{"key":"544_CR15","doi-asserted-by":"publisher","first-page":"328","DOI":"10.1214\/ss\/1042727943","volume":"17","author":"A Berchtold","year":"2002","unstructured":"Berchtold A, Raftery AE: The mixture transition distribution model for high-order Markov chains and non-Gaussian time series. Statistical Science 2002, 17: 328\u2013356. 10.1214\/ss\/1042727943","journal-title":"Statistical Science"},{"key":"544_CR16","doi-asserted-by":"publisher","first-page":"1211","DOI":"10.1006\/jmbi.1999.3206","volume":"293","author":"S Penel","year":"1999","unstructured":"Penel S, Morrison RG, Mortishire-Smith RJ, Doig AJ: Periodicity in alpha-helix lengths and C-capping preferences. J Mol Biol 1999, 293: 1211\u20131219. 10.1006\/jmbi.1999.3206","journal-title":"J Mol Biol"},{"key":"544_CR17","doi-asserted-by":"publisher","first-page":"480","DOI":"10.1214\/aos\/1018031204","volume":"27","author":"P Buhlmann","year":"1999","unstructured":"Buhlmann P, Wyner AJ: Variable length markov chains. The Annals of Statistics 1999, 27: 480\u2013513. 10.1214\/aos\/1018031204","journal-title":"The Annals of Statistics"},{"key":"544_CR18","doi-asserted-by":"publisher","first-page":"993","DOI":"10.1093\/bioinformatics\/bth028","volume":"20","author":"RK Azad","year":"2004","unstructured":"Azad RK, Borodovsky M: Effects of choice of DNA sequence model structure on gene identification accuracy. Bioinformatics 2004, 20: 993\u20131005. 10.1093\/bioinformatics\/bth028","journal-title":"Bioinformatics"},{"key":"544_CR19","doi-asserted-by":"publisher","first-page":"544","DOI":"10.1093\/nar\/26.2.544","volume":"26","author":"SL Salzberg","year":"1998","unstructured":"Salzberg SL, Delcher AL, Kasif S, White O: Microbial gene identification using interpolated Markov models. Nucleic Acids Res 1998, 26: 544\u2013548. 10.1093\/nar\/26.2.544","journal-title":"Nucleic Acids Res"},{"key":"544_CR20","doi-asserted-by":"publisher","first-page":"328","DOI":"10.1093\/nar\/30.1.328","volume":"30","author":"Y Suzuki","year":"2002","unstructured":"Suzuki Y, Yamashita R, Nakai K, Sugano S: DBTSS: DataBase of human Transcriptional Start Sites and full-length cDNAs. Nucleic Acids Res 2002, 30: 328\u2013331. 10.1093\/nar\/30.1.328","journal-title":"Nucleic Acids Res"},{"key":"544_CR21","doi-asserted-by":"publisher","first-page":"11995","DOI":"10.1073\/pnas.90.24.11995","volume":"90","author":"F Antequera","year":"1993","unstructured":"Antequera F, Bird A: Number of CpG islands and genes in human and mouse. Proc Natl Acad Sci U S A 1993, 90: 11995\u201311999.","journal-title":"Proc Natl Acad Sci U S A"},{"key":"544_CR22","first-page":"D67","volume":"33 (Database Is","author":"VG Levitsky","year":"2005","unstructured":"Levitsky VG, Katokhin AV, Podkolodnaya OA, Furman DP, Kolchanov NA: NPRD: Nucleosome Positioning Region Database. Nucleic Acids Res 2005, 33 (Database Issue): D67\u201370.","journal-title":"Nucleic Acids Res"},{"key":"544_CR23","doi-asserted-by":"publisher","first-page":"2891","DOI":"10.1073\/pnas.96.6.2891","volume":"96","author":"I Ioshikhes","year":"1999","unstructured":"Ioshikhes I, Trifonov EN, Zhang MQ: Periodical distribution of transcription factor sites in promoter regions and connection with chromatin structure. Proc Natl Acad Sci U S A 1999, 96: 2891\u20132895. 10.1073\/pnas.96.6.2891","journal-title":"Proc Natl Acad Sci U S A"},{"key":"544_CR24","doi-asserted-by":"publisher","first-page":"412","DOI":"10.1038\/ng780","volume":"29","author":"RV Davuluri","year":"2001","unstructured":"Davuluri RV, Grosse I, Zhang MQ: Computational identification of promoters and first exons in the human genome. Nat Genet 2001, 29: 412\u2013417. 10.1038\/ng780","journal-title":"Nat Genet"},{"key":"544_CR25","doi-asserted-by":"publisher","first-page":"S90","DOI":"10.1093\/bioinformatics\/17.suppl_1.S90","volume":"17 Suppl 1","author":"S Hannenhalli","year":"2001","unstructured":"Hannenhalli S, Levy S: Promoter prediction in the human genome. Bioinformatics 2001, 17 Suppl 1: S90\u20136.","journal-title":"Bioinformatics"},{"key":"544_CR26","doi-asserted-by":"publisher","first-page":"1467","DOI":"10.1038\/nbt1032","volume":"22","author":"VB Bajic","year":"2004","unstructured":"Bajic VB, Tan SL, Suzuki Y, Sugano S: Promoter prediction analysis on the whole human genome. Nat Biotechnol 2004, 22: 1467\u20131473. 10.1038\/nbt1032","journal-title":"Nat Biotechnol"},{"key":"544_CR27","doi-asserted-by":"publisher","first-page":"4772","DOI":"10.1073\/pnas.95.8.4772","volume":"95","author":"G Li","year":"1998","unstructured":"Li G, Chandler SP, Wolffe AP, Hall TC: Architectural specificity in chromatin structure at the TATA box in vivo: nucleosome displacement upon beta-phaseolin gene activation. Proc Natl Acad Sci U S A 1998, 95: 4772\u20134777. 10.1073\/pnas.95.8.4772","journal-title":"Proc Natl Acad Sci U S A"},{"key":"544_CR28","doi-asserted-by":"publisher","first-page":"78","DOI":"10.1006\/jmbi.1997.0951","volume":"268","author":"C Burge","year":"1997","unstructured":"Burge C, Karlin S: Prediction of complete gene structures in human genomic DNA. J Mol Biol 1997, 268: 78\u201394. 10.1006\/jmbi.1997.0951","journal-title":"J Mol Biol"},{"key":"544_CR29","first-page":"179","volume":"5","author":"A Krogh","year":"1997","unstructured":"Krogh A: Two methods for improving performance of an HMM and their application for gene finding. Proc Int Conf Intell Syst Mol Biol 1997, 5: 179\u2013186.","journal-title":"Proc Int Conf Intell Syst Mol Biol"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-6-219.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,3]],"date-time":"2025-01-03T22:20:51Z","timestamp":1735942851000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-6-219"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,9,6]]},"references-count":29,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2005,12]]}},"alternative-id":["544"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-6-219","relation":{},"ISSN":["1471-2105"],"issn-type":[{"type":"electronic","value":"1471-2105"}],"subject":[],"published":{"date-parts":[[2005,9,6]]},"assertion":[{"value":"29 April 2005","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 September 2005","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 September 2005","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"219"}}