{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T01:46:04Z","timestamp":1783734364799,"version":"3.55.0"},"reference-count":37,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,9,18]],"date-time":"2021-09-18T00:00:00Z","timestamp":1631923200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,9,18]],"date-time":"2021-09-18T00:00:00Z","timestamp":1631923200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100000060","name":"national institute of allergy and infectious diseases","doi-asserted-by":"publisher","award":["R01AI141810"],"award-info":[{"award-number":["R01AI141810"]}],"id":[{"id":"10.13039\/100000060","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000060","name":"national institute of allergy and infectious diseases","doi-asserted-by":"publisher","award":["R01AI145552"],"award-info":[{"award-number":["R01AI145552"]}],"id":[{"id":"10.13039\/100000060","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"national science foundation","doi-asserted-by":"publisher","award":["2013998"],"award-info":[{"award-number":["2013998"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Background<\/jats:title>\n                <jats:p>Identification of motifs and quantification of their occurrences are important for the study of genetic diseases, gene evolution, transcription sites, and other biological mechanisms. Exact formulae for estimating count distributions of motifs under Markovian assumptions have high computational complexity and are impractical to be used on large motif sets. Approximated formulae, e.g. based on compound Poisson, are faster, but reliable <jats:italic>p<\/jats:italic> value calculation remains challenging. Here, we introduce \u2018motif_prob\u2019, a fast implementation of an exact formula for motif count distribution through progressive approximation with arbitrary precision. Our implementation speeds up the exact calculation, usually impractical, making it feasible and posit to substitute currently employed heuristics.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>We implement motif_prob in both Perl and C+\u2009+ languages, using an efficient error-bound iterative process for the exact formula, providing comparison with state-of-the-art tools (e.g. MoSDi) in terms of precision, run time benchmarks, along with a real-world use case on bacterial motif characterization. Our software is able to process a million of motifs (13\u201331 bases) over genome lengths of 5 million bases within the minute on a regular laptop, and the run times for both the Perl and C+\u2009+ code are several orders of magnitude smaller (50\u20131000\u00d7\u2009faster) than MoSDi, even when using their fast compound Poisson approximation (60\u2013120\u00d7\u2009faster). In the real-world use cases, we first show the consistency of motif_prob with MoSDi, and then how the p-value quantification is crucial for enrichment quantification when bacteria have different GC content, using motifs found in antimicrobial resistance genes. The software and the code sources are available under the MIT license at <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/DataIntellSystLab\/motif_prob\">https:\/\/github.com\/DataIntellSystLab\/motif_prob<\/jats:ext-link>.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusions<\/jats:title>\n                <jats:p>The motif_prob software is a multi-platform and efficient open source solution for calculating exact frequency distributions of motifs. It can be integrated with motif discovery\/characterization tools for quantifying enrichment and deviation from expected frequency ranges with exact <jats:italic>p<\/jats:italic> values, without loss in data processing efficiency.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1186\/s12859-021-04355-6","type":"journal-article","created":{"date-parts":[[2021,9,18]],"date-time":"2021-09-18T18:02:28Z","timestamp":1631988148000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Fast and exact quantification of motif occurrences in biological sequences"],"prefix":"10.1186","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9021-5595","authenticated-orcid":false,"given":"Mattia","family":"Prosperi","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Simone","family":"Marini","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christina","family":"Boucher","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,9,18]]},"reference":[{"issue":"12","key":"4355_CR1","doi-asserted-by":"publisher","first-page":"2013","DOI":"10.1101\/gr.155960.113","volume":"23","author":"P-L Luu","year":"2013","unstructured":"Luu P-L, Sch\u00f6ler HR, Ara\u00fazo-Bravo MJ. Disclosing the crosstalk among DNA methylation, transcription factors, and histone marks in human pluripotent cells through discovery of DNA methylation motifs. Genome Res. 2013;23(12):2013\u201329.","journal-title":"Genome Res"},{"key":"4355_CR2","doi-asserted-by":"publisher","first-page":"743","DOI":"10.1038\/nrg1691","volume":"6","author":"JR Gatchel","year":"2005","unstructured":"Gatchel JR, Zoghbi HY. Diseases of unstable repeat expansion: mechanisms and common principles. Nat Rev Genet. 2005;6:743\u201355.","journal-title":"Nat Rev Genet"},{"key":"4355_CR3","doi-asserted-by":"publisher","first-page":"2013","DOI":"10.1101\/gr.155960.113","volume":"23","author":"PL Luu","year":"2013","unstructured":"Luu PL, Sch\u00f6ler HR, Ara\u00fazo-Bravo MJ. Disclosing the crosstalk among DNA methylation, transcription factors, and histone marks in human pluripotent cells through discovery of DNA methylation motifs. Genome Res. 2013;23:2013\u201329.","journal-title":"Genome Res"},{"issue":"1","key":"4355_CR4","doi-asserted-by":"publisher","first-page":"137","DOI":"10.1038\/nbt1053","volume":"23","author":"M Tompa","year":"2005","unstructured":"Tompa M, Li N, Bailey TL, Church GM, De Moor B, Eskin E, et al. Assessing computational tools for the discovery of transcription factor binding sites. Nat Biotechnol. 2005;23(1):137\u201344.","journal-title":"Nat Biotechnol"},{"issue":"466","key":"4355_CR5","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1016\/j.ins.2018.07.004","volume":"1","author":"NK Lee","year":"2018","unstructured":"Lee NK, Li X, Wang D. A comprehensive survey on genetic algorithms for DNA motif prediction. Inf Sci. 2018;1(466):25\u201343.","journal-title":"Inf Sci"},{"issue":"2","key":"4355_CR6","first-page":"130","volume":"11","author":"FA Hashim","year":"2019","unstructured":"Hashim FA, Mabrouk MS, Al-Atabany W. Review of different sequence motif finding algorithms. Avicenna J Med Biotechnol. 2019;11(2):130\u201348.","journal-title":"Avicenna J Med Biotechnol"},{"issue":"Web Server issu","key":"4355_CR7","doi-asserted-by":"publisher","first-page":"W199","DOI":"10.1093\/nar\/gkh465","volume":"32","author":"G Pavesi","year":"2004","unstructured":"Pavesi G, Mereghetti P, Mauri G, Pesole G. Weeder Web: discovery of transcription factor binding sites in a set of sequences from co-regulated genes. Nucleic Acids Res. 2004;32(Web Server issue):W199\u2013203.","journal-title":"Nucleic Acids Res"},{"issue":"7","key":"4355_CR8","doi-asserted-by":"publisher","first-page":"563","DOI":"10.1038\/nmeth1061","volume":"4","author":"L Ettwiller","year":"2007","unstructured":"Ettwiller L, Paten B, Ramialison M, Birney E, Wittbrodt J. Trawler: de novo regulatory motif discovery pipeline for chromatin immunoprecipitation. Nat Methods. 2007;4(7):563\u20135.","journal-title":"Nat Methods"},{"issue":"Web Server issu","key":"4355_CR9","doi-asserted-by":"publisher","first-page":"W202","DOI":"10.1093\/nar\/gkp335","volume":"37","author":"TL Bailey","year":"2009","unstructured":"Bailey TL, Boden M, Buske FA, Frith M, Grant CE, Clementi L, et al. MEME SUITE: tools for motif discovery and searching. Nucleic Acids Res. 2009;37(Web Server issue):W202\u20138.","journal-title":"Nucleic Acids Res"},{"issue":"12","key":"4355_CR10","doi-asserted-by":"publisher","first-page":"1653","DOI":"10.1093\/bioinformatics\/btr261","volume":"27","author":"TL Bailey","year":"2011","unstructured":"Bailey TL. DREME: motif discovery in transcription factor ChIP-seq data. Bioinformatics. 2011;27(12):1653\u20139.","journal-title":"Bioinformatics"},{"issue":"4","key":"4355_CR11","doi-asserted-by":"publisher","first-page":"e31","DOI":"10.1093\/nar\/gkr1104","volume":"40","author":"M Thomas-Chollier","year":"2012","unstructured":"Thomas-Chollier M, Herrmann C, Defrance M, Sand O, Thieffry D, van Helden J. RSAT peak-motifs: motif analysis in full-size ChIP-seq datasets. Nucleic Acids Res. 2012;40(4):e31\u2013e31.","journal-title":"Nucleic Acids Res"},{"issue":"1","key":"4355_CR12","doi-asserted-by":"publisher","first-page":"238","DOI":"10.1186\/s12864-018-4630-0","volume":"19","author":"LT Dang","year":"2018","unstructured":"Dang LT, Tondl M, Chiu MHH, Revote J, Paten B, Tano V, et al. TrawlerWeb: an online de novo motif discovery tool for next-generation sequencing datasets. BMC Genomics. 2018;19(1):238.","journal-title":"BMC Genomics"},{"issue":"1","key":"4355_CR13","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1186\/s12859-017-2005-1","volume":"19","author":"JM Caldonazzo Garbelini","year":"2018","unstructured":"Caldonazzo Garbelini JM, Kashiwabara AY, Sanches DS. Sequence motif finder using memetic algorithm. BMC Bioinform. 2018;19(1):4.","journal-title":"BMC Bioinform"},{"issue":"22","key":"4355_CR14","doi-asserted-by":"publisher","first-page":"4632","DOI":"10.1093\/bioinformatics\/btz290","volume":"35","author":"Y Li","year":"2019","unstructured":"Li Y, Ni P, Zhang S, Li G, Su Z. ProSampler: an ultrafast and accurate motif finder in large ChIP-seq datasets for combinatory motif discovery. Berger B, editor. Bioinformatics. 2019;35(22):4632\u20139.","journal-title":"Bioinformatics"},{"key":"4355_CR15","doi-asserted-by":"crossref","unstructured":"Bailey TL. STREME: accurate and versatile sequence motif discovery. bioRxiv. 2020;2020.11.23.394619.","DOI":"10.1101\/2020.11.23.394619"},{"issue":"W1","key":"4355_CR16","doi-asserted-by":"publisher","first-page":"W215","DOI":"10.1093\/nar\/gky431","volume":"46","author":"A Kiesel","year":"2018","unstructured":"Kiesel A, Roth C, Ge W, Wess M, Meier M, S\u00f6ding J. The BaMM web server for de-novo motif discovery and regulatory sequence analysis. Nucleic Acids Res. 2018;46(W1):W215\u201320.","journal-title":"Nucleic Acids Res"},{"issue":"2","key":"4355_CR17","doi-asserted-by":"publisher","first-page":"R24","DOI":"10.1186\/gb-2007-8-2-r24","volume":"8","author":"S Gupta","year":"2007","unstructured":"Gupta S, Stamatoyannopoulos JA, Bailey TL, Noble WS. Quantifying similarity between motifs. Genome Biol. 2007;8(2):R24.","journal-title":"Genome Biol"},{"key":"4355_CR18","doi-asserted-by":"publisher","unstructured":"Finding similar regions in many strings|Proceedings of the thirty-first annual ACM symposium on Theory of Computing [Internet]. [cited 2021 May 28]. https:\/\/doi.org\/10.1145\/301250.301376.","DOI":"10.1145\/301250.301376"},{"issue":"5","key":"4355_CR19","doi-asserted-by":"publisher","first-page":"531","DOI":"10.1093\/bioinformatics\/btl662","volume":"23","author":"J Zhang","year":"2007","unstructured":"Zhang J, Jiang B, Li M, Tromp J, Zhang X, Zhang MQ. Computing exact p values for DNA motifs. Bioinformatics. 2007;23(5):531\u20137.","journal-title":"Bioinformatics"},{"issue":"1","key":"4355_CR20","doi-asserted-by":"publisher","first-page":"35","DOI":"10.2307\/2532033","volume":"45","author":"JF Gentleman","year":"1989","unstructured":"Gentleman JF, Mullin RC. The distribution of the frequency of occurrence of nucleotide subsequences, based on their overlap capability. Biometrics. 1989;45(1):35\u201352.","journal-title":"Biometrics"},{"issue":"1","key":"4355_CR21","doi-asserted-by":"publisher","first-page":"259","DOI":"10.1016\/S0166-218X(00)00195-5","volume":"104","author":"M R\u00e9gnier","year":"2000","unstructured":"R\u00e9gnier M. A unified approach to word occurrence probabilities. Discrete Appl Math. 2000;104(1):259\u201380.","journal-title":"Discrete Appl Math"},{"issue":"2","key":"4355_CR22","doi-asserted-by":"publisher","first-page":"593","DOI":"10.1016\/S0304-3975(01)00264-X","volume":"287","author":"P Nicod\u00e8me","year":"2002","unstructured":"Nicod\u00e8me P, Salvy B, Flajolet P. Motif statistics. Theor Comput Sci. 2002;287(2):593\u2013617.","journal-title":"Theor Comput Sci"},{"issue":"6","key":"4355_CR23","doi-asserted-by":"publisher","first-page":"761","DOI":"10.1089\/10665270260518254","volume":"9","author":"S Robin","year":"2002","unstructured":"Robin S, Daudin J-J, Richard H, Sagot M-F, Schbath S. Occurrence probability of structured motifs in random sequences. J Comput Biol J Comput Mol Cell Biol. 2002;9(6):761\u201373.","journal-title":"J Comput Biol J Comput Mol Cell Biol"},{"issue":"1","key":"4355_CR24","doi-asserted-by":"publisher","first-page":"95","DOI":"10.1016\/S0097-3165(03)00123-7","volume":"104","author":"E Rivals","year":"2003","unstructured":"Rivals E, Rahmann S. Combinatorics of periods in strings. J Comb Theory Ser A. 2003;104(1):95\u2013113.","journal-title":"J Comb Theory Ser A"},{"issue":"5","key":"4355_CR25","doi-asserted-by":"publisher","first-page":"867","DOI":"10.1089\/cmb.2004.11.867","volume":"11","author":"G Bejerano","year":"2004","unstructured":"Bejerano G, Friedman N, Tishby N. Efficient exact p-value computation for small sample, sparse, and surprising categorical data. J Comput Biol J Comput Mol Cell Biol. 2004;11(5):867\u201386.","journal-title":"J Comput Biol J Comput Mol Cell Biol"},{"issue":"1","key":"4355_CR26","first-page":"51","volume":"56","author":"ME Lladser","year":"2008","unstructured":"Lladser ME, Betterton MD, Knight R. Multiple pattern matching: a Markov chain approach. J Math Biol. 2008;56(1):51\u201392.","journal-title":"J Math Biol"},{"issue":"12","key":"4355_CR27","doi-asserted-by":"publisher","first-page":"i356","DOI":"10.1093\/bioinformatics\/btp188","volume":"25","author":"T Marschall","year":"2009","unstructured":"Marschall T, Rahmann S. Efficient exact motif discovery. Bioinformatics. 2009;25(12):i356\u201364.","journal-title":"Bioinformatics"},{"key":"4355_CR28","doi-asserted-by":"publisher","first-page":"1250055","DOI":"10.1142\/S1793524512500556","volume":"5","author":"MCF Prosperi","year":"2012","unstructured":"Prosperi MCF, Prosperi L, Gray RR, Salemi M. On counting the frequency distribution of string motifs in molecular sequences. Int J Biomath. 2012;5:1250055.","journal-title":"Int J Biomath"},{"issue":"13","key":"4355_CR29","doi-asserted-by":"publisher","first-page":"3826","DOI":"10.1093\/nar\/gkh713","volume":"32","author":"GB Fogel","year":"2004","unstructured":"Fogel GB, Weekes DG, Varga G, Dow ER, Harlow HB, Onyia JE, et al. Discovery of sequence motifs related to coexpression of genes using evolutionary computation. Nucleic Acids Res. 2004;32(13):3826\u201335.","journal-title":"Nucleic Acids Res"},{"key":"4355_CR30","doi-asserted-by":"publisher","first-page":"337","DOI":"10.1007\/978-3-642-15294-8_28","volume-title":"Algorithms in bioinformatics. Lecture notes in computer science","author":"T Marschall","year":"2010","unstructured":"Marschall T, Rahmann S. Speeding up exact motif discovery by bounding the expected clump size. In: Moulton V, Singh M, editors. Algorithms in bioinformatics. Lecture notes in computer science. Berlin: Springer; 2010. p. 337\u201349."},{"key":"4355_CR31","unstructured":"Kopp W. motifcounter: R package for analysing TFBSs in DNA sequences [Internet]. Bioconductor version: Release (3.12); 2021 [cited 2021 Mar 17]. https:\/\/bioconductor.org\/packages\/motifcounter\/."},{"issue":"6","key":"4355_CR32","doi-asserted-by":"publisher","first-page":"547","DOI":"10.1089\/cmb.2007.0084","volume":"15","author":"UJ Pape","year":"2008","unstructured":"Pape UJ, Rahmann S, Sun F, Vingron M. Compound poisson approximation of the number of occurrences of a position frequency matrix (PFM) on both strands. J Comput Biol J Comput Mol Cell Biol. 2008;15(6):547\u201364.","journal-title":"J Comput Biol J Comput Mol Cell Biol"},{"key":"4355_CR33","unstructured":"DNA, Words and Models: Statistics of Exceptional Words by S. Robin, F. Rodolphe, S. Schbath | 9780521847292 | Hardcover | Barnes & Noble\u00ae [Internet]. [cited 2021 Mar 17]. https:\/\/www.barnesandnoble.com\/w\/dna-words-and-models-s-robin\/1110953123."},{"key":"4355_CR34","doi-asserted-by":"publisher","first-page":"2484","DOI":"10.1093\/jac\/dkw184","volume":"71","author":"PTLC Clausen","year":"2016","unstructured":"Clausen PTLC, Zankari E, Aarestrup FM, Lund O. Benchmarking of methods for identification of antimicrobial resistance genes in bacterial whole genome data. J Antimicrob Chemother. 2016;71:2484\u20138.","journal-title":"J Antimicrob Chemother"},{"key":"4355_CR35","doi-asserted-by":"publisher","first-page":"e1001107","DOI":"10.1371\/journal.pgen.1001107","volume":"6","author":"F Hildebrand","year":"2010","unstructured":"Hildebrand F, Meyer A, Eyre-Walker A. Evidence of selection upon genomic GC-content in bacteria. PLoS Genet. 2010;6:e1001107.","journal-title":"PLoS Genet"},{"key":"4355_CR36","doi-asserted-by":"publisher","first-page":"D561","DOI":"10.1093\/nar\/gkz1010","volume":"48","author":"E Doster","year":"2020","unstructured":"Doster E, Lakin SM, Dean CJ, Wolfe C, Young JG, Boucher C, et al. MEGARes 2.0: a database for classification of antimicrobial drug, biocide and metal resistance determinants in metagenomic sequence data. Nucleic Acids Res. 2020;48:D561\u20139.","journal-title":"Nucleic Acids Res"},{"key":"4355_CR37","doi-asserted-by":"publisher","first-page":"giaa038","DOI":"10.1093\/gigascience\/giaa038","volume":"9","author":"O Ibironke","year":"2020","unstructured":"Ibironke O, McGuinness LR, Lu S-E, Wang Y, Hussain S, Weisel CP, et al. Species-level evaluation of the human respiratory microbiome. GigaScience. 2020;9:giaa038. https:\/\/doi.org\/10.1093\/gigascience\/giaa038.","journal-title":"GigaScience"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-021-04355-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12859-021-04355-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-021-04355-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,18]],"date-time":"2021-09-18T18:03:06Z","timestamp":1631988186000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-021-04355-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,18]]},"references-count":37,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["4355"],"URL":"https:\/\/doi.org\/10.1186\/s12859-021-04355-6","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,9,18]]},"assertion":[{"value":"4 April 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 September 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 September 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"Dr. Christina Boucher is Associate Editor of BMC Bioinformatics. All other authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"445"}}