{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,26]],"date-time":"2026-08-26T03:20:24Z","timestamp":1787714424073,"version":"build-2784847793"},"reference-count":41,"publisher":"Oxford University Press (OUP)","issue":"3","license":[{"start":{"date-parts":[[2016,10,2]],"date-time":"2016-10-02T00:00:00Z","timestamp":1475366400000},"content-version":"vor","delay-in-days":3196,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/2.0\/uk\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2008,2,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Transcription factors (TFs) play a key role in gene regulation by binding to target sequences. In silico prediction of potential binding of a TF to a binding site is a well-studied problem in computational biology. The binding sites for one TF are represented by a position frequency matrix (PFM). The discovery of new PFMs requires the comparison to known PFMs to avoid redundancies. In general, two PFMs are similar if they occur at overlapping positions under a null model. Still, most existing methods compute similarity according to probabilistic distances of the PFMs. Here we propose a natural similarity measure based on the asymptotic covariance between the number of PFM hits incorporating both strands. Furthermore, we introduce a second measure based on the same idea to cluster a set of the Jaspar PFMs.<\/jats:p>\n               <jats:p>Results: We show that the asymptotic covariance can be efficiently computed by a two dimensional convolution of the score distributions. The asymptotic covariance approach shows strong correlation with simulated data. It outperforms three alternative methods. The Jaspar clustering yields distinct groups of TFs of the same class. Furthermore, a representative PFM is given for each class. In contrast to most other clustering methods, PFMs with low similarity automatically remain singletons.<\/jats:p>\n               <jats:p>Availability: A website to compute the similarity and to perform clustering, the source code and Supplementary Material are available at http:\/\/mosta.molgen.mpg.de<\/jats:p>\n               <jats:p>Contact: \u00a0utz.pape@molgen.mpg.de<\/jats:p>\n               <jats:p>Supplementary information: Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btm610","type":"journal-article","created":{"date-parts":[[2008,1,4]],"date-time":"2008-01-04T01:13:32Z","timestamp":1199409212000},"page":"350-357","source":"Crossref","is-referenced-by-count":41,"title":["Natural similarity measures between position frequency matrices with an application to clustering"],"prefix":"10.1093","volume":"24","author":[{"given":"Utz J.","family":"Pape","sequence":"first","affiliation":[{"name":"1 Computational Biology, Max Planck Institute f. Molecular Genetics, Ihnestr. 73, 14195 Berlin, 2Mathematics and Computer Science, Free University of Berlin, Takustr. 9, 14195 Berlin, 3COMET group, Genome Informatics, Universit\u00e4t Bielefeld, 33594 Bielefeld and 4Bioinformatics for High-Throughput Technologies, Computer Science 11, Dortmund University, 44221 Dortmund, Germany"},{"name":"1 Computational Biology, Max Planck Institute f. Molecular Genetics, Ihnestr. 73, 14195 Berlin, 2Mathematics and Computer Science, Free University of Berlin, Takustr. 9, 14195 Berlin, 3COMET group, Genome Informatics, Universit\u00e4t Bielefeld, 33594 Bielefeld and 4Bioinformatics for High-Throughput Technologies, Computer Science 11, Dortmund University, 44221 Dortmund, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sven","family":"Rahmann","sequence":"additional","affiliation":[{"name":"1 Computational Biology, Max Planck Institute f. Molecular Genetics, Ihnestr. 73, 14195 Berlin, 2Mathematics and Computer Science, Free University of Berlin, Takustr. 9, 14195 Berlin, 3COMET group, Genome Informatics, Universit\u00e4t Bielefeld, 33594 Bielefeld and 4Bioinformatics for High-Throughput Technologies, Computer Science 11, Dortmund University, 44221 Dortmund, Germany"},{"name":"1 Computational Biology, Max Planck Institute f. Molecular Genetics, Ihnestr. 73, 14195 Berlin, 2Mathematics and Computer Science, Free University of Berlin, Takustr. 9, 14195 Berlin, 3COMET group, Genome Informatics, Universit\u00e4t Bielefeld, 33594 Bielefeld and 4Bioinformatics for High-Throughput Technologies, Computer Science 11, Dortmund University, 44221 Dortmund, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Martin","family":"Vingron","sequence":"additional","affiliation":[{"name":"1 Computational Biology, Max Planck Institute f. Molecular Genetics, Ihnestr. 73, 14195 Berlin, 2Mathematics and Computer Science, Free University of Berlin, Takustr. 9, 14195 Berlin, 3COMET group, Genome Informatics, Universit\u00e4t Bielefeld, 33594 Bielefeld and 4Bioinformatics for High-Throughput Technologies, Computer Science 11, Dortmund University, 44221 Dortmund, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2008,1,2]]},"reference":[{"key":"2023061011442128600_B1","doi-asserted-by":"crossref","first-page":"ii5","DOI":"10.1093\/bioinformatics\/btg1052","article-title":"Computational detection of cis -regulatory modules","volume":"19","author":"Aerts","year":"2003","journal-title":"Bioinformatics"},{"key":"2023061011442128600_B2","volume-title":"Mathematics, Statistics and Systems for Health.","author":"Bailey","year":"1977"},{"key":"2023061011442128600_B3","doi-asserted-by":"crossref","first-page":"389","DOI":"10.1186\/1471-2105-7-389","article-title":"Fast index based algorithms and software for matching position specific scoring matrices","volume":"7","author":"Beckstette","year":"2006","journal-title":"BMC Bioinformatics"},{"key":"2023061011442128600_B4","doi-asserted-by":"crossref","first-page":"723","DOI":"10.1016\/0022-2836(87)90354-8","article-title":"Selection of DNA binding sites by regulatory proteins. Statistical-mechanical theory and application to operators and promoters","volume":"193","author":"Berg","year":"1987","journal-title":"J. Mol. Biol"},{"key":"2023061011442128600_B5","doi-asserted-by":"crossref","first-page":"3797","DOI":"10.1073\/pnas.0308656100","article-title":"Local feature frequency profile: A method to measure structural similarity in proteins","volume":"101","author":"Choi","year":"2004","journal-title":"PNAS"},{"key":"2023061011442128600_B6","first-page":"431","article-title":"The statistical significance of nucleotide position-weight matrix matches","volume":"12","author":"Claverie","year":"1996","journal-title":"Comput. Appl. Biosci"},{"key":"2023061011442128600_B7","doi-asserted-by":"crossref","first-page":"1188","DOI":"10.1101\/gr.849004","article-title":"Weblogo: a sequence logo generator","volume":"14","author":"Crooks","year":"2004","journal-title":"Genome Res"},{"key":"2023061011442128600_B8","doi-asserted-by":"crossref","DOI":"10.1002\/0471445428","volume-title":"Statistical Methods for Rates and Proportions.","author":"Fleiss","year":"2003"},{"key":"2023061011442128600_B9","doi-asserted-by":"crossref","first-page":"R24","DOI":"10.1186\/gb-2007-8-2-r24","article-title":"Quantifying similarity between motifs","volume":"8","author":"Gupta","year":"2007","journal-title":"Genome Biol"},{"key":"2023061011442128600_B10","first-page":"81","article-title":"Identification of consensus patterns in unaligned DNA sequences known to be functionally related","volume":"6","author":"Hertz","year":"1990","journal-title":"Comput. Appl. Biosci"},{"key":"2023061011442128600_B11","doi-asserted-by":"crossref","first-page":"D447","DOI":"10.1093\/nar\/gki138","article-title":"Ensembl 2005","volume":"33","author":"Hubbard","year":"2005","journal-title":"Nucleic Acids Res"},{"key":"2023061011442128600_B12","doi-asserted-by":"crossref","first-page":"237","DOI":"10.1186\/1471-2105-6-237","article-title":"Measuring similarities between transcription factor binding sites","volume":"6","author":"Kielbasa","year":"2005","journal-title":"BMC Bioinformatics"},{"key":"2023061011442128600_B13","volume-title":"Information Theory and Statistics.","author":"Kullback","year":"1959"},{"key":"2023061011442128600_B14","article-title":"Bayesian Models for Multiple Local Sequence Alignment and Gibbs Sampling Strategies","volume":"95","author":"Liu","year":"1990","journal-title":"J. Am. Stat. Assoc"},{"key":"2023061011442128600_B15","doi-asserted-by":"crossref","first-page":"i283","DOI":"10.1093\/bioinformatics\/bti1025","article-title":"Improved detection of DNA motifs using a self-organized clustering of familial binding profiles","volume":"21","author":"Mahony","year":"2005","journal-title":"Bioinformatics"},{"key":"2023061011442128600_B16","doi-asserted-by":"crossref","first-page":"e61","DOI":"10.1371\/journal.pcbi.0030061","article-title":"DNA familial binding profiles made easy: comparison of various motif alignment and clustering strategies","volume":"3","author":"Mahony","year":"2007","journal-title":"PLoS Comput. Biol"},{"key":"2023061011442128600_B17","doi-asserted-by":"crossref","first-page":"374","DOI":"10.1093\/nar\/gkg108","article-title":"TRANSFAC(R): transcriptional regulation, from patterns to profiles","volume":"31","author":"Matys","year":"2003","journal-title":"Nucleic Acids Res"},{"key":"2023061011442128600_B18","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1093\/bioinformatics\/bti731","article-title":"Sequence features of DNA binding sites reveal structural class of associated transcription factor","volume":"22","author":"Narlikar","year":"2006","journal-title":"Bioinformatics"},{"key":"2023061011442128600_B19","first-page":"134","article-title":"A new statistical model to select target sequences bound by transcription factors","volume":"17","author":"Pape","year":"2006","journal-title":"Genome Informatics"},{"key":"2023061011442128600_B20","article-title":"Compound Poisson approximation of DNA motif counts on both strands","author":"Pape","year":"2007"},{"key":"2023061011442128600_B21","doi-asserted-by":"crossref","first-page":"3836","DOI":"10.1093\/nar\/24.19.3836","article-title":"Searching databases of conserved sequence regions by aligning protein multiple-alignments published erratum appears in nucleic acids res 1996 nov 1;24(21):4372","volume":"24","author":"Pietrokovski","year":"1996","journal-title":"Nucleic Acids Res"},{"key":"2023061011442128600_B22","first-page":"151","article-title":"Dynamic programming algorithms for two statistical problems in computational biology","author":"Rahmann","year":"2003"},{"key":"2023061011442128600_B23","doi-asserted-by":"crossref","DOI":"10.2202\/1544-6115.1032","article-title":"On the power of profiles for transcription factor binding site detection","volume":"2","author":"Rahmann","year":"2003","journal-title":"Stat. Appl. Genet. Mol. Biol"},{"key":"2023061011442128600_B24","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1089\/10665270050081360","article-title":"Probabilistic and statistical properties of words: an overview","volume":"7","author":"Reinert","year":"2000","journal-title":"J. Comput. Biol"},{"key":"2023061011442128600_B25","doi-asserted-by":"crossref","first-page":"W438","DOI":"10.1093\/nar\/gki590","article-title":"T-Reg Comparator: an analysis tool for the comparison of position weight matrices","volume":"33","author":"Roepcke","year":"2005","journal-title":"Nucleic Acids Res"},{"key":"2023061011442128600_B26","doi-asserted-by":"crossref","first-page":"207","DOI":"10.1016\/j.jmb.2004.02.048","article-title":"Constrained binding site diversity within families of transcription factors enhances pattern discovery bioinformatics","volume":"338","author":"Sandelin","year":"2004","journal-title":"J. Mol. Biol"},{"key":"2023061011442128600_B27","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1186\/1471-2164-5-99","article-title":"Arrays of ultraconserved non-coding regions span the loci of key developmental genes in vertebrate genomes","volume":"5","author":"Sandelin","year":"2004","journal-title":"BMC Genomics"},{"key":"2023061011442128600_B28","doi-asserted-by":"crossref","first-page":"415","DOI":"10.1016\/0022-2836(86)90165-8","article-title":"Information content of binding sites on nucleotide sequences","volume":"188","author":"Schneider","year":"1986","journal-title":"J. Mol. Biol"},{"key":"2023061011442128600_B29","doi-asserted-by":"crossref","first-page":"307","DOI":"10.1093\/bioinformatics\/bth480","article-title":"Similarity of position frequency matrices for transcription factor binding sites","volume":"21","author":"Schones","year":"2005","journal-title":"Bioinformatics"},{"issue":"1 Pt 2","key":"2023061011442128600_B30","doi-asserted-by":"crossref","first-page":"505","DOI":"10.1093\/nar\/12.1Part2.505","article-title":"Computer methods to locate signals in nucleic acid sequences","volume":"12","author":"Staden","year":"1984","journal-title":"Nucleic Acids Res"},{"key":"2023061011442128600_B31","first-page":"89","article-title":"Methods for calculating the probabilities of finding patterns in sequences","volume":"5","author":"Staden","year":"1989","journal-title":"Comput. Appl. Biosci"},{"key":"2023061011442128600_B32","doi-asserted-by":"crossref","first-page":"1183","DOI":"10.1073\/pnas.86.4.1183","article-title":"Identifying protein-binding sites from unaligned DNA fragments","volume":"86","author":"Stormo","year":"1989","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023061011442128600_B33","doi-asserted-by":"crossref","first-page":"16","DOI":"10.1093\/bioinformatics\/16.1.16","article-title":"DNA binding sites: representation and discovery","volume":"16","author":"Stormo","year":"2000","journal-title":"Bioinformatics"},{"key":"2023061011442128600_B34","doi-asserted-by":"crossref","first-page":"2997","DOI":"10.1093\/nar\/10.9.2997","article-title":"Use of the \u201cPerceptron\u201d algorithm to distinguish translational initiation sites in E. coli","volume":"10","author":"Stormo","year":"1982","journal-title":"Nucleic Acids Res"},{"key":"2023061011442128600_B35","doi-asserted-by":"crossref","first-page":"12357","DOI":"10.1073\/pnas.91.26.12357","article-title":"DNA recognition code of transcription factors in the helix-turn-helix, probe helix, hormone receptor, and zinc finger families","volume":"91","author":"Suzuki","year":"1994","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023061011442128600_B36","doi-asserted-by":"crossref","first-page":"137","DOI":"10.1038\/nbt1053","article-title":"Assessing computational tools for the discovery of transcription factor binding sites","volume":"23","author":"Tompa","year":"2005","journal-title":"Nat. Biotechnol"},{"key":"2023061011442128600_B37","doi-asserted-by":"crossref","first-page":"2369","DOI":"10.1093\/bioinformatics\/btg329","article-title":"Combining phylogenetic data with co-regulated genes to identify regulatory motifs","volume":"19","author":"Wang","year":"2003","journal-title":"Bioinformatics"},{"key":"2023061011442128600_B38","doi-asserted-by":"crossref","first-page":"276","DOI":"10.1038\/nrg1315","article-title":"Applied Bioinformatics for the Identification of Regulatory Elements","volume":"5","author":"Wasserman","year":"2004","journal-title":"Nat. Rev. Genet"},{"key":"2023061011442128600_B39","volume-title":"Introduction to Computational Biology.","author":"Waterman","year":"2000"},{"key":"2023061011442128600_B40","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1093\/bioinformatics\/16.3.233","article-title":"Fast probabilistic analysis of sequence function using scoring matrices","volume":"16","author":"Wu","year":"2000","journal-title":"Bioinformatics"},{"key":"2023061011442128600_B41","doi-asserted-by":"crossref","first-page":"531","DOI":"10.1093\/bioinformatics\/btl662","article-title":"Computing exact P-values for DNA motifs","volume":"23","author":"Zhang","year":"2007","journal-title":"Bioinformatics"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/3\/350\/50567983\/bioinformatics_24_3_350.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/3\/350\/50567983\/bioinformatics_24_3_350.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,10]],"date-time":"2023-06-10T11:45:56Z","timestamp":1686397556000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/24\/3\/350\/254270"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,1,2]]},"references-count":41,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2008,2,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btm610","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2008,2,1]]},"published":{"date-parts":[[2008,1,2]]}}}