{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T23:57:57Z","timestamp":1773273477130,"version":"3.50.1"},"reference-count":11,"publisher":"Oxford University Press (OUP)","issue":"7","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2005,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: The discovery of patterns shared by several sequences that differ greatly is a basic task in sequence analysis, and still a challenge. Several methods have been developed for detecting patterns. Methods commonly used for motif search include the Gibbs sampler, Expectation-Maximization (EM) algorithm and some intuitive greedy approaches. One cannot guarantee the optimality of the result produced by the Gibbs sampler in a single run. The deterministic EM methods tend to get trapped by local optima. Solutions found by greedy approaches are rarely sufficiently good.<\/jats:p><jats:p>Results: A simple model describing a motif or a portion of local multiple sequence alignment is the weight matrix model, in which a motif is characterized with position-specific probabilities. Two substitution matrices are proposed to relate the sequence similarity with the weight matrix. Combining the substitution matrix and weight matrix, we examine three typical sets of protein sequences with increasing complexity. At a low score threshold for pair similarity, sliding windows are compared with a seed window to find the score sum, which provides a measure of statistical significance for multiple sequence comparison. Such a similarity analysis reveals many aspects of motifs. Blocks determined by similarity can be used to deduce a primary weight matrix or an improved substitution matrix. The algorithm successfully obtains the optimal solution for the test sets by just greedy iteration.<\/jats:p><jats:p>Availability: Softwares and sequence datasets are available on request from the author.<\/jats:p><jats:p>Contact: \u00a0zheng@itp.ac.cn<\/jats:p>","DOI":"10.1093\/bioinformatics\/bti090","type":"journal-article","created":{"date-parts":[[2004,10,29]],"date-time":"2004-10-29T00:51:45Z","timestamp":1099011105000},"page":"938-943","source":"Crossref","is-referenced-by-count":5,"title":["Relation between weight matrix and substitution matrix: motif search by similarity"],"prefix":"10.1093","volume":"21","author":[{"given":"Wei-Mou","family":"Zheng","sequence":"first","affiliation":[{"name":"Institute of Theoretical Physics, Academia Sinica Beijing 100080, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2004,10,28]]},"reference":[{"key":"2023013107280922400_B1","unstructured":"Bailey, T. and Elkan, C. 1995Unsupervised learning of multiple motifs in biopolymers using expectation maximization. Machine Learning2151\u201380"},{"key":"2023013107280922400_B2","doi-asserted-by":"crossref","unstructured":"Boguski, M.S., Ostell, J., States, D.J. In Ries, A.R., Sternberg, M.J.E., Wetzel, R. (Eds.). Protein engineering: a practical approach1992, Oxford IRL Press, pp. 57\u201388","DOI":"10.1093\/oso\/9780199631391.003.0003"},{"key":"2023013107280922400_B3","doi-asserted-by":"crossref","unstructured":"Conlon, E.M., Liu, X.S., Lieb, J.D., Liu, J.S. 2003Integrating regulatory motif discovery and genome-wide expression analysis. Proc. Natl Acad. Sci., USA1003339\u20133344","DOI":"10.1073\/pnas.0630591100"},{"key":"2023013107280922400_B4","unstructured":"Eddy, S. 1995Multiple alignment using hidden Markov models. Proceedings of the International Conference on Intelligent Systems for Molecular BiologyJuly 16\u201319Cambridge, UK , Cambridge AAAI\/MIT Press, pp. 114\u2013120"},{"key":"2023013107280922400_B5","doi-asserted-by":"crossref","unstructured":"Hastings, W.K. 1970Monte Carlo sampling methods using Markov chains and their applications. Biometrika5797\u2013109","DOI":"10.1093\/biomet\/57.1.97"},{"key":"2023013107280922400_B6","unstructured":"Helden, J.V., Andre, B., Collado-Vides, J. 1998Extracting regulatory sites from the upstream region of yeast genes by computational analysis of oligonucleotide frequencies. J. Mol. Biol.281827\u2013842"},{"key":"2023013107280922400_B7","doi-asserted-by":"crossref","unstructured":"Henikoff, S. and Henikoff, J.G. 1992Amino acid substitution matrices from protein blocks. Proc. Natl Acad. Sci., USA8910915\u201310919","DOI":"10.1073\/pnas.89.22.10915"},{"key":"2023013107280922400_B8","doi-asserted-by":"crossref","unstructured":"Hertz, G., Hartzell, G., III, Stormo, G. 1990Identification of consensus patterns in unaligned dna sequences known to be functionally related. Comput. Appl. Biosci.681\u201392","DOI":"10.1093\/bioinformatics\/6.2.81"},{"key":"2023013107280922400_B9","doi-asserted-by":"crossref","unstructured":"Hughey, R. and Krogh, A. 1996Hidden Markov models for sequence analysis: extension and analysis of the basic method. Comput. Appl. Biosci.1295\u2013107","DOI":"10.1093\/bioinformatics\/12.2.95"},{"key":"2023013107280922400_B10","doi-asserted-by":"crossref","unstructured":"Lawrence, C.E., Altshul, S., Boguski, M., Liu, J., Neuwald, A., Wootton, J. 1993Detecting subtle sequence signals: a Gibbs sampling strategy for multiple alignments. Science262208\u2013214","DOI":"10.1126\/science.8211139"},{"key":"2023013107280922400_B11","doi-asserted-by":"crossref","unstructured":"Thompson, W., Rouchka, E.C., Lawrence, C.E. 2003Gibbs recursive sampler: finding transcription factor binding sites. Nucleic Acids Res.313580\u20133585","DOI":"10.1093\/nar\/gkg608"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/7\/938\/48966414\/bioinformatics_21_7_938.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/7\/938\/48966414\/bioinformatics_21_7_938.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,1,14]],"date-time":"2024-01-14T19:52:05Z","timestamp":1705261925000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/21\/7\/938\/268849"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2004,10,28]]},"references-count":11,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2005,4,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bti090","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2005,4,1]]},"published":{"date-parts":[[2004,10,28]]}}}