{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,1,26]],"date-time":"2023-01-26T05:19:39Z","timestamp":1674710379220},"reference-count":15,"publisher":"Oxford University Press (OUP)","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2010,2,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Identification of motifs in biological sequences is a challenging problem because such motifs are often short, degenerate, and may contain gaps. Most algorithms that have been developed for motif-finding use the expectation-maximization (EM) algorithm iteratively. Although EM algorithms can converge quickly, they depend strongly on initialization parameters and can converge to local sub-optimal solutions. In addition, they cannot generate gapped motifs. The effectiveness of EM algorithms in motif finding can be improved by incorporating methods that choose different sets of initial parameters to enable escape from local optima, and that allow gapped alignments within motif models.<\/jats:p>\n               <jats:p>Results: We have developed HIGEDA, an algorithm that uses the hierarchical gene-set genetic algorithm (HGA) with EM to initiate and search for the best parameters for the motif model. In addition, HIGEDA can identify gapped motifs using a position weight matrix and dynamic programming to generate an optimal gapped alignment of the motif model with sequences from the dataset. We show that HIGEDA outperforms MEME and other motif-finding algorithms on both DNA and protein sequences.<\/jats:p>\n               <jats:p>Availability and implementation: Source code and test datasets are available for download at http:\/\/ouray.cudenver.edu\/\u223ctnle\/, implemented in C++ and supported on Linux and MS Windows.<\/jats:p>\n               <jats:p>Contact: \u00a0katheleen.gardiner@ucdenver.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/btp676","type":"journal-article","created":{"date-parts":[[2009,12,9]],"date-time":"2009-12-09T03:53:26Z","timestamp":1260330806000},"page":"302-309","source":"Crossref","is-referenced-by-count":6,"title":["HIGEDA: a hierarchical gene-set genetics based algorithm for finding subtle motifs in biological sequences"],"prefix":"10.1093","volume":"26","author":[{"given":"Thanh","family":"Le","sequence":"first","affiliation":[{"name":"1 Department of Computer Science and Engineering, 2 Department of Pediatrics, Computational Biosciences, Human Medical Genetics and Neuroscience Programs, University of Colorado, Denver, CO, USA"}]},{"given":"Tom","family":"Altman","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science and Engineering, 2 Department of Pediatrics, Computational Biosciences, Human Medical Genetics and Neuroscience Programs, University of Colorado, Denver, CO, USA"}]},{"given":"Katheleen","family":"Gardiner","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science and Engineering, 2 Department of Pediatrics, Computational Biosciences, Human Medical Genetics and Neuroscience Programs, University of Colorado, Denver, CO, USA"}]}],"member":"286","published-online":{"date-parts":[[2009,12,8]]},"reference":[{"key":"2023012511004295400_B1","first-page":"21","article-title":"The value of prior knowledge in discovering motifs with MEME","volume":"3","author":"Bailey","year":"1995","journal-title":"Proc. Intl. Conf. Intel. Syst. Mol. Biol."},{"key":"2023012511004295400_B2","first-page":"275","article-title":"A genetic-based EM motif-finding algorithm for biological sequence analysis","author":"Bi","year":"2007","journal-title":"Proc. IEEE Symp. Comput. Intel. Bioinfo. Comput. Biol."},{"key":"2023012511004295400_B3","first-page":"1","article-title":"Prediction of transcription factor binding sites using genetic algorithm","author":"Chang","year":"2006","journal-title":"1st Conf. Ind. Elec. Apps."},{"key":"2023012511004295400_B4","doi-asserted-by":"crossref","first-page":"e1000071","DOI":"10.1371\/journal.pcbi.1000071","article-title":"Discovering sequence motifs with arbitrary insertions and deletions","volume":"4","author":"Frith","year":"2008","journal-title":"PLoS Comput. Biol."},{"key":"2023012511004295400_B5","first-page":"135","article-title":"Using substitution probabilities to improve position-specific scoring matrices","volume":"12","author":"Henikoff","year":"1996","journal-title":"Comp. App. Biosci."},{"key":"2023012511004295400_B6","first-page":"67","article-title":"A hierarchical gene-set genetic algorithm","volume":"3","author":"Hong","year":"2008","journal-title":"J. Comp."},{"key":"2023012511004295400_B7","doi-asserted-by":"crossref","first-page":"1188","DOI":"10.1093\/bioinformatics\/btm080","article-title":"GAPWM: a genetic algorithm method for optimizing a position weight matrix","volume":"23","author":"Li","year":"2007","journal-title":"Bioinformatics"},{"key":"2023012511004295400_B8","doi-asserted-by":"crossref","first-page":"629","DOI":"10.1093\/bioinformatics\/btn009","article-title":"fdrMotif: identifying cis-elements by an EM algorithm coupled with false discovery rate control","volume":"24","author":"Li","year":"2008","journal-title":"Bioinformatics"},{"key":"2023012511004295400_B9","doi-asserted-by":"crossref","first-page":"919","DOI":"10.1109\/TNN.2006.875987","article-title":"Motif discoveries in unaligned molecular sequences using self-organizing neural networks","volume":"17","author":"Liu","year":"2006","journal-title":"IEEE Trans. Neural Networks"},{"key":"2023012511004295400_B10","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1016\/j.jbi.2006.07.001","article-title":"A new approach to the assessment of the quality of predictions of transcription factor binding sites","volume":"40","author":"Nowakowski","year":"2007","journal-title":"J. Biomed. Info."},{"key":"2023012511004295400_B11","doi-asserted-by":"crossref","first-page":"3516","DOI":"10.1093\/bioinformatics\/bth438","article-title":"Comparative analysis of methods for representing and searching for transcription factor binding sites","volume":"20","author":"Osada","year":"2004","journal-title":"Bioinformatics"},{"key":"2023012511004295400_B12","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1109\/TCBB.2005.5","article-title":"Bases of motifs for generating repeated patterns with wildcards","volume":"2","author":"Pisanti","year":"2005","journal-title":"IEEE\/ACM Trans. Comput. Biol. and Bioinfo."},{"key":"2023012511004295400_B13","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1186\/1748-7188-2-15","article-title":"Efficient and accurate P-value computation for position weight matrices","volume":"2","author":"Touzet","year":"2007","journal-title":"Algorithms for Mol. Biol."},{"key":"2023012511004295400_B14","doi-asserted-by":"crossref","first-page":"1577","DOI":"10.1093\/bioinformatics\/btl147","article-title":"GAME: Detecting cis-regulatory elements using a genetic algorithm","volume":"22","author":"Wei","year":"2006","journal-title":"Bioninformatics"},{"key":"2023012511004295400_B15","doi-asserted-by":"crossref","first-page":"409","DOI":"10.1198\/016214504000000377","article-title":"A Bayesian insertion\/deletion algorithm for distant protein motif searching via entropy filtering","volume":"99","author":"Xie","year":"2004","journal-title":"J. Am. Stat. Assoc."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/26\/3\/302\/48860660\/bioinformatics_26_3_302.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/26\/3\/302\/48860660\/bioinformatics_26_3_302.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,25]],"date-time":"2023-01-25T11:03:10Z","timestamp":1674644590000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/26\/3\/302\/215101"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,12,8]]},"references-count":15,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2010,2,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btp676","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2010,2,1]]},"published":{"date-parts":[[2009,12,8]]}}}