{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,1,29]],"date-time":"2025-01-29T05:49:33Z","timestamp":1738129773544,"version":"3.33.0"},"reference-count":28,"publisher":"Oxford University Press (OUP)","issue":"5","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2008,3,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: Most de novo motif identification methods optimize the motif model first and then separately test the statistical significance of the motif score. In the first stage, a motif abundance parameter needs to be specified or modeled. In the second stage, a Z-score or P-value is used as the test statistic. Error rates under multiple comparisons are not fully considered.<\/jats:p><jats:p>Methodology: We propose a simple but novel approach, fdrMotif, that selects as many binding sites as possible while controlling a user-specified false discovery rate (FDR). Unlike existing iterative methods, fdrMotif combines model optimization [e.g. position weight matrix (PWM)] and significance testing at each step. By monitoring the proportion of binding sites selected in many sets of background sequences, fdrMotif controls the FDR in the original data. The model is then updated using an expectation (E)- and maximization (M)-like procedure. We propose a new normalization procedure in the E-step for updating the model. This process is repeated until either the model converges or the number of iterations exceeds a maximum.<\/jats:p><jats:p>Results: Simulation studies suggest that our normalization procedure assigns larger weights to the binding sites than do two other commonly used normalization procedures. Furthermore, fdrMotif requires only a user-specified FDR and an initial PWM. When tested on 542 high confidence experimental p53 binding loci, fdrMotif identified 569 p53 binding sites in 505 (93.2%) sequences. In comparison, MEME identified more binding sites but in fewer ChIP sequences than fdrMotif. When tested on 500 sets of simulated \u2018ChIP\u2019 sequences with embedded known p53 binding sites, fdrMotif, compared to MEME, has higher sensitivity with similar positive predictive value. Furthermore, fdrMotif is robust to noise: it selected nearly identical binding sites in data adulterated with 50% added background sequences and the unadulterated data. We suggest that fdrMotif represents an improvement over MEME.<\/jats:p><jats:p>Availability: C code can be found at: http:\/\/www.niehs.nih.gov\/research\/resources\/software\/fdrMotif\/<\/jats:p><jats:p>Contact: \u00a0li3@niehs.nih.gov<\/jats:p><jats:p>Supplementary information: Supplementary data are available at http:\/\/www.niehs.nih.gov\/research\/resources\/software\/fdrMotif\/<\/jats:p>","DOI":"10.1093\/bioinformatics\/btn009","type":"journal-article","created":{"date-parts":[[2008,2,23]],"date-time":"2008-02-23T01:34:37Z","timestamp":1203730477000},"page":"629-636","source":"Crossref","is-referenced-by-count":6,"title":["fdrMotif: identifying<i>cis<\/i>-elements by an EM algorithm coupled with false discovery rate control"],"prefix":"10.1093","volume":"24","author":[{"given":"Leping","family":"Li","sequence":"first","affiliation":[{"name":"1 Biostatistics Branch and 2Computational Biology Facility, National Institute of Environmental Health Sciences, NIH, DHHS, Research Triangle Park, NC 27709, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Robert L.","family":"Bass","sequence":"additional","affiliation":[{"name":"1 Biostatistics Branch and 2Computational Biology Facility, National Institute of Environmental Health Sciences, NIH, DHHS, Research Triangle Park, NC 27709, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu","family":"Liang","sequence":"additional","affiliation":[{"name":"1 Biostatistics Branch and 2Computational Biology Facility, National Institute of Environmental Health Sciences, NIH, DHHS, Research Triangle Park, NC 27709, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2008,3,1]]},"reference":[{"key":"2023020210111801800_B1","first-page":"28","article-title":"Fitting a mixture model by expectation maximization to discover motifs in biopolymers","volume":"2","author":"Bailey","year":"1994","journal-title":"Proc. Int. Conf. Intell. Syst. Mol. Bol"},{"key":"2023020210111801800_B2","doi-asserted-by":"crossref","first-page":"289","DOI":"10.1111\/j.2517-6161.1995.tb02031.x","article-title":"Controlling the false discovery rate: a practical and powerful approach to multiple testing","volume":"57","author":"Benjamini","year":"1995","journal-title":"J. R. Stat. Soc. Ser. B"},{"key":"2023020210111801800_B3","doi-asserted-by":"crossref","first-page":"1165","DOI":"10.1214\/aos\/1013699998","article-title":"The control of the false discovery rate in multiple testing under dependency","volume":"29","author":"Benjamini","year":"2001","journal-title":"Ann. Stat"},{"key":"2023020210111801800_B4","doi-asserted-by":"crossref","first-page":"1188","DOI":"10.1101\/gr.849004","article-title":"WebLogo: a sequence logo generator","volume":"14","author":"Crooks","year":"2004","journal-title":"Genome Res"},{"key":"2023020210111801800_B5","doi-asserted-by":"crossref","first-page":"663","DOI":"10.1093\/bioinformatics\/15.7.563","article-title":"Identifying DNA and protein patterns with statistically significant alignments of multiple sequences","volume":"15","author":"Hertz","year":"1999","journal-title":"Bioinformatics"},{"key":"2023020210111801800_B6","doi-asserted-by":"crossref","first-page":"1284","DOI":"10.1371\/journal.pgen.0030127","article-title":"Divergent evolution of human p53 binding sites: cell cycle versus apoptosis","volume":"3","author":"Horvath","year":"2007","journal-title":"PLoS Genet"},{"key":"2023020210111801800_B7","first-page":"188","article-title":"Computational discovery of gene regulatory binding motifs: a Bayesian perspective","volume":"18","author":"Jensen","year":"2004","journal-title":"Stat. Sci"},{"key":"2023020210111801800_B8","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1016\/j.cell.2006.12.048","article-title":"Analysis of the vertebrate insulator protein CTCF-binding sites in the human genome","volume":"128","author":"Kim","year":"2007","journal-title":"Cell"},{"key":"2023020210111801800_B9","doi-asserted-by":"crossref","first-page":"191","DOI":"10.1089\/cmb.1994.1.191","article-title":"TRANSFAC retrieval program: a network model database of eukaryotic transcription regulating sequences and proteins","volume":"1","author":"Knuppel","year":"1994","journal-title":"J. Comput. Biol"},{"key":"2023020210111801800_B10","doi-asserted-by":"crossref","first-page":"1188","DOI":"10.1093\/bioinformatics\/btm080","article-title":"GAPWM: GAPWM: a genetic algorithm method for optimizing a position weight matrix","volume":"23","author":"Li","year":"2007","journal-title":"Bioinformatics"},{"key":"2023020210111801800_B11","doi-asserted-by":"crossref","first-page":"867","DOI":"10.1371\/journal.pgen.0030087","article-title":"Whole-genome cartography of estrogen receptor alpha binding sites","volume":"3","author":"Lin","year":"2007","journal-title":"PLoS Genet"},{"key":"2023020210111801800_B12","doi-asserted-by":"crossref","first-page":"1156","DOI":"10.1080\/01621459.1995.10476622","article-title":"Bayesian models for multiple local sequence alignment and gibbs sampling strategies","volume":"90","author":"Liu","year":"1995","journal-title":"J. Am. Stat. Assoc"},{"key":"2023020210111801800_B13","first-page":"127","article-title":"BioProspector: discovering conserved DNA motifs in upstream regulatory regions of co-expressed genes","volume":"6","author":"Liu","year":"2001","journal-title":"Pac. Symp. Biocomput"},{"key":"2023020210111801800_B14","doi-asserted-by":"crossref","first-page":"835","DOI":"10.1038\/nbt717","article-title":"An algorithm for finding protein-DNA binding sites with applications to chromatin-immunoprecipitation microarry experiments","volume":"20","author":"Liu","year":"2002","journal-title":"Nat. Biotechnol"},{"key":"2023020210111801800_B15","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1198\/004017005000000319","article-title":"Tuning variable selection procedures by adding noise","volume":"48","author":"Luo","year":"2004","journal-title":"Technometrics"},{"key":"2023020210111801800_B16","doi-asserted-by":"crossref","DOI":"10.1201\/9781420035933","volume-title":"Subset Selection in Regression.","author":"Miller","year":"2002"},{"key":"2023020210111801800_B17","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1137\/1026034","article-title":"Mixture densities maximum likelihood and EM algorithm","volume":"26","author":"Redner","year":"1984","journal-title":"SIAM Rev"},{"key":"2023020210111801800_B18","doi-asserted-by":"crossref","first-page":"939","DOI":"10.1038\/nbt1098-939","article-title":"Finding DNA regulatory motifs within unaligned noncoding sequences clustered by whole-genome mRNA quantitation","volume":"16","author":"Roth","year":"1998","journal-title":"Nat. Biotechnol"},{"key":"2023020210111801800_B19","doi-asserted-by":"crossref","first-page":"D91","DOI":"10.1093\/nar\/gkh012","article-title":"JASPAR: an open-access database for eukaryotic transcription factor binding profiles","volume":"32","author":"Sandelin","year":"2004","journal-title":"Nucleic Acids Res"},{"key":"2023020210111801800_B20","doi-asserted-by":"crossref","first-page":"1560","DOI":"10.1073\/pnas.0406123102","article-title":"Identifying tissue-selective transcription factor binding sites in vertebrate promoters","volume":"102","author":"Smith","year":"2005","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023020210111801800_B21","doi-asserted-by":"crossref","first-page":"479","DOI":"10.1111\/1467-9868.00346","article-title":"A direct approach to false discovery rate","volume":"64","author":"Storey","year":"2002","journal-title":"J. R. Stat. Soc. Ser. B"},{"key":"2023020210111801800_B22","article-title":"Estimating the positive false discovery rates under dependence, with applications to DNA microarrays","volume-title":"Technical Report.","author":"Storey","year":"2001"},{"key":"2023020210111801800_B23","doi-asserted-by":"crossref","first-page":"137","DOI":"10.1038\/nbt1053","article-title":"Assessing computational tools for the discovery of transcription factor binding sites","volume":"23","author":"Tompa","year":"2005","journal-title":"Nat. Biotechnol"},{"key":"2023020210111801800_B24","doi-asserted-by":"crossref","first-page":"1113","DOI":"10.1093\/bioinformatics\/17.12.1113","article-title":"A higher order background model improves the detection of promoter regulatory elements by Gibbs sampling","volume":"17","author":"Thijs","year":"2001","journal-title":"Bioinformatics"},{"key":"2023020210111801800_B25","doi-asserted-by":"crossref","first-page":"1071","DOI":"10.1111\/j.0006-341X.2003.00123.x","article-title":"Estimation of false discovery rates in multiple testing application to gene microarray data","volume":"59","author":"Tsai","year":"2003","journal-title":"Biometrics"},{"key":"2023020210111801800_B26","doi-asserted-by":"crossref","first-page":"207","DOI":"10.1016\/j.cell.2005.10.043","article-title":"A global map of p53 transcription-factor binding sites in the human genome","volume":"124","author":"Wei","year":"2006","journal-title":"Cell"},{"key":"2023020210111801800_B27","doi-asserted-by":"crossref","first-page":"235","DOI":"10.1198\/016214506000000843","article-title":"Controlling variable selection by the addition of pseudo variables","volume":"102","author":"Wu","year":"2007","journal-title":"J. Am. Stat. Assoc"},{"key":"2023020210111801800_B28","doi-asserted-by":"crossref","first-page":"170","DOI":"10.1002\/gepi.0042","article-title":"Truncated product method for combining P-values","volume":"22","author":"Zaykin","year":"2002","journal-title":"Genet. Epidemiol"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/5\/629\/49052842\/bioinformatics_24_5_629.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/5\/629\/49052842\/bioinformatics_24_5_629.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,28]],"date-time":"2025-01-28T18:18:09Z","timestamp":1738088289000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/24\/5\/629\/202275"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,3,1]]},"references-count":28,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2008,3,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btn009","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"type":"electronic","value":"1367-4811"},{"type":"print","value":"1367-4803"}],"subject":[],"published-other":{"date-parts":[[2008,3,1]]},"published":{"date-parts":[[2008,3,1]]}}}