{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,8,5]],"date-time":"2024-08-05T09:15:03Z","timestamp":1722849303527},"reference-count":25,"publisher":"Oxford University Press (OUP)","issue":"19","license":[{"start":{"date-parts":[[2016,10,26]],"date-time":"2016-10-26T00:00:00Z","timestamp":1477440000000},"content-version":"vor","delay-in-days":138,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2016,10,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: Chromatin immunoprecipitation followed by high-throughput sequencing (ChIP-seq) is the standard method to investigate chromatin protein composition. As the number of community-available ChIP-seq profiles increases, it becomes more common to use data from different sources, which makes joint analysis challenging. Issues such as lack of reproducibility, heterogeneous quality and conflicts between replicates become evident when comparing datasets, especially when they are produced by different laboratories.<\/jats:p><jats:p>Results: Here, we present Zerone, a ChIP-seq discretizer with built-in quality control. Zerone is powered by a Hidden Markov Model with zero-inflated negative multinomial emissions, which allows it to merge several replicates into a single discretized profile. To identify low quality or irreproducible data, we trained a Support Vector Machine and integrated it as part of the discretization process. The result is a classifier reaching 95% accuracy in detecting low quality profiles. We also introduce a graphical representation to compare discretization quality and we show that Zerone achieves outstanding accuracy. Finally, on current hardware, Zerone discretizes a ChIP-seq experiment on mammalian genomes in about 5\u2002min using less than 700\u2002MB of memory.<\/jats:p><jats:p>Availability and Implementation: Zerone is available as a command line tool and as an R package. The C source code and R scripts can be downloaded from https:\/\/github.com\/nanakiksc\/zerone. The information to reproduce the benchmark and the figures is stored in a public Docker image that can be downloaded from https:\/\/hub.docker.com\/r\/nanakiksc\/zerone\/.<\/jats:p><jats:p>Contact: guillaume.filion@gmail.com<\/jats:p><jats:p>Supplementary information: Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btw336","type":"journal-article","created":{"date-parts":[[2016,6,11]],"date-time":"2016-06-11T03:54:37Z","timestamp":1465617277000},"page":"2896-2902","source":"Crossref","is-referenced-by-count":11,"title":["Zerone: a ChIP-seq discretizer for multiple replicates with built-in quality control"],"prefix":"10.1093","volume":"32","author":[{"given":"Pol","family":"Cusc\u00f3","sequence":"first","affiliation":[{"name":"1 Genome Architecture, Gene Regulation, Stem Cells and Cancer Programme, Centre for Genomic Regulation (CRG), the Barcelona Institute of Science and Technology, Barcelona 08003, Spain"},{"name":"2 Universitat Pompeu Fabra (UPF), Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guillaume J.","family":"Filion","sequence":"additional","affiliation":[{"name":"1 Genome Architecture, Gene Regulation, Stem Cells and Cancer Programme, Centre for Genomic Regulation (CRG), the Barcelona Institute of Science and Technology, Barcelona 08003, Spain"},{"name":"2 Universitat Pompeu Fabra (UPF), Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2016,6,10]]},"reference":[{"key":"2023020113475928300_btw336-B1","doi-asserted-by":"crossref","first-page":"W202","DOI":"10.1093\/nar\/gkp335","article-title":"MEME SUITE: tools for motif discovery and searching","volume":"37","author":"Bailey","year":"2009","journal-title":"Nucleic Acids Res"},{"key":"2023020113475928300_btw336-B2","doi-asserted-by":"crossref","first-page":"1554","DOI":"10.1214\/aoms\/1177699147","article-title":"Statistical inference for probabilistic functions of finite state Markov chains","volume":"37","author":"Baum","year":"1966","journal-title":"Ann. Math. Stat"},{"key":"2023020113475928300_btw336-B3","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1961189.1961199","article-title":"LIBSVM. A Library for Support Vector Machines","volume":"2","author":"Chang","year":"2011","journal-title":"ACM Trans. Intell. Syst. Technol. (TIST)"},{"key":"2023020113475928300_btw336-B4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1111\/j.2517-6161.1977.tb01600.x","article-title":"Maximum likelihood from incomplete data via the EM algorithm","volume":"39","author":"Dempster","year":"1977","journal-title":"J. R. Stat. Soc. Ser. B"},{"key":"2023020113475928300_btw336-B5","doi-asserted-by":"crossref","first-page":"1017","DOI":"10.1093\/bioinformatics\/btr064","article-title":"FIMO: scanning for occurrences of a given motif","volume":"27","author":"Grant","year":"2011","journal-title":"Bioinformatics"},{"key":"2023020113475928300_btw336-B6","doi-asserted-by":"crossref","first-page":"948","DOI":"10.1038\/nature06947","article-title":"Domain organization of human chromosomes revealed by mapping of nuclear lamina interactions","volume":"453","author":"Guelen","year":"2008","journal-title":"Nature"},{"key":"2023020113475928300_btw336-B7","doi-asserted-by":"crossref","first-page":"48","DOI":"10.1093\/bioinformatics\/btu568","article-title":"JAMM: a peak finder for joint analysis of NGS replicates","volume":"31","author":"Ibrahim","year":"2015","journal-title":"Bioinformatics"},{"key":"2023020113475928300_btw336-B8","volume-title":"Pscl: Classes and Methods for R Developed in the Political Science Computational Laboratory","author":"Jackman","year":"2015"},{"key":"2023020113475928300_btw336-B9","doi-asserted-by":"crossref","first-page":"D493","DOI":"10.1093\/nar\/gkh103","article-title":"The UCSC Table Browser data retrieval tool","volume":"32","author":"Karolchik","year":"2004","journal-title":"Nucleic Acids Res"},{"key":"2023020113475928300_btw336-B10","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1016\/j.cell.2006.12.048","article-title":"Analysis of the vertebrate insulator protein CTCF-binding sites in the human genome","volume":"128","author":"Kim","year":"2007","journal-title":"Cell"},{"key":"2023020113475928300_btw336-B11","doi-asserted-by":"crossref","first-page":"439","DOI":"10.1038\/jhg.2013.66","article-title":"Histone modifications for human epigenome analysis","volume":"58","author":"Kimura","year":"2013","journal-title":"J. Hum. Genet"},{"key":"2023020113475928300_btw336-B12","doi-asserted-by":"crossref","first-page":"1752","DOI":"10.1214\/11-AOAS466","article-title":"Measuring reproducibility of high-throughput experiments","volume":"5","author":"Li","year":"2011","journal-title":"Ann. Appl. Stat"},{"key":"2023020113475928300_btw336-B13","doi-asserted-by":"crossref","first-page":"1185","DOI":"10.1038\/nmeth.2221","article-title":"The GEM mapper: fast, accurate and versatile alignment by filtration","volume":"9","author":"Marco-Sola","year":"2012","journal-title":"Nat. Methods"},{"key":"2023020113475928300_btw336-B14","doi-asserted-by":"crossref","first-page":"D142","DOI":"10.1093\/nar\/gkt997","article-title":"JASPAR 2014: an extensively expanded and updated open-access database of transcription factor binding profiles","volume":"42","author":"Mathelier","year":"2014","journal-title":"Nucleic Acids Res"},{"key":"2023020113475928300_btw336-B15","author":"Meyer","year":"2014"},{"key":"2023020113475928300_btw336-B16","doi-asserted-by":"crossref","first-page":"e83506.","DOI":"10.1371\/journal.pone.0083506","article-title":"Widespread misinterpretable ChIP-seq bias in yeast","volume":"8","author":"Park","year":"2013","journal-title":"PLoS ONE"},{"key":"2023020113475928300_btw336-B17","doi-asserted-by":"crossref","first-page":"517","DOI":"10.1016\/j.cell.2005.06.026","article-title":"Genome-wide map of nucleosome acetylation and methylation in yeast","volume":"122","author":"Pokholok","year":"2005","journal-title":"Cell"},{"key":"2023020113475928300_btw336-B18","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1109\/5.18626","article-title":"A tutorial on hidden Markov models and selected applications in speech recognition","volume":"77","author":"Rabiner","year":"1989","journal-title":"Proc. IEEE"},{"key":"2023020113475928300_btw336-B19","doi-asserted-by":"crossref","first-page":"R67.","DOI":"10.1186\/gb-2011-12-7-r67","article-title":"ZINBA integrates local covariates with DNA-seq data to identify broad and narrow regions of enrichment, even within amplified genomic regions","volume":"12","author":"Rashid","year":"2011","journal-title":"Genome Biol"},{"key":"2023020113475928300_btw336-B20","doi-asserted-by":"crossref","first-page":"299.","DOI":"10.1186\/1471-2105-10-299","article-title":"BayesPeak: Bayesian analysis of ChIP-seq data","volume":"10","author":"Spyrou","year":"2009","journal-title":"BMC Bioinformatics"},{"key":"2023020113475928300_btw336-B21","doi-asserted-by":"crossref","first-page":"626","DOI":"10.1093\/bib\/bbq068","article-title":"Rapid innovation in ChIP-seq peak-calling algorithms is outdistancing benchmarking efforts","volume":"12","author":"Szalkowski","year":"2011","journal-title":"Brief. Bioinf"},{"key":"2023020113475928300_btw336-B22","doi-asserted-by":"crossref","first-page":"18602","DOI":"10.1073\/pnas.1316064110","article-title":"Highly expressed loci are vulnerable to misleading ChIP localization of multiple unrelated proteins","volume":"110","author":"Teytelman","year":"2013","journal-title":"Proc. Natl. Acad. Sci. U. S. A"},{"key":"2023020113475928300_btw336-B23","doi-asserted-by":"crossref","first-page":"260","DOI":"10.1109\/TIT.1967.1054010","article-title":"Error bounds for convolutional codes and an asymptotically optimum decoding algorithm","volume":"13","author":"Viterbi","year":"1967","journal-title":"IEEE Trans. Inf. Theory"},{"key":"2023020113475928300_btw336-B24","doi-asserted-by":"crossref","DOI":"10.18637\/jss.v027.i08","article-title":"Regression models for count data in R","volume":"27","author":"Zeileis","year":"2008","journal-title":"J. Stat. Softw"},{"key":"2023020113475928300_btw336-B25","doi-asserted-by":"crossref","first-page":"R137.","DOI":"10.1186\/gb-2008-9-9-r137","article-title":"Model-based analysis of ChIP-Seq (MACS)","volume":"9","author":"Zhang","year":"2008","journal-title":"Genome Biol"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/32\/19\/2896\/49021909\/bioinformatics_32_19_2896.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/32\/19\/2896\/49021909\/bioinformatics_32_19_2896.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,17]],"date-time":"2024-06-17T15:03:53Z","timestamp":1718636633000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/32\/19\/2896\/2196396"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,6,10]]},"references-count":25,"journal-issue":{"issue":"19","published-print":{"date-parts":[[2016,10,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btw336","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2016,10,1]]},"published":{"date-parts":[[2016,6,10]]}}}