{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,31]],"date-time":"2025-10-31T14:05:21Z","timestamp":1761919521954},"reference-count":21,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Algorithms Mol Biol"],"published-print":{"date-parts":[[2012,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>High-throughput sequencing is becoming the standard tool for investigating protein-DNA interactions or epigenetic modifications. However, the data generated will always contain noise due to e.g. repetitive regions or non-specific antibody interactions. The noise will appear in the form of a background distribution of reads that must be taken into account in the downstream analysis, for example when detecting enriched regions (peak-calling). Several reported peak-callers can take experimental measurements of background tag distribution into account when analysing a data set. Unfortunately, the background is only used to adjust peak calling and not as a pre-processing step that aims at discerning the signal from the background noise. A normalization procedure that extracts the signal of interest would be of universal use when investigating genomic patterns.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>We formulated such a normalization method based on linear regression and made a proof-of-concept implementation in R and C++. It was tested on simulated as well as on publicly available ChIP-seq data on binding sites for two transcription factors, MAX and FOXA1 and two control samples, Input and IgG. We applied three different peak-callers to (i) raw (un-normalized) data using statistical background models and (ii) raw data with control samples as background and (iii) normalized data without additional control samples as background. The fraction of called regions containing the expected transcription factor binding motif was largest for the normalized data and evaluation with qPCR data for FOXA1 suggested higher sensitivity and specificity using normalized data over raw data with experimental background.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusions<\/jats:title>\n            <jats:p>The proposed method can handle several control samples allowing for correction of multiple sources of bias simultaneously. Our evaluation on both synthetic and experimental data suggests that the method is successful in removing background noise.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1748-7188-7-2","type":"journal-article","created":{"date-parts":[[2012,1,17]],"date-time":"2012-01-17T07:34:11Z","timestamp":1326785651000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["A strand specific high resolution normalization method for chip-sequencing data employing multiple experimental control measurements"],"prefix":"10.1186","volume":"7","author":[{"given":"Stefan","family":"Enroth","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Claes R","family":"Andersson","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Robin","family":"Andersson","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Claes","family":"Wadelius","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mats G","family":"Gustafsson","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jan","family":"Komorowski","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2012,1,16]]},"reference":[{"key":"138_CR1","doi-asserted-by":"publisher","first-page":"1497","DOI":"10.1126\/science.1141319","volume":"316","author":"DS Johnson","year":"2007","unstructured":"Johnson DS, Mortazavi A, Myers RM, Wold B: Genome-wide mapping of in vivo protein-DNA interactions. Science. 2007, 316: 1497-1502. 10.1126\/science.1141319","journal-title":"Science"},{"key":"138_CR2","doi-asserted-by":"publisher","first-page":"1351","DOI":"10.1038\/nbt.1508","volume":"26","author":"PV Kharchenko","year":"2008","unstructured":"Kharchenko PV, Tolstorukov MY, Park PJ: Design and analysis of ChIP-seq experiments for DNA-binding proteins. Nat Biotechnol. 2008, 26: 1351-1359. 10.1038\/nbt.1508","journal-title":"Nat Biotechnol"},{"key":"138_CR3","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1677\/JOE-08-0526","volume":"201","author":"BG Hoffman","year":"2009","unstructured":"Hoffman BG, Jones SJ: Genome-wide identification of DNA-protein interactions using chromatin immunoprecipitation coupled with flow cell sequencing. J Endocrinol. 2009, 201: 1-13. 10.1677\/JOE-08-0526","journal-title":"J Endocrinol"},{"key":"138_CR4","doi-asserted-by":"publisher","first-page":"618","DOI":"10.1186\/1471-2164-10-618","volume":"10","author":"TD Laajala","year":"2009","unstructured":"Laajala TD, Raghav S, Tuomela S, Lahesmaa R, Aittokallio T, Elo LL: A practical comparison of methods for detecting transcription factor binding sites in ChIP-seq experiments. BMC Genomics. 2009, 10: 618- 10.1186\/1471-2164-10-618","journal-title":"BMC Genomics"},{"key":"138_CR5","doi-asserted-by":"publisher","first-page":"2334","DOI":"10.1093\/bioinformatics\/btp384","volume":"25","author":"C Taslim","year":"2009","unstructured":"Taslim C, Wu J, Yan P, Singer G, Parvin J, Huang T, Lin S, Huang K: Comparative study on ChIP-seq data: normalization and binding pattern characterization. Bioinformatics. 2009, 25: 2334-2340. 10.1093\/bioinformatics\/btp384","journal-title":"Bioinformatics"},{"key":"138_CR6","doi-asserted-by":"publisher","first-page":"651","DOI":"10.1038\/nmeth1068","volume":"4","author":"G Robertson","year":"2007","unstructured":"Robertson G, Hirst M, Bainbridge M, Bilenky M, Zhao Y, Zeng T, Euskirchen G, Bernier B, Varhol R, Delaney A: Genome-wide profiles of STAT1 DNA association using chromatin immunoprecipitation and massively parallel sequencing. Nat Methods. 2007, 4: 651-657. 10.1038\/nmeth1068","journal-title":"Nat Methods"},{"key":"138_CR7","unstructured":"ENCODE Data Coordination Center at UCSC, Yale data.http:\/\/hgdownload.cse.ucsc.edu\/goldenPath\/hg18\/encodeDCC\/wgEncodeYaleChIPseq\/"},{"key":"138_CR8","doi-asserted-by":"publisher","first-page":"799","DOI":"10.1038\/nature05874","volume":"447","author":"E Birney","year":"2007","unstructured":"Birney E, Stamatoyannopoulos JA, Dutta A, Guigo R, Gingeras TR, Margulies EH, Weng Z, Snyder M, Dermitzakis ET, Thurman RE: Identification and analysis of functional elements in 1% of the human genome by the ENCODE pilot project. Nature. 2007, 447: 799-816. 10.1038\/nature05874","journal-title":"Nature"},{"key":"138_CR9","doi-asserted-by":"publisher","first-page":"R129","DOI":"10.1186\/gb-2009-10-11-r129","volume":"10","author":"M Motallebipour","year":"2009","unstructured":"Motallebipour M, Ameur A, Reddy Bysani MS, Patra K, Wallerman O, Mangion J, Barker MA, McKernan KJ, Komorowski J, Wadelius C: Differential binding and co-binding pattern of FOXA1 and FOXA3 and their relation to H3K4me3 in HepG2 cells revealed by ChIP-seq. Genome Biol. 2009, 10: R129- 10.1186\/gb-2009-10-11-r129","journal-title":"Genome Biol"},{"key":"138_CR10","doi-asserted-by":"publisher","first-page":"5221","DOI":"10.1093\/nar\/gkn488","volume":"36","author":"R Jothi","year":"2008","unstructured":"Jothi R, Cuddapah S, Barski A, Cui K, Zhao K: Genome-wide identification of in vivo protein-DNA binding sites from ChIP-Seq data. Nucleic Acids Res. 2008, 36: 5221-5231. 10.1093\/nar\/gkn488","journal-title":"Nucleic Acids Res"},{"key":"138_CR11","doi-asserted-by":"publisher","first-page":"1729","DOI":"10.1093\/bioinformatics\/btn305","volume":"24","author":"AP Fejes","year":"2008","unstructured":"Fejes AP, Robertson G, Bilenky M, Varhol R, Bainbridge M, Jones SJ: FindPeaks 3.1: a tool for identifying areas of enrichment from massively parallel short-read sequencing technology. Bioinformatics. 2008, 24: 1729-1730. 10.1093\/bioinformatics\/btn305","journal-title":"Bioinformatics"},{"key":"138_CR12","unstructured":"Findpeaks 4.0.http:\/\/sourceforge.net\/apps\/mediawiki\/vancouvershortr\/index.php?title=FindPeaks#FindPeaks_4.0"},{"key":"138_CR13","doi-asserted-by":"publisher","first-page":"R137","DOI":"10.1186\/gb-2008-9-9-r137","volume":"9","author":"Y Zhang","year":"2008","unstructured":"Zhang Y, Liu T, Meyer CA, Eeckhoute J, Johnson DS, Bernstein BE, Nusbaum C, Myers RM, Brown M, Li W, Liu XS: Model-based analysis of ChIP-Seq (MACS). Genome Biol. 2008, 9: R137- 10.1186\/gb-2008-9-9-r137","journal-title":"Genome Biol"},{"key":"138_CR14","doi-asserted-by":"publisher","first-page":"1231","DOI":"10.1093\/bioinformatics\/btp152","volume":"25","author":"S Anders","year":"2009","unstructured":"Anders S: Visualization of genomic data with the Hilbert curve. Bioinformatics. 2009, 25: 1231-1235. 10.1093\/bioinformatics\/btp152","journal-title":"Bioinformatics"},{"key":"138_CR15","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/S0378-1119(01)00697-7","volume":"277","author":"B Luscher","year":"2001","unstructured":"Luscher B: Function and regulation of the transcription factors of the Myc\/Max\/Mad network. Gene. 2001, 277: 1-14. 10.1016\/S0378-1119(01)00697-7","journal-title":"Gene"},{"key":"138_CR16","doi-asserted-by":"publisher","first-page":"D620","DOI":"10.1093\/nar\/gkp961","volume":"38","author":"KR Rosenbloom","year":"2010","unstructured":"Rosenbloom KR, Dreszer TR, Pheasant M, Barber GP, Meyer LR, Pohl A, Raney BJ, Wang T, Hinrichs AS, Zweig AS: ENCODE whole-genome data in the UCSC Genome Browser. Nucleic Acids Res. 2010, 38: D620-625. 10.1093\/nar\/gkp961","journal-title":"Nucleic Acids Res"},{"key":"138_CR17","doi-asserted-by":"publisher","first-page":"1692","DOI":"10.1109\/PROC.1975.10036","volume":"63","author":"B Widrow","year":"1975","unstructured":"Widrow B, Glover JR, McCool JM, Kaunitz J, Williams CS, Hearn RH, Zeidler JR, Dong E, Goodlin RC: ADAPTIVE NOISE CANCELLING - PRINCIPLES AND APPLICATIONS. Proc IEEE. 1975, 63: 1692-1716.","journal-title":"Proc IEEE"},{"key":"138_CR18","unstructured":"R Development Core Team. R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. ISBN 3-900051-07-0. 2009,  http:\/\/www.R-project.org"},{"key":"138_CR19","doi-asserted-by":"publisher","first-page":"841","DOI":"10.1093\/bioinformatics\/btq033","volume":"26","author":"AR Quinlan","year":"2010","unstructured":"Quinlan AR, Hall IM: BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010, 26: 841-842. 10.1093\/bioinformatics\/btq033","journal-title":"Bioinformatics"},{"key":"138_CR20","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1186\/1756-0381-3-4","volume":"3","author":"S Enroth","year":"2010","unstructured":"Enroth S, Andersson R, Wadelius C, Komorowski J: SICTIN: Rapid footprinting of massively parallel sequencing data. BioData Min. 2010, 3: 4- 10.1186\/1756-0381-3-4","journal-title":"BioData Min"},{"key":"138_CR21","unstructured":"M Galassi JD, Theiler J, Gough B, Jungman G, Alken P, Booth M, Rossi F: GNU Scientific Library Reference Manual. 3"}],"container-title":["Algorithms for Molecular Biology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1748-7188-7-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T18:18:14Z","timestamp":1630520294000},"score":1,"resource":{"primary":{"URL":"https:\/\/almob.biomedcentral.com\/articles\/10.1186\/1748-7188-7-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,1,16]]},"references-count":21,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2012,12]]}},"alternative-id":["138"],"URL":"https:\/\/doi.org\/10.1186\/1748-7188-7-2","relation":{},"ISSN":["1748-7188"],"issn-type":[{"value":"1748-7188","type":"electronic"}],"subject":[],"published":{"date-parts":[[2012,1,16]]},"assertion":[{"value":"18 August 2011","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 January 2012","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 January 2012","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"2"}}