{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,9]],"date-time":"2026-03-09T19:55:20Z","timestamp":1773086120449,"version":"3.50.1"},"reference-count":27,"publisher":"Oxford University Press (OUP)","issue":"13","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2014,7,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: Current high-throughput sequencing has greatly transformed genome sequence analysis. In the context of very low-coverage sequencing (&amp;lt;0.1\u00d7), performing \u2018binning\u2019 or \u2018windowing\u2019 on mapped short sequences (\u2018reads\u2019) is critical to extract genomic information of interest for further evaluation, such as copy-number alteration analysis. If the window size is too small, many windows will exhibit zero counts and almost no pattern can be observed. In contrast, if the window size is too wide, the patterns or genomic features will be \u2018smoothed out\u2019. Our objective is to identify an optimal window size in between the two extremes.<\/jats:p><jats:p>Results: We assume the reads density to be a step function. Given this model, we propose a data-based estimation of optimal window size based on Akaike\u2019s information criterion (AIC) and cross-validation (CV) log-likelihood. By plotting the AIC and CV log-likelihood curve as a function of window size, we are able to estimate the optimal window size that minimizes AIC or maximizes CV log-likelihood. The proposed methods are of general purpose and we illustrate their application using low-coverage next-generation sequence datasets from real tumour samples and simulated datasets.<\/jats:p><jats:p>Availability and implementation: An R package to estimate optimal window size is available at http:\/\/www1.maths.leeds.ac.uk\/\u223carief\/R\/win\/ .<\/jats:p><jats:p>Contact: \u00a0a.gusnanto@leeds.ac.uk<\/jats:p><jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btu123","type":"journal-article","created":{"date-parts":[[2014,3,7]],"date-time":"2014-03-07T02:57:42Z","timestamp":1394161062000},"page":"1823-1829","source":"Crossref","is-referenced-by-count":29,"title":["Estimating optimal window size for analysis of low-coverage next-generation sequence data"],"prefix":"10.1093","volume":"30","author":[{"given":"Arief","family":"Gusnanto","sequence":"first","affiliation":[{"name":"1 Department of Statistics, University of Leeds, Leeds LS2 9JT, United Kingdom, 2 Department of Statistics, Faculty of Science, King Saud University, Riyadh, Saudi Arabia, 3 Leeds Institute Cancer and Pathology, University of Leeds, Leeds LS9 7TF, UK and 4 Illumina UK Ltd., Chesterford Research Park, Saffron Walden, CB10 1XL, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Charles C.","family":"Taylor","sequence":"additional","affiliation":[{"name":"1 Department of Statistics, University of Leeds, Leeds LS2 9JT, United Kingdom, 2 Department of Statistics, Faculty of Science, King Saud University, Riyadh, Saudi Arabia, 3 Leeds Institute Cancer and Pathology, University of Leeds, Leeds LS9 7TF, UK and 4 Illumina UK Ltd., Chesterford Research Park, Saffron Walden, CB10 1XL, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ibrahim","family":"Nafisah","sequence":"additional","affiliation":[{"name":"1 Department of Statistics, University of Leeds, Leeds LS2 9JT, United Kingdom, 2 Department of Statistics, Faculty of Science, King Saud University, Riyadh, Saudi Arabia, 3 Leeds Institute Cancer and Pathology, University of Leeds, Leeds LS9 7TF, UK and 4 Illumina UK Ltd., Chesterford Research Park, Saffron Walden, CB10 1XL, UK"},{"name":"1 Department of Statistics, University of Leeds, Leeds LS2 9JT, United Kingdom, 2 Department of Statistics, Faculty of Science, King Saud University, Riyadh, Saudi Arabia, 3 Leeds Institute Cancer and Pathology, University of Leeds, Leeds LS9 7TF, UK and 4 Illumina UK Ltd., Chesterford Research Park, Saffron Walden, CB10 1XL, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Henry M.","family":"Wood","sequence":"additional","affiliation":[{"name":"1 Department of Statistics, University of Leeds, Leeds LS2 9JT, United Kingdom, 2 Department of Statistics, Faculty of Science, King Saud University, Riyadh, Saudi Arabia, 3 Leeds Institute Cancer and Pathology, University of Leeds, Leeds LS9 7TF, UK and 4 Illumina UK Ltd., Chesterford Research Park, Saffron Walden, CB10 1XL, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pamela","family":"Rabbitts","sequence":"additional","affiliation":[{"name":"1 Department of Statistics, University of Leeds, Leeds LS2 9JT, United Kingdom, 2 Department of Statistics, Faculty of Science, King Saud University, Riyadh, Saudi Arabia, 3 Leeds Institute Cancer and Pathology, University of Leeds, Leeds LS9 7TF, UK and 4 Illumina UK Ltd., Chesterford Research Park, Saffron Walden, CB10 1XL, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stefano","family":"Berri","sequence":"additional","affiliation":[{"name":"1 Department of Statistics, University of Leeds, Leeds LS2 9JT, United Kingdom, 2 Department of Statistics, Faculty of Science, King Saud University, Riyadh, Saudi Arabia, 3 Leeds Institute Cancer and Pathology, University of Leeds, Leeds LS9 7TF, UK and 4 Illumina UK Ltd., Chesterford Research Park, Saffron Walden, CB10 1XL, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2014,3,5]]},"reference":[{"key":"2023012711163583300_btu123-B1","doi-asserted-by":"crossref","first-page":"e72","DOI":"10.1093\/nar\/gks001","article-title":"Summarizing and correcting the GC content bias in high-throughput sequencing","volume":"40","author":"Benjamini","year":"2012","journal-title":"Nucleic Acids Res."},{"key":"2023012711163583300_btu123-B2","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1038\/nature07517","article-title":"Accurate whole human genome sequencing using reversible terminator chemistry","volume":"456","author":"Bentley","year":"2008","journal-title":"Nature"},{"key":"2023012711163583300_btu123-B3","doi-asserted-by":"crossref","first-page":"2537","DOI":"10.1093\/bioinformatics\/btn480","article-title":"F-Seq: a feature density estimator for high-throughput sequence tags","volume":"24","author":"Boyle","year":"2008","journal-title":"Bioinformatics"},{"key":"2023012711163583300_btu123-B4","doi-asserted-by":"crossref","first-page":"244","DOI":"10.1186\/1471-2164-11-244","article-title":"DNA copy number, including telomeres and mitochondria, assayed using next-generation sequencing","volume":"11","author":"Castle","year":"2010","journal-title":"BMC Genomics"},{"key":"2023012711163583300_btu123-B5","doi-asserted-by":"crossref","first-page":"R15","DOI":"10.1186\/gb-2011-12-2-r15","article-title":"A statistical framework for modelling gene expression using chromatin features and application to modENCODE datasets","volume":"12","author":"Cheng","year":"2011","journal-title":"Genome Biol."},{"key":"2023012711163583300_btu123-B6","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1038\/nmeth.1276","article-title":"High-resolution mapping of copy-number alterations with massively parallel sequencing","volume":"6","author":"Chiang","year":"2009","journal-title":"Nat. Methods"},{"key":"2023012711163583300_btu123-B7","doi-asserted-by":"crossref","first-page":"453","DOI":"10.1007\/BF01025868","article-title":"On the histogram as a density estimator: L2 theory","volume":"57","author":"Freedman","year":"1981","journal-title":"Z. Wahrsheinllchkeffstheorie Verwandte Gebeite"},{"key":"2023012711163583300_btu123-B8","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1093\/bioinformatics\/btr593","article-title":"Correcting for cancer genome size and tumour cell content enables better estimation of copy number alterations from next-generation sequence data","volume":"28","author":"Gusnanto","year":"2012","journal-title":"Bioinformatics"},{"key":"2023012711163583300_btu123-B9","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1016\/0167-7152(87)90083-6","article-title":"Estimation of integrated squared density derivatives","volume":"6","author":"Hall","year":"1987","journal-title":"Stat. Probab. Lett."},{"key":"2023012711163583300_btu123-B10","doi-asserted-by":"crossref","first-page":"2463","DOI":"10.1093\/bioinformatics\/btm359","article-title":"Robust smooth segmentation approach for array CGH data analysis","volume":"23","author":"Huang","year":"2007","journal-title":"Bioinformatics"},{"key":"2023012711163583300_btu123-B11","doi-asserted-by":"crossref","first-page":"1497","DOI":"10.1126\/science.1141319","article-title":"Genome-wide mapping of in vivo protein-DNA interactions","volume":"316","author":"Johnson","year":"2007","journal-title":"Science"},{"key":"2023012711163583300_btu123-B12","doi-asserted-by":"crossref","first-page":"511","DOI":"10.1016\/0167-7152(91)90116-9","article-title":"Using nonstochastic terms to advantage in kernel-based estimation of integrated squared density derivatives","volume":"11","author":"Jones","year":"1991","journal-title":"Stat. Probab. Lett."},{"key":"2023012711163583300_btu123-B13","doi-asserted-by":"crossref","first-page":"2097","DOI":"10.1093\/bioinformatics\/bts330","article-title":"Genomic dark matter: the reliability of short read mapping illustrated by the genome mappability score","volume":"28","author":"Lee","year":"2012","journal-title":"Bioinformatics"},{"key":"2023012711163583300_btu123-B14","doi-asserted-by":"crossref","first-page":"1754","DOI":"10.1093\/bioinformatics\/btp324","article-title":"Fast and accurate short read alignment with Burrows-Wheeler transform","volume":"25","author":"Li","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012711163583300_btu123-B15","doi-asserted-by":"crossref","first-page":"557","DOI":"10.1093\/biostatistics\/kxh008","article-title":"Circular binary segmentation for the analysis of array-based DNA copy number data","volume":"5","author":"Olshen","year":"2004","journal-title":"Biostatistics"},{"key":"2023012711163583300_btu123-B16","doi-asserted-by":"crossref","DOI":"10.1093\/oso\/9780198507659.001.0001","volume-title":"In All Likelihood: Statistical Modelling and Inference using Likelihood","author":"Pawitan","year":"2001"},{"key":"2023012711163583300_btu123-B17","doi-asserted-by":"crossref","first-page":"191","DOI":"10.1038\/nature08658","article-title":"A comprehensive catalogue of somatic mutations from a human cancer genome","volume":"463","author":"Pleasance","year":"2010","journal-title":"Nature"},{"key":"2023012711163583300_btu123-B18","doi-asserted-by":"crossref","first-page":"651","DOI":"10.1038\/nmeth1068","article-title":"Genome-wide profiles of STAT1 DNA association using chromatin immunoprecipitation and massively parallel sequencing","volume":"4","author":"Robertson","year":"2007","journal-title":"Nat. Methods"},{"key":"2023012711163583300_btu123-B19","doi-asserted-by":"crossref","first-page":"605","DOI":"10.1093\/biomet\/66.3.605","article-title":"On optimal and data-based histograms","volume":"66","author":"Scott","year":"1979","journal-title":"Biometrika"},{"key":"2023012711163583300_btu123-B20","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1111\/j.2517-6161.1974.tb00994.x","article-title":"Cross-validatory choice and assessment of statistical prediction","volume":"36","author":"Stone","year":"1974","journal-title":"J. R. Stat. Soc. B"},{"key":"2023012711163583300_btu123-B21","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1111\/j.2517-6161.1977.tb01603.x","article-title":"An asymptotic equivalence of choice of model by cross-validation and Akaike\u2019s criterion","volume":"39","author":"Stone","year":"1977","journal-title":"J. R. Stat. Soc. B"},{"key":"2023012711163583300_btu123-B22","doi-asserted-by":"crossref","first-page":"636","DOI":"10.1093\/biomet\/74.3.636","article-title":"Akaike\u2019s information criterion and the histogram","volume":"74","author":"Taylor","year":"1987","journal-title":"Biometrika"},{"key":"2023012711163583300_btu123-B23","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1080\/00031305.1997.10473591","article-title":"Data-based choice of histogram bin width","volume":"51","author":"Wand","year":"1997","journal-title":"Am. Stat."},{"key":"2023012711163583300_btu123-B24","doi-asserted-by":"crossref","first-page":"e151","DOI":"10.1093\/nar\/gkq510","article-title":"Using next-generation sequencing for high resolution multiplex analysis of copy number variation from nanogram quantities of DNA from formalin-fixed paraffin-embedded specimens","volume":"38","author":"Wood","year":"2010","journal-title":"Nucleic Acids Res."},{"key":"2023012711163583300_btu123-B25","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1093\/bfgp\/elq025","article-title":"Detecting structural variations in the human genome using next generation sequencing","volume":"9","author":"Xi","year":"2011","journal-title":"Brief. Funct. Genomics"},{"key":"2023012711163583300_btu123-B26","doi-asserted-by":"crossref","first-page":"80","DOI":"10.1186\/1471-2105-10-80","article-title":"CNV-seq, a new method to detect copy number variation using high-throughput sequencing","volume":"10","author":"Xie","year":"2009","journal-title":"BMC Bioinformatics"},{"key":"2023012711163583300_btu123-B27","doi-asserted-by":"crossref","first-page":"1586","DOI":"10.1101\/gr.092981.109","article-title":"Sensitive and accurate detection of copy number variants using read depth of coverage","volume":"19","author":"Yoon","year":"2009","journal-title":"Genome Res."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/13\/1823\/48926025\/bioinformatics_30_13_1823.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/13\/1823\/48926025\/bioinformatics_30_13_1823.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,24]],"date-time":"2024-05-24T21:27:23Z","timestamp":1716586043000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/30\/13\/1823\/2422194"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,3,5]]},"references-count":27,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2014,7,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btu123","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2014,7,1]]},"published":{"date-parts":[[2014,3,5]]}}}