{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,11]],"date-time":"2025-11-11T12:46:17Z","timestamp":1762865177067},"reference-count":28,"publisher":"Oxford University Press (OUP)","issue":"24","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2006,12,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Large-scale tiling array experiments are becoming increasingly common in genomics. In particular, the ENCODE project requires the consistent segmentation of many different tiling array datasets into \u2018active regions\u2019 (e.g. finding transfrags from transcriptional data and putative binding sites from ChIP-chip experiments). Previously, such segmentation was done in an unsupervised fashion mainly based on characteristics of the signal distribution in the tiling array data itself. Here we propose a supervised framework for doing this. It has the advantage of explicitly incorporating validated biological knowledge into the model and allowing for formal training and testing.<\/jats:p>\n               <jats:p>Methodology: In particular, we use a hidden Markov model (HMM) framework, which is capable of explicitly modeling the dependency between neighboring probes and whose extended version (the generalized HMM) also allows explicit description of state duration density. We introduce a formal definition of the tiling-array analysis problem, and explain how we can use this to describe sampling small genomic regions for experimental validation to build up a gold-standard set for training and testing. We then describe various ideal and practical sampling strategies (e.g. maximizing signal entropy within a selected region versus using gene annotation or known promoters as positives for transcription or ChIP-chip data, respectively).<\/jats:p>\n               <jats:p>Results: For the practical sampling and training strategies, we show how the size and noise in the validated training data affects the performance of an HMM applied to the ENCODE transcriptional and ChIP-chip experiments. In particular, we show that the HMM framework is able to efficiently process tiling array data as well as or better than previous approaches. For the idealized sampling strategies, we show how we can assess their performance in a simulation framework and how a maximum entropy approach, which samples sub-regions with very different signal intensities, gives the maximally performing gold-standard. This latter result has strong implications for the optimum way medium-scale validation experiments should be carried out to verify the results of the genome-scale tiling array experiments.<\/jats:p>\n               <jats:p>Supplementary information: The supplementary data are available at<\/jats:p>\n               <jats:p>Contact: \u00a0mark.gerstein@yale.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/btl515","type":"journal-article","created":{"date-parts":[[2006,10,13]],"date-time":"2006-10-13T07:47:15Z","timestamp":1160725635000},"page":"3016-3024","source":"Crossref","is-referenced-by-count":26,"title":["A supervised hidden markov model framework for efficiently segmenting tiling array data in transcriptional and chIP-chip experiments: systematically incorporating validated biological knowledge"],"prefix":"10.1093","volume":"22","author":[{"given":"Jiang","family":"Du","sequence":"first","affiliation":[{"name":"Department of Computer Science, Yale University 1 \u00a0 1 \u00a0 \u00a0 New Haven, CT 06520, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joel S.","family":"Rozowsky","sequence":"additional","affiliation":[{"name":"Department of Molecular Biophysics and Biochemistry, Yale University 2 \u00a0 2 \u00a0 \u00a0 New Haven, CT 06520, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jan O.","family":"Korbel","sequence":"additional","affiliation":[{"name":"Department of Molecular Biophysics and Biochemistry, Yale University 2 \u00a0 2 \u00a0 \u00a0 New Haven, CT 06520, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhengdong D.","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Molecular Biophysics and Biochemistry, Yale University 2 \u00a0 2 \u00a0 \u00a0 New Haven, CT 06520, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Thomas E.","family":"Royce","sequence":"additional","affiliation":[{"name":"Department of Molecular Biophysics and Biochemistry, Yale University 2 \u00a0 2 \u00a0 \u00a0 New Haven, CT 06520, USA"},{"name":"Program in Computational Biology and Bioinformatics, Yale University 3 \u00a0 3 \u00a0 \u00a0 New Haven, CT 06520, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Martin H.","family":"Schultz","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Yale University 1 \u00a0 1 \u00a0 \u00a0 New Haven, CT 06520, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael","family":"Snyder","sequence":"additional","affiliation":[{"name":"Department of Molecular Biophysics and Biochemistry, Yale University 2 \u00a0 2 \u00a0 \u00a0 New Haven, CT 06520, USA"},{"name":"Department of Molecular, Cellular and Developmental Biology, Yale University 4 \u00a0 4 \u00a0 \u00a0 New Haven, CT 06520, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mark","family":"Gerstein","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Yale University 1 \u00a0 1 \u00a0 \u00a0 New Haven, CT 06520, USA"},{"name":"Department of Molecular Biophysics and Biochemistry, Yale University 2 \u00a0 2 \u00a0 \u00a0 New Haven, CT 06520, USA"},{"name":"Program in Computational Biology and Bioinformatics, Yale University 3 \u00a0 3 \u00a0 \u00a0 New Haven, CT 06520, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2006,10,12]]},"reference":[{"key":"2023012408501675000_b1","doi-asserted-by":"crossref","first-page":"205","DOI":"10.1007\/BF00992677","article-title":"On the computational complexity of approximating distributions by probabilistic automata","volume":"9","author":"Abe","year":"1992","journal-title":"Mac. Learn."},{"key":"2023012408501675000_b2","doi-asserted-by":"crossref","first-page":"2242","DOI":"10.1126\/science.1103388","article-title":"Global identification of human transcribed sequences with genome tiling arrays","volume":"306","author":"Bertone","year":"2004","journal-title":"Science"},{"key":"2023012408501675000_b3","doi-asserted-by":"crossref","first-page":"349","DOI":"10.1016\/j.ygeno.2003.11.004","article-title":"Chip-chip: considerations for the design, analysis, and application of genome-wide chromatin immunoprecipitation experiments","volume":"83","author":"Buck","year":"2004","journal-title":"Genomics"},{"key":"2023012408501675000_b4","doi-asserted-by":"crossref","first-page":"499","DOI":"10.1016\/S0092-8674(04)00127-8","article-title":"Unbiased mapping of transcription factor binding sites along human chromosomes 21 and 22 points to widespread regulation of noncoding RNAs","volume":"116","author":"Cawley","year":"2004","journal-title":"Cell"},{"key":"2023012408501675000_b5","doi-asserted-by":"crossref","first-page":"1149","DOI":"10.1126\/science.1108625","article-title":"Transcriptional maps of 10 human chromosomes at 5-nucleotide resolution","volume":"308","author":"Cheng","year":"2005","journal-title":"Science"},{"key":"2023012408501675000_b6","doi-asserted-by":"crossref","first-page":"755","DOI":"10.1093\/bioinformatics\/14.9.755","article-title":"Profile hidden markov models","volume":"14","author":"Eddy","year":"1998","journal-title":"Bioinformatics"},{"key":"2023012408501675000_b7","doi-asserted-by":"crossref","first-page":"636","DOI":"10.1126\/science.1105136","article-title":"The ENCODE (ENCyclopedia Of DNA Elements) Project","volume":"306","author":"ENCODE Project Consortium","year":"2004","journal-title":"Science"},{"key":"2023012408501675000_b8","doi-asserted-by":"crossref","first-page":"R96","DOI":"10.1186\/gb-2005-6-11-r96","article-title":"Chipper: discovering transcription-factor targets from chromatin immunoprecipitation microarrays using variance stabilization","volume":"6","author":"Gibbons","year":"2005","journal-title":"Genome Biol."},{"key":"2023012408501675000_b9","doi-asserted-by":"crossref","first-page":"576","DOI":"10.1093\/bioinformatics\/18.4.576","article-title":"Making sense of microarray data distributions","volume":"18","author":"Hoyle","year":"2002","journal-title":"Bioinformatics"},{"key":"2023012408501675000_b10","doi-asserted-by":"crossref","first-page":"533","DOI":"10.1038\/35054095","article-title":"Genomic binding sites of the yeast cell-cycle transcription factors SBF and MBF","volume":"409","author":"Iyer","year":"2001","journal-title":"Nature"},{"key":"2023012408501675000_b11","doi-asserted-by":"crossref","first-page":"3629","DOI":"10.1093\/bioinformatics\/bti593","article-title":"TileMap: create chromosomal map of tiling array hybridizations","volume":"21","author":"Ji","year":"2005","journal-title":"Bioinformatics"},{"key":"2023012408501675000_b12","doi-asserted-by":"crossref","first-page":"331","DOI":"10.1101\/gr.2094104","article-title":"Novel RNAs identified from an in-depth analysis of the transcriptome of human chromosomes 21 and 22","volume":"14","author":"Kampa","year":"2004","journal-title":"Genome Res"},{"key":"2023012408501675000_b13","doi-asserted-by":"crossref","first-page":"916","DOI":"10.1126\/science.1068597","article-title":"Large-scale transcriptional activity in chromosomes 21 and 22","volume":"296","author":"Kapranov","year":"2002","journal-title":"Science"},{"key":"2023012408501675000_b14","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1002\/(SICI)1097-0134(1999)37:3+<121::AID-PROT16>3.0.CO;2-Q","article-title":"Predicting protein structure using only sequence information","author":"Karplus","year":"1999","journal-title":"Proteins"},{"key":"2023012408501675000_b15","doi-asserted-by":"crossref","first-page":"1501","DOI":"10.1006\/jmbi.1994.1104","article-title":"Hidden Markov models in computational biology. Applications to protein modeling","volume":"235","author":"Krogh","year":"1994","journal-title":"J. Mol. Biol."},{"key":"2023012408501675000_b16","doi-asserted-by":"crossref","first-page":"79","DOI":"10.1214\/aoms\/1177729694","article-title":"On information and sufficiency","volume":"22","author":"Kullback","year":"1951","journal-title":"Ann. Math. Stat."},{"key":"2023012408501675000_b17","first-page":"4781","article-title":"Genome-wide loss of heterozygosity analysis from laser capture microdissected prostate cancer using single nucleotide polymorphic allele (SNP) arrays and a novel bioinformatics platform dChipSNP","volume":"63","author":"Lieberfarb","year":"2003","journal-title":"Cancer Res."},{"key":"2023012408501675000_b18","doi-asserted-by":"crossref","first-page":"i274","DOI":"10.1093\/bioinformatics\/bti1046","article-title":"A hidden Markov model for analyzing ChIP-chip experiments on genome tiling arrays and its application to p53 binding sequences","volume":"21","author":"Li","year":"2005","journal-title":"Bioinformatics"},{"key":"2023012408501675000_b19","doi-asserted-by":"crossref","first-page":"1144","DOI":"10.1093\/bioinformatics\/btl089","article-title":"BioHMM: a heterogeneous hidden Markov model for segmenting array CGH data","volume":"22","author":"Marioni","year":"2006","journal-title":"Bioinformatics"},{"key":"2023012408501675000_b20","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1109\/91.824772","article-title":"Generalized hidden markov models \u2013 part i: Theoretical frameworks","volume":"8","author":"Mohamed","year":"2000","journal-title":"IEEE Transcations on Fuzzy Systems"},{"key":"2023012408501675000_b21","doi-asserted-by":"crossref","first-page":"1065","DOI":"10.1214\/aoms\/1177704472","article-title":"On estimation of a probability density function and mode","volume":"33","author":"Parzen","year":"1962","journal-title":"Ann. Math. Stat."},{"key":"2023012408501675000_b22","doi-asserted-by":"crossref","first-page":"D501","DOI":"10.1093\/nar\/gki025","article-title":"NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins","volume":"33","author":"Pruitt","year":"2005","journal-title":"Nucleic Acids Res."},{"key":"2023012408501675000_b23","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1109\/5.18626","article-title":"A tutorial on hidden markov models and selected applications in speech recognition","volume":"77","author":"Rabiner","year":"1989","journal-title":"Proc. IEEE"},{"key":"2023012408501675000_b24","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1101\/gad.1055203","article-title":"The transcriptional activity of human Chromosome 22","volume":"17","author":"Rinn","year":"2003","journal-title":"Genes Dev."},{"key":"2023012408501675000_b25","doi-asserted-by":"crossref","first-page":"466","DOI":"10.1016\/j.tig.2005.06.007","article-title":"Issues in the analysis of oligonucleotide tiling microarrays for transcript mapping","volume":"21","author":"Royce","year":"2005","journal-title":"Trends Genet."},{"key":"2023012408501675000_b26","doi-asserted-by":"crossref","first-page":"R73","DOI":"10.1186\/gb-2004-5-10-r73","article-title":"A comprehensive transcript index of the human genome generated using microarrays and computational approaches","volume":"5","author":"Schadt","year":"2004","journal-title":"Genome Biol."},{"key":"2023012408501675000_b27","doi-asserted-by":"crossref","first-page":"260","DOI":"10.1109\/TIT.1967.1054010","article-title":"Error bounds for convolutional codes and an asymptotically optimum decoding algorithm","volume":"13","author":"Viterbi","year":"1967","journal-title":"IEEE Trans. Inform. Theory"},{"key":"2023012408501675000_b28","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1002\/1097-0142(1950)3:1<32::AID-CNCR2820030106>3.0.CO;2-3","article-title":"Index for rating diagnostic tests","volume":"3","author":"Youden","year":"1950","journal-title":"Cancer"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/24\/3016\/48838786\/bioinformatics_22_24_3016.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/24\/3016\/48838786\/bioinformatics_22_24_3016.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,24]],"date-time":"2023-01-24T09:18:44Z","timestamp":1674551924000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/22\/24\/3016\/208870"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,10,12]]},"references-count":28,"journal-issue":{"issue":"24","published-print":{"date-parts":[[2006,12,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btl515","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2006,12,15]]},"published":{"date-parts":[[2006,10,12]]}}}