{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,25]],"date-time":"2025-09-25T16:02:01Z","timestamp":1758816121015},"reference-count":32,"publisher":"Oxford University Press (OUP)","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2006,2,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Background: In the post-genomic era, developing tools to decode biological information from genomic sequences is important. Inspired by affiliation network theory, we investigated gene sequences of two kinds of UniGene clusters (UCs): narrowly expressed transcripts (NETs), whose expression is confined to a few tissues; and prevalently expressed transcripts (PETs) that are expressed in many tissues.<\/jats:p>\n               <jats:p>Results: We explored the human and the mouse UniGene databases to compare NETs and PETs from different perspectives. We found that NETs were associated with smaller cluster size, shorter sequence length, a lower likelihood of having LocusLink annotations, and lower and more sporadic levels of expression. Significantly, the dinucleotide frequencies of NETs are similar to those of intergenic sequences in the genome, and they differ from those of PETs. We used these differences in dinucleotide frequencies to develop a discriminant analysis model to distinguish PETs from intergenic sequences.<\/jats:p>\n               <jats:p>Conclusions: Our results show that most NETs resemble intergenic sequences, casting doubts on the quality of such UniGene clusters. However, we also noted that a fraction of NETs resemble PETs in terms of dinucleotide frequencies and other features. Such NETs may have fewer quality problems. This work may be helpful in the studies of non-coding RNAs and in the validation of gene sequence databases.<\/jats:p>\n               <jats:p>Availability: \u00a0<\/jats:p>\n               <jats:p>Contact: \u00a0kcoombes@mdanderson.org<\/jats:p>\n               <jats:p>Supplementary information: \u00a0<\/jats:p>","DOI":"10.1093\/bioinformatics\/bti796","type":"journal-article","created":{"date-parts":[[2005,12,9]],"date-time":"2005-12-09T01:43:34Z","timestamp":1134092614000},"page":"385-391","source":"Crossref","is-referenced-by-count":5,"title":["Gene sequence signatures revealed by mining the UniGene affiliation network"],"prefix":"10.1093","volume":"22","author":[{"given":"Jiexin","family":"Zhang","sequence":"first","affiliation":[{"name":"Department of Biostatistics and Applied Mathematics, The University of Texas M.D. Anderson Cancer Center \u00a0 1515 Holcombe Boulevard, Box 447, Houston, TX 77030-4009, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Li","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Biostatistics and Applied Mathematics, The University of Texas M.D. Anderson Cancer Center \u00a0 1515 Holcombe Boulevard, Box 447, Houston, TX 77030-4009, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kevin R.","family":"Coombes","sequence":"additional","affiliation":[{"name":"Department of Biostatistics and Applied Mathematics, The University of Texas M.D. Anderson Cancer Center \u00a0 1515 Holcombe Boulevard, Box 447, Houston, TX 77030-4009, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2005,12,8]]},"reference":[{"key":"2023012408511854800_b1","doi-asserted-by":"crossref","first-page":"11995","DOI":"10.1073\/pnas.90.24.11995","article-title":"Number of CpG islands and genes in the human and mouse genomes","volume":"90","author":"Antiquera","year":"1993","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012408511854800_b2","doi-asserted-by":"crossref","first-page":"69","DOI":"10.1146\/annurev.genom.4.070802.110300","article-title":"Gene annotation: prediction and testing","volume":"4","author":"Ashurst","year":"2003","journal-title":"Annu. Rev. Genomics Hum. Genet."},{"key":"2023012408511854800_b3","doi-asserted-by":"crossref","first-page":"2242","DOI":"10.1126\/science.1103388","article-title":"Global identification of human transcribed sequences with genome tiling arrays","volume":"306","author":"Bertone","year":"2004","journal-title":"Science"},{"key":"2023012408511854800_b4","doi-asserted-by":"crossref","first-page":"78","DOI":"10.1006\/jmbi.1997.0951","article-title":"Prediction of complete gene structures in human genomic DNA","volume":"268","author":"Burge","year":"1997","journal-title":"J. Mol. Biol."},{"key":"2023012408511854800_b5","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1093\/emboj\/20.1.1","article-title":"The black cat\/white cat principle of signalintegration in bacterial promoters","volume":"20","author":"Cases","year":"2001","journal-title":"EMBO J."},{"key":"2023012408511854800_b6","doi-asserted-by":"crossref","first-page":"499","DOI":"10.1016\/S0092-8674(04)00127-8","article-title":"Unbiased mapping of transcription factor binding sites along human chromosomes 21 and 22 points to widespread regulation of noncoding RNAs","volume":"116","author":"Cawley","year":"2004","journal-title":"Cell"},{"key":"2023012408511854800_b7","doi-asserted-by":"crossref","first-page":"71","DOI":"10.1006\/geno.2001.6620","article-title":"Assessment of the total number of human transcription units","volume":"77","author":"Das","year":"2001","journal-title":"Genomics"},{"key":"2023012408511854800_b8","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1002\/prot.20147","article-title":"A unified representation of multiprotein complex data formodeling interaction networks","volume":"57","author":"Ding","year":"2004","journal-title":"Proteins"},{"key":"2023012408511854800_b9","doi-asserted-by":"crossref","first-page":"362","DOI":"10.1016\/S0168-9525(03)00140-9","article-title":"Human housekeeping genes are compact","volume":"19","author":"Eisenberg","year":"2003","journal-title":"Trends Genet."},{"key":"2023012408511854800_b10","doi-asserted-by":"crossref","first-page":"345","DOI":"10.1038\/ng0794-345","article-title":"How many genes in the human genome?","volume":"7","author":"Fields","year":"1994","journal-title":"Nat. Genet."},{"key":"2023012408511854800_b11","doi-asserted-by":"crossref","first-page":"261","DOI":"10.1016\/0022-2836(87)90689-9","article-title":"CpG islands in vertebrate genomes","volume":"196","author":"Gardiner-Garden","year":"1987","journal-title":"J. Mol. Biol."},{"key":"2023012408511854800_b12","doi-asserted-by":"crossref","first-page":"807","DOI":"10.1101\/gr.6.9.807","article-title":"Generation and analysis of 280\u2009000 human expressed sequence tags","volume":"6","author":"Hillier","year":"1996","journal-title":"Genome Res."},{"key":"2023012408511854800_b13","doi-asserted-by":"crossref","first-page":"931","DOI":"10.1038\/nature03001","article-title":"Finishing the euchromatic sequence of the human genome","volume":"431","author":"International Human Genome Sequencing Consortium (IHGSC)","year":"2004","journal-title":"Nature"},{"key":"2023012408511854800_b14","doi-asserted-by":"crossref","first-page":"331","DOI":"10.1101\/gr.2094104","article-title":"Novel RNAs identified from an in-depth analysis of thetranscriptome of human chromosomes 21 and 22","volume":"14","author":"Kampa","year":"2004","journal-title":"Genome Res."},{"key":"2023012408511854800_b15","doi-asserted-by":"crossref","first-page":"916","DOI":"10.1126\/science.1068597","article-title":"Large-scale transcriptional activity in chromosomes 21 and 22","volume":"296","author":"Kapranov","year":"2002","journal-title":"Science"},{"key":"2023012408511854800_b16","doi-asserted-by":"crossref","first-page":"598","DOI":"10.1016\/S1369-5274(98)80095-7","article-title":"Global dinucleotide signatures and analysis of genomic heterogeneity","volume":"1","author":"Karlin","year":"1998","journal-title":"Curr. Opin. Microbiol."},{"key":"2023012408511854800_b17","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1016\/S0168-9525(00)89076-9","article-title":"Dinucleotide relative abundance extremes: a genomic signature","volume":"11","author":"Karlin","year":"1995","journal-title":"Trends Genet."},{"key":"2023012408511854800_b18","volume-title":"Principles of Multivariate Analysis: A User's Perspective","author":"Krzanowski","year":"1988"},{"key":"2023012408511854800_b19","doi-asserted-by":"crossref","first-page":"860","DOI":"10.1038\/35057062","article-title":"Initial sequencing and analysis of the human genome","volume":"409","author":"Lander","year":"2001","journal-title":"Nature"},{"key":"2023012408511854800_b20","doi-asserted-by":"crossref","first-page":"1095","DOI":"10.1016\/0888-7543(92)90024-M","article-title":"CpG islands as gene markers in the human genome","volume":"13","author":"Larsen","year":"1992","journal-title":"Genomics"},{"key":"2023012408511854800_b21","doi-asserted-by":"crossref","first-page":"2566","DOI":"10.1073\/pnas.012582999","article-title":"Random graph models of social networks","volume":"99","author":"Newman","year":"2002","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012408511854800_b22","article-title":"UniGene: a unified view of the transcriptome","author":"Pontius","year":"2003"},{"key":"2023012408511854800_b23","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1016\/S0168-9525(99)01882-X","article-title":"Introducing RefSeq and LocusLink: curated human genome resources at the NCBI","volume":"16","author":"Pruitt","year":"2000","journal-title":"Trends Genet."},{"key":"2023012408511854800_b24","doi-asserted-by":"crossref","first-page":"694","DOI":"10.1007\/s001090050155","article-title":"Pieces of the puzzle: expressed sequence tags and the catalog of human genes","volume":"75","author":"Schuler","year":"1997","journal-title":"J. Mol. Med."},{"key":"2023012408511854800_b25","doi-asserted-by":"crossref","first-page":"1067","DOI":"10.1093\/nar\/gkg170","article-title":"A novel algorithm for computational identification of contaminated EST libraries","volume":"31","author":"Sorek","year":"2003","journal-title":"Nucleic Acids Res."},{"key":"2023012408511854800_b26","doi-asserted-by":"crossref","first-page":"268","DOI":"10.1038\/35065725","article-title":"Exploring complex networks","volume":"410","author":"Strogatz","year":"2001","journal-title":"Nature"},{"key":"2023012408511854800_b28","doi-asserted-by":"crossref","first-page":"176","DOI":"10.1353\/pbm.2004.0036","article-title":"A \u2018small-world\u2019 network hypothesis for memory and dreams","volume":"47","author":"Tsonis","year":"2004","journal-title":"Perspect. Biol. Med."},{"key":"2023012408511854800_b29","doi-asserted-by":"crossref","first-page":"1304","DOI":"10.1126\/science.1058040","article-title":"The sequence of the human genome","volume":"291","author":"Venter","year":"2001","journal-title":"Science"},{"key":"2023012408511854800_b30","volume-title":"Six Degrees: The Science of a Connected Age","author":"Watts","year":"2003"},{"key":"2023012408511854800_b31","doi-asserted-by":"crossref","first-page":"115","DOI":"10.1016\/S0163-7258(01)00151-6","article-title":"Genome analysis with gene-indexing databases","volume":"91","author":"Yuan","year":"2001","journal-title":"Pharmacol. Ther."},{"key":"2023012408511854800_b32","doi-asserted-by":"crossref","first-page":"818","DOI":"10.1038\/nbt836","article-title":"A model of molecular interactions on short oligonucleotide microarrays","volume":"21","author":"Zhang","year":"2003","journal-title":"Nat. Biotechnol."},{"key":"2023012408511854800_b33","doi-asserted-by":"crossref","first-page":"565","DOI":"10.1073\/pnas.94.2.565","article-title":"Identification of protein coding regions in the human genome by quadratic discriminant analysis","volume":"94","author":"Zhang","year":"1997","journal-title":"Proc. Natl Acad. Sci. USA"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/4\/385\/48839827\/bioinformatics_22_4_385.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/4\/385\/48839827\/bioinformatics_22_4_385.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,24]],"date-time":"2023-01-24T09:22:41Z","timestamp":1674552161000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/22\/4\/385\/183310"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,12,8]]},"references-count":32,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2006,2,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bti796","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2006,2,15]]},"published":{"date-parts":[[2005,12,8]]}}}