{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,21]],"date-time":"2026-01-21T12:09:14Z","timestamp":1768997354758,"version":"3.49.0"},"reference-count":21,"publisher":"Oxford University Press (OUP)","issue":"8","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2007,4,15]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: Complex genomes contain numerous repeated sequences, and genomic duplication is believed to be a main evolutionary mechanism to obtain new functions. Several tools are available for de novo repeat sequence identification, and many approaches exist for clustering homologous protein sequences. We present an efficient new approach to identify and cluster homologous DNA sequences with high accuracy at the level of whole genomes, excluding low-complexity repeats, tandem repeats and annotated interspersed repeats. We also determine the boundaries of each group member so that it closely represents a biological unit, e.g. a complete gene, or a partial gene coding a protein domain.<\/jats:p><jats:p>Results: We developed a program called HomologMiner to identify homologous groups applicable to genome sequences that have been properly marked for low-complexity repeats and annotated interspersed repeats. We applied it to the whole genomes of human (hg17), macaque (rheMac2) and mouse (mm8). Groups obtained include gene families (e.g. olfactory receptor gene family, zinc finger families), unannotated interspersed repeats and additional homologous groups that resulted from recent segmental duplications. Our program incorporates several new methods: a new abstract definition of consistent duplicate units, a new criterion to remove moderately frequent tandem repeats, and new algorithmic techniques. We also provide preliminary analysis of the output on the three genomes mentioned above, and show several applications including identifying boundaries of tandem gene clusters and novel interspersed repeat families.<\/jats:p><jats:p>Availability: All programs and datasets are downloadable from www.bx.psu.edu\/miller_lab<\/jats:p><jats:p>Contact: \u00a0mhou@cse.psu.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/btm048","type":"journal-article","created":{"date-parts":[[2007,2,19]],"date-time":"2007-02-19T01:05:35Z","timestamp":1171847135000},"page":"917-925","source":"Crossref","is-referenced-by-count":6,"title":["HomologMiner: looking for homologous genomic groups in whole genomes"],"prefix":"10.1093","volume":"23","author":[{"given":"Minmei","family":"Hou","sequence":"first","affiliation":[{"name":"Department of Computer Science & Engineering, Penn State University, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Piotr","family":"Berman","sequence":"additional","affiliation":[{"name":"Department of Computer Science & Engineering, Penn State University, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chih-Hao","family":"Hsu","sequence":"additional","affiliation":[{"name":"Department of Computer Science & Engineering, Penn State University, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Robert S.","family":"Harris","sequence":"additional","affiliation":[{"name":"Department of Computer Science & Engineering, Penn State University, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2007,2,18]]},"reference":[{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","article-title":"Gapped BLAST and PSI-BLAST: a new generation of protein database search programs","volume":"25","author":"Altschul","year":"1997","journal-title":"Nucleic Acids Res"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"789","DOI":"10.1101\/gr.2238404","article-title":"Analysis of segmental duplications and genome assembly in the mouse","volume":"14","author":"Bailey","year":"2004","journal-title":"Genome Res"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"1269","DOI":"10.1101\/gr.88502","article-title":"Automated de novo identification of repeat sequence families in sequenced genomes","volume":"12","author":"Bao","year":"2002","journal-title":"Genome Res"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"i40","DOI":"10.1093\/bioinformatics\/bth946","article-title":"Into the heat of darkness: large-scale clustering of human non-coding DNA","volume":"20","author":"Bejerano","year":"2004","journal-title":"Bioinformatics"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"708","DOI":"10.1101\/gr.1933104","article-title":"Aligning multiple genomic sequences with the threaded blockset aligner","volume":"14","author":"Blanchette","year":"2004","journal-title":"Genome Res"},{"key":"2023041107545543500_","first-page":"175","article-title":"Clustering near-identical sequences for fast homology search","volume":"3909","author":"Cameron","year":"2006","journal-title":"LNBI"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"582","DOI":"10.1093\/bioinformatics\/bti039","article-title":"RAP: a new computer program for de novo identification of repeated sequences in whole genomes","volume":"21","author":"Campagna","year":"2005","journal-title":"Bioinformatics"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"1529","DOI":"10.1126\/science.1116800","article-title":"Fewer genes, more noncoding RNA","volume":"309","author":"Claverie","year":"2005","journal-title":"Science"},{"key":"2023041107545543500_","volume-title":"Introduction to Algorithms.","author":"Cormen","year":"2001","edition":"2nd edn"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"i152","DOI":"10.1093\/bioinformatics\/bti1003","article-title":"PILER: identification and classification of genomic repeats","volume":"21","author":"Edgar","year":"2005","journal-title":"Bioinformatics"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"1575","DOI":"10.1093\/nar\/30.7.1575","article-title":"An efficient algorithm for large-scale detection of protein families","volume":"30","author":"Enright","year":"2002","journal-title":"Nucliec Acids Res"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"243","DOI":"10.1093\/bioinformatics\/btg397","article-title":"Graph-based clustering for finding distant relationships in a large set of protein sequences","volume":"20","author":"Kawaji","year":"2004","journal-title":"Bioinformatics"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"996","DOI":"10.1101\/gr.229102","article-title":"The human genome browser at UCSC","volume":"12","author":"Kent","year":"2002","journal-title":"Genome Res"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"178","DOI":"10.1504\/IJDMB.2006.010855","article-title":"BAG: a graph theoretic sequence clustering algorithm","volume":"1","author":"Kim","year":"2006","journal-title":"Int. J. Data Mining and Bioinformatics"},{"key":"2023041107545543500_","first-page":"545","article-title":"The olfactory receptor universe\u2014from whole genome analysis to structure and evolution","volume":"3","author":"Olender","year":"2004","journal-title":"Genet. Mol. Res"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"458","DOI":"10.1093\/bioinformatics\/16.5.458","article-title":"RSDB: representative protein sequence databases have high information content","volume":"16","author":"Park","year":"2000","journal-title":"Bioinformatics"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"1786","DOI":"10.1101\/gr.2395204","article-title":"De novo repeat classification and fragment assembly","volume":"14","author":"Pevzner","year":"2004","journal-title":"Genome Res"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"i351","DOI":"10.1093\/bioinformatics\/bti1018","article-title":"De novo identification of repeat families in large genomes","volume":"21","author":"Price","year":"2005","journal-title":"Bioinformatics"},{"key":"2023041107545543500_","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1101\/gr.809403","article-title":"Human-mouse alignments with BLASTZ","volume":"13","author":"Schwartz","year":"2004","journal-title":"Genome Res"},{"key":"2023041107545543500_","unstructured":"Smit AFA \u00a0et al. RepeatMasker Open-3.0, http:\/\/www.repeatmasker.org 1996"},{"key":"2023041107545543500_","first-page":"141","article-title":"A simple min cut algorithm","volume":"855","author":"Stoer","year":"1994","journal-title":"LNCS"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/8\/917\/49822144\/bioinformatics_23_8_917.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/8\/917\/49822144\/bioinformatics_23_8_917.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,2,11]],"date-time":"2024-02-11T01:02:38Z","timestamp":1707613358000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/23\/8\/917\/198388"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,2,18]]},"references-count":21,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2007,4,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btm048","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2007,4,15]]},"published":{"date-parts":[[2007,2,18]]}}}