{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T10:12:34Z","timestamp":1761559954760},"reference-count":13,"publisher":"Springer Science and Business Media LLC","issue":"S2","license":[{"start":{"date-parts":[[2005,7,1]],"date-time":"2005-07-01T00:00:00Z","timestamp":1120176000000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/2.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2005,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>The continuous flow of EST data remains one of the richest sources for discoveries in modern biology. The first step in EST data mining is usually associated with EST clustering, the process of grouping of original fragments according to their annotation, similarity to known genomic DNA or each other. Clustered EST data, accumulated in databases such as UniGene, STACK and TIGR Gene Indices have proven to be crucial in research areas from gene discovery to regulation of gene expression.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>We have developed a new nucleotide sequence matching algorithm and its implementation for clustering EST sequences. The program is based on the original CLU match detection algorithm, which has improved performance over the widely used d2_cluster. The CLU algorithm automatically ignores low-complexity regions like poly-tracts and short tandem repeats.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion<\/jats:title>\n            <jats:p>CLU represents a new generation of EST clustering algorithm with improved performance over current approaches. An early implementation can be applied in small and medium-size projects. The CLU program is available on an open source basis free of charge. It can be downloaded from <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"http:\/\/compbio.pbrc.edu\/pti\" ext-link-type=\"uri\">http:\/\/compbio.pbrc.edu\/pti<\/jats:ext-link>\n            <\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-6-s2-s3","type":"journal-article","created":{"date-parts":[[2005,7,16]],"date-time":"2005-07-16T09:00:41Z","timestamp":1121504441000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["CLU: A new algorithm for EST clustering"],"prefix":"10.1186","volume":"6","author":[{"given":"Andrey","family":"Ptitsyn","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Winston","family":"Hide","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2005,7,15]]},"reference":[{"key":"661_CR1","doi-asserted-by":"publisher","first-page":"95","DOI":"10.1007\/BF02789291","volume":"2","author":"T Kapros","year":"1994","unstructured":"Kapros T, Robertson AJ, Waterborg JH: A simple method to make better probes from short DNA fragments. Mol Biotechnol 1994, 2: 95\u20138.","journal-title":"Mol Biotechnol"},{"key":"661_CR2","doi-asserted-by":"publisher","first-page":"151","DOI":"10.1006\/geno.1996.0177","volume":"33","author":"GG Lennon","year":"1996","unstructured":"Lennon GG, Auffray C, Polymeropulos M, Soares MB: The I.M.A.G.E. Consortium: An integrated molecular analysis of genomes and their expression. Genomics 1996, 33: 151\u2013152. 10.1006\/geno.1996.0177","journal-title":"Genomics"},{"key":"661_CR3","doi-asserted-by":"publisher","first-page":"332","DOI":"10.1038\/ng0893-332","volume":"4","author":"MS Boguski","year":"1993","unstructured":"Boguski MS, Lowe TM, Tolstoshev CM: dbEST-database for \"expressed sequence tags\". Nature Genetics 1993, 4: 332\u20133. 10.1038\/ng0893-332","journal-title":"Nature Genetics"},{"key":"661_CR4","doi-asserted-by":"publisher","first-page":"141","DOI":"10.1093\/nar\/28.1.141","volume":"28","author":"J Quackenbush","year":"2000","unstructured":"Quackenbush J, Liang F, Holt I, Pertea G, Upton J: The TIGR Gene Indices: reconstruction and representation of expressed gene sequences. Nucleic Acids Research 2000, 28: 141\u2013145. 10.1093\/nar\/28.1.141","journal-title":"Nucleic Acids Research"},{"key":"661_CR5","doi-asserted-by":"publisher","first-page":"369","DOI":"10.1038\/ng0895-369","volume":"10","author":"MS Boguski","year":"1995","unstructured":"Boguski MS, Schuler GD: ESTablishing a human transcript map. Nature Genetics 1995, 10: 369\u2013371. 10.1038\/ng0895-369","journal-title":"Nature Genetics"},{"key":"661_CR6","doi-asserted-by":"publisher","first-page":"965","DOI":"10.1093\/bioinformatics\/15.12.965","volume":"15","author":"M Cariaso","year":"1999","unstructured":"Cariaso M, Folta P, Wagner M, Kuczmarski T, Lennon G: IMAGEne I: clustering and ranking of I.M.A.G.E. cDNA clones corresponding to known genes. Bioinformatics 1999, 15: 965\u2013973. 10.1093\/bioinformatics\/15.12.965","journal-title":"Bioinformatics"},{"key":"661_CR7","first-page":"61","volume":"7","author":"AR Williamson","year":"1995","unstructured":"Williamson AR, Elliston KO, Sturchio JL: The Merck Gene Index, a public resource for genomic research. J NIH Res 1995, 7: 61\u201363.","journal-title":"J NIH Res"},{"key":"661_CR8","doi-asserted-by":"publisher","first-page":"1143","DOI":"10.1101\/gr.9.11.1143","volume":"9","author":"RT Miller","year":"1999","unstructured":"Miller RT, Christoffels AG, Gopalakrishnan C, Burke J, Ptitsyn AA, Broveak TR, Hide WA: A comprehensive approach to clustering of expressed human gene sequence: the sequence tag alignment and consensus knowledge base. Genome Res 1999, 9: 1143\u201355. 10.1101\/gr.9.11.1143","journal-title":"Genome Res"},{"key":"661_CR9","doi-asserted-by":"publisher","first-page":"2963","DOI":"10.1093\/nar\/gkg379","volume":"31","author":"A Kalyanaraman","year":"2003","unstructured":"Kalyanaraman A, Aluru S, Kothari S, Brendel V: Efficient clustering of large EST data sets on parallel computers. Nucleic Acids Res 2003, 31: 2963\u201374. 10.1093\/nar\/gkg379","journal-title":"Nucleic Acids Res"},{"key":"661_CR10","first-page":"319","volume":"10","author":"VB Strelets","year":"1994","unstructured":"Strelets VB, Ptitsyn AA, Milanesi L, Lim HA: Data bank homology search algorithm with linear computation complexity. Comput Appl Biosci 1994, 10: 319\u201322.","journal-title":"Comput Appl Biosci"},{"key":"661_CR11","first-page":"529","volume":"8","author":"VB Streletc","year":"1992","unstructured":"Streletc VB, Shindyalov IN, Kolchanov NA, Milanesi L: Fast, statistically based alignment of amino acid sequences on the base of diagonal fragments of DOT-matrices. Comput Appl Biosci 1992, 8: 529\u201334.","journal-title":"Comput Appl Biosci"},{"key":"661_CR12","doi-asserted-by":"publisher","first-page":"1221","DOI":"10.1093\/bioinformatics\/btg138","volume":"19","author":"K Malde","year":"2003","unstructured":"Malde K, Coward E, Jonassen I: Fast sequence clustering using a suffix array algorithm. Bioinformatics 2003, 19: 1221\u20136. 10.1093\/bioinformatics\/btg138","journal-title":"Bioinformatics"},{"key":"661_CR13","doi-asserted-by":"publisher","first-page":"1135","DOI":"10.1101\/gr.9.11.1135","volume":"9","author":"J Burke","year":"1999","unstructured":"Burke J, Davison D, Hide W: d2_cluster: a validated method for clustering EST and full-length cDNA sequences. Genome Res 1999, 9: 1135\u201342. 10.1101\/gr.9.11.1135","journal-title":"Genome Res"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-6-S2-S3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/1471-2105-6-S2-S3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-6-S2-S3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,8,31]],"date-time":"2021-08-31T21:26:50Z","timestamp":1630445210000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-6-S2-S3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,7]]},"references-count":13,"journal-issue":{"issue":"S2","published-print":{"date-parts":[[2005,7]]}},"alternative-id":["661"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-6-s2-s3","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2005,7]]},"assertion":[{"value":"15 July 2005","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"S3"}}