{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T18:33:26Z","timestamp":1784140406384,"version":"3.55.0"},"reference-count":26,"publisher":"Oxford University Press (OUP)","issue":"20","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2009,10,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Over the last years a number of evidences have been accumulated about high incidence of tandem repeats in proteins carrying fundamental biological functions and being related to a number of human diseases. At the same time, frequently, protein repeats are strongly degenerated during evolution and, therefore, cannot be easily identified. To solve this problem, several computer programs which were based on different algorithms have been developed. Nevertheless, our tests showed that there is still room for improvement of methods for accurate and rapid detection of tandem repeats in proteins.<\/jats:p>\n               <jats:p>Results: We developed a new program called T-REKS for ab initio identification of the tandem repeats. It is based on clustering of lengths between identical short strings by using a K-means algorithm. Benchmark of the existing programs and T-REKS on several sequence datasets is presented. Our program being linked to the Protein Repeat DataBase opens the way for large-scale analysis of protein tandem repeats. T-REKS can also be applied to the nucleotide sequences.<\/jats:p>\n               <jats:p>Availability: The algorithm has been implemented in JAVA, the program is available upon request at http:\/\/bioinfo.montp.cnrs.fr\/?r=t-reks. Protein Repeat DataBase generated by using T-REKS is accessible at http:\/\/bioinfo.montp.cnrs.fr\/?r=repeatDB.<\/jats:p>\n               <jats:p>Contact: \u00a0julien.jorda@crbm.cnrs.fr; andrey.kajava@crbm.cnrs.fr<\/jats:p>\n               <jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btp482","type":"journal-article","created":{"date-parts":[[2009,8,12]],"date-time":"2009-08-12T03:55:08Z","timestamp":1250049308000},"page":"2632-2638","source":"Crossref","is-referenced-by-count":174,"title":["T-REKS: identification of Tandem REpeats in sequences with a K-meanS based algorithm"],"prefix":"10.1093","volume":"25","author":[{"given":"Julien","family":"Jorda","sequence":"first","affiliation":[{"name":"Centre de Recherches de Biochimie Macromol\u00e9culaire UMR 5237, CNRS, University of Montpellier 1 and 2, Montpellier, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrey V.","family":"Kajava","sequence":"additional","affiliation":[{"name":"Centre de Recherches de Biochimie Macromol\u00e9culaire UMR 5237, CNRS, University of Montpellier 1 and 2, Montpellier, France"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2009,8,11]]},"reference":[{"key":"2023013112122356900_B1","doi-asserted-by":"crossref","first-page":"521","DOI":"10.1006\/jmbi.2000.3684","article-title":"Homology-based method for identification of protein repeats using statistical significance estimates","volume":"298","author":"Andrade","year":"2000","journal-title":"J. Mol. Biol."},{"key":"2023013112122356900_B2","doi-asserted-by":"crossref","first-page":"125","DOI":"10.1016\/S0065-3233(06)73005-4","article-title":"Structure, function, and amyloidogenesis of fungal prions: filament polymorphism and prion variants","volume":"73","author":"Baxa","year":"2006","journal-title":"Adv. Protein Chem."},{"key":"2023013112122356900_B3","doi-asserted-by":"crossref","first-page":"573","DOI":"10.1093\/nar\/27.2.573","article-title":"Tandem repeats finder: a program to analyze DNA sequences","volume":"27","author":"Benson","year":"1999","journal-title":"Nucleic Acids Res."},{"key":"2023013112122356900_B4","doi-asserted-by":"crossref","first-page":"2812","DOI":"10.1093\/bioinformatics\/bth335","article-title":"STAR: an algorithm to search for Tandem Approximate Repeats","volume":"20","author":"Delgrange","year":"2004","journal-title":"Bioinformatics"},{"key":"2023013112122356900_B5","doi-asserted-by":"crossref","first-page":"1792","DOI":"10.1093\/nar\/gkh340","article-title":"MUSCLE: multiple sequence alignment with high accuracy and high throughput","volume":"32","author":"Edgar","year":"2004","journal-title":"Nucleic Acids Res."},{"key":"2023013112122356900_B6","doi-asserted-by":"crossref","first-page":"3784","DOI":"10.1093\/nar\/gkg563","article-title":"ExPASy: The proteomics server for in-depth protein knowledge and analysis","volume":"31","author":"Gasteiger","year":"2003","journal-title":"Nucleic Acids Res."},{"key":"2023013112122356900_B7","doi-asserted-by":"crossref","first-page":"515","DOI":"10.1016\/S0968-0004(00)01643-1","article-title":"The REPRO server: finding protein internal sequence repeats through the Web","volume":"25","author":"George","year":"2000","journal-title":"Trends Biochem Sci."},{"key":"2023013112122356900_B8","doi-asserted-by":"crossref","first-page":"4355","DOI":"10.1073\/pnas.84.13.4355","article-title":"Profile analysis: detection of distantly related proteins","volume":"84","author":"Gribskov","year":"1987","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023013112122356900_B9","doi-asserted-by":"crossref","first-page":"147","DOI":"10.1002\/j.1538-7305.1950.tb00463.x","article-title":"Error detecting and error correcting codes","volume":"29","author":"Hamming","year":"1950","journal-title":"Bell System Technical J."},{"key":"2023013112122356900_B10","doi-asserted-by":"crossref","first-page":"224","DOI":"10.1002\/1097-0134(20001101)41:2<224::AID-PROT70>3.0.CO;2-Z","article-title":"Rapid automatic detection and alignment of repeats in protein sequences","volume":"41","author":"Heger","year":"2000","journal-title":"Proteins"},{"key":"2023013112122356900_B11","doi-asserted-by":"crossref","first-page":"241","DOI":"10.1007\/BF02289588","article-title":"Hierarchical clustering schemes","volume":"32","author":"Johnson","year":"1967","journal-title":"Psychometrika"},{"key":"2023013112122356900_B12","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/S0065-3233(06)73001-7","article-title":"Beta-structures in fibrous proteins","volume":"73","author":"Kajava","year":"2006","journal-title":"Adv. Protein Chem."},{"key":"2023013112122356900_B13","doi-asserted-by":"crossref","first-page":"306","DOI":"10.1016\/j.jsb.2006.01.015","article-title":"The turn of the screw: variations of the abundant beta-solenoid motif in passenger domains of Type V secretory proteins","volume":"155","author":"Kajava","year":"2006","journal-title":"J. Struct. Biol."},{"key":"2023013112122356900_B14","doi-asserted-by":"crossref","first-page":"867","DOI":"10.1016\/S0969-2126(01)00222-2","article-title":"Modeling of the three-dimensional structure of proteins with the typical leucine-rich repeats","volume":"3","author":"Kajava","year":"1995","journal-title":"Structure"},{"key":"2023013112122356900_B15","doi-asserted-by":"crossref","first-page":"1203","DOI":"10.1110\/ps.9.6.1203","article-title":"Amino acid repeat patterns in protein sequences: their diversity and structural-functional implications","volume":"9","author":"Katti","year":"2000","journal-title":"Protein Sci."},{"key":"2023013112122356900_B16","doi-asserted-by":"crossref","first-page":"3672","DOI":"10.1093\/nar\/gkg617","article-title":"mreps: efficient and flexible detection of tandem repeats in DNA","volume":"31","author":"Kolpakov","year":"2003","journal-title":"Nucleic Acids Res."},{"key":"2023013112122356900_B17","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1089\/106652701300099038","article-title":"An algorithm for approximate tandem repeats","volume":"8","author":"Landau","year":"2001","journal-title":"J. Comput. Biol."},{"key":"2023013112122356900_B18","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1016\/S0968-0004(97)01058-X","article-title":"A repetitive sequence in subunits of the 26S proteasome and 20S cyclosome (anaphase-promoting complex)","volume":"22","author":"Lupas","year":"1997","journal-title":"Trends Biochem Sci."},{"key":"2023013112122356900_B19","article-title":"Some methods for classification and analysis of multivariate observations","volume-title":"Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability.","author":"MacQueen","year":"1967"},{"key":"2023013112122356900_B20","doi-asserted-by":"crossref","first-page":"151","DOI":"10.1006\/jmbi.1999.3136","article-title":"A census of protein repeats","volume":"293","author":"Marcotte","year":"1999","journal-title":"J. Mol. Biol."},{"key":"2023013112122356900_B21","doi-asserted-by":"crossref","first-page":"235","DOI":"10.1016\/S0065-3233(06)73008-X","article-title":"Structural models of amyloid-like fibrils","volume":"73","author":"Nelson","year":"2006","journal-title":"Adv. Protein Chem."},{"key":"2023013112122356900_B22","doi-asserted-by":"crossref","first-page":"382","DOI":"10.1186\/1471-2105-8-382","article-title":"XSTREAM: a practical algorithm for identification and architecture modeling of tandem repeats in protein sequences","volume":"8","author":"Newman","year":"2007","journal-title":"BMC Bioinformatics"},{"key":"2023013112122356900_B23","doi-asserted-by":"crossref","first-page":"276","DOI":"10.1016\/S0168-9525(00)02024-2","article-title":"EMBOSS: the European Molecular Biology Open Software Suite","volume":"16","author":"Rice","year":"2000","journal-title":"Trends Genet."},{"key":"2023013112122356900_B24","doi-asserted-by":"crossref","first-page":"e30","DOI":"10.1093\/bioinformatics\/btl309","article-title":"Tandem repeats over the edit distance","volume":"23","author":"Sokol","year":"2007","journal-title":"Bioinformatics"},{"issue":"Suppl. 1","key":"2023013112122356900_B25","doi-asserted-by":"crossref","first-page":"i311","DOI":"10.1093\/bioinformatics\/bth911","article-title":"Tracking repeats using significance and transitivity","volume":"20","author":"Szklarczyk","year":"2004","journal-title":"Bioinformatics"},{"key":"2023013112122356900_B26","doi-asserted-by":"crossref","first-page":"4673","DOI":"10.1093\/nar\/22.22.4673","article-title":"CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice","volume":"22","author":"Thompson","year":"1994","journal-title":"Nucleic Acids Res."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/25\/20\/2632\/48993465\/bioinformatics_25_20_2632.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/25\/20\/2632\/48993465\/bioinformatics_25_20_2632.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,31]],"date-time":"2023-01-31T21:37:10Z","timestamp":1675201030000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/25\/20\/2632\/193638"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,8,11]]},"references-count":26,"journal-issue":{"issue":"20","published-print":{"date-parts":[[2009,10,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btp482","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2009,10,15]]},"published":{"date-parts":[[2009,8,11]]}}}