{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T13:29:19Z","timestamp":1760707759757},"reference-count":2,"publisher":"Oxford University Press (OUP)","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2007,2,15]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: The size of current protein databases is a challenge for many Bioinformatics applications, both in terms of processing speed and information redundancy. It may be therefore desirable to efficiently reduce the database of interest to a maximally representative subset.<\/jats:p><jats:p>Results: The MinSet method employs a combination of a Suffix Tree and a Genetic Algorithm for the generation, selection and assessment of database subsets. The approach is generally applicable to any type of string-encoded data, allowing for a drastic reduction of the database size whilst retaining most of the information contained in the original set. We demonstrate the performance of the method on a database of protein domain structures encoded as strings. We used the SCOP40 domain database by translating protein structures into character strings by means of a structural alphabet and by extracting optimized subsets according to an entropy score that is based on a constant-length fragment dictionary. Therefore, optimized subsets are maximally representative for the distribution and range of local structures. Subsets containing only 10% of the SCOP structure classes show a coverage of &amp;gt;90% for fragments of length 1\u20134.<\/jats:p><jats:p>Availability: \u00a0<\/jats:p><jats:p>Contact: \u00a0jkleinj@nimr.mrc.ac.uk<\/jats:p><jats:p>Supplementary information: Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btl637","type":"journal-article","created":{"date-parts":[[2007,1,5]],"date-time":"2007-01-05T01:49:58Z","timestamp":1167961798000},"page":"515-516","source":"Crossref","is-referenced-by-count":11,"title":["MinSet: a general approach to derive maximally representative database subsets by using fragment dictionaries and its application to the SCOP database"],"prefix":"10.1093","volume":"23","author":[{"given":"Alessandro","family":"Pandini","sequence":"first","affiliation":[{"name":"Dipartimento di Scienze dell'Ambiente e del Territorio, Universit\u00e0 degli Studi di Milano-Bicocca 1 \u00a0 1 \u00a0 \u00a0 Milano, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Laura","family":"Bonati","sequence":"additional","affiliation":[{"name":"Dipartimento di Scienze dell'Ambiente e del Territorio, Universit\u00e0 degli Studi di Milano-Bicocca 1 \u00a0 1 \u00a0 \u00a0 Milano, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Franca","family":"Fraternali","sequence":"additional","affiliation":[{"name":"Bioinformatics Unit, King's College 2 \u00a0 2 \u00a0 \u00a0 London, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jens","family":"Kleinjung","sequence":"additional","affiliation":[{"name":"Division of Mathematical Biology, National Institute for Medical Research 3 \u00a0 3 \u00a0 \u00a0 The Ridgeway, London NW7 1AA, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2007,1,3]]},"reference":[{"key":"2023041109272593400_","doi-asserted-by":"crossref","first-page":"591","DOI":"10.1016\/j.jmb.2004.04.005","article-title":"A hidden markov model derived structural alphabet for proteins","volume":"339","author":"Camproux","year":"2004","journal-title":"J. Mol. Biol."},{"key":"2023041109272593400_","doi-asserted-by":"crossref","first-page":"D189","DOI":"10.1093\/nar\/gkh034","article-title":"The ASTRAL compendium in 2004","volume":"32","author":"Chandonia","year":"2004","journal-title":"Nucleic Acids Res."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/4\/515\/49830046\/bioinformatics_23_4_515.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/4\/515\/49830046\/bioinformatics_23_4_515.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,10]],"date-time":"2023-05-10T08:17:49Z","timestamp":1683706669000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/23\/4\/515\/182282"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,1,3]]},"references-count":2,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2007,2,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btl637","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2007,2,15]]},"published":{"date-parts":[[2007,1,3]]}}}