{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,8,23]],"date-time":"2024-08-23T09:17:33Z","timestamp":1724404653050},"reference-count":0,"publisher":"Oxford University Press (OUP)","issue":"2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2004,1,22]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Clustering of protein sequences is widely used for the functional characterization of proteins. However, it is still not easy to cluster distantly-related proteins, which have only regional similarity among their sequences. It is therefore necessary to develop an algorithm for clustering such distantly-related proteins.<\/jats:p>\n               <jats:p>Results: We have developed a time and space efficient clustering algorithm. It uses a graph representation where its vertices and edges denote proteins and their sequence similarities above a certain cutoff score, respectively. It repeatedly partitions the graph by removing edges that have small weights, which correspond to low sequence similarities. To find the appropriate partitions, we introduce a score combining the normalized cut and a locally minimal cut capacities. Our method is applied to the entire 40 703 human proteins in SWISS-PROT and TrEMBL. The resulting clusters shows a 76% recall (20 529 proteins) of the 26 917 classified by InterPro. It also finds relationships not found by other clustering methods.<\/jats:p>\n               <jats:p>Availability: The complete result of our algorithm for all the human proteins in SWISS-PROT and TrEMBL, and other supplementary information are available at http:\/\/motif.ics.es.osaka-u.ac.jp\/Ncut-KL\/<\/jats:p>","DOI":"10.1093\/bioinformatics\/btg397","type":"journal-article","created":{"date-parts":[[2004,1,20]],"date-time":"2004-01-20T21:07:58Z","timestamp":1074632878000},"page":"243-252","source":"Crossref","is-referenced-by-count":29,"title":["Graph-based clustering for finding distant relationships in a large set of protein sequences"],"prefix":"10.1093","volume":"20","author":[{"given":"Hideya","family":"Kawaji","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yoichi","family":"Takenaka","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hideo","family":"Matsuda","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2004,1,22]]},"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/20\/2\/243\/48905271\/bioinformatics_20_2_243.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/20\/2\/243\/48905271\/bioinformatics_20_2_243.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,25]],"date-time":"2023-01-25T18:46:44Z","timestamp":1674672404000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/20\/2\/243\/204884"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2004,1,22]]},"references-count":0,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2004,1,22]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btg397","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2004,1,22]]},"published":{"date-parts":[[2004,1,22]]}}}