{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,5,14]],"date-time":"2023-05-14T02:40:45Z","timestamp":1684032045263},"reference-count":35,"publisher":"Oxford University Press (OUP)","issue":"18","license":[{"start":{"date-parts":[[2016,10,2]],"date-time":"2016-10-02T00:00:00Z","timestamp":1475366400000},"content-version":"vor","delay-in-days":3305,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/2.0\/uk\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2007,9,15]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: Propagating functional annotations to sequence-similar, presumably homologous proteins lies at the heart of the bioinformatics industry. Correct propagation is crucially dependent on the accurate identification of subtle sequence motifs that are conserved in evolution. The evolutionary signal can be difficult to detect because functional sites may consist of non-contiguous residues while segments in-between may be mutated without affecting fold or function.<\/jats:p><jats:p>Results: Here, we report a novel graph clustering algorithm in which all known protein sequences simultaneously self-organize into hypothetical multiple sequence alignments. This eliminates noise so that non-contiguous sequence motifs can be tracked down between extremely distant homologues. The novel data structure enables fast sequence database searching methods which are superior to profile-profile comparison at recognizing distant homologues. This study will boost the leverage of structural and functional genomics and opens up new avenues for data mining a complete set of functional signature motifs.<\/jats:p><jats:p>Availability: \u00a0http:\/\/www.bioinfo.biocenter.helsinki.fi\/gtg<\/jats:p><jats:p>Contact: \u00a0liisa.holm@helsinki.fi<\/jats:p><jats:p>Supplementary information: Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btm358","type":"journal-article","created":{"date-parts":[[2007,9,7]],"date-time":"2007-09-07T00:14:25Z","timestamp":1189124065000},"page":"2361-2367","source":"Crossref","is-referenced-by-count":19,"title":["The global trace graph, a novel paradigm for searching protein sequence databases"],"prefix":"10.1093","volume":"23","author":[{"given":"Andreas","family":"Heger","sequence":"first","affiliation":[{"name":"1 Institute of Biotechnology and 2Department of Biological and Environmental Sciences, Division of Genetics, P.O. Box 56 (Viikinkaari 5), FI-00014 University of Helsinki, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Swapan","family":"Mallick","sequence":"additional","affiliation":[{"name":"1 Institute of Biotechnology and 2Department of Biological and Environmental Sciences, Division of Genetics, P.O. Box 56 (Viikinkaari 5), FI-00014 University of Helsinki, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christopher","family":"Wilton","sequence":"additional","affiliation":[{"name":"1 Institute of Biotechnology and 2Department of Biological and Environmental Sciences, Division of Genetics, P.O. Box 56 (Viikinkaari 5), FI-00014 University of Helsinki, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liisa","family":"Holm","sequence":"additional","affiliation":[{"name":"1 Institute of Biotechnology and 2Department of Biological and Environmental Sciences, Division of Genetics, P.O. Box 56 (Viikinkaari 5), FI-00014 University of Helsinki, Finland"},{"name":"1 Institute of Biotechnology and 2Department of Biological and Environmental Sciences, Division of Genetics, P.O. Box 56 (Viikinkaari 5), FI-00014 University of Helsinki, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2007,9,15]]},"reference":[{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"555","DOI":"10.1016\/0022-2836(91)90193-A","article-title":"Amino acid matrices from an information theoretic perspective","volume":"219","author":"Altschul","year":"1991","journal-title":"J. Mol. Biol"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"D226","DOI":"10.1093\/nar\/gkh039","article-title":"SCOP database in 2004: refinements integrate structure and sequence family data","volume":"32","author":"Andreeva","year":"2004","journal-title":"Nucleic Acids Res"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"D138","DOI":"10.1093\/nar\/gkh121","article-title":"The Pfam protein families database","volume":"32","author":"Bateman","year":"2004","journal-title":"Nucleic Acids Res"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"1456","DOI":"10.1093\/bioinformatics\/btl102","article-title":"A machine learning information retrieval approach to protein fold recognition","volume":"522","author":"Cheng","year":"2006","journal-title":"Bioinformatics"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"953","DOI":"10.1038\/nsb1101-953","article-title":"Identification of homology in protein structure classifiction","volume":"8","author":"Dietmann","year":"2001","journal-title":"Nat. Struct Biol"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"330","DOI":"10.1101\/gr.2821705","article-title":"ProbCons: probabilistic consistency-based multiple sequence alignment","volume":"15","author":"Do","year":"2005","journal-title":"Genome Res"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"755","DOI":"10.1093\/bioinformatics\/14.9.755","article-title":"Profile hidden Markov models","volume":"14","author":"Eddy","year":"1998","journal-title":"Bioinformatics"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"1243","DOI":"10.1093\/bioinformatics\/18.9.1243","article-title":"The use of structure information to increase alignment accuracy does not aid homologue detection with profile HMMs","volume":"18","author":"Griffith-Jones","year":"2002","journal-title":"Bioinformatics"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1016\/S0079-6107(00)00013-4","article-title":"Towards a covering set of protein family profiles","volume":"73","author":"Heger","year":"2000","journal-title":"Prog. Biophys"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1023\/A:1026145703834","article-title":"More for less in structural genomics","volume":"4","author":"Heger","year":"2003","journal-title":"J. Struct. Funct. Genomics"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"749","DOI":"10.1016\/S0022-2836(03)00269-9","article-title":"Exhaustive enumeration of protein domain families","volume":"328","author":"Heger","year":"2003","journal-title":"J. Mol. Biol"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"i130","DOI":"10.1093\/bioinformatics\/btg1017","article-title":"Sensitive pattern discovery with \u2018fuzzy\u2019 alignments of distantly related proteins","volume":"19","author":"Heger","year":"2003","journal-title":"Bioinformatics"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"843","DOI":"10.1089\/cmb.2004.11.843","article-title":"Accurate detection of very sparse sequence motifs","volume":"11","author":"Heger","year":"2004","journal-title":"J. Comput. Biol"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"D188","DOI":"10.1093\/nar\/gki096","article-title":"ADDA: a domain database with global coverage of the protein universe","volume":"33","author":"Heger","year":"2005","journal-title":"Nucl. Acids Res"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"566","DOI":"10.1093\/bioinformatics\/16.6.566","article-title":"DaliLite workbench for protein structure comparison","volume":"16","author":"Holm","year":"2000","journal-title":"Bioinformatics"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1002\/(SICI)1097-0134(199705)28:1<72::AID-PROT7>3.0.CO;2-L","article-title":"An evolutionary treasure: unification of a broad set of amidohydrolases related to urease","volume":"28","author":"Holm","year":"1997","journal-title":"Proteins"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"D216","DOI":"10.1093\/nar\/gki007","article-title":"ProtoNet 4.0: a hierarchical classification of one million protein sequences","volume":"33","author":"Kaplan","year":"2005","journal-title":"Nucleic Acids Res"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"641","DOI":"10.1093\/protein\/gzg081","article-title":"PROSPECT II: protein structure prediction program for the genome-scale","volume":"16","author":"Kim","year":"2003","journal-title":"Protein Eng"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"613","DOI":"10.1006\/jmbi.1999.3377","article-title":"Identification of related proteins on family, superfamily and fold level","volume":"295","author":"Lindahl","year":"2000","journal-title":"J. Mol. Biol"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"2466","DOI":"10.1093\/bioinformatics\/btl411","article-title":"Bayesian search of functionally divergent protein subgroups and their function specific residues","volume":"22","author":"Marttinen","year":"2006","journal-title":"Bioinformatics"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"627","DOI":"10.1016\/j.tibs.2004.10.006","article-title":"Patterns and clusters within the PSM column in TiBS, 1992\u20132004","volume":"29","author":"McEntyre","year":"2004","journal-title":"Trends Biochem. Sci"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"1665","DOI":"10.1093\/nar\/25.9.1665","article-title":"Extracting protein alignment models from the sequence database","volume":"25","author":"Neuwald","year":"1997","journal-title":"Nucleic Acids Res"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"407","DOI":"10.1093\/bioinformatics\/14.5.407","article-title":"COFFEE: an objective function for multiple sequence alignments","volume":"14","author":"Notredame","year":"1998","journal-title":"Bioinformatics"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"205","DOI":"10.1006\/jmbi.2000.4042","article-title":"T-Coffee: a novel method for fast and accurate multiple sequence alignment","volume":"302","author":"Notredame","year":"2000","journal-title":"J. Mol. Biol"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"1201","DOI":"10.1006\/jmbi.1998.2221","article-title":"Sequence comparisons using multiple sequences detect three times as many remote homologues as pairwise methods","volume":"284","author":"Park","year":"1998","journal-title":"J. Mol. Biol"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"818","DOI":"10.1093\/bioinformatics\/btg485","article-title":"Quality of alignment comparison by COMPASS improves with inclusion of diverse confident homologs","volume":"20","author":"Sadreyev","year":"2004","journal-title":"Bioinformatics"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1002\/prot.340090107","article-title":"Database of homology-derived protein structures and the structural meaning of sequence alignment","volume":"9","author":"Sander","year":"1991","journal-title":"Proteins"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"2994","DOI":"10.1093\/nar\/29.14.2994","article-title":"Improving the accuracy of PSI-BLAST protein database searches with composition-based statistics and other refinements","volume":"29","author":"Schaffer","year":"2001","journal-title":"Nucleic Acids Res"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"243","DOI":"10.1006\/jmbi.2001.4762","article-title":"FUGUE: sequence-structure homology recognition using environment-specific substitution tables and structure-dependent gap penalties","volume":"310","author":"Shi","year":"2001","journal-title":"J. Mol. Biol"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1152\/physiolgenomics.00166.2005","article-title":"From sequences to a functional unit","volume":"25","author":"Sivakumar","year":"2006","journal-title":"Physiol. Genomics"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"951","DOI":"10.1093\/bioinformatics\/bti125","article-title":"Protein homology detection by HMM-HMM comparison","volume":"21","author":"Soding","year":"2005","journal-title":"Bioinformatics"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"275","DOI":"10.1016\/j.sbi.2005.04.003","article-title":"Predicting protein function from sequence and structural data","volume":"15","author":"Watson","year":"2005","journal-title":"Curr. Opin. Struct. Biol"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"3986","DOI":"10.1093\/nar\/26.17.3986","article-title":"Protein sequence similarity searches using patterns as seeds","volume":"26","author":"Zhang","year":"1998","journal-title":"Nucleic Acids Res"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"1005","DOI":"10.1002\/prot.20007","article-title":"Single-body residue-level knowledge-based energy score combined with sequence-profile and secondary structure information for fold recognition","volume":"35","author":"Zhou","year":"2004","journal-title":"Proteins"},{"key":"2023041106220923100_","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1002\/prot.20308","article-title":"Fold recognition by combining sequence profiles derived from evolution and from depth-dependent structural alignment of fragments","volume":"58","author":"Zhou","year":"2005","journal-title":"Proteins"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/18\/2361\/49816913\/bioinformatics_23_18_2361.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/18\/2361\/49816913\/bioinformatics_23_18_2361.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,14]],"date-time":"2023-05-14T01:58:46Z","timestamp":1684029526000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/23\/18\/2361\/237883"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,9,15]]},"references-count":35,"journal-issue":{"issue":"18","published-print":{"date-parts":[[2007,9,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btm358","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2007,9,15]]},"published":{"date-parts":[[2007,9,15]]}}}