{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T13:43:21Z","timestamp":1760708601167},"reference-count":42,"publisher":"Oxford University Press (OUP)","issue":"2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2014,1,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Homology search methods are dominated by the central paradigm that sequence similarity is a proxy for common ancestry and, by extension, functional similarity. For determining sequence similarity in proteins, most widely used methods use models of sequence evolution and compare amino-acid strings in search for conserved linear stretches. Probabilistic models or sequence profiles capture the position-specific variation in an alignment of homologous sequences and can identify conserved motifs or domains. While profile-based search methods are generally more accurate than simple sequence comparison methods, they tend to be computationally more demanding. In recent years, several methods have emerged that perform protein similarity searches based on domain composition. However, few methods have considered the linear arrangements of domains when conducting similarity searches, despite strong evidence that domain order can harbour considerable functional and evolutionary signal.<\/jats:p>\n               <jats:p>Results: Here, we introduce an alignment scheme that uses a classical dynamic programming approach to the global alignment of domains. We illustrate that representing proteins as strings of domains (domain arrangements) and comparing these strings globally allows for a both fast and sensitive homology search. Further, we demonstrate that the presented methods complement existing methods by finding similar proteins missed by popular amino-acid\u2013based comparison methods.<\/jats:p>\n               <jats:p>Availability: An implementation of the presented algorithms, a web-based interface as well as a command-line program for batch searching against the UniProt database can be found at http:\/\/rads.uni-muenster.de. Furthermore, we provide a JAVA API for programmatic access to domain-string\u2013based search methods.<\/jats:p>\n               <jats:p>Contact: \u00a0terrapon.nicolas@gmail.com or ebb@uni-muenster.de<\/jats:p>\n               <jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btt379","type":"journal-article","created":{"date-parts":[[2013,7,5]],"date-time":"2013-07-05T01:56:51Z","timestamp":1372989411000},"page":"274-281","source":"Crossref","is-referenced-by-count":27,"title":["Rapid similarity search of proteins using alignments of domain arrangements"],"prefix":"10.1093","volume":"30","author":[{"given":"Nicolas","family":"Terrapon","sequence":"first","affiliation":[{"name":"1 Westfalian Wilhelms University, Institute of Evolution and Biodiversity, Huefferstr. 1, 48149 Muenster, Germany and 2Max Planck Institute for Infection Biology, Charit\u00e9platz 1, 10117 Berlin, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"January","family":"Weiner","sequence":"additional","affiliation":[{"name":"1 Westfalian Wilhelms University, Institute of Evolution and Biodiversity, Huefferstr. 1, 48149 Muenster, Germany and 2Max Planck Institute for Infection Biology, Charit\u00e9platz 1, 10117 Berlin, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sonja","family":"Grath","sequence":"additional","affiliation":[{"name":"1 Westfalian Wilhelms University, Institute of Evolution and Biodiversity, Huefferstr. 1, 48149 Muenster, Germany and 2Max Planck Institute for Infection Biology, Charit\u00e9platz 1, 10117 Berlin, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrew D.","family":"Moore","sequence":"additional","affiliation":[{"name":"1 Westfalian Wilhelms University, Institute of Evolution and Biodiversity, Huefferstr. 1, 48149 Muenster, Germany and 2Max Planck Institute for Infection Biology, Charit\u00e9platz 1, 10117 Berlin, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Erich","family":"Bornberg-Bauer","sequence":"additional","affiliation":[{"name":"1 Westfalian Wilhelms University, Institute of Evolution and Biodiversity, Huefferstr. 1, 48149 Muenster, Germany and 2Max Planck Institute for Infection Biology, Charit\u00e9platz 1, 10117 Berlin, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2013,7,4]]},"reference":[{"key":"2023012710391981500_btt379-B1","doi-asserted-by":"crossref","first-page":"D289","DOI":"10.1093\/nar\/gkq1238","article-title":"OMA 2011: orthology inference among 1000 complete genomes","volume":"39","author":"Altenhoff","year":"2011","journal-title":"Nucleic Acids Res."},{"key":"2023012710391981500_btt379-B2","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1016\/S0022-2836(05)80360-2","article-title":"Basic local alignment search tool","volume":"215","author":"Altschul","year":"1990","journal-title":"J. Mol. Biol."},{"key":"2023012710391981500_btt379-B3","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","article-title":"Gapped BLAST and PSI-BLAST: a new generation of protein database search programs","volume":"25","author":"Altschul","year":"1997","journal-title":"Nucleic Acids Res."},{"key":"2023012710391981500_btt379-B4","doi-asserted-by":"crossref","first-page":"1834","DOI":"10.1093\/bioinformatics\/btm240","article-title":"Automated improvement of domain annotations using context analysis of domain arrangements (AIDAN)","volume":"23","author":"Beaussart","year":"2007","journal-title":"Bioinformatics"},{"key":"2023012710391981500_btt379-B5","doi-asserted-by":"crossref","first-page":"911","DOI":"10.1016\/j.jmb.2005.08.067","article-title":"Domain rearrangements in protein evolution","volume":"353","author":"Bj\u00f6rklund","year":"2005","journal-title":"J. Mol. Biol."},{"key":"2023012710391981500_btt379-B6","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1186\/1745-6150-7-12","article-title":"Domain enhanced lookup time accelerated blast","volume":"7","author":"Boratyn","year":"2012","journal-title":"Biol. Direct"},{"key":"2023012710391981500_btt379-B7","doi-asserted-by":"crossref","first-page":"R74","DOI":"10.1186\/gb-2010-11-7-r74","article-title":"Quantifying the mechanisms of domain gain in animal proteins","volume":"11","author":"Buljan","year":"2010","journal-title":"Genome Biol."},{"key":"2023012710391981500_btt379-B8","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1042\/BJ20090122","article-title":"Genomic and structural aspects of protein evolution","volume":"419","author":"Chothia","year":"2009","journal-title":"Biochem. J."},{"key":"2023012710391981500_btt379-B9","doi-asserted-by":"crossref","first-page":"e1002195","DOI":"10.1371\/journal.pcbi.1002195","article-title":"Accelerated profile HMM searches","volume":"7","author":"Eddy","year":"2011","journal-title":"PLoS Comput. Biol."},{"key":"2023012710391981500_btt379-B10","doi-asserted-by":"crossref","first-page":"755","DOI":"10.1093\/bioinformatics\/14.9.755","article-title":"Profile hidden Markov models","volume":"14","author":"Eddy","year":"1998","journal-title":"Bioinformatics"},{"key":"2023012710391981500_btt379-B11","doi-asserted-by":"crossref","first-page":"451","DOI":"10.1093\/bioinformatics\/16.5.451","article-title":"GeneRAGE: a robust algorithm for sequence clustering and domain detection","volume":"16","author":"Enright","year":"2000","journal-title":"Bioinformatics"},{"key":"2023012710391981500_btt379-B12","first-page":"326","article-title":"Domain architecture conservation in orthologs","volume":"12","author":"Forslund","year":"2011","journal-title":"BMC Genomics"},{"key":"2023012710391981500_btt379-B13","doi-asserted-by":"crossref","first-page":"1619","DOI":"10.1101\/gr.278202","article-title":"CDART: protein homology by domain architecture","volume":"12","author":"Geer","year":"2002","journal-title":"Genome Res."},{"key":"2023012710391981500_btt379-B14","doi-asserted-by":"crossref","first-page":"1632","DOI":"10.1101\/gr.183801","article-title":"Annotation transfer for genomics: measuring functional divergence in multi-domain proteins","volume":"11","author":"Gerstein","year":"2001","journal-title":"Genome Res."},{"key":"2023012710391981500_btt379-B15","doi-asserted-by":"crossref","first-page":"2177","DOI":"10.1093\/nar\/gkp1219","article-title":"Homologous over-extension: a challenge for iterative similarity searches","volume":"38","author":"Gonzalez","year":"2010","journal-title":"Nucleic Acids Res."},{"key":"2023012710391981500_btt379-B16","doi-asserted-by":"crossref","first-page":"D306","DOI":"10.1093\/nar\/gkr948","article-title":"Interpro in 2011: new developments in the family and domain prediction database","volume":"40","author":"Hunter","year":"2012","journal-title":"Nucleic Acids Res."},{"key":"2023012710391981500_btt379-B17","doi-asserted-by":"crossref","first-page":"431","DOI":"10.1186\/1471-2105-11-431","article-title":"Hidden Markov model speed heuristic and iterative HMM search procedure","volume":"11","author":"Johnson","year":"2010","journal-title":"BMC Bioinformatics"},{"key":"2023012710391981500_btt379-B18","doi-asserted-by":"crossref","first-page":"5873","DOI":"10.1073\/pnas.90.12.5873","article-title":"Applications and statistics for multiple high-scoring segments in molecular sequences","volume":"90","author":"Karlin","year":"1993","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012710391981500_btt379-B19","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1186\/1471-2105-10-39","article-title":"Protein domain organisation: adding order","volume":"10","author":"Kummerfeld","year":"2009","journal-title":"BMC Bioinformatics"},{"key":"2023012710391981500_btt379-B20","doi-asserted-by":"crossref","first-page":"W60","DOI":"10.1093\/nar\/gkn172","article-title":"DAhunter: a web-based server that identifies homologous proteins by comparing domain architecture","volume":"36","author":"Lee","year":"2008","journal-title":"Nucleic Acids Res."},{"key":"2023012710391981500_btt379-B21","doi-asserted-by":"crossref","first-page":"S5","DOI":"10.1186\/1471-2105-10-S15-S5","article-title":"Protein comparison at the domain architecture level","volume":"10","author":"Lee","year":"2009","journal-title":"BMC Bioinformatics"},{"key":"2023012710391981500_btt379-B22","doi-asserted-by":"crossref","first-page":"2081","DOI":"10.1093\/bioinformatics\/btl366","article-title":"An initial strategy for comparing proteins at the domain architecture level","volume":"22","author":"Lin","year":"2006","journal-title":"Bioinformatics"},{"key":"2023012710391981500_btt379-B23","doi-asserted-by":"crossref","first-page":"D237","DOI":"10.1093\/nar\/gkl951","article-title":"CDD: a conserved domain database for interactive domain family analysis","volume":"35","author":"Marchler-Bauer","year":"2007","journal-title":"Nucleic Acids Res."},{"key":"2023012710391981500_btt379-B24","doi-asserted-by":"crossref","first-page":"444","DOI":"10.1016\/j.tibs.2008.05.008","article-title":"Arrangements in the modular evolution of proteins","volume":"33","author":"Moore","year":"2008","journal-title":"Trends Biochem. Sci."},{"key":"2023012710391981500_btt379-B42","article-title":"DoMosaics: software for domain arrangement visualization and domain-centric analysis of proteins","author":"Moore","year":"2013","journal-title":"Bioinformatics"},{"key":"2023012710391981500_btt379-B25","doi-asserted-by":"crossref","first-page":"443","DOI":"10.1016\/0022-2836(70)90057-4","article-title":"A general method applicable to the search for similarities in the amino acid sequence of two proteins","volume":"48","author":"Needleman","year":"1970","journal-title":"J. Mol. Biol."},{"key":"2023012710391981500_btt379-B26","doi-asserted-by":"crossref","first-page":"867","DOI":"10.1101\/gr.3638405","article-title":"Identification of genomic features using microsyntenies of domains: domain teams","volume":"15","author":"Pasek","year":"2005","journal-title":"Genome Res."},{"key":"2023012710391981500_btt379-B27","doi-asserted-by":"crossref","first-page":"2444","DOI":"10.1073\/pnas.85.8.2444","article-title":"Improved tools for biological sequence comparison","volume":"85","author":"Pearson","year":"1988","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012710391981500_btt379-B28","doi-asserted-by":"crossref","first-page":"D290","DOI":"10.1093\/nar\/gkr1065","article-title":"The pfam protein families database","volume":"40","author":"Punta","year":"2012","journal-title":"Nucleic Acids Res."},{"key":"2023012710391981500_btt379-B29","doi-asserted-by":"crossref","first-page":"2994","DOI":"10.1093\/nar\/29.14.2994","article-title":"Improving the accuracy of PSI-BLAST protein database searches with composition-based statistics and other refinements","volume":"29","author":"Schaffer","year":"2001","journal-title":"Nucleic Acids Res."},{"key":"2023012710391981500_btt379-B30","doi-asserted-by":"crossref","first-page":"413","DOI":"10.1093\/bib\/bbr036","article-title":"Ortholog identification in the presence of domain architecture rearrangement","volume":"12","author":"Sj\u00f6lander","year":"2011","journal-title":"Brief. Bioinform."},{"key":"2023012710391981500_btt379-B31","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1016\/0022-2836(81)90087-5","article-title":"Identification of common molecular subsequences","volume":"147","author":"Smith","year":"1981","journal-title":"J. Mol. Biol."},{"key":"2023012710391981500_btt379-B32","doi-asserted-by":"crossref","DOI":"10.1371\/journal.pcbi.1000063","article-title":"Sequence similarity network reveals common ancestry of multidomain proteins","volume":"4","author":"Song","year":"2008","journal-title":"PLoS Comput. Biol."},{"key":"2023012710391981500_btt379-B33","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1093\/bioinformatics\/14.3.279","article-title":"Statistics of large-scale sequence searching","volume":"14","author":"Spang","year":"1998","journal-title":"Bioinformatics"},{"key":"2023012710391981500_btt379-B34","doi-asserted-by":"crossref","first-page":"66","DOI":"10.1186\/1471-2105-6-66","article-title":"DIALIGN-T: an improved algorithm for segment-based multiple sequence alignment","volume":"6","author":"Subramanian","year":"2005","journal-title":"BMC Bioinformatics"},{"key":"2023012710391981500_btt379-B35","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1042\/BC20060086","article-title":"Current knowledge of the large rhogap family of proteins","volume":"99","author":"Tcherkezian","year":"2007","journal-title":"Biol. Cell"},{"key":"2023012710391981500_btt379-B36","doi-asserted-by":"crossref","first-page":"3077","DOI":"10.1093\/bioinformatics\/btp560","article-title":"Detection of new protein domains by co-occurrence: application to Plasmodium falciparum","volume":"23","author":"Terrapon","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012710391981500_btt379-B37","doi-asserted-by":"crossref","first-page":"D71","DOI":"10.1093\/nar\/gkr981","article-title":"Reorganizing the protein space at the universal protein resource (uniprot)","volume":"40","author":"UniProt Consortium","year":"2012","journal-title":"Nucleic Acids Res."},{"key":"2023012710391981500_btt379-B38","doi-asserted-by":"crossref","first-page":"66","DOI":"10.1016\/j.str.2008.11.008","article-title":"The evolutionary mechanics of domain organization in proteomes and the rise of modularity in the protein world","volume":"17","author":"Wang","year":"2009","journal-title":"Structure"},{"key":"2023012710391981500_btt379-B39","doi-asserted-by":"crossref","first-page":"932","DOI":"10.1093\/bioinformatics\/bti085","article-title":"Rapid motif-based prediction of circular permutations in multi-domain proteins","volume":"21","author":"Weiner","year":"2005","journal-title":"Bioinformatics"},{"key":"2023012710391981500_btt379-B40","doi-asserted-by":"crossref","first-page":"2037","DOI":"10.1111\/j.1742-4658.2006.05220.x","article-title":"Domain deletions and substitutions in the modular protein evolution","volume":"273","author":"Weiner","year":"2006","journal-title":"FEBS J."},{"key":"2023012710391981500_btt379-B41","doi-asserted-by":"crossref","first-page":"343","DOI":"10.1126\/science.1178028","article-title":"Functional and evolutionary insights from the genomes of three parasitoid nasonia species","volume":"327","author":"Werren","year":"2010","journal-title":"Science"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/2\/274\/48914637\/bioinformatics_30_2_274.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/2\/274\/48914637\/bioinformatics_30_2_274.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T10:46:20Z","timestamp":1674816380000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/30\/2\/274\/215388"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,7,4]]},"references-count":42,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2014,1,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btt379","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2014,1,15]]},"published":{"date-parts":[[2013,7,4]]}}}