{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T10:12:30Z","timestamp":1761559950302},"reference-count":35,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2004,10,25]],"date-time":"2004-10-25T00:00:00Z","timestamp":1098662400000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/2.0\/"},{"start":{"date-parts":[[2004,10,25]],"date-time":"2004-10-25T00:00:00Z","timestamp":1098662400000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/2.0\/"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                        <jats:title>Background<\/jats:title>\n                        <jats:p>Certain protein families are highly conserved across distantly related organisms and belong to large and functionally diverse superfamilies. The patterns of conservation present in these protein sequences presumably are due to selective constraints maintaining important but unknown structural mechanisms with some constraints specific to each family and others shared by a larger subset or by the entire superfamily. To exploit these patterns as a source of functional information, we recently devised a statistically based approach called <jats:underline>c<\/jats:underline> ontrast <jats:underline>h<\/jats:underline> ierarchical <jats:underline>a<\/jats:underline> lignment and <jats:underline>i<\/jats:underline> nteraction <jats:underline>n<\/jats:underline> etwork (CHAIN) analysis, which infers the strengths of various categories of selective constraints from co-conserved patterns in a multiple alignment. The power of this approach strongly depends on the quality of the multiple alignments, which thus motivated development of theoretical concepts and strategies to improve alignment of conserved motifs within large sets of distantly related sequences.<\/jats:p>\n                     <\/jats:sec><jats:sec>\n                        <jats:title>Results<\/jats:title>\n                        <jats:p>Here we describe a hidden Markov model (HMM), an algebraic system, and Markov chain Monte Carlo (MCMC) sampling strategies for alignment of multiple sequence motifs. The MCMC sampling strategies are useful both for alignment optimization and for adjusting position specific background amino acid frequencies for alignment uncertainties. Associated statistical formulations provide an objective measure of alignment quality as well as automatic gap penalty optimization. Improved alignments obtained in this way are compared with PSI-BLAST based alignments within the context of CHAIN analysis of three protein families: G<jats:sub>i<jats:italic>\u03b1<\/jats:italic>\n                           <\/jats:sub>subunits, prolyl oligopeptidases, and transitional endoplasmic reticulum (p97) AAA+ ATPases.<\/jats:p>\n                     <\/jats:sec><jats:sec>\n                        <jats:title>Conclusion<\/jats:title>\n                        <jats:p>While not entirely replacing PSI-BLAST based alignments, which likewise may be optimized for CHAIN analysis using this approach, these motif-based methods often more accurately align very distantly related sequences and thus can provide a better measure of selective constraints. In some instances, these new approaches also provide a better understanding of family-specific constraints, as we illustrate for p97 ATPases. Programs implementing these procedures and supplementary information are available from the authors.<\/jats:p>\n                     <\/jats:sec>","DOI":"10.1186\/1471-2105-5-157","type":"journal-article","created":{"date-parts":[[2004,11,3]],"date-time":"2004-11-03T07:24:13Z","timestamp":1099466653000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":27,"title":["Gapped alignment of protein sequence motifs through Monte Carlo optimization of a hidden Markov model"],"prefix":"10.1186","volume":"5","author":[{"given":"Andrew F","family":"Neuwald","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun S","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2004,10,25]]},"reference":[{"issue":"17","key":"273_CR1","doi-asserted-by":"publisher","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","volume":"25","author":"SF Altschul","year":"1997","unstructured":"Altschul SF, Madden TL, Schaffer AA, Zhang J, Zhang Z, Miller W, Lipman DJ: Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.\n                           Nucleic Acids Res 1997, 25(17):3389\u20133402. 10.1093\/nar\/25.17.3389","journal-title":"Nucleic Acids Res"},{"issue":"10","key":"273_CR2","doi-asserted-by":"publisher","first-page":"846","DOI":"10.1093\/bioinformatics\/14.10.846","volume":"14","author":"K Karplus","year":"1998","unstructured":"Karplus K, Barrett C, Hughey R: Hidden Markov models for detecting remote protein homologies.\n                           Bioinformatics 1998, 14(10):846\u2013856. 10.1093\/bioinformatics\/14.10.846","journal-title":"Bioinformatics"},{"issue":"4","key":"273_CR3","doi-asserted-by":"publisher","first-page":"673","DOI":"10.1101\/gr.862303","volume":"13","author":"AF Neuwald","year":"2003","unstructured":"Neuwald AF, Kannan N, Poleksic A, Hata N, Liu JS: Ran's C-terminal, basic patch and nucleotide exchange mechanisms in light of a canonical structure for Rab, Rho, Ras and Ran GTPases.\n                           Genome Res 2003, 13(4):673\u2013692. 10.1101\/gr.862303","journal-title":"Genome Res"},{"issue":"432","key":"273_CR4","doi-asserted-by":"publisher","first-page":"1156","DOI":"10.1080\/01621459.1995.10476622","volume":"90","author":"JS Liu","year":"1995","unstructured":"Liu JS, Neuwald AF, Lawrence CE: Bayesian models for multiple local sequence alignment and Gibbs sampling stragtegies.\n                           J Am Stat Assoc 1995, 90(432):1156\u20131170.","journal-title":"J Am Stat Assoc"},{"issue":"9","key":"273_CR5","doi-asserted-by":"publisher","first-page":"1665","DOI":"10.1093\/nar\/25.9.1665","volume":"25","author":"AF Neuwald","year":"1997","unstructured":"Neuwald AF, Liu JS, Lipman DJ, Lawrence CE: Extracting protein alignment models from the sequence database.\n                           Nucleic Acids Research 1997, 25(9):1665\u20131677. 10.1093\/nar\/25.9.1665","journal-title":"Nucleic Acids Research"},{"key":"273_CR6","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1080\/01621459.1999.10473814","volume":"94","author":"JS Liu","year":"1999","unstructured":"Liu JS, Neuwald AF, Lawrence CE: Markovian structures in biological sequence alignments.\n                           J Am Stat Assoc 1999, 94: 1\u201315.","journal-title":"J Am Stat Assoc"},{"issue":"1","key":"273_CR7","doi-asserted-by":"publisher","first-page":"88","DOI":"10.1002\/(SICI)1097-0134(19980701)32:1<88::AID-PROT10>3.0.CO;2-J","volume":"32","author":"SF Altschul","year":"1998","unstructured":"Altschul SF: Generalized affine gap costs for protein sequence alignment.\n                           Proteins 1998, 32(1):88\u201396.","journal-title":"Proteins"},{"issue":"1","key":"273_CR8","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1186\/1471-2105-5-113","volume":"5","author":"RC Edgar","year":"2004","unstructured":"Edgar RC: MUSCLE: a multiple sequence alignment method with reduced time and space complexity.\n                           BMC Bioinformatics 2004, 5(1):113. 10.1186\/1471-2105-5-113","journal-title":"BMC Bioinformatics"},{"issue":"5","key":"273_CR9","doi-asserted-by":"publisher","first-page":"1792","DOI":"10.1093\/nar\/gkh340","volume":"32","author":"RC Edgar","year":"2004","unstructured":"Edgar RC: MUSCLE: multiple sequence alignment with high accuracy and high throughput.\n                           Nucleic Acids Res 2004, 32(5):1792\u20131797. Print 2004 10.1093\/nar\/gkh340","journal-title":"Nucleic Acids Res"},{"issue":"14","key":"273_CR10","doi-asserted-by":"publisher","first-page":"3059","DOI":"10.1093\/nar\/gkf436","volume":"30","author":"K Katoh","year":"2002","unstructured":"Katoh K, Misawa K, Kuma K, Miyata T: MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform.\n                           Nucleic Acids Res 2002, 30(14):3059\u20133066. 10.1093\/nar\/gkf436","journal-title":"Nucleic Acids Res"},{"issue":"22","key":"273_CR11","doi-asserted-by":"publisher","first-page":"4673","DOI":"10.1093\/nar\/22.22.4673","volume":"22","author":"JD Thompson","year":"1994","unstructured":"Thompson JD, Higgins DG, Gibson TJ: CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.\n                           Nucleic Acids Res 1994, 22(22):4673\u20134680.","journal-title":"Nucleic Acids Res"},{"issue":"1","key":"273_CR12","doi-asserted-by":"publisher","first-page":"205","DOI":"10.1006\/jmbi.2000.4042","volume":"302","author":"C Notredame","year":"2000","unstructured":"Notredame C, Higgins DG, Heringa J: T-Coffee: A novel method for fast and accurate multiple sequence alignment.\n                           J Mol Biol 2000, 302(1):205\u2013217. 10.1006\/jmbi.2000.4042","journal-title":"J Mol Biol"},{"issue":"1","key":"273_CR13","doi-asserted-by":"publisher","first-page":"323","DOI":"10.1093\/nar\/29.1.323","volume":"29","author":"A Bahr","year":"2001","unstructured":"Bahr A, Thompson JD, Thierry JC, Poch O: BAliBASE (Benchmark Alignment dataBASE): enhancements for repeats, transmembrane sequences and circular permutations.\n                           Nucleic Acids Res 2001, 29(1):323\u2013326. 10.1093\/nar\/29.1.323","journal-title":"Nucleic Acids Res"},{"issue":"1","key":"273_CR14","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1101\/gr.9.1.27","volume":"9","author":"AF Neuwald","year":"1999","unstructured":"Neuwald AF, Aravind L, Spouge JL, Koonin EV: AAA+: A class of chaperone-like ATPases associated with the assembly, operation, and disassembly of protein complexes.\n                           Genome Res 1999, 9(1):27\u201343.","journal-title":"Genome Res"},{"issue":"10","key":"273_CR15","doi-asserted-by":"publisher","first-page":"1445","DOI":"10.1101\/gr.147400","volume":"10","author":"AF Neuwald","year":"2000","unstructured":"Neuwald AF, Hirano T: HEAT repeats associated with condensins, cohesins, and other complexes involved in chromosome-related functions.\n                           Genome Research 2000, 10(10):1445\u20131452. 10.1101\/gr.147400","journal-title":"Genome Research"},{"issue":"18","key":"273_CR16","doi-asserted-by":"publisher","first-page":"3570","DOI":"10.1093\/nar\/28.18.3570","volume":"28","author":"AF Neuwald","year":"2000","unstructured":"Neuwald AF, Poleksic A: PSI-BLAST searches using hidden markov models of structural repeats: prediction of an unusual sliding DNA clamp and of beta-propellers in UV-damaged DNA-binding protein.\n                           Nucleic Acids Res 2000, 28(18):3570\u20133580. 10.1093\/nar\/28.18.3570","journal-title":"Nucleic Acids Res"},{"volume-title":"GTPases","year":"2000","key":"273_CR17","unstructured":"Hall A, ed: GTPases:. Oxford University Press; 2000."},{"issue":"6","key":"273_CR18","doi-asserted-by":"publisher","first-page":"732","DOI":"10.1016\/S0959-440X(99)00037-8","volume":"9","author":"M Nardini","year":"1999","unstructured":"Nardini M, Dijkstra BW: Alpha\/beta hydrolase fold enzymes: the family keeps growing.\n                           Curr Opin Struct Biol 1999, 9(6):732\u2013737. 10.1016\/S0959-440X(99)00037-8","journal-title":"Curr Opin Struct Biol"},{"key":"273_CR19","doi-asserted-by":"publisher","first-page":"197","DOI":"10.1093\/protein\/5.3.197","volume":"5","author":"DL Ollis","year":"1992","unstructured":"Ollis DL, Cheah E, Cygler M, Dijkstra B, Frolow F, Franken SM, Harel M, Remington SJ, Silman I, Schrag J: The alpha\/beta hydrolase fold.\n                           Protein Eng 1992, 5: 197\u2013211.","journal-title":"Protein Eng"},{"issue":"1\u20132","key":"273_CR20","doi-asserted-by":"publisher","first-page":"44","DOI":"10.1016\/j.jsb.2003.11.014","volume":"146","author":"Q Wang","year":"2004","unstructured":"Wang Q, Song C, Li CC: Molecular perspectives on p97-VCP: progress in understanding its structure and diverse biological functions.\n                           J Struct Biol 2004, 146(1\u20132):44\u201357. 10.1016\/j.jsb.2003.11.014","journal-title":"J Struct Biol"},{"issue":"7","key":"273_CR21","doi-asserted-by":"publisher","first-page":"639","DOI":"10.1002\/bies.950170710","volume":"17","author":"F Confalonieri","year":"1995","unstructured":"Confalonieri F, Duguet M: A 200-amino acid ATPase module in search of a basic function.\n                           Bioessays 1995, 17(7):639\u2013650.","journal-title":"Bioessays"},{"issue":"6517","key":"273_CR22","doi-asserted-by":"publisher","first-page":"88","DOI":"10.1038\/374088a0","volume":"374","author":"JC Swaffield","year":"1995","unstructured":"Swaffield JC, Melcher K, Johnston SA: A highly conserved ATPase protein as a mediator between acidic activation domains and the TATA-binding protein.\n                           Nature 1995, 374(6517):88\u201391. 10.1038\/374088a0","journal-title":"Nature"},{"issue":"2","key":"273_CR23","doi-asserted-by":"publisher","first-page":"65","DOI":"10.1016\/S0962-8924(97)01212-9","volume":"8","author":"S Patel","year":"1998","unstructured":"Patel S, Latterich M: The AAA team: related ATPases with diverse functions.\n                           Trends Cell Biol 1998, 8(2):65\u201371. 10.1016\/S0962-8924(97)01212-9","journal-title":"Trends Cell Biol"},{"issue":"7","key":"273_CR24","doi-asserted-by":"publisher","first-page":"575","DOI":"10.1046\/j.1365-2443.2001.00447.x","volume":"6","author":"T Ogura","year":"2001","unstructured":"Ogura T, Wilkinson AJ: AAA+ superfamily ATPases: common structure \u2013 diverse function.\n                           Genes Cells 2001, 6(7):575\u2013597. 10.1046\/j.1365-2443.2001.00447.x","journal-title":"Genes Cells"},{"issue":"1\u20132","key":"273_CR25","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1016\/j.jsb.2003.10.010","volume":"146","author":"LM Iyer","year":"2004","unstructured":"Iyer LM, Leipe DD, Koonin EV, Aravind L: Evolutionary history and higher order classification of AAA+ ATPases.\n                           J Struct Biol 2004, 146(1\u20132):11\u201331. 10.1016\/j.jsb.2003.10.010","journal-title":"J Struct Biol"},{"issue":"9","key":"273_CR26","doi-asserted-by":"publisher","first-page":"755","DOI":"10.1093\/bioinformatics\/14.9.755","volume":"14","author":"SR Eddy","year":"1998","unstructured":"Eddy SR: Profile hidden Markov models.\n                           Bioinformatics 1998, 14(9):755\u2013763. 10.1093\/bioinformatics\/14.9.755","journal-title":"Bioinformatics"},{"issue":"2","key":"273_CR27","first-page":"95","volume":"12","author":"R Hughey","year":"1996","unstructured":"Hughey R, Krogh A: Hidden Markov models for sequence analysis: extension and analysis of the basic method.\n                           Comput Appl Biosci 1996, 12(2):95\u2013107.","journal-title":"Comput Appl Biosci"},{"key":"273_CR28","volume-title":"Monte Carlo Strategies in Scientific Computing","author":"JS Liu","year":"2001","unstructured":"Liu JS: Monte Carlo Strategies in Scientific Computing. New York Springer-Verlag; 2001."},{"key":"273_CR29","doi-asserted-by":"publisher","first-page":"671","DOI":"10.1126\/science.220.4598.671","volume":"220","author":"S Kirkpatrick","year":"1983","unstructured":"Kirkpatrick S, Gelatt CD, Vecchi MP: Optimization by simulated annealing.\n                           Science 1983, 220: 671\u2013680.","journal-title":"Science"},{"key":"273_CR30","doi-asserted-by":"publisher","first-page":"1618","DOI":"10.1002\/pro.5560040820","volume":"4","author":"AF Neuwald","year":"1995","unstructured":"Neuwald AF, Liu JS, Lawrence CE: Gibbs motif sampling: detection of bacterial outer membrane protein repeats.\n                           Protein Sci 1995, 4: 1618\u20131632.","journal-title":"Protein Sci"},{"issue":"6825","key":"273_CR31","doi-asserted-by":"publisher","first-page":"259","DOI":"10.1038\/35065704","volume":"410","author":"PG Debenedetti","year":"2001","unstructured":"Debenedetti PG, Stillinger FH: Supercooled liquids and the glass transition.\n                           Nature 2001, 410(6825):259\u2013267. 10.1038\/35065704","journal-title":"Nature"},{"issue":"15","key":"273_CR32","doi-asserted-by":"publisher","first-page":"4503","DOI":"10.1093\/nar\/gkg486","volume":"31","author":"AF Neuwald","year":"2003","unstructured":"Neuwald AF: Evolutionary clues to DNA polymerase III beta clamp structural mechanisms.\n                           Nucleic Acids Res 2003, 31(15):4503\u20134516. 10.1093\/nar\/gkg486","journal-title":"Nucleic Acids Res"},{"key":"273_CR33","doi-asserted-by":"publisher","first-page":"000","DOI":"10.1110\/ps.04637904","volume":"13","author":"N Kannan","year":"2004","unstructured":"Kannan N, Neuwald AF: Evolutionary constraints associated with functional specificity of the CMGC protein kinases MAPK, CDK, GSK, SRPK, DYRK, and CK2alpha.\n                           Protein Science 2004, 13: 000\u2013000. 10.1110\/ps.04637904","journal-title":"Protein Science"},{"issue":"6","key":"273_CR34","doi-asserted-by":"publisher","first-page":"1473","DOI":"10.1016\/S1097-2765(00)00143-X","volume":"6","author":"X Zhang","year":"2000","unstructured":"Zhang X, Shaw A, Bates PA, Newman RH, Gowen B, Orlova E, Gorman MA, Kondo H, Dokurno P, Lally J, Leonard G, Meyer H, van Heel M, Freemont PS: Structure of the AAA ATPase p97.\n                           Mol Cell 2000, 6(6):1473\u20131484. 10.1016\/S1097-2765(00)00143-X","journal-title":"Mol Cell"},{"key":"273_CR35","doi-asserted-by":"publisher","first-page":"574","DOI":"10.1016\/0022-2836(94)90032-9","volume":"243","author":"S Henikoff","year":"1994","unstructured":"Henikoff S, Henikoff JG: Position-based sequence weights.\n                           J Mol Biol 1994, 243: 574\u2013578. 10.1016\/0022-2836(94)90032-9","journal-title":"J Mol Biol"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-5-157.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/1471-2105-5-157\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-5-157.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,7]],"date-time":"2024-10-07T12:19:23Z","timestamp":1728303563000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-5-157"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2004,10,25]]},"references-count":35,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2004,12]]}},"alternative-id":["273"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-5-157","relation":{},"ISSN":["1471-2105"],"issn-type":[{"type":"electronic","value":"1471-2105"}],"subject":[],"published":{"date-parts":[[2004,10,25]]},"assertion":[{"value":"19 May 2004","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 October 2004","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 October 2004","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"157"}}