{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,14]],"date-time":"2026-03-14T08:27:46Z","timestamp":1773476866267,"version":"3.50.1"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2006,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>There have been many algorithms and software programs implemented for the inference of multiple sequence alignments of protein and DNA sequences. The \"true\" alignment is usually unknown due to the incomplete knowledge of the evolutionary history of the sequences, making it difficult to gauge the relative accuracy of the programs.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>We tested nine of the most often used protein alignment programs and compared their results using sequences generated with the simulation software Simprot which creates known alignments under realistic and controlled evolutionary scenarios. We have simulated more than 30000 alignment sets using various evolutionary histories in order to define strengths and weaknesses of each program tested. We found that alignment accuracy is extremely dependent on the number of insertions and deletions in the sequences, and that indel size has a weaker effect. We also considered benchmark alignments from the latest version of BAliBASE and the results relative to BAliBASE- and Simprot-generated data sets were consistent in most cases.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion<\/jats:title>\n            <jats:p>Our results indicate that employing Simprot's simulated sequences allows the creation of a more flexible and broader range of alignment classes than the usual methods for alignment accuracy assessment. Simprot also allows for a quick and efficient analysis of a wider range of possible evolutionary histories that might not be present in currently available alignment sets. Among the nine programs tested, the iterative approach available in Mafft (L-INS-i) and ProbCons were consistently the most accurate, with Mafft being the faster of the two.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-7-471","type":"journal-article","created":{"date-parts":[[2006,10,25]],"date-time":"2006-10-25T00:57:40Z","timestamp":1161737860000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":119,"title":["The accuracy of several multiple sequence alignment programs for proteins"],"prefix":"10.1186","volume":"7","author":[{"given":"Paulo AS","family":"Nuin","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhouzhi","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elisabeth RM","family":"Tillier","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2006,10,24]]},"reference":[{"key":"1210_CR1","doi-asserted-by":"publisher","first-page":"127","DOI":"10.1002\/prot.20527","volume":"61","author":"J Thompson","year":"2005","unstructured":"Thompson J, Koehl P, Ripp R, Poch O: BAliBASE 3.0: latest developments of the multiple sequence alignment benchmark. Proteins 2005, 61: 127\u201336. 10.1002\/prot.20527","journal-title":"Proteins"},{"key":"1210_CR2","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1186\/1471-2105-5-113","volume":"5","author":"R Edgar","year":"2004","unstructured":"Edgar R: MUSCLE: a multiple sequence alignment method with reduced time and space complexity. BMC Bioinformatics 2004, 5: 113. 10.1186\/1471-2105-5-113","journal-title":"BMC Bioinformatics"},{"issue":"7","key":"1210_CR3","doi-asserted-by":"publisher","first-page":"1267","DOI":"10.1093\/bioinformatics\/bth493","volume":"21","author":"IV Walle","year":"2005","unstructured":"Walle IV, Lasters I, Wyns L: SABmark-a benchmark for sequence alignment that covers the entire known fold space. Bioinformatics 2005, 21(7):1267\u20138. 10.1093\/bioinformatics\/bth493","journal-title":"Bioinformatics"},{"issue":"8","key":"1210_CR4","doi-asserted-by":"publisher","first-page":"713","DOI":"10.1093\/bioinformatics\/17.8.713","volume":"17","author":"K Karplus","year":"2001","unstructured":"Karplus K, Hu B: Evaluation of protein multiple alignments by SAM-T99 using the BAliBASE multiple alignment test set. Bioinformatics 2001, 17(8):713\u201320. 10.1093\/bioinformatics\/17.8.713","journal-title":"Bioinformatics"},{"key":"1210_CR5","first-page":"51","volume":"1","author":"M Rosenberg","year":"2005","unstructured":"Rosenberg M: MySSP: Non-stationary evolutionary sequence simulation, including indels. Evol Bioinformatics Online 2005, 1: 51\u201353.","journal-title":"Evol Bioinformatics Online"},{"issue":"Suppl 3","key":"1210_CR6","doi-asserted-by":"publisher","first-page":"iii31","DOI":"10.1093\/bioinformatics\/bti1200","volume":"21","author":"R Cartwright","year":"2005","unstructured":"Cartwright R: DNA assembly with gaps (Dawg): simulating sequence evolution. Bioinformatics 2005, 21(Suppl 3):iii31-iii38. 10.1093\/bioinformatics\/bti1200","journal-title":"Bioinformatics"},{"key":"1210_CR7","doi-asserted-by":"publisher","first-page":"102","DOI":"10.1186\/1471-2105-6-102","volume":"6","author":"M Rosenberg","year":"2005","unstructured":"Rosenberg M: Evolutionary distance estimation and fidelity of pair wise sequence alignment. BMC Bioinformatics 2005, 6: 102. 10.1186\/1471-2105-6-102","journal-title":"BMC Bioinformatics"},{"key":"1210_CR8","doi-asserted-by":"publisher","first-page":"278","DOI":"10.1186\/1471-2105-6-278","volume":"6","author":"M Rosenberg","year":"2005","unstructured":"Rosenberg M: Multiple sequence alignment accuracy and evolutionary distance estimation. BMC Bioinformatics 2005, 6: 278. 10.1186\/1471-2105-6-278","journal-title":"BMC Bioinformatics"},{"key":"1210_CR9","doi-asserted-by":"publisher","first-page":"126","DOI":"10.1016\/S0014-5793(02)03189-7","volume":"529","author":"T Lassmann","year":"2002","unstructured":"Lassmann T, Sonnhammer E: Quality assessment of multiple alignment programs. FEBS Lett 2002, 529: 126\u201330. 10.1016\/S0014-5793(02)03189-7","journal-title":"FEBS Lett"},{"issue":"2","key":"1210_CR10","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1093\/bioinformatics\/14.2.157","volume":"14","author":"J Stoye","year":"1998","unstructured":"Stoye J, Evers D, Meyer F: Rose: generating sequence families. Bioinformatics 1998, 14(2):157\u201363. 10.1093\/bioinformatics\/14.2.157","journal-title":"Bioinformatics"},{"key":"1210_CR11","doi-asserted-by":"publisher","first-page":"236","DOI":"10.1186\/1471-2105-6-236","volume":"6","author":"A Pang","year":"2005","unstructured":"Pang A, Smith A, Nuin P, Tillier E: SIMPROT: using an empirically determined indel distribution in simulations of protein evolution. BMC Bioinformatics 2005, 6: 236. 10.1186\/1471-2105-6-236","journal-title":"BMC Bioinformatics"},{"key":"1210_CR12","doi-asserted-by":"publisher","first-page":"102","DOI":"10.1002\/prot.1129","volume":"45","author":"B Qian","year":"2001","unstructured":"Qian B, Goldstein R: Distribution of Indel lengths. Proteins 2001, 45: 102\u20134. 10.1002\/prot.1129","journal-title":"Proteins"},{"issue":"6","key":"1210_CR13","first-page":"1396","volume":"10","author":"Z Yang","year":"1993","unstructured":"Yang Z: Maximum-likelihood estimation of phylogeny from DNA sequences when substitution rates differ over sites. Mol Biol Evol 1993, 10(6):1396\u2013401.","journal-title":"Mol Biol Evol"},{"issue":"22","key":"1210_CR14","doi-asserted-by":"publisher","first-page":"4673","DOI":"10.1093\/nar\/22.22.4673","volume":"22","author":"J Thompson","year":"1994","unstructured":"Thompson J, Higgins D, Gibson T: CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice. Nucleic Acids Res 1994, 22(22):4673\u201380.","journal-title":"Nucleic Acids Res"},{"issue":"3","key":"1210_CR15","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1093\/bioinformatics\/15.3.211","volume":"15","author":"B Morgenstern","year":"1999","unstructured":"Morgenstern B: DIALIGN 2: improvement of the segment-to-segment approach to multiple sequence alignment. Bioinformatics 1999, 15(3):211\u20138. 10.1093\/bioinformatics\/15.3.211","journal-title":"Bioinformatics"},{"issue":"22","key":"1210_CR16","doi-asserted-by":"publisher","first-page":"12098","DOI":"10.1073\/pnas.93.22.12098","volume":"93","author":"B Morgenstern","year":"1996","unstructured":"Morgenstern B, Dress A, Werner T: Multiple DNA and protein sequence alignment based on segment-to-segment comparison. Proc Natl Acad Sci USA 1996, 93(22):12098\u2013103. 10.1073\/pnas.93.22.12098","journal-title":"Proc Natl Acad Sci USA"},{"key":"1210_CR17","doi-asserted-by":"publisher","first-page":"205","DOI":"10.1006\/jmbi.2000.4042","volume":"302","author":"C Notredame","year":"2000","unstructured":"Notredame C, Higgins D, Heringa J: T-Coffee: A novel method for fast and accurate multiple sequence alignment. J Mol Biol 2000, 302: 205\u201317. 10.1006\/jmbi.2000.4042","journal-title":"J Mol Biol"},{"issue":"4","key":"1210_CR18","first-page":"373","volume":"6","author":"X Huang","year":"1990","unstructured":"Huang X, Hardison R, Miller W: A space-efficient algorithm for local similarities. Comput Appl Biosci 1990, 6(4):373\u201381.","journal-title":"Comput Appl Biosci"},{"key":"1210_CR19","doi-asserted-by":"publisher","first-page":"56","DOI":"10.1002\/prot.340090107","volume":"9","author":"C Sander","year":"1991","unstructured":"Sander C, Schneider R: Database of homology-derived protein structures and the structural meaning of sequence alignment. Proteins 1991, 9: 56\u201368. 10.1002\/prot.340090107","journal-title":"Proteins"},{"issue":"3","key":"1210_CR20","doi-asserted-by":"publisher","first-page":"452","DOI":"10.1093\/bioinformatics\/18.3.452","volume":"18","author":"C Lee","year":"2002","unstructured":"Lee C, Grasso C, Sharlow M: Multiple sequence alignment using partial order graphs. Bioinformatics 2002, 18(3):452\u201364. 10.1093\/bioinformatics\/18.3.452","journal-title":"Bioinformatics"},{"issue":"3","key":"1210_CR21","doi-asserted-by":"publisher","first-page":"443","DOI":"10.1016\/0022-2836(70)90057-4","volume":"48","author":"S Needleman","year":"1970","unstructured":"Needleman S, Wunsch C: A general method applicable to the search for similarities in the amino acid sequence of two proteins. J Mol Biol 1970, 48(3):443\u201353. 10.1016\/0022-2836(70)90057-4","journal-title":"J Mol Biol"},{"key":"1210_CR22","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1016\/0022-2836(81)90087-5","volume":"147","author":"T Smith","year":"1981","unstructured":"Smith T, Waterman M: Identification of common molecular subsequences. J Mol Biol 1981, 147: 195\u20137. 10.1016\/0022-2836(81)90087-5","journal-title":"J Mol Biol"},{"key":"1210_CR23","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1186\/1471-2105-5-113","volume":"5","author":"R Edgar","year":"2004","unstructured":"Edgar R: MUSCLE: a multiple sequence alignment method with reduced time and space complexity. BMC Bioinformatics 2004, 5: 113. 10.1186\/1471-2105-5-113","journal-title":"BMC Bioinformatics"},{"key":"1210_CR24","first-page":"13","volume":"11","author":"M Hirosawa","year":"1995","unstructured":"Hirosawa M, Totoki Y, Hoshida M, Ishikawa M: Comprehensive study on iterative algorithms of multiple sequence alignment. Comput Appl Biosci 1995, 11: 13\u20138.","journal-title":"Comput Appl Biosci"},{"issue":"2","key":"1210_CR25","doi-asserted-by":"publisher","first-page":"511","DOI":"10.1093\/nar\/gki198","volume":"33","author":"K Katoh","year":"2005","unstructured":"Katoh K, Kuma K, Toh H, Miyata T: MAFFT version 5: improvement in accuracy of multiple sequence alignment. Nucleic Acids Res 2005, 33(2):511\u20138. 10.1093\/nar\/gki198","journal-title":"Nucleic Acids Res"},{"issue":"5","key":"1210_CR26","first-page":"543","volume":"11","author":"O Gotoh","year":"1995","unstructured":"Gotoh O: A weighting system and algorithm for aligning many phylogenetically related sequences. Comput Appl Biosci 1995, 11(5):543\u201351.","journal-title":"Comput Appl Biosci"},{"issue":"14","key":"1210_CR27","doi-asserted-by":"publisher","first-page":"3059","DOI":"10.1093\/nar\/gkf436","volume":"30","author":"K Katoh","year":"2002","unstructured":"Katoh K, Misawa K, Kuma K, Miyata T: MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Res 2002, 30(14):3059\u201366. 10.1093\/nar\/gkf436","journal-title":"Nucleic Acids Res"},{"issue":"2","key":"1210_CR28","doi-asserted-by":"publisher","first-page":"330","DOI":"10.1101\/gr.2821705","volume":"15","author":"C Do","year":"2005","unstructured":"Do C, Mahabhashyam M, Brudno M, Batzoglou S: ProbCons: Probabilistic consistency-based multiple sequence alignment. Genome Res 2005, 15(2):330\u201340. 10.1101\/gr.2821705","journal-title":"Genome Res"},{"key":"1210_CR29","doi-asserted-by":"publisher","first-page":"66","DOI":"10.1186\/1471-2105-6-66","volume":"6","author":"A Subramanian","year":"2005","unstructured":"Subramanian A, Weyer-Menkhoff J, Kaufmann M, Morgenstern B: DIALIGN-T: an improved algorithm for segment-based multiple sequence alignment. BMC Bioinformatics 2005, 6: 66. 10.1186\/1471-2105-6-66","journal-title":"BMC Bioinformatics"},{"key":"1210_CR30","doi-asserted-by":"publisher","first-page":"298","DOI":"10.1186\/1471-2105-6-298","volume":"6","author":"T Lassmann","year":"2005","unstructured":"Lassmann T, Sonnhammer E: Kalign-an accurate and fast multiple sequence alignment algorithm. BMC Bioinformatics 2005, 6: 298. 10.1186\/1471-2105-6-298","journal-title":"BMC Bioinformatics"},{"key":"1210_CR31","doi-asserted-by":"publisher","first-page":"83","DOI":"10.1145\/135239.135244","volume":"35","author":"S Wu","year":"1992","unstructured":"Wu S, Manber U: Fast text searching allowing errors. Communications of the ACM 1992, 35: 83\u201391. 10.1145\/135239.135244","journal-title":"Communications of the ACM"},{"issue":"6","key":"1210_CR32","doi-asserted-by":"publisher","first-page":"997","DOI":"10.1089\/106652703322756195","volume":"10","author":"S Veerassamy","year":"2003","unstructured":"Veerassamy S, Smith A, Tillier E: A transition probability model for amino acid substitutions from blocks. J Comput Biol 2003, 10(6):997\u20131010. 10.1089\/106652703322756195","journal-title":"J Comput Biol"},{"key":"1210_CR33","doi-asserted-by":"publisher","first-page":"87","DOI":"10.1093\/bioinformatics\/15.1.87","volume":"15","author":"J Thompson","year":"1999","unstructured":"Thompson J, Plewniak F, Poch O: BAliBASE: a benchmark alignment database for the evaluation of multiple alignment programs. Bioinformatics 1999, 15: 87\u20138. 10.1093\/bioinformatics\/15.1.87","journal-title":"Bioinformatics"},{"key":"1210_CR34","doi-asserted-by":"crossref","unstructured":"Bateman A, Coin L, Durbin R, Finn R, Hollich V, Griffiths-Jones S, Khanna A, Marshall M, Moxon S, Sonnhammer E, Studholme D, Yeats C, Eddy S: The Pfam protein families database. Nucleic Acids Res 2004, (32 Database):D138\u201341. 10.1093\/nar\/gkh121","DOI":"10.1093\/nar\/gkh121"},{"key":"1210_CR35","doi-asserted-by":"publisher","first-page":"6","DOI":"10.1002\/(SICI)1097-0134(20000701)40:1<6::AID-PROT30>3.0.CO;2-7","volume":"40","author":"J Sauder","year":"2000","unstructured":"Sauder J, Arthur J, Dunbrack R: Large-scale comparison of protein sequence alignment algorithms with structure alignments. Proteins 2000, 40: 6\u201322. 10.1002\/(SICI)1097-0134(20000701)40:1<6::AID-PROT30>3.0.CO;2-7","journal-title":"Proteins"},{"issue":"3","key":"1210_CR36","doi-asserted-by":"publisher","first-page":"496","DOI":"10.1093\/bioinformatics\/18.3.496","volume":"18","author":"R Kahsay","year":"2002","unstructured":"Kahsay R, Wang G, Dongre N, Gao G, Dunbrack R: CASA: a server for the critical assessment of protein sequence alignment accuracy. Bioinformatics 2002, 18(3):496\u20137. 10.1093\/bioinformatics\/18.3.496","journal-title":"Bioinformatics"},{"issue":"2","key":"1210_CR37","doi-asserted-by":"publisher","first-page":"329","DOI":"10.1002\/prot.20299","volume":"58","author":"M Zachariah","year":"2005","unstructured":"Zachariah M, Crooks G, Holbrook S, Brenner S: A generalized affine gap model significantly improves protein sequence alignment accuracy. Proteins 2005, 58(2):329\u201338. 10.1002\/prot.20299","journal-title":"Proteins"},{"issue":"8","key":"1210_CR38","doi-asserted-by":"publisher","first-page":"1301","DOI":"10.1093\/bioinformatics\/bth090","volume":"20","author":"R Edgar","year":"2004","unstructured":"Edgar R, Sj\u00f6lander K: A comparison of scoring functions for protein sequence profile alignment. Bioinformatics 2004, 20(8):1301\u20138. 10.1093\/bioinformatics\/bth090","journal-title":"Bioinformatics"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-7-471.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T03:14:18Z","timestamp":1630466058000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-7-471"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,10,24]]},"references-count":38,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2006,12]]}},"alternative-id":["1210"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-7-471","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2006,10,24]]},"assertion":[{"value":"26 July 2006","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 October 2006","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 October 2006","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"471"}}