{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T21:12:02Z","timestamp":1777669922756,"version":"3.51.4"},"reference-count":33,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Background<\/jats:title>\n                <jats:p>Genomic sequence data cannot be fully appreciated in isolation. Comparative genomics \u2013 the practice of comparing genomic sequences from different species \u2013 plays an increasingly important role in understanding the genotypic differences between species that result in phenotypic differences as well as in revealing patterns of evolutionary relationships. One of the major challenges in comparative genomics is producing a high-quality alignment between two or more related genomic sequences. In recent years, a number of tools have been developed for aligning large genomic sequences. Most utilize heuristic strategies to identify a series of strong sequence similarities, which are then used as anchors to align the regions between the anchor points. The resulting alignment is globally correct, but in many cases is suboptimal locally. We describe a new program, GenAlignRefine, which improves the overall quality of global multiple alignments by using a genetic algorithm to improve local regions of alignment. Regions of low quality are identified, realigned using the program T-Coffee, and then refined using a genetic algorithm. Because a better COFFEE (Consistency based Objective Function For alignmEnt Evaluation) score generally reflects greater alignment quality, the algorithm searches for an alignment that yields a better COFFEE score. To improve the intrinsic slowness of the genetic algorithm, GenAlignRefine was implemented as a parallel, cluster-based program.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>We tested the GenAlignRefine algorithm by running it on a Linux cluster to refine sequences from a simulation, as well as refine a multiple alignment of 15 Orthopoxvirus genomic sequences approximately 260,000 nucleotides in length that initially had been aligned by Multi-LAGAN. It took approximately 150 minutes for a 40-processor Linux cluster to optimize some 200 fuzzy (poorly aligned) regions of the orthopoxvirus alignment. Overall sequence identity increased only slightly; but significantly, this occurred at the same time that the overall alignment length decreased \u2013 through the removal of gaps \u2013 by approximately 200 gapped regions representing roughly 1,300 gaps.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusion<\/jats:title>\n                <jats:p>We have implemented a genetic algorithm in parallel mode to optimize multiple genomic sequence alignments initially generated by various alignment tools. Benchmarking experiments showed that the refinement algorithm improved genomic sequence alignments within a reasonable period of time.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1186\/1471-2105-6-200","type":"journal-article","created":{"date-parts":[[2005,8,9]],"date-time":"2005-08-09T18:14:00Z","timestamp":1123611240000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":19,"title":["Genomic multiple sequence alignments: refinement using a genetic algorithm"],"prefix":"10.1186","volume":"6","author":[{"given":"Chunlin","family":"Wang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elliot J","family":"Lefkowitz","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2005,8,8]]},"reference":[{"issue":"Web Server issu","key":"525_CR1","doi-asserted-by":"publisher","first-page":"W280","DOI":"10.1093\/nar\/gkh355","volume":"32","author":"I Ovcharenko","year":"2004","unstructured":"Ovcharenko I, Nobrega MA, Loots GG, Stubbs L: ECR Browser: a tool for visualizing and accessing data from comparisons of multiple vertebrate genomes. Nucleic Acids Res 2004, 32(Web Server issue):W280\u20136.","journal-title":"Nucleic Acids Res"},{"issue":"Web Server issu","key":"525_CR2","doi-asserted-by":"publisher","first-page":"W41","DOI":"10.1093\/nar\/gkh361","volume":"32","author":"M Brudno","year":"2004","unstructured":"Brudno M, Steinkamp R, Morgenstern B: The CHAOS\/DIALIGN WWW server for multiple alignment of genomic sequences. Nucleic Acids Res 2004, 32(Web Server issue):W41\u20134.","journal-title":"Nucleic Acids Res"},{"issue":"11","key":"525_CR3","doi-asserted-by":"publisher","first-page":"2478","DOI":"10.1093\/nar\/30.11.2478","volume":"30","author":"AL Delcher","year":"2002","unstructured":"Delcher AL, Phillippy A, Carlton J, Salzberg SL: Fast algorithms for large-scale genome alignment and comparison. Nucleic Acids Res 2002, 30(11):2478\u20132483. 10.1093\/nar\/30.11.2478","journal-title":"Nucleic Acids Res"},{"issue":"8","key":"525_CR4","doi-asserted-by":"publisher","first-page":"1071","DOI":"10.1101\/gr.10.8.1071","volume":"10","author":"DL Baillie","year":"2000","unstructured":"Baillie DL, Rose AM: WABA success: a tool for sequence comparison between large genomes. Genome Res 2000, 10(8):1071\u20131073. 10.1101\/gr.10.8.1071","journal-title":"Genome Res"},{"issue":"Web Server issu","key":"525_CR5","doi-asserted-by":"publisher","first-page":"W273","DOI":"10.1093\/nar\/gkh458","volume":"32","author":"KA Frazer","year":"2004","unstructured":"Frazer KA, Pachter L, Poliakov A, Rubin EM, Dubchak I: VISTA: computational tools for comparative genomics. Nucleic Acids Res 2004, 32(Web Server issue):W273\u20139.","journal-title":"Nucleic Acids Res"},{"issue":"1","key":"525_CR6","doi-asserted-by":"publisher","first-page":"103","DOI":"10.1101\/gr.809403","volume":"13","author":"S Schwartz","year":"2003","unstructured":"Schwartz S, Kent WJ, Smit A, Zhang Z, Baertsch R, Hardison RC, Haussler D, Miller W: Human-mouse alignments with BLASTZ. Genome Res 2003, 13(1):103\u2013107. 10.1101\/gr.809403","journal-title":"Genome Res"},{"issue":"1","key":"525_CR7","doi-asserted-by":"publisher","first-page":"97","DOI":"10.1101\/gr.789803","volume":"13","author":"N Bray","year":"2003","unstructured":"Bray N, Dubchak I, Pachter L: AVID: A global alignment program. Genome Res 2003, 13(1):97\u2013102. 10.1101\/gr.789803","journal-title":"Genome Res"},{"issue":"4","key":"525_CR8","doi-asserted-by":"publisher","first-page":"721","DOI":"10.1101\/gr.926603","volume":"13","author":"M Brudno","year":"2003","unstructured":"Brudno M, Do CB, Cooper GM, Kim MF, Davydov E, Green ED, Sidow A, Batzoglou S: LAGAN and Multi-LAGAN: efficient tools for large-scale multiple alignment of genomic DNA. Genome Res 2003, 13(4):721\u2013731. 10.1101\/gr.926603","journal-title":"Genome Res"},{"key":"525_CR9","doi-asserted-by":"publisher","first-page":"212","DOI":"10.1017\/CBO9780511574931.013","volume-title":"Algorithm on Strings, Trees, and Sequences","author":"D Gusfield","year":"1997","unstructured":"Gusfield D: Algorithm on Strings, Trees, and Sequences. Cambridge University Press; 1997:212."},{"issue":"1","key":"525_CR10","first-page":"13","volume":"11","author":"M Hirosawa","year":"1995","unstructured":"Hirosawa M, Totoki Y, Hoshida M, Ishikawa M: Comprehensive study on iterative algorithms of multiple sequence alignment. Comput Appl Biosci 1995, 11(1):13\u201318.","journal-title":"Comput Appl Biosci"},{"issue":"3","key":"525_CR11","doi-asserted-by":"crossref","first-page":"572","DOI":"10.2144\/02323rv01","volume":"32","author":"HBJ Nicholas","year":"2002","unstructured":"Nicholas HBJ, Ropelewski AJ, Deerfield DW: Strategies for multiple sequence alignment. Biotechniques 2002, 32(3):572\u20134, 576, 578 passim.","journal-title":"Biotechniques"},{"issue":"5","key":"525_CR12","doi-asserted-by":"publisher","first-page":"407","DOI":"10.1093\/bioinformatics\/14.5.407","volume":"14","author":"C Notredame","year":"1998","unstructured":"Notredame C, Holm L, Higgins DG: COFFEE: an objective function for multiple sequence alignments. Bioinformatics 1998, 14(5):407\u2013422. 10.1093\/bioinformatics\/14.5.407","journal-title":"Bioinformatics"},{"issue":"3","key":"525_CR13","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1093\/bioinformatics\/15.3.211","volume":"15","author":"B Morgenstern","year":"1999","unstructured":"Morgenstern B: DIALIGN 2: improvement of the segment-to-segment approach to multiple sequence alignment. Bioinformatics 1999, 15(3):211\u2013218. 10.1093\/bioinformatics\/15.3.211","journal-title":"Bioinformatics"},{"issue":"1","key":"525_CR14","doi-asserted-by":"publisher","first-page":"205","DOI":"10.1006\/jmbi.2000.4042","volume":"302","author":"C Notredame","year":"2000","unstructured":"Notredame C, Higgins DG, Heringa J: T-Coffee: A novel method for fast and accurate multiple sequence alignment. J Mol Biol 2000, 302(1):205\u2013217. 10.1006\/jmbi.2000.4042","journal-title":"J Mol Biol"},{"issue":"3","key":"525_CR15","doi-asserted-by":"publisher","first-page":"443","DOI":"10.1016\/0022-2836(70)90057-4","volume":"48","author":"SB Needleman","year":"1970","unstructured":"Needleman SB, Wunsch CD: A general method applicable to the search for similarities in the amino acid sequence of two proteins. J Mol Biol 1970, 48(3):443\u2013453. 10.1016\/0022-2836(70)90057-4","journal-title":"J Mol Biol"},{"key":"525_CR16","first-page":"viii, 183 p.","volume-title":"Adaptation in natural and artificial systems : an introductory analysis with applications to biology, control, and artificial intelligence","author":"JH Holland","year":"1975","unstructured":"Holland JH: Adaptation in natural and artificial systems : an introductory analysis with applications to biology, control, and artificial intelligence. Ann Arbor , University of Michigan Press; 1975:viii, 183 p.."},{"issue":"8","key":"525_CR17","doi-asserted-by":"publisher","first-page":"1515","DOI":"10.1093\/nar\/24.8.1515","volume":"24","author":"C Notredame","year":"1996","unstructured":"Notredame C, Higgins DG: SAGA: sequence alignment by genetic algorithm. Nucleic Acids Res 1996, 24(8):1515\u20131524. 10.1093\/nar\/24.8.1515","journal-title":"Nucleic Acids Res"},{"issue":"22","key":"525_CR18","doi-asserted-by":"publisher","first-page":"4673","DOI":"10.1093\/nar\/22.22.4673","volume":"22","author":"JD Thompson","year":"1994","unstructured":"Thompson JD, Higgins DG, Gibson TJ: CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice. Nucleic Acids Res 1994, 22(22):4673\u20134680.","journal-title":"Nucleic Acids Res"},{"issue":"2","key":"525_CR19","doi-asserted-by":"publisher","first-page":"141","DOI":"10.1016\/j.compbiolchem.2004.02.001","volume":"28","author":"Y Wang","year":"2004","unstructured":"Wang Y, Li KB: An adaptive and iterative algorithm for refining multiple sequence alignment. Comput Biol Chem 2004, 28(2):141\u2013148. 10.1016\/j.compbiolchem.2004.02.001","journal-title":"Comput Biol Chem"},{"issue":"1-2","key":"525_CR20","doi-asserted-by":"publisher","first-page":"55","DOI":"10.1016\/S0168-1702(00)00210-0","volume":"70","author":"C Upton","year":"2000","unstructured":"Upton C, Hogg D, Perrin D, Boone M, Harris NL: Viral genome organizer: a system for analyzing complete viral genomes. Virus Res 2000, 70(1\u20132):55\u201364. 10.1016\/S0168-1702(00)00210-0","journal-title":"Virus Res"},{"key":"525_CR21","first-page":"113","volume-title":"Multiple Sequence Alignment Using SAGA: Investigating the Effects of Operator Scheduling, Population Seeding, and Crossover Operators.","author":"R Thomsen","year":"2004","unstructured":"Thomsen R, Boomsma W: Multiple Sequence Alignment Using SAGA: Investigating the Effects of Operator Scheduling, Population Seeding, and Crossover Operators. 2004, 113\u2013122."},{"issue":"10","key":"525_CR22","doi-asserted-by":"publisher","first-page":"1611","DOI":"10.1101\/gr.361602","volume":"12","author":"JE Stajich","year":"2002","unstructured":"Stajich JE, Block D, Boulez K, Brenner SE, Chervitz SA, Dagdigian C, Fuellen G, Gilbert JG, Korf I, Lapp H, Lehvaslaiho H, Matsalla C, Mungall CJ, Osborne BI, Pocock MR, Schattner P, Senger M, Stein LD, Stupka E, Wilkinson MD, Birney E: The Bioperl toolkit: Perl modules for the life sciences. Genome Res 2002, 12(10):1611\u20131618. 10.1101\/gr.361602","journal-title":"Genome Res"},{"key":"525_CR23","unstructured":"LAM\/MPI[http:\/\/charm.cs.uiuc.edu\/]"},{"issue":"2","key":"525_CR24","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1093\/bioinformatics\/14.2.157","volume":"14","author":"J Stoye","year":"1998","unstructured":"Stoye J, Evers D, Meyer F: Rose: generating sequence families. Bioinformatics 1998, 14(2):157\u2013163. 10.1093\/bioinformatics\/14.2.157","journal-title":"Bioinformatics"},{"issue":"1","key":"525_CR25","doi-asserted-by":"publisher","first-page":"6","DOI":"10.1186\/1471-2105-5-6","volume":"5","author":"DA Pollard","year":"2004","unstructured":"Pollard DA, Bergman CM, Stoye J, Celniker SE, Eisen MB: Benchmarking tools for the alignment of functional noncoding DNA. BMC Bioinformatics 2004, 5(1):6. 10.1186\/1471-2105-5-6","journal-title":"BMC Bioinformatics"},{"issue":"1","key":"525_CR26","doi-asserted-by":"publisher","first-page":"20","DOI":"10.1186\/1471-2148-4-20","volume":"4","author":"T Muller","year":"2004","unstructured":"Muller T, Rahmann S, Dandekar T, Wolf M: Accurate and robust phylogeny estimation based on profile distances: a study of the Chlorophyceae (Chlorophyta). BMC Evol Biol 2004, 4(1):20. 10.1186\/1471-2148-4-20","journal-title":"BMC Evol Biol"},{"issue":"2","key":"525_CR27","doi-asserted-by":"publisher","first-page":"160","DOI":"10.1007\/BF02101694","volume":"22","author":"M Hasegawa","year":"1985","unstructured":"Hasegawa M, Kishino H, Yano T: Dating of the human-ape splitting by a molecular clock of mitochondrial DNA. J Mol Evol 1985, 22(2):160\u2013174.","journal-title":"J Mol Evol"},{"key":"525_CR28","unstructured":"CHAOS\/DIALIGN[[http:\/\/dialign.gobics.de\/chaos-dialign-submission] http:\/\/dialign.gobics.de\/chaos-dialign-submission]."},{"key":"525_CR29","unstructured":"Ectromelia virus strain Naval - Sanger Institute[http:\/\/www.sanger.ac.uk\/Projects\/Ectromelia_virus\/]"},{"issue":"Pt 4","key":"525_CR30","doi-asserted-by":"publisher","first-page":"855","DOI":"10.1099\/0022-1317-83-4-855","volume":"83","author":"C Gubser","year":"2002","unstructured":"Gubser C, Smith GL: The sequence of camelpox virus shows it is most closely related to variola virus, the cause of smallpox. J Gen Virol 2002, 83(Pt 4):855\u2013872.","journal-title":"J Gen Virol"},{"key":"525_CR31","unstructured":"Open Source Initiative[http:\/\/www.opensource.org\/licenses\/artistic-license.php]"},{"key":"525_CR32","unstructured":"FTP Site[ftp:\/\/ftp.genome.uab.edu\/]"},{"key":"525_CR33","unstructured":"HPCL[http:\/\/ardra.hpcl.cis.uab.edu\/]"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-6-200.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,2,1]],"date-time":"2024-02-01T17:47:19Z","timestamp":1706809639000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-6-200"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,8,8]]},"references-count":33,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2005,12]]}},"alternative-id":["525"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-6-200","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2005,8,8]]},"assertion":[{"value":"21 January 2005","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 August 2005","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 August 2005","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"200"}}