{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T10:12:43Z","timestamp":1761559963564},"reference-count":27,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2006,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec><jats:title>Background<\/jats:title><jats:p>Single nucleotide polymorphisms (SNPs) as defined here are single base sequence changes or short insertion\/deletions between or within individuals of a given species. As a result of their abundance and the availability of high throughput analysis technologies SNP markers have begun to replace other traditional markers such as restriction fragment length polymorphisms (RFLPs), amplified fragment length polymorphisms (AFLPs) and simple sequence repeats (SSRs or microsatellite) markers for fine mapping and association studies in several species. For SNP discovery from chromatogram data, several bioinformatics programs have to be combined to generate an analysis pipeline. Results have to be stored in a relational database to facilitate interrogation through queries or to generate data for further analyses such as determination of linkage disequilibrium and identification of common haplotypes. Although these tasks are routinely performed by several groups, an integrated open source SNP discovery pipeline that can be easily adapted by new groups interested in SNP marker development is currently unavailable.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>We developed SNP-PHAGE (<jats:bold>SNP<\/jats:bold>discovery<jats:bold>P<\/jats:bold>ipeline with additional features for identification of common haplotypes within a sequence tagged site (<jats:bold>H<\/jats:bold>aplotype<jats:bold>A<\/jats:bold>nalysis) and<jats:bold>Ge<\/jats:bold>nBank (-dbSNP) submissions. This tool was applied for analyzing sequence traces from diverse soybean genotypes to discover over 10,000 SNPs. This package was developed on UNIX\/Linux platform, written in Perl and uses a MySQL database. Scripts to generate a user-friendly web interface are also provided with common queries for preliminary data analysis. A machine learning tool developed by this group for increasing the efficiency of SNP discovery is integrated as a part of this package as an optional feature. The SNP-PHAGE package is being made available open source at<jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"http:\/\/bfgl.anri.barc.usda.gov\/ML\/snp-phage\/\" ext-link-type=\"uri\">http:\/\/bfgl.anri.barc.usda.gov\/ML\/snp-phage\/<\/jats:ext-link>.<\/jats:p><\/jats:sec><jats:sec><jats:title>Conclusion<\/jats:title><jats:p>SNP-PHAGE provides a bioinformatics solution for high throughput SNP discovery, identification of common haplotypes within an amplicon, and GenBank (dbSNP) submissions. SNP selection and visualization are aided through a user-friendly web interface. This tool is useful for analyzing sequence tagged sites (STSs) of genomic sequences, and this software can serve as a starting point for groups interested in developing SNP markers.<\/jats:p><\/jats:sec>","DOI":"10.1186\/1471-2105-7-468","type":"journal-article","created":{"date-parts":[[2006,10,24]],"date-time":"2006-10-24T14:29:12Z","timestamp":1161700152000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":23,"title":["SNP-PHAGE \u2013 High throughput SNP discovery pipeline"],"prefix":"10.1186","volume":"7","author":[{"given":"Lakshmi K","family":"Matukumalli","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"John J","family":"Grefenstette","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David L","family":"Hyten","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ik-Young","family":"Choi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Perry B","family":"Cregan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Curtis P","family":"Van Tassell","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2006,10,23]]},"reference":[{"key":"1207_CR1","doi-asserted-by":"publisher","first-page":"452","DOI":"10.1038\/70570","volume":"23","author":"GT Marth","year":"1999","unstructured":"Marth GT, Korf I, Yandell MD, Yeh RT, Gu Z, Zakeri H, Stitziel NO, Hillier L, Kwok PY, Gish WR: A general approach to single-nucleotide polymorphism discovery. Nat Genet 1999, 23: 452\u2013456. 10.1038\/70570","journal-title":"Nat Genet"},{"key":"1207_CR2","volume-title":"Phrap","author":"http:","year":"2006","unstructured":"http:, [http:\/\/www.phrap.org] www.phrap.org: Phrap.2006."},{"key":"1207_CR3","doi-asserted-by":"publisher","first-page":"1725","DOI":"10.1101\/gr.194201","volume":"11","author":"Z Ning","year":"2001","unstructured":"Ning Z, Cox AJ, Mullikin JC: SSAHA: a fast search method for large DNA databases. Genome Res 2001, 11: 1725\u20131729. 10.1101\/gr.194201","journal-title":"Genome Res"},{"key":"1207_CR4","doi-asserted-by":"publisher","first-page":"513","DOI":"10.1038\/35035083","volume":"407","author":"D Altshuler","year":"2000","unstructured":"Altshuler D, Pollara VJ, Cowles CR, Van Etten WJ, Baldwin J, Linton L, Lander ES: An SNP map of the human genome generated by reduced representation shotgun sequencing. Nature 2000, 407: 513\u2013516. 10.1038\/35035083","journal-title":"Nature"},{"key":"1207_CR5","doi-asserted-by":"publisher","first-page":"94","DOI":"10.1016\/S1369-5266(02)00240-6","volume":"5","author":"A Rafalski","year":"2002","unstructured":"Rafalski A: Applications of single nucleotide polymorphisms in crop genetics. Curr Opin Plant Biol 2002, 5: 94\u2013100. 10.1016\/S1369-5266(02)00240-6","journal-title":"Curr Opin Plant Biol"},{"key":"1207_CR6","doi-asserted-by":"crossref","first-page":"167","DOI":"10.1101\/gr.9.2.167","volume":"9","author":"L Picoult-Newberg","year":"1999","unstructured":"Picoult-Newberg L, Ideker TE, Pohl MG, Taylor SL, Donaldson MA, Nickerson DA, Boyce-Jacino M: Mining SNPs from EST databases. Genome Res 1999, 9: 167\u2013174.","journal-title":"Genome Res"},{"key":"1207_CR7","doi-asserted-by":"publisher","first-page":"461","DOI":"10.1023\/B:PLAN.0000036376.11710.6f","volume":"54","author":"LL Dantec","year":"2004","unstructured":"Dantec LL, Chagne D, Pot D, Cantin O, Garnier-Gere P, Bedon F, Frigerio JM, Chaumeil P, Leger P, Garcia V, Laigret F, De Daruvar A, Plomion C: Automated SNP detection in expressed sequence tags: statistical considerations and application to maritime pine sequences. Plant Mol Biol 2004, 54: 461\u2013470. 10.1023\/B:PLAN.0000036376.11710.6f","journal-title":"Plant Mol Biol"},{"key":"1207_CR8","doi-asserted-by":"publisher","first-page":"2745","DOI":"10.1093\/nar\/25.14.2745","volume":"25","author":"DA Nickerson","year":"1997","unstructured":"Nickerson DA, Tobe VO, Taylor SL: PolyPhred: automating the detection and genotyping of single nucleotide substitutions using fluorescence-based resequencing. Nucleic Acids Res 1997, 25: 2745\u20132751. 10.1093\/nar\/25.14.2745","journal-title":"Nucleic Acids Res"},{"key":"1207_CR9","doi-asserted-by":"publisher","first-page":"375","DOI":"10.1038\/ng1746","volume":"38","author":"M Stephens","year":"2006","unstructured":"Stephens M, Sloan JS, Robertson PD, Scheet P, Nickerson DA: Automating sequence-based detection and genotyping of SNPs from diploid samples. Nat Genet 2006, 38: 375\u2013381. 10.1038\/ng1746","journal-title":"Nat Genet"},{"key":"1207_CR10","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1002\/humu.20188","volume":"26","author":"C Manaster","year":"2005","unstructured":"Manaster C, Zheng W, Teuber M, Wachter S, Doring F, Schreiber S, Hampe J: InSNP: a tool for automated detection and visualization of SNPs and InDels. Hum Mutat 2005, 26: 11\u201319. 10.1002\/humu.20188","journal-title":"Hum Mutat"},{"key":"1207_CR11","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1101\/gr.2754005","volume":"15","author":"S Weckx","year":"2005","unstructured":"Weckx S, Del Favero J, Rademakers R, Claes L, Cruts M, De Jonghe P, Van Broeckhoven C, De Rijk P: novoSNP, a novel computational tool for sequence variation discovery. Genome Res 2005, 15: 436\u2013442. 10.1101\/gr.2754005","journal-title":"Genome Res"},{"key":"1207_CR12","doi-asserted-by":"publisher","first-page":"e53","DOI":"10.1371\/journal.pcbi.0010053","volume":"1","author":"J Zhang","year":"2005","unstructured":"Zhang J, Wheeler DA, Yakub I, Wei S, Sood R, Rowe W, Liu PP, Gibbs RA, Buetow KH: SNPdetector: A Software Tool for Sensitive and Accurate SNP Detection. PLoS Comput Biol 2005, 1: e53. 10.1371\/journal.pcbi.0010053","journal-title":"PLoS Comput Biol"},{"key":"1207_CR13","doi-asserted-by":"publisher","first-page":"403","DOI":"10.1016\/S0022-2836(05)80360-2","volume":"215","author":"SF Altschul","year":"1990","unstructured":"Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ: Basic local alignment search tool. J Mol Biol 1990, 215: 403\u2013410. 10.1006\/jmbi.1990.9999","journal-title":"J Mol Biol"},{"key":"1207_CR14","doi-asserted-by":"publisher","first-page":"421","DOI":"10.1093\/bioinformatics\/btf881","volume":"19","author":"G Barker","year":"2003","unstructured":"Barker G, Batley J, O' Sullivan H, Edwards KJ, Edwards D: Redundancy based detection of sequence polymorphisms in expressed sequence tag data using autoSNP. Bioinformatics 2003, 19: 421\u2013422. 10.1093\/bioinformatics\/btf881","journal-title":"Bioinformatics"},{"key":"1207_CR15","unstructured":"Bioperl2006., [] http:\/\/www.bioperl.org:http:\/\/www.bioperl.orghttp:\/\/www.bioperl.org"},{"key":"1207_CR16","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1186\/1471-2164-5-60","volume":"5","author":"JA Aerts","year":"2004","unstructured":"Aerts JA, Jungerius BJ, Groenen MA: POSA: perl objects for DNA sequencing data analysis. BMC Genomics 2004, 5: 60. 10.1186\/1471-2164-5-60","journal-title":"BMC Genomics"},{"key":"1207_CR17","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1186\/1471-2105-7-4","volume":"7","author":"LK Matukumalli","year":"2006","unstructured":"Matukumalli LK, Grefenstette JJ, Hyten DL, Choi IY, Cregan PB, Van Tassell CP: Application of machine learning in SNP discovery. BMC Bioinformatics 2006, 7: 4. 10.1186\/1471-2105-7-4","journal-title":"BMC Bioinformatics"},{"key":"1207_CR18","doi-asserted-by":"publisher","first-page":"186","DOI":"10.1101\/gr.8.3.186","volume":"8","author":"B Ewing","year":"1998","unstructured":"Ewing B, Green P: Base-calling of automated sequencer traces using phred. II. Error probabilities. Genome Res 1998, 8: 186\u2013194.","journal-title":"Genome Res"},{"key":"1207_CR19","doi-asserted-by":"publisher","first-page":"175","DOI":"10.1101\/gr.8.3.175","volume":"8","author":"B Ewing","year":"1998","unstructured":"Ewing B, Hillier L, Wendl MC, Green P: Base-calling of automated sequencer traces using phred. I. Accuracy assessment. Genome Res 1998, 8: 175\u2013185.","journal-title":"Genome Res"},{"key":"1207_CR20","volume-title":"C4.5: Programs for Machine Learning","author":"RJ Quinlan","year":"1993","unstructured":"Quinlan RJ: C4.5: Programs for Machine Learning. Morgan Kaufmann; 1993."},{"key":"1207_CR21","doi-asserted-by":"publisher","first-page":"868","DOI":"10.1101\/gr.9.9.868","volume":"9","author":"X Huang","year":"1999","unstructured":"Huang X, Madan A: CAP3: A DNA sequence assembly program. Genome Res 1999, 9: 868\u2013877. 10.1101\/gr.9.9.868","journal-title":"Genome Res"},{"key":"1207_CR22","doi-asserted-by":"publisher","first-page":"233","DOI":"10.1038\/79981","volume":"26","author":"K Irizarry","year":"2000","unstructured":"Irizarry K, Kustanovich V, Li C, Brown N, Nelson S, Wong W, Lee CJ: Genome-wide analysis of single-nucleotide polymorphisms in human expressed sequences. Nat Genet 2000, 26: 233\u2013236. 10.1038\/79981","journal-title":"Nat Genet"},{"key":"1207_CR23","doi-asserted-by":"crossref","first-page":"1123","DOI":"10.1093\/genetics\/163.3.1123","volume":"163","author":"YL Zhu","year":"2003","unstructured":"Zhu YL, Song QJ, Hyten DL, Van Tassell CP, Matukumalli LK, Grimm DR, Hyatt SM, Fickus EW, Young ND, Cregan PB: Single-nucleotide polymorphisms in soybean. Genetics 2003, 163: 1123\u20131134.","journal-title":"Genetics"},{"key":"1207_CR24","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1101\/gr.8.3.195","volume":"8","author":"D Gordon","year":"1998","unstructured":"Gordon D, Abajian C, Green P: Consed: a graphical tool for sequence finishing. Genome Res 1998, 8: 195\u2013202.","journal-title":"Genome Res"},{"key":"1207_CR25","doi-asserted-by":"publisher","first-page":"274","DOI":"10.1186\/1479-7364-1-4-274","volume":"1","author":"MD Shriver","year":"2004","unstructured":"Shriver MD, Kennedy GC, Parra EJ, Lawson HA, Sonpar V, Huang J, Akey JM, Jones KW: The genomic distribution of population substructure in four populations using 8,525 autosomal SNPs. Hum Genomics 2004, 1: 274\u2013286.","journal-title":"Hum Genomics"},{"key":"1207_CR26","doi-asserted-by":"publisher","first-page":"978","DOI":"10.1086\/319501","volume":"68","author":"M Stephens","year":"2001","unstructured":"Stephens M, Smith NJ, Donnelly P: A new statistical method for haplotype reconstruction from population data. Am J Hum Genet 2001, 68: 978\u2013989. 10.1086\/319501","journal-title":"Am J Hum Genet"},{"key":"1207_CR27","doi-asserted-by":"publisher","first-page":"263","DOI":"10.1093\/bioinformatics\/bth457","volume":"21","author":"JC Barrett","year":"2005","unstructured":"Barrett JC, Fry B, Maller J, Daly MJ: Haploview: analysis and visualization of LD and haplotype maps. Bioinformatics 2005, 21: 263\u2013265. 10.1093\/bioinformatics\/bth457","journal-title":"Bioinformatics"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-7-468.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,9]],"date-time":"2023-05-09T13:36:45Z","timestamp":1683639405000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-7-468"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,10,23]]},"references-count":27,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2006,12]]}},"alternative-id":["1207"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-7-468","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2006,10,23]]},"assertion":[{"value":"18 April 2006","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 October 2006","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 October 2006","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"468"}}