{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,16]],"date-time":"2026-01-16T08:29:37Z","timestamp":1768552177691,"version":"3.49.0"},"reference-count":23,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2006,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>Single nucleotide polymorphisms (SNP) constitute more than 90% of the genetic variation, and hence can account for most trait differences among individuals in a given species. Polymorphism detection software PolyBayes and PolyPhred give high false positive SNP predictions even with stringent parameter values. We developed a machine learning (ML) method to augment PolyBayes to improve its prediction accuracy. ML methods have also been successfully applied to other bioinformatics problems in predicting genes, promoters, transcription factor binding sites and protein structures.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>The ML program C4.5 was applied to a set of features in order to build a SNP classifier from training data based on human expert decisions (True\/False). The training data were 27,275 candidate SNP generated by sequencing 1973 STS (sequence tag sites) (12 Mb) in both directions from 6 diverse homozygous soybean cultivars and PolyBayes analysis. Test data of 18,390 candidate SNP were generated similarly from 1359 additional STS (8 Mb). SNP from both sets were classified by experts. After training the ML classifier, it agreed with the experts on 97.3% of test data compared with 7.8% agreement between PolyBayes and experts. The PolyBayes positive predictive values (PPV) (i.e., fraction of candidate SNP being real) were 7.8% for all predictions and 16.7% for those with 100% posterior probability of being real. Using ML improved the PPV to 84.8%, a 5- to 10-fold increase. While both ML and PolyBayes produced a similar number of true positives, the ML program generated only 249 false positives as compared to 16,955 for PolyBayes. The complexity of the soybean genome may have contributed to high false SNP predictions by PolyBayes and hence results may differ for other genomes.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion<\/jats:title>\n            <jats:p>A machine learning (ML) method was developed as a supplementary feature to the polymorphism detection software for improving prediction accuracies. The results from this study indicate that a trained ML classifier can significantly reduce human intervention and in this case achieved a 5\u201310 fold enhanced productivity. The optimized feature set and ML framework can also be applied to all polymorphism discovery software. ML support software is written in Perl and can be easily integrated into an existing SNP discovery pipeline.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-7-4","type":"journal-article","created":{"date-parts":[[2006,1,21]],"date-time":"2006-01-21T19:14:22Z","timestamp":1137870862000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":30,"title":["Application of machine learning in SNP discovery"],"prefix":"10.1186","volume":"7","author":[{"given":"Lakshmi K","family":"Matukumalli","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"John J","family":"Grefenstette","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David L","family":"Hyten","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ik-Young","family":"Choi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Perry B","family":"Cregan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Curtis P","family":"Van Tassell","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2006,1,6]]},"reference":[{"key":"743_CR1","doi-asserted-by":"publisher","first-page":"1292","DOI":"10.1093\/bioinformatics\/bth085","volume":"20","author":"YD Cai","year":"2004","unstructured":"Cai YD, Doig AJ: Prediction of Saccharomyces cerevisiae protein functional class from functional domain composition. Bioinformatics 2004, 20: 1292\u20131300. 10.1093\/bioinformatics\/bth085","journal-title":"Bioinformatics"},{"key":"743_CR2","doi-asserted-by":"publisher","first-page":"2429","DOI":"10.1093\/bioinformatics\/bth267","volume":"20","author":"T Li","year":"2004","unstructured":"Li T, Zhang C, Ogihara M: A comparative study of feature selection and multiclass classification methods for tissue classification based on gene expression. Bioinformatics 2004, 20: 2429\u20132437. 10.1093\/bioinformatics\/bth267","journal-title":"Bioinformatics"},{"key":"743_CR3","doi-asserted-by":"publisher","first-page":"665","DOI":"10.1016\/S0196-9781(03)00133-5","volume":"24","author":"YD Cai","year":"2003","unstructured":"Cai YD, Liu XJ, Li YX, Xu XB, Chou KC: Prediction of beta-turns with learning machines. Peptides 2003, 24: 665\u2013669. 10.1016\/S0196-9781(03)00133-5","journal-title":"Peptides"},{"key":"743_CR4","first-page":"421","volume":"95","author":"PB Dobrokhotov","year":"2003","unstructured":"Dobrokhotov PB, Goutte C, Veuthey AL, Gaussier E: A probabilistic information retrieval approach to medical annotation in SWISS-PROT. Stud Health Technol Inform 2003, 95: 421\u2013426.","journal-title":"Stud Health Technol Inform"},{"key":"743_CR5","doi-asserted-by":"publisher","first-page":"38","DOI":"10.1186\/1471-2105-5-38","volume":"5","author":"LV Zhang","year":"2004","unstructured":"Zhang LV, Wong SL, King OD, Roth FP: Predicting co-complexed protein pairs using genomic and proteomic data integration. BMC Bioinformatics 2004, 5: 38. 10.1186\/1471-2105-5-38","journal-title":"BMC Bioinformatics"},{"key":"743_CR6","doi-asserted-by":"publisher","first-page":"355","DOI":"10.1261\/rna.5890304","volume":"10","author":"LY Han","year":"2004","unstructured":"Han LY, Cai CZ, Lo SL, Chung MC, Chen YZ: Prediction of RNA-binding proteins from primary sequence by a support vector machine approach. RNA 2004, 10: 355\u2013368. 10.1261\/rna.5890304","journal-title":"RNA"},{"key":"743_CR7","volume-title":"Bioinformatics","author":"E Frank","year":"2004","unstructured":"Frank E, Hall M, Trigg L, Holmes G, Witten IH: Data mining in bioinformatics using Weka. Bioinformatics 2004."},{"key":"743_CR8","volume-title":"C4.5: programs for machine learning","author":"JR Quinlan","year":"1993","unstructured":"Quinlan JR: C4.5: programs for machine learning. San Francisco, CA, USA, Morgan Kaufmann Publishers Inc; 1993."},{"key":"743_CR9","doi-asserted-by":"publisher","first-page":"586","DOI":"10.1093\/bioinformatics\/btg461","volume":"20","author":"P Pavlidis","year":"2004","unstructured":"Pavlidis P, Wapinski I, Noble WS: Support vector machine classification on the web. Bioinformatics 2004, 20: 586\u2013587. 10.1093\/bioinformatics\/btg461","journal-title":"Bioinformatics"},{"key":"743_CR10","doi-asserted-by":"publisher","first-page":"452","DOI":"10.1038\/70570","volume":"23","author":"GT Marth","year":"1999","unstructured":"Marth GT, Korf I, Yandell MD, Yeh RT, Gu Z, Zakeri H, Stitziel NO, Hillier L, Kwok PY, Gish WR: A general approach to single-nucleotide polymorphism discovery. Nat Genet 1999, 23: 452\u2013456. 10.1038\/70570","journal-title":"Nat Genet"},{"key":"743_CR11","doi-asserted-by":"publisher","first-page":"2745","DOI":"10.1093\/nar\/25.14.2745","volume":"25","author":"DA Nickerson","year":"1997","unstructured":"Nickerson DA, Tobe VO, Taylor SL: PolyPhred: automating the detection and genotyping of single nucleotide substitutions using fluorescence-based resequencing. Nucleic Acids Res 1997, 25: 2745\u20132751. 10.1093\/nar\/25.14.2745","journal-title":"Nucleic Acids Res"},{"key":"743_CR12","doi-asserted-by":"crossref","first-page":"1123","DOI":"10.1093\/genetics\/163.3.1123","volume":"163","author":"YL Zhu","year":"2003","unstructured":"Zhu YL, Song QJ, Hyten DL, Van Tassell CP, Matukumalli LK, Grimm DR, Hyatt SM, Fickus EW, Young ND, Cregan PB: Single-nucleotide polymorphisms in soybean. Genetics 2003, 163: 1123\u20131134.","journal-title":"Genetics"},{"key":"743_CR13","first-page":"97","volume-title":"Speciation and Cytogenetics","author":"HH Hadley","year":"1973","unstructured":"Hadley HH, Hymowitz T: Speciation and Cytogenetics. Madison, WI, Agron. Monogr; 1973:97\u2013116."},{"key":"743_CR14","doi-asserted-by":"publisher","first-page":"595","DOI":"10.2307\/2442301","volume":"67","author":"JA Lackey","year":"1980","unstructured":"Lackey JA: Chromosome numbers in the Phaseoleae (Fabaceae:Faboideae) and their relation to taxonomy. Am J Bot 1980, 67: 595\u2013602.","journal-title":"Am J Bot"},{"key":"743_CR15","doi-asserted-by":"publisher","first-page":"868","DOI":"10.1139\/g04-047","volume":"47","author":"JA Schlueter","year":"2004","unstructured":"Schlueter JA, Dixon P, Granger C, Grant D, Clark L, Doyle JJ, Shoemaker RC: Mining EST databases to resolve evolutionary events in major crop species. Genome 2004, 47: 868\u2013876. 10.1139\/g04-047","journal-title":"Genome"},{"key":"743_CR16","doi-asserted-by":"publisher","first-page":"513","DOI":"10.1038\/35035083","volume":"407","author":"D Altshuler","year":"2000","unstructured":"Altshuler D, Pollara VJ, Cowles CR, Van Etten WJ, Baldwin J, Linton L, Lander ES: An SNP map of the human genome generated by reduced representation shotgun sequencing. Nature 2000, 407: 513\u2013516. 10.1038\/35035083","journal-title":"Nature"},{"key":"743_CR17","doi-asserted-by":"publisher","first-page":"1679","DOI":"10.1101\/gr.287302","volume":"12","author":"Z Zhao","year":"2002","unstructured":"Zhao Z, Boerwinkle E: Neighboring-nucleotide effects on single nucleotide polymorphisms: a study of 2.6 million polymorphisms across the human genome. Genome Res 2002, 12: 1679\u20131686. 10.1101\/gr.287302","journal-title":"Genome Res"},{"key":"743_CR18","doi-asserted-by":"publisher","first-page":"421","DOI":"10.1093\/bioinformatics\/btf881","volume":"19","author":"G Barker","year":"2003","unstructured":"Barker G, Batley J, O' Sullivan H, Edwards KJ, Edwards D: Redundancy based detection of sequence polymorphisms in expressed sequence tag data using autoSNP. Bioinformatics 2003, 19: 421\u2013422. 10.1093\/bioinformatics\/btf881","journal-title":"Bioinformatics"},{"key":"743_CR19","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1104\/pp.102.019422","volume":"132","author":"J Batley","year":"2003","unstructured":"Batley J, Barker G, O'Sullivan H, Edwards KJ, Edwards D: Mining for single nucleotide polymorphisms and insertions\/deletions in maize expressed sequence tag data. Plant Physiol 2003, 132: 84\u201391. 10.1104\/pp.102.019422","journal-title":"Plant Physiol"},{"key":"743_CR20","volume-title":"Repeat Masker Open - 3.0","author":"AFA Smit","year":"1996","unstructured":"Smit AFA, Hubley R, Green P: Repeat Masker Open - 3.0.1996. [http:\/\/www.repeatmasker.org]"},{"key":"743_CR21","unstructured":"Supplementary Information[http:\/\/bfgl.anri.barc.usda.gov\/ML\/]"},{"key":"743_CR22","unstructured":"BioPerl[http:\/\/www.bioperl.org\/]"},{"key":"743_CR23","unstructured":"CPAN[http:\/\/www.cpan.org\/]"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-7-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T03:21:56Z","timestamp":1630466516000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-7-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,1,6]]},"references-count":23,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2006,12]]}},"alternative-id":["743"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-7-4","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2006,1,6]]},"assertion":[{"value":"11 August 2005","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 January 2006","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 January 2006","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"4"}}