{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,8]],"date-time":"2026-01-08T09:51:11Z","timestamp":1767865871964,"version":"3.49.0"},"reference-count":20,"publisher":"Springer Science and Business Media LLC","issue":"S2","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2009,2]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec><jats:title>Background<\/jats:title><jats:p>Bayesian networks are powerful instruments to learn genetic models from association studies data. They are able to derive the existing correlation between genetic markers and phenotypic traits and, at the same time, to find the relationships between the markers themselves. However, learning Bayesian networks is often non-trivial due to the high number of variables to be taken into account in the model with respect to the instances of the dataset. Therefore, it becomes very interesting to use an abstraction of the variable space that suitably reduces its dimensionality without losing information. In this paper we present a new strategy to achieve this goal by mapping the SNPs related to the same gene to one meta-variable. In order to assign states to the meta-variables we employ an approach based on classification trees.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>We applied our approach to data coming from a genome-wide scan on 288 individuals affected by arterial hypertension and 271 nonagenarians without history of hypertension. After pre-processing, we focused on a subset of 24 SNPs. We compared the performance of the proposed approach with the Bayesian network learned with SNPs as variables and with the network learned with haplotypes as meta-variables. The results were obtained by running a hold-out experiment five times. The mean accuracy of the new method was 64.28%, while the mean accuracy of the SNPs network was 58.99% and the mean accuracy of the haplotype network was 54.57%.<\/jats:p><\/jats:sec><jats:sec><jats:title>Conclusion<\/jats:title><jats:p>The new approach presented in this paper is able to derive a gene-based predictive model based on SNPs data. Such model is more parsimonious than the one based on single SNPs, while preserving the capability of highlighting predictive SNPs configurations. The prediction performance of this approach was consistently superior to the SNP-based and the haplotype-based one in all the test sets of the evaluation procedure. The method can be then considered as an alternative way to analyze the data coming from association studies.<\/jats:p><\/jats:sec>","DOI":"10.1186\/1471-2105-10-s2-s7","type":"journal-article","created":{"date-parts":[[2009,2,5]],"date-time":"2009-02-05T16:29:05Z","timestamp":1233851345000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Phenotype forecasting with SNPs data through gene-based Bayesian networks"],"prefix":"10.1186","volume":"10","author":[{"given":"Alberto","family":"Malovini","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Angelo","family":"Nuzzo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fulvia","family":"Ferrazzi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Annibale A","family":"Puca","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Riccardo","family":"Bellazzi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2009,2,5]]},"reference":[{"issue":"2","key":"3266_CR1","doi-asserted-by":"publisher","first-page":"95","DOI":"10.1038\/nrg1521","volume":"6","author":"JN Hirschhorn","year":"2005","unstructured":"Hirschhorn JN, Daly MJ: Genome-wide association studies for common diseases and complex traits. Nature reviews 2005, 6(2):95\u2013108.","journal-title":"Nature reviews"},{"issue":"5","key":"3266_CR2","doi-asserted-by":"publisher","first-page":"356","DOI":"10.1038\/nrg2344","volume":"9","author":"MI McCarthy","year":"2008","unstructured":"McCarthy MI, Abecasis GR, Cardon LR, Goldstein DB, Little J, Ioannidis JP, Hirschhorn JN: Genome-wide association studies for complex traits: consensus, uncertainty and challenges. Nature reviews 2008, 9(5):356\u2013369. 10.1038\/nrg2344","journal-title":"Nature reviews"},{"key":"3266_CR3","first-page":"205","volume-title":"Systems Bioinformatics: An Engineering Case-Based Approach","author":"P Sebastiani","year":"2007","unstructured":"Sebastiani P, Abad-Grau MM: Bayesian Networks for Genetic Analysis. In Systems Bioinformatics: An Engineering Case-Based Approach. Edited by: Alterovitz G, Ramoni MF. Artech House; 2007:205\u2013227."},{"issue":"4","key":"3266_CR4","doi-asserted-by":"publisher","first-page":"435","DOI":"10.1038\/ng1533","volume":"37","author":"P Sebastiani","year":"2005","unstructured":"Sebastiani P, Ramoni MF, Nolan V, Baldwin CT, Steinberg MH: Genetic dissection and prognostic modeling of overt stroke in sickle cell anemia. Nat Genet 2005, 37(4):435\u2013440. 10.1038\/ng1533","journal-title":"Nat Genet"},{"issue":"1","key":"3266_CR5","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1089\/cmb.2005.12.1","volume":"12","author":"A Rodin","year":"2005","unstructured":"Rodin A, Mosley TH Jr, Clark AG, Sing CF, Boerwinkle E: Mining genetic epidemiology data with Bayesian networks application to APOE gene variation and plasma lipid levels. J Comput Biol 2005, 12(1):1\u201311. 10.1089\/cmb.2005.12.1","journal-title":"J Comput Biol"},{"key":"3266_CR6","volume-title":"Machine learning","author":"TM Mitchell","year":"1997","unstructured":"Mitchell TM: Machine learning. McGraw-Hill; 1997."},{"key":"3266_CR7","volume-title":"Probabilistic reasoning in intelligent systems: networks of plausible inference","author":"J Pearl","year":"1988","unstructured":"Pearl J: Probabilistic reasoning in intelligent systems: networks of plausible inference. Morgan Kaufmann; 1988."},{"issue":"2","key":"3266_CR8","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1111\/j.2517-6161.1988.tb01721.x","volume":"50","author":"SL Lauritzen","year":"1988","unstructured":"Lauritzen SL, Spiegelhalter SJ: Local Computations with Probabilities on Graphical Structures and Their Application to Expert Systems. Journal of the Royal Statistical Society Series B 1988, 50(2):157\u2013224.","journal-title":"Journal of the Royal Statistical Society Series B"},{"key":"3266_CR9","volume-title":"Data mining: Practical machine learning tools and techniques with Java implementations","author":"IH Witten","year":"2000","unstructured":"Witten IH, Frank E: Data mining: Practical machine learning tools and techniques with Java implementations. San Francisco, CA, USA: Morgan Kaufman; 2000."},{"key":"3266_CR10","volume-title":"Progress in Machine Learning","author":"T Niblett","year":"1986","unstructured":"Niblett T, Bratko I: Constructing Decision Trees in Noisy Domains. In Progress in Machine Learning. Edited by: Bratko I, Lavrac N. England: Sigma Press; 1986."},{"key":"3266_CR11","volume-title":"White Paper,","author":"J Demsar","year":"2004","unstructured":"Demsar J, Zupan B, Leban G: Orange: From experimental machine learning to interactive data mining. White Paper, Faculty of Computer and Information Science, University of Ljubljana; 2004. [http:\/\/www.ailab.si\/orange]"},{"key":"3266_CR12","first-page":"309","volume":"9","author":"GF Cooper","year":"1992","unstructured":"Cooper GF, Herskovits E: A Bayesian method for the induction of probabilistic networks from data. Machine Learning 1992, 9: 309\u2013347.","journal-title":"Machine Learning"},{"key":"3266_CR13","unstructured":"Bayesware Discoverer[http:\/\/www.bayesware.com\/products\/discoverer\/discoverer.html]"},{"issue":"8915","key":"3266_CR14","doi-asserted-by":"publisher","first-page":"101","DOI":"10.1016\/S0140-6736(94)91285-8","volume":"344","author":"PK Whelton","year":"1994","unstructured":"Whelton PK: Epidemiology of hypertension. Lancet 1994, 344(8915):101\u2013106. 10.1016\/S0140-6736(94)91285-8","journal-title":"Lancet"},{"issue":"7","key":"3266_CR15","doi-asserted-by":"publisher","first-page":"572","DOI":"10.1161\/01.HYP.8.7.572","volume":"8","author":"SB Harrap","year":"1986","unstructured":"Harrap SB: Genetic analysis of blood pressure and sodium balance in spontaneously hypertensive rats. Hypertension 1986, 8(7):572\u2013582.","journal-title":"Hypertension"},{"issue":"3","key":"3266_CR16","doi-asserted-by":"publisher","first-page":"559","DOI":"10.1086\/519795","volume":"81","author":"S Purcell","year":"2007","unstructured":"Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MA, Bender D, Maller J, Sklar P, de Bakker PI, Daly MJ, et al.: PLINK: a tool set for whole-genome association and population-based linkage analyses. American journal of human genetics 2007, 81(3):559\u2013575. 10.1086\/519795","journal-title":"American journal of human genetics"},{"key":"3266_CR17","doi-asserted-by":"crossref","unstructured":"Nuzzo A, Riva A: A Knowledge Management Tool for Translational Research. AMIA Summit on Translational Bioinformatics: 2008 17.","DOI":"10.1186\/1471-2105-10-278"},{"key":"3266_CR18","first-page":"1","volume":"7","author":"J Demsar","year":"2006","unstructured":"Demsar J: Statistical Comparisons of Classifiers over Multiple Data Sets. Journal of Machine Learning Research 2006, 7: 1\u201330.","journal-title":"Journal of Machine Learning Research"},{"issue":"2","key":"3266_CR19","doi-asserted-by":"publisher","first-page":"263","DOI":"10.1093\/bioinformatics\/bth457","volume":"21","author":"JC Barrett","year":"2005","unstructured":"Barrett JC, Fry B, Maller J, Daly MJ: Haploview: analysis and visualization of LD and haplotype maps. Bioinformatics (Oxford, England) 2005, 21(2):263\u2013265. 10.1093\/bioinformatics\/bth457","journal-title":"Bioinformatics (Oxford, England)"},{"issue":"1","key":"3266_CR20","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1111\/j.2517-6161.1977.tb01600.x","volume":"39","author":"AP Dempster","year":"1977","unstructured":"Dempster AP, Laird NM, Rubin DB: Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society Series B 1977, 39(1):1\u201338.","journal-title":"Journal of the Royal Statistical Society Series B"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-10-S2-S7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,7]],"date-time":"2025-02-07T11:49:54Z","timestamp":1738928994000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-10-S2-S7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,2]]},"references-count":20,"journal-issue":{"issue":"S2","published-print":{"date-parts":[[2009,2]]}},"alternative-id":["3266"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-10-s2-s7","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2009,2]]},"assertion":[{"value":"5 February 2009","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"S7"}}