{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T23:03:21Z","timestamp":1781737401313,"version":"3.54.5"},"reference-count":30,"publisher":"Springer Science and Business Media LLC","issue":"10","license":[{"start":{"date-parts":[[2022,8,10]],"date-time":"2022-08-10T00:00:00Z","timestamp":1660089600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,8,10]],"date-time":"2022-08-10T00:00:00Z","timestamp":1660089600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100000266","name":"engineering and physical sciences research council","doi-asserted-by":"publisher","award":["EP\/S001646\/1"],"award-info":[{"award-number":["EP\/S001646\/1"]}],"id":[{"id":"10.13039\/501100000266","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2022,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Learning from data that contain missing values represents a common phenomenon in many domains. Relatively few Bayesian Network structure learning algorithms account for missing data, and those that do tend to rely on standard approaches that assume missing data are missing at random, such as the Expectation-Maximisation algorithm. Because missing data are often systematic, there is a need for more pragmatic methods that can effectively deal with data sets containing missing values not missing at random. The absence of approaches that deal with systematic missing data impedes the application of BN structure learning methods to real-world problems where missingness are not random. This paper describes three variants of greedy search structure learning that utilise pairwise deletion and inverse probability weighting to maximally leverage the observed data and to limit potential bias caused by missing values. The first two of the variants can be viewed as sub-versions of the third and best performing variant, but are important in their own in illustrating the successive improvements in learning accuracy. The empirical investigations show that the proposed approach outperforms the commonly used and state-of-the-art Structural EM algorithm, both in terms of learning accuracy and efficiency, as well as both when data are missing at random and not at random.<\/jats:p>","DOI":"10.1007\/s10994-022-06195-8","type":"journal-article","created":{"date-parts":[[2022,8,10]],"date-time":"2022-08-10T21:02:32Z","timestamp":1660165352000},"page":"3867-3896","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Greedy structure learning from data that contain systematic missing values"],"prefix":"10.1007","volume":"111","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4217-4194","authenticated-orcid":false,"given":"Yang","family":"Liu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anthony C.","family":"Constantinou","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,8,10]]},"reference":[{"issue":"1","key":"6195_CR1","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1002\/mpr.329","volume":"20","author":"MJ Azur","year":"2011","unstructured":"Azur, M. J., Stuart, E. A., Frangakis, C., & Leaf, P. J. (2011). Multiple imputation by chained equations: what is it and how does it work? International Journal of Methods in Psychiatric Research, 20(1), 40\u201349.","journal-title":"International Journal of Methods in Psychiatric Research"},{"key":"6195_CR2","doi-asserted-by":"publisher","first-page":"1047","DOI":"10.1214\/13-EJS802","volume":"7","author":"N Balov","year":"2013","unstructured":"Balov, N., et al. (2013). Consistent model selection of discrete Bayesian networks from incomplete data. Electronic Journal of Statistics, 7, 1047\u20131077.","journal-title":"Electronic Journal of Statistics"},{"key":"6195_CR3","doi-asserted-by":"publisher","first-page":"145","DOI":"10.1016\/j.ijar.2021.07.015","volume":"138","author":"T Bodewes","year":"2021","unstructured":"Bodewes, T., & Scutari, M. (2021). Learning Bayesian networks from incomplete data with the node-average likelihood. International Journal of Approximate Reasoning, 138, 145\u2013160.","journal-title":"International Journal of Approximate Reasoning"},{"issue":"Nov","key":"6195_CR4","first-page":"507","volume":"3","author":"DM Chickering","year":"2002","unstructured":"Chickering, D. M. (2002). Optimal structure identification with greedy search. Journal of Machine Learning Research, 3(Nov), 507\u2013554.","journal-title":"Journal of Machine Learning Research"},{"key":"6195_CR5","doi-asserted-by":"publisher","first-page":"75","DOI":"10.1016\/j.artmed.2016.01.002","volume":"67","author":"AC Constantinou","year":"2016","unstructured":"Constantinou, A. C., Fenton, N., Marsh, W., & Radlinski, L. (2016). From complex questionnaire and interviewing data to intelligent Bayesian network models for medical decision support. Artificial Intelligence in Medicine, 67, 75\u201393.","journal-title":"Artificial Intelligence in Medicine"},{"key":"6195_CR6","doi-asserted-by":"publisher","first-page":"151","DOI":"10.1016\/j.ijar.2021.01.001","volume":"131","author":"AC Constantinou","year":"2021","unstructured":"Constantinou, A. C., Liu, Y., Chobtham, K., Guo, Z., & Kitson, N. K. (2021). Large-scale empirical validation of Bayesian Network structure learning algorithms with noisy data. International Journal of Approximate Reasoning, 131, 151\u2013188.","journal-title":"International Journal of Approximate Reasoning"},{"key":"6195_CR7","unstructured":"Cussens, J. (2011). Bayesian network learning with cutting planes. In Proceedings of the 27th conference on uncertainty in artificial intelligence (UAI 2011), AUAI Press, pp. 153\u2013160."},{"key":"6195_CR8","unstructured":"Friedman, N., et al. (1997). Learning belief networks in the presence of missing values and hidden variables. In ICML, Citeseer, Vol. 97, pp. 125\u2013133."},{"key":"6195_CR9","unstructured":"Gain, A., & Shpitser, I. (2018). Structure learning under missing data. In International conference on probabilistic graphical models, PMLR, pp. 121\u2013132."},{"issue":"1","key":"6195_CR10","doi-asserted-by":"publisher","first-page":"106","DOI":"10.1007\/s10618-010-0178-6","volume":"22","author":"JA G\u00e1mez","year":"2011","unstructured":"G\u00e1mez, J. A., Mateo, J. L., & Puerta, J. M. (2011). Learning bayesian networks by hill climbing: Efficient methods based on progressive restriction of the neighborhood. Data Mining and Knowledge Discovery, 22(1), 106\u2013148.","journal-title":"Data Mining and Knowledge Discovery"},{"key":"6195_CR11","doi-asserted-by":"publisher","first-page":"549","DOI":"10.1146\/annurev.psych.58.110405.085530","volume":"60","author":"JW Graham","year":"2009","unstructured":"Graham, J. W. (2009). Missing data analysis: Making it work in the real world. Annual Review of Psychology, 60, 549\u2013576.","journal-title":"Annual Review of Psychology"},{"issue":"3","key":"6195_CR12","doi-asserted-by":"publisher","first-page":"197","DOI":"10.1007\/BF00994016","volume":"20","author":"D Heckerman","year":"1995","unstructured":"Heckerman, D., Geiger, D., & Chickering, D. M. (1995). Learning Bayesian networks: The combination of knowledge and statistical data. Machine Learning, 20(3), 197\u2013243.","journal-title":"Machine Learning"},{"issue":"260","key":"6195_CR13","doi-asserted-by":"publisher","first-page":"663","DOI":"10.1080\/01621459.1952.10483446","volume":"47","author":"DG Horvitz","year":"1952","unstructured":"Horvitz, D. G., & Thompson, D. J. (1952). A generalization of sampling without replacement from a finite universe. Journal of the American Statistical Association, 47(260), 663\u2013685.","journal-title":"Journal of the American Statistical Association"},{"issue":"1","key":"6195_CR14","first-page":"51","volume":"10","author":"C John","year":"2019","unstructured":"John, C., Ekpenyong, E. J., & Nworu, C. C. (2019). Imputation of missing values in economic and financial time series data using five principal component analysis approaches. CBN Journal of Applied Statistics, 10(1), 51\u201373.","journal-title":"CBN Journal of Applied Statistics"},{"key":"6195_CR15","doi-asserted-by":"crossref","unstructured":"Mohan, K., & Pearl, J. (2021). Graphical models for processing missing data. Journal of the American Statistical Association pp 1\u201316.","DOI":"10.1080\/01621459.2021.1874961"},{"key":"6195_CR16","unstructured":"Mohan, K., Pearl, J., & Tian, J. (2013). Graphical models for inference with missing data. In Burges, C.J.C., Bottou, L., Welling, M., Ghahramani, Z., & Weinberger, K.Q. (Eds.) Advances in Neural Information Processing Systems, Curran Associates, Inc., vol\u00a026, https:\/\/proceedings.neurips.cc\/paper\/2013\/file\/0ff8033cf9437c213ee13937b1c4c455-Paper.pdf."},{"key":"6195_CR17","doi-asserted-by":"publisher","first-page":"157","DOI":"10.2147\/CLEP.S129785","volume":"9","author":"AB Pedersen","year":"2017","unstructured":"Pedersen, A. B., Mikkelsen, E. M., Cronin-Fenton, D., Kristensen, N. R., Pham, T. M., Pedersen, L., & Petersen, I. (2017). Missing data and multiple imputation in clinical epidemiological research. Clinical Epidemiology, 9, 157.","journal-title":"Clinical Epidemiology"},{"issue":"3","key":"6195_CR18","doi-asserted-by":"publisher","first-page":"581","DOI":"10.1093\/biomet\/63.3.581","volume":"63","author":"DB Rubin","year":"1976","unstructured":"Rubin, D. B. (1976). Inference and missing data. Biometrika, 63(3), 581\u2013592.","journal-title":"Biometrika"},{"key":"6195_CR19","volume-title":"Multiple imputation for nonresponse in surveys","author":"DB Rubin","year":"2004","unstructured":"Rubin, D. B. (2004). Multiple imputation for nonresponse in surveys (Vol. 81). New York: Wiley."},{"issue":"12","key":"6195_CR20","doi-asserted-by":"publisher","first-page":"329","DOI":"10.3390\/a13120329","volume":"13","author":"A Ruggieri","year":"2020","unstructured":"Ruggieri, A., Stranieri, F., Stella, F., & Scutari, M. (2020). Hard and soft EM in Bayesian network learning from incomplete data. Algorithms, 13(12), 329.","journal-title":"Algorithms"},{"issue":"2","key":"6195_CR21","doi-asserted-by":"publisher","first-page":"461","DOI":"10.1214\/aos\/1176344136","volume":"6","author":"G Schwarz","year":"1978","unstructured":"Schwarz, G., et al. (1978). Estimating the dimension of a model. Annals of Statistics, 6(2), 461\u2013464.","journal-title":"Annals of Statistics"},{"key":"6195_CR22","doi-asserted-by":"crossref","unstructured":"Scutari, M. (2010). Learning Bayesian networks with the bnlearn R package. Journal of Statistical Software, 35(3).","DOI":"10.18637\/jss.v035.i03"},{"key":"6195_CR23","unstructured":"Silander, T., Lepp\u00e4-Aho, J., J\u00e4\u00e4saari, E., & Roos, T. (2018). Quotient normalized maximum likelihood criterion for learning Bayesian network structures. In International conference on artificial intelligence and statistics, PMLR, pp. 948\u2013957."},{"key":"6195_CR24","volume-title":"Causation, prediction, and search","author":"P Spirtes","year":"2000","unstructured":"Spirtes, P., Glymour, C. N., Scheines, R., & Heckerman, D. (2000). Causation, prediction, and search. Cambridge: MIT press."},{"issue":"1","key":"6195_CR25","doi-asserted-by":"publisher","first-page":"47","DOI":"10.1007\/s41060-017-0094-6","volume":"6","author":"EV Strobl","year":"2018","unstructured":"Strobl, E. V., Visweswaran, S., & Spirtes, P. L. (2018). Fast causal inference with non-random missingness by test-wise deletion. International Journal of Data Science and Analytics, 6(1), 47\u201362.","journal-title":"International Journal of Data Science and Analytics"},{"key":"6195_CR26","doi-asserted-by":"publisher","first-page":"297","DOI":"10.1016\/j.neucom.2018.08.067","volume":"318","author":"Y Tian","year":"2018","unstructured":"Tian, Y., Zhang, K., Li, J., Lin, X., & Yang, B. (2018). LSTM-based traffic flow prediction with missing data. Neurocomputing, 318, 297\u2013305.","journal-title":"Neurocomputing"},{"key":"6195_CR27","first-page":"376","volume":"2","author":"I Tsamardinos","year":"2003","unstructured":"Tsamardinos, I., Aliferis, C. F., Statnikov, A. R., & Statnikov, E. (2003). Algorithms for large scale Markov blanket discovery. FLAIRS conference, 2, 376\u2013380.","journal-title":"FLAIRS conference"},{"issue":"1","key":"6195_CR28","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1007\/s10994-006-6889-7","volume":"65","author":"I Tsamardinos","year":"2006","unstructured":"Tsamardinos, I., Brown, L. E., & Aliferis, C. F. (2006). The max-min hill-climbing Bayesian network structure learning algorithm. Machine Learning, 65(1), 31\u201378.","journal-title":"Machine Learning"},{"key":"6195_CR29","unstructured":"Tu, R., Zhang, C., Ackermann, P., Mohan, K., Kjellstr\u00f6m, H., & Zhang, K. (2019). Causal discovery in the presence of missing data. In The 22nd international conference on artificial intelligence and statistics, PMLR, pp. 1762\u20131770."},{"key":"6195_CR30","doi-asserted-by":"crossref","unstructured":"Zemicheal, T., & Dietterich, T.G. (2019). Anomaly detection in the presence of missing values for weather data quality control. In Proceedings of the 2nd ACM SIGCAS conference on computing and sustainable societies, pp. 65\u201373.","DOI":"10.1145\/3314344.3332490"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06195-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-022-06195-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06195-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,10,7]],"date-time":"2022-10-07T17:14:06Z","timestamp":1665162846000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-022-06195-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,8,10]]},"references-count":30,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2022,10]]}},"alternative-id":["6195"],"URL":"https:\/\/doi.org\/10.1007\/s10994-022-06195-8","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,8,10]]},"assertion":[{"value":"15 July 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 May 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 May 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 August 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}