{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T16:24:27Z","timestamp":1761582267206,"version":"build-2065373602"},"reference-count":35,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2023,3,27]],"date-time":"2023-03-27T00:00:00Z","timestamp":1679875200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Liverpool John Moores University"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>A lack of transparency in machine learning models can limit their application. We show that analysis of variance (ANOVA) methods extract interpretable predictive models from them. This is possible because ANOVA decompositions represent multivariate functions as sums of functions of fewer variables. Retaining the terms in the ANOVA summation involving functions of only one or two variables provides an efficient method to open black box classifiers. The proposed method builds generalised additive models (GAMs) by application of L1 regularised logistic regression to the component terms retained from the ANOVA decomposition of the logit function. The resulting GAMs are derived using two alternative measures, Dirac and Lebesgue. Both measures produce functions that are smooth and consistent. The term partial responses in structured models (PRiSM) describes the family of models that are derived from black box classifiers by application of ANOVA decompositions. We demonstrate their interpretability and performance for the multilayer perceptron, support vector machines and gradient-boosting machines applied to synthetic data and several real-world data sets, namely Pima Diabetes, German Credit Card, and Statlog Shuttle from the UCI repository. The GAMs are shown to be compliant with the basic principles of a formal framework for interpretability.<\/jats:p>","DOI":"10.3390\/a16040181","type":"journal-article","created":{"date-parts":[[2023,3,27]],"date-time":"2023-03-27T06:46:19Z","timestamp":1679899579000},"page":"181","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["How to Open a Black Box Classifier for Tabular Data"],"prefix":"10.3390","volume":"16","author":[{"given":"Bradley","family":"Walters","sequence":"first","affiliation":[{"name":"School of Computer Science and Mathematics, Liverpool John Moores University, Liverpool L3 2AF, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9927-3209","authenticated-orcid":false,"given":"Sandra","family":"Ortega-Martorell","sequence":"additional","affiliation":[{"name":"School of Computer Science and Mathematics, Liverpool John Moores University, Liverpool L3 2AF, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5679-7501","authenticated-orcid":false,"given":"Ivan","family":"Olier","sequence":"additional","affiliation":[{"name":"School of Computer Science and Mathematics, Liverpool John Moores University, Liverpool L3 2AF, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Paulo J. G.","family":"Lisboa","sequence":"additional","affiliation":[{"name":"School of Computer Science and Mathematics, Liverpool John Moores University, Liverpool L3 2AF, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,27]]},"reference":[{"key":"ref_1","first-page":"1","article-title":"Learning Certifiably Optimal Rule Lists for Categorical Data","volume":"18","author":"Angelino","year":"2018","journal-title":"J. Mach. Learn. Res."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"149","DOI":"10.1186\/1471-2105-10-149","article-title":"How to Find Simple and Accurate Rules for Viral Protease Cleavage Specificities","volume":"10","author":"Etchells","year":"2009","journal-title":"BMC Bioinform."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"581","DOI":"10.1111\/jgh.15384","article-title":"Opening the Black Box of AI-Medicine","volume":"36","author":"Poon","year":"2021","journal-title":"J. Gastroenterol. Hepatol."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1016\/j.jclinepi.2019.02.004","article-title":"A Systematic Review Shows No Performance Benefit of Machine Learning over Logistic Regression for Clinical Prediction Models","volume":"110","author":"Christodoulou","year":"2019","journal-title":"J. Clin. Epidemiol."},{"key":"ref_5","first-page":"93","article-title":"A Survey of Methods for Explaining Black Box Models","volume":"51","author":"Guidotti","year":"2018","journal-title":"ACM Comput. Surv."},{"key":"ref_6","unstructured":"Sarle, W.S. (1994, January 10\u201313). Neural Networks and Statistical Models. Proceedings of the Nineteenth Annual SAS Users Group International Conference, Dallas, TX, USA."},{"key":"ref_7","first-page":"3459","article-title":"Odds Ratio Function Estimation Using a Generalized Additive Neural Network","volume":"32","author":"Papoila","year":"2019","journal-title":"Neural Comput. Appl."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1038\/s41746-020-00377-1","article-title":"Development and Validation of an Interpretable Neural Network for Prediction of Postoperative In-Hospital Mortality","volume":"4","author":"Lee","year":"2021","journal-title":"NPJ Digit. Med."},{"key":"ref_9","unstructured":"Alvarez-Melis, D., and Jaakkola, T.S. (2018, January 2\u20138). Towards Robust Interpretability with Self-Explaining Neural Networks. Proceedings of the 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montr\u00e9al, QC, Canada."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"709","DOI":"10.1198\/106186007X237892","article-title":"Generalized Functional ANOVA Diagnostics for High-Dimensional Functions of Dependent Variables","volume":"16","author":"Hooker","year":"2007","journal-title":"J. Comput. Graph. Stat."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1189","DOI":"10.1214\/aos\/1013203451","article-title":"Greedy Function Approximation: A Gradient Boosting Machine","volume":"29","author":"Friedman","year":"2001","journal-title":"Ann. Stat."},{"key":"ref_12","first-page":"4699","article-title":"Neural Additive Models: Interpretable Machine Learning with Neural Nets","volume":"6","author":"Agarwal","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_13","unstructured":"Nori, H., Jenkins, S., Koch, P., and Caruana, R. (2019). InterpretML: A Unified Framework for Machine Learning Interpretability. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"108192","DOI":"10.1016\/j.patcog.2021.108192","article-title":"GAMI-Net: An Explainable Neural Network Based on Generalized Additive Models with Structured Interactions","volume":"120","author":"Yang","year":"2021","journal-title":"Pattern Recognit."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1009","DOI":"10.1111\/j.1467-9868.2009.00718.x","article-title":"Sparse Additive Models","volume":"71","author":"Ravikumar","year":"2009","journal-title":"J. R. Stat. Soc. Ser. B"},{"key":"ref_16","first-page":"97","article-title":"Group Sparse Additive Machine","volume":"30","author":"Chen","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1016\/j.artmed.2013.10.001","article-title":"White Box Radial Basis Function Classifiers with Component Selection for Clinical Prediction Models","volume":"60","author":"Lisboa","year":"2014","journal-title":"Artif. Intell. Med."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"259","DOI":"10.1016\/j.cpc.2009.09.018","article-title":"Variance Based Sensitivity Analysis of Model Output. Design and Estimator for the Total Sensitivity Index","volume":"181","author":"Saltelli","year":"2010","journal-title":"Comput. Phys. Commun."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"206","DOI":"10.1038\/s42256-019-0048-x","article-title":"Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead","volume":"1","author":"Rudin","year":"2019","journal-title":"Nat. Mach. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Walters, B., Ortega-Martorell, S., Olier, I., and Lisboa, P.J.G. (2022, January 18\u201323). Towards Interpretable Machine Learning for Clinical Decision Support. Proceedings of the International Joint Conference on Neural Networks, Padua, Italy.","DOI":"10.1109\/IJCNN55064.2022.9892114"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"19525","DOI":"10.1038\/s41598-022-23817-2","article-title":"Enhanced Survival Prediction Using Explainable Artificial Intelligence in Heart Transplantation","volume":"12","author":"Lisboa","year":"2022","journal-title":"Sci. Rep."},{"key":"ref_22","first-page":"4765","article-title":"A Unified Approach to Interpreting Model Predictions","volume":"30","author":"Lundberg","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Carvalho, D.V., Pereira, E.M., and Cardoso, J.S. (2019). Machine Learning Interpretability: A Survey on Methods and Metrics. Electronics, 8.","DOI":"10.3390\/electronics8080832"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1111\/j.2517-6161.1996.tb02080.x","article-title":"Regression Shrinkage and Selection Via the Lasso","volume":"58","author":"Tibshirani","year":"1996","journal-title":"J. R. Stat. Soc. Ser. B"},{"key":"ref_25","unstructured":"The MathWorks Inc. (1994). MATLAB, The MathWorks Inc."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"720","DOI":"10.1162\/neco.1992.4.5.720","article-title":"The Evidence Framework Applied to Classification Networks","volume":"4","author":"MacKay","year":"1992","journal-title":"Neural Comput."},{"key":"ref_27","unstructured":"Nabney, I. (2002). NETLAB: Algorithms for Pattern Recognitions, Springer."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"377","DOI":"10.1109\/72.839008","article-title":"Extracting Rules from Trained Neural Networks","volume":"11","author":"Tsukimoto","year":"2000","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Ripley, B.D. (1996). Pattern Recognition and Neural Networks, Cambridge University Press.","DOI":"10.1017\/CBO9780511812651"},{"key":"ref_30","unstructured":"Smith, J.W., Everhart, J.E., Dickson, W.C., Knowler, W.C., and Johannes, R.S. (1988, January 6\u20139). Using the ADAP Learning Algorithm to Forecast the Onset of Diabetes Mellitus. Proceedings of the Annual Symposium on Computer Application in Medical Care, Washington, DC, USA."},{"key":"ref_31","unstructured":"Newman, D.J., Hettich, S., Blake, C.L., and Merz, C.J. (2022, January 01). UCI Repository of Machine Learning Databases. Available online: http:\/\/www.ics.uci.edu\/~mlearn\/MLRepository.html."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Abe, N., Zadrozny, B., and Langford, J. (2006, January 20\u201323). Outlier Detection by Active Learning. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Philadelphia, PA, USA.","DOI":"10.1145\/1150402.1150459"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"e173","DOI":"10.1016\/S1470-2045(14)71116-7","article-title":"Nomograms in Oncology: More than Meets the Eye","volume":"16","author":"Balachandran","year":"2015","journal-title":"Lancet Oncol."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s12911-021-01569-9","article-title":"Explaining Multivariate Molecular Diagnostic Tests via Shapley Values","volume":"21","author":"Roder","year":"2021","journal-title":"BMC Med. Inform. Decis. Mak."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"1169","DOI":"10.1002\/(SICI)1097-0258(19980530)17:10<1169::AID-SIM796>3.0.CO;2-D","article-title":"Feed Forward Neural Networks for the Analysis of Censored Survival Data: A Partial Logistic Regression Approach","volume":"17","author":"Biganzoli","year":"1998","journal-title":"Stat. Med."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/16\/4\/181\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:04:03Z","timestamp":1760123043000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/16\/4\/181"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,27]]},"references-count":35,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2023,4]]}},"alternative-id":["a16040181"],"URL":"https:\/\/doi.org\/10.3390\/a16040181","relation":{},"ISSN":["1999-4893"],"issn-type":[{"type":"electronic","value":"1999-4893"}],"subject":[],"published":{"date-parts":[[2023,3,27]]}}}