{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T14:13:01Z","timestamp":1740147181026,"version":"3.37.3"},"reference-count":50,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2021,9,16]],"date-time":"2021-09-16T00:00:00Z","timestamp":1631750400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,9,16]],"date-time":"2021-09-16T00:00:00Z","timestamp":1631750400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100003246","name":"Nederlandse Organisatie voor Wetenschappelijk Onderzoek","doi-asserted-by":"publisher","award":["022.005.022"],"award-info":[{"award-number":["022.005.022"]}],"id":[{"id":"10.13039\/501100003246","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Adv Data Anal Classif"],"published-print":{"date-parts":[[2022,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The<jats:inline-formula><jats:alternatives><jats:tex-math>$$\\delta $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:mi>\u03b4<\/mml:mi><\/mml:math><\/jats:alternatives><\/jats:inline-formula>-machine is a statistical learning tool for classification based on dissimilarities or distances between profiles of the observations to profiles of a representation set, which was proposed by Yuan et al. (J Claasif 36(3): 442\u2013470, 2019). So far, the<jats:inline-formula><jats:alternatives><jats:tex-math>$$\\delta $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:mi>\u03b4<\/mml:mi><\/mml:math><\/jats:alternatives><\/jats:inline-formula>-machine was restricted to continuous predictor variables only. In this article, we extend the<jats:inline-formula><jats:alternatives><jats:tex-math>$$\\delta $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:mi>\u03b4<\/mml:mi><\/mml:math><\/jats:alternatives><\/jats:inline-formula>-machine to handle continuous, ordinal, nominal, and binary predictor variables. We utilized a tailored dissimilarity function for mixed type variables which was defined by Gower. This measure has properties of a Manhattan distance. We develop, in a similar vein, a Euclidean dissimilarity function for mixed type variables. In simulation studies we compare the performance of the two dissimilarity functions and we compare the predictive performance of the<jats:inline-formula><jats:alternatives><jats:tex-math>$$\\delta $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:mi>\u03b4<\/mml:mi><\/mml:math><\/jats:alternatives><\/jats:inline-formula>-machine to logistic regression models. We generated data according to two population distributions where the type of predictor variables, the distribution of categorical variables, and the number of predictor variables was varied. The performance of the<jats:inline-formula><jats:alternatives><jats:tex-math>$$\\delta $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:mi>\u03b4<\/mml:mi><\/mml:math><\/jats:alternatives><\/jats:inline-formula>-machine using the two dissimilarity functions and different types of representation set was investigated. The simulation studies showed that the adjusted Euclidean dissimilarity function performed better than the adjusted Gower dissimilarity function; that the<jats:inline-formula><jats:alternatives><jats:tex-math>$$\\delta $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:mi>\u03b4<\/mml:mi><\/mml:math><\/jats:alternatives><\/jats:inline-formula>-machine outperformed logistic regression; and that for constructing the representation set,<jats:italic>K<\/jats:italic>-medoids clustering achieved fewer active exemplars than the one using<jats:italic>K<\/jats:italic>-means clustering while maintaining the accuracy. We also applied the<jats:inline-formula><jats:alternatives><jats:tex-math>$$\\delta $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:mi>\u03b4<\/mml:mi><\/mml:math><\/jats:alternatives><\/jats:inline-formula>-machine to an empirical example, discussed its interpretation in detail, and compared the classification performance with five other classification methods. The results showed that the<jats:inline-formula><jats:alternatives><jats:tex-math>$$\\delta $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:mi>\u03b4<\/mml:mi><\/mml:math><\/jats:alternatives><\/jats:inline-formula>-machine has a good balance between accuracy and interpretability.<\/jats:p>","DOI":"10.1007\/s11634-021-00463-6","type":"journal-article","created":{"date-parts":[[2021,9,16]],"date-time":"2021-09-16T12:02:48Z","timestamp":1631793768000},"page":"875-907","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["A comparison of two dissimilarity functions for mixed-type predictor variables in the $$\\delta $$-machine"],"prefix":"10.1007","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8618-9366","authenticated-orcid":false,"given":"Beibei","family":"Yuan","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Willem","family":"Heiser","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mark","family":"de Rooij","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,9,16]]},"reference":[{"issue":"02","key":"463_CR1","doi-asserted-by":"publisher","first-page":"291","DOI":"10.1017\/S095457949700206X","volume":"9","author":"LR Bergman","year":"1997","unstructured":"Bergman LR, Magnusson D (1997) A person-oriented approach in research on developmental psychopathology. Dev Psychopathol 9(02):291\u2013319","journal-title":"Dev Psychopathol"},{"key":"463_CR2","volume-title":"Statistical learning from a regression perspective","author":"RA Berk","year":"2008","unstructured":"Berk RA (2008) Statistical learning from a regression perspective, 1st edn. Springer, New York","edition":"1"},{"key":"463_CR3","volume-title":"Modern multidimensional scaling: theory and applications","author":"I Borg","year":"2005","unstructured":"Borg I, Groenen PJ (2005) Modern multidimensional scaling: theory and applications. Springer Science & Business Media, Berlin"},{"issue":"2","key":"463_CR4","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1007\/BF00058655","volume":"24","author":"L Breiman","year":"1996","unstructured":"Breiman L (1996) Bagging predictors. Mach Learn 24(2):123\u2013140","journal-title":"Mach Learn"},{"issue":"1","key":"463_CR5","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman L (2001) Random forests. Mach Learn 45(1):5\u201332","journal-title":"Mach Learn"},{"key":"463_CR6","unstructured":"Breiman L, Friedman J, Olshen R, Stone C (1984) Classification and regression trees. Chapman and Hall\/CRC"},{"key":"463_CR7","unstructured":"Brown G (2004) Diversity in neural network ensembles. PhD thesis, University of Birmingham"},{"issue":"1","key":"463_CR8","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1177\/001316447303300111","volume":"33","author":"J Cohen","year":"1973","unstructured":"Cohen J (1973) Eta-squared and partial eta-squared in fixed factor anova designs. Educ Psychol Measure 33(1):107\u2013112","journal-title":"Educ Psychol Measure"},{"key":"463_CR9","volume-title":"Statistical power analysis for the behavioral sciences","author":"J Cohen","year":"1988","unstructured":"Cohen J (1988) Statistical power analysis for the behavioral sciences, 2nd edn. Academic press, New York","edition":"2"},{"issue":"3","key":"463_CR10","doi-asserted-by":"publisher","first-page":"273","DOI":"10.1007\/BF00994018","volume":"20","author":"C Cortes","year":"1995","unstructured":"Cortes C, Vapnik V (1995) Support-vector networks. Mach Learn 20(3):273\u2013297","journal-title":"Mach Learn"},{"key":"463_CR11","doi-asserted-by":"crossref","unstructured":"Cox TF, Cox M (2000) Multidimensional scaling. CRC Press","DOI":"10.1201\/9781420036121"},{"issue":"3","key":"463_CR12","doi-asserted-by":"publisher","first-page":"541","DOI":"10.1161\/01.CIR.69.3.541","volume":"69","author":"R Detrano","year":"1984","unstructured":"Detrano R, Yiannikas J, Salcedo EE, Rincon G, Go RT, Williams G, Leatherman J (1984) Bayesian probability analysis: a prospective demonstration of its clinical utility in diagnosing coronary disease. Circulation 69(3):541\u2013547","journal-title":"Circulation"},{"issue":"5","key":"463_CR13","doi-asserted-by":"publisher","first-page":"304","DOI":"10.1016\/0002-9149(89)90524-9","volume":"64","author":"R Detrano","year":"1989","unstructured":"Detrano R, Janosi A, Steinbrunn W, Pfisterer M, Schmid JJ, Sandhu S, Guppy KH, Lee S, Froelicher V (1989) International application of a new probability algorithm for the diagnosis of coronary artery disease. Am J Cardiol 64(5):304\u2013310","journal-title":"Am J Cardiol"},{"key":"463_CR14","unstructured":"Dheeru D, Karra\u00a0Taniskidou E (2017) UCI machine learning repository. http:\/\/archive.ics.uci.edu\/ml. University of California, Irvine, School of Information and Computer Sciences"},{"issue":"8","key":"463_CR15","doi-asserted-by":"publisher","first-page":"861","DOI":"10.1016\/j.patrec.2005.10.010","volume":"27","author":"T Fawcett","year":"2006","unstructured":"Fawcett T (2006) An introduction to roc analysis. Patt Recogn Lett 27(8):861\u2013874","journal-title":"Patt Recogn Lett"},{"key":"463_CR16","first-page":"148","volume-title":"Machine learning: proceedings of the thirteenth international conference","author":"Y Freund","year":"1996","unstructured":"Freund Y, Schapire RE et al (1996) Experiments with a new boosting algorithm. In: Saitta L (ed) Machine learning: proceedings of the thirteenth international conference, vol 96. Morgan Kauffman Inc., San Francisco, pp 148\u2013156"},{"issue":"5","key":"463_CR17","doi-asserted-by":"publisher","first-page":"1189","DOI":"10.1214\/aos\/1013203451","volume":"29","author":"J Friedman","year":"2001","unstructured":"Friedman J (2001) Greedy function approximation: a gradient boosting machine. Ann Stat 29(5):1189\u20131232","journal-title":"Ann Stat"},{"key":"463_CR18","volume-title":"The elements of statistical learning","author":"J Friedman","year":"2009","unstructured":"Friedman J, Hastie T, Tibshirani R (2009) The elements of statistical learning, 2nd edn. Springer series, New York (in Statistics)","edition":"2"},{"key":"463_CR19","doi-asserted-by":"crossref","unstructured":"Friedman J, Hastie T, Tibshirani R (2010a) glmnet: regularization paths for generalized linear models via coordinate descent. R package version 1.6-4","DOI":"10.18637\/jss.v033.i01"},{"issue":"1","key":"463_CR20","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18637\/jss.v033.i01","volume":"33","author":"J Friedman","year":"2010","unstructured":"Friedman J, Hastie T, Tibshirani R (2010) Regularization paths for generalized linear models via coordinate descent. J Stat Softw 33(1):1\u201322","journal-title":"J Stat Softw"},{"issue":"4","key":"463_CR21","doi-asserted-by":"publisher","first-page":"857","DOI":"10.2307\/2528823","volume":"27","author":"JC Gower","year":"1971","unstructured":"Gower JC (1971) A general coefficient of similarity and some of its properties. Biometrics 27(4):857\u2013871","journal-title":"Biometrics"},{"key":"463_CR22","unstructured":"Heiser WJ (1981) Unfolding analysis of proximity data. Ph.D. dissertation, Department of Data Theory, Leiden University"},{"key":"463_CR23","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1016\/j.lindif.2017.11.001","volume":"66","author":"M Hickendorff","year":"2018","unstructured":"Hickendorff M, Edelsbrunner PA, McMullen J, Schneider M, Trezise K (2018) Informative tools for characterizing individual differences in learning: latent class, latent profile, and latent transition analysis. Learn Indiv Diff 66:4\u201315","journal-title":"Learn Indiv Diff"},{"key":"463_CR24","unstructured":"Hsu CW, Chang CC, Lin CJ et\u00a0al (2003) A practical guide to support vector classification. Technical Report, Department of Computer Science, National Taiwan University"},{"key":"463_CR25","unstructured":"Huang Z (1997) Clustering large data sets with mixed numeric and categorical values. In: Motoda H (ed) Proceedings of the 1st Pacific-Asia conference on knowledge discovery and data mining,(PAKDD), Singapore. World Scientific Publishing Co., Inc., pp 21\u201334"},{"issue":"3","key":"463_CR26","doi-asserted-by":"publisher","first-page":"283","DOI":"10.1023\/A:1009769707641","volume":"2","author":"Z Huang","year":"1998","unstructured":"Huang Z (1998) Extensions to the k-means algorithm for clustering large data sets with categorical values. Data Mining Knowl Discov 2(3):283\u2013304","journal-title":"Data Mining Knowl Discov"},{"issue":"2","key":"463_CR27","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1111\/j.1469-8137.1912.tb05611.x","volume":"11","author":"P Jaccard","year":"1912","unstructured":"Jaccard P (1912) The distribution of the flora in the alpine zone. 1. New Phytol 11(2):37\u201350","journal-title":"New Phytol"},{"key":"463_CR28","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4614-7138-7","volume-title":"An introduction to statistical learning","author":"G James","year":"2013","unstructured":"James G, Witten D, Hastie T, Tibshirani R (2013) An introduction to statistical learning. Springer, New York"},{"key":"463_CR29","doi-asserted-by":"publisher","DOI":"10.1002\/9780470316801","volume-title":"Finding groups in data: an introduction to cluster analysis","author":"L Kaufman","year":"1990","unstructured":"Kaufman L, Rousseeuw PJ (1990) Finding groups in data: an introduction to cluster analysis. John Wiley & Sons, New York"},{"issue":"4","key":"463_CR30","first-page":"324","volume":"1","author":"S Kotsiantis","year":"2004","unstructured":"Kotsiantis S, Pintelas P (2004) Combining bagging and boosting. Int J Comput Intell 1(4):324\u2013333","journal-title":"Int J Comput Intell"},{"key":"463_CR31","first-page":"281","volume-title":"Proceedings of the fifth Berkeley symposium on mathematical statistics and probability","author":"J MacQueen","year":"1967","unstructured":"MacQueen J (1967) Some methods for classification and analysis of multivariate observations. In: LeCam LM, Neyman J (eds) Proceedings of the fifth Berkeley symposium on mathematical statistics and probability, vol 1. Oakland, CA, USA, pp 281\u2013297"},{"issue":"3","key":"463_CR32","doi-asserted-by":"publisher","first-page":"207","DOI":"10.1037\/0033-295X.85.3.207","volume":"85","author":"DL Medin","year":"1978","unstructured":"Medin DL, Schaffer MM (1978) Context theory of classification learning. Psychol Rev 85(3):207\u2013238","journal-title":"Psychol Rev"},{"key":"463_CR33","unstructured":"Melville P, Mooney RJ (2003) Constructing diverse classifier ensembles using artificial training examples. In: Gottlob G, Walsh T (eds) Proceedings of the eighteenth international joint conference on artificial intelligence, vol\u00a03, pp 505\u2013510"},{"issue":"3","key":"463_CR34","doi-asserted-by":"publisher","first-page":"361","DOI":"10.1214\/19-STS697","volume":"34","author":"JJ Meulman","year":"2019","unstructured":"Meulman JJ, van der Kooij AJ, Duisters KL et al (2019) Ros regression: integrating regularization with optimal scaling regression. Stat Sci 34(3):361\u2013390","journal-title":"Stat Sci"},{"key":"463_CR35","unstructured":"Meyer D, Dimitriadou E, Hornik K, Weingessel A, Leisch F (2014) e1071: misc functions of the department of statistics (e1071), TU Wien. R package version 1.6-4. http:\/\/CRAN.R-project.org\/package=e1071"},{"issue":"1","key":"463_CR36","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1023\/A:1007567018844","volume":"35","author":"B Mirkin","year":"1999","unstructured":"Mirkin B (1999) Concept learning and feature selection based on square-error clustering. Mach Learn 35(1):25\u201339","journal-title":"Mach Learn"},{"key":"463_CR37","unstructured":"Nosofsky R (1992) Exemplars, prototypes, and similarity rules. In: Healy AF, Estes WK, Kosslyn SM, Shiffrin RM (eds) From learning theory to connectionist theory, vol 1. Lawrence Erlbaum Associates Inc, pp 49\u2013167"},{"issue":"4","key":"463_CR38","doi-asserted-by":"publisher","first-page":"899","DOI":"10.3758\/BRM.42.4.899","volume":"42","author":"K Okada","year":"2010","unstructured":"Okada K, Shigemasu K (2010) Bayesian multidimensional scaling for the estimation of a minkowski exponent. Behav Res Methods 42(4):899\u2013905","journal-title":"Behav Res Methods"},{"key":"463_CR39","doi-asserted-by":"publisher","first-page":"169","DOI":"10.1613\/jair.614","volume":"11","author":"D Opitz","year":"1999","unstructured":"Opitz D, Maclin R (1999) Popular ensemble methods: an empirical study. J Artif Intell Res 11:169\u2013198","journal-title":"J Artif Intell Res"},{"key":"463_CR40","doi-asserted-by":"publisher","DOI":"10.1142\/5965","volume-title":"The dissimilarity representation for pattern recognition: foundations and applications","author":"E Pekalska","year":"2005","unstructured":"Pekalska E, Duin RP (2005) The dissimilarity representation for pattern recognition: foundations and applications. World Scientific, Singapore"},{"issue":"Dec","key":"463_CR41","first-page":"175","volume":"2","author":"E Pekalska","year":"2001","unstructured":"Pekalska E, Paclik P, Duin RP (2001) A generalized kernel approach to dissimilarity-based classification. J Mach Learn Res 2(Dec):175\u2013211","journal-title":"J Mach Learn Res"},{"key":"463_CR42","unstructured":"R Core Team (2015) R: a language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. http:\/\/www.R-project.org\/"},{"key":"463_CR43","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511812651","volume-title":"Pattern recognition and neural networks","author":"BD Ripley","year":"1996","unstructured":"Ripley BD (1996) Pattern recognition and neural networks. Cambridge University Press, New York"},{"key":"463_CR44","first-page":"205","volume-title":"The nature of cognition","author":"BH Ross","year":"1999","unstructured":"Ross BH, Makin VS (1999) Prototype versus exemplar models in cognition. In: Sternberg RJ (ed) The nature of cognition. MIT Press Cambridge, MA, pp 205\u2013241"},{"key":"463_CR45","doi-asserted-by":"crossref","unstructured":"\u015eahan S, Polat K, Kodaz H, G\u00fcne\u015f S (2005) The medical applications of attribute weighted artificial immune system (awais): diagnosis of heart and diabetes diseases. In: International conference on artificial immune systems, Springer, pp 456\u2013468","DOI":"10.1007\/11536444_35"},{"issue":"3","key":"463_CR46","doi-asserted-by":"publisher","first-page":"285","DOI":"10.1037\/a0023346","volume":"16","author":"D Steinley","year":"2011","unstructured":"Steinley D, Brusco MJ (2011) Choosing the number of clusters in k-means clustering. Psychol Methods 16(3):285","journal-title":"Psychol Methods"},{"key":"463_CR47","unstructured":"Therneau T, Atkinson B, Ripley B (2015) rpart: recursive partitioning and regression trees. https:\/\/CRAN.R-project.org\/package=rpart, r package version 4.1-10"},{"issue":"1","key":"463_CR48","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1111\/j.2517-6161.1996.tb02080.x","volume":"58","author":"R Tibshirani","year":"1996","unstructured":"Tibshirani R (1996) Regression shrinkage and selection via the lasso. J Roy Stat Soc Ser B (Methodological) 58(1):267\u2013288","journal-title":"J Roy Stat Soc Ser B (Methodological)"},{"issue":"2","key":"463_CR49","doi-asserted-by":"publisher","first-page":"159","DOI":"10.1023\/A:1007659514849","volume":"40","author":"GI Webb","year":"2000","unstructured":"Webb GI (2000) Multiboosting: A technique for combining boosting and wagging. Mach Learn 40(2):159\u2013196","journal-title":"Mach Learn"},{"issue":"3","key":"463_CR50","doi-asserted-by":"publisher","first-page":"442","DOI":"10.1007\/s00357-019-09338-0","volume":"36","author":"B Yuan","year":"2019","unstructured":"Yuan B, Heiser W, de Rooij M (2019) The $$\\delta $$-machine: classification based on distances towards prototypes. J Classif 36(3):442\u2013470","journal-title":"J Classif"}],"container-title":["Advances in Data Analysis and Classification"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11634-021-00463-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11634-021-00463-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11634-021-00463-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,8]],"date-time":"2024-09-08T13:12:30Z","timestamp":1725801150000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11634-021-00463-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,16]]},"references-count":50,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2022,12]]}},"alternative-id":["463"],"URL":"https:\/\/doi.org\/10.1007\/s11634-021-00463-6","relation":{},"ISSN":["1862-5347","1862-5355"],"issn-type":[{"type":"print","value":"1862-5347"},{"type":"electronic","value":"1862-5355"}],"subject":[],"published":{"date-parts":[[2021,9,16]]},"assertion":[{"value":"12 July 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 August 2021","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 September 2021","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 September 2021","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}