{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,15]],"date-time":"2025-11-15T01:49:48Z","timestamp":1763171388129,"version":"3.45.0"},"reference-count":30,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2025,4,12]],"date-time":"2025-04-12T00:00:00Z","timestamp":1744416000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,4,12]],"date-time":"2025-04-12T00:00:00Z","timestamp":1744416000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Classif"],"published-print":{"date-parts":[[2025,11]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Evaluation metrics provide a means for quantifying and comparing performances of supervised learning models, but drawing meaningful conclusions from acquired scores requires a contextual framework. Our paper addresses this by introducing the Dutch scaler (DS), a novel performance indicator for binary classification models. It quantifies a model\u2019s learning by contextualizing empirical metric scores with a baseline (Dutch draw) and a new instrument (Dutch oracle) representing the prediction quality of an \u201coptimal\u201d classifier. The DS performance indicator expresses the relative contribution of these components to obtain a model\u2019s score, specifying the actual learning quality. We derived closed-form expressions to map metric scores to DS scores for common evaluation metrics and categorized them by their functional form and second derivative. The DS enhances the assessment of classifiers and facilitates a framework to compare prediction quality differences between models with varying metric scores.<\/jats:p>","DOI":"10.1007\/s00357-025-09510-9","type":"journal-article","created":{"date-parts":[[2025,4,12]],"date-time":"2025-04-12T05:52:31Z","timestamp":1744437151000},"page":"639-659","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["The Dutch Scaler Performance Indicator: How Much Did My Model Actually Learn?"],"prefix":"10.1007","volume":"42","author":[{"given":"Etienne Pieter","family":"van de Bijl","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jan Gerard","family":"Klein","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joris","family":"Pries","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sandjai","family":"Bhulai","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Robert Douwe","family":"van der Mei","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,4,12]]},"reference":[{"key":"9510_CR1","unstructured":"Adams, R. A., & Essex, C. (2010). Calculus: A complete course, 7th ed edn. Pearson Canada, Toronto. OCLC: 749131682."},{"key":"9510_CR2","doi-asserted-by":"publisher","unstructured":"Antos, A., Devroye, L., & Gyorfi, L. (1999). Lower bounds for Bayes error estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence,21(7), 643\u2013645. https:\/\/doi.org\/10.1109\/34.777375. Accessed 02 Sept 2024.","DOI":"10.1109\/34.777375"},{"key":"9510_CR3","doi-asserted-by":"publisher","unstructured":"Brzezinski, D., Stefanowski, J., Susmaga, R., & Szczech, I. (2018). Visual-based analysis of classification measures and their properties for class imbalanced problems. Information Sciences,462, 242\u2013261. https:\/\/doi.org\/10.1016\/j.ins.2018.06.020. Accessed 04 May 2023.","DOI":"10.1016\/j.ins.2018.06.020"},{"key":"9510_CR4","doi-asserted-by":"publisher","unstructured":"Canbek, G., Sagiroglu, S., Temizel, T. T., & Baykal, N. (2017). Binary classification performance measures\/metrics: A comprehensive visualized roadmap to gain new insights. In 2017 International conference on computer science and engineering (UBMK) (pp. 821\u2013826). IEEE, Antalya. https:\/\/doi.org\/10.1109\/UBMK.2017.8093539. http:\/\/ieeexplore.ieee.org\/document\/8093539\/. Accessed 2024-04-05.","DOI":"10.1109\/UBMK.2017.8093539"},{"key":"9510_CR5","doi-asserted-by":"publisher","unstructured":"Canbek, G., Taskaya-Temizel, T., & Sagiroglu, S. (2020). Binary-classificationperformance evaluation reporting survey data with the findings. Mendeley. https:\/\/doi.org\/10.17632\/5C442VBJZG.3. https:\/\/data.mendeley.com\/datasets\/5c442vbjzg\/3. Accessed 29 Jan 2024.","DOI":"10.17632\/5C442VBJZG.3"},{"key":"9510_CR6","doi-asserted-by":"publisher","unstructured":"Canbek, G., Taskaya-Temizel, T., & Sagiroglu, S. (2021). BenchMetrics: A systematic benchmarking method for binary classification performance metrics. Neural Computing and Applications,33(21), 14623\u201314650. https:\/\/doi.org\/10.1007\/s00521-021-06103-6. Accessed 04 May 2023.","DOI":"10.1007\/s00521-021-06103-6"},{"key":"9510_CR7","doi-asserted-by":"publisher","unstructured":"Canbek, G., Temizel-Taskaya, T., & Sagiroglu, S. (2022a). Accuracy barrier (ACCBAR): A novel performance indicator for binary classification. In 2022 15th International conference on information security and cryptography (ISCTURKEY) (pp. 92\u201397). https:\/\/doi.org\/10.1109\/ISCTURKEY56345.2022.9931888","DOI":"10.1109\/ISCTURKEY56345.2022.9931888"},{"key":"9510_CR8","doi-asserted-by":"publisher","unstructured":"Canbek, G., Taskaya-Temizel, T., & Sagiroglu, S. (2022). PToPI: A comprehensive review, analysis, and knowledge representation of binary classification performance measures\/metrics. SN Computer Science,4(1), 13. https:\/\/doi.org\/10.1007\/s42979-022-01409-1. Accessed 04 May 2023.","DOI":"10.1007\/s42979-022-01409-1"},{"key":"9510_CR9","doi-asserted-by":"publisher","unstructured":"Chicco, D., & Jurman, G. (2020). The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics,21(1), 6. https:\/\/doi.org\/10.1186\/s12864-019-6413-7. Accessed 11 May 2023.","DOI":"10.1186\/s12864-019-6413-7"},{"key":"9510_CR10","doi-asserted-by":"publisher","unstructured":"Delgado, R., & Tibau, X.-A. (2019). Why Cohen\u2019s Kappa should be avoided as performance measure in classification. PLOS ONE,14(9), 0222916. https:\/\/doi.org\/10.1371\/journal.pone.0222916. Publisher: Public Library of Science. Accessed 12 May 2023.","DOI":"10.1371\/journal.pone.0222916"},{"key":"9510_CR11","doi-asserted-by":"publisher","unstructured":"Ferri, C., Hern\u00e1ndez-Orallo, J., & Modroiu, R. (2009). An experimental comparison of performance measures for classification. Pattern Recognition Letters,30(1), 27\u201338. https:\/\/doi.org\/10.1016\/j.patrec.2008.08.010. Accessed 04 May 2023.","DOI":"10.1016\/j.patrec.2008.08.010"},{"key":"9510_CR12","unstructured":"G\u00f6sgens, M., Zhiyanov, A., Tikhonov, A., & Prokhorenkova, L. (2021). Good classification measures and how to find them. In Advances in neural information processing systems, (vol. 21, pp. 17136\u201317147). Type: Conference paper. https:\/\/www.scopus.com\/inward\/record.uri?eid=2-s2.0-85129054764&partnerID=40 &md5=755934440b2394000bb3e6544c930dd7"},{"key":"9510_CR13","doi-asserted-by":"publisher","unstructured":"Gupta, N., Mujumdar, S., Patel, H., Masuda, S., Panwar, N., Bandyopadhyay, S., Mehta, S., Guttula, S., Afzal, S., Sharma\u00a0Mittal, R., & Munigala, V. (2021). Data quality for machine learning tasks. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, (pp. 4040\u20134041). ACM, Virtual Event Singapore. https:\/\/doi.org\/10.1145\/3447548.3470817. https:\/\/dl.acm.org\/doi\/10.1145\/3447548.3470817. Accessed 06 Jun 2023.","DOI":"10.1145\/3447548.3470817"},{"key":"9510_CR14","doi-asserted-by":"publisher","unstructured":"Hossin, M., & Sulaiman, M. N. (2015). A review on evaluation metrics for data classification evaluations. International Journal of Data Mining & Knowledge Management Process,5(2), 01\u201311. https:\/\/doi.org\/10.5121\/ijdkp.2015.5201. Accessed 16 May 2023.","DOI":"10.5121\/ijdkp.2015.5201"},{"key":"9510_CR15","unstructured":"Ishida, T., Yamane, I., Charoenphakdee, N., Niu, G., & Sugiyama, M. (2023). Is the performance of my deep network too good to be true? A direct approach to estimating the bayes error in binary classification. In The eleventh international conference on learning representations. https:\/\/openreview.net\/forum?id=FZdJQgy05rz"},{"key":"9510_CR16","doi-asserted-by":"publisher","unstructured":"Jain, A., Patel, H., Nagalapatti, L., Gupta, N., Mehta, S., Guttula, S., Mujumdar, S., Afzal, S., Sharma\u00a0Mittal, R., & Munigala, V. (2020). Overview and importance of data quality for machine learning tasks. In Proceedings of the 26th ACM IGKDD International Conference on Knowledge Discovery & Data Mining (pp. 3561\u20133562). ACM, Virtual Event CA USA. https:\/\/doi.org\/10.1145\/3394486.3406477. https:\/\/dl.acm.org\/doi\/10.1145\/3394486.3406477. Accessed 06 Jun 2023.","DOI":"10.1145\/3394486.3406477"},{"key":"9510_CR17","doi-asserted-by":"publisher","unstructured":"Kuncheva, L. I., Whitaker, C. J., Shipp, C. A., & Duin, R. P. W. (2003). Limits on the majority vote accuracy in classifier fusion. Pattern Analysis & Applications,6(1), 22\u201331. https:\/\/doi.org\/10.1007\/s10044-002-0173-7. Accessed 06 Jun 2023.","DOI":"10.1007\/s10044-002-0173-7"},{"key":"9510_CR18","doi-asserted-by":"publisher","unstructured":"Luque, A., Carrasco, A., Mart\u00edn, A., & De Las Heras, A. (2019). The impact of class imbalance in classification performance metrics based on the binary confusion matrix. Pattern Recognition,91, 216\u2013231. https:\/\/doi.org\/10.1016\/j.patcog.2019.02.023. Accessed 29 Jan 2024.","DOI":"10.1016\/j.patcog.2019.02.023"},{"key":"9510_CR19","volume-title":"Machine learning","author":"TM Mitchell","year":"1997","unstructured":"Mitchell, T. M. (1997). Machine learning. McGraw-Hill, New York: McGraw-Hill series in computer science."},{"key":"9510_CR20","unstructured":"Mohri, M., Rostamizadeh, A., & Talwalkar, A. (2018). Foundations of machine learning, second (edition). Adaptive computation and machine learning: The MIT Press, Cambridge, Massachusetts."},{"key":"9510_CR21","unstructured":"Powers, D. M. W. (2020). Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation. arXiv. arXiv:2010.16061 [cs, stat]. http:\/\/arxiv.org\/abs\/2010.16061. Accessed 04 May 2023."},{"key":"9510_CR22","doi-asserted-by":"publisher","unstructured":"Pries, J., van de Bijl, E. P., Klein, J. G., Bhulai, S., & van der Mei, R. D. (2023). The optimal input-independent baseline for binary classification: The Dutch Draw. Statistica Neerlandica,12297. https:\/\/doi.org\/10.1111\/stan.12297. Accessed 03 May 2023.","DOI":"10.1111\/stan.12297"},{"key":"9510_CR23","doi-asserted-by":"publisher","unstructured":"Raykar, V. C., Yu, S., Zhao, L. H., Jerebko, A., Florin, C., Valadez, G. H., Bogoni, L., & Moy, L. (2009). Supervised learning from multiple experts: Whom to trust when everyone lies a bit. In Proceedings of the 26th Annual international conference on machine learning (pp. 889\u2013896). ACM, Montreal Quebec Canada. https:\/\/doi.org\/10.1145\/1553374.1553488. https:\/\/dl.acm.org\/doi\/10.1145\/1553374.1553488. Accessed 04 May 2023.","DOI":"10.1145\/1553374.1553488"},{"key":"9510_CR24","doi-asserted-by":"publisher","unstructured":"Sidey-Gibbons, J. A. M., & Sidey-Gibbons, C. J. (2019). Machine learning in medicine: A practical introduction. BMC Medical Research Methodology,19(1), 64. https:\/\/doi.org\/10.1186\/s12874-019-0681-4. Accessed 27 Aug 2024.","DOI":"10.1186\/s12874-019-0681-4"},{"key":"9510_CR25","doi-asserted-by":"publisher","unstructured":"Sokolova, M., & Lapalme, G. (2009). A systematic analysis of performance measures for classification tasks. Information Processing & Management,45(4), 427\u2013437. https:\/\/doi.org\/10.1016\/j.ipm.2009.03.002. Accessed 04 May 2023.","DOI":"10.1016\/j.ipm.2009.03.002"},{"key":"9510_CR26","doi-asserted-by":"publisher","unstructured":"Texel, P. P. (2013). Measure, metric, and indicator: An object-oriented approach for consistent terminology. In 2013 Proceedings of IEEE Southeastcon (pp. 1\u20135). IEEE, Jacksonville, FL, USA. https:\/\/doi.org\/10.1109\/SECON.2013.6567438. http:\/\/ieeexplore.ieee.org\/document\/6567438\/. Accessed 04 May 2023.","DOI":"10.1109\/SECON.2013.6567438"},{"key":"9510_CR27","unstructured":"van de Bijl, E. P. (2023). Dutch Scaler. https:\/\/github.com\/etiennevandebijl\/Dutch-Scaler"},{"key":"9510_CR28","doi-asserted-by":"publisher","unstructured":"van de Bijl, E. P., Klein, J. G., Pries, J., Bhulai, S., Hoogendoorn, M., & van der Mei, R. D. (2024). The Dutch Draw: Constructing a universal baseline for binary classification problems. Journal of Applied Probability,1\u201319,. https:\/\/doi.org\/10.1017\/jpr.2024.52. Accessed 30 Jan 2025.","DOI":"10.1017\/jpr.2024.52"},{"key":"9510_CR29","doi-asserted-by":"publisher","unstructured":"Wolberg, W., Mangasarian, O., Street, N., & Street, W. (1993). Breast Cancer Wisconsin (Diagnostic). UCI Machine Learning Repository. UCI Machine Learning Repository. https:\/\/doi.org\/10.24432\/C5DW2B. https:\/\/archive.ics.uci.edu\/dataset\/17. Accessed 27 Aug 2024.","DOI":"10.24432\/C5DW2B"},{"key":"9510_CR30","doi-asserted-by":"publisher","unstructured":"Wolpert, D. H., & Macready, W. G. (1997). No free lunch theorems for optimization. IEEE Transactions on Evolutionary Computation,1(1), 67\u201382. https:\/\/doi.org\/10.1109\/4235.585893. Accessed 06 Jun 2023.","DOI":"10.1109\/4235.585893"}],"container-title":["Journal of Classification"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00357-025-09510-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00357-025-09510-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00357-025-09510-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,15]],"date-time":"2025-11-15T01:45:18Z","timestamp":1763171118000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00357-025-09510-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,12]]},"references-count":30,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,11]]}},"alternative-id":["9510"],"URL":"https:\/\/doi.org\/10.1007\/s00357-025-09510-9","relation":{},"ISSN":["0176-4268","1432-1343"],"issn-type":[{"type":"print","value":"0176-4268"},{"type":"electronic","value":"1432-1343"}],"subject":[],"published":{"date-parts":[[2025,4,12]]},"assertion":[{"value":"13 March 2025","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 April 2025","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical Approval"}},{"value":"The authors declare no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of Interest"}}]}}