{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T13:23:33Z","timestamp":1777555413074,"version":"3.51.4"},"reference-count":24,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2024,6,1]],"date-time":"2024-06-01T00:00:00Z","timestamp":1717200000000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Data Science"],"published-print":{"date-parts":[[2024,6,26]]},"abstract":"<jats:p>Measuring data drift is essential in machine learning applications where model scoring (evaluation) is done on data samples that differ from those used in training. The Kullback-Leibler divergence is a common measure of shifted probability distributions, for which discretized versions are invented to deal with binned or categorical data. We present the Unstable Population Indicator, a robust, flexible and numerically stable, discretized implementation of Jeffrey\u2019s divergence, along with an implementation in a Python package that can deal with continuous, discrete, ordinal and nominal data in a variety of popular data types. We show the numerical and statistical properties in controlled experiments. It is not advised to employ a common cut-off to distinguish stable from unstable populations, but rather to let that cut-off depend on the use case.<\/jats:p>","DOI":"10.3233\/ds-240059","type":"journal-article","created":{"date-parts":[[2024,3,1]],"date-time":"2024-03-01T11:12:38Z","timestamp":1709291558000},"page":"1-12","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":1,"title":["Measuring data drift with the unstable population indicator"],"prefix":"10.1177","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2581-8370","authenticated-orcid":false,"given":"Marcel R.","family":"Haas","sequence":"first","affiliation":[{"name":"Public Health and Primary Care, Leiden University Medical Center, Albinusdreef 2, The Netherlands and Business Intelligence, University of Amsterdam, Spui 21, 1012WX Amsterdam, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-5030-0108","authenticated-orcid":false,"given":"Lisette","family":"Sibbald","sequence":"additional","affiliation":[{"name":"Department of Methodology and Statistics, Tilburg University, Prof. Cobbenhagenlaan 125, 5037 DB Tilburg, The Netherlands and Department of Cognitive Neuropsychology, Tilburg University, Prof. Cobbenhagenlaan 125, 5037 DB Tilburg, The Netherlands and Business Intelligence, University of Amsterdam, Spui 21, 1012WX Amsterdam, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2024,6,1]]},"reference":[{"key":"ref001","unstructured":"S.\u00a0Ackerman, E.\u00a0Farchi, O.\u00a0Raz, M.\u00a0Zalmanovici and P.\u00a0Dube, Detection of Data Drift and Outliers Affecting Machine Learning Model Performance over Time,\n                      ArXiv\n                      (2020), arXiv:2012.09258."},{"key":"ref002","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2021.09.112"},{"key":"ref003","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-018-1257-z"},{"key":"ref004","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-020-2649-2"},{"key":"ref005","doi-asserted-by":"publisher","DOI":"10.1109\/MCSE.2007.55"},{"key":"ref006","unstructured":"G.\u00a0Karakoulas, Empirical Validation of Retail Credit-Scoring Models,\n                      The RMA Journal\n                      (2004), 56\u201360. https:\/\/cms.rmau.org\/uploadedFiles\/Credit_Risk\/Library\/RMA_Journal\/Other_Topics_(1998_to_present)\/Empirical%20Validation%20of%20Retail%20Credit-Scoring%20Models.pdf."},{"key":"ref007","unstructured":"M.\u00a0Kull and P.A.\u00a0Flach, Patterns of dataset shift, in: Learning over Multiple Contexts, at ECML 2014, 2014. https:\/\/www.semanticscholar.org\/paper\/Patterns-of-dataset-shift-Kull-Flach\/aa49eb379d55fd4c923f47efcd61b2090f58e54f."},{"key":"ref008","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177729694"},{"key":"ref009","doi-asserted-by":"publisher","DOI":"10.1109\/18.61115"},{"key":"ref010","doi-asserted-by":"publisher","DOI":"10.25080\/Majora-92bf1922-00a"},{"key":"ref011","doi-asserted-by":"publisher","DOI":"10.1109\/CISDA.2015.7208643"},{"key":"ref012","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.3509134"},{"key":"ref013","unstructured":"G.L.\u00a0Poe, K.L.\u00a0Giraud and J.B.\u00a0Loomis, Computational Methods for Measuring the Difference of Empirical Distributions, Econometric Modeling: Agriculture, 2005. https:\/\/www.jstor.org\/stable\/3697850."},{"key":"ref014","unstructured":"F.M.\u00a0Polo, R.\u00a0Izbicki, E.G.\u00a0Lacerda, J.P.\u00a0Ibieta-Jimenez and R.\u00a0Vicente, A Unified Framework for Dataset Shift Diagnostics,\n                      ArXiv\n                      (2022), arXiv:2205.08340."},{"key":"ref015","doi-asserted-by":"publisher","DOI":"10.7551\/mitpress\/9780262170055.001.0001"},{"key":"ref016","doi-asserted-by":"publisher","DOI":"10.1002\/j.1538-7305.1948.tb01338.x"},{"key":"ref017","doi-asserted-by":"publisher","DOI":"10.1002\/j.1538-7305.1948.tb00917.x"},{"key":"ref018","doi-asserted-by":"publisher","DOI":"10.7275\/tbfa-x148"},{"key":"ref019","doi-asserted-by":"publisher","DOI":"10.3390\/risks7020053"},{"key":"ref020","unstructured":"G.\u00a0Van Rossum and F.L.\u00a0Drake, in: Python 3 Reference Manual, CreateSpace, Scotts Valley, CA, 2009. ISBN 1441412697."},{"key":"ref021","doi-asserted-by":"publisher","DOI":"10.21105\/joss.03021"},{"key":"ref022","doi-asserted-by":"publisher","DOI":"10.1007\/s10618-018-0554-1"},{"key":"ref023","unstructured":"B.\u00a0Yurdakul, Statistical Properties of Population Stability Index (PSI), PhD thesis, Western Michigan University, 2018. https:\/\/scholarworks.wmich.edu\/dissertations\/3208\/."},{"key":"ref024","doi-asserted-by":"publisher","DOI":"10.21314\/JRMV.2020.227"}],"container-title":["Data Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/DS-240059","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.3233\/DS-240059","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/DS-240059","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T18:10:10Z","timestamp":1777399810000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.3233\/DS-240059"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,1]]},"references-count":24,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,6,26]]}},"alternative-id":["10.3233\/DS-240059"],"URL":"https:\/\/doi.org\/10.3233\/ds-240059","relation":{},"ISSN":["2451-8484","2451-8492"],"issn-type":[{"value":"2451-8484","type":"print"},{"value":"2451-8492","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,1]]}}}