{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,15]],"date-time":"2025-08-15T01:03:54Z","timestamp":1755219834596,"version":"3.43.0"},"reference-count":0,"publisher":"IOS Press","isbn-type":[{"type":"electronic","value":"9781643686080"}],"license":[{"start":{"date-parts":[[2025,8,7]],"date-time":"2025-08-07T00:00:00Z","timestamp":1754524800000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,8,7]]},"abstract":"<jats:p>The performance of prediction algorithms is typically measured using four metrics: sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). These metrics are usually calculated on samples drawn from patient populations. However, the performance metrics computed over a deliberately biased sample would not directly extend to its source population. Further, it is often necessary to infer the metric values for a population different from where the sample was drawn. In this paper, we illustrate methods to solve both challenges. Specifically, given the underlying patient distribution, we show corrections to the formula for these metrics based on two common inverse probability weighting methods: standard cell weighting and logistic regression weighting. We conduct simulation experiments to identify patients living with dementia and compare these methods in performance corrections with different sample sizes for different prevalence settings. We empirically show that weighting methods can correct the estimated values for algorithms\u2019 performance. Standard cell weighting is preferred over logistic regression weighting when the sample size is small and only the strata information is available in the populations of interest.<\/jats:p>","DOI":"10.3233\/shti250968","type":"book-chapter","created":{"date-parts":[[2025,8,7]],"date-time":"2025-08-07T11:36:43Z","timestamp":1754566603000},"source":"Crossref","is-referenced-by-count":0,"title":["Correcting Performance Metrics Bias During Generalization from Biased Samples to Populations"],"prefix":"10.3233","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4753-9057","authenticated-orcid":false,"given":"Peijin","family":"Han","sequence":"first","affiliation":[{"name":"University of Michigan, Ann Arbor, MI, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7615-4415","authenticated-orcid":false,"given":"Guanghao","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Michigan, Ann Arbor, MI, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3122-1936","authenticated-orcid":false,"given":"V.G. Vinod","family":"Vydiswaran","sequence":"additional","affiliation":[{"name":"University of Michigan, Ann Arbor, MI, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"7437","container-title":["Studies in Health Technology and Informatics","MEDINFO 2025 \u2014 Healthcare Smart \u00d7 Medicine Deep"],"original-title":[],"link":[{"URL":"https:\/\/ebooks.iospress.nl\/pdf\/doi\/10.3233\/SHTI250968","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,7]],"date-time":"2025-08-07T11:36:43Z","timestamp":1754566603000},"score":1,"resource":{"primary":{"URL":"https:\/\/ebooks.iospress.nl\/doi\/10.3233\/SHTI250968"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,7]]},"ISBN":["9781643686080"],"references-count":0,"URL":"https:\/\/doi.org\/10.3233\/shti250968","relation":{},"ISSN":["0926-9630","1879-8365"],"issn-type":[{"type":"print","value":"0926-9630"},{"type":"electronic","value":"1879-8365"}],"subject":[],"published":{"date-parts":[[2025,8,7]]}}}