{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T10:14:18Z","timestamp":1782209658959,"version":"3.54.5"},"reference-count":21,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,8,16]],"date-time":"2021-08-16T00:00:00Z","timestamp":1629072000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,8,16]],"date-time":"2021-08-16T00:00:00Z","timestamp":1629072000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100010767","name":"Innovative Medicines Initiative","doi-asserted-by":"publisher","award":["806968"],"award-info":[{"award-number":["806968"]}],"id":[{"id":"10.13039\/501100010767","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Big Data"],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Background<\/jats:title>\n                    <jats:p>The design used to create labelled data for training prediction models from observational healthcare databases (e.g., case-control and cohort) may impact the clinical usefulness. We aim to investigate hypothetical design issues and determine how the design impacts prediction model performance.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Aim<\/jats:title>\n                    <jats:p>To empirically investigate differences between models developed using a case-control design and a cohort design.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Methods<\/jats:title>\n                    <jats:p>Using a US claims database, we replicated two published prediction models (dementia and type 2 diabetes) which were developed using a case-control design, and trained models for the same prediction questions using cohort designs. We validated each model on data mimicking the point in time the models would be applied in clinical practice. We calculated the models\u2019 discrimination and calibration-in-the-large performances.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>The dementia models obtained area under the receiver operating characteristics of 0.560 and 0.897 for the case-control and cohort designs respectively. The type 2 diabetes models obtained area under the receiver operating characteristics of 0.733 and 0.727 for the case-control and cohort designs respectively. The dementia and diabetes case-control models were both poorly calibrated, whereas the dementia cohort model achieved good calibration. We show that careful construction of a case-control design can lead to comparable discriminative performance as a cohort design, but case-control designs over-represent the outcome class leading to miscalibration.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusions<\/jats:title>\n                    <jats:p>Any case-control design can be converted to a cohort design. We recommend that researchers with observational data use the less subjective and generally better calibrated cohort design when extracting labelled data. However, if a carefully constructed case-control design is used, then the model must be prospectively validated using a cohort design for fair evaluation and be recalibrated.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1186\/s40537-021-00501-2","type":"journal-article","created":{"date-parts":[[2021,8,16]],"date-time":"2021-08-16T06:03:01Z","timestamp":1629093781000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":21,"title":["Design matters in patient-level prediction: evaluation of a cohort vs. case-control design when developing predictive models in observational healthcare datasets"],"prefix":"10.1186","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2970-0778","authenticated-orcid":false,"given":"Jenna M.","family":"Reps","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Patrick B.","family":"Ryan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peter R.","family":"Rijnbeek","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Martijn J.","family":"Schuemie","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,8,16]]},"reference":[{"issue":"1","key":"501_CR1","doi-asserted-by":"publisher","first-page":"20","DOI":"10.1186\/s12916-014-0265-4","volume":"13","author":"P Croft","year":"2015","unstructured":"Croft P, Altman DG, Deeks JJ, Dunn KM, Hay AD, Hemingway H, LeResche L, Peat G, Perel P, Petersen SE, Riley RD. The science of clinical practice: disease diagnosis or patient prognosis? Evidence about \u201cwhat is likely to happen\u201d should shape clinical practice. BMC Med. 2015;13(1):20.","journal-title":"BMC Med"},{"key":"501_CR2","doi-asserted-by":"publisher","first-page":"2416","DOI":"10.1136\/bmj.i2416","volume":"353","author":"JA Damen","year":"2016","unstructured":"Damen JA, Hooft L, Schuit E, Debray TP, Collins GS, Tzoulaki I, Lassale CM, Siontis GC, Chiocchia V, Roberts C, Schl\u00fcssel MM. Prediction models for cardiovascular disease risk in the general population: systematic review. BMJ. 2016;353:2416.","journal-title":"BMJ"},{"issue":"5","key":"501_CR3","doi-asserted-by":"publisher","first-page":"303","DOI":"10.1097\/PAP.0000000000000072","volume":"22","author":"KG Moons","year":"2015","unstructured":"Moons KG, Altman DG, Reitsma JB, Collins GS. New guideline for the reporting of studies developing, validating, or updating a multivariable clinical prediction model: the TRIPOD statement. Adv Anat Pathol. 2015;22(5):303\u20135.","journal-title":"Adv Anat Pathol"},{"issue":"1","key":"501_CR4","doi-asserted-by":"publisher","first-page":"18","DOI":"10.1186\/s13054-017-1930-8","volume":"22","author":"R Haniffa","year":"2018","unstructured":"Haniffa R, Isaam I, De Silva AP, Dondorp AM, De Keizer NF. Performance of critical care prognostic scoring systems in low and middle-income countries: a systematic review. Crit Care. 2018;22(1):18.","journal-title":"Crit Care"},{"issue":"1","key":"501_CR5","doi-asserted-by":"publisher","first-page":"3","DOI":"10.23876\/j.krcp.2017.36.1.3","volume":"36","author":"CH Lee","year":"2017","unstructured":"Lee CH, Yoon HJ. Medical big data: promise and challenges. Kidney Res Clin Pract. 2017;36(1):3.","journal-title":"Kidney Res Clin Pract"},{"key":"501_CR6","doi-asserted-by":"publisher","first-page":"21","DOI":"10.1007\/978-3-540-75171-7_2","volume-title":"Machine learning techniques for multimedia","author":"P Cunningham","year":"2008","unstructured":"Cunningham P, Cord M, Delany SJ. Supervised learning. In: Machine learning techniques for multimedia. Berlin: Springer; 2008. pp.\u00a021\u201349."},{"key":"501_CR7","first-page":"227","volume-title":"August. A DaQL to monitor data quality in machine learning applications","author":"L Ehrlinger","year":"2019","unstructured":"Ehrlinger L, Haunschmid V, Palazzini D, Lettner C. August. A DaQL to monitor data quality in machine learning applications. Cham: Springer:; 2019. p. 227\u201337."},{"issue":"8","key":"501_CR8","doi-asserted-by":"publisher","first-page":"969","DOI":"10.1093\/jamia\/ocy032","volume":"25","author":"JM Reps","year":"2018","unstructured":"Reps JM, Schuemie MJ, Suchard MA, Ryan PB, Rijnbeek PR. Design and implementation of a standardized framework to generate and evaluate patient-level prediction models using observational healthcare data. J Am Med Inform Assoc. 2018;25(8):969\u201375.","journal-title":"J Am Med Inform Assoc"},{"key":"501_CR9","doi-asserted-by":"publisher","first-page":"1138","DOI":"10.18553\/jmcp.2018.24.11.1138","volume":"24","author":"JS Albrecht","year":"2018","unstructured":"Albrecht JS, Hanna M, Kim D, Perfetto EM. Predicting diagnosis of Alzheimer\u2019s disease and related dementias using administrative claims. J Manag Care Specialty Pharm. 2018;24:1138\u201345.","journal-title":"J Manag Care Specialty Pharm."},{"issue":"5","key":"501_CR10","doi-asserted-by":"publisher","first-page":"1896","DOI":"10.1111\/1475-6773.12461","volume":"51","author":"RG McCoy","year":"2016","unstructured":"McCoy RG, Nori VS, Smith SA, Hane CA. Development and validation of HealthImpact: an incident diabetes prediction model based on administrative data. Health Serv Res. 2016;51(5):1896\u2013918.","journal-title":"Health Serv Res"},{"issue":"1","key":"501_CR11","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1016\/j.chest.2020.03.009","volume":"158","author":"T Dey","year":"2020","unstructured":"Dey T, Mukherjee A, Chakraborty S. A practical overview of case-control studies in clinical practice. Chest. 2020;158(1):57\u201364.","journal-title":"Chest"},{"key":"501_CR12","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-77244-8","volume-title":"Clinical prediction models","author":"EW Steyerberg","year":"2009","unstructured":"Steyerberg EW. Clinical prediction models. Vol.\u00a0381. New York: Springer; 2009."},{"issue":"1","key":"501_CR13","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/s41467-020-20314-w","volume":"12","author":"W Yuan","year":"2021","unstructured":"Yuan W, Beaulieu-Jones BK, Yu KH, Lipnick SL, Palmer N, Loscalzo J, Cai T, Kohane IS. Temporal bias in case-control design: preventing reliable predictions of the future. Nat Commun. 2021;12(1):1\u201310.","journal-title":"Nat Commun"},{"issue":"11","key":"501_CR14","doi-asserted-by":"publisher","first-page":"1067","DOI":"10.1007\/s10654-016-0206-y","volume":"31","author":"K Ten Haaf","year":"2016","unstructured":"Ten Haaf K, Steyerberg EW. Methods for individualized assessment of absolute risk in case-control studies should be weighted carefully. Eur J Epidemiol. 2016;31(11):1067\u20138.","journal-title":"Eur J Epidemiol"},{"issue":"2","key":"501_CR15","doi-asserted-by":"crossref","first-page":"452","DOI":"10.1158\/1055-9965.EPI-19-1221","volume":"29","author":"LH Chien","year":"2020","unstructured":"Chien LH, Chen CH, Chen TY, Chang GC, Tsai YH, Hsiao CF, Chen KY, Su WC, Wang WC, Huang MS, Chen YM. Predicting lung cancer occurrence in never-smoking females in Asia: TNSF-SQ, a prediction model. Cancer Epidemiol Prev Biomark. 2020;29(2):452\u20139.","journal-title":"Cancer Epidemiol Prev Biomark"},{"issue":"11_Supplement_1","key":"501_CR16","doi-asserted-by":"publisher","first-page":"194","DOI":"10.1016\/S0735-1097(20)30821-4","volume":"75","author":"D Mandair","year":"2020","unstructured":"Mandair D, Tiwari P, Simon S, Rosenberg M. Development of a prediction model for incident Myocardial infraction using machine learning applied to harmonized electronic health record data. J Am Coll Cardiol. 2020;75(11_Supplement_1):194.","journal-title":"J Am Coll Cardiol."},{"issue":"1","key":"501_CR17","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/s41467-019-13993-7","volume":"11","author":"WK Ho","year":"2020","unstructured":"Ho WK, Tan MM, Mavaddat N, Tai MC, Mariapun S, Li J, Ho PJ, Dennis J, Tyrer JP, Bolla MK, Michailidou K. European polygenic risk score for prediction of breast cancer shows similar performance in Asian women. Nat Commun. 2020;11(1):1\u20131.","journal-title":"Nat Commun"},{"issue":"4","key":"501_CR18","doi-asserted-by":"publisher","first-page":"373","DOI":"10.1136\/jnnp-2018-318212","volume":"90","author":"XH Hou","year":"2019","unstructured":"Hou XH, Feng L, Zhang C, Cao XP, Tan L, Yu JT. Models for predicting risk of dementia: a systematic review. J Neurol Neurosurg Psychiatry. 2019;90(4):373\u20139.","journal-title":"J Neurol Neurosurg Psychiatry"},{"key":"501_CR19","doi-asserted-by":"publisher","first-page":"e5900","DOI":"10.1136\/bmj.e5900","volume":"345","author":"A Abbasi","year":"2012","unstructured":"Abbasi A, Peelen LM, Corpeleijn E, Van Der Schouw YT, Stolk RP, Spijkerman AM, Moons KG, Navis G, Bakker SJ, Beulens JW. Prediction models for risk of developing type 2 diabetes: systematic literature search and independent external validation study. BMJ. 2012;345:e5900.","journal-title":"BMJ"},{"key":"501_CR20","first-page":"10","volume":"231","author":"MA - Suchard","year":"2013","unstructured":"- Suchard MA, Simpson SE, Zorych I. al. Massive parallelization of serial inference algorithms for complex generalized linear models. ACM Transact Model Comput Simulation. 2013;231:10\u201332.","journal-title":"ACM Transact Model Comput Simulation"},{"issue":"6","key":"501_CR21","doi-asserted-by":"publisher","first-page":"1052","DOI":"10.1093\/jamia\/ocx030","volume":"24","author":"E Sharon","year":"2017","unstructured":"Sharon E, Davis TA, Lasko G, Chen ED, Siew ME, Matheny. Calibration drift in regression and machine learning models for acute kidney injury. J Am Med Inform Assoc. 2017;24(6):1052\u201361.","journal-title":"J Am Med Inform Assoc"}],"container-title":["Journal of Big Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-021-00501-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s40537-021-00501-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-021-00501-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,7]],"date-time":"2023-01-07T07:28:02Z","timestamp":1673076482000},"score":1,"resource":{"primary":{"URL":"https:\/\/journalofbigdata.springeropen.com\/articles\/10.1186\/s40537-021-00501-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,16]]},"references-count":21,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["501"],"URL":"https:\/\/doi.org\/10.1186\/s40537-021-00501-2","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-341457\/v1","asserted-by":"object"}]},"ISSN":["2196-1115"],"issn-type":[{"value":"2196-1115","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,8,16]]},"assertion":[{"value":"18 March 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 August 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 August 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The use of Optum Claims was reviewed by the New England Institutional Review Board (IRB) and were determined to be exempt from broad IRB approval.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"JMR, PBR and MJS are employees of Janssen R&D and shareholders of Johnson and Johnson.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"108"}}