{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"institution":[{"name":"medRxiv"}],"indexed":{"date-parts":[[2026,10,7]],"date-time":"2026-10-07T10:50:02Z","timestamp":1791370202275,"version":"4.3.1"},"posted":{"date-parts":[[2026,10,2]]},"group-title":"Health Informatics","reference-count":19,"publisher":"openRxiv","license":[{"start":{"date-parts":[[2026,10,2]],"date-time":"2026-10-02T00:00:00Z","timestamp":1790899200000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Human Genome Research Institute","award":["UG3HG014379"],"award-info":[{"award-number":["UG3HG014379"]}],"id":[{"id":"https:\/\/ror.org\/00baak391","id-type":"ROR","asserted-by":"crossref"}]},{"name":"Broad Next Generation Fund"},{"name":"Novo Nordisk Foundation","award":["NNF21SA0072102"],"award-info":[{"award-number":["NNF21SA0072102"]}],"id":[{"id":"https:\/\/ror.org\/04txyc737","id-type":"ROR","asserted-by":"crossref"}]},{"name":"Firuza Foundation"},{"name":"National Institutes of Health","award":["OT2OD038121"],"award-info":[{"award-number":["OT2OD038121"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"accepted":{"date-parts":[[2026,10,2]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                <jats:p>\n                  Advances in high-throughput proteomics technologies have enabled the assessment of dynamic health states across biobank-scale cohorts. Disease prediction models built on these data have higher accuracy than baseline clinical models for a broad range of diseases, and provide avenues to understand disease pathogenesis. However, the generalizability of these prediction models across cohorts has yet to be established at scale. Here, we train models for 15 diseases in the UK Biobank (UKB; n=53,026) based on Olink proteomics data, and find high accuracy for disease prediction (mean AUC=0.74; range 0.56-0.89). These models have improved accuracy over models using only clinical factors (mean \u0394AUC = 0.03), and do not depend on model architecture, with simpler models (e.g. L2) performing as well as complex models (e.g. transformer). We assess generalizability of the UKB-trained models in two external cohorts: FinnGen (FG; n=5,865) and\n                  <jats:italic>All of Us<\/jats:italic>\n                  (AoU; n=7,405), spanning two Olink platforms. We find that proteomics-based prevalent (classification) and incident (prediction over next 5 years) disease models largely generalize across cohorts, but performance varies across diseases. Specifically, 10 of 14 prevalent and 13 of 15 incident disease models show no significant decrease in performance across any biobank. Adjusting for demographic, ancestry, and technical covariates, we demonstrate that differences in phenotyping quality are likely the major drivers of variability across cohorts. These results indicate that proteomic risk models can be powerful and generalizable predictors of disease across multiple cohorts.\n                <\/jats:p>","DOI":"10.64898\/2026.10.01.26364038","type":"posted-content","created":{"date-parts":[[2026,10,3]],"date-time":"2026-10-03T00:05:13Z","timestamp":1790985913000},"source":"Crossref","is-referenced-by-count":0,"title":["Generalizability of proteomic risk prediction across biobanks reveals dependence on phenotype definitions"],"prefix":"10.64898","author":[{"given":"Sophie J.","family":"Parsa","sequence":"first","affiliation":[{"name":"Analytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA"},{"name":"Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"},{"name":"Harvard-MIT Program in Health Sciences and Technology, Harvard Medical School, Boston, MA 02446"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yon Ho","family":"Jee","sequence":"additional","affiliation":[{"name":"Analytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA"},{"name":"Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"M. Austin","family":"Argentieri","sequence":"additional","affiliation":[{"name":"Analytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA"},{"name":"Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jeremy","family":"Guez","sequence":"additional","affiliation":[{"name":"Analytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA"},{"name":"Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3048-9658","authenticated-orcid":false,"given":"Wenhan","family":"Lu","sequence":"additional","affiliation":[{"name":"Analytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA"},{"name":"Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Evin M.","family":"Padhi","sequence":"additional","affiliation":[{"name":"Department of Pathology, Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Elizabeth","family":"Kiernan","sequence":"additional","affiliation":[{"name":"Broad Genomics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Reza","family":"Jabal","sequence":"additional","affiliation":[{"name":"Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mitja I.","family":"Kurki","sequence":"additional","affiliation":[{"name":"Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"name":"FinnGen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1513-6077","authenticated-orcid":false,"given":"Benjamin M.","family":"Neale","sequence":"additional","affiliation":[{"name":"Analytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA"},{"name":"Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"},{"name":"Novo Nordisk Foundation Center for Genomic Mechanisms of Disease, Broad Institute of MIT and Harvard, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lee","family":"Lichtenstein","sequence":"additional","affiliation":[{"name":"Broad Genomics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Michael E.","family":"Talkowski","sequence":"additional","affiliation":[{"name":"Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"},{"name":"Center for Genomic Medicine, Massachusetts General Hospital, Boston, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stephen B.","family":"Montgomery","sequence":"additional","affiliation":[{"name":"Department of Pathology, Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2874-7371","authenticated-orcid":false,"given":"Niall J.","family":"Lennon","sequence":"additional","affiliation":[{"name":"Broad Genomics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0241-3522","authenticated-orcid":false,"given":"Alicia R.","family":"Martin","sequence":"additional","affiliation":[{"name":"Analytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA"},{"name":"Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0949-8752","authenticated-orcid":false,"given":"Mark J.","family":"Daly","sequence":"additional","affiliation":[{"name":"Analytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA"},{"name":"Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Konrad J.","family":"Karczewski","sequence":"additional","affiliation":[{"name":"Analytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA"},{"name":"Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA"},{"name":"Novo Nordisk Foundation Center for Genomic Mechanisms of Disease, Broad Institute of MIT and Harvard, Cambridge, MA, USA"},{"name":"Center for Genomic Medicine, Massachusetts General Hospital, Boston, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"54368","reference":[{"key":"2026100702550479000_2026.10.01.26364038v1.1","doi-asserted-by":"publisher","DOI":"10.1056\/NEJMoa2401389"},{"key":"2026100702550479000_2026.10.01.26364038v1.2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejca.2019.12.013"},{"key":"2026100702550479000_2026.10.01.26364038v1.3","doi-asserted-by":"crossref","first-page":"e242814","DOI":"10.1001\/jamahealthforum.2024.2814","article-title":"Redefining cancer screening coverage-screening to diagnosis","volume":"5","year":"2024","journal-title":"JAMA Health Forum"},{"key":"2026100702550479000_2026.10.01.26364038v1.4","doi-asserted-by":"publisher","DOI":"10.1056\/NEJMoa0810084"},{"key":"2026100702550479000_2026.10.01.26364038v1.5","unstructured":"Breast Cancer Risk Assessment and Screening in Average-Risk Women (2011). https:\/\/www.acog.org\/clinical\/clinical-guidance\/practice-bulletin\/articles\/2017\/07\/breast-cancer-risk-assessment-and-screening-in-average-risk-women."},{"key":"2026100702550479000_2026.10.01.26364038v1.6","doi-asserted-by":"publisher","DOI":"10.1161\/01.HYP.25.2.155"},{"key":"2026100702550479000_2026.10.01.26364038v1.7","doi-asserted-by":"publisher","DOI":"10.1101\/2025.08.27.25334571"},{"key":"2026100702550479000_2026.10.01.26364038v1.8","doi-asserted-by":"publisher","DOI":"10.1038\/s41591-024-03164-7"},{"key":"2026100702550479000_2026.10.01.26364038v1.9","doi-asserted-by":"crossref","first-page":"928","DOI":"10.1038\/s41467-025-67656-x","article-title":"Proteomic signatures of smoking and their associations with risk of incident diseases and mortality in diverse populations","volume":"17","year":"2025","journal-title":"Nat. Commun"},{"key":"2026100702550479000_2026.10.01.26364038v1.10","doi-asserted-by":"publisher","DOI":"10.1016\/j.cell.2024.10.045"},{"key":"2026100702550479000_2026.10.01.26364038v1.11","doi-asserted-by":"crossref","first-page":"20520","DOI":"10.1038\/s41598-025-06232-1","article-title":"Proteomic risk scores for predicting common diseases using linear and neural network models in the UK biobank","volume":"15","year":"2025","journal-title":"Sci. Rep"},{"key":"2026100702550479000_2026.10.01.26364038v1.12","doi-asserted-by":"publisher","DOI":"10.1101\/2025.02.19.25322536"},{"key":"2026100702550479000_2026.10.01.26364038v1.13","doi-asserted-by":"publisher","DOI":"10.1038\/s41591-024-03142-z"},{"key":"2026100702550479000_2026.10.01.26364038v1.14","first-page":"939","article-title":"Blood protein assessment of leading incident diseases and mortality in the UK Biobank. Nat","volume":"4","year":"2024","journal-title":"Aging"},{"key":"2026100702550479000_2026.10.01.26364038v1.15","first-page":"e470","article-title":"Proteomic prediction of diverse incident diseases: a machine learning-guided biomarker discovery study using data from a prospective cohort study. Lancet Digit","volume":"6","year":"2024","journal-title":"Health"},{"key":"2026100702550479000_2026.10.01.26364038v1.16","doi-asserted-by":"crossref","first-page":"e18","DOI":"10.1017\/S1463423625000180","article-title":"Validation of the Finnish Diabetes Risk Score and development of a country-specific diabetes prediction model for Turkey","volume":"26","year":"2025","journal-title":"Prim. Health Care Res. Dev"},{"key":"2026100702550479000_2026.10.01.26364038v1.17","doi-asserted-by":"crossref","first-page":"2010","DOI":"10.1038\/s42255-024-01133-5","article-title":"Mapping biological influences on the human plasma proteome beyond the genome","volume":"6","year":"2024","journal-title":"Nat. Metab"},{"key":"2026100702550479000_2026.10.01.26364038v1.18","doi-asserted-by":"publisher","DOI":"10.1038\/ng.3211"},{"key":"2026100702550479000_2026.10.01.26364038v1.19","doi-asserted-by":"publisher","DOI":"10.1016\/j.cell.2026.03.049"}],"container-title":[],"original-title":[],"link":[{"URL":"https:\/\/syndication.highwire.org\/content\/doi\/10.64898\/2026.10.01.26364038","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,10,7]],"date-time":"2026-10-07T09:55:17Z","timestamp":1791366917000},"score":1,"resource":{"primary":{"URL":"http:\/\/medrxiv.org\/lookup\/doi\/10.64898\/2026.10.01.26364038"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,10,2]]},"references-count":19,"URL":"https:\/\/doi.org\/10.64898\/2026.10.01.26364038","relation":{},"subject":[],"published":{"date-parts":[[2026,10,2]]},"subtype":"preprint"}}