{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T04:47:34Z","timestamp":1775018854552,"version":"3.50.1"},"reference-count":15,"publisher":"Oxford University Press (OUP)","issue":"7","license":[{"start":{"date-parts":[[2022,4,18]],"date-time":"2022-04-18T00:00:00Z","timestamp":1650240000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,6,14]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Objective<\/jats:title>\n                  <jats:p>The goals of this study were to harmonize data from electronic health records (EHRs) into common units, and impute units that were missing.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Materials and Methods<\/jats:title>\n                  <jats:p>The National COVID Cohort Collaborative (N3C) table of laboratory measurement data\u2014over 3.1 billion patient records and over 19\u00a0000 unique measurement concepts in the Observational Medical Outcomes Partnership (OMOP) common-data-model format from 55 data partners. We grouped ontologically similar OMOP concepts together for 52 variables relevant to COVID-19 research, and developed a unit-harmonization pipeline comprised of (1) selecting a canonical unit for each measurement variable, (2) arriving at a formula for conversion, (3) obtaining clinical review of each formula, (4) applying the formula to convert data values in each unit into the target canonical unit, and (5) removing any harmonized value that fell outside of accepted value ranges for the variable. For data with missing units for all the results within a lab test for a data partner, we compared values with pooled values of all data partners, using the Kolmogorov-Smirnov test.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>Of the concepts without missing values, we harmonized 88.1% of the values, and imputed units for 78.2% of records where units were absent (41% of contributors\u2019 records lacked units).<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Discussion<\/jats:title>\n                  <jats:p>The harmonization and inference methods developed herein can serve as a resource for initiatives aiming to extract insight from heterogeneous EHR collections. Unique properties of centralized data are harnessed to enable unit inference.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Conclusion<\/jats:title>\n                  <jats:p>The pipeline we developed for the pooled N3C data enables use of measurements that would otherwise be unavailable for analysis.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/jamia\/ocac054","type":"journal-article","created":{"date-parts":[[2022,4,14]],"date-time":"2022-04-14T11:11:57Z","timestamp":1649934717000},"page":"1172-1182","source":"Crossref","is-referenced-by-count":26,"title":["Harmonizing units and values of quantitative data elements in a very large nationally pooled electronic health record (EHR) dataset"],"prefix":"10.1093","volume":"29","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9730-1808","authenticated-orcid":false,"given":"Katie R","family":"Bradwell","sequence":"first","affiliation":[{"name":"Palantir Technologies , Denver, Colorado, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jacob T","family":"Wooldridge","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Stony Brook University , Stony Brook, New York, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Benjamin","family":"Amor","sequence":"additional","affiliation":[{"name":"Palantir Technologies , Denver, Colorado, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1483-4236","authenticated-orcid":false,"given":"Tellen D","family":"Bennett","sequence":"additional","affiliation":[{"name":"Section of Informatics and Data Science, Department of Pediatrics, University of Colorado School of Medicine, University of Colorado , Aurora, Colorado, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Adit","family":"Anand","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Stony Brook University , Stony Brook, New York, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Carolyn","family":"Bremer","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Stony Brook University , Stony Brook, New York, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yun Jae","family":"Yoo","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Stony Brook University , Stony Brook, New York, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhenglong","family":"Qian","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Stony Brook University , Stony Brook, New York, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Steven G","family":"Johnson","sequence":"additional","affiliation":[{"name":"Institute for Health Informatics, University of Minnesota , Minneapolis, Minnesota, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6840-9756","authenticated-orcid":false,"given":"Emily R","family":"Pfaff","sequence":"additional","affiliation":[{"name":"Department of Medicine, North Carolina Translational and Clinical Sciences Institute, University of North Carolina at Chapel Hill , Chapel Hill, North Carolina, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrew T","family":"Girvin","sequence":"additional","affiliation":[{"name":"Palantir Technologies , Denver, Colorado, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Amin","family":"Manna","sequence":"additional","affiliation":[{"name":"Palantir Technologies , Denver, Colorado, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Emily A","family":"Niehaus","sequence":"additional","affiliation":[{"name":"Palantir Technologies , Denver, Colorado, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stephanie S","family":"Hong","sequence":"additional","affiliation":[{"name":"School of Medicine, Section of Biomedical Informatics and Data Science, Johns Hopkins University School of Medicine , Baltimore, Maryland, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaohan Tanner","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Medicine, Johns Hopkins , Baltimore, Maryland, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4289-7632","authenticated-orcid":false,"given":"Richard L","family":"Zhu","sequence":"additional","affiliation":[{"name":"Department of Medicine, Johns Hopkins , Baltimore, Maryland, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mark","family":"Bissell","sequence":"additional","affiliation":[{"name":"Palantir Technologies , Denver, Colorado, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nabeel","family":"Qureshi","sequence":"additional","affiliation":[{"name":"Palantir Technologies , Denver, Colorado, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joel","family":"Saltz","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Stony Brook University , Stony Brook, New York, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9114-8737","authenticated-orcid":false,"given":"Melissa A","family":"Haendel","sequence":"additional","affiliation":[{"name":"Center for Health AI, University of Colorado , Aurora, Colorado, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5437-2545","authenticated-orcid":false,"given":"Christopher G","family":"Chute","sequence":"additional","affiliation":[{"name":"Schools of Medicine, Public Health, and Nursing, Johns Hopkins University , Baltimore, Maryland, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Harold P","family":"Lehmann","sequence":"additional","affiliation":[{"name":"Department of Medicine, Johns Hopkins , Baltimore, Maryland, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2723-5902","authenticated-orcid":false,"given":"Richard A","family":"Moffitt","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Stony Brook University , Stony Brook, New York, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"name":"the N3C Consortium","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2022,4,18]]},"reference":[{"issue":"3","key":"2022061415523996400_ocac054-B1","doi-asserted-by":"crossref","first-page":"427","DOI":"10.1093\/jamia\/ocaa196","article-title":"The National COVID Cohort Collaborative (N3C): rationale, design, infrastructure, and deployment","volume":"28","author":"Haendel","year":"2021","journal-title":"J Am Med Inform Assoc"},{"key":"2022061415523996400_ocac054-B2","year":"2018"},{"key":"2022061415523996400_ocac054-B3","first-page":"574","article-title":"Observational Health Data Sciences and Informatics (OHDSI): opportunities for observational researchers","volume":"216","author":"Hripcsak","year":"2015","journal-title":"Stud Health Technol Inform"},{"issue":"3","key":"2022061415523996400_ocac054-B4","doi-asserted-by":"crossref","first-page":"276","DOI":"10.1136\/jamia.1998.0050276","article-title":"Development of the Logical Observation Identifier Names and Codes (LOINC) vocabulary","volume":"5","author":"Huff","year":"1998","journal-title":"J Am Med Inform Assoc"},{"key":"2022061415523996400_ocac054-B5"},{"key":"2022061415523996400_ocac054-B6","first-page":"392","article-title":"Interoperability of medical databases: construction of mapping between hospitals laboratory results assisted by automated comparison of their distributions","volume":"2011","author":"Ficheur","year":"2011","journal-title":"AMIA Annu Symp Proc"},{"key":"2022061415523996400_ocac054-B7","first-page":"234","article-title":"Standardizing the unit of measurements in LOINC-coded laboratory tests can significantly improve semantic interoperability","volume":"275","author":"Rajput","year":"2020","journal-title":"Stud Health Technol Inform"},{"key":"2022061415523996400_ocac054-B8","first-page":"437","article-title":"The LOINC content model and its limitations of usage in the laboratory domain","volume":"270","author":"Drenkhahn","year":"2020","journal-title":"Stud Health Technol Inform"},{"key":"2022061415523996400_ocac054-B9","first-page":"108","article-title":"Aggregation and visualization of laboratory data by using ontological tools based on LOINC and SNOMED CT","volume":"264","author":"Drenkhahn","year":"2019","journal-title":"Stud Health Technol Inform"},{"issue":"2","key":"2022061415523996400_ocac054-B10","doi-asserted-by":"crossref","first-page":"192","DOI":"10.1093\/jamia\/ocx056","article-title":"Unit conversions between LOINC codes","volume":"25","author":"Hauser","year":"2018","journal-title":"J Am Med Inform Assoc"},{"key":"2022061415523996400_ocac054-B11","doi-asserted-by":"crossref","first-page":"614","DOI":"10.1093\/jamia\/ocx087","article-title":"Response to unit conversions between LOINC codes","author":"Vreeman","year":"2018","journal-title":"J Am Med Inform Assoc"},{"issue":"7","key":"2022061415523996400_ocac054-B12","doi-asserted-by":"crossref","first-page":"e2116901","DOI":"10.1001\/jamanetworkopen.2021.16901","article-title":"Clinical characterization and prediction of clinical severity of SARS-CoV-2 infection among US adults using data from the US National COVID Cohort Collaborative","volume":"4","author":"Bennett","year":"2021","journal-title":"JAMA Netw Open"},{"key":"2022061415523996400_ocac054-B13","volume-title":"Handbook of Methods of Applied Statistics","author":"Chakravarti","year":"1967"},{"issue":"2128","key":"2022061415523996400_ocac054-B14","article-title":"Improving reproducibility by using high-throughput observational studies with empirical calibration","volume":"376","author":"Schuemie","journal-title":"Philos Trans A Math Phys Eng Sci"},{"key":"2022061415523996400_ocac054-B15","year":"2021"}],"container-title":["Journal of the American Medical Informatics Association"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/29\/7\/1172\/44062146\/ocac054.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/29\/7\/1172\/44062146\/ocac054.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,6,14]],"date-time":"2022-06-14T16:49:19Z","timestamp":1655225359000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jamia\/article\/29\/7\/1172\/6569865"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,18]]},"references-count":15,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2022,4,18]]},"published-print":{"date-parts":[[2022,6,14]]}},"URL":"https:\/\/doi.org\/10.1093\/jamia\/ocac054","relation":{},"ISSN":["1527-974X"],"issn-type":[{"value":"1527-974X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022,7,1]]},"published":{"date-parts":[[2022,4,18]]}}}