{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,22]],"date-time":"2025-11-22T08:37:52Z","timestamp":1763800672058,"version":"3.45.0"},"reference-count":13,"publisher":"Oxford University Press (OUP)","license":[{"start":{"date-parts":[[2025,11,22]],"date-time":"2025-11-22T00:00:00Z","timestamp":1763769600000},"content-version":"vor","delay-in-days":325,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,1,18]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>The National Health and Nutrition Examination Survey (NHANES) provides extensive public data on demographics, health, and nutrition, collected in 2-year cycles since 1999. Although invaluable for epidemiological and health-related research, the complexity of NHANES data, involving numerous files and disjoint metadata, makes accessing, managing, and analysing these datasets challenging. This paper presents a reproducible computational environment built upon Docker containers, PostgreSQL databases, and R\/RStudio, designed to streamline NHANES data management, facilitate rigorous quality control, and simplify analyses across multiple survey cycles. We introduce specialized tools, such as the enhanced nhanesA R package and the phonto R package, to provide fast access to data, to help manage metadata, and to handle complexities arising from questionnaire design and cross-cycle data inconsistencies. Furthermore, we describe the Epiconnector platform, established to foster collaborative sharing of code, analytical scripts, and best practices, which taken together, can significantly enhance the reproducibility, extensibility, and robustness of scientific research using NHANES data.<\/jats:p>","DOI":"10.1093\/database\/baaf073","type":"journal-article","created":{"date-parts":[[2025,11,22]],"date-time":"2025-11-22T08:35:27Z","timestamp":1763800527000},"source":"Crossref","is-referenced-by-count":0,"title":["Enhancing statistical analysis of real world data"],"prefix":"10.1093","volume":"2025","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4070-5289","authenticated-orcid":false,"given":"Laha","family":"Ale","sequence":"first","affiliation":[{"name":"School of Computing and Artificial Intelligence, Southwest Jiaotong University , No. 999, Xi'an Rd, Chengdu, Sichuan, 611756 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4505-9893","authenticated-orcid":false,"given":"Robert","family":"Gentleman","sequence":"additional","affiliation":[{"name":"Dana Farber Cancer Institute Department of Data Science, , 450 Brookline Avenue, Boston, MA 02215 ,","place":["United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christopher","family":"Endres","sequence":"additional","affiliation":[{"name":"The Promenade Dance Studio, Inc. , 2605 Lord Baltimore Dr, Windsor Mill, MD 21244 ,","place":["United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sam","family":"Pullman","sequence":"additional","affiliation":[{"name":"Core for Computational Biomedicine, Harvard Medical School , 25 Shattuck Street, Boston, MA 02115 ,","place":["United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nathan","family":"Palmer","sequence":"additional","affiliation":[{"name":"Core for Computational Biomedicine, Harvard Medical School , 25 Shattuck Street, Boston, MA 02115 ,","place":["United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1255-0125","authenticated-orcid":false,"given":"Rafael","family":"Goncalves","sequence":"additional","affiliation":[{"name":"Center for Biomedical Informatics Research, Stanford University , 3180 Porter Dr, Palo Alto, CA 94304 ,","place":["United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4107-1553","authenticated-orcid":false,"given":"Deepayan","family":"Sarkar","sequence":"additional","affiliation":[{"name":"Statistical Sciences Unit, Indian Statistical Institute , 7 S.J.S. Sansanwal Marg New Delhi 110016 ,","place":["India"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2025,11,22]]},"reference":[{"key":"2025112203352252200_bib1","author":"CDC","journal-title":"National health and Nutrition Examination Survey"},{"key":"2025112203352252200_bib2","doi-asserted-by":"publisher","first-page":"e1001779","DOI":"10.1371\/journal.pmed.1001779","article-title":"UK biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age","volume":"12","author":"Sudlow","year":"2015","journal-title":"PLoS Med"},{"key":"2025112203352252200_bib3","article-title":"Harmonized US national health and nutrition examination survey 1988\u20132018 for high throughput exposome-health discovery","author":"Nguyen","year":"2023","journal-title":"medRxiv [Internet]"},{"key":"2025112203352252200_bib4","doi-asserted-by":"publisher","DOI":"10.1093\/database\/baae028","article-title":"nhanesA: achieving transparency and reproducibility in NHANES research","author":"Ale","year":"2024","journal-title":"Database: J Biol Databases Curation"},{"key":"2025112203352252200_bib5","volume-title":"Record layout of a SAS version 5 or 6 data set in SAS transport (xport) format","year":"2021"},{"key":"2025112203352252200_bib6","article-title":"BiocFileCache: manage files across sessions","author":"Shepherd","year":"2024"},{"key":"2025112203352252200_bib7","volume-title":"R Special Interest Group on Databases (R-SIG-DB)","author":"Wickham","year":"2024"},{"key":"2025112203352252200_bib8","article-title":"Dplyr: a grammar of data manipulation","author":"Wickham","year":"2023"},{"key":"2025112203352252200_bib9","author":"Ale"},{"key":"2025112203352252200_bib10","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1198\/106186007X178663","article-title":"Statistical analyses and reproducible research","volume":"16","author":"Gentleman","year":"2007","journal-title":"J Comput Graph Stat [Internet]"},{"key":"2025112203352252200_bib11","author":"Xie","year":"2024","journal-title":"Bookdown: Authoring books and technical documents with R markdown"},{"key":"2025112203352252200_bib12","first-page":"1","article-title":"Data Science at the Singularity","volume":"6","author":"Donoho","year":"2024","journal-title":"Harvard Data Sci Rev"},{"key":"2025112203352252200_bib13","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/sdata.2016.96","article-title":"A database of human exposomes and phenomes from the US National Health and Nutrition Examination Survey","volume":"3","author":"Patel","year":"2016","journal-title":"Sci Data"}],"container-title":["Database"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baaf073\/65468655\/baaf073.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baaf073\/65468655\/baaf073.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,22]],"date-time":"2025-11-22T08:35:29Z","timestamp":1763800529000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/database\/article\/doi\/10.1093\/database\/baaf073\/8340171"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025]]},"references-count":13,"URL":"https:\/\/doi.org\/10.1093\/database\/baaf073","relation":{},"ISSN":["1758-0463"],"issn-type":[{"value":"1758-0463","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2025]]},"published":{"date-parts":[[2025]]},"article-number":"baaf073"}}