{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,19]],"date-time":"2026-02-19T15:26:39Z","timestamp":1771514799095,"version":"3.50.1"},"reference-count":26,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2024,8,21]],"date-time":"2024-08-21T00:00:00Z","timestamp":1724198400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Big Data"],"abstract":"<jats:p>Time series data are recorded in various sectors, resulting in a large amount of data. However, the continuity of these data is often interrupted, resulting in periods of missing data. Several algorithms are used to impute the missing data, and the performance of these methods is widely varied. Apart from the choice of algorithm, the effective imputation depends on the nature of missing and available data. We conducted extensive studies using different types of time series data, specifically heart rate data and power consumption data. We generated the missing data for different time spans and imputed using different algorithms with binned data of different sizes. The performance was evaluated using the root mean square error (RMSE) metric. We observed a reduction in RMSE when using binned data compared to the entire dataset, particularly in the case of the expectation\u2013maximization (EM) algorithm. We found that RMSE was reduced when using binned data for 1-, 5-, and 15-min missing data, with greater reduction observed for 15-min missing data. We also observed the effect of data fluctuation. We conclude that the usefulness of binned data depends precisely on the span of missing data, sampling frequency of the data, and fluctuation within data. Depending on the inherent characteristics, quality, and quantity of the missing and available data, binned data can impute a wide variety of data, including biological heart rate data derived from the Internet of Things (IoT) device smartwatch and non-biological data such as household power consumption data.<\/jats:p>","DOI":"10.3389\/fdata.2024.1422650","type":"journal-article","created":{"date-parts":[[2024,8,22]],"date-time":"2024-08-22T17:16:26Z","timestamp":1724346986000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Efficient use of binned data for imputing univariate time series data"],"prefix":"10.3389","volume":"7","author":[{"given":"Jay","family":"Darji","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nupur","family":"Biswas","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vijay","family":"Padul","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jaya","family":"Gill","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Santosh","family":"Kesari","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shashaanka","family":"Ashili","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2024,8,21]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"e1873","DOI":"10.1002\/met.1873","article-title":"Missing data imputation of high-resolution temporal climate time series data","volume":"27","author":"Afrifa-Yamoah","year":"2020","journal-title":"Meteorol. Appl."},{"key":"B2","doi-asserted-by":"publisher","first-page":"767","DOI":"10.32604\/cmc.2022.019369","article-title":"Comparison of missing data imputation methods in time series forecasting","volume":"70","author":"Ahn","year":"2022","journal-title":"Comp. Mater. Cont."},{"key":"B3","doi-asserted-by":"publisher","first-page":"44483","DOI":"10.1109\/ACCESS.2022.3160841","article-title":"Systematic review of using machine learning in imputing missing values","volume":"10","author":"Alabadla","year":"2022","journal-title":"IEEE Access"},{"key":"B4","doi-asserted-by":"publisher","first-page":"1454","DOI":"10.3390\/s23031454","article-title":"Binned data provide better imputation of missing time series data from wearables","volume":"23","author":"Chakrabarti","year":"2023","journal-title":"Sensors"},{"key":"B5","article-title":"\u201cHandling missing data in the time-series data from wearables,\u201d","volume-title":"Time Series Analysis - Recent Advances, New Perspectives and Applications","author":"Darji","year":"2023"},{"key":"B6","doi-asserted-by":"publisher","first-page":"199","DOI":"10.1016\/0169-2070(91)90054-Y","article-title":"Seasonality, non-stationarity and the forecasting of monthly time series","volume":"7","author":"Franses","year":"1991","journal-title":"Int. J. Forecast."},{"key":"B7","unstructured":"HebrailG.\n            BerardA.\n          10.24432\/C58K5437860605Individual Household Electric Power Consumption2012"},{"key":"B8","doi-asserted-by":"publisher","first-page":"561","DOI":"10.1111\/j.1540-5907.2010.00447.x","article-title":"What to do about missing values in time-series cross-section data","volume":"54","author":"Honaker","year":"2010","journal-title":"Am. J. Pol. Sci."},{"key":"B9","doi-asserted-by":"publisher","first-page":"199","DOI":"10.1186\/s12874-020-01080-1","article-title":"Accuracy of random-forest-based imputation of missing data in the presence of non-normality, non-linearity, and interaction","volume":"20","author":"Hong","year":"2020","journal-title":"BMC Med. Res. Methodol."},{"key":"B10","doi-asserted-by":"publisher","first-page":"101","DOI":"10.1177\/1559827619878661","article-title":"Physical activity practice and healthy lifestyles related to resting heart rate in health sciences first-year students","volume":"16","author":"Hon\u00f3rio","year":"2022","journal-title":"Am. J. Lifestyle Med."},{"key":"B11","doi-asserted-by":"publisher","first-page":"96","DOI":"10.1016\/j.atmosenv.2014.11.049","article-title":"Imputation of missing data in time series for air pollutants","volume":"102","author":"Junger","year":"2015","journal-title":"Atmos. Environ."},{"key":"B12","doi-asserted-by":"publisher","first-page":"768","DOI":"10.14778\/3377369.3377383","article-title":"Mind the gap: an experimental evaluation of imputation of missing values techniques in time series","volume":"13","author":"Khayati","year":"2020","journal-title":"Proc. VLDB Endow."},{"key":"B13","unstructured":"\u201cThe effects of the irregular sample and missing data in time series analysis,\u201d135157\n            KreindlerD. M.\n            LumsdenC. J.\n          CRC PressNonlinear Dynamical Systems Analysis for the Behavioral Sciences Using Real Data2016"},{"key":"B14","doi-asserted-by":"publisher","first-page":"e0262131","DOI":"10.1371\/journal.pone.0262131","article-title":"Imputation by feature importance (IBFI): a methodology to envelop machine learning method for imputing missing patterns in time series data","volume":"17","author":"Mir","year":"2022","journal-title":"PLOS ONE"},{"key":"B15","first-page":"511","article-title":"\u201cMultiple imputation and the expectation-maximization algorithm,\u201d","volume-title":"Models for Discrete Longitudinal Data","author":"Molenberghs","year":"2005"},{"key":"B16","doi-asserted-by":"publisher","first-page":"107167","DOI":"10.1016\/j.asoc.2021.107167","article-title":"Modulo 9 model-based learning for missing data imputation","volume":"103","author":"Ngueilbaye","year":"2021","journal-title":"Appl. Soft Comput."},{"key":"B17","first-page":"2825","article-title":"Scikit-learn: machine learning in python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"B18","article-title":"\u201cA review of missing values handling methods on time-series data,\u201d","volume-title":"2016 International Conference on Information Technology Systems and Innovation, ICITSI 2016 \u2013 Proceedings","author":"Pratama","year":"2017"},{"key":"B19","volume-title":"Time Series Analysis and Its Applications. Springer Texts in Statistics.","author":"Shumway","year":"2017"},{"key":"B20","doi-asserted-by":"publisher","first-page":"112","DOI":"10.1093\/bioinformatics\/btr597","article-title":"MissForest\u2014non-parametric missing value imputation for mixed-type data","volume":"28","author":"Stekhoven","year":"2012","journal-title":"Bioinformatics"},{"key":"B21","doi-asserted-by":"publisher","first-page":"308","DOI":"10.1186\/1756-0500-4-308","article-title":"Simple parametric survival analysis with anonymized register data: a cohort study with truncated and interval censored event and censoring times","volume":"4","author":"St\u00f8vring","year":"2011","journal-title":"BMC Res. Notes"},{"key":"B22","doi-asserted-by":"crossref","first-page":"216","DOI":"10.1109\/FMEC.2019.8795338","article-title":"\u201cSmartwatches as IoT edge devices: a framework and survey,\u201d","volume-title":"2019 Fourth International Conference on Fog and Mobile Edge Computing (FMEC)","author":"Takiddeen","year":"2019"},{"key":"B23","doi-asserted-by":"publisher","first-page":"363","DOI":"10.1002\/sam.11348","article-title":"Random forest missing data algorithms","volume":"10","author":"Tang","year":"2017","journal-title":"Stat. Anal. Data Mining"},{"key":"B24","doi-asserted-by":"publisher","first-page":"2793","DOI":"10.1016\/j.csda.2011.04.012","article-title":"Iterative stepwise regression imputation using standard and robust methods","volume":"55","author":"Templ","year":"2011","journal-title":"Comput. Stat. Data Anal."},{"key":"B25","doi-asserted-by":"crossref","first-page":"595","DOI":"10.1016\/B978-0-12-818803-3.00023-4","article-title":"\u201cBayesian learning: inference and the EM algorithm,\u201d","volume-title":"Machine Learning : A Bayesian and Optimization Perspective, 2nd Edn","author":"Theodoridis","year":"2020"},{"key":"B26","doi-asserted-by":"publisher","first-page":"2541","DOI":"10.1016\/j.jss.2012.05.073","article-title":"Nearest neighbor selection for iteratively KNN imputation","volume":"85","author":"Zhang","year":"2012","journal-title":"J. Syst. Softw."}],"container-title":["Frontiers in Big Data"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdata.2024.1422650\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,22]],"date-time":"2024-08-22T17:16:55Z","timestamp":1724347015000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdata.2024.1422650\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,21]]},"references-count":26,"alternative-id":["10.3389\/fdata.2024.1422650"],"URL":"https:\/\/doi.org\/10.3389\/fdata.2024.1422650","relation":{},"ISSN":["2624-909X"],"issn-type":[{"value":"2624-909X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,8,21]]},"article-number":"1422650"}}