{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,7]],"date-time":"2026-05-07T16:10:13Z","timestamp":1778170213185,"version":"3.51.4"},"reference-count":24,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,6,4]],"date-time":"2024-06-04T00:00:00Z","timestamp":1717459200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,6,4]],"date-time":"2024-06-04T00:00:00Z","timestamp":1717459200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["RF1-AG062109"],"award-info":[{"award-number":["RF1-AG062109"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["U19-AG068753"],"award-info":[{"award-number":["U19-AG068753"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["RF1-AG062109"],"award-info":[{"award-number":["RF1-AG062109"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["RF1-AG062109"],"award-info":[{"award-number":["RF1-AG062109"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Med Inform Decis Mak"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Background<\/jats:title>\n                <jats:p>Machine learning (ML) has emerged as the predominant computational paradigm for analyzing large-scale datasets across diverse domains. The assessment of dataset quality stands as a pivotal precursor to the successful deployment of ML models. In this study, we introduce DREAMER (<jats:bold>D<\/jats:bold>ata <jats:bold>REA<\/jats:bold>diness for <jats:bold>M<\/jats:bold>achin<jats:bold>E<\/jats:bold> learning <jats:bold>R<\/jats:bold>esearch), an algorithmic framework leveraging supervised and unsupervised machine learning techniques to autonomously evaluate the suitability of tabular datasets for ML model development. DREAMER is openly accessible as a tool on GitHub and Docker, facilitating its adoption and further refinement within the research community..\n<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>The proposed model in this study was applied to three distinct tabular datasets, resulting in notable enhancements in their quality with respect to readiness for ML tasks, as assessed through established data quality metrics. Our findings demonstrate the efficacy of the framework in substantially augmenting the original dataset quality, achieved through the elimination of extraneous features and rows. This refinement yielded improved accuracy across both supervised and unsupervised learning methodologies.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusion<\/jats:title>\n                <jats:p>Our software presents an automated framework for data readiness, aimed at enhancing the integrity of raw datasets to facilitate robust utilization within ML pipelines. Through our proposed framework, we streamline the original dataset, resulting in enhanced accuracy and efficiency within the associated ML algorithms.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1186\/s12911-024-02544-w","type":"journal-article","created":{"date-parts":[[2024,6,4]],"date-time":"2024-06-04T04:31:15Z","timestamp":1717475475000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["DREAMER: a computational framework to evaluate readiness of datasets for machine learning"],"prefix":"10.1186","volume":"24","author":[{"given":"Meysam","family":"Ahangaran","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hanzhi","family":"Zhu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ruihui","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lingkai","family":"Yin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joseph","family":"Jang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arnav P.","family":"Chaudhry","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lindsay A.","family":"Farrer","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rhoda","family":"Au","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vijaya B.","family":"Kolachalama","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,6,4]]},"reference":[{"key":"2544_CR1","doi-asserted-by":"publisher","first-page":"160","DOI":"10.1007\/s42979-021-00592-x","volume":"2","author":"IH Sarker","year":"2021","unstructured":"Sarker IH. Machine learning: algorithms, real-world applications and research directions. SN Comput Sci. 2021;2:160.","journal-title":"SN Comput Sci"},{"key":"2544_CR2","unstructured":"Lawrence ND. Data readiness levels. arXiv preprint arXiv:170502245. 2017."},{"key":"2544_CR3","doi-asserted-by":"publisher","first-page":"18005","DOI":"10.1038\/s41598-021-97341-0","volume":"11","author":"MA Dakka","year":"2021","unstructured":"Dakka MA, Nguyen TV, Hall JMM, Diakiw SM, VerMilyea M, Linke R, et al. Automated detection of poor-quality data: case studies in healthcare. Sci Rep. 2021;11:18005.","journal-title":"Sci Rep"},{"key":"2544_CR4","doi-asserted-by":"crossref","unstructured":"Austin CC. A path to big data readiness. In: 2018 IEEE International Conference on Big Data (Big Data). IEEE; 2018. pp. 4844\u201353.","DOI":"10.1109\/BigData.2018.8622229"},{"key":"2544_CR5","doi-asserted-by":"publisher","first-page":"102233","DOI":"10.1016\/j.scs.2020.102233","volume":"60","author":"H Barham","year":"2020","unstructured":"Barham H, Daim T. The use of readiness assessment for big data projects. Sustain Cities Soc. 2020;60:102233.","journal-title":"Sustain Cities Soc"},{"key":"2544_CR6","doi-asserted-by":"publisher","first-page":"2","DOI":"10.1038\/s41746-021-00549-7","volume":"5","author":"AAH de Hond","year":"2022","unstructured":"de Hond AAH, Leeuwenberg AM, Hooft L, Kant IMJ, Nijman SWJ, van Os HJA, et al. Guidelines and quality criteria for artificial intelligence-based prediction models in healthcare: a scoping review. NPJ Digit Med. 2022;5:2.","journal-title":"NPJ Digit Med"},{"key":"2544_CR7","doi-asserted-by":"crossref","unstructured":"Castelijns LA, Maas Y, Vanschoren J. The abc of data: A classifying framework for data readiness. In: Machine Learning and Knowledge Discovery in Databases: International Workshops of ECML PKDD 2019, W\u00fcrzburg, Germany, September 16\u201320, 2019, Proceedings, Part I. Springer; 2020. pp. 3\u201316.","DOI":"10.1007\/978-3-030-43823-4_1"},{"key":"2544_CR8","unstructured":"Feurer M, Klein A, Eggensperger K, Springenberg J, Blum M, Hutter F. Efficient and Robust Automated Machine Learning. In Advances in neural information processing systems. 2015;28:2962\u20132970."},{"key":"2544_CR9","doi-asserted-by":"publisher","first-page":"86","DOI":"10.1145\/3458723","volume":"64","author":"T Gebru","year":"2021","unstructured":"Gebru T, Morgenstern J, Vecchione B, Vaughan JW, Wallach H, Iii HD, et al. Datasheets for datasets. Commun ACM. 2021;64:86\u201392.","journal-title":"Commun ACM"},{"key":"2544_CR10","doi-asserted-by":"publisher","first-page":"587","DOI":"10.1162\/tacl_a_00041","volume":"6","author":"EM Bender","year":"2018","unstructured":"Bender EM, Friedman B. Data statements for natural language processing: toward mitigating system bias and enabling better science. Trans Assoc Comput Linguist. 2018;6:587\u2013604.","journal-title":"Trans Assoc Comput Linguist"},{"issue":"4\/5","key":"2544_CR11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1147\/JRD.2019.2942288","volume":"63","author":"M Arnold","year":"2019","unstructured":"Arnold M, Bellamy RKE, Hind M, Houde S, Mehta S, Mojsilovi\u0107 A, et al. FactSheets: increasing trust in AI services through supplier\u2019s declarations of conformity. IBM J Res Dev. 2019;63(4\/5):1\u20136.","journal-title":"IBM J Res Dev"},{"key":"2544_CR12","doi-asserted-by":"crossref","unstructured":"Holland S, Hosny A, Newman S, Joseph J, Chmielinski K. The dataset nutrition label: A framework to drive higher data quality standards. arXiv preprint arXiv:180503677.\u00a0Hart Publishing. 2020;12(12):1.","DOI":"10.5040\/9781509932771.ch-001"},{"key":"2544_CR13","doi-asserted-by":"crossref","unstructured":"Mitchell M, Wu S, Zaldivar A, Barnes P, Vasserman L, Hutchinson B, et al. Model cards for model reporting. In: Proceedings of the conference on fairness, accountability, and transparency. 2019. pp. 220\u20139.","DOI":"10.1145\/3287560.3287596"},{"key":"2544_CR14","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18637\/jss.v090.i06","volume":"90","author":"AH Petersen","year":"2019","unstructured":"Petersen AH, Ekstr\u00f8m CT. dataMaid: your assistant for documenting supervised data quality screening in R. J Stat Softw. 2019;90:1\u201338.","journal-title":"J Stat Softw"},{"key":"2544_CR15","doi-asserted-by":"publisher","first-page":"169","DOI":"10.1177\/2515245919838783","volume":"2","author":"RC Arslan","year":"2019","unstructured":"Arslan RC. How to automatically document data with the codebook package to facilitate data reuse. Adv Methods Pract Psychol Sci. 2019;2:169\u201387.","journal-title":"Adv Methods Pract Psychol Sci"},{"key":"2544_CR16","doi-asserted-by":"crossref","unstructured":"Gupta N, Patel H, Afzal S, Panwar N, Mittal RS, Guttula S, et al. Data Quality Toolkit: automatic assessment of data quality and remediation for machine learning datasets. arXiv Preprint arXiv:210805935. 2021.","DOI":"10.1145\/3447548.3470817"},{"key":"2544_CR17","doi-asserted-by":"crossref","unstructured":"Afzal S, Rajmohan C, Kesarwani M, Mehta S, Patel H. Data Readiness Report. In: 2021 IEEE International Conference on Smart Data Services (SMDS). IEEE; 2021. pp. 42\u201351.","DOI":"10.1109\/SMDS53860.2021.00016"},{"key":"2544_CR18","doi-asserted-by":"publisher","first-page":"6039","DOI":"10.1038\/s41467-022-33128-9","volume":"13","author":"A Lavin","year":"2022","unstructured":"Lavin A, Gilligan-Lee CM, Visnjic A, Ganju S, Newman D, Ganguly S, et al. Technology readiness levels for machine learning systems. Nat Commun. 2022;13:6039.","journal-title":"Nat Commun"},{"key":"2544_CR19","doi-asserted-by":"crossref","unstructured":"Zhang A, Xing L, Zou J, Wu JC. Shifting machine learning for healthcare from development to deployment and from models to data. Nat Biomed Eng.\u00a0London: Nature Publishing Group; 2022;6(12):1330\u201345.","DOI":"10.1038\/s41551-022-00898-y"},{"key":"2544_CR20","unstructured":"ADNI Dataset. http:\/\/adni.loni.usc.edu. Accessed 28 May 2024."},{"key":"2544_CR21","unstructured":"FHS Dataset. https:\/\/www.framinghamheartstudy.org. Accessed 28\u00a0May  2024."},{"key":"2544_CR22","doi-asserted-by":"publisher","first-page":"861","DOI":"10.1117\/12.148698","volume-title":"Biomedical image processing and biomedical visualization","author":"WN Street","year":"1993","unstructured":"Street WN, Wolberg WH, Mangasarian OL. Nuclear feature extraction for breast tumor diagnosis. Biomedical image processing and biomedical visualization. SPIE; 1993. pp. 861\u201370."},{"key":"2544_CR23","doi-asserted-by":"publisher","first-page":"eaay5853","DOI":"10.1126\/sciadv.aay5853","volume":"6","author":"X-Y Xu","year":"2020","unstructured":"Xu X-Y, Huang X-L, Li Z-M, Gao J, Jiao Z-Q, Wang Y, et al. A scalable photonic computer solving the subset sum problem. Sci Adv. 2020;6:eaay5853.","journal-title":"Sci Adv"},{"key":"2544_CR24","doi-asserted-by":"publisher","first-page":"115","DOI":"10.1016\/j.compbiomed.2010.12.006","volume":"41","author":"S Oh","year":"2011","unstructured":"Oh S. A new dataset evaluation method based on category overlap. Comput Biol Med. 2011;41:115\u201322.","journal-title":"Comput Biol Med"}],"container-title":["BMC Medical Informatics and Decision Making"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12911-024-02544-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12911-024-02544-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12911-024-02544-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,4]],"date-time":"2024-06-04T04:31:24Z","timestamp":1717475484000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcmedinformdecismak.biomedcentral.com\/articles\/10.1186\/s12911-024-02544-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,4]]},"references-count":24,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["2544"],"URL":"https:\/\/doi.org\/10.1186\/s12911-024-02544-w","relation":{},"ISSN":["1472-6947"],"issn-type":[{"value":"1472-6947","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,4]]},"assertion":[{"value":"13 August 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 May 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 June 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"<i>Compliance with regulations:<\/i> We emphasize that the DREAMER framework is designed to align with key data protection regulations, ensuring that it does not violate patient confidentiality or privacy laws. DREAMER does not store or retain sensitive information and implements robust data anonymization and encryption techniques to protect personally identifiable information (PII).<i>Data handling and anonymization:<\/i> DREAMER employs strict data handling protocols to minimize exposure to sensitive data. All datasets used in the framework are anonymized before processing, removing direct identifiers to ensure patient privacy. Additionally, the framework includes secure data transmission methods to prevent unauthorized access during analysis.<i>Ethical guidelines:<\/i> We ensure that the DREAMER framework adheres to ethical guidelines for data usage in healthcare. This includes obtaining appropriate consents and ensuring that data usage aligns with the intended purpose without exploitation or misuse. Any case studies or real-world applications in the healthcare domain are conducted with ethical oversight to protect the rights and privacy of individuals.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"152"}}