{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T13:28:56Z","timestamp":1784208536333,"version":"3.55.0"},"reference-count":124,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2024,12,9]],"date-time":"2024-12-09T00:00:00Z","timestamp":1733702400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100012774","name":"Innovation Fund Denmark","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100012774","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2025,4,30]]},"abstract":"<jats:p>Sharing data with third parties is essential for advancing science, but it is becoming more and more difficult with the rise of data protection regulations, ethical restrictions, and growing fear of misuse. Fully synthetic data, which transcends anonymisation, may be the key to unlocking valuable untapped insights stored away in secured data vaults. This review examines current synthetic data generation methods and their utility measurement. We found that more traditional generative models such as Classification and Regression Tree models alongside Bayesian Networks remain highly relevant and are still capable of surpassing deep learning alternatives like Generative Adversarial Networks. However, our findings also display the same lack of agreement on metrics for evaluation, uncovered in earlier reviews, posing a persistent obstacle to advancing the field. We propose a tool for evaluating the utility of synthetic data and illustrate how it can be applied to three synthetic data generation models. By streamlining evaluation and promoting agreement on metrics, researchers can explore novel methods and generate compelling results that will convince data curators and lawmakers to embrace synthetic data. Our review emphasises the potential of synthetic data and highlights the need for greater collaboration and standardisation to unlock its full potential.<\/jats:p>","DOI":"10.1145\/3704437","type":"journal-article","created":{"date-parts":[[2024,11,14]],"date-time":"2024-11-14T10:04:39Z","timestamp":1731578679000},"page":"1-38","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":18,"title":["Systematic Review of Generative Modelling Tools and Utility Metrics for Fully Synthetic Tabular Data"],"prefix":"10.1145","volume":"57","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9228-2417","authenticated-orcid":false,"given":"Anton Danholt","family":"Lautrup","sequence":"first","affiliation":[{"name":"Department of Mathematics and Computer Science, University of Southern Denmark, Odense, Denmark"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4783-9893","authenticated-orcid":false,"given":"Tobias","family":"Hyrup","sequence":"additional","affiliation":[{"name":"Department of Mathematics and Computer Science, University of Southern Denmark, Odense, Denmark"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7713-4208","authenticated-orcid":false,"given":"Arthur","family":"Zimek","sequence":"additional","affiliation":[{"name":"Department of Mathematics and Computer Science, University of Southern Denmark, Odense, Denmark"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4000-5570","authenticated-orcid":false,"given":"Peter","family":"Schneider-Kamp","sequence":"additional","affiliation":[{"name":"Department of Mathematics and Computer Science, University of Southern Denmark, Odense, Denmark"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,12,9]]},"reference":[{"issue":"14","key":"e_1_3_2_2_2","doi-asserted-by":"crossref","first-page":"7075","DOI":"10.3390\/app12147075","article-title":"GAN-Based approaches for generating structured data in the medical domain","volume":"12","author":"Abedi Masoud","year":"2022","unstructured":"Masoud Abedi, Lars Hempel, Sina Sadeghi, and Toralf Kirsten. 2022. GAN-Based approaches for generating structured data in the medical domain. Applied Sciences 12, 14 (2022), 7075.","journal-title":"Applied Sciences"},{"issue":"4","key":"e_1_3_2_3_2","doi-asserted-by":"crossref","first-page":"212","DOI":"10.21307\/stattrans-2020-039","article-title":"Applying data synthesis for longitudinal business data across three countries","volume":"21","author":"Alam M. Jahangir","year":"2020","unstructured":"M. Jahangir Alam, Benoit Dostie, J\u00f6rg Drechsler, and Lars Vilhuber. 2020. Applying data synthesis for longitudinal business data across three countries. Statistics in Transition New Series 21, 4 (2020), 212\u2013236.","journal-title":"Statistics in Transition New Series"},{"key":"e_1_3_2_4_2","first-page":"73","volume-title":"ICCBD 2020: Proceedings of the 3rd International Conference on Computing and Big Data","author":"Alharbi Hanan Hammad","year":"2020","unstructured":"Hanan Hammad Alharbi and Masaomi Kimura. 2020. Missing data imputation using data generated by GAN. In ICCBD 2020: Proceedings of the 3rd International Conference on Computing and Big Data. ACM, 73\u201377."},{"key":"e_1_3_2_5_2","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1080\/00031305.1973.10478966","article-title":"Graphs in statistical analysis","volume":"27","author":"Anscombe Frank J.","year":"1973","unstructured":"Frank J. Anscombe. 1973. Graphs in statistical analysis. The American Statistician 27, 1 (1973), 17\u201321.","journal-title":"The American Statistician"},{"issue":"23","key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"12320","DOI":"10.3390\/app122312320","article-title":"Privacy and utility of private synthetic data for medical data analyses","volume":"12","author":"Appenzeller Arno","year":"2022","unstructured":"Arno Appenzeller, Moritz Leitner, Patrick Philipp, Erik Krempel, and J\u00fcrgen Beyerer. 2022. Privacy and utility of private synthetic data for medical data analyses. Applied Sciences 12, 23 (2022), 12320.","journal-title":"Applied Sciences"},{"key":"e_1_3_2_7_2","first-page":"214","volume-title":"Proceedings of the 34th International Conference on Machine Learning, ICML 2017","volume":"70","author":"Arjovsky Mart\u00edn","year":"2017","unstructured":"Mart\u00edn Arjovsky, Soumith Chintala, and L\u00e9on Bottou. 2017. Wasserstein generative adversarial networks. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017. Vol. 70, PMLR, 214\u2013223."},{"key":"e_1_3_2_8_2","first-page":"44:1\u201344:8","volume-title":"ICAIF \u201920: Proceedings of the 1st ACM International Conference on AI in Finance","author":"Assefa Samuel A.","year":"2020","unstructured":"Samuel A. Assefa, Danial Dervovic, Mahmoud Mahfouz, Robert E. Tillman, Prashant Reddy, and Manuela Veloso. 2020. Generating synthetic data in finance: Opportunities, challenges and pitfalls. In ICAIF \u201920: Proceedings of the 1st ACM International Conference on AI in Finance. ACM, New York, NY, USA, 44:1\u201344:8."},{"issue":"4","key":"e_1_3_2_9_2","doi-asserted-by":"crossref","first-page":"e043497","DOI":"10.1136\/bmjopen-2020-043497","article-title":"Can synthetic data be a proxy for real clinical trial data? A validation study","volume":"11","author":"Azizi Zahra","year":"2021","unstructured":"Zahra Azizi, Chaoyi Zheng, Lucy Mosquera, Louise Pilote, and Khaled El Emam. 2021. Can synthetic data be a proxy for real clinical trial data? A validation study. BMJ Open 11, 4 (2021), e043497.","journal-title":"BMJ Open"},{"issue":"1","key":"e_1_3_2_10_2","doi-asserted-by":"crossref","first-page":"190","DOI":"10.1016\/S0047-259X(03)00079-4","article-title":"On a new multivariate two-sample test","volume":"88","author":"Baringhaus Ludwig","year":"2004","unstructured":"Ludwig Baringhaus and Carsten Franz. 2004. On a new multivariate two-sample test. Journal of Multivariate Analysis 88, 1 (2004), 190\u2013206.","journal-title":"Journal of Multivariate Analysis"},{"issue":"9","key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"1165","DOI":"10.3390\/e23091165","article-title":"The problem of fairness in synthetic healthcare data","volume":"23","author":"Bhanot Karan","year":"2021","unstructured":"Karan Bhanot, Miao Qi, John S. Erickson, Isabelle Guyon, and Kristin P. Bennett. 2021. The problem of fairness in synthetic healthcare data. Entropy 23, 9 (2021), 1165.","journal-title":"Entropy"},{"issue":"3","key":"e_1_3_2_12_2","doi-asserted-by":"crossref","first-page":"71","DOI":"10.1007\/s00894-021-04674-8","article-title":"Generative chemistry: Drug discovery with deep learning generative models","volume":"27","author":"Bian Yuemin","year":"2021","unstructured":"Yuemin Bian and Xiang-Qun Xie. 2021. Generative chemistry: Drug discovery with deep learning generative models. Journal of Molecular Modeling 27, 3 (2021), 71.","journal-title":"Journal of Molecular Modeling"},{"issue":"1","key":"e_1_3_2_13_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1214\/aoms\/1177697800","article-title":"A distribution free version of the Smirnov two sample test in the p-variate case","volume":"40","author":"Bickel Peter J.","year":"1969","unstructured":"Peter J. Bickel. 1969. A distribution free version of the Smirnov two sample test in the p-variate case. The Annals of Mathematical Statistics 40, 1 (21969), 1\u201323.","journal-title":"The Annals of Mathematical Statistics"},{"issue":"11","key":"e_1_3_2_14_2","doi-asserted-by":"crossref","first-page":"7327","DOI":"10.1109\/TPAMI.2021.3116668","article-title":"Deep generative modelling: A comparative review of VAEs, GANs, normalizing flows, energy-based and autoregressive models","volume":"44","author":"Bond-Taylor Sam","year":"2022","unstructured":"Sam Bond-Taylor, Adam Leach, Yang Long, and Chris G. Willcocks. 2022. Deep generative modelling: A comparative review of VAEs, GANs, normalizing flows, energy-based and autoregressive models. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 11 (2022), 7327\u20137347.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_15_2","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1007\/978-3-030-57521-2_18","volume-title":"Proceedings of the International Conference on Privacy in Statistical Databases - UNESCO Chair in Data Privacy, PSD 2020.","volume":"12276","author":"Bowen Claire McKay","year":"2020","unstructured":"Claire McKay Bowen, Victoria Bryant, Leonard Burman, Surachai Khitatrakun, Robert McClelland, Philip Stallworth, Kyle Ueyama, and Aaron R. Williams. 2020. A synthetic supplemental public use file of low-income information return data: Methodology, utility, and privacy implications. In Proceedings of the International Conference on Privacy in Statistical Databases - UNESCO Chair in Data Privacy, PSD 2020.Lecture Notes in Computer Science, Vol. 12276, Springer, 257\u2013270."},{"issue":"1","key":"e_1_3_2_16_2","first-page":"32","article-title":"Comparative study of differentially private synthetic data algorithms from the NIST PSCR differential privacy synthetic data challenge","volume":"11","author":"Bowen Claire McKay","year":"2021","unstructured":"Claire McKay Bowen and Joshua Snoke. 2021. Comparative study of differentially private synthetic data algorithms from the NIST PSCR differential privacy synthetic data challenge. Journal of Privacy and Confidentiality 11, 1 (2021), 32 pages.","journal-title":"Journal of Privacy and Confidentiality"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","unstructured":"Amy Elise Braddon Suzanne Robinson Rosa Alati and Kim S. Betts. 2023. Exploring the utility of synthetic data to extract more value from sensitive health data assets: A focused example in perinatal epidemiology. Paediatric and Perinatal Epidemiology 37 4 (2023) 292--300.","DOI":"10.1111\/ppe.12942"},{"key":"e_1_3_2_18_2","unstructured":"Bauke Brenninkmeijer. 2021. Table Evaluator. GitHub code repository. Retrieved from https:\/\/github.com\/Baukebrenninkmeijer\/table-evaluator\/. Version 1.7.1. Accessed Nov. 2024."},{"issue":"1","key":"e_1_3_2_19_2","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1007\/s41781-021-00056-0","article-title":"Getting high: High fidelity simulation of high granularity calorimeters with high speed","volume":"5","author":"Buhmann Erik","year":"2021","unstructured":"Erik Buhmann, Sascha Diefenbacher, Engin Eren, Frank Gaede, Gregor Kasieczka, Anatolii Korol, and Katja Kr\u00fcger. 2021. Getting high: High fidelity simulation of high granularity calorimeters with high speed. Computing and Software for Big Science 5, 1 (2021), 13.","journal-title":"Computing and Software for Big Science"},{"issue":"1","key":"e_1_3_2_20_2","doi-asserted-by":"crossref","first-page":"777","DOI":"10.1093\/mnras\/stab294","article-title":"Survey2Survey: A deep learning generative model approach for cross-survey image mapping","volume":"503","author":"Buncher Brandon","year":"2021","unstructured":"Brandon Buncher, Awshesh N. Sharma, and Matias Carrasco-Kind. 2021. Survey2Survey: A deep learning generative model approach for cross-survey image mapping. Monthly Notices of the Royal Astronomical Society 503, 1 (2021), 777\u2013796.","journal-title":"Monthly Notices of the Royal Astronomical Society"},{"issue":"12","key":"e_1_3_2_21_2","doi-asserted-by":"crossref","first-page":"178","DOI":"10.3390\/data7120178","article-title":"Impacts of data synthesis: A metric for quantifiable data standards and performances","volume":"7","author":"Chandra Gunjan","year":"2022","unstructured":"Gunjan Chandra, Pekka Siirtola, Satu Tamminen, Mikael Knip, Riitta Veijola, and Juha R\u00f6ning. 2022. Impacts of data synthesis: A metric for quantifiable data standards and performances. Data 7, 12 (2022), 178.","journal-title":"Data"},{"key":"e_1_3_2_22_2","first-page":"26:1\u201326:6","volume-title":"BCB \u201920: Proceedings of the 11th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics","author":"Chen Junjie","year":"2020","unstructured":"Junjie Chen, Mohammad Erfan Mowlaei, and Xinghua Shi. 2020. Population-scale genomic data augmentation based on conditional generative adversarial networks. In BCB \u201920: Proceedings of the 11th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics. ACM, Virtual Event, USA, 26:1\u201326:6."},{"issue":"6","key":"e_1_3_2_23_2","doi-asserted-by":"crossref","first-page":"493","DOI":"10.1038\/s41551-021-00751-8","article-title":"Synthetic data in machine learning for medicine and healthcare","volume":"5","author":"Chen Richard J.","year":"2021","unstructured":"Richard J. Chen, Ming Y. Lu, Tiffany Y. Chen, Drew F. K. Williamson, and Faisal Mahmood. 2021. Synthetic data in machine learning for medicine and healthcare. Nature Biomedical Engineering 5, 6 (2021), 493\u2013497.","journal-title":"Nature Biomedical Engineering"},{"issue":"18","key":"e_1_3_2_24_2","doi-asserted-by":"crossref","first-page":"2938","DOI":"10.1093\/bioinformatics\/btx364","article-title":"UpSetR: An R package for the visualization of intersecting sets and their properties","volume":"33","author":"Conway Jake R.","year":"2017","unstructured":"Jake R. Conway, Alexander Lex, and Nils Gehlenborg. 2017. UpSetR: An R package for the visualization of intersecting sets and their properties. Bioinformatics 33, 18 (2017), 2938\u20132940.","journal-title":"Bioinformatics"},{"issue":"4","key":"e_1_3_2_25_2","doi-asserted-by":"crossref","first-page":"547","DOI":"10.1016\/j.dss.2009.05.016","article-title":"Modeling wine preferences by data mining from physicochemical properties","volume":"47","author":"Cortez Paulo","year":"2009","unstructured":"Paulo Cortez, Ant\u00f3nio Cerdeira, Fernando Almeida, Telmo Matos, and Jos\u00e9 Reis. 2009. Modeling wine preferences by data mining from physicochemical properties. Decision Support Systems 47, 4 (2009), 547\u2013553.","journal-title":"Decision Support Systems"},{"key":"e_1_3_2_26_2","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"crossref","first-page":"282","DOI":"10.1007\/978-3-030-88942-5_22","volume-title":"Proceedings of the 24th International Conference on Discovery Science, DS 2021.","volume":"12986","author":"Coutinho-Almeida Jo\u00e3o","year":"2021","unstructured":"Jo\u00e3o Coutinho-Almeida, Pedro Pereira Rodrigues, and Ricardo Jo\u00e3o Cruz Correia. 2021. GANs for tabular healthcare data generation: A review on utility and privacy. In Proceedings of the 24th International Conference on Discovery Science, DS 2021.Lecture Notes in Computer Science, Vol. 12986, Springer,, 282\u2013291."},{"key":"e_1_3_2_27_2","volume-title":"Elements of Information Theory (2nd ed.)","author":"Cover Thomas M.","year":"2006","unstructured":"Thomas M. Cover and Joy A. Thomas. 2006. Elements of Information Theory (2nd ed.). John Wiley & Sons, New York, NY, USA."},{"issue":"5","key":"e_1_3_2_28_2","doi-asserted-by":"crossref","first-page":"2158","DOI":"10.3390\/app11052158","article-title":"Fake it till you make it: Guidelines for effective synthetic data generation","volume":"11","author":"Dankar Fida K.","year":"2021","unstructured":"Fida K. Dankar and Mahmoud Ibrahim. 2021. Fake it till you make it: Guidelines for effective synthetic data generation. Applied Sciences 11, 5 (Feb.2021), 2158.","journal-title":"Applied Sciences"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/access.2022.3144765"},{"key":"e_1_3_2_30_2","unstructured":"DataCebo Inc. 2023. Synthetic Data Metrics. DataCebo Inc. Retrieved from https:\/\/docs.sdv.dev\/sdmetrics\/. Version 0.17.0. Accessed Nov. 2024."},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","first-page":"6","DOI":"10.1145\/3411170.3411243","volume-title":"GoodTechs \u201920: Proceedings of the 6th EAI International Conference on Smart Objects and Technologies for Social Good","author":"Deeva Irina","year":"2020","unstructured":"Irina Deeva, Petr D. Andriushchenko, Anna V. Kalyuzhnaya, and Alexander V. Boukhanovsky. 2020. Bayesian networks-based personal data synthesis. In GoodTechs \u201920: Proceedings of the 6th EAI International Conference on Smart Objects and Technologies for Social Good. ACM, 6\u201311."},{"issue":"3","key":"e_1_3_2_32_2","doi-asserted-by":"crossref","first-page":"523","DOI":"10.1093\/jssam\/smaa035","article-title":"Synthesizing geocodes to facilitate access to detailed geographical information in large-scale administrative data","volume":"9","author":"Drechsler J\u00f6rg","year":"2020","unstructured":"J\u00f6rg Drechsler and Jingchen Hu. 2020. Synthesizing geocodes to facilitate access to detailed geographical information in large-scale administrative data. Journal of Survey Statistics and Methodology 9, 3 (Dec.2020), 523\u2013548.","journal-title":"Journal of Survey Statistics and Methodology"},{"issue":"1","key":"e_1_3_2_33_2","doi-asserted-by":"crossref","first-page":"88","DOI":"10.3390\/e25010088","article-title":"HT-Fed-GAN: Federated generative model for decentralized tabular data synthesis","volume":"25","author":"Duan Shaoming","year":"2023","unstructured":"Shaoming Duan, Chuanyi Liu, Peiyi Han, Xiaopeng Jin, Xinyi Zhang, Tianyu He, Hezhong Pan, and Xiayu Xiang. 2023. HT-Fed-GAN: Federated generative model for decentralized tabular data synthesis. Entropy 25, 1 (2023), 88.","journal-title":"Entropy"},{"issue":"3","key":"e_1_3_2_34_2","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1561\/0400000042","article-title":"The algorithmic foundations of differential privacy","volume":"9","author":"Dwork Cynthia","year":"2013","unstructured":"Cynthia Dwork and Aaron Roth. 2013. The algorithmic foundations of differential privacy. Foundations and Trends\u00ae in Theoretical Computer Science 9, 3\u20134 (2013), 211\u2013487.","journal-title":"Foundations and Trends\u00ae in Theoretical Computer Science"},{"key":"e_1_3_2_35_2","doi-asserted-by":"crossref","first-page":"e2300071","DOI":"10.1200\/CCI.23.00071","article-title":"Status of synthetic data generation for structured health data","volume":"7","author":"Emam Khaled El","year":"2023","unstructured":"Khaled El Emam. 2023. Status of synthetic data generation for structured health data. JCO Clinical Cancer Informatics 7, 7 (62023), e2300071.","journal-title":"JCO Clinical Cancer Informatics"},{"issue":"11","key":"e_1_3_2_36_2","doi-asserted-by":"crossref","first-page":"e23139","DOI":"10.2196\/23139","article-title":"Evaluating identity disclosure risk in fully synthetic health data: Model development and validation","volume":"22","author":"Emam Khaled El","year":"2020","unstructured":"Khaled El Emam, Lucy Mosquera, and Jason Bass. 2020. Evaluating identity disclosure risk in fully synthetic health data: Model development and validation. Journal of Medical Internet Research 22, 11 (2020), e23139.","journal-title":"Journal of Medical Internet Research"},{"issue":"4","key":"e_1_3_2_37_2","doi-asserted-by":"crossref","first-page":"e35734","DOI":"10.2196\/35734","article-title":"Utility metrics for evaluating synthetic health data generation methods: Validation study","volume":"10","author":"Emam Khaled El","year":"2022","unstructured":"Khaled El Emam, Lucy Mosquera, Xi Fang, and Alaa El-Hussuna. 2022. Utility metrics for evaluating synthetic health data generation methods: Validation study. JMIR Medical Informatics 10, 4 (2022), e35734.","journal-title":"JMIR Medical Informatics"},{"issue":"1","key":"e_1_3_2_38_2","doi-asserted-by":"crossref","first-page":"ooab012","DOI":"10.1093\/jamiaopen\/ooab012","article-title":"Evaluating the utility of synthetic COVID-19 case data","volume":"4","author":"Emam Khaled El","year":"2021","unstructured":"Khaled El Emam, Lucy Mosquera, Elizabeth Jonker, and Harpreet Sood. 2021. Evaluating the utility of synthetic COVID-19 case data. JAMIA Open 4, 1 (2021), ooab012.","journal-title":"JAMIA Open"},{"issue":"1","key":"e_1_3_2_39_2","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1093\/jamia\/ocaa249","article-title":"Optimizing the synthesis of clinical trial data using sequential trees","volume":"28","author":"Emam Khaled El","year":"2021","unstructured":"Khaled El Emam, Lucy Mosquera, and Chaoyi Zheng. 2021. Optimizing the synthesis of clinical trial data using sequential trees. Journal of the American Medical Informatics Association 28, 1 (2021), 3\u201313.","journal-title":"Journal of the American Medical Informatics Association"},{"key":"e_1_3_2_40_2","doi-asserted-by":"crossref","first-page":"94","DOI":"10.1145\/3548785.3548793","volume-title":"IDEAS\u201922: Proceedings of the International Database Engineered Applications Symposium","author":"Endres Markus","year":"2022","unstructured":"Markus Endres, Asha Mannarapotta Venugopal, and Tung Son Tran. 2022. Synthetic data generation: A comparative study. In IDEAS\u201922: Proceedings of the International Database Engineered Applications Symposium. ACM, 94\u2013102."},{"issue":"11","key":"e_1_3_2_41_2","doi-asserted-by":"crossref","first-page":"1962","DOI":"10.14778\/3407790.3407802","article-title":"Relational data synthesis using generative adversarial networks: A design space exploration","volume":"13","author":"Fan Ju","year":"2020","unstructured":"Ju Fan, Tongyu Liu, Guoliang Li, Junyou Chen, Yuwei Shen, and Xiaoyong Du. 2020. Relational data synthesis using generative adversarial networks: A design space exploration. Proceedings of the VLDB Endowment 13, 11 (2020), 1962\u20131975.","journal-title":"Proceedings of the VLDB Endowment"},{"issue":"15","key":"e_1_3_2_42_2","doi-asserted-by":"crossref","first-page":"2733","DOI":"10.3390\/math10152733","article-title":"Survey on synthetic data generation, evaluation methods and GANs","volume":"10","author":"Figueira Alvaro","year":"2022","unstructured":"Alvaro Figueira and Bruno Vaz. 2022. Survey on synthetic data generation, evaluation methods and GANs. Mathematics 10, 15 (82022), 2733.","journal-title":"Mathematics"},{"issue":"3","key":"e_1_3_2_43_2","doi-asserted-by":"crossref","first-page":"e2019MS001896","DOI":"10.1029\/2019MS001896","article-title":"Machine learning for stochastic parameterization: Generative adversarial networks in the Lorenz \u201996 model","volume":"12","author":"Gagne David J.","year":"2020","unstructured":"David J. Gagne, Hannah M. Christensen, Aneesh C. Subramanian, and Adam H. Monahan. 2020. Machine learning for stochastic parameterization: Generative adversarial networks in the Lorenz \u201996 model. Journal of Advances in Modeling Earth Systems 12, 3 (2020), e2019MS001896.","journal-title":"Journal of Advances in Modeling Earth Systems"},{"key":"e_1_3_2_44_2","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1007\/978-3-030-62365-4_3","volume-title":"IDEAL 2020 - Proceedings of the 21st International Conference on Intelligent Data Engineering and Automated Learning.","volume":"12490","author":"Galloni Andrea","year":"2020","unstructured":"Andrea Galloni, Imre Lend\u00e1k, and Tom\u00e1s Horv\u00e1th. 2020. A novel evaluation metric for synthetic data generation. In IDEAL 2020 - Proceedings of the 21st International Conference on Intelligent Data Engineering and Automated Learning.Lecture Notes in Computer Science, Vol. 12490, Springer, 25\u201334."},{"issue":"10","key":"e_1_3_2_45_2","doi-asserted-by":"crossref","first-page":"1886","DOI":"10.14778\/3467861.3467876","article-title":"Kamino: Constraint-aware differentially private data synthesis","volume":"14","author":"Ge Chang","year":"2021","unstructured":"Chang Ge, Shubhankar Mohapatra, Xi He, and Ihab F. Ilyas. 2021. Kamino: Constraint-aware differentially private data synthesis. Proceedings of the VLDB Endowment 14, 10 (2021), 1886\u20131899.","journal-title":"Proceedings of the VLDB Endowment"},{"key":"e_1_3_2_46_2","unstructured":"Ian J. Goodfellow. 2016. NIPS 2016 Tutorial: Generative adversarial networks. arXiv:1701.00160. Retrieved from https:\/\/arxiv.org\/abs\/1701.00160"},{"key":"e_1_3_2_47_2","doi-asserted-by":"crossref","unstructured":"Ian J. Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2020. Generative adversarial networks. Commun. ACM 63 11 (2020) 139--144.","DOI":"10.1145\/3422622"},{"issue":"2","key":"e_1_3_2_48_2","first-page":"192","article-title":"Data privacy protection and utility preservation through Bayesian data synthesis: A case study on Airbnb listings","volume":"77","author":"Guo Shijie","year":"2022","unstructured":"Shijie Guo and Jingchen Hu. 2022. Data privacy protection and utility preservation through Bayesian data synthesis: A case study on Airbnb listings. The American Statistician 77, 2 (2022), 192\u2013200.","journal-title":"The American Statistician"},{"issue":"2","key":"e_1_3_2_49_2","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1109\/MIS.2009.36","article-title":"The unreasonable effectiveness of data","volume":"24","author":"Halevy Alon","year":"2009","unstructured":"Alon Halevy, Peter Norvig, and Fernando Pereira. 2009. The unreasonable effectiveness of data. IEEE Intelligent Systems 24, 2 (2009), 8\u201312.","journal-title":"IEEE Intelligent Systems"},{"key":"e_1_3_2_50_2","series-title":"Proceedings of Machine Learning Research","first-page":"1819","volume-title":"Proceedings of the 24th International Conference on Artificial Intelligence and Statistics, AISTATS 2021.","volume":"130","author":"Harder Frederik","year":"2021","unstructured":"Frederik Harder, Kamil Adamczewski, and Mijung Park. 2021. DP-MERF: Differentially private mean embeddings with randomfeatures for practical privacy-preserving data generation. In Proceedings of the 24th International Conference on Artificial Intelligence and Statistics, AISTATS 2021.Proceedings of Machine Learning Research, Vol. 130, PMLR, San Diego, California, USA, 1819\u20131827."},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.04.053"},{"issue":"01","key":"e_1_3_2_52_2","first-page":"e19\u2013e38","article-title":"Synthetic tabular data evaluation in the health domain covering resemblance, utility, and privacy dimensions","volume":"62","author":"Hernandez Mikel","year":"2023","unstructured":"Mikel Hernandez, Gorka Epelde, Ane Alberdi, Rodrigo Cilla, and Debbie Rankin. 2023. Synthetic tabular data evaluation in the health domain covering resemblance, utility, and privacy dimensions. Methods of Information in Medicine 62, S 01 (2023), e19\u2013e38.","journal-title":"Methods of Information in Medicine"},{"key":"e_1_3_2_53_2","unstructured":"Martin Heusel Hubert Ramsauer Thomas Unterthiner Bernhard Nessler and Sepp Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local nash equilibrium. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS'17). Curran Associates Inc. Red Hook NY USA 6629--6640."},{"key":"e_1_3_2_54_2","series-title":"Lecture Notes in Computer Science","first-page":"418","volume-title":"Proceedings of the 19th International Conference on Algorithms and Architectures for Parallel Processing, ICA3PP 2019.","volume":"11945","author":"Ho Stella","year":"2019","unstructured":"Stella Ho, Youyang Qu, Longxiang Gao, Jianxin Li, and Yong Xiang. 2019. Generative adversarial nets enhanced continual data release using differential privacy. In Proceedings of the 19th International Conference on Algorithms and Architectures for Parallel Processing, ICA3PP 2019.Lecture Notes in Computer Science, Vol. 11945, Springer, 418\u2013426."},{"key":"e_1_3_2_55_2","first-page":"28:1\u201328:6","volume-title":"ARES 2020: Proceedings of the 15th International Conference on Availability, Reliability and Security","author":"Holmes Michael","year":"2020","unstructured":"Michael Holmes and George Theodorakopoulos. 2020. Towards using differentially private synthetic data for machine learning in collaborative data science projects. In ARES 2020: Proceedings of the 15th International Conference on Availability, Reliability and Security. ACM, 28:1\u201328:6."},{"issue":"1","key":"e_1_3_2_56_2","first-page":"37","article-title":"Identification risks evaluation of partially synthetic data with the identificationriskcalculation R package","volume":"14","author":"Hornby Ryan","year":"2021","unstructured":"Ryan Hornby and Jingchen Hu. 2021. Identification risks evaluation of partially synthetic data with the identificationriskcalculation R package. Transactions on Data Privacy 14, 1 (2021), 37\u201352.","journal-title":"Transactions on Data Privacy"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","unstructured":"Allison Marie Horst Alison Presmanes Hill and Kristen B. Gorman. 2020. palmerpenguins: Palmer Archipelago Antarctica) penguin data. DOI:10.5281\/zenodo.3960218. R package version 0.1.0. Accessed Nov. 2024.","DOI":"10.5281\/zenodo.3960218"},{"issue":"3","key":"e_1_3_2_58_2","doi-asserted-by":"crossref","first-page":"651","DOI":"10.1198\/106186006X133933","article-title":"Unbiased recursive partitioning: A conditional inference framework","volume":"15","author":"Hothorn Torsten","year":"2006","unstructured":"Torsten Hothorn, Kurt Hornik, and Achim Zeileis. 2006. Unbiased recursive partitioning: A conditional inference framework. Journal of Computational and Graphical Statistics 15, 3 (2006), 651\u2013674.","journal-title":"Journal of Computational and Graphical Statistics"},{"key":"e_1_3_2_59_2","unstructured":"Bill Howe Julia Stoyanovich Haoyue Ping Bernease Herman and Matt Gee. 2017. Synthetic data for social good. arXiv:1710.08874. Retrieved from https:\/\/arxiv.org\/abs\/1710.08874"},{"issue":"5","key":"e_1_3_2_60_2","first-page":"1370","article-title":"Risk-efficient Bayesian data synthesis for privacy protection","volume":"10","author":"Hu Jingchen","year":"2021","unstructured":"Jingchen Hu, Terrance D. Savitsky, and Matthew R. Williams. 2021. Risk-efficient Bayesian data synthesis for privacy protection. Journal of Survey Statistics and Methodology 10, 5 (2021), 1370\u20131399.","journal-title":"Journal of Survey Statistics and Methodology"},{"issue":"65","key":"e_1_3_2_61_2","doi-asserted-by":"crossref","first-page":"3634","DOI":"10.21105\/joss.03634","article-title":"latentcor: An R package for estimating latent correlations from mixed data types","volume":"6","author":"Huang Mingze","year":"2021","unstructured":"Mingze Huang, Christian L. M\u00fcller, and Irina Gaynanova. 2021. latentcor: An R package for estimating latent correlations from mixed data types. Journal of Open Source Software 6, 65 (2021), 3634.","journal-title":"Journal of Open Source Software"},{"key":"e_1_3_2_62_2","doi-asserted-by":"crossref","unstructured":"Tobias Hyrup Anton Danholt Lautrup Arthur Zimek and Peter Schneider-Kamp. 2023. Sharing is CAIRing: Characterizing principles and assessing properties of universal privacy evaluation for synthetic tabular data. arXiv:2312.12216. Retrieved from https:\/\/arxiv.org\/abs\/2312.12216","DOI":"10.1016\/j.mlwa.2024.100608"},{"issue":"4","key":"e_1_3_2_63_2","doi-asserted-by":"crossref","first-page":"1613","DOI":"10.1111\/rssa.12876","article-title":"Using saturated count models for user-friendly synthesis of large confidential administrative databases","volume":"185","author":"Jackson James","year":"2022","unstructured":"James Jackson, Robin Mitra, Brian Francis, and Iain Dove. 2022. Using saturated count models for user-friendly synthesis of large confidential administrative databases. Journal of the Royal Statistical Society Series A: Statistics in Society 185, 4 (2022), 1613\u20131643.","journal-title":"Journal of the Royal Statistical Society Series A: Statistics in Society"},{"issue":"10","key":"e_1_3_2_64_2","doi-asserted-by":"crossref","first-page":"2318","DOI":"10.1109\/TKDE.2017.2720168","article-title":"Theory-guided data science: A new paradigm for scientific discovery from data","volume":"29","author":"Karpatne Anuj","year":"2017","unstructured":"Anuj Karpatne, Gowtham Atluri, James H. Faghmous, Michael Steinbach, Arindam Banerjee, Auroop Ganguly, Shashi Shekhar, Nagiza Samatova, and Vipin Kumar. 2017. Theory-guided data science: A new paradigm for scientific discovery from data. IEEE Transactions on Knowledge and Data Engineering 29, 10 (2017), 2318\u20132331.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"issue":"3","key":"e_1_3_2_65_2","doi-asserted-by":"crossref","first-page":"224","DOI":"10.1198\/000313006X124640","article-title":"A framework for evaluating the utility of data altered to protect confidentiality","volume":"60","author":"Karr Alan F.","year":"2006","unstructured":"Alan F. Karr, C. N. Kohnen, A. Oganian, J. P. Reiter, and A. P. Sanil. 2006. A framework for evaluating the utility of data altered to protect confidentiality. The American Statistician 60, 3 (2006), 224\u2013232.","journal-title":"The American Statistician"},{"key":"e_1_3_2_66_2","first-page":"8107","volume-title":"Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020","author":"Karras Tero","year":"2020","unstructured":"Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020. Analyzing and improving the image quality of StyleGAN. In Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020. Computer Vision Foundation\/IEEE, Seattle, WA, USA, 8107\u20138116."},{"issue":"4","key":"e_1_3_2_67_2","doi-asserted-by":"crossref","first-page":"801","DOI":"10.1093\/jamia\/ocaa303","article-title":"Application of Bayesian networks to generate synthetic health data","volume":"28","author":"Kaur Dhamanpreet","year":"2021","unstructured":"Dhamanpreet Kaur, Matthew Sobiesk, Shubham Patil, Jin Liu, Puran Bhagat, Amar Gupta, and Natasha Markuzon. 2021. Application of Bayesian networks to generate synthetic health data. Journal of the American Medical Informatics Association 28, 4 (2021), 801\u2013811.","journal-title":"Journal of the American Medical Informatics Association"},{"issue":"3","key":"e_1_3_2_68_2","doi-asserted-by":"crossref","first-page":"118","DOI":"10.1177\/014107680309600304","article-title":"Five steps to conducting a systematic review","volume":"96","author":"Khan Khalid S.","year":"2003","unstructured":"Khalid S. Khan, Regina Kunz, Jos Kleijnen, and Gerd Antes. 2003. Five steps to conducting a systematic review. Journal of the Royal Society of Medicine 96, 3 (2003), 118\u2013121.","journal-title":"Journal of the Royal Society of Medicine"},{"key":"e_1_3_2_69_2","first-page":"3581","volume-title":"Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2","author":"Kingma Diederik P.","year":"2014","unstructured":"Diederik P. Kingma, Shakir Mohamed, Danilo Jimenez Rezende, and Max Welling. 2014. Semi-supervised learning with deep generative models. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2. MIT Press, Cambridge, MA, USA, 3581\u20133589."},{"issue":"1","key":"e_1_3_2_70_2","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1016\/j.infsof.2008.09.009","article-title":"Systematic literature reviews in software engineering \u2013 A systematic literature review","volume":"51","author":"Kitchenham Barbara","year":"2009","unstructured":"Barbara Kitchenham, O. Pearl Brereton, David Budgen, Mark Turner, John Bailey, and Stephen Linkman. 2009. Systematic literature reviews in software engineering \u2013 A systematic literature review. Information and Software Technology 51, 1 (2009), 7\u201315.","journal-title":"Information and Software Technology"},{"key":"e_1_3_2_71_2","series-title":"Proceedings of Machine Learning Research","first-page":"17564","volume-title":"Proceedings of the International Conference on Machine Learning, ICML 2023.","volume":"202","author":"Kotelnikov Akim","year":"2023","unstructured":"Akim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, and Artem Babenko. 2023. TabDDPM: Modelling tabular data with diffusion models. In Proceedings of the International Conference on Machine Learning, ICML 2023.Proceedings of Machine Learning Research, Vol. 202, PMLR, 17564\u201317579."},{"issue":"2","key":"e_1_3_2_72_2","doi-asserted-by":"crossref","first-page":"341","DOI":"10.1007\/s10115-016-1004-2","article-title":"The (black) art of runtime evaluation: Are we comparing algorithms or implementations?","volume":"52","author":"Kriegel Hans-Peter","year":"2017","unstructured":"Hans-Peter Kriegel, Erich Schubert, and Arthur Zimek. 2017. The (black) art of runtime evaluation: Are we comparing algorithms or implementations? Knowledge and Information Systems 52, 2 (2017), 341\u2013378.","journal-title":"Knowledge and Information Systems"},{"issue":"2","key":"e_1_3_2_73_2","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1017\/S026988890200019X","article-title":"A review of explanation methods for Bayesian networks","volume":"17","author":"Lacave Carmen","year":"2002","unstructured":"Carmen Lacave and Francisco J. D\u00edez. 2002. A review of explanation methods for Bayesian networks. The Knowledge Engineering Review 17, 2 (2002), 107\u2013127.","journal-title":"The Knowledge Engineering Review"},{"key":"e_1_3_2_74_2","doi-asserted-by":"crossref","unstructured":"Anton D. Lautrup Tobias Hyrup Arthur Zimek and Peter Schneider-Kamp. 2024. SynthEval: A framework for detailed utility and privacy evaluation of tabular synthetic data. arXiv:2404.15821. Retrieved from https:\/\/arxiv.org\/abs\/2404.15821 Code available on GitHub v1.4.1.","DOI":"10.1007\/s10618-024-01081-4"},{"key":"e_1_3_2_75_2","doi-asserted-by":"crossref","unstructured":"Marta Lenatti Alessia Paglialonga Vanessa Orani Melissa Ferretti and Maurizio Mongelli. 2023. Characterization of synthetic health data using rule-based artificial intelligence models. IEEE Journal of Biomedical and Health Informatics 27 8 (2023) 1--9.","DOI":"10.1109\/JBHI.2023.3236722"},{"key":"e_1_3_2_76_2","doi-asserted-by":"crossref","unstructured":"Stefan Lenz Moritz Hess and Harald Binder. 2021. Deep generative models in DataSHIELD. BMC Medical Research Methodology 21 1 (2021) 16 pages.","DOI":"10.1186\/s12874-021-01237-6"},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2022.110239"},{"key":"e_1_3_2_78_2","first-page":"22","volume-title":"Proceedings of the 11th International Conference on Learning Representations, ICLR 2023","author":"Liu Tennison","year":"2023","unstructured":"Tennison Liu, Zhaozhi Qian, Jeroen Berrevoets, and Mihaela van der Schaar. 2023. GOGGLE: Generative modelling for tabular data by learning relational structure. In Proceedings of the 11th International Conference on Learning Representations, ICLR 2023. OpenReview.net, 22 pages."},{"key":"e_1_3_2_79_2","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"crossref","first-page":"306","DOI":"10.1007\/978-3-031-14463-9_20","volume-title":"Proceedings of the 6th IFIP TC 5, TC 12, WG 8.4, WG 8.9, WG 12.9 International Cross-Domain Conference on Machine Learning and Knowledge Extraction, CD-MAKE 2022.","volume":"13480","author":"Llugiqi Majlinda","year":"2022","unstructured":"Majlinda Llugiqi and Rudolf Mayer. 2022. An empirical analysis of synthetic-data-based anomaly detection. In Proceedings of the 6th IFIP TC 5, TC 12, WG 8.4, WG 8.9, WG 12.9 International Cross-Domain Conference on Machine Learning and Knowledge Extraction, CD-MAKE 2022.Lecture Notes in Computer Science, Vol. 13480, Springer, 306\u2013327."},{"key":"e_1_3_2_80_2","unstructured":"Tshilidzi Marwala Eleonore Fournier-Tombs and Serge Stinckwich. 2023. The use of synthetic data to train AI models: Opportunities and risks for sustainable development. arXiv:2309.00652. Retrieved from https:\/\/arxiv.org\/abs\/2309.00652"},{"key":"e_1_3_2_81_2","doi-asserted-by":"crossref","first-page":"1290","DOI":"10.1145\/3025453.3025912","volume-title":"Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems","author":"Matejka Justin","year":"2017","unstructured":"Justin Matejka and George W. Fitzmaurice. 2017. Same stats, different graphs: Generating datasets with varied appearance and identical statistics through simulated annealing. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems. ACM, 1290\u20131294."},{"key":"e_1_3_2_82_2","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1007\/978-3-030-49669-2_11","volume-title":"Proceedings of the 34th Annual IFIP WG 11.3 Conference on Data and Applications Security and Privacy, DBSec 2020.","volume":"12122","author":"Mayer Rudolf","year":"2020","unstructured":"Rudolf Mayer, Markus Hittmeir, and Andreas Ekelhart. 2020. Privacy-preserving anomaly detection using synthetic data. In Proceedings of the 34th Annual IFIP WG 11.3 Conference on Data and Applications Security and Privacy, DBSec 2020.Lecture Notes in Computer Science, Vol. 12122, Springer, 195\u2013207."},{"key":"e_1_3_2_83_2","unstructured":"Daniel McDuff Theodore Curran and Achuta Kadambi. 2023. Synthetic data in healthcare. arXiv:2304.03243. Retrieved from https:\/\/arxiv.org\/abs\/2304.03243"},{"key":"e_1_3_2_84_2","unstructured":"Mehdi Mirza and Simon Osindero. 2014. Conditional generative adversarial nets. arXiv:1411.1784. Retrieved from https:\/\/arxiv.org\/abs\/1411.1784"},{"key":"e_1_3_2_85_2","first-page":"1","volume-title":"Proceedings of the Winter Simulation Conference, WSC 2021","author":"Montevechi Jos\u00e9 Arnaldo Barra","year":"2021","unstructured":"Jos\u00e9 Arnaldo Barra Montevechi, Afonso Teberga Campos, Gustavo Teodoro Gabriel, and Carlos Henrique dos Santos. 2021. Input data modeling: An approach using generative adversarial networks. In Proceedings of the Winter Simulation Conference, WSC 2021. IEEE, Phoenix, AZ, USA, 1\u201312."},{"key":"e_1_3_2_86_2","doi-asserted-by":"crossref","first-page":"2772","DOI":"10.1109\/WSC57314.2022.10015375","volume-title":"Proceedings of the Winter Simulation Conference, WSC 2022","author":"Montevechi Jos\u00e9 Arnaldo Barra","year":"2022","unstructured":"Jos\u00e9 Arnaldo Barra Montevechi, Gustavo Teodoro Gabriel, Afonso Teberga Campos, Carlos Henrique dos Santos, Fabiano Leal, and Michael E. F. H. S. Machado. 2022. Using generative adversarial networks to validate discrete event simulation models. In Proceedings of the Winter Simulation Conference, WSC 2022. IEEE, 2772\u20132783."},{"key":"e_1_3_2_87_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cosrev.2023.100546"},{"issue":"11","key":"e_1_3_2_88_2","first-page":"1","article-title":"Synthpop: Bespoke creation of synthetic data in R","volume":"74","author":"Nowok Beata","year":"2016","unstructured":"Beata Nowok, Gillian M. Raab, and Chris Dibben. 2016. Synthpop: Bespoke creation of synthetic data in R. Journal of Statistical Software 74, 11 (2016), 1\u201326.","journal-title":"Journal of Statistical Software"},{"issue":"3","key":"e_1_3_2_89_2","doi-asserted-by":"crossref","first-page":"1126","DOI":"10.3390\/app11031126","article-title":"Synthesizing individual consumers\u2019 credit historical data using generative adversarial networks","volume":"11","author":"Park Nari","year":"2021","unstructured":"Nari Park, Yeong Hyeon Gu, and Seong Joon Yoo. 2021. Synthesizing individual consumers\u2019 credit historical data using generative adversarial networks. Applied Sciences 11, 3 (2021), 1126.","journal-title":"Applied Sciences"},{"key":"e_1_3_2_90_2","doi-asserted-by":"crossref","first-page":"108","DOI":"10.1109\/OJEMB.2022.3181796","article-title":"Bayesian inference-based Gaussian mixture models with optimal components estimation towards large-scale synthetic data generation for in silico clinical trials","volume":"3","author":"Pezoulas Vasileios C.","year":"2022","unstructured":"Vasileios C. Pezoulas, Nikolaos S. Tachos, George Gkois, Iacopo Olivotto, Fausto Barlocco, and Dimitrios I. Fotiadis. 2022. Bayesian inference-based Gaussian mixture models with optimal components estimation towards large-scale synthetic data generation for in silico clinical trials. IEEE Open Journal of Engineering in Medicine and Biology 3 (2022), 108\u2013114.","journal-title":"IEEE Open Journal of Engineering in Medicine and Biology"},{"key":"e_1_3_2_91_2","first-page":"42:1\u201342:5","volume-title":"Proceedings of the 29th International Conference on Scientific and Statistical Database Management","author":"Ping Haoyue","year":"2017","unstructured":"Haoyue Ping, Julia Stoyanovich, and Bill Howe. 2017. DataSynthesizer: Privacy-preserving synthetic datasets. In Proceedings of the 29th International Conference on Scientific and Statistical Database Management. ACM, 42:1\u201342:5."},{"key":"e_1_3_2_92_2","unstructured":"Zhaozhi Qian Bogdan-Constantin Cebere and Mihaela van der Schaar. 2023. Synthcity: Facilitating innovative use cases of synthetic data in different data modalities. arXiv:2301.07573. Retrieved from https:\/\/arxiv.org\/abs\/2301.07573"},{"issue":"3","key":"e_1_3_2_93_2","doi-asserted-by":"crossref","first-page":"596","DOI":"10.1093\/jssam\/smac007","article-title":"Improving the utility of poisson-distributed, differentially private synthetic data via prior predictive truncation with an application to CDC wonder","volume":"10","author":"Quick Harrison","year":"2022","unstructured":"Harrison Quick. 2022. Improving the utility of poisson-distributed, differentially private synthetic data via prior predictive truncation with an application to CDC wonder. Journal of Survey Statistics and Methodology 10, 3 (2022), 596\u2013617.","journal-title":"Journal of Survey Statistics and Methodology"},{"key":"e_1_3_2_94_2","unstructured":"Gillian M. Raab Beata Nowok and Chris Dibben. 2017. Guidelines for producing useful synthetic data. arXiv:1712.04078. Retrieved from https:\/\/arxiv.org\/abs\/1712.04078"},{"issue":"1","key":"e_1_3_2_95_2","doi-asserted-by":"crossref","first-page":"129","DOI":"10.1146\/annurev-statistics-040720-031848","article-title":"Synthetic Data","volume":"8","author":"Raghunathan Trivellore E.","year":"2021","unstructured":"Trivellore E. Raghunathan. 2021. Synthetic Data. Annual Review of Statistics and Its Application 8, 1 (2021), 129\u2013140.","journal-title":"Annual Review of Statistics and Its Application"},{"issue":"1","key":"e_1_3_2_96_2","first-page":"85","article-title":"A multivariate technique for multiply imputing missing values using a sequence of regression models","volume":"27","author":"Raghunathan Trivellore E.","year":"2001","unstructured":"Trivellore E. Raghunathan, James M. Lepkowski, John Van Hoewyk, Peter Solenberger, and John van Hoewyk. 2001. A multivariate technique for multiply imputing missing values using a sequence of regression models. Survey Methodology 27, 1 (2001), 85\u201396.","journal-title":"Survey Methodology"},{"issue":"7","key":"e_1_3_2_97_2","doi-asserted-by":"crossref","first-page":"e18910","DOI":"10.2196\/18910","article-title":"Reliability of supervised machine learning using synthetic data in health care: Model to preserve privacy for data sharing","volume":"8","author":"Rankin Debbie","year":"2020","unstructured":"Debbie Rankin, Michaela Black, Raymond Bond, Jonathan Wallace, Maurice Mulvenna, and Gorka Epelde. 2020. Reliability of supervised machine learning using synthetic data in health care: Model to preserve privacy for data sharing. JMIR Medical Informatics 8, 7 (2020), e18910.","journal-title":"JMIR Medical Informatics"},{"issue":"3","key":"e_1_3_2_98_2","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1080\/09332480.2004.10554907","article-title":"New approaches to data dissemination: A glimpse into the future?","volume":"17","author":"Reiter Jerome P.","year":"2004","unstructured":"Jerome P. Reiter. 2004. New approaches to data dissemination: A glimpse into the future? Chance 17, 3 (2004), 11\u201315.","journal-title":"Chance"},{"issue":"3","key":"e_1_3_2_99_2","first-page":"441","article-title":"Using CART to generate partially synthetic public use microdata","volume":"21","author":"Reiter Jerome P.","year":"2005","unstructured":"Jerome P. Reiter. 2005. Using CART to generate partially synthetic public use microdata. Journal of Official Statistics 21, 3 (2005), 441\u2013462.","journal-title":"Journal of Official Statistics"},{"issue":"1","key":"e_1_3_2_100_2","doi-asserted-by":"crossref","first-page":"3069","DOI":"10.1038\/s41467-019-10933-3","article-title":"Estimating the success of re-identifications in incomplete datasets using generative models","volume":"10","author":"Rocher Luc","year":"2019","unstructured":"Luc Rocher, Julien M. Hendrickx, and Yves-Alexandre de Montjoye. 2019. Estimating the success of re-identifications in incomplete datasets using generative models. Nature Communications 10, 1 (2019), 3069.","journal-title":"Nature Communications"},{"key":"e_1_3_2_101_2","doi-asserted-by":"crossref","unstructured":"Natsuki Sano. 2022. Utility and risk evaluation of synthetic data by orthogonal transformation. The Review of Socionetwork Strategies 16 1 (2022) 71--79.","DOI":"10.1007\/s12626-022-00107-x"},{"key":"e_1_3_2_102_2","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/978-3-030-00536-8_1","volume-title":"Proceedings of the 3rd International Workshop on Simulation and Synthesis in Medical Imaging, SASHIMI 2018, Held in Conjunction with MICCAI 2018.","volume":"11037","author":"Shin Hoo-Chang","year":"2018","unstructured":"Hoo-Chang Shin, Neil A. Tenenholtz, Jameson K. Rogers, Christopher G. Schwarz, Matthew L. Senjem, Jeffrey L. Gunter, Katherine P. Andriole, and Mark Michalski. 2018. Medical image synthesis for data augmentation and anonymization using generative adversarial networks. In Proceedings of the 3rd International Workshop on Simulation and Synthesis in Medical Imaging, SASHIMI 2018, Held in Conjunction with MICCAI 2018.Lecture Notes in Computer Science, Vol. 11037, Springer, 1\u201311."},{"issue":"2","key":"e_1_3_2_103_2","first-page":"14:1\u201320","article-title":"To link or synthesize? An approach to data quality comparison","volume":"15","author":"Smith Duncan","year":"2023","unstructured":"Duncan Smith, Mark Elliot, and Joseph W. Sakshaug. 2023. To link or synthesize? An approach to data quality comparison. Journal of Data and Information Quality 15, 2 (2023), 14:1\u201320.","journal-title":"Journal of Data and Information Quality"},{"issue":"3","key":"e_1_3_2_104_2","doi-asserted-by":"crossref","first-page":"663","DOI":"10.1111\/rssa.12358","article-title":"General and specific utility measures for synthetic data","volume":"181","author":"Snoke Joshua","year":"2018","unstructured":"Joshua Snoke, Gillian M. Raab, Beata Nowok, Chris Dibben, and Aleksandra Slavkovic. 2018. General and specific utility measures for synthetic data. Journal of the Royal Statistical Society: Series A (Statistics in Society) 181, 3 (2018), 663\u2013688.","journal-title":"Journal of the Royal Statistical Society: Series A (Statistics in Society)"},{"issue":"4","key":"e_1_3_2_105_2","doi-asserted-by":"crossref","first-page":"3367","DOI":"10.1109\/TKDE.2021.3130903","article-title":"Adversarial attacks against deep generative models on data: A survey","volume":"35","author":"Sun Hui","year":"2023","unstructured":"Hui Sun, Tianqing Zhu, Zhiqiu Zhang, Dawei Jin, Ping Xiong, and Wanlei Zhou. 2023. Adversarial attacks against deep generative models on data: A survey. IEEE Transactions on Knowledge and Data Engineering 35, 4 (2023), 3367\u20133388.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_3_2_106_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.trc.2015.10.010"},{"key":"e_1_3_2_107_2","first-page":"236","volume-title":"Proceedings of the 27th IEEE Pacific Rim International Symposium on Dependable Computing, PRDC 2022","author":"Tai Bo-Chen","year":"2022","unstructured":"Bo-Chen Tai, Szu-Chuang Li, Yennun Huang, and Pang-Chieh Wang. 2022. Examining the utility of differentially private synthetic data generated using variational autoencoder with TensorFlow privacy. In Proceedings of the 27th IEEE Pacific Rim International Symposium on Dependable Computing, PRDC 2022. IEEE, 236\u2013241."},{"key":"e_1_3_2_108_2","first-page":"169","volume-title":"Proceedings of the 37th IEEE International Conference on Data Engineering, ICDE 2021","author":"Takagi Shun","year":"2021","unstructured":"Shun Takagi, Tsubasa Takahashi, Yang Cao, and Masatoshi Yoshikawa. 2021. P3GM: Private high-dimensional data release via privacy preserving phased generative model. In Proceedings of the 37th IEEE International Conference on Data Engineering, ICDE 2021. IEEE, 169\u2013180."},{"issue":"1","key":"e_1_3_2_109_2","first-page":"1","article-title":"The impact of synthetic data generation on data utility with application to the 1991 UK samples of anonymised records","volume":"13","author":"Taub Jennifer","year":"2020","unstructured":"Jennifer Taub, Mark Elliot, and Joseph W. Sakshaug. 2020. The impact of synthetic data generation on data utility with application to the 1991 UK samples of anonymised records. Transactions on Data Privacy 13, 1 (2020), 1\u201323.","journal-title":"Transactions on Data Privacy"},{"issue":"1","key":"e_1_3_2_110_2","doi-asserted-by":"crossref","first-page":"147","DOI":"10.1038\/s41746-020-00353-9","article-title":"Generating high-fidelity synthetic patient data for assessing machine learning healthcare software","volume":"3","author":"Tucker Allan","year":"2020","unstructured":"Allan Tucker, Zhenchen Wang, Ylenia Rotalinti, and Puja Myles. 2020. Generating high-fidelity synthetic patient data for assessing machine learning healthcare software. npj Digital Medicine 3, 1 (2020), 147.","journal-title":"npj Digital Medicine"},{"key":"e_1_3_2_111_2","first-page":"22221","volume-title":"Advances in Neural Information Processing Systems 34: Proceedings of the Annual Conference on Neural Information Processing Systems 2021","author":"Breugel Boris van","year":"2021","unstructured":"Boris van Breugel, Trent Kyono, Jeroen Berrevoets, and Mihaela van der Schaar. 2021. DECAF: Generating fair synthetic data using causally-aware generative networks. In Advances in Neural Information Processing Systems 34: Proceedings of the Annual Conference on Neural Information Processing Systems 2021. Curran Associates, Inc., Virtual Event, 22221\u201322233."},{"key":"e_1_3_2_112_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2022.06.022"},{"key":"e_1_3_2_113_2","first-page":"1","volume-title":"Proceedings of the 9th IEEE International Conference on Data Science and Advanced Analytics, DSAA 2022","author":"Visani Giorgio","year":"2022","unstructured":"Giorgio Visani, Giacomo Graffi, Mattia Alfero, Enrico Bagli, Federico Chesani, and Davide Capuzzo. 2022. Enabling synthetic data adoption in regulated domains. In Proceedings of the 9th IEEE International Conference on Data Science and Advanced Analytics, DSAA 2022. IEEE, 1\u201310."},{"issue":"2","key":"e_1_3_2_114_2","doi-asserted-by":"crossref","first-page":"819","DOI":"10.1111\/coin.12427","article-title":"Generating and evaluating cross-sectional synthetic electronic healthcare data: Preserving data utility and patient privacy","volume":"37","author":"Wang Zhenchen","year":"2021","unstructured":"Zhenchen Wang, Puja Myles, and Allan Tucker. 2021. Generating and evaluating cross-sectional synthetic electronic healthcare data: Preserving data utility and patient privacy. Computational Intelligence 37, 2 (2021), 819\u2013851.","journal-title":"Computational Intelligence"},{"key":"e_1_3_2_115_2","unstructured":"Liyang Xie Kaixiang Lin Shu Wang Fei Wang and Jiayu Zhou. 2018. Differentially private generative adversarial network. arXiv:1802.06739. Retrieved from https:\/\/arxiv.org\/abs\/1802.06739"},{"key":"e_1_3_2_116_2","first-page":"7333","volume-title":"Advances in Neural Information Processing Systems 32: Proceedings of the Annual Conference on Neural Information Processing Systems 2019","author":"Xu Lei","year":"2019","unstructured":"Lei Xu, Maria Skoularidou, Alfredo Cuesta-Infante, and Kalyan Veeramachaneni. 2019. Modeling tabular data using conditional GAN. In Advances in Neural Information Processing Systems 32: Proceedings of the Annual Conference on Neural Information Processing Systems 2019. Curran Associates, Inc., 7333\u20137343."},{"key":"e_1_3_2_117_2","series-title":"Lecture Notes in Business Information Processing","doi-asserted-by":"crossref","first-page":"324","DOI":"10.1007\/978-3-030-61146-0_26","volume-title":"Proceedings of the Business Information Systems Workshops - BIS 2020 International Workshops.","volume":"394","author":"Yale Andrew","year":"2020","unstructured":"Andrew Yale, Saloni Dash, Karan Bhanot, Isabelle Guyon, John S. Erickson, and Kristin P. Bennett. 2020. Synthesizing quality open data assets from private health research studies. In Proceedings of the Business Information Systems Workshops - BIS 2020 International Workshops.Lecture Notes in Business Information Processing, Vol. 394, Springer, 324\u2013335."},{"key":"e_1_3_2_118_2","first-page":"10","volume-title":"Proceedings of the 27th European Symposium on Artificial Neural Networks, ESANN 2019","author":"Yale Andrew","year":"2019","unstructured":"Andrew Yale, Saloni Dash, Ritik Dutta, Isabelle Guyon, Adrien Pavao, and Kristin P. Bennett. 2019. Privacy preserving synthetic health data. In Proceedings of the 27th European Symposium on Artificial Neural Networks, ESANN 2019. i6doc.com, Bruges, Belgium, 10 pages."},{"key":"e_1_3_2_119_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2019.12.136"},{"issue":"1","key":"e_1_3_2_120_2","doi-asserted-by":"crossref","first-page":"7609","DOI":"10.1038\/s41467-022-35295-1","article-title":"A multifaceted benchmarking of synthetic electronic health record generation models","volume":"13","author":"Yan Chao","year":"2022","unstructured":"Chao Yan, Yao Yan, Zhiyu Wan, Ziqi Zhang, Larsson Omberg, Justin Guinney, Sean D. Mooney, and Bradley A. Malin. 2022. A multifaceted benchmarking of synthetic electronic health record generation models. Nature Communications 13, 1 (2022), 7609.","journal-title":"Nature Communications"},{"issue":"8","key":"e_1_3_2_121_2","doi-asserted-by":"crossref","first-page":"2378","DOI":"10.1109\/JBHI.2020.2980262","article-title":"Anonymization through data synthesis using generative adversarial networks (ADS-GAN)","volume":"24","author":"Yoon Jinsung","year":"2020","unstructured":"Jinsung Yoon, Lydia N. Drumright, and Mihaela van der Schaar. 2020. Anonymization through data synthesis using generative adversarial networks (ADS-GAN). IEEE Journal of Biomedical and Health Informatics 24, 8 (2020), 2378\u20132388.","journal-title":"IEEE Journal of Biomedical and Health Informatics"},{"issue":"3","key":"e_1_3_2_122_2","doi-asserted-by":"crossref","first-page":"618","DOI":"10.1093\/jssam\/smac016","article-title":"A semiparametric multiple imputation approach to fully synthetic data for complex surveys","volume":"10","author":"Yu Mandi","year":"2022","unstructured":"Mandi Yu, Yulei He, and Trivellore E. Raghunathan. 2022. A semiparametric multiple imputation approach to fully synthetic data for complex surveys. Journal of Survey Statistics and Methodology 10, 3 (2022), 618\u2013641.","journal-title":"Journal of Survey Statistics and Methodology"},{"issue":"4","key":"e_1_3_2_123_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3134428","article-title":"PrivBayes","volume":"42","author":"Zhang Jun","year":"2017","unstructured":"Jun Zhang, Graham Cormode, Cecilia M. Procopiuc, Divesh Srivastava, and Xiaokui Xiao. 2017. PrivBayes. ACM Transactions on Database Systems 42, 4 (2017), 1\u201341.","journal-title":"ACM Transactions on Database Systems"},{"key":"e_1_3_2_124_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2021.103977"},{"key":"e_1_3_2_125_2","first-page":"5855","volume-title":"Proceedings of the IEEE International Conference on Big Data, Big Data 2022","author":"Zhu Yujin","year":"2022","unstructured":"Yujin Zhu, Zilong Zhao, Robert Birke, and Lydia Y. Chen. 2022. Permutation-invariant tabular data synthesis. In Proceedings of the IEEE International Conference on Big Data, Big Data 2022. IEEE, 5855\u20135864."}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3704437","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3704437","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:05Z","timestamp":1750295885000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3704437"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,9]]},"references-count":124,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,4,30]]}},"alternative-id":["10.1145\/3704437"],"URL":"https:\/\/doi.org\/10.1145\/3704437","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,9]]},"assertion":[{"value":"2023-05-04","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-11-04","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-09","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}