How Useful Is Synthetic Data in Developing Predictive Models for Health?

1Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.

Abstract

Synthetic data, generated using generative AI techniques, closely mimics the characteristics of real data while enhancing privacy for sensitive health data. This study evaluates synthetic tabular data based on fidelity and utility for predictive models. Fidelity is measured through univariate distribution and bivariate differential pairwise correlations, while utility is measured by comparing machine learning model performance trained on synthetic and real data. Results show highly similar model performance on synthetic and real data. We also explore the potential of using synthetic data for hyperparameter tuning. Our findings reveal a strong correlation between prediction accuracy on synthetic and real data, suggesting that hyperparameters optimized using synthetic data can be effectively applied to models trained on real datasets for optimal results.

Cite

CITATION STYLE

APA

Basri, M. A., & Chen, H. (2025). How Useful Is Synthetic Data in Developing Predictive Models for Health? In Studies in Health Technology and Informatics (Vol. 327, pp. 552–556). IOS Press BV. https://doi.org/10.3233/SHTI250398

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free