Large Language Models for Synthetic Tabular Health Data: A Benchmark Study

5Citations
Citations of this article
14Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Synthetic tabular health data plays a crucial role in healthcare research, addressing privacy regulations and the scarcity of publicly available datasets. This is essential for diagnostic and treatment advancements. Among the most promising models are transformer-based Large Language Models (LLMs) and Generative Adversarial Networks (GANs). In this paper, we compare LLM models of the Pythia LLM Scaling Suite with varying model sizes ranging from 14M to 1B, against a reference GAN model (CTGAN). The generated synthetic data are used to train random forest estimators for classification tasks to make predictions on the real-world data. Our findings indicate that as the number of parameters increases, LLM models outperform the reference GAN model. Even the smallest 14M parameter models perform comparably to GANs. Moreover, we observe a positive correlation between the size of the training dataset and model performance. We discuss implications, challenges, and considerations for the real-world usage of LLM models for synthetic tabular data generation.

Cite

CITATION STYLE

APA

Miletic, M., & Sariyar, M. (2024). Large Language Models for Synthetic Tabular Health Data: A Benchmark Study. In Studies in Health Technology and Informatics (Vol. 316, pp. 963–967). IOS Press BV. https://doi.org/10.3233/SHTI240571

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free