30 Years of Synthetic Data

30Citations
Citations of this article
46Readers
Mendeley users who have this article in their library.

Abstract

The idea to generate synthetic data as a tool for broadening access to sensitive microdata has been proposed for the first time three decades ago. While first applications of the idea emerged around the turn of the century, the approach really gained momentum over the last ten years, stimulated at least in parts by some recent developments in computer science. We consider the 30th jubilee of Rubin’s seminal paper on synthetic data (J. Off. Stat. 9 (1993) 462-468) as an opportunity to look back at the historical developments but also to offer a review of the diverse approaches and methodological underpinnings proposed over the years. We will also discuss the various strategies that have been suggested to measure the utility and remaining risk of disclosure of the generated data.

Cite

CITATION STYLE

APA

Drechsler, J., & Haensch, A. C. (2024). 30 Years of Synthetic Data. Statistical Science, 39(2), 221–242. https://doi.org/10.1214/24-STS927

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free