Comparing the Utility and Disclosure Risk of Synthetic Data with Samples of Microdata

Claire Little; Mark Elliot; Richard Allmendinger

Conference Proceedings

Comparing the Utility and Disclosure Risk of Synthetic Data with Samples of Microdata

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2022) 13463 LNCS 234-249

DOI: 10.1007/978-3-031-13945-1_17

12Citations

4Readers

Get full text

Abstract

Most statistical agencies release randomly selected samples of Census microdata, usually with sample fractions under 10% and with other forms of statistical disclosure control (SDC) applied. An alternative to SDC is data synthesis, which has been attracting growing interest, yet there is no clear consensus on how to measure the associated utility and disclosure risk of the data. The ability to produce synthetic Census microdata, where the utility and associated risks are clearly understood, could mean that more timely and wider-ranging access to microdata would be possible. This paper follows on from previous work by the authors which mapped synthetic Census data on a risk-utility (R-U) map. The paper presents a framework to measure the utility and disclosure risk of synthetic data by comparing it to samples of the original data of varying sample fractions, thereby identifying the sample fraction which has equivalent utility and risk to the synthetic data. Three commonly used data synthesis packages are compared with some interesting results. Further work is needed in several directions but the methodology looks very promising.

Author supplied keywords

Cite

CITATION STYLE

APA

Little, C., Elliot, M., & Allmendinger, R. (2022). Comparing the Utility and Disclosure Risk of Synthetic Data with Samples of Microdata. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 13463 LNCS, pp. 234–249). Springer Science and Business Media Deutschland GmbH. https://doi.org/10.1007/978-3-031-13945-1_17

Comparing the Utility and Disclosure Risk of Synthetic Data with Samples of Microdata

Abstract

Author supplied keywords

Cite

Register to see more suggestions