An overview of synthetic administrative data for research

9Citations
Citations of this article
29Readers
Mendeley users who have this article in their library.

Abstract

Use of administrative data for research and for planning services has increased over recent decades due to the value of the large, rich information available. However, concerns about the release of sensitive or personal data and the associated disclosure risk can lead to lengthy approval processes and restricted data access. This can delay or prevent the production of timely evidence. A promising solution to facilitate more efficient data access is to create synthetic versions of the original datasets which are less likely to hold confidential information and can minimise disclosure risk. Such data may be used as an interim solution, allowing researchers to develop their analysis plans on non-disclosive data, whilst waiting for access to the real data. We aim to provide an overview of the background and uses of synthetic data and describe common methods used to generate synthetic data in the context of UK administrative research. We propose a simplified terminology for categories of synthetic data (univariate, multivariate, and complex modality synthetic data) as well as a more comprehensive description of the terminology used in the existing literature and illustrate challenges and future directions for research.

Cite

CITATION STYLE

APA

Kokosi, T., De Stavola, B., Mitra, R., Frayling, L., Doherty, A., Dove, I., … Harron, K. (2022). An overview of synthetic administrative data for research. International Journal of Population Data Science. Swansea University. https://doi.org/10.23889/ijpds.v7i1.1727

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free