Generation of Synthetic Continuous Numerical Data Using Generative Adversarial Networks

9Citations
Citations of this article
29Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Continuous numerical is a type of data which often used for unsupervised learning such as clustering. However, this valuable data often provided in a small amount because it is hard to obtain, expensive, required an expert to collect them, or not available because it contains confidential information that cannot be published. These limited data situations can be an obstacle for processing and analyzing data or restrain clustering related research in general. Therefore, there is a need to be an alternative that can replace or increase the amount of data. The proposed method is generating synthetic continuous numerical data using Generative Adversarial Networks (GANs). This study used two GAN architectures (GAN and CGAN) and focused on unlabeled continuous numerical data to provide replacement or additional data for the clustering task. The Quality of synthetic data was measured using the accuracy of the xgboost algorithm in classifying real and synthetic data. When the xgboost accuracy of perfectly realistic data is 50%, synthetic data based on CGAN achieving 63%. The result of this study shows that GAN can generate data similar enough and not significantly different from the real data.

Cite

CITATION STYLE

APA

Aziira, A. H., Setiawan, N. A., & Soesanti, I. (2020). Generation of Synthetic Continuous Numerical Data Using Generative Adversarial Networks. In Journal of Physics: Conference Series (Vol. 1577). Institute of Physics Publishing. https://doi.org/10.1088/1742-6596/1577/1/012027

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free