Random Forest Algorithm Based on GAN for Imbalanced Data Classification

7Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Data imbalance increases the difficulty of knowledge mining from data. Aiming at the problem of data imbalanced classification, a random forest algorithm based on GAN is proposed, which can effectively classify imbalanced data sets. The GAN model is used to generate a few class samples, which are mixed with the original data samples to form a new data set to reduce the imbalance of data. Then the new data set is divided into several data subsets with a balanced sample distribution. Finally, all the decision trees are collected to form the forest and the classification results are obtained. The random forest algorithm based on GAN improves the classification effect on minority classes, which makes the equilibrium rate of the dataset reach 50%. The algorithm in this paper has performed relevant experiments on multiple imbalanced datasets, the results show that the algorithm has good performance in data imbalance classification. The optimized algorithm after parallelization improves the running speed.

Cite

CITATION STYLE

APA

Shu, Q., Hu, T., & Liu, S. (2020). Random Forest Algorithm Based on GAN for Imbalanced Data Classification. In Journal of Physics: Conference Series (Vol. 1544). Institute of Physics Publishing. https://doi.org/10.1088/1742-6596/1544/1/012014

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free