SMOTE VS. RANDOM UNDERSAMPLING FOR IMBALANCED DATA-CAR OWNERSHIP DEMAND MODEL

12Citations
Citations of this article
34Readers
Mendeley users who have this article in their library.

Abstract

Because the numbers of cars reflect each person's travel behaviors for each specific location, the car ownership demand model plays a dominant role in analysis of the travel demand in order to understand each area's individual and household travel behaviors. However, the study project for the master plan of the Khon Kaen expressway represented imbalanced data; namely, the majority class and the minority class were not equal. Before developing a machine learning model, this study suggested a solution to balance the data by using oversampling and under-sampling techniques. The data, which had been improved with SMOTE (Synthetic Minority Oversampling Technique) and kNN (k-nearest neighbors) (k = 5), demonstrated a better effect than the other algorithms that were studied. The TPR (true positive rate) for the rural and suburban areas, which are types of regions with very different imbalance ratios, was calculated before balancing the data at 46.9% and 46.4%. As a result, the TPR values were 63.5% and 54.4%, respectively, following the data balancing.

Cite

CITATION STYLE

APA

Chaipanha, W., & Kaewwichian, P. (2022). SMOTE VS. RANDOM UNDERSAMPLING FOR IMBALANCED DATA-CAR OWNERSHIP DEMAND MODEL. Communications - Scientific Letters of the University of Žilina, 24(3), D105–D115. https://doi.org/10.26552/com.C.2022.3.D105-D115

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free