Abstract
This research is intended to explore and evaluate various predictive models for the classification performance of breast cancer risk factors. First, data acquisition is being carried out to obtained three datasets from Breast Cancer Surveillance Consortium (BCSC). After that, data integration is performed to combine the datasets into one. Then, data preprocessing is performed to do data cleaning. Feature selection is executed to eliminate unrelated attributes. Data resampling is applied to resolve imbalanced data. Four classifiers namely Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), and Multilayer Perceptron (MLP) are used in classifying the risk factors of breast cancer. These four classifiers undergo training and testing data with 80-20, 70-30, and 60-40 train test splits. RF performs the best performance with 82% of accuracy at 80-20 train test split.
Author supplied keywords
Cite
CITATION STYLE
Yee, W. S., Ng, H., Yap, T. T. V., Goh, V. T., Ng, K. H., & Cher, D. T. (2022). An Evaluation Study on the Predictive Models of Breast Cancer Risk Factor Classification. Journal of Logistics, Informatics and Service Science, 9(3), 129–145. https://doi.org/10.33168/LISS.2022.0310
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.