The performance comparison of the classifiers according to binary bow, count bow and TF-IDF feature vectors for malware detection

5Citations
Citations of this article
15Readers
Mendeley users who have this article in their library.
Get full text

Abstract

In this paper, we compared the performance of the classifiers according to feature vectors with Binary BOW, Count BOW and TF-IDF for malware detection. We used the feature of Opcode that extracted from PE file. For performance comparison, we measured the AUC score for the classifiers those are DT, KNN, MLP, MNB and SVM. As a result, we recommend neural network (MLP) and instance-based model (KNN) because they show the high AUC score and accuracy regardless of the unbalanced dataset and the feature vector. If you use classical classifiers, we recommend DT because it guarantees high AUC score and accuracy regardless of the same condition as the above. If you use SVM, you have to do Robust scaling to resolved outlier and unbalanced dataset. If you use MNB, you need to use N-gram technique to improve AUC score.

Cite

CITATION STYLE

APA

Kwon, Y. M., Jun, S. H., Gal, W. M., & Lim, M. J. (2018). The performance comparison of the classifiers according to binary bow, count bow and TF-IDF feature vectors for malware detection. International Journal of Engineering and Technology(UAE), 7(3), 15–22. https://doi.org/10.14419/ijet.v7i3.33.18515

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free