A hybrid naïve Bayes based on similarity measure to optimize the mixed-data classification

Fatima El Barakaz; Omar Boutkhoum; Abdelmajid El Moutaouakkil

Journal ArticleOPEN ACCESS

A hybrid naïve Bayes based on similarity measure to optimize the mixed-data classification

Telkomnika (Telecommunication Computing Electronics and Control) (2021) 19(1) 155-162

DOI: 10.12928/TELKOMNIKA.V19I1.18024

4Citations

6Readers

Abstract

In this paper, a hybrid method has been introduced to improve the classification performance of naïve Bayes (NB) for the mixed dataset and multi-class problems. This proposed method relies on a similarity measure which is applied to portions that are not correctly classified by NB. Since the data contains a multi-valued short text with rare words that limit the NB performance, we have employed an adapted selective classifier based on similarities (CSBS) classifier to exceed the NB limitations and included the rare words in the computation. This action has been achieved by transforming the formula from the product of the probabilities of the categorical variable to its sum weighted by numerical variable. The proposed algorithm has been experimented on card payment transaction data that contains the label of transactions: the multi-valued short text and the transaction amount. Based on K-fold cross validation, the evaluation results confirm that the proposed method achieved better results in terms of precision, recall, and F-score compared to NB and CSBS classifiers separately. Besides, the fact of converting a product form to a sum gives more chance to rare words to optimize the text classification, which is another advantage of the proposed method.

Author supplied keywords

Cite

CITATION STYLE

APA

Barakaz, F. E., Boutkhoum, O., & Moutaouakkil, A. E. (2021). A hybrid naïve Bayes based on similarity measure to optimize the mixed-data classification. Telkomnika (Telecommunication Computing Electronics and Control), 19(1), 155–162. https://doi.org/10.12928/TELKOMNIKA.V19I1.18024

A hybrid naïve Bayes based on similarity measure to optimize the mixed-data classification

Abstract

Author supplied keywords

Cite

Register to see more suggestions