Abstract
Text classification has seen a lot of research, especially after social media platforms came into existence. It involves categorizing text into a variety of classes, such as positive, negative, neutral, or any other class label. The primary issue is that the usual text datasets collected from social media platforms are skewed (i.e., imbalanced). Consequently, classifiers become less effective. However, this paper addresses the issue by developing a new technique called BalBERT, which is based on the BERT scheme and includes a new sublayer to improve performance when dealing with imbalanced text classification. This new sublayer introduces balancing techniques that take advantage of the BERT representation step, which is context-based. The results demonstrate the effectiveness of BalBERT as measured by two metrics: AVG-Recall and F1-score. It outperforms both the BERT baseline and the state-of-the-art results.
Author supplied keywords
Cite
CITATION STYLE
Mahmoudi, L., & Salem, M. (2023). BalBERT: A New Approach to Improving Dataset Balancing for Text Classification. Revue d’Intelligence Artificielle, 37(2), 425–431. https://doi.org/10.18280/ria.370219
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.