Abstract
Text classification is a task in natural language processing (NLP) in which text data is classified into one or more predefined categories or labels. Various techniques, including machine learning algorithms like SVMs, decision trees, and neural networks, can be used to perform this task. Other approaches involve a new model Bidirectional Encoder Representations from Transformers (BERT) which caused controversy in the machine learning community by presenting state-of-the-art results on various NLP tasks. We conducted an experiment to compare the performance of different natural language processing (NLP) pipelines and analysis models (traditional and new) of classification on two datasets. This study could shed significant light on improving the accuracy during text classification. We found that using lemmatization and knowledge-based n-gram features with LinearSVC classifier and BERT resulted in the high accuracies of 98% and 97% respectively surpassing other classification models used in the same corpus. This means that BERT, TF-IDF vectorization and LinearSVC classification model used Text categorization scores to get the best performance, with an advantage in favor of BERT, allowing the improvement of accuracy by increasing the number of epochs.
Author supplied keywords
Cite
CITATION STYLE
Sabiri, B., Khtira, A., El Asri, B., & Rhanoui, M. (2023). Analyzing BERT’s Performance Compared to Traditional Text Classification Models. In International Conference on Enterprise Information Systems, ICEIS - Proceedings (Vol. 1, pp. 572–582). Science and Technology Publications, Lda. https://doi.org/10.5220/0011983100003467
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.