Performance Assessment of ML and DL Models in Detecting Hate Speech from Mixed English–Roman Urdu Text with Small-Scale Datasets

1Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

This research evaluates hate speech detection across a minimal-size Mixed English and Roman Urdu language intersection through machine learning and deep learning model analysis. For traditional models, the data required text cleaning alongside tokenization and TF-IDF vectorization to participate in the same experiment, as deep learning models needed trainable embeddings. Experiment results between Naive Bayes, Logistic Regression, Linear SVM, Random Forest, LSTM, BiLSTM, and CNN demonstrated that Logistic Regression produced the greatest F1 score of 0.8073. CNN was the most effective choice among deep learning models, scoring an F1 score of 0.7786. The research demonstrates that traditional models perform well on small-scale datasets; however, deep learning has evolved as an effective 38 tool for processing code-mixed text. In the future, Researchers should examine pre-trained embeddings, larger datasets, and advanced models to raise detection abilities for hate speech in mixed-language text.

Cite

CITATION STYLE

APA

Saleem, H., Javed, M., Alahmadi, A., Aljubayri, I., Khan, M. Z., & Junaid. (2025). Performance Assessment of ML and DL Models in Detecting Hate Speech from Mixed English–Roman Urdu Text with Small-Scale Datasets. Advances in Artificial Intelligence and Machine Learning, 5(2), 3883–3899. https://doi.org/10.54364/AAIML.2025.52220

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free