Abstract
The prevalence of hate speech on digital platforms presents significant challenges, particularly in multilingual communities where code-mixing complicates detection. This study explores the use of XLM-RoBERTa, a transformer-based model with robust multilingual capabilities, to detect hate speech within code-mixed texts, focusing on Indonesian-English code-mixing. Traditional hate speech detection models rely on single-language datasets, limiting their effectiveness in such environments. We employ a dataset consisting of Indonesian, English, and code-mixed Indonesian-English texts to evaluate XLM-RoBERTa's performance. The dataset comprises 24,844 training samples, 2,760 test samples, and an additional 100 supplementary samples. Key hyperparameters included a batch size of 16 and 32, with a learning rate ranging from 1e-5 to 5e-5. The model achieved near-perfect accuracy (99.6%) on the primary test set and demonstrated strong generalization across realistic supplementary data, achieving an F1-score of 90.94%. These findings underscore the model's potential for application in complex linguistic contexts, contributing to the development of effective code-mixed hate speech detection.
Author supplied keywords
Cite
CITATION STYLE
Dinarta, F., & Wicaksana, A. (2025). Enhanced Hate Speech Detection in Indonesian-English Code-Mixed Texts Using XLM-RoBERTa. Informatica (Slovenia), 49(21), 45–56. https://doi.org/10.31449/inf.v49i21.7713
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.