Abstract
The rapid growth of digital communication has facilitated the spread of toxic speech, which can harm individuals or communities and often appears across multiple nuanced categories. These categories are difficult to detect in short texts due to semantic ambiguity, limited context, and label dependencies. This study introduces a Hybrid Semantic Enrichment with Convolutional Neural Network (HSE-CNN) approach to enhance multilabel toxic speech classification. The HSE-CNN model leverages semantic enrichment techniques such as back translation, text expansion, word sense disambiguation (WSD), and semantic similarity mapping to enrich the contextual meaning of input texts. Using an Indonesian social media dataset containing 13,169 entries labeled with 12 toxic speech categories, we conducted a series of experiments involving preprocessing, semantic enrichment, and classification using various deep learning models. The optimal configuration includes a learning rate of 0.001, batch size of 16, and training for 30 epochs. Our proposed model achieved an F1-score of 80%, accuracy of 93%, and AUC of 91%, demonstrating its superiority over non-enriched models. Compared to baseline models such as BiLSTM and BiGRU, the HSE-CNN yields a 6.7% improvement in accuracy and a 4.5% improvement in F1-score. These findings suggest that HSE-CNN offers a promising solution for toxic speech detection systems, especially in resource-limited languages, with potential applications in digital content moderation, online safety initiatives, and public awareness enhancement.
Author supplied keywords
Cite
CITATION STYLE
Muzakir, A., Suriani, U., & Ependi, U. (2025). A Hybrid Semantic Enrichment Approach for Multi-Label Toxic Speech Detection. International Journal of Safety and Security Engineering, 15(3), 509–520. https://doi.org/10.18280/ijsse.150310
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.