Abstract
The present research proposes a highly precise technique for detecting toxic comments in Assamese, a low-resource language that is a member of the Eastern Indo-Aryan family. As a low-resource language, Assamese encounters specific challenges such dialectal differences, incorrect spellings, frequent code-mixing with English, and a lack of resources with annotations. A custom-built dataset of 100,000 Assamese social media comments was gathered and split equally between toxic and non-toxic categories in order to solve the lack of labelled data. The article presents a multi-layer hybrid deep learning model that integrates multi-scale Convolutional Neural Networks (CNN), bidirectional LSTM (BiLSTM), and bidirectional GRU (BiGRU). A thorough grid search utilising several optimisers and activation functions was used to create and optimise two hybrid architectures: CNN-BiLSTM-BiGRU and BiGRU-BiLSTM-CNN. Among them, the CNN-BiLSTM-BiGRU model outperformed present approaches for Assamese and other structurally related languages, reaching 94.25% accuracy and an F1- score of 0.93%. This work not only produces strong analytical findings, but it also makes a novel contribution in developing and evaluating models for Assamese toxic comment detection under real-world linguistic and resource constraints, providing a solid foundation and valuable insights for furthering NLP research in other under resourced languages.
Cite
CITATION STYLE
Neog, M., & Baruah, N. (2025). A Multi-Layer Hybrid Deep Learning Model for Toxic Comment Detection in Assamese. IEEE Access, 13, 165718–165737. https://doi.org/10.1109/ACCESS.2025.3612497
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.