Abstract
This study investigates the application of advanced Transformer-based models, namely BERT, DistilBERT, BERT-multilingual, ALBERT, and BERT-CNN, for sentiment analysis in Bahasa Malaysia, addressing unique challenges such as mixed-language usage and abbreviated expressions in social media text. Using the Malaya dataset to ensure linguistic diversity and domain coverage, the research incorporates robust preprocessing techniques, including synonym mapping and sentiment-aware tokenization, to enhance feature extraction. Through rigorous evaluation, BERT-CNN exhibits the best accuracy (96.3%), followed by BERT-multilingual (89.84%) and BERT (89.5%). DistilBERT and ALBERT delivered competitive performance (88.96% and 88.76%, respectively) while offering reduced computational requirements, highlighting the trade-offs between performance and efficiency. The study emphasizes optimized strategies for handling challenges in positive sentiment classification and demonstrates the efficacy of transformer architectures in nuanced sentiment detection for low-resource languages. These findings contribute to advancing Natural Language Processing (NLP) for scalable sentiment analysis across domains.
Cite
CITATION STYLE
Zulkalnain, M. A., Syafeeza, A. R., Mohd Saad, W. H., & Rahaman, S. (2025). Evaluation of Transformer-Based Models for Sentiment Analysis in Bahasa Malaysia. Journal of Telecommunication, Electronic and Computer Engineering (JTEC), 17(1), 29–33. https://doi.org/10.54554/jtec.2025.17.01.004
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.