Abstract
The success of modern Natural Language Processing (NLP), underpinned by large-scale Transformer architectures, is contingent upon access to extensive benchmark datasets. This dependence has intensified the resource disparity between a handful of high-resource languages and others, such as Vietnamese, where data scarcity poses a significant barrier to scientific and technological progress. To address this, we present the most extensive unified survey of Vietnamese NLP to date, distinguishing itself from prior task-specific overviews by simultaneously synthesizing 68 publicly available datasets, a complete architectural evolution from statistical to generative models, modern LLM benchmarks (VLUE, VMLU), and a structured linguistic challenge analysis within a single, consolidated reference. Each dataset is evaluated with respect to size, domain, annotation scheme, and accessibility. Beyond the dataset review, we synthesize the current state of Vietnamese NLP by examining 21 prominent modeling techniques from neural representations to Transformer-based architectures while benchmarking their empirical performance. Our analysis reveals critical gaps in available resources, particularly in under-represented domains and task categories, and identifies methodological challenges distinctive to Vietnamese, including word segmentation and syntactic ambiguity. In response, we propose a targeted roadmap with actionable recommendations for dataset creation, model innovation, and evaluation strategies. This survey aims to catalyze robust, reproducible research in Vietnamese, contributing to greater inclusivity within the global NLP community.
Author supplied keywords
Cite
CITATION STYLE
Tran, K. V., Thai, T. M., Luu, S. T., Nguyen, K. V., & Nguyen, N. L. T. (2026, September 25). Natural language processing and computational linguistics for Vietnamese: A comprehensive review. Expert Systems with Applications. Elsevier Ltd. https://doi.org/10.1016/j.eswa.2026.132696
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.