Abstract
This study examines the impact of large language model-based data augmentation on Turkish sentiment analysis by analyzing how synthetic data influences model performance across different dataset configurations. A balanced corpus of hotel and movie reviews labeled with positive and negative sentiments was used to construct six dataset variants: Raw, Raw+GPT, Raw+Claude, Raw+GPT+Claude, GPT-only, and Claude-only. Synthetic samples were generated through zero-shot prompting using GPT-4o and Claude models and incorporated exclusively into the training set. Multiple classifiers - including Naive Bayes, SVM, Logistic Regression, CNN, LSTM, and BERT - were trained and evaluated using accuracy, precision, recall, and F1-score metrics. Experimental results reveal that while synthetic-only datasets lead to substantial performance degradation, hybrid datasets (Raw + Synthetic) achieve results statistically comparable to the Raw dataset, demonstrating that synthetic reviews can effectively complement real data. These findings highlight that large language model-generated synthetic data, when combined with authentic samples, preserves model robustness while expanding training diversity, offering a practical solution for low-resource language datasets.
Author supplied keywords
Cite
CITATION STYLE
Tekinay, B., & Alp Tocoglu, M. (2026). From Raw to Synthetic: Evaluating LLM-Based Data Augmentation for Sentiment Analysis. IEEE Access, 14, 25747–25769. https://doi.org/10.1109/ACCESS.2026.3664756
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.