LLM-Based Persona-Driven Text Data Augmentation

1Citations
Citations of this article
15Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Illicit online communication, such as drug-dealing dialogues, is increasingly conducted through covert, context dependent language patterns that evade traditional detection techniques in South Korea. However, developing reliable AI based detection systems remains challenging due to the scarcity of real world training data in such sensitive domains. This paper proposes a novel persona-driven data augmentation framework using Large Language Model(LLM) to generate realistic synthetic drug-dealing dialogues. By encoding domain specific buyer and seller personas along with linguistic behaviour rules, the method produces contextually coherent and semantically diverse dialogues that reflect authentic communication styles. Evaluation results demonstrate that the augmented data preserves key stylistic features (high cosine similarity), maintains lexical diversity (TTR), improves fluency (perplexity), and enhances coherence and lexical richness (ROUGE-L), outperforming traditional augmentation method. Furthermore, statistical validation confirms the semantic consistency and stability of the generated data. These findings highlight the viability of LLM-based augmentation in low-resource, high-risk domains and suggest its potential transferability to other specialized NLP applications requiring context-preserving generation.

Cite

CITATION STYLE

APA

Seong Jeong, H., Kyeong Ko, H., Yong Park, S., & Kim, T. (2025). LLM-Based Persona-Driven Text Data Augmentation. IEEE Access, 13, 167560–167577. https://doi.org/10.1109/ACCESS.2025.3611636

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free