Abstract
Lexical diversity, widely recognized as an essential indicator of language proficiency and cognitive ability, plays a pivotal role across various linguistic analyses. However, lexical diversity metrics are inherently sensitive to text length, introducing variability that affects their stability and interpretability. Traditional approaches typically examine metric behavior across different text lengths without adequately accounting for the influence of lexical content. To address this gap, this study investigated the minimum text length necessary to reliably estimate lexical diversity in genre-sensitive contexts. Representative lexical diversity metrics were systematically assessed in terms of internal consistency and genre discriminative power across varying text lengths, utilizing a Japanese corpus comprising four distinct genres: political speeches, natural conversations, news, and novels. Furthermore, synthetic data were used to assess metric behavior under controlled conditions. The findings indicated that shorter texts under 500 words exhibited significant variability, primarly reflecting sample size effects rather than genuine lexical differences. Stability in metric values emerged between 1,000 and 2,000 words. Furthermore, analyses using synthetic data suggested that artificially duplicating shorter texts may provide a practical workaround when longer samples are unavailable. This study enhances methodological rigor and offers a practical approach for determining appropriate text lengths in lexical diversity analyses.
Author supplied keywords
Cite
CITATION STYLE
Zheng, W. (2025). Text length requirements for stable and genre-sensitive lexical diversity measurement. Cogent Arts and Humanities, 12(1). https://doi.org/10.1080/23311983.2025.2584418
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.